Out-of-distribution Tests Reveal Compositionality in Chess Transformers
Anna Mészáros, Patrik Reizinger, Ferenc Huszár
TL;DR
The paper investigates whether chess Transformers demonstrate genuine compositional generalization by evaluating a ~270M-parameter policy on diverse out-of-distribution positions and chess variants. Using behavior cloning on ChessBench and Stockfish as oracle, it measures legal-move adherence, topK move quality, puzzle-sequence accuracy, and Elo across Standard, Chess960, and Horde. It finds strong rule extrapolation with near-perfect legality in ID and robust legality in many OOD cases, plus high-quality OOD puzzle play, but limited strategy adaptation in the Horde and highly divergent starts. The results indicate that purely data-driven Transformers encode substantial compositional structure of chess, though there remains a gap to explicit symbolic search, with implications for understanding reasoning and generalization in AI, as well as for robust deployment in varied game settings.
Abstract
Chess is a canonical example of a task that requires rigorous reasoning and long-term planning. Modern decision Transformers - trained similarly to LLMs - are able to learn competent gameplay, but it is unclear to what extent they truly capture the rules of chess. To investigate this, we train a 270M parameter chess Transformer and test it on out-of-distribution scenarios, designed to reveal failures of systematic generalization. Our analysis shows that Transformers exhibit compositional generalization, as evidenced by strong rule extrapolation: they adhere to fundamental syntactic rules of the game by consistently choosing valid moves even in situations very different from the training data. Moreover, they also generate high-quality moves for OOD puzzles. In a more challenging test, we evaluate the models on variants including Chess960 (Fischer Random Chess) - a variant of chess where starting positions of pieces are randomized. We found that while the model exhibits basic strategy adaptation, they are inferior to symbolic AI algorithms that perform explicit search, but gap is smaller when playing against users on Lichess. Moreover, the training dynamics revealed that the model initially learns to move only its own pieces, suggesting an emergent compositional understanding of the game.
