Language Models that Play Chess and Explain Their Moves
Queen is a 4B-parameter chess-language model that explains its moves and plans while playing at the level of a typical Grandmaster. A silent expert chess encoder feeds an instruction-tuned language model through cross-attention, trained with a question-answering curriculum. An iterative distillation step, a natural-language analog of the Bellman update, analyzes positions after top candidate moves and distills the consolidated explanation back into the model. Over seven iterations Elo rises from 1782 to 2697. The authors report that it beats all frontier models on strength and puzzles with three orders of magnitude fewer parameters.
Background and Problem Definition
Modern chess engines are silent experts. They play far beyond human strength, yet they cannot say why they chose a move. Language models have the opposite profile. They can write explanations that sound reasonable, but their own playing strength is weak, so those explanations carry little weight. A teacher who blunders cannot be trusted, however fluent the lecture.
The paper (arXiv:2610.03695v1, by Adithya Bhaskar, Jeffrey Cheng and Danqi Chen, category cs.CL) targets this gap. Can one model both play well and explain itself? The authors answer with Queen, a 4B-parameter chess-language model. According to the abstract, Queen explains its moves and plans while playing at the level of a typical Grandmaster.
A note on scope: this article is written from the abstract alone. Details such as training data size, the opponents used for rating, and the exact Elo measurement protocol are not stated there, so this article does not guess at them.
Architectural Core and Technical Principles
Queen combines two complementary components: an encoder-decoder architecture and an iterative distillation algorithm. The first component is the architecture. A silent expert chess encoder is joined to an instruction-tuned language model through cross-attention. The encoder reads the position. The language model produces the text. Cross-attention lets the language model consult the encoder's internal representations while it generates each token. Training follows a question-answering curriculum. The model must answer questions about positions, and in doing so it learns to extract chess concepts from the encoder's representations. Typical concepts of this kind are material balance, threats and tactical motifs, although the abstract does not list them.
The second component is iterative distillation. The authors describe it as a natural-language analog of the Bellman update. In reinforcement learning, a Bellman update corrects the value of a state using the values of its successor states. Queen follows the same shape in words. First, the model analyzes the positions that arise after its top candidate moves. Second, it consolidates those analyses into an explanation of the current position. Third, that explanation is distilled back into the model. In effect, the result of looking ahead is written down as text and then becomes training signal for the next round. The model can absorb what search would tell it without running that search at every inference step. The abstract reports that, over seven iterations, the model gains more than 900 Elo points, rising from 1782 to 2697.
Practical Evaluation and Applications
The abstract gives three main results. First, playing strength. Elo rises from 1782 to 2697, a gain of over 900 points. The authors state that the model substantially surpasses all frontier models on both playing strength and puzzle accuracy.
Second, size. Queen has 4B parameters, three orders of magnitude fewer than the frontier models it is compared against. This suggests that, in this domain, a specialized encoder plus targeted training can beat raw parameter count. Third, explanation quality. Using language-model-based evaluations, the authors find the explanations fluent and close to GPT-5.6-Sol (high) in coherence. Readers should keep several limits in mind. One: a language model scores the explanations. That measures fluency and coherence. It does not prove that an explanation faithfully reflects what drove the move. Two: "approaches in coherence" covers one dimension only, and the abstract does not claim parity on others. Three: this is a version 1 preprint and has not been peer reviewed. Four: the abstract gives no ablation details, so we cannot tell how much the encoder, the question-answering curriculum and the iterative distillation each contribute.
Industry Impact and Outlook
The authors argue that the architecture and training recipe are general. Wherever a silent expert encoder exists, the recipe could apply. They name games, robotics and computer use. This is a useful hypothesis, but at this point it is the authors' inference. The abstract offers evidence only from chess.
Two lessons stand out. The first is the division of labor: an expert system judges, and a language model speaks. Combined with a loop that distills search-like conclusions back into the model, it is a reusable route. The second is that evaluating explanations remains hard. A fluent explanation is not always a true one. Future work should test whether the stated reasons match the model's actual decisions, and should reproduce the results outside chess. Only then can the community decide whether this recipe truly generalizes.
Sources
FAQ
How large is Queen and how strong does it play?
Per the abstract, Queen has 4B parameters. Over seven iterations its Elo rises from 1782 to 2697, a gain above 900, which the authors place at the level of a typical Grandmaster.
How does the iterative distillation work?
The model analyzes positions after its top candidate moves, consolidates those analyses into an explanation of the current position, and distills that explanation back into itself. The authors call this a natural-language analog of the Bellman update.
How is the language model connected to the expert encoder?
A silent expert chess encoder is joined to an instruction-tuned language model through cross-attention. A question-answering curriculum teaches the model to extract chess concepts from the encoder's representations.