2026-10-11 16:38 UTC

Princeton PhD researcher Adithya Bhaskar claims architectural and algorithmic innovations โ€” a natural-language analogue of the AlphaZero algorithm โ€” trained a 4B LLM to 2700 Elo chess with accurate move explanations and no training plateau, and that the technique transfers to other games, robotics, and computer use; publication and independent replication of the Elo result and the cross-domain transfer would establish a new small-model RL training method, while failed replication deflates the claim.

state: watchingheat: lowuncertainty: mediumconvergesscott: highllm-training reinforcement-learning small-modelsAdithya BhaskarPrinceton University

What is this?

Adithya Bhaskar is a third-year Princeton NLP PhD student (advised by Danqi Chen, IIT Bombay undergrad) who announced QUEEN, a 4B-parameter chess-language model that plays at a level 'approaching a typical Grandmaster' and explains its moves in natural language. The bolder circulating specifics โ€” 2700 Elo with no training plateau, a natural-language analogue of AlphaZero, and transfer to other games, robotics, and computer use โ€” rest on the viral first-party posts; notably the author's own announcement phrasing ('approaching a typical Grandmaster') is more modest than the 2700 figure, and no published paper or independent replication appears in the supplied material (the Oct 5, 2026 arXiv listing snippet doesn't show it). Search confirms identity and the announcement only โ€” including a near-name collision with a different researcher (adhi.dev) that should not be conflated โ€” so the claim's standing rides entirely on the imminent paper and replication.

Why it matters to Scott

If replicated, a 4B model at GM-strength chess via verifiable-reward RL is a dated receipt for Scott's Model Barbell and verifier-first canon โ€” small specialists punching far above weight when the task has a real verifier โ€” and it lands in his own territory as a former computer-chess engine author (Chompster). His Explainability Trap canon supplies the distinctive test nobody else will run: the 'explains its moves accurately' claim is exactly where post-hoc confabulation should be probed before belief, and the robotics/computer-use transfer is precisely where the perfect chess verifier degrades.
ip:concept.model-barbellip:concept.explainability-trapip:concept.search-not-learningip:concept.mechanically-different-verifierswork:project.chompsterdev:concept.llm-self-play-refinementradar:concept.small-language-modelsradar:concept.post-trainingradar:concept.agentic-rlradar:microsoft-frognano-4b-releaseradar:qorl-small-model-postgres-planningradar:open-weight-masked-introspectionradar:dream-rsi-replay-exploration
queries asked of Scott's wikis
  • small specialized models vs frontier generality
  • RL with verifiable rewards for LLM post-training
  • AlphaZero-style self-play applied to language models
  • computer-use agent training environments
  • faithfulness of model-generated natural-language explanations
  • local inference economics of 4B-class models

Measured heat

now 0 pts/hpeak 37 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 171h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-04 13:00โญ origin echo-reconstructed"The most profound paper of my PhD so far. We truly did something special to make LMs play/explain chess, needing both architectural & algor
Adithya Bhaskar (@AdithyaNLP) on x (echo) ยท attributed from reddit.post.1wypjue
โ€”
10-06 01:02first on r/artificial ยท published ยท +36.0hPrinceton researchers train a 4B LLM to reach 2700 Elo in chess (with no signs of a plateau when they stopped training) and can explain its moves accurately. They say the training technique can also be applied to other games, robotics, and computer use
Eliv_nurotic
โ€”
10-06 10:03first on hacker news ยท published ยท +45.0hLanguage Models That Play Chess and Explain Their Moves
o4c
โ€”
10-06 01:02amplified on r/artificial ๐Ÿ‘‘reddit.post.1wypjue
Eliv_nurotic
peak 738 ยท 156 comments ยท 100% of case engagement
10-06 10:03amplified on hacker newshn.story.49976375
o4c
peak 2 ยท 0 comments ยท 0% of case engagement
10-06 03:20our radar first saw it ยท +38.4hdiscovery anchor: reddit.post.1wypjueโ€”
pace: p90 vs 1188 stories at the 168h mark (now 171h old) โ€” ahead of ai-graphene-simulator-claim (1.0x), behind swe-bench-pro-harness-cost-parity (1.0x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  redditPrinceton researchers train a 4B LLM to reach 2700 Elo in chess (with no signs of a plateau when they stopped training) and can explain its moves accurately. They say the training technique can also be applied to other games, robotics, and computer use
artificial
Retrieved article excerpt

Open article ยท Retrieved 2026-10-06T03:34:05.836462+00:00

[@AdithyaNLP](https://x.com/AdithyaNLP)

[Adithya Bhaskar](https://x.com/AdithyaNLP)[@AdithyaNLP](https://x.com/AdithyaNLP)

The most profound paper of my PhD so far. We truly did something special to make LMs play/explain chess, needing both architectural & algorithmic innovation (natural-language analogue of the Alphazero algorithm) + applicable to many other domains. Please read on and share!
1/10

[3:01 PM ยท Oct 5, 2026](https://x.com/AdithyaNLP/status/2107123924828049691)ยท[64.6K

Views](https://x.com/AdithyaNLP/status/2107123924828049691)

[36](https://x.com/compose/post?in_reply_to=2107123924828049691)

108

962

868
Eliv_nurotic738156
๐ŸŸง echo.x โญ"The most profound paper of my PhD so far. We truly did something special to make LMs play/explain chess, needing both architectural & algorAdithya Bhaskar (@AdithyaNLP)โ€”โ€”
๐ŸŸง hnLanguage Models That Play Chess and Explain Their Moveso4c20

Interpretation history

Decision trace