Independent evaluation will determine whether TTT-Discover enables language models to learn useful discovery procedures during inference and outperform fixed inference-time reasoning approaches.
state: expiredheat: lowuncertainty: highconvergesscott: hightest-time-training frontier-modelsUC Berkeley Sky Computing Lab
What is this?
TTT-Discover is a method that continues training an LLM during inference, using reinforcement learning, adaptive objectives, and search to specialize on a single difficult problem rather than merely sample from a frozen model. The authors report state-of-the-art results across mathematics, GPU-kernel engineering, algorithm design, and biology, including a new construction for an Erdős minimum-overlap bound, though the supplied evidence does not include independent replication. The work appears as a January 2026 arXiv paper with an MIT-licensed implementation; the snippets list authors including Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun, while the case associates it with UC Berkeley Sky Computing Lab.
Why it matters to Scott
TTT-Discover independently converges with Scott’s view that evaluated adaptive loops can improve capability during inference, but introduces weight updates as a materially different route from his frozen-model, search, and durable-external-state architectures. Independent results could therefore affect both his architectural argument about where learning should live and the economics of spending substantial compute to specialize on one problem.
ip:concept.inference-time-scalingip:source.the-third-substrate-ebookip:concept.search-not-learningip:concept.self-improving-loopsip:concept.cost-of-cognitionradar:concept.reinforcement-learningradar:concept.inference-efficiencyradar:concept.agentic-rlradar:concept.inference-economics
queries asked of Scott's wikis
- test-time learning versus fixed inference-time scaling
- agents that update model weights during task execution
- verifier-guided reinforcement learning for coding and discovery
- specialization after generalization in frontier models
- persistent agent memory versus weight-based adaptation
- economics of per-problem inference-time training
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
| source | object | author | score | comments |
| 🟧 hn | TTT-Discover: Learning to Discover at Test Time | matt_d | 1 | 0 |
| 🟧 echo.paper ⭐ | The original primary artifact is the authors’ arXiv paper, submitted January 22, 2026. It introduces TTT-Discover, which performs reinforcem | Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun | — | — |
Interpretation history
2026-08-17T20:31:49Z
More than six months after the primary release, no independent evaluation, replication, or substantive implementation has surfaced, and the weak resurfacing produced no new evidence. The claim remains unresolved but has faded as an active episode and can be reopened if external results appear.
2026-08-15T20:30:20Z
The staleness check found no independent evaluation, replication, or substantive implementation evidence, so the case remains a potentially important but untested architectural claim. The hot frontier-model context does not raise this episode’s maturity or urgency.
2026-08-13T19:42:45Z
The reobservation adds only negligible engagement and no independent evaluation, replication, or implementation result. The architectural implications remain potentially important, but the authors’ capability and economics claims are still untested externally.
2026-08-13T19:32:13Z
grounded: converges/high — TTT-Discover independently converges with Scott’s view that evaluated adaptive loops can improve capability during inference, but introduces weight updates as a
2026-08-13T19:30:02Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49290619 -> echo.paper.9017047455 by Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun
2026-08-13T19:29:09Z
case created — The first-party research release presents a bounded test-time-learning method with potentially material implications for inference-time adaptation.
Decision trace
- 08-18 06:31expireMore than six months after the primary release, no independent evaluation, replication, or substantive implementation has surfaced, and the weak resurfacing produced no new evidence. The claim remains
- 08-18 06:31alert_silentThe staleness trigger adds no consequential fact, and repeated observation without external validation does not merit Scott’s attention; reopen only on an independent benchmark, replication, or implem
- 08-18 06:31alert_routeThe staleness trigger adds no consequential fact, and repeated observation without external validation does not merit Scott’s attention; reopen only on an independent benchmark, replication, or implem
- 08-16 06:30repriceThe staleness check found no independent evaluation, replication, or substantive implementation evidence, so the case remains a potentially important but untested architectural claim. The hot frontier
- 08-16 06:30alert_silentNo consequential delta occurred; another observation of the authors’ existing claims would not justify interrupting Scott before an independent benchmark, replication, or implementation result appears
- 08-16 06:30alert_routeNo consequential delta occurred; another observation of the authors’ existing claims would not justify interrupting Scott before an independent benchmark, replication, or implementation result appears
- 08-14 05:42repriceThe reobservation adds only negligible engagement and no independent evaluation, replication, or implementation result. The architectural implications remain potentially important, but the authors’ ca
- 08-14 05:42alert_silentNo consequential new event occurred; this is repetitive exposure of the existing paper and can wait until an independent benchmark, replication, or substantive implementation appears.
- 08-14 05:42alert_routeNo consequential new event occurred; this is repetitive exposure of the existing paper and can wait until an independent benchmark, replication, or substantive implementation appears.
- 08-14 05:37alert_silentThe HN post only points to the authors’ January 2026 paper; it adds no independent evaluation, implementation release, or new result beyond the original TTT-Discover claims. Retain the case and wait f
- 08-14 05:37alert_routeThe HN post only points to the authors’ January 2026 paper; it adds no independent evaluation, implementation release, or new result beyond the original TTT-Discover claims. Retain the case and wait f
- 08-14 05:32groundTTT-Discover independently converges with Scott’s view that evaluated adaptive loops can improve capability during inference, but introduces weight updates as a materially different route from his fro
- 08-14 05:30promote_anchororigin walk conf 0.99
- 08-14 05:29createThe first-party research release presents a bounded test-time-learning method with potentially material implications for inference-time adaptation.