2026-10-11 17:11 UTC

Independent evaluation will determine whether TTT-Discover enables language models to learn useful discovery procedures during inference and outperform fixed inference-time reasoning approaches.

state: expiredheat: lowuncertainty: highconvergesscott: hightest-time-training frontier-modelsUC Berkeley Sky Computing Lab

What is this?

TTT-Discover is a method that continues training an LLM during inference, using reinforcement learning, adaptive objectives, and search to specialize on a single difficult problem rather than merely sample from a frozen model. The authors report state-of-the-art results across mathematics, GPU-kernel engineering, algorithm design, and biology, including a new construction for an Erdős minimum-overlap bound, though the supplied evidence does not include independent replication. The work appears as a January 2026 arXiv paper with an MIT-licensed implementation; the snippets list authors including Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun, while the case associates it with UC Berkeley Sky Computing Lab.

Why it matters to Scott

TTT-Discover independently converges with Scott’s view that evaluated adaptive loops can improve capability during inference, but introduces weight updates as a materially different route from his frozen-model, search, and durable-external-state architectures. Independent results could therefore affect both his architectural argument about where learning should live and the economics of spending substantial compute to specialize on one problem.
ip:concept.inference-time-scalingip:source.the-third-substrate-ebookip:concept.search-not-learningip:concept.self-improving-loopsip:concept.cost-of-cognitionradar:concept.reinforcement-learningradar:concept.inference-efficiencyradar:concept.agentic-rlradar:concept.inference-economics
queries asked of Scott's wikis
  • test-time learning versus fixed inference-time scaling
  • agents that update model weights during task execution
  • verifier-guided reinforcement learning for coding and discovery
  • specialization after generalization in frontier models
  • persistent agent memory versus weight-based adaptation
  • economics of per-problem inference-time training

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnTTT-Discover: Learning to Discover at Test Timematt_d10
🟧 echo.paper ⭐The original primary artifact is the authors’ arXiv paper, submitted January 22, 2026. It introduces TTT-Discover, which performs reinforcemMert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun——

Interpretation history

Decision trace