2026-10-11 17:10 UTC

Independent replication will determine whether Pathway's 150M-parameter recurrent latent-reasoning model achieves 29.5% on ARC-AGI-1 at roughly $0.0007 per task and materially outperforms transformer baselines on inference efficiency.

state: expiredheat: lowuncertainty: highconvergesscott: mediumrecurrent-reasoning inference-economics small-modelsPathway

What is this?

Pathway, which describes itself as an AI lab developing post-Transformer architectures, has announced BDH-CQ, a 150-million-parameter recurrent reasoning model that iterates in a latent workspace and updates memory from demonstrations rather than emitting chain-of-thought tokens. Pathway claims the model achieved 29.5% pass@2 on the public ARC-AGI-1 evaluation set at a computed inference cost of $0.0007 per task, with substantially lower cost than a higher-scoring comparison model. The supplied coverage largely repeats Pathway’s own announcement or a syndicated press release; it does not establish independent replication, and the snippets conflict on whether the claimed cost advantage is 11× or 57×.

Why it matters to Scott

Pathway’s claimed tokenless recurrent computation independently converges with Scott’s inference-time-scaling and machine-native-reasoning positions, while offering a potentially relevant architectural alternative to the explicit search used in AMA. If independently replicated, its unusually low task cost could affect his small-model routing and reasoning-system economics; for now, the seller-originated benchmark and conflicting cost multiplier keep it below high relevance.
ip:concept.inference-time-scalingip:framework.agent-native-computingdev:project.amaip:concept.evidence-class-ladderip:concept.ai-unit-economicsradar:nanbeige-4-2-3b-looped-transformerradar:ttt-discover-test-time-learningradar:concept.inference-economicsradar:concept.arc-agiradar:concept.benchmark-integrity
queries asked of Scott's wikis
  • recurrent latent reasoning vs token chain-of-thought
  • small-model reasoning and inference economics
  • test-time compute without generated reasoning tokens
  • ARC-AGI evaluation harnesses and replication
  • post-Transformer recurrent architectures
  • benchmark cost claims and reproducibility

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditA 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task
LocalLLaMA
juanviera2318639
🟠 redditA 150M param recurrent model scores 29.5% on ARC-AGI-1 at $0.0007 per task
singularity
juanviera2363962
🟧 echo.paper ⭐The paper’s abstract states: “A 150M-parameter configuration reaches 29.5% pass@2 at a computed inference cost of $0.0007 per task.” It intrBjörn Engdahl, Adrian Kosowski, Jan Chorowski, Zuzanna Stamirowska, Przemysław Uznański, Junlin Jiang, Rohan Phadke, Remigiusz Kinas, and Richard Zhong——
🟠 redditBDH-CQ: IN-CONTEXT LEARNING WITH RECURRENT LATENT REASONING [R]
MachineLearning
moschles6310
🟠 redditIntelligence per dollar is the new scaling law: A tiny reasoning model breaks the existing cost-accuracy Pareto frontier on Arc-AGI 1
artificial
dank_philosopher1112

Interpretation history

Decision trace