BAAI claims its open 27B AREX-2 (Qwen3.8-based) learns test-time self-improvement β propose, measure, reflect, revise on verifiable-feedback ML/algorithmic tasks β that transfers to deep research without new search trajectories; sustained independent adoption and measurements in real long-horizon agent workflows settle whether it is a durable open agent model rather than another release that fades.
state: watchingheat: lowuncertainty: highconvergesscott: mediumopen-models long-horizon-agentsBAAI
What is this?
AREX-2 is a 27B-parameter long-horizon agent model released openly on Hugging Face by the Beijing Academy of Artificial Intelligence (BAAI), built on Qwen3.8 27B and trained on machine-learning and algorithmic-programming tasks with verifiable feedback plus BAAI's earlier AREX deep-research data; its headline claim is that a learned test-time loop β propose, measure, reflect, revise β transfers to deep research without adding new search trajectories. It descends from the BAAI paper 'AREX: Towards a Recursively Self-Improving Agent for Deep Research' (arXiv 2607.21461), whose key mechanism is a learned context-update tool that compresses a growing interaction history into a compact improvement state holding verified evidence, constraint-satisfaction status, unresolved gaps and a next research plan. Reception per the case file has been tepid: the r/LocalLLaMA launch thread decayed to near-zero within two days, top comments flagged the missing base-model-in-same-loop comparison, and the only uptake datum is one unverified first-person endorsement for agentic coding offering no benchmarks. The supplied snippets contain no independent evaluation β the model card's numbers explicitly follow BAAI's own paper protocols β leaving two faint leads: a September arXiv page (2609.38288) whose only visible content is 'AREX-2 scored at 70' (no authorship or task context, so independence can't be judged from this material), and a third-party same-harness benchmark of Qwen3.8-27B itself against a compressed derivative (Prism ML's Bonsai 2), which shows the base-model class is actively benchmarked by the community but does not test AREX-2.
Why it matters to Scott
Converges where Scott already argues from the harness side: AREX's learned context-update tool β compressing a growing interaction history into a compact verified-evidence state β is his agent-authored context compaction, and its propose-measure-reflect-revise loop is his review-until-clear loop, but BAAI claims both as trained-in weights behavior that transfers, posing the harness-vs-weights question directly against his own patterns. Because AREX-2 is a 27B Qwen3.8 finetune it is benchable on his gamepc/Ollama stack, so the decisive missing measurement the community is asking for (base Qwen3.8-27B in the identical loop) is one he could produce himself with his trace-backed comparison fixture β a dated-receipts opportunity β but with no independent eval and only one unverified adoption report, it remains a watching-level bet.
dev:concept.agent-authored-context-compactiondev:concept.review-until-clear-loopdev:concept.trace-backed-agent-comparisondev:technology.ollamaradar:qwen38-27b-local-agent-capabilityradar:qwen38-27b-16gb-quant-benchmarkradar:concept.recursive-self-improvementradar:concept.self-improving-agentsradar:concept.long-horizon-agents
queries asked of Scott's wikis
- agent memory compaction learned state editing
- test-time compute self-improvement loop
- verifiable rewards open-ended task evaluation limits
- 27B local inference Ollama LiteLLM agent runtime
- Chinese open-weight lab release sovereignty
- long-horizon deep research agent harness benchmark
Measured heat
now 0 pts/hpeak 28 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 272h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
pace: p67 vs 1188 stories at the 168h mark (now 272h old) β ahead of agentic-retrieval-frames-benchmark (1.0x), behind anthropic-blocked-request-billing (1.0x)
Evidence (3) β β canonical anchor
Interpretation history
2026-10-03T20:59:23Z
grounded: converges/medium β Converges where Scott already argues from the harness side: AREX's learned context-update tool β compressing a growing interaction history into a compact verifi
2026-10-03T20:50:16Z
The trigger re-flagged the already-attached adoption post (only +1 pt/+2 comments since), so no new substance β but with the testimony now on record the case's meaning shifts from bare unverified claim to unverified claim with a first, faint uptake signal, justifying seedβwatching. The decisive watch item is unchanged: an independent benchmark, ideally base Qwen3.8-27B in the identical loop; not material since the endorsement was already priced.
2026-10-03T20:26:17Z
evidence attached: reddit.post.1wwwgq7 β Weak but on-hypothesis: a first-person adoption endorsement of AREX-2 for agentic coding, the exact uptake signal this case is watching for, though it offers no measurements.
2026-10-02T07:53:42Z
The early ~10x velocity spike was a brief burst that has fully decayed (0 pts/h at ~47h, 25th percentile); the comment tail added only shallow skepticism β missing base-model-in-same-loop comparison, no MTP head β and no independent measurement, adoption, or implementation has surfaced. Meaning unchanged: an unverified first-party transfer claim whose fate now rests entirely on future third-party evals.
2026-09-30T09:42:32Z
grounded: converges/high β BAAI's AREX design β an outer audit over a closed claim set and a learned context-update tool compressing long interaction histories into a compact improvement
2026-09-30T09:34:17Z
case created β Distinct first-party release with a specific test-time self-improvement transfer claim no open case covers; adoption, not the announcement, is the open question.
Decision trace
- 10-05 20:40review_screenjev screen: no material development (noul=0.17)
- 10-04 07:59repriceThe trigger re-flagged the already-attached adoption post (only +1 pt/+2 comments since), so no new substance β but with the testimony now on record the case's meaning shifts from bare unverified
- 10-04 07:59groundConverges where Scott already argues from the harness side: AREX's learned context-update tool β compressing a growing interaction history into a compact verified-evidence state β is his agent-au
- 10-04 07:26attachWeak but on-hypothesis: a first-person adoption endorsement of AREX-2 for agentic coding, the exact uptake signal this case is watching for, though it offers no measurements.
- 10-04 07:24propose_attachWeak but on-hypothesis: a first-person adoption endorsement of AREX-2 for agentic coding, the exact uptake signal this case is watching for, though it offers no measurements.
- 10-02 17:53repriceThe early ~10x velocity spike was a brief burst that has fully decayed (0 pts/h at ~47h, 25th percentile); the comment tail added only shallow skepticism β missing base-model-in-same-loop comparison,
- 10-01 01:21sensor_dirtycomment_update
- 09-30 20:21sensor_dirtyvelocity_spike
- 09-30 19:42groundBAAI's AREX design β an outer audit over a closed claim set and a learned context-update tool compressing long interaction histories into a compact improvement state β independently lands on his
- 09-30 19:34createDistinct first-party release with a specific test-time self-improvement transfer claim no open case covers; adoption, not the announcement, is the open question.