2026-10-11 16:38 UTC

BAAI claims its open 27B AREX-2 (Qwen3.8-based) learns test-time self-improvement β€” propose, measure, reflect, revise on verifiable-feedback ML/algorithmic tasks β€” that transfers to deep research without new search trajectories; sustained independent adoption and measurements in real long-horizon agent workflows settle whether it is a durable open agent model rather than another release that fades.

state: watchingheat: lowuncertainty: highconvergesscott: mediumopen-models long-horizon-agentsBAAI

What is this?

AREX-2 is a 27B-parameter long-horizon agent model released openly on Hugging Face by the Beijing Academy of Artificial Intelligence (BAAI), built on Qwen3.8 27B and trained on machine-learning and algorithmic-programming tasks with verifiable feedback plus BAAI's earlier AREX deep-research data; its headline claim is that a learned test-time loop β€” propose, measure, reflect, revise β€” transfers to deep research without adding new search trajectories. It descends from the BAAI paper 'AREX: Towards a Recursively Self-Improving Agent for Deep Research' (arXiv 2607.21461), whose key mechanism is a learned context-update tool that compresses a growing interaction history into a compact improvement state holding verified evidence, constraint-satisfaction status, unresolved gaps and a next research plan. Reception per the case file has been tepid: the r/LocalLLaMA launch thread decayed to near-zero within two days, top comments flagged the missing base-model-in-same-loop comparison, and the only uptake datum is one unverified first-person endorsement for agentic coding offering no benchmarks. The supplied snippets contain no independent evaluation β€” the model card's numbers explicitly follow BAAI's own paper protocols β€” leaving two faint leads: a September arXiv page (2609.38288) whose only visible content is 'AREX-2 scored at 70' (no authorship or task context, so independence can't be judged from this material), and a third-party same-harness benchmark of Qwen3.8-27B itself against a compressed derivative (Prism ML's Bonsai 2), which shows the base-model class is actively benchmarked by the community but does not test AREX-2.

Why it matters to Scott

Converges where Scott already argues from the harness side: AREX's learned context-update tool β€” compressing a growing interaction history into a compact verified-evidence state β€” is his agent-authored context compaction, and its propose-measure-reflect-revise loop is his review-until-clear loop, but BAAI claims both as trained-in weights behavior that transfers, posing the harness-vs-weights question directly against his own patterns. Because AREX-2 is a 27B Qwen3.8 finetune it is benchable on his gamepc/Ollama stack, so the decisive missing measurement the community is asking for (base Qwen3.8-27B in the identical loop) is one he could produce himself with his trace-backed comparison fixture β€” a dated-receipts opportunity β€” but with no independent eval and only one unverified adoption report, it remains a watching-level bet.
dev:concept.agent-authored-context-compactiondev:concept.review-until-clear-loopdev:concept.trace-backed-agent-comparisondev:technology.ollamaradar:qwen38-27b-local-agent-capabilityradar:qwen38-27b-16gb-quant-benchmarkradar:concept.recursive-self-improvementradar:concept.self-improving-agentsradar:concept.long-horizon-agents
queries asked of Scott's wikis
  • agent memory compaction learned state editing
  • test-time compute self-improvement loop
  • verifiable rewards open-ended task evaluation limits
  • 27B local inference Ollama LiteLLM agent runtime
  • Chinese open-weight lab release sovereignty
  • long-horizon deep research agent harness benchmark

Measured heat

now 0 pts/hpeak 28 pts/hcomments 0/hpeers p0momentum: steady2 platformsage 272h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-30 09:34 (minted)⭐ origin echo-reconstructed"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence... It learns to improve a solution o
BAAI (Beijing Academy of Artificial Intelligence) on github (echo) Β· attributed from reddit.post.1wtzedu Β· published time unknown
β€”
09-30 08:21first on r/LocalLLaMA Β· published Β· lag ?BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B
Skyline34rGt
β€”
09-30 08:21amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wtzedu
Skyline34rGt
peak 79 Β· 29 comments Β· 95% of case engagement
10-03 19:59amplified on r/LocalLLaMAreddit.post.1wwwgq7
Ok-Importance-3529
peak 1 Β· 5 comments Β· 5% of case engagement
09-30 09:20our radar first saw it Β· lag ?discovery anchor: reddit.post.1wtzeduβ€”
pace: p67 vs 1188 stories at the 168h mark (now 272h old) β€” ahead of agentic-retrieval-frames-benchmark (1.0x), behind anthropic-blocked-request-billing (1.0x)

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditBAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B
LocalLLaMA
Skyline34rGt7929
🟧 echo.github ⭐"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence... It learns to improve a solution oBAAI (Beijing Academy of Artificial Intelligence)β€”β€”
🟠 redditArex-2 vs swift vs qwen3.8 27b
LocalLLaMA
Ok-Importance-352905

Interpretation history

Decision trace