2026-10-11 17:12 UTC

Independent evaluations will determine whether Nanbeige4.2-3B's looped-transformer architecture delivers unusually strong agentic and coding performance for a 3B model.

state: expiredheat: lowuncertainty: mediumknownscott: mediumlooped-transformers small-models agentic-models coding-models open-modelsNanbeige

What is this?

Nanbeige is behind a family of open 3B-parameter language models aimed at strong reasoning, coding, and agentic performance despite their small size. Supplied sources describe Nanbeige4 and Nanbeige4.1-3B as decoder-only models using extensive training, multi-stage fine-tuning, preference distillation, and reward modeling, with reported results exceeding some larger models. However, the snippets do not independently establish the existence, architecture, or benchmark performance of Nanbeige4.2-3B; they only separately describe looped transformers as architectures that reuse blocks to trade additional compute for lower parameter memory.

Why it matters to Scott

This repeats Scott’s established Capability Audit position—and the radar’s existing looped-transformer validation case—that vendor-reported capability and efficiency claims require independent testing. A genuinely strong 3B coding/agentic model could affect his local-inference and agent-harness work, but the supplied evidence does not yet establish Nanbeige4.2-3B’s existence, architecture, availability, or performance, so there is currently nothing actionable beyond monitoring.
ip:concept.capability-auditip:concept.evaluation-driven-developmentdev:concept.hardware-aware-local-inferenceradar:looping-20b-token-efficient-pretrainingradar:concept.local-inferenceradar:concept.open-models
queries asked of Scott's wikis
  • looped transformers recurrent depth compute tradeoffs
  • small local models for coding agents
  • independent evaluation of agentic model claims
  • parameter count versus inference-time compute
  • open-weight model sovereignty and local inference
  • coding-agent harnesses for small models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (10) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditNew Model: Nanbeige4.2-3B (Looped Transformer, outperforms 4x size)
LocalLLaMA
Wooden-Deer-1276462164
🟧 echo.other ⭐The official model card introduces Nanbeige4.2-3B as a compact agentic model whose “Looped Transformer architecture reuses the transformer lNanbeige LLM Lab——
🟠 redditNanbeige4.2-3B drops: 3B params claiming to beat 9B/12B models on agentic tasks (atleast according to them)
LocalLLaMA
UsedMorning98865613
🟠 reddit[Model] Add support for Nanbeige4.2 by zqlcode · Pull Request #25994 · ggml-org/llama.cpp
LocalLLaMA
pmttyji395
🟠 redditNanbeige4.2-3B is sad really
LocalLLaMA
TechTefa011
🟠 redditNanbeige 4.2 3B Garbage Output
LocalLLaMA
LaurentPayot013
🟠 redditNanbeige4.2-3B: I'm not impressed
LocalLLaMA
crusaderky1831
🟧 hnAgentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1gmays10
🟠 redditTested Nanbeige4.2 3B vs Gemma 4 (12B) & Qwen3.5 9B | Coding with OpenCode, Tool Calling & Reasoning
LocalLLaMA
curiousily_21
🟠 redditFiguring out benchmaxxing.
LocalLLaMA
Witty_Mycologist_995018

Interpretation history

Decision trace