2026-10-11 17:09 UTC

Independent evaluations will determine whether Kwaipilot’s 35B-total, 3B-active KAT-Coder-V2.5-Dev delivers competitive agentic-coding and tool-use performance among similarly sized open-weight models.

state: resolvedheat: lowuncertainty: mediumknownscott: mediumopen-models coding-agents mixture-of-expertsKwaipilot

What is this?

Kwaipilot released KAT-Coder-V2.5-Dev on Hugging Face as an open-weight mixture-of-experts coding model with 35B total parameters and 3B active parameters, claiming state-of-the-art agentic-coding results. The supplied web results do not substantiate those claims, identify any independent evaluations, or establish comparative tool-use performance, so its competitiveness remains unverified in this material.

Why it matters to Scott

The need for independent, harness-level validation is already explicit in Scott’s Evaluation-Driven Development position, while Ask and gamepc make a capable 3B-active open coding model a plausible backend candidate. It matters operationally if real repository and tool-loop tests validate the claims, but the supplied material currently offers only an unverified launch rather than evidence that should change his model-routing choices.
ip:concept.evaluation-driven-developmentip:source.give-the-agent-a-workshop-ebookdev:project.askdev:project.gamepcdev:concept.hardware-aware-local-inferenceradar:concept.open-modelsradar:concept.coding-modelsradar:concept.coding-agentsradar:concept.ai-benchmarksradar:concept.inference-efficiency
queries asked of Scott's wikis
  • sparse MoE economics for local coding agents
  • independent evaluation of agentic coding models
  • coding-agent benchmarks versus real repository work
  • open-weight models for tool use and coding harnesses
  • active parameters versus total parameters in deployment
  • small open models as coding-agent backends

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditKwaipilot/KAT-Coder-V2.5-Dev · Hugging Face
LocalLLaMA
jacek202310549
🟧 echo.blog ⭐Kwaipilot released KAT-Coder-V2.5-Dev as a 35B-total, 3B-active open-weight MoE model and claimed state-of-the-art agentic-coding results amKwaipilot——
🟠 redditKat Coder 2.5 is insane. Especially considering I ran it at Q4_K_M
LocalLLaMA
ConfidentDinner664819469
🟠 redditI ran the 35B agentic comparison someone asked for (stock vs Ornith vs KAT-Coder, 120 runs)
LocalLLaMA
IvGranite3724
🟠 redditKAT Coder 2.5 dev: Do yourself a favor and try it!
LocalLLaMA
The_Paradoxy11689

Interpretation history

Decision trace