2026-10-11 17:10 UTC

Independent testing will determine whether Kimi Linear 48B-A3B provides practical 1M-context local inference at higher speed than comparable MoE models while retaining useful coding and frontend-generation quality.

state: expiredheat: lowuncertainty: highnovelscott: lowkimi-linear local-inference open-modelsKimi

What is this?

Kimi Linear 48B-A3B is a Kimi/Moonshot model described as having 48B total parameters but roughly 3B activated, with claims of up to 6× decoding throughput and support for contexts reaching 1 million tokens. One secondary source reports strong retrieval results beyond 512K tokens, but community discussion notes the absence of independent benchmarks, demos, and GGUF builds needed to verify practical local performance. The supplied snippets do not establish its coding or frontend-generation quality, so those remain testing questions rather than demonstrated capabilities.

Why it matters to Scott

No intersection found in Scott’s wikis or the radar’s accumulated pages. The model’s claimed long-context local-inference efficiency is broadly topical, but without independent benchmarks or evidence connecting it to Scott’s existing positions or projects, it is only a candidate for testing rather than a consequential update.
queries asked of Scott's wikis
  • linear attention for long-context agents
  • local inference economics for sparse MoE models
  • million-token context versus retrieval systems
  • GGUF quantization and Apple Silicon inference
  • coding-agent model evaluation harnesses
  • open-weight models and model sovereignty

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditKimi Linear 48B A3B?
LocalLLaMA
Atretador5819
🟧 echo.paper ⭐The original paper introduces Kimi Linear and states: “3B activated parameters and 48B total parameters,” with up to 6× decoding throughput Kimi Team——
🟧 hnKimi Linear: An Expressive, Efficient Attention Architectureronfriedhaber295125
🟠 redditKimi Linear: An Expressive, Efficient Attention Architecture
singularity
yogthos6315

Interpretation history

Decision trace