2026-10-11 17:12 UTC

Independent benchmarks will determine whether the C99 expert-streaming engine can run Kimi K3’s 1.56 TB checkpoint on a commodity CPU with 8 GB RAM and NVMe storage at practically usable speed.

state: expiredheat: lowuncertainty: highnovelscott: nonekimi-k3 local-inference expert-streaming

What is this?

Kimi K3 is a Moonshot AI model described as having roughly 2.78 trillion parameters and about 1.56 TB of weights, using a routed mixture-of-experts architecture. A repository claims a C99 expert-streaming engine can execute that checkpoint with one CPU, NVMe storage, and roughly 8 GB of RAM, but the supplied results provide no independent throughput or latency measurements establishing practical usability. The snippets also conflict on release status—one says the weights are not yet public and are promised for July 27, 2026, while others describe them as released—so both reproducibility and usable-speed claims remain unresolved.

Why it matters to Scott

No Scott wiki or radar intersection was found; the supplied hits provide no basis for connecting the unresolved expert-streaming performance claim to a position, project, or tracked case of Scott’s.
queries asked of Scott's wikis
  • expert streaming from NVMe for local inference
  • RAM versus storage bandwidth in mixture-of-experts inference
  • minimum viable speed for local LLM inference
  • commodity CPU inference economics
  • open-weight model hardware sovereignty
  • independent benchmarking of inference-engine claims

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (6) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditI pushed Kimi K3 onto one CPU with 8 GB of RAM
LocalLLaMA
FareedKhan557716129
🟧 echo.github ⭐The repository documents and implements the claim: “A 2.78-trillion-parameter model. One CPU. 8 GB of RAM.” It reports 8.24 GB peak RSS, a 1Fareed Khan——
🟧 hnAirLLM: Inference 2.8T Kimi K3 on a single 4GB GPUmaxloh10
🟠 redditGitHub - sqliteai/waste: Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.
LocalLLaMA
ab237713349
🟧 hnAirLLM 70B inference with single 4GB GPUAnon8420776
🟧 hnSmaller, faster, safer: running Kimi and GLM at scaleascorbic24360

Interpretation history

Decision trace