2026-10-11 17:10 UTC

llama.cpp will merge Kimi K3 support, and independent testing will determine whether it enables correct and practical local inference across common hardware configurations.

state: expiredheat: lowuncertainty: mediumknownscott: mediumkimi-k3 llama-cpp local-inferenceggml-orgpwilkin

What is this?

Contributor pwilkin has opened llama.cpp PR #26185 to add text-model support for Kimi K3; the supplied Hugging Face snippet says the code remains unmerged, while an Unsloth fork builds on the PR to add vision support. Earlier llama.cpp analysis found that K3’s architecture could not safely use the existing Kimi mapping because of its distinctive attention, MoE, metadata, and tensor structure. GGUF loading has reportedly been achieved, but the snippets provide no downstream benchmark validation, and K3’s enormous size—plus a reported recommendation of 64 or more accelerators—leaves practical inference on common local hardware unestablished.

Why it matters to Scott

Known via dev:concept.hardware-aware-local-inference and dev:project.gamepc: Scott already treats runtime support, hardware placement, memory pressure, precision, and measured output fidelity as separate requirements for practical local inference. The PR could extend his llama.cpp/GGUF model options, but it is unmerged and K3’s scale may keep it outside his workstation’s useful operating envelope, so independent capability testing matters before it changes what he builds.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.capability-auditip:concept.evaluation-driven-developmentip:concept.usable-mass-over-unusable-powerradar:concept.llama-cppradar:concept.local-inferenceradar:concept.ggufradar:concept.model-evaluationradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
  • local inference hardware economics and memory limits
  • llama.cpp and GGUF runtime strategy
  • open-weight models versus practical deployability
  • quantization fidelity and benchmark validation
  • local model support in coding-agent stacks
  • CPU/GPU hybrid inference across commodity hardware

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditmodel: add Kimi-K3 text model by pwilkin · Pull Request #26185 · ggml-org/llama.cpp
LocalLLaMA
pmttyji572
🟧 echo.github ⭐The pull request adds Kimi K3 text-model support to llama.cpp.pwilkin——
🟠 redditI hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens
LocalLLaMA
OtherRaisin342627476
🟠 redditImplementing Kimi K3 from scratch in PyTorch [P]
MachineLearning
Winter_Mistake_3185776

Interpretation history

Decision trace