2026-10-11 17:11 UTC

Independent deployments will determine whether vLLM's Kimi K3 support enables stable high-throughput serving at performance approaching the reported 370 tokens per second.

state: expiredheat: lowuncertainty: highnovelscott: nonekimi-k3 vllm llm-inferencevLLMKimi

What is this?

Kimi K3 is a 2.8-trillion-parameter, highly sparse mixture-of-experts model from Moonshot AI, featuring native vision, a one-million-token context window, Kimi Delta Attention, and Attention Residuals. vLLM announced preview serving support developed with Moonshot AI, NVIDIA, AMD, and the open-source community, with the evidence titles reporting performance of up to 370 tokens per second. However, the supplied snippets describe integration, validation, and optimization as still ongoing; they do not substantiate the web answer’s claim that independent deployments have already confirmed stable throughput at that rate.

Why it matters to Scott

No intersection found: neither Scott’s wikis nor the radar contain pages connecting his positions or projects to Kimi K3’s vLLM serving performance or the need for independent throughput validation.
queries asked of Scott's wikis
  • open-weight frontier model serving strategy
  • high-throughput LLM inference benchmarks
  • vLLM production deployment and optimization
  • independent validation of vendor performance claims
  • sparse MoE inference economics
  • self-hosted models and data sovereignty

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (7) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnKimi K3 on vLLM: Up to 370 Tokens/secwskwon10
🟧 echo.blog ⭐Announces Kimi K3 serving support in vLLM and reports performance of up to 370 tokens per second.vLLMβ€”β€”
🟠 redditKimi K3 text-only for llama.cpp
LocalLLaMA
ilintar7743
🟧 hnDay 0 Kimi-K3 Inference Deployment with Atom on AMD Instinct MI355X GPUsstevefan199940
🟠 redditThis may help y'all to run Kimi GGUFs on workstation class hardware
LocalLLaMA
Responsible_Fig_127108
🟧 hnRun Kimi K3 on a local computermarcobambini20
🟧 hnSelf-hosting Kimi K3: 20% more hardware cost, 20% better task resolutionflifenstein12141

Interpretation history

Decision trace