Kimi K3 is a 2.8-trillion-parameter, highly sparse mixture-of-experts model from Moonshot AI, featuring native vision, a one-million-token context window, Kimi Delta Attention, and Attention Residuals. vLLM announced preview serving support developed with Moonshot AI, NVIDIA, AMD, and the open-source community, with the evidence titles reporting performance of up to 370 tokens per second. However, the supplied snippets describe integration, validation, and optimization as still ongoing; they do not substantiate the web answerβs claim that independent deployments have already confirmed stable throughput at that rate.
2026-07-29T16:29:47Z
No independent vLLM deployment, stability report, or reproduction of the 370 tokens/sec claim has materialised after two days of adjacent tooling and local-serving interest. The core hypothesis remains untested and the window for timely validation has closed; further anticipation is repetitive.
2026-07-29T15:31:11Z
New evidence shows users attempting local deployment and reporting self-hosting cost/quality tradeoffs, but still no independent vLLM deployment, stability report, or reproduction of the 370 tokens/sec claim. The case remains a cold validation watch β the core hypothesis is untested.
2026-07-29T15:21:39Z
evidence attached: hn.story.49098130 β Independent self-hosting results on hardware cost and task resolution materially bear on whether Kimi K3 can be served practically outside hosted APIs.
2026-07-29T15:21:39Z
evidence attached: hn.story.49098395 β This is direct evidence that users are attempting practical local Kimi K3 deployment, bearing on the open serving-validation case.
2026-07-29T14:28:20Z
The new community inference runner (Colibri-based) adds tooling evidence for local Kimi K3 serving but still provides no benchmark, vLLM stability report, or reproduction of the 370 tokens/sec claim. The case remains a cold validation watch β no operator results have materialized to test the core hypothesis.
2026-07-29T14:21:43Z
evidence attached: reddit.post.1v9w5qw β A community inference runner provides additional evidence about the practical effort and tooling needed to serve Kimi K3 locally, though it offers no benchmark yet.
2026-07-28T20:28:13Z
The evidence still provides no independent vLLM deployment, stability report, or reproducible throughput near 370 tokens/sec; adjacent llama.cpp and Atom implementations do not test the core claim. This remains a cold validation watch, with further attention unwarranted until an operator benchmark appears.
2026-07-28T17:30:10Z
The evidence still lacks an independent vLLM deployment, stability report, or reproducible result approaching 370 tokens/sec; adjacent llama.cpp and Atom activity only supports broader serveability. Further engagement without operator benchmarks is repetitive amplification, so keep this cold pending substantive validation.
2026-07-28T10:26:15Z
No new evidence independently tests vLLM stability or reproduces the reported 370 tokens/sec; the AMD Atom deployment remains adjacent evidence of general serveability. The case stays a cold validation watch until a reproducible vLLM benchmark or operator report arrives.
2026-07-28T07:29:48Z
The independent Atom-on-AMD deployment supports Kimi K3βs broader serveability but remains orthogonal to vLLM stability and the reported 370 tokens/sec. No new reproducible vLLM benchmark changes the case, so keep it as a cold validation watch.
2026-07-28T05:24:42Z
The Atom deployment on AMD independently shows that Kimi K3 can run on high-end accelerator infrastructure, making the broader serving story more credible. It does not test vLLM stability or reproduce the reported 370 tokens/sec, so the core validation hypothesis remains open.
2026-07-28T05:20:45Z
evidence attached: hn.story.49079363 β Independent AMD deployment evidence strengthens the broader case that Kimi K3 can be served practically on high-end accelerator infrastructure, though it uses Atom rather than vLLM.
2026-07-28T04:24:53Z
The attached evidence still does not provide an independent vLLM deployment, stability report, or reproducible throughput benchmark; llama.cpp interest remains adjacent rather than validating the 370 tokens/sec claim. This is repetitive amplification, so the case remains a cold implementation watch pending substantive operator results.
2026-07-28T02:23:47Z
No substantive operator result has arrived to validate vLLM stability or reproduce the 370 tokens/sec claim; the llama.cpp activity remains adjacent implementation interest. Repeated anticipation adds no meaning, so wait for benchmarks or deployment reports.
2026-07-28T01:21:51Z
The attached evidence still contains no independent vLLM deployment, stability report, or reproducible throughput benchmark; llama.cpp interest remains adjacent anticipation. Further engagement without operator results is repetitive amplification, so defer review until substantive testing appears.
2026-07-28T00:22:34Z
No independent vLLM deployment, stability report, or reproducible throughput benchmark has arrived; the attached llama.cpp activity remains adjacent anticipation rather than validation. Repeated engagement adds no new meaning, so this stays a cold implementation watch.
2026-07-27T22:27:52Z
The newly attached material still offers no independent vLLM deployment, stability report, or reproducible throughput benchmark; llama.cpp testing intent is adjacent rather than validating the 370 tokens/sec claim. Repeated anticipation adds no new meaning, so this remains a cold implementation watch.
2026-07-27T21:27:22Z
The attached material still shows implementation availability and testing intent, not an independent vLLM deployment or reproducible throughput benchmark. The hot Kimi K3 neighborhood does not change this caseβs meaning; it remains a cold validation watch.
2026-07-27T20:27:04Z
The added activity remains anticipation around llama.cpp rather than an independent vLLM deployment or reproducible throughput result. It adds no validation of stability or the 370 tokens/sec claim, so the case stays a cold implementation watch.
2026-07-27T19:26:06Z
No independent deployment result yet tests vLLM stability or the reported 370 tokens/sec; the visible activity remains anticipation around the llama.cpp port rather than validation. Keep this as a cold implementation watch until reproducible benchmarks or operator reports arrive.
2026-07-27T18:22:08Z
The llama.cpp port broadens practical serving interest but does not validate vLLM stability or the reported 370 tokens/sec. With downloads and testing still pending and no fresh engagement, this is now an implementation watch rather than an active performance signal.
2026-07-27T18:21:31Z
evidence attached: reddit.post.1v87v71 β Llama.cpp support for text-only Kimi K3 materially advances the open local-serving episode, though independent performance validation is still pending.
2026-07-27T16:23:44Z
grounded: novel/none β No intersection found: neither Scottβs wikis nor the radar contain pages connecting his positions or projects to Kimi K3βs vLLM serving performance or the need
2026-07-27T16:23:09Z
case created β First-party serving support and a specific throughput claim create a distinct deployment-validation episode not covered by the existing Kimi K3 cases.