Kimi K3 is a Moonshot AI model described as having roughly 2.78 trillion parameters and about 1.56 TB of weights, using a routed mixture-of-experts architecture. A repository claims a C99 expert-streaming engine can execute that checkpoint with one CPU, NVMe storage, and roughly 8 GB of RAM, but the supplied results provide no independent throughput or latency measurements establishing practical usability. The snippets also conflict on release status—one says the weights are not yet public and are promised for July 27, 2026, while others describe them as released—so both reproducibility and usable-speed claims remain unresolved.
2026-08-04T04:32:57Z
No direct throughput benchmark or completed reproduction has emerged after sustained monitoring; the added activity only repeats established architectural feasibility. The practical-speed claim has faded without validation, so retire the episode unless a measured replication or materially faster implementation appears.
2026-08-04T03:24:00Z
The new attachment does not add an independent commodity CPU/NVMe throughput result or completed reproduction, so practical usability remains unsupported despite corroborated architectural feasibility. The case is informationally exhausted; revisit only if direct benchmarks or a materially faster implementation emerge.
2026-08-04T02:28:02Z
The latest attachment still adds no independent commodity CPU/NVMe throughput measurement or completed reproduction, so practical usability remains unsupported despite corroborated architectural feasibility. The case is informationally exhausted and should remain dormant until direct benchmarks or a materially faster implementation appear.
2026-08-04T01:23:07Z
No independent commodity CPU/NVMe throughput result or completed reproduction has emerged; the additional activity does not move the case beyond established architectural feasibility. Practical usability remains unsupported, and further amplification is informationally exhausted pending direct measurements.
2026-08-04T00:25:33Z
The latest attachment adds no independent commodity CPU/NVMe benchmark or completed reproduction; multiple implementations establish architectural feasibility but still leave practical throughput unsupported. The discussion is exhausted amplification, so revisit only on direct measurements or a materially faster implementation.
2026-08-03T23:24:18Z
The latest material still offers no independent reproduction or commodity CPU/NVMe throughput benchmark for the C99 engine. Multiple implementations establish architectural feasibility, but practical usability remains unsupported and further attention is repetitive amplification.
2026-08-03T22:26:00Z
The newly attached material still adds no independent commodity CPU/NVMe throughput benchmark or completed reproduction of the C99 engine. Architectural feasibility is established, but practical usability remains unsupported and further amplification is informationally exhausted.
2026-08-03T21:27:33Z
The latest attachment adds no independent reproduction or commodity CPU/NVMe throughput measurement, so it does not change the C99 engine’s status as an impractically slow feasibility demo. Multiple implementations corroborate expert-streaming architecture, but this practical-speed claim is informationally exhausted pending direct benchmarks.
2026-08-03T20:31:44Z
The added activity still provides no independent reproduction or commodity CPU/NVMe throughput measurement; multiple implementations corroborate architectural feasibility, not practically usable speed. The signal remains informationally exhausted pending a direct benchmark or materially faster implementation.
2026-08-03T19:25:44Z
No independent reproduction or throughput benchmark has appeared; the new activity remains amplification or adjacent inference context. Multiple implementations establish NVMe expert streaming as technically feasible, but the C99 engine’s practical usability remains unsupported.
2026-08-03T18:25:31Z
No new direct reproduction or independent throughput measurement changes the practical-usability claim; the additional activity remains adjacent context or repetitive amplification. Multiple implementations establish architectural feasibility, but the C99 engine still looks like an impractically slow demonstration pending measured benchmarks.
2026-08-03T17:30:59Z
Cloudflare’s at-scale serving report shows broader inference-engineering progress but does not test commodity CPU/NVMe execution or independently benchmark the C99 engine. Architectural feasibility is corroborated, while practical usability remains unsupported; wait for direct throughput measurements.
2026-08-03T17:21:38Z
evidence attached: hn.story.49158581 — Cloudflare's report on serving Kimi and GLM at scale provides useful independent context for whether large Chinese models can be made practical through efficiency and inference engineering.
2026-08-03T16:25:48Z
The attachment adds no independent Kimi K3 throughput measurement or completed reproduction, so the architecture is corroborated but practical speed remains unsupported. Repetitive amplification is exhausted; revisit only for direct benchmarks or a materially faster implementation.
2026-08-03T15:30:51Z
The latest attachment adds no independent Kimi K3 throughput measurement or completed reproduction, so multiple implementations still corroborate only architectural feasibility rather than practically usable speed. The signal is exhausted pending a direct benchmark or materially faster implementation.
2026-08-03T14:27:37Z
No independent Kimi K3 throughput benchmark or completed reproduction has appeared; multiple implementations corroborate NVMe expert streaming only as an architecture, not practically usable speed. The remaining activity is repetitive amplification, so the case should stay cold until direct measurements emerge.
2026-08-03T13:24:09Z
The latest material still provides no independent Kimi K3 throughput benchmark or completed reproduction of the C99 engine. Architectural feasibility is established by multiple implementations, but practical usability remains unsupported and further attention is repetitive amplification.
2026-08-03T11:26:27Z
The newly attached AirLLM item concerns 70B inference and adds no direct Kimi K3 reproduction or throughput measurement. Multiple implementations still establish only NVMe-streaming feasibility, leaving practical usability unresolved and the signal informationally exhausted pending real benchmarks.
2026-08-03T11:21:10Z
evidence attached: hn.story.49154228 — shared external link with case evidence
2026-08-03T10:22:40Z
The newly attached material still adds no independent throughput benchmark or completed reproduction, so multiple implementations establish architectural feasibility but not practically usable speed. The case is informationally exhausted until direct measurements appear and no longer merits frequent polling.
2026-08-03T09:23:13Z
The added material still does not provide an independent throughput benchmark or completed reproduction of the C99 engine; multiple implementations corroborate memory-constrained expert streaming, not practically usable speed. The signal is exhausted until direct measurements or a materially faster implementation appear.
2026-08-03T08:22:43Z
No independent throughput benchmark or completed reproduction has appeared; the additional activity only repeats evidence that NVMe expert streaming is technically feasible, not practically usable. Keep the case open but stop frequent polling until direct measurements or a materially faster implementation emerge.
2026-08-03T07:23:06Z
The attached material still adds no independent throughput benchmark or completed reproduction of the C99 engine; WASTE corroborates the expert-streaming architecture but not practical speed. Repeated amplification is exhausted, so defer review until direct measurements emerge.
2026-08-03T06:25:41Z
No new direct benchmark or completed reproduction establishes practically usable throughput; the independent WASTE implementation corroborates only the NVMe expert-streaming architecture. Further activity remains repetitive amplification, so the case should wait for measured performance rather than frequent polling.
2026-08-03T05:23:59Z
The latest attachment adds no independent throughput measurement or completed reproduction, so the second implementation still corroborates only the expert-streaming architecture, not practically usable speed. Continued attention is repetitive amplification; the case should remain cold until direct benchmark results appear.
2026-08-03T04:22:05Z
The added material still provides no independent throughput benchmark or completed reproduction of the C99 engine; the second implementation corroborates the architecture only. Practical usability remains unestablished, and continued attention is repetitive amplification rather than substantive movement.
2026-08-03T03:26:52Z
No new direct benchmark or completed reproduction establishes usable throughput; the second implementation corroborates NVMe expert streaming as an architecture, not the C99 engine’s practical-speed claim. Further attention remains repetitive amplification, so wait for measured performance.
2026-08-03T02:26:46Z
No new independent throughput measurement or completed reproduction changes the case; the second implementation corroborates the architecture, not practically usable speed. Further activity is repetitive amplification, so wait for direct benchmarks.
2026-08-03T01:21:36Z
The second implementation corroborates NVMe expert streaming as an architecture, but still provides no measured throughput or independent reproduction of the original engine’s performance. Practical usability remains unestablished, and the latest activity adds amplification rather than substance.
2026-08-03T00:24:34Z
A second independent C/NVMe expert-streaming implementation makes the architecture more than a one-off feasibility claim, warranting active watching. It still supplies no throughput benchmark or reproduction of the original engine, so practical usability remains unestablished.
2026-08-03T00:21:07Z
evidence attached: reddit.post.1vdy1nd — This is an independent implementation of NVMe expert streaming for running the full Kimi K3 beyond available RAM, directly bearing on the open validation case.
2026-08-02T23:23:29Z
The newly attached material still contains no independent reproduction or throughput benchmark, leaving the C99 engine an impractically slow feasibility demo rather than usable local inference. Repeated engagement and adjacent AirLLM evidence add no validation; wait for direct measurements.
2026-08-02T22:24:07Z
No new direct benchmark or completed reproduction changes the case: the C99 engine remains a low-memory feasibility demonstration whose reported throughput is impractical. Repeated engagement and the adjacent AirLLM implementation add no validation of this engine, so stop frequent polling until measured results emerge.
2026-08-02T21:23:20Z
The attachment adds no direct reproduction or independent throughput measurement, so the C99 engine remains an impractically slow feasibility demo rather than usable local inference. Repeated amplification is no longer informative; revisit only when a measured benchmark or materially faster implementation appears.
2026-08-02T20:24:00Z
The newly attached material still provides no independent reproduction or throughput benchmark for the C99 engine. The case remains a cold feasibility demonstration with impractical reported speed; adjacent implementations and repeated amplification do not validate practical usability.
2026-08-02T19:23:08Z
No independent reproduction or throughput measurement has appeared; the new attachment does not alter the C99 engine’s status as a low-memory feasibility demo with impractical reported speed. Further engagement is repetitive amplification, so wait for direct benchmarks rather than continuing frequent review.
2026-08-02T18:22:39Z
No independent benchmark or completed reproduction has emerged; the C99 engine remains a low-memory feasibility demonstration with impractical reported throughput. Adjacent implementations and repeated attention do not validate this engine’s usable-speed claim, so review only when direct measurements appear.
2026-08-02T17:22:33Z
The newly attached material adds no direct reproduction or independent throughput result, so the C99 engine remains an impressive memory-feasibility demo rather than practically usable inference. Adjacent low-memory implementations do not validate its performance claim; pause frequent review until measured replication appears.
2026-08-02T16:31:33Z
No substantive new evidence appears: there is still no independent reproduction or throughput benchmark of the C99 engine, and AirLLM remains adjacent rather than validating. Engagement continues to amplify a technically interesting but impractically slow feasibility demo, so defer further review until measured replication emerges.
2026-08-02T15:27:41Z
The newly attached material still adds no direct reproduction or independent throughput measurement, leaving the C99 engine as an impractically slow feasibility demo. Further attention is repetitive amplification; revisit only if a measured replication or materially faster implementation appears.
2026-08-02T14:23:48Z
No independent reproduction or throughput benchmark has emerged; the attached material remains repetitive or adjacent and does not establish practical usability for the C99 engine. Keep the case cold until a measured replication or materially faster implementation appears.
2026-08-02T13:23:58Z
No new direct benchmark or completed reproduction changes the case: the C99 engine remains a memory-feasibility demo with impractical reported throughput, while AirLLM is only adjacent evidence. Repetitive engagement no longer warrants hourly review; wait for an independent measurement or materially faster implementation.
2026-08-02T12:22:12Z
No direct replication or independent throughput measurement changes the C99 engine’s status: it remains a memory-feasibility demonstration whose reported speed is impractical. The adjacent AirLLM implementation supports only the broader low-memory premise, so further engagement is repetitive amplification rather than corroboration.
2026-08-02T11:25:25Z
The attached evidence adds no direct reproduction or independent throughput measurement for the C99 engine; AirLLM remains only adjacent support for memory-constrained execution. The reported ~32.69 seconds/token still frames this as a feasibility demo rather than practical local inference, so further attention without benchmarks is repetitive amplification.
2026-08-02T10:23:02Z
No direct replication or independent throughput benchmark has appeared; the adjacent AirLLM result supports memory-constrained feasibility but does not change the C99 engine’s status as an impractically slow demonstration. Further engagement is amplification rather than substantive movement.
2026-08-02T09:22:37Z
AirLLM provides an independent, adjacent implementation suggesting Kimi K3 can run under severe memory constraints, but it neither benchmarks the C99 engine nor establishes practically usable throughput. The case remains a cold feasibility claim awaiting direct reproduction or measured performance.
2026-08-02T09:21:00Z
evidence attached: hn.story.49142538 — Independent evidence that Kimi K3 can be run on unusually constrained hardware materially strengthens the broader hypothesis that practical low-resource K3 inference is feasible, though via a different implementation.
2026-08-02T08:22:20Z
The added material still provides no independent benchmark or completed replication, so the practical-usability hypothesis remains unresolved. Attention is repetitive amplification of the existing feasibility demo rather than substantive movement.
2026-08-02T07:21:20Z
The added material supplies no independent benchmark or completed replication; it only reinforces that the current ~32.69 seconds/token result is a feasibility demonstration, not practically usable inference. Keep the case open but cold pending measured reproduction or a materially faster implementation.
2026-08-02T06:21:05Z
The repository’s own ~32.69 seconds/token result makes the engine look like a technical feasibility demo rather than practically usable local inference, while no independent benchmark has appeared. New discussion is mostly reaction or unfulfilled replication intent, so the case cools pending real reproduction.
2026-08-02T05:24:12Z
grounded: novel/none — No Scott wiki or radar intersection was found; the supplied hits provide no basis for connecting the unresolved expert-streaming performance claim to a position
2026-08-02T05:23:56Z
origin walked (codex/luna, conf 0.98): anchor reddit.post.1vd874t -> echo.github.cd998f7203 by Fareed Khan
2026-08-02T05:21:35Z
case created — A concrete implementation makes an unusually large MoE locally runnable in principle, but its speed, hardware requirements, and storage costs remain unverified.