2026-10-11 17:11 UTC

Independent benchmarks will determine whether Kimi K3 can run interactively on a single consumer GPU with practical memory use and generation speed.

state: resolvedheat: lowuncertainty: mediumnovelscott: lowkimi-k3 local-inference consumer-gpu open-modelsAkashi203Moonshot AI

What is this?

Kimi K3 is presented as a frontier mixture-of-experts model from Moonshot AI, with open weights planned but substantial total-parameter memory requirements. A claim attributed to Akashi203 says it now runs on one consumer GPU, but the supplied snippets provide no independent local-hardware benchmark confirming practical memory use, time to first token, or interactive generation speed; instead, several sources argue that all experts must remain resident and therefore require hundreds of gigabytes even when quantized. The cited 62 tokens/s result appears to be an API measurement rather than a single-GPU local test, so the consumer-GPU claim remains unverified in the supplied evidence.

Why it matters to Scott

No intersection was found in Scottโ€™s wikis, and the radar does not already track this development or its actors. The unresolved single-consumer-GPU claim is broadly relevant to local inference, but without independent benchmarks or a connection to Scottโ€™s documented positions or projects, it is only topical.
queries asked of Scott's wikis
  • consumer-GPU local inference economics
  • MoE total parameters versus active parameters
  • quantization and memory-constrained inference
  • open weights versus practically runnable models
  • independent benchmarking of local model claims
  • interactive inference latency and tokens per second

Measured heat

no measured readings yet โ€” the hourly heat pass fills this in

How the heat travelled

no chain yet โ€” the hourly chain pass fills this in

Evidence (9) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnKimi k3 now runs on one consumer GPUOsamaJaber50
๐ŸŸง echo.x โญClaims that โ€œKimi K3 now runs on one consumer GPU.โ€Akashi203โ€”โ€”
๐ŸŸ  redditAnyone tested the IQ1_M 342GB Pruned Kimi K3? Is it usable?
LocalLLaMA
Hannibalj2ca7024
๐ŸŸง hnShow HN: Run Full Kimi K3 with 29 GB of RAMmarcobambini71
๐ŸŸง hnKimi k3 run on RTX 5090OsamaJaber20
๐ŸŸ  redditHow Kimi K3 Engineered Its Way to the Frontier [R]
MachineLearning
noninertialframe96176
๐ŸŸง hnRunning Kimi K3 on a local computermarcobambini20
๐ŸŸ  redditKIMI K3 IN REALITY
LocalLLaMA
techlatest_net04
๐ŸŸง hnRun Kimi K3 using 29 GB of RAM at 0.50 tok/smarcobambini314148

Interpretation history

Decision trace