Independent reproduction will determine whether minirun-app can run Kimi K3 on iPhone-class hardware by streaming its approximately 1.56 TB of weights from external SSD storage at practically useful performance.
state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference mobile-inference kimi-k3nanguoyu
What is this?
minirun-app is presented as an early implementation claiming to run Moonshot AI’s open-weight Kimi K3—a sparse mixture-of-experts model with an approximately 1.56 TB checkpoint—on an iPhone 16 Pro by streaming weights from external SSD storage. The supplied results establish that K3 activates only a subset of its experts per token and include other claims of low-memory execution through expert streaming, but conventional estimates call for roughly 1.6 TB or more of aggregate accelerator memory. The snippets neither independently verify minirun-app’s performance nor establish practically useful generation speeds, and they conflict on whether any consumer-hardware configuration has been verified, so independent reproduction is the central unresolved issue.
Why it matters to Scott
The radar already tracks the same unresolved Kimi K3 out-of-core inference claim on `radar:kimi-k3-nvme-expert-streaming`; minirun-app mainly changes the target to an iPhone and external SSD. It fits Scott’s hardware-aware local-inference and evidence-ceiling frameworks, but without independently reproduced token rates and output quality it is another unverified example rather than a development likely to change what he builds or argues.
dev:concept.hardware-aware-local-inferenceip:concept.evidence-class-ladderip:concept.latencyradar:kimi-k3-nvme-expert-streamingradar:concept.expert-streamingradar:concept.edge-inference
queries asked of Scott's wikis
- expert streaming from SSD for sparse MoE inference
- storage bandwidth as local-inference compute bottleneck
- out-of-core inference on mobile hardware
- practical token-rate thresholds for local models
- open-weight model sovereignty on consumer devices
- independent reproduction standards for AI systems demos
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-19T15:46:03Z
The claim has attracted no independent reproduction, technical discussion, or performance evidence within its horizon; the only movement is negligible engagement. With the reported 220 seconds per token already impractical, this episode has faded unless a reproduction later reopens it.
2026-08-17T14:41:49Z
No independent reproduction or performance evidence has appeared; the slight engagement increase is non-substantive, leaving this as an isolated implementation claim with impractical reported latency.
2026-08-17T14:35:43Z
grounded: known/low — The radar already tracks the same unresolved Kimi K3 out-of-core inference claim on `radar:kimi-k3-nvme-expert-streaming`; minirun-app mainly changes the target
2026-08-17T14:32:29Z
origin walked (codex/luna, conf 0.98): anchor hn.story.49331279 -> echo.github.057396f68f by Dong Wang
2026-08-17T14:31:11Z
case created — A first-party implementation presents a concrete and unusually hardware-constrained local-inference technique that can be independently tested.
Decision trace
- 08-20 01:46expireThe claim has attracted no independent reproduction, technical discussion, or performance evidence within its horizon; the only movement is negligible engagement. With the reported 220 seconds per tok
- 08-20 01:46alert_silentNo consequential new event occurred: a one-point score increase without comments or corroboration does not merit attention or change the evidence ceiling.
- 08-20 01:46alert_routeNo consequential new event occurred: a one-point score increase without comments or corroboration does not merit attention or change the evidence ceiling.
- 08-18 00:41repriceNo independent reproduction or performance evidence has appeared; the slight engagement increase is non-substantive, leaving this as an isolated implementation claim with impractical reported latency.
- 08-18 00:41alert_silentThe only new delta is a negligible score increase with no comments or technical corroboration, so there is nothing consequential to surface before the next briefing.
- 08-18 00:41alert_routeThe only new delta is a negligible score increase with no comments or technical corroboration, so there is nothing consequential to surface before the next briefing.
- 08-18 00:40alert_silentThe repository establishes that Minirun published an iPhone 16 Pro SSD-streaming implementation and claims Kimi K3 inference, but its reported approximately 220 seconds per token is not practically us
- 08-18 00:40alert_routeThe repository establishes that Minirun published an iPhone 16 Pro SSD-streaming implementation and claims Kimi K3 inference, but its reported approximately 220 seconds per token is not practically us
- 08-18 00:35groundThe radar already tracks the same unresolved Kimi K3 out-of-core inference claim on `radar:kimi-k3-nvme-expert-streaming`; minirun-app mainly changes the target to an iPhone and external SSD. It fits
- 08-18 00:32promote_anchororigin walk conf 0.98
- 08-18 00:31createA first-party implementation presents a concrete and unusually hardware-constrained local-inference technique that can be independently tested.