Contributor pwilkin has opened llama.cpp PR #26185 to add text-model support for Kimi K3; the supplied Hugging Face snippet says the code remains unmerged, while an Unsloth fork builds on the PR to add vision support. Earlier llama.cpp analysis found that K3’s architecture could not safely use the existing Kimi mapping because of its distinctive attention, MoE, metadata, and tensor structure. GGUF loading has reportedly been achieved, but the snippets provide no downstream benchmark validation, and K3’s enormous size—plus a reported recommendation of 64 or more accelerators—leaves practical inference on common local hardware unestablished.
Known via dev:concept.hardware-aware-local-inference and dev:project.gamepc: Scott already treats runtime support, hardware placement, memory pressure, precision, and measured output fidelity as separate requirements for practical local inference. The PR could extend his llama.cpp/GGUF model options, but it is unmerged and K3’s scale may keep it outside his workstation’s useful operating envelope, so independent capability testing matters before it changes what he builds.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.capability-auditip:concept.evaluation-driven-developmentip:concept.usable-mass-over-unusable-powerradar:concept.llama-cppradar:concept.local-inferenceradar:concept.ggufradar:concept.model-evaluationradar:llama-cpp-hot-expert-gpu-cache
queries asked of Scott's wikis
- local inference hardware economics and memory limits
- llama.cpp and GGUF runtime strategy
- open-weight models versus practical deployability
- quantization fidelity and benchmark validation
- local model support in coding-agent stacks
- CPU/GPU hybrid inference across commodity hardware
2026-09-01T09:30:59Z
The near-term implementation story has faded without first-party merge confirmation or reproducible llama.cpp correctness and hardware-practicality testing. Existing evidence still shows ecosystem experimentation, but Kimi K3’s scale leaves little reason to expect a consequential common-hardware result within this episode’s horizon.
2026-08-30T08:24:14Z
The added from-scratch PyTorch implementation modestly broadens evidence that Kimi K3 is inspectable outside its original stack, but it does not verify llama.cpp’s merge status, inference correctness, quantization fidelity, or practicality on common hardware. The case’s meaning is therefore unchanged and remains dormant pending an upstream record or reproducible runtime testing.
2026-08-30T08:22:29Z
evidence attached: reddit.post.1w2aupi — A first-party-style PyTorch implementation provides additional evidence that Kimi K3 is becoming practically inspectable and runnable.
2026-08-28T20:42:00Z
Another staleness-only observation confirms that discussion monitoring is exhausted, without establishing the reported merge or adding reproducible correctness and hardware-practicality results. Keep the case dormant until a direct upstream record or substantive independent test appears.
2026-08-26T20:36:04Z
The staleness check and negligible engagement change add no merge confirmation, reproducible correctness results, or evidence of practicality on common hardware. Comment monitoring is exhausted; retain the case only for an upstream record or substantive independent testing.
2026-08-24T19:56:20Z
The refreshed discussion is repetitive amplification, not evidence of a merge or new correctness and hardware-practicality results. The case remains supported by implementation and deployment lines but should wait on a first-party upstream record or reproducible testing.
2026-08-24T02:24:21Z
The refreshed comments add only repetitive discussion and the already-seen secondary merge claim, with no direct upstream record or reproducible correctness and hardware-practicality results. The case should remain on a slow cadence awaiting first-party merge confirmation or meaningful independent testing.
2026-08-24T01:28:21Z
The refreshed discussion adds no direct merge confirmation or new correctness and hardware-practicality results; it is repetitive amplification of the already-known deployment caveats. The case remains corroborated but should now wait on upstream records or reproducible testing rather than continued comment monitoring.
2026-08-24T00:22:12Z
The verification window expired without direct GitHub confirmation of PR #26185’s reported merge; the new observation is engagement-only. Upstream support remains credible but unverified, while correctness and practicality outside datacenter-class hardware still lack validation.
2026-08-23T22:25:51Z
The refreshed comments still do not directly verify PR #26185’s reported merge and add no new correctness or hardware-practicality evidence. Keep the case hot only through the short verification window; otherwise return it to a slow post-merge-validation cadence.
2026-08-23T18:31:59Z
The refreshed discussion provides no direct confirmation of the reported PR merge, so upstream Kimi K3 support remains a credible but secondary claim. Keep the case hot briefly for the GitHub record; correctness and practicality outside datacenter-class hardware remain unresolved.
2026-08-23T17:28:27Z
A specific new comment reports that llama.cpp PR #26185 merged last week, potentially completing the upstream-support milestone and shifting the case toward post-merge validation. Because the merge claim is not yet backed here by the GitHub record, verify it before treating support as established; correctness and practicality on common hardware remain unresolved.
2026-08-23T16:34:15Z
The refreshed comments and slight engagement decline add no evidence about an upstream merge, correctness, or practical inference beyond datacenter-class hardware. The case remains corroborated by implementation and deployment lines, but its implications for common local systems remain unproved.
2026-08-23T14:32:38Z
The refreshed discussion adds no substantive evidence beyond the known cost, throughput, and severe-quantization caveats. Upstream merge status, reproducible correctness, and practicality outside datacenter-class hardware remain unresolved.
2026-08-23T13:38:19Z
The refreshed discussion is repetitive amplification of already-known cost and quantization caveats, not new evidence about upstream support or practical inference. Keep the case open on a slow cadence for a merge, release, or reproducible correctness and hardware validation.
2026-08-23T11:25:00Z
The refreshed comments repeat already-established objections to single-stream cost claims and destructive 1-bit quantization; they add no merge, reproducible capability validation, or evidence of practicality beyond datacenter hardware.
2026-08-23T10:33:13Z
Refreshed discussion mainly reinforces the benchmark’s existing limitations: single-stream cost comparisons are misleading and extreme quantization may sacrifice capability. No merge, reproducible quality validation, or evidence of practicality on common hardware has appeared, so the case cools while remaining open.
2026-08-23T09:34:30Z
Independent deployment moves llama.cpp compatibility beyond a purely speculative PR, but the reported 594 GB quantization, eight-A100 footprint, and roughly 9 tok/s performance strongly narrow its practical role to datacenter-class hardware. Correctness remains weakly tested and upstream merge status is unchanged.
2026-08-23T09:22:43Z
evidence attached: reddit.post.1vw1j2p — Independent deployment evidence for Kimi K3 using llama.cpp materially informs practical compatibility and inference economics.
2026-08-21T18:31:21Z
The latest observation is again engagement-only, with no merge, release, implementation progress, or independent validation. Topic activity does not change this case’s meaning; practical Kimi K3 inference in llama.cpp remains speculative.
2026-08-19T17:51:21Z
The staleness check adds no evidence beyond the existing unmerged implementation claim; keep the case open on a slower cadence for an actual merge, release, or independent correctness and hardware-viability results.
2026-08-17T17:40:01Z
No merge, release, implementation update, or independent validation has appeared; repeated engagement-only observations add no substance. The PR remains a plausible path, but both correctness and practical hardware viability are unresolved.
2026-08-15T16:37:39Z
The added engagement is minor amplification without a merge, release, or independent correctness and performance testing. The case remains a plausible implementation path, but its practical value on common local hardware is still wholly unestablished.
2026-08-15T16:27:54Z
grounded: known/medium — Known via dev:concept.hardware-aware-local-inference and dev:project.gamepc: Scott already treats runtime support, hardware placement, memory pressure, precisio
2026-08-15T16:25:06Z
case created — A first-party implementation pull request could materially expand independent local access to Kimi K3 and warrants tracking through merge and testing.