llama.cpp contributor ynankani has proposed PR #27621 to extend existing CUDA fusion for MoE GLU and top-k routing from single-token execution to the multi-token batches used in speculative decoding. llama.cpp’s documentation explains that speculative decoding accelerates generation by drafting several tokens and verifying them as a batch, while the supplied reports indicate that MTP support can materially raise throughput for compatible local models. However, the snippets do not directly provide PR #27621’s benchmark figures or establish that the patch has merged, so its model-specific gains remain a contributor-reported proposal rather than a confirmed release result.
The radar already tracks llama.cpp’s adaptive MTP development in `radar:llama-cpp-adaptive-mtp`; this PR is an incremental CUDA/MoE implementation detail within that open speculative-decoding story rather than a new position. It still bears directly on Scott’s hardware-aware CUDA inference work and self-hosted GPU substrate because, if merged and independently validated, it could change throughput and model-selection decisions, but the supplied evidence does not establish actual benchmark gains or compatibility with his deployed models.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.cudaradar:llama-cpp-adaptive-mtpradar:concept.speculative-decodingradar:concept.llama-cppradar:concept.gpu-kernels
queries asked of Scott's wikis
- speculative decoding and multi-token prediction strategy
- CUDA kernel fusion for local MoE inference
- local inference throughput bottlenecks and benchmarks
- llama.cpp optimization projects and deployment choices
- draft-token acceptance versus verification cost
- sparse-model economics on consumer GPUs
2026-10-08T03:52:04Z
PR #27621 remains unmerged and independently unvalidated; the v0.6.0 release established its substrate (MTP for Qwen4Exp) but did not include the fusion patch. Velocity spikes and magnitude-valve eligibility belong to the release and ik_llama.cpp sibling episodes — patch-specific objects (23–40 points, 2–3 comments, dead HN thread) show no engagement movement. The case correctly stays at low heat while watching for merge into a v0.6.x release and independent benchmark reproduction.
2026-10-05T21:42:49Z
v0.6.0 shipping upstream MTP for Qwen4Exp changes the case's footing: the fusion patch is no longer a proposal ahead of its substrate but an optimization candidate against a released MTP baseline, and its fate in upcoming releases is now the thing to watch. The engagement surge (92nd-percentile rates, multi-platform valve) sits almost entirely in sibling MTP episodes — the release and ik_llama merge posts — while patch-specific objects stay quiet (23–40 points, 2–3 comments, dead HN thread), so heat stays low despite the loud neighbourhood.
2026-10-05T20:39:22Z
evidence attached: reddit.post.1wyh03u — Upstream llama.cpp v0.6.0 shipping MTP speculative decoding for Qwen4Exp is the substrate and momentum against which the proposed MoE-fusion specdec work resolves.
2026-09-16T19:31:36Z
The newly attached CUDA-graph MTP draft proposal targets a separate optimization, not independent validation of the MoE fusion patch. It adds implementation context but establishes neither additive gains nor a deployment change for this case.
2026-09-16T19:22:34Z
evidence attached: reddit.post.1wi5pwg — The llama.cpp pull request is a concrete additional optimization for CUDA-graph-backed multi-token speculative decoding.
2026-09-10T04:23:03Z
Refreshed discussion adds hardware-fit and prefill-performance questions around the adjacent ik_llama.cpp implementation, not results validating PR #27621. The fusion proposal remains plausible but unverified, with no established merge or deployment change; further routine discussion refreshes do not warrant closer monitoring.
2026-09-08T04:22:27Z
The staleness check adds no substantive evidence: PR #27621 remains an echoed contributor proposal with unconfirmed merge status and no independent validation of its gains. Adjacent MTP implementations support the broader approach, not this specific fusion patch; less frequent review is warranted.
2026-09-06T03:27:37Z
The newly attached 2025 paper link adds background, not evidence that PR #27621’s CUDA MoE fusion delivers the claimed gains or has merged. Adjacent MTP implementations remain relevant but do not independently validate this patch; its deployment implications are still unsettled.
2026-09-06T03:21:55Z
evidence attached: hn.story.49582850 — The paper provides relevant technical context on lossless speculative decoding and its potential inference-speed benefits.
2026-09-04T17:32:36Z
The refreshed comments are further hardware-fit questions about the adjacent ik_llama.cpp implementation, not evidence for PR #27621. The CUDA MoE fusion remains an unmerged or unconfirmed proposal without independent benchmark reproduction.
2026-09-04T14:34:48Z
The refreshed discussion remains focused on hardware fit and deployment of the adjacent ik_llama.cpp MTP implementation, with no new evidence about PR #27621’s benchmarks, reproducibility, or merge status. Broader MTP momentum therefore does not change the specific CUDA MoE fusion case.
2026-09-04T12:31:25Z
The refreshed comments remain hardware and deployment questions about the adjacent ik_llama.cpp MTP implementation, adding no validation, benchmark reproduction, or merge evidence for PR #27621. The specific CUDA MoE fusion remains a low-temperature contributor proposal despite broader MTP momentum.
2026-09-04T08:26:28Z
The refreshed comments remain deployment and hardware-compatibility interest around the adjacent merged MTP implementation, not validation of PR #27621. The CUDA MoE fusion is still an unconfirmed contributor proposal without independent benchmarks or established merge status.
2026-09-03T21:37:14Z
Refreshed discussion remains operator interest in hardware compatibility and deployment, not new evidence for PR #27621’s CUDA MoE fusion. The specific patch is still unmerged or otherwise unconfirmed and lacks independent, reproducible validation.
2026-09-03T19:45:05Z
The merged, independently tested ik_llama.cpp implementation strengthens the broader case that MTP can materially accelerate local sparse-model inference, but it still does not validate or establish the merge status of PR #27621’s CUDA MoE fusion. Refreshed comments add hardware interest rather than specific benchmarks or adoption evidence.
2026-09-03T17:24:35Z
evidence attached: reddit.post.1w6ccgs — This merged and independently tested MTP implementation is valuable corroboration that speculative decoding can materially improve local sparse-model throughput, with workload-dependent regressions.
2026-09-02T17:58:29Z
No merge, direct benchmark validation, or additional implementation evidence has appeared; the specific CUDA MoE fusion remains an unverified contributor proposal. The adjacent constrained-VRAM MTP result supports the broader optimization space but does not further substantiate this patch.
2026-08-31T17:36:20Z
The constrained-VRAM rollback implementation shows that MTP performance and memory trade-offs are actionable in practice, moving the broader optimization thread beyond a bare proposal. It does not validate the specific CUDA MoE fusion, and the refreshed discussion mainly highlights competing implementations and unresolved upstream status.
2026-08-31T17:24:46Z
evidence attached: reddit.post.1w3jp8j — A concrete llama.cpp MTP memory optimization reports faster speculative decoding on constrained local hardware, materially informing the open speculative-decoding performance case.
2026-08-31T12:34:56Z
No substantive evidence has arrived beyond minor Reddit engagement; the optimization remains an unmerged contributor proposal without visible benchmarks, independent validation, or deployment implications. It stays a low-temperature implementation detail within the broader adaptive-MTP story.
2026-08-31T12:29:49Z
grounded: known/medium — The radar already tracks llama.cpp’s adaptive MTP development in `radar:llama-cpp-adaptive-mtp`; this PR is an incremental CUDA/MoE implementation detail within
2026-08-31T12:27:07Z
case created — The linked first-party pull request presents a bounded CUDA optimization and benchmark claims relevant to sparse-model local inference.