2026-10-11 16:37 UTC

llama.cpp contributor ynankani claims the proposed CUDA MoE fusion for speculative decoding materially accelerates multi-token prediction across sparse models, potentially improving local draft-token throughput if merged.

state: watchingheat: lowuncertainty: highknownscott: mediumllama-cpp local-inference inference-optimizationynankaniggml-orgllama.cpp

What is this?

llama.cpp contributor ynankani has proposed PR #27621 to extend existing CUDA fusion for MoE GLU and top-k routing from single-token execution to the multi-token batches used in speculative decoding. llama.cpp’s documentation explains that speculative decoding accelerates generation by drafting several tokens and verifying them as a batch, while the supplied reports indicate that MTP support can materially raise throughput for compatible local models. However, the snippets do not directly provide PR #27621’s benchmark figures or establish that the patch has merged, so its model-specific gains remain a contributor-reported proposal rather than a confirmed release result.

Why it matters to Scott

The radar already tracks llama.cpp’s adaptive MTP development in `radar:llama-cpp-adaptive-mtp`; this PR is an incremental CUDA/MoE implementation detail within that open speculative-decoding story rather than a new position. It still bears directly on Scott’s hardware-aware CUDA inference work and self-hosted GPU substrate because, if merged and independently validated, it could change throughput and model-selection decisions, but the supplied evidence does not establish actual benchmark gains or compatibility with his deployed models.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.cudaradar:llama-cpp-adaptive-mtpradar:concept.speculative-decodingradar:concept.llama-cppradar:concept.gpu-kernels
queries asked of Scott's wikis
  • speculative decoding and multi-token prediction strategy
  • CUDA kernel fusion for local MoE inference
  • local inference throughput bottlenecks and benchmarks
  • llama.cpp optimization projects and deployment choices
  • draft-token acceptance versus verification cost
  • sparse-model economics on consumer GPUs

Measured heat

now 0 pts/hpeak 71 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 988h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-31 12:27 (minted)⭐ origin echo-reconstructedPull request #27621 extends MoE GLU and top-k router fusion beyond single-token execution to speculative decoding, with reported benchmarks
ynankani on github (echo) · attributed from reddit.post.1w3bh6f · published time unknown
—
08-31 11:50first on r/LocalLLaMA · published · lag ?CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani · Pull Request #27621 · ggml-org/llama.cpp
jacek2023
—
09-06 02:51first on hacker news · published · lag ?Accelerating LLM Inference with Lossless Speculative Decoding Algorithms (2025)
wslh
—
08-31 11:50amplified on r/LocalLLaMAreddit.post.1w3bh6f
jacek2023
peak 23 · 3 comments · 4% of case engagement
08-31 16:58amplified on r/LocalLLaMAreddit.post.1w3jp8j
ea_man
peak 25 · 13 comments · 5% of case engagement
09-03 16:30amplified on r/LocalLLaMAreddit.post.1w6ccgs
Alternative_Will5974
peak 191 · 89 comments · 39% of case engagement
09-06 02:51amplified on hacker newshn.story.49582850
wslh
peak 1 · 0 comments · 0% of case engagement
09-16 18:36amplified on r/LocalLLaMAreddit.post.1wi5pwg
jacek2023
peak 36 · 2 comments · 5% of case engagement
10-05 18:58amplified on r/LocalLLaMA 👑reddit.post.1wyh03u
vexatious-big
peak 260 · 66 comments · 46% of case engagement
08-31 12:20our radar first saw it · lag ?discovery anchor: reddit.post.1w3bh6f—
pace: p81 vs 519 stories at the 720h mark (now 988h old) — ahead of system76-thelio-mira-ai (1.0x), behind nvidia-sol-pi-harness-efficiency (1.0x)

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditCUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token by ynankani · Pull Request #27621 · ggml-org/llama.cpp
LocalLLaMA
jacek2023233
🟧 echo.github ⭐Pull request #27621 extends MoE GLU and top-k router fusion beyond single-token execution to speculative decoding, with reported benchmarks ynankani——
🟠 redditCompact Rollback MTP: a MTP version for QWEN models for those with little vRAM
LocalLLaMA
ea_man2513
🟠 redditQwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070
LocalLLaMA
Alternative_Will597418989
🟧 hnAccelerating LLM Inference with Lossless Speculative Decoding Algorithms (2025)wslh10
🟠 redditEnable CUDA graph for MTP draft by gaugarg-nv · Pull Request #28549 · ggml-org/llama.cpp
LocalLLaMA
jacek2023362
🟠 redditllama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more
LocalLLaMA
vexatious-big26066

Interpretation history

Decision trace