2026-10-11 17:10 UTC

Independent benchmarks will determine whether llama.cpp’s adaptive MTP mode selects speculative-decoding depth effectively enough to improve coding-agent throughput without manual tuning.

state: expiredheat: lowuncertainty: highknownscott: mediumllama-cpp speculative-decoding local-inference coding-agentsllama.cppggml-org

What is this?

llama.cpp, maintained under ggml-org, has added MTP-style speculative decoding for accelerating local inference, along with tooling to benchmark throughput, latency, and draft acceptance against a baseline. The cited PR reportedly introduces an adaptive MTP algorithm using a hysteresis state machine that raises draft depth after consecutive full acceptances, aiming to avoid manual depth tuning. The supplied snippets report substantial gains for MTP generally, but they do not provide accessible independent results isolating this adaptive mode on end-to-end coding-agent workloads; one snippet also flags potential memory and timeout problems.

Why it matters to Scott

The radar already tracks substantially the same open validation question in `radar:adaptive-speculative-decoding-300-gpu`, alongside llama.cpp MTP memory cases. Results on end-to-end coding-agent throughput would still bear directly on Scott’s hardware-aware local-inference policy and trace-backed model-plus-harness evaluation, potentially changing runtime configuration rather than merely illustrating a belief.
dev:concept.hardware-aware-local-inferencedev:concept.trace-backed-agent-comparisonip:concept.model-plus-harness-benchmark-unitip:concept.evaluation-driven-developmentradar:adaptive-speculative-decoding-300-gpuradar:concept.speculative-decodingradar:concept.llama-cppradar:llama-cpp-mtp-autofit-memoryradar:llama-cpp-mtp-default-memory-regression
queries asked of Scott's wikis
  • adaptive speculative-decoding depth policies
  • coding-agent end-to-end inference benchmarks
  • local inference throughput versus tool-call latency
  • llama.cpp agent harness integrations
  • MTP acceptance rates and memory tradeoffs
  • self-tuning inference optimization

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (11) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditllama.cpp adaptive MTP PR#27210
LocalLLaMA
Look_0ver_There17233
🟧 echo.github ⭐The PR author says he designed the adaptive MTP algorithm: a hysteresis state machine where consecutive full accepts increase draft depth anStew Forster (stew675)——
🟠 redditCombining MTP with ngram-mod worth it for coding?
LocalLLaMA
YetAnotherAnonymoose1114
🟠 redditDFlash2 speeds Qwen 3.8 27B up to 4 times
LocalLLaMA
Top-Eye-810426283
🟠 reddit3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned.
LocalLLaMA
fintip4826
🟠 redditWe should disable MTP when coding
LocalLLaMA
fbms2030
🟠 reddit[Benchmark] Optimal DFlash2 quants for speed and context size, 5090 RTX, llama.cpp, Qwen 3.8 27B Dynamic3 Unsloth. Comparison with MTP
LocalLLaMA
Opening-Broccoli919078
🟠 redditNew: Llama.cpp adaptive speculation for faster inference
LocalLLaMA
Dutchnamn10740
🟠 redditSupport for DFlash2 in llama.cpp has been merged! - spec : add DFlash2 support (local convolution + candidate selector) by SubSir · Pull Request #27342 · ggml-org/llama.cpp
LocalLLaMA
DjCanalex938
🟠 reddit27B great speed up for coding with draft-p-min 0.8
LocalLLaMA
Old-Sherbert-44951039
🟠 redditMTP causing tool calling issues - Qwen 3.6 9B
LocalLLaMA
OvertaxedOne78

Interpretation history

Decision trace