2026-10-11 17:09 UTC

llama.cpp maintainers will revert or gate default loading of bundled MTP tensors after reports that models consume extra RAM or VRAM even when MTP speculative decoding is disabled.

state: expiredheat: lowuncertainty: highnovelscott: lowllama-cpp speculative-decoding memory-efficiency local-inferenceggml-orgllama.cpp

What is this?

llama.cpp is a local LLM inference project maintained under ggml-org that is adding multi-token prediction (MTP) support for speculative decoding. Supplied reports describe excessive memory use in two different contexts: a SYCL-specific issue while MTP is active, and a merged fix for draft-side resources surviving server sleep/resume; neither directly establishes the claimed regression that bundled MTP tensors load when speculation is disabled. The snippets therefore do not establish that maintainers plan to revert or gate default tensor loading, although they show active work on MTP-related memory costs and cleanup.

Why it matters to Scott

No intersection found in Scott’s wikis, and no radar page already tracks this development. The supplied evidence also does not establish the predicted revert or gating change.
queries asked of Scott's wikis
  • local inference memory-efficiency budgets
  • speculative decoding cost-benefit thresholds
  • llama.cpp production deployment patterns
  • optional model features zero-cost when disabled
  • VRAM regressions in local model runtimes
  • bundled auxiliary weights and model packaging

Measured heat

no measured readings yet β€” the hourly heat pass fills this in

How the heat travelled

no chain yet β€” the hourly chain pass fills this in

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditPSA: llama.cpp now loads MTP tensors by default for any draft-mtp arch, even with MTP disabled
LocalLLaMA
Shoddy_Bed324016847
🟧 echo.github ⭐The PR explicitly documents the regression: β€œnextn tensors flip from never-loaded to loaded-when-present,” costing about one MoE layer even satindergrewalβ€”β€”
🟠 redditDoes MTP head get loaded in VRAM by default?
LocalLLaMA
xornullvoid27
🟠 redditOllama now auto-enables Qwen3.5 MTP on Macs, a previous Metal test was 11–27% slower
LocalLLaMA
BTA_Labs016

Interpretation history

Decision trace