2026-10-11 17:10 UTC

Upstream review and independent benchmarks will determine whether correcting llama.cpp’s inflated MTP buffer reservations materially expands usable context on memory-constrained AMD systems without inference regressions.

state: resolvedheat: lowuncertainty: highknownscott: lowllama-cpp local-inference context-window memory-fittingllama.cppAMDQwen

What is this?

The case concerns a reported llama.cpp memory-allocation issue in which the secondary context created for multi-token prediction (MTP) reserves more VRAM than necessary, reducing the context length that can fit on memory-constrained systems. A llama.cpp GitHub issue independently documents related behavior: the fitting logic reduces the main context, but the MTP draft context then initializes using the model’s larger native context and can cause an out-of-memory failure. The supplied evidence titles report an AMD ROCm benchmark improvement from 64,256 to 149,248 tokens for Qwen 27B after a patch, but the snippets do not directly verify those figures, upstream acceptance, or the absence of inference regressions.

Why it matters to Scott

The radar already tracks the same llama.cpp MTP memory-overhead development in `radar:llama-cpp-mtp-default-memory-regression`; this case adds a specific AMD/Qwen context-capacity claim awaiting validation rather than a new position. It touches Scott’s hardware-aware local-inference and evaluation-gated change practices, but the hits establish only an NVIDIA/CUDA local stack—not active AMD ROCm or llama.cpp use—so it is adjacent rather than action-changing.
dev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentradar:llama-cpp-mtp-default-memory-regressionradar:concept.rocmradar:concept.speculative-decodingradar:concept.long-context-inference
queries asked of Scott's wikis
  • local inference memory-fitting strategy
  • context length versus VRAM economics
  • speculative decoding memory overhead
  • AMD ROCm local model support
  • inference benchmark and regression methodology
  • runtime context allocation architecture

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditAMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B
LocalLLaMA
ea_man7528
🟧 echo.other ⭐Primary benchmark log containing the figures repeated by the Reddit post: “ROCm mainline 64,256” versus “ROCm patched 149,248” context tokeneaman (session attribution: “Sol”)——

Interpretation history

Decision trace