llama.cpp, maintained under ggml-org, has added MTP-style speculative decoding for accelerating local inference, along with tooling to benchmark throughput, latency, and draft acceptance against a baseline. The cited PR reportedly introduces an adaptive MTP algorithm using a hysteresis state machine that raises draft depth after consecutive full acceptances, aiming to avoid manual depth tuning. The supplied snippets report substantial gains for MTP generally, but they do not provide accessible independent results isolating this adaptive mode on end-to-end coding-agent workloads; one snippet also flags potential memory and timeout problems.
2026-08-31T06:28:56Z
No adaptive-versus-fixed benchmark, upstream decision, or coding-agent reliability result has emerged within the case horizon; minor engagement changes only revisit adjacent tuning anecdotes. Convergent implementations remain real, but this validation episode is dormant until materially new measurements appear.
2026-08-29T05:30:40Z
The new vLLM-specific prefix-cache bug explanation makes the reported agent failures more plausibly an implementation interaction than an inherent MTP quality regression. It further weakens any inference about llama.cpp’s adaptive selector, whose throughput and reliability still require controlled adaptive-versus-fixed coding-agent tests.
2026-08-29T01:33:17Z
A second vague user report makes MTP-related agent-quality regressions a somewhat stronger test lead, but it still lacks logs, configuration detail, reproduction, or a llama.cpp adaptive-selector connection. The core question remains unchanged pending controlled adaptive-versus-fixed coding-agent benchmarks that include tool-call and structured-output reliability.
2026-08-28T21:37:49Z
The new report expands the evaluation target from throughput and memory to agent reliability: MTP may coincide with tool-calling, JSON, MCP, and RAG failures. Because it is an unreproduced vLLM anecdote with unclear model/configuration details and does not test llama.cpp’s adaptive selector, it is a regression-test lead rather than evidence for or against adaptive depth selection.
2026-08-28T21:24:14Z
evidence attached: reddit.post.1w136y3 — Independent field evidence suggests MTP can break tool calling and RAG even when it improves decoding speed, adding an important reliability tradeoff.
2026-08-28T07:27:26Z
The refreshed comments remain hardware-specific fixed-depth and threshold tuning rather than evidence that adaptive MTP chooses effectively. No controlled adaptive-versus-fixed coding trace, upstream adaptive-MTP decision, or memory-aware benchmark changes the case’s meaning.
2026-08-28T01:32:40Z
The refreshed comments add another fixed-depth tuning sweep and repeat hardware-specific optima, but still do not test whether adaptive MTP selects those settings effectively. The case remains corroborated by convergent implementations while its coding-agent throughput advantage is unvalidated.
2026-08-28T00:27:40Z
The refreshed tuning discussion again suggests shallow fixed depths and hardware-specific optima, reinforcing the motivation for automatic selection without testing whether the adaptive selector actually finds them. No controlled adaptive-versus-fixed coding trace or upstream adaptive-MTP decision changes the case.
2026-08-27T20:44:40Z
Refreshed discussion adds only engagement and hardware-specific configuration anecdotes, with no controlled adaptive-versus-fixed benchmark, coding-agent trace, or upstream adaptive-MTP decision. DFlash2 remains the newly merged comparator, but the adaptive selector’s practical advantage is still unvalidated.
2026-08-27T19:00:15Z
DFlash2’s upstream merge turns a previously benchmarked alternative into an immediately testable mainline competitor, raising the bar and changing the evaluation from adaptive MTP in isolation to a coding-trace comparison across adaptive MTP, fixed MTP, and DFlash2. The new coding-speed report reinforces that speculative controls matter but remains hardware-dependent and does not validate adaptive depth selection.
2026-08-27T18:24:40Z
evidence attached: reddit.post.1w00dr3 — A user report with corroborating comments shows a material coding-throughput gain from speculative-decoding controls in llama.cpp, directly informing the open case.
2026-08-27T18:24:40Z
evidence attached: reddit.post.1w00ngm — The merged DFlash2 implementation is independent first-party evidence that llama.cpp is broadening practical speculative-decoding support.
2026-08-27T09:36:50Z
The refreshed discussion adds no measurements, upstream adoption, or controlled adaptive-versus-fixed coding-agent result; it remains repetitive amplification around configuration, overhead, and attribution. Convergent implementations sustain corroboration, but practical effectiveness remains unvalidated.
2026-08-26T09:24:50Z
The refreshed comments add no benchmark, upstream adoption, or controlled adaptive-versus-fixed coding-agent result; they remain repetitive questions about attribution, overhead, mainlining, and configuration. Two implementations sustain corroboration, but adaptive depth selection remains practically unvalidated.
2026-08-25T22:31:52Z
The refreshed comments add no measurements, upstream adoption, or controlled adaptive-versus-fixed coding-agent result; they only repeat attribution, mainline, overhead, and configuration questions. Convergent implementations sustain corroboration, but adaptive depth selection remains practically unvalidated.
2026-08-25T20:37:32Z
The refreshed discussion adds no measurements, upstream adoption, or controlled adaptive-versus-fixed coding-agent result. It remains repetitive amplification around attribution, overhead, and configuration, so convergent implementations sustain corroboration while practical effectiveness remains unvalidated.
2026-08-25T16:45:19Z
The refreshed comments add no benchmark, upstream adoption, or coding-agent workload result; they remain repetitive questions about overhead, attribution, and configuration. Two implementations sustain corroboration, but adaptive depth selection remains practically unvalidated.
2026-08-25T14:44:16Z
The latest comment refresh adds no measurements, upstream adoption, or adaptive-versus-fixed coding-workload result; it remains repetitive discussion of attribution, overhead, and configuration. Convergent implementations sustain corroboration, but the selector’s practical benefit remains unvalidated.
2026-08-25T13:37:48Z
Refreshed discussion focuses on attribution, mainline status, and overhead questions around the new fork, without adding measurements or adoption evidence. Two implementations still establish convergent interest, but adaptive-versus-fixed and coding-agent throughput benefits remain unvalidated.
2026-08-25T12:32:10Z
The case now has two concrete adaptive-speculation implementations and the first implementation-level performance claim, moving it beyond a single experimental PR. The reported structured-generation gain is still author-supplied and does not isolate adaptive versus fixed depth or demonstrate end-to-end coding-agent throughput, so practical effectiveness remains unsettled.
2026-08-25T12:24:08Z
evidence attached: reddit.post.1vxxa9x — A concrete fork implements adaptive speculation and reports a large structured-generation speedup, directly advancing the open llama.cpp validation case despite disputed attribution.
2026-08-24T13:24:21Z
The refreshed comments add only configuration anecdotes and questions about DFlash2 bugs, context limits, and fixed MTP depth. No adaptive-versus-fixed benchmark, upstream decision, or coding-agent workload result changes the selector’s unvalidated status.
2026-08-24T12:23:17Z
The new benchmark raises the competitive bar by showing DFlash2 can beat fixed-depth MTP on a specific Qwen/RTX 5090 speed-context tradeoff. It still does not evaluate adaptive MTP or coding-agent throughput, so the selector remains unvalidated rather than disproved.
2026-08-24T12:22:04Z
evidence attached: reddit.post.1vx14gl — A hands-on llama.cpp benchmark compares DFlash2 speculative decoding with MTP and reports materially better speed-context tradeoffs for Qwen workloads.
2026-08-23T02:24:50Z
The refreshed comments only continue the contested fixed-MTP slowdown discussion, with VRAM pressure remaining the likely confound. No controlled adaptive-versus-fixed benchmark, memory telemetry, upstream decision, or coding-agent workload result changes the selector’s unvalidated status.
2026-08-22T20:26:43Z
The refreshed comments remain contested anecdotes about fixed-depth MTP and VRAM pressure, not evidence about the adaptive selector. With no controlled adaptive-versus-fixed benchmark, memory telemetry, upstream decision, or coding-agent workload result, the case’s meaning is unchanged.
2026-08-22T17:32:24Z
The refreshed comments remain a contested, poorly specified fixed-MTP slowdown anecdote, with VRAM pressure still the leading confound. No adaptive-versus-fixed benchmark, memory telemetry, upstream decision, or coding-agent workload result changes the selector’s unvalidated status.
2026-08-22T15:32:40Z
Refreshed comments further contest the long-context slowdown anecdote with contrary configurations and a plausible VRAM-offload explanation; a linked benchmark is not itself available here as evidence. No adaptive-versus-fixed result, memory telemetry, or coding-agent workload benchmark changes the open validation question.
2026-08-22T14:37:59Z
The coding-workload anecdote adds a relevant failure mode—MTP memory overhead may trigger long-context offload and erase throughput gains—but refreshed comments make that confound more plausible rather than validating the claim. It still does not test adaptive depth, so the case remains pending controlled adaptive-versus-fixed benchmarks with memory telemetry.
2026-08-22T14:23:03Z
evidence attached: reddit.post.1vvdi3h — A concrete but anecdotal coding workload report suggests MTP can hurt long-context throughput, directly challenging the expected benefit and motivating controlled benchmarks.
2026-08-21T11:27:59Z
The refreshed comments add appreciation and further fixed-depth tuning anecdotes, but no controlled adaptive-versus-fixed result, upstream decision, or coding-agent workload benchmark. The adaptive selector therefore remains an open validation question rather than an emerging performance result.
2026-08-20T17:38:57Z
Refreshed comments add no controlled adaptive-versus-fixed comparison, upstream adoption decision, or coding-agent workload result. The discussion remains repetitive hardware-specific tuning evidence, leaving the adaptive selector’s practical value unvalidated.
2026-08-20T14:38:33Z
A hardware-specific comment suggests fixed MTP depth 2 can outperform depth 3, reinforcing that manual depth choice materially affects results. It still provides no adaptive-versus-fixed comparison or coding-agent workload benchmark, so the selector remains unvalidated.
2026-08-20T10:35:09Z
The refreshed DFlash2 discussion adds no controlled adaptive-versus-fixed comparison, upstream decision, or coding-agent workload result. It remains repetitive, hardware-specific context around adjacent speculative-decoding methods, leaving adaptive MTP’s value unvalidated.
2026-08-20T07:38:12Z
The new independent tuning report confirms that MTP performance is materially shaped by hardware placement, context settings, and implementation bugs, increasing the need for controlled workload-level evaluation. Because its gains combine multiple flags and fixed-depth MTP rather than testing the adaptive selector, it does not validate automatic depth selection or coding-agent throughput.
2026-08-20T07:22:46Z
evidence attached: reddit.post.1vtc0z7 — Detailed independent llama.cpp benchmarking reports substantial throughput and context gains while identifying an MTP bug, directly informing speculative-decoding validation.
2026-08-20T01:23:38Z
The refreshed discussion adds no adaptive-versus-fixed benchmark, upstream adoption decision, or coding-agent workload result. It remains repetitive amplification of adjacent DFlash2 and fixed-MTP tradeoffs, leaving the adaptive selector unvalidated.
2026-08-19T23:41:35Z
The refreshed comments remain hardware-specific discussion of DFlash2 and fixed-MTP tradeoffs, not evidence about adaptive depth selection. Without an adaptive-versus-fixed benchmark, upstream decision, or coding-agent workload result, the case’s meaning is unchanged.
2026-08-19T22:38:30Z
The refreshed DFlash2 comments add hardware-specific tradeoff discussion but no adaptive-versus-fixed benchmark, upstream decision, or coding-agent workload result. The case remains an open validation question and repetitive adjacent discussion does not justify promotion.
2026-08-19T21:44:11Z
Refreshed comments add no adaptive-versus-fixed benchmark, upstream adoption decision, or coding-agent workload result. The discussion remains repetitive amplification of adjacent DFlash2 and fixed-MTP tradeoffs, so the adaptive selector’s value is still unvalidated.
2026-08-19T18:35:09Z
The DFlash2 benchmark raises the performance bar for speculative-decoding alternatives and reinforces that gains are highly task- and hardware-dependent, but it does not test adaptive MTP selection. The case remains an open validation question pending adaptive-versus-fixed or end-to-end coding-agent benchmarks.
2026-08-19T18:23:40Z
evidence attached: reddit.post.1vsuaoj — An independent llama.cpp benchmark reports large but task-dependent gains from the related DFlash2 speculative-decoding path, materially contextualising local decoding-speed tradeoffs.
2026-08-19T14:36:05Z
The refreshed discussion adds only anecdotal configuration experience and repeats hardware- and workload-dependent MTP/ngram tradeoffs. No adaptive-versus-fixed benchmark, upstream decision, or end-to-end coding-agent result changes the open validation question.
2026-08-18T12:31:30Z
The refreshed discussion remains anecdotal configuration advice and amplification of fixed MTP/ngram tradeoffs; it adds no adaptive-versus-fixed benchmark, upstream adoption, or end-to-end coding-agent result. The case therefore remains an open validation question with no promotion signal.
2026-08-18T11:25:15Z
Hands-on fixed-MTP-plus-ngram results clarify that context growth, hardware, and mis-speculation can erase burst-throughput gains, making workload-level adaptive-depth evaluation more important. They do not test the adaptive selector itself or establish end-to-end coding-agent improvement, so the hypothesis remains open.
2026-08-18T11:22:36Z
evidence attached: reddit.post.1vrlgp5 — Hands-on llama.cpp results show MTP combined with ngram speculation can raise burst throughput while exposing tuning and mis-speculation tradeoffs.
2026-08-18T00:28:16Z
The refreshed comments remain discussion and implementation interest rather than validation; no independent benchmark, upstream adoption, or coding-agent workload result changes the case’s meaning. Keep watching, but wait for measurable throughput, acceptance, and memory comparisons.
2026-08-17T21:34:37Z
The refreshed discussion adds no independent benchmark, comparative implementation result, or upstream adoption signal; it remains repetitive interest around two experimental heuristics rather than validation of adaptive depth selection. Cool the case until measurable throughput, memory, or coding-agent evidence appears.
2026-08-17T18:39:33Z
A separate rolling-window adaptive-depth fork suggests convergent implementation interest beyond the original PR, moving the case into active watching. However, no independent throughput, memory, or end-to-end coding-agent results yet validate either heuristic.
2026-08-17T18:28:58Z
grounded: known/medium — The radar already tracks substantially the same open validation question in `radar:adaptive-speculative-decoding-300-gpu`, alongside llama.cpp MTP memory cases.
2026-08-17T18:26:30Z
origin walked (codex/luna, conf 0.97): anchor reddit.post.1vqzud4 -> echo.github.a4efe06825 by Stew Forster (stew675)
2026-08-17T18:25:43Z
case created — A concrete first-party llama.cpp pull request introduces a directly testable adaptive speculative-decoding mechanism.