llama.cpp is a local LLM inference project maintained under ggml-org that is adding multi-token prediction (MTP) support for speculative decoding. Supplied reports describe excessive memory use in two different contexts: a SYCL-specific issue while MTP is active, and a merged fix for draft-side resources surviving server sleep/resume; neither directly establishes the claimed regression that bundled MTP tensors load when speculation is disabled. The snippets therefore do not establish that maintainers plan to revert or gate default tensor loading, although they show active work on MTP-related memory costs and cleanup.
2026-08-07T19:36:19Z
The prediction has aged out without any upstream patch, maintainer intent, or implementation supporting a revert or gate; refreshed comments are repetitive discussion rather than new evidence. Automatic MTP adoption elsewhere further weakens the expectation of a near-term llama.cpp reversal, though it does not disprove one later.
2026-08-05T10:26:31Z
The Ollama report shows that automatic bundled-MTP use and its performance tradeoffs are spreading elsewhere, but it does not evidence llama.cpp maintainer intent or a patch to revert or gate loading. The prediction remains dormant and may now face a competing ecosystem trend toward automatic enablement.
2026-08-05T10:21:15Z
evidence attached: reddit.post.1vg2yh9 β Independent Ollama testing shows bundled MTP is now enabled automatically on Macs and reports a measurable performance tradeoff, relevant to default-loading behavior.
2026-08-03T02:25:32Z
After 48 hours, only negligible Reddit movement has appeared and there is still no upstream response, patch, or maintainer intent supporting a revert or gate. Treat the case as dormant until implementation activity emerges rather than continuing engagement-driven checks.
2026-07-31T01:24:25Z
No new upstream response, patch, or independent evidence supports the predicted revert or gate; the latest activity continues to amplify the known memory-cost premise rather than maintainer intent. Keep the case open, but revisit only if implementation activity appears.
2026-07-30T16:23:00Z
The added user report concerns memory behavior while MTP mode is active and does not establish default loading when MTP is disabled, much less maintainer intent to revert or gate it. The prediction remains open but should wait for an upstream issue, patch, or maintainer response.
2026-07-30T15:21:40Z
evidence attached: reddit.post.1vawdma β Direct user report supports the open hypothesis that MTP tensors consume system memory by default even when speculative decoding is not actively used.
2026-07-30T14:26:20Z
The latest attachment adds no upstream patch, maintainer intent, or independent corroboration for reverting or gating MTP tensor loading. The case remains an unfulfilled prediction, and repetitive amplification should not trigger another look without implementation activity.
2026-07-30T08:24:33Z
No upstream action, patch, or maintainer response has appeared since the regression was confirmed. The case remains a credible but unfulfilled prediction β repeated engagement-only updates have not advanced it.
2026-07-30T06:23:33Z
The newly attached material adds no upstream response, patch, or independent evidence that maintainers will revert or gate default MTP tensor loading. The regression premise remains credible, but repeated amplification no longer warrants frequent checks without implementation activity.
2026-07-30T04:22:01Z
The newly attached material still only repeats the confirmed disabled-path memory regression and adds no maintainer intent, patch, or independent corroboration for a revert or gate. Keep the prediction open but stop repricing engagement-only updates until upstream action appears.
2026-07-30T02:21:51Z
No maintainer response, patch, or independent evidence has appeared; the attachment again supports only the known memory regression, not the predicted revert or gate. Pause frequent checks until upstream action changes the case.
2026-07-30T01:22:52Z
The attachment adds no maintainer response, patch, or independent evidence for reverting or gating default MTP loading; it only repeats the established regression premise. Further engagement should not advance the case without upstream action.
2026-07-30T00:24:34Z
The attached evidence still only confirms the known disabled-path memory cost and adds no maintainer intent, patch, or independent corroboration for reverting or gating the behavior. Repetitive amplification does not advance the prediction.
2026-07-29T23:22:17Z
The added observation provides no new evidence of maintainer intent, a proposed gate, or a revert; attention remains repetitive amplification of the already-confirmed regression. Cool the case until upstream implementation or maintainer action appears.
2026-07-29T22:25:45Z
The new attachment still only confirms the known disabled-path memory regression; it provides no independent maintainer response, proposed gate, or reverting implementation. The prediction remains plausible but uncorroborated while the upstream issue is fresh.
2026-07-29T21:22:23Z
The latest attachment adds no independent evidence of maintainer intent or a concrete fix; it reiterates the already-confirmed disabled-path memory cost. The case remains a credible regression report but the predicted revert or gating is still speculative.
2026-07-29T20:22:50Z
The upstream PR disclosure now directly confirms the disabled-path memory regression, strengthening the premise beyond the Reddit report. There is still no maintainer response or implementation indicating that default loading will be reverted or gated.
2026-07-29T19:24:20Z
grounded: novel/none β No intersection found in Scottβs wikis, and no radar page already tracks this development. The supplied evidence also does not establish the predicted revert or
2026-07-29T19:23:46Z
origin walked (codex/luna, conf 0.97): anchor reddit.post.1va54em -> echo.github.43d7b04e23 by satindergrewal
2026-07-29T19:22:18Z
case created β The report identifies a bounded behavior change tied to a specific merged pull request and a measurable memory regression that maintainers can confirm or fix.