2026-10-11 16:37 UTC

sudoingX’s Ling-3.0-flash measurements reportedly show MTP n=1 raising short-prompt throughput on one Spark from about 23 tokens per second without drafting to 40.9 on code and 38.7 on prose, making speculative decoding a potentially substantial local-inference optimization.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: lowspeculative-decoding local-inferencesudoingX

What is this?

Ling-3.0-flash is an open-weight hybrid linear-attention MoE model from Ant Group's inclusionAI (July 2026 release, 262k context), trained with MTP heads that vendors recommend enabling at inference. The case's core datapoint is a Reddit account (sudoingX) reporting that MTP n=1 speculative decoding on a single DGX Spark raises short-prompt throughput from ~23 tok/s to 40.9 (code) / 38.7 (prose) — still a single unverified testimony chain. The web material now adds stronger independent corroboration for the direction: an LMSYS blog shows Ling-3.0-flash single-request decode on 4×Blackwell going from 288 to 606 tok/s with tuned NEXTN/MTP and 1120 tok/s with a DSpark confidence-scheduled drafter, and the official model card ships MTP/vLLM speculative configs as the default recommendation. The Spark-specific magnitudes remain unreplicated, but MTP speculative decoding as a first-class, workload-dependent local-inference optimization is now backed by the model vendor, a serving-stack team, and community results.

Why it matters to Scott

The independent consumer-GPU datapoint (Qwen3.8-Flash-Next MTP via a llama.cpp PR branch, CPU-resident draft experts, ~16.5 t/s at 131K context on 16GB VRAM + 64GB RAM) plus the vendor-default MTP configs converge with Scott's hardware-aware-local-inference position that runtime configuration is first-order — and it's a directly usable recipe reference for gamepc-scale hardware. But it remains the world confirming his position with another instance, not a challenge or extension of it: the Ling/Spark magnitudes are still one testimony chain, Scott runs neither Ling nor a Spark, and the n=1-shallow-drafting-sweet-spot datapoint is lineage for radar:llama-cpp-adaptive-mtp rather than news that would change what he builds or argues. Bump to medium only if the adaptive-MTP or llama.cpp MTP threads resolve against or beyond these numbers.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:llama-cpp-adaptive-mtpradar:concept.speculative-decodingradar:concept.local-inferenceradar:qwen38-flash-next-hybrid-inference
queries asked of Scott's wikis
  • hardware-aware local inference tuning position — runtime configuration and per-workload optimization
  • llama.cpp speculative decoding support and MTP draft-head branches
  • DGX Spark or unified-memory workstation experience and local model serving
  • hybrid linear-attention / Mamba-style MoE models in Scott's local stack
  • agent harness latency and tokens-per-second sensitivity for local coding agents
  • drafting depth / acceptance-rate tradeoff notes or benchmarks

Measured heat

now 7 pts/hpeak 85 pts/hcomments 2/hpeers p81momentum: accelerating2 platformsage 817h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-07 16:23 (minted)⭐ origin echo-reconstructedAccording to the Reddit account, the later benchmark graphic supplies a previously missing no-speculation baseline of about 23 tok/s and com
sudoingX on x (echo) · attributed from reddit.post.1w9v4yz · published time unknown
—
09-07 15:23first on r/LocalLLaMA · published · lag ?Higher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark
niacolhealth
—
09-07 15:23amplified on r/LocalLLaMAreddit.post.1w9v4yz
niacolhealth
peak 78 · 5 comments · 12% of case engagement
09-23 23:46amplified on r/LocalLLaMAreddit.post.1wom3fe
AvidCyclist250
peak 64 · 66 comments · 19% of case engagement
09-28 12:26amplified on r/LocalLLaMAreddit.post.1wsd7q9
Arany8
peak 0 · 15 comments · 2% of case engagement
10-01 11:18amplified on r/LocalLLaMA 👑reddit.post.1wuwrsk
jacek2023
peak 205 · 101 comments · 44% of case engagement
10-07 12:36amplified on r/LocalLLaMAreddit.post.1wzvdqe
jacek2023
peak 73 · 21 comments · 14% of case engagement
10-11 09:20amplified on r/LocalLLaMAreddit.post.1x33p4o
Dreeew84
peak 52 · 10 comments · 9% of case engagement
09-07 16:20our radar first saw it · lag ?discovery anchor: reddit.post.1w9v4yz—
pace: p84 vs 519 stories at the 720h mark (now 817h old) — ahead of minicpm5-2b-release (1.0x), behind mercury-25-diffusion-inference (1.0x)

Evidence (7) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditHigher acceptance length, slower prose: Ling’s n=1/2/3 MTP test on one Spark
LocalLLaMA
niacolhealth455
🟧 echo.x ⭐According to the Reddit account, the later benchmark graphic supplies a previously missing no-speculation baseline of about 23 tok/s and comsudoingX——
🟠 redditQwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080
LocalLLaMA
AvidCyclist2506466
🟠 redditX account claims high t/s setup, but thin on details
LocalLLaMA
Arany8015
🟠 redditQwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp
LocalLLaMA
jacek2023205101
🟠 redditfeat: add GLM5Next MTP, optimize by pwilkin · Pull Request #29928 · ggml-org/llama.cpp
LocalLLaMA
jacek20237321
🟠 redditReminder: try probabilistic MTP if you missed it. Decode +14% on prose
LocalLLaMA
Dreeew845210

Interpretation history

Decision trace