2026-10-11 17:09 UTC

speculative-decoding

band: hotmomentum: stable score: 1.0
temperature history

Episodes (19)

Independent testing will determine whether adaptive speculative decoding delivers substantial local-inference speedups on a $300 consumer GPU without unacceptable output-quality regressions.
expirednovelscott: low
Independent benchmarks will reproduce that Qwen3.6-27B receives larger speculative-decoding speedup multipliers at Q8 than Q6 and Q4 because draft-and-verify overhead scales less with weight size than base decoding.
expirednovelscott: low
llama.cpp maintainers will revert or gate default loading of bundled MTP tensors after reports that models consume extra RAM or VRAM even when MTP speculative decoding is disabled.
expirednovelscott: low
Independent benchmarks will determine whether tool-call-aware speculative decoding materially reduces agent inference latency or cost without degrading tool selection or argument correctness.
expiredconvergesscott: medium
Independent reproduction will determine whether the TwinSpark recipe can serve DeepSeek V4 Flash across two DGX Spark systems at roughly 75 tokens per second while retaining practical long-context operation.
expiredconvergesscott: medium
Independent testing will determine whether the v100-skinny kernels make NVFP4 weights and speculative decoding a practically high-throughput inference path for Qwen-class models on inexpensive V100 GPUs.
expiredknownscott: low
Independent benchmarks will determine whether mlx-dspark’s speculative decoding reproducibly accelerates Muse Glimmer 30B inference by roughly 2–3Γ— on Apple Silicon without changing model output.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s adaptive MTP mode selects speculative-decoding depth effectively enough to improve coding-agent throughput without manual tuning.
expiredknownscott: medium
Independent benchmarks will determine whether DFlash 2’s released parallel-drafting models and llama.cpp integration deliver practically useful speculative-decoding speedups over MTP for Qwen3.8 and Muse Glimmer local inference.
resolvedknownscott: medium
Independent benchmarks will determine whether DFlash2 roughly doubles Qwen3.8-27B decode throughput at 256k context on consumer GPUs while preserving output quality and reducing total wall time.
resolvedknownscott: medium
Independent benchmarks will determine whether Liquid AI’s released DSpark speculative-decoding support delivers up to 3.2Γ— faster practical local inference for LFM2.5 models.
expiredknownscott: medium
Independent benchmarks will determine whether llama.cpp’s GLM-4.5-Air MTP support materially improves local inference speed across memory-rich, compute-limited hardware without reducing output quality or stability.
expiredknownscott: medium
NVIDIA claims jointly designing speculative-decoding models and their serving systems yields materially better LLM inference throughput and economics than optimizing draft models and infrastructure separately.
expiredknownscott: low
Extension-Bid-639 claims a build combining quantization, expert caching, host-RAM offload, and multi-token prediction raises full-261K-context Qwen3.8-Flash-Next decode throughput from 25–29 to 37–41 tokens per second on two RTX 3090 GPUs, potentially making long-context local coding inference practical on commodity multi-GPU systems.
resolvedknownscott: medium
vLLM presents speculative decoding on AMD GPUs as an inference optimization, potentially reducing generation latency for AMD-based model serving.
expirednovelscott: low
sudoingX’s Ling-3.0-flash measurements reportedly show MTP n=1 raising short-prompt throughput on one Spark from about 23 tokens per second without drafting to 40.9 on code and 38.7 on prose, making speculative decoding a potentially substantial local-inference optimization.
corroboratedconvergesscott: low
Hayder Tirmazi claims four implementation changes to llama.cpp's n-gram caches (unnecessary map copies removed, flat hash maps, and related fixes) make prompt-lookup drafting up to 42x faster with up to 2.6x less memory at unchanged acceptance rates; an upstream llama.cpp merge or independent reproduction would establish lean n-gram-cache engineering as a standard local-inference optimization.
watchingconvergesscott: medium
llama.cpp contributor pratiknarola-t's merged PR #29869 claims few-row MMA Metal mat-mul and batched-copy kernels turn DFlash2 speculative decoding of Qwen3.8-27B from slower than serial decoding (~30 tok/s) into ~110 tok/s on an M3 Ultra β€” with the kernels, tests, and benchmarks disclosed as Claude Code-written β€” and the gains replicating across Apple GPUs, models, and specdec drafters would establish few-row matmul optimization as the standard enabler of speculative decoding on non-tensor-API Apple Silicon.
seedconvergesscott: high
A LocalLLaMA builder claims a custom CUDA megakernel fusing entire speculative-decoding cycles achieves 1.4–1.9Γ— speedups over llama.cpp for Qwen3.8-27B on a single RTX 3090, with the kernel written using Claude Opus 5.5.
watchingnovelscott: high

Trajectory notes