speculative-decoding
band: hotmomentum: stable
score: 1.0
Episodes (19)
Trajectory notes
- 2026-09-23T23:21:09Z: qwen38-flash-dual-3090-speedup closed (absorbed) β The radar already tracks this development through `radar:llama-cpp-hot-expert-offload` and `radar:llama-cpp-adaptive-mtp`; this update combines those techniques and claims a higher long-context throughput result rather than int
- 2026-09-10T01:26:25Z: vllm-amd-speculative-decoding closed (faded) β This is an adjacent optimization to Scottβs hardware-aware local inference practice, but the supplied hits establish CUDA/Ollama use rather than an AMD/vLLM deployment, and the grounding supplies no measured latency gains that woul
- 2026-09-04T23:31:14Z: nvidia-speculative-decoding-codesign closed (faded) β The radar already tracks speculative-decoding throughput as a joint function of draft strategy, model configuration, and serving implementation, notably on `radar:concept.speculative-decoding`, `radar:dflash-2-parallel-draft
- 2026-08-31T06:28:56Z: llama-cpp-adaptive-mtp closed (faded) β The radar already tracks substantially the same open validation question in `radar:adaptive-speculative-decoding-300-gpu`, alongside llama.cpp MTP memory cases. Results on end-to-end coding-agent throughput would still bear directly on Sc
- 2026-08-28T03:29:44Z: llama-cpp-glm45-air-mtp closed (faded) β The radar already tracks llama.cpp MTP as an open performance-and-regression story in `radar:llama-cpp-adaptive-mtp` and `radar:llama-cpp-mtp-default-memory-regression`; GLM-4.5-Air support is a model-specific extension rather than a new
- 2026-08-27T23:42:46Z: lfm25-dspark-speculative-decoding closed (faded) β Scott already holds the relevant position that vendor-reported performance claims require representative, hardware-specific validation, and the radar already tracks this model family in `radar:lfm2-5-2-6b-edge-agent-validation`
- 2026-08-23T17:26:32Z: qwen38-dflash2-long-context-speedup closed (disproved) β The radar already tracks this development in `radar:dflash-2-parallel-drafting-validation`, with the 256K Qwen3.8 hardware claim also overlapping `radar:qwen38-27b-24gb-long-context-throughput`. It matters to Scottβs acti
- 2026-08-19T16:54:39Z: dflash-2-parallel-drafting-validation closed (absorbed) β The radar already tracks the same practical decision surface in `radar:llama-cpp-adaptive-mtp`, with model-specific open cases for Qwen3.8-27B MTP and Muse Glimmer speculative decoding. DFlash 2 adds a competing drafting
- 2026-08-17T03:29:07Z: v100-skinny-nvfp4-speculative-decoding closed (faded) β This is another unverified throughput-claim benchmark case in a pattern the radar already tracks extensively β radar:qwen36-quant-specdecode-scaling covers the same Qwen3.6-27B quant/speculative-decoding scaling story, and
- 2026-08-16T15:38:00Z: mlx-dspark-muse-glimmer-speedup closed (faded) β The radar already tracks Muse Glimmer 30Bβs practical local inference in `radar:meta-muse-open-weights-local-inference` and separately tracks speculative decoding and MLX on Apple Silicon. The claimed mlx-dspark gain bears on Sco