2026-10-11 17:12 UTC

Independent benchmarks will determine whether DFlash2 roughly doubles Qwen3.8-27B decode throughput at 256k context on consumer GPUs while preserving output quality and reducing total wall time.

state: resolvedheat: lowuncertainty: mediumknownscott: mediumspeculative-decoding qwen local-inferenceQwenllama.cpp

What is this?

DFlash2 is a speculative-decoding method published by z-lab for accelerating models including Alibaba’s open-weight Qwen3.8-27B; a Reddit snippet says support was added through llama.cpp PR #27342. Supplied benchmark snippets report substantial throughput gains, while one implementation account describes the CUDA, verification, recurrent-state, and attention patches needed for long-context operation. However, the specific claim of roughly 2× decode throughput at full 256k context on a single RTX 5090—without quality loss and with lower end-to-end wall time—is not independently established by the snippets: they include project benchmarks, a video test, and community reports with differing hardware and speedups.

Why it matters to Scott

The radar already tracks this development in `radar:dflash-2-parallel-drafting-validation`, with the 256K Qwen3.8 hardware claim also overlapping `radar:qwen38-27b-24gb-long-context-throughput`. It matters to Scott’s active `gamepc` local-model substrate and hardware-aware inference practice because independent end-to-end benchmarks—not decode-only figures—could guide whether to adopt DFlash2 for long-context local workloads, but the supplied evidence does not yet establish that result.
dev:project.gamepcdev:concept.hardware-aware-local-inferenceip:concept.model-plus-harness-benchmark-unitip:concept.latencyradar:dflash-2-parallel-drafting-validationradar:qwen38-27b-24gb-long-context-throughputradar:concept.speculative-decodingradar:concept.long-context-inference
queries asked of Scott's wikis
  • speculative decoding verification and lossless acceleration
  • local inference performance harnesses and benchmarking
  • 256k context consumer GPU economics
  • llama.cpp optimization and deployment strategy
  • open-weight models for local agent systems
  • decode throughput versus end-to-end wall time

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (11) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Qwen3.8-27B-UD-Q4_K_XL - full 256k context + ~2x decode from DFlash2 on single RTX 5090
LocalLLaMA
maddie-lovelace1010
🟠 redditI pushed Qwen3.8-27B to 381 tps for a single request on a RTX 3090
LocalLLaMA
iamMess8272
🟠 redditOptimizing Qwen3.8-27B on one MI300X with an open-source agent toolkit: 311 to 495 tok/s
LocalLLaMA
cheptsov1515
🟠 reddit[Benchmark] DFlash2 vs MTP comparison. 5090RTX, Qwen 3.8 27B, Dynamic v3 GGUF, llama.cpp. Token generation, latency and available context.
LocalLLaMA
Opening-Broccoli9190629
🟠 redditQwen3.8-27B: 3x long-context decode speed (BF16, lossless)
LocalLLaMA
stargate425213
🟠 redditStrix Halo (8060S / gfx1151), Qwen-3.8-27B @ Q8 and Q6 UD v3, up to 256K ctx, llama.cpp, DFlash2, vision, real workloads quality and steady performances, optimized recipes, ...
LocalLLaMA
PieBru2317
🟠 redditSingle RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K
LocalLLaMA
Fz1zz2746
🟠 redditI benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
LocalLLaMA
FantasticNature75908924
🟠 reddit~3,200 tok/min on Two 3090s, No NVLink — Local Qwen Coding Agent at Full 262K Context
LocalLLaMA
CryptographerLow781703
🟠 redditThought I'd share my custom quant for RTX Pro 6000 cards. Qwen3.8-27B-heretic-ara-MXFP6-MXFP8-DFlash2
LocalLLaMA
FaustAg43
🟠 redditQwen3.8-27B NVFP4 with vision + 451K token KV-cache on one RTX 5090 (power limited to 400W) at 120 tokens/s average
LocalLLaMA
t4a89454443

Interpretation history

Decision trace