2026-10-11 17:12 UTC

Independent reproduction will determine whether DeepSeek V4 Flash can sustain useful million-token inference at practical speeds on a single RTX 5090 using CPU-offloaded experts and adaptive speculative decoding.

state: resolvedheat: lowuncertainty: highconvergesscott: mediumlocal-inference open-models long-contextBlackBeardAIDeepSeek

What is this?

The case centers on a BlackBeardAI-reported repository patch and benchmark claiming that DeepSeek V4 Flash can run at a full one-million-token context on one RTX 5090 by CPU-offloading mixture-of-experts weights and adapting speculative decoding by generation phase. The cited title reports roughly 13.8 tokens/s during reasoning and 17.0 tokens/s for final or code output, while a KTransformers snippet warns that speculative decoding can become counterproductive when CPU expert bandwidth is the bottleneck. The supplied results support the general feasibility of single-workstation CPU-offloaded inference, but they do not independently reproduce the exact setup or throughput, and other snippets report materially different speeds on different hardware and configurations.

Why it matters to Scott

The claimed result extends Scott’s hardware-aware local-inference work by proposing a concrete operating point—CPU-offloaded experts plus phase-adaptive speculative decoding—for useful 1M-context inference on one consumer GPU. If reproduced it could inform his gamepc runtime policy and test his Residence Dividend thesis in a locally deployable setting, but the result remains unverified and the radar already tracks closely adjacent expert-streaming and speculative-decoding cases.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:source.the-inference-field-ebookip:concept.residence-dividendradar:adaptive-speculative-decoding-300-gpuradar:llama-cpp-hot-expert-gpu-cacheradar:kimi-linear-local-validationradar:concept.long-context-inferenceradar:concept.expert-streaming
queries asked of Scott's wikis
  • single-GPU local inference economics and CPU offload
  • mixture-of-experts CPU bandwidth bottlenecks
  • adaptive speculative decoding by workload phase
  • million-token context utility and benchmark validity
  • long-context models versus RAG and agent memory
  • open-weight model sovereignty on consumer hardware

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (21) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit[Update] DeepSeek-V4-Flash-0731 on a single RTX 5090: phase-adaptive DSpark K1/K2 with dual CUDA graphs — ~13.8 tok/s reasoning, ~17.0 tok/s final/code at full 1M context
LocalLLaMA
BlackBeardAI65
🟧 echo.github ⭐The earliest primary artifact is the repository's first commit, containing the adaptive DSpark patch and benchmark README. It reports K1 forblackbeardlabs——
🟠 redditDeepSeek-V4-Flash on SM89 4x48gb 4090s with DSpark
LocalLLaMA
dangerous_inference2932
🟠 redditDeepSeek-v4-Flash-Mini 54GB GGUF running at ~20.5 t/s
LocalLLaMA
giveen1916
🟠 redditI updated my localy run benchmark with DeepSeek V4 Flash 0731
LocalLLaMA
WonderRico5538
🟠 reddit[DSV4-0731] 1MM lossless ctx on 3x3090+DDR5 300PP 15TG
LocalLLaMA
Important_Quote_118028
🟠 redditjabbatheduck/DeepSeek-v4-flash-mini · Hugging Face
LocalLLaMA
giveen2715
🟠 redditDeepSeek V4 Flash 0731 at 10–17 t/s (nothink) on MacBook M5 Pro **64GB***, partly via SSD streaming
LocalLLaMA
vogelvogelvogelvogel2833
🟠 redditRecommendations for optimizing an agentic Deepseek V4 Flash setup
LocalLLaMA
neverbyte19
🟠 redditDeepseek V4 Flash just hit Colibri, does anyone have numbers?
LocalLLaMA
schaka2720
🟠 redditFinal optimization: from ~10 tok/s to ~15 tok/s on DeepSeek-V4-Flash-0731 at 128K ctx - 1 RTX 3090
LocalLLaMA
Ok_Ninja75261834
🟧 hnDeepSeek V4 Flash 0731tosh781468
🟠 redditDeepSeek V4 Flash 0731 - ARC-AGI Results
LocalLLaMA
johnnyApplePRNG16053
🟠 redditDeepSeek v4 Flash 0731 on H100 node
LocalLLaMA
SlipperyCorruptor017
🟠 redditServing Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
LocalLLaMA
StartupTim1958
🟠 redditExtremely slow DSpark draft model performance (1-2 t/s) with DeepSeek-V4-Flash on llama-server compared to MTP?
LocalLLaMA
Easy_Werewolf79031718
🟠 redditUpdated benchmark: Deepseek V4 Flash on SlopCodeBench (local)
LocalLLaMA
corruptbytes1516
🟠 reddit300b on 32gb MoE-streaming findings + optimisations
LocalLLaMA
maddie-lovelace156
🟠 redditDeepSeek-V4-Flash-0731 Q8_K_XL sometimes stops mid-task in OpenCode - anyone else seeing this?
LocalLLaMA
dieSpaghettiCarbona1215
🟧 hnDeepSeekV4SSD: DeepSeek-V4-Flash-0731 on an M-series MacJKCalhoun20
🟠 redditDeepSeek V4 Flash 0731 is the ‘killer app’ that is going to sell A LOT of DGX Sparks
LocalLLaMA
Porespellar238252

Interpretation history

Decision trace