2026-10-11 17:12 UTC

Independent reproduction will determine whether the TwinSpark recipe can serve DeepSeek V4 Flash across two DGX Spark systems at roughly 75 tokens per second while retaining practical long-context operation.

state: expiredheat: lowuncertainty: highconvergesscott: mediumdeepseek-v4-flash local-inference dgx-spark speculative-decodingDeepSeekNVIDIA

What is this?

TwinSpark is a community-published serving recipe for running DeepSeek’s official 284B mixture-of-experts DeepSeek V4 Flash model across two NVIDIA DGX Spark systems using tensor parallelism, MTP speculative decoding, FP8 KV cache, and a 200Gbps interconnect. The supplied repository reports about 41 tokens/s single-stream decode at up to 200K context, while an NVIDIA forum report measures roughly 44 tokens/s and warns that cold long-context prefill reaches about 53 seconds at 32K and 250 seconds at 128K. Other supplied snippets describe separate dual-Spark runs reaching 1M-token context, but they do not establish the hypothesized 75 tokens/s; on present evidence, independent reproduction supports practical long-context operation with significant prefill costs, not the claimed speed.

Why it matters to Scott

The recipe converges with Scott’s hardware-aware local-inference practice by making accelerator placement, precision, interconnect, KV-cache policy, and speculative decoding explicit—and its reproduced 41–44 tok/s performance plus steep long-context prefill costs are actionable constraints for evaluating local model-serving substrates. It is not yet a 75 tok/s breakthrough, but it extends the radar’s existing DeepSeek V4 Flash investigation from a single RTX 5090 to distributed DGX Spark serving.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.latencyip:concept.evaluation-driven-developmentradar:deepseek-v4-flash-1m-rtx5090radar:concept.distributed-inferenceradar:concept.speculative-decodingradar:concept.long-context-inference
queries asked of Scott's wikis
  • local inference economics versus cloud APIs
  • multi-node inference on prosumer AI hardware
  • speculative decoding performance and acceptance rates
  • long-context latency versus usable context windows
  • open-model sovereignty and on-prem deployment
  • coding-agent integration with local model servers

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐DeepSeek-V4-Flash-0731 (284B MoE) at 75 tok/s on 2× DGX Spark — full recipe, 11 gotchas, reboot-proof cluster, Codex CLI integration
LocalLLaMA
Striking-Swim67021712

Interpretation history

Decision trace