Independent reproduction will determine whether the TwinSpark recipe can serve DeepSeek V4 Flash across two DGX Spark systems at roughly 75 tokens per second while retaining practical long-context operation.
state: expiredheat: lowuncertainty: highconvergesscott: mediumdeepseek-v4-flash local-inference dgx-spark speculative-decodingDeepSeekNVIDIA
What is this?
TwinSpark is a community-published serving recipe for running DeepSeek’s official 284B mixture-of-experts DeepSeek V4 Flash model across two NVIDIA DGX Spark systems using tensor parallelism, MTP speculative decoding, FP8 KV cache, and a 200Gbps interconnect. The supplied repository reports about 41 tokens/s single-stream decode at up to 200K context, while an NVIDIA forum report measures roughly 44 tokens/s and warns that cold long-context prefill reaches about 53 seconds at 32K and 250 seconds at 128K. Other supplied snippets describe separate dual-Spark runs reaching 1M-token context, but they do not establish the hypothesized 75 tokens/s; on present evidence, independent reproduction supports practical long-context operation with significant prefill costs, not the claimed speed.
Why it matters to Scott
The recipe converges with Scott’s hardware-aware local-inference practice by making accelerator placement, precision, interconnect, KV-cache policy, and speculative decoding explicit—and its reproduced 41–44 tok/s performance plus steep long-context prefill costs are actionable constraints for evaluating local model-serving substrates. It is not yet a 75 tok/s breakthrough, but it extends the radar’s existing DeepSeek V4 Flash investigation from a single RTX 5090 to distributed DGX Spark serving.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.latencyip:concept.evaluation-driven-developmentradar:deepseek-v4-flash-1m-rtx5090radar:concept.distributed-inferenceradar:concept.speculative-decodingradar:concept.long-context-inference
queries asked of Scott's wikis
- local inference economics versus cloud APIs
- multi-node inference on prosumer AI hardware
- speculative decoding performance and acceptance rates
- long-context latency versus usable context windows
- open-model sovereignty and on-prem deployment
- coding-agent integration with local model servers
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-08-14T18:37:13Z
No independent reproduction or benchmark artifact arrived within the case horizon, and repeated reobservations yielded only commentary. The claimed 75 tok/s result remains an isolated report, so the episode has faded without changing the grounded 41–44 tok/s assessment.
2026-08-12T17:39:40Z
The expanded discussion remains commentary rather than reproduction: it raises valid KV-cache quality and power questions but supplies no measurements or benchmark artifacts. The 75 tok/s and practical long-context claims therefore remain a single-author result against grounded 41–44 tok/s evidence.
2026-08-11T16:49:49Z
The refreshed comments add praise, a qualitative Spark-versus-Strix-Halo observation, and a power-use question, but no independent benchmark or operational evidence. The 75 tok/s and practical long-context claims remain a single-author result, so the case stays cool and uncorroborated.
2026-08-11T14:55:14Z
The only new signal is modest engagement without comments, benchmark artifacts, or an independent 75 tok/s reproduction. Existing grounded reports still point to roughly 41–44 tok/s with costly long-context prefill, so the headline performance claim remains unsettled and the case can cool.
2026-08-11T14:36:54Z
grounded: converges/medium — The recipe converges with Scott’s hardware-aware local-inference practice by making accelerator placement, precision, interconnect, KV-cache policy, and specula
2026-08-11T14:31:56Z
case created — The detailed deployment recipe and benchmark constitute a concrete, independently testable local-inference claim distinct from existing DeepSeek evaluation and memory-constrained serving cases.
Decision trace
- 08-15 04:37expireNo independent reproduction or benchmark artifact arrived within the case horizon, and repeated reobservations yielded only commentary. The claimed 75 tok/s result remains an isolated report, so the e
- 08-15 04:37alert_silentThe staleness trigger carries no consequential new fact; without a reproduction, artifact, or quantified operational result, there is nothing Scott needs before a future fresh episode.
- 08-15 04:37alert_routeThe staleness trigger carries no consequential new fact; without a reproduction, artifact, or quantified operational result, there is nothing Scott needs before a future fresh episode.
- 08-13 03:39repriceThe expanded discussion remains commentary rather than reproduction: it raises valid KV-cache quality and power questions but supplies no measurements or benchmark artifacts. The 75 tok/s and practica
- 08-13 03:39alert_silentNo consequential fact has changed; refreshed comments neither reproduce the throughput claim nor quantify long-context quality, prefill, or power, so this can wait for routine review.
- 08-13 03:39alert_routeNo consequential fact has changed; refreshed comments neither reproduce the throughput claim nor quantify long-context quality, prefill, or power, so this can wait for routine review.
- 08-12 02:49repriceThe refreshed comments add praise, a qualitative Spark-versus-Strix-Halo observation, and a power-use question, but no independent benchmark or operational evidence. The 75 tok/s and practical long-co
- 08-12 02:49alert_silentNo consequential new fact arrived; the comments neither reproduce the throughput claim nor provide power, prefill, or long-context measurements, so this can wait for a normal briefing.
- 08-12 02:49alert_routeNo consequential new fact arrived; the comments neither reproduce the throughput claim nor provide power, prefill, or long-context measurements, so this can wait for a normal briefing.
- 08-12 02:21sensor_dirtycomment_update
- 08-12 00:55repriceThe only new signal is modest engagement without comments, benchmark artifacts, or an independent 75 tok/s reproduction. Existing grounded reports still point to roughly 41–44 tok/s with costly long-c
- 08-12 00:55alert_silentNo consequential new fact arrived; engagement alone does not validate the reported throughput or change a decision before the next briefing.
- 08-12 00:55alert_routeNo consequential new fact arrived; engagement alone does not validate the reported throughput or change a decision before the next briefing.
- 08-12 00:53alert_silentThe published recipe is directly relevant to local-inference engineering, but the consequential 74.8 tok/s and practical 1M-context claims currently rest on a single low-engagement self-report. With n
- 08-12 00:53surface_candidateThe published recipe is directly relevant to local-inference engineering, but the consequential 74.8 tok/s and practical 1M-context claims currently rest on a single low-engagement self-report. With n
- 08-12 00:53alert_routeThe published recipe is directly relevant to local-inference engineering, but the consequential 74.8 tok/s and practical 1M-context claims currently rest on a single low-engagement self-report. With n
- 08-12 00:36groundThe recipe converges with Scott’s hardware-aware local-inference practice by making accelerator placement, precision, interconnect, KV-cache policy, and speculative decoding explicit—and its reproduce
- 08-12 00:31createThe detailed deployment recipe and benchmark constitute a concrete, independently testable local-inference claim distinct from existing DeepSeek evaluation and memory-constrained serving cases.