A community-built deployment runs DeepSeek-V4-Flash across two NVIDIA DGX Spark/GB10 systems using a development build of vLLM, with reports of roughly 60–67 tok/s sustained single-stream decoding and peaks near 84 tok/s; concurrency claims reach roughly 220 aggregate tok/s. The supplied snippets include configuration details and independent engineering discussion, but they do not firmly establish koalfied-coder’s identity or ASUS’s role, and results conflict: another dual-Spark test reports stable sustained generation of only about 41 tok/s. The headline performance is therefore plausible and partly reproducible, but not uniformly confirmed across workloads and configurations.
The radar already tracks this exact development in `radar:deepseek-v4-twinspark-dgx-serving`. Its disputed dual-node throughput bears directly on Scott’s hardware-aware local-inference work and could inform whether scaling his self-hosted model substrate is operationally and economically worthwhile, but the conflicting results do not yet establish a reliable new conclusion.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsradar:deepseek-v4-twinspark-dgx-servingradar:concept.local-inferenceradar:concept.distributed-inferenceradar:concept.inference-economics
queries asked of Scott's wikis
- local inference throughput economics versus API inference
- dual-node unified-memory model serving
- local coding-agent inference latency requirements
- open-model quantization and speculative decoding
- high-throughput local inference concurrency
- DGX Spark or GB10 deployment experience
2026-09-02T16:46:10Z
After repeated checks, no matched reproduction, logs, or coding-agent benchmark has narrowed the throughput discrepancy, and no confirming evidence is now expected within the episode’s horizon. Preserve dual-Spark operability as credible, but let the specific 67–84 tok/s and inference-economics claim fade pending a genuinely new reproduction.
2026-08-31T16:37:57Z
The refreshed discussion adds no matched reproduction, logs, or coding-agent benchmark, so it does not narrow the gap between the headline 67–84 tok/s claim and the slower independent result. Dual-Spark operability remains credible, but dependable throughput and inference economics remain unsettled.
2026-08-30T16:31:35Z
The refreshed comments add no matched configuration, logs, or workload-relevant benchmark, so they do not resolve the gap between the reported 67–84 tok/s and the slower independent result. Dual-Spark operability remains credible, while headline throughput and economics stay unsettled.
2026-08-30T15:34:36Z
The refreshed discussion still supplies no matched configuration, logs, coding benchmark, or independent reproduction of the claimed 67–84 tok/s range. Dual-Spark operability remains credible, but the headline throughput and inference-economics conclusion are unchanged and unsettled.
2026-08-30T05:31:35Z
The refreshed discussion adds only an unsupported claim that GLM 5.3 Flash is preferable, not a matched reproduction or new benchmark evidence. DeepSeek V4 Flash is credibly operable on two DGX Sparks, but the claimed 67–84 tok/s performance and resulting inference economics remain unresolved.
2026-08-30T01:24:03Z
Refreshed comments add requests for coding benchmarks and alternative quantization tests, but no matched configuration, logs, or reproduction of the claimed 67–84 tok/s range. The setup’s operability remains credible while its headline throughput and practical economics remain unsettled.
2026-08-29T23:25:54Z
The case has moved beyond a single-operator anecdote because a second user independently runs and benchmarks the same model on two DGX Sparks. However, the available results do not reproduce the headline 67–84 tok/s and instead reinforce that throughput is configuration- and workload-sensitive, so the practical-performance claim remains unsettled.
2026-08-29T23:22:37Z
evidence attached: reddit.post.1w215qm — Independent two-DGX-Spark benchmarking directly tests the case’s claim about practical local throughput and compares it with a competing model.
2026-08-29T19:42:03Z
No substantive evidence arrived: modest engagement growth without new comments, logs, or independent reproduction leaves the conflicting throughput claim unresolved. The setup remains worth tracking, but this delta adds no confidence or urgency.
2026-08-29T19:32:33Z
grounded: known/medium — The radar already tracks this exact development in `radar:deepseek-v4-twinspark-dgx-serving`. Its disputed dual-node throughput bears directly on Scott’s hardwa
2026-08-29T19:27:10Z
case created — This is a concrete performance claim tied to a public setup repository, but it currently has only one operator report.