2026-10-11 17:12 UTC

koalfied-coder claims a reproducible two-DGX-Spark setup runs DeepSeek Flash v4 at a sustained 67–84 tokens per second with fast prompt evaluation, making the model practically usable for high-throughput local inference.

state: expiredheat: lowuncertainty: highknownscott: mediumdeepseek local-inference inference-economicskoalfied-codertonyd2wildNVIDIAASUS

What is this?

A community-built deployment runs DeepSeek-V4-Flash across two NVIDIA DGX Spark/GB10 systems using a development build of vLLM, with reports of roughly 60–67 tok/s sustained single-stream decoding and peaks near 84 tok/s; concurrency claims reach roughly 220 aggregate tok/s. The supplied snippets include configuration details and independent engineering discussion, but they do not firmly establish koalfied-coder’s identity or ASUS’s role, and results conflict: another dual-Spark test reports stable sustained generation of only about 41 tok/s. The headline performance is therefore plausible and partly reproducible, but not uniformly confirmed across workloads and configurations.

Why it matters to Scott

The radar already tracks this exact development in `radar:deepseek-v4-twinspark-dgx-serving`. Its disputed dual-node throughput bears directly on Scott’s hardware-aware local-inference work and could inform whether scaling his self-hosted model substrate is operationally and economically worthwhile, but the conflicting results do not yet establish a reliable new conclusion.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsradar:deepseek-v4-twinspark-dgx-servingradar:concept.local-inferenceradar:concept.distributed-inferenceradar:concept.inference-economics
queries asked of Scott's wikis
  • local inference throughput economics versus API inference
  • dual-node unified-memory model serving
  • local coding-agent inference latency requirements
  • open-model quantization and speculative decoding
  • high-throughput local inference concurrency
  • DGX Spark or GB10 deployment experience

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐67-84 t/s DeepSeek flash v4 off 2x GX10s
LocalLLaMA
koalfied-coder4528
🟠 redditHumaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup
LocalLLaMA
serige2311

Interpretation history

Decision trace