2026-10-11 17:12 UTC

Independent testing will determine whether the ds4-8gb-cpu implementation can run DeepSeek V4 Flash with about 7.7 GiB of RAM through NVMe demand paging at practically useful speed and quality.

state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference nvme-demand-paging memory-constrained-servingbaker27727DeepSeek

What is this?

The case concerns a proposed `ds4-8gb-cpu` implementation, attributed in the case to baker27727, that would attempt to run DeepSeek V4 Flash in roughly 7.7 GiB of RAM by demand-paging model data from NVMe storage. The supplied results do not document this implementation directly or provide an independent benchmark; instead, they describe conventional 2-bit builds around 87 GB with a 92–102 GB memory floor, while other reported deployments use substantially larger systems. Whether NVMe paging can overcome that gap at useful generation speed and acceptable quality therefore remains an unverified experimental claim.

Why it matters to Scott

The radar already tracks essentially the same unresolved NVMe/out-of-core inference proposition in `radar:hotpin-lossless-moe-streaming`, `radar:kimi-k3-nvme-expert-streaming`, and `radar:slipstream-ssd-moe-streaming`. It touches Scott’s hardware-aware local-inference work and capability-audit discipline, but without an implementation or independent speed-and-quality benchmarks it is another unverified example of an established pattern, not yet information that changes what he would build or argue.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditradar:hotpin-lossless-moe-streamingradar:kimi-k3-nvme-expert-streamingradar:slipstream-ssd-moe-streamingradar:concept.local-inference
queries asked of Scott's wikis
  • NVMe demand paging for LLM inference
  • memory-constrained local model serving
  • storage bandwidth bottlenecks in token generation
  • CPU-only inference practical throughput
  • extreme quantization quality tradeoffs
  • out-of-core model execution architecture

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐DeepSeek V4 Flash with 7.7 GiB RAM using NVMe demand pagingniqoslo10
🟧 hnRunning DeepSeek V4 Flash on an RTX 5080 with 16GB VRAM Under Linux/WSL2 via DS4peppe20017523

Interpretation history

Decision trace