Independent testing will determine whether the ds4-8gb-cpu implementation can run DeepSeek V4 Flash with about 7.7 GiB of RAM through NVMe demand paging at practically useful speed and quality.
state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference nvme-demand-paging memory-constrained-servingbaker27727DeepSeek
What is this?
The case concerns a proposed `ds4-8gb-cpu` implementation, attributed in the case to baker27727, that would attempt to run DeepSeek V4 Flash in roughly 7.7 GiB of RAM by demand-paging model data from NVMe storage. The supplied results do not document this implementation directly or provide an independent benchmark; instead, they describe conventional 2-bit builds around 87 GB with a 92–102 GB memory floor, while other reported deployments use substantially larger systems. Whether NVMe paging can overcome that gap at useful generation speed and acceptable quality therefore remains an unverified experimental claim.
Why it matters to Scott
The radar already tracks essentially the same unresolved NVMe/out-of-core inference proposition in `radar:hotpin-lossless-moe-streaming`, `radar:kimi-k3-nvme-expert-streaming`, and `radar:slipstream-ssd-moe-streaming`. It touches Scott’s hardware-aware local-inference work and capability-audit discipline, but without an implementation or independent speed-and-quality benchmarks it is another unverified example of an established pattern, not yet information that changes what he would build or argue.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditradar:hotpin-lossless-moe-streamingradar:kimi-k3-nvme-expert-streamingradar:slipstream-ssd-moe-streamingradar:concept.local-inference
queries asked of Scott's wikis
- NVMe demand paging for LLM inference
- memory-constrained local model serving
- storage bandwidth bottlenecks in token generation
- CPU-only inference practical throughput
- extreme quantization quality tradeoffs
- out-of-core model execution architecture
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-14T17:38:52Z
No independent CPU/NVMe benchmark or reproducible implementation arrived within the case’s horizon, and the adjacent RTX 5080 work still does not test the 7.7 GiB claim. This has faded into an unvalidated duplicate of the broader out-of-core inference pattern and can be rediscovered if real results emerge.
2026-08-12T16:51:20Z
The RTX 5080/16GB DS4 implementation establishes an adjacent consumer-hardware test path, but it neither tests nor validates the much stronger 7.7 GiB CPU/NVMe demand-paging claim. The case now has a concrete lead worth watching, while practical speed, quality, and memory behavior remain unknown.
2026-08-12T16:24:03Z
evidence attached: hn.story.49273916 — This is independent evidence that DeepSeek V4 Flash can be made usable on constrained consumer hardware, materially informing the case about practical low-memory serving.
2026-08-11T12:51:03Z
Re-evaluation found no implementation details, benchmarks, independent tests, or discussion beyond the original title claim. The case remains an unverified duplicate of the broader NVMe/out-of-core inference pattern and cools despite the hot surrounding topic.
2026-08-11T12:37:36Z
grounded: known/low — The radar already tracks essentially the same unresolved NVMe/out-of-core inference proposition in `radar:hotpin-lossless-moe-streaming`, `radar:kimi-k3-nvme-ex
2026-08-11T12:34:55Z
case created — The released implementation makes a specific, testable claim about extending large-model inference to severely memory-constrained systems.
Decision trace
- 08-15 03:38expireNo independent CPU/NVMe benchmark or reproducible implementation arrived within the case’s horizon, and the adjacent RTX 5080 work still does not test the 7.7 GiB claim. This has faded into an unvalid
- 08-15 03:38alert_silentThe only change is elapsed staleness; there is no new consequential evidence to report. A reproducible 7.7 GiB CPU/NVMe test with throughput, memory, and quality measurements would justify reopening t
- 08-15 03:38alert_routeThe only change is elapsed staleness; there is no new consequential evidence to report. A reproducible 7.7 GiB CPU/NVMe test with throughput, memory, and quality measurements would justify reopening t
- 08-13 02:51repriceThe RTX 5080/16GB DS4 implementation establishes an adjacent consumer-hardware test path, but it neither tests nor validates the much stronger 7.7 GiB CPU/NVMe demand-paging claim. The case now has a
- 08-13 02:51alert_silentThe implementation lead lacks reproducible throughput, quality, and memory-behavior results and uses materially different hardware from the claimed 7.7 GiB CPU/NVMe setup. It can wait for the next bri
- 08-13 02:51alert_routeThe implementation lead lacks reproducible throughput, quality, and memory-behavior results and uses materially different hardware from the claimed 7.7 GiB CPU/NVMe setup. It can wait for the next bri
- 08-13 02:26alert_silentA new RTX 5080/16GB DS4 repository is a concrete implementation lead, but the visible evidence provides no throughput, quality, memory-behavior, or reproducibility results. It also does not validate t
- 08-13 02:26alert_routeA new RTX 5080/16GB DS4 repository is a concrete implementation lead, but the visible evidence provides no throughput, quality, memory-behavior, or reproducibility results. It also does not validate t
- 08-13 02:24attachThis is independent evidence that DeepSeek V4 Flash can be made usable on constrained consumer hardware, materially informing the case about practical low-memory serving.
- 08-13 02:23propose_attachThis is independent evidence that DeepSeek V4 Flash can be made usable on constrained consumer hardware, materially informing the case about practical low-memory serving.
- 08-11 22:51repriceRe-evaluation found no implementation details, benchmarks, independent tests, or discussion beyond the original title claim. The case remains an unverified duplicate of the broader NVMe/out-of-core in
- 08-11 22:51alert_silentThere is no new consequential delta to report; wait for a reproducible implementation or independent speed, memory, and quality results.
- 08-11 22:51alert_routeThere is no new consequential delta to report; wait for a reproducible implementation or independent speed, memory, and quality results.
- 08-11 22:48alert_silentThe only visible evidence is a low-engagement link title asserting 7.7 GiB NVMe-paged inference; no README, implementation details, benchmarks, or quality results are supplied. It duplicates an alread
- 08-11 22:48alert_routeThe only visible evidence is a low-engagement link title asserting 7.7 GiB NVMe-paged inference; no README, implementation details, benchmarks, or quality results are supplied. It duplicates an alread
- 08-11 22:37groundThe radar already tracks essentially the same unresolved NVMe/out-of-core inference proposition in `radar:hotpin-lossless-moe-streaming`, `radar:kimi-k3-nvme-expert-streaming`, and `radar:slipstream-s
- 08-11 22:34createThe released implementation makes a specific, testable claim about extending large-model inference to severely memory-constrained systems.