2026-10-11 17:11 UTC

Independent benchmarks will determine whether dsv4-streaming can run the 284B DeepSeek V4 Flash checkpoint on 64GB Apple Silicon with near-lossless quality and practically useful expert-streaming throughput.

state: expiredheat: lowuncertainty: highknownscott: mediumlocal-inference moe expert-streamingDeepSeekkk-r

What is this?

DeepSeek V4 Flash is described as a 284B-parameter mixture-of-experts model with roughly 13B active parameters, aimed at coding and agentic workloads. The supplied snippets claim it can technically run on 64GB Apple Silicon by streaming experts from SSD, but cache misses sharply reduce generation speed; other material demonstrates or recommends 128GB configurations. The specific dsv4-streaming repository, kk-r’s role, and its claimed cache and near-lossless-quality results are not independently established by these snippets, so the case remains a benchmark claim awaiting verification.

Why it matters to Scott

The radar already tracks the same DeepSeek V4 Flash local-inference validation problem in `radar:deepseek-v4-flash-57gb-local-quant` and `radar:deepseek-v4-nvme-demand-paging`, alongside closely related expert-streaming cases. Verified throughput, cache behavior, and quality measurements would bear directly on Scott’s hardware-aware local-inference policy and evaluation discipline, but the presently unverified repository claim adds no established result yet.
dev:concept.hardware-aware-local-inferenceip:concept.evaluation-driven-developmentradar:deepseek-v4-flash-57gb-local-quantradar:deepseek-v4-nvme-demand-pagingradar:concept.expert-streaming
queries asked of Scott's wikis
  • MoE expert streaming and cache design
  • Apple Silicon local inference memory limits
  • SSD-backed inference throughput economics
  • quantization quality versus model fit
  • open-weight frontier models on consumer hardware
  • local coding-agent model deployment

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditDeepSeek-V4-Flash (284B) on a 64 GB MacBook: measurements, failures, and a cache result that surprised us
LocalLLaMA
cowboy-bebob20
🟧 echo.github ⭐The linked repository is the primary artifact: it contains the code, raw logs, measurement sweeps, failure analyses, and the cache finding. kk (GitHub account kk-r)——
🟠 redditI ran DeepSeek-V4-Flash (284B) on a 64 GB MacBook. notes and numbers
LocalLLaMA
cowboy-bebob13
🟠 redditHas anyone run Deepseek V4 Flash on two 128gb Macs using Exo?
LocalLLaMA
shveddy39
🟠 redditDeepSeek V4 Flash on an M2 Ultra: repacked to 141 GiB losslessly, smaller than the Q4 GGUF, at 25.8 t/s (42 t/s peak)
LocalLLaMA
Agusx12111310

Interpretation history

Decision trace