2026-10-11 17:13 UTC

LocalLLaMA builder fechyyy claims a 21M model with a 6.4B-parameter product-key memory table (16.8M rows) memory-mapped from SSD matches a 114M dense model while running ~140 tok/s from NVMe on 0.4GB VRAM (RX 9070) โ€” replication would establish SSD-resident sparse lookup memory as a practical route to dense-class capacity on small local models.

state: seedheat: lowuncertainty: highconvergesscott: mediumproduct-key-memory local-inference ssd-resident-weightsfechyyy

What is this?

A builder posting as fechyyy on r/LocalLLaMA reports a self-benchmarked experiment: a 21M-parameter model augmented with a 6.4B-parameter product-key memory โ€” a sparse lookup table of 16.8M rows โ€” that is memory-mapped from SSD rather than resident in VRAM. Per the post, the augmented model matches a 114M-parameter dense model while running at ~140 tok/s with the table streamed from NVMe, using only 0.4GB of VRAM on an AMD RX 9070. The supplied web search returned nothing usable (only a Reddit human-verification page), so the claim cannot be independently corroborated here and rests on the author's self-reported numbers and setup; replication by others is the open question.

Why it matters to Scott

Converges with the external-memory-over-parameter-count thesis Scott argues in Walk a Wiki ('retrieval shape, not merely model quality') and The Model Is Not the Memory โ€” a hobbyist existence proof that a 21M model plus a 6.4B SSD-resident sparse lookup table matches a 5x denser model โ€” but via an opaque learned product-key table, not legible claims+edges, which extends rather than repeats the three-substrates taxonomy and lands squarely in the weight-streaming lineage the radar already tracks (Percepta's indexed memory, the PLE hot-swap table). Self-reported, toy-scale and uncorroborated, so it stays at medium until replication; a confirmed result would be a dated receipt for the thesis and a live SSD-vs-VRAM placement technique for his hardware-aware local-inference stack on gamepc.
ip:source.walk-a-wiki-cant-drive-a-rag-ebookip:source.the-model-is-not-the-memory-ebookip:concept.three-substratesdev:concept.hardware-aware-local-inferenceradar:concept.weight-streamingradar:concept.local-inferenceradar:concept.small-modelsradar:percepta-spotlight-architectureradar:llama-cpp-hot-swappable-ple-memory
queries asked of Scott's wikis
  • agent memory: learned/parametric weights vs external retrieval
  • RAG and knowledge systems: lookup table as a memory substrate
  • local inference: VRAM offloading, NVMe/SSD weight streaming
  • small local models matching larger dense models via external memory
  • mmap / memory-mapped weights in inference harness projects
  • sparse key-value knowledge store experiments and benchmarks

Measured heat

now 0 pts/hpeak 17 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 119h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-06 16:57โญ origin directly observedI gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)
fechyyy on r/LocalLLaMA
โ€”
10-06 16:57amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wz7tvs
fechyyy
peak 275 ยท 44 comments ยท 100% of case engagement
10-06 19:20our radar first saw it ยท +2.4hdiscovery anchor: reddit.post.1wz7tvsโ€”
pace: p80 vs 1247 stories at the 96h mark (now 119h old) โ€” ahead of inco-splash-apple-silicon-inference (1.0x), behind forgejo-1604-critical-rce (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญI gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)
LocalLLaMA
fechyyy27544

Interpretation history

Decision trace