2026-10-11 17:10 UTC

System evaluation will determine whether the High Bandwidth Flash design presented at Hot Chips 2026 can expand AI memory capacity at useful bandwidth and materially lower cost than HBM-only configurations.

state: corroboratedheat: lowuncertainty: highconvergesscott: mediumhigh-bandwidth-flash ai-infrastructure inference-economics

What is this?

High Bandwidth Flash (HBF) is a proposed AI memory tier that stacks NAND flash in HBM-style packages to put terabyte-scale capacity beside accelerators — a SanDisk concept now being standardized with SK hynix under an Open Compute Project workstream (first open spec announced at FMS 2026), with first samples targeted for H2 2026, a pilot production line expected by year-end, and commercialization aimed at 2027. The material new development in the supplied coverage is a system-level workload analysis presented at Hot Chips 2026 by GPU IP firm OXMIQ Labs (presented with PRAXMATI per Jason's Chips; Chips and Cheese names the speakers as Anurag Agarwal and Radhakrishna Giduthuri): modeling a 72-GPU rack running Kimi-K2 at FP4, an HBF-only configuration delivers ~14x HBM capacity (294.9 TB vs 20.7 TB) but cuts aggregate bandwidth (922 vs 1,584 TB/s), so HBF wins on cost-per-token only in low-batch, capacity-bound serving — positioned as a complement to HBM, not a replacement. All such figures are models, not hardware measurements: the coverage consistently lists sustained bandwidth, latency, endurance, realized pricing, and software integration (HBF is DMA-accessed in large aligned chunks, closer to an on-package SSD than a memory pool) as unresolved, and competing approaches (Qualcomm's High Bandwidth Compute, Kioxia's PCIe 6.0 GPU-direct SSDs, future zHBM) will contest the same niche.

Why it matters to Scott

OXMIQ's Hot Chips system model independently arrives where Scott's hardware-aware-local-inference practice already starts — memory-tier placement is a regime-dependent runtime decision, not a fixed config — and its result that HBF pays on cost-per-token only in low-batch, capacity-bound serving describes exactly the regime Scott himself runs (Ollama on gamepc, wait-not-downgrade job queue, VRAM-capped open-weight serving). Still medium rather than high: every figure is simulation with no hardware, pricing, or endurance data, so nothing yet changes what he builds, but it sharpens his cost-per-token unit-economics lens and connects directly to the 2027 memory-capacity commitment lineage.
dev:concept.hardware-aware-local-inferenceip:concept.ai-unit-economicsdev:technology.ollamaradar:2027-memory-capacity-selloutradar:adaptive-kv-cache-streaming
queries asked of Scott's wikis
  • KV cache versus weights tier placement in hardware-aware inference
  • inference unit economics: cost per token and memory bandwidth ceilings
  • local serving of large open-weight models within VRAM capacity limits
  • flash or SSD tiers in RAG and knowledge-system storage design
  • non-volatile memory as substrate for persistent agent memory
  • batch size versus throughput trade-offs in model serving notes

Measured heat

now 0 pts/hpeak 3 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1153h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

08-24 19:39 (minted)⭐ origin echo-reconstructedThe article reports on applying High Bandwidth Flash as presented at Hot Chips 2026.
Chips and Cheese on blog (echo) · attributed from hn.story.49420592 · published time unknown
—
08-24 14:48first on hacker news · published · lag ?Hot Chips 2026: Applying High Bandwidth Flash (HBF)
ksec
—
08-28 10:19first on r/LocalLLaMA · published · lag ?Micron: HBM Requires Three Times More Wafer Area Than DDR5
FullstackSensei
—
08-24 14:48amplified on hacker newshn.story.49420592
ksec
peak 51 · 17 comments · 10% of case engagement
08-27 12:27amplified on hacker newshn.story.49463693
metrofun
peak 2 · 0 comments · 0% of case engagement
08-28 10:19amplified on r/LocalLLaMA 👑reddit.post.1w0mmk7
FullstackSensei
peak 366 · 107 comments · 40% of case engagement
08-30 12:45amplified on r/LocalLLaMAreddit.post.1w2gmn5
9r4n4y
peak 18 · 5 comments · 2% of case engagement
09-03 17:35amplified on r/LocalLLaMAreddit.post.1w6e6u2
giveen
peak 84 · 12 comments · 8% of case engagement
09-14 02:56amplified on r/LocalLLaMAreddit.post.1wfrk63
Summit-Star001
peak 79 · 10 comments · 7% of case engagement
3 more amplifiers in ainews.case_chain
08-24 19:21our radar first saw it · lag ?discovery anchor: hn.story.49420592—

Evidence (10) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnHot Chips 2026: Applying High Bandwidth Flash (HBF)ksec5117
🟧 echo.blog ⭐The article reports on applying High Bandwidth Flash as presented at Hot Chips 2026.Chips and Cheese——
🟧 hnFlint: Efficiently Leveraging High Bandwidth Flash for LLM Inferencemetrofun20
🟠 redditMicron: HBM Requires Three Times More Wafer Area Than DDR5
LocalLLaMA
FullstackSensei358104
🟠 reddit{INTRESTING PAPER BASED ON HBF}2607.10186] FlashAccel: Leveraging High-Bandwidth Flash (HBF) for High-Throughput LLM Inference
LocalLLaMA
9r4n4y185
🟠 redditMicron Explores Near-GPU NAND Flash to Run Bigger LLMs
LocalLLaMA
giveen8412
🟠 redditMicron's memory wall chart. Compute up ~3x every two years, HBM bandwidth under 2x
LocalLLaMA
Summit-Star0017910
🟠 redditAsianometry: High Bandwidth Flash: What Is It Good For?
LocalLLaMA
asssuber6134
🟠 redditFormer Intel CEO: "HBM is lousy". High Bandwidth Flash Is Coming
LocalLLaMA
Glittering_Depth_72218283
🟧 hnHBM: High-Bandwidth Mistakegok110

Interpretation history

Decision trace