2026-10-11 17:15 UTC

Fractal-BLT’s publisher claims its released .NET 10 MoE runtime streams weights from NVMe to GPU with zero allocation, potentially enabling local inference on models whose weights exceed GPU memory.

state: watchingheat: lowuncertainty: highconvergesscott: mediumlocal-inference inference-economics moe-servingH4ZEY86

What is this?

Fractal-BLT is a claimed open-source .NET 10 inference runtime for Mixture-of-Experts models, published on GitHub under the account h4zey86 and announced via a title-only Hacker News submission (1 point, zero comments, by Ctrl_Alt_Haze) that advertises zero-allocation streaming of model weights from NVMe directly to GPU — i.e., running MoE models whose total weights exceed GPU memory. The supplied web results contain only that HN listing and its title: no repo content, release artifact, benchmark, or third-party measurement of Fractal-BLT itself appears anywhere, so the claim remains exactly as advertised — unverified. The surrounding results do show the broader storage-tiered-inference class is now crowded and maturing: consumer expert-streaming runtimes report real numbers (LayerStoRm: 186 GiB MoE on 96 GB VRAM across four consumer GPUs at ~24.5 tok/s; a pure-C 744B model on 32 GB RAM), while Lightbits Labs demonstrates NVMe tiering entering production-grade serving (KV-page streaming over RDMA, ~1,154× faster time-to-first-token claims on L40S clusters), even as generic MoE deployment guidance still repeats the 'all experts must reside in VRAM' doctrine.

Why it matters to Scott

The new third-party NVFP4 reproduction independently arrives where Scott's canon already sits: its 'prefers the API' verdict is a dated receipt for his task-aware local-vs-API threshold, and the decode-feasible/prefill-impractical split extends hardware-aware local inference with a phase boundary — storage-tiered weights buy decode capacity, not serving viability. For gamepc this is a decision-relevant negative (streaming runtimes remain unattractive for prefill-heavy agentic work), while Fractal-BLT itself stays untouched by any inspectable evidence.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.task-aware-model-routingradar:slipstream-ssd-moe-streamingradar:hotpin-lossless-moe-streamingradar:kimi-k3-nvme-expert-streamingradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
  • memory-constrained local serving gamepc workstation hardware-aware inference position
  • zero-allocation runtime engineering .NET LLM inference implementation
  • local inference vs API economics decode prefill throughput threshold
  • MoE expert streaming NVMe offload running models beyond VRAM
  • agentic workload prefill-heavy latency requirements local model serving
  • GGUF llama.cpp runtime alternatives custom inference engine projects

Measured heat

now 0 pts/hpeak 18 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 797h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-08 11:22 (minted)⭐ origin echo-reconstructedThe linked HN submission describes Fractal-BLT as a “Zero-allocation .NET 10 MoE runtime streaming NVMe to GPU.”
H4ZEY86 on github (echo) · attributed from hn.story.49608430 · published time unknown
—
09-08 10:39first on hacker news · published · lag ?Fractal-BLT – Zero-allocation .NET 10 MoE runtime streaming NVMe to GPU
Ctrl_Alt_Haze
—
09-18 07:39first on r/LocalLLaMA · published · lag ?Flyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors.
Main-Wolverine-1042
—
09-08 10:39amplified on hacker newshn.story.49608430
Ctrl_Alt_Haze
peak 1 · 0 comments · 0% of case engagement
09-08 20:07amplified on hacker news 👑hn.story.49616257
Argonautlabs
peak 279 · 156 comments · 80% of case engagement
09-18 07:39amplified on r/LocalLLaMAreddit.post.1wjjmj9
Main-Wolverine-1042
peak 37 · 36 comments · 7% of case engagement
09-27 22:13amplified on r/LocalLLaMAreddit.post.1wrxap8
TypicalPudding6190
peak 68 · 51 comments · 12% of case engagement
09-08 11:21our radar first saw it · lag ?discovery anchor: hn.story.49608430—
pace: p85 vs 519 stories at the 720h mark (now 797h old) — ahead of deepseek-v4-flash-vision-release (1.0x), behind rp2350-local-image-generation (1.0x)

Evidence (5) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnFractal-BLT – Zero-allocation .NET 10 MoE runtime streaming NVMe to GPUCtrl_Alt_Haze10
🟧 echo.github ⭐The linked HN submission describes Fractal-BLT as a “Zero-allocation .NET 10 MoE runtime streaming NVMe to GPU.”H4ZEY86——
🟧 hnKimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDsArgonautlabs279156
🟠 redditFlyweight: open-source C++/CUDA engine for running MoE models bigger than your VRAM on one GPU + system RAM. First PyPI release, looking for contributors.
LocalLLaMA
Main-Wolverine-10423736
🟠 redditQwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM
LocalLLaMA
TypicalPudding61906851

Interpretation history

Decision trace