Fractal-BLT is a claimed open-source .NET 10 inference runtime for Mixture-of-Experts models, published on GitHub under the account h4zey86 and announced via a title-only Hacker News submission (1 point, zero comments, by Ctrl_Alt_Haze) that advertises zero-allocation streaming of model weights from NVMe directly to GPU — i.e., running MoE models whose total weights exceed GPU memory. The supplied web results contain only that HN listing and its title: no repo content, release artifact, benchmark, or third-party measurement of Fractal-BLT itself appears anywhere, so the claim remains exactly as advertised — unverified. The surrounding results do show the broader storage-tiered-inference class is now crowded and maturing: consumer expert-streaming runtimes report real numbers (LayerStoRm: 186 GiB MoE on 96 GB VRAM across four consumer GPUs at ~24.5 tok/s; a pure-C 744B model on 32 GB RAM), while Lightbits Labs demonstrates NVMe tiering entering production-grade serving (KV-page streaming over RDMA, ~1,154× faster time-to-first-token claims on L40S clusters), even as generic MoE deployment guidance still repeats the 'all experts must reside in VRAM' doctrine.
The new third-party NVFP4 reproduction independently arrives where Scott's canon already sits: its 'prefers the API' verdict is a dated receipt for his task-aware local-vs-API threshold, and the decode-feasible/prefill-impractical split extends hardware-aware local inference with a phase boundary — storage-tiered weights buy decode capacity, not serving viability. For gamepc this is a decision-relevant negative (streaming runtimes remain unattractive for prefill-heavy agentic work), while Fractal-BLT itself stays untouched by any inspectable evidence.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.task-aware-model-routingradar:slipstream-ssd-moe-streamingradar:hotpin-lossless-moe-streamingradar:kimi-k3-nvme-expert-streamingradar:picchio-llama-cpp-bottleneck-diagnostics
queries asked of Scott's wikis
- memory-constrained local serving gamepc workstation hardware-aware inference position
- zero-allocation runtime engineering .NET LLM inference implementation
- local inference vs API economics decode prefill throughput threshold
- MoE expert streaming NVMe offload running models beyond VRAM
- agentic workload prefill-heavy latency requirements local model serving
- GGUF llama.cpp runtime alternatives custom inference engine projects
2026-09-28T08:33:40Z
The velocity spike is upvote drift on the 3-week-old NVFP4 thread — a 52x multiple off a near-zero peer baseline, but an absolute 3.3 pts/h, cooling, against a 29.8 peak long past — plus two failed reobservations; no new fact since the prefill counter-data-point already priced, and still nothing touching Fractal-BLT itself. The magnitude-valve spread reading measures the SSD-streaming class's attention, which sibling cases (Kimi K3, Slipstream, HotPin) already price; heating this case would double-deliver a subject whose own objects have zero current attention (1 pt, 0 comments, no inspectable artifact). Case meaning unchanged: unverified single-publisher claim inside a multiply-demonstrated class with an open, gamepc-relevant prefill question.
2026-09-28T06:50:48Z
A first-hand counter-data-point on the NVFP4 thread (12 tok/s decode, 160 tok/s prefill on an RTX 3060 reading ~4.3 bpw q4km weights from NVMe) is >3x the prefill figure that anchored the class-level 'prefill-impractical' boundary, reframing NVMe-streamed MoE prefill as quant/config-dependent rather than uniformly impractical. Fractal-BLT itself remains unverified and untouched by any of this; the case's meaning shifts only in that the class boundary Scott's gamepc decision leans on is now an open question, not a settled negative.
2026-09-28T03:49:41Z
The velocity spike is engagement drift, not news: +4 points and one comment on the peripheral NVFP4 Reddit post (0.5→3.3 pts/h off a near-zero baseline) and one failed reobservation, with nothing new touching Fractal-BLT itself. The case's meaning is unchanged — an unverified single-claim lead inside a multiply-demonstrated but prefill-impractical approach class — so no state, relevance, or belief movement is warranted.
2026-09-27T22:48:29Z
grounded: converges/medium — The new third-party NVFP4 reproduction independently arrives where Scott's canon already sits: its 'prefers the API' verdict is a dated receipt for his task-awa
2026-09-27T22:41:56Z
A third-party reproduction of the independent NVFP4 SSD-streaming result confirms it runs on comparable hardware but reports prefill 'unbearably bad' and prefers an API, downgrading NVMe streaming from 'working path' to 'decode-feasible, prefill-impractical' — a caveat that also tempers Fractal-BLT's prospect. Fractal-BLT itself remains untouched by independent evidence; the seed→watching move reflects that the approach class it embodies now has multiple measured independent implementations, not any new Fractal-BLT fact.
2026-09-27T22:24:43Z
evidence attached: reddit.post.1wrxap8 — Independent SSD-streaming MoE result (177B NVFP4 at 9-10 tok/s on a 16GB GPU + 32GB RAM) corroborates NVMe-streaming as a working path for models exceeding GPU memory.
2026-09-18T08:22:50Z
Flyweight adds a separate RAM/CPU-offload candidate, not corroboration of Fractal-BLT’s NVMe-to-GPU streaming or zero-allocation claims. Its reported release and commenter performance figures do not change this case’s status as an unverified implementation lead.
2026-09-18T08:22:06Z
evidence attached: reddit.post.1wjjmj9 — A separate released runtime supports the same developing approach of running models larger than VRAM through CPU or system-memory offload.
2026-09-10T06:30:19Z
The refreshed sibling discussion adds no substantive technical evidence about Fractal-BLT or independently measured performance for the separate MacBook experiment. Fractal-BLT remains an unverified runtime lead; further sibling commentary should not trigger frequent reassessment absent inspectable implementation evidence or reproducible measurements.
2026-09-09T05:28:06Z
The refreshed sibling discussion adds no implementation evidence or measured result beyond previously assessed author testimony. Fractal-BLT remains an unverified local-inference lead; its next meaningful update requires directly inspectable runtime evidence or reproducible measurements, not further discussion of the separate MacBook experiment.
2026-09-09T04:28:36Z
The refreshed comments add no new technical finding beyond previously assessed testimony about a separate MacBook experiment. Fractal-BLT remains an unverified runtime lead, and further sibling discussion does not establish its implementation, zero-allocation behavior or practical performance.
2026-09-09T02:26:38Z
The refreshed sibling discussion adds no substantive evidence about Fractal-BLT; the separate MacBook author's storage-streaming testimony remains a broader evaluation lead, not validation of this runtime. Keep the case open at a slower cadence pending a directly inspectable implementation or reproducible measurements.
2026-09-09T01:22:52Z
The refreshed sibling discussion remains repetitive commentary around previously assessed author testimony, not new evidence about Fractal-BLT. Its claimed release, zero-allocation behavior and practical NVMe-to-GPU performance remain unverified; the separate MacBook experiment cannot establish them.
2026-09-09T00:26:17Z
The refreshed discussion supplies no new technical evidence beyond the separate MacBook experiment’s previously assessed author testimony. Fractal-BLT remains an unverified runtime lead; neither sibling attention nor storage-speed speculation establishes its release, zero-allocation behavior or practical performance.
2026-09-08T23:24:31Z
The refreshed discussion adds skepticism about usability and storage-speed speculation, not new implementation or performance evidence. The separate MacBook experiment remains a capacity-versus-speed evaluation lead; it does not validate Fractal-BLT’s release or zero-allocation NVMe-to-GPU claims.
2026-09-08T21:46:13Z
The separate MacBook submission now includes author testimony describing expert-file layout and uncached disk reads, making its storage-streaming claim more concrete but not independently validating the reported throughput. This strengthens the broader capacity-versus-speed evaluation lead, not Fractal-BLT’s release, zero-allocation claim or suitability for Scott’s workstation.
2026-09-08T20:31:36Z
The new MacBook/four-SSD submission supplies a separate claim about storage-streamed inference, not corroboration of Fractal-BLT’s release, zero-allocation design or performance. It strengthens the broader evaluation theme only; calling the title-only evidence a demonstrated implementation overstates what is established.
2026-09-08T20:23:18Z
evidence attached: hn.story.49616257 — This is an independent artifact demonstrating extreme SSD-streamed local inference beyond device-memory capacity.
2026-09-08T11:26:54Z
No substantive new evidence has arrived: the GitHub echo repeats the HN title rather than independently establishing a released implementation. Fractal-BLT remains a relevant validation lead for local MoE capacity, not yet evidence of usable NVMe-to-GPU inference or a performance advantage.
2026-09-08T11:25:23Z
grounded: converges/medium — Fractal-BLT’s claimed weight-streaming approach converges with Scott’s Hardware-aware local inference position and offers a concrete candidate to evaluate for e
2026-09-08T11:22:46Z
case created — A concrete runtime artifact addresses GPU-memory limits through a distinct storage-streaming approach, but the lone title-only observation supplies no performance or usability evidence.