2026-10-11 16:37 UTC

KnownAd4832 claims a purpose-built single-model inference engine sustains ~65 tok/s decode of Qwen3.8-Flash-Next at 128K context on a 12GB RTX 5070 (~430 tok/s prompt processing, versus ~15 tok/s on llama.cpp), and replication would establish custom model-specific engines as a practical path for low-VRAM long-context local inference.

state: significantheat: lowuncertainty: mediumconvergesscott: highlocal-inference inference-engines consumer-gpu-optimization long-context
Surfaced 2026-09-25T23:56:26Z β€” The repo README is the primary artifact behind the post's claim. It announces "Strata: an inference engine built for exactly one model and e β€” The claim graduated from a lone first-party post to a front-page r/LocalLLaMA thread (228 pts/165 comments) behind a named one-click artifact (Strata), but verification has not moved: the sharpest comments are control demands (same-quants llama.cpp baseline, logit equivalence) and there is no independent replication. New independent datapoint β€” ExLlamaV3 on a 5070 Ti with 96GB does ~21.5 tok/s at ~200K β€” makes Strata's 65 tok/s a claimed ~3x over the best documented alternative stack if it holds, not just 4x over llama.cpp.

What is this?

Strata is a hobbyist inference engine purpose-built for exactly one model β€” Alibaba's Qwen3.8-Flash-Next, a sparse mixture-of-experts LLM whose architecture (36 of 48 layers are fixed-state Gated DeltaNet, and a 51B N-gram embedding table can live in host memory per the official vLLM/SGLang docs) means only a small slice of the model is touched per token. Its author, Reddit user KnownAd4832, exploits that property by keeping always-used weights on the GPU and paging the rest from system RAM, claiming ~65 tok/s decode at 128K context on a 12GB RTX 5070 versus ~15 tok/s for llama.cpp; per the case's own evidence trail, community members on similar 12GB-class rigs have since replicated 32-60 tok/s, but the engine reportedly forces greedy decoding and its output fidelity and low-quant quality remain untested. The web snippets independently corroborate the surrounding ecosystem rather than the headline number: upstream stacks (vLLM merged, SGLang day-0, DGX Spark and Jetson forum threads) are racing on the same 'huge MoE on small memory' problem, and the same recipe is now being marketed directly ('GPU poor rejoice': one 24-32GB GPU + 64GB RAM + NVMe at ~64 tok/s decode). The release has also seeded a self-compounding genre of single-model engines β€” Ninfer, MoEspresso, Gem16, Halogen, Slipstream, Inco AI's Splash, TensorFold β€” while llama.cpp lands official multi-token-prediction support that could close the gap in exactly this regime (with one reported negative datapoint under expert offload).

Why it matters to Scott

Practitioners and ggml-org are independently compiling Scott's hardware-aware-local-inference principle (placement, precision, memory pressure as explicit runtime policy) into shipped single-model engines, with throughput now replicated in class β€” and the case's two open gates (temp-0 token identity vs llama.cpp; 128K usability at IQ3_XXS) are literally his deterministic-verification and dumb-zone tests, cheap to run on his own 12GB gamepc/Ollama rig. The negative offload-MTP datapoint plus a self-compounding genre of Codex-written one-model engines (mission-shaped software at ecosystem scale) make his test-Strata-now-vs-wait decision live rather than academic, and hand him a dated-receipts publishing angle while the community's decisive fidelity check remains unanswered.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaip:concept.deterministic-verification-before-assertionip:concept.dumb-zoneip:concept.mission-shaped-softwareradar:qwen38-flash-next-commodity-local-inferenceradar:concept.local-inferenceradar:concept.moe-offloadingradar:concept.inference-runtimesradar:concept.long-contextradar:magnitude-self-optimizing-inference-engineradar:slipstream-ssd-moe-streaming
queries asked of Scott's wikis
  • hardware-aware local inference: placement, precision, memory pressure as runtime policy
  • MoE expert offload / expert paging to CPU RAM on small-VRAM consumer GPUs
  • deterministic output verification: temp-0 token identity vs reference engine
  • long-context quality gates at aggressive quantization (128K, IQ3-class)
  • gamepc 12GB local runtime choice (Ollama) β€” test-now vs wait decision
  • coding agents producing bespoke real software (Codex-written hobbyist engines)

Measured heat

now 0 pts/hpeak 341 pts/hcomments 0/hpeers p25momentum: steady3 platformsage 407h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-24 16:40⭐ origin echo-reconstructedThe repo README is the primary artifact behind the post's claim. It announces "Strata: an inference engine built for exactly one model and e
Niko1221 on github (echo) Β· attributed from reddit.post.1wp7zyb
β€”
09-24 17:30first on r/LocalLLaMA Β· published Β· +0.8hQwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
KnownAd4832
β€”
09-28 15:07first on hacker news Β· published Β· +94.5hStrata: Qwen3.8-Flash-Next (125B Moe) on a 8GB+ Nvidia GPU
simonpure
β€”
09-24 17:30amplified on r/LocalLLaMAreddit.post.1wp7zyb
KnownAd4832
peak 350 Β· 335 comments Β· 15% of case engagement
09-27 17:53amplified on r/LocalLLaMAreddit.post.1wrqql8
marcobaldo
peak 10 Β· 23 comments Β· 1% of case engagement
09-27 22:02amplified on r/LocalLLaMAreddit.post.1wrx15j
Danmoreng
peak 16 Β· 7 comments Β· 0% of case engagement
09-28 15:07amplified on hacker newshn.story.49879253
simonpure
peak 5 Β· 0 comments Β· 0% of case engagement
09-29 01:52amplified on r/LocalLLaMAreddit.post.1wsxgbo
pubudeux
peak 84 Β· 65 comments Β· 3% of case engagement
09-30 04:01amplified on r/LocalLLaMAreddit.post.1wtv43r
MLDataScientist
peak 156 Β· 139 comments Β· 6% of case engagement
16 more amplifiers in ainews.case_chain
09-24 18:20our radar first saw it Β· +1.7hdiscovery anchor: reddit.post.1wp7zybβ€”
09-25 23:54reached heat=high Β· +31.2h Β· via ledgerβ€”β€”
pace: p96 vs 1032 stories at the 336h mark (now 407h old) β€” ahead of alibaba-qwen4-announcement (1.0x), behind meta-muse-spark-13-release (1.0x)

Evidence (23) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second
LocalLLaMA
KnownAd4832348335
🟧 echo.github ⭐The repo README is the primary artifact behind the post's claim. It announces "Strata: an inference engine built for exactly one model and eNiko1221β€”β€”
🟠 redditQwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max
LocalLLaMA
marcobaldo823
🟠 redditGem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs
LocalLLaMA
Danmoreng167
🟧 hnStrata: Qwen3.8-Flash-Next (125B Moe) on a 8GB+ Nvidia GPUsimonpure50
🟠 redditFirst few days of qwen3.8-flash-next on 4x R9700 - it's been really interesting so far
LocalLLaMA
pubudeux8465
🟠 redditQwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine
LocalLLaMA
MLDataScientist153139
🟧 hnHalogen – Fastest Way to Run Qwen3.8-Flash-Next on AMD Strix Halordslw30
🟧 hnStrata: Run a 125B-parameter AI model on a normal gaming PCmaille20
🟠 redditQwen Flash Next MTP work restarted
LocalLLaMA
jacek20236324
🟠 redditRunning 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant
LocalLLaMA
SnooPredictions5153324
🟠 redditNinfer Stahp
LocalLLaMA
swagonflyyyy10
🟠 redditStrata on a power limited 5090 and 96GB of DDR5-6400 is cranking out 150-200 tok/s decode and 5-6k prefill! Qwen3.8-Flash-Next at IQ3_S, CTX at 128k tokens (8-bit).
LocalLLaMA
z0_o664131
🟠 reddit2.3x faster Qwen3.8 27B on a 5090: ninfer vs llama.cpp, 4 setups, same prompt - speed and quality tested
LocalLLaMA
theexile133706
🟠 redditI built Ninfer 4080 for 16GB class GPUs
LocalLLaMA
roofkid6155
🟠 redditRunning Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD
LocalLLaMA
fuzhongkai7364
🟧 hnRun Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/ssnehesht934419
🟠 redditQwen3.8 27B | 1 x R9700: 262K context, half a million tokens of reusable cache, ~180 tok/s. And yes, let's talk about the "3-bit" :)
LocalLLaMA
evp-cloud025
🟠 redditPractical limit hit. Decoding so fast that tool calls (cpu) starting to become real limit not decode or prefill. Single RTX5090. Porting Kenshi to Godot project.
LocalLLaMA
BringTea_666013
🟠 redditNInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s
LocalLLaMA
lkarlslund2218
🟠 redditFollow up: Qwen 3.8 27B at ~96t/s decode with NInfer on a 16GB RTX 5080, 110k context
LocalLLaMA
Kernoriordan23
🟠 redditBasalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)
LocalLLaMA
jesdga9566113
🟠 redditQwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix
LocalLLaMA
zipzak1361

Interpretation history

Decision trace