2026-10-11 16:37 UTC

LocalLLaMA user EmPips reports the Strata runtime streams ~80GB Qwen3.8-Next (IQ3_XXS) from DDR4 to a 7900 XTX at 45-50 tok/s β€” roughly 2x tuned llama.cpp on the same rig β€” with similar results reported by users on 12-16GB cards, and independent replication would establish system-RAM weight streaming as a practical route to large-MoE local inference on modest GPUs.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highlocal-inference weight-streaming moe-serving
Surfaced 2026-10-04T13:50:37Z β€” Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next. β€” The first off-platform echo appeared β€” the ayobluestarr benchmark reshared to HN β€” but it is inert (1 point, 0 comments), so the case's settled meaning holds: approach-level RAM streaming validated by independent rigs, Strata engine claims and provenance still open. The magnitude-valve spread reading is a false positive: the top-decile engagement is the single-platform bot-fight thread, and the only cross-platform object is dead, so heat stays low despite eligibility.

What is this?

The Strata runtime (by developer Niko1221) streams Qwen3.8-Flash-Next (177B MoE, ~80GB IQ3_XXS) from system DDR4 to an AMD 7900 XTX, with the original poster reporting 45–70 tok/s. Independent replications across 12+ hardware tiers (6–96GB VRAM, DDR4/DDR5, Mac SSD, Apple Silicon via strata-mlx, Strix Halo APU) confirm the approach: system-RAM/SSD weight streaming makes 80GB-class MoEs runnable on modest GPUs. Strata's specific claim of '2x over tuned llama.cpp' remains unproven β€” the SOTAAZ benchmark shows Strata faster within 12 GiB VRAM but uses different expert-caching vs llama.cpp's --n-cpu-moe, and no same-quant tuned-vs-tuned 177B run exists. Heavy-quant (IQ3_XXS) agentic quality is degraded; a first independent coding test attributes correctness errors to Strata's engine (rapid PR churn, 'squeezing performance blindly') not just quantization. A suspected bot/astroturf campaign on Reddit taints Strata-specific testimony; the approach thesis stands on non-Strata lines (llama.cpp expert-streaming, Mac SSD streaming).

Why it matters to Scott

Independent non-Strata replications (llama.cpp expert-streaming on 12GB+DDR4, Mac Mini SSD-streaming, 3x3060 A/B) validate Scott's weight-streaming playbook and memory-bandwidth-as-decode-bottleneck position across DDR4/DDR5/SSD/APU/MLX tiers β€” exactly the cross-platform corroboration his framework predicts. The case adds actionable evidence: heavy-quant agentic degradation now attributed to Strata's engine (not just quant), and a concrete same-quant streaming-vs-offload test on his gamepc/Ollama stack becomes the logical adjudication. This isn't mere illustration; it moves the open question from 'does RAM streaming work' to 'does Strata's engine add value and at what quality cost'.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.weight-streamingradar:qwen38-flash-next-commodity-local-inferenceradar:adaptive-kv-cache-streamingradar:airllm-low-vram-model-streaming
queries asked of Scott's wikis
  • weight-streaming playbook DDR4 DDR5 SSD NVMe tiering
  • memory-bandwidth-as-decode-bottleneck MoE expert caching
  • local-inference economics consumer GPU system RAM offload
  • MoE serving 177B-class models modest VRAM
  • open-weights strategy model sovereignty local inference
  • llama.cpp expert streaming --n-cpu-moe vs Strata hot-expert cache

Measured heat

now 0 pts/hpeak 347 pts/hcomments 0/hpeers p25momentum: steady2 platformsage 205h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

10-03 03:12⭐ origin directly observedAnyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.
EmPips on r/LocalLLaMA
β€”
10-03 06:49first on r/LocalLLaMA Β· published Β· +3.6hQwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
ayobluestarr
β€”
10-04 06:22first on hacker news Β· published Β· +27.2hQwen3.8-Flash-Next 177B running at 11–15 tok/s on 1 RTX 5070 12GB and 32GB RAM
maziianobeatz58
β€”
10-03 03:12amplified on r/LocalLLaMAreddit.post.1wwcqas
EmPips
peak 89 Β· 118 comments Β· 6% of case engagement
10-03 06:49amplified on r/LocalLLaMAreddit.post.1wwggwt
ayobluestarr
peak 43 Β· 41 comments Β· 3% of case engagement
10-03 11:11amplified on r/LocalLLaMAreddit.post.1wwkn3x
soyalemujica
peak 0 Β· 47 comments Β· 1% of case engagement
10-03 14:05amplified on r/LocalLLaMAreddit.post.1wwo3ps
dampflokfreund
peak 4 Β· 42 comments Β· 1% of case engagement
10-03 14:15amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wwobfg
Mayion
peak 686 Β· 426 comments Β· 34% of case engagement
10-03 17:36amplified on r/LocalLLaMAreddit.post.1wwt19n
MD_Reptile
peak 26 Β· 13 comments Β· 1% of case engagement
25 more amplifiers in ainews.case_chain
10-03 03:20our radar first saw it Β· +0.1hdiscovery anchor: reddit.post.1wwcqasβ€”
10-04 07:46reached heat=high Β· +28.6h Β· via ledgerβ€”β€”
pace: p97 vs 1188 stories at the 168h mark (now 205h old) β€” ahead of openai-research-acceleration (1.0x), behind opus-55-behavior-shift (1.0x)

Evidence (31) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Anyone sitting on a lot of slow system memory and a modest GPU.. try Strata + Qwen3.8 Next.
LocalLLaMA
EmPips89118
🟠 redditQwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4
LocalLLaMA
ayobluestarr4341
🟠 redditThanks to Strata I have quit 27b for Qwen Flash (24gb VRAM plus 64gb ram)
LocalLLaMA
soyalemujica047
🟠 redditYes bots we get it, Strata is good now please stop
LocalLLaMA
Mayion684424
🟠 redditQwen 3.8 Flash Next q2_0 running on a 2060 laptop (32 GB RAM + 6 GB VRAM) using Strata!
LocalLLaMA
dampflokfreund042
🟠 redditFlash next rig born from mining parts.
LocalLLaMA
MD_Reptile2613
🟠 redditStrata Qwen 3.8 flash next is the biggest thing since the release of Qwen 3.8 27b
LocalLLaMA
inthesearchof2785
🟠 redditThe curse of 64GB system RAM
LocalLLaMA
Cautious_Chicken_604102258
🟧 hnQwen3.8-Flash-Next 177B running at 11–15 tok/s on 1 RTX 5070 12GB and 32GB RAMmaziianobeatz5810
🟠 reddit4090 48G +128G+strata test
LocalLLaMA
Shot-Ad-4147016
🟠 redditPSA: if you're on an Intel hybrid CPU, run Strata's calibrate - it nearly tripled my decode speed (IQ3_S at 256K, 16 GB card)
LocalLLaMA
MoonsvnLyn022
🟠 redditA Strata fork for IBM AC922 running Qwen3.8-FN UD-Q4_K_XL is doing up to 7,357 tk/s prefill and 113 tk/s decode
LocalLLaMA
okoyl32320
🟠 redditBlackwell 6000 96GB Optimization https://github.com/Niko1221/Strata/pull/813
LocalLLaMA
LegacyRemaster022
🟠 redditstrata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts
LocalLLaMA
Studio2714456
🟠 redditI got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming
LocalLLaMA
turtleninja9924
🟠 redditWhich model, which harness? I have data for you.
LocalLLaMA
dh7net6347
🟠 redditStrata is amazing and all but can we actually see what you’re building with it that you couldn’t do before
LocalLLaMA
vinigrae034
🟠 reddit~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.
LocalLLaMA
sdfprwggv54
🟠 redditNot another β€œs engine is Amazing” Post. Thermals Q
LocalLLaMA
nonproductive20
🟠 redditI turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card
LocalLLaMA
ExxploreCraft015
🟠 redditStrata - RTX 3060 Error
LocalLLaMA
Ordinary-Mango946210
🟠 redditQwen3.8-Flash-Next on Strata
LocalLLaMA
KnownAd483217385
🟠 redditQwen3.8 Flash Next on 5060 Ti 16GB - 55 tok/s average, and a few demos
LocalLLaMA
bobaburger06
🟠 redditCan I run Swift 1.5 flash next iq3xxs in 16GB VRAM and 32GB RAM?
LocalLLaMA
royalflash417017
🟠 redditSingle 3090 Qwen 27B user, considering buying 128GB of RAM because of the hype
LocalLLaMA
regunakyle39172
🟠 redditStrata 0.1.40.1 left ~10 GB of VRAM unused on my 4-GPU rig β€” I thought it was a bug. It isn't.
LocalLLaMA
Critical-Entry337709
🟠 redditTested in Coding: Strata
LocalLLaMA
PathfinderTactician68132
🟠 redditI was doing some testing on Strata. vs llama vs. runner and was missing something
LocalLLaMA
ZenZombie117019
🟠 redditSharing my latest project: strata-mlx
LocalLLaMA
yibie510
🟠 redditStrata with Qwen3.8 Flash Next UD-Q4_K_XL
LocalLLaMA
hyudryu2760
🟠 redditQuestion about flash next +strata+ 32gb sadness.
LocalLLaMA
ironicstatistic024

Interpretation history

Decision trace