2026-10-11 16:37 UTC

LocalLLaMA builder Ok-Breadfruit-3523 claims a ~$800 rig of five ex-mining BC-250 boards exposes ~71GB VRAM and serves Qwen3-Coder-Next Q4 at 40 tok/s (30k context, ~30 at 100k) over 1GbE with headless worker boards β€” replication or wider adoption of ex-mining BC-250 rigs would establish salvaged mining hardware as a cheap ~70GB route to local large-MoE coding inference despite its power inefficiency.

state: watchingheat: lowuncertainty: mediumconvergesscott: mediumlocal-inference mining-hardware-reuse moe-modelsOk-Breadfruit-3523

What is this?

The BC-250 is an ASRock crypto-mining board built around AMD's Cyan Skillfish APU β€” Zen 2 CPU plus a GFX1013 GPU from PlayStation 5-lineage silicon, with 16GB of GDDR6 β€” that flooded the second-hand market (~$50–170/board) after mining collapsed; AMD never sold it as a consumer part. An independent community ecosystem documents single-board LLM inference on it: Vulkan is the only working compute path (ROCm unsupported, Linux mandatory), requiring community kernel/firmware patches to raise memory limits and re-enable 16 of 40 fused-off compute units, with 35B-class MoEs at ~37–75 tok/s on one board reported in the akandr/bc250 and H6-Technologies guides. The case's builder (Ok-Breadfruit-3523) extends this to a 6-board cluster (~$800 of boards in an ASRock 4U12G chassis) serving Qwen MoE quants at ~28–60 tok/s at 100k context over llama.cpp Vulkan+RPC, partitioning rather than pooling the boards' VRAM. The supplied web material corroborates only single-board results β€” every multi-board throughput figure is the builder's own, the rig's power draw is admitted to be poor but unmeasured, and current board prices have risen enough that the ~$800 total is plausible but unverified.

Why it matters to Scott

The 6-board follow-up independently implements the exact policy dev:concept.hardware-aware-local-inference prescribes β€” partitioned-not-pooled VRAM, per-board quant choice, Vulkan-because-ROCm-is-unsupported on off-label AMD silicon β€” on the llama.cpp Vulkan+RPC stack that would fill gamepc's serving role, and it extends the ex-mining VRAM-per-dollar lineage already on the radar (CMP 170HX unlock, DumpsterCluster retired-GPU serving). It stays at medium rather than high because every multi-board figure remains builder-only, and the unmeasured draw on a rig its own builder calls 'wildly inefficient' is precisely the ip:concept.ai-unit-economics question that decides whether ~$800-for-~70GB salvage actually beats his existing single-CUDA-box substrate.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsradar:nvidia-cmp-vram-unlockradar:cmp-170hx-unlock-verificationradar:dumpstercluster-low-cost-70b-servingradar:dumpstercluster-retired-gpu-inferenceradar:llama-cpp-rpc-parallel-loadingradar:lemonade-vulkan-rocm-dropradar:concept.amd-gpuradar:concept.vulkan
queries asked of Scott's wikis
  • VRAM-per-dollar thresholds for local model serving hardware
  • llama.cpp RPC distributed multi-node inference setups
  • Vulkan vs ROCm gaps on AMD consumer and off-label silicon
  • ex-mining and salvage hardware reuse patterns
  • local coding model backend for agent harness work
  • power draw and electricity cost tradeoffs in home inference rigs

Measured heat

now 2 pts/hpeak 78 pts/hcomments 0/hpeers p76momentum: cooling1 platformsage 324h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-28 03:52⭐ origin directly observedI’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s
Ok-Breadfruit-3523 on r/LocalLLaMA
β€”
10-02 02:49first on r/LocalLLaMA Β· published Β· +95.0hUpdate on the β€œMonstrosity”. 6 BC-250 board cluster
Ok-Breadfruit-3523
β€”
09-28 03:52amplified on r/LocalLLaMA πŸ‘‘reddit.post.1ws49si
Ok-Breadfruit-3523
peak 170 Β· 66 comments Β· 51% of case engagement
10-02 02:49amplified on r/LocalLLaMAreddit.post.1wviphh
Ok-Breadfruit-3523
peak 76 Β· 41 comments Β· 25% of case engagement
10-07 22:59amplified on r/LocalLLaMAreddit.post.1x0b1ws
Prudent_Appearance71
peak 16 Β· 49 comments Β· 14% of case engagement
10-11 04:27amplified on r/LocalLLaMAreddit.post.1x2yv1b
Ok-Breadfruit-3523
peak 37 Β· 12 comments Β· 10% of case engagement
09-28 04:20our radar first saw it Β· +0.5hdiscovery anchor: reddit.post.1ws49siβ€”
pace: p80 vs 1188 stories at the 168h mark (now 324h old) β€” ahead of safa-frontier-safety-authority (1.0x), behind google-weathernext-3-release (1.0x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-09-28T04:23:01.267462+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
Ok-Breadfruit-352317066
🟠 redditUpdate on the β€œMonstrosity”. 6 BC-250 board cluster
LocalLLaMA
Ok-Breadfruit-35237641
🟠 reddit2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison
LocalLLaMA
Prudent_Appearance711549
🟠 redditRunning Next Flash IQ3_XXS at ~70 tok/s with 100k context or 2 instances of Qwen 3.6 35B A3B IQ4 at ~145 tok/s with 256k all on $500 of ex mining BC-250 boards
LocalLLaMA
Ok-Breadfruit-35233712

Interpretation history

Decision trace