The BC-250 is an ASRock crypto-mining board built around AMD's Cyan Skillfish APU β Zen 2 CPU plus a GFX1013 GPU from PlayStation 5-lineage silicon, with 16GB of GDDR6 β that flooded the second-hand market (~$50β170/board) after mining collapsed; AMD never sold it as a consumer part. An independent community ecosystem documents single-board LLM inference on it: Vulkan is the only working compute path (ROCm unsupported, Linux mandatory), requiring community kernel/firmware patches to raise memory limits and re-enable 16 of 40 fused-off compute units, with 35B-class MoEs at ~37β75 tok/s on one board reported in the akandr/bc250 and H6-Technologies guides. The case's builder (Ok-Breadfruit-3523) extends this to a 6-board cluster (~$800 of boards in an ASRock 4U12G chassis) serving Qwen MoE quants at ~28β60 tok/s at 100k context over llama.cpp Vulkan+RPC, partitioning rather than pooling the boards' VRAM. The supplied web material corroborates only single-board results β every multi-board throughput figure is the builder's own, the rig's power draw is admitted to be poor but unmeasured, and current board prices have risen enough that the ~$800 total is plausible but unverified.
The 6-board follow-up independently implements the exact policy dev:concept.hardware-aware-local-inference prescribes β partitioned-not-pooled VRAM, per-board quant choice, Vulkan-because-ROCm-is-unsupported on off-label AMD silicon β on the llama.cpp Vulkan+RPC stack that would fill gamepc's serving role, and it extends the ex-mining VRAM-per-dollar lineage already on the radar (CMP 170HX unlock, DumpsterCluster retired-GPU serving). It stays at medium rather than high because every multi-board figure remains builder-only, and the unmeasured draw on a rig its own builder calls 'wildly inefficient' is precisely the ip:concept.ai-unit-economics question that decides whether ~$800-for-~70GB salvage actually beats his existing single-CUDA-box substrate.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:concept.ai-unit-economicsradar:nvidia-cmp-vram-unlockradar:cmp-170hx-unlock-verificationradar:dumpstercluster-low-cost-70b-servingradar:dumpstercluster-retired-gpu-inferenceradar:llama-cpp-rpc-parallel-loadingradar:lemonade-vulkan-rocm-dropradar:concept.amd-gpuradar:concept.vulkan
queries asked of Scott's wikis
- VRAM-per-dollar thresholds for local model serving hardware
- llama.cpp RPC distributed multi-node inference setups
- Vulkan vs ROCm gaps on AMD consumer and off-label silicon
- ex-mining and salvage hardware reuse patterns
- local coding model backend for agent harness work
- power draw and electricity cost tradeoffs in home inference rigs
now 2 pts/hpeak 78 pts/hcomments 0/hpeers p76momentum: cooling1 platformsage 324h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
2026-10-11T05:43:50Z
Third builder post (same author) shows significant throughput gains β 4 BC-250 boards on 2.5GbE now hit ~70 tok/s (Next Flash IQ3_XXS, 100k) and ~145 tok/s (2Γ Qwen 3.6 35B IQ4, 256k) for ~$500 β but remains single-source; CMP 170HX track has genuine cross-builder corroboration (two independent builders) for a different salvage hardware family. BC-250 still lacks independent replication, power measurements, and verified board pricing.
2026-10-11T05:29:22Z
evidence attached: reddit.post.1x2yv1b β Independent builder replication of ex-mining BC-250 cluster achieving better-than-claimed throughput with Claude Code optimization.
2026-10-07T23:30:14Z
The new evidence is a second salvaged-mining-hardware build β 2ΓCMP 170HX (128GB HBM) serving GLM-5.3-Flash at 384K/~90 tok/s, cross-corroborated in comments by a second CMP owner running 4 cards at 524K on a vLLM backport β which genuinely supports the cheap-salvage-large-MoE tier but through a hardware family already tracked by the CMP-unlock radar refs, not a replication of the BC-250 cluster. The case's own subject has gone quiet (no third BC-250 post, no power data, no independent multi-board verification since ~Oct 3), so it stays watching at low heat with expiry and supersession-by-CMP as the live failure modes.
2026-10-07T23:27:06Z
evidence attached: reddit.post.1x0b1ws β A dual 64GB CMP 170HX ex-mining-card build serving GLM-5.3-Flash at 384K context is a second salvaged-mining-hardware data point for the cheap large-MoE local-inference tier the case tracks.
2026-10-03T16:52:14Z
Another round of velocity spikes on the 6-board follow-up are threshold artifacts (3.67 pts/h off sub-0.5 baselines on a 70-pt post whose score actually slipped 71β70); the only new content is a comment trickle, including a skeptic claiming a single RTX 3060 + 64GB box matches the Qwen Flash throughput β commentary that sharpens the unit-economics question but is not evidence. No replication, no power data, no third build: the case keeps its meaning but is one quiet cycle from expiry-candidate unless the promised 7-board / 3.8-flash-next test appears.
2026-10-02T03:37:46Z
grounded: converges/medium β The 6-board follow-up independently implements the exact policy dev:concept.hardware-aware-local-inference prescribes β partitioned-not-pooled VRAM, per-board q
2026-10-02T03:27:53Z
The promised follow-up materialized: the same builder now runs a 6-board cluster with a partitioned preferred config (4Γ Qwen-Next-Flash IQ2_XS ~28 tok/s @100k/~115 ppt; 2Γ 35B q4 60 tok/s/~450 ppt, llama.cpp Vulkan+RPC) β a consistent, more detailed progression that pulls the case back from expiry-candidate into an active build program, though every multi-board number is still self-reported (no power draw, no replication, and the original 40 tok/s Qwen3-Coder-Next Q4 headline is not restated). Attention is negligible: the follow-up has 15 pts/4 comments and its 83rd-percentile peer rate reflects a tiny young post, not spread, so heat stays low.
2026-10-02T03:26:03Z
evidence attached: reddit.post.1wviphh β Same builder's follow-up with concrete updated numbers (Qwen at 100k context, 28-60 tok/s over 6 boards) β direct progression of the BC-250 rig hypothesis.
2026-09-30T10:52:12Z
Second post-peak velocity_spike is another threshold artifact: the score trickled ~135β154 on upvote residue while measured rate sits at 0.0 pts/h with comments frozen at 62 β not renewed spread. ~55h in there is still no builder follow-up, no replication, no second source; the case's meaning is unchanged but it has shifted from 'quiet watch' toward 'expiry candidate' if the promised 7-board and '3.8 flash' follow-ups don't surface within a few days.
2026-09-29T09:54:40Z
The attention episode crested and receded with zero evidentiary movement: peak ~22 pts/h has decayed to ~1 pt/h, comments are flat at 62, the post remains Reddit-only, and the newest top comments are the same appreciation plus unanswered prefill/power/model questions. The velocity-spike triggers were threshold-crossing artifacts (score crossing p90=135 at its peak), not renewed spread. The case reverts from 'hot unverified claim' to a quiet watch on the builder's promised 7-board and '3.8 flash' follow-ups.
2026-09-28T10:48:54Z
The velocity spike is real β the post is running at top-decile rate for its Reddit cohort (93.6 pct) and still accelerating β but the discussion so far is appreciation plus open technical questions (prefill over 1GbE, power draw, model choice), with no replication or second source. The claim's evidentiary standing is unchanged; its attention price rises to medium.
2026-09-28T04:32:14Z
grounded: novel/medium β Lands directly on Scott's hardware-aware local-inference practice and the gamepc/ollama serving stack: a validated ~$800-for-~71GB salvage path to serving large
2026-09-28T04:23:49Z
case created β A concrete first-party build report with measured throughput and expansion plans on a distinctive ex-mining hardware path, resolvable by the builder's promised follow-ups or independent replication of BC-250 rigs.