A LocalLLaMA builder demonstrates a $2,800 rig using eight ex-cloud-gaming Radeon Pro V620 GPUs (256GB VRAM) with a custom vLLM fork to serve Qwen3.8-Flash-Next at 60โ100 tok/s decode, establishing a distinct cheap high-VRAM path from salvaged enterprise hardware.
state: seedheat: mediumuncertainty: mediumconvergesscott: mediumlocal-inference-economics salvaged-gpu-hardware vllm-forks_TheWolfOfWalmart_
What is this?
Reddit user _TheWolfOfWalmart_ on r/LocalLLaMA built a $2,800 inference server using eight decommissioned AMD Radeon Pro V620 GPUs (32GB each, 256GB total VRAM) sourced from cloud-gaming infrastructure, paired with a custom vLLM fork to serve Qwen3.8-Flash-Next at 60โ100 tokens/sec decode and 3000+ tokens/sec prefill. The V620 is a cut-down MI100 (CDNA1) with no display outputs, originally sold to cloud-gaming providers; the builder sourced them for ~$350 each. This demonstrates a distinct cheap high-VRAM path from salvaged enterprise AMD hardware, separate from the ex-mining BC-250 (Hopper) route. The web search returned no organic results (deadline reached), so this grounding relies solely on the case hypothesis and evidence title.
Why it matters to Scott
Converges on Scott's hardware-aware local inference and cheap high-VRAM serving positions โ a builder independently validates the economics of salvaged enterprise AMD CDNA1 (V620/MI100) with a custom vLLM fork, delivering 256GB VRAM at $2,800 and 60โ100 tok/s decode. This extends the hardware option space beyond Scott's current NVIDIA/CUDA/Ollama stack (gamepc) and directly bears on his hardware-aware inference concept and single-tenant appliance work (OpenClaw), but runs on ROCm/vLLM rather than his CUDA/Ollama substrate.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.cheap-model-front-doordev:project.openclawradar:amd-llama-cpp-prefill-speedupradar:amd-machine-readable-isa-kernelsradar:airllm-low-vram-model-streamingradar:adaptive-kv-cache-streamingradar:agent-bottling-benchmark
queries asked of Scott's wikis
- local-inference-economics salvaged GPU VRAM-per-dollar
- AMD ROCm vLLM forks CDNA1 MI100 V620 serving
- model-sovereignty open-weights local serving hardware
- hardware-hacking enterprise GPU repurposing inference
- inference-economics high-VRAM cheap decode throughput
Measured heat
now 0 pts/hpeak 95 pts/hcomments 0/hpeers p25momentum: steady1 platformsage 71h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion
How the heat travelled
pace: p89 vs 1204 stories at the 48h mark (now 71h old) โ ahead of bfl-flux-3-image-release (1.0x), behind sanotts-microcontroller-tts (1.0x)
Evidence (1) โ โญ canonical anchor
Interpretation history
2026-10-08T20:50:20Z
grounded: converges/medium โ Converges on Scott's hardware-aware local inference and cheap high-VRAM serving positions โ a builder independently validates the economics of salvaged enterpri
2026-10-08T20:35:13Z
case created โ Substantive builder result with concrete hardware, performance numbers, and custom serving fork; parallels but is distinct from the ex-mining BC-250 case.
Decision trace
- 10-10 15:35sensor_dirtyvelocity_spike
- 10-10 06:37sensor_dirtyvelocity_spike
- 10-09 20:31sensor_dirtyvelocity_spike
- 10-09 15:39sensor_dirtycomment_update
- 10-09 09:34sensor_dirtyvelocity_spike
- 10-09 07:59attention_routeThe editor compared this story and chose to keep watching.
- 10-09 07:54attention_candidatecreate
- 10-09 07:50groundConverges on Scott's hardware-aware local inference and cheap high-VRAM serving positions โ a builder independently validates the economics of salvaged enterprise AMD CDNA1 (V620/MI100) with a cu
- 10-09 07:35createSubstantive builder result with concrete hardware, performance numbers, and custom serving fork; parallels but is distinct from the ex-mining BC-250 case.