2026-10-11 16:37 UTC

A LocalLLaMA builder demonstrates a $2,800 rig using eight ex-cloud-gaming Radeon Pro V620 GPUs (256GB VRAM) with a custom vLLM fork to serve Qwen3.8-Flash-Next at 60โ€“100 tok/s decode, establishing a distinct cheap high-VRAM path from salvaged enterprise hardware.

state: seedheat: mediumuncertainty: mediumconvergesscott: mediumlocal-inference-economics salvaged-gpu-hardware vllm-forks_TheWolfOfWalmart_

What is this?

Reddit user _TheWolfOfWalmart_ on r/LocalLLaMA built a $2,800 inference server using eight decommissioned AMD Radeon Pro V620 GPUs (32GB each, 256GB total VRAM) sourced from cloud-gaming infrastructure, paired with a custom vLLM fork to serve Qwen3.8-Flash-Next at 60โ€“100 tokens/sec decode and 3000+ tokens/sec prefill. The V620 is a cut-down MI100 (CDNA1) with no display outputs, originally sold to cloud-gaming providers; the builder sourced them for ~$350 each. This demonstrates a distinct cheap high-VRAM path from salvaged enterprise AMD hardware, separate from the ex-mining BC-250 (Hopper) route. The web search returned no organic results (deadline reached), so this grounding relies solely on the case hypothesis and evidence title.

Why it matters to Scott

Converges on Scott's hardware-aware local inference and cheap high-VRAM serving positions โ€” a builder independently validates the economics of salvaged enterprise AMD CDNA1 (V620/MI100) with a custom vLLM fork, delivering 256GB VRAM at $2,800 and 60โ€“100 tok/s decode. This extends the hardware option space beyond Scott's current NVIDIA/CUDA/Ollama stack (gamepc) and directly bears on his hardware-aware inference concept and single-tenant appliance work (OpenClaw), but runs on ROCm/vLLM rather than his CUDA/Ollama substrate.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:concept.cheap-model-front-doordev:project.openclawradar:amd-llama-cpp-prefill-speedupradar:amd-machine-readable-isa-kernelsradar:airllm-low-vram-model-streamingradar:adaptive-kv-cache-streamingradar:agent-bottling-benchmark
queries asked of Scott's wikis
  • local-inference-economics salvaged GPU VRAM-per-dollar
  • AMD ROCm vLLM forks CDNA1 MI100 V620 serving
  • model-sovereignty open-weights local serving hardware
  • hardware-hacking enterprise GPU repurposing inference
  • inference-economics high-VRAM cheap decode throughput

Measured heat

now 0 pts/hpeak 95 pts/hcomments 0/hpeers p25momentum: steady1 platformsage 71h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-08 17:12โญ origin directly observed$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill
_TheWolfOfWalmart_ on r/LocalLLaMA
โ€”
10-08 17:12amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1x0wnz1
_TheWolfOfWalmart_
peak 405 ยท 175 comments ยท 100% of case engagement
10-08 17:34our radar first saw it ยท +0.4hdiscovery anchor: reddit.post.1x0wnz1โ€”
pace: p89 vs 1204 stories at the 48h mark (now 71h old) โ€” ahead of bfl-flux-3-image-release (1.0x), behind sanotts-microcontroller-tts (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญ$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill
LocalLLaMA
_TheWolfOfWalmart_405175

Interpretation history

Decision trace