2026-10-11 17:11 UTC

Independent reproduction will determine whether storing model weights on a low-cost AMD FPGA can deliver approximately 60,000 tokens per second with practically useful LLM behavior.

state: expiredheat: lowuncertainty: highconvergesscott: mediumspecialized-inference fpga-inference inference-economics local-inferenceMike AylesAMDTaalas

What is this?

A Show HN author identified as Mike Ayles reportedly built a Taalas-inspired design that stores LLM weights on-chip in an approximately $250 AMD FPGA and claims throughput near 60,000 tokens per second. The supplied snippets support the general premise that decode performance is often constrained by weight movement and that FPGA designs can exploit specialized memory paths, but they do not document this implementation, identify the model or workload, establish practically useful language behavior, or provide an independent reproduction. The 60,000-token figure is also ambiguous without batch size and per-user interactivity measurements, so the central performance and economics claims remain unverified here.

Why it matters to Scott

If independently reproduced with disclosed model quality, batch size, latency and power data, this would materially extend Scott’s hardware-aware local-inference work and his claim that cheap, deployable capability can outperform premium but costly infrastructure. It is not yet high relevance because the supplied evidence is only a headline-level performance claim; the radar already tracks AMD–Taalas specialized inference, but not this distinct low-cost FPGA implementation.
dev:concept.hardware-aware-local-inferenceip:concept.usable-mass-over-unusable-powerip:concept.capability-auditip:concept.latencyradar:amd-taalas-silicon-etched-inferenceradar:concept.inference-economicsradar:concept.local-inferenceradar:concept.inference-efficiency
queries asked of Scott's wikis
  • on-chip weights versus memory-bandwidth-bound decoding
  • FPGA and custom-silicon inference economics
  • local inference hardware specialization
  • throughput versus per-user token latency
  • extreme quantization and practically useful model behavior
  • independent benchmarking of inference claims

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnShow HN: Taalas-style on-chip LLM weights on a $250 AMD FPGA (60k tok/s)mikeayles7730
🟧 echo.blog ⭐The author reports a Taalas-inspired implementation using on-chip LLM weights on a roughly $250 AMD FPGA and claims throughput around 60,000Mike Ayles——

Interpretation history

Decision trace