2026-10-11 17:10 UTC

fpga-inference

band: coolmomentum: stable score: 0.18
temperature history

Episodes (3)

Andrew Fan claims his published INT4 TPU design and programmable firmware provide a low-cost Cmod A7 testbed for simplified transformer inference, enabling hands-on kernel and memory-bottleneck experiments without a GPU.
seedknownscott: low
Independent reproduction will determine whether storing model weights on a low-cost AMD FPGA can deliver approximately 60,000 tokens per second with practically useful LLM behavior.
expiredconvergesscott: medium
LocalLLaMA builder I_am_purrfect claims a from-scratch FPGA implementation of the Qwen3.5 architecture โ€” which he describes as largely built with Claude Opus 4.8 โ€” runs 9B/27B INT4 models on ~$280 used SQRL FK33 mining cards (8GB HBM2 ~400GB/s, multi-card for 27B); independent replication, an opened repo, or published tok/s numbers would establish ex-mining FPGA fabric as a genuinely new cheap local-inference tier, while debunking, unusable performance, or a repo-less fade closes it as an unrealized build.
corroboratedconvergesscott: medium