LocalLLaMA builder I_am_purrfect claims a from-scratch FPGA implementation of the Qwen3.5 architecture β which he describes as largely built with Claude Opus 4.8 β runs 9B/27B INT4 models on ~$280 used SQRL FK33 mining cards (8GB HBM2 ~400GB/s, multi-card for 27B); independent replication, an opened repo, or published tok/s numbers would establish ex-mining FPGA fabric as a genuinely new cheap local-inference tier, while debunking, unusable performance, or a repo-less fade closes it as an unrealized build.
state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumfpga-inference local-inference-hardware agent-assisted-hardware-designI_am_purrfect
What is this?
I_am_purrfect, a builder on r/LocalLLaMA, reports a from-scratch FPGA implementation of the Qwen3.5 architecture β written largely with Claude Opus 4.8 as the coding agent β that runs INT4-quantized 9B/27B Qwen3.5 models on used SQRL FK33 mining cards (~$280 each, 8GB HBM2 at ~400GB/s, multiple cards for 27B). The supplied web results contain no trace of this post, the builder, or the FK33 card itself: they only confirm the surrounding context β Qwen3.5 is a real Apache-2.0 open-weights family (released February 2026; the 27B dense variant ships as a ~20GB 4-bit build with 262k context) normally served on 24GB consumer GPUs or Apple Silicon via the GGUF/llama.cpp/Ollama stack. The build therefore rests entirely on the builder's own testimony: no repo, tok/s figures, card specs, or independent replication appear anywhere in the supplied material.
Why it matters to Scott
If verified (repo, tok/s, replication), this is a two-front dated receipt for positions Scott's canon already holds: Opus-written RTL for a full model architecture extends his agentic-coding thesis that disciplined agent loops ship real artifacts β from software into silicon β and a ~$280 8GB-HBM2/~400GB/s card serving 27B INT4 would open a new tier directly under hardware-aware local inference and his gamepc serving substrate. Until then it is unverified first-party testimony, so the value is conditional on the artifact landing; what's genuinely new versus the radar's lineage (Cmod A7 testbed, BC-250 mining rig, Opus-RISC-V build) is the three-way synthesis of agent-designed RTL + ex-mining fabric + frontier-model serving.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookradar:concept.fpga-inferenceradar:andrew-fan-fpga-transformer-tpuradar:bc250-mining-rig-local-inferenceradar:opus55-hwe-riscv-beats-vexriscvradar:concept.chip-designradar:concept.memory-bandwidth
queries asked of Scott's wikis
- FPGA inference for local LLMs
- repurposing used mining GPUs/cards for AI workloads
- coding agent writing Verilog or hardware RTL
- local inference hardware cost per token amortization
- memory bandwidth bound token generation decode speed
- INT4 quantization for local model serving
Measured heat
now 0 pts/hpeak 75 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1250h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion
How the heat travelled
Evidence (3) β β canonical anchor
Interpretation history
2026-10-08T01:39:45Z
Minor engagement uptick (score 343β354, comments 40β43) on the aging Reddit post triggered velocity spikes, but measured_heat shows current rate ~0 pts/h, momentum steady, peer percentile 50 β the spikes are noise on a cooling curve. No benchmarks, replication, RTL review, or debunk have appeared. The case remains a promising first-party build with a public MIT-licensed repo awaiting its performance numbers.
2026-10-05T12:48:04Z
The attention wave crested and is now cooling (peak ~75 pts/h β ~5.7, comments down to cheerleading) while HN uptake stayed dead β the episode now reads as a promising first-party build awaiting its numbers, not a spreading story; the magnitude-valve spread reading is inflated because the three 'platforms' are one hot Reddit post, the artifact's own repo echo, and a dead HN submission. The only new substance is an unverified technical reply citing a 75MHz clock from the gallery screenshots β suggestive (design likely runs far under HBM potential) but not a benchmark, review, or replication, so no material change and heat cools to low.
2026-10-04T22:02:26Z
Graduates from unverified testimony to a checkable public artifact: the MIT-licensed full-fabric VHDL repo is now the canonical anchor (linked from the Reddit selftext, origin-walked to Nero7991 at conf 0.85) and jacquesm independently surfaced it on HN β though with zero traction. The open question narrows from 'does the artifact exist' to 'does it perform': still no tok/s, RTL review, or replication; the magnitude-valve spread reading is inflated because the echo counts as a platform, so medium heat rather than high.
2026-10-04T21:24:13Z
evidence attached: hn.story.49957671 β Independent parallel full-fabric VHDL transformer-inference engine running Qwen3.5-class models materially contextualises whether FPGA fabric is emerging as a real local-inference tier, though zero traction and no perf numbers yet.
2026-10-04T17:36:59Z
origin walked (opencode/cheap-glm, conf 0.85): anchor reddit.post.1wxken1 -> echo.github.e11005881c by Nero7991 (GitHub) β evidently the same individual as Reddit user u/I_am_purrfect
2026-10-04T17:33:37Z
grounded: converges/medium β If verified (repo, tok/s, replication), this is a two-front dated receipt for positions Scott's canon already holds: Opus-written RTL for a full model architect
2026-10-04T17:25:10Z
case created β The builder's own post claims frontier-class 9B/27B models running on $280 ex-mining FPGA cards β a surprising, checkable first-party artifact claim squarely in Scott's local-inference and agent-designed-hardware interests, distinct from the existing FPGA (Andrew Fan Cmod A7 testbed) and ex-mining GPU (BC250 rig) cases.
Decision trace
- 10-08 12:40attention_routeThe editor compared this story and chose to keep watching.
- 10-08 12:39attention_candidatecoverage review: magnitude valve eligible (multi-platform, top-decile engagement); not yet communicated
- 10-08 12:39repriceMinor engagement uptick (score 343β354, comments 40β43) on the aging Reddit post triggered velocity spikes, but measured_heat shows current rate ~0 pts/h, momentum steady, peer percentile 50 β the spi
- 10-06 10:22sensor_dirtyvelocity_spike
- 10-06 02:23sensor_dirtyvelocity_spike
- 10-05 23:48repriceThe attention wave crested and is now cooling (peak ~75 pts/h β ~5.7, comments down to cheerleading) while HN uptake stayed dead β the episode now reads as a promising first-party build awaiting its n
- 10-05 19:21sensor_dirtyvelocity_spike
- 10-05 15:20sensor_dirtycomment_update
- 10-05 12:21sensor_dirtyvelocity_spike
- 10-05 10:21sensor_dirtycomment_update
- 10-05 09:02repriceGraduates from unverified testimony to a checkable public artifact: the MIT-licensed full-fabric VHDL repo is now the canonical anchor (linked from the Reddit selftext, origin-walked to Nero7991 at co
- 10-05 08:24attachIndependent parallel full-fabric VHDL transformer-inference engine running Qwen3.5-class models materially contextualises whether FPGA fabric is emerging as a real local-inference tier, though zero tr
- 10-05 08:24propose_attachIndependent parallel full-fabric VHDL transformer-inference engine running Qwen3.5-class models materially contextualises whether FPGA fabric is emerging as a real local-inference tier, though zero tr
- 10-05 06:21sensor_dirtyvelocity_spike
- 10-05 05:21sensor_dirtycomment_update
- 10-05 04:36promote_anchororigin walk conf 0.85
- 10-05 04:33groundIf verified (repo, tok/s, replication), this is a two-front dated receipt for positions Scott's canon already holds: Opus-written RTL for a full model architecture extends his agentic-coding thes
- 10-05 04:25createThe builder's own post claims frontier-class 9B/27B models running on $280 ex-mining FPGA cards β a surprising, checkable first-party artifact claim squarely in Scott's local-inference and a