2026-10-11 16:38 UTC

LocalLLaMA builder I_am_purrfect claims a from-scratch FPGA implementation of the Qwen3.5 architecture β€” which he describes as largely built with Claude Opus 4.8 β€” runs 9B/27B INT4 models on ~$280 used SQRL FK33 mining cards (8GB HBM2 ~400GB/s, multi-card for 27B); independent replication, an opened repo, or published tok/s numbers would establish ex-mining FPGA fabric as a genuinely new cheap local-inference tier, while debunking, unusable performance, or a repo-less fade closes it as an unrealized build.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: mediumfpga-inference local-inference-hardware agent-assisted-hardware-designI_am_purrfect

What is this?

I_am_purrfect, a builder on r/LocalLLaMA, reports a from-scratch FPGA implementation of the Qwen3.5 architecture β€” written largely with Claude Opus 4.8 as the coding agent β€” that runs INT4-quantized 9B/27B Qwen3.5 models on used SQRL FK33 mining cards (~$280 each, 8GB HBM2 at ~400GB/s, multiple cards for 27B). The supplied web results contain no trace of this post, the builder, or the FK33 card itself: they only confirm the surrounding context β€” Qwen3.5 is a real Apache-2.0 open-weights family (released February 2026; the 27B dense variant ships as a ~20GB 4-bit build with 262k context) normally served on 24GB consumer GPUs or Apple Silicon via the GGUF/llama.cpp/Ollama stack. The build therefore rests entirely on the builder's own testimony: no repo, tok/s figures, card specs, or independent replication appear anywhere in the supplied material.

Why it matters to Scott

If verified (repo, tok/s, replication), this is a two-front dated receipt for positions Scott's canon already holds: Opus-written RTL for a full model architecture extends his agentic-coding thesis that disciplined agent loops ship real artifacts β€” from software into silicon β€” and a ~$280 8GB-HBM2/~400GB/s card serving 27B INT4 would open a new tier directly under hardware-aware local inference and his gamepc serving substrate. Until then it is unverified first-party testimony, so the value is conditional on the artifact landing; what's genuinely new versus the radar's lineage (Cmod A7 testbed, BC-250 mining rig, Opus-RISC-V build) is the three-way synthesis of agent-designed RTL + ex-mining fabric + frontier-model serving.
dev:concept.hardware-aware-local-inferencedev:project.gamepcip:source.your-ai-can-code-you-just-don-t-know-how-to-drive-it-ebookradar:concept.fpga-inferenceradar:andrew-fan-fpga-transformer-tpuradar:bc250-mining-rig-local-inferenceradar:opus55-hwe-riscv-beats-vexriscvradar:concept.chip-designradar:concept.memory-bandwidth
queries asked of Scott's wikis
  • FPGA inference for local LLMs
  • repurposing used mining GPUs/cards for AI workloads
  • coding agent writing Verilog or hardware RTL
  • local inference hardware cost per token amortization
  • memory bandwidth bound token generation decode speed
  • INT4 quantization for local model serving

Measured heat

now 0 pts/hpeak 75 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 1250h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

08-20 13:50⭐ origin echo-reconstructedPrimary artifact is the MIT-licensed repo itself, not the Reddit posts. Repo description: "Full-fabric VHDL Qwen3.5 9B/ Qwen3.8 27B LLM infe
Nero7991 (GitHub) β€” evidently the same individual as Reddit user u/I_am_purrfect on github (echo) Β· attributed from reddit.post.1wxken1
β€”
10-04 16:51first on r/LocalLLaMA Β· published Β· +1083.0hQwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware
I_am_purrfect
β€”
10-04 20:48first on hacker news Β· published Β· +1087.0hFull-fabric VHDL LLM inference engine. Runs Qwen3.5-class transformer inference
jacquesm
β€”
10-04 16:51amplified on r/LocalLLaMA πŸ‘‘reddit.post.1wxken1
I_am_purrfect
peak 354 Β· 43 comments Β· 99% of case engagement
10-04 20:48amplified on hacker newshn.story.49957671
jacquesm
peak 2 Β· 0 comments Β· 1% of case engagement
10-04 17:20our radar first saw it Β· +1083.5hdiscovery anchor: reddit.post.1wxken1β€”

Evidence (3) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditQwen3.5 arch implementation in FPGA fabric for 9B/27B INT4 models on relatively cheap eBay mining hardware
LocalLLaMA
Retrieved article excerpt

Open article Β· Retrieved 2026-10-04T17:23:47.229420+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
I_am_purrfect35443
🟧 echo.github ⭐Primary artifact is the MIT-licensed repo itself, not the Reddit posts. Repo description: "Full-fabric VHDL Qwen3.5 9B/ Qwen3.8 27B LLM infeNero7991 (GitHub) β€” evidently the same individual as Reddit user u/I_am_purrfectβ€”β€”
🟧 hnFull-fabric VHDL LLM inference engine. Runs Qwen3.5-class transformer inferencejacquesm20

Interpretation history

Decision trace