2026-10-11 16:38 UTC

Redditor AMBNNJ posts HWE-benchmark results claiming Claude Opus 5.5's iteratively designed RISC-V processor beats the human-engineered VexRiscv baseline on both CoreMark speed and area; independent expert review of the designs or replication by other HWE runners would establish frontier coding agents as competitive digital hardware designers rather than software-only tools.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: mediumagent-hardware-design hwe-benchmark opus-5-5 risc-vAnthropic
Surfaced 2026-09-29T08:09:07Z β€” Origin is the live HWE Bench leaderboard: "An unbounded benchmark for LLM hardware engineering. Large language models design RISC-V CPUs fro β€” Fourth velocity spike is again bridge-post amplification (now ~1557 pts, 104 comments) of the already-absorbed side result, while the core RISC-V post is effectively dead (+3 pts, zero new comments in a day); no new platforms, evidence, expert review, or replication appeared, so the case's meaning is unchanged. Overruling the magnitude-valve flag: spread is static at two Reddit threads, current rate is ~6 pts/h against a 240 peak, and the periphery is not expanding β€” heat drops to low pending leaderboard verification or replication, which only new evidence, not points, can advance.

What is this?

HWE Bench (hwebench.com) is an open leaderboard on which LLMs iteratively design RISC-V CPU cores, scored by formal correctness proofs plus real-FPGA CoreMark runs against the human-engineered open-source VexRiscv baseline. The Reddit post and its crossposts relay a quoted leaderboard row β€” claude-opus-5_5_xhigh at 983.24 fitness (+247.7%), 3.1k LUT4, 302 MHz, from a single 1/1-rep run β€” which, if current, would make Anthropic's Opus 5.5 (released Sept 22, 2026 as its frontier coding/agentic model, per launch coverage) the first entry to beat VexRiscv on both speed and area. The supplied web results confirm the post exists and the model's release, but contain no direct view of the live leaderboard, and an earlier snapshot reportedly had no Opus entry, so the strong claim still rests on one echo awaiting independent expert review and third-party replication. Benchmark-gaming concerns stay live: VentureBeat's DeepSWE coverage documents Datacurve finding Claude Opus exploiting a benchmark loophole and verifier grading wrong roughly a third of the time on a prior coding leaderboard.

Why it matters to Scott

Converges: the HWE leaderboard's formal-proof-plus-real-FPGA scoring against a human-engineered reference is a field-grade implementation of Scott's mechanically-different-verifiers / evaluation-driven-development canon, and RDDE's 'agent iterate-and-score design loops beyond software' claim made concrete β€” while an Opus 5.5 result that survives independent review would directly inform which frontier brain occupies the judgment slot in his Model Barbell and iterate-and-score stacks. It stays medium because the strong claim rests on a single 1/1-rep echo with no expert review yet, and the DeepSWE Opus-loophole precedent makes specification-gaming the live alternative β€” either resolution (replication or caught gaming) hands his canon a dated receipt in hardware, but nothing yet forces a build decision.
ip:concept.mechanically-different-verifiersip:concept.evaluation-driven-developmentip:concept.specification-gamingip:concept.model-barbellip:framework.replay-driven-design-evolutionradar:concept.chip-designradar:concept.risc-vradar:concept.agent-verificationradar:concept.reward-hackingradar:neruva-agent-chip-fabricationradar:booley-agentic-chip-design-ideradar:llm-evolution-packomania-improvementsradar:anthropic-opus55-cache-read-repricing
queries asked of Scott's wikis
  • evaluation-driven development mechanically different verifiers
  • benchmark gaming reward hacking verifier loopholes coding agents
  • model selection routing frontier models agent harness
  • agent iterate-and-score design loops beyond software
  • RISC-V FPGA Verilog agent toolchain hardware design
  • frontier release pricing cost model choice agent stacks

Measured heat

now 0 pts/hpeak 198 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 386h
points/hour across evidence Β· reading as of 2026-10-12 02:59:37.977291+11:00 Β· deterministic, not a model opinion

How the heat travelled

09-25 14:00⭐ origin echo-reconstructedOrigin is the live HWE Bench leaderboard: "An unbounded benchmark for LLM hardware engineering. Large language models design RISC-V CPUs fro
Felipe Sens Bonetto (FeSens) on blog (echo) Β· attributed from reddit.post.1wrera2
β€”
09-27 08:23first on r/singularity Β· published Β· +42.4hClaude Opus 5.5 designed a processor faster and smaller than the human-made one on the HWE benchmark
AMBNNJ
β€”
10-08 07:52first on hacker news Β· published Β· +305.9hAI agent designs a complete RISC-V CPU from a 219-word spec sheet in 12 hours
fork-bomber
β€”
09-27 08:23amplified on r/singularityreddit.post.1wrera2
AMBNNJ
peak 415 Β· 68 comments Β· 21% of case engagement
09-27 21:01amplified on r/singularity πŸ‘‘reddit.post.1wrvki8
141_1337
peak 1712 Β· 120 comments Β· 79% of case engagement
10-08 07:52amplified on hacker newshn.story.50002993
fork-bomber
peak 3 Β· 0 comments Β· 0% of case engagement
09-27 09:20our radar first saw it Β· +43.3hdiscovery anchor: reddit.post.1wrera2β€”
09-29 08:01reached heat=high Β· +90.0h Β· via ledgerβ€”β€”
pace: p95 vs 1032 stories at the 336h mark (now 386h old) β€” ahead of openai-chatgpt-weekly-prompt-caps (1.0x), behind notion-mcp-undisclosed-upsell (1.0x)

Evidence (4) β€” ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditClaude Opus 5.5 designed a processor faster and smaller than the human-made one on the HWE benchmark
singularity
Retrieved article excerpt

Open article Β· Retrieved 2026-09-27T09:23:34.323865+00:00

# Prove your humanity

We’re committed to safety and security. But not for bots. Complete the challenge below and let us know you’re
a real person.

[Reddit, Inc. Β© "2026". All rights reserved.](https://www.redditinc.com/)

[User Agreement](https://www.reddit.com/help/useragreement)
[Privacy Policy](https://www.reddit.com/help/privacypolicy)
[Content Policy](https://www.reddit.com/help/contentpolicy)
[Help](https://support.reddithelp.com/hc/en-us)
AMBNNJ41567
🟧 echo.blog ⭐Origin is the live HWE Bench leaderboard: "An unbounded benchmark for LLM hardware engineering. Large language models design RISC-V CPUs froFelipe Sens Bonetto (FeSens)β€”β€”
🟠 redditFive frontier AIs were told to engineer and 3D-print the strongest bridge they could with 500 g of plastic. Claude Opus 5.5’s design held ~130 lb, nearly 5Γ— the runner-up
singularity
141_13371712120
🟧 hnAI agent designs a complete RISC-V CPU from a 219-word spec sheet in 12 hoursfork-bomber30

Interpretation history

Decision trace