2026-10-11 17:12 UTC

Vero evaluations will determine whether AI agents can autonomously produce complete software repositories whose required behavior is established through machine-checked formal verification.

state: expiredheat: lowuncertainty: mediumconvergesscott: highcoding-agents formal-verification agent-harnessesVeroVerina

What is this?

Vero is presented as a repository-level benchmark for evaluating whether AI coding agents can jointly synthesize software implementations and machine-checked proofs that the implementations satisfy their specifications. Unlike function-level evaluations, it uses multi-module repository tasks intended to test cross-module dependencies and long-horizon reasoning. The supplied snippets do not identify Vero’s authors or establish Verina’s role, and Vero should not be confused with Scale Labs’ separately described VeRO agent-optimization harness.

Why it matters to Scott

Vero independently operationalizes Scott’s Evaluation-Driven Development and Spec-Driven Development positions by testing whether agents can turn machine-checkable specifications into repository-scale implementations with accompanying proofs. Its results could validate or constrain his load-bearing claim that generated code can become a replaceable derived artifact when governed by binding verification, while extending radar coverage beyond separate repository-scale coding and proof-synthesis evaluations.
ip:concept.evaluation-driven-developmentip:concept.spec-driven-developmentip:concept.verification-loopsip:concept.spec-as-assetradar:concept.coding-agent-benchmarksradar:concept.formal-verificationradar:mirrorcode-autonomous-project-scoperadar:mathcode-mathematical-coding-agent
queries asked of Scott's wikis
  • repository-level coding-agent evaluation harnesses
  • formal verification as an agent-generated correctness layer
  • joint code and proof synthesis
  • machine-checkable specifications for autonomous coding
  • long-horizon multi-module agent benchmarks
  • verified software generation and proof-carrying code

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnVero: Can AI Agents Build Formally Verified Software Repositories?matt_d10
🟧 echo.paper ⭐The original paper introduces Vero as “the first benchmark to evaluate joint implementation and proof synthesis at the repository level,” usZhe Ye, Hantao Lou, Yuechun Sun, Peiyang Song, Zhengxu Yan, Timothe Kasriel, Qingyang Zhang, Kaiyu Yang, Soonho Kong, Jingxuan He, and Dawn Song——

Interpretation history

Decision trace