2026-10-11 18:00 UTC

Bez's maintainer claims a generate-and-verify pipeline โ€” model-written candidates checked against Chromium, Firefox, and WebKit plus WPT โ€” can grow from its current 0.6% browser-compat coverage into a usable generated web engine; coverage climbing toward usable rendering versus stalling resolves whether spec-driven generation can build large multi-component systems.

state: corroboratedheat: mediumuncertainty: mediumconvergesscott: highcoding-agents spec-driven-generation

What is this?

Bez is an AI-generated browser engine project: per its maintainer's write-up ('Bez: Generating a browser engine from specs and tests'), an LLM writes implementation candidates that are accepted or rejected by checking them against the actual behavior of Chromium, Firefox, and WebKit plus the Web Platform Tests (WPT) suite โ€” generate-and-verify against existing engines rather than human implementation from scratch. The project reports roughly 0.6% browser-compat coverage today and, unusually for this genre, documents that ~93% of the target surface remains unreached, framing a resolvable bet: either coverage climbs toward a usable engine or the approach stalls. Caveat: none of the supplied web results mention Bez itself; they only ground the surrounding landscape (Chromium at ~75% share, Servo 0.4 in July 2026 as the first credible fourth engine built by closely following spec terminology, and WPT as the shared conformance suite even human-built engines verify against), so Bez's specifics and trajectory rest entirely on the maintainer's own account.

Why it matters to Scott

Bez independently runs the pipeline Scott's canon argues for โ€” generate-and-verify against a characterisation oracle of three engines + WPT, i.e. the AI Legacy Takeover convergence loop (OHBVP, measured by characterisation tests) aimed at the largest possible target โ€” and because that oracle is near-perfect, whether coverage climbs from 0.6% or stalls cleanly discriminates between his mechanisms: failure would isolate upstream decomposition judgment (Hora's Watchmaker) rather than verification as the binding constraint on spec-driven generation at engine scale, sharpening the boundary of the verification-cost/spec-as-asset thesis. The maintainer's honest 93%-unreached reporting is a live receipt for Discussed Is Not Deployed, and the explicitly resolvable bet is dated-receipts material โ€” Scott can stake his frameworks' predictions on the record now.
ip:concept.verification-loopsip:concept.characterisation-testingip:source.ai-legacy-takeoverip:concept.spec-driven-developmentip:framework.horas-watchmakerip:concept.verification-costip:framework.discussed-is-not-deployedradar:mirrorcode-long-horizon-reimplementationradar:mirrorcode-autonomous-project-scoperadar:shumer-logic-gate-computerradar:doltlite-2000-agent-pr-buildradar:five-bugs-test-spec-blindspotradar:hwatu-verification-browser
queries asked of Scott's wikis
  • generate-and-verify loop verification-gated coding agent harness
  • spec-driven generation limits large multi-component systems
  • tests and specs as ground truth automated oracle for code generation
  • long-horizon autonomous coding AI builds big systems claims evidence
  • browser engine diversity Servo WPT web platform infrastructure

Measured heat

now 0 pts/hpeak 24 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 242h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

10-01 14:00โญ origin echo-reconstructedThe repo README is the primary artifact. It states: \"Bez - a generated web engine. Generate a web rendering engine from specs and tests.\"
burrito.space (Dietrich Ayala, ex-Mozilla) โ€” HN submitted by third party nerdypepper on github (echo) ยท attributed from hn.story.49925036
โ€”
10-01 18:08first on hacker news ยท published ยท +4.1hBez: Generating a browser engine from specs and tests
nerdypepper
โ€”
10-01 18:08amplified on hacker news ๐Ÿ‘‘hn.story.49925036
nerdypepper
peak 126 ยท 50 comments ยท 97% of case engagement
10-02 02:09amplified on hacker newshn.story.49929194
albert-yu
peak 2 ยท 3 comments ยท 3% of case engagement
10-01 20:21our radar first saw it ยท +6.4hdiscovery anchor: hn.story.49925036โ€”
pace: p72 vs 1188 stories at the 168h mark (now 242h old) โ€” ahead of microsoft-vibevoice-streaming-asr (1.0x), behind magic-v5-pretraining-efficiency (1.0x)

Evidence (3) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸง hnBez: Generating a browser engine from specs and tests
Retrieved article excerpt

Open article ยท Retrieved 2026-10-01T20:35:57.630792+00:00

# Bez - a generated web engine[#](https://tangled.org/burrito.space/bez#bez---a-generated-web-engine)

Generate a web rendering engine from specs and tests.

## Why[#](https://tangled.org/burrito.space/bez#why)

Building a web engine by hand costs hundreds of engineers and many years, so only a few
companies can have one, and only they decide how the web works. Bez generates the engine
instead: from the specifications, with the three shipping browsers (checked against each
other) and WPT as the tests. Once the pipeline exists, each extra engine costs very
little.

## Goals[#](https://tangled.org/burrito.space/bez#goals)

- **Complete, or tree-shaken.** The default build is a complete engine. For uses
  other than a web browser, a build can instead contain only the web features its
  content uses: analyse a site or an app, and every feature it never touches is left out
  of the binary, not just switched off (`docs/content-scoped-engine.md`).
- **Fast on any device.** A smaller engine does less work and needs less memory.
- **Easy to embed.** A small engine with a clear API, for apps and devices that today ship
  all of Chromium or go without.
- **Fastest to update.** When a spec or test changes, the affected code is regenerated and
  re-verified, not rewritten by hand.

## How it works[#](https://tangled.org/burrito.space/bez#how-it-works)

```
flowchart LR
  spec["spec text"] --> model["model writes<br/>many candidates"]
  model --> engine["candidate runs<br/>inside the engine"]
  browsers["Chromium, Firefox,<br/>WebKit (cached)"] --> check{"same boxes,<br/>same places?"}
  engine --> check
  check -- no --> model
  check -- yes --> code["committed as<br/>ordinary Rust"]
```

- [`docs/generation-loop.md`](https://tangled.org/burrito.space/bez/tree/main/docs/generation-loop.md) โ€” how one rule goes from spec text
  to committed code.
- [`docs/geometry-verification.md`](https://tangled.org/burrito.space/bez/tree/main/docs/geometry-verification.md) โ€” how the engine's
  layout is checked against three browsers.
- [`docs/generated-functions.md`](https://tangled.org/burrito.space/bez/tree/main/docs/generated-functions.md) โ€” the layout rules generated so far and what each
  batch showed.
- [`roadmap.md`](https://tangled.org/burrito.space/bez/tree/main/roadmap.md) โ€” the plan, the decisions and the open questions.

## Status[#](https://tangled.org/burrito.space/bez#status)

Coverage history: share of browser-compat-data leaf keys per status, one bar per recorded run, with a line behind the bars for the number of leaf keys reached out of the total

Computed against browser-compat-data 8.0.4 (17259 leaf keys, 74 rules in [`docs/feature-map.json`](https://tangled.org/burrito.space/bez/tree/main/docs/feature-map.json)) on 2026-09-25.

| area | generated | hand-written | linked | oracle only | unreached |
| --- | --- | --- | --- | --- | --- |
| overall | 0.6% | 0.3% | 0.5% | 5.7% | 93.0% |
| css | 2.5% | 1.1% | 2.0% | 0.0% | 94.4% |
| html | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |
| api | 0.0% | 0.0% | 0.0% | 9.9% | 90.1% |
| javascript | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |
| svg | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |
| webassembly | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |
| http | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |
| mathml | 0.0% | 0.0% | 0.0% | 0.0% | 100.0% |

See [`docs/dashboard.md`](https://tangled.org/burrito.space/bez/tree/main/docs/dashboard.md) for what these statuses mean and how this table is kept in sync.

DOM, style, box tree and fragment tree are hand-written in [`crates/dom`](https://tangled.org/burrito.space/bez/tree/main/crates/dom) and
[`crates/layout`](https://tangled.org/burrito.space/bez/tree/main/crates/layout) and pass their browser checks
([roadmap.md โ†’ "Phase 1 โ€” Engine bootstrap"](https://tangled.org/burrito.space/bez/tree/main/roadmap.md#phase-1--engine-bootstrap)).
Nine CSS 2.1 layout rules live in
[`crates/layout/src/generated`](https://tangled.org/burrito.space/bez/tree/main/crates/layout/src/generated); eight were written by
models and admitted against the three-browser vote, and block height keeps its
hand-written rule because no model candidate beat it. Together they pass all 227 recipe
cases and all 11 usable WPT normal-flow pages. Which rules the engine actually calls:
[roadmap.md โ†’ "Which resolver each generated function runs"](https://tangled.org/burrito.space/bez/tree/main/roadmap.md#which-resolver-each-generated-function-runs).
What would prove the thesis:
[roadmap.md โ†’ "Proof of concept"](https://tangled.org/burrito.space/bez/tree/main/roadmap.md#proof-of-concept).

## Findings so far[#](https://tangled.org/burrito.space/bez#findings-so-far)

- **Three browsers mostly agree, and the third names the odd one out.** 699 of 705
  browser-pair comparisons agreed across 235 documents (commit `3ca7323`); all six
  disagreements were twelve-deep percentage nesting, and majority voting named Firefox
  the outlier every time. [`docs/firefox-app-units.md`](https://tangled.org/burrito.space/bez/tree/main/docs/firefox-app-units.md).
- **That Firefox difference is a real web-compat bug.** Gecko rounds lengths to 1/60 px
  where Blink and WebKit use 1/64 px, which makes a flex item wrap only in Firefox or
  `offsetWidth` differ by a pixel. Reproductions match Mozilla-diagnosed breakage on
  Slack, Google Store and Samsung, and Mozilla bug 1719314.
  [`experiments/firefox-compat/`](https://tangled.org/burrito.space/bez/tree/main/experiments/firefox-compat).
- **The WPT vote table agrees across most of the suite.** A stable three-engine majority
  covers 2,162,676 of 2,282,301 test/subtest keys (94.8%). [`docs/wpt-votes.md`](https://tangled.org/burrito.space/bez/tree/main/docs/wpt-votes.md).
- **Script-observable behaviour agrees too.** Trace probes under a fake-media profile
  agree on 15 of 18 engine pairs; the three disagreements are a real platform difference.
  [docs/dump-format.md โ†’ "Trace protocol"](https://tangled.org/burrito.space/bez/tree/main/docs/dump-format.md#trace-protocol).
- **Conformance suites are usable oracles.** Counting a test usable when two of three
  engines agree: WPT canvas 82.6% (92.7% excluding tentative), Khronos WebGL with dEQP
  99.7%, WPT Web Audio 74.4% (85.6% excluding tentative); the WebGPU CTS about 85% for
  validation and about 50% for numeric execution.
  [`docs/conformance-oracles.md`](https://tangled.org/burrito.space/bez/tree/main/docs/conformance-oracles.md), [`docs/webgpu-cts.md`](https://tangled.org/burrito.space/bez/tree/main/docs/webgpu-cts.md).
- **Logic programming earns a narrow place.** Margin collapsing written as Datalog agreed
  with the browsers on 1195 of 1195 offsets and caught a seeded bug the geometry check
  missed in all 227 cases. [`docs/logic-layer.md`](https://tangled.org/burrito.space/bez/tree/main/docs/logic-layer.md).
- **The economics hold for a subset of the platform.** About 55โ€“60% of engine-relevant
  compat entries have a usable automated oracle and generatable spec prose; roughly 8โ€“18%
  have none. [`docs/platform-economics.md`](https://tangled.org/burrito.space/bez/tree/main/docs/platform-economics.md),
  [`docs/oracle-coverage.md`](https://tangled.org/burrito.space/bez/tree/main/docs/oracle-coverage.md).
- **An engine scoped to one site's content is designed**, and its content analyser has
  landed. [`docs/content-scoped-engine.md`](https://tangled.org/burrito.space/bez/tree/main/docs/content-scoped-engine.md).

## Open questions[#](https://tangled.org/burrito.space/bez#open-questions)

What building each rule costs by hand, how wide a feature's generated part should be,
how much of WPT is reachable without JavaScript, and where IDL-generated code ends:
[roadmap.md โ†’ "Open questions"](https://tangled.org/burrito.space/bez/tree/main/roadmap.md#open-questions). Area-by-area status:
[`docs/platform-areas.md`](https://tangled.org/burrito.space/bez/tree/main/docs/platform-areas.md).

## Sources[#](https://tangled.org/burrito.space/bez#sources)

- browser-specs, every web spec in one repo: <https://github.com/w3c/browser-specs>
- web-features, the web platform grouped into features: <https://github.com/web-platform-dx/web-features/>
- browser-compat-data, every piece of the web surface: <https://github.com/mdn/browser-compat-data>
- WPT, the web platform tests: <https://web-platform-tests.org/>
- wpt-gen, WPT tests generated from specs: <https://github.com/GoogleChromeLabs/wpt-gen>
- webref, machine-readable terms from web specs: <https://github.com/w3c/webref>
nerdypepper12650
๐ŸŸง echo.github โญThe repo README is the primary artifact. It states: \"Bez - a generated web engine. Generate a web rendering engine from specs and tests.\" burrito.space (Dietrich Ayala, ex-Mozilla) โ€” HN submitted by third party nerdypepperโ€”โ€”
๐ŸŸง hnShow HN: Visi โ€“ Excel as a CLIalbert-yu23

Interpretation history

Decision trace