Back
Bez: Generating a browser engine from specs and tests
SiTech AI Team3 min read

Bez: Generating a browser engine from specs and tests

Bez is a project that generates a web rendering engine from specifications and tests: models write many candidate implementations, real browsers check the result, and the code that passes is committed as ordinary Rust.

What Bez is

Bez is a project that generates a web rendering engine from specifications and tests rather than by hand. Its authors argue that hand-building an engine costs hundreds of engineers and years of work, so only a few companies can have one, and only they decide how the web works. In the pipeline, a model writes many candidate implementations from spec text; each runs inside the engine and its result is checked against Chromium, Firefox and WebKit: same boxes, same places? Rejected candidates go back to the model; passing ones are committed as ordinary Rust. WPT serves as the test suite, and once the pipeline exists, extra engines cost little.

Where the project stands

Computed against browser-compat-data 8.0.4 (17,259 leaf keys, 74 rules) on 2026-09-25: 0.6% of the surface generated, 0.3% hand-written, 0.5% linked, 5.7% oracle-only and 93.0% unreached. CSS is the most advanced area at 2.5% generated, while HTML, JavaScript, SVG, WebAssembly, HTTP and MathML remain 100% unreached. In crates/layout/src/generated, nine CSS 2.1 layout rules pass all 227 recipe cases and 11 usable WPT normal-flow pages: eight came from models and were admitted by the three-browser vote; block height stayed hand-written because no candidate beat it.

Coverage history: browser-compat-data leaf keys per status, one bar per recorded run

What the experiments found

Across 235 documents, 699 of 705 browser-pair comparisons agreed. All six disagreements were twelve-deep percentage nesting; majority voting named Firefox the outlier every time. The difference is a real web-compat bug: Gecko rounds lengths to 1/60 px where Blink and WebKit use 1/64 px, and reproductions match Mozilla-diagnosed breakage on Slack, Google Store and Samsung, plus Mozilla bug 1719314. The WPT vote table: a three-engine majority covers 2,162,676 of 2,282,301 test and subtest keys (94.8%), and conformance suites work as oracles too: WebGL with dEQP 99.7%, WPT canvas 82.6%, Web Audio 74.4%, WebGPU CTS about 85% for validation. Margin collapsing written as Datalog matched the browsers on all 1195 offsets and caught a seeded bug the geometry check missed.

Goals and open questions

The default build is a complete engine, but for non-browser uses, a content-scoped build can leave out every feature the content never touches rather than merely switch it off, which also means less work and memory. When a spec or test changes, the affected code is regenerated and re-verified, not rewritten by hand. The economics hold for part of the platform: about 55-60% of engine-relevant entries have a usable automated oracle and generatable spec prose; roughly 8-18% have neither. Open questions include what each rule costs by hand, how much of WPT is reachable without JavaScript and where IDL-generated code ends.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.