
Can AI design circuit boards yet? Inside the EEBench benchmark
OpenAI's GPT-6 Astra demo showed a model working on a PCB in KiCad. EEBench, a public benchmark from the atopile team, grades circuit designs in simulation with real manufacturer parts and worst-case tolerances.
OpenAI put a demo of GPT-6 Astra working on a circuit board in KiCad on the front page of its launch post, a rare appearance for electronics in a major model release — and a reminder of the question the team behind the EEBench benchmark keeps asking: how do you measure whether the electronics an AI produces are actually any good?
EEBench is a public benchmark for electronics design built and funded by the team behind atopile. Their experience is that current models know far more about electronics than their output in conventional design tools shows: they have read textbooks, datasheets and application notes.
Agents work on declarative code, not in a GUI
An agent driving a graphical CAD tool spends much of its time clicking around and tracking what is on screen, with coordinates, menus and application state filling its context. EEBench uses atopile instead: the circuit lives in declarative code, so the agent works directly on components, connections and electrical constraints, and can change a design, build it, run a simulation and inspect what failed without leaving the project. The team says that beats asking a model to draw lines in a GUI.
The real world is messy
One public task is based on a residential energy meter. When its 5 V supply disappears, the circuit must keep the processor alive for another 20 ms to save the accumulated reading, and the protected rail must stay above the processor's 3.0 V brownout threshold throughout. Most models jump to the right base conclusion: add a capacitor. A real capacitor makes it harder. A ceramic part may deliver much less than its advertised capacitance once voltage is across it, parts carry tolerances, and more capacitance costs more and slows the rail's recovery when power returns. In the published example the rail falls from 4.55 volts and crosses the 3 V threshold after 0.85 ms, far short of the required 20 ms.
Deterministic grading and the leaderboard
Checks are fully deterministic: the harness builds the submitted design, constructs the circuit graph and bill of materials and runs SPICE simulations, where each requirement produces a measurement with a limit. The technical score is combined with cost efficiency against a reference bill of materials, and cost counts only once the circuit works.
In the September 1 results, Claude Opus 5 leads with 61.6% across the 13 tasks in EEBench V1, ahead of Grok 4.6 at 57.1% and Claude Fable 5.1 at 56.4%. xAI included EEBench in the Grok 4.6 model card under engineering acceleration, and its own published run put the model at 60.0% with high reasoning effort. OpenAI's tested models sit lower: GPT-5.5 at 42.3% and GPT-5.6 Sol at 39.4%. EEBench V1 does not yet test whether a model can lay out, manufacture and bring up a complete product.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.