
PAC-Bench: How well can models one-shot a Pac-Man game?
An engineer asked 30 model and harness combinations to build a Pac-Man game from one short prompt. The best entry scored 99 out of 100, while the weakest recorded run managed just 2 points.
An engineer has published PAC-Bench, a public benchmark that puts one question to leading AI models: can they build a working Pac-Man game from a single short prompt, with no follow-up and no human fixes? The task is identical for every competitor, "Create a Pac-Man game in a single html page", and the results can be played side by side on the project page.
What the benchmark measures
The write-up scores 30 entries drawn from different models and harnesses. Each game was judged out of 100 points across five checks: controls (20), ghosts (25), Pac-Man getting stuck (20), the maze (20) and sound (15). The author notes that scoring was done by Opus 5.5 reviewing the live games on 2026-09-28, not by the models themselves.
Entries came from first-party harnesses such as Claude Code, Codex, Cursor Cloud and Grok Build, from Claude Code on OpenRouter, and from Antigravity with Gemini. Wall time and token counts are taken from harness transcripts; where a run recorded no number, the table shows a dash.
The scoreboard
The top of the table is dominated by Anthropic models. Claude Opus 5.5 finished first with 99 points, credited by the review with the best audio in the set: a full two-phrase intro with a bass line, a continuous siren that rises with progress, and an arcade-style death sound. Claude Fable 5.1 followed on 96, while Claude Fable 5 and a Claude Sonnet 5.5 cell on OpenRouter shared 95. xAI's Grok 4.7 took 94, and OpenAI's GPT-5.6 Sol reached 90.
The spread below that is wide. Only six entries reached 90 points or more, while the median entry scored 58 and the weakest run, MiniMax M3, managed 2. Several games were not merely rough but impossible to finish: Grok 4.5 left 10 pellets unreachable, and Moonshot's Kimi K2.7 Code left 0 of 270 pellets reachable from the spawn point.
Where the models fail
Ghost behaviour is the most common weak point. In 17 of the 30 entries the review recorded a major defect on the ghosts check: all four ghosts sharing a single AI, eaten ghosts that never return home, or ghosts drawn on top of walls. Nine entries had a major maze defect, from large numbers of dead ends to layouts that are not really Pac-Man mazes at all.
Cost and speed vary as much as quality. GPT-6 Luna came in cheapest at about $0.013, while Claude Fable 5 topped the bill at roughly $17.28. Qwen 3.8 Max was the slowest run at 44 minutes of wall time, while the in-chat Grok Bot entry took under a minute.
Why it matters
Such scores are a narrow signal. A playable arcade game is not a general measure of reasoning, and the author is explicit that a dash means the run did not record a number. What the benchmark does show is how much a one-shot task depends on the harness around the model, and how often generated code produces something that looks right but cannot be played to the end.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.