Back
Tiny AI Arena: four language models fight to the death on a pixel grid
SiTech AI Team3 წთ. საკითხავი

Tiny AI Arena: four language models fight to the death on a pixel grid

Tiny AI Arena is a spectator project in which four language models fight to the death on a pixel grid, with an Elo leaderboard and frame-by-frame replays. Its Hacker News thread debated what a win really proves.

On 27 September 2026 a project called Tiny AI Arena appeared on Hacker News in a Show HN post by a developer posting as hp6. Rather than another static benchmark table, it lets visitors watch four language models fight to the death on a small pixel grid and then replay every move frame by frame.

A leaderboard and 43 fights

The front page combines a ranking with a match history. When checked it listed 43 finished matches, none running, averaging 10.3 rounds each. All models start at an Elo of 1000. claude-sonnet-5 led with 1063 after 21 fights, ahead of grok-4.6 (1030) and gemini-3.6-flash (1029), while deepseek-v4-flash-0731 was last on 890 with no win in 25 attempts.

The Tiny AI Arena start screen with the leaderboard and match history

Clicking any row opens a spectator view in which a match can be replayed: arrow keys step between frames and Space starts autoplay. A fight runs on an 8x8 board with four fighters drawn at random from the pool of models.

How the fight is decided

The project's README sets out the rules. The winner is the last fighter alive. Every round each fighter takes one turn and the turn order is reshuffled. A move goes one cell up, down, left or right, an attack hits an adjacent enemy for 15-24 damage, and waiting is allowed; each action costs one action point. Four randomly placed rocks block movement, the gold power-up adds one point per turn, and a kill grants one more point per turn plus a 50 HP heal.

Fighters may also talk, up to 50 characters per turn, and the last six messages enter every model's prompt. A referee on the server checks each action: an illegal one is wasted but still costs a point, and an unusable reply is retried once before the fighter waits. The code is a monorepo where an Express and SQLite server calls the models through OpenRouter and a Phaser 4 client draws the replay; every action is stored as a frame, so recordings stay exact.

The doubts on Hacker News

The post collected 60 points and 31 comments, and the tone was more critical than celebratory. One commenter asked whether first place for claude-sonnet-5 says anything about intelligence; hp6 answered that from a benchmarking standpoint it is debatable, and that a real signal would need far more games, a way to separate luck from skill and a more complex game, which he said would kill the fun.

Others went after the details. Several users reported that the page would not scroll on Android phones, and the author said it should be fixed. A spectator saw agents stand still while attacked. The dialogue took the hardest hit: lines such as "I'm coming for you, Crimson!" looked to one commenter like proof that models tuned for agentic work produce flat text, though another blamed the 50-character cap. Among the requests were shareable replay links for comparing models.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.