Back
Ataraxos AI beats the best Stratego player of all time, trained on just 16 GPUs
SiTech AI Team3 min read

Ataraxos AI beats the best Stratego player of all time, trained on just 16 GPUs

A team from Carnegie Mellon, MIT, New York University and Stanford built Ataraxos, an AI that beat Pim Niemeijer, arguably the best Stratego player of all time, 15-1 with four draws. Training took 16 GPUs and a few thousand dollars.

Another classic game has fallen to AI. Researchers from Carnegie Mellon, MIT, New York University and Stanford built Ataraxos, which beat Pim Niemeijer, arguably the best Stratego player ever, 15-1 with four draws, trained on just 16 GPUs and a few thousand dollars; the work appears in the journal Nature.

Hidden armies

In Stratego, each player gets 40 pieces, ranks from marshal down to spy plus bombs and a flag; you win by capturing the opponent's flag. The opponent sees where your pieces are, but not what they are: identities are revealed only when two pieces collide. Stratego is thus an imperfect-information game like poker.

The scale differs: Texas Hold'em hides two cards and 1,326 hands; Stratego's 40 pieces can stand in any order, more than a decillion setups. Games also run long: chess takes about 40 moves, Stratego up to 2,000.

Then there is the bluffing: players sometimes move a weak piece as if it were a marshal. Bluff too often and your threats mean nothing; never bluff and you become predictable. Such balance stumped earlier systems, including DeepMind's DeepNash.

Learning to guess

Ataraxos learned by playing against itself, 163 million games in total. Its key innovation was thinking ahead: a second neural network, a belief model, infers the opponent's hidden pieces from their movement, so Ataraxos samples plausible arrangements and plays out candidate moves in each.

The name Ataraxos comes from the ancient Greek word for calm. "For humans, it's very hard when you know a secret to make decisions ignoring that fact. For machines, it's easy," Farina said.

The match

Over three weeks, Niemeijer played 20 online games against Ataraxos, earning $100 per win, and won once. The researchers argue that loss was no flaw: playing Stratego well requires randomizing your setup, so luck always matters.

At the 2025 Stratego World Championship, attendees who challenged the bot fared worse: it won 38 of 40 games. Players were surprised by how often it hid its flag in a corner behind two bombs, a rarely played setup.

The price tag

Ataraxos's biggest feat may be its price tag. DeepNash was trained for two to three months on 1,024 of Google's specialized chips, a run estimated to cost $3 million to $4.5 million at 2025 prices; Ataraxos needed 16 GPUs for a week, plus four more for four days. A simulator written by Samuel Sokota and Farina plays millions of moves per second; Ataraxos used 34 times fewer games than DeepNash but ended up stronger.

The same architecture beat three world champions at Barrage Stratego, mastered the card game Hanabi and beat the best bots at dou dizhu. The team now aims beyond board games: negotiations, financial markets, military conflicts. They are also making the AI more understandable: Ataraxos cannot yet explain its moves.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.