Ai Engineering 4 min read

AI Beats the World's Best Stratego Player on a Few Thousand Dollars of Compute

Ataraxos, built by CMU, MIT, NYU and Stanford researchers, beat four-time Stratego world champion Pim Niemeijer 15-1 in games published in Nature this October, trained for a few thousand dollars where DeepMind's 2022 attempt cost millions.

An AI system called Ataraxos beat Pim Niemeijer, the most dominant Stratego player in history, 15 games to 1 (with four draws) across three weeks of online play, and the result was published in Nature this October. The team spans Carnegie Mellon, MIT, NYU, and Stanford, with lead author Samuel Sokota alongside Eugene Vinitsky and MIT’s Gabriele Farina. The number that rewrites the economics: Ataraxos trained on 16 GPUs for a week plus four GPUs for four days, a few thousand dollars. DeepMind’s DeepNash, which cracked the same game in 2022, reportedly cost $3 to 4.5 million using over a thousand specialized chips for two to three months. The frontier moved by roughly three orders of magnitude in cost.

Why Stratego Was the Hard One

Chess and Go are perfect-information games; Stratego is the classic imperfect-information benchmark, and it is brutal in specific ways. Each player’s 40 pieces have hidden identities, giving over a decillion possible initial setups, games can run to roughly 2,000 moves, and bluffing is core strategy. DeepNash solved it with massive self-play plus a game-theoretic equilibrium strategy. Ataraxos got there with two ideas. First, self-play reinforcement learning over 163 million games, with the observation that hidden information makes self-play unstable, large strategic swings early and smaller adjustments later. Second, the innovation that made the budget possible: a separate belief model, a neural network that infers opponent piece identities from how they move, then samples plausible arrangements and plays out candidate moves instead of searching every possibility. That pre-move search is what DeepMind could not make work.

What the Numbers Say

Against human opposition, the record is one-sided: at the 2025 World Championship, Ataraxos won 38 of 40 games against challengers, and it learned roughly 34 times faster than DeepNash. The single loss to Niemeijer was attributed to the luck inherent in Stratego’s piece randomization rather than a capability gap. The behavioral details are what make the result feel like a milestone rather than a benchmark: the bot’s name comes from the Greek word for untroubled, and Vinitsky reported watching it “bluff its way back from like a two percent victory probability.” It even shifted the human metagame of the world’s most competitive Stratego scene. The architecture generalized beyond one game: the same approach mastered Barrage Stratego, Hanabi, and dou dizhu, all imperfect-information games.

Why Imperfect Information Is the Real Frontier

The results that made headlines this decade, chess, Go, and now mathematics, are mostly perfect-information problems, where the AI can in principle see everything the rules allow. Stratego, poker, negotiation, and markets share the property that defines most real decision-making: the opponent’s state is hidden, information is revealed through behavior, and optimal play requires reasoning about what another mind knows. Farina’s summary of the core difficulty: “For humans, it’s very hard when you know a secret to make decisions ignoring the fact that you know that secret.” A belief model that infers hidden state from behavior, then reasons over sampled possibilities, is the architecture that transfers to negotiation, market modeling, and, the team’s explicit mention, military wargaming, which is a phrase that will get attention in exactly the rooms the FTC’s new AI probe is watching.

What to Watch

Three things. First, the interpretability work the team says is next: Farina concedes “we’re not quite there yet” on explaining the strategy, and an unbeatable game-playing system that cannot explain its bluffs is a soft demonstration of the interpretability gap in agentic AI generally. Second, the cost curve: a three-orders-of-magnitude training-cost drop in four years means every remaining expensive research benchmark should be re-priced, and the same belief-model architecture will be tried on everything from auctions to adversarial security. Third, the human response: Niemeijer’s reaction and whether Stratego’s federation changes tournament rules for bot-assisted play, because the metagame shift has already started. The pattern from chess is repeating, faster and cheaper each time.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading