Back to list
Oct 02, 2026
111
0
0
ResearchNEW

Ataraxos: Stratego AI Trained for Under $8,000, per Nature

Nature reports Ataraxos won 15 of 20 Stratego games against a four-time world champion, on a training run costing under $8,000 at 2025 prices.

#Ataraxos#Stratego#Nature#Reinforcement Learning#Imperfect Information
Ataraxos: Stratego AI Trained for Under $8,000, per Nature
AI Summary

Nature reports Ataraxos won 15 of 20 Stratego games against a four-time world champion, on a training run costing under $8,000 at 2025 prices.

Introduction

On September 30, 2026, Nature published "Scalable decision-making for games of imperfect information" by Samuel Sokota, Eugene Vinitsky, Hengyuan Hu, Zhiyuan Fan, J. Zico Kolter, and Gabriele Farina. The paper introduces Ataraxos, an AI for the board game Stratego built on general techniques for self-play reinforcement learning and test-time search under hidden information. In a 20-game series over three weeks against Pim Niemeijer, Ataraxos won 15, lost 1, and drew 4, which the paper reports as an 85% effective win rate with draws counted as half wins. The authors call this, "to our knowledge," the first superhuman result in the game's history. Ars Technica's coverage, published October 1, adds context on why the game was hard and why the budget stands out: the paper states the training run cost less than US$8,000 at 2025 prices.

Feature Overview

Why Stratego is hard. Per Ars Technica, each player has 40 pieces, from the marshal down to the spy, plus bombs and a flag, and a piece's identity is revealed only when pieces collide. Farina (MIT) compared it to Texas Hold'em, where two hidden cards allow 1,326 possible hands, while Stratego has "more than a decillion possible setups." He also noted that a chess game usually lasts around 40 moves, while a Stratego game can easily last 2,000.

Three-part design. The paper describes a pattern with three components. First, a policy-value network trained by self-play. Second, a belief network trained on self-play data from the final policy-value network; it models the hidden information and acts as a generative model, sampling realizations of the hidden pieces by likelihood. Third, a test-time search procedure that refines the policy. The paper also describes an update-size schedule with large updates early and small updates later, giving rapid early improvement and continued gains afterward.

Training scale. The reinforcement learning run used 16 NVIDIA H100 GPUs for one week, and the belief network used 4 H100s for 4 days. The RL run covered 163 million finished games and 208 billion environment steps. The team wrote a custom GPU-accelerated Stratego simulator in CUDA C++ that executes millions of state updates per second. According to the paper, all training data were generated from scratch by the agents, with no external data.

Comparison with DeepNash. DeepMind's DeepNash (2022, Science) trained on 1,024 TPU v3 nodes for two to three months, which the paper estimates at US$3.0 million to US$4.5 million at 2025 pricing, and consumed about 5.5 billion games and 5 to 10 trillion training examples. Ataraxos used about 160 million games and about 50 billion examples. The paper summarizes this as roughly 1/500th of the compute cost, 1/30th of the self-play games, and 1/100th of the training examples. Ars Technica reports that bluffing balance stumped earlier systems such as DeepNash, and that DeepMind could not make pre-move search work in Stratego.

Generality beyond one game. The paper reports that the same techniques beat three multi-time world champions at Barrage Stratego, winning each of four series, again described "to our knowledge" as a first superhuman result for that variant. It also reports a new state of the art in Hanabi, a cooperative card game, and wins over the bots PerfectDou and DouZero at dou dizhu.

ResultReported figure
20-game series vs Pim Niemeijer15 wins, 1 loss, 4 draws (85% effective win rate)
2025 World Championship demo (Aug 1-3, 2025)38 wins, 2 losses, 0 draws in 40 games (95% effective win rate)
RL training hardware16 H100 GPUs for 1 week; belief network on 4 H100s for 4 days
Estimated training costUnder US$8,000 at 2025 prices

Usability Analysis

Ataraxos is a research system rather than a product, so usability here means reproducibility and what it reveals about play. The paper states that code for Stratego, Hanabi, and dou dizhu is public at github.com/AtaraxosAI under MIT licenses, and that records of the 20 games against Niemeijer are available at ataraxosai.github.io. Researchers can therefore inspect the approach and replay the series.

The evaluation conditions matter. Niemeijer's accolades, per the paper, include four world championships, 15 Dutch national championships, two online world championships, and more than 600 weeks ranked No. 1. He was paid US$1,000 to take part, plus $100 per win and $50 per draw. Ataraxos did not adapt to him between games, which the paper notes, citing three-time world champion Vincent de Boer, was a handicap for the AI. Ars Technica reported that the bot often tucked its flag in a corner behind just two bombs, a rarely played setup, and Farina said it "skewed the metagame a little bit." Vinitsky described watching the bot "bluff" its way back from a two percent victory probability.

Pros and Cons

On the positive side, the compute and data footprint is small relative to the 2022 system, the design is described in general terms and was reused across three games, and the code is public under MIT licenses. On the limiting side, the system cannot explain its moves; Farina told Ars Technica that building machines with strong and also interpretable strategies is something "we're not quite there yet" on. The headline match was a single 20-game series against one opponent, and the paper's own DeepNash comparison relies on cost estimates at 2025 pricing. As for the one loss, Farina's explanation is that Stratego requires randomizing setups, so luck plays a role; that is the team's claim rather than an independent finding.

Outlook

The authors frame the work as general techniques for hidden-information decision-making. Vinitsky pointed to war gaming as a possible application, saying the derived techniques could be used to play a scenario forward and see how a strong opponent might respond. That is his stated view, not a demonstrated deployment. The near-term items to watch are whether outside groups reproduce the results from the public code, and whether the belief-network-plus-search pattern carries to other settings with long horizons and heavy hidden information, which Vinitsky identified as the distinctive feature of Stratego.

Conclusion

Ataraxos is a well-documented research result: a three-part architecture, a disclosed training budget under US$8,000, public MIT-licensed code, and match records that anyone can check. The superhuman claim is hedged by the authors themselves with "to our knowledge," and the lack of explainability is acknowledged. It is most relevant to researchers in game AI, reinforcement learning, and decision-making under uncertainty. Rating: 4 out of 5.

Editor's Verdict

Ataraxos: Stratego AI Trained for Under $8,000, per Nature earns a solid recommendation within the research space.

The strongest case for paying attention: compute and data footprint is a small fraction of the 2022 DeepNash system, per the paper's own comparison. That alone raises the bar for what readers should expect in this space. Reinforcing that, design pattern is described in general terms and was applied to three different games — practical value rather than just headline appeal. The broader signal worth registering is straightforward: a belief network that samples hidden pieces by likelihood gives search something concrete to work with, which the paper pairs with self-play to handle hidden information. On the other side of the ledger, one constraint is real rather than a marketing footnote: moves cannot be explained by the system, and an author says interpretable strategies are something the team is not quite there on yet. It should factor into any serious decision. Layered on top of that, headline result comes from a single 20-game series against one opponent, so the sample is small — which narrows the set of teams for whom this is an obvious yes.

For ML researchers, technical leads, and readers tracking the underlying science behind new capabilities, this is a serious evaluation candidate, not just a curiosity to bookmark. For everyone else, the safer posture is to monitor coverage and revisit once the use cases that matter to your team are demonstrated in the wild.

Advertisement

Pros

  • Compute and data footprint is a small fraction of the 2022 DeepNash system, per the paper's own comparison.
  • Design pattern is described in general terms and was applied to three different games.
  • Code for Stratego, Hanabi, and dou dizhu is public under MIT licenses, with records of the 20 games online.
  • Results are backed by matches against a multi-time world champion and a public demo with 40 games.

Cons

  • Moves cannot be explained by the system, and an author says interpretable strategies are something the team is not quite there on yet.
  • Headline result comes from a single 20-game series against one opponent, so the sample is small.
  • Cost comparison with DeepNash rests on estimates at 2025 pricing rather than reported budgets.
  • Real-world uses such as war gaming are one author's suggestion and have not been demonstrated.
Advertisement

Comments0

Key Features

1. Architecture: a self-play policy-value network, a belief network that samples hidden pieces by likelihood, and a test-time search procedure that refines the policy. 2. Match result: 15 wins, 1 loss, 4 draws over 20 games against Pim Niemeijer (85% effective win rate, draws counted as half wins). 3. Training cost: 16 H100 GPUs for 1 week plus 4 H100s for 4 days for the belief network; under US$8,000 at 2025 prices. 4. Efficiency vs DeepNash: roughly 1/500th of the compute cost, 1/30th of the self-play games, 1/100th of the training examples, per the paper. 5. Generality: also beat three multi-time world champions at Barrage Stratego, set a new state of the art in Hanabi, and beat PerfectDou and DouZero at dou dizhu; code is MIT-licensed.

Key Insights

  • A belief network that samples hidden pieces by likelihood gives search something concrete to work with, which the paper pairs with self-play to handle hidden information.
  • Compute cost is the sharpest contrast with DeepNash, since the paper reports roughly 1/500th of the cost and 1/30th of the self-play games.
  • Training from scratch on a disclosed budget of under US$8,000 at 2025 prices makes the result easier for outside labs to reason about and reproduce.
  • Test-time search is credited as part of the design, and Ars Technica reports that DeepMind could not make pre-move search work in Stratego.
  • Reuse of the same techniques across Barrage Stratego, Hanabi, and dou dizhu suggests the method is not tuned to a single game.
  • Evaluation conditions favored the human side in one respect, because the AI did not adapt to Niemeijer between games, which the paper notes was a handicap.
  • Unusual play such as a flag tucked behind just two bombs shows the agent departing from established setups, which Farina says skewed the metagame a little.

Was this review helpful?

Share

Twitter/X
Advertisement