Brood War Bench Finds AI Agents Can Win Matches but Not Master Strategy
Codex Astra / xhigh went 18–0 in the new agent benchmark, yet its creator says no tested system moved beyond beginner-level play or reliably coordinated the work needed to build a lasting advantage.
Listen to this story
The audio brief
Story brief
3 key pointsBrood War Bench’s leaderboard shows a sharp distinction between winning individual games and sustaining a coherent strategy. Codex Astra / xhigh went 18–0 and Astra / medium went 16–2, yet benchmark creator Ben Swerdlow judged every tested system beginner-level. Agents often exploited hesitation with early worker harassment, while failing at economic management, continuous production, upgrades, and coordinated...
- 01
Codex Astra / xhigh finished first at 18–0; Astra / medium followed at 16–2.
- 02
Claude Fable placed third at 15–3, while Grok 4.6 / xhigh managed 2–15.
- 03
Probe harassment exposed opponent indecision but did not translate into advanced strategic play.
A clean leaderboard win can make AI agents look more capable than they are. Brood War Bench puts Codex Astra / xhigh at 18 wins and zero losses, but its author says none of the evaluated systems played StarCraft: Brood War beyond a beginner level. The gap is revealing: agents could exploit an opponent’s hesitation, yet struggled to maintain an economy, produce forces steadily, and coordinate their own work.
The newly released benchmark evaluates models through agents playing Brood War, the real-time strategy game. It grew out of creator Ben Swerdlow’s experiment building a version that could only be played through agents, then asking how far those agents could go independently.
A leaderboard with a narrow lesson
Codex Astra / xhigh led the published table, followed by Codex Astra / medium at 16–2. Claude Fable placed third at 15–3. At the other end, Grok 4.6 / xhigh recorded two wins and 15 losses, while its medium setting recorded one win and 16 losses.
Those records establish who performed best in this evaluation, not that any entrant reached advanced competitive play. Swerdlow’s broader conclusion was that every tested model remained at a beginner level, even as the strongest systems won consistently against the other evaluated agents.
Why disruption beat sustained play
The benchmark’s most useful finding may be about how the agents won. Swerdlow observed Codex systems repeatedly using disruption: sending a Probe, a basic Protoss worker, across the map to attack enemy workers or buildings. That tactic could work because opposing agents sometimes spent long stretches deciding how to respond instead of continuing other tasks.
That is a very different capability from managing a full game over time. The same systems were weaker at sustained production and macroeconomic play: they delayed technology upgrades, sent small numbers of basic units into defended bases, and used workers in desperate last stands, according to the evaluation.
The recurring breakdowns
- Short-term disruption was stronger than maintaining production and an economy over time.
- Separate subagents handled the economy, army production, and army control with limited communication.
- Units could be sent into attacks one at a time rather than held for a coordinated push.
The coordination problem inside the agent
Swerdlow observed Codex creating separate subagents for the economy, unit production, and army control. But the agents did not communicate much, leaving an army controller to send each newly created unit forward without awareness of a larger force that other agents were trying to assemble.
The result is less a verdict on one model than a constraint exposed by this particular task. Brood War requires quick choices, but also a shared plan that persists across many choices. This benchmark suggests the tested agents could identify a local opening and punish indecision, while still failing at the slower work of turning many specialized actions into one strategy.
Sources
- bw.swerdlow.devNew benchmark dropped
Reader comments
Newest comments first. Replies stay oldest first.