Brood War Bench Finds AI Agents Can Win Matches but Not Master Strategy

Codex Astra / xhigh went 18–0 in the new agent benchmark, yet its creator says no tested system moved beyond beginner-level play or reliably coordinated the work needed to build a lasting advantage.

By 3 min read
Brood War Bench Finds AI Agents Can Win Matches but Not Master Strategy
Brood War Bench Finds AI Agents Can Win Matches but Not Master Strategy

Listen to this story

The audio brief

About 1:26
0:001:26
Read transcript
Codex Astra at ex-high went eighteen and zero in a new StarCraft: Brood War benchmark, but the result comes with a sharp qualification: its creator, Ben Swerdlow, says every system tested still played at a beginner level. The benchmark measures agents against other evaluated agents, not against top human competitors. That distinction explains how a perfect record can coexist with weak strategy. The strongest systems were good at spotting hesitation and exploiting it. Codex repeatedly sent a Probe, a basic Protoss worker, across the map to attack enemy workers or buildings. When an opposing agent paused to decide what to do, that disruption could win the game. But the same systems struggled with the slower work that creates a lasting advantage: maintaining the economy, producing units continuously, researching upgrades, and combining forces for a coordinated attack. Codex Astra at medium finished second, with sixteen wins and two losses. Claude Fable placed third at fifteen and three. Grok four point six at ex-high managed just two wins and fifteen losses. Swerdlow also found a coordination problem inside the agents themselves. Separate subagents handled the economy, production, and army control, but communicated too little. Units were often sent forward one at a time instead of as a planned push. The key constraint to watch is whether future systems can turn many specialized actions into one shared strategy that persists beyond the first opening.

Story brief

3 key points

Brood War Bench’s leaderboard shows a sharp distinction between winning individual games and sustaining a coherent strategy. Codex Astra / xhigh went 18–0 and Astra / medium went 16–2, yet benchmark creator Ben Swerdlow judged every tested system beginner-level. Agents often exploited hesitation with early worker harassment, while failing at economic management, continuous production, upgrades, and coordinated...

  1. 01

    Codex Astra / xhigh finished first at 18–0; Astra / medium followed at 16–2.

  2. 02

    Claude Fable placed third at 15–3, while Grok 4.6 / xhigh managed 2–15.

  3. 03

    Probe harassment exposed opponent indecision but did not translate into advanced strategic play.

A clean leaderboard win can make AI agents look more capable than they are. Brood War Bench puts Codex Astra / xhigh at 18 wins and zero losses, but its author says none of the evaluated systems played StarCraft: Brood War beyond a beginner level. The gap is revealing: agents could exploit an opponent’s hesitation, yet struggled to maintain an economy, produce forces steadily, and coordinate their own work.

The newly released benchmark evaluates models through agents playing Brood War, the real-time strategy game. It grew out of creator Ben Swerdlow’s experiment building a version that could only be played through agents, then asking how far those agents could go independently.

A leaderboard with a narrow lesson

Codex Astra / xhigh led the published table, followed by Codex Astra / medium at 16–2. Claude Fable placed third at 15–3. At the other end, Grok 4.6 / xhigh recorded two wins and 15 losses, while its medium setting recorded one win and 16 losses.

Those records establish who performed best in this evaluation, not that any entrant reached advanced competitive play. Swerdlow’s broader conclusion was that every tested model remained at a beginner level, even as the strongest systems won consistently against the other evaluated agents.

Why disruption beat sustained play

The benchmark’s most useful finding may be about how the agents won. Swerdlow observed Codex systems repeatedly using disruption: sending a Probe, a basic Protoss worker, across the map to attack enemy workers or buildings. That tactic could work because opposing agents sometimes spent long stretches deciding how to respond instead of continuing other tasks.

That is a very different capability from managing a full game over time. The same systems were weaker at sustained production and macroeconomic play: they delayed technology upgrades, sent small numbers of basic units into defended bases, and used workers in desperate last stands, according to the evaluation.

The recurring breakdowns

  • Short-term disruption was stronger than maintaining production and an economy over time.
  • Separate subagents handled the economy, army production, and army control with limited communication.
  • Units could be sent into attacks one at a time rather than held for a coordinated push.

The coordination problem inside the agent

Swerdlow observed Codex creating separate subagents for the economy, unit production, and army control. But the agents did not communicate much, leaving an army controller to send each newly created unit forward without awareness of a larger force that other agents were trying to assemble.

The result is less a verdict on one model than a constraint exposed by this particular task. Brood War requires quick choices, but also a shared plan that persists across many choices. This benchmark suggests the tested agents could identify a local opening and punish indecision, while still failing at the slower work of turning many specialized actions into one strategy.

Sources

  1. bw.swerdlow.devNew benchmark dropped

Loading discussion...

YOUR READING SPACE

Notifications

Brood War Bench Finds AI Agents Can Win Matches but Not Master Strategy | Superpower Daily