Brood War Bench
Summary
The article reports a Brood War benchmark where AI agents from multiple models play StarCraft: Brood War, revealing Codex Astra as the leader and Grok as less capable. It provides a leaderboard, observations on how agents think, plan, and coordinate subagents, and discusses ongoing AI benchmarking in RTS tasks. Useful for readers tracking AI tool capabilities and real-world benchmarking approaches.