Tsinghua's AAArena tests AI coding agents improving game bots from replays
Overview
Tsinghua researchers introduce AAArena, a game-agent benchmark built from 12 games in the university's yearly bot-building contest, with 1,920 archived human programs as rivals.
Coding agents rewrite their bots from match replays without changing model weights. Per Rohan Paul's report of the paper, detailed replays beat win/loss-only feedback in all three games tested, and a Pacman bot reached rank 1 with replays versus rank 11 without.
The paper reports that agents mostly stall on games with complex rules. Tripling the match budget did not lift any of four stuck bots to rank 1, according to the same report.
Written by AI from the articles below · updated Oct 11, 9:16 AM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXTsinghua study finds AI agents reach top ranks in game bot contest using replaysAIA Tsinghua paper introduces AAArena, built from 12 games in the university's bot-building contest with 1,920 archived human programs as rivals. Coding agents, with model weights unchanged, rewrite their bots from match replays, and detailed replays beat win/loss feedback in all three games tested. A Pacman bot reached rank 1 with replays versus rank 11 without, but tripling the match budget did not lift any of four stuck bots to rank 1.

Heat trend
Not enough continuous observations to show a trend yet.