Skip to content
Trending storyDeveloping

Tsinghua's AAArena tests AI coding agents improving game bots from replays

1 article1 sourcesince Oct 11Last article 2h ago ·

Overview

AISummary of 1 article

Tsinghua researchers introduce AAArena, a game-agent benchmark built from 12 games in the university's yearly bot-building contest, with 1,920 archived human programs as rivals.

Coding agents rewrite their bots from match replays without changing model weights. Per Rohan Paul's report of the paper, detailed replays beat win/loss-only feedback in all three games tested, and a Pacman bot reached rank 1 with replays versus rank 11 without.

The paper reports that agents mostly stall on games with complex rules. Tripling the match budget did not lift any of four stuck bots to rank 1, according to the same report.

Written by AI from the articles below · updated Oct 11, 9:16 AM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 11
  1. Rohan PaulX
    Tsinghua study finds AI agents reach top ranks in game bot contest using replays

    AIA Tsinghua paper introduces AAArena, built from 12 games in the university's bot-building contest with 1,920 archived human programs as rivals. Coding agents, with model weights unchanged, rewrite their bots from match replays, and detailed replays beat win/loss feedback in all three games tested. A Pacman bot reached rank 1 with replays versus rank 11 without, but tripling the match budget did not lift any of four stuck bots to rank 1.

    Image from @rohanpaul_ai's post

Heat trend

Not enough continuous observations to show a trend yet.