Skip to content
View original post on X: Tri Dao· 36/100AI score36/100

Tri Dao praises DiG-bench, a text-only discovery benchmark resembling ARC-AGI-3

AISummary

Tri Dao praised DiG-bench, a new text-only benchmark for discovery that resembles ARC-AGI-3 without requiring vision capability. The benchmark, built by researchers from Princeton, MIT, KAUST, and Inria, tests frontier models on text-based discovery games. Their early findings indicate frontier models have improved substantially but still struggle with some surprisingly simple problems.

Post on XView on X
@tri_dao

I really like this new benchmark. Has the flavor of ARC-AGI3 but it’s pure text so you don’t have to worry about the vision capability

James Whittington@jcrwhittington
We’re excited to announce DiG-bench, a new benchmark for discovery! Over the last few weeks we’ve been testing frontier AI models on our novel discovery games and seeing how they score. Each game is a text-based environment, so they probe discovery capabilities in the natural domain of language models, rather than requiring additional, potentially confounding, visual understanding. TL;DR frontier models have improved a lot over the last few months. But they are still stumped by some surprisingly simple problems, even in their native text domain. With @cocosci_lab (@Princeton) @MITCoCoSci (@MIT) @SchmidhuberAI (@KAUST_News) @misovalko (@Inria) @tri_dao (@PrincetonCS) @RMBattleday @zebkDotCom @FraserGreenlee @akaijsa @ClareMaguire @TimMuller1 @kubicek_ales @physicscat0x7d @SukritSumant @thoughtchannel_ (1/5)
View quoted post on X

Source: Tri Dao · x.comPublished · added here