EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle
Original title🚀 EvolveScaler is here.
AISummary
Tencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.
Source: Tencent Hunyuan · x.comPublished · added here