Skip to content
Read the original: Tencent Hunyuan· Published 38/100AI score38/100

EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

Original title🚀 EvolveScaler is here.

AISummary

Tencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

Read the original x.com

Source: Tencent Hunyuan · x.comPublished · added here