Tencent Hunyuan, Fudan and Tsinghua release ExplorationBench to test AI rule discovery
Overview
Tencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that measures whether AI systems can discover rules through experiments in verifiable alien worlds.
According to the release, feedback from experiments raised the best AlienCode score across 10 frontier systems to 89.0% after four rounds, while closed-book runs without feedback stayed between 0.5% and 11.0%.
The researchers say every answer is graded by an interpreter or a proof checker rather than an LLM judge, so scores do not depend on another model's opinion.
Written by AI from the articles below · updated Oct 8, 10:56 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Tencent HunyuanTencent Hunyuan releases ExplorationBench to test AI rule discovery
AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.
Heat trend
Not enough continuous observations to show a trend yet.