Skip to content
Trending storyDeveloping

Tencent Hunyuan, Fudan and Tsinghua release ExplorationBench to test AI rule discovery

1 article1 sourcesince Oct 8Last article Yesterday ·

Overview

AISummary of 1 article

Tencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that measures whether AI systems can discover rules through experiments in verifiable alien worlds.

According to the release, feedback from experiments raised the best AlienCode score across 10 frontier systems to 89.0% after four rounds, while closed-book runs without feedback stayed between 0.5% and 11.0%.

The researchers say every answer is graded by an interpreter or a proof checker rather than an LLM judge, so scores do not depend on another model's opinion.

Written by AI from the articles below · updated Oct 8, 10:56 PM ET

Check the sources:

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Tencent Hunyuan
    Tencent Hunyuan releases ExplorationBench to test AI rule discovery

    AITencent Hunyuan, with Fudan and Tsinghua researchers, released ExplorationBench, a benchmark that tests whether AI systems can discover rules through experiments in verifiable alien worlds. Across 10 frontier systems, feedback from experiments raised the best AlienCode score to 89.0% after four rounds, while closed-book runs without feedback stayed at 0.5–11.0%.

Heat trend

Not enough continuous observations to show a trend yet.