Skip to content
Trending storyDeveloping

Epoch AI finds six frontier models cannot yet automate its own research and operations work

3 articles2 sourcesLast article 12h ago ·

Overview

AISummary of 3 articles

Epoch AI reports that six AI models, given 11 real tasks from its own operations, cannot yet fully automate its work.

Tasks included graphic design, data insights, and research design; Epoch ran the models at the highest available reasoning settings without intervention, then graded outputs manually against its employee standards. Fable 5.1 and GPT-6 Astra led on average performance, reliably handling well-defined work such as coding and computational analysis.

All models still failed on open-ended judgment, including meeting Epoch's standards, designing informative experiments, and generating diverse ideas. On that basis, Epoch concludes AI cannot yet replace its workers. The findings are presented as the first Epoch Automation Reports, and Epoch says it will expand the task suite and publish updated results as new models are released.

Written by AI from the articles below · updated Oct 8, 9:04 PM ET

Check the sources:

Developments

3 developments

  1. Oct 8, 1:37 PM ET · 1 article
    Epoch AI publishes report on whether AI can automate research-style tasks, to be expanded as models are released
    Epoch AI: Over time, we’ll expand this task suite, retire saturated tasks, and publish updated findings as new models are released.
  2. Oct 8, 1:36 PM ET · 1 article
    Epoch AI introduces Epoch Automation Reports evaluating frontier models on tasks from its own work
    Epoch AI: Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our…
  3. Oct 7, 8:00 PM ET · 1 article
    Epoch AI evaluates six models on real Epoch work tasks to test automation readiness
    Epoch AI: Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

Article timeline

Follow the coverage from different perspectives. Times are ET.

Oct 8
  1. Epoch AI
    Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our…

    AI…own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.

Oct 7
  1. Epoch AIPick
    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

Heat trend

Not enough continuous observations to show a trend yet.