Epoch AI finds six frontier models cannot yet automate its own research and operations work
Overview
Epoch AI reports that six AI models, given 11 real tasks from its own operations, cannot yet fully automate its work.
Tasks included graphic design, data insights, and research design; Epoch ran the models at the highest available reasoning settings without intervention, then graded outputs manually against its employee standards. Fable 5.1 and GPT-6 Astra led on average performance, reliably handling well-defined work such as coding and computational analysis.
All models still failed on open-ended judgment, including meeting Epoch's standards, designing informative experiments, and generating diverse ideas. On that basis, Epoch concludes AI cannot yet replace its workers. The findings are presented as the first Epoch Automation Reports, and Epoch says it will expand the task suite and publish updated results as new models are released.
Written by AI from the articles below · updated Oct 8, 9:04 PM ET
Check the sources:
Developments
3 developments
- Oct 8, 1:37 PM ET · 1 articleEpoch AI publishes report on whether AI can automate research-style tasks, to be expanded as models are releasedEpoch AI: Over time, we’ll expand this task suite, retire saturated tasks, and publish updated findings as new models are released.
- Oct 8, 1:36 PM ET · 1 articleEpoch AI introduces Epoch Automation Reports evaluating frontier models on tasks from its own workEpoch AI: Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our…
- Oct 7, 8:00 PM ET · 1 articleEpoch AI evaluates six models on real Epoch work tasks to test automation readinessEpoch AI: Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Epoch AIOver time, we’ll expand this task suite, retire saturated tasks, and publish updated findings as new models are released.
AIRead @kellyhongsn’s full report:
- Epoch AICan AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our…
AI…own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.
- Epoch AIPickEpoch tests six AI models on real Epoch work and finds they cannot yet fully automate it
AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.
Heat trend
Not enough continuous observations to show a trend yet.