Skip to content
Read the original: METR Blog· Published Pick62/100AI score62/100

METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

Original titleSummary of METR's predeployment evaluation of Claude Opus 5.5

AISummary

METR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

AIWhy it matters

The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Read the original metr.org

Source: METR Blog · metr.orgPublished · added here