Opus 5 and GPT-6 Astra beat Montezuma's Revenge; Metaculus to recreate Turing-style test
Overview
Ethan Mollick says Anthropic's Opus 5 beat Montezuma's Revenge, as OpenAI's GPT-6 Astra did.
He notes that one criterion in Metaculus's AGI bet requires AI to win a discontinued, weak Turing-style prize. Metaculus has decided to recreate that test to confirm whether the criterion is resolved. The report is a post by Mollick on X and does not include the test results or Metaculus's own statement.
Written by AI from the articles below · updated Oct 9, 10:42 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Ethan Mollick@emollickXOpus 5 also beats Montezuma's Revenge, and Metaculus recreates a testAIEthan Mollick says Anthropic's Opus 5 beat Montezuma's Revenge, as well as OpenAI's GPT-6 Astra. He notes that one criterion in the AGI bet is for AI to win a discontinued, weak Turing-style prize, and Metaculus has decided to recreate that test to confirm whether the criterion is resolved.
Heat trend
Not enough continuous observations to show a trend yet.