OpenAI's Mark Chen says GPT-5.5 performs like Mythos in UK AISI cyber range
Original titleThis is just one eval, but it's an important one - UK AISI’s cyber range tests long-horizon, agentic capability. 5.5 performs similarly t...
AISummary
Mark Chen says GPT-5.5 performs similarly to Mythos on UK AISI's cyber range, which tests long-horizon, agentic capability, and calls it one eval rather than a full picture.
He adds that frontier model risks are real and that OpenAI aims to deploy AI people can actually use through mitigations.
The attached chart shows completed steps per cumulative token spent for GPT-5.5, Mythos Preview, and several Claude and GPT models, from M1 reconnaissance up to M9 full network takeover.
Source: Mark Chen · x.comPublished · added here