Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality
Original titleIntroducing FrontierCode
AISummary
Cognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge.
On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%.
The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.
AIWhy it matters
The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.
Source: Cognition Blog (Devin, Windsurf) · cognition.comPublished · added here