Skip to content
Read the original: Cognition Blog (Devin, Windsurf)· Published Pick70/100AI score70/100

Cognition Introduces FrontierCode, a Benchmark for Mergeable Code Quality

Original titleIntroducing FrontierCode

AISummary

Cognition introduced FrontierCode, a coding benchmark built with open-source maintainers that measures whether models produce code a maintainer would merge.

On FrontierCode Diamond, the hardest 50 tasks, Claude Opus 4.8 scored 13.4%, GPT-5.5 scored 6.3%, and Gemini 3.1 Pro scored 4.7%.

The authors report 81% fewer misclassification errors than SWE-Bench Pro, though this figure comes from their own analysis of agent trajectories.

AIWhy it matters

The benchmark's blocker and rubric design shows how code quality can be measured beyond unit-test correctness, which matters for judging coding agents.

Read the original cognition.com

Source: Cognition Blog (Devin, Windsurf) · cognition.comPublished · added here