Epoch finds 23 false negatives in DeepSWE, capping its score near 79.6%
Original titleSuper cool work! This also explains why some benches have an "artificial ceiling".
AISummary
Epoch AI found 23 false negatives among 131 DeepSWE tasks, implying a ceiling of roughly 79.6%. This suggests the current near-saturated score of about 74% is constrained by artificial benchmark limits rather than model capability.
Source: wh · x.comPublished · added here