Skip to content
Read the original: wh· Published 38/100AI score38/100

Epoch finds 23 false negatives in DeepSWE, capping its score near 79.6%

Original titleSuper cool work! This also explains why some benches have an "artificial ceiling".

AISummary

Epoch AI found 23 false negatives among 131 DeepSWE tasks, implying a ceiling of roughly 79.6%. This suggests the current near-saturated score of about 74% is constrained by artificial benchmark limits rather than model capability.

Read the original x.com

Source: wh · x.comPublished · added here