Skip to content
Trending storyDeveloping

Microsoft's CABRA benchmark finds coding agents fail on code comprehension, not edit volume

1 article1 sourcesince Oct 9Last article 3h ago ·

Overview

AISummary of 1 article

Microsoft researchers built CABRA, a framework that generates synthetic coding tasks and varies one difficulty dimension at a time.

They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing. Plain LLMs degraded as tasks grew harder, while agents stayed near-perfect by offloading work to tools such as grep.

On SWE-bench Verified, the count of reading and analysis calls correlated with agent failures at -0.200, compared with -0.159 for lines edited, according to the researchers' reported figures. The source is a secondary social media post summarizing the paper, not the paper itself.

Written by AI from the articles below · updated Oct 9, 6:59 PM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 9
  1. Rohan PaulX
    Microsoft paper finds coding agents struggle more with code understanding than editing

    AIMicrosoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time. Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep. On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

    Image from @rohanpaul_ai's post

Heat trend

Not enough continuous observations to show a trend yet.