Skip to content
View original post on X: Rohan PaulX· 57/100AI score57/100

Microsoft paper finds coding agents struggle more with code understanding than editing

AISummary

Microsoft researchers introduce CABRA, a framework that generates synthetic coding tasks with one difficulty dimension varied at a time.

Across 6,840 tasks, plain LLMs degraded as tasks grew, while agents stayed near-perfect by offloading work to tools such as grep.

On SWE-bench Verified, counts of reading and analysis calls correlated with agent failures at -0.200, versus -0.159 for lines edited.

Post on XView on X
Rohan PaulVerified on X
@rohanpaul_ai

New Microsoft paper finds that coding agents trip up when they have to understand a lot of code, not when they have to edit a lot of it, so test them on reading and comparing code instead of diff size.

Microsoft researchers built CABRA, which generates synthetic coding tasks and raises 1 kind of difficulty at a time. They ran 8 LLMs and 6 agents on 6,840 tasks and labeled each tool call as reading, analyzing, searching, editing, or testing.

Plain LLMs got worse as tasks grew, but agents stayed near-perfect by using tools like grep. On SWE-bench Verified, the count of reading and analysis calls tracked agent failures better than lines edited, with correlations of -0.200 versus -0.159.

Source: Rohan Paul · x.comPublished · added here