EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes
Original titleEvaluating Agents Beyond the First Prompt
AISummary
EvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction.
The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase.
Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.
Source: Philipp Schmid · philschmid.dePublished · added here