Skip to content
Read the original: Philipp Schmid· Published 62/100AI score62/100

EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

Original titleEvaluating Agents Beyond the First Prompt

AISummary

EvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction.

The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase.

Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

Read the original philschmid.de

Source: Philipp Schmid · philschmid.dePublished · added here