Anthropic shows a three-agent harness for long-running app development
Harness design for long-running application development
Anthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.
The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.
Source: Anthropic Engineering · anthropic.com