Stanford paper: harness edits fix agent process failures, weight training fixes poor plans
Overview
A Stanford-affiliated paper says agent failures split into process failures, such as loops and exhausted step budgets, and content failures, where the agent delivers a poor plan.
On a travel-planning benchmark, harness evolution raised Qwen3.5-4B's held-out score from 0.16 to 0.30 and plan delivery from 55% to 90%, but did not reduce poor plans. A LoRA adapter cut poor plans for Qwen3.5-9B from 28% to 5% of held-out runs, according to a post by Rohan Paul on X summarizing the paper.
Written by AI from the articles below · updated Oct 10, 4:44 PM ET
Check the sources:
Article timeline
The articles in this story. Times are ET.
Rohan Paul@rohanpaul_aiXStanford paper says agent loops need harness fixes and bad plans need weight trainingAIA new Stanford-affiliated paper says agent failures should be sorted into process failures, such as loops and exhausted step budgets, and content failures, where a poor plan is delivered. On a travel-planning benchmark, harness evolution raised Qwen3.5-4B's held-out score from 0.16 to 0.30 and plan delivery from 55% to 90%, but it did not reduce poor plans. A LoRA adapter cut poor plans for Qwen3.5-9B from 28% to 5% of held-out runs.

Heat trend
Not enough continuous observations to show a trend yet.