Skip to content
View original post on X: Rohan PaulX· 60/100AI score60/100

Stanford paper says agent loops need harness fixes and bad plans need weight training

AISummary

A new Stanford-affiliated paper says agent failures should be sorted into process failures, such as loops and exhausted step budgets, and content failures, where a poor plan is delivered.

On a travel-planning benchmark, harness evolution raised Qwen3.5-4B's held-out score from 0.16 to 0.30 and plan delivery from 55% to 90%, but it did not reduce poor plans. A LoRA adapter cut poor plans for Qwen3.5-9B from 28% to 5% of held-out runs.

Post on XView on X
Rohan PaulVerified on X
@rohanpaul_ai

New Stanford paper finds that harness changes fix agents that loop or stall, while agents that deliver bad plans need weight training instead.

An agent can be improved by editing its harness, the prompts, tools, and checks around the model, or by fine-tuning its weights.

They sorted failed runs into process failures, such as loops and used-up step budgets, and content failures, where a poor plan was delivered. On a travel-planning benchmark, an LLM-driven loop rewrote the harness, and its best runs then fine-tuned the model.

Harness evolution lifted Qwen3.5-4B from 0.16 to 0.30 on held-out tasks, as plan delivery rose from 55% to 90%. Harness edits never shrank the share of poor plans, but a LoRA adapter cut them from 28% to 5% of Qwen3.5-9B's held-out runs.

Before improving an agent, label why its runs fail, then fix process failures in the harness and content failures in the weights.

Source: Rohan Paul · x.comPublished