Skip to content
Trending storyDeveloping

Stanford paper: harness edits fix agent process failures, weight training fixes poor plans

1 article1 sourcesince Oct 10Last article 1h ago ·

Overview

AISummary of 1 article

A Stanford-affiliated paper says agent failures split into process failures, such as loops and exhausted step budgets, and content failures, where the agent delivers a poor plan.

On a travel-planning benchmark, harness evolution raised Qwen3.5-4B's held-out score from 0.16 to 0.30 and plan delivery from 55% to 90%, but did not reduce poor plans. A LoRA adapter cut poor plans for Qwen3.5-9B from 28% to 5% of held-out runs, according to a post by Rohan Paul on X summarizing the paper.

Written by AI from the articles below · updated Oct 10, 4:44 PM ET

Check the sources:

Article timeline

The articles in this story. Times are ET.

Oct 10
  1. Rohan PaulX
    Stanford paper says agent loops need harness fixes and bad plans need weight training

    AIA new Stanford-affiliated paper says agent failures should be sorted into process failures, such as loops and exhausted step budgets, and content failures, where a poor plan is delivered. On a travel-planning benchmark, harness evolution raised Qwen3.5-4B's held-out score from 0.16 to 0.30 and plan delivery from 55% to 90%, but it did not reduce poor plans. A LoRA adapter cut poor plans for Qwen3.5-9B from 28% to 5% of held-out runs.

    Image from @rohanpaul_ai's post

Heat trend

Not enough continuous observations to show a trend yet.