HERMES harness lifts GPT-5.6 Sol whole-repo migration score from 6.5% to 31.0%
Overview
A paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis.
The authors report that, with the same model and effort setting, replacing Codex with HERMES raised GPT-5.6 Sol's whole-repository migration score from 6.5% to 31.0%. Across four software engineering benchmarks, they report HERMES beats matched baseline harnesses by 12.4 points on average.
The authors also report that, with strong activation and diagnosis models, Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%. These are the paper's claims as reported by Elvis Saravia (@omarsar0); the figures come from the authors' own evaluation.
AIWritten by AI from the articles below · overview updated Oct 8, 9:06 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Elvis SaraviaHERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%
A paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.
Heat trend
Not enough continuous observations to show a trend yet.