Skip to content
Read the original: Clément Delangue· Published 62/100AI score62/100

Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

Original titleWe turned Claude Code, Codex, Hermes, Pi, @opencode and other coding harnesses into RL environments. No changes to the harnesses, no chan...

AISummary

Hugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training.

On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

Read the original x.com

Source: Clément Delangue · x.comPublished · added here