RSIGym: research environment for self-improving agents with RSI-Index benchmark
Overview
Evolvent AI introduces RSIGym, a research environment that gives research agents training, inference, evaluation, and sandboxes as callable services, and an RSI-Index benchmark for measuring how model weights and the agent harness co-improve.
According to Elvis Saravia, who reports the result on X, the approach lets agents spend their budget on experiments rather than rebuilding infrastructure. Saravia says that with Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. These figures come from a single social media post and have not been independently verified in the reporting available.
Written by AI from the articles below · updated Oct 8, 9:05 PM ET
Check the sources:
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Elvis SaraviaRSIGym gives research agents services, lifting SWE-bench Verified to 50.33%
AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.
Heat trend
Not enough continuous observations to show a trend yet.