SPADE uses self-play to generate training environments that improve Qwen3 models
Original titleImport AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye
AISummary
Researchers from several universities introduced SPADE, a framework in which an LLM alternates between generating executable training environments and solving them to generate synthetic training data.
Tested on Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507 using GRPO, SPADE lifted the 30B-A3B model's game suite average to 58.3, 8.1 points above base and 5.3 above the strongest fixed-environment baseline.
The authors note that it cannot push models far beyond the capabilities of the model generating the environments.
Source: Import AI · importai.substack.comPublished · added here