Skip to content
Read the original: Import AI· Published 46/100AI score46/100

SPADE uses self-play to generate training environments that improve Qwen3 models

Original titleImport AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

AISummary

Researchers from several universities introduced SPADE, a framework in which an LLM alternates between generating executable training environments and solving them to generate synthetic training data.

Tested on Qwen3-4B-Instruct-2507, Qwen3-8B, and Qwen3-30B-A3B-Instruct-2507 using GRPO, SPADE lifted the 30B-A3B model's game suite average to 58.3, 8.1 points above base and 5.3 above the strongest fixed-environment baseline.

The authors note that it cannot push models far beyond the capabilities of the model generating the environments.

Read the original importai.substack.com

Source: Import AI · importai.substack.comPublished · added here