Miles runs the full RL loop on Rubin (SGLang rollout, Megatron training, one container image). On one 4-GPU tray:
- Qwen3-30B-A3B on GSM8K: ~45% → ~95% reward over 50 rollouts, matching the GB300 curve.
- DeepSeek-V4-Flash (4-layer, FP8): rollout and training end to end.
- Qwen3.5-35B-A3B agentic RL: 64 concurrent mini-SWE-agent sandboxes on the Vera CPU, on SWE-bench Verified. Reward holds ~0.6, median response length down ~30%.
