Miles brings Day-0 RL support to DeepSeek-V4.1-Flash.
Miles keeps the trainer close to what @sgl_project samples. Parallelism and shared state let the new architecture scale intact across GPUs. Quantization-aware training mirrors SGLang's FP4/FP8 rounding, while routing replay reuses the rollout's expert choices. Numerical consistency comes from FP32 and deterministic reductions. Colocated training and rollout fit full-parameter RL on 16 GPUs.
Over steps 0–80 of a DAPO run, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.
Blog and cookbook in the comments. 📚

