Skip to content
View original post on X: RadixArk· 60/100AI score60/100

Miles adds day-0 RL support for DeepSeek-V4.1-Flash

AISummary

RadixArk says Miles brings day-0 RL support to DeepSeek-V4.1-Flash, with SGLang providing inference support. The post says quantization-aware training mirrors SGLang's FP4/FP8 rounding, and that colocated training and rollout fit full-parameter RL on 16 GPUs. In a DAPO run over steps 0–80, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.

Post on XView on X
@radixark

Miles brings Day-0 RL support to DeepSeek-V4.1-Flash.

Miles keeps the trainer close to what @sgl_project samples. Parallelism and shared state let the new architecture scale intact across GPUs. Quantization-aware training mirrors SGLang's FP4/FP8 rounding, while routing replay reuses the rollout's expert choices. Numerical consistency comes from FP32 and deterministic reductions. Colocated training and rollout fit full-parameter RL on 16 GPUs.

Over steps 0–80 of a DAPO run, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.

Blog and cookbook in the comments. 📚

SGLang@sgl_project
DeepSeek V4.1 Flash weights are out! We are shipping day-0 inference and RL support in SGLang and Miles. V4.1 extends the V4 stack with compressed KV shared across layers, a two-stage sparse indexer, and a 196B Engram lookup memory. It is natively multimodal with 552B backbone parameters, 16B active decode and 8B prefill, and supports up to 1M context. There will be some very exciting performance upgrades for this model in the next few days. Stay tuned! Supported features, blog and cookbook in the comments ⬇️
View quoted post on X

Source: RadixArk · x.comPublished · added here