Fireworks details numerical mismatch fixes for stable RL training
Original titleTL;DR
AISummary
Fireworks reports that numerical mismatch between training and rollout engines can destabilize reinforcement learning, with a GLM 5.2 experiment showing collapsing reward without alignment and stable reward with it over 25 steps.
The post notes that MoE models add further alignment challenges, as Qwen3.5-MoE differences in expert output combination caused disagreement even when one implementation used higher precision.
Fireworks says it co-develops its trainer and rollout engine to keep frontier RL training aligned across numerics, kernels, and MoEs.
Source: Sophia Yang · x.comPublished · added here