Fireworks explains how numerical mismatch and MoE routing can derail RL training
Original titleReinforcement learning: Why alignment of numerics and MoE routing matter
AISummary
Numerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical.
In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps.
A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.
Source: Fireworks AI Blog · fireworks.aiPublished · added here