vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X
Original titleTraining and rollout logprobs matched bit for bit on ROCm. The @RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime. A 200...
AISummary
The RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.
Source: vLLM · x.comPublished · added here