Diffusion Reward Models learn full human preference distributions, not single scores
Original titleA reward of “3” can mean two completely different things.
AISummary
OpenBMB introduces Diffusion Reward Models (DRM), which learn the full reward distribution of human preferences instead of collapsing them into one scalar score.
The approach preserves disagreement and uncertainty, enabling distribution-aware Best-of-N ranking and a new test-time scaling axis by sampling more reward outputs.
DRM also improves downstream policy performance over scalar reward baselines when used as the reward in RLHF, according to the post.
Source: OpenBMB · x.comPublished · added here