Skip to content
Read the original: OpenBMB (MiniCPM) · new models on Hugging Face· Published 45/100AI score45/100

openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

Original titleopenbmb/JustRL-II-base-model

AISummary

OpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning.

The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline.

The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.

Read the original huggingface.co

Source: OpenBMB (MiniCPM) · new models on Hugging Face · huggingface.coPublished · added here