Xiaomi releases MiMo-V2.6-Pro-MOPD, a 1.02T-parameter sparse MoE model
Original titleXiaomiMiMo/MiMo-V2.6-Pro-MOPD
AISummary
Xiaomi has released MiMo-V2.6-Pro-MOPD, an upgrade of the MiMo-V2.6-Pro-RL checkpoint that fuses several domain-specialized teachers into one model via MOPD2 and targets tool-call repetition.
The sparse MoE model has 1.02T total and 42B activated parameters, a 1M-token context length, and accepts text, image, video, and audio inputs. Weights are available on Hugging Face and ModelScope, with deployment recipes for SGLang and vLLM.
Source: Xiaomi MiMo · new models on Hugging Face · huggingface.coPublished · added here