Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks
Original titleXiaomi MiMo-V2-Omni
AISummary
Xiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding.
The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold.
It also states the model supports over 10 hours of continuous audio understanding.
AIWhy it matters
The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.
Source: Xiaomi MiMo · mimo.xiaomi.comPublished · added here