Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks
AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.
Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.