Skip to content

Areas · Latest news

Multimodal AI

Capabilities beyond text: visual understanding, mixed text and images, and audio-video input and output.

98 top picks all-time · 46 in the past 30 days · chosen from 508 items collected all-time

Latest pick

Top picks archive · Page 5

Top picks 81–98 of 98

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceOfficialAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

May 31

May 31Sun
  1. MiniMax BlogOfficialAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat
  1. Xiaomi MiMoOfficialAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 26

May 26Tue
  1. MiniMax BlogOfficialAI score67

    MiniMax Agent Team Adds Parallel Multi-Agent Collaboration for Long Tasks

    AIMiniMax has upgraded its Agent, renamed Mavis, and introduced Agent Teams that run multiple role-based Agents in parallel on desktop. The team uses Leader, Worker, and Verifier roles so complex tasks can be split, checked, and reported at key checkpoints, and it merges TokenPlan and Agent Plan into one subscription with credits shared between Agent and API. The post also discusses the added token, handoff, and retry costs of multi-Agent work, and says the Agent will be open-sourced alongside MiniMax M3.

    Why it matters: The post explains why multi-Agent helps long tasks and where its verification, token, and aggregation costs come from, useful for judging when a team setup beats a single Agent.

May 10

May 10Sun
  1. Thinking Machines LabOfficialAI score67

    Thinking Machines Lab previews interaction models for real-time human-AI collaboration

    AIThinking Machines Lab announced a research preview of interaction models that take in audio, video, and text continuously and respond in real time without external turn-detection harnesses. The model, TML-Interaction-Small, is a 276B-parameter MoE with 12B active parameters, paired with an asynchronous background model for sustained reasoning and tool use. The post reports competitive intelligence scores and lower turn-taking latency against GPT-realtime and Gemini Live models, along with new interactivity benchmarks where baseline models largely failed.

    Why it matters: The post explains a time-aligned, full-duplex design and benchmarks against turn-based models, showing how interaction and background reasoning can be split across two cooperating models.

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

Apr 21

Apr 21Tue
  1. Xiaomi MiMoOfficialAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceOfficialAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 17

Mar 17Tue
  1. Xiaomi MiMoOfficialAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

Mar 11

Mar 11Wed
  1. Nano Banana 2.1OfficialAI score62

    How to get the most out of Nano Banana 2 for image generation

    AINano Banana 2, also called Gemini 3.1 Flash Image, adds visual grounding with Google Search, 512px resolutions, and extreme aspect ratios of 1:8 and 1:4. The guide advises using it as the default for new projects, with Nano Banana Pro reserved for complex prompts it fails, and keeping Thinking mode off by default.

    Why it matters: The guide compares Nano Banana 1, 2, and Pro with concrete routing advice, which helps developers decide which model to default to and how to control cost.

Mar 4

Mar 4Wed
  1. Mistral AI · new models on Hugging FaceOfficialAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 combines instruct, reasoning, and Devstral capabilities in one multimodal model with 119B total parameters, 6.5B active per token, and a 256k context window. The source reports a 40% reduction in latency-optimized end-to-end completion time and 3x more requests per second in throughput-optimized setups versus Mistral Small 3. It is released under Apache 2.0 and supports reasoning mode toggling per request.

    Why it matters: The source lists architecture, context length, and mode-switching controls, letting readers compare this release's design with earlier Mistral Small models.

Feb 17

Feb 17Tue
  1. Eugene YanXAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.

    Why it matters: The post places Sonnet 4.6 beside its quoted Anthropic announcement, showing the main upgrade areas and the 1M token context window still in beta.

Feb 4

Feb 4Wed
  1. Intern Large ModelsOfficialAI score60

    Intern-S1-Pro: 1T MoE open-source multimodal scientific reasoning model released

    AIIntern Large Models introduces Intern-S1-Pro, a 1T-parameter MoE open-source multimodal scientific reasoning model with 1T-A22B active configuration. The post claims competitive scientific reasoning against leading closed-source models and supports vLLM and SGLang, with weights on Hugging Face and code on GitHub.

    Why it matters: The post pairs a 1T MoE open-source scientific reasoning model with benchmark tables against closed models, letting readers compare claimed strengths directly.

    Image from @intern_lm's post

Jan 23

Jan 23Fri
  1. Mistral AI · new models on Hugging FaceOfficialAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 is a 119B-parameter MoE model with 6.5B active per token and a 256k context window, combining instruct, reasoning, and Devstral-style coding in one model. It accepts text and image input, lets users set reasoning_effort per request, and is released under Apache 2.0. The model card reports a 40% latency reduction and 3x throughput versus Mistral Small 3 in its tested setups, and its benchmark chart shows reasoning scores on GPQA Diamond, MMLU Pro, AIME-style text tasks, and MMMU-Pro.

    Why it matters: The model card names concrete architecture, context, and licensing details, letting readers compare its reasoning toggle and efficiency claims against other open models.

Jan 1

Jan 1Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score75

    Moonshot AI releases open-source multimodal agent model Kimi K2.5

    AIMoonshot AI released Kimi K2.5, an open-source native multimodal agentic model built by continual pretraining on about 15 trillion mixed visual and text tokens. The model card reports a 1T-parameter Mixture-of-Experts architecture with 32B activated parameters and a 256K context length, and it lists benchmark results against GPT-5.2, Claude 4.5 Opus, Gemini 3 Pro, DeepSeek V3.2, and Qwen3-VL-235B-A22B-Thinking. Weights and code are released under a Modified MIT License, with API access on the Moonshot platform.

    Why it matters: The model card gives a full benchmark table against GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro, useful for comparing open multimodal agent models.

Dec 11, 2025

Dec 11, 2025Thu
  1. Runway ResearchOfficialAI score62

    Runway Introduces GWM-1, a Real-Time General World Model Family

    AIRunway announced GWM-1, its first general world model family, built on Gen-4.5 and generating frames autoregressively in real time under interactive control. It comes in three variants: GWM Worlds for explorable environments, GWM Avatars for conversational characters, and GWM Robotics for robotic manipulation. Runway also says it is working toward unifying these domains under a single base world model, and GWM Robotics includes a Python SDK.

    Why it matters: The post separates three GWM-1 variants and ties each to a concrete use, which clarifies where a general world model would fit compared with a single model.