Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

Jun 9

Jun 9Tue
  1. ByteDance · new models on Hugging FaceAI score28

    ByteDance releases Sa2VA-Qwen3-VL-4B-SAM3 for image and video referring segmentation

    AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.

Jun 8

Jun 8Mon
  1. ByteDance · new models on Hugging FaceAI score46

    ByteDance Open-Sources Bernini-R 1.3B Video Diffusion Renderer on Hugging Face

    AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.

Jun 5

Jun 5Fri

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

  3. ByteDance · new models on Hugging FaceAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. PaddlePaddleAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

    Image from @PaddlePaddle's post

May 31

May 31Sun
  1. MiniMax BlogAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat
  1. Xiaomi MiMoAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. PaddlePaddleAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

    Image from @PaddlePaddle's post

May 27

May 27Wed
  1. Google LabsAI score22

    Google I/O creators discuss human imagination shaping AI creative tools

    AIAt Google I/O, creators behind Flow, Project Genie, and Google Flow Music said human imagination, not the technology itself, shapes new storytelling. Designers Khyati Trehan and Kaloyan blend traditional design knowledge with vibe-coding to build Google Flow Tools, with Trehan saying that if the right tool doesn't exist, she can make it. Google Labs points users to Google Flow, Project Genie, and Google Flow Music at labs.google.

    Image from @GoogleLabs's post

May 26

May 26Tue
  1. MiniMax BlogAI score67

    MiniMax Agent Team Adds Parallel Multi-Agent Collaboration for Long Tasks

    AIMiniMax has upgraded its Agent, renamed Mavis, and introduced Agent Teams that run multiple role-based Agents in parallel on desktop. The team uses Leader, Worker, and Verifier roles so complex tasks can be split, checked, and reported at key checkpoints, and it merges TokenPlan and Agent Plan into one subscription with credits shared between Agent and API. The post also discusses the added token, handoff, and retry costs of multi-Agent work, and says the Agent will be open-sourced alongside MiniMax M3.

    Why it matters: The post explains why multi-Agent helps long tasks and where its verification, token, and aggregation costs come from, useful for judging when a team setup beats a single Agent.

May 15

May 15Fri
  1. Intern Large ModelsAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

    Image from @intern_lm's post

May 13

May 13Wed

May 11

May 11Mon
  1. Soumith ChintalaAI score22

    Thinky previews real-time interaction models for human-AI collaboration

    AISoumith Chintala, a Thinky-linked voice, said the company is at step one of a plan to increase human-AI bandwidth and raise the ceiling of joint intelligence. He shared a preview of interaction models, described as real-time collaborative tools that talk, listen, watch, and think alongside people. A linked Thinking Machines post describes the approach and early results.

  2. Andrej KarpathyAI score34

    Karpathy urges AI outputs shift from text toward HTML and interactive visuals

    AIAndrej Karpathy says asking an LLM to structure its response as HTML and viewing it in a browser works well, and that slideshows have also worked for him. He argues vision is the preferred AI output channel, outlining a progression from raw text and markdown toward HTML and eventually interactive neural videos, while input methods like pointing and gesturing still need improvement.

May 10

May 10Sun
  1. Thinking Machines LabAI score67

    Thinking Machines Lab previews interaction models for real-time human-AI collaboration

    AIThinking Machines Lab announced a research preview of interaction models that take in audio, video, and text continuously and respond in real time without external turn-detection harnesses. The model, TML-Interaction-Small, is a 276B-parameter MoE with 12B active parameters, paired with an asynchronous background model for sustained reasoning and tool use. The post reports competitive intelligence scores and lower turn-taking latency against GPT-realtime and Gemini Live models, along with new interactivity benchmarks where baseline models largely failed.

    Why it matters: The post explains a time-aligned, full-duplex design and benchmarks against turn-based models, showing how interaction and background reasoning can be split across two cooperating models.

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  2. Mistral AI · new models on Hugging FaceAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

Apr 23

Apr 23Thu

Apr 21

Apr 21Tue
  1. Xiaomi MiMoAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

Apr 17

Apr 17Fri

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 6

Apr 6Mon
  1. Z.ai Release NotesAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 25

Mar 25Wed

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 17

Mar 17Tue
  1. Xiaomi MiMoAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

  2. Tri DaoAI score49

    Mamba-3 linear model released, outperforming Mamba-2 and Gated DeltaNet

    AITri Dao announced Mamba-3, which he described as the most powerful linear sequence model to date, as hybrid architectures increasingly rely on strong linear models. The post cites Qwen, Kimi-Linear, and NVIDIA's Nemotron-3 Super as examples of this trend. According to co-author Albert Gu, Mamba-3 shows noticeable performance gains over Mamba-2 and Gated DeltaNet at all sizes while maintaining speed.

Mar 13

Mar 13Fri
  1. Berkeley AI ResearchAI score34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    AIBerkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

Mar 12

Mar 12Thu
  1. Intern Large ModelsAI score47

    InternVL-U: Open-Source 4B Unified Model for Reasoning, Generation, and Editing

    AIInternVL-U is a lightweight 4B unified multimodal model that combines reasoning, generation, and editing in one framework, according to Intern Large Models. The post says it uses unified contextual modeling, modality-specific modular design, and decoupled visual representations to balance performance and efficiency. It reportedly outperforms unified baselines more than 3× its size on text rendering, scientific reasoning, and spatially grounded generation and editing, and is open-source on GitHub and Hugging Face.

    Image from @intern_lm's post

Mar 11

Mar 11Wed

Mar 4

Mar 4Wed
  1. Mistral AI · new models on Hugging FaceAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 combines instruct, reasoning, and Devstral capabilities in one multimodal model with 119B total parameters, 6.5B active per token, and a 256k context window. The source reports a 40% reduction in latency-optimized end-to-end completion time and 3x more requests per second in throughput-optimized setups versus Mistral Small 3. It is released under Apache 2.0 and supports reasoning mode toggling per request.

    Why it matters: The source lists architecture, context length, and mode-switching controls, letting readers compare this release's design with earlier Mistral Small models.

Mar 2

Mar 2Mon

Feb 19

Feb 19Thu
  1. Guillaume Lample @ NeurIPS 2024AI score40

    Mistral releases Voxtral Realtime paper, Apache 2.0 speech model

    AIMistral has published the technical report for Voxtral Realtime, a speech transcription model released under the Apache 2.0 license. The model reportedly achieves state-of-the-art transcription performance at sub-500ms latency. Mistral also launched a Realtime playground in Mistral Studio and made the model available in Hugging Face Transformers.

    Image from @GuillaumeLample's post

Feb 17

Feb 17Tue
  1. Eugene YanAI score72

    Claude Sonnet 4.6 released with upgrades and 1M token context window

    AIAnthropic's Claude Sonnet 4.6 is announced as its most capable Sonnet model, with full upgrades across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. It also features a 1M token context window in beta. The author notes that the model is versatile across classification, coding, computer use, and autonomous agents by adjusting effort and thinking modes.