Skip to contentSkip to stories
Updated

#Video

Sep 1

  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score60

    Shanghai AI Lab releases Intern Lumina U2 unified multimodal model on Hugging Face

    AIShanghai AI Lab's InternLM has published Intern Lumina U2, a 16B-parameter MoE model with 1B active parameters that handles text QA, image generation and editing, and image, video, and 3D understanding. The model uses an 8-codebook fully-discrete visual representation built on AToken. Checkpoints are provided for Huawei Ascend NPUs and NVIDIA GPUs under Apache 2.0, with the technical report still listed as coming soon.

    Why it matters: The model unifies text, image, video, and 3D understanding with image generation in one framework, a broader scope than single-modality releases.

Jul 30

  1. MiniMax BlogAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.

Jul 28

  1. MiniMax · new models on Hugging FaceAI score76

    MiniMax H3 releases open-weight omni-modal video model with native stereo audio

    AIMiniMax released H3, an open-weights omni-modal model that generates video with native stereo audio up to 2K and 15 seconds. The system combines H3-Context-IR preprocessing, the H3-Base generator at 768p, and H3-Regenerate-2K for 2K output, with the Context-IR and 2K modules available only through API.

    Why it matters: The source details a three-module pipeline and open weights with deployment paths, showing how a video model is served and reproduced locally.

Jul 7

  1. Meta AI BlogAI score75

    Meta launches Muse Image, an agentic image model with search and code tools

    AIMeta Superintelligence Labs has released Muse Image, which can invoke search and coding tools and self-refine its generations before output. It is available today in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook coming soon. Meta also previewed Muse Video, which is coming soon to creators and Meta AI and is reported as ranking No. 3 on Arena for text-to-video at the time of writing.

    Why it matters: The source describes how search, code execution, and self-refinement change image generation, which matters to anyone comparing agentic media models with plain prompt-to-image systems.

Mar 17

  1. Xiaomi MiMoAI score71

    Xiaomi releases MiMo-V2-Omni, an omni-modal model for agentic tasks

    AIXiaomi introduces MiMo-V2-Omni, a single model that fuses image, video, and audio encoders into a shared backbone with native tool calling and UI grounding. The company reports benchmark results against Gemini 3 Pro, Claude Opus 4.6, and GPT 5.2, and demonstrates browser-based shopping and video-publishing workflows run through the OpenClaw agent scaffold. It also states the model supports over 10 hours of continuous audio understanding.

    Why it matters: The page gives benchmark comparisons, a driving-risk demo, and browser-task walkthroughs, letting readers check how far the omni-modal claims extend into agent use.

Dec 11, 2025

  1. Runway ResearchAI score62

    Runway Introduces GWM-1, a Real-Time General World Model Family

    AIRunway announced GWM-1, its first general world model family, built on Gen-4.5 and generating frames autoregressively in real time under interactive control. It comes in three variants: GWM Worlds for explorable environments, GWM Avatars for conversational characters, and GWM Robotics for robotic manipulation. Runway also says it is working toward unifying these domains under a single base world model, and GWM Robotics includes a Python SDK.

    Why it matters: The post separates three GWM-1 variants and ties each to a concrete use, which clarifies where a general world model would fit compared with a single model.

Nov 1, 2025

  1. Runway ResearchAI score72

    Runway releases Gen-4.5, ranked first on the Text-to-Video benchmark

    AIRunway announced Gen-4.5, a video generation model that it says holds the top position on the Artificial Analysis Text-to-Video benchmark with 1,247 Elo points. The model is available across all paid Runway plans at comparable pricing, and the post lists limitations including causal reasoning errors, object permanence failures, and success bias.

    Why it matters: The post separates Runway's own ranking claim from the listed limitations, such as causal reasoning and object permanence errors, which helps judge where the model is reliable.

That’s everything