Skip to contentSkip to stories

Updated

#Video

Showing low-relevance items too. Hide low-relevance items

Sep 24

Sep 24Thu
  1. OdysseyOfficialAI score18

    Odyssey introduces Agora-2, a multi-agent world model

    AIOdyssey launched Agora-2, a multi-agent world model that the company is making available for public experimentation. The post predicts such models will increasingly power applications in AI training, AI safety, robotics, autonomous vehicles, defense, energy, cybersecurity, and gaming.

  2. Google · Gemini appOfficialAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  3. Google Cloud · AI & Machine LearningOfficialAI score55

    Gemini 3.8 Live with Live Avatar becomes generally available in Gemini Enterprise

    AIGoogle says Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, with US and EU endpoints, provisioned throughput, and enterprise compliance. Its video avatars use synchronized lip-syncing, custom avatars are limited to an allowlist, and generated audio and video carry SynthID watermarks. The model also understands and speaks 97 languages and can run tool calls in the background while the conversation continues.

  4. AI at MetaOfficialAI score42

    Meta unveils Muse Realtime Voice and Avatar with shared speech-token streaming

    AIMeta's Muse Realtime Voice generates speech tokens encoding both content and prosody, and Muse Realtime Avatar consumes that shared stream to produce streaming video. Using a fixed-length history as motion context keeps computation bounded regardless of conversation length while synchronizing voice, lip motion, and expressions.

    Image from @AIatMeta's post
  5. AI at MetaOfficialAI score22

    Meta distills 40-step video diffusion into a 2-step live streaming model

    AIMeta distilled a 40-step diffusion teacher using 3-way CFG, requiring 120 evaluations per video chunk, into an unguided 2-step causal student with a fixed-length KV cache. The student uses self-forcing to resist drift and keep near-teacher quality while needing 60x fewer evaluations, enabling instant responses in live video streaming.

    Image from @AIatMeta's post
  6. MiniMax (official)OfficialAI score34

    MiniMax-H3 video generation accelerated on AMD MI355X by Nunchux

    AINunchux runs MiniMax-H3 on AMD MI355X GPUs, generating 5 seconds of video in 1.3 seconds with up to 26.7x faster inference than SGLang on 8 GPUs. The stack supports streaming generation, letting users change prompts while the video plays. Free access to MiniMax-H3 through Nunchux is coming soon, with a waitlist open.

  7. Kling AI BlogOfficialAI score12

    Kling AI outlines six AI video limitations and workarounds for consistency and control

    AIKling AI's blog identifies six limitations of current AI video generation, including temporal consistency, character consistency across shots, unrealistic physics, long-form generation, fine details and text, and prompt control. It recommends workarounds such as reference images, shorter single-action clips, storyboards, and adding text or logos in post. The article says Kling VIDEO 3.0 and VIDEO 3.0 Omni offer reference-based subject consistency to help reduce these problems.

Sep 23

Sep 23Wed
  1. WanOfficialAI score15

    Wan3.0 generates cinematic 1080p video with sound from one prompt

    AIAlibaba Cloud and Venice AI demonstrated Wan3.0 producing a finished scene from a single text prompt, with cinematic motion, 1080p output, and generated sound. The post presents the end-to-end workflow as a sign that the gap between an idea and a finished video is shrinking.

    Video from @Alibaba_Wan's post
  2. Philipp SchmidBlogAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  3. Black Forest LabsOfficialAI score16

    Black Forest Labs links to FLUX 3 action model details

    AIBlack Forest Labs points readers to a dedicated page on its website describing the FLUX 3 action model. The post itself provides no specifics about the model's capabilities, benchmarks, or availability.

  4. Black Forest LabsOfficialAI score40

    FLUX 3 Action uses a smaller architecture to predict actions and frames

    AIBlack Forest Labs says FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3 but uses a smaller architecture. The company attributes this smaller size to more efficient representations learned through its Self-Flow research. During midtraining, the model was trained to predict actions and future frames together.

    Video from @bfl_ai's post
  5. World LabsOfficialAI score27

    World Labs previews Chisel, a world-building feature in Atlas beta

    AIWorld Labs has previewed Chisel, a feature in its Atlas beta that lets users block out a world and have Atlas bring it to life. The post offers no further technical details, pricing, or availability information.

    Video from @theworldlabs's post
  6. Google LabsOfficialAI score29

    Google Labs Releases Six Flow Tools Built by Creatives in Sound, Design, and Content

    AIGoogle Labs released six new Google Flow Tools built by creatives across architecture, sound design, and digital content, including Mondo Sónico, CaptionCast, ThumbnailForge, Surface, CollageMotion Pro, and SwissFlow Studio. Each tool targets a specific workflow, such as generating synchronized audio stems, transcribing and styling captions, or producing animated collages from text prompts. Users can try the tools, duplicate and remix them, or build their own by describing a task in Google Flow.

  7. Comfy BlogOfficialAI score62

    Comfy Router launches one API for frontier image, video, 3D, and audio models

    AIComfy Router is now live on the Comfy Developer Platform, giving developers one API to call frontier image, video, 3D, and audio models. Day one models include Seedance 2.5, MiniMax H3, Nano Banana Pro, GPT Image 2, Kling, and Black Forest Labs, and the provider for each job is selectable. Requests fail rather than silently switching providers, and inputs and outputs are deleted after 24 hours.

    Why it matters: The post shows how one API key and a provider parameter let developers swap routes for media models without rewriting calls, with failed requests reporting the provider.

  8. howie.seriousXAI score18

    Opus 5.5 at medium thinking produces a near-professional promo video

    AIHowie.serious says Opus 5.5 with only medium thinking enabled produced a video of impressive quality, suggesting a grim outlook for human video work. The quoted post says the model made a promo video for the Applore app in a few minutes, and the author calls the result remarkable.

    Video from @howie_serious's post

Sep 22

Sep 22Tue
  1. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

Sep 21

Sep 21Mon
  1. Xiaomi MiMoOfficialAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

    Image from @XiaomiMiMo's post
  2. LumaOfficialAI score13

    Luma shows hybrid production making underwater shoots more accessible

    AILuma says underwater filming usually demands specialized gear, costly set construction, and unique green-screen lighting challenges. It argues hybrid production makes underwater shoots much more accessible. The post is promotional and cites its own video made with Luma.

    Video from @LumaLabsAI's post
  3. Xiaomi MiMoOfficialAI score34

    Xiaomi's MiMo Gallery showcases outputs generated entirely by MiMo-V2.6

    AIXiaomi's MiMo team released MiMo Gallery, a showcase where every 3D model, game, slide deck, video, music piece, and image was generated by MiMo-V2.6. The post previews MiMo-V2.6 as an upcoming model and points readers to the gallery for a first look.

    Video from @XiaomiMiMo's post
  4. MiniMax Design (H3)OfficialAI score22

    Community speeds up MiniMax H3 video generation with sparse attention

    AIA community developer integrated the Jev method into MiniMax H3 to sparsify attention, deciding per layer which parts to keep. On an RTX 4070, video generation time dropped from 6 min 7 sec to 3 min 34 sec, a 41.7% reduction. The post notes that Jev selected sparsity rates of 1%, 3%, 5%, and 10% across 49 layers in 4-step generation.

Sep 20

Sep 20Sun
  1. WanOfficialAI score13

    Wan 3.0 video shows strong scene consistency in a game-style clip

    AIA creator made a game-style video with Wan 3.0 and Nadou Pro, highlighting its quality and especially the consistency between scenes. The main post jokingly suggests the creator may have hit record mid-game, implying the clip is a real gameplay capture.

  2. WanOfficialAI score10

    Wan posts motion control and animation demo clips

    AIWan (@Alibaba_Wan) promoted its motion control and animation capabilities with a short post, calling the results solid. The post includes no specifications, benchmarks, or availability details. A related Wan 3.0 experiment shared by another account shows a combat sport motion demonstration.

  3. WanOfficialAI score7

    Wan 3.0 praised alongside GPT Image 2.5 visual demo

    AIAlibaba's Wan account replied to a post praising Wan 3.0, which was shown producing a video in a Higgsfield demo alongside a GPT Image 2.5 "Sunburst" image. The main post itself only says "nice work!" and gives no specs, benchmarks, or availability details.

  4. WanOfficialAI score12

    Wan Lab offers a new skill for trying Wan video creation

    AIWan has released a skill on its create.wan.video lab platform that users can try now. The post links directly to the skill page and gives no further details on its features, capabilities, or pricing.

    Image from @Alibaba_Wan's post
  5. WanOfficialAI score22

    Wan 3.0 turns a single photo into an animated scene swap

    AIWan 3.0 lets users place their own photo into any scene and animate it using the Peel-Off Sticker skill on wan.video. The post presents this as a single-photo workflow but gives no further technical details.

    Video from @Alibaba_Wan's post

Sep 19

Sep 19Sat
  1. WanOfficialAI score12

    Story maker's OpenArt AI Ad Awards entry made with Wan 3.0

    AIA creator submitted a second entry to the OpenArt AI Ad Awards, made exclusively with Wan 3.0. The post presents the work as a genre-shifting story that showcases the model's range across styles.

Sep 18

Sep 18Fri
  1. Kling AIOfficialAI score20

    Diane Shorthouse Explores AI Filmmaking with Kling AI on MINIBOTS

    AIVeteran film and TV producer Diane Shorthouse is exploring how AI can expand what filmmakers create, sharing her experience making MINIBOTS with Kling AI. The post says AI can support safer production and help bring emotional performances to AI characters.

    Video from @Kling_ai's post
  2. MiniMax Design (H3)OfficialAI score18

    Hailuo AI speeds up storyboarding with 3x3 panel-to-video generation

    AIHailuo AI's post suggests that a single 3x3 storyboard image can be converted into a video using the MiniMax H3 Max r2v model at 480p for 15 seconds. The quoted post describes settings with Quality prompt tuning and standard reference strength, and asks for a 2D animation with panel-to-panel cuts while excluding multiple panels and BGM.

  3. WanOfficialAI score6

    Wan 3.0 raises output resolution by one level

    AIWan 3.0's output resolution has been raised by one level, according to a quoted post from @roco_kn_roco. The main post from Wan (@Alibaba_Wan) contains only emoji and adds no further details.