Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 24

Sep 24Thu
  1. Google DeepMindAI score62

    Google DeepMind adds Live Avatar to Gemini 3.8 Live for enterprise

    AIGoogle DeepMind has launched Gemini 3.8 Live with Live Avatar, which adds near real-time visual presence to its native live dialogue models. The feature is available today in Gemini Enterprise, supports 97 languages with adaptive lip-sync, and allows custom avatars through enterprise allowlisting. All output carries an imperceptible SynthID watermark.

    Why it matters: The post specifies the new avatar capabilities, the Gemini Enterprise access path, and the SynthID watermark, which helps readers judge its enterprise deployment fit.

  2. Google · Gemini appAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  3. Google Cloud · AI & Machine LearningAI score55

    Gemini 3.8 Live with Live Avatar becomes generally available in Gemini Enterprise

    AIGoogle says Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise, with US and EU endpoints, provisioned throughput, and enterprise compliance. Its video avatars use synchronized lip-syncing, custom avatars are limited to an allowlist, and generated audio and video carry SynthID watermarks. The model also understands and speaks 97 languages and can run tool calls in the background while the conversation continues.

  4. ModelScopeAI score38

    Qwen-Image-2.1-Fun-Controlnet-Union adds eight controls and inpainting

    AIModelScope released Qwen-Image-2.1-Fun-Controlnet-Union, a single checkpoint adding eight structural controls, including Canny, Depth, Pose, and Scribble, plus inpainting to Qwen-Image 2.1. Control and inpainting share one branch with 16 injection points across every second Transformer block, keeping the base model frozen and requiring no checkpoint switching. It runs at guidance scale 1.0 with CFG-distilled sampling and prefix KV caching, and is available under the Qwen Research License with base Qwen-Image 2.1 weights required.

  5. Goodfire ResearchAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    AIGoodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

  6. Anthropic ResearchAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

Sep 23

Sep 23Wed
  1. Liquid AI BlogAI score46

    LFM2.5-VL-DSpark speeds up vision-language model decoding on GPUs and edge devices

    AILiquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, delivering decoding throughput gains of up to 2.66× on GPUs and 3.13× on edge devices. The drafter adds about 280M parameters, an 8.9% increase in the deployed model's parameter count, and is available on Hugging Face with support in llama.cpp, SGLang, and MLX-VLM.

  2. Philipp SchmidAI score62

    Gemini 3.8 Flash TTS guide shows how to create and reuse your own voice

    AIGemini 3.8 Flash TTS and Flash-Lite TTS are now available in the Gemini API and AI Studio, with a new feature to replicate your own voice or create one from a sentence. The guide shows recording two clips, one of 15-20 seconds of natural speech and one reading a required consent sentence, then creating a reusable voice ID. It also explains that input text is now spoken word for word, so delivery belongs in speech_metadata.style and short sounds inline.

  3. AnthropicAI score62

    Claude finds a previously unknown enzyme system in bacteriophage DNA

    AIClaude has identified a previously unknown enzyme system in bacteriophage DNA, located beside a long array of repeating DNA that somewhat resembles CRISPR. Anthropic says its function is not yet understood, but only a handful of known systems share its features, all of which can cut, copy, and paste DNA. The source notes that programmable systems like CRISPR have been important to medicine, but more work is needed to learn what this system does and whether it can be used similarly.

  4. Google AI StudioAI score62

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle introduces Gemini 3.8 Flash TTS for creative voice design and Gemini 3.8 Flash-Lite TTS for high-volume, cost-efficient speech generation. Flash TTS supports voice creation from natural language prompts across more than 100 languages and dialects, and both models are rolling out today in the Gemini API and Google AI Studio, with enterprise access coming soon via Gemini Enterprise.

  5. Google DeepMindAI score60

    Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS models

    AIGoogle DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models offering custom voice design, line-by-line performance control, and multilingual support across more than 100 languages. Flash TTS is rolling out to developers in the Gemini API and Google AI Studio and to everyone in Gemini Notebook, while Flash-Lite TTS is available to developers and in Google Vids. Voice replication requires consent verification, and generated audio carries SynthID watermarking.

    Why it matters: The source details the voice design, performance direction, and consent safeguards, showing how the model covers creative and high-volume use cases with access across several Google products.

  6. ModelScopeAI score44

    NVIDIA releases Nemotron 3 Diarization for live speaker attribution

    AINVIDIA's Nemotron 3 Diarization is now available on ModelScope, labeling speakers and timestamps in streaming audio for up to eight speaker slots per conversation. The 99.2M-parameter model uses an end-to-end streaming architecture built on NVIDIA's Streaming Sortformer, running on Ampere, Hopper, and Blackwell GPUs via NeMo Speech C++. It is designed to pair with existing ASR systems such as Nemotron ASR, Parakeet, Canary, or Whisper to produce speaker-attributed transcripts.

  7. ModelScopeAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

  8. ModelScopeAI score62

    Xiaomi MiMo-V2.6 open-sourced as a multimodal agent model family under MIT License

    AIXiaomi has released MiMo-V2.6 as an open model family under the MIT License, designed for large-scale reinforcement learning. MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, with 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. The 1.02T-parameter MoE activates 42B parameters and supports text, image, video, and audio input with a 1M-token context.

Sep 22

Sep 22Tue

Sep 21

Sep 21Mon
  1. Xiaomi MiMoAI score36

    Xiaomi MiMo-V2.6 unifies code, design, and tool use across creative outputs

    AIXiaomi's MiMo-V2.6 combines code, design, and tool use to build frontend interfaces, presentations, Figma-linked visual assets, and videos. The post says MiMo-V2.5-TTS supports narration in video production, and that the model can compose music, including an orchestral piece for around ten instruments that can be converted to MIDI. On Design Arena, the Pro version reportedly performs comparably to Claude Opus 5 and GPT-5.6 Sol.

  2. Xiaomi MiMoAI score44

    MiMo-V2.6 builds and interacts with 3D worlds from text, images, or video

    AIXiaomi's MiMo-V2.6 combines 3D spatial reasoning, multimodal perception, and computer use to turn text, images, or video into playable 3D worlds. The model coordinates agents to build scenes, write interaction logic, and refine results, and can create Blender objects for animation, 3D printing, and games. It also controls a Franka Panda arm in simulation via visual feedback and uses desktop tools to process data, inspecting results to adjust its next actions.

  3. Xiaomi MiMoAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

  4. Apple · new models on Hugging FaceAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.

  5. Xiaomi MiMo · new models on Hugging FaceAI score67

    Xiaomi releases MiMo-V2.6-Flash-RL, a 309B sparse MoE model with 1M context

    AIXiaomi released MiMo-V2.6-Flash-RL, an efficiency-balanced checkpoint in its MiMo-V2.6 series, on Hugging Face. The model is a sparse MoE with 309B total and 15B activated parameters, supports text, image, video, and audio input, and offers a 1M-token context. The technical report says it was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs its benchmark tables with the RL training method, which helps readers judge how the checkpoint's scores relate to its training approach.

  6. Xiaomi MiMo · new models on Hugging FaceAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

Sep 20

Sep 20Sun
  1. OpenBMBAI score44

    MiniCPM-o Booking Desk: open-source real-time voice appointment agent built on MiniCPM-o 4.5

    AIDeveloper @mrgoodmantweets built MiniCPM-o Booking Desk, an open-source appointment booking agent that uses MiniCPM-o 4.5 for real-time, full-duplex voice and audio-visual interaction. The agent listens, speaks, and reads live booking status from an operator screen, while deterministic state control keeps execution reliable. An appointment is only booked after user confirmation.