Skip to contentSkip to stories
Updated

#Multimodal

Oct 8

Oct 8Thu
  1. Xiaomi MiMoOfficialAI score63

    Xiaomi releases MiMo-V2.5-TTS series of speech synthesis models

    AIXiaomi released the MiMo-V2.5-TTS Series, three speech synthesis models for stock voices, voice design, and voice cloning. The models accept natural-language style instructions and inline audio tags, and the source says the three models are free of charge for a limited time on the Xiaomi MiMo API platform. Xiaomi also open-sourced integration Skills for agent applications on GitHub.

    Why it matters: The release shows how a TTS family adds style instructions, inline audio tags, and voice design or cloning to speech synthesis, which matters for agent and creative workflows.

  2. Sundar PichaiXAI score65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

    Video from @sundarpichai's post
  3. Anthropic ResearchOfficialAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    AIJohns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    Why it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

Oct 7

Oct 7Wed
  1. Aravind SrinivasXAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

    Why it matters: Two open-weight multi-vector models share one space for text and images, with a 0.6B on-device option, a useful comparison for building multimodal retrieval.

  2. Google DeepMind · The KeywordOfficialAI score62

    Google expands SynthID Detector globally to check AI-generated media

    AIGoogle is making its SynthID Detector available globally in English, letting anyone check whether an image, video, or audio file was made with AI from Google or partners including OpenAI, NVIDIA, Kakao, and soon Apple. The tool joins built-in verification in Search, the Gemini app, and Chrome, which now handle over 1 million requests daily. Google says SynthID has watermarked over 180 billion images and videos and 240,000 years of audio.

    Why it matters: The source specifies which vendors' AI media the detector checks, helping readers judge how far the verification covers content they encounter online.

Oct 6

Oct 6Tue
  1. OpenAIOfficialAI score72

    GPT-6 and Intelligent UI roll out to everyone in ChatGPT

    AIOpenAI announced that GPT-6 and Intelligent UI are now rolling out in ChatGPT for all users. The company says Intelligent UI provides fast, interactive answers, visual explanations of complex topics, and interactive tools for tasks.

    Why it matters: The post names GPT-6 and Intelligent UI rolling out to all ChatGPT users, which matters for anyone tracking how the interface changes.

    Video from @OpenAI's post
  2. Liquid AI BlogOfficialAI score62

    Liquid AI releases open d1-3B and d1-omni-600M decision models for edge devices

    AILiquid AI released two open-weight d1 decision models, d1-3B and d1-omni-600M, on Hugging Face. d1-3B scores 48.57 on the Decision Index v0.2.1 public split and answers a single question in 8 ms on an NVIDIA GeForce RTX 4090 and 50 ms on a Jetson Orin Nano. d1-omni-600M is an experimental checkpoint that handles text with images or audio and scores 15.95 on the same index.

    Why it matters: The release pairs open-weight decision models with measured latency across Apple, NVIDIA, and Jetson hardware, showing how edge deployment changes what is practical.

  3. Google DeepMindOfficialAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    Why it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  4. Philipp SchmidXAI score70

    EmbeddingGemma 2 releases native multimodal embeddings built on Gemma 4

    AIGoogle releases EmbeddingGemma 2, its first native multimodal embedding model, built on Gemma 4 under Apache 2.0. It embeds over 100 languages, code, images, audio, and video into one vector, with an 8,192-token context and four sizes from 270M to 740M parameters. Matryoshka output dimensions of 768, 512, 256, or 128 are supported, and the model is available in Sentence Transformers and LiteRT-LM, with a reported 14% gain on MTEB Code.

    Why it matters: The release extends an embedding model to text, code, images, audio, and video in one vector, a useful option for retrieval systems that mix media types.

  5. vLLMOfficialAI score60

    vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings

    AIvLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.

    Why it matters: The post gives a runnable serve command and day-0 vLLM support, showing how to deploy the new multimodal embedding model locally.

    Image from @vllm_project's post
  6. Unsloth AIOfficialAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle released EmbeddingGemma 2, a 740M-parameter open embedding model under Apache 2.0 that combines a 270M text model with vision (170M) and audio (300M) encoders. The 270M text model can run locally with 0.5GB of RAM, and the full multimodal model with 1GB, and Unsloth provides GGUF files and fine-tuning support.

    Why it matters: The post pairs the model's parameter split and local memory footprint with a benchmark table, showing how the multimodal embedding model compares with other embedding models.

    Image from @UnslothAI's post
  7. Google GemmaOfficialAI score62

    Google Gemma introduces EmbeddingGemma 2, a multimodal on-device embedding model

    AIGoogle Gemma announces EmbeddingGemma 2, a lightweight embedding model that maps text, code, images, video, and audio into a single unified embedding space. The model has a 740M parameter form factor with modular encoders, Matryoshka Representation Learning dimensions from 768 down to 128, and an 8K context window that is 4x larger than the text-only EmbeddingGemma. It is released under the commercially permissive Apache 2.0 license.

    Why it matters: The post gives concrete specs for an on-device multimodal embedding model, including parameter count, dimension options, context window, and license, useful for judging deployment fit.

    Video from @googlegemma's post
  8. Google DeepMindOfficialAI score62

    Google DeepMind releases EmbeddingGemma 2, a natively multimodal open embedding model

    AIGoogle DeepMind introduced EmbeddingGemma 2, its first natively multimodal open model for on-device embeddings. The model expands beyond text to unify code, images, audio, and video in a shared embedding space.

    Why it matters: The release extends an on-device embedding model from text to code, images, audio, and video, which matters for teams building cross-modal search or retrieval.

    Video from @GoogleDeepMind's post
  9. Sundar PichaiXAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

    Why it matters: The post gives the parameter count and modalities, and notes that weights are on Hugging Face, which helps readers assess its fit for offline RAG.

    Video from @sundarpichai's post
  10. Google DeepMind · The KeywordOfficialAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  11. merveXAI score72

    Mistral Large 4 will open its weights at the end of October

    AIMistral announced Mistral Large 4, which it describes as a natively multimodal model with 1T parameters and 49B active. Mistral says it is available via API now, with open weights to follow at the end of October, and a Hugging Face page is listed for the release.

    Why it matters: The quoted Mistral announcement gives specific size, activation, and API details, and the open-weights timing matters for teams weighing open model options.

    Image from @mervenoyann's post
  12. Thomas WolfXAI score62

    Mistral Large 4 open weights are set for release at end of October

    AIMistral announced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active, now available via API. Open weights are scheduled for release at the end of October, with a countdown page on Hugging Face showing October 31, 2026.

    Why it matters: The post pairs a Mistral Large 4 announcement with a dated open-weights release, giving a concrete timeline for readers tracking European open models.

    Image from @Thom_Wolf's post
  13. Julien ChaumondXAI score70

    Mistral Large 4 announced with open weights due end of October

    AIJulien Chaumond reposted Mistral's announcement of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters. Mistral says it is available via API today, with open weights scheduled for release at the end of October, and is working privately with cybersecurity partners.

    Why it matters: The post lays out Mistral Large 4's scale, multimodal design, and availability timeline, which helps readers gauge the open-weights landscape outside China.

  14. Guillaume Lample @ NeurIPS 2024XAI score78

    Mistral launches Large 4 preview with 1T parameters and open weights due October

    AIMistral has launched a preview of Mistral Large 4 (ML4), a 1T-parameter multimodal model with 49B active parameters. The company says it is the strongest open-weight model from the US or Europe on aggregated benchmarks and is available via API now, with open weights planned for the end of October.

    Why it matters: The post gives parameter counts, a preview timeline, and an open-weights release date, which help readers judge how Mistral's model compares with other open-weight options.

    Image from @GuillaumeLample's post
  15. Mistral AIOfficialAI score62

    Mistral AI unveils Mistral Large 4, a 1T-parameter natively multimodal model

    AIMistral AI introduced Mistral Large 4, a natively multimodal model with 1T parameters and 49B active parameters. The company says it is the best open-weights model from the US or Europe on aggregated benchmarks and is available via API today, with open weights due at the end of October.

    Why it matters: The post gives concrete scale, active parameter, and deployment details for a model claimed as the best US or European open-weights model on aggregated benchmarks.

    Video from @MistralAI's post