Skip to contentSkip to stories

Updated

#Multimodal

Oct 6

Oct 6Tue
  1. Philipp SchmidAI score70

    EmbeddingGemma 2 releases native multimodal embeddings built on Gemma 4

    AIGoogle releases EmbeddingGemma 2, its first native multimodal embedding model, built on Gemma 4 under Apache 2.0. It embeds over 100 languages, code, images, audio, and video into one vector, with an 8,192-token context and four sizes from 270M to 740M parameters. Matryoshka output dimensions of 768, 512, 256, or 128 are supported, and the model is available in Sentence Transformers and LiteRT-LM, with a reported 14% gain on MTEB Code.

    Why it matters: The release extends an embedding model to text, code, images, audio, and video in one vector, a useful option for retrieval systems that mix media types.

  2. vLLMAI score60

    vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings

    AIvLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.

  3. Google · Innovation & AIAI score42

    Google Study Tests AI-Guided Blind Sweep Ultrasounds for Pregnant Women in Kenya and Chicago

    AIGoogle researchers, working with Northwestern Medicine and Jacaranda Health, trained healthcare workers to perform "blind sweep" ultrasounds analyzed by machine learning models. The models estimated gestational age and fetal presentation as accurately as a trained sonographer in a study of 1,000 mothers each in Nairobi and Chicago. The AI processes results on the device, so it needs no electricity supply or Wi-Fi.

  4. Paige BaileyAI score54

    EmbeddingGemma 2 launches as an Apache 2.0 multimodal embeddings model

    AIGoogle's EmbeddingGemma 2 is an open embeddings model for on-device use that covers code, image, video, audio, and text. It comes in modular sizes from 270M text/code to 740M full multimodal, supports Matryoshka truncation down to 128 dimensions, and reports a 14% gain on MTEB Code over v1 under an Apache 2.0 license. The author's post highlights the release and a Hugging Face demo, while the benchmark table compares it with several models.

  5. Google for DevelopersAI score40

    Google's multimodal embedding toolkit runs fully offline on device

    AIGoogle's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.

  6. Google DeepMindAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    AIGoogle DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

  7. Sundar PichaiAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

  8. Google DeepMind · The KeywordAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  9. Merve NoyanAI score72

    Mistral Large 4 will open its weights at the end of October

    AIMistral announced Mistral Large 4, which it describes as a natively multimodal model with 1T parameters and 49B active. Mistral says it is available via API now, with open weights to follow at the end of October, and a Hugging Face page is listed for the release.

    Why it matters: The quoted Mistral announcement gives specific size, activation, and API details, and the open-weights timing matters for teams weighing open model options.

  10. Simon WillisonAI score36

    Mistral's Pelican SVG Test Passes, Tied to Mistral Large 4 Context

    AISimon Willison reports that Mistral can now generate his pelican SVG test, shared via a Markdown SVG renderer. The post links to a rendered result but gives no benchmark or scoring details. Background from Mistral's own announcement describes Mistral Large 4 as a 1T-parameter, natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

  11. Julien ChaumondAI score70

    Mistral Large 4 announced with open weights due end of October

    AIJulien Chaumond reposted Mistral's announcement of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters. Mistral says it is available via API today, with open weights scheduled for release at the end of October, and is working privately with cybersecurity partners.

    Why it matters: The post lays out Mistral Large 4's scale, multimodal design, and availability timeline, which helps readers gauge the open-weights landscape outside China.

  12. Guillaume LampleAI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

  13. Guillaume LampleAI score78

    Mistral launches Large 4 preview with 1T parameters and open weights due October

    AIMistral has launched a preview of Mistral Large 4 (ML4), a 1T-parameter multimodal model with 49B active parameters. The company says it is the strongest open-weight model from the US or Europe on aggregated benchmarks and is available via API now, with open weights planned for the end of October.

    Why it matters: The post gives parameter counts, a preview timeline, and an open-weights release date, which help readers judge how Mistral's model compares with other open-weight options.

  14. Mistral AIAI score80

    Mistral Large 4 launches as a public preview with weights due end of month

    AIMistral AI launched a public preview API for Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters, and says it will release the weights by the end of the month. The company reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench. The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's datacenters in Europe.

    Why it matters: The post gives benchmark figures and a weights timeline for an open-weight model, letting readers compare it with other open models and judge its access terms.

  15. Luma AI NewsAI score22

    Cyberpunk AI Prompts Guide Covers Video and Image Generation Workflows

    AIThe guide offers a prompt structure for cyberpunk visuals built from subject, environment, lighting, camera, style, and quality modifiers, with magenta and cyan neon, rain-slicked reflections, and fog named as key mood elements. It presents 15 ready-to-use prompts and argues that free tools suit testing directions, while full access is needed for commercial campaigns.

Oct 5

Oct 5Mon
  1. Google Developers BlogAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    AIGoogle released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    Why it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  2. Google Developers BlogAI score67

    Google releases EmbeddingGemma 2, a multimodal embedding model for on-device search

    AIGoogle DeepMind launched EmbeddingGemma 2, an open-weight 740M parameter model that maps text, images, video frames, and audio into one vector space. The model can run on-device, with about 567MB active RAM for the full multimodal model on a Google Pixel 11 Pro, and is available through Google AI Edge Gallery, Google AI Edge Foresight on Mac, and MediaPipe Tasks, with ML Kit support coming in the weeks ahead.

    Why it matters: The post names concrete on-device apps, memory footprints, and latency figures, showing how a multimodal embedding model can power local search without cloud calls.

  3. Liquid AIAI score37

    Liquid AI's d1 decision model adds vision, rivaling GPT-6.1 Sol at lower cost

    AILiquid AI released d1 with vision support, accepting images, text, or both as inputs. In tests on six real applications, d1 matched or beat GPT-6.1 Sol on four while costing 19x to 200x less than both GPT-6.1 Sol and Claude Opus 5.5. It returns probabilities for yes/no, choice, or score questions in one forward pass, with text decisions in 200 to 300 ms.

  4. Liquid AIAI score36

    Liquid AI's d1 model inspects parts from camera images with 85-97% accuracy

    AILiquid AI's vision-enabled decision model d1 inspects parts directly from camera images and is described as the best such model currently on the market. It reaches 85% to 97% accuracy across four VisA inspection tasks covering circuit boards, candles, cashews, and chewing gum. It understands each task from a short description without task-specific training.

  5. Liquid AI · new models on Hugging FaceAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.