Skip to contentSkip to stories

Updated

#Multimodal

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 6

Oct 6Tue
  1. Google DeepMindAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    Why it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  2. Philipp SchmidAI score70

    EmbeddingGemma 2 releases native multimodal embeddings built on Gemma 4

    AIGoogle releases EmbeddingGemma 2, its first native multimodal embedding model, built on Gemma 4 under Apache 2.0. It embeds over 100 languages, code, images, audio, and video into one vector, with an 8,192-token context and four sizes from 270M to 740M parameters. Matryoshka output dimensions of 768, 512, 256, or 128 are supported, and the model is available in Sentence Transformers and LiteRT-LM, with a reported 14% gain on MTEB Code.

    Why it matters: The release extends an embedding model to text, code, images, audio, and video in one vector, a useful option for retrieval systems that mix media types.

  3. vLLMAI score60

    vLLM Adds Day-0 Support for Google's EmbeddingGemma 2 Multimodal Embeddings

    AIvLLM announced day-0 support for EmbeddingGemma 2 from Google DeepMind, a bidirectional omni-modal embedding model that maps text, image, audio, video, and interleaved inputs into one vector space. Users can try it with the latest vLLM nightly build using the command vllm serve google/embeddinggemma-2 --runner pooling. The quoted Google post says the model is built on the Gemma 4 architecture and released under Apache 2.0.

    Image from @vllm_project's post
  4. Google · Innovation & AIAI score42

    Google Study Tests AI-Guided Blind Sweep Ultrasounds for Pregnant Women in Kenya and Chicago

    AIGoogle researchers, working with Northwestern Medicine and Jacaranda Health, trained healthcare workers to perform "blind sweep" ultrasounds analyzed by machine learning models. The models estimated gestational age and fetal presentation as accurately as a trained sonographer in a study of 1,000 mothers each in Nairobi and Chicago. The AI processes results on the device, so it needs no electricity supply or Wi-Fi.

  5. 👩‍💻 Paige BaileyAI score54

    EmbeddingGemma 2 launches as an Apache 2.0 multimodal embeddings model

    AIGoogle's EmbeddingGemma 2 is an open embeddings model for on-device use that covers code, image, video, audio, and text. It comes in modular sizes from 270M text/code to 740M full multimodal, supports Matryoshka truncation down to 128 dimensions, and reports a 14% gain on MTEB Code over v1 under an Apache 2.0 license. The author's post highlights the release and a Hugging Face demo, while the benchmark table compares it with several models.

    Video from @DynamicWebPaige's post
  6. Google GemmaAI score62

    Google Gemma introduces EmbeddingGemma 2, a multimodal on-device embedding model

    AIGoogle Gemma announces EmbeddingGemma 2, a lightweight embedding model that maps text, code, images, video, and audio into a single unified embedding space. The model has a 740M parameter form factor with modular encoders, Matryoshka Representation Learning dimensions from 768 down to 128, and an 8K context window that is 4x larger than the text-only EmbeddingGemma. It is released under the commercially permissive Apache 2.0 license.

    Video from @googlegemma's post
  7. Google for DevelopersAI score40

    Google's multimodal embedding toolkit runs fully offline on device

    AIGoogle's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.

    Video from @googledevs's post
  8. Google DeepMindAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    AIGoogle DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

    Image from @GoogleDeepMind's post
  9. Sundar PichaiAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

    Video from @sundarpichai's post
  10. Google DeepMind · The KeywordAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  11. merveAI score72

    Mistral Large 4 will open its weights at the end of October

    AIMistral announced Mistral Large 4, which it describes as a natively multimodal model with 1T parameters and 49B active. Mistral says it is available via API now, with open weights to follow at the end of October, and a Hugging Face page is listed for the release.

    Why it matters: The quoted Mistral announcement gives specific size, activation, and API details, and the open-weights timing matters for teams weighing open model options.

    Image from @mervenoyann's post
  12. Simon WillisonAI score36

    Mistral's Pelican SVG Test Passes, Tied to Mistral Large 4 Context

    AISimon Willison reports that Mistral can now generate his pelican SVG test, shared via a Markdown SVG renderer. The post links to a rendered result but gives no benchmark or scoring details. Background from Mistral's own announcement describes Mistral Large 4 as a 1T-parameter, natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

    Image from @simonw's post
  13. Julien ChaumondAI score70

    Mistral Large 4 announced with open weights due end of October

    AIJulien Chaumond reposted Mistral's announcement of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters. Mistral says it is available via API today, with open weights scheduled for release at the end of October, and is working privately with cybersecurity partners.

    Why it matters: The post lays out Mistral Large 4's scale, multimodal design, and availability timeline, which helps readers gauge the open-weights landscape outside China.

  14. Guillaume Lample @ NeurIPS 2024AI score42

    Mistral's ML4 matches top open-weight models on coding and agentic benchmarks

    AIMistral's ML4 model matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and reaches state-of-the-art results on finance and legal workflows and complex multimodal grounding benchmarks. The post says it can navigate terminal workflows, work across spreadsheets, slides, and PDFs, and reason over scientific and multimodal tasks.

    Image from @GuillaumeLample's post
  15. Guillaume Lample @ NeurIPS 2024AI score78

    Mistral launches Large 4 preview with 1T parameters and open weights due October

    AIMistral has launched a preview of Mistral Large 4 (ML4), a 1T-parameter multimodal model with 49B active parameters. The company says it is the strongest open-weight model from the US or Europe on aggregated benchmarks and is available via API now, with open weights planned for the end of October.

    Why it matters: The post gives parameter counts, a preview timeline, and an open-weights release date, which help readers judge how Mistral's model compares with other open-weight options.

    Image from @GuillaumeLample's post
  16. Mistral AIAI score80

    Mistral Large 4 launches as a public preview with weights due end of month

    AIMistral AI launched a public preview API for Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters, and says it will release the weights by the end of the month. The company reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA, 28.3% on Terminal-Bench 4, and 59.9% on AutomationBench. The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's datacenters in Europe.

    Why it matters: The post gives benchmark figures and a weights timeline for an open-weight model, letting readers compare it with other open models and judge its access terms.

  17. GeekParkAI score46

    Huawei Mate 90 Pro Max starts at 9,499 yuan, with camera upgrades leading

    AIHuawei's Mate 90 Pro Max, launched October 1, starts at 9,499 yuan for the 12GB+512GB model, and the Collector's Edition starts at 10,999 yuan. The phone's main upgrade is its camera, including a new "Portrait Original" mode that preserves skin tone and makeup, a 200-megapixel telephoto lens with about 4x optical zoom, and generative-AI glare removal for night shots.