Skip to contentSkip to stories

Updated

#RAG

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Sakana AIOfficialAI score36

    Sakana AI's technology powers Iris's physician evidence search tool

    AIIris Inc.'s medical evidence search tool Evidence Finder has adopted Sakana AI's technology for answering physicians' questions. The system searches the literature and generates answers that cite their sources, handling literature comparison, synthesis, and answer generation.

    Image from @SakanaAILabs's post
  2. Tessl BlogOfficialAI score29

    One Brain Means Owning Your Organizational Memory

    AILeapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.

  3. Google Cloud TechOfficialAI score40

    Google Cloud's borderless Lakehouse lets Gemini query multicloud data directly

    AIGoogle Cloud's borderless Lakehouse lets Gemini query data on AWS and Azure without variable egress fees. It reads directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data. It also federates open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue.

    Video from @GoogleCloudTech's post
  4. Google GemmaOfficialAI score44

    Google publishes a developer guide for EmbeddingGemma 2 multimodal embeddings

    AIGoogle Gemma announces a developer guide showing how to embed text, code, images, video, audio, and interleaved inputs with EmbeddingGemma 2 using the sentence-transformers library. The guide outlines a four-step workflow: loading the model, embedding text and code with task prompts, embedding multimodal inputs, and optionally truncating dimensions with Matryoshka.

    Image from @googlegemma's post
  5. CohereOfficialAI score20

    Cohere hosts live webinar on future of search and retrieval

    AICohere is hosting a live webinar on the future of search and retrieval, covering Embed 5, Parse 5, and its new retrieval methodology, RCP-nDCG. The post is a livestream announcement and does not include details of the methodology or model performance.

Oct 7

Oct 7Wed
  1. Aravind SrinivasXAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

  2. PerplexityOfficialAI score32

    Perplexity reports 92.4% on MADQA document QA benchmark

    AIPerplexity reports its system reached 92.4% accuracy on MADQA, a benchmark of 500 questions over 800 PDFs. The post describes this as the top result among retrievers on that benchmark.

    Image from @perplexity_ai's post
  3. PerplexityOfficialAI score41

    Perplexity's 0.6B and 9B embedding models share one embedding space

    AIPerplexity's 0.6B and 9B models are both distilled token by token from one 18B teacher, so they share a single embedding space. A corpus indexed with the 9B model can be searched using 0.6B queries, raising ViDoRe v3 from 62.3% to 63.5% with no added query cost.

    Image from @perplexity_ai's post
  4. PerplexityOfficialAI score36

    Perplexity's models search PDFs and slides directly, without OCR

    AIPerplexity's models embed images and rendered pages directly, so PDFs, slides, and scans can be searched without OCR. This preserves tables, figures, and layout that text extraction typically drops.

    Image from @perplexity_ai's post
  5. PerplexityOfficialAI score36

    Perplexity's pplx-embed-v2-late keeps per-token vectors for retrieval

    AIPerplexity's pplx-embed-v2-late retains a 128-dimensional vector for each token rather than compressing a document into one vector. It scores matches with MaxSim, pairing each query token with its closest document token, which the post presents as preserving detail in long or visually dense pages.

    Image from @perplexity_ai's post
  6. AWS Machine Learning BlogOfficialAI score44

    Qlik Builds Grounded Enterprise AI Answers Using Amazon Bedrock

    AIQlik built Qlik Answers, a natural-language assistant that returns sourced answers from knowledge bases, analytics apps, glossaries, and documents, using Amazon Bedrock for model access. The system routes each question through specialist agents and retrieval on Amazon OpenSearch Service, with Amazon Bedrock Guardrails applied to every request and response. Qlik serves more than 40,000 customers across regions, using Amazon SageMaker AI as an in-Region fallback when models are not yet available on Bedrock.

Oct 6

Oct 6Tue
  1. meng shaoXAI score62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.

    Image from @shao__meng's post
  2. Google DeepMindOfficialAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    Why it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  3. GoogleOfficialAI score44

    EmbeddingGemma 2 pairs with Gemma 4 for on-device RAG

    AIGoogle says EmbeddingGemma 2, paired with Gemma 4, enables efficient on-device retrieval-augmented generation with a lower memory footprint. In this setup, EmbeddingGemma 2 retrieves local files and Gemma 4 reasons over them to produce grounded answers while keeping data private.

    Video from @Google's post
  4. Google for DevelopersOfficialAI score40

    Google's multimodal embedding toolkit runs fully offline on device

    AIGoogle's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.

    Video from @googledevs's post
  5. Google DeepMindOfficialAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    AIGoogle DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

    Image from @GoogleDeepMind's post
  6. Sundar PichaiXAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

    Video from @sundarpichai's post
  7. The Next PlatformNewsAI score38

    Dell Adds Data Context, Prep, and Storage Features to Its AI Data Platform

    AIDell is adding agentic AI capabilities to its AI Data Platform, including a Unified Semantic Layer with a searchable glossary and an Enterprise Knowledge Graph built with Nvidia's Auto-Ontology open source library. The features are designed to give agents shared context, reducing repeated token generation and compute costs. The platform's layers include the Data Orchestration Engine, Data Engines, and Storage Engines such as PowerScale, ObjectScale, and the Lightning File System.

Oct 5

Oct 5Mon
  1. Google Developers BlogOfficialAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    AIGoogle released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    Why it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  2. MIT Technology Review · AINewsAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    AIA survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

Oct 2

Oct 2Fri
  1. Cloudflare Blog · AIOfficialAI score41

    Cloudflare Launches Web Search API via AI Gateway for Live Agent Grounding

    AICloudflare introduced a Web Search API through AI Gateway, partnering with Ceramic.ai, Exa, and Linkup to give agents fresh web results instead of guessed URLs. Requests appear in AI Gateway logs and draw from AI Gateway credits, with partners committing to Cloudflare's Verified bots crawling standards and including source links in results. Partners at list API pricing without markup are available via a REST endpoint or a Workers binding, with native server tools planned.

Oct 1

Oct 1Thu
  1. Cloudflare Blog · AIOfficialAI score52

    Cloudflare AI Search reaches general availability with image embeddings and OCR

    AICloudflare's AI Search is now generally available, adding native image embeddings for visual retrieval, OCR for scanned PDFs, and a 10 MiB file limit up from 4 MiB. Billing begins November 1, 2026, charged per ingested token, stored GB-month, and query, with a free monthly allotment on all Workers plans.

Sep 30

Sep 30Wed
  1. PerplexityOfficialAI score20

    Perplexity's embedding preview tops ConTEB benchmark average nDCG@10

    AIPerplexity's embedding preview achieves the highest average nDCG@10 among tested models on ConTEB, though not on every task. It also outperforms voyage-context-4 on chunk retrieval while using 8x less storage per vector, at 1 KB (1024 dims, int8) versus 8 KB (2048 dims, float32).

    Image from @perplexity_ai's post
  2. PerplexityOfficialAI score23

    Perplexity submits turbopuffer's context-bench for blind evaluation

    AIPerplexity says context-bench, a context-aware retrieval benchmark created and privately held by turbopuffer, has 2,099 queries and 38,894 documents. Perplexity submitted it for blind evaluation, with queries and capabilities inspired by turbopuffer customer conversations.

    Image from @perplexity_ai's post
  3. PerplexityOfficialAI score20

    Perplexity distills query-aware compression scores to train retrieval embedders

    AIPerplexity overcomes gold-chunk supervision limits by distilling relevance from its query-aware context compression model. The model scores every document token against the query, and those scores, aggregated into chunk-level targets, train the embedder to retrieve answer and supporting chunks.

    Image from @perplexity_ai's post
  4. Aidan GomezXAI score38

    Cohere launches Embed 5 Pro and Embed 5 Fast embedding models

    AICohere has introduced Embed 5, a new family of state-of-the-art embedding models, with Embed 5 Pro for frontier capabilities and Embed 5 Fast for low-latency performance. The post says the models are extremely scalable, offer SOTA accuracy, and can be deployed privately and accessed through Model Vault.

Sep 29

Sep 29Tue
  1. Microsoft Foundry BlogOfficialAI score30

    Why content extraction still matters in the GenAI era

    AIMicrosoft's Azure AI team argues that better models do not eliminate the need for a dedicated content extraction layer, since agents need trustworthy, structured, and auditable inputs. The post notes that building extraction directly on an LLM quickly demands chunking, layout parsing, grounding, normalization, and evaluation infrastructure. Microsoft positions Azure Document Intelligence and Azure Content Understanding in Foundry Tools as managed options for that layer.

Sep 28

Sep 28Mon
  1. LlamaIndex 🦙OfficialAI score30

    LlamaIndex says frontier VLMs still struggle parsing tax and W-series forms

    AILlamaIndex argues that frontier vision-language models still fail on real forms such as W-2s, 1040s, W-9s, and scanned W-4s, because forms require detecting every field, preserving section hierarchy, linking values to their exact boxes, and reading handwriting and checkmarks. The company's blog post details these failure modes and presents a custom cookbook for LlamaParse as a cheaper way to handle such forms.

    Image from @llama_index's post

Sep 18

Sep 18Fri
  1. GitHub Blog · AI & MLOfficialAI score34

    Should You Read AI Code, Is RAG Dead, and Did Skills Kill MCP?

    AIGitHub's latest podcast episode examines five common AI hot takes, including whether developers must still read AI-generated code. It argues review effort should match risk, and that Skills and MCP solve different problems. It also says retrieval-augmented generation (RAG) remains useful and works alongside agents, skills, and MCP.