Skip to content

#RAG

Oct 8

TodayOct 8Thu6 items
  1. Sakana AIAI score36

    Sakana AI technology adopted in Iris's Evidence Finder tool

    Sakana AI's technology has been adopted in "Evidence Finder," an evidence search tool for physicians from Iris Inc. (@ShoOkiyama). In response to physicians' questions, it searches the literature and answers while explicitly citing its sources. Sakana AI's technology handles the comparison and integration of literature and the generation of answers. For details, see here: https://sakana.ai/namazu-aillis/ 🐟

  2. Tessl BlogAI score29

    One Brain Means Owning Your Organizational Memory

    Leapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.

  3. Google Cloud TechAI score40

    Our borderless Lakehouse lets Gemini query AWS and Azure with no variable egress fees, read directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data, and federate open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue → https://goo.gle/3TT9y4X

    Our borderless Lakehouse lets Gemini query AWS and Azure with no variable egress fees, read directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data, and federate open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue → https://goo.gle/3TT9y4X

  4. LlamaIndexAI score8

    Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️

    Markdown is all you need. (Mostly.) A parser can get every word on the page right and still lose which column a number belongs to. Then your model has to guess. Markdown keeps headings, lists, and tables intact, and it stays readable when you're debugging a bad answer. For tables with merged headers, we switch to HTML. Read our breakdown on why it's our default output for parsing below! ⬇️

Oct 7

Oct 7Wed
  1. Microsoft Foundry BlogAI score22

    Azure Document Intelligence vs. Content Understanding: Choosing the Right Document Service

    Microsoft's Foundry blog guide advises keeping existing Azure Document Intelligence workloads that meet production requirements. It recommends evaluating Azure Content Understanding for high-variation, unstructured, reasoning, RAG, or multimodal document scenarios, and for new cloud OCR or layout workloads.

  2. Aravind SrinivasAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    Perplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

  3. PerplexityAI score41

    Both sizes are distilled token by token from one 18B teacher, so they share one embedding space. A corpus indexed with the 9B model can be searched with 0.6B queries. That lifts ViDoRe v3 from 62.3% to 63.5% with no added query cost.

    Both sizes are distilled token by token from one 18B teacher, so they share one embedding space. A corpus indexed with the 9B model can be searched with 0.6B queries. That lifts ViDoRe v3 from 62.3% to 63.5% with no added query cost.

  4. PerplexityAI score36

    Dense embedding models compress each document into one vector, which loses detail as pages get longer or more visual. pplx-embed-v2-late keeps a 128-dimensional vector per token and scores with MaxSim, so each query token is matched to its closest token in the document.

    Dense embedding models compress each document into one vector, which loses detail as pages get longer or more visual. pplx-embed-v2-late keeps a 128-dimensional vector per token and scores with MaxSim, so each query token is matched to its closest token in the document.

  5. PerplexityAI score45

    We're releasing pplx-embed-v2-late, two late-interaction embedding models that retrieve text, images, and pages with a shared embedding space for cross-model querying. Both models achieve frontier performance and are publicly available on Hugging Face. https://www.perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector

    We're releasing pplx-embed-v2-late, two late-interaction embedding models that retrieve text, images, and pages with a shared embedding space for cross-model querying. Both models achieve frontier performance and are publicly available on Hugging Face. https://www.perplexity.ai/hub/blog/multimodal-embeddings-beyond-a-single-vector

  6. AWS Machine Learning BlogAI score44

    Qlik Builds Grounded Enterprise AI Answers Using Amazon Bedrock

    Qlik built Qlik Answers, a natural-language assistant that returns sourced answers from knowledge bases, analytics apps, glossaries, and documents, using Amazon Bedrock for model access. The system routes each question through specialist agents and retrieval on Amazon OpenSearch Service, with Amazon Bedrock Guardrails applied to every request and response. Qlik serves more than 40,000 customers across regions, using Amazon SageMaker AI as an in-Region fallback when models are not yet available on Bedrock.

Oct 6

Oct 6Tue
  1. meng shaoAI score62

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, video, and audio into one 768-dimensional space. Text-only use needs a 270M-parameter footprint, about 191MB active RAM when quantized on a Pixel 11 Pro, while loading all modalities takes about 567MB. The reported MTEB Code NDCG@10 score is 78.68, about 14% above the first generation, and MTEB Multilingual v2 is 61.36, roughly flat.

  2. Google DeepMindAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    AIWhy it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  3. GoogleAI score44

    When paired with Gemma 4, EmbeddingGemma 2 powers efficient, on-device retrieval augmented generation (RAG) with a lower memory footprint. EmbeddingGemma 2 retrieves local files, and Gemma 4 reasons over them for grounded answers with complete privacy.

    When paired with Gemma 4, EmbeddingGemma 2 powers efficient, on-device retrieval augmented generation (RAG) with a lower memory footprint. EmbeddingGemma 2 retrieves local files, and Gemma 4 reasons over them for grounded answers with complete privacy.

  4. Google for DevelopersAI score40

    Google's multimodal embedding toolkit runs fully offline on device

    Google's new multimodal embedding setup processes image, audio, and video entirely offline with zero server calls. Its modular design lets developers drop unused vision and audio components to save memory, and flexible dimension sizes cut local database storage by up to 6x. It can also pair with Gemma 4 to build RAG pipelines with minimal memory and processing requirements.

  5. Google DeepMindAI score58

    Google DeepMind releases EmbeddingGemma 2 with 740M parameters under Apache 2.0

    Google DeepMind released EmbeddingGemma 2, a 740M-parameter embedding model, under an Apache 2.0 license. The post says it is competitive across benchmarks and outperforms some specialist models more than twice its size, and that developers can use it for multimodal search or pair it with Gemma 4 for on-device RAG. Weights are available on Hugging Face and Kaggle.

  6. Sundar PichaiAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google introduces EmbeddingGemma 2, its first open, natively multimodal embedding model, covering text, code, image, video, and audio tasks. It has a 740M parameter form factor, is positioned for offline, privacy-first RAG when paired with Gemma 4, and the post claims it outperforms some specialist models more than twice its size. Weights are available now on Hugging Face.

  7. The Next PlatformAI score38

    Dell Adds Data Context, Prep, and Storage Features to Its AI Data Platform

    Dell is adding agentic AI capabilities to its AI Data Platform, including a Unified Semantic Layer with a searchable glossary and an Enterprise Knowledge Graph built with Nvidia's Auto-Ontology open source library. The features are designed to give agents shared context, reducing repeated token generation and compute costs. The platform's layers include the Data Orchestration Engine, Data Engines, and Storage Engines such as PowerScale, ObjectScale, and the Lightning File System.

Oct 5

Oct 5Mon
  1. Google Developers BlogAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    AIWhy it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  2. MIT Technology Review · AIAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    A survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

Oct 2

Oct 2Fri
  1. Cloudflare Blog · AIAI score41

    Cloudflare Launches Web Search API via AI Gateway for Live Agent Grounding

    Cloudflare introduced a Web Search API through AI Gateway, partnering with Ceramic.ai, Exa, and Linkup to give agents fresh web results instead of guessed URLs. Requests appear in AI Gateway logs and draw from AI Gateway credits, with partners committing to Cloudflare's Verified bots crawling standards and including source links in results. Partners at list API pricing without markup are available via a REST endpoint or a Workers binding, with native server tools planned.

Oct 1

Oct 1Thu

Sep 30

Sep 30Wed
  1. PerplexityAI score20

    On ConTEB, the preview has the highest average nDCG@10 of the models tested, though not on every task. It also beats voyage-context-4 on chunk retrieval while using 8x less storage per vector: 1 KB (1024 dims, int8) vs 8 KB (2048, float32).

    On ConTEB, the preview has the highest average nDCG@10 of the models tested, though not on every task. It also beats voyage-context-4 on chunk retrieval while using 8x less storage per vector: 1 KB (1024 dims, int8) vs 8 KB (2048, float32).

  2. PerplexityAI score23

    context-bench is a benchmark for context-aware retrieval, created and privately held by @turbopuffer. Its queries, documents and capabilities are inspired by conversations with turbopuffer customers. It has 2,099 queries and 38,894 documents. We submitted for blind evaluation.

    context-bench is a benchmark for context-aware retrieval, created and privately held by @turbopuffer. Its queries, documents and capabilities are inspired by conversations with turbopuffer customers. It has 2,099 queries and 38,894 documents. We submitted for blind evaluation.

  3. PerplexityAI score20

    We overcome gold-chunk supervision limits by distilling relevance from our query-aware context compression model. It scores every document token against the query. Aggregated into chunk-level targets, those scores train the embedder to retrieve answer and supporting chunks.

    We overcome gold-chunk supervision limits by distilling relevance from our query-aware context compression model. It scores every document token against the query. Aggregated into chunk-level targets, those scores train the embedder to retrieve answer and supporting chunks.

  4. PerplexityAI score25

    Retrieval systems often split long documents into chunks, but this strips away the surrounding context. Contextual embedding models fix this by encoding the whole document once and pooling chunk vectors afterward. They are usually trained on one gold chunk per query.

    Retrieval systems often split long documents into chunks, but this strips away the surrounding context. Contextual embedding models fix this by encoding the whole document once and pooling chunk vectors afterward. They are usually trained on one gold chunk per query.

  5. PerplexityAI score34

    We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

    We built a new way to train contextual embedding models, which encode each chunk of a document with the whole document in view. pplx-embed-v2-context-9b-preview sets a new state of the art on ConTEB and @turbopuffer's new, privately held context-bench. https://www.perplexity.ai/hub/blog/contextual-embedding-beyond-the-gold-passage

Sep 29

Sep 29Tue
  1. Microsoft Foundry BlogAI score30

    Why content extraction still matters in the GenAI era

    Microsoft's Azure AI team argues that better models do not eliminate the need for a dedicated content extraction layer, since agents need trustworthy, structured, and auditable inputs. The post notes that building extraction directly on an LLM quickly demands chunking, layout parsing, grounding, normalization, and evaluation infrastructure. Microsoft positions Azure Document Intelligence and Azure Content Understanding in Foundry Tools as managed options for that layer.

  2. AMDAI score18

    More users shouldn't have to mean slower AI. In @Signal_65 RAG testing hosted via @TensorWave, AMD Instinct MI355X delivered ~42% lower p99 latency and roughly 2x the concurrent-user headroom before breaching the evaluated SLA: https://bit.ly/45hAvl3

    More users shouldn't have to mean slower AI. In @Signal_65 RAG testing hosted via @TensorWave, AMD Instinct MI355X delivered ~42% lower p99 latency and roughly 2x the concurrent-user headroom before breaching the evaluated SLA: https://bit.ly/45hAvl3

Sep 28

Sep 28Mon
  1. LlamaIndexAI score30

    LlamaIndex says frontier VLMs still struggle parsing tax and W-series forms

    LlamaIndex argues that frontier vision-language models still fail on real forms such as W-2s, 1040s, W-9s, and scanned W-4s, because forms require detecting every field, preserving section hierarchy, linking values to their exact boxes, and reading handwriting and checkmarks. The company's blog post details these failure modes and presents a custom cookbook for LlamaParse as a cheaper way to handle such forms.

Sep 18

Sep 18Fri