Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. SantiagoXAI score22

    Model infers derived values from document data, computing yearly costs from monthly figures

    AIA new model extracts values absent from a document by computing them from figures that are present, such as deriving a yearly product cost from a monthly price. Santiago says the video shows examples of inferring complex formulas. The background post describes this as Higher-Order Extraction, which deterministically computes needed numbers from raw page values.

  2. GitHub Copilot ChangelogOfficialAI score30

    GitHub launches purpose-built AI model for leaked secret detection across developer workflows

    AIGitHub is rolling out a fine-tuned, purpose-built model for secret detection that reads surrounding code to identify likely credentials, including passwords without recognizable token formats. Existing AI-detected Password alerts have been upgraded automatically, and AI-detected secrets in push protection is in private preview. New opt-in checks in push protection and the GitHub Copilot /security-review command will consume GitHub AI Credits.

  3. PerplexityOfficialAI score32

    Perplexity reports 92.4% on MADQA document QA benchmark

    AIPerplexity reports its system reached 92.4% accuracy on MADQA, a benchmark of 500 questions over 800 PDFs. The post describes this as the top result among retrievers on that benchmark.

    Image from @perplexity_ai's post
  4. Andrew CurranXAI score42

    Andrew Curran says OpenAI's 722 math results omit cryptography breakthroughs

    AIAndrew Curran notes OpenAI's 722 published mathematical results show striking under-representation of cryptographic breakthroughs, and he says he has personally witnessed US government censorship of academic quantum cryptanalysis results. He calls backroom government interventionism his base case and says rumors suggest yesterday's OpenAI math release was only the first of three batches.

    Image from @AndrewCurran_'s post
  5. PerplexityOfficialAI score41

    Perplexity's 0.6B and 9B embedding models share one embedding space

    AIPerplexity's 0.6B and 9B models are both distilled token by token from one 18B teacher, so they share a single embedding space. A corpus indexed with the 9B model can be searched using 0.6B queries, raising ViDoRe v3 from 62.3% to 63.5% with no added query cost.

    Image from @perplexity_ai's post
  6. PerplexityOfficialAI score36

    Perplexity's models search PDFs and slides directly, without OCR

    AIPerplexity's models embed images and rendered pages directly, so PDFs, slides, and scans can be searched without OCR. This preserves tables, figures, and layout that text extraction typically drops.

    Image from @perplexity_ai's post
  7. PerplexityOfficialAI score36

    Perplexity's pplx-embed-v2-late keeps per-token vectors for retrieval

    AIPerplexity's pplx-embed-v2-late retains a 128-dimensional vector for each token rather than compressing a document into one vector. It scores matches with MaxSim, pairing each query token with its closest document token, which the post presents as preserving detail in long or visually dense pages.

    Image from @perplexity_ai's post
  8. Microsoft ResearchOfficialAI score24

    Agent Lightning connects existing AI agents to reinforcement learning training

    AIMicrosoft Research introduced Agent Lightning, a tool that connects existing AI agents to reinforcement learning training. It aims to make agents easier to improve without rebuilding them, since their tools, context, and decision-making are typically managed by complex frameworks.

    Video from @MSFTResearch's post
  9. Jerry LiuXAI score47

    OpenDocRouter offers one API for document parsing across many models

    AILlamaIndex has launched OpenDocRouter, a unified API that transcribes documents to markdown across frontier and open-weight models at cost plus a small transaction fee. The service handles rate limits, adds bounding boxes and layout as a service, and will benchmark new OCR models on ParseBench before adding them.

    Video from @jerryjliu0's post
  10. LlamaIndex 🦙OfficialAI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

    Video from @llama_index's post
  11. Codex · GitHub ReleasesOfficialAI score34

    Codex 0.161.0 makes GPT-6.1 Sol default and adds Daybreak opt-in

    AIOpenAI's Codex 0.161.0 release makes GPT-6.1 Sol the default model in the bundled and Amazon Bedrock catalogs. It adds Amazon Bedrock support for multi-agent V2 and Ultra reasoning on compatible models, plus opt-in Daybreak routing enabled through --enable cli_daybreak or features.cli_daybreak=true.

  12. Microsoft ResearchOfficialAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    AIMicrosoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    Why it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

  13. NVIDIA Technical BlogOfficialAI score22

    Validate AI Factory Changes with Digital Twins and AI Agents

    AINVIDIA describes using digital twins and AI agents to validate changes to AI factory infrastructure, which combines GPUs, CPUs, switches, DPUs, and SuperNICs with schedulers, orchestration services, security controls, and a fast-changing software stack. The source frames the challenge as confirming that hardware, software, and policies work together for target workloads before deployment. The available excerpt does not give further detail on specific tools or results.

  14. Ai2OfficialAI score36

    Ai2 adapts existing models to bytes with a short training run

    AIAi2 describes a recipe that adapts existing models to process raw bytes through a brief additional training run. The approach keeps the model's core intact and adds components that group bytes into variable-length patches for processing, then expand outputs back to byte-level predictions of the next byte.

  15. Demis HassabisXAI score34

    Isomorphic Labs joins Virtual Biology Initiative as founding member

    AIIsomorphic Labs has joined the Virtual Biology Initiative (VBI) as a founding member, helping build an open resource for global researchers to accelerate understanding and treatment of disease. Demis Hassabis, owner of the account and Google/Gemini-affiliated, called applying AI to medicine and human health the most important thing AI can be used for and welcomed the partnership with CZI and VBI.

  16. AWS Machine Learning BlogOfficialAI score38

    Agentic Automation Business Cases Need to Count More Than Saved Hours

    AIAWS Machine Learning Blog argues that the traditional hours-saved ROI model, built for rule-based RPA, misses most of the value of agentic automation. It proposes an Agentic Value Model covering time savings, exception handling, decision quality, and change resilience, with value counted only when tied to a defined P&L mechanism and owner.

  17. AWS Machine Learning BlogOfficialAI score44

    Qlik Builds Grounded Enterprise AI Answers Using Amazon Bedrock

    AIQlik built Qlik Answers, a natural-language assistant that returns sourced answers from knowledge bases, analytics apps, glossaries, and documents, using Amazon Bedrock for model access. The system routes each question through specialist agents and retrieval on Amazon OpenSearch Service, with Amazon Bedrock Guardrails applied to every request and response. Qlik serves more than 40,000 customers across regions, using Amazon SageMaker AI as an in-Region fallback when models are not yet available on Bedrock.

  18. AWS Machine Learning BlogOfficialAI score53

    Automate remediation after AWS DevOps Agent investigations with Lambda and Bedrock

    AIThe AWS Machine Learning Blog describes an automated remediation workflow that acts on AWS DevOps Agent investigation results. Amazon EventBridge triggers a Lambda durable function that uses Amazon Bedrock to propose fixes from an allowlist of tools, running read-only actions autonomously and pausing for human approval before infrastructure changes. The post demonstrates the flow with a Lambda function whose 3-second timeout is raised to 30 seconds after a single approval.

  19. Lucas Beyer (bl16)XAI score36

    Reality Check: a public leaderboard for robot manipulation VLA models

    AILucas Beyer praises Reality Check, a new leaderboard for benchmarking VLA and related robot manipulation models. Half of its tasks are fully open, while the other half are held out to detect benchmaxxing by future model versions. The companion post from Nicolas Keller describes the launch as the first public robot manipulation benchmark, built on 14,400 real-world rollouts across four models.

  20. GitHub Copilot ChangelogOfficialAI score58

    GitHub Copilot local sandboxing now generally available across CLI, app, and VS Code

    AIGitHub has made local sandboxing for GitHub Copilot generally available in GitHub Copilot CLI, the GitHub Copilot app, and VS Code sessions using Agent Host. Sandboxes restrict the filesystem, network, and credentials that Copilot-initiated tools and commands can access, based on developer or organization policies. The feature is powered by Microsoft eXecution Container (MXC), supports Windows, macOS, and Linux, and is included at no additional cost.

  21. GitHub Copilot ChangelogOfficialAI score42

    GitHub Copilot CLI adds discovery of local Ollama models via /model

    AIGitHub Copilot CLI version 1.0.94-0 lets users run /model to discover supported models from a running local Ollama instance alongside configured and GitHub Copilot cloud models. Discovered models are not added automatically; users choose one, review its provider and endpoint, then confirm Add and use for this session or Add without switching, and models must support tool calling and streaming. Choosing a local model does not enable offline mode or disable GitHub telemetry, and COPILOT_OFFLINE=true remains a separate explicit setting.

  22. AWS Machine Learning BlogOfficialAI score32

    AWS playbook: six-week program closes AI builder gap for non-engineers

    AIAWS ran a six-week program pairing non-engineering professionals with mentors and tools like Amazon Bedrock AgentCore and the Strands Agents SDK to build working AI prototypes. Four participants with no engineering background built WealthWise, a multi-agent financial advisory tool with five agents on Amazon Nova models, which won first place. The article says participants who completed the phased program retained three times more practical skills than those in two-day intensive formats.

  23. 🚨 AI News | TestingCatalogXAI score20

    Google launches new Playground experiment in the US for paid users

    AIGoogle has rolled out a new Playground experiment in the US, letting paid-plan users generate content while all users can play games in its gallery. The post does not specify what kind of content can be generated or which plans qualify as paid.

    Video from @testingcatalog's post
  24. Wired · AINewsAI score46

    Pentagon's Tradewinds Program Uses Five-Minute Videos to Speed AI Purchases

    AIThe Department of Defense's Chief Digital and Artificial Intelligence Office runs the Tradewinds program, which grants "post-competitive" status to AI vendors based on videos of five minutes or less. That status can let government buyers use other transaction agreements and sometimes make awards in less than a week. OpenAI, Anthropic, and Google are listed as participants, though none commented.

  25. indigoXAI score34

    Grok Bot acts as a model router, using Gemini and Opus together

    AIThe poster says they already use Grok Bot as a model router, citing last weekend's personal agent livestream. In the demo, Gemini produced an infographic inside Grok Bot, and Claude Opus then checked the content. This follows Elon Musk's announcement that Grok Bot will use the best backend model for each task, including Claude Opus 5.5, MidJourney, and Suno.

    Video from @indigox's post
  26. WaymoOfficialAI score27

    Waymo releases framework for AV incident-management exercises and drills

    AIWaymo has introduced a first-of-its-kind framework for autonomous vehicle incident-management exercises, ranging from tabletop scenarios to full-scale drills. Adapted from emergency management best practices, it is designed to help AV developers, operational partners, and first responders test plans and strengthen coordination together.

    Image from @Waymo's post
  27. 🚨 AI News | TestingCatalogXAI score41

    Google releases Foresight macOS app using Gemma 4 for voice notes

    AIGoogle released the Google AI Edge Foresight app for macOS, powered by Gemma 4 and EmbeddingGemma 2. The app can connect to Google Drive to build a knowledge graph and, when transcription is active, uses local Gemma 4 E4B or Gemma 4 12B models to transcribe voice notes into new documents. EmbeddingGemma 2 is an open-weight, Apache 2 licensed 740M-parameter multimodal embedding model with an 8K context window.

    Video from @testingcatalog's post
  28. Stanford HAIOfficialAI score31

    Stanford HAI report examines which AI-for-good partnerships succeed or fail

    AIStanford HAI has released a new report examining partnerships in which funders have invested hundreds of millions of dollars to help nonprofits deploy AI for social good with frontier AI labs. Drawing on interviews, the report investigates who benefits from these partnerships and what drives their success or failure.

    Image from @StanfordHAI's post
  29. OpenRouterOfficialAI score34

    ElevenLabs joins OpenRouter with 9 TTS and 2 STT models

    AIElevenLabs is now available on OpenRouter, offering 9 Text to Speech models supporting up to 90+ languages and 2 Speech to Text models with speaker labels and word-level timestamps. All ElevenLabs models are 50% off on OpenRouter through October 19, 8am PT.

    Video from @OpenRouter's post