Skip to contentSkip to stories

Updated

All AI news

Oct 5

Oct 5Mon
  1. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  2. IEEE Spectrum · AIAI score49

    Human Oversight of AI Agents Could Fail as Approval Processes Push People Out

    AIResearchers Avijit Ghosh, Margaret Mitchell, and Samir Passi argue in a September 6 arXiv paper that current human-in-the-loop designs for AI agents push humans out of meaningful oversight. They say agents are tuned for speed, accuracy, and volume, overwhelming reviewers, and recommend adding friction, such as requiring users to state their own choice first, to counter automation bias and fatigue.

  3. O'Reilly RadarAI score38

    Zero to Agent in 30 Minutes: Building Your First Agent with MCP

    AIBruce Hopkins shows how to wrap an existing stock-data REST API, the Twelve Data API, in a Model Context Protocol (MCP) server so an MCP client can discover and call it. The demo uses Python with FastMCP, exposing current and historical stock-price functions as tools and resources with descriptive prompts. Developers can add an MCP interface around existing capabilities without replacing their underlying application logic.

  4. MIT Technology Review · AIAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    AIA survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

  5. SantiagoAI score34

    This is the first AI-led prescription and treatment app on the market (at least that I know of).

    AIIt helps treat acne. And it's happening here in the United States (Utah)! • You scan your face and answer a few questions • The app gives you a personalized plan • It gives you a prescription if necessary • It tracks your progress and adjusts over time You also have access to a physician at no extra cost.

  6. DeedyAI score13

    Menlo has tried to rebuild from the ground up how we think about investing and we are fortunate to be lead investors in 5 of the top 15…

    AI…consumer AI apps by monthly revenue despite our low volume strategy. Thanks for the shout, Shaun! Have a ton of admiration for what you have done at Sequoia. Too many people know you for your firebrand tweets and too few know your excellent investments and that you have a PhD in physics from Caltech (and it’s fun to have your cofounder be a partner!)

  7. Tibor BlahoAI score62

    OpenAI adds opt-in text watermarking for API and EU ChatGPT and Codex output

    AIOpenAI is rolling out text watermarking for EU AI Act compliance, with opt-in access for API customers globally on select models starting today. Watermarking stays off by default in the API, while an invisible watermark will be added to eligible ChatGPT and Codex text in the European Union over the coming weeks. Access to the text watermark detector is initially limited to approved researchers and expert organizations, and the image and audio verification tools remain publicly accessible.

  8. SantiagoAI score47

    This is a very comprehensive way to test agents: You describe a business (any business, in one line), and this app generates a complete…

    AI…synthetic company and writes it across multiple systems: CRM, tickets, Slack, files, emails, call recordings, etc. You can then test your agents against all this connected data and see how they behave. If something breaks, you can reset the company to the initial state and start again.

  9. Karl's AI WattsAI score23

    Chinese creator shares AI Skills workflow for batch video editing and content production

    AIThe post, by 卡尔的AI沃茨, publicly shares an AI workflow previously presented internally at 影视飓风, covering finding suitable Skills, packaging experience into Skills, combining Skills into workflows, and batching and scheduling them. The author says this lets readers build a local batch video-editing Skill and a full-chain content pipeline spanning copy, posters, video, and web pages.

  10. Clément DelangueAI score62

    Hugging Face turns 10 coding harnesses into RL environments via a capture proxy

    AIHugging Face says a capture proxy lets reinforcement learning train open models inside unmodified coding harnesses such as Claude Code, Codex, and OpenCode. The proxy records the exact token IDs and logprobs vLLM samples and hands them to TRL for training. On LFM2.5-2.6B, training in four harnesses at once raised OpenCode results from 34% to 58%, while SFT on 3,189 Qwen3.8-27B rollouts plateaued at 47.5%.

  11. Guillermo RauchAI score44

    gdp-ts brings compile-time authorization proofs to TypeScript APIs

    AIGuillermo Rauch introduced gdp-ts, a library, linter, and AI skill that uses "proofs" to enforce that sensitive functions are called only after an authorization check. The TypeScript typechecker verifies these proofs at compile time, aiming to stop security bugs from shipping, including those written by AI agents. The README models a Vercel API constraint requiring a role and entitlement proof to change a Project's password.

  12. DatabricksAI score31

    Databricks makes IP Functions generally available for network analytics in SQL

    AIDatabricks has made IP Functions generally available, letting users parse, validate, and join IPv4 and IPv6 addresses and CIDR blocks with built-in SQL functions optimized in Photon. In benchmarks versus another leading cloud data warehouse, CIDR joins ran up to 3.1x faster and cost up to 6.4x less. The functions support its Security Lakehouse vision for threat detection, investigation, and network analytics on one governed copy of data.

  13. a16z NewsAI score62

    Consumer AI usage is broad but paid use is concentrated, a16z ranking finds

    AIa16z's seventh Top 100 Consumer AI Apps report adds a spending ranking based on YipitData card panels, showing usage is wide but shallow. Only 4.5% of U.S. consumers had an active paid personal subscription to ChatGPT, Gemini, or Claude as of August, while the top 1% of payers accounted for 19.5% of observed consumer AI spend. The report also notes ChatGPT still leads, Claude has moved into the third position, and personal agents are emerging as a possible new monetization path.

  14. Baidu Inc.AI score13

    "Let the agent handle it" sounds simple, but there's a full AI stack working behind the scenes.

    AIIn our latest edition of AI Pulse, we explore: • Why the full stack matters for useful, cost-efficient AI • What productivity, commerce and industrial agents need from the technology behind them • How those needs help shape the stack itself And there's plenty more, from our new podcast AI, Evolving to Miaoda upgrades and Kooko AI, our all-in-one AI workspace. Get the details ↓

  15. Exponential ViewAI score36

    AI Helps Self-Represented Litigants Argue Cases, Including an Australian Win

    AIIn Australia, computing academic Greg Baker used AI to challenge his employer's refusal to make his casual job permanent, and the Fair Work Commission ruled in his favor. In England and Wales, 60% of defendants in defended county court claims this year had no lawyer, and in the US more than nine in ten consumers sued for debt face cases without one. Around 0.6% of all Claude use in May was for lawyers' tasks, with four-fifths of those queries from people asking about their rights or what the law means.

  16. MIT Technology Review · AIAI score20

    Predictive analytics moves toward autonomous, agentic AI decision making in enterprises

    AIEnterprises are shifting from backward-looking analytics to forward-looking predictive systems that can act on their own conclusions, according to Everest Group partner Vishal Gupta. The source credits deep learning and generative AI with enabling real-time model training and the use of unstructured data alongside numerical records. Gupta says the word "analytics" is giving way to AI.

  17. PyTorch BlogAI score24

    PyTorch's Accelerator Working Group Standardizes Hardware Backend Integration in H1 2026

    AIThe PyTorch Accelerator Integration Working Group released updates on its H1 2026 progress toward standardizing how new hardware connects to the framework. Key workstreams include the Cross-Repository CI Relay (CRCR), which automatically reports downstream backend test results to a shared dashboard, and refactored test suites that decouple PyTorch's 600,000-plus tests from specific accelerators.