Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 5

Oct 5Mon
  1. Liquid AIAI score37

    Liquid AI's d1 decision model adds vision, rivaling GPT-6.1 Sol at lower cost

    AILiquid AI released d1 with vision support, accepting images, text, or both as inputs. In tests on six real applications, d1 matched or beat GPT-6.1 Sol on four while costing 19x to 200x less than both GPT-6.1 Sol and Claude Opus 5.5. It returns probabilities for yes/no, choice, or score questions in one forward pass, with text decisions in 200 to 300 ms.

    Image from @liquidai's post
  2. Liquid AIAI score36

    Liquid AI's d1 model inspects parts from camera images with 85-97% accuracy

    AILiquid AI's vision-enabled decision model d1 inspects parts directly from camera images and is described as the best such model currently on the market. It reaches 85% to 97% accuracy across four VisA inspection tasks covering circuit boards, candles, cashews, and chewing gum. It understands each task from a short description without task-specific training.

    Video from @liquidai's post
  3. Understanding AI (Timothy B. Lee)AI score62

    Agent swarms may be the next scaling law, but speed may matter more than capability

    AIThe article examines whether multi-agent swarms could become a new scaling law, comparing them with inference scaling from o1. OpenAI researcher Noam Brown said its models are now sometimes trained with other agents, while the cited Anthropic data suggests gains beyond 10 agents are smaller and mainly speed-related. The article also raises the risks of groupthink and misaligned agents, and it notes that a Microsoft Research and UC Berkeley paper found teams sometimes solved tasks solo agents could not.

  4. Liquid AI · new models on Hugging FaceAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.

  5. Google AIAI score46

    Gemma 4 and BOTANIC-1 pinpoint crop-yield DNA mutations in minutes

    AILiving Models paired Google's Gemma 4 with BOTANIC-1, a plant-DNA model trained on 320 species, to identify causal genetic variants. In a melon yield test, the pipeline ranked the target mutation first out of 2,494 possibilities in under four minutes. The approach aims to speed up breeding of climate-resilient crops that would otherwise take years of field trials.

  6. Elad GilAI score40

    Era launches free simulated enterprises for testing AI agents

    AIEra, launched by Ofir Ehrlich's team, generates a complete simulated company spanning Salesforce, Slack, Jira, Zendesk, Gong, and Deel, plus cloud databases and storage. Agents interact with it through live MCP and API interfaces, and because Era generated the company, it knows the exact ground truth for testing and benchmarking. The post says the product is live today and free.

  7. X.PINAI score50

    Huawei and Qualcomm reach multiyear cross-licensing patent deal

    AIHuawei's new multiyear patent agreement with Qualcomm would make Qualcomm the net payer for the first time, according to Nikkei Asia. The deal cross-licenses patents in 5G, computing, AI, and networking, and Qualcomm will also buy some of Huawei's U.S. patents outright. Financial terms have not been disclosed, and the transaction still requires regulatory approval.

  8. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  9. IEEE Spectrum · AIAI score49

    Human Oversight of AI Agents Could Fail as Approval Processes Push People Out

    AIResearchers Avijit Ghosh, Margaret Mitchell, and Samir Passi argue in a September 6 arXiv paper that current human-in-the-loop designs for AI agents push humans out of meaningful oversight. They say agents are tuned for speed, accuracy, and volume, overwhelming reviewers, and recommend adding friction, such as requiring users to state their own choice first, to counter automation bias and fatigue.

  10. O'Reilly RadarAI score38

    Zero to Agent in 30 Minutes: Building Your First Agent with MCP

    AIBruce Hopkins shows how to wrap an existing stock-data REST API, the Twelve Data API, in a Model Context Protocol (MCP) server so an MCP client can discover and call it. The demo uses Python with FastMCP, exposing current and historical stock-price functions as tools and resources with descriptive prompts. Developers can add an MCP interface around existing capabilities without replacing their underlying application logic.

  11. MIT Technology Review · AIAI score30

    Enterprise AI agents need organizational knowledge to reach production, survey finds

    AIA survey of 300 data, AI, and technology executives found only 34% of organizations' agentic AI projects reach production, with legacy systems, security concerns, and missing knowledge context as main obstacles. Production leaders, who advance 61% of projects beyond pilot, show stronger semantic knowledge capabilities. Most firms plan to invest in retrieval pipelines, AI-ready APIs, retrieval-augmented generation, and knowledge graphs.

  12. SantiagoAI score34

    Utah approves Nolla Health's AI app to issue acne prescriptions

    AINolla Health has reportedly become the first U.S. organization to receive regulatory approval for an AI system to issue initial prescriptions, starting with acne treatment in Utah. The app scans a user's face, asks a few questions, creates a personalized plan, prescribes medication when needed, and tracks progress over time. Users also have access to a physician at no extra cost.

  13. Tibor BlahoAI score62

    OpenAI adds opt-in text watermarking for API and EU ChatGPT and Codex output

    AIOpenAI is rolling out text watermarking for EU AI Act compliance, with opt-in access for API customers globally on select models starting today. Watermarking stays off by default in the API, while an invisible watermark will be added to eligible ChatGPT and Codex text in the European Union over the coming weeks. Access to the text watermark detector is initially limited to approved researchers and expert organizations, and the image and audio verification tools remain publicly accessible.

    Image from @btibor91's post
  14. SantiagoAI score47

    Tool generates synthetic companies to test AI agents across business systems

    AIA tool can turn a one-line business description into a complete synthetic company spread across CRM, ticketing, Slack, files, emails, and call recordings. Developers can test agents against this connected data, then reset the company to its initial state and rerun the test when something breaks. The background post describes the product as Era, a free simulated enterprise that connects to Salesforce, Slack, Jira, Zendesk, Gong, and Deel through live MCP and API interfaces.

  15. Karl's AI WattsAI score23

    Karl's AI Watts shares a full AI Skills workflow tutorial

    AIKarl's AI Watts publishes the AI workflow he previously shared internally at Tim Studio, covering finding Skills, packaging experience into Skills, combining them into workflows, and batching and scheduling them. The post says viewers could build a local batch video-editing Skill and an end-to-end content pipeline spanning copy, posters, video, and web pages.

    Video from @aiwarts's post