Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 21

Jul 21Tue
  1. JetBrains AI BlogAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  2. JetBrains AI BlogAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

Jul 20

Jul 20Mon

Jul 17

Jul 17Fri
  1. OpenAI NewsroomAI score22

    Foreguard, built with ChatGPT and Codex, helps families plan care and benefits early

    AISekhar and Katie Brandt built Foreguard, a free tool built with ChatGPT and Codex, to help families claim public benefits they are entitled to and plan private insurance coverage. The tool shows that modest budgets of $100 a month can create millions of dollars of day-one financial protection. The goal is to help families prepare earlier, before care decisions become urgent.

Jul 16

Jul 16Thu
  1. Google LabsAI score36

    NotebookLM Renamed Gemini Notebook, Same App, Expanded Role

    AIGoogle Labs announced that NotebookLM, which began as the Project Tailwind experiment, is now called Gemini Notebook. The app remains the same and keeps its mission of helping users learn faster, with the name change reflecting its role in Google's AI portfolio. Notebooks are already accessible in the Gemini app and are coming to Google Search, with folders promised as a future feature.

Jul 15

Jul 15Wed

Jul 13

Jul 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 12

Jul 12Sun

Jul 10

Jul 10Fri

Jul 9

Jul 9Thu

Jul 8

Jul 8Wed

Jul 7

Jul 7Tue

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)AI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.

Jul 1

Jul 1Wed
  1. Cognition Blog (Devin, Windsurf)AI score57

    Cognition launches Devin Security Swarm to find, verify, and patch vulnerabilities

    AICognition has launched Devin Security Swarm, which uses parallel agents to find vulnerabilities across a codebase, confirms exploitability in isolated sandboxes, and opens remediation PRs. In an evaluation on 50 real-world GitHub Security Advisory vulnerabilities, Devin reached 72% recall at about $90.23 per run, compared with 68% for Claude Security at $131.87 per run. The product is available starting today, with scan profiles and incremental scans that process only changed code after the first full baseline.

Jun 30

Jun 30Tue

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

Jun 28

Jun 28Sun
  1. PaddlePaddleAI score46

    PaddlePaddle announces Unlimited-OCR now runs in vLLM

    AIUnlimited-OCR, Baidu's long-context OCR model, now runs in vLLM, with a recipe provided for developers to try it. The background post says it parses entire books in one pass using Reference Sliding Window Attention (R-SWA), which keeps the KV cache fixed during decoding, and claims 35% faster throughput than DeepSeek-OCR at 6K output tokens.

Jun 27

Jun 27Sat
  1. PaddlePaddleAI score36

    PaddleFormers 1.2 adds DeepSeek-V4 training with 128K+ context support

    AIPaddleFormers 1.2 is released with support for training DeepSeek-V4 and 128K+ long-context training. The update adds Context Parallel, Packing, Document Mask Attention, and the Muon optimizer, plus ultra-fused mHC, CSA, and HCA operators, DeepEP/HybridEP communication, and lossless FP8 training with AutoSubbatch memory balancing. The project is presented as fully open-source and is available on GitHub.

Jun 25

Jun 25Thu

Jun 23

Jun 23Tue

Jun 19

Jun 19Fri

Jun 17

Jun 17Wed
  1. Jim FanAI score64

    ENPIRE lets Codex agents run autonomous research on a robot fleet

    AINVIDIA GEAR's ENPIRE gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks with minimal human oversight. The author reports tasks such as tying zip-ties, organizing fine pins, and installing GPUs, and a faster time-to-solution with eight parallel robots than with fewer. Safety uses a kinematic limit that resets a robot leaving its envelope, a torque-limited gripper, and a frozen reward function classifier. The team says everything will be open-sourced.

Jun 16

Jun 16Tue
  1. Xiaomi MiMoAI score38

    Xiaomi launches MiMo Claw, an agent integrated with Kingsoft Office

    AIXiaomi has launched MiMo Claw, an agent built on its flagship MiMo model and integrated with Kingsoft Office for Word, Excel, PowerPoint, and PDF workflows. The company says it consumes 40–60% fewer tokens than comparable solutions, and daily usage has been expanded from 1 hour to 4 hours, with free access and no deployment required. A limited-time subscription is priced at ¥14.9 per month.

Jun 12

Jun 12Fri

Jun 11

Jun 11Thu
  1. OpenRouter BlogAI score74

    OpenRouter Fusion panels beat individual models on the DRACO deep research benchmark

    AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.

    Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.