Skip to contentSkip to stories

Updated

All AI news

Apr 17

Apr 17Fri

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

  2. BAAIAI score40

    ClawKeeper v1.0 released as open-source security framework for AI agents

    AIBAAI has released ClawKeeper v1.0, an open-source security framework for AI agents built around OpenClaw. It combines Skill-based command-level policies, Plugin-based runtime monitoring, and an independent Watcher that intervenes against high-risk operations such as prompt injections, key leaks, rogue commands, and remote code execution, even if the agent is compromised.

Apr 10

Apr 10Fri

Apr 9

Apr 9Thu

Apr 8

Apr 8Wed
  1. Stability AIAI score44

    Stability AI launches Brand Studio, a creative production platform built around brand identity

    AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    AIAnthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    Why it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

Apr 6

Apr 6Mon

Apr 2

Apr 2Thu
  1. Werner VogelsAI score25

    AWS launches Sustainability console for carbon emissions measurement

    AIAWS has introduced the AWS Sustainability console, which puts emissions data in the hands of users without exposing billing data. The post frames carbon and cost as similar signals that must be measured to be optimized. It links to AWS's announcement covering programmatic access, configurable CSV reports, and Scope 1–3 reporting in one place.

Apr 1

Apr 1Wed
  1. Jim FanAI score62

    CaP-X open-sources agentic robotics toolkit, benchmark, and RL setup

    AIJim Fan announced the open-source release of CaP-X, an agentic robotics framework in which LLM-driven agents control robot arms and humanoids through perception and actuation APIs. The release includes CaP-Gym with 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR, and CaP-Bench, which evaluates 12 frontier LLMs and VLMs across 8 tiers. The post also reports that a 7B open-source model rose from 20% to 72% success after 50 RL training iterations, with synthesized programs transferring to real robots.

Mar 31

Mar 31Tue
  1. Intern Large ModelsAI score52

    Intern Large Models unveils Kernel-Smith for generating GPU kernels and operators

    AIIntern Large Models introduced Kernel-Smith, a framework for generating high-performance GPU kernels and operators using an evolutionary agent and post-training recipe. The post says it outperforms Gemini-3.0-pro and Claude-4.6-opus on Kernel-Bench, and that optimized kernels have been merged into SGLang and LMDeploy. The accompanying figure compares best program score trajectories across evolution steps, with Kernel-Smith-235B-RL reaching the highest peak.

Mar 21

Mar 21Sat

Mar 19

Mar 19Thu
  1. Cognition Blog (Devin, Windsurf)AI score50

    Devin can now schedule recurring sessions that carry state between runs

    AIDevin can now schedule its own recurring sessions from a plain-language description, such as running a weekly feature-flag cleanup every Monday at 9am. Devin keeps its own notes across runs, so each scheduled session builds on earlier results rather than starting over. The feature can also be combined with Managed Devins to run parallel recurring tasks, such as a weekly QA pass reported to Slack.

Mar 18

Mar 18Wed
  1. Cognition Blog (Devin, Windsurf)AI score72

    Devin can now break tasks down and run a team of managed Devins

    AIDevin can now break large tasks into scoped pieces and delegate them to a team of managed Devins that run in parallel. Each managed Devin runs in its own isolated virtual machine with its own terminal, browser, and development environment, and has its own session link. The main coordinator session monitors progress, resolves conflicts, and compiles results, and managed Devins are available now for all users.

    Why it matters: The post explains how a coordinator session splits work across isolated managed sessions, giving readers a concrete pattern for running agent tasks in parallel.

Mar 13

Mar 13Fri
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score44

    Fun-CineForge Releases Open-Source Dubbing Pipeline, Model, and CineDub-CN Dataset

    AIFun-CineForge, from FunAudioLLM, is an open-source toolkit with an end-to-end dataset pipeline and an MLLM-based model for zero-shot movie dubbing across diverse cinematic scenes. The team built CineDub-CN, described as the first large-scale Chinese television dubbing dataset, and reports that its model outperforms state-of-the-art methods on audio quality, lip-sync, timbre transition, and instruction following. Inference code and checkpoints were released on March 16, 2026, and the model runs on a consumer-grade GPU.

Mar 10

Mar 10Tue

Mar 5

Mar 5Thu
  1. Nick TurleyAI score62

    GPT-5.4 Thinking rolls out to ChatGPT with mid-response interrupts

    AIGPT-5.4 Thinking is rolling out to ChatGPT, and users can now interrupt it before it produces the final answer. Users can steer the response while it is still working rather than sending multiple follow-up turns. The update also improves deep web research and long-context reasoning, which the post says helps specific questions arrive faster and stay focused.

Mar 4

Mar 4Wed

Mar 2

Mar 2Mon

Feb 26

Feb 26Thu

Feb 24

Feb 24Tue
  1. Cognition Blog (Devin, Windsurf)AI score46

    Cognition Launches Cognition for Government to Modernize Federal Software With Devin and Windsurf

    AICognition launched Cognition for Government on February 25, 2026, offering its Devin autonomous software engineering agent and Windsurf AI IDE to modernize U.S. government legacy systems. Devin, available in AWS GovCloud with a FedRAMP High version forthcoming, can complete migrations 5-40x faster than human engineers, while Windsurf is the only FedRAMP High AI IDE and holds DoD IL4/5/6 accreditation.

  2. Replit BlogAI score43

    Replit Pro launches at $100/month as Core drops to $20/month

    AIReplit launched a $100/month Pro plan with Turbo Mode, pooled credits for up to 15 builders, and priority support, while cutting Core from $25 to $20 per month and letting it invite up to 5 collaborators. The Teams plan is being sunset, with Teams users automatically upgraded to Pro at no additional cost for the rest of their term. Economy and Power Modes for Agent are available on all paid plans.

Feb 23

Feb 23Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 2.2 adds desktop testing, self-review autofix, and 3x faster startup

    AICognition released Devin 2.2, which gives Devin full access to its own Linux desktop so it can launch and test desktop applications, not just browser-based web apps. Devin can also plan, code, review its own output, and fix issues before opening a PR, and it now starts up 3x faster. New users get $10 in free credits, and Desktop support is enabled by default for new sessions as of February 24, 2026.

Feb 20

Feb 20Fri
  1. Jim FanAI score75

    DreamDojo: Open-source world model trained on 44K hours of human video

    AIJim Fan announced DreamDojo, an open-source interactive world model that takes robot motor controls and generates future frames in pixels. It is pre-trained on 44K hours of human egocentric video using latent actions, then post-trained onto specific robot hardware, and a real-time version runs at 10 FPS for live teleoperation, policy evaluation, and model-based planning. The author reports a +17% real-world success gain on a fruit packing task, and weights, code, datasets, and the whitepaper are released.

Feb 12

Feb 12Thu

Feb 11

Feb 11Wed

Feb 9

Feb 9Mon
  1. Cognition Blog (Devin, Windsurf)AI score43

    Devin Can Now Autofix Review Comments from Devin Review and Other Bots

    AICognition has configured Devin to automatically autofix incoming review comments from Devin Review and other PR review bots, as well as lint and CI/CD issues. Devin resolves flagged problems and feeds the fixes back into the pull request without human intervention for mechanical fixes. Users can select which bots Devin responds to in Settings > Customization > Autofix settings.

Feb 4

Feb 4Wed

Feb 2

Feb 2Mon

Jan 29

Jan 29Thu
  1. Chip HuyenAI score18

    Chip Huyen launches GoodAIList.com to track trending open-source AI repos

    AIChip Huyen built GoodAIList.com, which tracks 14K open-source AI repositories with contributions from over 145K developers. Each day it searches for new repos using 123 keywords and topics, surfaces those gaining traction, and categorizes them with AI-generated annotations that she notes are not highly accurate. The site also maps contributor locations, which she uses to find people doing interesting AI work when she travels.

Jan 22

Jan 22Thu
  1. BAAIAI score38

    BAAI releases RoboCOIN, a large bimanual robot manipulation dataset

    AIBAAI's RoboCOIN is a bimanual robot dataset with more than 180,000 trajectories across 421 tasks, collected from 15 robot platforms in 16 real-world scenarios. Its three-tier annotations at trajectory, segment, and frame levels help robots learn both what to do and how to do it. The post says integrating these annotations raised success rates on complex tasks by up to 50% for models such as π₀.

Jan 20

Jan 20Tue
  1. Cognition Blog (Devin, Windsurf)AI score54

    Cognition launches Devin Review to help humans review AI-generated code

    AICognition introduced Devin Review, a free early-release code review tool that works on any public or private GitHub PR, with features for organizing diffs, chatting about changes, and flagging AI-detected bugs. The company says code review, not code generation, is now the bottleneck as coding agents increase the volume and size of pull requests.

Jan 19

Jan 19Mon
  1. Factory NewsAI score47

    Factory Introduces Agent Readiness to Score Codebases for Autonomous Coding Agents

    AIFactory's new Agent Readiness tool evaluates repositories across eight technical pillars and five maturity levels, using 60+ binary criteria run via the /readiness-report command. The company says it can also open pull requests to fix foundational gaps such as missing AGENTS.md files, linter configuration, and pre-commit hooks. Factory says scores are now more consistent, with variance dropping from an average of 7% to 0.6%.