Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Apr 24

Apr 24Fri
  1. Mckay WrigleyXAI score25

    Wrigley says GPT-5.5 leads coding while Claude leads agents

    AIMckay Wrigley says his coding split moved from 80/20 Claude/GPT to 80/20 GPT/Claude within three months, and he trusts GPT 5.5 for engineering. He still finds Claude better for non-coding agent work, calls Opus 4.7 underwhelming, and attributes Anthropic's issues to compute constraints.

  2. Ahmad Al-DahleXAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Apr 22

Apr 22Wed
  1. Factory NewsOfficialAI score38

    Factory's Automated QA Skill Tests Apps Like Real Users and Posts Reports to PRs

    AIFactory has released an Automated QA skill that drives an app as a real user would, filling forms, typing into terminals, and calling endpoints, then posts a structured report with screenshots, terminal snapshots, and API traces as a single updating comment on each pull request. Teams can run it on every push or make it an optional CI check triggered by a PR label, comment command, or manual dispatch, and developers can run /qa locally in any Droid session. Automated QA is available today in all Factory plans.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Xiaomi MiMoOfficialAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

  3. Michael TruellXAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

  4. NVIDIA AI DeveloperOfficialAI score29

    NVIDIA OpenShell v0.0.34 adds live sandbox policy updates and VM installs

    AINVIDIA's OpenShell v0.0.34 release lets users update sandbox policy without restarting the runtime. The update also adds install-vm, which installs the gateway and VM driver with new --driver-dir support, and sandbox get, which shows the active runtime policy. Supervisor seccomp improvements and HTTP normalization are included as well.

Apr 20

Apr 20Mon
  1. NVIDIA AI DeveloperOfficialAI score35

    OpenShell v0.0.33 adds hardened sandboxing and a standalone libkrun driver

    AINVIDIA released OpenShell v0.0.33, which adds seccomp and process-limit hardening, inference routing, and a standalone libkrun compute driver. The libkrun driver provides a lightweight VM backend as a second compute path for agents. The release also includes bug fixes, a docs refresh, and improved test stability.

Apr 17

Apr 17Fri

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

  2. BAAIOfficialAI score40

    ClawKeeper v1.0 releases open-source security framework for OpenClaw AI agents

    AIBAAI announces ClawKeeper v1.0, an open-source security framework for OpenClaw AI agents, combining Skill-based command policies, Plugin-based runtime monitoring, and a Watcher system-level observer. The independent Watcher is designed to block high-risk operations such as prompt injections, key leaks, rogue commands, and remote code execution, even if the agent is compromised. The paper is available on arXiv and the project code is hosted on GitHub.

Apr 9

Apr 9Thu
  1. Andrej KarpathyXAI score45

    Karpathy says AI capability gap stems from uneven use and training

    AIAndrej Karpathy argues that people judging AI from free-tier ChatGPT or Advanced Voice Mode miss the strong capabilities of current agentic models like OpenAI Codex and Claude Code. He says gains are "peaky," concentrated in verifiable technical domains like programming and math that suit reinforcement learning and attract B2B investment, while writing and everyday advice improve less. Those who use frontier agentic tools professionally in these fields see far greater capability, which is why the two groups talk past each other.

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

  2. Stability AIOfficialAI score44

    Stability AI launches Brand Studio, a creative production platform built around brand identity

    AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringOfficialAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    AIAnthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    Why it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

Apr 6

Apr 6Mon
  1. Z.ai Release NotesOfficialAI score34

    Z.ai's GLM-5.3 and GLM-5.2 Lead Open-Source Coding and Long-Context Models

    AIZ.ai's GLM-5.3 delivers a 50% coding gain over GLM-5.2 on Z.ai Code Bench, reaching open-source state-of-the-art on public benchmarks including Terminal Bench 3.0. GLM-5.3-Flash uses 320B total parameters with 18B activated, combining linear and sparse attention to reduce compute and KV-cache needs. GLM-5.2 supports a 1M lossless context window for long-horizon tasks.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score44

    Windsurf releases SWE-1.6, a software engineering model optimized for speed and user experience

    AIWindsurf has made SWE-1.6, its model for software engineering agents, generally available, with the company saying it improves on the SWE-1.6 Preview by reducing overthinking, looping, and sequential tool calls. The model is free for three months, with a free version offered at 200 tok/s through Fireworks and a faster paid version at 950 tok/s through Cerebras.

Apr 4

Apr 4Sat
  1. Andrej KarpathyXAI score43

    Karpathy praises Farzapedia, a personal wiki for AI agents

    AIAndrej Karpathy highlights Farzapedia, a personal wiki Farza built from 2,500 diary entries, Apple Notes, and iMessage conversations, as a good example of his proposed LLM-wiki approach. He argues this file-based memory is explicit, user-owned, interoperable, and usable with any AI tool, letting users control how AI knows them.

  2. Andrej KarpathyXAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Apr 2

Apr 2Thu
  1. Andrej KarpathyXAI score49

    Karpathy shares an LLM-maintained personal knowledge base workflow

    AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.

  2. AI Futures ProjectBlogAI score62

    AI Futures Project shortens Automated Coder timelines to mid 2028

    AIAI Futures Project moved Daniel Kokotajlo's Automated Coder median from late 2029 to mid 2028 and Eli's from early 2032 to mid 2030. The main reasons cited are a faster METR time horizon doubling time and the impressive results of Claude Opus 4.6. The authors also say progress in agentic coding has been faster than expected over the past 3 to 5 months.

Apr 1

Apr 1Wed
  1. Jim FanXAI score62

    CaP-X open-sources agentic robotics toolkit, benchmark, and RL setup

    AIJim Fan announced the open-source release of CaP-X, an agentic robotics framework in which LLM-driven agents control robot arms and humanoids through perception and actuation APIs. The release includes CaP-Gym with 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR, and CaP-Bench, which evaluates 12 frontier LLMs and VLMs across 8 tiers. The post also reports that a 7B open-source model rose from 20% to 72% success after 50 RL training iterations, with synthesized programs transferring to real robots.

    Video from @DrJimFan's post

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceOfficialAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

Mar 30

Mar 30Mon
  1. Mckay WrigleyXAI score22

    AI tools may soon use, clone, and extend any software autonomously

    AIMckay Wrigley predicts AI tools will within 6-12 months autonomously use any software, clone it in a weekend, monitor it for updates, and add custom features. He frames this as a future where users never need to operate their computer themselves. The prediction follows a referenced Claude Code update adding computer use in research preview for Pro and Max plans.

Mar 27

Mar 27Fri

Mar 26

Mar 26Thu
  1. Andrej KarpathyXAI score47

    Karpathy wants agents to handle full app DevOps from one command

    AIAndrej Karpathy argues that the hardest part of building a deployed app is not the code but the DevOps work of assembling services, API keys, payments, auth, and deployment. He says the goal is for agents to handle this entire lifecycle as code, with agent-native CLI and API access instead of manual web clicking. He calls it a from-scratch redesign that is only now barely technically possible.

Mar 25

Mar 25Wed

Mar 24

Mar 24Tue
  1. ARC PrizeOfficialAI score70

    ARC Prize announces ARC-AGI-3, an interactive benchmark for frontier agents

    AIARC Prize has released ARC-AGI-3, a set of hundreds of interactive, turn-based environments with thousands of game-style levels, with no instructions or stated goals. Humans score 100% while frontier AI scores 0.51%. ARC Prize 2026 offers over $2 million in prizes for open-source solutions to ARC-AGI-2 and ARC-AGI-3.

    Why it matters: The benchmark's human versus frontier AI gap and its interactive design show how agent evaluation is shifting from instruction-following toward exploration and adaptation.

  2. Anthropic EngineeringOfficialAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

  3. Jim FanXAI score62

    Jim Fan warns that compromised LiteLLM package shows risks for AI agents

    AIJim Fan reposted a report that LiteLLM PyPI release 1.82.8 was compromised and contained a litellm_init.pth file that sends credentials to a remote server and self-replicates. He argues agents make this worse, since files like skills, configs, or PDFs read into context could spread malicious instructions. He concludes that agentic frameworks need guardrails and audited tooling.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringOfficialAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

  2. Artificial IgnoranceBlogAI score20

    AI Agents Now Read Documentation and Create Dashboard Cells More Than Humans Do

    AIHex CEO Barry McCardel posted a graph showing AI agents now create more Hex cells than humans do. Mintlify launched an analytics feature in February to track AI agent traffic to documentation, saying agents may read docs more often than humans. The article argues that writers should consider AI systems as a primary audience.

Mar 21

Mar 21Sat