Skip to contentSkip to stories

Updated

All AI news

Jun 17

Jun 17Wed
  1. PromptArmor Threat IntelligenceAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

Jun 12

Jun 12Fri

Jun 9

Jun 9Tue

Jun 8

Jun 8Mon

May 30

May 30Sat
  1. Xiaomi MiMoAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. Cognition Blog (Devin, Windsurf)AI score62

    Devin Tests Its Own Code Changes in the Cloud and Returns Proof

    AICognition describes autonomous testing in Devin, where the agent writes a source-grounded test plan, operates the app through computer use, and returns labeled screenshots and an annotated video. Login steps are handled by a deterministic testing skill, and the company says test runs approved per day more than doubled in recent months. Known limits include timing errors with transient UI elements and models sometimes triggering states through JavaScript instead of clicking the interface.

    Why it matters: The post explains how computer use, test plans, deterministic login scripts, and annotated recordings let Devin verify its own code changes end to end.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens

    AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.

    Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 21

May 21Thu
  1. Tri DaoAI score44

    Transformers reduce to GEMM-plus-epilogue, enabling LLM-written fast kernels

    AITri Dao says that after a mathematical rewrite, all transformer operations can be expressed as a series of GEMMs with epilogues. Given a few optimized primitives, LLMs and novice humans can write near speed-of-light kernels for transformer ops. The related CODA work fuses memory-bound surrounding ops into the matmul epilogue, and LLMs can also write CODA kernels approaching speed-of-light.

May 18

May 18Mon
  1. Eugene YanAI score62

    Cloudflare outlines an eight-stage agent harness for vulnerability discovery

    AIEugene Yan shares Cloudflare's description of a vulnerability discovery harness that runs eight stages, from reconnaissance to report writing. The pipeline uses about 50 concurrent agents to hunt for bugs, independent agents to try to disprove findings, and a trace step to confirm whether attacker input reaches each bug. Reachable findings feed back into new hunt tasks before a report is written against a predefined schema.

May 16

May 16Sat
  1. Ahead of AI (Sebastian Raschka)AI score62

    Recent LLM architecture changes that cut long-context KV cache and attention cost

    AISebastian Raschka reviews recent open-weight LLM architecture changes aimed at reducing long-context memory and compute costs. He covers KV sharing and per-layer embeddings in Gemma 4, per-layer query-head budgeting in Laguna XS.2, Compressed Convolutional Attention in ZAYA1-8B, and mHC with CSA/HCA compressed attention in DeepSeek V4. The article reports that DeepSeek V4-Pro uses 27% of single-token inference FLOPs and 10% of the KV cache size of DeepSeek V3.2 at a 1M-token context.

May 10

May 10Sun
  1. Cognition Blog (Devin, Windsurf)AI score39

    Devin Automates HIL/SIL Failure Triage and Scales Test Generation at Automotive Firms

    AICognition reports that deploying its Devin agent on hardware-in-the-loop and software-in-the-loop workflows cut failure triage time and multiplied test generation at automotive customers. One team reclaimed 2K–4K engineering hours monthly across about 4,000 tickets, while RV Tech rose from 1–2 to 10–15 generated tests per day. Devin also helps convert bottlenecked HIL tests into SIL equivalents to catch failures earlier.

May 3

May 3Sun

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

Apr 17

Apr 17Fri

Apr 16

Apr 16Thu

Apr 4

Apr 4Sat
  1. Andrej KarpathyAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. Andrej KarpathyAI score49

    Karpathy shares an LLM-maintained personal knowledge base workflow

    AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.

Mar 24

Mar 24Tue
  1. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

Mar 23

Mar 23Mon
  1. Anthropic EngineeringAI score78

    Anthropic shows a three-agent harness for long-running app development

    AIAnthropic's Labs team describes a three-agent harness with planner, generator, and evaluator agents for building full-stack applications over multi-hour autonomous coding sessions. The evaluator uses Playwright to test the running app against sprint contracts, and a retro game maker built with the harness worked end to end where a single-agent run's core feature did not. The author later removed the sprint construct and kept only the components still needed on Opus 4.6.

    Why it matters: The post shows how a generator-evaluator loop, with explicit grading criteria and a tuned QA agent, turned a solo run's broken output into a working app, and how the harness was pruned as models improved.

Mar 1

Mar 1Sun
  1. Artificial IgnoranceAI score46

    Build Your Own Benchmark: Why Public AI Evals Are Saturating and What Replaces Them

    AIPublic AI benchmarks such as MMLU, SWE-bench Verified, and GPQA Diamond are saturating or showing contamination, prompting OpenAI to call SWE-bench Verified "no longer suitable" in late February and recommend SWE-bench Pro. OpenAI's audit found 59.4% of the problems its best model failed had flawed test cases, and GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash could reproduce original fixes from memory. The article argues that behavioral tests, such as Vending-Bench's simulated vending machine business, may be more useful for everyday model choice.

Feb 27

Feb 27Fri

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)AI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 22

Feb 22Sun
  1. Artificial IgnoranceAI score62

    Harness engineering emerges as a playbook for managing coding agents

    AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 19

Feb 19Thu

Feb 6

Feb 6Fri

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

Jan 20

Jan 20Tue
  1. Anthropic EngineeringAI score67

    Anthropic redesigns its performance engineering take-home as Claude models improve

    AIAnthropic's performance engineering lead Tristan Hume describes how a take-home test for hiring performance engineers was repeatedly defeated by successive Claude models. Claude Opus 4 outperformed most human applicants within the 4-hour limit, and Claude Opus 4.5 matched the best candidates in 2 hours. Anthropic is releasing the original take-home as an open challenge, with the best known Claude result at 1487 cycles.

    Why it matters: The post traces how each Claude model defeated the take-home test, showing concrete redesign tradeoffs for evaluating engineers when AI assistance is available.

Jan 13

Jan 13Tue
  1. Tim DettmersAI score36

    Tim Dettmers Argues Agents Should Automate Most Personal Work, Not Just Code

    AITim Dettmers, a professor who has used Claude Code for eight months to automate his own work, argues that more than 90% of code and text should be written by agents. He says the coding-focused hype on Twitter overstates parallel sessions and autonomy, which translate poorly to most non-software tasks. The post offers a balanced guide to what actually works in agent-based automation.

Oct 27, 2025

Oct 27, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score36

    Devin Automates .NET Framework to .NET Core Migration in Weeks, Not Months

    AICognition says its autonomous coding agent Devin can complete a .NET Framework to .NET Core migration in as little as two weeks, using a Strangler Fig approach adapted from Jimmy Bogard's guide. The post says Devin handles planning via Ask Devin and DeepWiki, dependency sharing, controller and view conversion, and session state adaptation through a remote app.

Sep 28, 2025

Sep 28, 2025Sun
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition rebuilds Devin around Claude Sonnet 4.5 for 2x speed

    AICognition rebuilt its Devin coding agent for Claude Sonnet 4.5, reporting 2x faster performance and 12% better results on its Junior Developer Evals, now available in Agent Preview. The team found the model is aware of its context window, which led to premature wrap-up behavior that they countered with repeated prompts and a 200k usage cap within a 1M token beta.

    Why it matters: The post explains which agent behaviors changed under Sonnet 4.5, such as context-window awareness and note-taking, that forced a rebuild rather than a simple model swap.