Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)BlogAI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 25

Jun 25Thu
  1. OpenAI NewsroomOfficialAI score43

    OpenAI paper previews how Codex agents may reshape future work

    AIOpenAI's Economic Research team published a new paper on the shift from chat to delegation, where people hand longer, more complex work to agents. The post draws on Codex usage inside OpenAI as a preview of what agentic work may look like in the future.

Jun 23

Jun 23Tue
  1. Eugene YanXAI score34

    Eugene Yan praises Claude Tag's multiplayer thread for team input

    AIEugene Yan says he likes Claude Tag's multiplayer form factor because other people can reply in the thread to give Claude context and direction. Claude Tag, per the background post, lets teams in Slack tag Claude in as a team member with access to chosen channels and tools.

Jun 20

Jun 20Sat
  1. BAAIOfficialAI score13

    BAAI roundtable explores frontier models, self-evolution, and world models

    AIAt BAAI's "Reconstructing the World — Large Model Summit," President Wang Zhongyuan joined Zhu Jun, Luo Fuli, Liu Zhiyuan, An Bo, and other industry leaders in a roundtable. They discussed frontier model capability evolution, AI self-evolution, multimodality, and world models, with a focus on bridging the digital and physical worlds.

Jun 18

Jun 18Thu
  1. Andrew NgXAI score15

    DeepLearning.AI launches course on adding voice to AI agents

    AIDeepLearning.AI has launched a course, taught by VocalBridge CEO Ashwyn, on adding voice to AI agents and applications. It covers building voice agents that are both reliable and fast, with three projects: a voice-interactive game, an agent that gains a voice in about 10 lines of code, and an agent that places outbound calls via a make_phone_call function.

    Video from @AndrewYNg's post

Jun 17

Jun 17Wed
  1. PromptArmor Threat IntelligenceOfficialAI score62

    PromptArmor shows Codex auto-review agent approved malware install via prompt injection

    AIPromptArmor demonstrated that OpenAI's Approve-for-me agent approved a malicious NPM install with elevated privileges after a hidden prompt injection in an external GitHub issue influenced the main Codex agent. The malicious package's post-install script then ran unsandboxed with the user's full privileges. The report also gives steps for organizations to disable agentic auto-review in Claude Code and Codex.

    Why it matters: The report shows a prompt-injected GitHub issue leading an approval agent to permit a malicious NPM install, a concrete test of agent-in-the-loop guardrails.

  2. Jim FanXAI score64

    ENPIRE lets Codex agents run autonomous research on a robot fleet

    AINVIDIA GEAR's ENPIRE gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks with minimal human oversight. The author reports tasks such as tying zip-ties, organizing fine pins, and installing GPUs, and a faster time-to-solution with eight parallel robots than with fewer. Safety uses a kinematic limit that resets a robot leaving its envelope, a torque-limited gripper, and a frozen reward function classifier. The team says everything will be open-sourced.

    Video from @DrJimFan's post

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogOfficialAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    AIOpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    Why it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

  2. Jim FanXAI score62

    Jim Fan's ENPIRE lets Codex agents run autonomous research on robot fleets

    AIJim Fan introduces ENPIRE, which gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks autonomously. The post reports that the system can tie zip-ties, organize fine pins, and install GPUs, and that eight robots exploring in parallel improve faster than fewer. The team plans to open-source everything.

    Video from @DrJimFan's post
  3. Xiaomi MiMoOfficialAI score38

    Xiaomi launches MiMo Claw, an agent integrated with Kingsoft Office

    AIXiaomi has launched MiMo Claw, an agent built on its flagship MiMo model and integrated with Kingsoft Office for Word, Excel, PowerPoint, and PDF workflows. The company says it consumes 40–60% fewer tokens than comparable solutions, and daily usage has been expanded from 1 hour to 4 hours, with free access and no deployment required. A limited-time subscription is priced at ¥14.9 per month.

  4. Michael TruellXAI score62

    SpaceX acquires Cursor in all-stock deal to build AI models

    AISpaceX has exercised its option to acquire Cursor in an all-stock transaction, with the stated goal of building the world's most useful AI models. Over the past few months, SpaceXAI has been jointly training a model with Cursor, which will be released in Cursor and Grok Build soon.

    Why it matters: The deal combines Cursor's coding tools with SpaceX's AI team, and a jointly trained model is due to ship in Cursor and Grok Build.

  5. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.2 with 1M-token context and MIT open-source license

    AIZ.ai has released GLM-5.2, its flagship model for long-horizon tasks, which it says substantially improves on GLM-5.1 and supports a 1M-token context. The model adds IndexShare, which cuts per-token FLOPs by 2.9× at 1M context, and is released under the MIT open-source license.

    Why it matters: The source gives benchmark tables against named rival models and deployment settings, useful for judging where GLM-5.2 sits among current flagship models.

Jun 15

Jun 15Mon
  1. Z.ai Release NotesOfficialAI score62

    Z.ai Release Notes: GLM-5.2 Adds 1M Lossless Context for Long Tasks

    AIZ.ai's release notes list GLM-5.2 as supporting 1M lossless context, with improved long-horizon task performance and reduced context drift and goal forgetting. The company says GLM-5.2 achieves open-source SOTA performance on coding and long-horizon task benchmarks. The page also includes the newer GLM-5.3 and GLM-5.3-Flash entries, which are listed above GLM-5.2.

    Why it matters: The page lists a dated series of Z.ai model releases, showing how the coding and long-horizon agent line has evolved from GLM-4.5 through GLM-5.2.

  2. BAAIOfficialAI score22

    Turing Award winners Diffie and Barto keynote BAAI Conference on AI security and RL

    AITuring Award winners Whitfield Diffie and Andrew Barto delivered keynotes at the BAAI Conference on AI security and reinforcement learning. Diffie argued that today's feedback-based approach only patches programs after they fail, and that formal methods offer a path to substantially more reliable intended behavior. Barto framed reinforcement learning around control, search, and associative memory, describing its core insight as caching search results rather than searching continuously.

    Image from @BAAIBeijing's post

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 12

Jun 12Fri

Jun 11

Jun 11Thu
  1. OpenRouter BlogOfficialAI score74

    OpenRouter Fusion panels beat individual models on the DRACO deep research benchmark

    AIOpenRouter introduced Fusion, a tool that sends a prompt to a panel of models and has a judge model fuse their results into one answer. On 100 DRACO deep research tasks, a Fable 5 and GPT-5.5 panel scored 69.0%, above Fable 5 alone at 65.3%, and a budget panel of Gemini 3 Flash, Kimi K2.6, and DeepSeek V4 Pro reached 64.7% at about half the cost of Fable 5.

    Why it matters: The source gives benchmark scores, panel compositions, and contamination controls, letting readers judge how much of the gain comes from model diversity versus self-synthesis.

  2. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Jun 10

Jun 10Wed
  1. Zed BlogOfficialAI score48

    Zed Unveils DeltaDB, Version Control Built Around Agent Conversations Instead of Commits

    AIZed is building DeltaDB, a version control system that records every operation as a fine-grained delta, linking agent conversations to the code they produce. The company says a beta will arrive in a few weeks, and it invites users to join a waitlist. The system is designed so teammates can collaborate on work in progress without waiting for commits, pull requests, or pushes.

  2. Xiaomi MiMoOfficialAI score67

    Xiaomi releases open-source MiMo Code V0.1 terminal coding assistant

    AIXiaomi MiMo has released MiMo Code V0.1, an open-source AI coding assistant for the terminal under the MIT license. It ships with MiMo V2.5, a multimodal model offered free for a limited time with a million-token context window. The tool automatically loads existing Claude Code skills, MCP servers and commands, and reuses API configuration, and it supports providers including Anthropic, OpenAI, DeepSeek, Kimi and GLM.

    Why it matters: The post specifies MiMo Code's Claude Code compatibility and MIT license, which bear directly on whether existing coding-agent setups can migrate without rework.

    Image from @XiaomiMiMo's post
  3. Xiaomi MiMoOfficialAI score82

    MiMo Code open-sources a terminal coding agent for long-horizon tasks

    AIXiaomi's MiMo team released MiMo Code, an MIT-licensed terminal coding agent built on OpenCode for long-horizon programming tasks. The design centers on three areas: Max Mode parallel sampling that generates five candidates per turn, Goal-based completion verification, and a memory system that checkpoints session state and rebuilds context. The article reports offline benchmark results and a double-blind A/B test with 1,213 pairs in which MiMo Code's win rate exceeded 65% beyond 200 execution steps.

    Why it matters: The article explains how MiMo Code handles long-horizon coding through computation, checkpointed memory, and cross-session evolution, useful for judging design tradeoffs in coding agents.

Jun 9

Jun 9Tue
  1. One Useful Thing (Ethan Mollick)BlogAI score72

    Ethan Mollick tests Claude 5 Fable and finds it runs long projects with little user input

    AIEthan Mollick, who had early access to Claude 5 Fable, reports that it outperformed other public models in his tests, including an isochrone travel-time map and a nine-and-a-half-hour software build called Concord. He says the model delegated work to other agents and made many design choices he could not see or weigh in on, leaving him closer to a client than a hands-on operator. He also notes high token usage, frequent fallback to Claude 4.8 Opus under security guardrails, and persistent quirks in its writing style.

    Why it matters: The author's hands-on tests show how much work the model now completes without user steering, which shapes how people should think about their role with AI tools.

Jun 5

Jun 5Fri
  1. Michael TruellXAI score34

    Cursor's vision: agents you collaborate with like a colleague

    AIMichael Truell says working with agents should feel like collaborating with a colleague, not just exchanging text chats. He envisions interacting with them through gestures on a shared screen and live conversation. The post builds on Cursor's Design Mode, which lets users point, draw, or talk to update a UI.

  2. BAAIOfficialAI score20

    BAAI's 8th Conference set for June 12–13 in Beijing

    AIThe 8th BAAI Conference will be held June 12–13 at the Zhongguancun Innovation Center in Beijing, with Turing Award laureates and China's large-model leaders attending. Core focuses include world models and agents, plus two new flagship sessions on AI-native education and the token economy. The event will feature 25 forums, over 200 speeches, and a first-ever on-site AI agent conference companion for real-time listening and summarization.

    Image from @BAAIBeijing's post

Jun 4

Jun 4Thu
  1. Cohere · new models on Hugging FaceOfficialAI score60

    Cohere releases North Mini Code 1.0, a 30B-A3B open-weights coding model

    AICohere and Cohere Labs released North Mini Code 1.0, an open-weights 30B-A3B mixture-of-experts model for code generation and agentic terminal tasks, under Apache 2.0. The model has 256K context and 64K max output, and is trained for tool use. Its benchmark table lists Terminal-Bench v2 at 36.0, SWE-Bench Verified at 67.6, and LiveCodeBench v6 at 70.3, below Qwen3.6 on several tasks.

    Why it matters: The card lists benchmark results against Qwen3.6, Gemma4, and other models, showing where North Mini Code trails on some coding and agentic tasks.

  2. One Useful Thing (Ethan Mollick)BlogAI score44

    Ethan Mollick Announces Co-Existence, a Sequel Book on Working Alongside AI

    AIEthan Mollick is releasing Co-Existence on October 20, a new book about working with AI systems that are sometimes, but not always, better than humans. The book follows his 2024 title Co-Intelligence, which he says was written about an era of chatbots rather than autonomous agents. Mollick also reports writing every chapter draft himself while using AI readers and fact-checkers, and building the book's website with Claude Code using Opus 4.8.

  3. Varun MohanXAI score42

    Antigravity enables /teamwork-preview for parallel subagent teams

    AIAntigravity has enabled the /teamwork-preview command for all paid plans, letting users run parallel implementation and verification agents on complex tasks. The post says the team built a working OS with it, but warns that it can consume a large number of tokens.

Jun 3

Jun 3Wed
  1. PaddlePaddleOfficialAI score31

    Baidu CoBuddy, a free code-focused model, now live on Novita

    AIBaidu CoBuddy is now available for free on Novita AI, a code-focused model aimed at developers and AI agents. It offers a 131K context window, up to 65K output tokens, and native tool calling. The model is served through Novita's serverless API for high-throughput, low-latency inference.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score60

    Cognition launches $10M AI Productivity Guarantee for enterprise Devin customers

    AICognition introduced the AI Productivity Guarantee, under which it will issue credits up to $10M if Devin delivers less engineering value than enterprise customers pay for. The company uses an AI estimator to measure hours of productive output, validated against engineers' own estimates of how long the same work would have taken by hand. Value is converted to dollars at a standard global rate and compared against each customer's consumption near the end of the annual contract.

    Why it matters: The post explains how Cognition estimates Devin's output in hours and backs the estimate with a $10M credit commitment, a concrete model for measuring AI vendor value.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition Estimates Engineering Hours Saved by Its Devin Coding Agent

    AICognition built an automated agent that classifies Devin sessions as productive and estimates the human engineering hours each one would have taken. On 233 held-out sessions the estimator reached an rlog of 0.74, with individual errors often 2 to 3 times in either direction but roughly unbiased in aggregate. The system is calibrated to underestimate and is currently running with Devin customers.

    Why it matters: The post shows how the measurement design, from hours-based metrics to conservative calibration, determines whether agent productivity estimates can be trusted in aggregate.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceOfficialAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

Jun 1

Jun 1Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score50

    Cognition launches Devin Desktop, the next generation of Windsurf

    AICognition has announced Devin Desktop, the next generation of Windsurf, which makes the Agent Command Center the default IDE surface for managing local and cloud agents, PRs, and context. Spaces let related agents share context, and Agent Client Protocol (ACP) support lets any ACP-compatible agent run alongside Devin. The IDE remains fully backwards-compatible with Windsurf, including editor extensions, keybindings, LSPs, and terminal workflows.