Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Jul 31

Jul 31Fri
  1. DeepSeekOfficialAI score42

    DeepSeek-V4-Flash API launches in public beta with stronger agent performance

    AIDeepSeek has released the official DeepSeek-V4-Flash API in public beta, with substantially upgraded agent capabilities. The company says its benchmark scores now far surpass those of V4-Pro-Preview. The official V4-Flash natively supports the Responses API format and is adapted for Codex, with configuration details in DeepSeek's API docs.

    Image from @deepseek_ai's post
  2. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. Thinking MachinesOfficialAI score38

    Thinking Machines' Inkling-Small gains performance per FLOP over Inkling

    AIThinking Machines says its Inkling-Small model delivers more performance per FLOP than Inkling on Terminal-Bench 2.1 agentic tool use, HLE reasoning, and IFBench instruction following. Variable thinking effort lets users choose their own point on the cost-performance curve.

Jul 29

Jul 29Wed
  1. Air Street PressBlogAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

    Why it matters: The piece ties Laguna S 2.1's open weights and published trajectories to Poolside's release cadence, showing how its model factory compounds gains across successive releases.

Jul 28

Jul 28Tue
  1. Tri DaoXAI score42

    Putting LLM brains on robots yields 4x SOTA gains without extra training

    AITri Dao reports that connecting an LLM as the "brain" to robot control policies quadruples state-of-the-art performance with no extra training. He says he was surprised by how well it works and expects agents running on robots to arrive soon. Background from a quoted post reports real-robot success rising from 16.7% to 97.3% and simulated LIBERO-PRO success from 12.8% to 53.3%.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score28

    LTM Partners with Cognition to Deploy Devin for Cybersecurity Risk Reduction

    AILTM has partnered with Cognition to deploy Devin, the AI software engineer, through BlueVerse RightLogic, a managed, outcome-based service that clears customers' vulnerability backlogs. RightLogic is designed to clear 80 percent of an enterprise's CVE backlog, up from the 60 percent previously delivered, and will focus first on banking, financial services, and insurance. The service is the first of five joint offerings the companies plan to bring to market.

  3. Rowan CheungXAI score40

    Zuckerberg says Meta's superintelligence lab should stay small and elite

    AIMark Zuckerberg said Meta's superintelligence lab should have 50 to 100 people who can keep the whole project in their heads at once. He said he personally recruits top AI researchers because underperformers have an outsized negative effect, and he rejects top-down deadlines and non-technical management layers.

    Video from @rowancheung's post
  4. JetBrains AI BlogOfficialAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 27

Jul 27Mon
  1. Sequoia CapitalBlogAI score24

    Cyera to Acquire Oasis Security to Combine Data and Identity Security for AI

    AICyera is joining with Oasis Security, which builds agentic access management for non-human identities such as API keys, service accounts, OAuth tokens, and agent credentials. The combination pairs Cyera's knowledge of where sensitive data lives with Oasis's visibility into which identities and agents can reach it. Sequoia Capital, which backed both companies since their Series A rounds, says the pairing covers the full path an AI agent takes through an enterprise.

  2. Kimi.aiOfficialAI score52

    Kimi K3 launches on Together AI as a Day 0 partner

    AIKimi K3 is now available on Together AI, which is a Day 0 launch partner for the model. Together AI offers developers immediate access to K3 through high-throughput inference aimed at coding agents and production workloads.

    Image from @Kimi_Moonshot's post
  3. Kimi.aiOfficialAI score38

    Kimi K3 now available on DigitalOcean Serverless Inference

    AIMoonshot AI's Kimi K3 is now available on DigitalOcean's Serverless Inference, letting developers start building in minutes. DigitalOcean describes K3 as supporting a 1M-token context, native vision, and multi-hour agentic tasks, and it is accessible through the Inference Router.

    Image from @Kimi_Moonshot's post
  4. Kimi.aiOfficialAI score38

    Kimi and kvcache-ai open-source AgentENV for scalable agent environments

    AIMoonshot AI's Kimi, in collaboration with kvcache-ai, has open-sourced AgentENV, a distributed system for running agent environments at scale. Its components power agentic RL training for Kimi K3, supporting fast snapshot, resume, and fork for large-scale parallel agent workflows. The project is available on GitHub at

Jul 26

Jul 26Sun
  1. Fireworks AI BlogOfficialAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

  2. Philipp SchmidBlogAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

  3. Berkeley AI ResearchOfficialAI score44

    Berkeley AI Research Trains LLMs to Update Beliefs for Long Tasks

    AIBerkeley AI Research introduces ABBEL, a framework that replaces full interaction histories with natural-language belief states that models update as new observations arrive. On CollabBench collaborative coding, belief grading closes about half the performance gap to full-context models while using fewer peak tokens and training in 50 steps instead of 100.

Jul 25

Jul 25Sat
  1. Ali GhodsiXAI score26

    Longer-running AI agents often perform worse than faster ones, says Ghodsi

    AIAli Ghodsi argues that AI agents which take longer to work through a task are often worse, while Genie reaches results faster. He adds that ontology will be key to giving agents the context they need to answer correctly and quickly. The related post reports that Genie Code outperformed three general-purpose coding agents on more than 400 real user data tasks.

  2. LangChain BlogOfficialAI score39

    What does it mean for companies to "own their intelligence" with AI?

    AILangChain Blog argues that companies need to own their AI intelligence rather than rely on generic models, because general models do not know company-specific policies, workflows, or risk tolerances. Ownership means controlling the agent system (model optionality, harness, and context), the economics, quality, and risk of AI work, and how intelligence compounds over time. The post uses an insurer's claims processing as an example of why off-the-shelf models fall short.

Jul 24

Jul 24Fri
  1. Alex AlbertXAI score34

    Opus 5 now produces consultant-grade spreadsheets and slide decks, Alex Albert says

    AIAlex Albert, of Anthropic, says Opus 5 now produces near-superhuman spreadsheets and slide decks that match what a consultant would make, just over six months after its predecessor. He also notes that finance professionals are reporting strong reactions to Claude for Excel, and he expects agentic progress seen in coding to extend to other fields in 2026.

    Video from @alexalbert__'s post
  2. Mike KriegerXAI score22

    Mike Krieger says models now build games from brief, dynamic prompts

    AIMike Krieger, who is associated with Anthropic, says two games were built from prompts of about four sentences that used dynamic /workflows extensively. He contrasts this with earlier in the year, when he relied on a bespoke harness and verification system, noting that current models accomplish much more with far less instruction.

  3. catXAI score66

    Claude Opus 5 released as strong option for long-running autonomous work

    AIAnthropic introduces Claude Opus 5 as a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price, according to the quoted announcement. The author, who works on the product, says Claude Opus 5 is great at long-running autonomous work and invites users to try it and share feedback.

    Why it matters: The post pairs a new model's long-running autonomous strength with a pricing claim, letting readers weigh capability against cost for agentic workloads.

  4. Mike KriegerXAI score46

    Mike Krieger says Claude Opus 5 became his daily driver

    AIAnthropic co-founder Mike Krieger says Claude Opus 5 has become his daily driver at work and on weekends. He reports it can work for hours on complex tasks and consistently gets to the bottom of tricky problems, and he has also built some games with it. Anthropic's announcement describes Opus 5 as close to the frontier intelligence of Fable 5 at half the price.

Jul 23

Jul 23Thu
  1. Matei ZahariaXAI score36

    Berkeley STAR Lab packages AI research optimizers into one GEPA API

    AIBerkeley's STAR Lab packaged multiple LLM-based "autoresearch" algorithms into a single API within the GEPA package, letting users mix and match them. The optimizers can be applied to tasks including prompt writing, agent design, and code optimization. The quoted thread adds that GEPA, AutoResearch, and Meta-Harness each win on different tasks, and that the new optimize_anything omni meta-optimizer beats every standalone optimizer at a matched budget.

  2. One Useful Thing (Ethan Mollick)BlogAI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

  3. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

  4. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Cognition Acquires The Interaction Company, Maker of the Poke Texting AI Agent

    AICognition has acquired The Interaction Company of California, the maker of Poke, a personal AI agent that texts users proactively and is approved to text natively on Apple Messages. Poke has exchanged more than 100 million messages in the last three months, and Poke users can keep using the product as before. Cognition says its models and infrastructure will make Poke faster and more reliable.

  5. Andrew NgXAI score65

    Andrew Ng announces OpenWorker, an open-source agent that delivers finished work

    AIAndrew Ng and Rohit Prasad announced OpenWorker, an open-source agent that produces deliverables such as documents, Slack messages, and calendar updates across files and everyday tools. It checks in before consequential actions, runs on Mac with Windows support coming soon, and works with user-supplied API keys for models including GPT 5.6 Sol, Claude Fable, Gemini 3.6, open-weight models, or local Ollama models. Source code is available on GitHub, and the tool requires the user's own API key.

    Video from @AndrewYNg's post

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

Jul 21

Jul 21Tue
  1. OpenAI NewsroomOfficialAI score14

    Ben Kalkman uses ChatGPT to plan a family treehouse build

    AIBen Kalkman, a father of nine, used ChatGPT as a construction consultant to help shape the structural design, source materials, and turn his kids' wild ideas into buildable plans for a family treehouse. The project continues to evolve on no-screen Sundays.

    Image from @OpenAINewsroom's post
  2. Eugene YanXAI score36

    Eugene Yan argues evals should weigh tail tasks, not median performance

    AIEugene Yan argues that model evals anchor on median tasks, but tail tasks determine project completion, making reliable models like Fable and Opus the difference between success and failure. He recommends treating models as collaborators who handle multi-hour or multi-day work with intent and success criteria, not as narrow-spec tools. Steve Yegge adds that Fable's carefulness is the dimension that matters most for production work.

    Image from @eugeneyan's post
  3. Soumith ChintalaXAI score45

    Soumith Chintala says Poolside's Laguna S 2.1 suits agentic work on DGX Spark

    AISoumith Chintala praised Poolside's Laguna S 2.1 as looking strong for agentic use and said it fits on a single NVIDIA DGX Spark. The quoted Poolside release describes it as a 118B total-parameter Mixture-of-Experts model with 8B active per token, up to 1M-token context, and thinking and no-thinking modes, with weights openly available under OpenMDW-1.1.

  4. Bryan CatanzaroXAI score57

    Poolside releases open-weight Laguna S 2.1 for agentic coding

    AIPoolside released Laguna S 2.1, an open-weight model with 118B total parameters and 8B active per token. The author says it performs strongly on agentic coding and long-horizon tasks, and it can run on a single NVIDIA DGX Spark. Weights are on Hugging Face under the OpenMDW-1.1 license, with access also available through OpenRouter and Poolside's API.

  5. JetBrains AI BlogOfficialAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  6. Andrej KarpathyXAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.