Skip to contentSkip to stories

Updated

#Agent

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 25

Jul 25Sat
  1. Ali GhodsiXAI score26

    Longer-running AI agents often perform worse than faster ones, says Ghodsi

    AIAli Ghodsi argues that AI agents which take longer to work through a task are often worse, while Genie reaches results faster. He adds that ontology will be key to giving agents the context they need to answer correctly and quickly. The related post reports that Genie Code outperformed three general-purpose coding agents on more than 400 real user data tasks.

  2. LangChain BlogOfficialAI score39

    What does it mean for companies to "own their intelligence" with AI?

    AILangChain Blog argues that companies need to own their AI intelligence rather than rely on generic models, because general models do not know company-specific policies, workflows, or risk tolerances. Ownership means controlling the agent system (model optionality, harness, and context), the economics, quality, and risk of AI work, and how intelligence compounds over time. The post uses an insurer's claims processing as an example of why off-the-shelf models fall short.

Jul 24

Jul 24Fri
  1. Alex AlbertXAI score34

    Opus 5 now produces consultant-grade spreadsheets and slide decks, Alex Albert says

    AIAlex Albert, of Anthropic, says Opus 5 now produces near-superhuman spreadsheets and slide decks that match what a consultant would make, just over six months after its predecessor. He also notes that finance professionals are reporting strong reactions to Claude for Excel, and he expects agentic progress seen in coding to extend to other fields in 2026.

    Video from @alexalbert__'s post
  2. Mike KriegerXAI score22

    Mike Krieger says models now build games from brief, dynamic prompts

    AIMike Krieger, who is associated with Anthropic, says two games were built from prompts of about four sentences that used dynamic /workflows extensively. He contrasts this with earlier in the year, when he relied on a bespoke harness and verification system, noting that current models accomplish much more with far less instruction.

  3. catXAI score66

    Claude Opus 5 released as strong option for long-running autonomous work

    AIAnthropic introduces Claude Opus 5 as a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price, according to the quoted announcement. The author, who works on the product, says Claude Opus 5 is great at long-running autonomous work and invites users to try it and share feedback.

    Why it matters: The post pairs a new model's long-running autonomous strength with a pricing claim, letting readers weigh capability against cost for agentic workloads.

  4. Mike KriegerXAI score46

    Mike Krieger says Claude Opus 5 became his daily driver

    AIAnthropic co-founder Mike Krieger says Claude Opus 5 has become his daily driver at work and on weekends. He reports it can work for hours on complex tasks and consistently gets to the bottom of tricky problems, and he has also built some games with it. Anthropic's announcement describes Opus 5 as close to the frontier intelligence of Fable 5 at half the price.

Jul 23

Jul 23Thu
  1. Matei ZahariaXAI score36

    Berkeley STAR Lab packages AI research optimizers into one GEPA API

    AIBerkeley's STAR Lab packaged multiple LLM-based "autoresearch" algorithms into a single API within the GEPA package, letting users mix and match them. The optimizers can be applied to tasks including prompt writing, agent design, and code optimization. The quoted thread adds that GEPA, AutoResearch, and Meta-Harness each win on different tasks, and that the new optimize_anything omni meta-optimizer beats every standalone optimizer at a matched budget.

  2. One Useful Thing (Ethan Mollick)BlogAI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.

  3. BAAI · new models on Hugging FaceOfficialAI score62

    BAAI releases AREX-Base, a 122B deep research agent model

    AIBAAI has released AREX-Base, a 122B-total, 10B-activated Mixture-of-Experts deep research agent built on Qwen3.5-122B-A10B with a 262,144-token context. The model uses an inner research loop and an outer self-improvement loop, and the source reports it scoring 82.5 on BrowseComp and 85.4 on GAIA, under Apache 2.0.

    Why it matters: The release pairs a 122B-parameter deep research agent with benchmark tables against frontier and open models, letting readers compare its search-agent results directly.

  4. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Cognition Acquires The Interaction Company, Maker of the Poke Texting AI Agent

    AICognition has acquired The Interaction Company of California, the maker of Poke, a personal AI agent that texts users proactively and is approved to text natively on Apple Messages. Poke has exchanged more than 100 million messages in the last three months, and Poke users can keep using the product as before. Cognition says its models and infrastructure will make Poke faster and more reliable.

  5. Andrew NgXAI score65

    Andrew Ng announces OpenWorker, an open-source agent that delivers finished work

    AIAndrew Ng and Rohit Prasad announced OpenWorker, an open-source agent that produces deliverables such as documents, Slack messages, and calendar updates across files and everyday tools. It checks in before consequential actions, runs on Mac with Windows support coming soon, and works with user-supplied API keys for models including GPT 5.6 Sol, Claude Fable, Gemini 3.6, open-weight models, or local Ollama models. Source code is available on GitHub, and the tool requires the user's own API key.

    Video from @AndrewYNg's post

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

Jul 21

Jul 21Tue
  1. Eugene YanXAI score36

    Eugene Yan argues evals should weigh tail tasks, not median performance

    AIEugene Yan argues that model evals anchor on median tasks, but tail tasks determine project completion, making reliable models like Fable and Opus the difference between success and failure. He recommends treating models as collaborators who handle multi-hour or multi-day work with intent and success criteria, not as narrow-spec tools. Steve Yegge adds that Fable's carefulness is the dimension that matters most for production work.

    Image from @eugeneyan's post
  2. Soumith ChintalaXAI score45

    Soumith Chintala says Poolside's Laguna S 2.1 suits agentic work on DGX Spark

    AISoumith Chintala praised Poolside's Laguna S 2.1 as looking strong for agentic use and said it fits on a single NVIDIA DGX Spark. The quoted Poolside release describes it as a 118B total-parameter Mixture-of-Experts model with 8B active per token, up to 1M-token context, and thinking and no-thinking modes, with weights openly available under OpenMDW-1.1.

  3. Bryan CatanzaroXAI score57

    Poolside releases open-weight Laguna S 2.1 for agentic coding

    AIPoolside released Laguna S 2.1, an open-weight model with 118B total parameters and 8B active per token. The author says it performs strongly on agentic coding and long-horizon tasks, and it can run on a single NVIDIA DGX Spark. Weights are on Hugging Face under the OpenMDW-1.1 license, with access also available through OpenRouter and Poolside's API.

  4. JetBrains AI BlogOfficialAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  5. Andrej KarpathyXAI score30

    Karpathy suggests long voice rambles help LLMs understand your intent

    AIAndrej Karpathy describes using /voice to ramble for about 10 minutes, sometimes as a short interview, to give an LLM context that would be tedious to type. He says LLMs reconstruct these messy streams of thought remarkably well, often returning a cleaner version than the speaker started with, which improves shared understanding and reduces later corrections.

  6. Rowan CheungXAI score34

    Frontier AI models raise growing cybersecurity challenges, Demis warns

    AIRowan Cheung says AI models pushing the frontier are creating a growing challenge for cybersecurity. Quoting Demis, he reports that security must be addressed alongside the agentic era, with cyber worries about some models being just the beginning. Demis suggests this may be the time to push for standards and international cooperation.

    Video from @rowancheung's post
  7. koray kavukcuogluXAI score72

    Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

    AIGoogle introduces Gemini 3.6 Flash as its workhorse model, with better coding, knowledge work, and multimodal performance while reducing token usage. It also launches Gemini 3.5 Flash-Lite, described as the fastest and most cost-effective 3.5-class model for high-throughput applications, and 3.5 Flash Cyber, a version of 3.5 Flash fine-tuned to find and fix cybersecurity vulnerabilities.

    Why it matters: The post lists three distinct models, each aimed at a different job, so readers can map which one fits coding, high-volume, or security workloads.

    Video from @koraykv's post
  8. JetBrains AI BlogOfficialAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

Jul 20

Jul 20Mon

Jul 18

Jul 18Sat

Jul 16

Jul 16Thu
  1. Soumith ChintalaXAI score60

    Kimi K3 launches as a 2.8 trillion parameter open-weight model

    AIMoonshot AI announced Kimi K3, a native multimodal model with 2.8 trillion parameters and a 1 million token context window. The announcement cites up to 6.3x faster decoding in million-token contexts and about 25% higher training efficiency, and says open weights arrive by July 27, 2026. The author, Soumith Chintala, reposted it with a brief note of congratulations.

Jul 15

Jul 15Wed
  1. Sequoia CapitalBlogAI score23

    Sable Builds Aidan, an AI Employee That Runs Customer Calls

    AISable has created "Aidan," an AI employee that leads customer calls using vision, voice, video and real-time browser interaction. Frontier companies including Notion and Decagon are already using Sable to explain their products, and more than 150 companies are on its waitlist. Sequoia Capital led Sable's seed round and co-led its Series A.

  2. Sequoia CapitalBlogAI score32

    Bunkerhill Health's Carebricks Lets Health Systems Deploy AI Agents for Patient Care

    AIBunkerhill Health's Carebricks platform lets health systems create and deploy AI agents across clinical and operational use cases using data hospitals already generate. Sequoia Capital backed the company at seed and is continuing to invest. At UTMB Health, Bunkerhill grew from one agent in production to more than twenty, consolidating multiple vendors' point solutions into one platform.

Jul 14

Jul 14Tue
  1. Cognition Blog (Devin, Windsurf)OfficialAI score44

    Cognition Marks One Year Since Windsurf Merger With Devin and SWE Model Gains

    AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.

Jul 13

Jul 13Mon
  1. AI Snake OilBlogAI score57

    Narayanan argues AI job change will unfold over decades, not with one model release

    AIArvind Narayanan's ICML keynote argues that AI's labor impact will depend on slow organizational adaptation rather than a single lab milestone. He cites reliability measurements showing agent accuracy rose much faster than reliability over the last 24 months, and points to software engineering and past technologies like electricity and ATMs. He concludes that evaluation work and human judgment will become more central as building tasks are increasingly automated.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 10

Jul 10Fri

Jul 9

Jul 9Thu
  1. Fidji SimoXAI score62

    OpenAI launches ChatGPT Work, an agent powered by Codex and GPT-5.6

    AIOpenAI introduced ChatGPT Work, a new agent inside ChatGPT powered by Codex and GPT-5.6. The quoted announcement says it can take action across apps and files, stay with a project for hours if needed, and turn a goal into finished work. Fidji Simo's own post adds that the team has worked to make Chat more agentic for a while.

  2. Meta AI BlogOfficialAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.