Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Jul 20

Jul 20Mon

Jul 19

Jul 19Sun

Jul 18

Jul 18Sat

Jul 16

Jul 16Thu
  1. Soumith ChintalaAI score60

    Kimi K3 launches as a 2.8 trillion parameter open-weight model

    AIMoonshot AI announced Kimi K3, a native multimodal model with 2.8 trillion parameters and a 1 million token context window. The announcement cites up to 6.3x faster decoding in million-token contexts and about 25% higher training efficiency, and says open weights arrive by July 27, 2026. The author, Soumith Chintala, reposted it with a brief note of congratulations.

Jul 15

Jul 15Wed
  1. Sequoia CapitalAI score32

    Bunkerhill Health's Carebricks Lets Health Systems Deploy AI Agents for Patient Care

    AIBunkerhill Health's Carebricks platform lets health systems create and deploy AI agents across clinical and operational use cases using data hospitals already generate. Sequoia Capital backed the company at seed and is continuing to invest. At UTMB Health, Bunkerhill grew from one agent in production to more than twenty, consolidating multiple vendors' point solutions into one platform.

Jul 14

Jul 14Tue
  1. Cognition Blog (Devin, Windsurf)AI score44

    Cognition Marks One Year Since Windsurf Merger With Devin and SWE Model Gains

    AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.

Jul 13

Jul 13Mon
  1. AI Snake OilAI score57

    Narayanan argues AI job change will unfold over decades, not with one model release

    AIArvind Narayanan's ICML keynote argues that AI's labor impact will depend on slow organizational adaptation rather than a single lab milestone. He cites reliability measurements showing agent accuracy rose much faster than reliability over the last 24 months, and points to software engineering and past technologies like electricity and ATMs. He concludes that evaluation work and human judgment will become more central as building tasks are increasingly automated.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

  3. Cognition Blog (Devin, Windsurf)AI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 12

Jul 12Sun
  1. OpenAI NewsroomAI score12

    James Costello uses ChatGPT to run his demolition business

    AIStructural engineer James Costello, who oversees complex New York City high-rise demolitions, uses ChatGPT to review lengthy contracts, organize compliance documents, and create construction plans for his family-rooted firm DEMTEC. The post says the tool helps him move through these workflows faster and with more confidence, freeing time to grow the business and support his team.

    Image from @OpenAINewsroom's post

Jul 10

Jul 10Fri

Jul 9

Jul 9Thu
  1. Meta AI BlogAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Michael TruellAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

  2. Meta AI BlogAI score75

    Meta launches Muse Image, an agentic image model with search and code tools

    AIMeta Superintelligence Labs has released Muse Image, which can invoke search and coding tools and self-refine its generations before output. It is available today in the Meta AI app, meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, with Facebook coming soon. Meta also previewed Muse Video, which is coming soon to creators and Meta AI and is reported as ranking No. 3 on Arena for text-to-video at the time of writing.

    Why it matters: The source describes how search, code execution, and self-refinement change image generation, which matters to anyone comparing agentic media models with plain prompt-to-image systems.

Jul 5

Jul 5Sun
  1. ARC PrizeAI score47

    ARC Prize Awards First ARC-AGI-3 Milestone Prize to Tufa Labs' Open-Source Agent

    AITufa Labs won the first $37.5K ARC-AGI-3 milestone prize with "The Duck," a small open-source LLM that plays the games by writing and running Python in a live REPL. Reki placed second with a vision-language agent using Gemma-4-31B, and md Boktiar Mahbub Murad placed third with the "forge" framework. The second and final milestone prize ends September 30.

Jul 4

Jul 4Sat

Jul 3

Jul 3Fri
  1. Lil'Log (Lilian Weng)AI score62

    Lilian Weng surveys harness engineering as a path to recursive self-improvement

    AIThe post argues that the system surrounding a base model, called the harness, increasingly determines how well AI agents deploy and improve. It reviews research where harness components such as workflows, context, and code are optimized automatically through evolutionary search and meta-agent loops. The author concludes that evaluators, memory management, and human oversight remain open bottlenecks.

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)AI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.

Jul 1

Jul 1Wed
  1. PromptArmor Threat IntelligenceAI score58

    Copilot Cowork Skills Still Reach DeepSeek After Admin Opt-Out

    AIPromptArmor reports that Skills in Microsoft Copilot Cowork can call DeepSeek even when an organization has not opted into the DeepSeek Preview. The calls use the agent's own access path, so users need no API key, and a Skill built this way received a 100/100 score from Microsoft's Skill Scanner. After Microsoft removed the DeepSeek Preview setting on June 25, the report says admins had no remaining setting to block DeepSeek through the Cowork code environment, leaving disabling Cowork entirely as the only option.

  2. Cognition Blog (Devin, Windsurf)AI score57

    Cognition launches Devin Security Swarm to find, verify, and patch vulnerabilities

    AICognition has launched Devin Security Swarm, which uses parallel agents to find vulnerabilities across a codebase, confirms exploitability in isolated sandboxes, and opens remediation PRs. In an evaluation on 50 real-world GitHub Security Advisory vulnerabilities, Devin reached 72% recall at about $90.23 per run, compared with 68% for Claude Security at $131.87 per run. The product is available starting today, with scan profiles and incremental scans that process only changed code after the first full baseline.

  3. Jim FanAI score51

    Jim Fan introduces ASPIRE, a self-evolving robot skills library for continual learning

    AIJim Fan announces ASPIRE, a system where coding agents use multimodal sensory traces from simulation and real robots to run evolutionary search over control programs and add the results to a growing skills library. The post claims up to a roughly 10x reduction in transfer learning tokens for sim-to-real and single-arm to bimanual transfer, and says the full stack will be open-sourced.

Jun 30

Jun 30Tue
  1. One Useful Thing (Ethan Mollick)AI score62

    Ethan Mollick argues AI is shifting from chatbots to long-running agents

    AIMollick argues AI capability is improving at a better-than-exponential rate, citing METR, GDPval, Epoch, and his own tests showing models working autonomously for hours. He says work is shifting from co-working with chatbots to assigning tasks to agents, with OpenAI workers managing multiple agents and experts getting the most from them. He adds that open-weights Chinese models trail the American frontier by roughly 6-12 months.

  2. Jim FanAI score60

    ASPIRE lets robots build an evolving skills library that transfers across tasks

    AIJim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.

    Video from @DrJimFan's post
  3. Andrew NgAI score50

    Andrew Ng outlines three loops for building 0-to-1 AI products

    AIAndrew Ng describes three loops he uses to build 0-to-1 products with AI agents: an agentic coding loop, a developer feedback loop, and an external feedback loop. He says the agentic coding loop runs every few minutes, letting coding agents build, test, and iterate on software for around an hour without human intervention. The developer feedback loop operates over tens of minutes to hours, with humans steering product decisions because they hold a context advantage over AI systems.

    Image from @AndrewYNg's post

Jun 29

Jun 29Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition's Devin Fusion routes coding work between two models to cut cost

    AICognition has released a preview of Devin Fusion, a multi-model harness that runs a frontier main agent alongside a cheaper sidekick agent. On FrontierCode 1.1 Extended, the company reports scores near frontier models at up to 60% lower cost per task, and 41% lower cost when paired with Fable 5, which access was suspended from June 12, 2026.

    Why it matters: The post explains a sidekick architecture with cached persistent contexts, which contrasts with advisor-style tools and shows how cost cuts depend on the main model's delegation behavior.

Jun 27

Jun 27Sat
  1. Ahead of AI (Sebastian Raschka)AI score37

    Local Coding Agents: Setting Up Qwen3.6 with Open-Source Harnesses

    AISebastian Raschka's tutorial shows how to build a fully local coding agent by pairing an open-weight LLM served through an inference runtime with an open-source harness that can read files, edit code, and run commands. He recommends Qwen-Code for Qwen3.6, citing Nvidia's Polar paper, which found Qwen models performed best in Qwen-Code. The Qwen3.6 35B-A3B model is about 22 GB to download and needs roughly 30–40 GB of RAM.

Jun 25

Jun 25Thu