Skip to contentSkip to stories

Updated

#Agent

Showing low-relevance items too. Hide low-relevance items

Aug 4

Aug 4Tue

Aug 3

Aug 3Mon
  1. Liquid AI BlogAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

  2. Amanda AskellAI score62

    Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents

    AIAmanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.

    Image from @AmandaAskell's post
  3. JetBrains AI BlogAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

  4. Intern Large ModelsAI score34

    Legal and AI meanings of "agent" diverge over accountability for machines

    AIThe post contrasts AI agents, systems that perceive, plan, and act, with legal agents who receive authority and assume fiduciary duties and accountability. Mark Nitzberg of Berkeley AI Research says closing this gap requires AI that is well-founded, legible, and steerable, while Lan Xue of Tsinghua notes that because machines cannot be punished, responsibility must be redistributed across design, development, deployment, and use.

    Video from @intern_lm's post
  5. Manus BlogAI score38

    Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps

    AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.

Aug 2

Aug 2Sun

Aug 1

Aug 1Sat
  1. Andrej KarpathyAI score66

    Karpathy tests Opus 5 by rendering Lord of the Rings opening in 3D

    AIAndrej Karpathy gave Claude Opus 5 the first paragraph of Lord of the Rings with a 1M token budget and asked for a Three.js render. Opus spent about two hours writing 5500 lines of code that procedurally renders the story, which Karpathy calls janky but fun. He notes the model struggled to audit its work because it cannot efficiently perceive video or play the resulting game, relying on slow screenshots that led to several errors.

    Video from @karpathy's post
  2. Werner VogelsAI score22

    Werner Vogels praises conversation with Clare Liguori on Kiro and agent support

    AIWerner Vogels called his conversation with Clare Liguori an excellent discussion of developer support for agents and Kiro. The quoted InfoQ podcast covers moving agents from demo to production, including why extra if statements can hurt agent performance, achieving high accuracy and low cost with small models, and observability within agent hops.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. SkyworkAI score35

    Skywork AI Hardware Family's first Skywork Note batch sells out in one week

    AISkywork's first batch of its Skywork Note AI hardware device sold out one week after launch, prompting an accelerated rollout of the wider family, including the recording clip, the Recall pendant, and the TriRing AI ring. The company says the device is meant to capture real-world conversations and moments outside the screen, so users spend less time typing and more time away from it.

  3. DeepSeek API NewsAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu

Jul 29

Jul 29Wed
  1. Air Street PressAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

Jul 28

Jul 28Tue
  1. Tri DaoAI score42

    Putting LLM brains on robots yields 4x SOTA gains without extra training

    AITri Dao reports that connecting an LLM as the "brain" to robot control policies quadruples state-of-the-art performance with no extra training. He says he was surprised by how well it works and expects agents running on robots to arrive soon. Background from a quoted post reports real-robot success rising from 16.7% to 97.3% and simulated LIBERO-PRO success from 12.8% to 53.3%.

  2. Cognition Blog (Devin, Windsurf)AI score28

    LTM Partners with Cognition to Deploy Devin for Cybersecurity Risk Reduction

    AILTM has partnered with Cognition to deploy Devin, the AI software engineer, through BlueVerse RightLogic, a managed, outcome-based service that clears customers' vulnerability backlogs. RightLogic is designed to clear 80 percent of an enterprise's CVE backlog, up from the 60 percent previously delivered, and will focus first on banking, financial services, and insurance. The service is the first of five joint offerings the companies plan to bring to market.

  3. JetBrains AI BlogAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 27

Jul 27Mon
  1. Sequoia CapitalAI score24

    Cyera to Acquire Oasis Security to Combine Data and Identity Security for AI

    AICyera is joining with Oasis Security, which builds agentic access management for non-human identities such as API keys, service accounts, OAuth tokens, and agent credentials. The combination pairs Cyera's knowledge of where sensitive data lives with Oasis's visibility into which identities and agents can reach it. Sequoia Capital, which backed both companies since their Series A rounds, says the pairing covers the full path an AI agent takes through an enterprise.

Jul 26

Jul 26Sun
  1. Fireworks AI BlogAI score60

    Fireworks AI adds open-weight Kimi K3 with US-only serverless endpoints

    AIFireworks AI made the open-weight Kimi K3 available for inference and training on its platform, with US-only serverless endpoints and Zero Data Retention. In its own head-to-head with Opus 5, the post reports K3 at 92.7% accuracy and $0.52 per task on SWE (480) against Opus 5's 94.8% and $1.05, with the vendor claiming up to 5x better cost efficiency per task.

    Why it matters: The post compares Kimi K3 with Opus 5 on accuracy and cost per task, giving readers concrete figures to judge the open model against closed alternatives for their own workloads.

  2. Philipp SchmidAI score62

    EvoCode-Bench Tests Coding Agents Across Multi-Turn Iterative Specification Changes

    AIEvoCode-Bench is a multi-turn coding benchmark with 26 tasks spanning 227 sequential rounds, where agents keep a persistent workspace and must pass cumulative tests after each evolving instruction. The results show that agents perform much worse when building on their own prior work than when starting from a clean, human-completed codebase. Regressions, not failure to implement new features, are the main bottleneck, and agents that maintained a persistent requirements document more than doubled their success rates.

  3. Berkeley AI ResearchAI score44

    Berkeley AI Research Trains LLMs to Update Beliefs for Long Tasks

    AIBerkeley AI Research introduces ABBEL, a framework that replaces full interaction histories with natural-language belief states that models update as new observations arrive. On CollabBench collaborative coding, belief grading closes about half the performance gap to full-context models while using fewer peak tokens and training in 50 steps instead of 100.

Jul 25

Jul 25Sat
  1. Ali GhodsiAI score26

    Longer-running AI agents often perform worse than faster ones, says Ghodsi

    AIAli Ghodsi argues that AI agents which take longer to work through a task are often worse, while Genie reaches results faster. He adds that ontology will be key to giving agents the context they need to answer correctly and quickly. The related post reports that Genie Code outperformed three general-purpose coding agents on more than 400 real user data tasks.

  2. LangChain BlogAI score39

    What does it mean for companies to "own their intelligence" with AI?

    AILangChain Blog argues that companies need to own their AI intelligence rather than rely on generic models, because general models do not know company-specific policies, workflows, or risk tolerances. Ownership means controlling the agent system (model optionality, harness, and context), the economics, quality, and risk of AI work, and how intelligence compounds over time. The post uses an insurer's claims processing as an example of why off-the-shelf models fall short.

Jul 24

Jul 24Fri
  1. Alex AlbertAI score34

    Opus 5 now produces consultant-grade spreadsheets and slide decks, Alex Albert says

    AIAlex Albert, of Anthropic, says Opus 5 now produces near-superhuman spreadsheets and slide decks that match what a consultant would make, just over six months after its predecessor. He also notes that finance professionals are reporting strong reactions to Claude for Excel, and he expects agentic progress seen in coding to extend to other fields in 2026.

    Video from @alexalbert__'s post
  2. catAI score66

    Claude Opus 5 released as strong option for long-running autonomous work

    AIAnthropic introduces Claude Opus 5 as a thoughtful and proactive model that comes close to the frontier intelligence of Fable 5 at half the price, according to the quoted announcement. The author, who works on the product, says Claude Opus 5 is great at long-running autonomous work and invites users to try it and share feedback.

    Why it matters: The post pairs a new model's long-running autonomous strength with a pricing claim, letting readers weigh capability against cost for agentic workloads.

  3. Mike KriegerAI score46

    Mike Krieger says Claude Opus 5 became his daily driver

    AIAnthropic co-founder Mike Krieger says Claude Opus 5 has become his daily driver at work and on weekends. He reports it can work for hours on complex tasks and consistently gets to the bottom of tricky problems, and he has also built some games with it. Anthropic's announcement describes Opus 5 as close to the frontier intelligence of Fable 5 at half the price.

Jul 23

Jul 23Thu
  1. Matei ZahariaAI score36

    Berkeley STAR Lab packages AI research optimizers into one GEPA API

    AIBerkeley's STAR Lab packaged multiple LLM-based "autoresearch" algorithms into a single API within the GEPA package, letting users mix and match them. The optimizers can be applied to tasks including prompt writing, agent design, and code optimization. The quoted thread adds that GEPA, AutoResearch, and Meta-Harness each win on different tasks, and that the new optimize_anything omni meta-optimizer beats every standalone optimizer at a matched budget.

  2. One Useful Thing (Ethan Mollick)AI score67

    Ethan Mollick's guide to choosing AI tools for agentic work

    AIEthan Mollick's guide says ChatGPT and Claude are the main choices for real work, since their agent modes can act on a computer. He separates agent modes that run on the company's computers from those that access the user's own computer. He recommends keeping approval settings on for sending, spending, or deleting, because of prompt injection risk. He also notes that Gemini currently lags for agentic work, though its Notebook and video tools are useful.