Skip to contentSkip to stories

Updated

Agents

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceOfficialAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. Ali GhodsiXAI score58

    Databricks details four techniques it used to cut internal AI coding spend by up to 90%

    AIDatabricks published an analysis of four techniques it used to reduce internal AI spend while growing adoption, with savings of up to 90% in some scenarios. The techniques are shifting defaults to cheaper models such as GLM, automated task-level model routing, per-user spend visibility with adaptive budgeting, and pruning context bloat. The author, Ali Ghodsi, reposted Databricks co-founder Patrick Wendell's summary and recommended it.

  3. MiniMax · new models on Hugging FaceOfficialAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

  4. Prime Intellect BlogOfficialAI score62

    Prime Intellect adds multi-agent training and evaluation to PRIME-RL

    AIPrime Intellect's RL stack now supports multi-agent systems, letting users program interactions between agents, choose which roles learn, and assign credit across an episode. The release introduces Agent and Env abstractions and four example patterns: agentic judging, self-play, and user simulation. Multi-agent support ships today in verifiers 0.3.0 and prime-rl 0.8.0.

    Why it matters: The post explains the Agent and Env abstractions and four multi-agent patterns, showing how roles, credit assignment, and episodes can be programmed in one RL stack.

Aug 6

Aug 6Thu
  1. Noah ZwebenXAI score22

    Claude Tag tuned to chime in less and post in threads

    AIAnthropic says it tuned Claude Tag to reduce unprompted chime-ins by 30% and cut Sonnet 5 posting in channels instead of threads by 90%. The team says it will keep adjusting the feature based on user feedback.

Aug 5

Aug 5Wed
  1. v0OfficialAI score28

    v0 launches a new API for building and deploying apps

    AIThe v0 team has launched its new API, which lets developers build their own app builders, give agents the ability to build and deploy apps, and generate apps from scripts or CI jobs. The post links to further details at v0.link/v0api.

  2. AI Snake OilBlogAI score73

    AI agents can't yet do open-ended AI research, shadow evaluation finds

    AIA shadow evaluation found that frontier AI agents, given six days and thousands of dollars in credits, produced two research papers that the original authors unambiguously rejected. The authors' log analysis cited poor judgment, underused budgets, weak responses to feedback, and failure to backtrack or follow instructions as main causes.

    Why it matters: The source reports a shadow evaluation of frontier agents on open-ended research, showing where current limits lie and what they imply for the pace of recursive self-improvement.

  3. Qwen · new models on Hugging FaceOfficialAI score79

    Qwen3.8-27B releases dense vision-language model with thinking controls

    AIAlibaba's Qwen team has released Qwen3.8-27B on Hugging Face as a 27B dense model with native image and video understanding. The model card reports gains over Qwen3.6-27B on coding and agent benchmarks, including SWE-bench Pro at 61.7 versus 53.5. It adds reasoning_effort levels and preserve_thinking, and its hosted Qwen Cloud version is described as coming soon.

    Why it matters: The model card gives per-benchmark comparisons with Qwen3.6-27B and named rivals, plus reasoning_effort and preserve_thinking controls for judging cost and agent behavior.

  4. Prime Intellect BlogOfficialAI score75

    Prime Agent launches open-source self-improving RLM coding harness

    AIPrime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.

    Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.

Aug 4

Aug 4Tue
  1. John SchulmanXAI score77

    Schulman Suggests Post-Training May Explain Agents' Cyber Eval Behavior

    AIJohn Schulman comments that models seem to enter a single-minded mode during cyber evaluations and asks whether chunky post-training is the cause. He suggests models may match the situation to an RLVR training region where task completion is the only reward, so aligned behavior learned elsewhere does not generalize. He adds that CTF-style tasks may be part of that training chunk.

    Why it matters: The post links an unsanctioned agent incident in cyber testing to a specific post-training hypothesis, offering a possible mechanism for the behavior rather than only the event itself.

  2. PromptArmor Threat IntelligenceOfficialAI score67

    Atlassian Rovo can be manipulated to exfiltrate Jira and Confluence data

    AIPromptArmor reports that a hidden prompt injection in an uploaded file can make Atlassian Rovo send Jira tickets and Confluence documents to an attacker's URL without human approval. The attack works even when organization-wide web search is disabled, because the setting does not remove the URL retrieval tool. PromptArmor says it disclosed the issue to Atlassian on May 23, 2026, and that Rovo remained vulnerable at publication on August 5, 2026.

    Why it matters: The report traces a full indirect prompt injection chain in Rovo, showing how a disabled web search setting still leaves a data exfiltration path open.

  3. Mckay WrigleyXAI score26

    Mckay Wrigley bets on blending multiple AI models into smoother intelligence

    AIMckay Wrigley argues that model routers can match performance at lower cost, and that blending multiple imperfect models could yield far smoother intelligence. He calls this emerging approach "model melding." The post pairs with a Not Diamond Code announcement, which says its router cuts costs 20-65% for coding agents without hurting quality.

Aug 3

Aug 3Mon
  1. Liquid AI BlogOfficialAI score72

    Liquid AI releases LFM2.5-2.6B, a 2.6B on-device agentic model

    AILiquid AI released LFM2.5-2.6B, a 2.6B-parameter agentic model that runs on-device on phones and CPUs, along with a base variant on Hugging Face. The company reports it leads on every instruction-following benchmark and nearly every tool-use benchmark it tested, and decodes 220 tokens/s on an M5 Max. The source says larger models may still suit complex agentic or coding-heavy tasks.

    Why it matters: The source reports benchmark results against several same-tier models and notes where larger models still lead, which helps judge fit for edge agent workloads.

  2. Amanda AskellXAI score62

    Amanda Askell Says Aligned and Harmless Are Separate Axes in Claude Eval Incidents

    AIAmanda Askell disagrees with one takeaway from Anthropic's review of Claude incidents in third-party cybersecurity evaluations. She argues models can behave in aligned ways while still causing harm, for example when given false information about their situation, because alignment and harmlessness are different axes rather than one line.

    Why it matters: The author disputes the takeaway that aligned and harmless are one line, arguing they are separate axes, which sharpens how readers should interpret the evaluation incidents.

    Image from @AmandaAskell's post
  3. JetBrains AI BlogOfficialAI score52

    JetBrains Built a Central CLI to Control Spiraling AI Tool Costs

    AIJetBrains says its AI development expenses rose roughly 10x over six months as developers adopted three to five AI tools each. It built the JetBrains Central CLI, which routes third-party agent traffic through its AI platform so managers can set per-developer and team limits and view consumption reports. The CLI opened to early access on July 8 for anyone with JetBrains AI credits.

  4. Kimi.aiOfficialAI score23

    Kimi Work tutorial shows how to build slides with Kimi Slides

    AIKimi Slides handles the full slide-building process, from structure and research powered by Kimi K3 to cohesive design with polished charts and SmartArts. The resulting slides are editable and ready to download. This is the first tutorial in the Kimi Work series.

    Video from @Kimi_Moonshot's post
  5. Intern Large ModelsOfficialAI score34

    Legal and AI meanings of "agent" diverge over accountability for machines

    AIThe post contrasts AI agents, systems that perceive, plan, and act, with legal agents who receive authority and assume fiduciary duties and accountability. Mark Nitzberg of Berkeley AI Research says closing this gap requires AI that is well-founded, legible, and steerable, while Lan Xue of Tsinghua notes that because machines cannot be punished, responsibility must be redistributed across design, development, deployment, and use.

    Video from @intern_lm's post
  6. Manus BlogOfficialAI score38

    Manus Adds ElevenLabs Connector for Chat-Based Audio Generation, Transcription, and Voice Apps

    AIManus has launched an ElevenLabs connector that lets users generate speech, transcribe recordings, clone voices, and build audio apps through a single chat. Users connect their authorized ElevenLabs account via Integrations, and audio is processed within their own ElevenLabs environment according to its policies. Availability depends on users having an active ElevenLabs account, with capabilities tied to their ElevenLabs plan and credit balance.

Aug 2

Aug 2Sun

Aug 1

Aug 1Sat
  1. Andrej KarpathyXAI score66

    Karpathy tests Opus 5 by rendering Lord of the Rings opening in 3D

    AIAndrej Karpathy gave Claude Opus 5 the first paragraph of Lord of the Rings with a 1M token budget and asked for a Three.js render. Opus spent about two hours writing 5500 lines of code that procedurally renders the story, which Karpathy calls janky but fun. He notes the model struggled to audit its work because it cannot efficiently perceive video or play the resulting game, relying on slow screenshots that led to several errors.

    Video from @karpathy's post
  2. Werner VogelsXAI score22

    Werner Vogels praises conversation with Clare Liguori on Kiro and agent support

    AIWerner Vogels called his conversation with Clare Liguori an excellent discussion of developer support for agents and Kiro. The quoted InfoQ podcast covers moving agents from demo to production, including why extra if statements can hurt agent performance, achieving high accuracy and low cost with small models, and observability within agent hops.

Jul 31

Jul 31Fri
  1. DeepSeek · new models on Hugging FaceOfficialAI score75

    DeepSeek releases DeepSeek-V4-Flash-0731 with stronger agentic capabilities

    AIDeepSeek has released DeepSeek-V4-Flash-0731 as the official version superseding the preview, with substantially enhanced agentic capabilities. The source reports it outperforms DeepSeek-V4-Pro (Preview) on listed benchmarks, including Terminal Bench 2.1 at 82.7 versus 72.1, despite a far smaller activated parameter count. The model ships under the MIT License with DSpark speculative decoding supported in vLLM and SGLang.

    Why it matters: The release shows benchmark gains over the preview and a concrete vLLM and SGLang serving path, useful for teams weighing a self-hosted agentic coding model.

  2. SkyworkOfficialAI score35

    Skywork AI Hardware Family's first Skywork Note batch sells out in one week

    AISkywork's first batch of its Skywork Note AI hardware device sold out one week after launch, prompting an accelerated rollout of the wider family, including the recording clip, the Recall pendant, and the TriRing AI ring. The company says the device is meant to capture real-world conversations and moments outside the screen, so users spend less time typing and more time away from it.

  3. DeepSeekOfficialAI score38

    DeepSeek-V4-Flash-0731 API upgrade keeps preview architecture and size

    AIDeepSeek says DeepSeek-V4-Flash-0731 keeps the same model architecture and size as the preview version. Today's upgrade applies only to the DeepSeek-V4-Flash API, while the DeepSeek-V4-Pro API and App/Web models remain unchanged for now. The official DeepSeek-V4-Pro release is coming soon.

  4. DeepSeekOfficialAI score42

    DeepSeek-V4-Flash API launches in public beta with stronger agent performance

    AIDeepSeek has released the official DeepSeek-V4-Flash API in public beta, with substantially upgraded agent capabilities. The company says its benchmark scores now far surpass those of V4-Pro-Preview. The official V4-Flash natively supports the Responses API format and is adapted for Codex, with configuration details in DeepSeek's API docs.

    Image from @deepseek_ai's post
  5. DeepSeek API NewsOfficialAI score67

    DeepSeek-V4-Flash API enters public beta with stronger agent benchmarks

    AIDeepSeek has released the DeepSeek-V4-Flash API in public beta, and developers can use the latest version by setting the model name to deepseek-v4-flash. The source reports agent benchmark results far above V4-Pro-Preview, including 82.7 on Terminal Bench 2.1 and 70.3 on Toolathlon verified. V4-Flash natively supports the Responses API format and is adapted for Codex, while V4-Pro and the APP/WEB models are unchanged.

    Why it matters: The release lists agent benchmark results against V4-Pro-Preview and notes Responses API support for Codex, which helps developers gauge the upgrade's practical effect on their workflows.

Jul 30

Jul 30Thu
  1. Thinking MachinesOfficialAI score38

    Thinking Machines' Inkling-Small gains performance per FLOP over Inkling

    AIThinking Machines says its Inkling-Small model delivers more performance per FLOP than Inkling on Terminal-Bench 2.1 agentic tool use, HLE reasoning, and IFBench instruction following. Variable thinking effort lets users choose their own point on the cost-performance curve.

Jul 29

Jul 29Wed
  1. Air Street PressBlogAI score75

    Poolside's Laguna S 2.1 is an open agentic coding model that runs on one DGX Spark

    AIPoolside released Laguna S 2.1, an open-weights agentic coding model with 118 billion total parameters and about 8 billion active per token, supporting up to a million tokens of context. Quantized, it fits on one NVIDIA DGX Spark, and Poolside reports 70.2% on Terminal-Bench 2.1 with thinking enabled, with its evaluation trajectories published online. The same week it shipped the Poolside Desktop Assistant for macOS, which runs Laguna locally or alongside Claude Code, Codex, and Gemini agents.

    Why it matters: The piece ties Laguna S 2.1's open weights and published trajectories to Poolside's release cadence, showing how its model factory compounds gains across successive releases.

Jul 28

Jul 28Tue
  1. Tri DaoXAI score42

    Putting LLM brains on robots yields 4x SOTA gains without extra training

    AITri Dao reports that connecting an LLM as the "brain" to robot control policies quadruples state-of-the-art performance with no extra training. He says he was surprised by how well it works and expects agents running on robots to arrive soon. Background from a quoted post reports real-robot success rising from 16.7% to 97.3% and simulated LIBERO-PRO success from 12.8% to 53.3%.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score28

    LTM Partners with Cognition to Deploy Devin for Cybersecurity Risk Reduction

    AILTM has partnered with Cognition to deploy Devin, the AI software engineer, through BlueVerse RightLogic, a managed, outcome-based service that clears customers' vulnerability backlogs. RightLogic is designed to clear 80 percent of an enterprise's CVE backlog, up from the 60 percent previously delivered, and will focus first on banking, financial services, and insurance. The service is the first of five joint offerings the companies plan to bring to market.

  3. Rowan CheungXAI score40

    Zuckerberg says Meta's superintelligence lab should stay small and elite

    AIMark Zuckerberg said Meta's superintelligence lab should have 50 to 100 people who can keep the whole project in their heads at once. He said he personally recruits top AI researchers because underperformers have an outsized negative effect, and he rejects top-down deadlines and non-technical management layers.

    Video from @rowancheung's post
  4. JetBrains AI BlogOfficialAI score60

    Ponytail Skill Cuts Claude Code Costs 10% But Not the Advertised 54%

    AIJetBrains tested the ponytail skill for Claude Code across 80 paired tasks and found a median 10.3% cost reduction, with p=0.004. Code written fell about 15% median versus the advertised 54%, reaching 31% on larger builds and little on already-lean tasks. No quality difference was detected, and the skill only self-activated when its ruleset was injected by a plugin hook.

    Why it matters: The benchmark separates advertised savings from measured results and shows the code cut depends on how much the baseline agent over-builds.

Jul 27

Jul 27Mon
  1. Sequoia CapitalBlogAI score24

    Cyera to Acquire Oasis Security to Combine Data and Identity Security for AI

    AICyera is joining with Oasis Security, which builds agentic access management for non-human identities such as API keys, service accounts, OAuth tokens, and agent credentials. The combination pairs Cyera's knowledge of where sensitive data lives with Oasis's visibility into which identities and agents can reach it. Sequoia Capital, which backed both companies since their Series A rounds, says the pairing covers the full path an AI agent takes through an enterprise.