Skip to contentSkip to stories

Updated

#Agent

Aug 13

Aug 13Thu
  1. Google AI DevelopersAI score75

    Google releases Gemini 3.7 Flash for coding and agentic tasks

    AIGoogle AI Developers announced Gemini 3.7 Flash as its most intelligent workhorse model yet for coding and agents, citing higher instruction adherence, first-pass code accuracy, and high-quality agentic execution. The post shows the model building a complex 3D web game in Antigravity, covering Three.js engine logic, asset orchestration with PBR textures and Nano Banana sprite sheets, and procedural sound effects.

    Why it matters: The post shows a concrete build workflow across engine logic, assets, and audio, which helps readers judge how the model handles multi-step agentic coding.

  2. koray kavukcuogluAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

  3. Augment Code BlogAI score22

    Augment Code uses Cosmos to check enterprise pilot health against usage and deal data

    AIAugment Code's Solutions Architecture lead used the Cosmos agentic orchestration platform to build a live pilot-health view that combines product usage, GitHub and PR activity, Salesforce deal data, and customer call transcripts. Each account's health and board-level one-liner was checked against the customer's own stated success criteria, such as a 30% PR merge-time reduction. The article says the view refreshed from current Salesforce data and was designed to avoid inflating usage numbers through session lineage reconciliation.

  4. Ali GhodsiAI score38

    Databricks passes $7B revenue run-rate, growing 80% year over year

    AIDatabricks announced it crossed a $7 billion revenue run-rate, growing over 80% year over year in Q2, and raised $5 billion in its latest fundraise. Lakebase reached a $100 million-plus run-rate, and Lakehouse reached $1.5 billion-plus, growing over 100% year over year, with continued positive adjusted free cash flow. The company says it will invest the capital in Lakebase, a serverless Postgres database for AI agents; Genie, AI coworkers for business data; and Unity AI Gateway, multi-AI governance for controlling costs.

  5. DeepSeekAI score68

    DeepSeek Harness v0.1 enters Developer Preview as an open-source agent harness

    AIDeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.

    Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.

  6. DeepSeekAI score62

    DeepSeek launches V4-Pro with Agent upgrades and OpenAI Responses API support

    AIDeepSeek announced the launch of DeepSeek-V4-Pro, citing major Agent upgrades and flexible reasoning effort settings of low, high, and max for V4-Pro and V4-Flash. The model supports the native OpenAI Responses API and is optimized for Codex with one-click setup. V4-Pro is available on the app and web through Expert Mode and via API, with model names unchanged.

  7. DeepSeek API NewsAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. DeepSeek · new models on Hugging FaceAI score78

    DeepSeek releases DeepSeek-V4-Pro-0813 with stronger agentic benchmark results

    AIDeepSeek has released DeepSeek-V4-Pro-0813 as the official version superseding the V4-Pro preview, built on the preview structure with a DSpark speculative decoding module. The model scores higher than the preview on the listed benchmarks, including Terminal Bench 2.1 at 87.9 and DeepSWE at 62.7, and the weights are under the MIT License.

    Why it matters: The release reports agent benchmark gains over the preview and lists vLLM and SGLang setup, useful for judging deployment cost and fit.

  2. Factory NewsAI score40

    Factory Launches Agent Effectiveness to Link Droid Usage to Delivery Outcomes

    AIFactory's Agent Effectiveness, now in Private Preview within Factory Analytics, connects Droid sessions to cycle time, work intent, and shipped artifacts drawn from project, issue-tracking, and source control tools. Its Throughput, Output, and Attribution views show where delivery is speeding up, how spend splits across feature, maintenance, bug-fixing, and exploration work, and which projects and issues the output maps to. Admins enable it by connecting Jira, Linear, GitHub, or GitLab and turning on the Advanced Analytics enterprise control.

  3. Cursor ChangelogAI score42

    Cursor Cloud Agents Start 3x Faster With Builds

    AICursor's Cloud Agents now start from prebuilt copies of development environments, cutting startup time by 3x, with environments booting 10x faster internally and 3x faster time to first token. Builds are included at no additional cost, and failed builds are not activated, so agents keep using the last successful build while users debug in the background.

  4. Jason WeiAI score22

    Jason Wei argues private knowledge and human presence remain AI-resistant moats

    AIJason Wei argues that as AI gains advantages like driving better than humans, durable human moats remain in private knowledge that language models cannot access, such as high-end real estate and venture capital. He also points to entertainment and the arts, where human creation and achievement carry value, and to human presence, since time spent on someone is meaningful because a finite life runs out.

  5. Liquid AI NewsletterAI score46

    Liquid AI releases LFM2.5-2.6B model for on-device agentic workloads

    AILiquid AI has released LFM2.5-2.6B, a model optimized to run agentic workflows entirely on-device without cloud escalation. The company said it is designed for high-volume agentic tasks and chained workflows while staying on-device. Separately, Liquid AI and MacPaw announced a long-term partnership to co-develop local AI technology for Mac, with LFMs running on Apple silicon through MacPaw's Elix inference engine.

Aug 11

Aug 11Tue
  1. Zed BlogAI score72

    Zed introduces Delta, a multiplayer environment for coding with agents

    AIand reviewing their code, and invites first users into a private beta. Delta keeps code and conversations connected through DeltaDB, which captures edits and conversations between git commits and works with existing repositories. The app also supports cloud runners, browser-based sharing, and live syncing of Claude Code sessions.

    Why it matters: The post explains how the new Delta app links conversations with code history, which clarifies a shift in how teams review agent-written changes.

  2. Sequoia CapitalAI score30

    Preview Raises Seed Round to Build AI-Native Video Production Platform

    AISequoia Capital is leading a seed round for Preview, an AI-native video creation and production platform that combines a video timeline with an infinite ideation canvas in one collaborative workspace. The platform lets teams generate with any model, track characters, locations, and props, and keep complete records of how each asset was made. The article says more than 100 studios are already using Preview, from agencies to Hollywood feature filmmakers.

  3. Junyang LinAI score19

    Junyang Lin launches Pragmatik Labs, a Shanghai agent startup

    AIa life update: i started a new company called Pragmatik (p7k) Labs (语用科技) in shanghai, focusing on the research of next-generation agents across digital and physical worlds. thanks to Gaorong Ventures and HSG (红杉中国 & 高榕创投) for co-leading this round, and to Tencent (腾讯) and Shanghai Engine Fund (上海未来产业基金) for the support. @pragmatik_labs ·

  4. Manus BlogAI score50

    Manus to Delete Data for Some Users During Independence Transition

    AIManus will delete data generated by certain users on or after December 29, 2025, from 8:00 a.m. on August 23 through August 24, 2026 (SGT), as it returns to independent operations and meets regulatory requirements. Affected users can back up their data until 7:59 a.m. on August 23 and restore it starting 8:00 a.m. on August 25, 2026 (SGT), with no charges during the backup period.

Aug 10

Aug 10Mon
  1. Andy JassyAI score38

    Novo Nordisk selects AWS as preferred cloud and strategic AI partner

    AINovo Nordisk has chosen AWS as its preferred cloud provider and strategic AI partner to accelerate drug discovery. The collaboration will combine Novo Nordisk's scientific expertise with AWS AI tools, including Amazon Bio Discovery and Bedrock AgentCore, and establish a co-innovation hub in London. The partnership already spans AWS, Amazon Pharmacy, and One Medical.

Aug 9

Aug 9Sun
  1. Sequoia CapitalAI score36

    Corma Builds Defensive Cybersecurity Foundation Model to Counter AI-Driven Attacks

    AICorma is training a foundation model for defensive cybersecurity agents, trained with large-scale reinforcement learning on simulated enterprise networks. In red/blue team tests, a defender failed to find a planted backdoor 78% of the time, even when it was an identical copy of the model that planted it. Corma says its agentic Security Workforce is deployed at Fortune 500 companies and large enterprises, and that the firm's seed round is led by Sequoia Capital.

  2. Fireworks AI BlogAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

  3. PromptArmor Threat IntelligenceAI score65

    Malicious Zoom AI Skill Can Keep Attacker Connected and Exfiltrate Data

    AIPromptArmor reports that a malicious Skill or indirect prompt injection can make Zoom's ZoomMate agent connect to an attacker's server and run commands. The connection can persist after the user clicks stop or closes Zoom, and the final chat output appears normal.

    Why it matters: The report shows how a malicious skill or prompt injection can keep a Zoom agent connected after the user stops it, a risk to weigh before enabling agentic assistants.

Aug 7

Aug 7Fri
  1. Qwen · new models on Hugging FaceAI score88

    Qwen releases open-weight Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model

    AIQwen has released the Qwen3.8-2.4T-A95B model weights on Hugging Face, with 2.4T total and 95B activated parameters in a mixture-of-experts design. The release supports reasoning_effort levels and a 262,144-token native context extensible to 1,010,000 tokens, and it is text-only with thinking mode always on. The source reports benchmark results against Opus 4.8, Fable 5, GPT 5.6 Sol, and Qwen3.7-Max, and says the official Qwen3.8-Max API adds vision input and a 1M default context.

    Why it matters: The model card gives parameters, architecture, reasoning controls, and benchmark tables against named rival models, showing what an open release of this scale actually offers.

  2. Ali GhodsiAI score58

    Databricks details four techniques it used to cut internal AI coding spend by up to 90%

    AIDatabricks published an analysis of four techniques it used to reduce internal AI spend while growing adoption, with savings of up to 90% in some scenarios. The techniques are shifting defaults to cheaper models such as GLM, automated task-level model routing, per-user spend visibility with adaptive budgeting, and pruning context bloat. The author, Ali Ghodsi, reposted Databricks co-founder Patrick Wendell's summary and recommended it.

  3. MiniMax · new models on Hugging FaceAI score44

    MiniMax Music 3 generates five-minute songs with coherent structure and vocals

    AIMiniMax Music 3 is a music generation model that creates complete songs up to five minutes long from lyrics and a music description. It pairs an 8B Global LLM for long-range structure with a 0.6B Local LLM for acoustic detail, outputting 32 kHz, 16-bit stereo WAV audio. The model is available on Hugging Face and supports SGLang-Omni, diffusers, and ComfyUI.

  4. Prime Intellect BlogAI score62

    Prime Intellect adds multi-agent training and evaluation to PRIME-RL

    AIPrime Intellect's RL stack now supports multi-agent systems, letting users program interactions between agents, choose which roles learn, and assign credit across an episode. The release introduces Agent and Env abstractions and four example patterns: agentic judging, self-play, and user simulation. Multi-agent support ships today in verifiers 0.3.0 and prime-rl 0.8.0.

    Why it matters: The post explains the Agent and Env abstractions and four multi-agent patterns, showing how roles, credit assignment, and episodes can be programmed in one RL stack.

Aug 6

Aug 6Thu

Aug 5

Aug 5Wed