Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Aug 18

Aug 18Tue
  1. VercelOfficialAI score42

    Vercel launches $1M hacker challenge to test Sandbox security

    AIVercel is offering up to $1,000,000 in a public hacker challenge testing its Vercel Sandbox against escapes from the Firecracker microVM and bypasses of the host-side network boundary. Rewards reach $50,000 per report, administered through HackerOne (@Hacker0x01). The company says agents can now exploit vulnerable sandbox boundaries, so it is testing its own defenses in the open.

Aug 17

Aug 17Mon
  1. Z.ai Release NotesOfficialAI score63

    Z.ai releases GLM-5.3 with stronger coding and vulnerability discovery

    AIZ.ai's release notes announce GLM-5.3, which the company says delivers a 50% gain over GLM-5.2 on Z.ai Code Bench and reaches open-source SOTA on public benchmarks including Terminal Bench 3.0. The company also reports that GLM-5.3 matches Mythos 5 in white-box code review and vulnerability discovery, identifying 2,436 vulnerabilities in real-world targets, 1,097 of them medium- or high-severity. A separate GLM-5.3-Flash entry describes native visual capabilities and a hybrid architecture with 320B total and 18B activated parameters.

    Why it matters: The release notes show GLM-5.3's coding and cybersecurity gains, with a vulnerability count, letting readers compare it against Z.ai's prior GLM-5.x line and other coding models.

  2. Chip HuyenXAI score22

    Chip Huyen asks for a model tiering system for agent orchestration

    AIChip Huyen asks what a good model tiering system looks like, since she is tired of naming specific models per vendor for her agent orchestrator. She wants to instruct the orchestrator by task tier, such as "use models tier ..." for a given kind of task, instead of listing Claude, OpenAI, and other models individually.

  3. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry adds five Claude agent features to Azure-hosted deployments

    AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.

    Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.

  4. Jason WeiXAI score45

    Jason Wei argues tool use cannot replace larger language models

    AIJason Wei now believes a small 1B-parameter "cognitive core" relying on tools is insufficient, because fast, natural recall without tool use matters. He cites speed, knowledge better learned through backpropagation than retrieved from search, and the greater reliability of already-known facts over repeated lookups. Since a 1B model has an information limit, he argues that demanding AI will still need larger models, not just tool access.

  5. Replit BlogOfficialAI score60

    Replit adds black-box pen tests that probe apps like external attackers

    AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.

    Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.

  6. Import AIBlogAI score44

    DiG-bench Tests AI Rule Discovery as Opus 5 and Fable 5 Lead

    AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.

Aug 16

Aug 16Sun
  1. Philipp SchmidBlogAI score58

    Controlling Android with Gemini 3.7 Flash and 150 lines of Python

    AIThe author built a Python agent that uses Gemini 3.7 Flash to control an Android emulator from raw screenshots, returning normalized 0–999 coordinates that are scaled to 1080x1920 pixels over ADB. In a test, the agent opened Chrome, closed popups, and solved one round of Wordle in two guesses without accessibility IDs or DOM access. The article presents the loop as usable for UI testing and task automation across native apps, webviews, and canvas interfaces, with code in an open-source quickstart repository.

Aug 15

Aug 15Sat
  1. Prime Intellect BlogOfficialAI score73

    Prime Intellect tests frontier models on 153 autonomous nanoGPT research runs

    AIPrime Intellect ran 153 autonomous runs on the nanoGPT optimizer speedrun across 18 frontier models, with runs lasting up to eight days on 8xH200s. The results show a large gap between models at every stage of the research process, though none of the runs produced a fundamentally new method.

    Why it matters: The experiment measures how frontier models conduct autonomous research, showing large gaps between models in experiment choice, execution, and result interpretation.

Aug 14

Aug 14Fri
  1. Augment Code BlogOfficialAI score62

    Augment rebuilds its Auggie CLI harness on Pi, cutting SWE-bench Pro task cost 53%

    AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.

    Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.

  2. Epoch AI · The Epoch BriefOfficialAI score42

    Nine big questions AI benchmarks can help answer

    AIEpoch AI lays out nine open questions about AI capability progress, including whether AI can take over open-ended jobs and whether benchmark scores share a single underlying factor. The post is an opinion newsletter piece in which the author says the questions shape Epoch's benchmarking work. It names examples such as the Andon Café agent-run cafe, Remote Labor Index, MirrorCode, and the Epoch Capabilities Index (ECI).

  3. Andrew NgXAI score38

    Andrew Ng maps the four key skills for AI engineering

    AIAndrew Ng's team released an AI Engineering Skills Map, built from analysis of over 10,000 job postings and expert interviews, identifying four priority skills. The skills are building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. Ng says these skills matter for all developers, not only those with the AI Engineer title.

  4. Ali GhodsiXAI score46

    Databricks Smart Routing cuts AI coding task costs about 30% in Unity Gateway

    AIAli Ghodsi says Smart Routing on Databricks' AI Gateway lowers costs by about 30% without sacrificing quality. The quoted Databricks post says it matches each coding task to the right model and harness based on task needs, so higher-cost models focus on intelligence while lower-cost models compete on cost and performance.

Aug 13

Aug 13Thu
  1. Meituan LongCatOfficialAI score46

    LongCat-2.0 Free for One Week on Nous Portal with Hermes Agent

    AILongCat-2.0, Meituan's model, is now live on the Nous Portal and free to try with Hermes Agent for one week. Nous describes it as a 1.6T-parameter MoE with a 1M context built for agentic coding, scoring 70.8 on Terminal-Bench 2.1. It can ingest an entire codebase in one pass, and the Portal is at

  2. Ali GhodsiXAI score22

    Databricks CEO says AI agents with enterprise context drive 80% growth

    AIAli Ghodsi attributes Databricks' 80% growth at $7B to enterprise AI agents becoming usable, now that Genie Ontology automates the capture of organizational context. He says over 70% of queries on the platform are now generated by Genie agents, and that this usage drives consumption and revenue.

  3. Matei ZahariaXAI score44

    Databricks adds Smart Routing to Unity AI Gateway for coding agents

    AIDatabricks has made Smart Routing available in Unity AI Gateway to improve coding agent quality and cost. It matches each coding task to a suitable model and harness based on task needs while preserving good cache hit rates. Databricks says this can match frontier quality while cutting task costs by 30% or more.

  4. Varun MohanXAI score52

    Gemini 3.7 Flash Goes Live in Google Antigravity for Coding and Agents

    AIGoogle has made Gemini 3.7 Flash available in Google Antigravity, described as its most intelligent workhorse model yet for coding and agents. Users can download or upgrade Antigravity to try the model. The author, Varun Mohan, says the model brings a big capability improvement at half the API cost.

  5. Augment Code BlogOfficialAI score44

    Augment Code Expands AI Review Loop to Automate PR-to-Merge Workflow

    AIAugment Code describes an expanded AI-native review system in which specialized agents handle review, repair, and verification from pull request to merge. Humans still make judgment calls and the final merge decision, with the company claiming a 3× increase in code output in its earlier review system.

  6. Google AI DevelopersOfficialAI score75

    Google releases Gemini 3.7 Flash for coding and agentic tasks

    AIGoogle AI Developers announced Gemini 3.7 Flash as its most intelligent workhorse model yet for coding and agents, citing higher instruction adherence, first-pass code accuracy, and high-quality agentic execution. The post shows the model building a complex 3D web game in Antigravity, covering Three.js engine logic, asset orchestration with PBR textures and Nano Banana sprite sheets, and procedural sound effects.

    Why it matters: The post shows a concrete build workflow across engine logic, assets, and audio, which helps readers judge how the model handles multi-step agentic coding.

    Video from @googleaidevs's post
  7. koray kavukcuogluXAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

    Image from @koraykv's post
  8. Augment Code BlogOfficialAI score22

    Augment Code uses Cosmos to check enterprise pilot health against usage and deal data

    AIAugment Code's Solutions Architecture lead used the Cosmos agentic orchestration platform to build a live pilot-health view that combines product usage, GitHub and PR activity, Salesforce deal data, and customer call transcripts. Each account's health and board-level one-liner was checked against the customer's own stated success criteria, such as a 30% PR merge-time reduction. The article says the view refreshed from current Salesforce data and was designed to avoid inflating usage numbers through session lineage reconciliation.

  9. Ali GhodsiXAI score38

    Databricks passes $7B revenue run-rate, growing 80% year over year

    AIDatabricks announced it crossed a $7 billion revenue run-rate, growing over 80% year over year in Q2, and raised $5 billion in its latest fundraise. Lakebase reached a $100 million-plus run-rate, and Lakehouse reached $1.5 billion-plus, growing over 100% year over year, with continued positive adjusted free cash flow. The company says it will invest the capital in Lakebase, a serverless Postgres database for AI agents; Genie, AI coworkers for business data; and Unity AI Gateway, multi-AI governance for controlling costs.

  10. DeepSeekOfficialAI score68

    DeepSeek Harness v0.1 enters Developer Preview as an open-source agent harness

    AIDeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.

    Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.

  11. DeepSeekOfficialAI score62

    DeepSeek launches V4-Pro with Agent upgrades and OpenAI Responses API support

    AIDeepSeek announced the launch of DeepSeek-V4-Pro, citing major Agent upgrades and flexible reasoning effort settings of low, high, and max for V4-Pro and V4-Flash. The model supports the native OpenAI Responses API and is optimized for Codex with one-click setup. V4-Pro is available on the app and web through Expert Mode and via API, with model names unchanged.

    Why it matters: The post lists reasoning effort levels and OpenAI Responses API support, giving developers concrete settings to weigh against their existing Agent and coding workflows.

    Image from @deepseek_ai's post
  12. DeepSeek API NewsOfficialAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

Aug 12

Aug 12Wed
  1. DeepSeek · new models on Hugging FaceOfficialAI score78

    DeepSeek releases DeepSeek-V4-Pro-0813 with stronger agentic benchmark results

    AIDeepSeek has released DeepSeek-V4-Pro-0813 as the official version superseding the V4-Pro preview, built on the preview structure with a DSpark speculative decoding module. The model scores higher than the preview on the listed benchmarks, including Terminal Bench 2.1 at 87.9 and DeepSWE at 62.7, and the weights are under the MIT License.

    Why it matters: The release reports agent benchmark gains over the preview and lists vLLM and SGLang setup, useful for judging deployment cost and fit.

  2. Factory NewsOfficialAI score40

    Factory Launches Agent Effectiveness to Link Droid Usage to Delivery Outcomes

    AIFactory's Agent Effectiveness, now in Private Preview within Factory Analytics, connects Droid sessions to cycle time, work intent, and shipped artifacts drawn from project, issue-tracking, and source control tools. Its Throughput, Output, and Attribution views show where delivery is speeding up, how spend splits across feature, maintenance, bug-fixing, and exploration work, and which projects and issues the output maps to. Admins enable it by connecting Jira, Linear, GitHub, or GitLab and turning on the Advanced Analytics enterprise control.

  3. Cursor ChangelogOfficialAI score42

    Cursor Cloud Agents Start 3x Faster With Builds

    AICursor's Cloud Agents now start from prebuilt copies of development environments, cutting startup time by 3x, with environments booting 10x faster internally and 3x faster time to first token. Builds are included at no additional cost, and failed builds are not activated, so agents keep using the last successful build while users debug in the background.

  4. Jason WeiXAI score22

    Jason Wei argues private knowledge and human presence remain AI-resistant moats

    AIJason Wei argues that as AI gains advantages like driving better than humans, durable human moats remain in private knowledge that language models cannot access, such as high-end real estate and venture capital. He also points to entertainment and the arts, where human creation and achievement carry value, and to human presence, since time spent on someone is meaningful because a finite life runs out.

  5. Liquid AI NewsletterOfficialAI score46

    Liquid AI releases LFM2.5-2.6B model for on-device agentic workloads

    AILiquid AI has released LFM2.5-2.6B, a model optimized to run agentic workflows entirely on-device without cloud escalation. The company said it is designed for high-volume agentic tasks and chained workflows while staying on-device. Separately, Liquid AI and MacPaw announced a long-term partnership to co-develop local AI technology for Mac, with LFMs running on Apple silicon through MacPaw's Elix inference engine.

Aug 11

Aug 11Tue
  1. Zed BlogOfficialAI score72

    Zed introduces Delta, a multiplayer environment for coding with agents

    AIand reviewing their code, and invites first users into a private beta. Delta keeps code and conversations connected through DeltaDB, which captures edits and conversations between git commits and works with existing repositories. The app also supports cloud runners, browser-based sharing, and live syncing of Claude Code sessions.

    Why it matters: The post explains how the new Delta app links conversations with code history, which clarifies a shift in how teams review agent-written changes.

  2. Sequoia CapitalBlogAI score30

    Preview Raises Seed Round to Build AI-Native Video Production Platform

    AISequoia Capital is leading a seed round for Preview, an AI-native video creation and production platform that combines a video timeline with an infinite ideation canvas in one collaborative workspace. The platform lets teams generate with any model, track characters, locations, and props, and keep complete records of how each asset was made. The article says more than 100 studios are already using Preview, from agencies to Hollywood feature filmmakers.

  3. Philipp SchmidBlogAI score60

    Gemini API now combines Google Search and Google Maps in one call

    AIGoogle Search and Google Maps can now be used in the same Gemini API call with Gemini 3.5 Flash and 3.6 Flash. Custom functions and MCP servers can be added to the same request through Tool Combination. The author says Gemini handles the search, place lookup, and function call in one interaction without extra roundtrips from the developer's side.

  4. Aman SangerXAI score22

    Aman Sanger says SpaceXAI will lead general knowledge work next

    AIAman Sanger of Cursor says each AI product wave produced a dominant player, naming OpenAI for chat, Anthropic for coding, and welcoming SpaceXAI for general knowledge work. The post links to Grok Bot, described in quoted context as an early-beta AI teammate that signs into tools, uses them like a person, and returns finished work.