Skip to contentSkip to stories

Updated

Agents

Showing low-relevance items too. Hide low-relevance items

Aug 26

Aug 26Wed
  1. Michael TruellXAI score60

    Grok Bot opens to all Grok and Cursor subscribers

    AIGrok Bot is now available to all standard Grok and Cursor subscribers, with SuperGrok and Cursor Pro subscribers included. Cursor's Michael Truell says users are delegating tasks ranging from running small e-commerce businesses to testing production software. Weekly usage limits are also being reset for all users.

    Why it matters: The post reports broader availability and the range of delegated tasks users run, showing how an agent product is being used in practice.

  2. Google AI DevelopersOfficialAI score47

    Google launches Gemini 3.5 Transcribe, a speech-to-text model for developers

    AIGoogle has released Gemini 3.5 Transcribe, a speech-to-text model that filters out spoken hesitations and accurately grounds technical terms, file names, and code variables against the active context. The model uses visual biasing to incorporate screen-aware context into developer workflows, as demonstrated in Antigravity.

    Video from @googleaidevs's post
  3. LMSYS OrgOfficialAI score65

    Zhipu's GLM-5.3-Flash adds native vision with day-0 SGLang support

    AIZ.ai released GLM-5.3-Flash, a 320B-A18B model, with day-0 support in SGLang, after appearing earlier as ox-alpha. The post calls it the first native multimodal model in the GLM-5 series and says it outperforms GLM-5.2 at one-tenth the cost, with stable 1M-token long-context performance.

    Why it matters: The post reports GLM-5.3-Flash's native multimodal design, its efficiency claims, and day-0 SGLang support, which bear on running it in practice.

Aug 25

Aug 25Tue
  1. Fireworks AI BlogOfficialAI score40

    DeepSeek V4 Pro 0813 Tops SWE-Bench and Cuts Cost per Solved Task

    AIDeepSeek V4 Pro 0813 scored 95.2% on SWE-Bench Verified, ahead of Kimi K3 at 92.6% and Fable 5 at 85.4%, in Fireworks AI's eval runs. It costs $0.309 per solved task on SWE-bench versus $0.808 for Fable 5, and it is available through Fireworks serverless and dedicated endpoints, with SFT, DPO, and RFT training support. Its 1M-token context window and native tool calling target long-horizon agentic workloads, though its Java accuracy on Aider Polyglot (48.9%) trails Fable 5 (74.5%).

  2. Fireworks AI BlogOfficialAI score46

    DeepSeek V4 Pro Solves Security Tasks at Half the Cost Per Success

    AIDeepSeek V4 Pro 0813 recorded zero refusals across 840 adversarial security tasks in CyberGym testing, solving them at about half the cost per success of the top-scoring model tested, Kimi K3. In the 697-task common cohort, V4 Pro reached a 53.7% reward rate at $2.50 per solved task, versus 47.6% and $9.64 for GPT-5.5 and 5.9% and $33.28 for Claude Opus 4.8.

  3. Fireworks AI BlogOfficialAI score52

    Harvey Tenet, a legal model post-trained from Kimi K3 with Fireworks

    AIHarvey and Fireworks post-trained Tenet from the Kimi K3 base using asynchronous reinforcement learning on the Fireworks Training API for long-horizon legal work. On the Legal Agent Benchmark, Tenet reached 19.7% all-pass versus 10.8% for base Kimi K3, and its cost per task was $5.92 versus $5.62.

  4. Z.ai Release NotesOfficialAI score62

    Z.ai releases GLM-5.3-Flash with native visual capabilities and hybrid architecture

    AIZ.ai has released GLM-5.3-Flash, a model with native visual capabilities that observe interfaces, rendering results, and interaction feedback across code, browsers, and GUIs. It uses a hybrid linear and sparse attention architecture with 320B total parameters and 18B activated, which the company says significantly reduces compute and KV-cache requirements. The release notes also describe support for office document and financial research workflows.

    Why it matters: The release notes give GLM-5.3-Flash's architecture, parameter counts, and cybersecurity findings, which make the model's scope concrete for comparison with earlier GLM releases.

  5. Andrew NgXAI score46

    OpenWorker adds built-in security agents for code, dependencies, and cloud

    AIOpenWorker, an open source agent that completes tasks on a laptop, has released a new version with built-in cybersecurity agents. The agents scan code for vulnerabilities, scan dependencies for supply chain injections, and check cloud security configurations for attack surfaces. Users can run open weight models locally so sensitive code stays on their machine.

  6. Z.ai (GLM) · new models on Hugging FaceOfficialAI score72

    Z.ai releases GLM-5.3 open weights with gains from post-training

    AIZ.ai released GLM-5.3 on Hugging Face, built on the same base model as GLM-5.2, with all gains coming from post-training. The source reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam, with a benchmark table comparing it against Kimi K3, DeepSeek-V4 Pro-0813, Qwen3.8-Max, and others.

    Why it matters: The source gives benchmark tables against GLM-5.2 and rival models, showing where the post-training gains concentrate in coding and cyber tasks.

Aug 24

Aug 24Mon
  1. InferactOfficialAI score58

    Inferact details vLLM optimizations for AgentX agentic coding benchmark

    AIInferact, working with vLLM and SemiAnalysis, reports vLLM throughput results on the AgentX multi-turn agentic coding benchmark for DeepSeek V4 Pro, MiniMax M3, and Kimi K3. The thread attributes gains to sparse prefix-cache retention, a distributed KV pool with Mooncake Store, and prefill-decode disaggregation via NIXL, reporting 4.45x higher throughput for DeepSeek V4 Pro on GB300 Dynamo compared to B300 at 60 tok/s interactivity. A full technical blog is promised later this week.

  2. Microsoft AI BlogOfficialAI score14

    Five Signals Show How Organizations Scale AI Through Security, Governance, and Observability

    AIMicrosoft's AI Blog outlines five signals that trust, not speed alone, lets organizations scale AI from pilots to enterprise-wide use. Its first signal is observability, citing Microsoft's Cyber Pulse AI Security Report finding that 29% of employees use unsanctioned AI agents their security teams cannot see. The post also says security should be built into AI systems by design and governance should be continuous rather than a one-time approval.

Aug 21

Aug 21Fri
  1. Ian Johnson 🔬🤖XAI score12

    CurieOS helps agents and people accelerate science and engineering collaboration

    AIIan Johnson says putting many disciplines on one platform speeds up all of them as agents and people collaborate. The example given is CurieOS, which handled literature review and calculations for a V1 jet impingement lid and proposed a funneled jet geometry in V2 that cut pressure drop with minimal engineer steering. The V3 design is being validated and built in parallel with other work on the platform.

  2. swyxXAI score39

    Swyx Says Simulating Humans Could Be Last Barrier to Automated AI Research

    AISwyx argues that simulating humans and their feedback is likely the final barrier to recursive self-improvement, where models automate increasingly large parts of ML research. He says Simile, which builds human simulations, is already finding product-market fit with Fortune 100 companies despite its early stage.

  3. Andrew NgXAI score31

    Andrew Ng outlines six core skills for building and deploying AI applications

    AIAndrew Ng's AI Engineering Skills Map ranks building and deploying AI applications as the top skill tier, spanning LLM foundations, data grounding, agentic systems, evaluation-driven development, production operations, and machine learning foundations. He explains that AI outputs are less predictable than traditional software, so skilled engineers build iteratively, examining results and deciding next steps based on intermediate outcomes. The skills map was derived from job postings, expert interviews, and survey responses.

  4. Amazon ScienceOfficialAI score50

    SOP-Bench Tests AI Agents on Real Business Procedures Across 12 Industries

    AIAmazon Science released SOP-Bench, an open benchmark that measures how well AI agents execute standard operating procedures written by domain experts. It covers 12 business areas, including healthcare intake and dangerous-goods classification, with more than 2,000 tasks, working tools, and ground-truth answers. The benchmark was presented at the 2026 KDD conference.

  5. DeepSeekOfficialAI score62

    DeepSeek releases experimental multimodal model V4-Flash-Vision-Exp on its API

    AIDeepSeek has made its experimental multimodal model DeepSeek-V4-Flash-Vision-Exp available on the DeepSeek API Platform. The company says it matches DeepSeek-V4-Flash on text tasks, including agents, reasoning, and world knowledge. On multimodal agent benchmarks it improves substantially over V4-Flash and approaches Opus-4.8, and DeepSeek Harness 0.1.1 was released the same day with support for the new model.

    Why it matters: The post pairs a new multimodal model with a benchmark table against V4-Flash and Opus-4.8, showing where the gains and remaining gaps sit.

    Image from @deepseek_ai's post
  6. DeepSeek API NewsOfficialAI score60

    DeepSeek releases experimental vision model DeepSeek-V4-Flash-Vision-Exp on its API

    AIDeepSeek has made DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding model, available on its API platform via model='deepseek-v4-flash-vision-exp'. The source says its pure-text capabilities are on par with DeepSeek-V4-Flash, while it shows a significant leap on agent benchmarks requiring visual understanding, which it says brings multimodal agent capabilities close to Opus-4.8.

    Why it matters: The source gives benchmark scores and a model identifier, so readers can compare the experimental vision model against the text-only DeepSeek-V4-Flash on agent tasks.

Aug 20

Aug 20Thu
  1. swyxXAI score34

    Matt Pocock's /wayfinder skill navigates unclear projects with research and grilling

    AIMatt Pocock's /wayfinder skill is designed for "fog of war" situations where the end state of a project is unclear. It orchestrates research and other grill sessions to help users discover what they don't yet know, building on his popular /grill-me skill. Latent Space is featuring the skill in an exclusive interview as the first in a series of Skills coverage.

    Image from @swyx's post
  2. Mistral AIOfficialAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Jazzyear · InsightsNewsAI score29

    Jazzyear's 2026 tech investment conference maps where capital is flowing in AI and hard tech

    AIAt the 2026 Jiazi Gravity Tech Industry Investment Conference in Beijing, Jiazi Guangnian's CEO Zhang Yijia said first-half 2026 saw investment amounts rise 91.6% year on year, IPOs rise 39.2%, and M&A transaction value double. The report said AI absorbed over 70% of global venture investment, with OpenAI and Anthropic together raising $217 billion, roughly 40%.

  2. Matei ZahariaXAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  3. JetBrains AI BlogOfficialAI score31

    Air Adds Multiproject View, Markdown Rendering, and Windows IME Fixes

    AIAir's latest release lets users open several projects in one window and run agents across them in parallel, with tasks grouped by project in the sidebar. Markdown files now render as formatted text while editing, with syntax shown only when editing, and Chinese, Japanese, Korean, and other IMEs now work on Windows. The release also adds a Customize screen for keymap, theme, and accent color, and lets users choose the agent and model for Agent Review.

  4. TinkerOfficialAI score38

    Qwen3.8-27B is now available on Tinker

    AITinker has made Qwen3.8-27B available today. The model is natively multimodal, handling images and video, with flexible thinking control. Tinker says it performs meaningfully better at coding, professional work, research, and long-horizon agentic tasks.

  5. Kimi.aiOfficialAI score24

    Kimi Work tutorial shows financial analysts three research workflows

    AIKimi publishes Tutorial #2 for its Kimi Work product, showing financial analysts how to use it for three investment research tasks. The tasks are building a live investor dashboard, updating financial models in spreadsheets, and processing and generating reports in batch. The post promises more Kimi Work workflows to follow.

    Video from @Kimi_Moonshot's post

Aug 18

Aug 18Tue
  1. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.