Skip to content

Formats · Latest news

Product updates

New features, redesigns, and commercial changes in AI products and applications.

225 top picks all-time · 140 in the past 30 days · chosen from 2,778 items collected all-time

Latest pick

Top picks archive · Page 9

Top picks 161–180 of 225

Aug 26

Aug 26Wed
  1. Michael TruellXAI score60

    Grok Bot opens to all Grok and Cursor subscribers

    AIGrok Bot is now available to all standard Grok and Cursor subscribers, with SuperGrok and Cursor Pro subscribers included. Cursor's Michael Truell says users are delegating tasks ranging from running small e-commerce businesses to testing production software. Weekly usage limits are also being reset for all users.

    Why it matters: The post reports broader availability and the range of delegated tasks users run, showing how an agent product is being used in practice.

  2. LMSYS OrgOfficialAI score60

    SGLang adds Day-0 support for Qwen3.8-Flash-Next with an NVFP4 checkpoint

    AISGLang announced Day-0 support for Qwen3.8-Flash-Next, a 125B MoE model with 6B active parameters and 51B N-gram embeddings, in collaboration with Alibaba Qwen, NVIDIA, and AMD. The post reports 540 tok/s decode speed at BS=1 on NVIDIA B200 (TP4) with an NVFP4 checkpoint, and says N-gram host offloading saves 23.5 GiB VRAM per GPU and raises KV capacity by 78.5%.

    Why it matters: The source gives concrete serving details, including N-gram host offloading, hybrid attention, and a 540 tok/s B200 decode figure, showing how the model is deployed in practice.

Aug 24

Aug 24Mon
  1. Engineering at MetaOfficialAI score72

    Meta details MetaRoCE, an RDMA transport designed for AI-scale Ethernet

    AIMeta designed MetaRoCE, a clean-sheet RDMA transport for AI workloads on commodity Ethernet, and is releasing its specification, reference software and compliance test suite through the Open Compute Project. On a 64-node AMD GPU cluster running RCCL collectives, the post reports MetaRoCE delivering higher throughput and lower flow completion times than RoCEv2, with about 86% throughput maintained at 1% packet loss.

    Why it matters: The post explains how per-path endpoint intelligence replaces lossless fabric assumptions, with measured throughput and loss results against RoCEv2 on a 64-node AMD cluster.

Aug 19

Aug 19Wed
  1. Liquid AI BlogOfficialAI score60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    AILiquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

Aug 18

Aug 18Tue
  1. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.

  2. Cursor BlogOfficialAI score68

    Cursor explains Continuity, a WAL-based Git storage system

    AICursor's blog describes Continuity, its Git storage system, which stores each push as a write-ahead log entry in S3-compatible object storage. The article contrasts this design with GitHub's earlier Spokes system, which used three-phase commit replication across local disks. Continuity uses stateless replicas that catch up from the log, and the article reports write throughput of up to 120 pushes/s on S3 Standard and over 300 pushes/s on S3 Express One Zone.

    Why it matters: The article explains why hosting Git at scale is hard and how Continuity's WAL-based design compares with the earlier Spokes approach, which is useful background for infrastructure work.

Aug 17

Aug 17Mon
  1. Microsoft Foundry BlogOfficialAI score62

    Microsoft Foundry adds five Claude agent features to Azure-hosted deployments

    AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.

    Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.

  2. Replit BlogOfficialAI score60

    Replit adds black-box pen tests that probe apps like external attackers

    AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.

    Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.

Aug 16

Aug 16Sun
  1. Cursor ChangelogOfficialAI score60

    Cursor launches Origin, a code hosting service with GitHub sync

    AICursor begins rolling out Origin, its code hosting feature, in early beta to all paid plans, excluding enterprise orgs whose admins opt out. Repos can be hosted on Origin, where Origin is the source of truth, or synced from GitHub, where GitHub stays the source of truth and pull requests sync both ways. Vercel, Depot, and Buildkite integrations are already available, and agent-native features are slated to ship soon.

    Why it matters: The source specifies how Origin hosts repos alongside GitHub sync, showing how the hosting source of truth differs between the two types of repo.

Aug 14

Aug 14Fri
  1. Augment Code BlogOfficialAI score62

    Augment rebuilds its Auggie CLI harness on Pi, cutting SWE-bench Pro task cost 53%

    AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.

    Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.

Aug 13

Aug 13Thu
  1. koray kavukcuogluXAI score72

    Google launches Gemini 3.7 Flash for coding and agentic workflows

    AIGoogle launches Gemini 3.7 Flash, its latest Flash model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. The post reports gains from 3.5 to 3.7 Flash, including DeepSWE v1.1 rising from 37.0% to 65.3%, Code Arena Elo from 1506 to 1588, and AutomationBench from 13.4% to 30.4%.

    Why it matters: The post pairs a launch with specific before-and-after benchmark gains and an introductory price, letting readers weigh capability against cost for coding and agent work.

    Image from @koraykv's post
  2. DeepSeekOfficialAI score68

    DeepSeek Harness v0.1 enters Developer Preview as an open-source agent harness

    AIDeepSeek has released DeepSeek Harness v0.1 in Developer Preview, opening the codebase under the MIT license for developers building agent harnesses. The harness is built on the Cordis meta-framework and treats models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI as plugins that can be mixed, matched, replaced, and extended.

    Why it matters: The source specifies the MIT license and a plugin-based architecture covering models, tools, and sessions, which helps developers assess extensibility before adopting it.

  3. DeepSeek API NewsOfficialAI score62

    DeepSeek-V4-Pro Reaches GA with Agent Gains and Peak/Off-Peak API Pricing

    AIDeepSeek has made DeepSeek-V4-Pro generally available on its app, web, and API, with the API model name set to deepseek-v4-pro. The release reports agent benchmark results, including 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified. It also adds native OpenAI Responses API support, low/high/max thinking effort levels, and off-peak API prices set at half of peak prices starting 16:00 UTC on August 16, 2026.

    Why it matters: The update pairs new agent benchmark results with API format and pricing changes, so developers can judge both capability and cost impact before migrating.

  4. Prime Intellect BlogOfficialAI score70

    Prime Intellect releases Prime Flash MoE kernels for faster Blackwell inference

    AIPrime Intellect has released Prime Flash MoE, a set of Blackwell-optimized CUDA kernels for mixture-of-experts feed-forward layers. The kernels are up to 2.4× faster than PyTorch grouped GEMM and deliver about 2.3× speedup across the 4k–128k token range, and are integrated into its prime-rl framework. Two pipelines are offered: a fused single-kernel path for small problem sizes and a split three-kernel path for larger ones, supporting both bf16 and MXFP8.

    Why it matters: The post explains how fusing MoE expert computation on Blackwell hardware avoids intermediate memory traffic, with benchmarks showing where fused and split pipelines each win.

Aug 11

Aug 11Tue
  1. Zed BlogOfficialAI score72

    Zed introduces Delta, a multiplayer environment for coding with agents

    AIand reviewing their code, and invites first users into a private beta. Delta keeps code and conversations connected through DeltaDB, which captures edits and conversations between git commits and works with existing repositories. The app also supports cloud runners, browser-based sharing, and live syncing of Claude Code sessions.

    Why it matters: The post explains how the new Delta app links conversations with code history, which clarifies a shift in how teams review agent-written changes.

Aug 9

Aug 9Sun
  1. Fireworks AI BlogOfficialAI score60

    Meta releases Muse Glimmer 30B, available on Fireworks for always-on agents

    AIMeta's Muse Glimmer is a 30B dense model with a 128K+ token context window, now available on Fireworks in serverless and on-demand deployments. Meta reports it leads its size class on MCP Atlas (75.5) and DeepSearch QA (74.6) against Gemma 4 31B and Qwen 3.6 27B, with its sliding-window attention and two KV heads keeping the cache small for concurrent agent sessions.

    Why it matters: The post pairs an architecture explained through KV cache size with benchmark tables against two rival models, which helps readers judge whether it fits their agent workload.

Aug 7

Aug 7Fri
  1. Prime Intellect BlogOfficialAI score62

    Prime Intellect adds multi-agent training and evaluation to PRIME-RL

    AIPrime Intellect's RL stack now supports multi-agent systems, letting users program interactions between agents, choose which roles learn, and assign credit across an episode. The release introduces Agent and Env abstractions and four example patterns: agentic judging, self-play, and user simulation. Multi-agent support ships today in verifiers 0.3.0 and prime-rl 0.8.0.

    Why it matters: The post explains the Agent and Env abstractions and four multi-agent patterns, showing how roles, credit assignment, and episodes can be programmed in one RL stack.

Aug 5

Aug 5Wed
  1. Prime Intellect BlogOfficialAI score75

    Prime Agent launches open-source self-improving RLM coding harness

    AIPrime Agent is a new open-source coding harness built on a persistent IPython kernel, a Recursive Language Model design, and Continual Harness state that the agent can create, read, update, and delete. Prime Intellect reports ARC-AGI-3 results of 95.5% RHAE Best@1 with Opus 5 and competitive long-context scores with the open-weights GLM-5.2 model.

    Why it matters: The post explains how the RLM and Continual Harness designs let an agent write code against its own context, sub-agents, and harness state, with benchmark evidence.

Aug 4

Aug 4Tue
  1. Zed BlogOfficialAI score65

    Zed Enables OS-Level Sandboxing by Default for Its Agent Panel

    AIZed's agent panel now sandboxes its terminal and fetch tools by default, starting in release 1.14, and the restrictions are enforced by the operating system rather than by agent instructions. By default the sandbox blocks writes outside project directories, writes to .git, and network requests, and agents can request temporary escalation with a stated reason. The post also notes that sandboxing covers only those tools and does not protect against other tools, external programs, or the regular built-in terminal.

    Why it matters: The post explains how OS-enforced sandboxing limits agent terminal and fetch access, and why fine-grained command rules fall short of it.

Jul 30

Jul 30Thu
  1. MiniMax BlogOfficialAI score72

    MiniMax H3 unifies text, image, video, and audio generation in one model

    AIMiniMax launches H3, a general-purpose multimodal generation model that understands text, images, video, and audio as unified context. It generates video up to 15 seconds at 2K resolution with native stereo sound, and the company says model weights will be opened in the coming days, subject to applicable laws and regulations. MiniMax also says H3 is priced below mainstream models at 2K and 768p.

    Why it matters: The post explains how a unified multimodal design and training choices enable 2K video with native stereo sound, useful for comparing against closed video generators.