Skip to content

Areas · Latest news

Deployment & engineering

Engineering practice for running models: inference optimization, memory and cost, serving architecture, and infrastructure choices.

210 top picks · 101 in the past 30 days · chosen from 2,088 items collected

Latest pick

Top picks archive · Page 5

Top picks 81–100 of 210

Sep 21

Sep 21Mon
  1. Amazon ScienceAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  2. LMSYS OrgAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

Sep 18

Sep 18Fri
  1. Anthropic NewsroomAI score62

    Anthropic partners with Accenture on embedded AI model evaluation

    AIAnthropic is partnering with Accenture, through its specialist AI business Faculty, on independent evaluation of frontier models, including red-teaming, alignment assessments, and safeguard testing. Anthropic and Accenture each expect to invest at least $1 billion in this capacity over five years. The source says embedded evaluators would have employee-comparable access, but standards for access and reporting, and a settled funding system, do not yet exist.

    Why it matters: The source ties a new evaluation arrangement to an unresolved question of who funds and sets standards for independent AI evaluators, which is useful context for governance debates.

Sep 17

Sep 17Thu
  1. Google AI StudioAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

Sep 15

Sep 15Tue
  1. Zed BlogAI score72

    Zed launches Delta public beta to replace pull requests with agent threads

    AIZed has launched the public beta of Delta, a multiplayer environment for coding with agents and reviewing their work, which replaces pull requests with shared threads. Delta is built on DeltaDB, which records edits and messages between Git commits, and it is free during the beta, with paid plans for individuals and teams to follow.

    Why it matters: The post explains how Delta replaces pull requests with shared agent threads and DeltaDB, showing a concrete alternative to the GitHub review workflow.

  2. Claude Apps Release NotesAI score72

    Claude Cowork moves into every conversation, adding designs, slides, and docs

    AIClaude now makes Cowork capabilities available from any conversation without choosing a mode first, with chats, tasks, projects, connectors, and skills carrying over. Users can also create designs, decks, and docs in any conversation, including Claude Code and the Artifacts tab, and edit them with Claude.

    Why it matters: The release merges Cowork tasks into ordinary chats and adds design, slide, and doc creation, changing how Claude users start larger work.

  3. Google AIAI score72

    Google rolls out Gemini 3.8 Live and Extended Thinking across consumer, developer, and enterprise channels

    AIGoogle is rolling out Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking across several channels. Consumers get them in Search Live and Gemini Live, developers get public preview access through the Gemini API, and enterprises get private preview through Gemini Enterprise, with Customer Experience support coming soon.

    Why it matters: The post lays out where each Gemini 3.8 Live variant reaches consumers, developers, and enterprises, which clarifies access paths for a voice model release.

  4. Google DeepMindAI score72

    Google DeepMind releases Gemini 3.8 Live models for real-time voice agents

    AIGoogle DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two live dialogue models for voice agents. Extended Thinking scores 82.6 on Artificial Analysis' Speech to Speech Quality Index, 68.6% on τ-Voice, and 97.7% on Big Bench Audio. Gemini 3.8 Live is rolling out now in the Gemini API, Google AI Studio, and Search Live, with enterprise access in private preview.

    Why it matters: The release covers a voice model's benchmark results and availability across developer, enterprise, and consumer products, useful for judging voice agent options.

  5. Cognition Blog (Devin, Windsurf)AI score60

    Cognition and AWS sign multi-year deal to deploy Devin for enterprise modernization

    AICognition and AWS have entered a multi-year Strategic Collaboration Agreement to help enterprises deploy the Devin autonomous engineer in production. Devin can be purchased through AWS Marketplace, and the companies are exploring deeper engineering integrations within customers' AWS environments. Mercedes-Benz reportedly used Devin to analyze more than 200,000 lines of COBOL, reducing an estimated eight-month modernization project to eight days.

    Why it matters: The collaboration shows how an autonomous coding agent is being packaged for enterprise legacy modernization inside existing AWS environments, with concrete customer migration figures.

  6. Gemini API ChangelogAI score62

    Google makes Gemini 3.8 Live models generally available for real-time voice

    AIGoogle has made two audio-to-audio models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, generally available through the Live API. Gemini 3.8 Live, model ID gemini-3.8-live, is the default for low-latency voice agents, with interleaved reasoning and asynchronous function calling. Gemini 3.8 Live Extended Thinking, model ID gemini-3.8-live-extended-thinking, supports background reasoning during live audio and is recommended when more reasoning is needed.

    Why it matters: The changelog names two model IDs and their intended use, showing how Live API developers can choose between low-latency voice and higher background reasoning.

Sep 14

Sep 14Mon
  1. Google Developers BlogAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  2. vLLM BlogAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  3. Google · new models on Hugging FaceAI score62

    Google releases EmbeddingGemma 2, an open multimodal embedding model

    AIGoogle DeepMind released EmbeddingGemma 2, an open model under Apache 2.0 that maps text, images, video, and audio into one shared 768-dimensional vector space. The model has 740M total parameters and supports 8,192-token context, with Matryoshka truncation to 128d, 256d, and 512d. The source reports 14% better code-task performance than EmbeddingGemma 1 and says it is designed for consumer hardware such as phones and laptops.

    Why it matters: The release combines text, image, video, and audio retrieval in one 768-dimensional space at 740M parameters, a useful reference for on-device multimodal search design.

Sep 11

Sep 11Fri
  1. Augment Code BlogAI score80

    Augment Code details how its software factory raised output per developer 4.5×

    AIAugment Code reports that size-adjusted output per active developer rose from 12.3 to 55.7 between November 2025 and July 2026, while median time to merge fell from 11.2 to 3.1 hours. The post says the company added specialized agents wherever work was piling up, across planning, review, verification, feedback, and incident response, and kept engineers responsible for product decisions, architecture, and production risk.

    Why it matters: The post pairs internal productivity and quality metrics with the order in which agents were added, showing how review and verification bottlenecks shaped a software delivery pipeline.

Sep 10

Sep 10Thu
  1. Sherwin WuAI score72

    OpenAI launches Agents API for building cloud agents on Codex harness

    AIOpenAI has launched the Agents API, a cloud-based way to build agents backed by the Codex harness. Developers can connect their favorite tools and connectors and attach agents to any sandbox. The author says the API lets firms build scaled agents and expose them inside their own internal AI applications.

    Why it matters: The source describes an API for building cloud agents on the Codex harness, useful for teams planning to embed agents in internal applications.

  2. Cognition Blog (Devin, Windsurf)AI score66

    Cognition releases SWE-2, a coding model trained with cost-penalized RL

    AICognition introduces SWE-2, a coding model post-trained from Kimi K3 that scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while costing 64% less. The post attributes the gains to an RL algorithm that trains all reasoning-effort levels in one run, with cost penalties tuned to the base model's Pareto frontier. SWE-2 is available starting today in Devin Desktop and CLI, with rollout to Devin Web and Fusion.

    Why it matters: The post explains how the cost penalty and length-weighted baseline are derived, which helps readers judge the tradeoffs in coding model post-training.

  3. DeepSeekAI score72

    DeepSeek V4.1-Flash goes live on its API with native multimodal support

    AIDeepSeek says V4.1-Flash is now live on its API with native multimodal support, accessed through the model name deepseek-flash. The older V4-Flash and V4-Flash-Vision-Exp are retired, while deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. Requests to deepseek-v4-pro will route to V4.1-Flash at V4.1-Flash rates starting 04:00 UTC on Sept 14, 2026, until V4.1-Pro launches.

  4. DeepSeek API NewsAI score72

    DeepSeek releases V4.1-Flash with native multimodal support and API updates

    AIDeepSeek officially released DeepSeek-V4.1-Flash, the smallest model in its new architecture family, with native multimodal visual understanding. The API now serves it under the model name deepseek-flash, while V4 Flash and V4 Flash Vision Exp were retired and routed to V4.1 Flash. API prices were reduced with the release, and V4 Pro remains available after September 14, 2026.

    Why it matters: The release lists benchmark results alongside API model-name changes and retirements, so developers can check both capability claims and migration steps.

Sep 9

Sep 9Wed
  1. Cursor ChangelogAI score73

    Cursor launches Projects for long-running, multi-agent coding work

    AICursor is launching Projects, a beta feature for larger work such as a feature, migration, or full app, rolling out to all users starting today. A coordinator agent plans the work, delegates it to implementing agents that can run in parallel, and runs on a cloud computer so it continues when the laptop is closed. Each Project keeps shared context files synced across cloud and local machines, and subscriptions let the coordinator act on Slack channels, schedules, or PRs without a prompt.

    Why it matters: The source details how a coordinator agent plans, delegates, and syncs shared context across cloud and local machines, useful for judging how long-running agent work might fit a team's workflow.

  2. Fireworks AI BlogAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.