Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Apr 25

Apr 25Sat

Apr 24

Apr 24Fri
  1. DeepSeek API NewsAI score67

    DeepSeek API adds V4-Pro and V4-Flash, retiring legacy model names in July 2026

    AIThe DeepSeek API now supports V4-Pro and V4-Flash through both the OpenAI ChatCompletions and Anthropic interfaces. Developers keep the same base_url and set the model parameter to deepseek-v4-pro or deepseek-v4-flash. The legacy names deepseek-chat and deepseek-reasoner will be discontinued on 2026-07-24, and until then they map to the non-thinking and thinking modes of deepseek-v4-flash, respectively.

    Why it matters: The source gives exact model names, an unchanged base URL, and a July 2026 discontinuation date, so developers can plan their migration from legacy names.

Apr 23

Apr 23Thu

Apr 22

Apr 22Wed
  1. Factory NewsAI score38

    Factory's Automated QA Skill Tests Apps Like Real Users and Posts Reports to PRs

    AIFactory has released an Automated QA skill that drives an app as a real user would, filling forms, typing into terminals, and calling endpoints, then posts a structured report with screenshots, terminal snapshots, and API traces as a single updating comment on each pull request. Teams can run it on every push or make it an optional CI check triggered by a PR label, comment command, or manual dispatch, and developers can run /qa locally in any Droid session. Automated QA is available today in all Factory plans.

  2. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

  3. Anthropic EngineeringAI score78

    Anthropic traces Claude Code quality complaints to three product changes

    AIAnthropic says three changes to Claude Code, the Claude Agent SDK, and Claude Cowork caused recent quality complaints, and the API was not affected. The fixes were resolved by April 20 (v2.1.116), and the company is resetting usage limits for all subscribers as of April 23.

    Why it matters: The postmortem traces three separate changes to specific dates and versions, showing how a bug in context management can look like broad degradation to users.

  4. koray kavukcuogluAI score49

    Google unveils 8th-generation TPUs, with 8t for training and 8i for inference

    AIAt Google Cloud Next this week, Google introduced its 8th-generation TPUs, split into two variants: 8t for massive-scale training and 8i for low-latency inference. Google presents the launch as a milestone in its accelerator roadmap, aimed at optimizing the full AI stack. The post links to a blog with further details on the systems architecture.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Michael TruellAI score62

    Cursor partners with SpaceX to scale up Composer, with an option to acquire

    AICursor's Michael Truell says the company is partnering with the SpaceX team to scale up Composer, calling it a meaningful step toward building the best place to code with AI. The quoted SpaceX post says Cursor gives SpaceX the right to acquire Cursor later this year for $60 billion, or pay $10 billion for the work together. It also cites SpaceX's Colossus training supercomputer, described as a million H100-equivalent system, as a source of training capacity.

  3. NVIDIA AI DeveloperAI score29

    NVIDIA OpenShell v0.0.34 adds live sandbox policy updates and VM installs

    AINVIDIA's OpenShell v0.0.34 release lets users update sandbox policy without restarting the runtime. The update also adds install-vm, which installs the gateway and VM driver with new --driver-dir support, and sandbox get, which shows the active runtime policy. Supervisor seccomp improvements and HTTP normalization are included as well.

Apr 20

Apr 20Mon
  1. NVIDIA AI DeveloperAI score35

    OpenShell v0.0.33 adds hardened sandboxing and a standalone libkrun driver

    AINVIDIA released OpenShell v0.0.33, which adds seccomp and process-limit hardening, inference routing, and a standalone libkrun compute driver. The libkrun driver provides a lightweight VM backend as a second compute path for agents. The release also includes bug fixes, a docs refresh, and improved test stability.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue

Apr 13

Apr 13Mon
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

  2. Cognition Blog (Devin, Windsurf)AI score49

    Devin Introduces New Self-Serve Plans and Charges for Ask Devin and Devin Review

    AIDevin is retiring its Core and Team plans for a new lineup of Free, Pro at $20/month, Max at $200/month, Teams with usage-based billing and an $80/month minimum, and custom-priced Enterprise. Ask Devin's Deep Mode, Devin Review after a 2-week free trial, and higher-quality DeepWiki generation will move to usage-based billing, with DeepWiki's existing generation and open-source Devin Review remaining free. Self-serve usage beyond included quota will be billed in dollars rather than ACUs.

Apr 10

Apr 10Fri

Apr 9

Apr 9Thu

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

  2. Cognition Blog (Devin, Windsurf)AI score31

    Cognition Expands to Japan, Appoints Takumi Masai to Lead Devin Launch

    AICognition is expanding into Japan, its first step into Asia, and has appointed Takumi Masai as Japan President and General Manager to lead a local team working with Japanese enterprises. DeNA has used Devin to more than double operational efficiency across multiple engineering functions, and Mizuho Securities has deployed it as one of the first large-scale financial institutions in Japan.

  3. Stability AIAI score44

    Stability AI launches Brand Studio, a creative production platform built around brand identity

    AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.

Apr 7

Apr 7Tue
  1. Cognition Blog (Devin, Windsurf)AI score70

    How Devin Is Modernizing COBOL at Fortune 500 Companies

    AICognition describes how Devin handles COBOL modernization at several Fortune 500 companies, citing a shortage of COBOL developers and 68% failure rates for such efforts. The post identifies three obstacles for agents: untraceable data across copybooks, little COBOL in model training, and no way to run code on Linux-based VMs. It says Devin succeeds on documentation, batch migrations, and large-scale refactoring, while transactional workloads remain out of reach.

    Why it matters: The post explains why agents struggle with COBOL and which workloads they can migrate, giving a framework for judging where automation fits legacy systems.

  2. Anthropic EngineeringAI score67

    Anthropic decouples agent brain, hands, and session in Managed Agents

    AIAnthropic's Managed Agents separates the harness, sandbox, and session into independently replaceable interfaces. The source says this design let failed containers be replaced, kept tokens out of the sandbox, and reduced p50 time-to-first-token by roughly 60% and p95 by over 90%.

    Why it matters: The post explains how decoupling the harness, sandbox, and session changed failure recovery, credential security, and latency, offering a reusable architecture pattern for long-running agents.

  3. Werner VogelsAI score62

    Amazon S3 Files lets users mount any S3 bucket as a filesystem

    AIWerner Vogels announced S3 Files, which lets users mount any S3 bucket as a filesystem without making copies, running sync scripts, or choosing between file and object storage. He linked to a detailed post by Andy Warfield on the feature and its design history, including the filerectories concept that did not make the final release.

Apr 6

Apr 6Mon
  1. Cognition Blog (Devin, Windsurf)AI score44

    Windsurf releases SWE-1.6, a software engineering model optimized for speed and user experience

    AIWindsurf has made SWE-1.6, its model for software engineering agents, generally available, with the company saying it improves on the SWE-1.6 Preview by reducing overthinking, looping, and sequential tool calls. The model is free for three months, with a free version offered at 200 tok/s through Fireworks and a faster paid version at 950 tok/s through Cerebras.

Apr 2

Apr 2Thu
  1. Werner VogelsAI score25

    AWS launches Sustainability console for carbon emissions measurement

    AIAWS has introduced the AWS Sustainability console, which puts emissions data in the hands of users without exposing billing data. The post frames carbon and cost as similar signals that must be measured to be optimized. It links to AWS's announcement covering programmatic access, configurable CSV reports, and Scope 1–3 reporting in one place.

    Image from @Werner's post

Mar 31

Mar 31Tue
  1. Intern Large ModelsAI score52

    Intern Large Models unveils Kernel-Smith for generating GPU kernels and operators

    AIIntern Large Models introduced Kernel-Smith, a framework for generating high-performance GPU kernels and operators using an evolutionary agent and post-training recipe. The post says it outperforms Gemini-3.0-pro and Claude-4.6-opus on Kernel-Bench, and that optimized kernels have been merged into SGLang and LMDeploy. The accompanying figure compares best program score trajectories across evolution steps, with Kernel-Smith-235B-RL reaching the highest peak.

    Image from @intern_lm's post

Mar 27

Mar 27Fri

Mar 26

Mar 26Thu
  1. Andrej KarpathyAI score47

    Karpathy wants agents to handle full app DevOps from one command

    AIAndrej Karpathy argues that the hardest part of building a deployed app is not the code but the DevOps work of assembling services, API keys, payments, auth, and deployment. He says the goal is for agents to handle this entire lifecycle as code, with agent-native CLI and API access instead of manual web clicking. He calls it a from-scratch redesign that is only now barely technically possible.

  2. Hamel HusainAI score38

    Data Scientists Face New Pressures as LLM APIs Let Teams Ship AI Without Them

    AIHamel Husain argues data scientists remain essential as foundation-model APIs let teams ship AI without them, because much of the work lies in evaluation, debugging, and metric design. He says teams often rely on generic off-the-shelf metrics and unverified LLM judges instead of examining their own data. He lists five eval pitfalls, starting with generic metrics, and recommends looking at traces and doing error analysis.

Mar 24

Mar 24Tue
  1. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.

Only the first 50 pages are available. Search or browse topics for older items.