Skip to contentSkip to stories

Updated

#Deployment/Engineering

Jul 23

Jul 23Thu

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

Jul 21

Jul 21Tue
  1. JetBrains AI BlogAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  2. koray kavukcuogluAI score72

    Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

    AIGoogle introduces Gemini 3.6 Flash as its workhorse model, with better coding, knowledge work, and multimodal performance while reducing token usage. It also launches Gemini 3.5 Flash-Lite, described as the fastest and most cost-effective 3.5-class model for high-throughput applications, and 3.5 Flash Cyber, a version of 3.5 Flash fine-tuned to find and fix cybersecurity vulnerabilities.

    Why it matters: The post lists three distinct models, each aimed at a different job, so readers can map which one fits coding, high-volume, or security workloads.

  3. JetBrains AI BlogAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

  4. Air Street PressAI score67

    DeepMind's Raia Hadsell argues AI should move beyond language to world models and robotics

    AIAt RAAIS, DeepMind VP of Research Raia Hadsell argued that the field focuses too much on language and should apply large-model training to worlds, robots, biology, and weather. The article cites DeepMind's DiffusionGemma, a 26-billion-parameter open text model that generates blocks by denoising rather than one token at a time, and the Genie-3 world model, which runs in real time for several minutes. It also describes world models as a source of synthetic training data for robots.

  5. Meta AI BlogAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    AISYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 20

Jul 20Mon
  1. Bryan CatanzaroAI score28

    Open models enable forensic analysis that commercial guardrails blocked

    AIA security team found commercial frontier model APIs blocked their incident-response log analysis, which required submitting real attack commands and exploit payloads. They ran the forensic analysis on GLM 5.2, an open-weight model, on their own infrastructure, which also kept attacker data and referenced credentials inside their environment.

Jul 18

Jul 18Sat
  1. Ahead of AI (Sebastian Raschka)AI score52

    How Reasoning Effort Settings Are Built Into LLMs Through Training

    AIThe article explains how reasoning models can offer multiple effort modes, separating training-time methods from inference-time controls such as system prompts and chat templates. It compares six open-weight models, including DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling, noting that their reports disclose different levels of detail. It also shows how GPT-5.6's model selection and effort settings act as two separate scaling axes.

Jul 17

Jul 17Fri
  1. OpenAI NewsroomAI score22

    Foreguard, built with ChatGPT and Codex, helps families plan care and benefits early

    AISekhar and Katie Brandt built Foreguard, a free tool built with ChatGPT and Codex, to help families claim public benefits they are entitled to and plan private insurance coverage. The tool shows that modest budgets of $100 a month can create millions of dollars of day-one financial protection. The goal is to help families prepare earlier, before care decisions become urgent.

  2. Andrew NgAI score28

    DeepLearning.AI launches course on fast LLM inference with Cerebras

    AIDeepLearning.AI has launched a short course, built with Cerebras, on building LLM applications that respond quickly using inference-optimized hardware. The course compares how GPUs, TPUs, and Cerebras' Wafer-Scale Engine handle the memory-to-compute bottleneck, which keeps model weights close to compute units to speed token generation. It covers real-time applications such as live translation and voice agents, plus habits for agentic coding.

Jul 16

Jul 16Thu

Jul 15

Jul 15Wed
  1. Sequoia CapitalAI score32

    Bunkerhill Health's Carebricks Lets Health Systems Deploy AI Agents for Patient Care

    AIBunkerhill Health's Carebricks platform lets health systems create and deploy AI agents across clinical and operational use cases using data hospitals already generate. Sequoia Capital backed the company at seed and is continuing to invest. At UTMB Health, Bunkerhill grew from one agent in production to more than twenty, consolidating multiple vendors' point solutions into one platform.

Jul 14

Jul 14Tue
  1. OpenAI NewsroomAI score22

    Vishal used ChatGPT to train for elite para cycling after amputation

    AIVishal, who lost a leg at age six, used ChatGPT to plan his cycling training, nutrition, recovery, and prosthetic design research. After one year of competitive racing, he qualified for elite competition and earned a chance to represent India at the 2026 Asian Para Road Cycling Championships in Saudi Arabia and the Para Cycling Road World Cup in Thailand.

  2. Cognition Blog (Devin, Windsurf)AI score44

    Cognition Marks One Year Since Windsurf Merger With Devin and SWE Model Gains

    AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.

Jul 13

Jul 13Mon
  1. AI Snake OilAI score57

    Narayanan argues AI job change will unfold over decades, not with one model release

    AIArvind Narayanan's ICML keynote argues that AI's labor impact will depend on slow organizational adaptation rather than a single lab milestone. He cites reliability measurements showing agent accuracy rose much faster than reliability over the last 24 months, and points to software engineering and past technologies like electricity and ATMs. He concludes that evaluation work and human judgment will become more central as building tasks are increasingly automated.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

  3. Cognition Blog (Devin, Windsurf)AI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 11

Jul 11Sat
  1. OpenAI NewsroomAI score15

    Emma Dahl used ChatGPT to help design and build her custom wedding dress

    AIEmma Dahl wanted a historically inspired wedding dress that incorporated pearls from her grandmother's necklace, so she used ChatGPT over months to troubleshoot niche sewing and corset construction. The chatbot also helped her choose a sewing machine upgrade, the shape of her veil, and how to pack and travel with the orchids for her bouquet.

Jul 10

Jul 10Fri

Jul 9

Jul 9Thu
  1. AI Snake OilAI score62

    AI labs may escape the commodity trap by moving up the stack

    AIThe essay argues that AI labs selling model inference face commodity pricing pressure, but may achieve durable profits by moving into products, enterprise deployments, and switching-cost moats. It cites historical infrastructure industries and the Bertrand paradox to support the view that value capture depends on climbing the stack. The authors also warn that successful lock-in could raise enterprise costs and concentrate power, making early interoperability and portability standards important.

  2. Meta AI BlogAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Michael TruellAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.