Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Jul 22

Jul 22Wed
  1. Cognition Blog (Devin, Windsurf)OfficialAI score41

    Cognition signs MOU with U.S. Department of Energy to join Genesis Mission

    AICognition has signed a memorandum of understanding with the U.S. Department of Energy to join the Genesis Mission, a national AI initiative launched by executive order in November 2025. Cognition will contribute its Devin autonomous AI software engineer in four areas: software and data security, modernizing legacy scientific code, expanding scientific workforce capacity, and cloud modernization. Devin Desktop and CLI are listed as FedRAMP Class D (High) Authorized, and the company has offered in-kind code security scans for national laboratory codebases.

  2. Lisa SuXAI score62

    AMD Helios to power Anthropic's Claude at gigawatt scale

    AILisa Su says AMD Helios will help power Anthropic's Claude at gigawatt scale. The quoted AMD announcement states the partnership expands to up to 2 GW of AMD Instinct MI450 Series GPUs in AMD Helios, with AMD committing up to $5B in strategic equity investment in Anthropic.

Jul 21

Jul 21Tue
  1. JetBrains AI BlogOfficialAI score55

    JetBrains Air adds ACP agents, local models, and Java/Kotlin code intelligence

    AIJetBrains Air now connects to ACP-compatible coding agents, including GitHub Copilot CLI, OpenCode, Pi, and Cline, through the Agent Client Protocol. The release also adds Beta Java and Kotlin navigation and diagnostics powered by the IntelliJ IDEA code engine, local model support through Ollama or LM Studio, and Docker-based agent tasks on Windows.

  2. koray kavukcuogluXAI score72

    Google releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

    AIGoogle introduces Gemini 3.6 Flash as its workhorse model, with better coding, knowledge work, and multimodal performance while reducing token usage. It also launches Gemini 3.5 Flash-Lite, described as the fastest and most cost-effective 3.5-class model for high-throughput applications, and 3.5 Flash Cyber, a version of 3.5 Flash fine-tuned to find and fix cybersecurity vulnerabilities.

    Why it matters: The post lists three distinct models, each aimed at a different job, so readers can map which one fits coding, high-volume, or security workloads.

    Video from @koraykv's post
  3. JetBrains AI BlogOfficialAI score62

    JetBrains Context adds repository indexing to coding agents in early access

    AIJetBrains has launched JetBrains Context in early access, a repository intelligence layer that builds a semantic index so coding agents can retrieve relevant code without repeated searching. In tests on 205 SWE-bench tasks, 175 production-monorepo tasks, and 1,953 code-localization tasks, it reduced agent turns by up to 68%, latency by up to 59%, and execution cost by up to 48%. It works with Claude Code, Codex CLI, and Junie CLI at no additional cost for JetBrains AI subscribers, and it does not store source code on JetBrains Context servers.

    Why it matters: The source gives benchmark figures for turns, latency, and cost, showing how repository indexing might change agent workflows on large codebases.

  4. Air Street PressBlogAI score67

    DeepMind's Raia Hadsell argues AI should move beyond language to world models and robotics

    AIAt RAAIS, DeepMind VP of Research Raia Hadsell argued that the field focuses too much on language and should apply large-model training to worlds, robots, biology, and weather. The article cites DeepMind's DiffusionGemma, a 26-billion-parameter open text model that generates blocks by denoising rather than one token at a time, and the Genie-3 world model, which runs in real time for several minutes. It also describes world models as a source of synthetic training data for robots.

  5. Meta AI BlogOfficialAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    AISYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 20

Jul 20Mon
  1. Lisa SuXAI score40

    AMD and Azure expand partnership to deploy Helios AI systems

    AIAMD and Microsoft Azure are expanding their partnership to deploy AMD Helios with MI455X GPUs, Venice CPUs, and Pensando networking. According to AMD's linked background, Microsoft will use Helios to power frontier model AI inference for itself, its AI customers, and Azure AI services.

Jul 18

Jul 18Sat
  1. Ahead of AI (Sebastian Raschka)BlogAI score52

    How Reasoning Effort Settings Are Built Into LLMs Through Training

    AIThe article explains how reasoning models can offer multiple effort modes, separating training-time methods from inference-time controls such as system prompts and chat templates. It compares six open-weight models, including DeepSeek V4, Nemotron 3 Ultra, Kimi K2.5, GLM-5, Qwen3, and Inkling, noting that their reports disclose different levels of detail. It also shows how GPT-5.6's model selection and effort settings act as two separate scaling axes.

Jul 17

Jul 17Fri
  1. OpenAI NewsroomOfficialAI score22

    Foreguard, built with ChatGPT and Codex, helps families plan care and benefits early

    AISekhar and Katie Brandt built Foreguard, a free tool built with ChatGPT and Codex, to help families claim public benefits they are entitled to and plan private insurance coverage. The tool shows that modest budgets of $100 a month can create millions of dollars of day-one financial protection. The goal is to help families prepare earlier, before care decisions become urgent.

    Image from @OpenAINewsroom's post
  2. Andrew NgXAI score28

    DeepLearning.AI launches course on fast LLM inference with Cerebras

    AIDeepLearning.AI has launched a short course, built with Cerebras, on building LLM applications that respond quickly using inference-optimized hardware. The course compares how GPUs, TPUs, and Cerebras' Wafer-Scale Engine handle the memory-to-compute bottleneck, which keeps model weights close to compute units to speed token generation. It covers real-time applications such as live translation and voice agents, plus habits for agentic coding.

    Video from @AndrewYNg's post

Jul 16

Jul 16Thu

Jul 15

Jul 15Wed
  1. Sequoia CapitalBlogAI score32

    Bunkerhill Health's Carebricks Lets Health Systems Deploy AI Agents for Patient Care

    AIBunkerhill Health's Carebricks platform lets health systems create and deploy AI agents across clinical and operational use cases using data hospitals already generate. Sequoia Capital backed the company at seed and is continuing to invest. At UTMB Health, Bunkerhill grew from one agent in production to more than twenty, consolidating multiple vendors' point solutions into one platform.

Jul 14

Jul 14Tue
  1. OpenAI NewsroomOfficialAI score22

    Vishal used ChatGPT to train for elite para cycling after amputation

    AIVishal, who lost a leg at age six, used ChatGPT to plan his cycling training, nutrition, recovery, and prosthetic design research. After one year of competitive racing, he qualified for elite competition and earned a chance to represent India at the 2026 Asian Para Road Cycling Championships in Saudi Arabia and the Para Cycling Road World Cup in Thailand.

    Image from @OpenAINewsroom's post
  2. Cognition Blog (Devin, Windsurf)OfficialAI score44

    Cognition Marks One Year Since Windsurf Merger With Devin and SWE Model Gains

    AICognition says its one-year-old merger with Windsurf has produced a more capable Devin, which now manages other Devins at a mid-to-senior engineering level, and new SWE-1.7 model, described as its most capable and efficient to date. The company reports growing from 44 to 350 people and revenue run rate from $73M to $500M+ since merging the brands.

Jul 13

Jul 13Mon
  1. AI Snake OilBlogAI score57

    Narayanan argues AI job change will unfold over decades, not with one model release

    AIArvind Narayanan's ICML keynote argues that AI's labor impact will depend on slow organizational adaptation rather than a single lab milestone. He cites reliability measurements showing agent accuracy rose much faster than reliability over the last 24 months, and points to software engineering and past technologies like electricity and ATMs. He concludes that evaluation work and human judgment will become more central as building tasks are increasingly automated.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Fable 5 with a sidekick costs less than Opus 4.8 on FrontierCode

    AICognition found that Fable 5 led runs cost less than Opus 4.8 led runs on FrontierCode 1.1 when both used the same sidekick, $1.86 versus $2.04 per run. Fable 5 scored 60.7 against 54.6 for Opus 4.8 in those configurations, and it took fewer lead turns, delegated earlier, and rarely edited code itself. The post attributes the difference to delegation style rather than per-token price, and notes that the approach gives little benefit on short or serial debugging tasks.

    Why it matters: The source compares lead-model delegation habits on a coding benchmark, showing how a pricier model can lower total agent cost through fewer turns and better handoffs.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score39

    Cognition's Devin Reaches FedRAMP High In-Process for Federal Engineering Teams

    AICognition's entire platform, including Devin Cloud, is now FedRAMP Class D (High) In-Process and listed on the FedRAMP Marketplace, extending FedRAMP High authorization beyond Devin Desktop (formerly Windsurf). Devin Desktop and CLI are already FedRAMP High Authorized for workloads with ITAR and DoW IL4, IL5, and IL6 requirements. The company says Devin Security Swarm can find and validate vulnerabilities and open remediation pull requests, and that fleets of Devins can upgrade legacy software 5-40x faster than humans alone.

Jul 10

Jul 10Fri
  1. OpenAI NewsroomOfficialAI score22

    Restaurant owner's ChatGPT playbook drives 200% sales growth

    AIPatrick Cheng used ChatGPT to redesign menus, optimize pricing, and improve marketing after acquiring a struggling restaurant in 2021. He reports more than 200% sales growth and a tripling of catering business, and is now launching NextGen, a nonprofit helping family-owned businesses adopt AI through education and micro-grants.

    Image from @OpenAINewsroom's post

Jul 9

Jul 9Thu
  1. AI Snake OilBlogAI score62

    AI labs may escape the commodity trap by moving up the stack

    AIThe essay argues that AI labs selling model inference face commodity pricing pressure, but may achieve durable profits by moving into products, enterprise deployments, and switching-cost moats. It cites historical infrastructure industries and the Bertrand paradox to support the view that value capture depends on climbing the stack. The authors also warn that successful lock-in could raise enterprise costs and concentrate power, making early interoperability and portability standards important.

  2. Meta AI BlogOfficialAI score72

    Meta releases Muse Spark 1.1 with agent and coding gains

    AIMeta Superintelligence Labs has introduced Muse Spark 1.1, a multimodal reasoning model aimed at agentic tasks, with gains in tool use, computer use, coding, and multimodal understanding. It supports a 1 million token context window and is available in Thinking mode in the Meta AI app and on meta.ai, with developers able to access it through a public preview of the Meta Model API.

    Why it matters: The post specifies Muse Spark 1.1's agent, coding, and multimodal gains and its Meta Model API preview access, which helps developers judge its fit for their workflows.

Jul 8

Jul 8Wed
  1. Michael TruellXAI score57

    Cursor and SpaceXAI release Grok 4.5, a coding-focused model

    AICursor co-founder Michael Truell announced Grok 4.5, a model trained with SpaceXAI that the post calls Opus-class, fast, and low cost. He says it is a significant step up over Composer 2.5 and has become the daily driver for many on the Cursor team. A benchmark table shows Grok 4.5 at 83.3% on Terminal-Bench 2.1 and 78.0% on SWE-Bench Multilingual, with the post saying more releases will follow.

  2. Cognition Blog (Devin, Windsurf)OfficialAI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

  3. Cognition Blog (Devin, Windsurf)OfficialAI score47

    Cognition Tests Trustworthiness of SWE-1.7, Built on Kimi K2.7 Code

    AICognition says its SWE-1.7 model, developed from the open-source Kimi K2.7 Code base, performs as well as or better than leading U.S. frontier models on its new trustworthiness evaluation suite. The suite combines 145 politically sensitive questions, sampled in English and Chinese, with realistic coding scenarios to measure propaganda, censorship, and security behavior. Cognition says SWE-1.7 improves substantially over the base Kimi K2.7 Code model, though the company says the benchmarks are still in development.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchOfficialAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

  2. Lilian WengXAI score34

    Lilian Weng on harness engineering's role in AI self-improvement

    AILilian Weng published a new post on harness engineering for AI self-improvement. She expects harnesses to evolve toward self-improvement and enable auto-research, while smarter models keep harnesses simple. Even if many harness gains are later internalized into core models, specifying goals and context will remain necessary.

Jul 3

Jul 3Fri
  1. Lil'Log (Lilian Weng)BlogAI score62

    Lilian Weng surveys harness engineering as a path to recursive self-improvement

    AIThe post argues that the system surrounding a base model, called the harness, increasingly determines how well AI agents deploy and improve. It reviews research where harness components such as workflows, context, and code are optimized automatically through evolutionary search and meta-agent loops. The author concludes that evaluators, memory management, and human oversight remain open bottlenecks.

  2. Arthur MenschXAI score34

    Mistral argues enterprises need open models and their own data for AI growth

    AIMistral CEO Arthur Mensch says enterprises should use open-source models because closed providers that force data retention gain leverage over their business. He argues companies should store data in open systems, control AI access rules, and build continuous training loops to shrink costs and create hard-to-copy systems. Mistral offers its Studio control plane and Forge training platform, deployed on customer infrastructure or through zero-data-retention hosting.

Jul 2

Jul 2Thu
  1. Cognition Blog (Devin, Windsurf)OfficialAI score38

    Cognition launches Devin Security Vulnerability Remediation Program for enterprise backlogs

    AICognition launched the Devin Security Vulnerability Remediation Program, in which its forward-deployed engineers embed with customer teams to deploy Devin to find, validate, and fix vulnerabilities. The program first works through existing scanner backlogs from tools such as Snyk, SonarQube, and Semgrep, shipping validated fixes as pull requests, then adds Devin Security Swarm for continuous discovery of logic flaws. Most engagements run about six weeks, and eligibility is limited to enterprise Devin Cloud customers meeting the program's requirements.