Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google Cloud API Gateway can now expose REST APIs as MCP tools in preview

    AIGoogle Cloud API Gateway now acts as a remote MCP server in Public Preview, making REST operations in an annotated OpenAPI 3.0.x or 3.1.x spec available as agent-ready MCP tools. Existing JWT or API-key authentication, quotas, and logging apply to MCP calls, so teams do not need a separate MCP server. Current limits include no support for OpenAPI 2.0, a maximum of 1,000 tools per gateway, and no MCP and model routing in the same API config.

    Why it matters: The post shows how an existing OpenAPI spec becomes agent-callable MCP tools, with the same auth and quota policies applied, which helps teams avoid building a separate MCP server.

  2. Amp NewsAI score34

    Amp's macOS app now runs threads on your Mac without a terminal

    AIThe Amp macOS app now starts a runner automatically, so threads can run on your Mac without keeping amp --no-tui open in a terminal. Users add folders or projects under Runner in App Settings, then select "This Mac" when starting threads from ampcode.com, a phone, or Puck. A Keep This Mac Awake option prevents sleep while the runner is on and the Mac is plugged in, though the screen still turns off and locks.

  3. eric zakariassonAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  4. GitHub Blog · AI & MLAI score46

    Copilot app rebuilds pull request view to render a 2,200-file diff smoothly

    AIGitHub rebuilt the pull request view in the GitHub Copilot app to keep review fast on very large diffs, testing it on an open source pull request with 2,200 files, over a million changed lines, and more than 400 inline review comments. The core difficulty is that review comment heights can only be measured at render time, which breaks the fixed-geometry virtualization used for code-only diffs. GitHub split the document height into a deterministic code domain and a separately measured domain for comment blocks.

  5. InferactAI score49

    Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

    AIInferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline. The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding. Inferact says it is open-sourcing the kernel today.

    Video from @inferact's post
  6. Greg BrockmanAI score67

    ChatGPT Voice gains plugins and ChatGPT Work access on web and mobile

    AIChatGPT Voice can now use plugins such as email, calendar, and Slack, and it can be powered by GPT-6 Astra, Sol, and Luna. It is also available in ChatGPT Work on web and mobile for creating docs, decks, sites, and spreadsheets by voice, rolling out globally in the latest app version.

    Why it matters: The quoted OpenAI post names the new voice tool access, supported models, and Work integration, showing how voice now acts across workflows.

  7. Google · AI blogAI score36

    Google Beam expands to six countries, adds Industrious flexible-workspace network

    AIGoogle is shipping Google Beam units to customers in the U.S., Canada, U.K., France, Germany, and Japan, supported by 18 channel partners, with HP Dimension with Google Beam as the flagship hardware. Starting in October, users can book Beam at select Industrious locations in Atlanta, Chicago, New York City, and Palo Alto, and an internal eight-week Google study reported 50% more connection and 21% fewer follow-up meetings.

  8. Black Forest LabsAI score67

    Black Forest Labs releases FLUX 3 Action, an open 7B world action model for robots

    AIBlack Forest Labs says FLUX 3 Action is an open-weights 7B world action model that ranks first on the RoboLab benchmark. The company says it outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. The model predicts video and actions together, and the company is releasing the weights, code, fine-tuning recipe, benchmarks, and examples. It also integrated the model into Hugging Face's LeRobot with NVIDIA, with edge deployment on NVIDIA Jetson.

    Why it matters: The release pairs benchmark results with the trade-off it claims to remove between world action model performance and VLA speed, which is useful context for robotics teams weighing open models.

    Video from @bfl_ai's post
  9. eric zakariassonAI score36

    Optimizing reading for AI agents cuts context-gathering costs

    AIEric Zakariasson argues that agents spend heavily on reading context before and after work, so optimizing that reading makes a major difference. He recommends the linked guide to builders, or handing it to an agent to implement its findings. Cursor's related post reports 7% lower token costs with no drop in agent quality, achieved through tighter prompts, selective tool loading, better caching, and compressed file reads.

    Image from @ericzakariasson's post
  10. Microsoft ResearchAI score60

    Microsoft Research shows offloading robot AI inference improves performance and battery life

    AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.

    Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.

  11. Google DeepMindAI score62

    Google DeepMind details server-side memory for Private AI Compute

    AIGoogle DeepMind describes a persistent memory layer for its Private AI Compute platform that stores user context encrypted in the cloud. The encryption keys are held on the user's devices, and data is decrypted only inside hardware-isolated secure enclaves before being re-encrypted. The company says it is publishing a tamper-proof public record of its server software and an independent audit.

    Why it matters: The post explains how persistent cloud memory can keep personal AI context encrypted under keys held on the user's device, a concrete privacy design.

  12. Google · Gemini appAI score44

    Gemini adds Airtable, Adobe, Peloton and more Connected Apps in new rollout

    AIGoogle is rolling out a new wave of Connected Apps to Gemini, letting users link tools such as Airtable, Linear, monday.com, Adobe, Webflow, Peloton and SeatGeek directly in the chat. The apps span productivity, creativity and lifestyle categories, and can be connected in Gemini settings or invoked by typing @ mention in a chat.

  13. Azure BlogAI score40

    Azure resilience now requires continuous validation, not just architecture diagrams

    AIMicrosoft's Azure Blog argues that resilience drifts as workloads change, so architecture diagrams cannot prove a system is resilient. It says roughly 70 percent of cloud outages are related to change, and that teams need health modeling and resiliency goals measured against live signals. The article is the first in a series on validating resilience at scale.

  14. Google LabsAI score29

    Google Labs Releases Six Flow Tools Built by Creatives in Sound, Design, and Content

    AIGoogle Labs released six new Google Flow Tools built by creatives across architecture, sound design, and digital content, including Mondo Sónico, CaptionCast, ThumbnailForge, Surface, CollageMotion Pro, and SwissFlow Studio. Each tool targets a specific workflow, such as generating synchronized audio stems, transcribing and styling captions, or producing animated collages from text prompts. Users can try the tools, duplicate and remix them, or build their own by describing a task in Google Flow.

  15. Comfy BlogAI score62

    Comfy Router launches one API for frontier image, video, 3D, and audio models

    AIComfy Router is now live on the Comfy Developer Platform, giving developers one API to call frontier image, video, 3D, and audio models. Day one models include Seedance 2.5, MiniMax H3, Nano Banana Pro, GPT Image 2, Kling, and Black Forest Labs, and the provider for each job is selectable. Requests fail rather than silently switching providers, and inputs and outputs are deleted after 24 hours.

    Why it matters: The post shows how one API key and a provider parameter let developers swap routes for media models without rewriting calls, with failed requests reporting the provider.

  16. Google for DevelopersAI score52

    Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS text-to-speech models

    AIGoogle announced two new Gemini 3.8 text-to-speech models, positioned as its most expressive yet. Gemini 3.8 Flash TTS targets creative work, letting developers use natural language to define vocal personas, cues, pacing, and dialects, while Gemini 3.8 Flash-Lite TTS is built for high-volume pipelines such as bulk audiobook production and audio dubbing. Both are available now through the Gemini API in Google AI Studio.