Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 30

Sep 30Wed
  1. Google · Innovation & AIOfficialAI score46

    Google AI Flu Model Ranks First in CDC FluSight Hospitalization Forecasts

    AIA flu forecasting model built with Google AI ranked first among 39 eligible models in the CDC's FluSight 2025-26 season evaluation for predicting U.S. flu-related hospital admissions. The model was developed using Empirical Research Assistance (ERA), an AI tool that generates optimization algorithms, and ERA's underlying technology is now available to trusted testers.

  2. SGLangOfficialAI score16

    SGLang engineers to present and take questions at Modal Runtime

    AISGLang will appear at Modal Runtime in San Francisco tomorrow, with a talk by Banghua Zhu on building frontier AI infrastructure with SGLang and Miles at 11:05am. SGLang engineers will also be available from 8:30am to 6:30pm to discuss inference, serving, and RL questions.

  3. Guillermo RauchXAI score30

    Vercel Connect invites services to reach developers and AI agents

    AIGuillermo Rauch invites service providers to add themselves to Vercel Connect to reach over 20 million developers and the agents they build. He argues that connecting services is now the main challenge in building, and that Connect makes it easier and more secure for both agents and apps. Services submit by describing themselves, adding OAuth or API key auth, verifying with a real token, and sending it for review.

  4. Microsoft CopilotOfficialAI score23

    Microsoft's new Copilot combines Home, Code, and Autopilot in one app

    AIMicrosoft's new Copilot brings Home, Code, and Autopilot together in one place for creating custom apps, building decks, automating workflows, and resuming work. Users can start using the Copilot app now and try new features as they become available in Frontier.

    Image from @MSFTCopilot's post
  5. GammaOfficialAI score22

    Gamma launches Salesforce connector for generating presentations from live data

    AIGamma has released a Salesforce connector that lets users link their Salesforce account and describe what they need, with Gamma building presentations from live data. The post cites use cases including weekly pipeline reviews, client-ready QBR decks, rep coaching one-pagers, admin onboarding docs, and marketing performance reports.

    Video from @GammaApp's post
  6. Vercel DevelopersOfficialAI score16

    Vercel Connect now accepts service submissions for review

    AIVercel says developers can submit their service to Vercel Connect by describing it and adding OAuth or API key authentication. The process includes verifying the setup with a real token before sending the submission for review.

  7. MiniMax (official)OfficialAI score44

    HeyGen Video launches on MiniMax H3 at $0.01 per second

    AIHeyGen has released HeyGen Video, a production-quality video product built on MiniMax H3 and post-trained by HeyGen. Pricing starts at $0.01 per second through October, a 50% discount, aimed at businesses that need video without production-level costs.

  8. FireworksOfficialAI score34

    GLM 5.3 Flash now available for training on Fireworks' Serverless API

    AIFireworks AI has made GLM 5.3 Flash available for training through its Serverless Training API, open to all users. The model supports both vision and text inputs. Fireworks says it performs well on its benchmarks for agentic coding, document analysis, and tool use while remaining cost-efficient to serve.

  9. ClineOfficialAI score34

    Cline desktop app can run agents on remote Linux servers over SSH

    AICline says its desktop app can keep running on a laptop while the agent executes on any Linux machine reachable via SSH, configured under Settings → Remote. The setup requires no root access, no npm, and no public port, uploading a self-contained helper and tunneling only the authenticated Cline protocol.

  10. Google Cloud TechOfficialAI score15

    Google Cloud tips for capping GPU and replica settings to control costs

    AIGoogle Cloud recommends limiting accelerator count to a single GPU, setting replica count to 1-1, and avoiding capacity reservations to keep monthly bills predictable. These strict hardware limits apply to auto-scaling configurations for AI workloads.

  11. Google Cloud TechOfficialAI score20

    Google Cloud Model Garden adds scale-to-zero to power down idle GPUs

    AIGoogle Cloud Model Garden now lets users enable scale-to-zero to automatically shut down GPU instances when no incoming requests are active. The feature is presented as a way to avoid paying for idle GPU capacity.

  12. OpenClaw🦞OfficialAI score34

    OpenClaw v2026.9.7 adds OpenAI Agents API and ChatGPT sign-in

    AIOpenClaw released v2026.9.7 with faster performance under load, smoother long chats, and update backup and rollback improvements. The release adds OpenAI Agents API and ChatGPT sign-in in Beta, along with better Apple chat and restart recovery. It includes 2,818 PRs from 344 contributors.

  13. TypeSafe AIOfficialAI score27

    Jev reranking beats GPT-5 Mini on sales data retrieval

    AIJev reranking retrieves Rox sales data 20x faster, 10x cheaper, and 12% more accurate than GPT-5 Mini. The benchmark compared Jev classification against LLM-based reranking for pulling transcripts, emails, CRM notes, news, and documents.

  14. Nathan LambertXAI score47

    “Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coor...

    AI“Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service” It’s the API company’s problem if their model can be manipulated like this. Add KYC

  15. Hacker News · Launch HN, YC launches (10+ points)BlogAI score62

    Magnitude launches an open source inference engine that tunes kernels to local hardware

    AIMagnitude is an open source inference engine for agents that compiles and tunes its kernels on the user's device before running a model. The source claims up to 2x faster decoding than llama.cpp, citing 92% faster decode on Metal and 19% on CUDA, and says one click connects agents such as Pi, OpenCode, Codex, and Claude Code. It supports macOS, Windows, and Linux, and the source states that prompts and models stay on the user's machine.

  16. NVIDIA AIOfficialAI score40

    NVIDIA Shows Visual AI Agent Built in Under 30 Minutes

    AINVIDIA says a single prompt can build and deploy a visual AI agent for a manufacturing line in under 30 minutes, with alerts, video search, and incident reports. The method uses the new Build Vision AI skill in NVIDIA VSS Blueprint 3.3, and a tutorial is available for readers who want to build one.

    Video from @NVIDIAAI's post
  17. Comfy BlogOfficialAI score60

    Comfy API launches to deploy ComfyUI workflows as autoscaling endpoints

    AIComfy API is now available to all users on a paid Comfy plan, letting them package a ComfyUI workflow with its custom nodes, LoRAs, models, and Python dependencies and deploy it as an autoscaling API endpoint. Builds capture the ComfyUI version and dependencies, and each immutable release gets its own URL, so the tested environment is the deployed one. Usage is billed separately, with GPU time charged by the second and storage prorated hourly.

    Why it matters: The post explains how a ComfyUI workflow is packaged into immutable releases and deployed as an autoscaling endpoint, showing a path from local graph to production service.

  18. Stanford HAIOfficialAI score4

    Stanford HAI hosts fall series on AI in science research

    AIStanford HAI is running a fall lecture series, held Wednesdays at 4:30 PM PT, where different researchers present how AI and data-driven methods are used across the sciences. Topics mentioned include decoding the genome and modeling the brain. Many speakers use Stanford's Marlowe GPU cluster, operated by HAI with UIT and VPDoR, and the talks are open to the public in person.

  19. Google WorkspaceOfficialAI score22

    Bci reaches 90% AI adoption, saving over 7,000 hours with Gemini

    AIJuan Burgueño says Bci has reached over 90% AI adoption, saving more than 7,000 hours by using Google Workspace and Gemini. The bank has scaled to over 2,000 custom Gems to accelerate innovation and build an AI-first bank.

    Video from @GoogleWorkspace's post
  20. FireworksOfficialAI score61

    Fireworks adds GLOBAL multi-region deployments under one endpoint

    AIFireworks has added a GLOBAL option that lets one deployment run across regions behind a single endpoint and identity. The scheduler draws compatible capacity from the broadest allowed pool while respecting hardware, quota, reliability, and data-residency constraints. In a seven-day observational study, multi-region deployments showed a 99.992% request success rate versus 99.269% for single-region deployments, though the authors say this is not causal.

    Why it matters: The article explains how one deployment draws capacity across regions, showing how scheduling and routing changed to remove a single-region capacity ceiling.

  21. Ant LingOfficialAI score28

    Ant Group's Tiger Agent runs Ling-3.1-flash for desktop task automation

    AIAnt Group's internal Tiger Agent uses Ling-3.1-flash to plan tasks, while its desktop agent provides browser, file, terminal, and live preview capabilities. A GitHub Trending demo shows the system turning research and semantic grouping into a concise brief.

    Video from @AntLingAGI's post
  22. GammaOfficialAI score34

    Gamma adds Ideogram 4.5 for precise image editing

    AIGamma has added Ideogram 4.5, an image edit model that Ideogram says avoids the artifact buildup, pixel shifts, and color changes that accumulate across repeated edits in leading models. Ideogram says this makes multi-turn editing possible, and the model is available in Ideogram, its API, and launch partners, with open weights promised later.

  23. NVIDIA AIOfficialAI score27

    NVIDIA NeMo Relay Traces Hermes Agent Runs in Arize Phoenix

    AINVIDIA and Nous Research published a hands-on walkthrough of NVIDIA NeMo Relay for collecting traces from Hermes Agent. The guide runs two example scenarios and shows the agent's calls and retries in Arize Phoenix. It also covers how Nous used traces and task results to evaluate fixes across repeated runs.

    Video from @NVIDIAAI's post
  24. DeepSeek HarnessXAI score62

    DeepSeek Harness v0.2 preview launches as a desktop app for macOS and Windows

    AIDeepSeek releases the DeepSeek Harness v0.2 preview with a desktop app for macOS and Windows. The release adds a plugin manager for installing, disabling, and uninstalling plugins without terminal commands, plus an experimental creator mode that generates plugins from user descriptions. The company says DeepSeek Harness is now the most widely used coding agent among users of the official DeepSeek API by DAU and daily sessions.