Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 18

Sep 18Fri
  1. LMSYS OrgOfficialAI score52

    LMSYS blog shows DeepSeek-V4-Flash and Kimi-K3 running on consumer hardware via SSD Expert Pack

    AILMSYS Org announced a blog on running DeepSeek-V4-Flash and Kimi-K3 on consumer hardware using SSD Expert Pack, built by WiCi AI and the SGLang team. Routed experts stay on an NVMe SSD, and the runtime loads only router-selected experts into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD, DeepSeek-V4-Flash MXFP4 decoded at 1.85–1.99 tokens/sec and Kimi-K3 community Q2_K (text-only) at about 0.29 tokens/sec.

    Image from @lmsysorg's post
  2. SemiAnalysisBlogAI score52

    Engram offloading to DRAM beats SSD for DeepSeek-V4.1-Flash serving on B200

    AISemiAnalysis tested offloading DeepSeek-V4.1-Flash's Engram embedding table from HBM to host DRAM and to local SSD. On B200 configurations, DRAM delivered more total tokens per dollar and higher P90 interactivity than SSD at every measured point. The report concludes SSD offloading is likely not worth the tradeoff for production serving in its unoptimized setup.

  3. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  4. Liquid AI · new models on Hugging FaceOfficialAI score55

    Liquid AI releases LFM2.5-VL-3B-DSpark drafter for faster vision-language decoding

    AILiquid AI released LFM2.5-VL-3B-DSpark, a speculative-decoding draft model for its LFM2.5-VL-3B vision-language model. The source reports decoding up to 2.66× faster on a single H100 with SGLang, up to 3.13× on Apple M5 Max with MLX-VLM, and up to 2.14× on Apple M3 Ultra with llama.cpp, with output unchanged under greedy decoding.

  5. KrASIA · Big TechNewsAI score47

    Huawei unveils Atlas 960E superpod linking 4,096 NPUs with near-packaged optics

    AIHuawei unveiled the Atlas 960E superpod, which links up to 4,096 NPUs using near-packaged optics (NPO) and claims eight exaflops at FP8 precision and up to one petabyte of high-bandwidth memory. Its Hi-ONE engine provides 7.2 terabits per second of transmission capacity, and Huawei says 5,500 engines replace 48,000 conventional 800G optical modules, cutting power consumption by more than 550 kilowatts. Huawei has proposed its NPO implementation agreement to the Optical Internetworking Forum, though it has not yet become a standard.

  6. InferactOfficialAI score46

    Kimi K3 serving in vLLM is now 2.2–2.8× faster

    AIInferact, with Red Hat AI, NVIDIA, and Huawei, co-led an optimization effort that makes Kimi K3 on vLLM 2.2–2.8× faster. The work spans scheduling, KDA state handling, and custom MoE kernels. The vLLM project's background post cites those throughput gains on a B300 benchmark against v0.27.1 and links a technical deep dive.

  7. Hamel HusainBlogAI score62

    Hamel Husain's FAQ on AI evals: error analysis, judges, and trace review

    AIHamel Husain and Shreya Shankar's FAQ explains AI evals as tests of whether an AI system does what users and the business want. It recommends starting with error analysis on at least 30 traces, then turning recurring failures into binary code-based checks or LLM judges validated against human labels.

  8. Anthropic NewsroomOfficialAI score62

    Anthropic partners with Accenture on embedded AI model evaluation

    AIAnthropic is partnering with Accenture, through its specialist AI business Faculty, on independent evaluation of frontier models, including red-teaming, alignment assessments, and safeguard testing. Anthropic and Accenture each expect to invest at least $1 billion in this capacity over five years. The source says embedded evaluators would have employee-comparable access, but standards for access and reporting, and a settled funding system, do not yet exist.

    Why it matters: The source ties a new evaluation arrangement to an unresolved question of who funds and sets standards for independent AI evaluators, which is useful context for governance debates.

  9. Gemini API ChangelogOfficialAI score26

    Google limits Gemini 2.5 model API access to users who used them recently

    AIGoogle is restricting access to its Gemini 2.5 models to users who have actively used them in the past, to keep performance reliable. The models are not deprecated and remain available through the API until further notice. For new projects, Google recommends its latest models, 3.5 Flash-Lite or 3.8 Flash.

Sep 17

Sep 17Thu
  1. Together AI BlogOfficialAI score31

    Fintech Scales Coding Agent Traffic on Together's Dedicated Model Inference

    AIA global fintech scaled its AI coding agent traffic by running the GLM-5.2 model on Together AI's Dedicated Model Inference, after capacity planning failed to keep pace with unpredictable engineering-hour bursts. The customer gained self-service endpoint provisioning, a metrics API for diagnosing queuing, and live configuration changes that shipped with zero downtime. The setup runs dozens of B200 GPUs at 256K context across multiple replicas.

  2. vLLM BlogOfficialAI score38

    vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

    AIvLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs. In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas. The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

  3. xAI News (Grok)OfficialAI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

  4. Sherwin WuXAI score44

    ChatGPT for Word launches, bringing ChatGPT and Codex into Microsoft Word

    AIOpenAI has released ChatGPT for Word, letting users access ChatGPT and Codex directly inside Microsoft Word. The post says ChatGPT for Excel and PowerPoint has been growing rapidly, and Word completes that set. The quoted ChatGPT post adds that the tool can draft from notes, rewrite paragraphs, proofread, suggest edits, and flag formatting issues.

  5. Sherwin WuXAI score34

    OpenAI Launches 47 Community-Built Legal Plugins in ChatGPT

    AIOpenAI launched 47 community-built plugins for legal work in ChatGPT, created by legal experts from LegalQuants, Skills.law, and LECG rather than white-labeled by OpenAI. The plugins are already live in the ChatGPT plugin store, alongside 26 partner-built plugins from companies including Thomson Reuters, Harvey, Legora, and iManage.

  6. AI at MetaOfficialAI score44

    Meta's Muse agent now available on Mac for local tasks

    AIMeta is rolling out Muse for Mac today, a personal agent that can complete tasks directly on the user's computer with explicit permission. Examples include organizing the downloads folder, finding lost files, and summarizing messages and notes, with more capabilities coming soon.

  7. Greg BrockmanXAI score60

    OpenAI launches Astra for Law, a GPT-6 Astra offering for law firms

    AIOpenAI has introduced Astra for Law, a new offering powered by GPT-6 Astra with tools, settings, and context for lawyers and legal technology firms. The offering includes data privacy as a core feature, plus 26 partner-built plugins and 47 community plugins for legal work in ChatGPT.

    Why it matters: The offering targets legal work specifically, with partner and community plugins and data privacy as a core feature, which shows how frontier models are packaged for professional sectors.

  8. Google ResearchOfficialAI score52

    Google Research enables teachers to create generative UI learning interactives

    AIGoogle Research is sharing an experiment that lets educators generate custom, guided STEM simulations tailored to their curriculum using generative UI. It is releasing a sample library of over 30 English interactives for physics, chemistry, biology, and math, all AI-generated and reviewed by teachers. Schools using Google Workspace for Education can sign up through the Google for Education Pilot Program to give feedback.

  9. Google AI StudioOfficialAI score58

    Google AI Studio open-sources Speakeasy's OpenAPI SDK generator suite

    AIGoogle AI Studio announced that Speakeasy is open sourcing its OpenAPI client generation suite under AGPLv3, following a May 2026 vendor shutdown that disrupted Google's SDK pipeline. The suite covers SDK generation for 7 languages, an agent-native CLI generator, and a documentation MCP server generator. Google says its pipeline now serves six targets with roughly one engineer maintaining it.

  10. Noah ZwebenXAI score62

    Claude Code adds Projects that run parallel threads from one conversation

    AIAnthropic's Claude Code now runs projects from a single conversation, where Claude directs parallel threads that keep working after the user closes their laptop. The feature is in beta for select Pro and Max users in cloud sessions, with wider availability for all Claude users promised soon.

    Why it matters: The quoted launch replaces scattered sessions with one coordinator that runs parallel threads in the background, a workflow change worth weighing for complex projects.

  11. catXAI score62

    Claude Code adds Projects that coordinate multiple parallel sessions

    AIAnthropic's Claude Code is rolling out Projects on desktop and web, in beta for select users. A project splits work into threads, runs them as parallel cloud sessions, passes context between them, and keeps running after the user leaves. The author says Claude keeps context across tasks and can give an aggregated status update on request.

    Why it matters: The post explains how Projects shifts work from managing single sessions to coordinating many parallel tasks, a change that affects how Claude Code users plan and track their work.

  12. Google AI StudioOfficialAI score80

    Google updates Gemini managed agents with Files and Credentials APIs

    AIGoogle AI Studio released antigravity-preview-09-2026, an updated harness for Gemini managed agents, now live in the Interactions API and AI Studio and running on Gemini 3.8 Flash. The release adds a Files API for moving data into and out of the agent's sandbox and a Credentials API that stores secrets encrypted so the model never sees them.

    Why it matters: The post shows what changed in the agent harness and how the new Files and Credentials APIs keep secrets out of the model's context, useful for developers building agents.

  13. World LabsOfficialAI score32

    World Labs' Atlas generates real-time flythrough from 32 input images

    AIWorld Labs says its Atlas model turns 32 input images into a real-time flight through NVIDIA's Voyager headquarters. Trained on NVIDIA Blackwell GPUs, Atlas uses the images as 3D spatial context to generate new views with pixel-perfect camera control.

    Video from @theworldlabs's post
  14. Hacker News · Launch HN, YC launches (10+ points)BlogAI score58

    Skillsync launches tool to move AI chat sessions across coding agents

    AISkillsync, a Y Combinator W26 company, launched a tool that converts AI chat sessions between coding agents, including messages, reasoning and tool calls. The conversion engine txcript is open source, and the local-first app runs on Mac with a CLI and MCP support. Only sessions shared into team workspaces leave the user's machine.

  15. Daniel HanXAI score44

    Unsloth Desktop adds multi-user accounts and faster GRPO training

    AIUnsloth Desktop now supports multi-user accounts, alongside a revamped Docker image and custom Jupyter Notebook with custom themes, titles, and expandable cells. The update adds RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference.

  16. Unsloth AIOfficialAI score60

    Unsloth Docker image lets users train and run 500+ models locally

    AIUnsloth announced that its Docker image now lets users train and run more than 500 models locally with no setup required. The image works on NVIDIA and AMD hardware and supports a new GUI or notebook workflow. The post links to an installation guide and the GitHub repository.

    Why it matters: The post explains how to run and train 500+ models locally with a no-setup Docker image, supported on NVIDIA and AMD, with a new GUI or notebook workflow.

    Image from @UnslothAI's post
  17. Sierra BlogOfficialAI score38

    Sierra Achieves AIUC-1 Certification for Its AI Agent Platform

    AISierra has become AIUC-1 certified after an independent audit by Schellman and testing by the Artificial Intelligence Underwriting Company (AIUC), a new standard for AI agents that tests resistance to manipulation and unauthorized access. Schellman found that Sierra met all applicable AIUC-1 requirements, and the technical evaluations recur at least quarterly with a full audit each year. The certification complements Sierra's existing SOC 2 Type II, ISO 27001, and ISO 42001 attestations.

  18. Google DeepMindOfficialAI score33

    Researchers use AlphaGenome Atlas to identify disease-causing DNA variants

    AIResearchers at the Broad Institute, the University of Exeter, and other institutions are already using AlphaGenome Atlas to identify potential disease-causing DNA variants and interpret their role. The post is a thread announcement, and it gives no further details on methods or results.

    Video from @GoogleDeepMind's post
  19. OpenBMBOfficialAI score40

    OpenMed and MiniCPM5-2B demo local agentic clinical AI workflow

    AIOpenMed paired with MiniCPM5-2B to demonstrate a local clinical AI workflow combining privacy-preserving data processing with a compact model's tool use and long-context reasoning. OpenMed masks sensitive identifiers and extracts clinical context before MiniCPM5-2B calls tools, compares lab results, and generates clinical handoffs with source references. The post presents this as an example of keeping inference on local, resource-constrained hardware.

    Image from @OpenBMB's post