Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri
  1. RadixArkOfficialAI score60

    RadixArk's Miles runs end-to-end RL on NVIDIA Vera Rubin with SGLang

    AIRadixArk says Miles runs reinforcement learning end to end on NVIDIA Vera Rubin, using SGLang rollouts, Megatron training and one container image. Agentic RL runs 64 concurrent sandboxes on the Vera CPU next to the GPUs. The linked SGLang post reports that Kimi K3 inference gained up to 20% faster FP8 MLA at batch 1 with 128K context, and a 5.9% end-to-end speedup from MoE tail fusion.

    Why it matters: The post gives concrete speedup figures and a specific RL setup on early-access Rubin hardware, useful for engineers comparing inference and training stacks.

  2. WorkBuddyOfficialAI score22

    WorkBuddy Co-writing adds HTML editing with AI Edit

    AIWorkBuddy says its Co-writing feature now supports HTML editing, joining Word, Markdown, PPT, and Excel files. Users can open an HTML page, make changes through Co-writing, and adjust content directly with AI Edit.

    Video from @WorkBuddy_AI's post
  3. LMSYS OrgOfficialAI score50

    SGLang brings Kimi K3 inference speedups to NVIDIA Vera Rubin

    AISGLang optimized attention, MoE, and speculative verification kernels for Kimi K3 NVFP4 on early-access NVIDIA Vera Rubin hardware. Reported gains include up to 20% faster FP8 MLA at batch 1 and 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion. Miles, from RadixArk, uses SGLang for rollouts in end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

  4. InferactOfficialAI score52

    vLLM on NVIDIA Vera Rubin NVL72 reaches 7.8x GB200 throughput on MiniMax M3

    AIInferact says vLLM now supports NVIDIA Vera Rubin, with early results showing more than 7.8x the throughput of GB200 on MiniMax M3 at matched interactivity on AgentX. The team says it integrated a Rubin-optimized MSA prefill kernel and locality-aware optimizations, and describes the results as early.

    Image from @inferact's post
  5. InferactOfficialAI score13

    Inferact runs open models in production, upstreaming vLLM tuning

    AIInferact, founded by the creators and core maintainers of vLLM, says its platform runs open models in production on customer compute or its own. The company says changes for hardware such as NVIDIA Vera Rubin are landing in open-source vLLM and that the platform benefits directly from that upstream tuning.

    Image from @inferact's post
  6. vLLMOfficialAI score34

    vLLM reports 7.8x and 3.7x throughput gains on two benchmarks

    AIvLLM says it delivered 7.8x or more the throughput of GB200 on SemiAnalysis AgentX, serving MiniMax M3 on Vera Rubin NVL72 at matched interactivity. In MLPerf Inference v6.1, vLLM with Dynamo reached up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL-235B-A22B.

    Image from @vllm_project's post
  7. vLLMOfficialAI score33

    vLLM says Rubin GPUs run its Blackwell kernels unmodified

    AIvLLM says a Rubin GPU has 3.5x the FP4 FLOPS, about 2.4x the HBM bandwidth, and 1.7x the NVLink bandwidth of a GB200. The faster NVLink lowers the cost of AllReduce and all-to-all operations in MoE serving. vLLM's Blackwell kernels run on Rubin unmodified, and FlashInfer 0.7.0 adds Rubin-tuned attention, GEMM, and MoE kernels.

    Image from @vllm_project's post
  8. vLLMOfficialAI score28

    vLLM uses NVIDIA locality domains to speed MoE decode up to 1.2x

    AIvLLM's locality-aware MoE shards FC1 and FC2 weights column-wise and launches one kernel per CUDA 13.4 locality domain using Green Contexts, so each SM reads only local HBM. Early results show MiniMax M3 MoE layers up to 1.2x faster in decode. This is post 4 of a 5-post thread.

    GIF from @vllm_project's post
  9. vLLMOfficialAI score44

    vLLM now runs DeepSeek, Kimi, GLM, and MiniMax on Vera Rubin

    AIvLLM says users can try its Vera Rubin support today with the vllm/vllm-openai:cu134-nightly image, which runs DeepSeek, Kimi, GLM, and MiniMax models. The project thanked contributors and said it expects further improvements for vLLM on Vera Rubin.

  10. vLLM BlogOfficialAI score62

    vLLM adds support for NVIDIA Vera Rubin NVL72 with 7.8x throughput over GB200

    AIvLLM now supports NVIDIA Vera Rubin NVL72, with daily container builds and support for models from DeepSeek, Moonshot AI, Z.ai, and MiniMax. In early AgentX benchmarks, vLLM running MiniMax M3 delivered up to 7.84x the throughput per GPU of GB200 NVL72 at matched interactivity. The post is an early look, and the team expects further gains from ongoing optimizations.

    Why it matters: The post gives specific hardware specs, kernel configurations, and benchmark figures for running vLLM on Vera Rubin NVL72, useful for teams planning deployments.

  11. SemiAnalysisXAI score62

    Nvidia's off-balance-sheet guarantees rise to $530B in latest 10-Q

    AISemiAnalysis reports that Nvidia's latest 10-Q discloses $530B of gross off-balance-sheet guarantees across six line items, up from $184B in the prior quarter. The main drivers are higher supply commitments, mainly memory purchases, and datacenter backstops for SB Energy's PORTS-Pike campus in Ohio for OpenAI. Two new items also appear: $36B of AI Cloud Agreements for Neocloud backstops and $20B of datacenter leases Nvidia has taken on to assign to Neocloud offtakers.

    Image from @SemiAnalysis_'s post
  12. SGLangOfficialAI score34

    SGLang reports inference speedups for MLA, MoE, and KDA kernels

    AISGLang says restructured MLA decode kernels on Rubin, which fit a deeper pipeline in 327 KiB of shared memory, deliver a 16% speedup at batch 16 with 128K context and bit-identical output. The post also reports 20% faster full FP8 MLA at batch 1 and 20% faster KDA verify kernels after keeping weights in registers and reducing synchronization. It additionally covers fusing MoE finalization, the shared expert, 8-GPU all-reduce, and RMSNorm into one collective kernel.

  13. SGLangOfficialAI score37

    Miles runs full RL loop on Rubin GPUs with SGLang and Megatron

    AIMiles runs the full reinforcement learning loop on Nvidia Rubin, using SGLang for rollout, Megatron for training, and one container image. On a single 4-GPU tray, Qwen3-30B-A3B's GSM8K reward rises from about 45% to about 95% over 50 rollouts, matching the GB300 curve. The post also reports DeepSeek-V4-Flash end-to-end rollout and training, and Qwen3.5-35B-A3B agentic RL with 64 concurrent mini-SWE-agent sandboxes on SWE-bench Verified, where reward holds near 0.6 and median response length falls about 30%.

  14. SGLangOfficialAI score52

    SGLang adds Rubin optimizations that speed up Kimi K3 inference

    AISGLang says it worked with NVIDIA to optimize attention, MoE, and speculative verification kernels for Kimi K3 inference on early-access Rubin hardware. It reports up to 20% faster FP8 MLA at batch 1 with 128K context, 20% faster KDA verification with bitwise-identical output, and a 5.9% end-to-end speedup from MoE tail fusion that removes 276 kernel launches per decode step. The post also says SGLang powers Miles' end-to-end RL training on Rubin, including agentic RL with 64 concurrent sandboxes on the Vera CPU.

  15. Bertholomus AIXAI score22

    DeepSeek TP=2 and TP=4 kernel kits released on GitHub

    AIA GitHub post announces that kernel kits for DeepSeek's TP=2 and TP=4 recipes are now live. The post links to a repository named deepseek-v4.1-tensorfold-tp2-2xgb10 and gives no further details on performance or features.

  16. Soumith ChintalaXAI score22

    Tinker cuts prices up to 70% as efficiency improves

    AITinker, an API for training and fine-tuning models, is cutting prices by up to 70% after engineering efficiency gains. The company says the savings are passed on to customers, and that buying more produces greater savings. GLM-5.3-Flash and DeepSeek-v4.1-Flash are also now available on Tinker for long-context work.

  17. Rohan PaulXAI score46

    Pine launches cloud computer for AI agents, reports 1/20 token cost

    AIPine has launched a cloud computer built for AI agents, which developers create through an SDK and give jobs in plain language. Running GPT-5.6 Luna, Pine reports about 1/20 the model-token cost of GPT-5.6 Sol with Codex on SaaS-Bench v1.1, scoring 78.3%, the highest in the published comparison. Pine also reports 1/26 the token cost of Opus 5 with Claude Code and 2 to 5 times faster speed in selected preliminary internal tests.

    Image from @rohanpaul_ai's post
  18. Hacker News · AI (150+ points)BlogAI score38

    Show HN: big-arrow-on-the-screen lets AI agents draw arrows and text on macOS

    AIbig-arrow-on-the-screen (bigarrow) is a MIT-licensed macOS command-line tool and skill for Claude Code and Codex that draws arrows, boxes and text over any window. Clicks pass through, keyboard focus stays put, and each arrow removes itself after a set duration or when its agent process ends. The tool only points; it never clicks, types or captures the screen, and it requires no macOS permission to draw.

  19. dexXAI score43

    Dex Horthy posts a one-word teaser, "he cook"

    AIDex Horthy (@dexhorthy) posted only the words "he cook" on X, with no further detail. The post is a short reaction and does not describe a product, release, or result on its own. Background from the quoted post by @0xblacklight describes a serverless background agent that created a GitHub pull request from an issue.

  20. SiliconANGLE · AINewsAI score26

    Seismora builds a control plane to route AI workloads across devices and clouds

    AISeismora Inc. is developing a control plane that routes AI workloads across devices, edge infrastructure and cloud providers, according to founder and CEO Vito Palermo. The company calls its approach "cognitive routing," choosing where work runs based on capability, cost, latency and policy constraints. Its intended customers are developers building agentic applications, not end consumers.

  21. RadixArkOfficialAI score22

    RadixArk praises Proximal for training coding agents with Miles

    AIRadixArk says Proximal is using Miles to train coding agents and calls it a flexible, scalable foundation for teams running their own training workloads. Proximal says its training framework is built on Miles, with runs on Modal's on-demand GPU clusters and serverless GPUs for inference. Its sandboxing infrastructure runs on Kubernetes and can handle millions of concurrent rollouts.

  22. dexXAI score38

    HumanLayer releases teleport and orchestrate commands with a minimalist UI

    AIHumanLayer announces a new release with /hl:teleport, which moves a local session to any remote host the user owns or launches without losing context. The release also adds /hl:orchestrate, which lets HumanLayer drive its own tasks, including splitting work, forking workflows, and moving artifacts, and it ships a minimalist UI with rounded corners and less visual noise.

    Image from @dexhorthy's post
  23. Kevin LiuXAI score32

    Jarhead, an open-source voice assistant that operates your Mac for you

    AIKevin Liu releases Jarhead, an open-source voice assistant that uses a Mac the way a person would, opening apps, clicking through Spotify, browsing, typing, running git, and marking up the screen. It runs on GPT-Live-1, which listens and talks at the same time so users can interrupt it mid-task, and its reasoning can use a ChatGPT plan with Codex, Claude Code, or an API key.

    Video from @kevskgs's post
  24. SGLangOfficialAI score12

    SGLang opens agenda for November summit at Fort Mason, San Francisco

    AISGLang opens the agenda for its SGLang Summit, held November 12–13 at Fort Mason in San Francisco, with early-bird tickets until October 20. The two-day program covers open models, AI hardware, recursive self-improvement, agents, and robotics, with speakers listed from Intel, OpenAI, Thinking Machines, and Perplexity. It also includes hands-on workshops on model training and serving, a hardware Expo, and community events.

    Video from @sgl_project's post