Skip to content
View original post on X: SGLangOfficial· 34/100AI score34/100

SGLang reports inference speedups for MLA, MoE, and KDA kernels

AISummary

SGLang says restructured MLA decode kernels on Rubin, which fit a deeper pipeline in 327 KiB of shared memory, deliver a 16% speedup at batch 16 with 128K context and bit-identical output.

The post also reports 20% faster full FP8 MLA at batch 1 and 20% faster KDA verify kernels after keeping weights in registers and reducing synchronization.

It additionally covers fusing MoE finalization, the shared expert, 8-GPU all-reduce, and RMSNorm into one collective kernel.

Post on XView on X
SGLangVerified on X
@sgl_project

Part of a thread · earlier post

Highlights of how we made inference faster:
MLA: Rubin’s 327 KiB of shared memory per CTA fits a deeper decode pipeline, delivering a 16% speedup at batch 16 / 128K context with bit-identical output. At batch 1 / 128K, restructuring split-KV reduction makes full FP8 MLA 20% faster.

MoE: Finalization, shared expert, 8-GPU all-reduce and RMSNorm fused into one collective kernel.

KDA: Keeping weights in registers and simplifying synchronization cuts barrier stalls, making the verify kernel 20% faster.

The post also covers FP8 conversion, all-reduce tuning, tensor-core GEMM, and eliminating hidden copies.

Source: SGLang · x.comPublished · added here