3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect.
Original title3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect.
AISummary
Plus vLLM fusions: a CuTe-DSL fused WO-A (up to ~6–7% lower ITL), mHC coefficients on a side stream, and sparse MQA logits (14–23× faster/layer at 512K).
Source: vLLM · x.com