Skip to content
Read the original: vLLM· vllm_project·Published · YesterdayAI score34/100

3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect.

Original title3/ Kernels: we integrated @deepseek_ai's MegaAttention (NVFP4 KV, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect.

AISummary

Plus vLLM fusions: a CuTe-DSL fused WO-A (up to ~6–7% lower ITL), mHC coefficients on a side stream, and sparse MQA logits (14–23× faster/layer at 512K).

Read the original x.com

Source: vLLM · x.com