Read the original: PyTorch Blog· Sean Chen (Red Hat) and Shangdi Yu (PyTorch, Meta Platforms)·Published· 6d agoAI score47
Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM
Building a High-Performance and Portable vLLM Linear Backend with Helion
AISummary
The vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.
Source: PyTorch Blog · pytorch.org