Read the original: PyTorch Blog· Han Xu, Jacky Zhou, Jackie (Jiaqi) Xu, Hongtao Yu, Peng Chen (Dev Infra), Darren Liu, Dev (Devashish) Shankar, Max Leung, Nick Riasanovsky, Hao Yan, Manman Ren, Yuanwei (Kevin) Fang· Published · added 38/100AI score38/100
TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM
Original titleOptimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell
AISummary
Meta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.
Source: PyTorch Blog · pytorch.orgPublished · added here