Skip to content
Read the original: PyTorch Blog· Published 38/100AI score38/100

TLX-Optimized Jagged Flash Attention Beats FA4 on Blackwell B200 for Meta GEM

Original titleOptimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell

AISummary

Meta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.

Read the original pytorch.org

Source: PyTorch Blog · pytorch.orgPublished · added here