Skip to content
Read the original: Tri Dao· Published 62/100AI score62/100

FlashAttention-4 paper: attention on Blackwell GPUs nears matmul speed

Original titleThe FA4 paper is finally out after a year of work. On Blackwell GPUs, attention now goes about as fast as matmul even though the bottlene...

AISummary

The FlashAttention-4 paper is out, reporting that attention on Blackwell GPUs now runs at roughly matmul speed, reaching about 1600 TFLOPs.

The forward pass is bottlenecked by exponential computation and the backward pass by shared memory bandwidth, and the redesign uses polynomial exponential emulation, a new online softmax that avoids 90% of softmax rescaling, and 2CTA MMA instructions that let two thread blocks share operands to cut shared memory traffic.

Read the original x.com

Source: Tri Dao · x.comPublished · added here