Skip to content
Read the original: vLLM· Published 28/100AI score28/100

vLLM bounded SWA replay cuts prefill compute 30–40% after prefix hits

Original title2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns onl...

AISummary

vLLM's bounded sliding-window attention (SWA) replay reruns only each request's last 128 tokens after a prefix hit, instead of replaying 40 × 128 tokens. With CUDA graphs, prefill compute drops 30–40% for layers 21–39.

Read the original x.com

Source: vLLM · x.comPublished · added here