vLLM bounded SWA replay cuts prefill compute 30–40% after prefix hits
Original title2/ SWA bounded replay: rebuilding sliding-window KV exactly after a prefix hit means replaying 40 × 128 tokens. Bounded replay reruns onl...
AISummary
vLLM's bounded sliding-window attention (SWA) replay reruns only each request's last 128 tokens after a prefix hit, instead of replaying 40 × 128 tokens. With CUDA graphs, prefill compute drops 30–40% for layers 21–39.
Source: vLLM · x.comPublished · added here