Meta distills 40-step video diffusion into a 2-step live streaming model
Original titleLive video streaming needs to respond instantly while remaining visually consistent over long conversations.
AISummary
Meta distilled a 40-step diffusion teacher using 3-way CFG, requiring 120 evaluations per video chunk, into an unguided 2-step causal student with a fixed-length KV cache. The student uses self-forcing to resist drift and keep near-teacher quality while needing 60x fewer evaluations, enabling instant responses in live video streaming.
Source: AI at Meta · x.comPublished · added here