Read the original: vLLM Blog· NVIDIA Computer Vision Team (NVCV)· Published · added 38/100AI score38/100
vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning
Original titleScaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM
AISummary
vLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs.
In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas.
The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.
Source: vLLM Blog · vllm.aiPublished · added here