Skip to content
Read the original: vLLM Blog· Published 38/100AI score38/100

vLLM Adds NVIDIA Hardware Video Decoding to Scale Multi-GPU Video Captioning

Original titleScaling Multi-GPU Video Captioning with PyNvVideoCodec and vLLM

AISummary

vLLM now supports NVIDIA hardware video decoding through PyNvVideoCodec, moving video decoding off the CPU so multi-GPU video captioning can scale to 8 GPUs.

In benchmarks on 8xH100 GPUs, GPU-based decoding provides more than double the throughput of the CPU-based decoder for Qwen/Qwen3-VL-8B-Instruct with 8 single-GPU vLLM replicas.

The functionality is included in standard CUDA vLLM releases, and PyNvVideoCodec==2.0.4 is required for custom installations.

Read the original vllm.ai

Source: vLLM Blog · vllm.aiPublished · added here