Impact
The vulnerability allows a malicious request to specify a PyNvVideoCodec decoder backend for video input, bypassing the engine’s static GPU memory reservation. This causes the creation of CUDA contexts, decoder surfaces, and decoded frame allocations that are not accounted for in the allocated KV-cache budget, enabling an attacker to deplete the shared GPU memory. As a result, subsequent requests may fail or workers may crash, effectively causing service disruption. The weakness is a form of resource exhaustion mediated through an improper configuration handling deficiency.
Affected Systems
The flaw affects older releases of the vllm-project's vLLM inference engine before version 0.28.0. Deployments that expose chat completion or response endpoints and have a GPU capable of video decoding with PyNvVideoCodec installed are vulnerable. The issue is tied to the vllm-project:vllm product; no other vendors or versions are listed.
Risk and Exploitability
The CVSS score of 6.5 classifies it as medium severity. The EPSS score is less than 1%, indicating low current exploit probability. It is not listed in the CISA Known Exploit Vulnerabilities catalog. The attack requires the ability to submit a crafted request containing a video payload to a server running a vulnerable vLLM instance with GPU video decoding support and with PyNvVideoCodec enabled. Once executed, the attacker can exhaust GPU memory, producing denial of service by crashing worker processes or causing request failures.
OpenCVE Enrichment
Github GHSA