Description
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.
Published: 2026-09-16
Score: 6.5 Medium
EPSS: < 1% Very Low
KEV: No
Impact: GPU Memory Exhaustion Leading to Denial of Service
Action: Apply Patch
AI Analysis

Impact

The vulnerability allows a malicious request to specify a PyNvVideoCodec decoder backend for video input, bypassing the engine’s static GPU memory reservation. This causes the creation of CUDA contexts, decoder surfaces, and decoded frame allocations that are not accounted for in the allocated KV-cache budget, enabling an attacker to deplete the shared GPU memory. As a result, subsequent requests may fail or workers may crash, effectively causing service disruption. The weakness is a form of resource exhaustion mediated through an improper configuration handling deficiency.

Affected Systems

The flaw affects older releases of the vllm-project's vLLM inference engine before version 0.28.0. Deployments that expose chat completion or response endpoints and have a GPU capable of video decoding with PyNvVideoCodec installed are vulnerable. The issue is tied to the vllm-project:vllm product; no other vendors or versions are listed.

Risk and Exploitability

The CVSS score of 6.5 classifies it as medium severity. The EPSS score is less than 1%, indicating low current exploit probability. It is not listed in the CISA Known Exploit Vulnerabilities catalog. The attack requires the ability to submit a crafted request containing a video payload to a server running a vulnerable vLLM instance with GPU video decoding support and with PyNvVideoCodec enabled. Once executed, the attacker can exhaust GPU memory, producing denial of service by crashing worker processes or causing request failures.

Generated by OpenCVE AI on September 18, 2026 at 01:07 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Upgrade vllm to version 0.28.0 or later to apply the fixed memory reservation logic.
  • Disable or remove the PyNvVideoCodec decoder backend from the VideoLoader registry if GPU video support is not required.
  • Monitor GPU memory usage and set alerts for unusually high allocation patterns to detect potential exploitation attempts.

Generated by OpenCVE AI on September 18, 2026 at 01:07 UTC.

Tracking

Sign in to view the affected projects.

Advisories
Source ID Title
Github GHSA Github GHSA GHSA-8pw2-6jv3-mj5j vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
History

Wed, 07 Oct 2026 14:30:00 +0000

Type Values Removed Values Added
First Time appeared Vllm
Vllm vllm
CPEs cpe:2.3:a:vllm:vllm:*:*:*:*:*:*:*:*
Vendors & Products Vllm
Vllm vllm

Fri, 18 Sep 2026 04:15:00 +0000

Type Values Removed Values Added
First Time appeared Vllm-project
Vllm-project vllm
Vendors & Products Vllm-project
Vllm-project vllm

Thu, 17 Sep 2026 12:15:00 +0000

Type Values Removed Values Added
References
Metrics threat_severity

None

threat_severity

Moderate


Wed, 16 Sep 2026 19:30:00 +0000

Type Values Removed Values Added
Metrics ssvc

{'options': {'Automatable': 'no', 'Exploitation': 'poc', 'Technical Impact': 'partial'}, 'version': '2.0.3'}


Wed, 16 Sep 2026 18:00:00 +0000

Type Values Removed Values Added
Description vLLM is an inference and serving engine for large language models. Prior to 0.28.0, request bodies for Chat Completions and Responses can set media_io_kwargs.video.video_backend to pynvvideocodec, and MediaConnector.fetch_video forwards that choice to VideoMediaIO even when startup configuration selected a software decoder. The engine's _reserve_mm_ipc_gpu_memory logic budgets decoder memory only from static configuration, so the request-selected VIDEO_LOADER_REGISTRY backend can create a CUDA context, decoder surfaces, and decoded-frame allocations that were not removed from the engine's KV-cache budget. An attacker able to submit video requests to a video-capable GPU deployment with PyNvVideoCodec installed can exhaust shared GPU memory, causing request failures, worker crashes, or denial of service. The first release containing the fix is version 0.28.0.
Title vLLM: Request-selected PyNvVideoCodec GPU decode bypasses static VRAM reservation
Weaknesses CWE-400
CWE-770
References
Metrics cvssV3_1

{'score': 6.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H'}


cve-icon MITRE

Status: PUBLISHED

Assigner: GitHub_M

Published:

Updated: 2026-09-16T18:37:14.118Z

Reserved: 2026-08-03T15:20:30.218Z

Link: CVE-2026-69147

cve-icon Vulnrichment

Updated: 2026-09-16T18:36:52.431Z

cve-icon NVD

Status : Analyzed

Published: 2026-09-16T18:17:11.770

Modified: 2026-10-07T14:16:57.827

Link: CVE-2026-69147

cve-icon Redhat

Severity : Moderate

Publid Date: 2026-09-16T17:49:20Z

Links: CVE-2026-69147 - Bugzilla

cve-icon OpenCVE Enrichment

Updated: 2026-09-18T04:00:03Z

Weaknesses
  • CWE-400

    Uncontrolled Resource Consumption

  • CWE-770

    Allocation of Resources Without Limits or Throttling