Description
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
Published: 2026-09-17
Score: 8.7 High
EPSS: < 1% Very Low
KEV: No
Impact: Denial of Service via memory exhaustion
Action: Immediate Patch
AI Analysis

Impact

vLLM through 0.29.0 contains a memory leak due to improper cleanup of metadata for rejected inference requests. Attackers can submit requests with max_tokens=0 to force continual allocation of memory on decode workers until the service restarts, resulting in denial of service. The weakness is a classic memory exhaustion vulnerability (CWE-401 and CWE-770).

Affected Systems

The vulnerability affects all vLLM releases up to and including 0.29.0, running on any platform supported by the project (Linux, Windows). Users employing these versions should verify that their deployment uses the latest release.

Risk and Exploitability

The CVSS score of 8.7 reflects a high severity impact, but the EPSS score of less than 1% indicates a very low probability of exploitation at present. The exploit does not provide arbitrary code execution; it requires remote submission of specific request payloads to an exposed vLLM endpoint and is a resource‑based attack that forces worker restarts. Because the attack is purely denial of service and hinges on a small, predictable payload, it remains manageable through patching or defensive filtering.

Generated by OpenCVE AI on September 23, 2026 at 01:43 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Apply the latest vLLM release that includes the memory‑cleanup fix, at minimum v0.29.1 or newer.
  • If immediate upgrade is not possible, configure the application to reject or deny any inference request where max_tokens equals zero before it reaches the decode worker.
  • Implement runtime monitoring of memory usage for decode workers and set alerts to trigger if memory consumption exceeds expected thresholds, ensuring prompt remediation if the issue resurfaces.

Generated by OpenCVE AI on September 23, 2026 at 01:43 UTC.

Tracking

Sign in to view the affected projects.

Advisories

No advisories yet.

History

Wed, 23 Sep 2026 00:15:00 +0000

Type Values Removed Values Added
Weaknesses CWE-770
References
Metrics threat_severity

None

threat_severity

Important


Tue, 22 Sep 2026 02:30:00 +0000

Type Values Removed Values Added
Metrics ssvc

{'options': {'Automatable': 'yes', 'Exploitation': 'none', 'Technical Impact': 'partial'}, 'version': '2.0.3'}


Thu, 17 Sep 2026 22:30:00 +0000

Type Values Removed Values Added
Description vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
Title vLLM through 0.29.0 Memory Exhaustion via Rejected Requests
First Time appeared Vllm
Vllm vllm
Weaknesses CWE-401
CPEs cpe:2.3:a:vllm:vllm:*:*:*:*:*:*:*:*
Vendors & Products Vllm
Vllm vllm
References
Metrics cvssV3_1

{'score': 7.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H'}

cvssV4_0

{'score': 8.7, 'vector': 'CVSS:4.0/AV:N/AC:L/AT:N/PR:N/UI:N/VC:N/VI:N/VA:H/SC:N/SI:N/SA:N'}


cve-icon MITRE

Status: PUBLISHED

Assigner: VulnCheck

Published:

Updated: 2026-09-24T14:23:01.535Z

Reserved: 2026-09-17T21:50:02.601Z

Link: CVE-2026-93436

cve-icon Vulnrichment

Updated: 2026-09-22T01:58:58.003Z

cve-icon NVD

Status : Analyzed

Published: 2026-09-17T23:18:54.710

Modified: 2026-09-28T18:37:42.427

Link: CVE-2026-93436

cve-icon Redhat

Severity : Important

Publid Date: 2026-09-17T22:15:31Z

Links: CVE-2026-93436 - Bugzilla

cve-icon OpenCVE Enrichment

Updated: 2026-09-23T01:45:19Z

Weaknesses
  • CWE-401

    Missing Release of Memory after Effective Lifetime

  • CWE-770

    Allocation of Resources Without Limits or Throttling