Impact
vLLM through 0.29.0 contains a memory leak due to improper cleanup of metadata for rejected inference requests. Attackers can submit requests with max_tokens=0 to force continual allocation of memory on decode workers until the service restarts, resulting in denial of service. The weakness is a classic memory exhaustion vulnerability (CWE-401 and CWE-770).
Affected Systems
The vulnerability affects all vLLM releases up to and including 0.29.0, running on any platform supported by the project (Linux, Windows). Users employing these versions should verify that their deployment uses the latest release.
Risk and Exploitability
The CVSS score of 8.7 reflects a high severity impact, but the EPSS score of less than 1% indicates a very low probability of exploitation at present. The exploit does not provide arbitrary code execution; it requires remote submission of specific request payloads to an exposed vLLM endpoint and is a resource‑based attack that forces worker restarts. Because the attack is purely denial of service and hinges on a small, predictable payload, it remains manageable through patching or defensive filtering.
OpenCVE Enrichment