Impact
The vulnerability exists in vLLM versions prior to 0.28.0, where the engine does not verify that token IDs sent to the /v1/embeddings or /pooling endpoints are non‑negative. An attacker who can access these public endpoints can supply a negative token ID, causing a CUDA device‑side assertion that corrupts the GPU context. Once triggered, all subsequent requests to the affected endpoints fail, effectively denying service until the process is restarted.
Affected Systems
Affected products are vllm-project’s vLLM component versions older than 0.28.0. Modern installations that have been updated to 0.28.0 or later are not impacted. The information might lack precise sub‑version ranges beyond the stated cutoff, so any deployment using a pre‑0.28.
Risk and Exploitability
The CVSS score of 8.7 reflects a high severity denial of service. The flaw can be exploited through unauthenticated HTTP requests to the endpoints, so privileged access is not required. The EPSS score of <1% indicates a very low probability of exploitation in the general population. The vulnerability is not listed in the CISA KEV catalog, yet it could still be leveraged in targeted attacks. No special network reachability is required beyond access to the service’s REST API.
OpenCVE Enrichment