Impact
The vLLM inference engine’s /inference/v1/generate endpoint, in versions prior to 0.30.0, accepts tensors and multimodal transport parameters supplied by callers. Maliciously crafted inputs such as forged grid geometry, field types, or non‑positive placeholder lengths can terminate the shared EngineCore, leading to a denial of service. Additionally, when a request’s content hash is known or can be induced, forged cache identifiers can poison or retrieve cross‑request encoder‑cache state, exposing sensitive data. Dropping sparse placeholder masks allows an attacker to alter replayed transport semantics, potentially causing incorrect inference results and further compromising integrity. The combination of service disruption and state leakage represents a significant risk to both availability and confidentiality.
Affected Systems
vLLM Project’s vllm engine, disaggregated scale‑out implementation, in any release earlier than version 0.30.0. The vulnerability exists in the API endpoint /inference/v1/generate which is exposed when the service is operable and reachable.
Risk and Exploitability
The CVSS score of 6.5 indicates moderate severity, and the vulnerability is not listed in CISA’s KEV catalog; EPSS data is unavailable. Exploitation requires remote access to the inference endpoint, typically over HTTP or HTTPS, and the ability to construct bespoke POST requests containing crafted features.kwargs_data, features.mm_hashes, and features.mm_placeholders fields. If the endpoint is publicly exposed or lacks proper authentication, a threat actor can trigger the server‑side instability and pursue cache‑based information theft without needing privileged system access. The attack path is straightforward for anyone with network reachability to the service; rate limiting or strict input validation could mitigate the risk.
OpenCVE Enrichment
Github GHSA