Impact
vLLM before 0.29.0 contains a validation flaw in SamplingParams._validate_allowed_token_ids that checks the length of allowed_token_ids against the tokenizer length instead of the model output logits width. The flaw allows attackers to supply token IDs that exceed the model’s vocabulary size, which pass validation but corrupt the GPU logits state via LogitBiasState. This corruption enables concurrent requests to sample tokens that lie outside their intended allowlists, potentially revealing data or behaving unpredictably. The weakness is a classic example of CWE‑129, an “Improper Validation of Array Index” scenario, and also introduces a memory corruption issue (CWE-787) by writing beyond array bounds during logits processing.
Affected Systems
The vulnerability affects the vLLM project’s vLLM product on all releases earlier than version 0.29.0. Users running any vLLM instance built from source or installed from packaging before the 0.29.0 release are susceptible, regardless of the deployment environment.
Risk and Exploitability
The CVSS score of 6.3 places the issue in the moderate severity range. The EPSS score of less than 1 % indicates a very low probability of exploitation in the near term, and the vulnerability is not listed in CISA’s KEV catalog. Likely attack vectors involve an attacker with network access to a vLLM server sending crafted requests that set allowed_token_ids beyond the logits width; no privileged or local execution is required. This scenario suggests the risk is moderate but the actual likelihood of exploitation remains low under normal exposure conditions.
OpenCVE Enrichment