Impact
vLLM versions up to 0.29.0 contain a memory corruption flaw in the Triton _bincount_kernel where prompt token IDs are used without bounds checking against the vocabulary size. An attacker can craft multimodal audio requests containing tokens equal to the vocabulary size, causing out‑of‑bounds writes to the penalty prompt‑presence bitset. This corrupts the sampler state used by concurrent requests, altering repetition‑penalty behavior and producing incorrect or unpredictable responses. The vulnerability does not provide direct code execution but can affect the integrity of generated output.
Affected Systems
The affected software is the open‑source vLLM library from the vllm‑project, with versions 0.29.0 and earlier. Any deployment that enables multimodal audio requests and relies on the Triton sampling kernel is susceptible.
Risk and Exploitability
The CVSS score of 6.3 classifies the issue as moderate severity. The EPSS score is below 1%, indicating a very low probability of exploitation. The vulnerability is not listed in the CISA KEV catalog. The likely attack vector is via network access to an exposed vLLM API that accepts multimodal audio requests; an attacker can craft requests containing token IDs equal to the vocabulary size. Based on the description, it is inferred that the attacker must be able to submit such crafted requests to trigger the out‑of‑bounds write. Exploitation does not provide arbitrary code execution but can corrupt sampler state and alter repetition‑penalty behavior for concurrent requests.
OpenCVE Enrichment