Impact
vLLM is an inference engine for large language models. The flaw allows a caller‑supplied X-Request-Id header to be used as a cache key for query embeddings at the /score and /rerank endpoints. A malicious user can send a request that shares a victim’s request identifier and overwrite the cached embedding, causing the victim’s documents to be evaluated against the attacker’s query. This defect also can trigger late‑interaction cache‑miss errors that may disrupt scoring operation. The impact is a loss of integrity for query scoring and potential service disruption, but it does not grant arbitrary code execution or data exfiltration.
Affected Systems
The vulnerability exists in the vllm project’s vllm engine. All releases prior to version 0.30.0 are affected; the problem is fixed starting with v0.30.0.
Risk and Exploitability
The CVSS score of 4.2 places the issue in the medium risk range. The EPSS score is not available and the vulnerability is not listed in the CISA KEV catalog. Exploitation requires the attacker to have network access to the /score or /rerank endpoint and to know or guess a victim’s X-Request-Id value. A concurrent request bearing the same identifier can overwrite the cache and trigger the vulnerability. Given these prerequisites, the likelihood of exploitation is moderate and the damage is limited to scoring integrity and possible availability problems.
OpenCVE Enrichment
Github GHSA