Impact
vLLM is an inference engine for large language models that exposes OpenAI-compatible request models. In versions earlier than 0.30.0, the validation of the cache_salt parameter is too permissive. Its characters and length are allowed to exceed the restrictions required by the IPCCacheServerKey consumer in the LMCache-MP connector. When a request supplies a cache_salt containing a forbidden character or a length that is too large, the scheduler cache lookup throws an uncaught ValueError. This exception terminates the EngineCore process, which causes all concurrent users to experience a denial of service. The weakness is an input validation flaw that can lead to application misbehaviour and availability loss.
Affected Systems
vLLM is the affected product. Any deployment of vLLM using the LMCache-MP connector and running a version earlier than 0.30.0 is vulnerable. The fix is available in release 0.30.0.
Risk and Exploitability
The CVSS score of 6.5 indicates a moderate severity, and the vulnerability is not listed in the CISA KEV catalog. Because the flaw is triggered by a single crafted request and can crash the core service, the practical impact on availability can be significant for any user of the affected deployment. However, no public exploit is known, and there is no stated EPSS score. The attack vector is inferred to be network via an OpenAI-compatible API call, as that is how the cache_salt is supplied.
OpenCVE Enrichment
Github GHSA