Impact
vLLM, an inference engine for large language models, has a resource exhaustion flaw that allows an unauthenticated client to send a small compressed audio file that decompresses into a huge PCM allocation, bypassing the duration limiter and causing an out‑of‑memory worker crash. This results in a denial of service against the /v1/chat/completions endpoint. The flaw is a classic example of misuse of asyn memory allocation (CWE‑770).
Affected Systems
The vulnerability impacts deployments of the vllm-project vllm engine that serve an audio‑capable model and are running a version earlier than v0.24.0. The issue is fixed in the 0.24.0 release and later. No authentication is required to trigger the crash; however authentication can reduce reachability to the endpoint.
Risk and Exploitability
The CVSS score of 6.5 indicates moderate severity, while the EPSS score of less than 1% suggests a very low probability of exploitation in the wild. The vulnerability is listed outside of the CISA KEV catalog. Attackers can trigger it by making unauthenticated requests to the /v1/chat/completions endpoint of an exposed vllm service, optionally bypassing request timeouts. Because the failure manifests as a worker crash rather than direct data exposure, the impact is primarily availability.
OpenCVE Enrichment
Github GHSA