Description
vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can therefore submit a small compressed audio input that expands into a very large float32 PCM allocation, bypassing the duration guard already used by /v1/audio/transcriptions and causing an out-of-memory worker crash. Inline data URLs reach this path without being bounded by VLLM_AUDIO_FETCH_TIMEOUT. The issue affects deployments serving an audio-capable model, and authentication changes only the deployment-specific reachability. This issue is fixed in version 0.24.0.
Published: 2026-09-16
Score: 6.5 Medium
EPSS: < 1% Very Low
KEV: No
Impact: Denial of Service
Action: Immediate Patch
AI Analysis

Impact

vLLM, an inference engine for large language models, has a resource exhaustion flaw that allows an unauthenticated client to send a small compressed audio file that decompresses into a huge PCM allocation, bypassing the duration limiter and causing an out‑of‑memory worker crash. This results in a denial of service against the /v1/chat/completions endpoint. The flaw is a classic example of misuse of asyn memory allocation (CWE‑770).

Affected Systems

The vulnerability impacts deployments of the vllm-project vllm engine that serve an audio‑capable model and are running a version earlier than v0.24.0. The issue is fixed in the 0.24.0 release and later. No authentication is required to trigger the crash; however authentication can reduce reachability to the endpoint.

Risk and Exploitability

The CVSS score of 6.5 indicates moderate severity, while the EPSS score of less than 1% suggests a very low probability of exploitation in the wild. The vulnerability is listed outside of the CISA KEV catalog. Attackers can trigger it by making unauthenticated requests to the /v1/chat/completions endpoint of an exposed vllm service, optionally bypassing request timeouts. Because the failure manifests as a worker crash rather than direct data exposure, the impact is primarily availability.

Generated by OpenCVE AI on September 18, 2026 at 01:13 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Upgrade the vllm engine to version 0.24.0 or newer where the audio decoder restraint is enforced.
  • Restrict access to the /v1/chat/completions endpoint, requiring authentication or routing through an API gateway that blocks high‑volume or oversized requests.
  • Configure the audio decoder or server to enforce a maximum decoded duration or size limit, ensuring that large decompression requests cannot allocate excessive memory.

Generated by OpenCVE AI on September 18, 2026 at 01:13 UTC.

Tracking

Sign in to view the affected projects.

Advisories
Source ID Title
Github GHSA Github GHSA GHSA-hcwq-8wjf-3gcr vLLM: Unauthenticated audio decompression-bomb DoS in /v1/chat/completions
History

Wed, 07 Oct 2026 14:30:00 +0000

Type Values Removed Values Added
First Time appeared Vllm
Vllm vllm
CPEs cpe:2.3:a:vllm:vllm:*:*:*:*:*:*:*:*
Vendors & Products Vllm
Vllm vllm

Fri, 18 Sep 2026 21:30:00 +0000

Type Values Removed Values Added
Metrics ssvc

{'options': {'Automatable': 'no', 'Exploitation': 'none', 'Technical Impact': 'partial'}, 'version': '2.0.3'}


Fri, 18 Sep 2026 04:15:00 +0000

Type Values Removed Values Added
First Time appeared Vllm-project
Vllm-project vllm
Vendors & Products Vllm-project
Vllm-project vllm

Thu, 17 Sep 2026 12:15:00 +0000

Type Values Removed Values Added
References
Metrics threat_severity

None

threat_severity

Moderate


Wed, 16 Sep 2026 16:45:00 +0000

Type Values Removed Values Added
Description vLLM is an inference and serving engine for large language models. Prior to 0.24.0, the input_audio handling path for /v1/chat/completions calls AudioMediaIO.load_bytes or AudioMediaIO.load_file without passing VLLM_MAX_AUDIO_DECODE_DURATION_S to the shared audio decoder. An unauthenticated client can therefore submit a small compressed audio input that expands into a very large float32 PCM allocation, bypassing the duration guard already used by /v1/audio/transcriptions and causing an out-of-memory worker crash. Inline data URLs reach this path without being bounded by VLLM_AUDIO_FETCH_TIMEOUT. The issue affects deployments serving an audio-capable model, and authentication changes only the deployment-specific reachability. This issue is fixed in version 0.24.0.
Title vLLM: Unauthenticated audio decompression-bomb DoS in /v1/chat/completions
Weaknesses CWE-770
References
Metrics cvssV3_1

{'score': 6.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H'}


cve-icon MITRE

Status: PUBLISHED

Assigner: GitHub_M

Published:

Updated: 2026-09-18T18:11:54.721Z

Reserved: 2026-06-24T01:47:55.285Z

Link: CVE-2026-57173

cve-icon Vulnrichment

Updated: 2026-09-18T18:11:49.597Z

cve-icon NVD

Status : Analyzed

Published: 2026-09-16T17:17:24.603

Modified: 2026-10-07T14:24:22.100

Link: CVE-2026-57173

cve-icon Redhat

Severity : Moderate

Publid Date: 2026-09-16T16:34:39Z

Links: CVE-2026-57173 - Bugzilla

cve-icon OpenCVE Enrichment

Updated: 2026-09-18T04:00:03Z

Weaknesses
  • CWE-770

    Allocation of Resources Without Limits or Throttling