Description
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_hashes field, ranges in the features.mm_placeholders field, and wire-selected multimodal field processors without rebinding them to the active model renderer contract. Forged grid geometry, field types, or non-positive placeholder lengths can terminate the shared EngineCore; when an attacker knows or can induce a victim's content hash, forged cache hashes can poison or retrieve cross-request encoder-cache state; and dropped sparse placeholder masks can alter replayed transport semantics. This issue is fixed in version 0.30.0.
Published: 2026-10-05
Score: 6.5 Medium
EPSS: n/a
KEV: No
Impact: Denial of Service and Confidentiality Breach
Action: Immediate Patch
AI Analysis

Impact

The vLLM inference engine’s /inference/v1/generate endpoint, in versions prior to 0.30.0, accepts tensors and multimodal transport parameters supplied by callers. Maliciously crafted inputs such as forged grid geometry, field types, or non‑positive placeholder lengths can terminate the shared EngineCore, leading to a denial of service. Additionally, when a request’s content hash is known or can be induced, forged cache identifiers can poison or retrieve cross‑request encoder‑cache state, exposing sensitive data. Dropping sparse placeholder masks allows an attacker to alter replayed transport semantics, potentially causing incorrect inference results and further compromising integrity. The combination of service disruption and state leakage represents a significant risk to both availability and confidentiality.

Affected Systems

vLLM Project’s vllm engine, disaggregated scale‑out implementation, in any release earlier than version 0.30.0. The vulnerability exists in the API endpoint /inference/v1/generate which is exposed when the service is operable and reachable.

Risk and Exploitability

The CVSS score of 6.5 indicates moderate severity, and the vulnerability is not listed in CISA’s KEV catalog; EPSS data is unavailable. Exploitation requires remote access to the inference endpoint, typically over HTTP or HTTPS, and the ability to construct bespoke POST requests containing crafted features.kwargs_data, features.mm_hashes, and features.mm_placeholders fields. If the endpoint is publicly exposed or lacks proper authentication, a threat actor can trigger the server‑side instability and pursue cache‑based information theft without needing privileged system access. The attack path is straightforward for anyone with network reachability to the service; rate limiting or strict input validation could mitigate the risk.

Generated by OpenCVE AI on October 6, 2026 at 00:29 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Apply the v0.30.0 or later release of vllm
  • Restrict the /inference/v1/generate endpoint to authenticated and authorized clients only
  • Implement input validation or size limits on features.kwargs_data, features.mm_hashes, and features.mm_placeholders to prevent malformed requests

Generated by OpenCVE AI on October 6, 2026 at 00:29 UTC.

Tracking

Sign in to view the affected projects.

Advisories
Source ID Title
Github GHSA Github GHSA GHSA-ph72-cqr5-qpp7 vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
History

Tue, 06 Oct 2026 00:45:00 +0000

Type Values Removed Values Added
First Time appeared Vllm-project
Vllm-project vllm
Vendors & Products Vllm-project
Vllm-project vllm

Mon, 05 Oct 2026 23:00:00 +0000

Type Values Removed Values Added
Description vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_hashes field, ranges in the features.mm_placeholders field, and wire-selected multimodal field processors without rebinding them to the active model renderer contract. Forged grid geometry, field types, or non-positive placeholder lengths can terminate the shared EngineCore; when an attacker knows or can induce a victim's content hash, forged cache hashes can poison or retrieve cross-request encoder-cache state; and dropped sparse placeholder masks can alter replayed transport semantics. This issue is fixed in version 0.30.0.
Title vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Weaknesses CWE-1284
CWE-20
CWE-617
CWE-639
CWE-668
CWE-704
References
Metrics cvssV3_1

{'score': 6.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H'}


Subscriptions

Vllm-project Vllm
cve-icon MITRE

Status: PUBLISHED

Assigner: GitHub_M

Published:

Updated: 2026-10-05T22:46:03.163Z

Reserved: 2026-10-05T19:11:07.947Z

Link: CVE-2026-105754

cve-icon Vulnrichment

No data.

cve-icon NVD

Status : Received

Published: 2026-10-05T23:17:02.017

Modified: 2026-10-05T23:17:02.017

Link: CVE-2026-105754

cve-icon Redhat

No data.

cve-icon OpenCVE Enrichment

Updated: 2026-10-06T00:30:18Z

Weaknesses
  • CWE-1284

    Improper Validation of Specified Quantity in Input

  • CWE-20

    Improper Input Validation

  • CWE-617

    Reachable Assertion

  • CWE-639

    Authorization Bypass Through User-Controlled Key

  • CWE-668

    Exposure of Resource to Wrong Sphere

  • CWE-704

    Incorrect Type Conversion or Cast