Impact
NVIDIA TensorRT‑LLM has a flaw in its OpenAI‑compatible inference API that permits an attacker to trigger unlimited GPU memory allocation. The vulnerability is a classic resource exhaustion weakness, which can consume all available GPU resources and render the inference service inoperable. Consequently, a successful exploit could lead to a denial‑of‑service scenario affecting any client relying on the TensorRT‑LLM endpoint for predictions.
Affected Systems
The affected product is NVIDIA TensorRT‑LLM. No specific version range is provided in the data, so all releases of TensorRT‑LLM that implement the OpenAI‑compatible inference API may be impacted.
Risk and Exploitability
The CVSS score of 6.2 indicates a moderate severity, and the EPSS score of less than 1% suggests exploitation is currently unlikely to be observed in the wild. The vulnerability is not listed in CISA's KEV catalog. Based on the description, the attack vector is inferred to be remote via the public inference API. An adversary would need to make repeated inference requests that trigger excessive GPU memory allocation, but no additional authentication or privilege escalation is required.
OpenCVE Enrichment