Impact
pypdf processes PDF files that contain a /ToUnicode dictionary with an unusually long source‑code or destination‑string token. The parse_bfchar routine decodes and retains these oversized tokens, allocating memory proportionate to the token length. An attacker can craft a PDF to trigger this path, causing the interpreter to consume excessive memory during text extraction and potentially crash or become unresponsive. This is an unbounded memory consumption flaw (CWE‑400).
Affected Systems
The flaw exists in the py‑pdf:pypdf library in all releases before version 6.18.1. Any application importing pypdf that can receive arbitrary PDFs is at risk. Versions 6.18.1 and later are not affected.
Risk and Exploitability
The CVSS score of 8.7 classifies the issue as high severity. While EPSS data is unavailable, the low technical barriers to supply a malicious PDF suggest that exploitation is reasonably feasible, especially in environments where PDF parsing is performed automatically. The vulnerability is not listed in the CISA KEV catalog, but the lack of a published exploit does not reduce the theoretical importance. The likely attack vector is the delivery of a crafted PDF through email, a web page, or any interface that triggers pypdf to parse the file, leading to denial of service via memory exhaustion.
OpenCVE Enrichment