Impact
A crafted NT_LSTM layer that sets the gate-matrix dimension greater than the deserialized na_ field in Tesseract’s default LSTM engine causes a heap out-of-bounds write during the first recognition step. The resulting heap corruption can lead to a crash or, if the attacker controls the model data, potentially controlled alteration of nearby memory. This flaw resides in the low-level LSTM implementation and is triggered by an improperly validated model input. The weakness is a type of out-of-bounds buffer overwrite (CWE-787).
Affected Systems
The vulnerability impacts the Tesseract OCR engine provided by tesseract-ocr for versions 5.5.3 and earlier. Any system that processes custom .traineddata files or LSTM models containing the vulnerable NT_LSTM layer is affected, regardless of whether the input originates internally or externally.
Risk and Exploitability
The CVSS score of 8.6 indicates high severity. Because the flaw is locally exploitable, an attacker must supply a specially crafted LSTM model file; no public exploits are documented and the EPSS score is unavailable. The likely attack vector is the injection of a malicious traineddata file into the OCR pipeline; this is inferred from the description. The vulnerability is not listed in CISA’s KEV catalog, suggesting that there is currently no widespread automated exploitation. However, systems that accept untrusted model data should still treat the vulnerability as a serious risk potential code execution if the attacker can influence the corrupted memory region.
OpenCVE Enrichment