Impact
Tesseract OCR’s ReadNormProtos parses a NORMPROTO token from a .traineddata file into a fixed 61‑byte stack buffer using operator>>(char*) without setting a stream width. Tokens longer than 60 characters result in a stack buffer overflow, allowing up to 39 bytes of attacker‑controlled memory to be written during TessBaseAPI::Init. The resulting corruption can cause the engine to crash or, on standard‑library implementations that do not enforce bounds, enable control‑flow hijacking. The flaw corresponds to buffer‑overflow weaknesses (CWE‑120, CWE‑121).
Affected Systems
All builds of Tesseract OCR 5.5.3 and earlier that use libstdc++ or other standard libraries lacking bounds enforcement are affected. Builds that employ Apple’s libc++ with C++20 bounded array overload are incidentally protected. The vulnerability is triggered when the engine loads an untrusted or malicious .traineddata file, so any installation that processes externally sourced traineddata is at risk.
Risk and Exploitability
The CVSS score of 8.6 indicates high severity. EPSS is not available and the vulnerability is not listed in CISA’s KEV catalog, so known exploitation prevalence is unknown. An attacker only needs to supply a crafted traineddata file that drives TessBaseAPI::Init; the stack overflow can lead to denial of service and, in vulnerable standard‑library builds, may allow control‑flow hijacking. Until a patched version or a reliable mitigation is deployed, the risk remains high.
OpenCVE Enrichment