Description
OOM Denial of Service via Unbounded Map Pre-Sizing in Apache OpenNLP SymSpellModelSerializer

Versions Affected:

- 3.0.0-M4
- 3.0.0-M5

(The opennlp-spellcheck extension was introduced in 3.0.0-M4. Releases 1.x and 2.x do not contain the affected code.)

Description:

The SymSpellModelSerializer.create() method reads two 32-bit signed integer count fields (unigramCount and bigramCount) from a binary SymSpell model stream and passes each value directly to LinkedHashMap.newLinkedHashMap() after validating only that it is non-negative. No upper bound is applied, so the count is fully attacker-controlled when the model file originates from an untrusted source.

A crafted .bin model file in which either count field is set to Integer.MAX_VALUE (or any value large enough to exhaust the available heap) causes the map to be pre-sized to a capacity of 2^30 entries. The oversized backing array is allocated on the first put() into that map, requesting 4–8 GB depending on whether compressed oops are in effect, and the load fails with an OutOfMemoryError. Because the count fields sit immediately after a fixed-size header (magic, format version, three UTF strings, the configuration fields, and the edit-distance identifier) the attacker pays no meaningful size cost to weaponize a payload: a file of well under 100 bytes plus a single real entry is sufficient to crash a JVM that loads it.

Any code path that deserializes a SymSpell model is affected, including SymSpellModels.deserialize(InputStream), SymSpellModels.fromBytes(byte[]), classpath model loading via SymSpellModelResolver.resolveByLanguage(String), the CorrectTextTool command-line tool, and model-archive loading through the registered ArtifactSerializer. The opennlp-spellcheck extension ships in the official OpenNLP binary distribution.

The practical impact is denial of service against processes that load SymSpell model files from untrusted or semi-trusted origins.

Mitigation:

- 3.x users should upgrade to 3.0.0-M6.

Note: The fix applies an upper bound to both count fields, checked before the map is pre-sized; counts that are negative or exceed the bound cause an IOException to be thrown and the read to fail fast with no large allocation. The bound is the existing AbstractModelReader.MAX_ENTRIES limit introduced earlie, which the current change promotes to public visibility so that serializers implementing their own binary format can share it. The default bound is 10,000,000, which is well above the entry counts of legitimate SymSpell dictionaries but far below any value that would threaten heap exhaustion. Deployments that legitimately need to load larger dictionaries can raise the limit at JVM startup by setting the OPENNLP_MAX_ENTRIES system property to the desired positive integer (e.g. -DOPENNLP_MAX_ENTRIES=50000000); invalid or non-positive values fall back to the default. Note that this property is shared with the model-reader limit and raising it relaxes both.

Users who cannot upgrade immediately should treat all SymSpell .bin model files as untrusted input unless their provenance is verified, and should avoid loading models supplied by end users or fetched from third-party repositories without integrity checks.
Published: 2026-09-11
Score: 7.5 High
EPSS: < 1% Very Low
KEV: No
Impact: Denial of Service (Out-of-Memory)
Action: Upgrade
AI Analysis

Impact

Apache OpenNLP’s SymSpellModelSerializer creates a LinkedHashMap sized directly from two counts that are only verified to be non‑negative and can be set to a very large value, triggering an attempt to allocate up to several gigabytes of backing array, causing an Out‑of‑MemoryError and a crash. This results in a denial‑of‑service condition to any process that loads such a file. The weakness is a classic unbounded memory allocation flaw, identified as CWE-789.

Affected Systems

The vulnerability exists in the opennlp-spellcheck extension shipped with Apache OpenNLP 3.0.0‑M4 and 3.0.0‑M5. Any code path that deserializes a SymSpell model—including SymSpellModels.deserialize, SymSpellModels.fromBytes, classpath model loading via SymSpellModelResolver, the CorrectTextTool command‑line tool, and model‑archive loading—can be exercised by an attacker.

Risk and Exploitability

The EPSS score is < 1% and the vulnerability is not listed in KEV. The exploitation scenario is straightforward: a crafted .bin file can be supplied via file upload, configuration, or any untrusted source that feeds a model into the extension. Because the exploit requires only a minimal payload of less than 100 bytes, the likelihood of successful deployment is high in environments that load models from arbitrary locations. The CVSS score is 7.5; nonetheless the impact and breadth of affected deployments indicate a severe threat.

Generated by OpenCVE AI on September 21, 2026 at 04:13 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Upgrade Apache OpenNLP to version 3.0.0‑M6 or later, which implements an upper bound on the map sizes and fails fast on invalid counts.
  • If an upgrade is not immediately possible, treat all SymSpell .bin model files as untrusted: avoid loading models supplied by end users or third‑party repositories unless their integrity and provenance are verified.
  • When larger dictionaries are legitimately required, configure the OPENNLP_MAX_ENTRIES system property to a safe upper limit (e.g., -DOPENNLP_MAX_ENTRIES=50000000) before launch; do not set a value that would allow more entries than the default 10,000,000.

Generated by OpenCVE AI on September 21, 2026 at 04:13 UTC.

Tracking

Sign in to view the affected projects.

Advisories

No advisories yet.

History

Wed, 16 Sep 2026 14:15:00 +0000

Type Values Removed Values Added
CPEs cpe:2.3:a:apache:opennlp:3.0.0:m4:*:*:*:*:*:*
cpe:2.3:a:apache:opennlp:3.0.0:m5:*:*:*:*:*:*

Mon, 14 Sep 2026 21:00:00 +0000

Type Values Removed Values Added
Metrics cvssV3_1

{'score': 7.5, 'vector': 'CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H'}

ssvc

{'options': {'Automatable': 'yes', 'Exploitation': 'none', 'Technical Impact': 'partial'}, 'version': '2.0.3'}


Sun, 13 Sep 2026 01:30:00 +0000

Type Values Removed Values Added
First Time appeared Apache
Apache opennlp
Vendors & Products Apache
Apache opennlp

Fri, 11 Sep 2026 23:45:00 +0000

Type Values Removed Values Added
Description OOM Denial of Service via Unbounded Map Pre-Sizing in Apache OpenNLP SymSpellModelSerializer Versions Affected: - 3.0.0-M4 - 3.0.0-M5 (The opennlp-spellcheck extension was introduced in 3.0.0-M4. Releases 1.x and 2.x do not contain the affected code.) Description: The SymSpellModelSerializer.create() method reads two 32-bit signed integer count fields (unigramCount and bigramCount) from a binary SymSpell model stream and passes each value directly to LinkedHashMap.newLinkedHashMap() after validating only that it is non-negative. No upper bound is applied, so the count is fully attacker-controlled when the model file originates from an untrusted source. A crafted .bin model file in which either count field is set to Integer.MAX_VALUE (or any value large enough to exhaust the available heap) causes the map to be pre-sized to a capacity of 2^30 entries. The oversized backing array is allocated on the first put() into that map, requesting 4–8 GB depending on whether compressed oops are in effect, and the load fails with an OutOfMemoryError. Because the count fields sit immediately after a fixed-size header (magic, format version, three UTF strings, the configuration fields, and the edit-distance identifier) the attacker pays no meaningful size cost to weaponize a payload: a file of well under 100 bytes plus a single real entry is sufficient to crash a JVM that loads it. Any code path that deserializes a SymSpell model is affected, including SymSpellModels.deserialize(InputStream), SymSpellModels.fromBytes(byte[]), classpath model loading via SymSpellModelResolver.resolveByLanguage(String), the CorrectTextTool command-line tool, and model-archive loading through the registered ArtifactSerializer. The opennlp-spellcheck extension ships in the official OpenNLP binary distribution. The practical impact is denial of service against processes that load SymSpell model files from untrusted or semi-trusted origins. Mitigation: - 3.x users should upgrade to 3.0.0-M6. Note: The fix applies an upper bound to both count fields, checked before the map is pre-sized; counts that are negative or exceed the bound cause an IOException to be thrown and the read to fail fast with no large allocation. The bound is the existing AbstractModelReader.MAX_ENTRIES limit introduced earlie, which the current change promotes to public visibility so that serializers implementing their own binary format can share it. The default bound is 10,000,000, which is well above the entry counts of legitimate SymSpell dictionaries but far below any value that would threaten heap exhaustion. Deployments that legitimately need to load larger dictionaries can raise the limit at JVM startup by setting the OPENNLP_MAX_ENTRIES system property to the desired positive integer (e.g. -DOPENNLP_MAX_ENTRIES=50000000); invalid or non-positive values fall back to the default. Note that this property is shared with the model-reader limit and raising it relaxes both. Users who cannot upgrade immediately should treat all SymSpell .bin model files as untrusted input unless their provenance is verified, and should avoid loading models supplied by end users or fetched from third-party repositories without integrity checks.
Title Apache OpenNLP: OOM DoS via Unbounded Array Allocation in SymSpellModelSerializer
Weaknesses CWE-789
References

cve-icon MITRE

Status: PUBLISHED

Assigner: apache

Published:

Updated: 2026-09-14T18:50:04.110Z

Reserved: 2026-07-28T16:59:40.025Z

Link: CVE-2026-67211

cve-icon Vulnrichment

Updated: 2026-09-11T21:07:08.953Z

cve-icon NVD

Status : Analyzed

Published: 2026-09-11T18:16:57.470

Modified: 2026-09-16T14:04:32.393

Link: CVE-2026-67211

cve-icon Redhat

No data.

cve-icon OpenCVE Enrichment

Updated: 2026-09-21T04:15:08Z

Weaknesses
  • CWE-789

    Memory Allocation with Excessive Size Value