Description
In the Linux kernel, the following vulnerability has been resolved:

mm/huge_memory: transfer the pmd dirty bit to the folio on zap

zap_huge_pmd_folio() propagates the pmd young bit to the folio for the
file case, but not the dirty bit. The pte path does propagate it, in
zap_present_folio_ptes() and so does the pmd split path, in
__split_huge_pmd_locked().

For most file mappings the omission is harmless, because writing to a
shared file mapping goes through page_mkwrite(), which dirties the folio.
tmpfs is different: it has no page_mkwrite(), and vma_wants_writenotify()
is false for it, so a *read* fault on a MAP_SHARED tmpfs mapping installs
a writable pmd via do_read_fault(). do_read_fault() does not call
fault_dirty_shared_page(), so subsequent stores through that mapping set
only the hardware dirty bit in the pmd and never call folio_mark_dirty().
A shmem folio allocated by a fault is marked uptodate but not dirty (see
the clear: block in shmem_get_folio_gfp()), so PG_dirty is never set at
all.

Unmapping such a folio - munmap(), or exit_mmap() when the process dies -
then loses the only record that it was written, because zap_huge_pmd()
drops the pmd without transferring the dirty bit. Reclaim afterwards sees
a clean shmem folio: the whole swap-out block in shrink_folio_list() is
inside "if (folio_test_dirty(folio))", so pageout() is skipped and the
folio falls into __remove_mapping(). There, folio_is_file_lru() is false
for a swapbacked folio, so no shadow entry is created and
__filemap_remove_folio(folio, NULL) simply empties the i_pages slot. The
data is freed without ever being written to swap, and the next fault on
that index returns a freshly zeroed folio.

This is silent data loss for any process that keeps state in a MAP_SHARED
tmpfs segment across an unmap - for example a cache handed from one
process generation to the next through /dev/shm. It requires the folio to
be PMD-mapped, so it only shows up once shmem THP is enabled (which is
what we did in Meta fleet and started noticing crashes); with THP off the
pte path transfers the dirty bit correctly. It also only becomes visible
when swap is enabled, because with no swap device shmem folios (which are
on the anon LRU) are not scanned by reclaim at all, so the clean folio is
never dropped.

Reproduced on x86_64 with a tmpfs mounted huge=within_size: read-fault a
2MB-backed region, write a known pattern through the resulting mapping,
munmap, force reclaim of the cgroup, then re-map and read back. Without
this patch the region reads back as zeros and vmstat shows zswpout 0 - the
data was discarded rather than swapped. With this patch the region reads
back correctly and the pages are swapped out as expected. With
huge=never, or when the first touch is a write, the test passes either
way.
Published: 2026-09-16
Score: n/a
EPSS: < 1% Very Low
KEV: No
Impact: Silent data loss due to missing dirty bit transfer during page zap
Action: Upgrade kernel
AI Analysis

Impact

The Linux kernel bug caused the dirty state of pages that are part of a MAP_SHARED tmpfs mapping to be lost when the range is unmapped. The dirty flag is normally propagated so that reclaim can write the page to swap. Because the pmd dirty bit was not copied to the folio in certain cases, reclaim treated the page as clean, skipped swapping it out, and freed the data, resulting in the next access returning zeroed memory. This silent data loss can affect any process that relies on a shared tmpfs segment across an unmap, such as a cache handed between process generations via /dev/shm. The vulnerability does not allow arbitrary code execution but compromises data integrity.

Affected Systems

All Linux kernel versions released before the patch that fixed mm/huge_memory, particularly those that enable shmem Transparent Huge Pages and use swap space. Any system that creates MAP_SHARED tmpfs mappings with THP turned on is susceptible.

Risk and Exploitability

The CVSS score is not provided, but the bug leads to serious silent data loss. The EPSS score is below 1 %, indicating a very low probability of exploitation in the wild. The vulnerability is not listed in CISA KEV, suggesting it has not been observed in active attacks yet. The likely attack vector is local: an attacker who can manipulate shared tmpfs memory and trigger unmapping and reclaim on a host with THP and swap enabled could cause data loss. Overall, the risk is moderate to high in environments where the conditions above are present, but the exploitation probability remains low.

Generated by OpenCVE AI on September 18, 2026 at 03:56 UTC.

Remediation

No solution or workaround provided in the CVE record.

OpenCVE Recommended Actions

  • Upgrade the Linux kernel to a version that contains the fix for CVE-2026-89987
  • Disable shmem Transparent Huge Pages (e.g., echo "never" > /sys/kernel/mm/transparent_hugepage/enabled) if a patch cannot be applied promptly
  • Avoid using MAP_SHARED tmpfs for persistent state or switch to private mappings; ensure critical data are not stored in shared tmpfs segments that may be unmapped

Generated by OpenCVE AI on September 18, 2026 at 03:56 UTC.

Tracking

Sign in to view the affected projects.

Advisories

No advisories yet.

History

Fri, 18 Sep 2026 04:15:00 +0000

Type Values Removed Values Added
Weaknesses CWE-665

Wed, 16 Sep 2026 10:45:00 +0000

Type Values Removed Values Added
Description In the Linux kernel, the following vulnerability has been resolved: mm/huge_memory: transfer the pmd dirty bit to the folio on zap zap_huge_pmd_folio() propagates the pmd young bit to the folio for the file case, but not the dirty bit. The pte path does propagate it, in zap_present_folio_ptes() and so does the pmd split path, in __split_huge_pmd_locked(). For most file mappings the omission is harmless, because writing to a shared file mapping goes through page_mkwrite(), which dirties the folio. tmpfs is different: it has no page_mkwrite(), and vma_wants_writenotify() is false for it, so a *read* fault on a MAP_SHARED tmpfs mapping installs a writable pmd via do_read_fault(). do_read_fault() does not call fault_dirty_shared_page(), so subsequent stores through that mapping set only the hardware dirty bit in the pmd and never call folio_mark_dirty(). A shmem folio allocated by a fault is marked uptodate but not dirty (see the clear: block in shmem_get_folio_gfp()), so PG_dirty is never set at all. Unmapping such a folio - munmap(), or exit_mmap() when the process dies - then loses the only record that it was written, because zap_huge_pmd() drops the pmd without transferring the dirty bit. Reclaim afterwards sees a clean shmem folio: the whole swap-out block in shrink_folio_list() is inside "if (folio_test_dirty(folio))", so pageout() is skipped and the folio falls into __remove_mapping(). There, folio_is_file_lru() is false for a swapbacked folio, so no shadow entry is created and __filemap_remove_folio(folio, NULL) simply empties the i_pages slot. The data is freed without ever being written to swap, and the next fault on that index returns a freshly zeroed folio. This is silent data loss for any process that keeps state in a MAP_SHARED tmpfs segment across an unmap - for example a cache handed from one process generation to the next through /dev/shm. It requires the folio to be PMD-mapped, so it only shows up once shmem THP is enabled (which is what we did in Meta fleet and started noticing crashes); with THP off the pte path transfers the dirty bit correctly. It also only becomes visible when swap is enabled, because with no swap device shmem folios (which are on the anon LRU) are not scanned by reclaim at all, so the clean folio is never dropped. Reproduced on x86_64 with a tmpfs mounted huge=within_size: read-fault a 2MB-backed region, write a known pattern through the resulting mapping, munmap, force reclaim of the cgroup, then re-map and read back. Without this patch the region reads back as zeros and vmstat shows zswpout 0 - the data was discarded rather than swapped. With this patch the region reads back correctly and the pages are swapped out as expected. With huge=never, or when the first touch is a write, the test passes either way.
Title mm/huge_memory: transfer the pmd dirty bit to the folio on zap
First Time appeared Linux
Linux linux Kernel
CPEs cpe:2.3:o:linux:linux_kernel:*:*:*:*:*:*:*:*
Vendors & Products Linux
Linux linux Kernel
References

Subscriptions

Linux Linux Kernel
cve-icon MITRE

Status: PUBLISHED

Assigner: Linux

Published:

Updated: 2026-09-16T10:33:02.270Z

Reserved: 2026-09-11T19:38:34.779Z

Link: CVE-2026-89987

cve-icon Vulnrichment

No data.

cve-icon NVD

Status : Received

Published: 2026-09-16T11:17:09.767

Modified: 2026-09-16T11:17:09.767

Link: CVE-2026-89987

cve-icon Redhat

No data.

cve-icon OpenCVE Enrichment

Updated: 2026-09-18T05:15:03Z

Weaknesses