PT-2026-106495 · Linux · Linux
CVE-2026-98166
·
Published
2026-10-06
·
Updated
2026-10-06
None
No severity ratings or metrics are available. When they are, we'll update the corresponding info on the page.
In the Linux kernel, the following vulnerability has been resolved:
drm/ttm: fix swapped-out resources never leaving their bulk move range
ttm tt swapout() returns the number of pages swapped out on success and
a negative error code on failure; for a populated ttm it never returns
zero. Commit b2ed01e7ad3d ("drm/ttm: Fix ttm bo swapout() infinite LRU
walk on swapout failure") moved the bulk move bookkeeping in
ttm bo swapout cb() under "if (!ret)", so the
ttm resource del bulk move unevictable() / ttm resource move to lru tail()
pair is now skipped on every successful swapout. The equivalent change
for the shrinker in commit 1d59f36e95f7 ("drm/ttm: Fix ttm bo shrink()
infinite LRU walk on backup failure") tests "lret > 0", which is what
was intended here as well.
Before b2ed01e7ad3d the resource was taken off the bulk move before the
swapout; since then a swapped-out resource stays inside its BO's
bulk move range (and on the manager LRU) although it is unevictable.
When it is later freed or the BO leaves the bulk move
(ttm resource free(), ttm bo set bulk move() via amdgpu vm bo del()),
ttm resource del bulk move() skips it because of its
!ttm resource unevictable() guard, so a range endpoint in pos->first /
pos->last is left pointing at freed memory. The next
ttm lru bulk move tail() or ttm resource add bulk move() on that cursor
is a use-after-free, seen as the resv WARN in ttm lru bulk move add(),
"list del corruption" in ttm resource move to lru tail() or a NULL
dereference in ttm resource manager next() -- minutes to hours after a
hibernation, or at process exit / reboot following one. Samuel
Ainsworth's analysis of drm/amd issue 5387 (see Link) identified the
dangling cursor; the missing removal at swapout time is the reason it
dangles.
Testing the condition for success restores the removal. On an AMD
Phoenix APU (ASUS UM3406GA, gfx1103) running suspend-then-hibernate on
a 7.0.y stable kernel carrying the backport (Ubuntu 7.0.0-31) the bug
crashed 5 of 18 hibernation cycles; a function profile of one
hibernation showed 336 ttm tt swapout() calls and zero
ttm resource del bulk move unevictable() calls. With this change the
removal happens for every swapped-out resource and 12 further cycles
were clean.
Found an issue in the description? Have something to add? Feel free to write us 👾
Related Identifiers
Affected Products
Linux