PT-2026-104915 · Azure Linux · Kernel

Publicado

2026-09-24

·

Atualizado

2026-09-24

Nenhuma

Não há classificações de severidade ou métricas disponíveis. Quando houver, atualizaremos as informações correspondentes na página.
In the Linux kernel, the following vulnerability has been resolved:
memcg: bypass the reclaim and oom killer for dying tasks once oom reaper is done
At Meta, we are seeing instances where an OOM killed job is stuck in the exit path for several hours. In one particular case, the job was stuck for more than 8 hours and I had to manually remove the memory.max limits to allow the process to exit.
The job was a single process job and had ~55 GiB memory.max and zswap enabled. It had almost 0 anon in memory and ~111 GiB in zswap compressed to ~51 GiB zswap pool (i.e. almost all of memory.current was zswap). Nothing was left on the LRUs to reclaim.
On further inspection, I observed ~20k threads of that process stuck with the following stack:
[<0>] mem cgroup out of memory+0x4e/0xa0 [<0>] charge memcg+0x8bf/0x990 [<0>] mem cgroup swapin charge folio+0x4e/0x80 [<0>] read swap cache async+0x10c/0x260 [<0>] swapin readahead+0x116/0x3f0 [<0>] do swap page+0x13c/0x1ce0 [<0>] handle mm fault+0x61d/0x11f0 [<0>] do user addr fault+0x3e7/0x6d0 [<0>] exc page fault+0x8f/0x110 [<0>] asm exc page fault+0x22/0x30 [<0>] get user 8+0x14/0x20 [<0>] futex cleanup+0x27/0x1c0 [<0>] futex exit release+0x47/0x60 [<0>] do exit+0x107/0x940 [<0>] do group exit+0x81/0xa0 [<0>] get signal+0x2b1/0x6e0 [<0>] arch do signal or restart+0x1a/0x1c0 [<0>] exit to user mode loop+0xa8/0x1c0 [<0>] do syscall 64+0x152/0x250 [<0>] entry SYSCALL 64 after hwframe+0x4b/0x53
In addition the dmesg was filled with "Out of memory and no killable processes..." messages.
I have no idea why oom reaper was not able to reap/unmap the process. My guess is that since oom reaper tries to acquire mmap lock in read mode limited number of times and then gives up, there might be a thread of that process which had mmap lock in write mode at that time.
My initial suspicion was the futex cleanup and kernel page fault causing infinite fault and charge retries but that was put to rest in previous discussions happened on similar problem [1].
My current theory is that it is just a simple slow serialization behind the oom lock. Unlike page allocator, memcg charge code takes the oom lock without the "try". Though memcg oom code uses mutex lock killable(), note that in the call stack get signal() consumes SIGKILL (or sigdelset(SIGKILL)) before calling do group exit(). So this mutex lock killable() is just a mutex lock() here. Therefore 10s of thousands of threads are waiting on oom lock and one by one they get -EFAULT from get user() in the futex cleanup code and bails out.
Discussion from [1] led to commit a75ffa26122b ("memcg, oom: do not bypass oom killer for dying tasks") which routes dying tasks into the OOM path precisely so the oom reaper can reap their mm and free the memory asynchronously. But the reaper is best-effort and one-shot: if it cannot take mmap lock for read (e.g. a sibling thread holds it for write) it sets MMF OOM SKIP and never retries, leaving only the glacial oom lock-serialized synchronous drain.
Once MMF OOM SKIP is set there is no more asynchronous reclaim coming for the mm, so a dying task charging against it has nothing left to wait for: it frees its memory only once it finishes exiting. Running reclaim and the (no-victim) OOM killer for it is then pointless, and doing it for 10s of thousands of exiting threads is what serializes them behind oom lock. So before reclaim, if current is an OOM victim whose reaper is done, fail the charge.
Reproduced with 20k threads, each parking a robust futex head on its own zswapped page, OOM-group-killed while a sibling holds mmap lock for write so the reaper gives up and sets MMF OOM SKIP. Tested on next-20260728 and baseline show ~90 seconds exit time while with the patch the exit time reduced to ~3 seconds.
Encontrou algum problema na descrição? Tem algo a acrescentar? Fique à vontade para nos escrever 👾

Identificadores relacionados

AZL-103754

Produtos afetados

Kernel