PT-2026-98732 · Linux · Linux
CVE-2026-98069
·
Publicado
2026-09-25
·
Atualizado
2026-09-26
CVSS v3.1
8.1
Alta
| Vetor | AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H |
In the Linux kernel, the following vulnerability has been resolved:
net/rds: acquire the fastpath locks in rds conn shutdown()
rds conn shutdown() quiesces the transmit and receive-refill paths by
waiting for RDS IN XMIT and RDS RECV REFILL to be sampled clear, and
then runs the transport shutdown and rds conn path reset(). Sampling
the bits clear is not the same as owning them: the moment after the
wait event() returns, rds send xmit() can re-acquire RDS IN XMIT (or
rds ib recv refill() can re-acquire RDS RECV REFILL) and run
concurrently with the teardown.
The sender does recheck the connection state after taking the lock,
but that recheck is a classic store-buffering pattern: teardown writes
the state and reads the bit while the sender writes the bit and reads
the state. acquire in xmit() is only an acquire operation, so on
weakly ordered architectures both sides can miss each other's write,
and the transmit path then runs while the transport zeroes its rings
(e.g. rds ib ring init()) and rds send path reset() rewrites the
transmit state under it.
Oracle UEK fixed the same class of crashes - a 14-year tail of
BUG ON()s in rds ib sub signaled(), unexpected op-codes and NULL
dereferences in rds ib send cqe handler() during failover testing -
by making the teardown path acquire the fastpath bit locks instead
of testing them ("rds: Make sure transmit path and connection
tear-down does not run concurrently"). Ownership of a single word is
decided by RMW atomicity, so no cross-variable ordering is needed.
Do the same here: take both locks before calling the transport
shutdown, hold them across rds conn path reset(), and release them
explicitly with a wake-up afterwards. Both are released with
clear bit unlock(), so that the ring re-initialization done by the
transport shutdown and the transmit state rewritten by
rds send path reset() are ordered before either bit is seen clear by
the next acquire in xmit() or acquire refill().
The fastpath users of these bits - rds send xmit() and
rds ib recv refill() - are trylock style and back off while teardown
owns the locks, so no new lock dependency is introduced for them.
rds tcp reset callbacks() is different: since the previous patch it
acquires RDS IN XMIT as well, and it blocks doing so, so its wait now
spans the teardown instead of at most one send batch. That waiter
runs from rds tcp accept one() on the single-threaded krdsd workqueue
and holds rds tcp accept lock and t conn path lock while it waits, so
a duelling SYN accepted while its path is being torn down parks
accept processing for the duration of the teardown - for TCP bounded
by the (up to 5 s) drain loop in rds tcp conn path shutdown(). An IB
path's drain in rds ib conn path shutdown() has no round cap, but no
blocking waiter either: rds tcp reset callbacks() is the only blocking
acquirer of these bits and waits only on its own TCP path, and the
fastpaths are trylock-and-back-off on both transports, so a long IB
drain lengthens only that path's own quiesce. The
window is narrow: the accept-side state check has to pass before the
teardown moves the path to RDS CONN DISCONNECTING.
Because krdsd is a single global workqueue, everything else queued
there - accept processing for other connections and network
namespaces, and the flush workqueue(rds wq) in rds tcp listen stop()
during namespace teardown - waits behind the parked accept worker for
that time. It cannot deadlock, although the waits do point at each
other: the teardown blocks until the bit's holder releases it, and
the holder may be that krdsd accept worker. The holder finishes
without needing anything the teardown owns: the sync cancels
rds tcp reset callbacks() issues target cp send w and cp recv w on
the path's ordered cp wq, whose only execution slot is occupied by
the blocked cp down w itself, so they are pending at most and cancel
without flushing - a reliance on cp wq being ordered that is now
noted next to those cancels (on
---truncated---
Correção
Encontrou algum problema na descrição? Tem algo a acrescentar? Fique à vontade para nos escrever 👾
Identificadores relacionados
Produtos afetados
Linux