PT-2026-106549 · Linux · Linux
CVE-2026-98220
·
Published
2026-10-06
·
Updated
2026-10-06
None
No severity ratings or metrics are available. When they are, we'll update the corresponding info on the page.
In the Linux kernel, the following vulnerability has been resolved:
sched ext: Fix NULL sched deref in kfunc sub-sched error paths
When the root scheduler has sub-scheds attached, the COMPAT kfunc
wrappers scx bpf select cpu and() and scx bpf dsq insert vtime() refuse
the call and report to @p's scheduler:
scx error(scx task sched(p), "... must be used");The wrappers are reachable with tasks that have no scheduler.
scx bpf select cpu and() is in the select cpu kfunc group, which
scx kfunc context filter() opens to BPF PROG TYPE SYSCALL programs;
scx bpf dsq insert vtime() is in the enqueue dispatch group, which
ops.enqueue() and ops.dispatch() may call with any KF RCU task -- the
group has no kf tasks validation, and scx dsq insert preamble() checks
task ownership with scx task on sched() precisely because @p may be an
arbitrary task.
scx task sched(p) is p->scx.sched, which is NULL for tasks past
sched ext dead() -- which clears it via scx disable and exit task() on
exit -- and for idle tasks, which the enable paths skip as they are
never scheduled through SCX. It is also an rcu dereference protected()
that expects @p's pi lock or rq lock, which neither wrapper holds.
Passing NULL to scx error() reaches scx vexit(), which dereferences
sch->exit info, oopsing the kernel.
One concrete trigger exercised while developing the fix: a
BPF PROG TYPE SYSCALL program calling the select cpu and wrapper on an
exited-but-not-reaped task while a sub-scheduler was attached (its pid
stays findable while the zombie is unreaped; faulting instruction is
the scx vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
the offset of sch->exit info):
sched ext: BPF scheduler "kfunc subsched null" enabled
sched ext: BPF sub-scheduler "kfunc subsched null" enabled
sched ext: Unassociated program run select cpu (id 76)
BUG: kernel NULL pointer dereference, address: 0000000000000398
#PF: supervisor read access in kernel mode
#PF: error code(0x0000) - not-present page
Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 7 UID: 0 PID: 8201 Comm: kfunc test runn Tainted: G W
RIP: 0010:scx vexit+0x25/0xa0
Code: ... <4c> 8b bf 98 03 00 00 ...
CR2: 0000000000000398
Call Trace:
scx exit+0x4f/0x70
scx bpf select cpu and+0xab/0xb0
bpf prog 430ed61a7b66e03a run select cpu and+0x9c/0xe7
? x64 sys bpf+0x2c/0x40
bpf prog test run syscall+0x130/0x2f0
sys bpf+0x930/0x10d0
? x64 sys bpf+0x2c/0x40
x64 sys bpf+0x2c/0x40
do syscall 64+0xbc/0x460
entry SYSCALL 64 after hwframe+0x76/0x7e
Read @p's scheduler under RCU instead, which the wrappers can do from
their guard(rcu)(): fault it when it can be determined, and when it
can't be determined -- @p is a task past sched ext dead() or an idle
task -- there is nothing obviously wrong to report, so just refuse the
call as before without faulting any scheduler.
These COMPAT wrappers are scheduled for eventual removal once the
deprecation grace period elapses, but until then -- and regardless of
their removal timeline -- they must not oops the kernel on a task they
are handed.
Found an issue in the description? Have something to add? Feel free to write us 👾
Related Identifiers
Affected Products
Linux