PT-2026-106545 · Linux · Linux
CVE-2026-98216
·
Published
2026-10-06
·
Updated
2026-10-06
None
No severity ratings or metrics are available. When they are, we'll update the corresponding info on the page.
In the Linux kernel, the following vulnerability has been resolved:
IB/hfi1: Fix the PIO CRED credit-return mmap
hfi1 file mmap()'s PIO CRED case must hand user space the single
credit-return page that holds this context's entry. That page is the
second or third page of the per-node credit-return allocation once the
hardware send context index reaches 64 or 128, so the failure below is
intermittent: when the entry lands on the first page the offset is zero
and everything works.
Two things are wrong.
First, cr page offset is a byte offset but .va is a struct
credit return *, so adding it is pointer arithmetic and scales the offset
by sizeof(struct credit return) == 64. memvirt then lands 256 KiB or
512 KiB past a 10240-byte allocation. With an IOMMU translating, that
address is inside the vmalloc range but in no vm area, so
dma mmap coherent() -> iommu dma mmap() finds no pages, vmalloc to pfn()
returns page to pfn(NULL), and remap pfn range() installs a frame above
MAXPHYADDR. The first user read then takes:
psm2 ep open pr: Corrupted page table at address 7a14d007e000
PGD 800000013886a067 P4D 800000013886a067 PUD 13886b067 PMD 13886c067
PTE 800049168e911235
Oops: Bad pagetable: 000d [#1] SMP PTI
Second, and still wrong once the arithmetic is corrected,
dma mmap coherent() describes a whole coherent buffer and selects the
page within it with vma->vm pgoff. Offsetting cpu addr has no effect:
for a vmap'd allocation iommu dma mmap() uses cpu addr only to locate the
vm area and then maps pages[vm pgoff], which hfi1 file mmap() has just
set to 0. User space therefore always receives the first credit-return
page, every credit read is for the wrong context, and send PIO stalls
forever.
Use the DMA API as intended: pass the base of the allocation with its
full length and select the page with vm pgoff. A separate length is
needed because memlen must keep describing the VMA for the existing size
check. The dma-direct path stays correct as well, since dma direct mmap()
adds the same vm pgoff to the base pfn.
Tested on a Dell T7610 (Xeon E5-2650 v2, Intel IOMMU in DMA-FQ mode)
against a Threadripper PRO 3995WX peer, both Omni-Path 100. Before this
change psm2 ep open() Oopses the kernel; with only the arithmetic
corrected psm2 ep open() succeeds but any transfer that uses send PIO
hangs, PSM2 SDMA=2 (send PIO disabled) completing normally while
PSM2 SDMA=0 (send PIO only) hangs every time. With this change send PIO,
send DMA and the default mixed mode all work.
Found an issue in the description? Have something to add? Feel free to write us 👾
Related Identifiers
Affected Products
Linux