π¨ CVE-2026-97924
In the Linux kernel, the following vulnerability has been resolved:
tracing/user_events: Don't destroy fields when event removal fails
destroy_user_event() destroys the event's fields before attempting to
remove the trace event call. If user_event_set_call_visible() fails,
e.g. because the event is still enabled and trace_remove_event_call()
returns -EBUSY, the event is left registered with an irreversibly
destroyed field list. Any subsequent interaction with the event then
operates on an empty field list while it is still fully visible in
tracefs.
Move the field destruction after the call removal, and splice the
field list back onto the event when the removal fails so the event
remains in a consistent state.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tracing/user_events: Don't destroy fields when event removal fails
destroy_user_event() destroys the event's fields before attempting to
remove the trace event call. If user_event_set_call_visible() fails,
e.g. because the event is still enabled and trace_remove_event_call()
returns -EBUSY, the event is left registered with an irreversibly
destroyed field list. Any subsequent interaction with the event then
operates on an empty field list while it is still fully visible in
tracefs.
Move the field destruction after the call removal, and splice the
field list back onto the event when the removal fails so the event
remains in a consistent state.
π@cveNotify
π¨ CVE-2026-97925
In the Linux kernel, the following vulnerability has been resolved:
tick/broadcast: Plug clockevents replacement race
ζ±ζΊδΉΎ reported and decoded the following race condition when a broadcast
device is replaced:
CPUA CPUB
__tick_broadcast_oneshot_control()
bc = tick_broadcast_device.evtdev;
tick_install_broadcast_device(dev)
clockevents_exchange_device(cur, dev)
shutdown(cur);
detach(cur);
cur->handler = noop;
tick_broadcast_device.evtdev = dev;
tick_broadcast_set_event(bc, next_event); <- FAIL: arms a detached device.
If the original broadcast device has a restricted interrupt affinity mask
and the last CPU in that mask goes offline then the BUG() in
tick_cleanup_dead_cpu() triggers because the clockevent device is not in
detached state.
The reason for this is that tick_install_broadcast_device() is not
serialized vs. tick broadcast operations.
The obvious cure is to serialize tick_install_broadcast_device() with
tick_broadcast_lock against a concurrent tick broadcast operation.
That requires to split clockevents_exchange_device() into two parts, one
which does the exchange, shutdown and detach operation and the other which
drops the module reference count. This is required because the module
reference cannot be dropped while holding tick_broadcast_lock.
Let clockevents_exchange_device() do both operations as before, but let the
broadcast device code take the two step approach and do the device
exchange under tick_broadcast_lock and drop the module reference count
after releasing it.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tick/broadcast: Plug clockevents replacement race
ζ±ζΊδΉΎ reported and decoded the following race condition when a broadcast
device is replaced:
CPUA CPUB
__tick_broadcast_oneshot_control()
bc = tick_broadcast_device.evtdev;
tick_install_broadcast_device(dev)
clockevents_exchange_device(cur, dev)
shutdown(cur);
detach(cur);
cur->handler = noop;
tick_broadcast_device.evtdev = dev;
tick_broadcast_set_event(bc, next_event); <- FAIL: arms a detached device.
If the original broadcast device has a restricted interrupt affinity mask
and the last CPU in that mask goes offline then the BUG() in
tick_cleanup_dead_cpu() triggers because the clockevent device is not in
detached state.
The reason for this is that tick_install_broadcast_device() is not
serialized vs. tick broadcast operations.
The obvious cure is to serialize tick_install_broadcast_device() with
tick_broadcast_lock against a concurrent tick broadcast operation.
That requires to split clockevents_exchange_device() into two parts, one
which does the exchange, shutdown and detach operation and the other which
drops the module reference count. This is required because the module
reference cannot be dropped while holding tick_broadcast_lock.
Let clockevents_exchange_device() do both operations as before, but let the
broadcast device code take the two step approach and do the device
exchange under tick_broadcast_lock and drop the module reference count
after releasing it.
π@cveNotify
π¨ CVE-2026-97926
In the Linux kernel, the following vulnerability has been resolved:
ufs: validate cylinder group metadata before caching it
ufs_read_cylinder() copies the cylinder group index and the rotor
positions straight from the on-disk group and caches them without any
check:
ucpi->c_cgx = fs32_to_cpu(sb, ucg->cg_cgx);
ucpi->c_rotor = fs32_to_cpu(sb, ucg->cg_rotor);
ucpi->c_frotor = fs32_to_cpu(sb, ucg->cg_frotor);
ucpi->c_irotor = fs32_to_cpu(sb, ucg->cg_irotor);
They are then used as indices during allocation and free:
- c_cgx indexes the cylinder summary array as
UFS_SB(sb)->fs_cs(ucpi->c_cgx), so a value past s_ncg writes a 32
bit count outside the s_csp allocation.
- c_frotor becomes a bitmap scan start, start = c_frotor >> 3, and
then length = ((s_fpg + 7) >> 3) - start. A start beyond the block
bitmap wraps the unsigned length to a huge value, so ubh_scanc()
walks far past the cylinder group buffers. c_irotor drives the
inode bitmap the same way.
A crafted image can set any of these freely, turning an ordinary
allocation into an out of bounds access.
Reject a cylinder group whose recorded index does not match the group
being read, or whose rotors fall outside the group, before the metadata
is cached. Valid filesystems keep cg_cgx equal to the group number and
the rotors within the group, so only malformed images are rejected.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ufs: validate cylinder group metadata before caching it
ufs_read_cylinder() copies the cylinder group index and the rotor
positions straight from the on-disk group and caches them without any
check:
ucpi->c_cgx = fs32_to_cpu(sb, ucg->cg_cgx);
ucpi->c_rotor = fs32_to_cpu(sb, ucg->cg_rotor);
ucpi->c_frotor = fs32_to_cpu(sb, ucg->cg_frotor);
ucpi->c_irotor = fs32_to_cpu(sb, ucg->cg_irotor);
They are then used as indices during allocation and free:
- c_cgx indexes the cylinder summary array as
UFS_SB(sb)->fs_cs(ucpi->c_cgx), so a value past s_ncg writes a 32
bit count outside the s_csp allocation.
- c_frotor becomes a bitmap scan start, start = c_frotor >> 3, and
then length = ((s_fpg + 7) >> 3) - start. A start beyond the block
bitmap wraps the unsigned length to a huge value, so ubh_scanc()
walks far past the cylinder group buffers. c_irotor drives the
inode bitmap the same way.
A crafted image can set any of these freely, turning an ordinary
allocation into an out of bounds access.
Reject a cylinder group whose recorded index does not match the group
being read, or whose rotors fall outside the group, before the metadata
is cached. Valid filesystems keep cg_cgx equal to the group number and
the rotors within the group, so only malformed images are rejected.
π@cveNotify
π¨ CVE-2026-97927
In the Linux kernel, the following vulnerability has been resolved:
ufs: create the root dentry after loading cylinder metadata
ufs_fill_super() installed sb->s_root before it loaded the cylinder
group structures for a writable mount:
sb->s_root = d_make_root(inode);
...
if (!sb_rdonly(sb))
if (!ufs_read_cylinder_structures(sb))
goto failed;
When ufs_read_cylinder_structures() failed, the error path freed the
in-core superblock information and set sb->s_fs_info to NULL while
sb->s_root stayed installed. get_tree_bdev() then reached
deactivate_locked_super(), and because s_root was present,
generic_shutdown_super() called sync_filesystem() and the put_super
operation. Both dereference UFS_SB(sb), which is now NULL, so a mount
that fails only while reading the cylinder groups oopses during
teardown. A crafted image whose first cylinder group cannot be read
reaches this path.
Load the cylinder group metadata first and create the root dentry last,
so the superblock is published to the VFS only once it is fully set up.
ufs_setup_cstotal() and ufs_read_cylinder_structures() take only the
super_block and do not use the root inode, so the reordering is safe.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ufs: create the root dentry after loading cylinder metadata
ufs_fill_super() installed sb->s_root before it loaded the cylinder
group structures for a writable mount:
sb->s_root = d_make_root(inode);
...
if (!sb_rdonly(sb))
if (!ufs_read_cylinder_structures(sb))
goto failed;
When ufs_read_cylinder_structures() failed, the error path freed the
in-core superblock information and set sb->s_fs_info to NULL while
sb->s_root stayed installed. get_tree_bdev() then reached
deactivate_locked_super(), and because s_root was present,
generic_shutdown_super() called sync_filesystem() and the put_super
operation. Both dereference UFS_SB(sb), which is now NULL, so a mount
that fails only while reading the cylinder groups oopses during
teardown. A crafted image whose first cylinder group cannot be read
reaches this path.
Load the cylinder group metadata first and create the root dentry last,
so the superblock is published to the VFS only once it is fully set up.
ufs_setup_cstotal() and ufs_read_cylinder_structures() take only the
super_block and do not use the root inode, so the reordering is safe.
π@cveNotify
π¨ CVE-2026-97928
In the Linux kernel, the following vulnerability has been resolved:
drm/amdgpu: skip the VMID 0 flush for VRAM
Clear-on-release only runs on VRAM, which amdgpu_ttm_map_buffer() reaches
via its direct MC address without programming a GART window, yet the wipe
still forces a VMID 0 flush. On GFX11 (e.g. Navi33) that spurious SDMA
flush can wedge the engine; only flush when a GART window is actually used.
v2: Let amdgpu_ttm_map_buffer() return whether the VMID 0 flush is needed,
and drive the clear and copy paths from that. (Christian)
v3: Make the vm_needs_flush output parameter mandatory instead of
allowing NULL. (Christian)
(cherry picked from commit a306e406e570b74318ff7d80e5b07b540ca1d3a9)
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
drm/amdgpu: skip the VMID 0 flush for VRAM
Clear-on-release only runs on VRAM, which amdgpu_ttm_map_buffer() reaches
via its direct MC address without programming a GART window, yet the wipe
still forces a VMID 0 flush. On GFX11 (e.g. Navi33) that spurious SDMA
flush can wedge the engine; only flush when a GART window is actually used.
v2: Let amdgpu_ttm_map_buffer() return whether the VMID 0 flush is needed,
and drive the clear and copy paths from that. (Christian)
v3: Make the vm_needs_flush output parameter mandatory instead of
allowing NULL. (Christian)
(cherry picked from commit a306e406e570b74318ff7d80e5b07b540ca1d3a9)
π@cveNotify
π¨ CVE-2026-97929
In the Linux kernel, the following vulnerability has been resolved:
ALSA: usbusx2y: validate URB actual_length in interrupt callback
i_usx2y_in04_int() processes the interrupt URB data without checking
urb->actual_length. A short transfer from a malfunctioning device
would cause the handler to process uninitialized heap data from the
kmalloc-allocated in04_buf, which is then copied to the mmap-accessible
ctl_snapshot[] array.
Fix by using kzalloc() for in04_buf to zero-initialize the buffer,
and adding an actual_length check to skip processing on short
transfers while still resubmitting the URB.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ALSA: usbusx2y: validate URB actual_length in interrupt callback
i_usx2y_in04_int() processes the interrupt URB data without checking
urb->actual_length. A short transfer from a malfunctioning device
would cause the handler to process uninitialized heap data from the
kmalloc-allocated in04_buf, which is then copied to the mmap-accessible
ctl_snapshot[] array.
Fix by using kzalloc() for in04_buf to zero-initialize the buffer,
and adding an actual_length check to skip processing on short
transfers while still resubmitting the URB.
π@cveNotify
π¨ CVE-2026-97930
In the Linux kernel, the following vulnerability has been resolved:
ALSA: usbusx2y: fix in04_last array size mismatch with in04_buf
The in04_last array in struct usx2ydev is declared as char[24], but
in04_buf is allocated as sizeof(struct us428_ctls) which is 21 bytes.
In i_usx2y_in04_int(), when ctl_snapshot_last == -2 (initialization
path):
memcpy(usx2y->in04_last, usx2y->in04_buf, sizeof(usx2y->in04_last));
This copies 24 bytes from a 21-byte slab allocation, reading 3 bytes
past the end of the source object.
Introduce a USX2Y_IN04_SIZE constant defined as sizeof(struct
us428_ctls) and use it consistently for the in04_last array, the
in04_buf allocation, the URB transfer length, and the comparison loop,
replacing the bare 24 and 21 literals throughout.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ALSA: usbusx2y: fix in04_last array size mismatch with in04_buf
The in04_last array in struct usx2ydev is declared as char[24], but
in04_buf is allocated as sizeof(struct us428_ctls) which is 21 bytes.
In i_usx2y_in04_int(), when ctl_snapshot_last == -2 (initialization
path):
memcpy(usx2y->in04_last, usx2y->in04_buf, sizeof(usx2y->in04_last));
This copies 24 bytes from a 21-byte slab allocation, reading 3 bytes
past the end of the source object.
Introduce a USX2Y_IN04_SIZE constant defined as sizeof(struct
us428_ctls) and use it consistently for the in04_last array, the
in04_buf allocation, the URB transfer length, and the comparison loop,
replacing the bare 24 and 21 literals throughout.
π@cveNotify
π¨ CVE-2026-97931
In the Linux kernel, the following vulnerability has been resolved:
ALSA: us122l: Prevent write upgrades for read mappings
The hwdep mmap callback rejects read-buffer mappings that are initially
writable, but leaves VM_MAYWRITE set on mappings created with PROT_READ.
A process that can open the hwdep node O_RDWR can later use mprotect() to
make the mapping writable.
The read allocation begins with struct usb_stream. Its read_size member is
used by the fault handler to decide which pages belong to the read buffer.
The read VMA intentionally remains expandable because pcm_usb_stream uses
mremap() after reading that size. Changing read_size first can therefore
map and access pages beyond the allocation. The same member is also
consumed by usb_stream_free(), where changing it can make
free_pages_exact() release pages outside the allocation.
Clear VM_MAYWRITE for read-buffer mappings after rejecting an initially
writable VMA. This keeps the separate output-buffer mapping writable while
preventing later permission upgrades.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ALSA: us122l: Prevent write upgrades for read mappings
The hwdep mmap callback rejects read-buffer mappings that are initially
writable, but leaves VM_MAYWRITE set on mappings created with PROT_READ.
A process that can open the hwdep node O_RDWR can later use mprotect() to
make the mapping writable.
The read allocation begins with struct usb_stream. Its read_size member is
used by the fault handler to decide which pages belong to the read buffer.
The read VMA intentionally remains expandable because pcm_usb_stream uses
mremap() after reading that size. Changing read_size first can therefore
map and access pages beyond the allocation. The same member is also
consumed by usb_stream_free(), where changing it can make
free_pages_exact() release pages outside the allocation.
Clear VM_MAYWRITE for read-buffer mappings after rejecting an initially
writable VMA. This keeps the separate output-buffer mapping writable while
preventing later permission upgrades.
π@cveNotify
π¨ CVE-2026-97932
In the Linux kernel, the following vulnerability has been resolved:
tracing: Don't dereference trace_event_file in deferred trigger free
The enable_event trigger defers trace_event_put_ref() to the
trigger free kthread, but the trace_event_file can already be freed
when the instance is removed.
Keep the trace_event_call directly in enable_trigger_data so the
deferred free does not access the freed trace_event_file.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tracing: Don't dereference trace_event_file in deferred trigger free
The enable_event trigger defers trace_event_put_ref() to the
trigger free kthread, but the trace_event_file can already be freed
when the instance is removed.
Keep the trace_event_call directly in enable_trigger_data so the
deferred free does not access the freed trace_event_file.
π@cveNotify
π¨ CVE-2026-97934
In the Linux kernel, the following vulnerability has been resolved:
tracing: Fix memory corruption from a "STACKTRACE" histogram key
"cpu", "CPU", "stacktrace" and "STACKTRACE" are generic fields, defined
with an offset and a size of zero so that the filter code can match them
by name. parse_field() maps them onto their common_* equivalents for
backward compatibility, but unlike the common_* names it hands the
placeholder back to the caller instead of NULL.
create_hist_field() takes a non-NULL field as a promise that the record
carries a stacktrace and picks HIST_FIELD_FN_STACK, so the __data_loc
word is read from offset 0, that is from common_type, and its low 16
bits are followed as an offset into the record. What is found there
becomes the length of an unbounded memcpy. Pick an event whose id is
small enough that the offset stays inside its own record and the length
is a kernel text address:
# cd /sys/kernel/tracing
# echo 'hist:keys=STACKTRACE' > events/ftrace/print/trigger
# echo hello > trace_marker
Oops: general protection fault, probably for non-canonical address
RIP: 0010:rb_next+0x23/0x60
</IRQ>
RIP: 0010:memcpy+0xc/0x30
event_hist_trigger+0x2e7/0x12c0
Kernel panic - not syncing: Fatal exception in interrupt
Leave the field NULL, which is what the comment above the branch says
the code does and what common_stacktrace already does. FILTER_CPU and
FILTER_COMM are left alone, their create_hist_field() branches never
look at the field.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tracing: Fix memory corruption from a "STACKTRACE" histogram key
"cpu", "CPU", "stacktrace" and "STACKTRACE" are generic fields, defined
with an offset and a size of zero so that the filter code can match them
by name. parse_field() maps them onto their common_* equivalents for
backward compatibility, but unlike the common_* names it hands the
placeholder back to the caller instead of NULL.
create_hist_field() takes a non-NULL field as a promise that the record
carries a stacktrace and picks HIST_FIELD_FN_STACK, so the __data_loc
word is read from offset 0, that is from common_type, and its low 16
bits are followed as an offset into the record. What is found there
becomes the length of an unbounded memcpy. Pick an event whose id is
small enough that the offset stays inside its own record and the length
is a kernel text address:
# cd /sys/kernel/tracing
# echo 'hist:keys=STACKTRACE' > events/ftrace/print/trigger
# echo hello > trace_marker
Oops: general protection fault, probably for non-canonical address
RIP: 0010:rb_next+0x23/0x60
</IRQ>
RIP: 0010:memcpy+0xc/0x30
event_hist_trigger+0x2e7/0x12c0
Kernel panic - not syncing: Fatal exception in interrupt
Leave the field NULL, which is what the comment above the branch says
the code does and what common_stacktrace already does. FILTER_CPU and
FILTER_COMM are left alone, their create_hist_field() branches never
look at the field.
π@cveNotify
π¨ CVE-2026-97935
In the Linux kernel, the following vulnerability has been resolved:
tracing: Set the trace clock before registering the histogram trigger
hist_register_trigger() puts the trigger on the global named_triggers
list in cmd_ops->init(), and only then sets the trace clock:
if (data->cmd_ops->init) {
ret = data->cmd_ops->init(data);
if (ret < 0)
goto out;
}
if (hist_data->enable_timestamps) {
ret = tracing_set_clock(file->tr, hist_data->attrs->clock);
if (ret) {
hist_err(tr, HIST_ERR_SET_CLOCK_FAIL, errpos(clock));
goto out;
}
The clock string is not checked anywhere before that call, so a named
trigger using common_timestamp with an unknown clock fails after it has
already become findable. event_hist_trigger_parse() then frees it
without taking it off the list, and the next lookup by name reads the
freed object:
~# cd /sys/kernel/tracing/events/sched/sched_switch
~# echo 'hist:name=foo:keys=common_pid:ts=common_timestamp:clock=bogus' > trigger
bash: echo: write error: Invalid argument
~# echo 'hist:name=foo:keys=common_pid' > trigger
BUG: KASAN: slab-use-after-free in find_named_trigger+0xac/0xc0
Read of size 8 at addr ffff88800915d760 by task init/1
find_named_trigger+0xac/0xc0
hist_register_trigger+0xc1/0x900
event_hist_trigger_parse+0x3146/0x6af0
event_trigger_write+0xce/0x160
Freed by task 63:
kfree+0x154/0x420
trigger_kthread_fn+0xfd/0x160
Set the clock before the trigger is registered, so that nothing which
can fail runs after it is published, the way commit 6f86bdeab633
("tracing: Fix bad hist from corrupting named_triggers list") moved the
registration below the rest of the setup.
tracing_set_filter_buffering() is reference counted, so the init failure
path has to drop the reference that the clock block now takes first.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tracing: Set the trace clock before registering the histogram trigger
hist_register_trigger() puts the trigger on the global named_triggers
list in cmd_ops->init(), and only then sets the trace clock:
if (data->cmd_ops->init) {
ret = data->cmd_ops->init(data);
if (ret < 0)
goto out;
}
if (hist_data->enable_timestamps) {
ret = tracing_set_clock(file->tr, hist_data->attrs->clock);
if (ret) {
hist_err(tr, HIST_ERR_SET_CLOCK_FAIL, errpos(clock));
goto out;
}
The clock string is not checked anywhere before that call, so a named
trigger using common_timestamp with an unknown clock fails after it has
already become findable. event_hist_trigger_parse() then frees it
without taking it off the list, and the next lookup by name reads the
freed object:
~# cd /sys/kernel/tracing/events/sched/sched_switch
~# echo 'hist:name=foo:keys=common_pid:ts=common_timestamp:clock=bogus' > trigger
bash: echo: write error: Invalid argument
~# echo 'hist:name=foo:keys=common_pid' > trigger
BUG: KASAN: slab-use-after-free in find_named_trigger+0xac/0xc0
Read of size 8 at addr ffff88800915d760 by task init/1
find_named_trigger+0xac/0xc0
hist_register_trigger+0xc1/0x900
event_hist_trigger_parse+0x3146/0x6af0
event_trigger_write+0xce/0x160
Freed by task 63:
kfree+0x154/0x420
trigger_kthread_fn+0xfd/0x160
Set the clock before the trigger is registered, so that nothing which
can fail runs after it is published, the way commit 6f86bdeab633
("tracing: Fix bad hist from corrupting named_triggers list") moved the
registration below the rest of the setup.
tracing_set_filter_buffering() is reference counted, so the init failure
path has to drop the reference that the clock block now takes first.
π@cveNotify
π¨ CVE-2026-97936
In the Linux kernel, the following vulnerability has been resolved:
tracing: Fix memory corruption from the histogram stacktrace modifier
parse_field() sets HIST_FIELD_FL_STACKTRACE from the ".stacktrace"
modifier before it looks the field name up, and nothing afterwards
checks that the name resolved to a field which holds a stacktrace.
create_hist_field() picks HIST_FIELD_FN_STACK on the strength of the
field pointer alone, which reads a __data_loc word from the record and
follows its low 16 bits as an offset into the same record.
event_hist_trigger() takes the first word there as an entry count and
copies that many longs into a 31 entry array:
n_entries = *stack;
memcpy(entries, ++stack, n_entries * sizeof(unsigned long));
Neither end of that copy is bounded, and the count is whatever the event
holds at the offset, so any field will do:
# cd /sys/kernel/tracing/events/sched/sched_process_fork
# echo 'hist:keys=parent_pid.stacktrace' > trigger
# (true)
BUG: kernel NULL pointer dereference, address: 0000000000000008
RIP: 0010:rb_insert_color+0x18/0x130
timerqueue_linked_add+0x7e/0xd0
enqueue_hrtimer+0x39/0xb0
__hrtimer_run_queues+0x10f/0x1f0
</IRQ>
RIP: 0010:memcpy+0xc/0x30
event_hist_trigger+0x165/0x690
The timer interrupt landed on the rbtree the copy had already run over.
No debug options are needed for this; KASAN reports the same write as an
out-of-bounds read of 13835058055416381440 bytes.
Documentation/trace/histogram.rst already states the rule, "must be a
long[] type", so enforce it once the name has been resolved. Names which
resolve to no field at all, "hitcount.stacktrace" and the common_*
pseudo-fields, are refused for the same reason: they hold no stacktrace
to read.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
tracing: Fix memory corruption from the histogram stacktrace modifier
parse_field() sets HIST_FIELD_FL_STACKTRACE from the ".stacktrace"
modifier before it looks the field name up, and nothing afterwards
checks that the name resolved to a field which holds a stacktrace.
create_hist_field() picks HIST_FIELD_FN_STACK on the strength of the
field pointer alone, which reads a __data_loc word from the record and
follows its low 16 bits as an offset into the same record.
event_hist_trigger() takes the first word there as an entry count and
copies that many longs into a 31 entry array:
n_entries = *stack;
memcpy(entries, ++stack, n_entries * sizeof(unsigned long));
Neither end of that copy is bounded, and the count is whatever the event
holds at the offset, so any field will do:
# cd /sys/kernel/tracing/events/sched/sched_process_fork
# echo 'hist:keys=parent_pid.stacktrace' > trigger
# (true)
BUG: kernel NULL pointer dereference, address: 0000000000000008
RIP: 0010:rb_insert_color+0x18/0x130
timerqueue_linked_add+0x7e/0xd0
enqueue_hrtimer+0x39/0xb0
__hrtimer_run_queues+0x10f/0x1f0
</IRQ>
RIP: 0010:memcpy+0xc/0x30
event_hist_trigger+0x165/0x690
The timer interrupt landed on the rbtree the copy had already run over.
No debug options are needed for this; KASAN reports the same write as an
out-of-bounds read of 13835058055416381440 bytes.
Documentation/trace/histogram.rst already states the rule, "must be a
long[] type", so enforce it once the name has been resolved. Names which
resolve to no field at all, "hitcount.stacktrace" and the common_*
pseudo-fields, are refused for the same reason: they hold no stacktrace
to read.
π@cveNotify
π¨ CVE-2026-97937
In the Linux kernel, the following vulnerability has been resolved:
ftrace: fork: Initialize function graph state before copy_exec_state()
dup_task_struct() copies the parent's task_struct, including ret_stack.
ftrace_graph_init_task() clears the copied function graph state, but it
currently runs after copy_exec_state().
For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state.
If that allocation fails, copy_process() reaches bad_fork_free and
free_task() calls ftrace_graph_exit_task(). Since the child still carries
the parent's ret_stack pointer, the unwind frees the parent's active
function graph return stack. The parent subsequently accesses freed memory
from function_graph_enter_regs().
KASAN reports:
[ 22.190920] ==================================================================
[ 22.195899] BUG: KASAN: slab-use-after-free in function_graph_enter_regs+0xa76/0xb90
[ 22.200747] Write of size 8 at addr ff110000054dc0a8 by task repro/1
[ 22.205134]
[ 22.210770] CPU: 0 UID: 0 PID: 1 Comm: repro Not tainted 7.2.0-07732-g9328b3b03bdc-dirty #3 PREEMPT(lazy)
[ 22.212576] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
[ 22.213750] Call Trace:
[ 22.215271] <TASK>
[ 22.216242] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.217774] dump_stack_lvl+0x4e/0x70
[ 22.220531] print_report+0x157/0x4b4
[ 22.223202] ? fixup_red_left+0x9/0x30
[ 22.224407] ? complete_report_info+0x83/0x110
[ 22.226679] ? function_graph_enter_regs+0xa76/0xb90
[ 22.228084] kasan_report+0xce/0x100
[ 22.230109] ? function_graph_enter_regs+0xa76/0xb90
[ 22.232860] ? stack_trace_save+0x4/0xd0
[ 22.234156] function_graph_enter_regs+0xa76/0xb90
[ 22.236090] ? kasan_save_stack+0x30/0x50
[ 22.237752] ? __pfx_function_graph_enter_regs+0x10/0x10
[ 22.238694] ? ring_buffer_lock_reserve+0x345/0xf80
[ 22.239628] ? stack_trace_save+0x4/0xd0
[ 22.242121] ? stack_trace_save+0x4/0xd0
[ 22.243588] ftrace_graph_func+0xda/0x160
[ 22.245362] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.246520] 0xffffffffa0000095
[ 22.250528] ? stack_trace_save+0x9/0xd0
[ 22.251757] ? ring_buffer_unlock_commit+0x11d/0x5c0
[ 22.253152] stack_trace_save+0x9/0xd0
[ 22.254264] kasan_save_stack+0x30/0x50
[ 22.273631] kasan_save_track+0x14/0x30
[ 22.276763] kasan_save_free_info+0x3b/0x70
[ 22.278296] __kasan_slab_free+0x43/0x70
[ 22.280157] kmem_cache_free+0xbf/0x3b0
[ 22.282963] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.284001] free_task+0xa2/0x160
[ 22.285699] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.286752] copy_process+0x2aae/0x7bc0
Initialize the child function graph state immediately after
dup_task_struct(), before the first fallible operation.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ftrace: fork: Initialize function graph state before copy_exec_state()
dup_task_struct() copies the parent's task_struct, including ret_stack.
ftrace_graph_init_task() clears the copied function graph state, but it
currently runs after copy_exec_state().
For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state.
If that allocation fails, copy_process() reaches bad_fork_free and
free_task() calls ftrace_graph_exit_task(). Since the child still carries
the parent's ret_stack pointer, the unwind frees the parent's active
function graph return stack. The parent subsequently accesses freed memory
from function_graph_enter_regs().
KASAN reports:
[ 22.190920] ==================================================================
[ 22.195899] BUG: KASAN: slab-use-after-free in function_graph_enter_regs+0xa76/0xb90
[ 22.200747] Write of size 8 at addr ff110000054dc0a8 by task repro/1
[ 22.205134]
[ 22.210770] CPU: 0 UID: 0 PID: 1 Comm: repro Not tainted 7.2.0-07732-g9328b3b03bdc-dirty #3 PREEMPT(lazy)
[ 22.212576] Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
[ 22.213750] Call Trace:
[ 22.215271] <TASK>
[ 22.216242] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.217774] dump_stack_lvl+0x4e/0x70
[ 22.220531] print_report+0x157/0x4b4
[ 22.223202] ? fixup_red_left+0x9/0x30
[ 22.224407] ? complete_report_info+0x83/0x110
[ 22.226679] ? function_graph_enter_regs+0xa76/0xb90
[ 22.228084] kasan_report+0xce/0x100
[ 22.230109] ? function_graph_enter_regs+0xa76/0xb90
[ 22.232860] ? stack_trace_save+0x4/0xd0
[ 22.234156] function_graph_enter_regs+0xa76/0xb90
[ 22.236090] ? kasan_save_stack+0x30/0x50
[ 22.237752] ? __pfx_function_graph_enter_regs+0x10/0x10
[ 22.238694] ? ring_buffer_lock_reserve+0x345/0xf80
[ 22.239628] ? stack_trace_save+0x4/0xd0
[ 22.242121] ? stack_trace_save+0x4/0xd0
[ 22.243588] ftrace_graph_func+0xda/0x160
[ 22.245362] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.246520] 0xffffffffa0000095
[ 22.250528] ? stack_trace_save+0x9/0xd0
[ 22.251757] ? ring_buffer_unlock_commit+0x11d/0x5c0
[ 22.253152] stack_trace_save+0x9/0xd0
[ 22.254264] kasan_save_stack+0x30/0x50
[ 22.273631] kasan_save_track+0x14/0x30
[ 22.276763] kasan_save_free_info+0x3b/0x70
[ 22.278296] __kasan_slab_free+0x43/0x70
[ 22.280157] kmem_cache_free+0xbf/0x3b0
[ 22.282963] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.284001] free_task+0xa2/0x160
[ 22.285699] ? ftrace_stub_direct_tramp+0x10/0x10
[ 22.286752] copy_process+0x2aae/0x7bc0
Initialize the child function graph state immediately after
dup_task_struct(), before the first fallible operation.
π@cveNotify
π¨ CVE-2026-97938
In the Linux kernel, the following vulnerability has been resolved:
reboot: fix cad_pid use-after-free race
cad_pid is a single kernel-wide struct pid pointer. proc_do_cad_pid()
reads it and passes it to pid_vnr() without protecting the lifetime of
the referenced struct pid. A concurrent writer can replace cad_pid and
drop the final reference to the old struct pid after the reader has
loaded the pointer but before pid_vnr() has finished dereferencing it,
causing a use-after-free.
kill_cad_pid() has the same lifetime race when it passes cad_pid to
kill_pid().
At the time this issue was reported, an unprivileged user could reach the
sysctl through user and PID namespaces because cad_pid was registered in
pid_table[]. Moving cad_pid back to the global reboot sysctl table
corrected that namespace and permission mismatch, but did not fix the
underlying lifetime race.
Fix this by treating cad_pid as an RCU-protected pointer at both read
sites and by waiting for a grace period before dropping the old reference
on the write side.
call_rcu(&old_pid->rcu, ...) cannot be used here because free_pid()
also queues pid->rcu; queueing the same rcu_head twice can corrupt the
RCU callback list.
Original KASAN crash stack:
kernel/pid.c:545 pid_nr_ns() # reads freed pid->level
kernel/pid.c:556 pid_vnr() # calls pid_nr_ns()
kernel/pid.c:775 proc_do_cad_pid() # calls pid_vnr(cad_pid)
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
reboot: fix cad_pid use-after-free race
cad_pid is a single kernel-wide struct pid pointer. proc_do_cad_pid()
reads it and passes it to pid_vnr() without protecting the lifetime of
the referenced struct pid. A concurrent writer can replace cad_pid and
drop the final reference to the old struct pid after the reader has
loaded the pointer but before pid_vnr() has finished dereferencing it,
causing a use-after-free.
kill_cad_pid() has the same lifetime race when it passes cad_pid to
kill_pid().
At the time this issue was reported, an unprivileged user could reach the
sysctl through user and PID namespaces because cad_pid was registered in
pid_table[]. Moving cad_pid back to the global reboot sysctl table
corrected that namespace and permission mismatch, but did not fix the
underlying lifetime race.
Fix this by treating cad_pid as an RCU-protected pointer at both read
sites and by waiting for a grace period before dropping the old reference
on the write side.
call_rcu(&old_pid->rcu, ...) cannot be used here because free_pid()
also queues pid->rcu; queueing the same rcu_head twice can corrupt the
RCU callback list.
Original KASAN crash stack:
kernel/pid.c:545 pid_nr_ns() # reads freed pid->level
kernel/pid.c:556 pid_vnr() # calls pid_nr_ns()
kernel/pid.c:775 proc_do_cad_pid() # calls pid_vnr(cad_pid)
π@cveNotify
π¨ CVE-2026-97939
In the Linux kernel, the following vulnerability has been resolved:
ipmr: account multicast table and route memory
A netadmin in a user+net namespace can create many IPv4 and IPv6
multicast routing tables with MRT_TABLE and MRT6_TABLE. Each unseen
id allocates an mr_table via the shared mr_table_alloc(), links it
into the per-net list, and leaves it until netns teardown. Those
objects were not charged to memcg, so the host unreclaimable slab
grows with the table count.
Account mr_table allocations with GFP_KERNEL_ACCOUNT and mark the
IPv4/IPv6 MFC caches SLAB_ACCOUNT. This matches the established
handling of IP addresses, routes and alternate interface names.
Unresolved MFC entries are still allocated from softIRQ with
GFP_ATOMIC and are not charged. They expire after 10 seconds and are
bounded by the socket receive queue; see commit 0079ad8e8dc3
("ipmr: remove hard code cache_resolve_queue_len limit").
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ipmr: account multicast table and route memory
A netadmin in a user+net namespace can create many IPv4 and IPv6
multicast routing tables with MRT_TABLE and MRT6_TABLE. Each unseen
id allocates an mr_table via the shared mr_table_alloc(), links it
into the per-net list, and leaves it until netns teardown. Those
objects were not charged to memcg, so the host unreclaimable slab
grows with the table count.
Account mr_table allocations with GFP_KERNEL_ACCOUNT and mark the
IPv4/IPv6 MFC caches SLAB_ACCOUNT. This matches the established
handling of IP addresses, routes and alternate interface names.
Unresolved MFC entries are still allocated from softIRQ with
GFP_ATOMIC and are not charged. They expire after 10 seconds and are
bounded by the socket receive queue; see commit 0079ad8e8dc3
("ipmr: remove hard code cache_resolve_queue_len limit").
π@cveNotify
π¨ CVE-2026-97940
In the Linux kernel, the following vulnerability has been resolved:
ipv6: fix fib6 walker UAF on seq stop
ipv6_route_iter_active() treats a walker in FWS_U at the table root as
already unlinked. fib6_del_route() can move a still-linked walker into
that same state when the current leaf is the last route at the root,
so ipv6_route_native_seq_stop() skips fib6_walker_unlink(). The seq
private object can then be freed while it remains on
net->ipv6.fib6_walkers. A later route deletion walks the dangling list
and uses the freed walker.
Use the list head as membership state and reinitialize it when
unlinking. Keep the existing w->node check so a never-started iterator
with a zeroed private object is not treated as linked.
The same stop helper is used by /proc/net/ipv6_route and by the BPF
ipv6_route iterator. The BPF show path only widens the race.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
ipv6: fix fib6 walker UAF on seq stop
ipv6_route_iter_active() treats a walker in FWS_U at the table root as
already unlinked. fib6_del_route() can move a still-linked walker into
that same state when the current leaf is the last route at the root,
so ipv6_route_native_seq_stop() skips fib6_walker_unlink(). The seq
private object can then be freed while it remains on
net->ipv6.fib6_walkers. A later route deletion walks the dangling list
and uses the freed walker.
Use the list head as membership state and reinitialize it when
unlinking. Keep the existing w->node check so a never-started iterator
with a zeroed private object is not treated as linked.
The same stop helper is used by /proc/net/ipv6_route and by the BPF
ipv6_route iterator. The BPF show path only widens the race.
π@cveNotify
π¨ CVE-2026-97941
In the Linux kernel, the following vulnerability has been resolved:
mm/slab: take n->list_lock in __slab_try_return_freelist() to avoid race
Commit ba7425312607 ("mm, slab: add an optimistic
__slab_try_return_freelist()") incorrectly assumed that nobody has freed
an object to the slab as long as slab->freelist is NULL and cmpxchg
succeeds.
However, as reported by Hyunwoo Kim [1], other CPUs might have freed
an object to the slab, insert the slab to the partial list, then
allocated an object from the slab, and be in the middle of removing
the slab from the list under n->list_lock.
Since __refill_objects_node() puts the slab back on pc.slabs
outside n->list_lock, it might insert the slab into that list while
the slab is concurrently being removed from n->partial.
This led to a list corruption [1]:
list_add corruption. next->prev should be prev
(ffff888100000248), but was dead000000000122.
(next=ffffea000416e410).
kernel BUG at lib/list_debug.c:29!
Oops: invalid opcode: 0000 [#1] SMP NOPTI
CPU: 1 UID: 65534 PID: 144 Comm: poc Not tainted
7.2.0-16172-gcf72cbb39da8-dirty #1 PREEMPT(lazy)
RIP: 0010:__list_add_valid_or_report+0x80/0xd0
...
Call Trace:
alloc_from_new_slab+0x183/0x300
___slab_alloc+0x31c/0x890
__kmalloc_noprof+0x3d4/0x800
lsm_blob_alloc+0x2d/0x50
security_msg_msg_alloc+0x26/0x90
load_msg+0x1aa/0x210
do_msgsnd+0x91/0x800
do_syscall_64+0x109/0x5d0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
...
Kernel panic - not syncing: Fatal exception
This is a classic ABA problem where cmpxchg succeeds but the state has
changed since __refill_objects_node() took the freelist from the slab.
As Vlastimil Babka mentioned [2], it should be rare to return more than
one slab (due to the racy read of slab->counters in
get_partial_node_bulk()). Therefore, instead of introducing additional
complexity, acquire and release n->list_lock twice in the worst case.
Return the slab directly to the partial list and hold n->list_lock
across the cmpxchg and add_partial(). This is similar to the initial
version of commit ba7425312607 [3]. This is enough to avoid the race as
the list manipulation is serialized by n->list_lock. While at it,
bring back unlikely() hint now that the condition is unlikely.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
mm/slab: take n->list_lock in __slab_try_return_freelist() to avoid race
Commit ba7425312607 ("mm, slab: add an optimistic
__slab_try_return_freelist()") incorrectly assumed that nobody has freed
an object to the slab as long as slab->freelist is NULL and cmpxchg
succeeds.
However, as reported by Hyunwoo Kim [1], other CPUs might have freed
an object to the slab, insert the slab to the partial list, then
allocated an object from the slab, and be in the middle of removing
the slab from the list under n->list_lock.
Since __refill_objects_node() puts the slab back on pc.slabs
outside n->list_lock, it might insert the slab into that list while
the slab is concurrently being removed from n->partial.
This led to a list corruption [1]:
list_add corruption. next->prev should be prev
(ffff888100000248), but was dead000000000122.
(next=ffffea000416e410).
kernel BUG at lib/list_debug.c:29!
Oops: invalid opcode: 0000 [#1] SMP NOPTI
CPU: 1 UID: 65534 PID: 144 Comm: poc Not tainted
7.2.0-16172-gcf72cbb39da8-dirty #1 PREEMPT(lazy)
RIP: 0010:__list_add_valid_or_report+0x80/0xd0
...
Call Trace:
alloc_from_new_slab+0x183/0x300
___slab_alloc+0x31c/0x890
__kmalloc_noprof+0x3d4/0x800
lsm_blob_alloc+0x2d/0x50
security_msg_msg_alloc+0x26/0x90
load_msg+0x1aa/0x210
do_msgsnd+0x91/0x800
do_syscall_64+0x109/0x5d0
entry_SYSCALL_64_after_hwframe+0x77/0x7f
...
Kernel panic - not syncing: Fatal exception
This is a classic ABA problem where cmpxchg succeeds but the state has
changed since __refill_objects_node() took the freelist from the slab.
As Vlastimil Babka mentioned [2], it should be rare to return more than
one slab (due to the racy read of slab->counters in
get_partial_node_bulk()). Therefore, instead of introducing additional
complexity, acquire and release n->list_lock twice in the worst case.
Return the slab directly to the partial list and hold n->list_lock
across the cmpxchg and add_partial(). This is similar to the initial
version of commit ba7425312607 [3]. This is enough to avoid the race as
the list manipulation is serialized by n->list_lock. While at it,
bring back unlikely() hint now that the condition is unlikely.
π@cveNotify
π¨ CVE-2026-97942
In the Linux kernel, the following vulnerability has been resolved:
x86/alternatives: Exclude text poking against change_page_attr()
From time to time, the following BUG can be observed
in the x86 alternatives patching code [0]:
> kernel BUG at arch/x86/kernel/alternative.c:2576!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 355 Comm: (udev-worker) Not tainted 7.1.3-1-default #1 PREEMPT(full) openSUSE Tumbleweed 8c1795b03ec64f997e57a8ad38b1161e3b98da64
> Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS unknown 02/02/2022
> RIP: 0010:__text_poke+0x2aa/0x450
> Call Trace:
> <TASK>
> smp_text_poke_batch_finish+0x2a7/0x320
> __static_call_transform+0xb7/0x220
> arch_static_call_transform+0x5b/0xb0
> __static_call_init+0xe9/0x270
> static_call_module_notify+0x11f/0x150
> notifier_call_chain+0x61/0xe0
> blocking_notifier_call_chain_robust+0x63/0xc0
> load_module+0x1c92/0x20c0
> init_module_from_file+0xd8/0x140
> idempotent_init_module+0x100/0x2f0
> __x64_sys_finit_module+0x71/0xe0
> do_syscall_64+0xe1/0x610
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
which matches the following BUG_ON() in alternative.c:
/*
* If something went wrong, crash and burn since recovery paths are not
* implemented.
*/
BUG_ON(!pages[0] || (cross_page_boundary && !pages[1]));
This can happen if vmalloc_to_page() fails, for any reason. Such can happen
if text poking races with CPA, which can possibly result in the collapsing
of page tables (or breaking of PMD hugepages). It is not a problem for most
users of vmalloc_to_page() (they solely own the vmalloc'd range) but, when
CONFIG_ARCH_HAS_EXECMEM_ROX=y, various modules own a single execmem vmalloc
range, and can call set_memory_*() in parallel on it. This can happen to
race against __text_poke and cause havoc in vmalloc_to_page().
Fix it by excluding against CPA using the init_mm mmap read lock.
[ dhansen: Fix up SoB ordering. The actual code flow here was:
Pedro=>Lorenzo=>Mike=>Me which is reflected in the SoB chain
now. I *believe* Mike simply picked up Lorenzo's update to
Pedro's post from the Link ]
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
x86/alternatives: Exclude text poking against change_page_attr()
From time to time, the following BUG can be observed
in the x86 alternatives patching code [0]:
> kernel BUG at arch/x86/kernel/alternative.c:2576!
> Oops: invalid opcode: 0000 [#1] SMP NOPTI
> CPU: 0 UID: 0 PID: 355 Comm: (udev-worker) Not tainted 7.1.3-1-default #1 PREEMPT(full) openSUSE Tumbleweed 8c1795b03ec64f997e57a8ad38b1161e3b98da64
> Hardware name: QEMU Standard PC (i440FX + PIIX, 1996), BIOS unknown 02/02/2022
> RIP: 0010:__text_poke+0x2aa/0x450
> Call Trace:
> <TASK>
> smp_text_poke_batch_finish+0x2a7/0x320
> __static_call_transform+0xb7/0x220
> arch_static_call_transform+0x5b/0xb0
> __static_call_init+0xe9/0x270
> static_call_module_notify+0x11f/0x150
> notifier_call_chain+0x61/0xe0
> blocking_notifier_call_chain_robust+0x63/0xc0
> load_module+0x1c92/0x20c0
> init_module_from_file+0xd8/0x140
> idempotent_init_module+0x100/0x2f0
> __x64_sys_finit_module+0x71/0xe0
> do_syscall_64+0xe1/0x610
> entry_SYSCALL_64_after_hwframe+0x76/0x7e
which matches the following BUG_ON() in alternative.c:
/*
* If something went wrong, crash and burn since recovery paths are not
* implemented.
*/
BUG_ON(!pages[0] || (cross_page_boundary && !pages[1]));
This can happen if vmalloc_to_page() fails, for any reason. Such can happen
if text poking races with CPA, which can possibly result in the collapsing
of page tables (or breaking of PMD hugepages). It is not a problem for most
users of vmalloc_to_page() (they solely own the vmalloc'd range) but, when
CONFIG_ARCH_HAS_EXECMEM_ROX=y, various modules own a single execmem vmalloc
range, and can call set_memory_*() in parallel on it. This can happen to
race against __text_poke and cause havoc in vmalloc_to_page().
Fix it by excluding against CPA using the init_mm mmap read lock.
[ dhansen: Fix up SoB ordering. The actual code flow here was:
Pedro=>Lorenzo=>Mike=>Me which is reflected in the SoB chain
now. I *believe* Mike simply picked up Lorenzo's update to
Pedro's post from the Link ]
π@cveNotify
π¨ CVE-2026-97943
In the Linux kernel, the following vulnerability has been resolved:
x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
x86 implements page attribute modification using its Change Page
Attributes (CPA) mechanism.
This tracks properties of ranges such as cache mode through x86 page
attributes, and as part of that logic manipulates kernel page tables.
Since commit:
41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
ranges of kernel page table entries can be collapsed into
huge page table entries as part of this logic.
As part of this collapse, it frees the page tables which the collapsed
entries previously pointed to, and it does so without any relevant locks
being held to preclude concurrent kernel page table walkers.
The only way this code can be reached is if CPA_COLLAPSE is specified, and
this is only set in set_memory_rox() via:
set_memory_rox()
-> change_page_attr_set_clr()
-> cpa_flush()
-> cpa_collapse_large_pages()
Notable users of this are execmem and BPF when manipulating executable
mappings.
However, this is problematic for ptdump as it walks ranges it does not own
and thus runs the risk of a use-after-free on page tables freed underneath
it.
In addition, concurrent CPA collapse operations are possible which can also
cause races.
Resolve the issue by acquiring the mmap write lock on init_mm across the
whole operation.
It is safe to acquire a sleeping lock as all the callers invoke
set_memory_rox() from process context and in any case,
change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a
mutex, disallowing atomic context here.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF
x86 implements page attribute modification using its Change Page
Attributes (CPA) mechanism.
This tracks properties of ranges such as cache mode through x86 page
attributes, and as part of that logic manipulates kernel page tables.
Since commit:
41d88484c71c ("x86/mm/pat: restore large ROX pages after fragmentation")
ranges of kernel page table entries can be collapsed into
huge page table entries as part of this logic.
As part of this collapse, it frees the page tables which the collapsed
entries previously pointed to, and it does so without any relevant locks
being held to preclude concurrent kernel page table walkers.
The only way this code can be reached is if CPA_COLLAPSE is specified, and
this is only set in set_memory_rox() via:
set_memory_rox()
-> change_page_attr_set_clr()
-> cpa_flush()
-> cpa_collapse_large_pages()
Notable users of this are execmem and BPF when manipulating executable
mappings.
However, this is problematic for ptdump as it walks ranges it does not own
and thus runs the risk of a use-after-free on page tables freed underneath
it.
In addition, concurrent CPA collapse operations are possible which can also
cause races.
Resolve the issue by acquiring the mmap write lock on init_mm across the
whole operation.
It is safe to acquire a sleeping lock as all the callers invoke
set_memory_rox() from process context and in any case,
change_page_attr_set_clr() calls vm_unmap_alias() which ultimately takes a
mutex, disallowing atomic context here.
π@cveNotify
π¨ CVE-2026-97944
In the Linux kernel, the following vulnerability has been resolved:
x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash()
The switch of the FineIBT preamble from "subl $hash, %r10d" to the
shorter "subl $hash, %eax" moved the hash immediate from offset 7 to
offset 5 of the preamble. fineibt_preamble_hash was updated to match,
but the open-coded offset in cfi_get_func_hash() was missed and it
still reads the hash at offset 7.
cfi_get_func_hash() is used by the BPF JIT to give a struct_ops
trampoline the CFI hash of the stub function it stands in for. With
FineIBT the trampoline now gets the upper half of the real hash
followed by the first two bytes of the next instruction, so the first
indirect call from the kernel into a struct_ops program,
tcp_init_congestion_control() calling ->init() of a BPF congestion
control for example, fails the FineIBT check and the kernel dies with
a CFI failure.
Move the FineIBT preamble template and its offset defines above
cfi_get_func_hash() and use fineibt_preamble_hash there, so every
reader of the preamble shares one definition of its layout. The
CFI_FINEIBT arm is only built with CONFIG_FINEIBT, the only
configuration in which cfi_mode can take that value.
cfi_get_func_arity() does not need the same treatment: the __bhi_args
call whose displacement it reads still ends at the function address.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
x86/cfi: Fix FineIBT hash offset in cfi_get_func_hash()
The switch of the FineIBT preamble from "subl $hash, %r10d" to the
shorter "subl $hash, %eax" moved the hash immediate from offset 7 to
offset 5 of the preamble. fineibt_preamble_hash was updated to match,
but the open-coded offset in cfi_get_func_hash() was missed and it
still reads the hash at offset 7.
cfi_get_func_hash() is used by the BPF JIT to give a struct_ops
trampoline the CFI hash of the stub function it stands in for. With
FineIBT the trampoline now gets the upper half of the real hash
followed by the first two bytes of the next instruction, so the first
indirect call from the kernel into a struct_ops program,
tcp_init_congestion_control() calling ->init() of a BPF congestion
control for example, fails the FineIBT check and the kernel dies with
a CFI failure.
Move the FineIBT preamble template and its offset defines above
cfi_get_func_hash() and use fineibt_preamble_hash there, so every
reader of the preamble shares one definition of its layout. The
CFI_FINEIBT arm is only built with CONFIG_FINEIBT, the only
configuration in which cfi_mode can take that value.
cfi_get_func_arity() does not need the same treatment: the __bhi_args
call whose displacement it reads still ends at the function address.
π@cveNotify
π¨ CVE-2026-97945
In the Linux kernel, the following vulnerability has been resolved:
x86/mm: Fix user-space data loss with MADV_FREE and THP
Some of users of Polars (a data analytics library) have lost production
data from this bug. They seem to have just the right combination of
huge pages, MADV_FREE and heavy reclaim pressure.
pmd_modify() masks the old value with (_HPAGE_CHG_MASK & ~_PAGE_DIRTY),
silently discarding the hardware dirty bit. The subsequent
pmd_mksaveddirty() call is supposed to transfer _PAGE_DIRTY into
_PAGE_SAVED_DIRTY when write-protecting, but the dirty bit was already
stripped from the value, so there is nothing left to transfer.
Contrast with pte_modify(), which keeps _PAGE_DIRTY_BITS in its mask,
and pud_modify(), which keeps _HPAGE_CHG_MASK untouched: pmd_modify()
is the odd one out. Any pmd_modify() on a writable, dirty PMD loses
the dirty state.
One visible consequence is data loss with MADV_FREE on PMD-mapped THP:
memset(buf, 0x5A, size); // PMD-mapped THP, PMD dirty
madvise(buf, size, MADV_FREE); // PMD cleaned but left writable,
// folio marked lazyfree
memset(buf, 0x5A, size); // hardware sets _PAGE_DIRTY again
mprotect(buf, size, PROT_READ); // pmd_modify() drops the dirty bit
mprotect(buf, size, PROT_READ|PROT_WRITE);
// ... memory pressure ...
Reclaim (e.g. under memcg pressure) then finds the lazyfree folio with
no dirty bit set anywhere and frees it in
__discard_anon_folio_pmd_locked(), even though the data was rewritten
after MADV_FREE; subsequent reads fault in fresh zero pages. NUMA
hinting alone can trigger the same loss, as do_huge_pmd_numa_page()
restores the PMD through pmd_modify() as well.
PMD-mapped file THPs are affected too: mprotect()/NUMA hinting dropping
the dirty bit means rewritten data is never written back.
Fix it by keeping _PAGE_DIRTY in the preserved mask, exactly like
pte_modify() and pud_modify() do. The existing
pmd_mksaveddirty()/pmd_clear_saveddirty() pair then performs the
hardware-dirty <-> saved-dirty transition based on the write bit,
preserving the shadow-stack encoding rules.
π@cveNotify
In the Linux kernel, the following vulnerability has been resolved:
x86/mm: Fix user-space data loss with MADV_FREE and THP
Some of users of Polars (a data analytics library) have lost production
data from this bug. They seem to have just the right combination of
huge pages, MADV_FREE and heavy reclaim pressure.
pmd_modify() masks the old value with (_HPAGE_CHG_MASK & ~_PAGE_DIRTY),
silently discarding the hardware dirty bit. The subsequent
pmd_mksaveddirty() call is supposed to transfer _PAGE_DIRTY into
_PAGE_SAVED_DIRTY when write-protecting, but the dirty bit was already
stripped from the value, so there is nothing left to transfer.
Contrast with pte_modify(), which keeps _PAGE_DIRTY_BITS in its mask,
and pud_modify(), which keeps _HPAGE_CHG_MASK untouched: pmd_modify()
is the odd one out. Any pmd_modify() on a writable, dirty PMD loses
the dirty state.
One visible consequence is data loss with MADV_FREE on PMD-mapped THP:
memset(buf, 0x5A, size); // PMD-mapped THP, PMD dirty
madvise(buf, size, MADV_FREE); // PMD cleaned but left writable,
// folio marked lazyfree
memset(buf, 0x5A, size); // hardware sets _PAGE_DIRTY again
mprotect(buf, size, PROT_READ); // pmd_modify() drops the dirty bit
mprotect(buf, size, PROT_READ|PROT_WRITE);
// ... memory pressure ...
Reclaim (e.g. under memcg pressure) then finds the lazyfree folio with
no dirty bit set anywhere and frees it in
__discard_anon_folio_pmd_locked(), even though the data was rewritten
after MADV_FREE; subsequent reads fault in fresh zero pages. NUMA
hinting alone can trigger the same loss, as do_huge_pmd_numa_page()
restores the PMD through pmd_modify() as well.
PMD-mapped file THPs are affected too: mprotect()/NUMA hinting dropping
the dirty bit means rewritten data is never written back.
Fix it by keeping _PAGE_DIRTY in the preserved mask, exactly like
pte_modify() and pud_modify() do. The existing
pmd_mksaveddirty()/pmd_clear_saveddirty() pair then performs the
hardware-dirty <-> saved-dirty transition based on the write bit,
preserving the shadow-stack encoding rules.
π@cveNotify