** Description changed: SRU Justification: [ Impact ] Noble 6.8.0-136 backported 340cea84f691 ("cifs: open files should not hold ref on superblock", v7.0) via LP: #2154496. Since then an open cifs file pins only its dentry, so umount(2) can complete while a read-ahead (cifs_readdata on cifsiod_wq), an uncached read/write (cifs_aio_ctx) or a writeback (cifs_writedata) still owns a cifsFileInfo. generic_shutdown_super() finds the inode busy, poisons i_sb with VFS_PTR_POISON, and the later _cifsFileInfo_put() from cifs_readahead_complete() faults on 0xdead0000000000f5 + 0x390. With panic_on_oops=1 the host panics; with panic_on_oops=0 every later umount of a cifs filesystem hangs forever in cifs_kill_sb(). Workloads that mount and unmount SMB shares while a reader is killed (containers, per-job mounts) can hit it. 6.8.0-139 (LP: #2160250) added c68337442f03 ("cifs: Fix busy dentry used after unmounting"), which flushes deferredclose_wq only and covers the deferred-close variant (1 of our 9 production panics). The read-ahead variant (8 of 9) is fixed upstream by 75f5c412fa86 ("smb: client: fix busy dentry warning on unmount after DIO", v7.2, CVE-2026-72315), which is in no Noble 6.8 kernel up to 6.8.0-146 and does not apply as-is because the 6.8 cifs read/write path predates the netfs conversion. Observed: 6.8.0-137 (9 panics / 8 hosts / 30 h on ~850 hosts; lab reproducer panics in 5-90 s) and 6.8.0-146 (lab, ~10 s). Not affected: 6.8.0-111 (same workload, ~3,800 hosts, 0 in 17 days). [ Fix ] Backport of 75f5c412fa86 to the pre-netfs 6.8 code: a per-superblock counter (cifs_sb->outstanding_rreq, as upstream) is taken where a cifs_readdata / cifs_writedata / cifs_aio_ctx acquires its cifsFileInfo reference (including the two writeback sites that transfer an already-held reference) and released after the put in the three release functions. cifs_kill_sb() waits for the counter to reach zero, then flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(), exactly as upstream. A flush-only alternative (flush cifsiod_wq, serverclose_wq, fileinfo_put_wq) was built and tested and still panics: the read request is still on the socket when umount runs, so there is nothing queued to flush. [ Test Plan ] Mount an SMB3 share, start a sequential read of a 2-3 MB file (read-ahead queued), SIGKILL the reader after 1-90 ms, open/close another file, umount immediately; repeat in 4-8 parallel workers on separate mount points (script attached). Stock 6.8.0-137 panics within ~90 s at 8 workers; stock 6.8.0-146 within ~10 s. With the patch on 6.8.0-146.146: ~76,000 cycles against a Dell PowerScale (Isilon) share (krb5, ro) and a Samba share (ro and rw, buffered and O_DIRECT writers, 40-150 ms added server delay) with 0 "Dentry still in use" warnings, 0 faults, 0 hung umounts. We can test a -proposed kernel within minutes. [ Where problems could occur ] The change is confined to fs/smb/client. The counter must balance at every cifsFileInfo acquisition and release of cifs_readdata, cifs_writedata and cifs_aio_ctx; an unbalanced path would make umount(2) wait forever in cifs_kill_sb() (an earlier revision of this port missed the two writeback transfer sites; code review caught it before any write test, which is why they are counted explicitly and the write path was tested separately). The wait runs only at unmount, after the VFS has detached the superblock, so no new I/O can start on it; steady-state I/O paths gain one atomic increment and decrement per request. [ Other Info ] Related but separate: 5520e89a5a4f ("smb: client: fix cifsFileInfo reference leak in deferred close") fixes a refcount leak with a different Fixes: tag; it is not part of this bug. == Evidence (details) == Release: Ubuntu 24.04 LTS (Noble) Package: linux, tested at 6.8.0-137.137 and 6.8.0-146.146 Expected: umount(2) returns and the host continues running. Actual: _cifsFileInfo_put() dereferences an inode after superblock teardown and faults. The host panics with panic_on_oops=1, or later CIFS unmounts hang with panic_on_oops=0. Affected kernels: * Tested: 6.8.0-137.137. We saw 9 production panics on 8 hosts in about - 30 hours across about 850 nodes. The lab reproducer panics in 5–90 seconds. + 30 hours across about 850 nodes. The lab reproducer panics in 5–90 seconds. * Tested: 6.8.0-146.146 (noble-proposed). The lab reproducer panics in about - 10 seconds. + 10 seconds. * Code-inferred: 6.8.0-136.136, which introduced 340cea84f691c, and - 6.8.0-138.138, which predates c68337442f03. + 6.8.0-138.138, which predates c68337442f03. * Partial fix from 6.8.0-139.139: c68337442f03 flushes deferredclose_wq only. * Not affected: 6.8.0-111.111. The same workload ran on about 3,800 nodes - for 17 days without an occurrence. This kernel predates 340cea84f691c. + for 17 days without an occurrence. This kernel predates 340cea84f691c. == Summary == Noble 6.8.0-136 introduced upstream 340cea84f691c ("cifs: open files should not hold ref on superblock", mainline v7.0) via the upstream-stable patchset tracked by LP: #2154496. After this change, a cifs open file holds only a dentry reference. umount(2) can complete while asynchronous work still owns a cifsFileInfo. When that work later calls _cifsFileInfo_put(), the superblock is gone and the kernel faults on the VFS_PTR_POISON value written into the surviving inode. We observed these paths: * (a) deferredclose: smb2_deferred_work_close -> _cifsFileInfo_put. - Fixed by c68337442f03 in 6.8.0-139 and later. + Fixed by c68337442f03 in 6.8.0-139 and later. * (b) cifsiod: cifs_readahead_complete -> _cifsFileInfo_put. - Not fixed in any Noble 6.8 kernel through 6.8.0-146. + Not fixed in any Noble 6.8 kernel through 6.8.0-146. * (c) cifsoplockd: cifs_oplock_break -> _cifsFileInfo_put. - This secondary path appeared only after path (b) had already oopsed with - panic_on_oops=0. We have not demonstrated it as an independent race. + This secondary path appeared only after path (b) had already oopsed with + panic_on_oops=0. We have not demonstrated it as an independent race. In production on 6.8.0-137, path (b) caused 8 panics and path (a) caused 1. On 6.8.0-146 the reproducer triggers path (b), followed by path (c) when panic_on_oops=0. The proposed fix eliminated this complete reproduced sequence. We do not claim that it fixes a separate oplock race. == Impact == With panic_on_oops=1, the whole host panics. With panic_on_oops=0, the oopsed kworkers do not complete and later CIFS unmounts hang in D state at: - __flush_workqueue <- cifs_kill_sb + __flush_workqueue <- cifs_kill_sb Workloads that mount and unmount SMB shares for each job, such as container workloads, exercise this path continuously. == Kernel log, variant (b) — 6.8.0-137-generic #137-Ubuntu, production == [24042.343670] BUG: Dentry 000000002c909471{i=c34b9,n=<file>} still in use (1) [unmount of cifs cifs] [24042.343678] WARNING: CPU: 28 PID: 1352709 at fs/dcache.c:1528 umount_check+0x64/0x90 [24042.343780] CPU: 28 PID: 1352709 Comm: umount Kdump: loaded Tainted: P OE 6.8.0-137-generic #137-Ubuntu [24042.343783] RIP: 0010:umount_check+0x64/0x90 [24042.343802] d_walk+0xc0/0x2a0 [24042.343810] generic_shutdown_super+0x21/0x180 [24042.343815] cifs_kill_sb+0x5b/0x70 [cifs] [24042.343853] cleanup_mnt+0xc3/0x170 [24042.343938] WARNING: CPU: 28 PID: 1352709 at fs/super.c:649 generic_shutdown_super+0x120/0x180 - VFS: Busy inodes after unmount of cifs (cifs) + VFS: Busy inodes after unmount of cifs (cifs) [24043.843507] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#1] PREEMPT SMP NOPTI [24043.854484] CPU: 4 PID: 1352168 Comm: kworker/4:1 Kdump: loaded Tainted: P W OE 6.8.0-137-generic #137-Ubuntu [24043.875232] Workqueue: cifsiod cifs_readahead_complete [cifs] [24043.881106] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] [24043.910734] RAX: dead0000000000f5 RBX: ffff8e6ee73a1ea8 RCX: 000000000000000a [24043.983676] cifs_readahead_complete+0x23e/0x2f0 [cifs] [24043.988976] process_one_work+0x181/0x3a0 [24043.993014] worker_thread+0x18b/0x330 [24044.001087] kthread+0xef/0x120 [24044.245052] Kernel panic - not syncing: Fatal exception == Kernel log, variant (a) — 6.8.0-137, production == [17180.976882] BUG: Dentry 00000000e23b32d6{i=34fc,n=<file>} still in use (1) [unmount of cifs cifs] [17180.977008] RIP: 0010:umount_check+0x64/0x90 [17180.977052] cifs_kill_sb+0x5b/0x70 [cifs] [17180.977268] VFS: Busy inodes after unmount of cifs (cifs) [17181.900781] Workqueue: deferredclose smb2_deferred_work_close [cifs] [17181.907243] RIP: 0010:_raw_spin_lock+0x13/0x60 [17182.014078] cifsFileInfo_put_final+0xed/0x120 [cifs] [17182.019221] _cifsFileInfo_put+0x350/0x4a0 [cifs] [17182.028127] smb2_deferred_work_close+0x5f/0x70 [cifs] - Kernel panic - not syncing: Fatal exception + Kernel panic - not syncing: Fatal exception == Kernel log, variants (b) and (c) == Kernel: 6.8.0-146-generic #146-Ubuntu, lab, panic_on_oops=0 [ 142.574284] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#1] PREEMPT SMP NOPTI [ 142.605742] Workqueue: cifsiod cifs_readahead_complete [cifs] [ 142.611633] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] [ 142.857981] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#2] PREEMPT SMP NOPTI [ 142.868925] Workqueue: cifsoplockd cifs_oplock_break [cifs] [ 142.868997] RIP: 0010:cifs_oplock_break+0x43/0x620 [cifs] [ 142.888932] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] (preceded by "BUG: Dentry ... still in use (1) [unmount of cifs cifs]" from umount_check) Afterwards on 6.8.0-146: 8 umount processes in D state, all at - __flush_workqueue+0x14a/0x3e0 - <- cifs_kill_sb+0x3c/0x70 [cifs] - <- deactivate_locked_super - <- cleanup_mnt + __flush_workqueue+0x14a/0x3e0 + <- cifs_kill_sb+0x3c/0x70 [cifs] + <- deactivate_locked_super + <- cleanup_mnt == vmcore analysis (6.8.0-146.146, variant b) == Analysis used crash with linux-image-unsigned-6.8.0-146-generic-dbgsym. The faulting instruction at _cifsFileInfo_put+0x77 is the inlined CIFS_SB(inode->i_sb), which reads sb->s_fs_info at offset 0x390. The inode is d_inode(cifs_file->dentry). Objects in the dump: * inode ffff8bb592233680: - i_ino=0xec6, matching the "i=ec6" umount_check line; i_nlink=1; - i_count=1; i_state=0; still on sb->s_inodes; and - i_op = i_sb = i_mapping = 0xdead0000000000f5 (VFS_PTR_POISON). + i_ino=0xec6, matching the "i=ec6" umount_check line; i_nlink=1; + i_count=1; i_state=0; still on sb->s_inodes; and + i_op = i_sb = i_mapping = 0xdead0000000000f5 (VFS_PTR_POISON). * dentry ffff8bb5861c7380: - "<file>"; d_lockref.count=1, held by the cifsFileInfo. + "<file>"; d_lockref.count=1, held by the cifsFileInfo. * superblock ffff8ab5ae3e9800: - type "cifs"; s_count=0; s_active=0; s_root=NULL. + type "cifs"; s_count=0; s_active=0; s_root=NULL. * cifsFileInfo ffff8bb551338a00: - allocated from kmalloc-512. + allocated from kmalloc-512. * cifs_tcon ffff8ab51c791800: - allocated. + allocated. generic_shutdown_super() writes this VFS_PTR_POISON value when CHECK_DATA_CORRUPTION(!list_empty(&sb->s_inodes), "VFS: Busy inodes after unmount") fires at fs/super.c:649-663. The read-ahead cifsFileInfo kept the dentry and inode alive across unmount. The superblock was torn down, and the completion work then dereferenced inode->i_sb. This matches the mechanism addressed by 75f5c412fa86: wait for in-flight requests and drain final-put work before kill_anon_super(). The deferredclose_wq flush from c68337442f03 does not cover cifsiod or cifsoplockd work. The vmcore and vmlinux are available on request. == Reproducer (lab, both 6.8.0-137 and 6.8.0-146) == Server: Dell PowerScale (Isilon) SMB3 share, mounted read-only with - -o vers=3.0,sec=krb5,dir_mode=0755,file_mode=0755,noperm, - noserverino,nosharesock,cruid=0,nobrl,ro + -o vers=3.0,sec=krb5,dir_mode=0755,file_mode=0755,noperm, + noserverino,nosharesock,cruid=0,nobrl,ro mount.cifs reports: - cache=strict,soft,nounix,mapposix,rsize=1048576,wsize=1048576 + cache=strict,soft,nounix,mapposix,rsize=1048576,wsize=1048576 Loop, N parallel workers, each on its own mount point: 1. Mount the share. 2. Start a sequential read, which queues read-ahead: - dd if=<2–3 MB file on the share> of=/dev/null bs=1M & + dd if=<2–3 MB file on the share> of=/dev/null bs=1M & 3. Sleep for 0.01–0.09 seconds, then kill the dd process while I/O is in - flight. + flight. 4. Open and close another file to leave a deferred close pending: - head -c 65536 <another file> >/dev/null + head -c 65536 <another file> >/dev/null 5. Immediately unmount the mount point. 6. Repeat. Results: * 6.8.0-137, 8 workers: panic within about 90 seconds (variant b), with a - kdump fingerprint on the BMC. + kdump fingerprint on the BMC. * 6.8.0-137, 4 workers, panic_on_oops=0: oops (b), then (c), within about - 5 seconds. + 5 seconds. * 6.8.0-146, 4 workers for 45 seconds: 5 "Dentry ... still in use" - warnings and no fault. + warnings and no fault. * 6.8.0-146, 8 workers: oops (b), then (c), within about 10 seconds. The "still in use" warning occurs several times per minute with 4 workers on 6.8.0-146. The fault requires the deferred work to run after the superblock is freed, so it becomes more likely with parallel workers. == Regression boundary (from the Noble changelog) == * linux 6.8.0-136.136, "Noble update: upstream stable patchset 2026-05-28" - (LP: #2154496), added: - - 340cea84f691c (v7.0) - cifs: open files should not hold ref on superblock + (LP: #2154496), added: + + 340cea84f691c (v7.0) + cifs: open files should not hold ref on superblock * linux 6.8.0-139.139, "Noble update: upstream stable patchset 2026-07-09" - (LP: #2160250), added: - - c68337442f03 (v7.1, Cc: stable, Fixes: 340cea84f691c) - cifs: Fix busy dentry used after unmounting - - This commit adds flush_workqueue(deferredclose_wq) in cifs_kill_sb(). It - covers variant (a) only. + (LP: #2160250), added: + + c68337442f03 (v7.1, Cc: stable, Fixes: 340cea84f691c) + cifs: Fix busy dentry used after unmounting + + This commit adds flush_workqueue(deferredclose_wq) in cifs_kill_sb(). It + covers variant (a) only. * No Noble 6.8 kernel through 6.8.0-146.146 contains: - 75f5c412fa86 (v7.2, Fixes: 340cea84f691c) - smb: client: fix busy dentry warning on unmount after DIO - - This commit adds cifs_sb->outstanding_rreq, waits for in-flight requests, - and flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(). - This is the mechanism that covers variants (b) and (c). + 75f5c412fa86 (v7.2, Fixes: 340cea84f691c) + smb: client: fix busy dentry warning on unmount after DIO + + This commit adds cifs_sb->outstanding_rreq, waits for in-flight requests, + and flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(). + This is the mechanism that covers variants (b) and (c). == Request == Track this under CVE-2026-72315, the outstanding-I/O defect fixed by 75f5c412fa86. 1. Apply the attached Noble 6.8 backport of 75f5c412fa86. Noble predates - the cifs netfs conversion, so the backport accounts for the cifsFileInfo - references held by cifs_readdata, cifs_writedata and cifs_aio_ctx instead - of netfs requests. cifs_kill_sb() waits for that per-superblock count to - reach zero, then flushes serverclose_wq and fileinfo_put_wq before - kill_anon_super(). The patch survived about 76,000 tested cycles. The - flush-only alternative still panicked within about 30 seconds with the - same cifsiod/cifs_readahead_complete trace. + the cifs netfs conversion, so the backport accounts for the cifsFileInfo + references held by cifs_readdata, cifs_writedata and cifs_aio_ctx instead + of netfs requests. cifs_kill_sb() waits for that per-superblock count to + reach zero, then flushes serverclose_wq and fileinfo_put_wq before + kill_anon_super(). The patch survived about 76,000 tested cycles. The + flush-only alternative still panicked within about 30 seconds with the + same cifsiod/cifs_readahead_complete trace. 2. We can validate a candidate kernel within minutes with the attached - reproducer. + reproducer. 3. Consider noting in the 6.8.0-136, 6.8.0-137 and 6.8.0-138 release notes - that 340cea84f691c shipped without its stable follow-up c68337442f03. + that 340cea84f691c shipped without its stable follow-up c68337442f03. Related but separate: 5520e89a5a4f ("smb: client: fix cifsFileInfo reference leak in deferred close", Fixes: c3f207ab29f7) fixes a refcount leak when queue_delayed_work() finds work already pending. It has a different Fixes commit and is not asserted to cause this panic. Please track it separately. == Environment / attachments == Ubuntu 24.04 (Noble), x86_64, HPE ProLiant XL225n Gen10 Plus, in-tree cifs.ko 2.47, and SMB3 to Dell PowerScale. Production mounts and unmounts the share for each cri-o container lifecycle. The lab uses a hand-driven loop. - Attached evidence: - - * Production panic context and kworker trace. - * Firmware pstore record from a lab 6.8.0-137 panic. - * Live and vmcore dmesg from stock 6.8.0-146. - * Vmcore analysis. - * Reproducer. - * Version, PCI, release and package metadata. - * Tested SRU patch. /proc/version_signature reported: - Ubuntu 6.8.0-137.137-generic 6.8.12 - Ubuntu 6.8.0-146.146-generic 6.8.12 + Ubuntu 6.8.0-137.137-generic 6.8.12 + Ubuntu 6.8.0-146.146-generic 6.8.12 A 3.4 GB kdump vmcore of the 6.8.0-146.146 variant-(b) panic was captured on 2026-09-27 with makedumpfile -c -d 31. It is available on request with - the matching vmlinux. Its dmesg and object analysis are attached. + the matching vmlinux. Its dmesg is attached. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168697 Title: linux (Noble): CIFS unmount use-after-free and panic introduced in 6.8.0-136 (CVE-2026-72315) Status in linux package in Ubuntu: New Bug description: SRU Justification: [ Impact ] Noble 6.8.0-136 backported 340cea84f691 ("cifs: open files should not hold ref on superblock", v7.0) via LP: #2154496. Since then an open cifs file pins only its dentry, so umount(2) can complete while a read-ahead (cifs_readdata on cifsiod_wq), an uncached read/write (cifs_aio_ctx) or a writeback (cifs_writedata) still owns a cifsFileInfo. generic_shutdown_super() finds the inode busy, poisons i_sb with VFS_PTR_POISON, and the later _cifsFileInfo_put() from cifs_readahead_complete() faults on 0xdead0000000000f5 + 0x390. With panic_on_oops=1 the host panics; with panic_on_oops=0 every later umount of a cifs filesystem hangs forever in cifs_kill_sb(). Workloads that mount and unmount SMB shares while a reader is killed (containers, per-job mounts) can hit it. 6.8.0-139 (LP: #2160250) added c68337442f03 ("cifs: Fix busy dentry used after unmounting"), which flushes deferredclose_wq only and covers the deferred-close variant (1 of our 9 production panics). The read-ahead variant (8 of 9) is fixed upstream by 75f5c412fa86 ("smb: client: fix busy dentry warning on unmount after DIO", v7.2, CVE-2026-72315), which is in no Noble 6.8 kernel up to 6.8.0-146 and does not apply as-is because the 6.8 cifs read/write path predates the netfs conversion. Observed: 6.8.0-137 (9 panics / 8 hosts / 30 h on ~850 hosts; lab reproducer panics in 5-90 s) and 6.8.0-146 (lab, ~10 s). Not affected: 6.8.0-111 (same workload, ~3,800 hosts, 0 in 17 days). [ Fix ] Backport of 75f5c412fa86 to the pre-netfs 6.8 code: a per-superblock counter (cifs_sb->outstanding_rreq, as upstream) is taken where a cifs_readdata / cifs_writedata / cifs_aio_ctx acquires its cifsFileInfo reference (including the two writeback sites that transfer an already-held reference) and released after the put in the three release functions. cifs_kill_sb() waits for the counter to reach zero, then flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(), exactly as upstream. A flush-only alternative (flush cifsiod_wq, serverclose_wq, fileinfo_put_wq) was built and tested and still panics: the read request is still on the socket when umount runs, so there is nothing queued to flush. [ Test Plan ] Mount an SMB3 share, start a sequential read of a 2-3 MB file (read-ahead queued), SIGKILL the reader after 1-90 ms, open/close another file, umount immediately; repeat in 4-8 parallel workers on separate mount points (script attached). Stock 6.8.0-137 panics within ~90 s at 8 workers; stock 6.8.0-146 within ~10 s. With the patch on 6.8.0-146.146: ~76,000 cycles against a Dell PowerScale (Isilon) share (krb5, ro) and a Samba share (ro and rw, buffered and O_DIRECT writers, 40-150 ms added server delay) with 0 "Dentry still in use" warnings, 0 faults, 0 hung umounts. We can test a -proposed kernel within minutes. [ Where problems could occur ] The change is confined to fs/smb/client. The counter must balance at every cifsFileInfo acquisition and release of cifs_readdata, cifs_writedata and cifs_aio_ctx; an unbalanced path would make umount(2) wait forever in cifs_kill_sb() (an earlier revision of this port missed the two writeback transfer sites; code review caught it before any write test, which is why they are counted explicitly and the write path was tested separately). The wait runs only at unmount, after the VFS has detached the superblock, so no new I/O can start on it; steady-state I/O paths gain one atomic increment and decrement per request. [ Other Info ] Related but separate: 5520e89a5a4f ("smb: client: fix cifsFileInfo reference leak in deferred close") fixes a refcount leak with a different Fixes: tag; it is not part of this bug. == Evidence (details) == Release: Ubuntu 24.04 LTS (Noble) Package: linux, tested at 6.8.0-137.137 and 6.8.0-146.146 Expected: umount(2) returns and the host continues running. Actual: _cifsFileInfo_put() dereferences an inode after superblock teardown and faults. The host panics with panic_on_oops=1, or later CIFS unmounts hang with panic_on_oops=0. Affected kernels: * Tested: 6.8.0-137.137. We saw 9 production panics on 8 hosts in about 30 hours across about 850 nodes. The lab reproducer panics in 5–90 seconds. * Tested: 6.8.0-146.146 (noble-proposed). The lab reproducer panics in about 10 seconds. * Code-inferred: 6.8.0-136.136, which introduced 340cea84f691c, and 6.8.0-138.138, which predates c68337442f03. * Partial fix from 6.8.0-139.139: c68337442f03 flushes deferredclose_wq only. * Not affected: 6.8.0-111.111. The same workload ran on about 3,800 nodes for 17 days without an occurrence. This kernel predates 340cea84f691c. == Summary == Noble 6.8.0-136 introduced upstream 340cea84f691c ("cifs: open files should not hold ref on superblock", mainline v7.0) via the upstream-stable patchset tracked by LP: #2154496. After this change, a cifs open file holds only a dentry reference. umount(2) can complete while asynchronous work still owns a cifsFileInfo. When that work later calls _cifsFileInfo_put(), the superblock is gone and the kernel faults on the VFS_PTR_POISON value written into the surviving inode. We observed these paths: * (a) deferredclose: smb2_deferred_work_close -> _cifsFileInfo_put. Fixed by c68337442f03 in 6.8.0-139 and later. * (b) cifsiod: cifs_readahead_complete -> _cifsFileInfo_put. Not fixed in any Noble 6.8 kernel through 6.8.0-146. * (c) cifsoplockd: cifs_oplock_break -> _cifsFileInfo_put. This secondary path appeared only after path (b) had already oopsed with panic_on_oops=0. We have not demonstrated it as an independent race. In production on 6.8.0-137, path (b) caused 8 panics and path (a) caused 1. On 6.8.0-146 the reproducer triggers path (b), followed by path (c) when panic_on_oops=0. The proposed fix eliminated this complete reproduced sequence. We do not claim that it fixes a separate oplock race. == Impact == With panic_on_oops=1, the whole host panics. With panic_on_oops=0, the oopsed kworkers do not complete and later CIFS unmounts hang in D state at: __flush_workqueue <- cifs_kill_sb Workloads that mount and unmount SMB shares for each job, such as container workloads, exercise this path continuously. == Kernel log, variant (b) — 6.8.0-137-generic #137-Ubuntu, production == [24042.343670] BUG: Dentry 000000002c909471{i=c34b9,n=<file>} still in use (1) [unmount of cifs cifs] [24042.343678] WARNING: CPU: 28 PID: 1352709 at fs/dcache.c:1528 umount_check+0x64/0x90 [24042.343780] CPU: 28 PID: 1352709 Comm: umount Kdump: loaded Tainted: P OE 6.8.0-137-generic #137-Ubuntu [24042.343783] RIP: 0010:umount_check+0x64/0x90 [24042.343802] d_walk+0xc0/0x2a0 [24042.343810] generic_shutdown_super+0x21/0x180 [24042.343815] cifs_kill_sb+0x5b/0x70 [cifs] [24042.343853] cleanup_mnt+0xc3/0x170 [24042.343938] WARNING: CPU: 28 PID: 1352709 at fs/super.c:649 generic_shutdown_super+0x120/0x180 VFS: Busy inodes after unmount of cifs (cifs) [24043.843507] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#1] PREEMPT SMP NOPTI [24043.854484] CPU: 4 PID: 1352168 Comm: kworker/4:1 Kdump: loaded Tainted: P W OE 6.8.0-137-generic #137-Ubuntu [24043.875232] Workqueue: cifsiod cifs_readahead_complete [cifs] [24043.881106] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] [24043.910734] RAX: dead0000000000f5 RBX: ffff8e6ee73a1ea8 RCX: 000000000000000a [24043.983676] cifs_readahead_complete+0x23e/0x2f0 [cifs] [24043.988976] process_one_work+0x181/0x3a0 [24043.993014] worker_thread+0x18b/0x330 [24044.001087] kthread+0xef/0x120 [24044.245052] Kernel panic - not syncing: Fatal exception == Kernel log, variant (a) — 6.8.0-137, production == [17180.976882] BUG: Dentry 00000000e23b32d6{i=34fc,n=<file>} still in use (1) [unmount of cifs cifs] [17180.977008] RIP: 0010:umount_check+0x64/0x90 [17180.977052] cifs_kill_sb+0x5b/0x70 [cifs] [17180.977268] VFS: Busy inodes after unmount of cifs (cifs) [17181.900781] Workqueue: deferredclose smb2_deferred_work_close [cifs] [17181.907243] RIP: 0010:_raw_spin_lock+0x13/0x60 [17182.014078] cifsFileInfo_put_final+0xed/0x120 [cifs] [17182.019221] _cifsFileInfo_put+0x350/0x4a0 [cifs] [17182.028127] smb2_deferred_work_close+0x5f/0x70 [cifs] Kernel panic - not syncing: Fatal exception == Kernel log, variants (b) and (c) == Kernel: 6.8.0-146-generic #146-Ubuntu, lab, panic_on_oops=0 [ 142.574284] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#1] PREEMPT SMP NOPTI [ 142.605742] Workqueue: cifsiod cifs_readahead_complete [cifs] [ 142.611633] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] [ 142.857981] general protection fault, probably for non-canonical address 0xdead000000000485: 0000 [#2] PREEMPT SMP NOPTI [ 142.868925] Workqueue: cifsoplockd cifs_oplock_break [cifs] [ 142.868997] RIP: 0010:cifs_oplock_break+0x43/0x620 [cifs] [ 142.888932] RIP: 0010:_cifsFileInfo_put+0x77/0x4a0 [cifs] (preceded by "BUG: Dentry ... still in use (1) [unmount of cifs cifs]" from umount_check) Afterwards on 6.8.0-146: 8 umount processes in D state, all at __flush_workqueue+0x14a/0x3e0 <- cifs_kill_sb+0x3c/0x70 [cifs] <- deactivate_locked_super <- cleanup_mnt == vmcore analysis (6.8.0-146.146, variant b) == Analysis used crash with linux-image-unsigned-6.8.0-146-generic-dbgsym. The faulting instruction at _cifsFileInfo_put+0x77 is the inlined CIFS_SB(inode->i_sb), which reads sb->s_fs_info at offset 0x390. The inode is d_inode(cifs_file->dentry). Objects in the dump: * inode ffff8bb592233680: i_ino=0xec6, matching the "i=ec6" umount_check line; i_nlink=1; i_count=1; i_state=0; still on sb->s_inodes; and i_op = i_sb = i_mapping = 0xdead0000000000f5 (VFS_PTR_POISON). * dentry ffff8bb5861c7380: "<file>"; d_lockref.count=1, held by the cifsFileInfo. * superblock ffff8ab5ae3e9800: type "cifs"; s_count=0; s_active=0; s_root=NULL. * cifsFileInfo ffff8bb551338a00: allocated from kmalloc-512. * cifs_tcon ffff8ab51c791800: allocated. generic_shutdown_super() writes this VFS_PTR_POISON value when CHECK_DATA_CORRUPTION(!list_empty(&sb->s_inodes), "VFS: Busy inodes after unmount") fires at fs/super.c:649-663. The read-ahead cifsFileInfo kept the dentry and inode alive across unmount. The superblock was torn down, and the completion work then dereferenced inode->i_sb. This matches the mechanism addressed by 75f5c412fa86: wait for in-flight requests and drain final-put work before kill_anon_super(). The deferredclose_wq flush from c68337442f03 does not cover cifsiod or cifsoplockd work. The vmcore and vmlinux are available on request. == Reproducer (lab, both 6.8.0-137 and 6.8.0-146) == Server: Dell PowerScale (Isilon) SMB3 share, mounted read-only with -o vers=3.0,sec=krb5,dir_mode=0755,file_mode=0755,noperm, noserverino,nosharesock,cruid=0,nobrl,ro mount.cifs reports: cache=strict,soft,nounix,mapposix,rsize=1048576,wsize=1048576 Loop, N parallel workers, each on its own mount point: 1. Mount the share. 2. Start a sequential read, which queues read-ahead: dd if=<2–3 MB file on the share> of=/dev/null bs=1M & 3. Sleep for 0.01–0.09 seconds, then kill the dd process while I/O is in flight. 4. Open and close another file to leave a deferred close pending: head -c 65536 <another file> >/dev/null 5. Immediately unmount the mount point. 6. Repeat. Results: * 6.8.0-137, 8 workers: panic within about 90 seconds (variant b), with a kdump fingerprint on the BMC. * 6.8.0-137, 4 workers, panic_on_oops=0: oops (b), then (c), within about 5 seconds. * 6.8.0-146, 4 workers for 45 seconds: 5 "Dentry ... still in use" warnings and no fault. * 6.8.0-146, 8 workers: oops (b), then (c), within about 10 seconds. The "still in use" warning occurs several times per minute with 4 workers on 6.8.0-146. The fault requires the deferred work to run after the superblock is freed, so it becomes more likely with parallel workers. == Regression boundary (from the Noble changelog) == * linux 6.8.0-136.136, "Noble update: upstream stable patchset 2026-05-28" (LP: #2154496), added: 340cea84f691c (v7.0) cifs: open files should not hold ref on superblock * linux 6.8.0-139.139, "Noble update: upstream stable patchset 2026-07-09" (LP: #2160250), added: c68337442f03 (v7.1, Cc: stable, Fixes: 340cea84f691c) cifs: Fix busy dentry used after unmounting This commit adds flush_workqueue(deferredclose_wq) in cifs_kill_sb(). It covers variant (a) only. * No Noble 6.8 kernel through 6.8.0-146.146 contains: 75f5c412fa86 (v7.2, Fixes: 340cea84f691c) smb: client: fix busy dentry warning on unmount after DIO This commit adds cifs_sb->outstanding_rreq, waits for in-flight requests, and flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(). This is the mechanism that covers variants (b) and (c). == Request == Track this under CVE-2026-72315, the outstanding-I/O defect fixed by 75f5c412fa86. 1. Apply the attached Noble 6.8 backport of 75f5c412fa86. Noble predates the cifs netfs conversion, so the backport accounts for the cifsFileInfo references held by cifs_readdata, cifs_writedata and cifs_aio_ctx instead of netfs requests. cifs_kill_sb() waits for that per-superblock count to reach zero, then flushes serverclose_wq and fileinfo_put_wq before kill_anon_super(). The patch survived about 76,000 tested cycles. The flush-only alternative still panicked within about 30 seconds with the same cifsiod/cifs_readahead_complete trace. 2. We can validate a candidate kernel within minutes with the attached reproducer. 3. Consider noting in the 6.8.0-136, 6.8.0-137 and 6.8.0-138 release notes that 340cea84f691c shipped without its stable follow-up c68337442f03. Related but separate: 5520e89a5a4f ("smb: client: fix cifsFileInfo reference leak in deferred close", Fixes: c3f207ab29f7) fixes a refcount leak when queue_delayed_work() finds work already pending. It has a different Fixes commit and is not asserted to cause this panic. Please track it separately. == Environment / attachments == Ubuntu 24.04 (Noble), x86_64, HPE ProLiant XL225n Gen10 Plus, in-tree cifs.ko 2.47, and SMB3 to Dell PowerScale. Production mounts and unmounts the share for each cri-o container lifecycle. The lab uses a hand-driven loop. /proc/version_signature reported: Ubuntu 6.8.0-137.137-generic 6.8.12 Ubuntu 6.8.0-146.146-generic 6.8.12 A 3.4 GB kdump vmcore of the 6.8.0-146.146 variant-(b) panic was captured on 2026-09-27 with makedumpfile -c -d 31. It is available on request with the matching vmlinux. Its dmesg is attached. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168697/+subscriptions
Комментариев нет:
Отправить комментарий