суббота

[Bug 2148534] Re: [Ubuntu 26.04] Failed install OS onto JBOD disk on B540d-2HS M.2 controller

I also made a kernel image base on Ko boon Lin's comments #10, https://ratatoskr.run/linux-scsi/2026/03/6765435 https://drive.google.com/file/d/11jqds5MYkHfl0faO1FlfhJYNVfFf8Iqk/view?usp=drive_link, Yuri Zhang Could you like try both of them kernel. In fact, both of fix looks good. but need a real test to confirm. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2148534 Title: [Ubuntu 26.04] Failed install OS onto JBOD disk on B540d-2HS M.2 controller Status in linux package in Ubuntu: Confirmed Status in linux source package in Resolute: Confirmed Bug description: The system hangs during storage initialization when using a B540 RAID Kit with a single Micron 7450 Pro NVMe SSD configured in JBOD mode. Environment: Hardware: Lenovo ThinkSystem SR650 V4 Controller: B540d-2HS M.2 controller Disk: ThinkSystem M.2 7450 PRO 960GB Read Intensive NVMe PCIe 4.0 x4 NHS SSD Kernel (broken): 7.0.0-13-generic (Ubuntu) Kernel (working): 6.8.x (Ubuntu 24.04 LTS) Reproduce steps: 1. Set the B540 RAID controller to allow JBOD or configure the Micron 7450P as a JBOD drive. 2. Start Ubuntu 26.04 installation. 3. The installer/kernel attempts to scan scsi0. 4. The driver hangs at "waiting for commands" and eventually triggers a firmware reset loop. The megaraid_sas driver reports multiple command timeouts followed by a controller firmware crash. Even though the logs show Secure JBOD support-: Yes and NVMe passthru support-: Yes, It seems like kernel oops in megasas_make_prp_nvme due to a page fault on an unmapped virtual address while building the NVMe, which can be refered in logs below with Call Trace. [ 271.793726] #PF: supervisor write access in kernel mode [ 271.794694] #PF: error_code(0x0002) - not-present page [ 271.795648] PGD 100000067 P4D 101b42067 PUD 101b43067 PMD 10219a067 PTE 0 [ 271.796887] Oops: Oops: 0002 [#1] SMP NOPTI [ 271.797675] CPU: 1 UID: 0 PID: 1905 Comm: kworker/u1025:3 Tainted: P S O 7.0.0-13-generic #13-Ubuntu PREEMPT(lazy) [ 271.799778] Tainted: [P]=PROPRIETARY_MODULE, [S]=CPU_OUT_OF_SPEC, [O]=OOT_MODULE [ 271.801123] Hardware name: Lenovo ThinkSystem SR650 V4/SB27B70076, BIOS IHE111A-1.21 03/28/2025 [ 271.801140] sd 0:0:2:0: [sda] tag#244 page boundary ptr_sgl: 0x0000000040507511 [ 271.802682] Workqueue: writeback wb_workfn [ 271.802733] sd 0:0:2:0: [sda] tag#245 page boundary ptr_sgl: 0x000000004c4d3310 [ 271.806138] (flush-8:0) [ 271.806643] RIP: 0010:megasas_make_prp_nvme.isra.0+0x12f/0x220 [megaraid_sas] [ 271.807955] Code: 20 49 83 c7 20 48 89 d1 48 83 e1 fc 83 e2 01 4c 0f 45 f9 49 8b 5f 10 45 8b 67 18 4c 89 e9 4c 8d 69 08 45 85 eb 74 52 45 29 ce <48> 89 19 83 c0 01 45 85 f6 7f af 4c 8b 7d c8 c1 e0 03 b9 01 00 00 [ 271.811246] RSP: 0018:ff50bfb988f8b298 EFLAGS: 00010206 [ 271.812215] RAX: 0000000000000200 RBX: 00000000f6a00000 RCX: ff50bfb983b9a000 [ 271.813520] RDX: ff50bfb983b9a008 RSI: 0000000000000000 RDI: 0000000000000000 [ 271.814822] RBP: ff50bfb988f8b2f8 R08: 00000000ffbb0000 R09: 0000000000001000 [ 271.816365] R10: 0000000000001000 R11: 0000000000000fff R12: 0000000000200000 [ 271.817860] R13: ff50bfb983b9a008 R14: 00000000001ff000 R15: ff2b1ee1e1f2cb08 [ 271.819321] FS: 0000000000000000(0000) GS:ff2b1ee542880000(0000) knlGS:0000000000000000 [ 271.820923] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 [ 271.822134] CR2: ff50bfb983b9a000 CR3: 0000060144844004 CR4: 0000000000f73ef0 [ 271.823587] PKRU: 55555554 [ 271.824274] Call Trace: [ 271.824920] <TASK> [ 271.825514] megasas_build_io_fusion+0x2e2/0x330 [megaraid_sas] [ 271.826772] megasas_build_and_issue_cmd_fusion+0xa3/0x280 [megaraid_sas] [ 271.828168] ? sd_init_command+0x138/0x4a0 [ 271.829110] megasas_queue_command+0x11b/0x200 [megaraid_sas] [ 271.830347] ? scsi_init_command+0x74/0xc0 [ 271.831291] ? __rq_qos_issue+0x29/0x50 [ 271.832199] scsi_dispatch_cmd+0x95/0x290 [ 271.833129] scsi_queue_rq+0x62e/0x8e0 [ 271.834017] blk_mq_dispatch_rq_list+0x131/0x510 [ 271.835064] __blk_mq_do_dispatch_sched+0x2d9/0x360 [ 271.836154] ? mod_memcg_lruvec_state+0x101/0x2f0 [ 271.837206] __blk_mq_sched_dispatch_requests+0x157/0x1a0 [ 271.838369] ? elv_rb_add+0x70/0x90 [ 271.839210] blk_mq_sched_dispatch_requests+0x2d/0x80 [ 271.840316] blk_mq_run_hw_queue+0x2c0/0x330 [ 271.841295] blk_mq_dispatch_list+0x159/0x350 [ 271.842319] blk_mq_flush_plug_list+0x59/0x1e0 [ 271.843328] blk_add_rq_to_plug+0xe7/0x220 [ 271.844262] ? blk_account_io_start+0xd2/0x210 [ 271.845258] blk_mq_submit_bio+0x677/0x920 [ 271.846182] __submit_bio+0xad/0x250 [ 271.847022] submit_bio_noacct_nocheck+0x102/0x1d0 [ 271.848068] submit_bio_noacct+0x131/0x430 [ 271.848982] submit_bio+0xb3/0x110 [ 271.849782] ext4_io_submit+0x40/0x70 [ 271.850624] ext4_do_writepages+0x60a/0x970 [ 271.851546] ? psi_group_change+0x195/0x460 [ 271.852466] ext4_writepages+0xc8/0x1b0 [ 271.853325] ? ext4_writepages+0xc8/0x1b0 [ 271.854209] do_writepages+0xcb/0x180 [ 271.855038] ? write_inode+0x78/0x150 [ 271.855863] __writeback_single_inode+0x45/0x270 [ 271.856852] ? wbc_detach_inode+0x115/0x2c0 [ 271.857756] writeback_sb_inodes+0x25e/0x5e0 [ 271.858685] __writeback_inodes_wb+0x54/0x100 [ 271.859621] ? queue_io+0x13b/0x150 [ 271.860416] wb_writeback+0x2e0/0x370 [ 271.861223] wb_workfn+0x39e/0x470 [ 271.861987] process_one_work+0x1ac/0x3d0 [ 271.862851] worker_thread+0x1b8/0x360 [ 271.863671] ? __pfx_worker_thread+0x10/0x10 [ 271.864586] kthread+0xf7/0x130 [ 271.865304] ? __pfx_kthread+0x10/0x10 [ 271.866118] ret_from_fork+0x195/0x2a0 [ 271.866934] ? __pfx_kthread+0x10/0x10 [ 271.867743] ? __pfx_kthread+0x10/0x10 [ 271.868551] ret_from_fork_asm+0x1a/0x30 [ 271.869394] </TASK> To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2148534/+subscriptions

[Bug 2165806] Re: Ubuntu 26.04.1 Server ISO cannot install on MAXIO MAP1602 NVMe due to kernel 7.0.0-30 I/O regression; 7.0.0-31 works

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165806 Title: Ubuntu 26.04.1 Server ISO cannot install on MAXIO MAP1602 NVMe due to kernel 7.0.0-30 I/O regression; 7.0.0-31 works Status in linux package in Ubuntu: New Bug description: Ubuntu 26.04.1 Server cannot be installed on an ACEMAGIC M1A PRO+ with its factory HOGE H820 2 TB / MAXIO MAP1602 NVMe SSD. The Ubuntu 26.04.1 live-server ISO boots kernel 7.0.0-30. With this kernel, reproducible NVMe READ and WRITE failures occur during installation, and the installation eventually fails during grub-install. The same physical machine, SSD, M.2 slot, and SSD firmware operate normally with kernels 7.0.0-14, 6.8, 6.14, and 7.0.0-31. The regression is also independently reproducible on installed Ubuntu 24.04 and Ubuntu 26.04 systems by switching between 7.0.0-30 and 7.0.0-31. PRIMARY IMPACT The current Ubuntu 26.04.1 Server ISO cannot complete installation on this hardware. During installation, kernel 7.0.0-30 reports NVMe block I/O failures. The installer then eventually fails with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The grub-install failure occurs in the same installation in which the kernel is reporting NVMe READ/WRITE failures. I believe the storage regression is likely contributing to the grub-install failure, but I have not independently proven that the NVMe errors are the direct cause of the grub-install error. HARDWARE System: ACEMAGIC M1A PRO+ AMD Ryzen AI Max+ 395 AMD Radeon 8060S 128 GB memory, 8000 MT/s BIOS: American Megatrends Project version: P10_F11_20_IEC0008_BI0010_AMI_120W Build date: 2026-01-13 NVMe: Model: HOGE H820 2TB Firmware: H240313a PCI vendor ID: 0x1e4b PCI device/controller: 0x1602 Controller family: MAXIO MAP1602 NVMe SMART data observed during troubleshooting: critical_warning: 0 available_spare: 100% percentage_used: 0% media_errors: 0 num_err_log_entries: 0 temperature: approximately 39 C UBUNTU 26.04.1 SERVER INSTALL FAILURE Installation method: - Ubuntu 26.04.1 live-server ISO - ISO written directly to USB using dd - UEFI boot - Secure Boot disabled - entire NVMe disk - no LVM - ext4 root filesystem - 1 GiB FAT EFI System Partition The installer successfully partitions the disk, creates the filesystems, and mounts the target. During installation, kernel 7.0.0-30 begins producing NVMe block I/O failures. Examples of WRITE failures observed: invalid error, dev nvme0n1, sector 11896 op 0x1:(WRITE) ... invalid error, dev nvme0n1, sector 2080 op 0x1:(WRITE) ... followed by messages including: Buffer I/O error on dev nvme0n1p1, logical block 32, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 33, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 34, lost async page write READ failures were also observed, including: invalid error, dev nvme0n1, sector 3104 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3144 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3184 op 0x0:(READ) ... ... invalid error, dev nvme0n1, sector 3480 op 0x0:(READ) ... The failing READ requests repeatedly showed multi-segment I/O, including phys_seg 33. Installation ultimately fails at grub-install with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The Ubuntu installer offered to send an error report to Canonical after the failure, and I selected that option. Therefore an installer-side failure report from this machine should already exist in Canonical's error-reporting infrastructure, although I do not currently have an identifier for that report. REGRESSION TESTING All of the following tests used the same physical ACEMAGIC M1A PRO+, same HOGE H820 SSD, same M.2 slot, and same SSD firmware. Ubuntu 26.04 ------------ 1. Ubuntu 26.04 base installation, kernel 7.0.0-14 PASS Ubuntu 26.04 installs and boots successfully. No NVMe "invalid error", Buffer I/O errors, or related read/write failures are present. 2. Upgrade the installed Ubuntu 26.04 system to kernel 7.0.0-30 FAIL The same NVMe "invalid error" messages return. 3. Upgrade the same Ubuntu 26.04 installation from 7.0.0-30 to 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS The NVMe errors disappear again. Ubuntu 26.04.1 live-server ISO ------------------------------ 4. Ubuntu 26.04.1 live-server ISO, kernel 7.0.0-30 FAIL The same NVMe READ/WRITE errors occur during installation. Unlike the Ubuntu 24.04.4 installer case below, the failures are severe enough here that grub-install fails and Ubuntu 26.04.1 Server cannot be installed. Ubuntu 24.04 ------------ 5. Ubuntu 24.04.4 installation using kernel 7.0.0-30 KERNEL REGRESSION PRESENT, BUT INSTALLATION COMPLETES Ubuntu 24.04.4 installs successfully. After boot, the same NVMe "invalid error" messages are present in the kernel logs. Thus the storage regression is present with 7.0.0-30 on Ubuntu 24.04 as well, even though that installer happens to complete. 6. Downgrade the same Ubuntu 24.04 system to kernel 6.8 PASS The NVMe errors disappear. Sustained fio testing completes without NVMe errors. The Radeon 8060S does not initialize correctly on this older kernel, which is unrelated to the NVMe regression. 7. Upgrade the same Ubuntu 24.04 system to kernel 6.14 PASS NVMe remains clean. Radeon 8060S / amdgpu initializes successfully. Sustained fio testing completes without NVMe errors. 8. Upgrade the same Ubuntu 24.04 system to kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS NVMe errors remain absent. A 10-minute mixed direct-I/O fio workload using libaio, iodepth=32, numjobs=4, and direct=1 completes without NVMe errors. SUMMARY Ubuntu 26.04 + 7.0.0-14 -> PASS Ubuntu 26.04 + 7.0.0-30 -> FAIL Ubuntu 26.04 + 7.0.0-31 -> PASS Ubuntu 26.04.1 ISO + -30 -> FAIL, installation cannot complete Ubuntu 24.04.4 + 7.0.0-30 -> regression present, install completes Ubuntu 24.04 + 6.8 -> PASS Ubuntu 24.04 + 6.14 -> PASS Ubuntu 24.04 + 7.0.0-31 -> PASS The -30 -> -31 comparison has been reproduced on both Ubuntu 24.04 and Ubuntu 26.04. This strongly isolates the problem to the kernel code present in 7.0.0-30 and absent/fixed in 7.0.0-31, rather than to the Ubuntu userspace release or SSD hardware. APST TESTING NVMe APST was tested as a possible cause. The Ubuntu 26.04.1 installer was booted with: nvme_core.default_ps_max_latency_us=0 The runtime value was verified as: /sys/module/nvme_core/parameters/default_ps_max_latency_us = 0 The same NVMe WRITE / Buffer I/O errors still occurred. Disabling NVMe APST therefore does not resolve the failure. STORAGE VALIDATION The SSD does not appear to contain defective media at the sector range where errors were reported. On a working kernel, a direct read spanning the affected region completed successfully: dd if=/dev/nvme0n1 of=/dev/null bs=512 skip=3000 count=600 iflag=direct 600+0 records in 600+0 records out No new NVMe errors were produced. The same SSD also completed substantial mixed read/write fio testing on working kernels without media errors or kernel block-I/O failures. This included sustained testing on: - 6.8 - 6.14 - 7.0.0-31 LIKELY UPSTREAM REGRESSION / FIX I have not performed a source-level git bisect, so the following is a suspected cause rather than independently proven. Linux 7.0.11 included: iommupt: Avoid rewalking during map upstream commit: d6c65b0fd6218bd21ed0be7a8d3218e8f6dc91de Linux 7.0.13 later included: iommu/dma: Do not try to iommu_map a 0 length region in swiotlb upstream commit: 6ec91df8aff77e2e8fe3179c1f3fc15b43a40ba3 The latter fix describes an IOMMU DMA mapping path in which an unaligned mapping can produce a zero-length middle region. iommu_map() then rejects the zero-length mapping as illegal. The upstream description explicitly notes NVMe as a frequent trigger because NVMe can issue unusually aligned buffers in some paths. Ubuntu kernel 7.0.0-31 contains the Linux 7.0.13 and 7.0.14 upstream stable updates, and 7.0.0-31 is the first tested Ubuntu 7.0 kernel in this environment where the problem disappears. Given the observed: 7.0.0-30 -> FAIL 7.0.0-31 -> PASS and the I/O errors being reported as "invalid error", this IOMMU/DMA fix appears to be a strong candidate for the fix. Please confirm whether this system is hitting that regression, or another change incorporated between Ubuntu 7.0.0-30 and 7.0.0-31. EXPECTED RESULT Ubuntu 26.04.1 Server should install successfully on this hardware. The HOGE H820 / MAXIO MAP1602 NVMe controller should operate without kernel block-I/O failures. IMPACT The practical issue is not merely log noise. The kernel shipped on the Ubuntu 26.04.1 Server ISO causes storage errors during installation and prevents the current 26.04.1 live-server image from completing installation on this system. The same regression can also be reproduced after installation by booting 7.0.0-30 on both Ubuntu 24.04 and Ubuntu 26.04. Kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa does not reproduce the problem on either release. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165806/+subscriptions

[Bug 2165828] Re: ASUS Vivobook S 15 (x1e80100, Samsung ATNA33XC20 eDP): only half of internal panel displays after DPMS off/on; VT switch (full modeset) recovers

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165828 Title: ASUS Vivobook S 15 (x1e80100, Samsung ATNA33XC20 eDP): only half of internal panel displays after DPMS off/on; VT switch (full modeset) recovers Status in linux package in Ubuntu: New Bug description: On the ASUS Vivobook S 15 (Snapdragon X / X1E-78-100), when the internal eDP panel is powered off by the screensaver (DRM DPMS off) and then woken up, only half of the panel displays content — the other half stays black. The session keeps running normally; only the panel output is broken. A full display re-initialization restores the panel completely: - Switch to a text VT (Ctrl+Alt+F3) and back (Ctrl+Alt+F2), or - `chvt 3 && chvt <previous>` from a root shell After the VT switch the panel is fully functional again — no reboot or logout needed. This proves the failure is in the display pipeline's power-on path (encoder/panel re-init after DPMS), not in the userspace session. Affected hardware / software ---------------------------- - ASUS Vivobook S 15 S5507QA (Snapdragon X X1E-78-100), internal eDP panel Samsung ATNA33XC20 (driver: panel_samsung_atna33xc20) - Display controller: msm_dpu ae01000.display-controller (drm/msm, DPU 9.0.2: "dpu hardware revision:0x90020000") - Ubuntu 26.04.1 "Resolute", arm64 - Kernel: 7.0.0-30-generic (latest at time of report, resolute-updates) - GNOME Shell on Wayland Reproduction ------------ 1. Boot normally, log in (Wayland session) 2. Let the screensaver blank the panel (GNOME Screen Blank / DPMS power-off) 3. Wake the display (keypad/touch) 4. Only half of the panel shows output; the other half remains black 5. Ctrl+Alt+F3 / Ctrl+Alt+F2 → panel fully restored Notes / observations -------------------- - No drm/dpu error messages are logged at the moment the half-screen state appears — the failure is silent (journalctl -k shows only the boot-time "dpu hardware revision:0x90020000" line). - Unrelated-but-possibly-relevant error seen earlier this boot under heavy software-rendering load: [drm:dpu_crtc_frame_event_cb [msm]] *ERROR* crtc109 event 1 overflow (observed while the GPU was unaccelerated due to missing firmware — see separate linux-firmware report; the half-screen issue persists independently of GPU acceleration state.) - The panel is driven with DSC; a plausible cause is incorrect DSC/tile reconfiguration on the DPMS power-up path in dpu1 encoder code, since a full modeset (VT switch) reprograms everything and recovers. - Not the same as the known x1e80100 eDP HPD pinctrl issue (Stephan Gerhold's Aug 2025 series, display never comes up at all) — here the display works fine from boot and only breaks on DPMS resume. Suggested next steps for maintainers ------------------------------------ - Reproduce with dpms off/on (e.g. `sleep 5 &&swaymsg ...` equivalent: `modetest -M msm_dpu -w` DPMS cycles, or GNOME screen-blank timer) - Instrument dpu1 encoder enable/power-on path around DSC config restore - Compare against the DPMS handling for the same panel on other x1e80100 devices (Lenovo T14s Gen 6 uses the same ATNA33XC20 panel and may be affected identically) Environment (from affected machine) ----------------------------------- DistroRelease: Ubuntu 26.04 Architecture: arm64 MachineType: ASUSTeK COMPUTER INC. Vivobook S 15 S5507QA_S5507QAD Kernel: 7.0.0-30-generic SourcePackage: linux (version 7.0.0-30.30) Tags: arm64 wayland-session kernel-daily-bug Happy to test patches, capture mode dumps (modetest), or edid on request. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165828/+subscriptions

[Bug 2165828] [NEW] ASUS Vivobook S 15 (x1e80100, Samsung ATNA33XC20 eDP): only half of internal panel displays after DPMS off/on; VT switch (full modeset) recovers

Public bug reported: On the ASUS Vivobook S 15 (Snapdragon X / X1E-78-100), when the internal eDP panel is powered off by the screensaver (DRM DPMS off) and then woken up, only half of the panel displays content — the other half stays black. The session keeps running normally; only the panel output is broken. A full display re-initialization restores the panel completely: - Switch to a text VT (Ctrl+Alt+F3) and back (Ctrl+Alt+F2), or - `chvt 3 && chvt <previous>` from a root shell After the VT switch the panel is fully functional again — no reboot or logout needed. This proves the failure is in the display pipeline's power-on path (encoder/panel re-init after DPMS), not in the userspace session. Affected hardware / software ---------------------------- - ASUS Vivobook S 15 S5507QA (Snapdragon X X1E-78-100), internal eDP panel Samsung ATNA33XC20 (driver: panel_samsung_atna33xc20) - Display controller: msm_dpu ae01000.display-controller (drm/msm, DPU 9.0.2: "dpu hardware revision:0x90020000") - Ubuntu 26.04.1 "Resolute", arm64 - Kernel: 7.0.0-30-generic (latest at time of report, resolute-updates) - GNOME Shell on Wayland Reproduction ------------ 1. Boot normally, log in (Wayland session) 2. Let the screensaver blank the panel (GNOME Screen Blank / DPMS power-off) 3. Wake the display (keypad/touch) 4. Only half of the panel shows output; the other half remains black 5. Ctrl+Alt+F3 / Ctrl+Alt+F2 → panel fully restored Notes / observations -------------------- - No drm/dpu error messages are logged at the moment the half-screen state appears — the failure is silent (journalctl -k shows only the boot-time "dpu hardware revision:0x90020000" line). - Unrelated-but-possibly-relevant error seen earlier this boot under heavy software-rendering load: [drm:dpu_crtc_frame_event_cb [msm]] *ERROR* crtc109 event 1 overflow (observed while the GPU was unaccelerated due to missing firmware — see separate linux-firmware report; the half-screen issue persists independently of GPU acceleration state.) - The panel is driven with DSC; a plausible cause is incorrect DSC/tile reconfiguration on the DPMS power-up path in dpu1 encoder code, since a full modeset (VT switch) reprograms everything and recovers. - Not the same as the known x1e80100 eDP HPD pinctrl issue (Stephan Gerhold's Aug 2025 series, display never comes up at all) — here the display works fine from boot and only breaks on DPMS resume. Suggested next steps for maintainers ------------------------------------ - Reproduce with dpms off/on (e.g. `sleep 5 &&swaymsg ...` equivalent: `modetest -M msm_dpu -w` DPMS cycles, or GNOME screen-blank timer) - Instrument dpu1 encoder enable/power-on path around DSC config restore - Compare against the DPMS handling for the same panel on other x1e80100 devices (Lenovo T14s Gen 6 uses the same ATNA33XC20 panel and may be affected identically) Environment (from affected machine) ----------------------------------- DistroRelease: Ubuntu 26.04 Architecture: arm64 MachineType: ASUSTeK COMPUTER INC. Vivobook S 15 S5507QA_S5507QAD Kernel: 7.0.0-30-generic SourcePackage: linux (version 7.0.0-30.30) Tags: arm64 wayland-session kernel-daily-bug Happy to test patches, capture mode dumps (modetest), or edid on request. ** Affects: linux (Ubuntu) Importance: Undecided Status: New -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165828 Title: ASUS Vivobook S 15 (x1e80100, Samsung ATNA33XC20 eDP): only half of internal panel displays after DPMS off/on; VT switch (full modeset) recovers Status in linux package in Ubuntu: New Bug description: On the ASUS Vivobook S 15 (Snapdragon X / X1E-78-100), when the internal eDP panel is powered off by the screensaver (DRM DPMS off) and then woken up, only half of the panel displays content — the other half stays black. The session keeps running normally; only the panel output is broken. A full display re-initialization restores the panel completely: - Switch to a text VT (Ctrl+Alt+F3) and back (Ctrl+Alt+F2), or - `chvt 3 && chvt <previous>` from a root shell After the VT switch the panel is fully functional again — no reboot or logout needed. This proves the failure is in the display pipeline's power-on path (encoder/panel re-init after DPMS), not in the userspace session. Affected hardware / software ---------------------------- - ASUS Vivobook S 15 S5507QA (Snapdragon X X1E-78-100), internal eDP panel Samsung ATNA33XC20 (driver: panel_samsung_atna33xc20) - Display controller: msm_dpu ae01000.display-controller (drm/msm, DPU 9.0.2: "dpu hardware revision:0x90020000") - Ubuntu 26.04.1 "Resolute", arm64 - Kernel: 7.0.0-30-generic (latest at time of report, resolute-updates) - GNOME Shell on Wayland Reproduction ------------ 1. Boot normally, log in (Wayland session) 2. Let the screensaver blank the panel (GNOME Screen Blank / DPMS power-off) 3. Wake the display (keypad/touch) 4. Only half of the panel shows output; the other half remains black 5. Ctrl+Alt+F3 / Ctrl+Alt+F2 → panel fully restored Notes / observations -------------------- - No drm/dpu error messages are logged at the moment the half-screen state appears — the failure is silent (journalctl -k shows only the boot-time "dpu hardware revision:0x90020000" line). - Unrelated-but-possibly-relevant error seen earlier this boot under heavy software-rendering load: [drm:dpu_crtc_frame_event_cb [msm]] *ERROR* crtc109 event 1 overflow (observed while the GPU was unaccelerated due to missing firmware — see separate linux-firmware report; the half-screen issue persists independently of GPU acceleration state.) - The panel is driven with DSC; a plausible cause is incorrect DSC/tile reconfiguration on the DPMS power-up path in dpu1 encoder code, since a full modeset (VT switch) reprograms everything and recovers. - Not the same as the known x1e80100 eDP HPD pinctrl issue (Stephan Gerhold's Aug 2025 series, display never comes up at all) — here the display works fine from boot and only breaks on DPMS resume. Suggested next steps for maintainers ------------------------------------ - Reproduce with dpms off/on (e.g. `sleep 5 &&swaymsg ...` equivalent: `modetest -M msm_dpu -w` DPMS cycles, or GNOME screen-blank timer) - Instrument dpu1 encoder enable/power-on path around DSC config restore - Compare against the DPMS handling for the same panel on other x1e80100 devices (Lenovo T14s Gen 6 uses the same ATNA33XC20 panel and may be affected identically) Environment (from affected machine) ----------------------------------- DistroRelease: Ubuntu 26.04 Architecture: arm64 MachineType: ASUSTeK COMPUTER INC. Vivobook S 15 S5507QA_S5507QAD Kernel: 7.0.0-30-generic SourcePackage: linux (version 7.0.0-30.30) Tags: arm64 wayland-session kernel-daily-bug Happy to test patches, capture mode dumps (modetest), or edid on request. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165828/+subscriptions

пятница

[Bug 2165806] [NEW] Ubuntu 26.04.1 Server ISO cannot install on MAXIO MAP1602 NVMe due to kernel 7.0.0-30 I/O regression; 7.0.0-31 works

Public bug reported: Ubuntu 26.04.1 Server cannot be installed on an ACEMAGIC M1A PRO+ with its factory HOGE H820 2 TB / MAXIO MAP1602 NVMe SSD. The Ubuntu 26.04.1 live-server ISO boots kernel 7.0.0-30. With this kernel, reproducible NVMe READ and WRITE failures occur during installation, and the installation eventually fails during grub-install. The same physical machine, SSD, M.2 slot, and SSD firmware operate normally with kernels 7.0.0-14, 6.8, 6.14, and 7.0.0-31. The regression is also independently reproducible on installed Ubuntu 24.04 and Ubuntu 26.04 systems by switching between 7.0.0-30 and 7.0.0-31. PRIMARY IMPACT The current Ubuntu 26.04.1 Server ISO cannot complete installation on this hardware. During installation, kernel 7.0.0-30 reports NVMe block I/O failures. The installer then eventually fails with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The grub-install failure occurs in the same installation in which the kernel is reporting NVMe READ/WRITE failures. I believe the storage regression is likely contributing to the grub-install failure, but I have not independently proven that the NVMe errors are the direct cause of the grub-install error. HARDWARE System: ACEMAGIC M1A PRO+ AMD Ryzen AI Max+ 395 AMD Radeon 8060S 128 GB memory, 8000 MT/s BIOS: American Megatrends Project version: P10_F11_20_IEC0008_BI0010_AMI_120W Build date: 2026-01-13 NVMe: Model: HOGE H820 2TB Firmware: H240313a PCI vendor ID: 0x1e4b PCI device/controller: 0x1602 Controller family: MAXIO MAP1602 NVMe SMART data observed during troubleshooting: critical_warning: 0 available_spare: 100% percentage_used: 0% media_errors: 0 num_err_log_entries: 0 temperature: approximately 39 C UBUNTU 26.04.1 SERVER INSTALL FAILURE Installation method: - Ubuntu 26.04.1 live-server ISO - ISO written directly to USB using dd - UEFI boot - Secure Boot disabled - entire NVMe disk - no LVM - ext4 root filesystem - 1 GiB FAT EFI System Partition The installer successfully partitions the disk, creates the filesystems, and mounts the target. During installation, kernel 7.0.0-30 begins producing NVMe block I/O failures. Examples of WRITE failures observed: invalid error, dev nvme0n1, sector 11896 op 0x1:(WRITE) ... invalid error, dev nvme0n1, sector 2080 op 0x1:(WRITE) ... followed by messages including: Buffer I/O error on dev nvme0n1p1, logical block 32, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 33, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 34, lost async page write READ failures were also observed, including: invalid error, dev nvme0n1, sector 3104 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3144 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3184 op 0x0:(READ) ... ... invalid error, dev nvme0n1, sector 3480 op 0x0:(READ) ... The failing READ requests repeatedly showed multi-segment I/O, including phys_seg 33. Installation ultimately fails at grub-install with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The Ubuntu installer offered to send an error report to Canonical after the failure, and I selected that option. Therefore an installer-side failure report from this machine should already exist in Canonical's error-reporting infrastructure, although I do not currently have an identifier for that report. REGRESSION TESTING All of the following tests used the same physical ACEMAGIC M1A PRO+, same HOGE H820 SSD, same M.2 slot, and same SSD firmware. Ubuntu 26.04 ------------ 1. Ubuntu 26.04 base installation, kernel 7.0.0-14 PASS Ubuntu 26.04 installs and boots successfully. No NVMe "invalid error", Buffer I/O errors, or related read/write failures are present. 2. Upgrade the installed Ubuntu 26.04 system to kernel 7.0.0-30 FAIL The same NVMe "invalid error" messages return. 3. Upgrade the same Ubuntu 26.04 installation from 7.0.0-30 to 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS The NVMe errors disappear again. Ubuntu 26.04.1 live-server ISO ------------------------------ 4. Ubuntu 26.04.1 live-server ISO, kernel 7.0.0-30 FAIL The same NVMe READ/WRITE errors occur during installation. Unlike the Ubuntu 24.04.4 installer case below, the failures are severe enough here that grub-install fails and Ubuntu 26.04.1 Server cannot be installed. Ubuntu 24.04 ------------ 5. Ubuntu 24.04.4 installation using kernel 7.0.0-30 KERNEL REGRESSION PRESENT, BUT INSTALLATION COMPLETES Ubuntu 24.04.4 installs successfully. After boot, the same NVMe "invalid error" messages are present in the kernel logs. Thus the storage regression is present with 7.0.0-30 on Ubuntu 24.04 as well, even though that installer happens to complete. 6. Downgrade the same Ubuntu 24.04 system to kernel 6.8 PASS The NVMe errors disappear. Sustained fio testing completes without NVMe errors. The Radeon 8060S does not initialize correctly on this older kernel, which is unrelated to the NVMe regression. 7. Upgrade the same Ubuntu 24.04 system to kernel 6.14 PASS NVMe remains clean. Radeon 8060S / amdgpu initializes successfully. Sustained fio testing completes without NVMe errors. 8. Upgrade the same Ubuntu 24.04 system to kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS NVMe errors remain absent. A 10-minute mixed direct-I/O fio workload using libaio, iodepth=32, numjobs=4, and direct=1 completes without NVMe errors. SUMMARY Ubuntu 26.04 + 7.0.0-14 -> PASS Ubuntu 26.04 + 7.0.0-30 -> FAIL Ubuntu 26.04 + 7.0.0-31 -> PASS Ubuntu 26.04.1 ISO + -30 -> FAIL, installation cannot complete Ubuntu 24.04.4 + 7.0.0-30 -> regression present, install completes Ubuntu 24.04 + 6.8 -> PASS Ubuntu 24.04 + 6.14 -> PASS Ubuntu 24.04 + 7.0.0-31 -> PASS The -30 -> -31 comparison has been reproduced on both Ubuntu 24.04 and Ubuntu 26.04. This strongly isolates the problem to the kernel code present in 7.0.0-30 and absent/fixed in 7.0.0-31, rather than to the Ubuntu userspace release or SSD hardware. APST TESTING NVMe APST was tested as a possible cause. The Ubuntu 26.04.1 installer was booted with: nvme_core.default_ps_max_latency_us=0 The runtime value was verified as: /sys/module/nvme_core/parameters/default_ps_max_latency_us = 0 The same NVMe WRITE / Buffer I/O errors still occurred. Disabling NVMe APST therefore does not resolve the failure. STORAGE VALIDATION The SSD does not appear to contain defective media at the sector range where errors were reported. On a working kernel, a direct read spanning the affected region completed successfully: dd if=/dev/nvme0n1 of=/dev/null bs=512 skip=3000 count=600 iflag=direct 600+0 records in 600+0 records out No new NVMe errors were produced. The same SSD also completed substantial mixed read/write fio testing on working kernels without media errors or kernel block-I/O failures. This included sustained testing on: - 6.8 - 6.14 - 7.0.0-31 LIKELY UPSTREAM REGRESSION / FIX I have not performed a source-level git bisect, so the following is a suspected cause rather than independently proven. Linux 7.0.11 included: iommupt: Avoid rewalking during map upstream commit: d6c65b0fd6218bd21ed0be7a8d3218e8f6dc91de Linux 7.0.13 later included: iommu/dma: Do not try to iommu_map a 0 length region in swiotlb upstream commit: 6ec91df8aff77e2e8fe3179c1f3fc15b43a40ba3 The latter fix describes an IOMMU DMA mapping path in which an unaligned mapping can produce a zero-length middle region. iommu_map() then rejects the zero-length mapping as illegal. The upstream description explicitly notes NVMe as a frequent trigger because NVMe can issue unusually aligned buffers in some paths. Ubuntu kernel 7.0.0-31 contains the Linux 7.0.13 and 7.0.14 upstream stable updates, and 7.0.0-31 is the first tested Ubuntu 7.0 kernel in this environment where the problem disappears. Given the observed: 7.0.0-30 -> FAIL 7.0.0-31 -> PASS and the I/O errors being reported as "invalid error", this IOMMU/DMA fix appears to be a strong candidate for the fix. Please confirm whether this system is hitting that regression, or another change incorporated between Ubuntu 7.0.0-30 and 7.0.0-31. EXPECTED RESULT Ubuntu 26.04.1 Server should install successfully on this hardware. The HOGE H820 / MAXIO MAP1602 NVMe controller should operate without kernel block-I/O failures. IMPACT The practical issue is not merely log noise. The kernel shipped on the Ubuntu 26.04.1 Server ISO causes storage errors during installation and prevents the current 26.04.1 live-server image from completing installation on this system. The same regression can also be reproduced after installation by booting 7.0.0-30 on both Ubuntu 24.04 and Ubuntu 26.04. Kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa does not reproduce the problem on either release. ** Affects: linux (Ubuntu) Importance: Undecided Status: New -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165806 Title: Ubuntu 26.04.1 Server ISO cannot install on MAXIO MAP1602 NVMe due to kernel 7.0.0-30 I/O regression; 7.0.0-31 works Status in linux package in Ubuntu: New Bug description: Ubuntu 26.04.1 Server cannot be installed on an ACEMAGIC M1A PRO+ with its factory HOGE H820 2 TB / MAXIO MAP1602 NVMe SSD. The Ubuntu 26.04.1 live-server ISO boots kernel 7.0.0-30. With this kernel, reproducible NVMe READ and WRITE failures occur during installation, and the installation eventually fails during grub-install. The same physical machine, SSD, M.2 slot, and SSD firmware operate normally with kernels 7.0.0-14, 6.8, 6.14, and 7.0.0-31. The regression is also independently reproducible on installed Ubuntu 24.04 and Ubuntu 26.04 systems by switching between 7.0.0-30 and 7.0.0-31. PRIMARY IMPACT The current Ubuntu 26.04.1 Server ISO cannot complete installation on this hardware. During installation, kernel 7.0.0-30 reports NVMe block I/O failures. The installer then eventually fails with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The grub-install failure occurs in the same installation in which the kernel is reporting NVMe READ/WRITE failures. I believe the storage regression is likely contributing to the grub-install failure, but I have not independently proven that the NVMe errors are the direct cause of the grub-install error. HARDWARE System: ACEMAGIC M1A PRO+ AMD Ryzen AI Max+ 395 AMD Radeon 8060S 128 GB memory, 8000 MT/s BIOS: American Megatrends Project version: P10_F11_20_IEC0008_BI0010_AMI_120W Build date: 2026-01-13 NVMe: Model: HOGE H820 2TB Firmware: H240313a PCI vendor ID: 0x1e4b PCI device/controller: 0x1602 Controller family: MAXIO MAP1602 NVMe SMART data observed during troubleshooting: critical_warning: 0 available_spare: 100% percentage_used: 0% media_errors: 0 num_err_log_entries: 0 temperature: approximately 39 C UBUNTU 26.04.1 SERVER INSTALL FAILURE Installation method: - Ubuntu 26.04.1 live-server ISO - ISO written directly to USB using dd - UEFI boot - Secure Boot disabled - entire NVMe disk - no LVM - ext4 root filesystem - 1 GiB FAT EFI System Partition The installer successfully partitions the disk, creates the filesystems, and mounts the target. During installation, kernel 7.0.0-30 begins producing NVMe block I/O failures. Examples of WRITE failures observed: invalid error, dev nvme0n1, sector 11896 op 0x1:(WRITE) ... invalid error, dev nvme0n1, sector 2080 op 0x1:(WRITE) ... followed by messages including: Buffer I/O error on dev nvme0n1p1, logical block 32, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 33, lost async page write Buffer I/O error on dev nvme0n1p1, logical block 34, lost async page write READ failures were also observed, including: invalid error, dev nvme0n1, sector 3104 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3144 op 0x0:(READ) ... invalid error, dev nvme0n1, sector 3184 op 0x0:(READ) ... ... invalid error, dev nvme0n1, sector 3480 op 0x0:(READ) ... The failing READ requests repeatedly showed multi-segment I/O, including phys_seg 33. Installation ultimately fails at grub-install with: grub-install: error: cannot copy '/usr/lib/grub/x86_64-efi-signed/grubx64.efi.signed' to '/boot/efi/EFI/ubuntu/grubx64.efi': Invalid argument. The Ubuntu installer offered to send an error report to Canonical after the failure, and I selected that option. Therefore an installer-side failure report from this machine should already exist in Canonical's error-reporting infrastructure, although I do not currently have an identifier for that report. REGRESSION TESTING All of the following tests used the same physical ACEMAGIC M1A PRO+, same HOGE H820 SSD, same M.2 slot, and same SSD firmware. Ubuntu 26.04 ------------ 1. Ubuntu 26.04 base installation, kernel 7.0.0-14 PASS Ubuntu 26.04 installs and boots successfully. No NVMe "invalid error", Buffer I/O errors, or related read/write failures are present. 2. Upgrade the installed Ubuntu 26.04 system to kernel 7.0.0-30 FAIL The same NVMe "invalid error" messages return. 3. Upgrade the same Ubuntu 26.04 installation from 7.0.0-30 to 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS The NVMe errors disappear again. Ubuntu 26.04.1 live-server ISO ------------------------------ 4. Ubuntu 26.04.1 live-server ISO, kernel 7.0.0-30 FAIL The same NVMe READ/WRITE errors occur during installation. Unlike the Ubuntu 24.04.4 installer case below, the failures are severe enough here that grub-install fails and Ubuntu 26.04.1 Server cannot be installed. Ubuntu 24.04 ------------ 5. Ubuntu 24.04.4 installation using kernel 7.0.0-30 KERNEL REGRESSION PRESENT, BUT INSTALLATION COMPLETES Ubuntu 24.04.4 installs successfully. After boot, the same NVMe "invalid error" messages are present in the kernel logs. Thus the storage regression is present with 7.0.0-30 on Ubuntu 24.04 as well, even though that installer happens to complete. 6. Downgrade the same Ubuntu 24.04 system to kernel 6.8 PASS The NVMe errors disappear. Sustained fio testing completes without NVMe errors. The Radeon 8060S does not initialize correctly on this older kernel, which is unrelated to the NVMe regression. 7. Upgrade the same Ubuntu 24.04 system to kernel 6.14 PASS NVMe remains clean. Radeon 8060S / amdgpu initializes successfully. Sustained fio testing completes without NVMe errors. 8. Upgrade the same Ubuntu 24.04 system to kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa PASS NVMe errors remain absent. A 10-minute mixed direct-I/O fio workload using libaio, iodepth=32, numjobs=4, and direct=1 completes without NVMe errors. SUMMARY Ubuntu 26.04 + 7.0.0-14 -> PASS Ubuntu 26.04 + 7.0.0-30 -> FAIL Ubuntu 26.04 + 7.0.0-31 -> PASS Ubuntu 26.04.1 ISO + -30 -> FAIL, installation cannot complete Ubuntu 24.04.4 + 7.0.0-30 -> regression present, install completes Ubuntu 24.04 + 6.8 -> PASS Ubuntu 24.04 + 6.14 -> PASS Ubuntu 24.04 + 7.0.0-31 -> PASS The -30 -> -31 comparison has been reproduced on both Ubuntu 24.04 and Ubuntu 26.04. This strongly isolates the problem to the kernel code present in 7.0.0-30 and absent/fixed in 7.0.0-31, rather than to the Ubuntu userspace release or SSD hardware. APST TESTING NVMe APST was tested as a possible cause. The Ubuntu 26.04.1 installer was booted with: nvme_core.default_ps_max_latency_us=0 The runtime value was verified as: /sys/module/nvme_core/parameters/default_ps_max_latency_us = 0 The same NVMe WRITE / Buffer I/O errors still occurred. Disabling NVMe APST therefore does not resolve the failure. STORAGE VALIDATION The SSD does not appear to contain defective media at the sector range where errors were reported. On a working kernel, a direct read spanning the affected region completed successfully: dd if=/dev/nvme0n1 of=/dev/null bs=512 skip=3000 count=600 iflag=direct 600+0 records in 600+0 records out No new NVMe errors were produced. The same SSD also completed substantial mixed read/write fio testing on working kernels without media errors or kernel block-I/O failures. This included sustained testing on: - 6.8 - 6.14 - 7.0.0-31 LIKELY UPSTREAM REGRESSION / FIX I have not performed a source-level git bisect, so the following is a suspected cause rather than independently proven. Linux 7.0.11 included: iommupt: Avoid rewalking during map upstream commit: d6c65b0fd6218bd21ed0be7a8d3218e8f6dc91de Linux 7.0.13 later included: iommu/dma: Do not try to iommu_map a 0 length region in swiotlb upstream commit: 6ec91df8aff77e2e8fe3179c1f3fc15b43a40ba3 The latter fix describes an IOMMU DMA mapping path in which an unaligned mapping can produce a zero-length middle region. iommu_map() then rejects the zero-length mapping as illegal. The upstream description explicitly notes NVMe as a frequent trigger because NVMe can issue unusually aligned buffers in some paths. Ubuntu kernel 7.0.0-31 contains the Linux 7.0.13 and 7.0.14 upstream stable updates, and 7.0.0-31 is the first tested Ubuntu 7.0 kernel in this environment where the problem disappears. Given the observed: 7.0.0-30 -> FAIL 7.0.0-31 -> PASS and the I/O errors being reported as "invalid error", this IOMMU/DMA fix appears to be a strong candidate for the fix. Please confirm whether this system is hitting that regression, or another change incorporated between Ubuntu 7.0.0-30 and 7.0.0-31. EXPECTED RESULT Ubuntu 26.04.1 Server should install successfully on this hardware. The HOGE H820 / MAXIO MAP1602 NVMe controller should operate without kernel block-I/O failures. IMPACT The practical issue is not merely log noise. The kernel shipped on the Ubuntu 26.04.1 Server ISO causes storage errors during installation and prevents the current 26.04.1 live-server image from completing installation on this system. The same regression can also be reproduced after installation by booting 7.0.0-30 on both Ubuntu 24.04 and Ubuntu 26.04. Kernel 7.0.0-31 from ppa:canonical-kernel-team/ppa does not reproduce the problem on either release. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165806/+subscriptions

[Bug 2165732] Re: [UBUNTU 24.04] kernel: CPU hotplug unsupported by CPUMF

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165732 Title: [UBUNTU 24.04] kernel: CPU hotplug unsupported by CPUMF Status in Ubuntu on IBM z Systems: New Status in linux package in Ubuntu: New Status in linux source package in Noble: New Bug description: Description: kernel: CPU hotplug unsupported by CPUMF Symptom: The kernel crashes with a panic when CPU hotplug add is triggered during execution of command 'perf stat -- <command>' on LPAR. z/VM is not affected. Problem: CPU hotplug add does not allocate memory and does not initialise per-CPU variables required by the PMU device driver. Reproduction: Run the following commands # echo 0 > /sys/devices/system/cpu/cpu1/online # perf stat -e cycles -i -- stress-ng -t10s --matrix X # sleep 1 # echo 1 > /sys/devices/system/cpu/cpu1/online Solution: Install CPU hotplug handler function to support CPU hotplug add and delete operations. Upstream-ID: ddd52d6c635a2dc628a238c637d928425b0e3f53 To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu-z-systems/+bug/2165732/+subscriptions

[Bug 2165740] Re: No audio (SOF DSP fails to boot, -110) after suspend/resume with an external DisplayPort monitor — needs upstream fix 0c0e418dbcf0 backported to the 7.0 kernel

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165740 Title: No audio (SOF DSP fails to boot, -110) after suspend/resume with an external DisplayPort monitor — needs upstream fix 0c0e418dbcf0 backported to the 7.0 kernel Status in linux package in Ubuntu: New Bug description: **Summary** On Ubuntu 26.04 (Resolute, kernel 7.0.0-30-generic) the SOF audio DSP crashes on resume from s2idle whenever an external DisplayPort monitor is connected, leaving **all audio dead — both playback and the internal microphone — until a full reboot**. This is fixed upstream by commit `0c0e418dbcf0`, which is **not** in the 7.0.y series, so it will not reach 26.04 via the normal upstream-stable rollups. Requesting a cherry-pick of that commit into the Resolute kernel. **Impact** - Total loss of audio (output + capture) after a routine suspend/resume — a very common action on a docked laptop. Only recovery is a full reboot; module reload / PCI rebind do not recover it. - Hits repeatedly (roughly every 1–4 suspend/resume cycles while an external DP monitor is attached). **Steps to reproduce** 1. Connect an external DisplayPort monitor (via USB-C dock or direct DP). 2. Suspend (s2idle) and resume — e.g. `sudo rtcwake -m freeze -s 25`. No audio needs to be playing. 3. After resume, audio is gone. Deterministic within a few cycles. Not reproducible with no external monitor connected (bare laptop is fine, on battery and AC). **Kernel log signature (on the failing resume)** ``` sof-audio-pci-intel-mtl 0000:00:1f.3: DSP panic! / error: DSP Firmware Oops / EXCCAUSE 0x5 sof-audio-pci-intel-mtl 0000:00:1f.3: IMR restore failed, trying to cold boot sof-audio-pci-intel-mtl 0000:00:1f.3: 0xd000001c: ROM_EXT, REMOVE_ACCESS_CONTROL / error code: 0x2328 sof-audio-pci-intel-mtl 0000:00:1f.3: dsp init failed after 3 attempts with err: -110 ``` **Upstream fix (please backport)** - Commit `0c0e418dbcf0582bf80d8dbfd9b306607c065992` ("chain-dma: fix potential NULL dereference"), first in mainline **v7.2-rc7**, Cc: stable — already backported to **7.1, 6.18, and 6.12** stable. - It is **not** present in **7.0.y**, which is EOL upstream, so 26.04's 7.0 kernel will not pick it up through the usual "Resolute update: vX upstream stable release" SRUs. Hence this cherry-pick request. **Upstream confirmation / discussion** - thesofproject/sof#10955 — an Intel SOF maintainer confirmed this exact case is the bug fixed by the commit above. Our report, logs, and a deterministic reproducer are in that thread: https://github.com/thesofproject/sof/issues/10955#issuecomment-5371357276 **Environment** - Lenovo ThinkPad P14s Gen6 (21QT0016MX); BIOS R2WET43W 1.25 - Intel Core Ultra 9 285H (Arrow Lake); audio `00:1f.3` Intel "Arrow Lake cAVS" [8086:7728], driver `sof-audio-pci-intel-mtl` - Ubuntu 26.04, kernel 7.0.0-30-generic; `firmware-sof-signed` 2025.12.2-1 (ADSPFW 2.14.1.1) - Suspend type: s2idle. Likely affects all MTL/ARL SOF laptops, not just this model. **Workaround** Reboot to recover audio. No reliable pre-suspend avoidance (undocking before suspend does not help; disabling PCI runtime PM does not help). ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-30-generic 7.0.0-30.30 ProcVersionSignature: Ubuntu 7.0.0-30.30-generic 7.0.12 Uname: Linux 7.0.0-30-generic x86_64 ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 CasperMD5CheckResult: unknown CurrentDesktop: KDE Date: Fri Aug 28 17:32:27 2026 InstallationDate: Installed on 2026-06-24 (65 days ago) InstallationMedia: Kubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423) MachineType: LENOVO 21QT0016MX ProcFB: 0 i915drmfb ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-30-generic root=UUID=d5d6fa86-81cb-4904-925b-70634bd44d5f ro quiet cryptdevice=UUID=8011d035-e532-4f50-b0ed-881177e835e4:luks-8011d035-e532-4f50-b0ed-881177e835e4 root=/dev/mapper/luks-8011d035-e532-4f50-b0ed-881177e835e4 splash PulseList: Error: command ['pacmd', 'list'] failed with exit code 1: No PulseAudio daemon running, or not running as session daemon. SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 06/30/2026 dmi.bios.release: 1.26 dmi.bios.vendor: LENOVO dmi.bios.version: R2WET44W (1.26 ) dmi.board.asset.tag: Not Available dmi.board.name: 21QT0016MX dmi.board.vendor: LENOVO dmi.board.version: SDK0T76528 WIN dmi.chassis.asset.tag: No Asset Information dmi.chassis.type: 10 dmi.chassis.vendor: LENOVO dmi.chassis.version: None dmi.ec.firmware.release: 1.16 dmi.modalias: dmi:bvnLENOVO:bvrR2WET44W(1.26):bd06/30/2026:br1.26:efr1.16:svnLENOVO:pn21QT0016MX:pvrThinkPadP14sGen6:rvnLENOVO:rn21QT0016MX:rvrSDK0T76528WIN:cvnLENOVO:ct10:cvrNone:skuLENOVO_MT_21QT_BU_Think_FM_ThinkPadP14sGen6:pfaThinkPadP14sGen6: dmi.product.family: ThinkPad P14s Gen 6 dmi.product.name: 21QT0016MX dmi.product.sku: LENOVO_MT_21QT_BU_Think_FM_ThinkPad P14s Gen 6 dmi.product.version: ThinkPad P14s Gen 6 dmi.sys.vendor: LENOVO To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165740/+subscriptions

[Bug 2165791] Re: linux 7.0.0-30-generic: list_del corruption + fatal Oops in ep_remove_safe/remove_wait_queue, triggered by the snapd 2.76 -> 2.76.3 upgrade

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165791 Title: linux 7.0.0-30-generic: list_del corruption + fatal Oops in ep_remove_safe/remove_wait_queue, triggered by the snapd 2.76 -> 2.76.3 upgrade Status in linux package in Ubuntu: New Bug description: Ubuntu 26.04, linux-image-7.0.0-30-generic 7.0.0-30.30, on a LENOVO 82VG (IdeaPad 1, BIOS KSCN31WW 05/09/2024), amd64. The machine hard-locked in the middle of an unattended-upgrades run. The crash reproduced a second time during manual recovery. Both crashes happened in snapd, in the epoll teardown path, and both were tied to the snapd 2.76 -> 2.76.3 package transition -- not to steady-state operation. TIMELINE Aug 20 14:25 kernel 7.0.0-30 installed. Ran 8 days, zero panics. Aug 28 05:45 unattended-upgrades starts a ~120 package transaction. Aug 28 05:47 term.log ends at "Setting up snapd (2.76.3+ubuntu26.04)". Machine dies here. history.log has no End-Date for this transaction. dpkg left 51 packages in iU and snapd in iF. Aug 28 11:03 Manual recovery: `dpkg --configure -a` reconfigures snapd. -> list_del corruption WARNING, kernel tainted G W. Aug 28 11:14 Reboot requested. snapd receives SIGTERM. -> fatal Oops, kdump captured a 252 MB vmcore. Aug 28 11:20+ Three subsequent clean shutdowns, plus a multi-hour memtest86+ run. No further crashes. /proc/sys/kernel/tainted back to 0. RAM was ruled out: memtest86+ passed clean, and the same kernel had already run 8 days without incident before snapd 2.76.3 arrived. The failure is deterministic and tied to the package transition, not random. FIRST EVENT (warning, kernel still alive) list_del corruption. next->prev should be ffff8f518099a328, but was 0000000000000000. (next=ffff8f4fa4969b50) WARNING: lib/list_debug.c:65 at __list_del_entry_valid_or_report+0xc0/0x10b, CPU#4: snapd/284305 Not tainted 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Call Trace: remove_wait_queue.cold+0x9/0x12 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 A second, symmetric warning followed immediately (lib/list_debug.c:62, prev->next should be ffff8f4f2004b068, but was 0000000000000000), i.e. both directions of the list entry had already been zeroed. SECOND EVENT (fatal, ~11 minutes later, on snapd termination) BUG: unable to handle page fault for address: ffffffff3226cb80 Oops: Oops: 0002 [#1] SMP NOPTI CPU: 3 UID: 0 PID: 284675 Comm: snapd Kdump: loaded Tainted: G W 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Hardware name: LENOVO 82VG/LNVNB161216, BIOS KSCN31WW 05/09/2024 RIP: 0010:native_queued_spin_lock_slowpath+0x2f5/0x370 Call Trace: __raw_spin_lock_irqsave+0x57/0x80 _raw_spin_lock_irqsave+0xe/0x20 remove_wait_queue+0x1b/0x80 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 ? __memcg_slab_free_hook+0x113/0x180 ? kmem_cache_free+0x266/0x3f0 ? __fput+0x1a2/0x2d0 ? fput_close_sync+0x40/0xc0 ? __x64_sys_close+0x3e/0x90 ANALYSIS The fatal trace shows epoll_ctl(EPOLL_CTL_DEL) racing a concurrent close() on the same descriptor: the kmem_cache_free / __fput / fput_close_sync / __x64_sys_close frames sit alongside the ep_remove_safe path. This matches the known eventpoll use-after-free shape, where ep_remove() drops file->f_ep under the lock but keeps using the file object, while a concurrent __fput() frees the struct eventpoll underneath it. The subsequent list operation then writes into freed memory -- which is exactly what the two list_debug warnings reported (both list pointers zeroed), and what the page fault at ffffffff3226cb80 in the spinlock slowpath is the consequence of. The zeroed pointers in the warning, and the fact that the machine survived 11 more minutes in a tainted state before dying on the next snapd termination, are consistent with memory that was freed and then reused. IMPACT Total loss of the machine mid-upgrade, with dpkg left in a broken state (51 packages unconfigured). Recovering required a second boot and a manual `dpkg --configure -a`, which itself re-triggered the bug. ATTACHED / AVAILABLE /var/crash/linux-image-7.0.0-30-generic-202608281115.crash (48 KB apport report, contains the full VmCoreDmesg, 1675 lines) /var/crash/202608281115/dump.202608281115 (252 MB vmcore, preserved, available on request) /var/crash/202608281115/dmesg.202608281115 (165 KB) WORKAROUND None applied to the kernel itself. 7.0.0-29-generic is still installed and available as a fallback from the GRUB advanced menu. unattended-upgrades was disabled locally right after the crash, then re-enabled the same day, once it was established that snapd cannot in fact be upgraded unattended on this system: the installed snapd (2.76.3+ubuntu26.04) comes from resolute-updates, which is not among Unattended-Upgrade::Allowed-Origins, and resolute-security only carries the older 2.76+ubuntu26.04.3. ADDITIONAL EVIDENCE (same day, after recovery) Steady-state snapd operation does NOT reproduce this bug. On the same kernel, after recovery, all of the following completed with no warning and no Oops: - 8x `snap remove --purge <snap>` - 23x `snap remove <snap> --revision=<rev>` - 2x `snap set system <option>` (daemon reconfiguration) - 1x `snap forget <snapshot set>` - 1 full clean reboot, i.e. snapd receiving SIGTERM at shutdown -- which is precisely the operation that produced the fatal Oops on Aug 28 11:14. /proc/sys/kernel/tainted stayed at 0 throughout, and the current boot logs contain no list_del, BUG: or Oops entries. This narrows the trigger considerably: the fault appears to require the package transition itself -- dpkg reconfiguring and restarting snapd across the 2.76 -> 2.76.3 version change -- rather than epoll teardown during ordinary snapd shutdown, which by itself is not sufficient to reproduce it. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165791/+subscriptions

[Bug 2092985] Re: UBSAN: Shift-Out-of-Bounds in soc-dapm.c (Linux Kernel 6.8.0 on Ubuntu 24.04)

Confirming this is still present on a much newer kernel, with additional detail that may help narrow down the cause. **System:** Lenovo N42-20 Chromebook, MrChromebox UEFI Full ROM firmware (coreboot), DMI board name "GOOGLE Reks" — same board as the original report. **Kernel:** 7.0.0-30-generic (Linux Mint, Ubuntu 24.04 base) — so this is not specific to the 6.8.x kernel series, it's still present in current mainline-derived kernels. **Symptom:** Audio plays normally for roughly 5-10 minutes (reproducible faster with `speaker-test -D hw:rt5650 -c 2 -t sine`, which typically fails within ~30 seconds), then playback locks into a continuous single tone that persists until the playing application is paused/stopped: ``` Write error: -5,Input/output error xrun_recovery failed: -5,Input/output error Transfer failed: Input/output error ``` **At boot**, before any playback issue occurs, dmesg shows the same UBSAN shift-out-of-bounds traces reported here, in `soc-dapm.c` during `cht-bsw-rt5645` / `snd_soc_sst_cht_bsw_rt5645` probe and during `alsactl`'s subsequent mixer writes: ``` UBSAN: shift-out-of-bounds in .../sound/soc/soc-dapm.c:429:15 shift exponent 16384 is too large for 32-bit type 'unsigned int' Call Trace: dapm_add_path.cold+0x1f/0x6c [snd_soc_core] snd_soc_dapm_add_route+0x373/0x660 [snd_soc_core] ... snd_cht_mc_probe+0x499/0x840 [snd_soc_sst_cht_bsw_rt5645] ``` plus further UBSAN hits in `snd_soc_dapm_put_volsw` triggered by `alsactl` via `snd_ctl_elem_write`/`snd_ctl_ioctl`. Also worth noting: the firmware/topology ABI versions don't match exactly: ``` sof-audio-acpi-intel-byt 808622A8:00: Firmware: ABI 3:22:1 Kernel ABI 3:23:1 sof-audio-acpi-intel-byt 808622A8:00: Topology: ABI 3:22:1 Kernel ABI 3:23:1 ``` This ABI drift between the kernel's SOF core and the shipped `sof-cht-rt5645.tplg`/firmware seems like a plausible source of the corrupted register value that produces the out-of-range shift, though I can't confirm that's the actual root cause. **Attempted workarounds (neither resolved it):** 1. `snd_intel_dspcfg.dsp_driver=1` kernel boot parameter to force the legacy (non-SOF) driver — on this kernel it's ignored outright: ``` intel_sst_acpi 808622A8:00: dsp_driver parameter 1 not supported, using automatic detection sof-audio-acpi-intel-byt 808622A8:00: dsp_driver parameter 1 not supported, using automatic detection ``` 2. Blacklisting all `snd_sof*` modules to force the legacy stack. This successfully prevents SOF from loading at all, and the RT5645 codec driver still probes fine over I2C (`rt5645 i2c-10EC5650:00: Detected Google Chrome platform`), but the `snd_soc_sst_cht_bsw_rt5645` machine driver never binds — even after manually `modprobe`-ing it, no sound card is created (`aplay -l` shows only HDMI outputs, no rt5650 card). It appears the platform device this driver needs is normally only instantiated as a byproduct of the SOF ACPI driver's own matching, and nothing else on this board creates it, so blacklisting SOF results in no usable audio device at all rather than falling back cleanly to the legacy stack. Given both workaround paths are closed off, this does look like it needs an actual fix in the topology/DAPM parsing path (or a DMI quirk enabling proper legacy fallback on this board) rather than something resolvable via boot parameters or module blacklisting. Happy to test patches, provide additional logs (`alsa-info.sh`, full `dmesg`, `acpidump`), or run bisection if that would help. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2092985 Title: UBSAN: Shift-Out-of-Bounds in soc-dapm.c (Linux Kernel 6.8.0 on Ubuntu 24.04) Status in linux package in Ubuntu: Confirmed Bug description: Bug Description: Summary: UBSAN detected a shift-out-of-bounds error in the Linux kernel source file sound/soc/soc-dapm.c at line 814. Issue Details: The code attempts a bit-shift operation with an exponent of 16384 on a 32-bit unsigned int type, which exceeds the maximum allowable range (0–31). This triggers undefined behavior and may result in unpredictable system behavior. Reproducibility: Consistently observed during boot initialization, specifically while udev-worker was running. Hardware: Google Reks/Reks (Chromebox BIOS MrChromebox-2408.1, dated 09/14/2024). Kernel Version: 6.8.0-51-generic #52-Ubuntu. Steps to Reproduce: 1) Boot a system with indicated Google Chromebook hardware and coreboot BIOS with Ubuntu LTS 24.04.1 and kernel version 6.8.0-51-generic. 2) Monitor dmesg logs for UBSAN warnings. Observed Behavior: The system logs the following error in dmesg: UBSAN: shift-out-of-bounds in /build/linux-vCyKs5/linux-6.8.0/sound/soc/soc-dapm.c:814:15 shift exponent 16384 is too large for 32-bit type 'unsigned int' Expected Behavior: No UBSAN warnings or undefined behavior in kernel operations during boot. Additional Information: Log Snippet: [ 14.206658] UBSAN: shift-out-of-bounds in /build/linux-vCyKs5/linux-6.8.0/sound/soc/soc-dapm.c:814:15 [ 14.206671] shift exponent 16384 is too large for 32-bit type 'unsigned int' [ 14.206678] CPU: 0 PID: 380 Comm: (udev-worker) Not tainted 6.8.0-51-generic #52-Ubuntu [ 14.206683] Hardware name: GOOGLE Reks/Reks, BIOS MrChromebox-2408.1 09/14/2024 Potential Impact: Undefined behavior in kernel modules can lead to system instability or incorrect operation. Suggested Fix: Review and modify the bit-shift logic in soc-dapm.c to ensure the shift exponent remains within the valid range for the data type. Consider masking or clamping the exponent to a value between 0 and 31 for 32-bit integers. ProblemType: Bug DistroRelease: Ubuntu 24.04 Package: linux-image-6.8.0-51-generic 6.8.0-51.52 ProcVersionSignature: Ubuntu 6.8.0-51.52-generic 6.8.12 Uname: Linux 6.8.0-51-generic x86_64 ApportVersion: 2.28.1-0ubuntu3.3 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/seq: chris 1567 F.... pipewire /dev/snd/controlC1: chris 1567 F.... pipewire chris 1570 F.... wireplumber CRDA: N/A CasperMD5CheckResult: unknown CurrentDesktop: LXQt Date: Sat Jan 4 08:43:20 2025 InstallationDate: Installed on 2024-12-23 (12 days ago) InstallationMedia: Lubuntu 24.04.1 LTS "Noble Numbat" - Release amd64 (20240827) Lsusb: Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 001 Device 002: ID 046d:c52f Logitech, Inc. Unifying Receiver Bus 001 Device 003: ID 0408:2040 Quanta Computer, Inc. Lenovo EasyCamera Bus 001 Device 004: ID 8087:0a2a Intel Corp. Bluetooth wireless interface Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub MachineType: GOOGLE Reks ProcFB: 0 i915drmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-6.8.0-51-generic root=UUID=a7cf1589-b7fe-4151-a70b-4ef90c746255 ro quiet splash vt.handoff=7 RelatedPackageVersions: linux-restricted-modules-6.8.0-51-generic N/A linux-backports-modules-6.8.0-51-generic N/A linux-firmware 20240318.git3b128b60-0ubuntu2.6 SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 09/14/2024 dmi.bios.release: 24.8 dmi.bios.vendor: coreboot dmi.bios.version: MrChromebox-2408.1 dmi.board.name: Reks dmi.board.vendor: GOOGLE dmi.board.version: 1.0 dmi.chassis.type: 9 dmi.chassis.vendor: GOOGLE dmi.ec.firmware.release: 0.0 dmi.modalias: dmi:bvncoreboot:bvrMrChromebox-2408.1:bd09/14/2024:br24.8:efr0.0:svnGOOGLE:pnReks:pvr1.0:rvnGOOGLE:rnReks:rvr1.0:cvnGOOGLE:ct9:cvr:sku: dmi.product.family: Intel_Strago dmi.product.name: Reks dmi.product.version: 1.0 dmi.sys.vendor: GOOGLE To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2092985/+subscriptions

[Bug 2163682] Re: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W

Second attachment: dmesg from the same 7.0.0-29.29 boot after unbinding and rebinding snd_hda_intel on 0000:02:00.1, with the GPU in D3cold. Same kernel, healthy. ** Attachment added: "dmesg from the same 7.0.0-29.29 boot after the HDA rebind, GPU in D3cold" https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+attachment/5996189/+files/dmesg-7.0.0-29-generic-after-hda-rebind.log -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2163682 Title: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W Status in linux package in Ubuntu: Incomplete Bug description: Ubuntu 26.04 LTS, desktop workstation (Dell Precision 3280 CFF) used headless as an LXD host. GPU: NVIDIA GB203GL [RTX PRO 4000 Blackwell SFF Edition] [10de:2c33] (rev a1) audio function [10de:22e9], upstream bridge 00:01.1 [8086:462d] Driver: nvidia-driver-595-open 595.84-0ubuntu0.26.04.1 (NVIDIA open kernel modules 595.84) NVreg_DynamicPowerManagement=0x02, nvidia-drm modeset=0 fbdev=0 With linux-image-7.0.0-28-generic (7.0.0-28.28) the dGPU cycles into runtime suspend normally. After booting linux-image-7.0.0-29-generic (7.0.0-29.29) the card stays in D0/active permanently with power/runtime_usage=1 and never suspends again. Package idle power measured at the wall rises from 10-14 W to about 22 W. The reference (runtime_usage=1) is held by the nvidia driver itself: it survives logging out of the GNOME session, unloading nvidia_drm and nvidia_modeset, and unloading all nvidia modules is worse still (no driver = no P8, +14 W). Cross-check isolating the kernel (same machine, same configuration): kernel nvidia result 7.0.0-28.28 595.71.05 suspends normally (140 and 172 wakeups over two boots) 7.0.0-29.29 595.84 never suspends (3 wakeups, all within the first 15 s of boot) 7.0.0-28.28 595.84 suspends normally <-- same driver as the failing case Last row is the decisive one: identical driver 595.84, identical module set (nvidia, nvidia_uvm, nvidia_modeset, nvidia_drm), identical configuration; only the kernel differs, and the GPU suspends again. The reverse pairing (7.0.0-29 with 595.71.05) is not testable, that driver version is no longer in the archive. Steps to reproduce: 1. Boot 7.0.0-29-generic with nvidia-driver-595-open and NVreg_DynamicPowerManagement=0x02, no CUDA workload, no display attached. 2. Leave the machine idle. 3. cat /sys/bus/pci/devices/0000:02:00.0/power_state -> D0 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_status -> active cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_usage -> 1 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_suspended_time -> frozen Expected (and observed on 7.0.0-28.28, ~150 s after boot): power_state D3cold, runtime_status suspended, runtime_usage 0, runtime_suspended_time 142242 ms vs runtime_active_time 8864 ms. Audio function 0000:02:00.1 also D3cold/suspended. Ruled out by measurement, not assumption: GUI session and its GPU clients (gnome-shell/Xwayland/nautilus/gnome-remote-desktop), KMS modules, CUDA/UVM context, VRAM threshold (memory.used 2 MiB against the 200 MB threshold), missing or changed config files, nvidia-persistenced, GPU containers (stopped), d3cold_allowed (1 on GPU, audio function and bridge), and running without the driver at all. The changelog between 7.0.0-28.28 and 7.0.0-29.29 shows no PCI-PM, D3cold, pcieport or ASPM change (CVE fixes plus one amdgpu HMM fix), so the regression is presumably a side effect rather than an intended change. Note on the attached apport data: it was collected while running the *working* kernel 7.0.0-28-generic, because the machine is remote and 7.0.0-29 costs the extra power. Happy to reboot into 7.0.0-29 and attach a second set on request. ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-28-generic 7.0.0-28.28 ProcVersionSignature: Ubuntu 7.0.0-28.28-generic 7.0.12 Uname: Linux 7.0.0-28-generic x86_64 NonfreeKernelModules: zfs ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC1: gdm-greeter 3041 F.... wireplumber /dev/snd/controlC0: gdm-greeter 3041 F.... wireplumber /dev/snd/seq: gdm-greeter 3003 F.... pipewire CasperMD5CheckResult: pass CurrentDesktop: ubuntu:GNOME Date: Mon Aug 17 20:25:43 2026 InstallationDate: Installed on 2026-07-05 (43 days ago) InstallationMedia: Ubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423.1) Lsusb: Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub Lsusb-t: /: Bus 001.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/16p, 480M /: Bus 002.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/9p, 20000M/x2 MachineType: Dell Inc. Precision 3280 Compact ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color ProcFB: ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-28-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro quiet crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M RfKill: SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 05/22/2026 dmi.bios.release: 1.24 dmi.bios.vendor: Dell Inc. dmi.bios.version: 1.24.1 dmi.board.name: 0H1DC6 dmi.board.vendor: Dell Inc. dmi.board.version: A00 dmi.chassis.type: 3 dmi.chassis.vendor: Dell Inc. dmi.ec.firmware.release: 1.16 dmi.modalias: dmi:bvnDellInc.:bvr1.24.1:bd05/22/2026:br1.24:efr1.16:svnDellInc.:pnPrecision3280Compact:pvr:rvnDellInc.:rn0H1DC6:rvrA00:cvnDellInc.:ct3:cvr:sku0C81:pfaPrecision: dmi.product.family: Precision dmi.product.name: Precision 3280 Compact dmi.product.sku: 0C81 dmi.sys.vendor: Dell Inc. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+subscriptions

[Bug 2163682] Re: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W

Follow-up, and a correction. First, what you asked for: attached is a dmesg captured on 7.0.0-29.29 with the GPU in the failing state. Second, and more important — while producing it I found that the isolation in my original report was confounded, and I now believe this bug is misattributed to the kernel. This machine has a second, intermittent defect. At boot, the NVIDIA HDA function 0000:02:00.1 sometimes fails its codec probe: snd_hda_intel 0000:02:00.1: azx_get_response timeout, switching to polling mode: last cmd=0x000f0000 snd_hda_intel 0000:02:00.1: Codec #0 probe error; disabling it... snd_hda_intel 0000:02:00.1: no codecs initialized snd_hda_intel 0000:02:00.1: GPU sound probed, but not operational: please add a quirk to driver_denylist When that happens, snd_hda_intel stays bound holding a runtime-PM usage reference it never drops: azx_probe_continue() takes the -ENXIO error path and never reaches the pm_runtime_use_autosuspend/allow/put_autosuspend block, while azx_probe() has already returned 0. Through the quirk_gpu_hda device link (DL_FLAG_PM_RUNTIME — "pci 0000:02:00.1: D0 power state depends on 0000:02:00.0") the reference propagates to the GPU function and pins it at D0/active/usage=1. That is exactly the symptom I reported. I re-checked every boot still held in the persistent journal: boot date kernel probe failed wakeups GPU outcome -7 2026-08-14 -29.29 yes 3 stuck at D0 for 44 h <- the boot this report is based on -6 2026-08-16 -28.28 no 32 healthy -5 2026-08-19 -28.28 no 13 healthy -4 2026-08-20 -28.28 no 5 healthy -3 2026-08-20 -28.28 yes 9 stuck at D0 -1 2026-08-28 -28.28 yes 2 not stuck 0 2026-08-28 -29.29 yes 4 stuck at D0, healthy after rebind Every boot I used as evidence for "-29.29 is broken" carries the probe failure. Every boot I used as evidence for "-28.28 is healthy" does not. Boot -3 shows the same failure producing the same D0 pin on 7.0.0-28.28, so the effect is not specific to -29.29 at all. Direct test today: I booted 7.0.0-29.29 and confirmed the GPU stuck at D0/active/usage=1 with three wakeups inside the first 17 seconds (this is the attached dmesg). I then unbound and rebound only the audio driver, changing nothing else: echo 0000:02:00.1 | sudo tee /sys/bus/pci/drivers/snd_hda_intel/unbind echo 0000:02:00.1 | sudo tee /sys/bus/pci/drivers/snd_hda_intel/bind The codec probed cleanly, and within about 90 seconds both functions reached D3cold — on the same running 7.0.0-29.29 kernel. It has sustained runtime suspend since (runtime_suspended_time 72 s -> 517 s, runtime_active_time frozen, wake counter steady). The second attachment is a dmesg from that healthy state on the same kernel. One thing I cannot explain and am not going to leave out: boot -1 had the probe failure and the GPU still suspended normally. So the failure does not deterministically pin the card. It appears to pin it only when it coincides with the GPU's first RTD3 suspend, which is consistent with the timing in the boots where it did pin. On the suspend/resume step you asked for: I have not performed one. This host is headless, s2idle has never been exercised on it, and the defect is PCI runtime PM (RTD3/D3cold) rather than ACPI system suspend — so I am not certain suspend/resume is the data you need. If it is, tell me and I will arrange physical access and run it. Given all of the above, please close this report as invalid if that is cleanest, or retitle it to the actual defect — snd_hda_intel leaving a runtime-PM reference after a failed codec probe on the NVIDIA HDA function, pinning the GPU at D0 through quirk_gpu_hda. I am glad to gather whatever logs help for the latter; the driver itself suggests a driver_denylist quirk in its own message. Apologies for sending you in the wrong direction. ** Attachment added: "dmesg from 7.0.0-29.29 with the GPU stuck at D0 (HDA codec probe failed on this boot)" https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+attachment/5996188/+files/dmesg-7.0.0-29-generic.log -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2163682 Title: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W Status in linux package in Ubuntu: Incomplete Bug description: Ubuntu 26.04 LTS, desktop workstation (Dell Precision 3280 CFF) used headless as an LXD host. GPU: NVIDIA GB203GL [RTX PRO 4000 Blackwell SFF Edition] [10de:2c33] (rev a1) audio function [10de:22e9], upstream bridge 00:01.1 [8086:462d] Driver: nvidia-driver-595-open 595.84-0ubuntu0.26.04.1 (NVIDIA open kernel modules 595.84) NVreg_DynamicPowerManagement=0x02, nvidia-drm modeset=0 fbdev=0 With linux-image-7.0.0-28-generic (7.0.0-28.28) the dGPU cycles into runtime suspend normally. After booting linux-image-7.0.0-29-generic (7.0.0-29.29) the card stays in D0/active permanently with power/runtime_usage=1 and never suspends again. Package idle power measured at the wall rises from 10-14 W to about 22 W. The reference (runtime_usage=1) is held by the nvidia driver itself: it survives logging out of the GNOME session, unloading nvidia_drm and nvidia_modeset, and unloading all nvidia modules is worse still (no driver = no P8, +14 W). Cross-check isolating the kernel (same machine, same configuration): kernel nvidia result 7.0.0-28.28 595.71.05 suspends normally (140 and 172 wakeups over two boots) 7.0.0-29.29 595.84 never suspends (3 wakeups, all within the first 15 s of boot) 7.0.0-28.28 595.84 suspends normally <-- same driver as the failing case Last row is the decisive one: identical driver 595.84, identical module set (nvidia, nvidia_uvm, nvidia_modeset, nvidia_drm), identical configuration; only the kernel differs, and the GPU suspends again. The reverse pairing (7.0.0-29 with 595.71.05) is not testable, that driver version is no longer in the archive. Steps to reproduce: 1. Boot 7.0.0-29-generic with nvidia-driver-595-open and NVreg_DynamicPowerManagement=0x02, no CUDA workload, no display attached. 2. Leave the machine idle. 3. cat /sys/bus/pci/devices/0000:02:00.0/power_state -> D0 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_status -> active cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_usage -> 1 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_suspended_time -> frozen Expected (and observed on 7.0.0-28.28, ~150 s after boot): power_state D3cold, runtime_status suspended, runtime_usage 0, runtime_suspended_time 142242 ms vs runtime_active_time 8864 ms. Audio function 0000:02:00.1 also D3cold/suspended. Ruled out by measurement, not assumption: GUI session and its GPU clients (gnome-shell/Xwayland/nautilus/gnome-remote-desktop), KMS modules, CUDA/UVM context, VRAM threshold (memory.used 2 MiB against the 200 MB threshold), missing or changed config files, nvidia-persistenced, GPU containers (stopped), d3cold_allowed (1 on GPU, audio function and bridge), and running without the driver at all. The changelog between 7.0.0-28.28 and 7.0.0-29.29 shows no PCI-PM, D3cold, pcieport or ASPM change (CVE fixes plus one amdgpu HMM fix), so the regression is presumably a side effect rather than an intended change. Note on the attached apport data: it was collected while running the *working* kernel 7.0.0-28-generic, because the machine is remote and 7.0.0-29 costs the extra power. Happy to reboot into 7.0.0-29 and attach a second set on request. ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-28-generic 7.0.0-28.28 ProcVersionSignature: Ubuntu 7.0.0-28.28-generic 7.0.12 Uname: Linux 7.0.0-28-generic x86_64 NonfreeKernelModules: zfs ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC1: gdm-greeter 3041 F.... wireplumber /dev/snd/controlC0: gdm-greeter 3041 F.... wireplumber /dev/snd/seq: gdm-greeter 3003 F.... pipewire CasperMD5CheckResult: pass CurrentDesktop: ubuntu:GNOME Date: Mon Aug 17 20:25:43 2026 InstallationDate: Installed on 2026-07-05 (43 days ago) InstallationMedia: Ubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423.1) Lsusb: Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub Lsusb-t: /: Bus 001.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/16p, 480M /: Bus 002.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/9p, 20000M/x2 MachineType: Dell Inc. Precision 3280 Compact ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color ProcFB: ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-28-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro quiet crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M RfKill: SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 05/22/2026 dmi.bios.release: 1.24 dmi.bios.vendor: Dell Inc. dmi.bios.version: 1.24.1 dmi.board.name: 0H1DC6 dmi.board.vendor: Dell Inc. dmi.board.version: A00 dmi.chassis.type: 3 dmi.chassis.vendor: Dell Inc. dmi.ec.firmware.release: 1.16 dmi.modalias: dmi:bvnDellInc.:bvr1.24.1:bd05/22/2026:br1.24:efr1.16:svnDellInc.:pnPrecision3280Compact:pvr:rvnDellInc.:rn0H1DC6:rvrA00:cvnDellInc.:ct3:cvr:sku0C81:pfaPrecision: dmi.product.family: Precision dmi.product.name: Precision 3280 Compact dmi.product.sku: 0C81 dmi.sys.vendor: Dell Inc. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+subscriptions

[Bug 2165140] Re: qrtr: ns: node limit of 64 breaks QRTR routing on large multi-node deployments

** Description changed: SRU Justification: - [ Impact ] + [Impact] -  * Commit 27d5e84e810b ("net: qrtr: ns: Limit the total number of nodes") -    introduced a hard cap of 64 on the number of QRTR nodes the namespace -    server (qrtr_ns) will track, including the host node itself. -    That limit was set as a DoS mitigation but is too low for legitimate -    large-scale deployments, such as Qualcomm AI200 systems, which register -    up to 384 QRTR nodes across their SoCs/cards. + * Commit 27d5e84e810b ("net: qrtr: ns: Limit the total number of nodes") + introduced a hard cap of 64 on the number of QRTR nodes the namespace + server (qrtr_ns) will track, including the host node itself. + That limit was set as a DoS mitigation but is too low for legitimate + large-scale deployments, such as Qualcomm AI200 systems, which register + up to 384 QRTR nodes across their SoCs/cards. -  * This patch fixes this issue by increasing the limit to 512. + * This patch fixes this issue by increasing the limit to 512. -  * This patch should be backported because it affects deployments with 64 -    or more Qualcomm AI100/AI200 targets. - - [ Test Plan ] - -  * Checking dmesg on an Ubuntu 24.04 system with at least 6.8.0-136-generic, -    having at least 64 remote QRTR nodes, it will show "QRTR clients exceed max -    node limit!" for every extra node. -  * Also, by running a QRTR client then sending a QRTR_TYPE_NEW_LOOKUP packet, -    only 64 nodes will show up (63 remote node + the local node). -  * After applying the fix, dmesg should no longer show the error message, and -    the lookup client will return a full list of nodes. + * This patch should be backported because it affects deployments with 64 + or more Qualcomm AI100/AI200 targets. [Fix] -  * Cherry-pick upstream commit: -    ff194cffd586 ("net: qrtr: ns: Raise node count limit to 512") + * Backport of upstream commit: + ff194cffd586 ("net: qrtr: ns: Raise node count limit to 512") - [ Where problems could occur ] + [Test Plan] -  * If a deployment exceeds 512 nodes, the original symptom returns identically -    (same log message, same rejection behavior). + * Checking dmesg on an Ubuntu 24.04 system with at least 6.8.0-136-generic, + having at least 64 remote QRTR nodes, it will show "QRTR clients exceed max + node limit!" for every extra node. + * Also, by running a QRTR client then sending a QRTR_TYPE_NEW_LOOKUP packet, + only 64 nodes will show up (63 remote node + the local node). + * After applying the fix, dmesg should no longer show the error message, and + the lookup client will return a full list of nodes. + + [Where problems could occur] + + * If a deployment exceeds 512 nodes, the original symptom returns identically + (same log message, same rejection behavior). -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165140 Title: qrtr: ns: node limit of 64 breaks QRTR routing on large multi-node deployments Status in linux package in Ubuntu: New Bug description: SRU Justification: [Impact] * Commit 27d5e84e810b ("net: qrtr: ns: Limit the total number of nodes") introduced a hard cap of 64 on the number of QRTR nodes the namespace server (qrtr_ns) will track, including the host node itself. That limit was set as a DoS mitigation but is too low for legitimate large-scale deployments, such as Qualcomm AI200 systems, which register up to 384 QRTR nodes across their SoCs/cards. * This patch fixes this issue by increasing the limit to 512. * This patch should be backported because it affects deployments with 64 or more Qualcomm AI100/AI200 targets. [Fix] * Backport of upstream commit: ff194cffd586 ("net: qrtr: ns: Raise node count limit to 512") [Test Plan] * Checking dmesg on an Ubuntu 24.04 system with at least 6.8.0-136-generic, having at least 64 remote QRTR nodes, it will show "QRTR clients exceed max node limit!" for every extra node. * Also, by running a QRTR client then sending a QRTR_TYPE_NEW_LOOKUP packet, only 64 nodes will show up (63 remote node + the local node). * After applying the fix, dmesg should no longer show the error message, and the lookup client will return a full list of nodes. [Where problems could occur] * If a deployment exceeds 512 nodes, the original symptom returns identically (same log message, same rejection behavior). To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165140/+subscriptions

[Bug 2165791] [NEW] linux 7.0.0-30-generic: list_del corruption + fatal Oops in ep_remove_safe/remove_wait_queue, triggered by the snapd 2.76 -> 2.76.3 upgrade

Public bug reported: Ubuntu 26.04, linux-image-7.0.0-30-generic 7.0.0-30.30, on a LENOVO 82VG (IdeaPad 1, BIOS KSCN31WW 05/09/2024), amd64. The machine hard-locked in the middle of an unattended-upgrades run. The crash reproduced a second time during manual recovery. Both crashes happened in snapd, in the epoll teardown path, and both were tied to the snapd 2.76 -> 2.76.3 package transition -- not to steady-state operation. TIMELINE Aug 20 14:25 kernel 7.0.0-30 installed. Ran 8 days, zero panics. Aug 28 05:45 unattended-upgrades starts a ~120 package transaction. Aug 28 05:47 term.log ends at "Setting up snapd (2.76.3+ubuntu26.04)". Machine dies here. history.log has no End-Date for this transaction. dpkg left 51 packages in iU and snapd in iF. Aug 28 11:03 Manual recovery: `dpkg --configure -a` reconfigures snapd. -> list_del corruption WARNING, kernel tainted G W. Aug 28 11:14 Reboot requested. snapd receives SIGTERM. -> fatal Oops, kdump captured a 252 MB vmcore. Aug 28 11:20+ Three subsequent clean shutdowns, plus a multi-hour memtest86+ run. No further crashes. /proc/sys/kernel/tainted back to 0. RAM was ruled out: memtest86+ passed clean, and the same kernel had already run 8 days without incident before snapd 2.76.3 arrived. The failure is deterministic and tied to the package transition, not random. FIRST EVENT (warning, kernel still alive) list_del corruption. next->prev should be ffff8f518099a328, but was 0000000000000000. (next=ffff8f4fa4969b50) WARNING: lib/list_debug.c:65 at __list_del_entry_valid_or_report+0xc0/0x10b, CPU#4: snapd/284305 Not tainted 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Call Trace: remove_wait_queue.cold+0x9/0x12 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 A second, symmetric warning followed immediately (lib/list_debug.c:62, prev->next should be ffff8f4f2004b068, but was 0000000000000000), i.e. both directions of the list entry had already been zeroed. SECOND EVENT (fatal, ~11 minutes later, on snapd termination) BUG: unable to handle page fault for address: ffffffff3226cb80 Oops: Oops: 0002 [#1] SMP NOPTI CPU: 3 UID: 0 PID: 284675 Comm: snapd Kdump: loaded Tainted: G W 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Hardware name: LENOVO 82VG/LNVNB161216, BIOS KSCN31WW 05/09/2024 RIP: 0010:native_queued_spin_lock_slowpath+0x2f5/0x370 Call Trace: __raw_spin_lock_irqsave+0x57/0x80 _raw_spin_lock_irqsave+0xe/0x20 remove_wait_queue+0x1b/0x80 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 ? __memcg_slab_free_hook+0x113/0x180 ? kmem_cache_free+0x266/0x3f0 ? __fput+0x1a2/0x2d0 ? fput_close_sync+0x40/0xc0 ? __x64_sys_close+0x3e/0x90 ANALYSIS The fatal trace shows epoll_ctl(EPOLL_CTL_DEL) racing a concurrent close() on the same descriptor: the kmem_cache_free / __fput / fput_close_sync / __x64_sys_close frames sit alongside the ep_remove_safe path. This matches the known eventpoll use-after-free shape, where ep_remove() drops file->f_ep under the lock but keeps using the file object, while a concurrent __fput() frees the struct eventpoll underneath it. The subsequent list operation then writes into freed memory -- which is exactly what the two list_debug warnings reported (both list pointers zeroed), and what the page fault at ffffffff3226cb80 in the spinlock slowpath is the consequence of. The zeroed pointers in the warning, and the fact that the machine survived 11 more minutes in a tainted state before dying on the next snapd termination, are consistent with memory that was freed and then reused. IMPACT Total loss of the machine mid-upgrade, with dpkg left in a broken state (51 packages unconfigured). Recovering required a second boot and a manual `dpkg --configure -a`, which itself re-triggered the bug. ATTACHED / AVAILABLE /var/crash/linux-image-7.0.0-30-generic-202608281115.crash (48 KB apport report, contains the full VmCoreDmesg, 1675 lines) /var/crash/202608281115/dump.202608281115 (252 MB vmcore, preserved, available on request) /var/crash/202608281115/dmesg.202608281115 (165 KB) WORKAROUND None applied to the kernel itself. 7.0.0-29-generic is still installed and available as a fallback from the GRUB advanced menu. unattended-upgrades was disabled locally right after the crash, then re-enabled the same day, once it was established that snapd cannot in fact be upgraded unattended on this system: the installed snapd (2.76.3+ubuntu26.04) comes from resolute-updates, which is not among Unattended-Upgrade::Allowed-Origins, and resolute-security only carries the older 2.76+ubuntu26.04.3. ADDITIONAL EVIDENCE (same day, after recovery) Steady-state snapd operation does NOT reproduce this bug. On the same kernel, after recovery, all of the following completed with no warning and no Oops: - 8x `snap remove --purge <snap>` - 23x `snap remove <snap> --revision=<rev>` - 2x `snap set system <option>` (daemon reconfiguration) - 1x `snap forget <snapshot set>` - 1 full clean reboot, i.e. snapd receiving SIGTERM at shutdown -- which is precisely the operation that produced the fatal Oops on Aug 28 11:14. /proc/sys/kernel/tainted stayed at 0 throughout, and the current boot logs contain no list_del, BUG: or Oops entries. This narrows the trigger considerably: the fault appears to require the package transition itself -- dpkg reconfiguring and restarting snapd across the 2.76 -> 2.76.3 version change -- rather than epoll teardown during ordinary snapd shutdown, which by itself is not sufficient to reproduce it. ** Affects: linux (Ubuntu) Importance: Undecided Status: New ** Attachment added: "apport report with full VmCoreDmesg, kernel 7.0.0-30 panic" https://bugs.launchpad.net/bugs/2165791/+attachment/5996187/+files/linux-image-7.0.0-30-generic-202608281115.crash -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2165791 Title: linux 7.0.0-30-generic: list_del corruption + fatal Oops in ep_remove_safe/remove_wait_queue, triggered by the snapd 2.76 -> 2.76.3 upgrade Status in linux package in Ubuntu: New Bug description: Ubuntu 26.04, linux-image-7.0.0-30-generic 7.0.0-30.30, on a LENOVO 82VG (IdeaPad 1, BIOS KSCN31WW 05/09/2024), amd64. The machine hard-locked in the middle of an unattended-upgrades run. The crash reproduced a second time during manual recovery. Both crashes happened in snapd, in the epoll teardown path, and both were tied to the snapd 2.76 -> 2.76.3 package transition -- not to steady-state operation. TIMELINE Aug 20 14:25 kernel 7.0.0-30 installed. Ran 8 days, zero panics. Aug 28 05:45 unattended-upgrades starts a ~120 package transaction. Aug 28 05:47 term.log ends at "Setting up snapd (2.76.3+ubuntu26.04)". Machine dies here. history.log has no End-Date for this transaction. dpkg left 51 packages in iU and snapd in iF. Aug 28 11:03 Manual recovery: `dpkg --configure -a` reconfigures snapd. -> list_del corruption WARNING, kernel tainted G W. Aug 28 11:14 Reboot requested. snapd receives SIGTERM. -> fatal Oops, kdump captured a 252 MB vmcore. Aug 28 11:20+ Three subsequent clean shutdowns, plus a multi-hour memtest86+ run. No further crashes. /proc/sys/kernel/tainted back to 0. RAM was ruled out: memtest86+ passed clean, and the same kernel had already run 8 days without incident before snapd 2.76.3 arrived. The failure is deterministic and tied to the package transition, not random. FIRST EVENT (warning, kernel still alive) list_del corruption. next->prev should be ffff8f518099a328, but was 0000000000000000. (next=ffff8f4fa4969b50) WARNING: lib/list_debug.c:65 at __list_del_entry_valid_or_report+0xc0/0x10b, CPU#4: snapd/284305 Not tainted 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Call Trace: remove_wait_queue.cold+0x9/0x12 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 A second, symmetric warning followed immediately (lib/list_debug.c:62, prev->next should be ffff8f4f2004b068, but was 0000000000000000), i.e. both directions of the list entry had already been zeroed. SECOND EVENT (fatal, ~11 minutes later, on snapd termination) BUG: unable to handle page fault for address: ffffffff3226cb80 Oops: Oops: 0002 [#1] SMP NOPTI CPU: 3 UID: 0 PID: 284675 Comm: snapd Kdump: loaded Tainted: G W 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy) Hardware name: LENOVO 82VG/LNVNB161216, BIOS KSCN31WW 05/09/2024 RIP: 0010:native_queued_spin_lock_slowpath+0x2f5/0x370 Call Trace: __raw_spin_lock_irqsave+0x57/0x80 _raw_spin_lock_irqsave+0xe/0x20 remove_wait_queue+0x1b/0x80 ep_remove_safe+0x3c/0xe0 do_epoll_ctl+0x599/0x840 __x64_sys_epoll_ctl+0x6c/0xb0 x64_sys_call+0x1be6/0x2390 do_syscall_64+0x105/0x5a0 ? __memcg_slab_free_hook+0x113/0x180 ? kmem_cache_free+0x266/0x3f0 ? __fput+0x1a2/0x2d0 ? fput_close_sync+0x40/0xc0 ? __x64_sys_close+0x3e/0x90 ANALYSIS The fatal trace shows epoll_ctl(EPOLL_CTL_DEL) racing a concurrent close() on the same descriptor: the kmem_cache_free / __fput / fput_close_sync / __x64_sys_close frames sit alongside the ep_remove_safe path. This matches the known eventpoll use-after-free shape, where ep_remove() drops file->f_ep under the lock but keeps using the file object, while a concurrent __fput() frees the struct eventpoll underneath it. The subsequent list operation then writes into freed memory -- which is exactly what the two list_debug warnings reported (both list pointers zeroed), and what the page fault at ffffffff3226cb80 in the spinlock slowpath is the consequence of. The zeroed pointers in the warning, and the fact that the machine survived 11 more minutes in a tainted state before dying on the next snapd termination, are consistent with memory that was freed and then reused. IMPACT Total loss of the machine mid-upgrade, with dpkg left in a broken state (51 packages unconfigured). Recovering required a second boot and a manual `dpkg --configure -a`, which itself re-triggered the bug. ATTACHED / AVAILABLE /var/crash/linux-image-7.0.0-30-generic-202608281115.crash (48 KB apport report, contains the full VmCoreDmesg, 1675 lines) /var/crash/202608281115/dump.202608281115 (252 MB vmcore, preserved, available on request) /var/crash/202608281115/dmesg.202608281115 (165 KB) WORKAROUND None applied to the kernel itself. 7.0.0-29-generic is still installed and available as a fallback from the GRUB advanced menu. unattended-upgrades was disabled locally right after the crash, then re-enabled the same day, once it was established that snapd cannot in fact be upgraded unattended on this system: the installed snapd (2.76.3+ubuntu26.04) comes from resolute-updates, which is not among Unattended-Upgrade::Allowed-Origins, and resolute-security only carries the older 2.76+ubuntu26.04.3. ADDITIONAL EVIDENCE (same day, after recovery) Steady-state snapd operation does NOT reproduce this bug. On the same kernel, after recovery, all of the following completed with no warning and no Oops: - 8x `snap remove --purge <snap>` - 23x `snap remove <snap> --revision=<rev>` - 2x `snap set system <option>` (daemon reconfiguration) - 1x `snap forget <snapshot set>` - 1 full clean reboot, i.e. snapd receiving SIGTERM at shutdown -- which is precisely the operation that produced the fatal Oops on Aug 28 11:14. /proc/sys/kernel/tainted stayed at 0 throughout, and the current boot logs contain no list_del, BUG: or Oops entries. This narrows the trigger considerably: the fault appears to require the package transition itself -- dpkg reconfiguring and restarting snapd across the 2.76 -> 2.76.3 version change -- rather than epoll teardown during ordinary snapd shutdown, which by itself is not sufficient to reproduce it. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165791/+subscriptions