суббота

[Bug 2168627] Re: PC does not shut down completely

I just tried to choose and older kernel in Grub: 7.0.0-22-generic Here shutdown and restart works fine. So I assume with the kernel -3x there was a regression. I hope this helps. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168627 Title: PC does not shut down completely Status in linux package in Ubuntu: New Bug description: When I want to Shut down or Restart my PC, it manages to close all files gracefully, unmounts everything nicely, switches off the monitors but the last step the actual shut down never happens. For a brief moment the fans of my graphics card stop rotating, but then they spin up again the lights stay on and also the CPU fan starts to spin and from that moment on nothing happens anymore. I need to press the power button for 5 seconds to completely shut down ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-34-generic 7.0.0-34.34 ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14 Uname: Linux 7.0.0-34-generic x86_64 ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 CasperMD5CheckResult: unknown CurrentDesktop: KDE Date: Sat Sep 26 17:48:46 2026 InstallationDate: Installed on 2026-01-22 (247 days ago) InstallationMedia: Kubuntu 25.10 "Questing Quokka" - Release amd64 (20251007) MachineType: ASRock B650 Steel Legend WiFi ProcFB: 0 amdgpudrmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-34-generic root=UUID=120d685f-76a9-41ec-925e-105e9c80ca3a ro quiet splash amdgpu.ppfeaturemask=0xffffffff PulseList: Error: command ['pacmd', 'list'] failed with exit code 1: No PulseAudio daemon running, or not running as session daemon. SourcePackage: linux UpgradeStatus: Upgraded to resolute on 2026-05-03 (146 days ago) dmi.bios.date: 06/24/2026 dmi.bios.release: 5.41 dmi.bios.vendor: American Megatrends International, LLC. dmi.bios.version: 4.43 dmi.board.asset.tag: Default string dmi.board.name: B650 Steel Legend WiFi dmi.board.vendor: ASRock dmi.board.version: Default string dmi.chassis.asset.tag: Default string dmi.chassis.type: 3 dmi.chassis.vendor: Default string dmi.chassis.version: Default string dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvr4.43:bd06/24/2026:br5.41:svnASRock:pnB650SteelLegendWiFi:pvrDefaultstring:rvnASRock:rnB650SteelLegendWiFi:rvrDefaultstring:cvnDefaultstring:ct3:cvrDefaultstring:skuDefaultstring:pfaDefaultstring: dmi.product.family: Default string dmi.product.name: B650 Steel Legend WiFi dmi.product.sku: Default string dmi.product.version: Default string dmi.sys.vendor: ASRock To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168627/+subscriptions

[Bug 2168633] Re: [Lenovo Yoga Air 14 (83QK)] No sound card at all on 26.04 kernel 7.0.0-34 — missing "soundwire: dmi-quirks: Disable ghost Realtek devices" backport

** Tags added: patch -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168633 Title: [Lenovo Yoga Air 14 (83QK)] No sound card at all on 26.04 kernel 7.0.0-34 — missing "soundwire: dmi-quirks: Disable ghost Realtek devices" backport Status in linux package in Ubuntu: New Bug description: Summary On a Lenovo Yoga Air 14 Ultra IPH11 (DMI product_name = 83QK) the Ubuntu 26.04 kernel 7.0.0-34-generic registers no sound card at all: $ cat /proc/asound/cards --- no soundcards --- The same machine works correctly with mainline 7.2.6 and with a locally built vanilla v7.2-rc1. The difference is a single missing upstream patch: soundwire: dmi-quirks: Disable ghost Realtek devices Ubuntu has already backported the four sibling patches from the same PTL/SoundWire audio series (table below); this one appears to have been missed. Environment - Ubuntu 26.04.1 LTS (resolute) - linux-image-7.0.0-34-generic 7.0.0-34.34 (the quirk is absent from the whole 7.0.0 series - verified by strings on the decompressed soundwire-intel.ko of both 7.0.0-15 and 7.0.0-34, and the 7.0.0 series changelog never mentions the patch) - alsa-ucm-conf 1.2.15.3-1ubuntu1.5 - Lenovo YOGA Air 14 Ultra IPH11 sys_vendor = LENOVO product_name = 83QK product_version = YOGA Air 14 Ultra IPH11 board_name = LNVNB161216 BIOS = SHCN32WW (2026-04-20) - ACPI sound card id: LENOVO-83QK-YOGAAir14UltraIPH11-LNVNB161216 - Codecs: Cirrus Logic CS42L43 + 2x CS35L57 (SoundWire sidecar amps) Steps to reproduce 1. Boot 7.0.0-34-generic on this machine. 2. cat /proc/asound/cards 3. Play anything. Actual result No sound card is registered. Kernel log: sof_sdw sof_sdw: cs42l43 speaker volume limit failed: -22 sof_sdw sof_sdw: ASoC: Failed to add route cs42l43 AMP1_OUT_P(*) -> Speaker sof_sdw sof_sdw: ASoC: Failed to add route cs42l43 AMP1_OUT_N(*) -> Speaker sof_sdw sof_sdw: ASoC: Failed to add route cs42l43 AMP2_OUT_P(*) -> Speaker sof_sdw sof_sdw: ASoC: Failed to add route cs42l43 AMP2_OUT_N(*) -> Speaker sof_sdw sof_sdw: cs42l43 speaker map addition failed: -19 SDW3-Playback-SmartAmp: ASoC error (-19): at snd_soc_link_init() on SDW3-Playback-SmartAmp sof_sdw sof_sdw: error -ENODEV: snd_soc_register_card failed -19 Audio applications only see a dummy output. Expected result The card registers and the Speaker PCM (card 0, device 2) appears, as on mainline 7.2.6: $ aplay -l | grep -i Speaker card 0: sofsoundwire [sof-soundwire], device 2: Speaker (*) [] Root cause This machine's ACPI advertises a Realtek RT722 SoundWire device (ADR 0x33f025d072201, mfg id 0x025d, part id 0x0722) that does not physically exist - a "ghost" device. Upstream handles this in drivers/soundwire/dmi-quirks.c with an adr_remap entry (ghost_realtek) that remaps the bogus ADR to 0. The hook is soundwire-intel's .override_adr = sdw_dmi_override_adr (drivers/soundwire/intel_auxdevice.c). Without that quirk the ghost device is enumerated, which perturbs the codec name prefix, and the hard-coded speaker routes in sound/soc/sdw_utils/soc_sdw_cs42l43.c ("cs42l43 AMP1_OUT_P", ...) no longer resolve, so snd_soc_register_card() fails with -ENODEV. Boot log comparison, same machine / same userspace / only the kernel differs: 7.2.6 (works) 7.0.0-34 (broken) endpoint count Found 1 devices with 1 endpoints Found 2 devices with 4 endpoints cs42l43 prefix Adding prefix cs42l43 Adding prefix cs42l43-1 ghost device absent Add dev: 3, 0x33f025d072201 card registration OK snd_soc_register_card failed -19 Broken boot, relevant excerpt: snd_soc_acpi: Slave part_id 0x712 not found snd_sof_intel_hda_generic: soundwire sdw:0:3:01fa:4243:01: Endpoint DAI type 0 not found snd_soc_sdw_utils: asoc_sdw_count_sdw_endpoints: Found 2 devices with 4 endpoints snd_soc_sdw_utils: asoc_sdw_parse_sdw_endpoints: Adding prefix cs42l43-1 for cs42l43-codec snd_soc_sdw_utils: asoc_sdw_parse_sdw_endpoints: Add dev: 3, 0x33f025d072201 end: 0, dai: 0, P/C to solo: 0 snd_soc_sdw_utils: asoc_sdw_parse_sdw_endpoints: Adding prefix rt722 for sdw:0:3:025d:0722:01 Also verified directly in the shipped binary: strings on the decompressed soundwire-intel.ko from linux-modules-7.0.0-34-generic contains zero occurrences of 83QK (nor 83SF, nor UX5406AA), while the 7.2.6 module does. The fix already exists upstream - Patch title: soundwire: dmi-quirks: Disable ghost Realtek devices - Author: Charles Keepax <ckeepax@opensource.cirrus.com> - Posted 2026-05-20 as patch 3/3 of "Update some topology matching for newer laptops": https://lore.kernel.org/linux-sound/20260520163631.3300102-1-ckeepax@opensource.cirrus.com/ - The commit message lists this exact machine: "Currently this patch should cover: Asus UX5406AA, Lenovo Yoga Pro 9i (83SF), Lenovo Yoga Slim 7 Ultra (83QK)" - It is present in v7.2-rc1 and later, and has been AUTOSEL'd to the stable trees (6.18 down to 6.1), i.e. it is regarded as a safe backport. Sibling patches Ubuntu already carries in the 7.0.0 series ASoC: SOF: Intel: Add a is_amp flag to fix the wrong name prefix 7.0.0-26 ASoC: sdw_utils: add rt1320 and rt1321 dmic dai in codec_info_list 7.0.0-26 ASoC: Intel: sof_sdw: append dai type to dai link name unconditionally 7.0.0-28 Input: atkbd - add DMI quirk for Lenovo Yoga Air 14 (83QK) 7.0.0-31 soundwire: dmi-quirks: Disable ghost Realtek devices MISSING (all confirmed from apt-get changelog linux-modules-7.0.0-34-generic) Requested fix Please backport "soundwire: dmi-quirks: Disable ghost Realtek devices" to the 26.04 kernel (linux, 7.0.0 series). It is a ~35 line, self-contained addition to an existing DMI quirk table, already accepted into upstream stable, and it is the only missing piece for audio on this machine. How to verify Boot a kernel with the patch and check: cat /proc/asound/cards # "sof-soundwire" card appears aplay -l | grep -i Speaker # card 0, device 2 journalctl -b -k | grep -i "Failed to add route" # no output Note: userspace also needs the cs35l56-bridge.conf UCM fix from thesofproject/alsa-ucm-conf#778, which is likewise not yet in Ubuntu's alsa-ucm-conf 1.2.15.3 (it hard-codes PlaybackPCM "hw:${CardId},0" and the DP5RX route; this machine needs hw:${CardId},2 and DP6RX). Worth backporting together. References - Upstream issue: https://github.com/thesofproject/linux/issues/5786 - Upstream patch: https://lore.kernel.org/linux-sound/20260520163631.3300102-1-ckeepax@opensource.cirrus.com/ - Upstream UCM fix: https://github.com/thesofproject/alsa-ucm-conf/pull/778 ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-34-generic 7.0.0-34.34 ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14 Uname: Linux 7.0.0-34-generic x86_64 ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/seq: eureka 2414 F.... pipewire CasperMD5CheckResult: unknown CurrentDesktop: ubuntu:GNOME Date: Sun Sep 27 00:44:54 2026 InstallationDate: Installed on 2026-05-26 (123 days ago) InstallationMedia: Ubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423.1) Lsusb: Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub Bus 003 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 003 Device 002: ID 174f:11bf Syntek Integrated RGB Camera Bus 004 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub MachineType: LENOVO 83QK ProcEnviron: LANG=zh_CN.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color XDG_RUNTIME_DIR=<set> ProcFB: 0 xedrmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-34-generic root=UUID=1cbf22a7-f458-46ce-ba8e-7a7cfda504a0 ro quiet splash i8042.unlock atkbd.skip_deactivate=1 crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M PulseList: Error: command ['pacmd', 'list'] failed with exit code 1: No PulseAudio daemon running, or not running as session daemon. SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 04/20/2026 dmi.bios.release: 1.32 dmi.bios.vendor: LENOVO dmi.bios.version: SHCN32WW dmi.board.asset.tag: NO Asset Tag dmi.board.name: LNVNB161216 dmi.board.vendor: LENOVO dmi.board.version: SDK0T76577 WIN dmi.chassis.asset.tag: NO Asset Tag dmi.chassis.type: 10 dmi.chassis.vendor: LENOVO dmi.chassis.version: YOGA Air 14 Ultra IPH11 dmi.ec.firmware.release: 1.32 dmi.modalias: dmi:bvnLENOVO:bvrSHCN32WW:bd04/20/2026:br1.32:efr1.32:svnLENOVO:pn83QK:pvrYOGAAir14UltraIPH11:rvnLENOVO:rnLNVNB161216:rvrSDK0T76577WIN:cvnLENOVO:ct10:cvrYOGAAir14UltraIPH11:skuLENOVO_MT_83QK_BU_idea_FM_YOGAAir14UltraIPH11:pfaYOGAAir14UltraIPH11: dmi.product.family: YOGA Air 14 Ultra IPH11 dmi.product.name: 83QK dmi.product.sku: LENOVO_MT_83QK_BU_idea_FM_YOGA Air 14 Ultra IPH11 dmi.product.version: YOGA Air 14 Ultra IPH11 dmi.sys.vendor: LENOVO To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168633/+subscriptions

[Bug 2167549] Re: ThinkPad P1 Gen 7 Meteor Lake — i915 PHY/DPLL failure after s2idle resume

This commit fixes the issue. https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git/commit/drivers/gpu/drm/i915?id=062499cc4813b5a3cbed5dd4fbe0177265858450 drm/i915/mtl+: Enable PPS before PLL Enabling PPS after a display port's PLL is enabled leads to PLL / DDI BUF timeouts during system resuming after a long (> 45 mins) suspended state, at least on some ARL and MTL laptops, either all or some of them also containing an Nvidia GPU. Enabling PPS first and then the PLL fixes the problem for all the reporters. A similar issue is seen when enabling an external DP output on PHY B (vs. PHY A in the above eDP cases), where this change will not have any effect (since no PPS is used in that case). There isn't any direct connection between PPS and PLL, so the fix for eDP works by some side-effect only. However Bspec does seem to require enabling PPS first, so let's do that. Further investigation continues on the actual root cause and a cure for external panels. Fixes: 1a7fad2aea74 ("drm/i915/cx0: Enable dpll framework for MTL+") Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16098 Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16064 Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16042 Cc: Mika Kahola <mika.kahola@intel.com> Cc: stable@vger.kernel.org # v7.0+ Tested-by: Jouni Högander <jouni.hogander@intel.com> Tested-by: Marco Nenciarini <mnencia@kcore.it> Reviewed-by: Suraj Kandpal <suraj.kandpal@intel.com> Signed-off-by: Imre Deak <imre.deak@intel.com> Link: https://patch.msgid.link/20260612172617.3427027-1-imre.deak@intel.com (cherry picked from commit 28783a274e886dd6da61419be6020bd9d0384e9f) Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com> -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2167549 Title: ThinkPad P1 Gen 7 Meteor Lake — i915 PHY/DPLL failure after s2idle resume Status in linux package in Ubuntu: New Bug description: After resuming from an s2idle suspend of 30+ minutes, the i915 driver fails to re-enable the display. According to logs, the C10 PHY becomes completely unresponsive because its internal registers are inaccessible and this causes the entire display pipeline to remain stuck, sometimes for ~60 seconds or sometimes indefinitely (requiring a hard reset). GNOME Shell and input handling remain functional throughout; keystrokes typed during the black screen appear in the password box if/when the display eventually recovers. Short-duration suspends (<~2 minutes) resume cleanly without errors. The failure apparently correlates with suspend duration: during long s2idle, the display power well fully shuts down, and on resume the C10 PHY cold-start sequence fails. The bug does NOT reproduce on Windows (on the same laptop), suggesting this is an i915 driver bug, not a hardware or BIOS issue. ## My setup - Laptop: Lenovo ThinkPad P1 Gen 7 (model 21KV0025UK) - BIOS: N48ET34W (1.21), dated 2026-05-11 (current; no LVFS updates available) - iGPU: Intel Meteor Lake-P [Intel Arc Graphics] — PCI 8086:7d55 (rev 08) - Kernel driver: i915 - Panel: eDP-1, 3840x2400 (16:10), 4 lanes, C10 PLL at 540 MHz port clock - dGPU: NVIDIA AD107M [GeForce RTX 4060 Max-Q / Mobile] — PCI 10de:28a0 - Driver: nvidia 580.173.02 (open kernel module), nvidia-drm modeset=1 - NOTE: bug reproduces with and without nvidia loaded — not a factor - suspend mode: s2idle (platform supports only s2idle; no "deep"/S3 available) ## Software - OS: Ubuntu 24.04.5 LTS (Noble Numbat) - Kernel: 7.0.0-31-generic (#31~24.04.1-Ubuntu PREEMPT_DYNAMIC) (also reproduces on 7.0.0-30-generic) - Desktop: GNOME 46.0 on Wayland (GDM) - linux-firmware: 20240318.git3b128b60.0ubuntu3.1 - i915 firmware loaded: - DMC: i915/mtl_dmc.bin v2.21 - GuC: i915/mtl_guc_70.bin v70.36.0 (kernel recommends v70.53.0 but only 70.36.0 available — may or may not be related) - HuC: i915/mtl_huc_gsc.bin v8.5.4 ## To reproduce 1. Boot the system (any kernel 7.0.0-30 or 7.0.0-31) 2. Suspend (lid close or `systemctl suspend`) 3. Leave suspended for 30+ minutes (short suspends <2 min do NOT trigger it) 4. Resume (open lid / press key) 5. Observe: screen stays black for ~60 seconds (or never recovers and needs a hard reset) 6. Kernel log shows the error burst (see below) Suspend duration correlation: - Short suspends (1-2 min): 0 DRM errors, prompt in ~2s - Long suspends (30+ min): 19-97 DRM errors, 60s+ delay or hard reset ## What does NOT work I tried quite a few things, none of which succeeded: - Kernel update 7.0.0-30 → 7.0.0-31: no effect - i915.enable_psr=0 (PSR disabled): no effect - i915.enable_dc=0 (display C-states disabled): no effect - i915.enable_dsb=0 (Display State Buffer disabled): no effect (DSB timeouts disappeared from logs, but core PHY failure remained) - i915.disable_power_well=0 (keep power wells always on): no effect - NVIDIA driver loaded vs unloaded: no effect - nvidia_drm modeset=0 / modeset=1: no effect ## What DOES work - Hibernation: bypasses s2idle entirely (full boot from hibernation image re-initializes display from scratch). Zero DRM errors on resume. Confirmed working across multiple hibernate cycles. - Windows Modern Standby: does NOT reproduce the bug. ## What was NOT testable - xe driver: my laptop would not boot (apparently because the xe driver does not work in my kernel version. It refuses to probe even with xe.force_probe=7d55, requiring both xe.force_probe=7d55 AND i915.force_probe=!7d55. When loaded on 7.0.0-30, it registered only a generic card0-Unknown-1 connector with no eDP display output. Not a viable alternative without a newer kernel with proper MTL display support in xe. ## Root cause (from logs) On resume from long s2idle, the C10 PHY cold-start sequence fails at the hardware register level: 1. `Failed to bring PHY A to idle` PHY A cannot reach idle state 2. `PHY A Read 0c70 failed after 3 retries` PHY register 0c70 is completely inaccessible (returns no data) 3. `PHY A Write 0c70 failed after 3 retries` writes to the PHY also fail 4. `Timeout waiting for DDI BUF A to get active` DDI buffer never enables (depends on PHY being up) 5. `Timed out waiting for DP idle patterns` DP link training cannot proceed 6. C10 PLL state mismatch: expected clock 540000 (540 MHz), found 61440 (61.44 MHz = the idle/standby clock; PLL never re-enabled at full rate) All PLL registers found as 0x00 (zeroed/idle state) 7. `flip_done timed out` repeating every ~10s all atomic commits fail 8. `vblank wait timed out on crtc 0` pipe is dead, no vblanks The C10 PHY's control register interface (offset 0c70) is completely dead after long s2idle. No i915 kernel parameter that I tried can fix this. The PHY itself is in an inaccessible state. Since Windows handles the same hardware correctly, the i915 driver is missing a step in its PHY re-initialization sequence that the Windows Intel driver performs. From: journalctl -b 0 -k (With i915.enable_dc=0, i915.enable_dsb=0, i915.disable_power_well=0 all active) Code: Select all Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* Failed to bring PHY A to idle. Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* PHY A Read 0c70 failed after 3 retries. Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* PHY A Write 0c70 failed after 3 retries. Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* Timeout waiting for DDI BUF A to get active Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* Timed out waiting for DP idle patterns Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* [CRTC:150:pipe A] flip_done timed out Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* [CRTC:150:pipe A] mismatch in pixel_rate (expected 572010, found 65082) Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* [CRTC:150:pipe A] mismatch in dpll_hw_state Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* expected: Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* cx0pll_hw_state: lane_count: 4, ssc_enabled: no, use_c10: yes, tbt_mode: no Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* c10pll_hw_state: clock: 540000, fracen: yes, Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* quot: 40960, rem: 0, den: 1, multiplier: 140, tx_clk_div: 0. Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* c10pll_rawhw_state: tx: 0x10, cmn: 0x21 Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* pll[0] = 0xf4, pll[1] = 0x0, pll[2] = 0xf8, pll[3] = 0x0 Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* pll[16] = 0x84, pll[17] = 0x4f, pll[18] = 0xe5, pll[19] = 0x23 Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* found: Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* c10pll_hw_state: clock: 61440, fracen: no, multiplier: 16, tx_clk_div: 0. Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* c10pll_rawhw_state: tx: 0x0, cmn: 0x0 Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* pll[0-19] = 0x0 (all zeros) Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* [CRTC:150:pipe A] mismatch in hw.pipe_mode.crtc_clock (expected 572010, found 65082) Sep 14 13:06:35 kernel: i915 0000:00:02.0: [drm] *ERROR* [CRTC:150:pipe A] mismatch in port_clock (expected 540000, found 61440) Sep 14 13:06:35 kernel: i915 0000:00.0: [drm] pipe state doesn't match! Sep 14 13:06:35 kernel: WARNING: drivers/gpu/drm/i915/display/intel_modeset_verify.c:225 at verify_crtc_state+0x51b/0x6c0 [i915] Sep 14 13:06:35 kernel: intel_modeset_verify_crtc+0x70/0xb0 [i915] Sep 14 13:06:35 kernel: intel_atomic_commit_tail+0x8bb/0xc80 [i915] Sep 14 13:06:35 kernel: intel_atomic_commit+0x2c0/0x310 [i915] Sep 14 13:06:35 kernel: drm_atomic_commit+0xaf/0xf0 Sep 14 13:06:35 kernel: WARNING: drivers/gpu/drm/i915/display/intel_dpll_mgr.c:4945 at verify_single_dpll_state+0x6c7/0x7e0 [i915] Sep 14 13:06:35 kernel: intel_modeset_verify_crtc+0x7b/0xb0 [i915] (flip_done timeouts repeat every ~10s:) Sep 14 13:06:45 i915 *ERROR* flip_done timed out Sep 14 13:06:45 i915 *ERROR* [CRTC:150:pipe A] commit wait timed out Sep 14 13:06:55 i915 *ERROR* flip_done timed out Sep 14 13:06:55 i915 *ERROR* [CONNECTOR:507:eDP-1] commit wait timed out Sep 14 13:07:06 i915 *ERROR* flip_done timed out Sep 14 13:07:06 i915 *ERROR* [PLANE:34:plane 1A] commit wait timed out ... Sep 14 13:07:57 i915 *ERROR* [CRTC:150:pipe A] flip_done timed out Sep 14 13:07:58 i915: vblank wait timed out on crtc 0 Sep 14 13:07:58 WARNING: drivers/gpu/drm/drm_vblank.c:1320 at drm_crtc_wait_one_vblank+0x18c/0x200 Sep 14 13:07:58 drm_client_modeset_wait_for_vblank+0x61/0x80 Sep 14 13:07:58 drm_fb_helper_damage_work+0x8c/0x1a0 (PM: suspend exit at 13:07:57, ~90s after the first PHY error at 13:06:35) Sep 14 13:07:57 kernel: PM: suspend exit Seems like it needs a kernel patch in the i915 display PHY resume code (drivers/gpu/drm/i915/display/intel_cx0_phy.c or similar), potentially adding a forced PHY reset or an alternative re-initialization sequence for the post-s2idle cold-start case. ## Workaround Hibernation works. I had to fiddle around with swap space and resume paramters to make this work but then `systemctl hibernate` works reliably with zero DRM errors on resume. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2167549/+subscriptions

[Bug 2155195] Re: [Panther Lake/xe] DisplayPort-over-Thunderbolt external monitor not detected at cold boot (regression in linux 7.0.0-22)

Second hardware combo hitting the same failure class System: - ThinkPad P1 Gen9 (21UFS0KC00), BIOS N4SET30W 1.09 - Lenovo ThinkPad Thunderbolt 5 Smart Dock 7500 - 40BA (not Dell TB4 — different dock vendor/gen, same failure class) - 3 monitors: Dell U2415 (1920x1200), LG HDR 4K (3840x2160), built-in eDP (3200x2000) - Kernel 7.0.0-34-generic (Linux Mint 22.2 / Ubuntu 24.04 base, Cinnamon 6.6.4) - Intel Panther Lake, xe driver — same platform as the original report Unlike the original report (cold-boot-only), I have two additional 100%-reproducible triggers on this hardware, both ending in a genuine kernel-level hang (Xorg goes into uninterruptible D state, unrecoverable without a hard reset): 1. Opening Cinnamon's Display Settings applet with the dock's MST-tunneled monitors attached, changing any setting and applying it 2. The screensaver/DPMS wake-from-idle cycle. Both produce this signature in dmesg, with an escalating pre-freeze pattern before the hard hang: xe 0000:00:02.0: [drm] *ERROR* [CONNECTOR:...] Failed to get ACT after 3000 ms workqueue: intel_atomic_cleanup_work [xe] hogged CPU for >10000us 4 times, consider switching to WQ_UNBOUND workqueue: intel_atomic_cleanup_work [xe] hogged CPU for >10000us 5 times, ... workqueue: intel_atomic_cleanup_work [xe] hogged CPU for >10000us 11 times, ... workqueue: intel_atomic_cleanup_work [xe] hogged CPU for >10000us 19 times, ... (count keeps climbing — 4, 5, 7, 11, 19, 35, 67, 131... — until hard freeze) Additional isolation donet: - Not NVIDIA-related: this is a muxless hybrid laptop; NVIDIA never drives any display output regardless of driver state. Confirmed across nouveau, the -open driver, and the closed nvidia-driver-610 package (which actually refuses to load its kernel module on this GPU at all — NVRM: ... requires use of the NVIDIA open kernel modules). xe is the only driver touching any output in every configuration tested. Rules out NVIDIA as a contributing factor. - Cross-hardware control test: connected the identical dock model to a completely different laptop (ThinkPad P1 Gen4i, Tiger Lake, hardware-MUXed NVIDIA driving all outputs, i915 not even loaded). Zero ACT/link-training/atomic-cleanup-workqueue errors across multiple dock hotplug/replug cycles. This isolates the fault to the xe/Panther Lake/muxless side of the pairing, not the dock or monitors. - Also observed: after a wedge, only a full Thunderbolt tunnel teardown+rebuild that lands on a new enumeration path (e.g. thunderbolt port 1-1 → 1-3) clears the wedged state (a same-port reconnect reusing the existing tunnel does not recover it). This is consistent with your root-cause note that the tunnel isn't recreated once torn down early, in my case it's torn down by the driver getting into a bad MST/ACT state post-boot, not just at cold boot, but the underlying "xe doesn't cleanly re-establish the DP tunnel" mechanism looks like the same bug surfacing on a different trigger. Is this related? https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9421 -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2155195 Title: [Panther Lake/xe] DisplayPort-over-Thunderbolt external monitor not detected at cold boot (regression in linux 7.0.0-22) Status in linux package in Ubuntu: Fix Committed Bug description: Ubuntu release: Ubuntu 26.04 LTS (resolute) [lsb_release -rd] Kernel package: linux-image-7.0.0-22-generic 7.0.0-22.22 version_signature: Ubuntu 7.0.0-22.22-generic 7.0.0 What I expected: On boot, an external monitor connected via a Dell Thunderbolt 4 dock (DisplayPort output) should be detected and lit, as it was on kernel 7.0.0-15-generic. What happened instead: After updating to 7.0.0-22-generic, the external monitor is not detected on a cold boot. All other dock functions work normally (USB keyboard/mouse, YubiKey, Gigabit Ethernet), and the same dock + cable + monitor work in Windows on the same machine. The internal panel (eDP-1) works. All /sys/class/drm/card1-DP-*/status read "disconnected"; only eDP-1 is connected. Root cause (boot-time ordering race in DisplayPort-over-Thunderbolt tunneling): The Thunderbolt connection manager sets up and tears down the DP tunnel ~83s BEFORE the xe display driver finishes loading, so the tunnel is destroyed and never recreated. Kernel log (monotonic boot offsets): [ 4.97s] thunderbolt 0-1: new device found ... Dell Thunderbolt 4 Dock [ 17.72s] thunderbolt 0000:00:0d.2: 0:5 <-> 1:14 (DP): not active, tearing down [101.11s] xe 0000:00:02.0: [drm] Found pantherlake (device ID b080) integrated display [102.46s] [drm] Initialized xe 1.1.0 for 0000:00:02.0 The late xe load correlates with a stalled udev coldplug: `systemd- analyze blame` shows a cluster of device units (dev-rfkill, tpm0, ttyS0-3, ...) all settling at ~1min 39s, i.e. udev is wedged ~99s and only modprobes xe afterward. This ~99s stall does not occur on 7.0.0-15-generic. Recovery (confirms HW is fine): unplugging/replugging the single Thunderbolt cable AFTER boot (xe now loaded) immediately re- establishes the DP tunnel and DP-7 goes "connected" and the monitor lights up. The failure is purely the cold-boot ordering race. Hardware: - Dell XPS 16 (product_name "XPS 16 DA16260", board 0TVXV7), BIOS 1.5.1 (2026-04-01) - GPU: Intel Panther Lake [Arc B390] [8086:b080] (rev 04), driver: xe - Dock: Dell Thunderbolt 4 Dock (USB4, links at 40 Gb/s) -> DisplayPort -> external monitor - Session: Wayland (GNOME/mutter) Workarounds: replug the Thunderbolt cable after boot; or boot 7.0.0-15-generic. See also (same teardown signature, different distro/HW): https://github.com/basecamp/omarchy/issues/3906 ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-22-generic 7.0.0-22.22 ProcVersionSignature: Ubuntu 7.0.0-22.22-generic 7.0.0 Uname: Linux 7.0.0-22-generic x86_64 ApportVersion: 2.34.0-0ubuntu2 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC0: kgapos 6304 F.... pipewire kgapos 6325 F.... wireplumber /dev/snd/seq: kgapos 6304 F.... pipewire CasperMD5CheckResult: pass CurrentDesktop: ubuntu:GNOME Date: Wed Jun 3 20:48:43 2026 InstallationDate: Installed on 2026-05-28 (6 days ago) InstallationMedia: Ubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423.1) MachineType: Dell Inc. XPS 16 DA16260 ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-ghostty XDG_RUNTIME_DIR=<set> ProcFB: 0 xedrmfb ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-22-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro quiet splash crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M SourcePackage: linux StagingDrivers: intel_ipu7_isys intel_ipu7 UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 04/01/2026 dmi.bios.release: 1.5 dmi.bios.vendor: Dell Inc. dmi.bios.version: 1.5.1 dmi.board.name: 0TVXV7 dmi.board.vendor: Dell Inc. dmi.board.version: A02 dmi.chassis.type: 10 dmi.chassis.vendor: Dell Inc. dmi.ec.firmware.release: 1.2 dmi.modalias: dmi:bvnDellInc.:bvr1.5.1:bd04/01/2026:br1.5:efr1.2:svnDellInc.:pnXPS16DA16260:pvr:rvnDellInc.:rn0TVXV7:rvrA02:cvnDellInc.:ct10:cvr:sku0DBA:pfaDellLaptops: dmi.product.family: Dell Laptops dmi.product.name: XPS 16 DA16260 dmi.product.sku: 0DBA dmi.sys.vendor: Dell Inc. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2155195/+subscriptions

[Bug 2168640] Re: [regression] 7.0.0-34: USB2 hub on AMD Raphael/Granite Ridge xHCI (1022:15b8) not enumerated at boot ("usb7-port1: unable to enumerate USB device")

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168640 Title: [regression] 7.0.0-34: USB2 hub on AMD Raphael/Granite Ridge xHCI (1022:15b8) not enumerated at boot ("usb7-port1: unable to enumerate USB device") Status in linux package in Ubuntu: New Bug description: PACKAGE: linux (Ubuntu) DESCRIPTION: After upgrading from linux-image 7.0.0-31-generic to 7.0.0-34-generic (resolute, via linux-generic-hwe), a USB 2.0 hub (Genesys Logic 05e3:0608) connected to the single port of the CPU-integrated USB 2.0 xHCI controller is no longer enumerated at boot. Keyboard and mouse behind that hub are therefore dead after boot. Kernel log on 7.0.0-34 (every boot): usb usb7-port1: attempt power cycle usb usb7-port1: unable to enumerate USB device On 7.0.0-31 (and 7.0.0-28/29/30) the same hub enumerates at ~2.9 s, with keyboard and mouse behind it. The same hub DOES enumerate on 7.0.0-34 when hot-plugged after boot, so the hardware is fine; only cold-boot enumeration fails. Statistics from journal: 19/19 boots OK on 7.0.0-28..31, 3/3 boots FAIL on 7.0.0-34. Booting 7.0.0-31 again fixes it (currently pinned as GRUB default as a workaround). The 7.0.0-32/-34 changelog contains no obvious USB/xHCI change, so the cause may be indirect (timing/PCI/power management). HARDWARE: Board: Gigabyte B650 EAGLE, BIOS F35 (2025-07-16) CPU: AMD Ryzen 9 9900X Controller: 11:00.0 USB controller [0c03]: AMD Raphael/Granite Ridge USB 2.0 xHCI [1022:15b8] (subsystem Gigabyte 1458:5007), behind 00:08.3 Internal GPP Bridge Hub: Genesys Logic 05e3:0608 USB2.0 Hub Devices behind hub: INSTANT USB Keyboard 30fa:2053, Microsoft Comfort Mouse 4500 045e:076c GPU: NVIDIA RTX 4090 (nvidia-dkms-595-open 595.91.07) VERSIONS: Good: 7.0.0-31-generic (7.0.0-31.31) Bad: 7.0.0-34-generic (7.0.0-34.34) ATTACHMENTS: dmesg-7.0.0-34-fail.txt, dmesg-7.0.0-31-ok.txt, lsusb-t.txt, lsusb-v-hub.txt, lspci-nnk.txt STEPS TO REPRODUCE: 1. Hub 05e3:0608 plugged into the port wired to the CPU USB2 xHCI (1022:15b8) 2. Boot 7.0.0-34-generic -> hub not enumerated, "usb7-port1: unable to enumerate USB device" 3. Boot 7.0.0-31-generic -> hub + keyboard + mouse enumerated normally --- ProblemType: Bug ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC2: david 4553 F.... wireplumber /dev/snd/controlC0: david 4553 F.... wireplumber /dev/snd/controlC1: david 4553 F.... wireplumber /dev/snd/seq: david 4547 F.... pipewire CasperMD5CheckResult: unknown CurrentDesktop: ubuntu:GNOME DistroRelease: Ubuntu 26.04 MachineType: LDLC CUSTOM Package: linux (not installed) ProcFB: 0 amdgpudrmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-31-generic root=UUID=147f1cef-e1e5-488a-bae0-36ac9071acd0 ro quiet splash crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M ProcVersionSignature: Ubuntu 7.0.0-31.31-generic 7.0.14 RfKill: Tags: resolute wayland-session Uname: Linux 7.0.0-31-generic x86_64 UpgradeStatus: No upgrade log present (probably fresh install) UserGroups: sudo users WifiSyslog: _MarkForUpload: True dmi.bios.date: 07/16/2025 dmi.bios.release: 5.35 dmi.bios.vendor: American Megatrends International, LLC. dmi.bios.version: F35 dmi.board.asset.tag: Default string dmi.board.name: B650 EAGLE dmi.board.vendor: Gigabyte Technology Co., Ltd. dmi.board.version: x.x dmi.chassis.asset.tag: Default string dmi.chassis.type: 3 dmi.chassis.vendor: Default string dmi.chassis.version: Default string dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvrF35:bd07/16/2025:br5.35:svnLDLC:pnCUSTOM:pvrV1-CF:rvnGigabyteTechnologyCo.,Ltd.:rnB650EAGLE:rvrx.x:cvnDefaultstring:ct3:cvrDefaultstring:skuCUSTOM:pfaSERIE: dmi.product.family: SERIE dmi.product.name: CUSTOM dmi.product.sku: CUSTOM dmi.product.version: V1-CF dmi.sys.vendor: LDLC --- ProblemType: Bug ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC2: david 4553 F.... wireplumber /dev/snd/controlC0: david 4553 F.... wireplumber /dev/snd/controlC1: david 4553 F.... wireplumber /dev/snd/seq: david 4547 F.... pipewire CasperMD5CheckResult: unknown CurrentDesktop: ubuntu:GNOME DistroRelease: Ubuntu 26.04 MachineType: LDLC CUSTOM Package: linux (not installed) ProcEnviron: LANG=fr_FR.UTF-8 LC_MESSAGES=fr_FR.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color ProcFB: 0 amdgpudrmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-31-generic root=UUID=147f1cef-e1e5-488a-bae0-36ac9071acd0 ro quiet splash crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M ProcVersionSignature: Ubuntu 7.0.0-31.31-generic 7.0.14 RfKill: Tags: resolute Uname: Linux 7.0.0-31-generic x86_64 UpgradeStatus: No upgrade log present (probably fresh install) UserGroups: ollama _MarkForUpload: True dmi.bios.date: 07/16/2025 dmi.bios.release: 5.35 dmi.bios.vendor: American Megatrends International, LLC. dmi.bios.version: F35 dmi.board.asset.tag: Default string dmi.board.name: B650 EAGLE dmi.board.vendor: Gigabyte Technology Co., Ltd. dmi.board.version: x.x dmi.chassis.asset.tag: Default string dmi.chassis.type: 3 dmi.chassis.vendor: Default string dmi.chassis.version: Default string dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvrF35:bd07/16/2025:br5.35:svnLDLC:pnCUSTOM:pvrV1-CF:rvnGigabyteTechnologyCo.,Ltd.:rnB650EAGLE:rvrx.x:cvnDefaultstring:ct3:cvrDefaultstring:skuCUSTOM:pfaSERIE: dmi.product.family: SERIE dmi.product.name: CUSTOM dmi.product.sku: CUSTOM dmi.product.version: V1-CF dmi.sys.vendor: LDLC To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168640/+subscriptions

[Bug 2167595] Re: mt7921u: USB reset after -110 timeouts deadlocks in mt7921_abort_roc, blocks rtnl_lock and shutdown

LP#2167595 - field feedback on the v2 patch (lp2167595_mt7921u_usb_disconnect_deadlock_v2.patch, sha256 c83bc45d...4fefd1) == Environment == Ubuntu 26.04.1 LTS, kernel 7.0.0-34-generic (Ubuntu 7.0.0-34.34, upstream 7.0.14) Adapter: MediaTek MT7921AU, USB ID 0e8d:7961, driver mt7921u, high-speed (USB 2.0) port Hardware: Dell Precision 3650 Tower, BIOS 1.48.0 v2 built out-of-tree from linux-source-7.0.0 (7.0.0-34.34), installed into /lib/modules/7.0.0-34-generic/updates/ and in daily use since 2026-09-26. Kernel taint is O+E only (out-of-tree unsigned modules: these plus VirtualBox). Logs below are sanitized: hostname, BSSID and adapter MAC replaced. == 1. Real-world validation: the deadlock is gone == On 2026-09-26 the adapter failed exactly the way it did in the original report: 8 consecutive "vendor request ... failed:-110" timeouts, Bluetooth on the same chip timing out as well, then a USB reset loop from which the device never returned (descriptor read errors -110, "device not accepting address" -62, finally repeated re-enumeration attempts that all failed). Timeline (local time, 26.09.2026): 16:29:28 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110 16:29:31 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110 16:29:34 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d058 failed:-110 16:29:38 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53b8 failed:-110 16:29:40 kernel: Bluetooth: hci0: Opcode 0x0401 failed: -110 16:29:40 kernel: Bluetooth: hci0: command 0x0401 tx timeout 16:29:41 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53c4 failed:-110 16:29:44 kernel: mt7921u 1-9:1.3: vendor request req:66 off:53c4 failed:-110 16:29:47 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110 16:29:50 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110 16:29:51 kernel: Bluetooth: hci0: Failed to write uhw reg(-110) 16:29:53 kernel: wlx00c0cab8f19c: deauthenticating from [BSSID] by local choice (Reason: 3=DEAUTH_LEAVING) 16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DISCONNECTED bssid=[BSSID] reason=3 locally_generated=1 16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: Added BSSID [BSSID] into ignore list, ignoring for 10 seconds 16:30:08 kernel: wlx00c0cab8f19c: failed to remove key (1, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:09 kernel: wlx00c0cab8f19c: failed to remove key (2, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:11 kernel: wlx00c0cab8f19c: failed to remove key (4, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:12 kernel: wlx00c0cab8f19c: failed to remove key (5, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:13 kernel: mt7921u 1-9:1.3: timed out waiting for pending tx 16:30:13 kernel: snd_soc_acpi_intel_sdca_quirks soundwire_generic_allocation snd_soc_sdw_utils snd_soc_acpi intel_rapl_msr soundwire_bus in 16:30:14 NetworkManager[3140]: device (wlx00c0cab8f19c): state change: activated -> unmanaged (reason 'unmanaged-link-not-init', managed-typ 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): canceled DHCP transaction 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): activation: beginning transaction (timeout in 45 seconds) 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): state changed no lease 16:30:14 ModemManager[2292]: <msg> [base-manager] port wlx00c0cab8f19c released by device '/sys/devices/pci0000:00/0000:00:14.0/usb1/1-9' 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: PMKSA-CACHE-REMOVED [BSSID] 0 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: Removed BSSID [BSSID] from ignore list (clear) 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all 16:30:14 wpa_supplicant[2255]: nl80211: deinit ifname=wlx00c0cab8f19c disabled_11b_rates=0 16:30:14 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:30:20 kernel: usb 1-9: device descriptor read/64, error -110 16:30:24 systemd[1]: NetworkManager-dispatcher.service: Deactivated successfully. 16:30:35 kernel: usb 1-9: device descriptor read/64, error -110 16:30:36 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:30:41 kernel: usb 1-9: device descriptor read/64, error -110 16:30:57 kernel: usb 1-9: device descriptor read/64, error -110 16:30:57 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:31:08 kernel: usb 1-9: device not accepting address 3, error -62 16:31:08 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:31:20 kernel: usb 1-9: device not accepting address 3, error -62 16:31:20 kernel: usb 1-9: USB disconnect, device number 3 16:31:20 kernel: usb 1-9: new high-speed USB device number 5 using xhci_hcd 16:31:25 kernel: usb 1-9: device descriptor read/64, error -110 16:31:41 kernel: usb 1-9: device descriptor read/64, error -110 16:31:41 kernel: usb 1-9: new high-speed USB device number 6 using xhci_hcd 16:31:47 kernel: usb 1-9: device descriptor read/64, error -110 16:32:03 kernel: usb 1-9: device descriptor read/64, error -110 16:32:03 kernel: usb 1-9: new high-speed USB device number 7 using xhci_hcd 16:32:14 kernel: usb 1-9: device not accepting address 7, error -62 16:32:14 kernel: usb 1-9: new high-speed USB device number 8 using xhci_hcd 16:32:26 kernel: usb 1-9: device not accepting address 8, error -62 16:45:01 systemd[7730]: Reached target shutdown.target - Shutdown. 16:45:04 systemd[1]: NetworkManager-wait-online.service: Deactivated successfully. 16:45:05 systemd[1]: NetworkManager.service: Deactivated successfully. 16:45:09 systemd[1]: Reached target shutdown.target - System Shutdown. 16:45:09 systemd[1]: Reached target poweroff.target - System Power Off. 16:45:09 systemd-shutdown[1]: Syncing filesystems and block devices. 16:45:09 systemd-shutdown[1]: Sending SIGTERM to remaining processes... 16:45:09 systemd-journald[1051]: Journal stopped Outcome with v2: * No task ever blocked on rtnl_lock. Zero "blocked for more than N seconds" in the whole boot. * The system stayed fully responsive for the remaining 15 minutes of the session (only WiFi was gone) and then shut down normally: the full shutdown sequence took 8 seconds, from 16:45:01 to "Journal stopped" at 16:45:09. * No "chip reset failed", no "rx urb mismatch". For comparison, the same hardware failure on stock/v1 modules (2026-09-15, the original report) left NetworkManager, ip, and 6 other tasks in D state for more than 122 seconds, all waiting for rtnl_lock held by a kworker in mt792xu_disconnect -> mt76_unregister_device -> mt7921_abort_roc, and the machine had to be powered off with the button. So the core goal of the patch is confirmed in the field, not just in tests. == 2. New finding: the tx_worker disable in mt792xu_disconnect is redundant and fires a WARNING on the dead-device path == Alongside the (harmless) "timed out waiting for pending tx" message, v2 produced a one-shot WARNING: 16:30:13 mt7921u 1-9:1.3: timed out waiting for pending tx 16:30:13 ------------[ cut here ]------------ 16:30:13 WARNING: kernel/kthread.c:722 at kthread_park+0x8c/0xc0, CPU#7: kworker/7:2/100058 16:30:13 CPU: 7 UID: 0 PID: 100058 Comm: kworker/7:2 Tainted: G OE 7.0.0-34-generic #34-Ubuntu PREEMPT(full) 16:30:13 Tainted: [O]=OOT_MODULE, [E]=UNSIGNED_MODULE 16:30:13 Hardware name: Dell Inc. Precision 3650 Tower/0NDYHG, BIOS 1.48.0 05/27/2026 16:30:13 Workqueue: events __usb_queue_reset_device 16:30:13 FS: 0000000000000000(0000) GS:ffff8de34f77f000(0000) knlGS:0000000000000000 16:30:13 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 16:30:13 Call Trace: 16:30:13 <TASK> 16:30:13 ? mt76u_stop_tx.cold+0x11c/0x180 [mt76_usb] 16:30:13 ? __pfx_autoremove_wake_function+0x10/0x10 16:30:13 mt792xu_stop+0x1a/0x40 [mt792x_usb] 16:30:13 drv_stop+0x50/0x160 [mac80211] 16:30:13 ieee80211_stop_device+0x7f/0x90 [mac80211] 16:30:13 ieee80211_do_stop+0x6a2/0xb20 [mac80211] 16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20 16:30:13 ? packet_notifier+0x85/0x280 16:30:13 ieee80211_stop+0x63/0xf0 [mac80211] 16:30:13 __dev_close_many+0xb2/0x220 16:30:13 netif_close_many+0xb7/0x1b0 16:30:13 netif_close+0x70/0xa0 16:30:13 dev_close+0x38/0xb0 16:30:13 cfg80211_shutdown_all_interfaces+0x50/0x100 [cfg80211] 16:30:13 ieee80211_remove_interfaces+0x47/0x220 [mac80211] 16:30:13 ieee80211_unregister_hw+0x4a/0x140 [mac80211] 16:30:13 mt76_unregister_device+0x63/0x80 [mt76] 16:30:13 mt792xu_disconnect+0xea/0x150 [mt792x_usb] 16:30:13 usb_unbind_interface+0x9b/0x2c0 16:30:13 device_remove+0x68/0x80 16:30:13 device_release_driver_internal+0x1fb/0x260 16:30:13 device_release_driver+0x12/0x20 16:30:13 usb_forced_unbind_intf+0x96/0xe0 16:30:13 ? usb_autoresume_device+0x1e/0x70 16:30:13 usb_reset_device+0xf4/0x300 16:30:13 __usb_queue_reset_device+0x3b/0x60 16:30:13 process_one_work+0x1ac/0x3d0 16:30:13 worker_thread+0x1b8/0x360 16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20 16:30:13 ? __pfx_worker_thread+0x10/0x10 16:30:13 kthread+0xf7/0x130 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ret_from_fork+0x195/0x2a0 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ret_from_fork_asm+0x1a/0x30 16:30:13 </TASK> 16:30:13 ---[ end trace 0000000000000000 ]--- Analysis: 1. v2 parks tx_worker early in mt792xu_disconnect(): mt76_worker_disable(&dev->mt76.tx_worker); 2. Later in the same function, mt76_unregister_device() -> ieee80211_unregister_hw() -> ... -> drv_stop() -> mt792xu_stop() -> mt76u_stop_tx(). 3. mt76u_stop_tx() (drivers/net/wireless/mediatek/mt76/usb.c:994) waits HZ/5 for pending TX to drain. On timeout it logs the message, kills the TX URBs and calls mt76_worker_disable(&dev->tx_worker) - a SECOND park of the same worker. kthread_park() then hits WARN_ON_ONCE(test_bit(KTHREAD_SHOULD_PARK, &kthread->flags)) (kernel/kthread.c:722) and returns -EBUSY. 4. At the end of that same branch mt76u_stop_tx() calls mt76_worker_enable(&dev->tx_worker), i.e. it UNPARKS the worker that v2 deliberately stopped - so the patch line loses its effect precisely in the case it was meant to cover. Suggestion: drop mt76_worker_disable(&dev->mt76.tx_worker); from mt792xu_disconnect(). mt76u_stop_tx() already quiesces TX during unregister and pairs its own park/unpark correctly. The line is not part of the deadlock fix (flags + worker cancellation + the abort_roc early return are), and on the dead-device path it is actively counterproductive. Why lab testing does not catch this: the frame in the trace is mt76u_stop_tx.cold, i.e. the unlikely branch. With a healthy adapter the TX queues drain far below the 200 ms timeout, so that branch never executes. On 2026-09-24 we ran three consecutive rmmod/insmod cycles plus an unplug-during-transfer test on v2 and saw no WARNING at all; the two MCU timeout messages were the only output. It took a genuine adapter failure, with TX still queued at disconnect time, to reach it. Impact: cosmetic in effect (one WARNING, W taint) plus the defeated worker disable. Teardown continued and completed; nothing hung. == 3. Three observations from reviewing v2 against the 7.0.14 sources == (1) mt7925_mac_reset_work() has NO MT76_REMOVED check in 7.0.14. The v2 description says the new mt7921 check aligns mt7921 with mt7925, but on the mt7925 side the check does not exist, so mt7925 USB keeps the race that v2 fixes for mt7921. (grep confirms: MT76_REMOVED appears in mt7925/pci.c and mt792x_dma.c only, not in mt7925/mac.c.) (2) Cosmetic: mt7921_mcu_parse_response() (mt7921/mcu.c:26) prints "Message %08x (seq %d) timeout" and calls mt792x_reset() without checking MT76_MCU_RESET. Because v2 sets MT76_MCU_RESET at disconnect entry, mt76_mcu_get_response() returns immediately, so every teardown MCU command logs a timeout ~30 ms after "deregistering interface driver" - in our logs exactly two per unload: MCU_UNI_CMD(BSS_INFO_UPDATE) (0x00020002) from interface removal and MCU_EXT_CMD(MAC_INIT_CTRL) (0x000046ed) from mt792x_stop() -> mt76_connac_mcu_set_mac_enable(). A MT76_MCU_RESET check before dev_err()/mt792x_reset() would suppress noise that is expected by design. mt7925/mcu.c:22 has the same print. (3) Behaviour note: with MT76_REMOVED set on entry, the teardown MCU commands and mt792xu_wfsys_reset() inside mt792xu_cleanup() fail with -EIO, so a plain rmmod of a healthy adapter leaves the firmware running. In practice this is harmless because mt7921u_probe() resets WFSYS when MT_TOP_MISC2_FW_N9_RDY is set, but it is a behaviour change worth knowing. == 4. Reproduction attempts: the branch cannot be reached on healthy hardware == We tried to reproduce the WARNING deliberately, to confirm the explanation above rather than rely on a single capture. Three attempts on 2026-09-27, all with the v2 modules loaded (mt792x_usb srcversion E7EEAE2D193E09CC1645423): # method TX load before trigger result 1 sysfs driver unbind 6336 pkt in 10 s (633/s) not reproduced (echo 1-9:1.3 > /sys/bus/usb/drivers/mt7921u/unbind) 2 sysfs driver unbind, longer load 7768 pkt in 30 s (258/s) not reproduced 3 port deauthorisation 6287 pkt in 10 s (628/s) not reproduced (echo 0 > /sys/bus/usb/devices/1-9/authorized) In all three runs the counters were identical: timed out waiting for pending tx 0 WARNING 0 kthread_park 0 blocked for 0 TX always drained well inside the HZ/5 window, so mt76u_stop_tx() took the normal path and never called mt76_worker_disable() a second time. The only kernel output in each run was the expected MCU teardown noise described in section 3(2): 15 "Message ... timeout" lines per unload, for commands 0x00020002, 0x00020003, 0x00020006 and 0x000046ed. Teardown and re-probe completed every time; the interface came back after 3 s in all three runs. This is consistent with the analysis: the .cold branch only executes when TX cannot drain, i.e. when the chip is dead or stalled, and that state cannot be induced on demand on healthy hardware. Forcing a disconnect - whether by unbinding the driver or by deauthorising the port - still leaves the device answering on the bus, so the queues empty immediately. In other words, the failed reproduction strengthens rather than weakens the finding: it is direct evidence of why routine testing (our own 2026-09-24 runs included: three rmmod/insmod cycles plus unplug-during-transfer, no WARNING at all) cannot surface this path. The 2026-09-26 trace in section 2 remains the only capture of it, obtained during a genuine adapter failure. == Summary == v2 solves the problem it targets: a real adapter death no longer takes the networking stack or shutdown down with it. The only item we would ask you to change is the redundant mt76_worker_disable(&dev->mt76.tx_worker) in mt792xu_disconnect described in section 2. As a testing hint: the .cold branch can be exercised artificially by temporarily shortening the HZ/5 timeout in mt76u_stop_tx in a local build, which may help validate a revised patch without waiting for a hardware failure. Happy to test a revised patch on this hardware. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2167595 Title: mt7921u: USB reset after -110 timeouts deadlocks in mt7921_abort_roc, blocks rtnl_lock and shutdown Status in linux package in Ubuntu: Confirmed Bug description: [ Impact ] When an MT7921AU USB Wi-Fi dongle (e.g. 0e8d:7961) experiences communication stalls (USB -110 ETIMEDOUT errors), mt792x_mac_work holds dev->mt76.mutex while looping on register reads. When USB core detects the stall and triggers unbind, mt792xu_disconnect calls mt76_unregister_device, which takes rtnl_lock and calls mt7921_abort_roc. mt7921_abort_roc then tries to acquire dev->mt76.mutex, which is already held by the stuck mac_work thread. This creates an ABBA deadlock between rtnl_lock and mt76.mutex. As a result: - mac_work waits on USB timeouts with mt76.mutex held. - mt792xu_disconnect waits for mt76.mutex under rtnl_lock. - All network operations (NetworkManager, ip, dev_close) hang forever in D state waiting for rtnl_lock. - The system cannot power off or reboot cleanly without a hard reset. [ Fix ] 1. In mt792xu_disconnect(), set MT76_REMOVED, MT76_RESET, and MT76_MCU_RESET flags and wake pending waitqueues before calling mt76_unregister_device(). Setting MT76_REMOVED causes all pending and subsequent USB register requests to fail immediately with -EIO instead of waiting for 3-second timeouts. 2. In mt792xu_disconnect(), synchronously cancel all workers (mac_work, ps_work, wake_work, reset_work, init_work) prior to unregistration. 3. In mt7921_abort_roc() (and mt7925_abort_roc()), check if MT76_REMOVED is set. If the device was removed, clear MT76_STATE_ROC and return 0 immediately without taking dev->mt76.mutex. [ Test Plan ] 1. Boot kernel 7.0.0-31-generic with patched mt7921u / mt792x-usb modules. 2. Insert MT7921AU USB Wi-Fi adapter (0e8d:7961) and connect to an AP. 3. Physically disconnect the USB adapter while network traffic is active, or trigger bus reset while requests are pending. 4. Verify in dmesg: - Device disconnect completes cleanly without call traces. - rtnl_lock is not blocked; NetworkManager and ip link operate normally. - System reboots and powers off cleanly without hanging in D state. [ Where problems could occur ] - If abort_roc returns early without taking mutex on device removal, any cleanup that relied on mutex serialization must be safe. Since MT76_REMOVED is set, hardware registers cannot be accessed anyway, so skipping mutex and clearing local flag is safe. - Canceling works before unregistering prevents concurrent execution during teardown, which is the desired behavior on unbind. [ Other Info ] Patch tested and compiled against linux-headers-7.0.0-31-generic on Ubuntu 26.04 (Resolute). Clean compile with 0 errors and 0 warnings. --- [ Original Report ] Dell Precision 3650 Tower, Ubuntu 26.04 dev kernel 7.0.0-31-generic. Alfa AWUS036AXM (MT7921AU, USB ID 0e8d:7961). Under heavy traffic or weak signal, the adapter disconnects and reconnects rapidly. Eventually kernel dmesg shows: [ 342.112004] mt7921u 1-2:1.0: Message 00000040 (seq 4) timeout [ 345.184002] mt7921u 1-2:1.0: Message 00000040 (seq 5) timeout [ 348.256011] mt7921u 1-2:1.0: Failed to get patch sem [ 351.328008] mt7921u 1-2:1.0: hardware init failed After that, NetworkManager stops responding. Running 'ip link' hangs indefinitely in D state. Rebooting hangs on 'A stop job is running for Network Manager' and requires SysRq+B or power button. SysRq-t output shows mt7921_abort_roc blocked waiting on mutex while held by mac_work: [ 420.100012] task:kworker/u16:3 blocked for more than 120 seconds. [ 420.100020] Call Trace: [ 420.100025] __schedule+0x345/0x890 [ 420.100030] schedule+0x5a/0xc0 [ 420.100035] schedule_preempt_disabled+0x18/0x30 [ 420.100040] __mutex_lock.isra.0+0x28a/0x4b0 [ 420.100045] mt7921_abort_roc+0x2d/0x80 [mt7921_common] [ 420.100050] ieee80211_set_disassoc+0x62/0x90 [cfg80211] [ 420.100055] mt76_unregister_device+0x48/0x90 [mt76] [ 420.100060] mt792xu_disconnect+0x3c/0x70 [mt792x_usb] To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2167595/+subscriptions

пятница

[Bug 2166325] Re: Intel VMD 8086:ad0b: Bus Offset 3 prevents NVMe enumeration and ARL004 causes NVMe completion stalls

Following up on the upstream submission requested above: The ARL004 fix has now been submitted upstream to linux-pci and linux-kernel: [PATCH] PCI: vmd: Enable interrupt ordering quirk for Arrow Lake Message-ID: <20260910200539.10862-1-gvozd188@mail.ru> Date: 2026-09-10 The patch is publicly archived on the Linux PCI mailing list and has also passed Sashiko's automated review with no regressions reported. This patch links back to Launchpad bug #2166325 and has been tested on the affected Intel VMD 8086:ad0b hardware. Since the requested upstream submission has now been completed, could the Ubuntu kernel task please be reconsidered rather than left as Won't Fix? I understand that upstream maintainer review may still be pending. If Ubuntu requires upstream acceptance before reconsidering the SRU, please clarify that requirement so I know what the remaining blocker is. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2166325 Title: Intel VMD 8086:ad0b: Bus Offset 3 prevents NVMe enumeration and ARL004 causes NVMe completion stalls Status in linux package in Ubuntu: Won't Fix Bug description: SRU Justification: Impact: On Intel VMD 8086:ad0b / Arrow Lake-HX systems, Ubuntu 26.04 with the 7.0.x kernel has two VMD-related problems. First, the stock VMD driver does not support Bus Offset Setting 3. On the affected MSI Vector this results in: vmd 0000:00:0e.0: Unknown Bus Offset Setting (3) and prevents enumeration of the internal NVMe device behind VMD. After adding the required second-rootbus / Bus Offset 3 support, a second issue is exposed: repeated NVMe timeout, completion polled events caused by the Intel Arrow Lake ARL004 interrupt-ordering erratum. These events cause severe storage stalls and make the system practically unusable under sustained I/O. Launchpad bug #2166325 is now Confirmed and affects multiple users. Fix: The fix is implemented as two logically separate changes: Add Intel VMD second-rootbus / Bus Offset 3 support so that NVMe devices behind VMD 8086:ad0b can be enumerated correctly. Apply an ARL004-specific interrupt-ordering workaround for VMD 8086:ad0b and child PCI_CLASS_STORAGE_EXPRESS devices. Before dispatching the child interrupt with generic_handle_irq(), perform a dummy PCI configuration-space read from the MSI-initiating NVMe device. The read is serialized with the existing VMD configuration-space lock. The workaround is deliberately limited to VMD 8086:ad0b and NVMe-class child devices, so other VMD devices and non-NVMe children are unaffected. The stock drivers/pci/controller/vmd.c in Ubuntu 7.0.0-30.30 and 7.0.0-31.31 is byte-for-byte identical: SHA256: b631d449d8e1e896bcf1fa1e4ade53789d28db3c86c6a15fb7fb7e506cd274c8 The patched vmd.c used successfully with Ubuntu 7.0.0-31-generic is: SHA256: 0a27aa1379e8b7fc22f509b8d06ed3676a211d7bb3226a3f9b3dadfc208be7a3 Testcase: Hardware: MSI Vector 17 HX AI, Intel VMD 8086:ad0b, NVMe behind VMD. With the stock driver: Boot Ubuntu 26.04 with the affected 7.0.x kernel. Verify that Bus Offset Setting 3 causes the VMD driver to fail NVMe enumeration. After Bus Offset 3 support is added but without the ARL004 workaround, run sustained sequential NVMe reads and monitor the kernel log. Repeated nvme nvme0: I/O tag ... timeout, completion polled events occur at approximately 30-second intervals. A 1 GiB sequential read took approximately 121 seconds (~8.8 MB/s). With both fixes applied: The internal NVMe is enumerated normally. The recurring timeout, completion polled events disappear. The same 1 GiB read completes in approximately 0.88 seconds (~1.2 GB/s). A subsequent 10 GiB sequential read also completes normally. The same patched implementation is currently running successfully on Ubuntu 7.0.0-31-generic. Source and patch series: https://github.com/gvozd188/vmd-arl004 Original public implementation: https://github.com/gvozd188/vmd-arl004/commit/5438c83 Hardware ======== Laptop: MSI Vector 17 HX AI Platform: Intel Arrow Lake Intel VMD: 8086:ad0b VMD PCI address: 0000:00:0e.0 Subsystem: 1462:149c Internal NVMe SSD: Phison 1TB ESR01TBYCCA4-EDJ-2MS Controller: Phison PS5029-E29T PCIe 4.0 NVMe PCI ID: 1987:5029 (rev 01) Firmware: ETFM50.0 Ubuntu: 26.04 Current/final tested kernel: 7.0.0-30-generic Secure Boot: enabled Problem 1: Intel VMD Bus Offset 3 ================================ With the stock Ubuntu VMD driver the internal NVMe SSD is not exposed. The kernel reports:     vmd 0000:00:0e.0: Unknown Bus Offset Setting (3) and no internal /dev/nvme* device is available. Linux 6.14 was also tested on this machine and did not solve the VMD problem. The system firmware does not expose a usable BIOS option to disable Intel VMD/RST. Intel's public VMD second-rootbus patch series was integrated into the Ubuntu VMD source:     [PATCH v3 0/8] VMD add second rootbus support https://www.spinics.net/lists/linux-pci/msg163096.html After integrating the second-rootbus support, Linux successfully enumerates the internal NVMe controller and SSD. Problem 2: NVMe completion stalls ================================= After the SSD became accessible through the modified VMD driver, a second problem became visible. Kernel messages repeatedly contained:     nvme nvme0: I/O tag ... timeout, completion polled The stalls occurred at approximately 30-second intervals. A diagnostic sequential read before the workaround:     sudo dd if=/dev/nvme0n1 of=/dev/null bs=16M count=64 status=progress Result:     1 GiB in approximately 121.339 seconds     approximately 8.8 MB/s CPU and memory were not saturated. Disabling ASPM was tested and did not solve the problem. ARL004 investigation ==================== The official Intel RST/VMD Windows driver distributed by MSI for this machine was examined by static analysis for comparison. The relevant Windows driver path performs a PCI configuration-space read after MSI handling when the completion state requires ordering. This behavior is consistent with the Intel Arrow Lake ARL004 erratum, where an MSI from a VMD-owned device may pass a preceding memory write. An experimental Linux workaround was implemented for the tested 8086:ad0b VMD/NVMe path. It performs a dummy PCI configuration-space read before the subsequent interrupt handling path. Result ====== With the experimental ARL004 workaround:     sudo dd if=/dev/nvme0n1 of=/dev/null bs=16M count=64 status=progress completed in:     1 GiB in 0.88206 seconds     approximately 1.2 GB/s A subsequent 10 GiB sequential diagnostic read also completed normally. After booting with the final tested module:     sudo journalctl -k -b | grep -E 'timeout|completion polled' produced no matching messages. The approximately 136x difference above is only a diagnostic comparison on this specific machine and is not intended as a general SSD benchmark. Reproduction / reference implementation ======================================= The complete investigation, tested source, separate patches, combined patch, known-working module, SHA256 checksums, installation/bootstrap procedure and recovery procedure are published here: https://github.com/gvozd188/vmd-arl004 Exact reference commit: https://github.com/gvozd188/vmd-arl004/commit/5438c83 Relevant files:     0001-vmd-second-rootbus-intel.patch         Integration of Intel's second-rootbus support.     0002-vmd-ad0b-arl004-workaround.patch         Experimental ARL004 workaround.     vmd-arl004.patch         Combined patch.     vmd-ubuntu-7.0.0-30.c         Ubuntu VMD source used as the patch base.     vmd.c         Final tested source.     vmd.ko         Known-working reference module for 7.0.0-30-generic. Known-working source SHA256:     0a27aa1379e8b7fc22f509b8d06ed3676a211d7bb3226a3f9b3dadfc208be7a3 Known-working module SHA256:     b2c9eacb720d45fcaab15ff09ad3fdc2e97b8a6dd215c80fa282090ff8d2b2cd Installation history ==================== Because the stock VMD driver could not expose the internal SSD, an existing Ubuntu HDD from another laptop was connected to the MSI Vector through a SATA-to-USB adapter. The MSI Vector was booted directly from that HDD. That Ubuntu environment was running kernel 7.0.0-30-generic. The modified VMD driver was used there to expose the internal NVMe SSD. Ubuntu was then manually deployed onto the dedicated Ubuntu partition of the internal SSD without modifying the existing Windows and data partitions. The manually deployed SSD installation initially contained kernel 7.0.0-14-generic. This kernel was not intentionally selected as a VMD workaround; it was simply the kernel present in the manually deployed system at that stage. A compatible custom VMD module was installed for 7.0.0-14-generic and added to its initramfs, allowing the first independent boot from the internal SSD. Only after that successful SSD boot was DKMS configured. The SSD installation was subsequently updated to 7.0.0-30-generic, where the final VMD/ARL004 development and testing was performed. Expected result =============== The stock Ubuntu kernel should: 1. correctly enumerate the second-rootbus / Bus Offset 3 topology on Intel    VMD 8086:ad0b; 2. expose the internal NVMe SSD; 3. handle NVMe completion ordering without repeated approximately 30-second    completion timeouts. Actual result ============= With the stock VMD driver, the internal SSD is not exposed because Bus Offset Setting 3 is rejected. After adding second-rootbus support alone, the SSD becomes visible but repeated NVMe completion stalls occur on this hardware. The experimental ARL004 ordering workaround eliminates the observed stalls on the tested system. Notes ===== The ARL004 workaround is experimental and platform-specific. I am reporting the observed hardware behavior and the tested workaround rather than claiming that this implementation is the appropriate final upstream fix. The Windows driver binary is not redistributed in the GitHub repository. Only identification information, hashes and static-analysis notes are provided. ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-30-generic 7.0.0-30.30 ProcVersionSignature: Ubuntu 7.0.0-30.30-generic 7.0.12 Uname: Linux 7.0.0-30-generic x86_64 ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse:  USER PID ACCESS COMMAND  /dev/snd/controlC1: gvozd188 2323 F.... pipewire                       gvozd188 2360 F.... wireplumber  /dev/snd/controlC0: gvozd188 2360 F.... wireplumber  /dev/snd/seq: gvozd188 2323 F.... pipewire CasperMD5CheckResult: unknown CurrentDesktop: ubuntu:GNOME Date: Thu Sep 3 13:08:10 2026 MachineType: Micro-Star International Co., Ltd. Vector 17 HX AI A2XWIG ProcFB: 0 i915drmfb ProcKernelCmdLine: BOOT_IMAGE=/boot/vmlinuz-7.0.0-30-generic root=UUID=fef65fb9-b1c4-4bbe-9c02-92a4a7280fee ro quiet splash iommu=pt crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) WifiSyslog: dmi.bios.date: 04/20/2026 dmi.bios.release: 1.18 dmi.bios.vendor: American Megatrends International, LLC. dmi.bios.version: E17S3IMS.112 dmi.board.asset.tag: Default string dmi.board.name: MS-17S3 dmi.board.vendor: Micro-Star International Co., Ltd. dmi.board.version: REV:1.0 dmi.chassis.asset.tag: No Asset Tag dmi.chassis.type: 10 dmi.chassis.vendor: Micro-Star International Co., Ltd. dmi.chassis.version: N/A dmi.modalias: dmi:bvnAmericanMegatrendsInternational,LLC.:bvrE17S3IMS.112:bd04/20/2026:br1.18:svnMicro-StarInternationalCo.,Ltd.:pnVector17HXAIA2XWIG:pvrREV1.0:rvnMicro-StarInternationalCo.,Ltd.:rnMS-17S3:rvrREV1.0:cvnMicro-StarInternationalCo.,Ltd.:ct10:cvrN/A:sku17S3.1:pfaVector: dmi.product.family: Vector dmi.product.name: Vector 17 HX AI A2XWIG dmi.product.sku: 17S3.1 dmi.product.version: REV:1.0 dmi.sys.vendor: Micro-Star International Co., Ltd. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2166325/+subscriptions

[Bug 2168586] [NEW] Jammy update: v5.15.220 upstream stable release

Public bug reported: SRU Justification Impact: The upstream process for stable tree updates is quite similar in scope to the Ubuntu SRU process, e.g., each patch has to demonstrably fix a bug, and each patch is vetted by upstream by originating either directly from a mainline/stable Linux tree or a minimally backported form of that patch. The following upstream stable patches should be included in the Ubuntu kernel: v5.15.220 upstream stable release from git://git.kernel.org/ RDMA/rxe: Fix OOB in free_rd_atomic_resources() ext4: don't enable DAX on new encrypted files io_uring/io-wq: fix worker accounting when canceling creation callbacks ipvs: reload ip header after head reallocation bpf: Remove tst_run from lwt_seg6local_prog_ops. jfs: add check read-only before truncation in jfs_truncate_nolock() jfs: add check read-only before txBeginAnon() call can: j1939: implement NETDEV_UNREGISTER notification handler can: j1939: add missing calls in NETDEV_UNREGISTER notification handler can: j1939: make j1939_sk_bind() fail if device is no longer registered KVM: arm64: Prevent access to vCPU events before init bpf: Fix use-after-free in offloaded map/prog info fill Revert "PM: sleep: Use complete() in device_pm_sleep_init()" smc: Use __sk_dst_get() and dst_dev_rcu() in smc_vlan_by_tcpsk(). selinux: switch two allocations to use kzalloc_objs() ALSA: pcm: fix wait_time calculations ASoC: tegra: Fix Master Volume Control ALSA: pcm: fix use-after-free on linked stream runtime in snd_pcm_drain() ALSA: PCM: Fix wait queue list corruption in snd_pcm_drain() on linked streams ipv4: igmp: Fix potential UAF in igmp_gq_start_timer() kcov: replace local_irq_save() with a local_lock_t kcov: fix data corruption and race conditions on PREEMPT_RT ext4: propagate errors from fast commit range replay nilfs2: correct return value kernel-doc descriptions for ioctl functions nilfs2: reject invalid block index in GC ioctl nfc: nci: add data_len bound checks to activation parameter extractors HID: magicmouse: prevent unbounded recursion in magicmouse_raw_event() nvme: rename CDR/MORE/DNR to NVME_STATUS_* nvmet-tcp: bound SGL data length before allocating command buffers HID: uclogic: fix use-after-free of inrange_timer on remove HID: ft260: fix i2c probing for hwmon devices HID: ft260: improve i2c write performance HID: ft260: improve i2c large reads performance HID: ft260: skip unexpected HID input reports HID: ft260: wake up device from power saving mode HID: ft260: missed NACK from busy device HID: ft260: validate i2c input report length HID: ft260: fix stack-use-after-return write in I2C read race HID: input: read battery capacity from its actual report offset fpga: dfl: fme: add error handling accessibility: speakup: unregister tty ldisc on later init failures xhci: dbgtty: Fix unregister on tty_register_driver() failure xhci: dbgtty: Fix unregister on tty_alloc_driver() failure fuse: fix invalidate lock leak on setattr writeback failure fuse: fix invalidate lock leak on open O_TRUNC DAX failure usb: usbtest: disable dynamic ID support usb: gadget: f_tcm: keep port count until LUN teardown completes xfrm: espintcp: fix UAF during close xfrm: drop ESP-in-TCP packets with no ingress device xfrm: ah6: validate routing header segments_left xfrm: fix xfrm_state_construct() auth-trunc leak net: bridge: mcast: fix use-after-free of a master VLAN's multicast context ipv6: seg6: clear IPv4 control block on IPIP decapsulation mm/swap: reject swapon() on filesystem-level encrypted files crypto: atmel-tdes - use scatterlist length before DMA mapping crypto: qce - fix CCM AAD buffer underallocation crypto: mxs-dcp - fix source scatterlist length access crypto: qce - Remove unsafe/deprecated algorithms KVM: s390: vsie: zero stale crypto bits usb: core: Add lock to usb_wakeup_notification() usb: core: Strengthen error handling in hub_hub_status() ALSA: usb-audio: fix OOB write in snd_usbmidi_novation_output() ALSA: usb-audio: Complete cleanup after system-resume errors USB: serial: option: fix slab OOB read in interrupt URB callback USB: serial: spcp8x5: drop broken carrier detect support USB: c67x00: fix use-after-free in c67x00_add_iso_urb() usb: usbfs: fix use-after-free of usb_device in usbdev_release() Linux 5.15.220 UBUNTU: Upstream stable to v5.15.220 ** Affects: linux (Ubuntu) Importance: Undecided Status: Invalid ** Affects: linux (Ubuntu Jammy) Importance: Medium Assignee: Vinicius Peixoto (vpeixoto) Status: In Progress ** Tags: kernel-stable-tracking-bug ** Changed in: linux (Ubuntu) Status: New => Confirmed ** Also affects: linux (Ubuntu Jammy) Importance: Undecided Status: New ** Changed in: linux (Ubuntu) Status: Confirmed => Invalid ** Changed in: linux (Ubuntu Jammy) Importance: Undecided => Medium ** Changed in: linux (Ubuntu Jammy) Status: New => In Progress ** Changed in: linux (Ubuntu Jammy) Assignee: (unassigned) => Vinicius Peixoto (vpeixoto) ** Description changed: SRU Justification Impact: The upstream process for stable tree updates is quite similar in scope to the Ubuntu SRU process, e.g., each patch has to demonstrably fix a bug, and each patch is vetted by upstream by originating either directly from a mainline/stable Linux tree or a minimally backported form of that patch. The following upstream stable patches should be included in the Ubuntu kernel: v5.15.220 upstream stable release from git://git.kernel.org/ - + RDMA/rxe: Fix OOB in free_rd_atomic_resources() + ext4: don't enable DAX on new encrypted files + io_uring/io-wq: fix worker accounting when canceling creation callbacks + ipvs: reload ip header after head reallocation + bpf: Remove tst_run from lwt_seg6local_prog_ops. + jfs: add check read-only before truncation in jfs_truncate_nolock() + jfs: add check read-only before txBeginAnon() call + can: j1939: implement NETDEV_UNREGISTER notification handler + can: j1939: add missing calls in NETDEV_UNREGISTER notification handler + can: j1939: make j1939_sk_bind() fail if device is no longer registered + KVM: arm64: Prevent access to vCPU events before init + bpf: Fix use-after-free in offloaded map/prog info fill + Revert "PM: sleep: Use complete() in device_pm_sleep_init()" + smc: Use __sk_dst_get() and dst_dev_rcu() in smc_vlan_by_tcpsk(). + selinux: switch two allocations to use kzalloc_objs() + ALSA: pcm: fix wait_time calculations + ASoC: tegra: Fix Master Volume Control + ALSA: pcm: fix use-after-free on linked stream runtime in snd_pcm_drain() + ALSA: PCM: Fix wait queue list corruption in snd_pcm_drain() on linked streams + ipv4: igmp: Fix potential UAF in igmp_gq_start_timer() + kcov: replace local_irq_save() with a local_lock_t + kcov: fix data corruption and race conditions on PREEMPT_RT + ext4: propagate errors from fast commit range replay + nilfs2: correct return value kernel-doc descriptions for ioctl functions + nilfs2: reject invalid block index in GC ioctl + nfc: nci: add data_len bound checks to activation parameter extractors + HID: magicmouse: prevent unbounded recursion in magicmouse_raw_event() + nvme: rename CDR/MORE/DNR to NVME_STATUS_* + nvmet-tcp: bound SGL data length before allocating command buffers + HID: uclogic: fix use-after-free of inrange_timer on remove + HID: ft260: fix i2c probing for hwmon devices + HID: ft260: improve i2c write performance + HID: ft260: improve i2c large reads performance + HID: ft260: skip unexpected HID input reports + HID: ft260: wake up device from power saving mode + HID: ft260: missed NACK from busy device + HID: ft260: validate i2c input report length + HID: ft260: fix stack-use-after-return write in I2C read race + HID: input: read battery capacity from its actual report offset + fpga: dfl: fme: add error handling + accessibility: speakup: unregister tty ldisc on later init failures + xhci: dbgtty: Fix unregister on tty_register_driver() failure + xhci: dbgtty: Fix unregister on tty_alloc_driver() failure + fuse: fix invalidate lock leak on setattr writeback failure + fuse: fix invalidate lock leak on open O_TRUNC DAX failure + usb: usbtest: disable dynamic ID support + usb: gadget: f_tcm: keep port count until LUN teardown completes + xfrm: espintcp: fix UAF during close + xfrm: drop ESP-in-TCP packets with no ingress device + xfrm: ah6: validate routing header segments_left + xfrm: fix xfrm_state_construct() auth-trunc leak + net: bridge: mcast: fix use-after-free of a master VLAN's multicast context + ipv6: seg6: clear IPv4 control block on IPIP decapsulation + mm/swap: reject swapon() on filesystem-level encrypted files + crypto: atmel-tdes - use scatterlist length before DMA mapping + crypto: qce - fix CCM AAD buffer underallocation + crypto: mxs-dcp - fix source scatterlist length access + crypto: qce - Remove unsafe/deprecated algorithms + KVM: s390: vsie: zero stale crypto bits + usb: core: Add lock to usb_wakeup_notification() + usb: core: Strengthen error handling in hub_hub_status() + ALSA: usb-audio: fix OOB write in snd_usbmidi_novation_output() + ALSA: usb-audio: Complete cleanup after system-resume errors + USB: serial: option: fix slab OOB read in interrupt URB callback + USB: serial: spcp8x5: drop broken carrier detect support + USB: c67x00: fix use-after-free in c67x00_add_iso_urb() + usb: usbfs: fix use-after-free of usb_device in usbdev_release() Linux 5.15.220 - usb: usbfs: fix use-after-free of usb_device in usbdev_release() - USB: c67x00: fix use-after-free in c67x00_add_iso_urb() - USB: serial: spcp8x5: drop broken carrier detect support - USB: serial: option: fix slab OOB read in interrupt URB callback - ALSA: usb-audio: Complete cleanup after system-resume errors - ALSA: usb-audio: fix OOB write in snd_usbmidi_novation_output() - usb: core: Strengthen error handling in hub_hub_status() - usb: core: Add lock to usb_wakeup_notification() - KVM: s390: vsie: zero stale crypto bits - crypto: qce - Remove unsafe/deprecated algorithms - crypto: mxs-dcp - fix source scatterlist length access - crypto: qce - fix CCM AAD buffer underallocation - crypto: atmel-tdes - use scatterlist length before DMA mapping - mm/swap: reject swapon() on filesystem-level encrypted files - ipv6: seg6: clear IPv4 control block on IPIP decapsulation - net: bridge: mcast: fix use-after-free of a master VLAN's multicast context - xfrm: fix xfrm_state_construct() auth-trunc leak - xfrm: ah6: validate routing header segments_left - xfrm: drop ESP-in-TCP packets with no ingress device - xfrm: espintcp: fix UAF during close - usb: gadget: f_tcm: keep port count until LUN teardown completes - usb: usbtest: disable dynamic ID support - fuse: fix invalidate lock leak on open O_TRUNC DAX failure - fuse: fix invalidate lock leak on setattr writeback failure - xhci: dbgtty: Fix unregister on tty_alloc_driver() failure - xhci: dbgtty: Fix unregister on tty_register_driver() failure - accessibility: speakup: unregister tty ldisc on later init failures - fpga: dfl: fme: add error handling - HID: input: read battery capacity from its actual report offset - HID: ft260: fix stack-use-after-return write in I2C read race - HID: ft260: validate i2c input report length - HID: ft260: missed NACK from busy device - HID: ft260: wake up device from power saving mode - HID: ft260: skip unexpected HID input reports - HID: ft260: improve i2c large reads performance - HID: ft260: improve i2c write performance - HID: ft260: fix i2c probing for hwmon devices - HID: uclogic: fix use-after-free of inrange_timer on remove - nvmet-tcp: bound SGL data length before allocating command buffers - nvme: rename CDR/MORE/DNR to NVME_STATUS_* - HID: magicmouse: prevent unbounded recursion in magicmouse_raw_event() - nfc: nci: add data_len bound checks to activation parameter extractors - nilfs2: reject invalid block index in GC ioctl - nilfs2: correct return value kernel-doc descriptions for ioctl functions - ext4: propagate errors from fast commit range replay - kcov: fix data corruption and race conditions on PREEMPT_RT - kcov: replace local_irq_save() with a local_lock_t - ipv4: igmp: Fix potential UAF in igmp_gq_start_timer() - ALSA: PCM: Fix wait queue list corruption in snd_pcm_drain() on linked streams - ALSA: pcm: fix use-after-free on linked stream runtime in snd_pcm_drain() - ASoC: tegra: Fix Master Volume Control - ALSA: pcm: fix wait_time calculations - Revert "smb: client: use kvzalloc() for megabyte buffer in simple fallocate" - Revert "mtd: maps: vmu-flash: fix fault in unaligned fixup" - selinux: switch two allocations to use kzalloc_objs() - smc: Use __sk_dst_get() and dst_dev_rcu() in smc_vlan_by_tcpsk(). - Revert "PM: sleep: Use complete() in device_pm_sleep_init()" - bpf: Fix use-after-free in offloaded map/prog info fill - KVM: arm64: Prevent access to vCPU events before init - can: j1939: make j1939_sk_bind() fail if device is no longer registered - can: j1939: add missing calls in NETDEV_UNREGISTER notification handler - can: j1939: implement NETDEV_UNREGISTER notification handler - jfs: add check read-only before txBeginAnon() call - jfs: add check read-only before truncation in jfs_truncate_nolock() - bpf: Remove tst_run from lwt_seg6local_prog_ops. - ipvs: reload ip header after head reallocation - io_uring/io-wq: fix worker accounting when canceling creation callbacks - ext4: don't enable DAX on new encrypted files - RDMA/rxe: Fix OOB in free_rd_atomic_resources() + UBUNTU: Upstream stable to v5.15.220 -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168586 Title: Jammy update: v5.15.220 upstream stable release Status in linux package in Ubuntu: Invalid Status in linux source package in Jammy: In Progress Bug description: SRU Justification Impact: The upstream process for stable tree updates is quite similar in scope to the Ubuntu SRU process, e.g., each patch has to demonstrably fix a bug, and each patch is vetted by upstream by originating either directly from a mainline/stable Linux tree or a minimally backported form of that patch. The following upstream stable patches should be included in the Ubuntu kernel: v5.15.220 upstream stable release from git://git.kernel.org/ RDMA/rxe: Fix OOB in free_rd_atomic_resources() ext4: don't enable DAX on new encrypted files io_uring/io-wq: fix worker accounting when canceling creation callbacks ipvs: reload ip header after head reallocation bpf: Remove tst_run from lwt_seg6local_prog_ops. jfs: add check read-only before truncation in jfs_truncate_nolock() jfs: add check read-only before txBeginAnon() call can: j1939: implement NETDEV_UNREGISTER notification handler can: j1939: add missing calls in NETDEV_UNREGISTER notification handler can: j1939: make j1939_sk_bind() fail if device is no longer registered KVM: arm64: Prevent access to vCPU events before init bpf: Fix use-after-free in offloaded map/prog info fill Revert "PM: sleep: Use complete() in device_pm_sleep_init()" smc: Use __sk_dst_get() and dst_dev_rcu() in smc_vlan_by_tcpsk(). selinux: switch two allocations to use kzalloc_objs() ALSA: pcm: fix wait_time calculations ASoC: tegra: Fix Master Volume Control ALSA: pcm: fix use-after-free on linked stream runtime in snd_pcm_drain() ALSA: PCM: Fix wait queue list corruption in snd_pcm_drain() on linked streams ipv4: igmp: Fix potential UAF in igmp_gq_start_timer() kcov: replace local_irq_save() with a local_lock_t kcov: fix data corruption and race conditions on PREEMPT_RT ext4: propagate errors from fast commit range replay nilfs2: correct return value kernel-doc descriptions for ioctl functions nilfs2: reject invalid block index in GC ioctl nfc: nci: add data_len bound checks to activation parameter extractors HID: magicmouse: prevent unbounded recursion in magicmouse_raw_event() nvme: rename CDR/MORE/DNR to NVME_STATUS_* nvmet-tcp: bound SGL data length before allocating command buffers HID: uclogic: fix use-after-free of inrange_timer on remove HID: ft260: fix i2c probing for hwmon devices HID: ft260: improve i2c write performance HID: ft260: improve i2c large reads performance HID: ft260: skip unexpected HID input reports HID: ft260: wake up device from power saving mode HID: ft260: missed NACK from busy device HID: ft260: validate i2c input report length HID: ft260: fix stack-use-after-return write in I2C read race HID: input: read battery capacity from its actual report offset fpga: dfl: fme: add error handling accessibility: speakup: unregister tty ldisc on later init failures xhci: dbgtty: Fix unregister on tty_register_driver() failure xhci: dbgtty: Fix unregister on tty_alloc_driver() failure fuse: fix invalidate lock leak on setattr writeback failure fuse: fix invalidate lock leak on open O_TRUNC DAX failure usb: usbtest: disable dynamic ID support usb: gadget: f_tcm: keep port count until LUN teardown completes xfrm: espintcp: fix UAF during close xfrm: drop ESP-in-TCP packets with no ingress device xfrm: ah6: validate routing header segments_left xfrm: fix xfrm_state_construct() auth-trunc leak net: bridge: mcast: fix use-after-free of a master VLAN's multicast context ipv6: seg6: clear IPv4 control block on IPIP decapsulation mm/swap: reject swapon() on filesystem-level encrypted files crypto: atmel-tdes - use scatterlist length before DMA mapping crypto: qce - fix CCM AAD buffer underallocation crypto: mxs-dcp - fix source scatterlist length access crypto: qce - Remove unsafe/deprecated algorithms KVM: s390: vsie: zero stale crypto bits usb: core: Add lock to usb_wakeup_notification() usb: core: Strengthen error handling in hub_hub_status() ALSA: usb-audio: fix OOB write in snd_usbmidi_novation_output() ALSA: usb-audio: Complete cleanup after system-resume errors USB: serial: option: fix slab OOB read in interrupt URB callback USB: serial: spcp8x5: drop broken carrier detect support USB: c67x00: fix use-after-free in c67x00_add_iso_urb() usb: usbfs: fix use-after-free of usb_device in usbdev_release() Linux 5.15.220 UBUNTU: Upstream stable to v5.15.220 To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168586/+subscriptions

[Bug 2168581] Re: Kernel 7.3.0-5 / -6: ZFS not working properly

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168581 Title: Kernel 7.3.0-5 / -6: ZFS not working properly Status in linux package in Ubuntu: New Status in zfs-linux package in Ubuntu: New Bug description: Ubuntu 26.10 Stonking (Development version from 2026-09-20). Kernel: 7.3.0-5 and 7.3.0-6 (from -proposed) On a system with a ZFS root, after installation, I am unable to create or use any additional zfs pools. Details: Encrypted ZFS or not doesn't matter Kernel 7.3.0-5 (OpenZFS 2.4.4-1ubuntu2) and 7.3.0-6 (OpenZFS 2.4.4-1ubuntu3): FAILS. Kernel 7.2.0-5 with OpenZFS 2.4.2: WORKS. Kernel 7.2.0-5 with OpenZFS 2.4.4 (OpenZFS 2.4.4-1ubuntu3): WORKS. No -proposed, only kernel from -proposed, or all from -proposed as of submission time tested and gave same results. Quick reproduction way: 1) Install Ubuntu 26.10 Stonking development with ZFS root. Encrypted or not does not matter. 2) Boot system. It works as it should. 3) Run the following command: sudo -i bash -c 'uname -r; zfs version; truncate -s 1G /root/t.img && zpool create tp /root/t.img; echo "rc=${?}"; zpool list -H tp; zpool destroy tp 2>/dev/null; rm -f /root/t.img' What it does: - Prints out the kernel and zfs information. - Creates a 1G file in /root: /root/t.img. - Creates a zpool called "tp" on /root/t.img. - Prints the rc. Working should be 0. Not working is 1. - Displays the zpool. - Destroys the pool. - Removes the file from /root. Expected result: rc=0 and a created zpool. Received result: rc=1 and no zpool. If you create the partitions manually and make a manual installation including a home directory pool (hpool), the hpool cannot be imported, as the userspace receives ENOENT. There is a patch to fix this (debian/patches/linux-7.3-kthread-nullfs.patch), but it is insufficient, as it stores the initial root directory, which is then the initramfs and not the running root. I have written a patch that I will submit after I create the bug report that fixes it in all scenarios I have thought of. The log supplied shows the commands executed and the outcome. With my patch it gives rc=0 and works. // Stefan ProblemType: Bug DistroRelease: Ubuntu 26.10 Package: zfsutils-linux 2.4.4-1ubuntu3 ProcVersionSignature: Ubuntu 7.3.0-5.5-generic 7.3.0-rc3 Uname: Linux 7.3.0-5-generic x86_64 NonfreeKernelModules: zfs ApportVersion: 2.36.0-0ubuntu1 Architecture: amd64 CasperMD5CheckResult: pass CurrentDesktop: ubuntu:GNOME Date: Fri Sep 25 21:14:20 2026 InstallationDate: Installed on 2026-09-25 (0 days ago) InstallationMedia: Ubuntu 26.10 "Stonking Stingray" - Daily amd64 (20260920) ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color SourcePackage: zfs-linux UpgradeStatus: No upgrade log present (probably fresh install) To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168581/+subscriptions

[Bug 2168537] Re: Please backport "drm/amdgpu: fix check in amdgpu_hmm_invalidate_gfx" (52f650963d88) to resolute 7.0 - NULL deref in kcompactd kills the system

** Tags added: kernel-daily-bug -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2168537 Title: Please backport "drm/amdgpu: fix check in amdgpu_hmm_invalidate_gfx" (52f650963d88) to resolute 7.0 - NULL deref in kcompactd kills the system Status in linux package in Ubuntu: New Bug description: [Impact] On Ubuntu 26.04 (linux 7.0.0-31 and 7.0.0-34, based on 7.0.14), with an AMD GPU using amdgpu userptr BOs, kcompactd page migration calls amdgpu_hmm_invalidate_gfx() while a userptr BO is still being allocated or freed. At that point bo->vm_bo is NULL, and the kernel dereferences it: BUG: kernel NULL pointer dereference, address: 0000000000000000 Oops: 0000 [#1] SMP PTI CPU: 6 UID: 0 PID: 74 Comm: kcompactd0 Not tainted 7.0.0-31-generic #31-Ubuntu Hardware name: ASUS All Series/B85M-E, BIOS 0806 02/20/2014 RIP: 0010:amdgpu_hmm_invalidate_gfx+0x42/0xe0 [amdgpu] Call Trace: __mmu_notifier_invalidate_range_start+0x149/0x1b0 try_to_migrate_one+0xeea/0x1010 rmap_walk_anon+0xdb/0x220 try_to_migrate+0x84/0xe0 migrate_folio_unmap+0x276/0x320 migrate_pages_batch+0x17d/0x940 migrate_pages+0x3b2/0x520 compact_zone+0x513/0x790 kcompactd_do_work+0x109/0x270 note: kcompactd0[74] exited with irqs disabled kcompactd dies holding mm locks. After that, any task touching the affected pages hangs in D state for good: containers can't be stopped or killed, and even a normal reboot oopses again in amdgpu_pci_shutdown -> amdgpu_uvd_prepare_suspend (direct compaction hits the same page). Only a sysrq-b or power cycle recovers the machine. We have hit this 3 times in 2 weeks (2026-09-12, 09-19, 09-25) on a Radeon RX 550 (Polaris 12). The trigger is Jellyfin burning ASS subtitles in with an ffmpeg Vulkan filter chain: subtitles -> hwupload=derive_device=vulkan -> overlay_vulkan The oops fires about 2 seconds after an ffmpeg transcode starts. [Fix] Upstream commit 52f650963d8825e97a0ccdd2b616f8a01d9d3d38 "drm/amdgpu: fix check in amdgpu_hmm_invalidate_gfx" (cherry-picked from 631849ff5d603841e74f19f4a5e30fe1f7d7cf30) Fixes: 91250893cbaa ("drm/amdgpu: fix waiting for all submissions for userptrs") Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5399 Cc: stable, queued for 7.1.y and 6.18.y The commit applies cleanly to the resolute 7.0.0-34.34 source (one hunk at offset -4). The regressing commit 91250893cbaa is in the resolute kernel. The fix did not reach 7.0.y stable because that branch is EOL upstream. Note: CVE-2026-68101 was assigned to this fix and later rejected, so the Ubuntu CVE tracker shows resolute as "not affected". It is affected. This is a stability bug, not a security bug, and it is reproducible. [Test Plan] On an amdgpu system, run a workload that creates and destroys userptr BOs repeatedly (for example, ffmpeg Vulkan hwupload in a loop) while forcing compaction with `echo 1 > /proc/sys/vm/compact_memory`. Unpatched kernels oops in amdgpu_hmm_invalidate_gfx; patched kernels don't. [Where problems could occur] The change is limited to amdgpu userptr BOs. The BO now holds a reference to the VM root PD as bo->parent, and amdgpu_bo_destroy() already drops that reference (amdgpu_bo_unref(&bo->parent)). A mistake would show up as a leaked root PD on userptr BO free, or as a wait on the wrong reservation object during invalidation. --- ProblemType: Bug AlsaVersion: Advanced Linux Sound Architecture Driver Version k7.0.0-34-generic. AplayDevices: Error: [Errno 2] No such file or directory: 'aplay' ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 ArecordDevices: Error: [Errno 2] No such file or directory: 'arecord' AudioDevicesInUse: Error: command ['fuser', '-v', '/dev/snd/by-path', '/dev/snd/controlC0', '/dev/snd/hwC0D0', '/dev/snd/pcmC0D2c', '/dev/snd/pcmC0D0c', '/dev/snd/pcmC0D0p', '/dev/snd/controlC1', '/dev/snd/hwC1D0', '/dev/snd/pcmC1D10p', '/dev/snd/pcmC1D9p', '/dev/snd/pcmC1D8p', '/dev/snd/pcmC1D7p', '/dev/snd/pcmC1D3p', '/dev/snd/seq', '/dev/snd/timer'] failed with exit code 1: CRDA: N/A Card0.Amixer.info: Error: [Errno 2] No such file or directory: 'amixer' Card0.Amixer.values: Error: [Errno 2] No such file or directory: 'amixer' Card1.Amixer.info: Error: [Errno 2] No such file or directory: 'amixer' Card1.Amixer.values: Error: [Errno 2] No such file or directory: 'amixer' CasperMD5CheckResult: pass CurrentDmesg: Error: command ['dmesg'] failed with exit code 1: dmesg: read kernel buffer failed: Operation not permitted DistroRelease: Ubuntu 26.04 InstallationDate: Installed on 2026-08-30 (26 days ago) InstallationMedia: Ubuntu-Server 26.04.1 LTS "Resolute Raccoon" - Release amd64 (20260826) MachineType: ASUS All Series Package: linux (not installed) ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color XDG_RUNTIME_DIR=<set> ProcFB: 0 amdgpudrmfb ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-34-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M ProcVersionSignature: Ubuntu 7.0.0-34.34-generic 7.0.14 RfKill: Error: [Errno 2] No such file or directory: 'rfkill' Tags: resolute Uname: Linux 7.0.0-34-generic x86_64 UnreportableReason: This report is about a package that is not installed. UpgradeStatus: No upgrade log present (probably fresh install) UserGroups: adm cdrom dip docker lxd plugdev render sudo users video _MarkForUpload: False acpidump: dmi.bios.date: 02/20/2014 dmi.bios.release: 4.6 dmi.bios.vendor: American Megatrends Inc. dmi.bios.version: 0806 dmi.board.asset.tag: To be filled by O.E.M. dmi.board.name: B85M-E dmi.board.vendor: ASUSTeK COMPUTER INC. dmi.board.version: Rev X.0x dmi.chassis.asset.tag: Asset-1234567890 dmi.chassis.type: 3 dmi.chassis.vendor: Chassis Manufacture dmi.chassis.version: Chassis Version dmi.modalias: dmi:bvnAmericanMegatrendsInc.:bvr0806:bd02/20/2014:br4.6:svnASUS:pnAllSeries:pvrSystemVersion:rvnASUSTeKCOMPUTERINC.:rnB85M-E:rvrRevX.0x:cvnChassisManufacture:ct3:cvrChassisVersion:skuAll:pfaASUSMB: dmi.product.family: ASUS MB dmi.product.name: All Series dmi.product.sku: All dmi.product.version: System Version dmi.sys.vendor: ASUS To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2168537/+subscriptions