LP#2167595 - field feedback on the v2 patch (lp2167595_mt7921u_usb_disconnect_deadlock_v2.patch, sha256 c83bc45d...4fefd1) == Environment == Ubuntu 26.04.1 LTS, kernel 7.0.0-34-generic (Ubuntu 7.0.0-34.34, upstream 7.0.14) Adapter: MediaTek MT7921AU, USB ID 0e8d:7961, driver mt7921u, high-speed (USB 2.0) port Hardware: Dell Precision 3650 Tower, BIOS 1.48.0 v2 built out-of-tree from linux-source-7.0.0 (7.0.0-34.34), installed into /lib/modules/7.0.0-34-generic/updates/ and in daily use since 2026-09-26. Kernel taint is O+E only (out-of-tree unsigned modules: these plus VirtualBox). Logs below are sanitized: hostname, BSSID and adapter MAC replaced. == 1. Real-world validation: the deadlock is gone == On 2026-09-26 the adapter failed exactly the way it did in the original report: 8 consecutive "vendor request ... failed:-110" timeouts, Bluetooth on the same chip timing out as well, then a USB reset loop from which the device never returned (descriptor read errors -110, "device not accepting address" -62, finally repeated re-enumeration attempts that all failed). Timeline (local time, 26.09.2026): 16:29:28 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110 16:29:31 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110 16:29:34 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d058 failed:-110 16:29:38 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53b8 failed:-110 16:29:40 kernel: Bluetooth: hci0: Opcode 0x0401 failed: -110 16:29:40 kernel: Bluetooth: hci0: command 0x0401 tx timeout 16:29:41 kernel: mt7921u 1-9:1.3: vendor request req:63 off:53c4 failed:-110 16:29:44 kernel: mt7921u 1-9:1.3: vendor request req:66 off:53c4 failed:-110 16:29:47 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d02c failed:-110 16:29:50 kernel: mt7921u 1-9:1.3: vendor request req:63 off:d054 failed:-110 16:29:51 kernel: Bluetooth: hci0: Failed to write uhw reg(-110) 16:29:53 kernel: wlx00c0cab8f19c: deauthenticating from [BSSID] by local choice (Reason: 3=DEAUTH_LEAVING) 16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DISCONNECTED bssid=[BSSID] reason=3 locally_generated=1 16:30:07 wpa_supplicant[2255]: wlx00c0cab8f19c: Added BSSID [BSSID] into ignore list, ignoring for 10 seconds 16:30:08 kernel: wlx00c0cab8f19c: failed to remove key (1, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:09 kernel: wlx00c0cab8f19c: failed to remove key (2, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:11 kernel: wlx00c0cab8f19c: failed to remove key (4, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:12 kernel: wlx00c0cab8f19c: failed to remove key (5, ff:ff:ff:ff:ff:ff) from hardware (-110) 16:30:13 kernel: mt7921u 1-9:1.3: timed out waiting for pending tx 16:30:13 kernel: snd_soc_acpi_intel_sdca_quirks soundwire_generic_allocation snd_soc_sdw_utils snd_soc_acpi intel_rapl_msr soundwire_bus in 16:30:14 NetworkManager[3140]: device (wlx00c0cab8f19c): state change: activated -> unmanaged (reason 'unmanaged-link-not-init', managed-typ 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): canceled DHCP transaction 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): activation: beginning transaction (timeout in 45 seconds) 16:30:14 NetworkManager[3140]: dhcp4 (wlx00c0cab8f19c): state changed no lease 16:30:14 ModemManager[2292]: <msg> [base-manager] port wlx00c0cab8f19c released by device '/sys/devices/pci0000:00/0000:00:14.0/usb1/1-9' 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: PMKSA-CACHE-REMOVED [BSSID] 0 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: Removed BSSID [BSSID] from ignore list (clear) 16:30:14 wpa_supplicant[2255]: wlx00c0cab8f19c: CTRL-EVENT-DSCP-POLICY clear_all 16:30:14 wpa_supplicant[2255]: nl80211: deinit ifname=wlx00c0cab8f19c disabled_11b_rates=0 16:30:14 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:30:20 kernel: usb 1-9: device descriptor read/64, error -110 16:30:24 systemd[1]: NetworkManager-dispatcher.service: Deactivated successfully. 16:30:35 kernel: usb 1-9: device descriptor read/64, error -110 16:30:36 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:30:41 kernel: usb 1-9: device descriptor read/64, error -110 16:30:57 kernel: usb 1-9: device descriptor read/64, error -110 16:30:57 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:31:08 kernel: usb 1-9: device not accepting address 3, error -62 16:31:08 kernel: usb 1-9: reset high-speed USB device number 3 using xhci_hcd 16:31:20 kernel: usb 1-9: device not accepting address 3, error -62 16:31:20 kernel: usb 1-9: USB disconnect, device number 3 16:31:20 kernel: usb 1-9: new high-speed USB device number 5 using xhci_hcd 16:31:25 kernel: usb 1-9: device descriptor read/64, error -110 16:31:41 kernel: usb 1-9: device descriptor read/64, error -110 16:31:41 kernel: usb 1-9: new high-speed USB device number 6 using xhci_hcd 16:31:47 kernel: usb 1-9: device descriptor read/64, error -110 16:32:03 kernel: usb 1-9: device descriptor read/64, error -110 16:32:03 kernel: usb 1-9: new high-speed USB device number 7 using xhci_hcd 16:32:14 kernel: usb 1-9: device not accepting address 7, error -62 16:32:14 kernel: usb 1-9: new high-speed USB device number 8 using xhci_hcd 16:32:26 kernel: usb 1-9: device not accepting address 8, error -62 16:45:01 systemd[7730]: Reached target shutdown.target - Shutdown. 16:45:04 systemd[1]: NetworkManager-wait-online.service: Deactivated successfully. 16:45:05 systemd[1]: NetworkManager.service: Deactivated successfully. 16:45:09 systemd[1]: Reached target shutdown.target - System Shutdown. 16:45:09 systemd[1]: Reached target poweroff.target - System Power Off. 16:45:09 systemd-shutdown[1]: Syncing filesystems and block devices. 16:45:09 systemd-shutdown[1]: Sending SIGTERM to remaining processes... 16:45:09 systemd-journald[1051]: Journal stopped Outcome with v2: * No task ever blocked on rtnl_lock. Zero "blocked for more than N seconds" in the whole boot. * The system stayed fully responsive for the remaining 15 minutes of the session (only WiFi was gone) and then shut down normally: the full shutdown sequence took 8 seconds, from 16:45:01 to "Journal stopped" at 16:45:09. * No "chip reset failed", no "rx urb mismatch". For comparison, the same hardware failure on stock/v1 modules (2026-09-15, the original report) left NetworkManager, ip, and 6 other tasks in D state for more than 122 seconds, all waiting for rtnl_lock held by a kworker in mt792xu_disconnect -> mt76_unregister_device -> mt7921_abort_roc, and the machine had to be powered off with the button. So the core goal of the patch is confirmed in the field, not just in tests. == 2. New finding: the tx_worker disable in mt792xu_disconnect is redundant and fires a WARNING on the dead-device path == Alongside the (harmless) "timed out waiting for pending tx" message, v2 produced a one-shot WARNING: 16:30:13 mt7921u 1-9:1.3: timed out waiting for pending tx 16:30:13 ------------[ cut here ]------------ 16:30:13 WARNING: kernel/kthread.c:722 at kthread_park+0x8c/0xc0, CPU#7: kworker/7:2/100058 16:30:13 CPU: 7 UID: 0 PID: 100058 Comm: kworker/7:2 Tainted: G OE 7.0.0-34-generic #34-Ubuntu PREEMPT(full) 16:30:13 Tainted: [O]=OOT_MODULE, [E]=UNSIGNED_MODULE 16:30:13 Hardware name: Dell Inc. Precision 3650 Tower/0NDYHG, BIOS 1.48.0 05/27/2026 16:30:13 Workqueue: events __usb_queue_reset_device 16:30:13 FS: 0000000000000000(0000) GS:ffff8de34f77f000(0000) knlGS:0000000000000000 16:30:13 CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033 16:30:13 Call Trace: 16:30:13 <TASK> 16:30:13 ? mt76u_stop_tx.cold+0x11c/0x180 [mt76_usb] 16:30:13 ? __pfx_autoremove_wake_function+0x10/0x10 16:30:13 mt792xu_stop+0x1a/0x40 [mt792x_usb] 16:30:13 drv_stop+0x50/0x160 [mac80211] 16:30:13 ieee80211_stop_device+0x7f/0x90 [mac80211] 16:30:13 ieee80211_do_stop+0x6a2/0xb20 [mac80211] 16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20 16:30:13 ? packet_notifier+0x85/0x280 16:30:13 ieee80211_stop+0x63/0xf0 [mac80211] 16:30:13 __dev_close_many+0xb2/0x220 16:30:13 netif_close_many+0xb7/0x1b0 16:30:13 netif_close+0x70/0xa0 16:30:13 dev_close+0x38/0xb0 16:30:13 cfg80211_shutdown_all_interfaces+0x50/0x100 [cfg80211] 16:30:13 ieee80211_remove_interfaces+0x47/0x220 [mac80211] 16:30:13 ieee80211_unregister_hw+0x4a/0x140 [mac80211] 16:30:13 mt76_unregister_device+0x63/0x80 [mt76] 16:30:13 mt792xu_disconnect+0xea/0x150 [mt792x_usb] 16:30:13 usb_unbind_interface+0x9b/0x2c0 16:30:13 device_remove+0x68/0x80 16:30:13 device_release_driver_internal+0x1fb/0x260 16:30:13 device_release_driver+0x12/0x20 16:30:13 usb_forced_unbind_intf+0x96/0xe0 16:30:13 ? usb_autoresume_device+0x1e/0x70 16:30:13 usb_reset_device+0xf4/0x300 16:30:13 __usb_queue_reset_device+0x3b/0x60 16:30:13 process_one_work+0x1ac/0x3d0 16:30:13 worker_thread+0x1b8/0x360 16:30:13 ? _raw_spin_lock_irqsave+0xe/0x20 16:30:13 ? __pfx_worker_thread+0x10/0x10 16:30:13 kthread+0xf7/0x130 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ret_from_fork+0x195/0x2a0 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ? __pfx_kthread+0x10/0x10 16:30:13 ret_from_fork_asm+0x1a/0x30 16:30:13 </TASK> 16:30:13 ---[ end trace 0000000000000000 ]--- Analysis: 1. v2 parks tx_worker early in mt792xu_disconnect(): mt76_worker_disable(&dev->mt76.tx_worker); 2. Later in the same function, mt76_unregister_device() -> ieee80211_unregister_hw() -> ... -> drv_stop() -> mt792xu_stop() -> mt76u_stop_tx(). 3. mt76u_stop_tx() (drivers/net/wireless/mediatek/mt76/usb.c:994) waits HZ/5 for pending TX to drain. On timeout it logs the message, kills the TX URBs and calls mt76_worker_disable(&dev->tx_worker) - a SECOND park of the same worker. kthread_park() then hits WARN_ON_ONCE(test_bit(KTHREAD_SHOULD_PARK, &kthread->flags)) (kernel/kthread.c:722) and returns -EBUSY. 4. At the end of that same branch mt76u_stop_tx() calls mt76_worker_enable(&dev->tx_worker), i.e. it UNPARKS the worker that v2 deliberately stopped - so the patch line loses its effect precisely in the case it was meant to cover. Suggestion: drop mt76_worker_disable(&dev->mt76.tx_worker); from mt792xu_disconnect(). mt76u_stop_tx() already quiesces TX during unregister and pairs its own park/unpark correctly. The line is not part of the deadlock fix (flags + worker cancellation + the abort_roc early return are), and on the dead-device path it is actively counterproductive. Why lab testing does not catch this: the frame in the trace is mt76u_stop_tx.cold, i.e. the unlikely branch. With a healthy adapter the TX queues drain far below the 200 ms timeout, so that branch never executes. On 2026-09-24 we ran three consecutive rmmod/insmod cycles plus an unplug-during-transfer test on v2 and saw no WARNING at all; the two MCU timeout messages were the only output. It took a genuine adapter failure, with TX still queued at disconnect time, to reach it. Impact: cosmetic in effect (one WARNING, W taint) plus the defeated worker disable. Teardown continued and completed; nothing hung. == 3. Three observations from reviewing v2 against the 7.0.14 sources == (1) mt7925_mac_reset_work() has NO MT76_REMOVED check in 7.0.14. The v2 description says the new mt7921 check aligns mt7921 with mt7925, but on the mt7925 side the check does not exist, so mt7925 USB keeps the race that v2 fixes for mt7921. (grep confirms: MT76_REMOVED appears in mt7925/pci.c and mt792x_dma.c only, not in mt7925/mac.c.) (2) Cosmetic: mt7921_mcu_parse_response() (mt7921/mcu.c:26) prints "Message %08x (seq %d) timeout" and calls mt792x_reset() without checking MT76_MCU_RESET. Because v2 sets MT76_MCU_RESET at disconnect entry, mt76_mcu_get_response() returns immediately, so every teardown MCU command logs a timeout ~30 ms after "deregistering interface driver" - in our logs exactly two per unload: MCU_UNI_CMD(BSS_INFO_UPDATE) (0x00020002) from interface removal and MCU_EXT_CMD(MAC_INIT_CTRL) (0x000046ed) from mt792x_stop() -> mt76_connac_mcu_set_mac_enable(). A MT76_MCU_RESET check before dev_err()/mt792x_reset() would suppress noise that is expected by design. mt7925/mcu.c:22 has the same print. (3) Behaviour note: with MT76_REMOVED set on entry, the teardown MCU commands and mt792xu_wfsys_reset() inside mt792xu_cleanup() fail with -EIO, so a plain rmmod of a healthy adapter leaves the firmware running. In practice this is harmless because mt7921u_probe() resets WFSYS when MT_TOP_MISC2_FW_N9_RDY is set, but it is a behaviour change worth knowing. == 4. Reproduction attempts: the branch cannot be reached on healthy hardware == We tried to reproduce the WARNING deliberately, to confirm the explanation above rather than rely on a single capture. Three attempts on 2026-09-27, all with the v2 modules loaded (mt792x_usb srcversion E7EEAE2D193E09CC1645423): # method TX load before trigger result 1 sysfs driver unbind 6336 pkt in 10 s (633/s) not reproduced (echo 1-9:1.3 > /sys/bus/usb/drivers/mt7921u/unbind) 2 sysfs driver unbind, longer load 7768 pkt in 30 s (258/s) not reproduced 3 port deauthorisation 6287 pkt in 10 s (628/s) not reproduced (echo 0 > /sys/bus/usb/devices/1-9/authorized) In all three runs the counters were identical: timed out waiting for pending tx 0 WARNING 0 kthread_park 0 blocked for 0 TX always drained well inside the HZ/5 window, so mt76u_stop_tx() took the normal path and never called mt76_worker_disable() a second time. The only kernel output in each run was the expected MCU teardown noise described in section 3(2): 15 "Message ... timeout" lines per unload, for commands 0x00020002, 0x00020003, 0x00020006 and 0x000046ed. Teardown and re-probe completed every time; the interface came back after 3 s in all three runs. This is consistent with the analysis: the .cold branch only executes when TX cannot drain, i.e. when the chip is dead or stalled, and that state cannot be induced on demand on healthy hardware. Forcing a disconnect - whether by unbinding the driver or by deauthorising the port - still leaves the device answering on the bus, so the queues empty immediately. In other words, the failed reproduction strengthens rather than weakens the finding: it is direct evidence of why routine testing (our own 2026-09-24 runs included: three rmmod/insmod cycles plus unplug-during-transfer, no WARNING at all) cannot surface this path. The 2026-09-26 trace in section 2 remains the only capture of it, obtained during a genuine adapter failure. == Summary == v2 solves the problem it targets: a real adapter death no longer takes the networking stack or shutdown down with it. The only item we would ask you to change is the redundant mt76_worker_disable(&dev->mt76.tx_worker) in mt792xu_disconnect described in section 2. As a testing hint: the .cold branch can be exercised artificially by temporarily shortening the HZ/5 timeout in mt76u_stop_tx in a local build, which may help validate a revised patch without waiting for a hardware failure. Happy to test a revised patch on this hardware. -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2167595 Title: mt7921u: USB reset after -110 timeouts deadlocks in mt7921_abort_roc, blocks rtnl_lock and shutdown Status in linux package in Ubuntu: Confirmed Bug description: [ Impact ] When an MT7921AU USB Wi-Fi dongle (e.g. 0e8d:7961) experiences communication stalls (USB -110 ETIMEDOUT errors), mt792x_mac_work holds dev->mt76.mutex while looping on register reads. When USB core detects the stall and triggers unbind, mt792xu_disconnect calls mt76_unregister_device, which takes rtnl_lock and calls mt7921_abort_roc. mt7921_abort_roc then tries to acquire dev->mt76.mutex, which is already held by the stuck mac_work thread. This creates an ABBA deadlock between rtnl_lock and mt76.mutex. As a result: - mac_work waits on USB timeouts with mt76.mutex held. - mt792xu_disconnect waits for mt76.mutex under rtnl_lock. - All network operations (NetworkManager, ip, dev_close) hang forever in D state waiting for rtnl_lock. - The system cannot power off or reboot cleanly without a hard reset. [ Fix ] 1. In mt792xu_disconnect(), set MT76_REMOVED, MT76_RESET, and MT76_MCU_RESET flags and wake pending waitqueues before calling mt76_unregister_device(). Setting MT76_REMOVED causes all pending and subsequent USB register requests to fail immediately with -EIO instead of waiting for 3-second timeouts. 2. In mt792xu_disconnect(), synchronously cancel all workers (mac_work, ps_work, wake_work, reset_work, init_work) prior to unregistration. 3. In mt7921_abort_roc() (and mt7925_abort_roc()), check if MT76_REMOVED is set. If the device was removed, clear MT76_STATE_ROC and return 0 immediately without taking dev->mt76.mutex. [ Test Plan ] 1. Boot kernel 7.0.0-31-generic with patched mt7921u / mt792x-usb modules. 2. Insert MT7921AU USB Wi-Fi adapter (0e8d:7961) and connect to an AP. 3. Physically disconnect the USB adapter while network traffic is active, or trigger bus reset while requests are pending. 4. Verify in dmesg: - Device disconnect completes cleanly without call traces. - rtnl_lock is not blocked; NetworkManager and ip link operate normally. - System reboots and powers off cleanly without hanging in D state. [ Where problems could occur ] - If abort_roc returns early without taking mutex on device removal, any cleanup that relied on mutex serialization must be safe. Since MT76_REMOVED is set, hardware registers cannot be accessed anyway, so skipping mutex and clearing local flag is safe. - Canceling works before unregistering prevents concurrent execution during teardown, which is the desired behavior on unbind. [ Other Info ] Patch tested and compiled against linux-headers-7.0.0-31-generic on Ubuntu 26.04 (Resolute). Clean compile with 0 errors and 0 warnings. --- [ Original Report ] Dell Precision 3650 Tower, Ubuntu 26.04 dev kernel 7.0.0-31-generic. Alfa AWUS036AXM (MT7921AU, USB ID 0e8d:7961). Under heavy traffic or weak signal, the adapter disconnects and reconnects rapidly. Eventually kernel dmesg shows: [ 342.112004] mt7921u 1-2:1.0: Message 00000040 (seq 4) timeout [ 345.184002] mt7921u 1-2:1.0: Message 00000040 (seq 5) timeout [ 348.256011] mt7921u 1-2:1.0: Failed to get patch sem [ 351.328008] mt7921u 1-2:1.0: hardware init failed After that, NetworkManager stops responding. Running 'ip link' hangs indefinitely in D state. Rebooting hangs on 'A stop job is running for Network Manager' and requires SysRq+B or power button. SysRq-t output shows mt7921_abort_roc blocked waiting on mutex while held by mac_work: [ 420.100012] task:kworker/u16:3 blocked for more than 120 seconds. [ 420.100020] Call Trace: [ 420.100025] __schedule+0x345/0x890 [ 420.100030] schedule+0x5a/0xc0 [ 420.100035] schedule_preempt_disabled+0x18/0x30 [ 420.100040] __mutex_lock.isra.0+0x28a/0x4b0 [ 420.100045] mt7921_abort_roc+0x2d/0x80 [mt7921_common] [ 420.100050] ieee80211_set_disassoc+0x62/0x90 [cfg80211] [ 420.100055] mt76_unregister_device+0x48/0x90 [mt76] [ 420.100060] mt792xu_disconnect+0x3c/0x70 [mt792x_usb] To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2167595/+subscriptions
Комментариев нет:
Отправить комментарий