пятница

[Bug 2163682] Re: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W

Second attachment: dmesg from the same 7.0.0-29.29 boot after unbinding and rebinding snd_hda_intel on 0000:02:00.1, with the GPU in D3cold. Same kernel, healthy. ** Attachment added: "dmesg from the same 7.0.0-29.29 boot after the HDA rebind, GPU in D3cold" https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+attachment/5996189/+files/dmesg-7.0.0-29-generic-after-hda-rebind.log -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2163682 Title: Regression 7.0.0-28.28 -> 7.0.0-29.29: NVIDIA dGPU never enters runtime suspend (D3cold), idle power +8 W Status in linux package in Ubuntu: Incomplete Bug description: Ubuntu 26.04 LTS, desktop workstation (Dell Precision 3280 CFF) used headless as an LXD host. GPU: NVIDIA GB203GL [RTX PRO 4000 Blackwell SFF Edition] [10de:2c33] (rev a1) audio function [10de:22e9], upstream bridge 00:01.1 [8086:462d] Driver: nvidia-driver-595-open 595.84-0ubuntu0.26.04.1 (NVIDIA open kernel modules 595.84) NVreg_DynamicPowerManagement=0x02, nvidia-drm modeset=0 fbdev=0 With linux-image-7.0.0-28-generic (7.0.0-28.28) the dGPU cycles into runtime suspend normally. After booting linux-image-7.0.0-29-generic (7.0.0-29.29) the card stays in D0/active permanently with power/runtime_usage=1 and never suspends again. Package idle power measured at the wall rises from 10-14 W to about 22 W. The reference (runtime_usage=1) is held by the nvidia driver itself: it survives logging out of the GNOME session, unloading nvidia_drm and nvidia_modeset, and unloading all nvidia modules is worse still (no driver = no P8, +14 W). Cross-check isolating the kernel (same machine, same configuration): kernel nvidia result 7.0.0-28.28 595.71.05 suspends normally (140 and 172 wakeups over two boots) 7.0.0-29.29 595.84 never suspends (3 wakeups, all within the first 15 s of boot) 7.0.0-28.28 595.84 suspends normally <-- same driver as the failing case Last row is the decisive one: identical driver 595.84, identical module set (nvidia, nvidia_uvm, nvidia_modeset, nvidia_drm), identical configuration; only the kernel differs, and the GPU suspends again. The reverse pairing (7.0.0-29 with 595.71.05) is not testable, that driver version is no longer in the archive. Steps to reproduce: 1. Boot 7.0.0-29-generic with nvidia-driver-595-open and NVreg_DynamicPowerManagement=0x02, no CUDA workload, no display attached. 2. Leave the machine idle. 3. cat /sys/bus/pci/devices/0000:02:00.0/power_state -> D0 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_status -> active cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_usage -> 1 cat /sys/bus/pci/devices/0000:02:00.0/power/runtime_suspended_time -> frozen Expected (and observed on 7.0.0-28.28, ~150 s after boot): power_state D3cold, runtime_status suspended, runtime_usage 0, runtime_suspended_time 142242 ms vs runtime_active_time 8864 ms. Audio function 0000:02:00.1 also D3cold/suspended. Ruled out by measurement, not assumption: GUI session and its GPU clients (gnome-shell/Xwayland/nautilus/gnome-remote-desktop), KMS modules, CUDA/UVM context, VRAM threshold (memory.used 2 MiB against the 200 MB threshold), missing or changed config files, nvidia-persistenced, GPU containers (stopped), d3cold_allowed (1 on GPU, audio function and bridge), and running without the driver at all. The changelog between 7.0.0-28.28 and 7.0.0-29.29 shows no PCI-PM, D3cold, pcieport or ASPM change (CVE fixes plus one amdgpu HMM fix), so the regression is presumably a side effect rather than an intended change. Note on the attached apport data: it was collected while running the *working* kernel 7.0.0-28-generic, because the machine is remote and 7.0.0-29 costs the extra power. Happy to reboot into 7.0.0-29 and attach a second set on request. ProblemType: Bug DistroRelease: Ubuntu 26.04 Package: linux-image-7.0.0-28-generic 7.0.0-28.28 ProcVersionSignature: Ubuntu 7.0.0-28.28-generic 7.0.12 Uname: Linux 7.0.0-28-generic x86_64 NonfreeKernelModules: zfs ApportVersion: 2.34.1-0ubuntu0.1 Architecture: amd64 AudioDevicesInUse: USER PID ACCESS COMMAND /dev/snd/controlC1: gdm-greeter 3041 F.... wireplumber /dev/snd/controlC0: gdm-greeter 3041 F.... wireplumber /dev/snd/seq: gdm-greeter 3003 F.... pipewire CasperMD5CheckResult: pass CurrentDesktop: ubuntu:GNOME Date: Mon Aug 17 20:25:43 2026 InstallationDate: Installed on 2026-07-05 (43 days ago) InstallationMedia: Ubuntu 26.04 "Resolute Raccoon" - Release amd64 (20260423.1) Lsusb: Bus 001 Device 001: ID 1d6b:0002 Linux Foundation 2.0 root hub Bus 002 Device 001: ID 1d6b:0003 Linux Foundation 3.0 root hub Lsusb-t: /: Bus 001.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/16p, 480M /: Bus 002.Port 001: Dev 001, Class=root_hub, Driver=xhci_hcd/9p, 20000M/x2 MachineType: Dell Inc. Precision 3280 Compact ProcEnviron: LANG=en_US.UTF-8 PATH=(custom, no user) SHELL=/bin/bash TERM=xterm-256color ProcFB: ProcKernelCmdLine: BOOT_IMAGE=/vmlinuz-7.0.0-28-generic root=/dev/mapper/ubuntu--vg-ubuntu--lv ro quiet crashkernel=2G-4G:320M,4G-32G:512M,32G-64G:1024M,64G-128G:2048M,128G-:4096M RfKill: SourcePackage: linux UpgradeStatus: No upgrade log present (probably fresh install) dmi.bios.date: 05/22/2026 dmi.bios.release: 1.24 dmi.bios.vendor: Dell Inc. dmi.bios.version: 1.24.1 dmi.board.name: 0H1DC6 dmi.board.vendor: Dell Inc. dmi.board.version: A00 dmi.chassis.type: 3 dmi.chassis.vendor: Dell Inc. dmi.ec.firmware.release: 1.16 dmi.modalias: dmi:bvnDellInc.:bvr1.24.1:bd05/22/2026:br1.24:efr1.16:svnDellInc.:pnPrecision3280Compact:pvr:rvnDellInc.:rn0H1DC6:rvrA00:cvnDellInc.:ct3:cvr:sku0C81:pfaPrecision: dmi.product.family: Precision dmi.product.name: Precision 3280 Compact dmi.product.sku: 0C81 dmi.sys.vendor: Dell Inc. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163682/+subscriptions

Комментариев нет:

Отправить комментарий