Hi Aaron, Attached is the latest test result, please help to investigate. Thanks a lot! Anyway, something needing to mention first: Due to our test machine resource allocation reason, we conduct this batch of tests on a light-configuration machine (same server, same firmware, but only NVME disks installed, much less other adapters being installed),and we found a little bit different symptoms from last batch: 1. The original (missing disks) issue could be reproduced (very possibly) only when re-plug the disk to a different slot location from where it is plugged. 2. There is no failure-to-shut-down, nor failure-to-reboot symptom found, after these hot-plugs iterations. ------------------------------------------------------------------------------- Below is the procedure that we perform this batch of tests and logs collection: 1.Use ipmitool to capture the serial console log from before the tests being started. 2.Add parameters to the GRUB configuration file vim /etc/default/grub log_buf_len=64M ignore_loglevel nvme.dyndbg=+p pci.dyndbg=+p vmd.dyndbg=+p 3.Enable Persistent Logging a.Create the journal logs persistent directory: sudo mkdir -p /var/log/journal b.Set the correct permissions: sudo chown root:systemd-journal /var/log/journal sudo chmod 2755 /var/log/journal c.Restart the journal service: sudo systemctl restart systemd-journald 4.reboot,Install the provided deb package: linux-image-7.1.0-rc4+_7.1.0~rc4-00098-gf4790477726f-43_Bmd64.deb 5. reboot uname -r to check the kernel version 6.Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 1-config_B.before_unplug.log 7.Test hot-plugging of VROC NVMe drives and collect logs. Note, please leave 120 seconds between each unplug / re-plug actions: a. unplug (slot1 2) b. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 2-config_B.after_unplug.log c. re-plug(slot1 2) d. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 3-config_B.after_replug.log e. unplug(slot1 2) f. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 4-config_B.after_unplug.log g. re-plug Other slot (slot3 4) h. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 5-config_B.after_replug.log i. unplug(slot3 4) j. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 6-config_B.after_unplug.log k. re-plug (slot1 2) l. Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 7-config_B.after_replug.log 8. Reboot the system 9.Collect dmesg, lspci -vvv, lsblk, nvme list,journalctl -f,journalctl -f -u systemd-udevd, zip and name this set of log file as 8-config_B.after_reboot.log save serial log to config_B_2026_7_23.zip -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2148638 Title: ubuntu 26.04 - When the NVME hard drives were hot-swapped, the OS reported an error and some of the drives were not recognized. Status in linux package in Ubuntu: New Status in linux source package in Resolute: New Bug description: When the CD8P NVMe hard drives were hot-swapped, the OS reported an error and some of the drives were not recognized. This issue happens on CD8P NVMe disk and does not happens on bm1743 NVMe disk. CD8P drives are Kaoxia drives The BM1743 drives that work are Samsung devices. Nvme disk detail: CD8P https://lenovopress.lenovo.com/lp1904-thinksystem-cd8p-read-intensive-nvme-pcie-50-ssd bm1743 https://lenovopress.lenovo.com/lp2156-thinksystem-bm1743-read-intensive-nvme-pcie-50-x4-ssd Steps to reproduce: 1.In the UEFI, enable VMD without created RAID disk. 2.Install ubuntu26.04 on M.2 sata disk. 3.use command lsblk to check all NVMEe disk 4.Unplug all NVMe device, then check NVMe device information again via lsblk. 5.plug all NVMe SSD 6.OS reported an error and some of the drives were not recognized. Compare with Ubuntu 24.04: There is no errors messages and all NVMe disks can be recognized after re-plug all NVMe SSD Info: The issue also happens on latest daily build: 0415 This issue only happens on VMD enabled, there is no this issue when vmd feature disable. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2148638/+subscriptions
Комментариев нет:
Отправить комментарий