вторник

[Bug 2167995] [NEW] HMC error triggers cascading PF resets and tears down RDMA connections

Public bug reported: [ Impact ] A single HMC error triggers cascading PF resets and tears down RDMA connections on the device: kernel: ice 0000:17:00.1 irdma1: HMC Error kernel: ice 0000:17:00.1 irdma1: Requesting a reset kernel: DMAR: DRHD: handling fault status reg 2 kernel: DMAR: [DMA Read NO_PASID] Request device [17:00.1] fault addr 0xfddce000 [fault reason 0x06] PTE Read access is not set kernel: bond1: (slave enp23s0f1np1): link status definitely down, disabling slave kernel: bond1: active interface up! kernel: ice 0000:17:00.1: PTP reset successful ice 0000:17:00.1: VSI rebuilt. VSI index 0, type ICE_VSI_PF ice 0000:17:00.1: VSI rebuilt. VSI index 1, type ICE_VSI_CTRL bond1: (slave enp23s0f1np1): link status definitely up, 25000 Mbps full duplex This causes a short network outage on intel ice NIC, but not i40e which already has a fix. upstream commit: https://github.com/torvalds/linux/commit/4dc9c884c0aafac1e8d37536d88d485135e26883 applys the same fix to ice. [ Test Plan ] Make sure machine with intel ice works fine by running basic network traffic test and then reproduce the issue in a lab environment. Confirm the issue can not be reproduced with the above commit. [ Where problems could occur ] The same fix has been applied to intel i40e, the commit does the same thing. If something goes wrong in the path, and PF is not reset, it may hit the same network outage when using intel ice ** Affects: linux (Ubuntu) Importance: Undecided Status: Fix Released ** Affects: linux (Ubuntu Jammy) Importance: High Assignee: gerald.yang (gerald-yang-tw) Status: In Progress ** Affects: linux (Ubuntu Noble) Importance: High Assignee: gerald.yang (gerald-yang-tw) Status: In Progress ** Affects: linux (Ubuntu Resolute) Importance: High Assignee: gerald.yang (gerald-yang-tw) Status: In Progress ** Also affects: linux (Ubuntu Jammy) Importance: Undecided Status: New ** Also affects: linux (Ubuntu Resolute) Importance: Undecided Status: New ** Also affects: linux (Ubuntu Noble) Importance: Undecided Status: New ** Changed in: linux (Ubuntu Jammy) Status: New => In Progress ** Changed in: linux (Ubuntu Noble) Status: New => In Progress ** Changed in: linux (Ubuntu Resolute) Status: New => In Progress ** Changed in: linux (Ubuntu) Status: New => Fix Released ** Changed in: linux (Ubuntu Jammy) Assignee: (unassigned) => gerald.yang (gerald-yang-tw) ** Changed in: linux (Ubuntu Noble) Assignee: (unassigned) => gerald.yang (gerald-yang-tw) ** Changed in: linux (Ubuntu Resolute) Assignee: (unassigned) => gerald.yang (gerald-yang-tw) ** Changed in: linux (Ubuntu Jammy) Importance: Undecided => High ** Changed in: linux (Ubuntu Noble) Importance: Undecided => High ** Changed in: linux (Ubuntu Resolute) Importance: Undecided => High -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2167995 Title: HMC error triggers cascading PF resets and tears down RDMA connections Status in linux package in Ubuntu: Fix Released Status in linux source package in Jammy: In Progress Status in linux source package in Noble: In Progress Status in linux source package in Resolute: In Progress Bug description: [ Impact ] A single HMC error triggers cascading PF resets and tears down RDMA connections on the device: kernel: ice 0000:17:00.1 irdma1: HMC Error kernel: ice 0000:17:00.1 irdma1: Requesting a reset kernel: DMAR: DRHD: handling fault status reg 2 kernel: DMAR: [DMA Read NO_PASID] Request device [17:00.1] fault addr 0xfddce000 [fault reason 0x06] PTE Read access is not set kernel: bond1: (slave enp23s0f1np1): link status definitely down, disabling slave kernel: bond1: active interface up! kernel: ice 0000:17:00.1: PTP reset successful ice 0000:17:00.1: VSI rebuilt. VSI index 0, type ICE_VSI_PF ice 0000:17:00.1: VSI rebuilt. VSI index 1, type ICE_VSI_CTRL bond1: (slave enp23s0f1np1): link status definitely up, 25000 Mbps full duplex This causes a short network outage on intel ice NIC, but not i40e which already has a fix. upstream commit: https://github.com/torvalds/linux/commit/4dc9c884c0aafac1e8d37536d88d485135e26883 applys the same fix to ice. [ Test Plan ] Make sure machine with intel ice works fine by running basic network traffic test and then reproduce the issue in a lab environment. Confirm the issue can not be reproduced with the above commit. [ Where problems could occur ] The same fix has been applied to intel i40e, the commit does the same thing. If something goes wrong in the path, and PF is not reset, it may hit the same network outage when using intel ice To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2167995/+subscriptions

Комментариев нет:

Отправить комментарий