I generated the attached kernel module + userspace utility reproducer to verify that the issue is indeed resolved in noble/linux 6.8.0-147.147 and resolute/linux 7.0.0-39.39. With the reproducer it can be seen that the fast path is taken in the NULL-mapping order-0 folio case now that this patch is applied. The fast path is usable 0/512 attempts on the unpatched kernels, but can be used on 512/512 attempts on the patched kernels. Running the reproducer on noble/linux and resolute/linux yields the following: noble/linux 6.8.0-146-generic $ ./gupbench -p 512 -r 5 -i 200 1 8 32 128 page size: 4K, pages/thread: 512, reps: 5, iters: 200 get_user_pages_fast_only() verdict: gupfast (alloc_page+vm_insert_page): 0/512 anon (control) : 512/512 pin_user_pages_fast(FOLL_WRITE) median ns/page: threads gupfast anon 1 187.1 110.2 8 171.2 73.1 32 310.4 108.7 128 1212.5 195.2 RESULT: FAIL - GUP-fast rejects NULL-mapping order-0 folios (bug present) noble/linux 6.8.0-147-generic $ ./gupbench -p 512 -r 5 -i 200 1 8 32 128 page size: 4K, pages/thread: 512, reps: 5, iters: 200 get_user_pages_fast_only() verdict: gupfast (alloc_page+vm_insert_page): 512/512 anon (control) : 512/512 pin_user_pages_fast(FOLL_WRITE) median ns/page: threads gupfast anon 1 110.2 113.3 8 89.9 74.1 32 114.7 106.6 128 177.9 209.3 RESULT: PASS - GUP-fast accepts NULL-mapping order-0 folios (fixed) resolute/linux 7.0.0-38-generic $ ./gupbench -p 512 -r 5 -i 200 1 8 32 128 page size: 4K, pages/thread: 512, reps: 5, iters: 200 get_user_pages_fast_only() verdict: gupfast (alloc_page+vm_insert_page): 0/512 anon (control) : 512/512 pin_user_pages_fast(FOLL_WRITE) median ns/page: threads gupfast anon 1 177.3 115.7 8 153.7 117.0 32 257.4 119.5 128 752.7 213.2 RESULT: FAIL - GUP-fast rejects NULL-mapping order-0 folios (bug present) resolute/linux 7.0.0-39-generic $ ./gupbench -p 512 -r 5 -i 200 1 8 32 128 page size: 4K, pages/thread: 512, reps: 5, iters: 200 get_user_pages_fast_only() verdict: gupfast (alloc_page+vm_insert_page): 512/512 anon (control) : 512/512 pin_user_pages_fast(FOLL_WRITE) median ns/page: threads gupfast anon 1 121.9 117.1 8 93.8 84.2 32 124.2 115.6 128 155.0 212.5 RESULT: PASS - GUP-fast accepts NULL-mapping order-0 folios (fixed) ** Attachment added: "lp2162917-gup-fast-reproducer.tar.xz" https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2162917/+attachment/6007403/+files/lp2162917-gup-fast-reproducer.tar.xz ** Tags removed: verification-needed-noble-linux verification-needed-resolute-linux verification-needed-resolute-linux-nvidia verification-needed-resolute-linux-nvidia-bos ** Tags added: verification-done-noble-linux verification-done-resolute-linux verification-done-resolute-linux-nvidia verification-done-resolute-linux-nvidia-bos -- You received this bug notification because you are subscribed to linux in Ubuntu. Matching subscriptions: Bgg, Bmail, Nb https://bugs.launchpad.net/bugs/2162917 Title: Backport: "mm/gup: fix GUP-fast fallback for NULL-mapping order-0 folios" Status in linux package in Ubuntu: Fix Released Status in linux-nvidia package in Ubuntu: Invalid Status in linux-nvidia-bos package in Ubuntu: Invalid Status in linux source package in Noble: Fix Committed Status in linux-nvidia source package in Noble: Fix Committed Status in linux-nvidia-bos source package in Noble: Invalid Status in linux source package in Resolute: Fix Committed Status in linux-nvidia source package in Resolute: Fix Released Status in linux-nvidia-7.0 source package in Resolute: Invalid Status in linux-nvidia-bos source package in Resolute: Fix Released Status in linux source package in Stonking: Fix Released Status in linux-nvidia source package in Stonking: Invalid Status in linux-nvidia-bos source package in Stonking: Invalid Bug description: SRU Justification [Impact] There is a performance regression seen for calls to get_user_pages() on single-page NULL-mapping folios, introduced by commit f002882ca369 ("mm: merge folio_is_secretmem() and folio_fast_pin_allowed() into gup_fast_folio_allowed()"). When either CONFIG_SECRETMEM=y, or when the folio is a long-term writable pin, the current implementation requires that the mapping be checked as either a secretmem mapping or a file-backed one. If it is either, or if the mapping is NULL, the slow path is forced. However, when the mapping is NULL, the folio cannot be a secretmem mapping. Thus, as long as the folio is NOT a long-term writable pin, the fast path can still be used. This performance regression was initially observed during GPU Direct Storage (GDS) workloads. [Fix] The upstream commit c494788faffe ("mm/gup: fix GUP-fast fallback for NULL-mapping order-0 folios") resolves the performance regression by changing gup_fast_folio_allowed() to allow the fast path for the case described above, where the folio is NOT a long-term writable pin and its mapping field is NULL. [Test Plan] Build and boot tested. The performance regression and subsequent fix can be verified with a GDS workload. The original bug report also describes a test kernel module that uses the `alloc_page` + `vm_insert_page` + `pin_user_pages_fast(..., FOLL_WRITE, ...)` functions to emulate GDS behavior, and using `get_user_pages_fast_only()` on that to validate that the fast path is now allowed in this scenario. [Where problems could occur] The fix affects the get_user_pages*() path of the mm subsystem, which is used heavily. The fix has been reviewed and applied in upstream Linux. ------------------------------------------------------- Original report: Clean cherry-pick from linux-next: ``` (cherry picked from commit ae75e88d8c258fd849de594e7d468b5263e7b3e3 linux-next) ``` John Hubbard, Acked-by David Hildenbrand, applied by Andrew Morton. [Lore thread](https://lore.kernel.org/all/20260708005745.164928-1-jhubbard@nvidia.com/). ## Problem `f002882ca369` (present on this branch) made `gup_fast_folio_allowed()` bail to the slow path for *any* order-0 folio with a NULL `->mapping` when `CONFIG_SECRETMEM=y`. Pages from `alloc_page()` + `vm_insert_page()` legitimately have a NULL mapping, so every `pin_user_pages_fast()` over such a range misses the fast path — nvidia-fs (GPUDirect Storage) allocates its shadow buffers exactly this way. The NULL check was meant to catch truncated file-backed pages, not secretmem. Secretmem folios are published via `filemap_add_folio()`, which always sets `->mapping`, so a NULL mapping proves the folio is *not* secretmem. Returning `!reject_file_backed` keeps long-term writable pins on the slow path and restores the fast path otherwise — exactly the pre-`f002882ca369` behaviour. `CONFIG_SECRETMEM=y` on amd64 and arm64, so this is live on every flavour we ship. ## Measurement (GH200, 288 cores) A module reproducing the nvidia-fs pattern (`alloc_page` + `vm_insert_page`, then `pin_user_pages_fast(..., FOLL_WRITE, ...)`), with an anonymous-memory control the patch cannot affect. `get_user_pages_fast_only()` gives the GUP-fast verdict directly: **0/N unpatched, N/N patched**. Median ns/page, 512 pages/thread, 5 reps: | threads | 4K unpatched → patched | 64K unpatched → patched | |--------:|-----------------------:|------------------------:| | 1 | 34 → 34 (1.0x) | 33 → 30 (1.1x) | | 8 | 74 → 33 (**2.2x**) | 1920 → 31 (**62x**) | | 32 | 95 → 33 (**2.9x**) | 4792 → 33 (**145x**) | | 128 | 500 → 218 (noisy) | 16673 → 76 (**219x**) | The anon control held at 30-35 ns/page across all four kernels, so only the affected range moved. Stock `7.0.0-1015-nvidia-64k` independently reproduces the unpatched 64K numbers (8 threads: 1176 vs 37), so this isn't a test-config artefact. Single-threaded it's a wash; the win is under concurrency. `perf` on the unpatched 64K kernel shows the slow path is ~90% lock contention (`queued_spin_lock_slowpath` 72%), gone entirely once patched — contention that scales with thread count, not a fixed per-page cost. ## Risk Low. One line in a static function with three callers, all in GUP-fast. Only `reject_file_backed == false && check_secretmem && mapping == NULL` changes behaviour; long-term writable pins are untouched. A secretmem folio caught mid-truncate can't happen here: `secretmem_setattr()` refuses to shrink, there's no `.fallocate` (so no punch-hole), and `secretmem_migrate_folio()` returns `-EBUSY`. Only inode eviction remains, which requires every VMA gone — no VMA, no PTE for GUP-fast to walk. Raised by David Hildenbrand on v1 and resolved before he Acked. ## Testing Applies cleanly; built arm64 4K/64K and x86_64, no new warnings; verified in the binary (unpatched `mov w0, #0x0` vs patched `eor w0, w0, #0x1`); benchmarked as above. To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2162917/+subscriptions
Комментариев нет:
Отправить комментарий