Uploaded image for project: 'Lustre'
  1. Lustre
  2. LU-20524

sanity-sec test_72a/72b: kernel panic in range_delete_generic() deleting a NID range that has sub-ranges

XMLWordPrintable

    • Icon: Bug Bug
    • Resolution: Fixed
    • Icon: Critical Critical
    • Lustre 2.18.0
    • Lustre 2.17.0, Lustre 2.18.0
    • None
    • 2
    • 9223372036854775807

      AI generated report

      Symptom

      sanity-sec test_72a ("dynamic nodemap properties on OSS") and
      test_72b ("dynamic nodemap properties on MDS") sporadically panic the
      server node. The subtest never completes, so the suite is truncated at
      test_72a/test_72b in results.yml.

      BUG: unable to handle kernel paging request at ffff9ba0876d7460
      [exception RIP: range_delete_generic+128]
         include/linux/rbtree_augmented.h: 326
       range_delete            lustre/ptlrpc/nodemap_range.c:489  [ptlrpc]
       nodemap_del             lustre/ptlrpc/nodemap_handler.c:4542 [ptlrpc]
       nodemap_del             lustre/ptlrpc/nodemap_handler.c:4516 [ptlrpc]
       cfg_nodemap_cmd         lustre/ptlrpc/nodemap_handler.c:5426 [ptlrpc]
       server_iocontrol_nodemap  [ptlrpc]
       mgs_iocontrol_nodemap   [mgs]
       mgs_iocontrol           [mgs]
       class_handle_ioctl      [obdclass]
      

      Seen on plain master, unrelated to any patch under test:

      • Janitor 65796, sanity-sec-zfs-rocky8.10 (2026-07-04) – crash in test_72a
      • Janitor 66688, sanity-sec-zfs-rocky8.10 retry2 (2026-07-25) – crash in {{test_72b}}

      Both are the same fault: the faulting address is exactly range->rn_tree + 8, i.e. the read of{{{} rn_tree->nmrt_range_interval_root.rb_leftmostof a freed
      struct lu_nid_range{}}}.

      Root cause

      The NID range of a dynamic sub-nodemap is stored in the interval tree embedded
      in the range of its parent nodemap (lu_nid_range.rn_subtree), and points
      back to that tree through rn_tree.

      range_delete_generic() (lustre/ptlrpc/nodemap_range.c) frees a range
      without looking at its subtree. So deleting the parent nodemap's range – what
      lctl nodemap_del_range on the parent does, and what {{dyn_nm_helper()}} does just before lctl nodemap_del – leaves every sub-range linked into
      freed memory with a dangling rn_tree. Two consequences:

      1. The sub-ranges are no longer reachable from range_search(), so clients in
          those ranges silently fall back to the default nodemap.
      2. The next deletion of one of those sub-ranges – e.g. the recursive
          nodemap_del() of the sub-nodemaps that immediately follows in the test –
          writes into the freed parent range.

      Confirmed in the janitor vmcore (66688). The crashing range is
      {{[1.1.1.51-1.1.1.100] of nodemap {{nm_test72_2; its sibling
      [1.1.1.2-1.1.1.50]}}}} of nm_test72_1 is still live and carries the same
      {{rn_tree value, and {{crash's {{kmem shows the page that pointer refers
      to has been returned to the page allocator (CNT 0).}}}}}}

      {{{{{{The panic is sporadic because it only happens when the whole slab page has been
      freed and unmapped – the janitor kernel is built with
      CONFIG_DEBUG_PAGEALLOC}}}}}} On a production kernel the same code silently
      corrupts a freed object instead, which is arguably worse.

      Introduced by

      • e9278b8da1 "LU-17431 nodemap: make dynamic nodemaps hierarchical"
          (in master 2025-04-10) added rn_subtree and never taught range deletion
          about it – this alone already orphans sub-ranges.
      • 5bde32348f  "LU-17431 nodemap: find innermost nid range"
          (in master 2025-08-12) added the rn_tree back-pointer that range deletion
          now dereferences, turning the orphaning into a use-after-free.

      Reproducer

      No crash needed – the functional half is deterministic. On any server node:

      lctl nodemap_activate 1
      lctl nodemap_add -d -p default nmrepro
      lctl nodemap_add_range --name nmrepro --range 1.1.1.[1-100]@tcp
      lctl nodemap_add -d -p nmrepro nmrepro_1
      lctl nodemap_add_range --name nmrepro_1 --range 1.1.1.[2-50]@tcplctl nodemap_test_nid 1.1.1.10@tcp        # -> nmrepro_1  (correct)
      lctl nodemap_del_range --name nmrepro --range 1.1.1.[1-100]@tcp
      lctl nodemap_test_nid 1.1.1.10@tcp        # -> default    (BUG)lctl nodemap_del nmrepro                  # writes into the freed range

      To see the use-after-free itself, boot the server with slub_debug=FZPU, run
      the above, then echo 1 > /sys/kernel/slab/kmalloc-192/validate:

       BUG kmalloc-192: Poison overwritten
        First byte 0x0 instead of 0x6b
        Allocated in range_create_generic+0x131/0x680 [ptlrpc]
        Freed in nodemap_del_range+0x2d0/0x5c0 [ptlrpc]
        Object ...: 6b 6b 6b 6b 6b 6b 6b 6b 00 00 00 00 00 00 00 00
      

      The overwritten bytes are at object offset 0x98, which is
      rn_subtree.nmrt_range_interval_root of the freed parent range.

      Proposed fix

      In range_delete_generic(), move the sub-ranges up into the tree the deleted
      range itself belongs to before freeing it. They cannot conflict there: they were
      included in the deleted range, which did not overlap any other range of that
      tree. This covers every deletion path (nodemap_del_range(), nodemap_del(), nodemap_config_dealloc()) regardless of orderig.

            green Oleg Drokin
            green Oleg Drokin
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

              Created:
              Updated:
              Resolved: