Uploaded image for project: 'Lustre'
  1. Lustre
  2. LU-20715

replay-single/sanity-pfl: (client.c:1861:ptlrpc_send_new_req()) ASSERTION( list_empty(&req->rq_list) ) failed

XMLWordPrintable

    • Icon: Bug Bug
    • Resolution: Unresolved
    • Icon: Medium Medium
    • Lustre 2.18.0
    • Lustre 2.18.0
    • None
    • 3
    • 9223372036854775807

      Description

      A client LBUGs in ptlrpc_send_new_req() during target failover when the
      import reconnects while a request replay is still in flight. The recovery
      state machine re-enters LUSTRE_IMP_REPLAY and replays the same request a
      second time, putting it back into RQ_PHASE_NEW while it is still linked on
      imp_sending_list / imp_delayed_list.

      LustreError: 1017627:0:(client.c:1861:ptlrpc_send_new_req()) ASSERTION( list_empty(&req->rq_list) ) failed:
      LustreError: 1017627:0:(client.c:1861:ptlrpc_send_new_req()) LBUG
      CPU: 9 PID: 1017627 Comm: ptlrpcd_rcv Kdump: loaded Tainted: G           O      -------- -  - 4.18.0rocky8.10-debug #2
      Call Trace:
       dump_stack+0x99/0xca
       lbug_with_loc.cold.4+0xd/0x86 [libcfs]
       ptlrpc_send_new_req+0xe86/0x1230 [ptlrpc]
       ptlrpc_check_set+0xa92/0x2ff0 [ptlrpc]
       ptlrpcd+0xaeb/0xe90 [ptlrpc]
       kthread+0x1de/0x210
       ret_from_fork+0x1f/0x30
      Kernel panic - not syncing: LBUG
      

      The console prelude is the same in every recorded hit – a reply with
      -ENOTCONN, the connection dropping, and the client reconnecting into the
      target's recovery, all within tens of milliseconds:

      LustreError: lustre-MDT0000-mdc-ffff8f262355c000: operation mds_statfs to node 0@lo failed: rc = -107
      Lustre: lustre-MDT0000-mdc-ffff8f262355c000: Connection to lustre-MDT0000 (at 0@lo) was lost
      Lustre: lustre-MDT0000: Will be in recovery for at least 1:00, or until 1 client reconnects
      Lustre: lustre-MDT0000: Client 9c0615e3-... (at 0@lo) reconnected, waiting for 1 clients in recovery for 1:10
      LustreError: 1017627:0:(client.c:1861:ptlrpc_send_new_req()) ASSERTION( list_empty(&req->rq_list) ) failed:
      

      Occurrences

      Long standing and test-agnostic – 7 reports in the crash DB spanning
      2023-01-19 .. 2026-09-08, always on the client, always ptlrpcd_rcv:

      date test artifact
      2026-09-08 sanity-pfl test_22d boilpot-bigmem2-43
      2026-08-29 replay-dual test_0a boilpot-bigmem2-81
      2026-08-26 replay-single test_3c boilpot-bigmem2-61
      2026-05-22 replay-single test_112f/112k/70f (review-dne, review-dne-zfs) build 125454 / 125445, review 66081
      2023-01-19 recovery-mds-scale test_failover_ost lustre-master build 4380

      Crash DB signature: newid=67941

      Root cause

      ptlrpc_import_recovery_state_machine() runs ptlrpc_replay_next() in the
      LUSTRE_IMP_REPLAY arm with no imp_replay_inflight check, unlike the
      LUSTRE_IMP_REPLAY_LOCKS and LUSTRE_IMP_REPLAY_WAIT arms immediately
      below it. When a connect completes while a replay is outstanding, the import
      re-enters LUSTRE_IMP_REPLAY; imp_last_replay_transno has not advanced
      (no reply yet), so ptlrpc_replay_next() selects the same request and
      ptlrpc_replay_req() resets it to RQ_PHASE_NEW and re-queues it, while it
      is still linked on an import list.

      The whole sequence is in the client debug log of the 2026-08-26 dump
      (vmcore-lustredebug.txt), 9 ms end to end:

      .171635 import ffff995652a5b800: changing import state from CONNECTING to REPLAY
      .171652 (recover.c:53:ptlrpc_replay_next())  import ... committed 47244640259 last 0
      .171654 (client.c:3553:ptlrpc_replay_req())  @@@ REPLAY req@ffff9956565d4980 x1874569864843008/t51539607553 o36  fl New:RQU/204/0
      .171673 (client.c:1919:ptlrpc_send_new_req()) Sending RPC req@ffff9956565d4980 ... o36  <- now on imp_sending_list
      .180586 import ffff995652a5b800: changing import state from CONNECTING to REPLAY
      .180593 (client.c:3169:ptlrpc_free_committed()) @@@ stopping search req@ffff9956565d4980 ... fl Rpc:Qr/204/ffffffff  <- still in flight
      .180598 (recover.c:53:ptlrpc_replay_next())  import ... last 0
      .180601 (client.c:3553:ptlrpc_replay_req())  @@@ REPLAY req@ffff9956565d4980 ... fl New:Qr/206/ffffffff  <- same req re-armed
      .180608 (client.c:1860:ptlrpc_send_new_req()) ASSERTION( list_empty(&req->rq_list) ) failed
      

      The server had not yet dequeued that RPC (no ptlrpc_server_handle_request
      for x1874569864843008 anywhere in the log), so the request was alive and
      would have been replied to – the second replay was both unnecessary and fatal.

      When introduced

      git blame -w -M -C on the changed line points at *0343ecb7de2d ("Landing
      b_recovery")*, and git show confirms that commit added the
      LUSTRE_IMP_REPLAY arm without the imp_replay_inflight guard that its two
      siblings received. All branches are affected.

            wc-triage WC Triage
            green Oleg Drokin
            Votes:
            0 Vote for this issue
            Watchers:
            2 Start watching this issue

              Created:
              Updated: