Uploaded image for project: 'Lustre'
  1. Lustre
  2. LU-20584

tgt_import_update() reads uninitialized nfd_skip_ipv6, dropping all large NIDs

XMLWordPrintable

    • Icon: Bug Bug
    • Resolution: Fixed
    • Icon: Medium Medium
    • Lustre 2.18.0
    • None
    • None
    • 3
    • 9223372036854775807

      Summary

      tgt_import_update() in lustre/target/tgt_mount.c declares an on-stack
      struct nid_fetch_data nfd and initializes only nfd_lmd and nfd_pos.
      nfd_skip_ipv6 and nfd_has_ipv6 are left uninitialized.

      server_nid2radix(), the callback passed to LNetFetchNIDs(), drops every
      non-nid4 NID when nfd_skip_ipv6 is set:

      	/* skip IPv6 NIDs for an old server */
      	if (!nid_is_nid4(nid)) {
      		if (nfd->nfd_skip_ipv6)
      			return 0;
      		nfd->nfd_has_ipv6 = true;
      	}
      

      On a server whose NIDs are all IPv6, an uninitialized non-zero
      nfd_skip_ipv6 therefore drops every NID, nfd_pos stays 0, and the
      target logs:

      tgt_mount.c:1612:tgt_import_update()) MGC: can't get local NIDs from LNet, rc = -100
      

      The bug is invisible on IPv4, because the flag is only consulted for non-nid4
      NIDs.

      Impact

      tgt_import_update() is the upcall that re-reports a target's NID list to the
      MGS after the target's MGC reconnects. It is how the MGS nidtbl is restored
      after an MGS restart for targets that did not restart with it (LU-19692).

      When it fails, the MGS never relearns those NIDs. A client that can only reach
      that target over one of them gets no usable import connection:

      LustreError: import.c:553:import_select_connection()) lustre-OST0000-osc-...: no connections available: rc = -22
      LustreError: lov_obd.c:150:lov_connect_osc()) lustre-OST0000_UUID: target connect error: rc = -22
      

      Reproducer

      conf-sanity test_73e in an IPv6 environment. The test adds nets tcp10..tcp50 to
      both servers, moves the client to tcp50, then restarts only the MGS/MDS. The
      OSS is not restarted, so its NIDs must come back through
      tgt_import_update(). They do not, and the test fails at:

      cp: cannot fstat '/mnt/lustre/a': Permission denied
       conf-sanity test_73e: @@@@@@ FAIL: check client /mnt/lustre failed
      

      Verified on a 3-node VM cluster (2.17.57): test_73e fails on IPv6, passes on
      IPv4, and passes on IPv6 once both flags are initialized.

      Fix

      Initialize both fields in tgt_import_update(). This notification always
      carries large NIDs (tgt_nids_notify() sets LDD_F_LARGE_NID and the RPC
      requires OBD_CONNECT_MGS_NIDLIST), so nfd_skip_ipv6 should be false.

      Regressing commit

      The two fields were added by 2867efe066c8 ("LU-20054 mgs: more checks for
      incoming buffers"), which initialized them in server_lsi2mti() but not in
      tgt_import_update().

      Related: LU-19692 (added tgt_import_update()), LU-19700 (making MGS NID tables
      reliably in sync).

            hornc Chris Horn
            hornc Chris Horn
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

              Created:
              Updated:
              Resolved: