Uploaded image for project: 'Lustre'
  1. Lustre
  2. LU-20749

osd_ldiskfs backend fails to load at mount time

XMLWordPrintable

    • Icon: Bug Bug
    • Resolution: Unresolved
    • Icon: Medium Medium
    • None
    • Lustre 2.18.0
    • 3
    • 9223372036854775807

      When a server target is mounted, the OSD type is looked up by class_get_type() in lustre/obdclass/genops.c, which asks the kernel to load the matching module with request_module() if that type is not registered yet. And nothing else loads that module: mount.lustre does not, and an HA resource agent simply runs mount, so request_module() is the normal and only path by which osd_ldiskfs get loaded.

      That path reports almost nothing when it fails. request_module() runs modprobe through call_usermodehelper(), which discards modprobe's error output, and class_get_type() discards the return code as well. All that is left of a failure is:

      LustreError: Can't load module 'osd-ldiskfs'
      LustreError: 423309:0:(genops.c:379:class_newdev()) OBD: unknown type: osd-ldiskfs
      LustreError: 423309:0:(obd_config.c:646:class_attach()) Cannot create device lustre-MDT0000-osd of type osd-ldiskfs : -19
      LustreError: 423309:0:(tgt_mount.c:2554:server_fill_super()) Unable to start osd on /dev/mapper/mds1_flakey: -19
      LustreError: 423309:0:(super25.c:206:lustre_fill_super()) llite: Unable to mount <unknown>: rc = -19
      

      plus ENODEV from the mount itself, which mount.lustre turns into "Are the lustre modules loaded? Check /etc/modprobe.conf and /proc/filesystems". Every possible cause produces exactly the same output: the module not being installed, a kABI or relocation mismatch, a rejected signature, an error while inserting one of its dependencies, modprobe not being runnable at that moment, or a transient shortage of memory.

      Proposed work:

      • Report the return code of request_module() in class_get_type(), instead of discarding it. This is the only place where the error code of the module load still exists.
      • Retry the load once before giving up, since the failure can be transient: modprobe may fail to be run, or the previous instance of the module may still be going away. A load that only succeeds on the second attempt should leave a warning behind, so that it can still be found in the logs afterwards.
      • Load the backend OSD module from mount.lustre, in parse_ldd(), where osd_is_lustre() has already determined the mount type from the on-disk data and where modprobe can report on its own why it failed, to the caller of mount(8), to the HA resource agent log and to the journal. This mirrors what is already done for ZFS in zfs_init(), which loads the zfs module from userspace but not osd_zfs.
      • Name that module in the ENODEV hint of mount.lustre, so that it no longer points only at the Lustre modules, which are loaded in that case.

            sebastien Sebastien Buisson
            sebastien Sebastien Buisson
            Votes:
            0 Vote for this issue
            Watchers:
            3 Start watching this issue

              Created:
              Updated: