-
Bug
-
Resolution: Unresolved
-
Medium
-
None
-
None
-
3
-
9223372036854775807
target_bulk_io() in lustre/ldlm/ldlm_lib.c :
time64_t start = ktime_get_seconds(); /* MONOTONIC */
...
deadline = start + bulk_timeout;
if (deadline > req->rq_deadline) /* rq_deadline is REALTIME */
deadline = req->rq_deadline;
do
while (rc == -ETIMEDOUT && deadline > ktime_get_seconds());
if (rc == -ETIMEDOUT)
DEBUG_REQ(..., deadline - start, ktime_get_real_seconds() - deadline);
start / deadline are built from ktime_get_seconds() (seconds since boot), but clamped against and logged against req->rq_deadline / ktime_get_real_seconds() (seconds since epoch — differs by billions of seconds). Effects:
• The if (deadline > req->rq_deadline) deadline = req->rq_deadline clamp essentially never fires in practice (boot-relative deadline is always numerically far smaller than the epoch-relative rq_deadline ), so the intended "don't let bulk exceed the request's own deadline" bound is silently defeated — bulk transfers can run the full bulk_timeout even if the request's rq_deadline (client-side expectation) has already passed.
• The final DEBUG_REQ error message computes ktime_get_real_seconds() - deadline , which — since deadline is boot-relative — prints a nonsensical, huge elapsed-time value in logs.