Skip to content

drbd 9.3.3: infinite reconnect loop after node reboot — handshake picks rule=sync-source-missed-finish, making the Outdated node sync source, which it refuses ("Resync skipped: sync-source (Outdated)", rv=-19); recovery only via connect --discard-my-data #143

Description

@abzholdings

:

Summary

After a clean reboot of a diskful node, the resync handshake with the surviving diskful UpToDate Primary peer deterministically selects uuid_compare()=source-use-bitmap by rule=sync-source-missed-finish, i.e. it makes the rebooted, Outdated node the sync source. The node then (correctly) refuses to be a sync source with an Outdated disk — Resync (rule=sync-source-missed-finish) skipped: sync-source (Outdated) — and the cluster-wide state change is aborted with rv = -19. The connection never leaves Connecting, and the handshake retries immediately with no backoff: 15,943 aborted cluster-wide state changes (127,619 kernel log lines) in under 20 minutes, ~13 retries/second.

The natural resolution — the Outdated node becoming SyncTarget of the UpToDate Primary — is exactly what a manual drbdadm disconnect + connect --discard-my-data produced (a 16 MiB bitmap-based resync that completed in 1 second). It seems the handshake should either fall back to that direction automatically when the chosen sync source is Outdated and the peer is UpToDate, or at minimum stop retrying the same doomed plan at full speed.

Versions

Component Version
DRBD kernel module 9.3.3 (/sys/module/drbd/version)
drbd-utils 9.34.3 (GIT 54d5a64e, 2026-04-17)
Kernel / OS 7.0.0-28-generic, Ubuntu 26.04 LTS
Managed by LINSTOR 1.33.3 / piraeus-operator v2.10.2 (Kubernetes v1.36.3)

Topology

Resource pvc-6e9129c9-7a19-4d43-ba3e-a0d28b2ab787, one 10 GiB volume (minor 1001), protocol C, quorum majority, 3 nodes:

  • dc01-wrk-01 (node-id 0) — diskful; the node that rebooted; comes back with its data Outdated (a resync it was source for "missed its finish" while the node was down)
  • dc01-wrk-04 (node-id 3) — diskful, UpToDate, Primary (volume open / in use the whole time)
  • dc01-wrk-02 (node-id 2) — diskless tie-breaker, Secondary

Kernel log (dc01-wrk-01, single boot; lines lightly trimmed for width)

Boot at 05:36:33; resource brought up by the LINSTOR satellite:

05:38:24 drbd pvc-6e9129c9-...: Starting worker thread (node-id 0)
05:38:24 drbd pvc-6e9129c9-.../0 drbd1001: disk( Attaching -> UpToDate ) [attach]
05:38:24 drbd pvc-6e9129c9-.../0 drbd1001: attached to current UUID: 45FB4CC9C2E63D32
05:38:26 drbd pvc-6e9129c9-... dc01-wrk-02: conn( Connecting -> Connected ) peer( Unknown -> Secondary ) [connected]
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001: disk( UpToDate -> Outdated ) [connected]
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-02: pdsk( DUnknown -> Diskless ) repl( Off -> Established ) [connected]

Handshake with the UpToDate Primary (dc01-wrk-04) — this block then repeats ~13x/second for 20 minutes:

05:38:26 drbd pvc-6e9129c9-...: Preparing cluster-wide state change 2479592503: 0->3 role( Secondary ) conn( Connected )
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: Missed end of resync as sync-source
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: drbd_sync_handshake:
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: self 45FB4CC9C2E63D32:A349C42E41991F5C:D4C6E1172B9D0BB0:94564B13F93C3BE0 bits:4096 flags:20
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: peer 45FB4CC9C2E63D33:0000000000000000:A4A33D080DA23274:F3E3144E74904D32 bits:0 flags:1120
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: uuid_compare()=source-use-bitmap by rule=sync-source-missed-finish
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: Resync (rule=sync-source-missed-finish) skipped: sync-source (Outdated)
05:38:26 drbd pvc-6e9129c9-...: Aborting cluster-wide state change 2479592503 (24ms) rv = -19

Loop metrics for this boot: first abort 05:38:26, last abort 05:58:11; journalctl -k -b 0 | grep <rsc> | grep -c "Aborting cluster-wide state change" = 15,943; total kernel lines for the resource: 127,619. (An intermediate invalidate attempt at 05:57:37, repl( StartingSyncT ) toward the diskless peer, had no effect.)

Manual recovery — disconnect, then connect with --discard-my-data on the Outdated node:

05:58:11 ... dc01-wrk-04: conn( Connecting -> Disconnecting ) [disconnect]
05:58:11 ... dc01-wrk-04: conn( StandAlone -> Unconnected ) [connect]
05:58:13 ... drbd1001 dc01-wrk-04: Resync direction reversed by --discard-my-data. Reverting to older data!
05:58:13 ... dc01-wrk-04: conn( Connecting -> Connected ) peer( Unknown -> Primary ) [connected]
05:58:13 ... drbd1001 dc01-wrk-04: pdsk( DUnknown -> UpToDate ) repl( Off -> WFBitMapT ) [connected]
05:58:13 ... drbd1001: disk( Outdated -> Inconsistent ) [receive-bitmap]
05:58:13 ... drbd1001 dc01-wrk-04: repl( WFBitMapT -> SyncTarget ) [receive-bitmap]
05:58:13 ... drbd1001 dc01-wrk-04: Began resync as SyncTarget (will sync 16384 KB [4096 bits set]).
05:58:14 ... drbd1001 dc01-wrk-04: Resync done (total 1 sec; paused 0 sec; 16384 K/sec)
05:58:14 ... drbd1001: disk( Inconsistent -> UpToDate ) [resync-finished]

UUID situation at handshake time: self current 45FB4CC9C2E63D32 with bits:4096 toward the peer; peer current 45FB4CC9C2E63D33 (differs only in the low bit) with a zeroed bitmap UUID slot for this node and bits:0 — i.e. the peer finished the earlier resync (as target) and rotated, while this node was away; hence "Missed end of resync as sync-source".

Expected vs actual

Expected: the connect resolves automatically — with the peer UpToDate and Primary, and the local disk Outdated, the local node should end up SyncTarget (which the manual --discard-my-data proved is a trivial 16 MiB bitmap resync). Or, failing that, the state change should not be retried at full speed forever: the same refusal is recomputed ~13 times per second with no backoff, flooding the log (127k lines) and keeping the connection permanently in Connecting.

Actual: livelock — rule=sync-source-missed-finish insists the Outdated node is the source, the source role is refused because the disk is Outdated (rv = -19), and the handshake repeats indefinitely. Only operator intervention (disconnect + connect --discard-my-data) recovers.

Question

When sync-source-missed-finish selects a sync source whose disk is Outdated and the peer is UpToDate, is falling back to the sync-target direction (the effect of --discard-my-data) safe/intended? If the current behaviour is deliberate, could the retry at least back off instead of looping at ~13 attempts/second?

================================================================================
ORCHESTRATOR NOTES

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions