:
Summary
After a clean reboot of a diskful node, the resync handshake with the surviving diskful UpToDate Primary peer deterministically selects uuid_compare()=source-use-bitmap by rule=sync-source-missed-finish, i.e. it makes the rebooted, Outdated node the sync source. The node then (correctly) refuses to be a sync source with an Outdated disk — Resync (rule=sync-source-missed-finish) skipped: sync-source (Outdated) — and the cluster-wide state change is aborted with rv = -19. The connection never leaves Connecting, and the handshake retries immediately with no backoff: 15,943 aborted cluster-wide state changes (127,619 kernel log lines) in under 20 minutes, ~13 retries/second.
The natural resolution — the Outdated node becoming SyncTarget of the UpToDate Primary — is exactly what a manual drbdadm disconnect + connect --discard-my-data produced (a 16 MiB bitmap-based resync that completed in 1 second). It seems the handshake should either fall back to that direction automatically when the chosen sync source is Outdated and the peer is UpToDate, or at minimum stop retrying the same doomed plan at full speed.
Versions
| Component |
Version |
| DRBD kernel module |
9.3.3 (/sys/module/drbd/version) |
| drbd-utils |
9.34.3 (GIT 54d5a64e, 2026-04-17) |
| Kernel / OS |
7.0.0-28-generic, Ubuntu 26.04 LTS |
| Managed by |
LINSTOR 1.33.3 / piraeus-operator v2.10.2 (Kubernetes v1.36.3) |
Topology
Resource pvc-6e9129c9-7a19-4d43-ba3e-a0d28b2ab787, one 10 GiB volume (minor 1001), protocol C, quorum majority, 3 nodes:
- dc01-wrk-01 (node-id 0) — diskful; the node that rebooted; comes back with its data Outdated (a resync it was source for "missed its finish" while the node was down)
- dc01-wrk-04 (node-id 3) — diskful,
UpToDate, Primary (volume open / in use the whole time)
- dc01-wrk-02 (node-id 2) — diskless tie-breaker, Secondary
Kernel log (dc01-wrk-01, single boot; lines lightly trimmed for width)
Boot at 05:36:33; resource brought up by the LINSTOR satellite:
05:38:24 drbd pvc-6e9129c9-...: Starting worker thread (node-id 0)
05:38:24 drbd pvc-6e9129c9-.../0 drbd1001: disk( Attaching -> UpToDate ) [attach]
05:38:24 drbd pvc-6e9129c9-.../0 drbd1001: attached to current UUID: 45FB4CC9C2E63D32
05:38:26 drbd pvc-6e9129c9-... dc01-wrk-02: conn( Connecting -> Connected ) peer( Unknown -> Secondary ) [connected]
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001: disk( UpToDate -> Outdated ) [connected]
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-02: pdsk( DUnknown -> Diskless ) repl( Off -> Established ) [connected]
Handshake with the UpToDate Primary (dc01-wrk-04) — this block then repeats ~13x/second for 20 minutes:
05:38:26 drbd pvc-6e9129c9-...: Preparing cluster-wide state change 2479592503: 0->3 role( Secondary ) conn( Connected )
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: Missed end of resync as sync-source
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: drbd_sync_handshake:
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: self 45FB4CC9C2E63D32:A349C42E41991F5C:D4C6E1172B9D0BB0:94564B13F93C3BE0 bits:4096 flags:20
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: peer 45FB4CC9C2E63D33:0000000000000000:A4A33D080DA23274:F3E3144E74904D32 bits:0 flags:1120
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: uuid_compare()=source-use-bitmap by rule=sync-source-missed-finish
05:38:26 drbd pvc-6e9129c9-.../0 drbd1001 dc01-wrk-04: Resync (rule=sync-source-missed-finish) skipped: sync-source (Outdated)
05:38:26 drbd pvc-6e9129c9-...: Aborting cluster-wide state change 2479592503 (24ms) rv = -19
Loop metrics for this boot: first abort 05:38:26, last abort 05:58:11; journalctl -k -b 0 | grep <rsc> | grep -c "Aborting cluster-wide state change" = 15,943; total kernel lines for the resource: 127,619. (An intermediate invalidate attempt at 05:57:37, repl( StartingSyncT ) toward the diskless peer, had no effect.)
Manual recovery — disconnect, then connect with --discard-my-data on the Outdated node:
05:58:11 ... dc01-wrk-04: conn( Connecting -> Disconnecting ) [disconnect]
05:58:11 ... dc01-wrk-04: conn( StandAlone -> Unconnected ) [connect]
05:58:13 ... drbd1001 dc01-wrk-04: Resync direction reversed by --discard-my-data. Reverting to older data!
05:58:13 ... dc01-wrk-04: conn( Connecting -> Connected ) peer( Unknown -> Primary ) [connected]
05:58:13 ... drbd1001 dc01-wrk-04: pdsk( DUnknown -> UpToDate ) repl( Off -> WFBitMapT ) [connected]
05:58:13 ... drbd1001: disk( Outdated -> Inconsistent ) [receive-bitmap]
05:58:13 ... drbd1001 dc01-wrk-04: repl( WFBitMapT -> SyncTarget ) [receive-bitmap]
05:58:13 ... drbd1001 dc01-wrk-04: Began resync as SyncTarget (will sync 16384 KB [4096 bits set]).
05:58:14 ... drbd1001 dc01-wrk-04: Resync done (total 1 sec; paused 0 sec; 16384 K/sec)
05:58:14 ... drbd1001: disk( Inconsistent -> UpToDate ) [resync-finished]
UUID situation at handshake time: self current 45FB4CC9C2E63D32 with bits:4096 toward the peer; peer current 45FB4CC9C2E63D33 (differs only in the low bit) with a zeroed bitmap UUID slot for this node and bits:0 — i.e. the peer finished the earlier resync (as target) and rotated, while this node was away; hence "Missed end of resync as sync-source".
Expected vs actual
Expected: the connect resolves automatically — with the peer UpToDate and Primary, and the local disk Outdated, the local node should end up SyncTarget (which the manual --discard-my-data proved is a trivial 16 MiB bitmap resync). Or, failing that, the state change should not be retried at full speed forever: the same refusal is recomputed ~13 times per second with no backoff, flooding the log (127k lines) and keeping the connection permanently in Connecting.
Actual: livelock — rule=sync-source-missed-finish insists the Outdated node is the source, the source role is refused because the disk is Outdated (rv = -19), and the handshake repeats indefinitely. Only operator intervention (disconnect + connect --discard-my-data) recovers.
Question
When sync-source-missed-finish selects a sync source whose disk is Outdated and the peer is UpToDate, is falling back to the sync-target direction (the effect of --discard-my-data) safe/intended? If the current behaviour is deliberate, could the retry at least back off instead of looping at ~13 attempts/second?
================================================================================
ORCHESTRATOR NOTES
:
Summary
After a clean reboot of a diskful node, the resync handshake with the surviving diskful
UpToDatePrimarypeer deterministically selectsuuid_compare()=source-use-bitmap by rule=sync-source-missed-finish, i.e. it makes the rebooted, Outdated node the sync source. The node then (correctly) refuses to be a sync source with an Outdated disk —Resync (rule=sync-source-missed-finish) skipped: sync-source (Outdated)— and the cluster-wide state change is aborted withrv = -19. The connection never leavesConnecting, and the handshake retries immediately with no backoff: 15,943 aborted cluster-wide state changes (127,619 kernel log lines) in under 20 minutes, ~13 retries/second.The natural resolution — the Outdated node becoming SyncTarget of the UpToDate Primary — is exactly what a manual
drbdadm disconnect+connect --discard-my-dataproduced (a 16 MiB bitmap-based resync that completed in 1 second). It seems the handshake should either fall back to that direction automatically when the chosen sync source is Outdated and the peer is UpToDate, or at minimum stop retrying the same doomed plan at full speed.Versions
/sys/module/drbd/version)Topology
Resource
pvc-6e9129c9-7a19-4d43-ba3e-a0d28b2ab787, one 10 GiB volume (minor 1001), protocol C,quorum majority, 3 nodes:UpToDate,Primary(volume open / in use the whole time)Kernel log (dc01-wrk-01, single boot; lines lightly trimmed for width)
Boot at 05:36:33; resource brought up by the LINSTOR satellite:
Handshake with the UpToDate Primary (dc01-wrk-04) — this block then repeats ~13x/second for 20 minutes:
Loop metrics for this boot: first abort 05:38:26, last abort 05:58:11;
journalctl -k -b 0 | grep <rsc> | grep -c "Aborting cluster-wide state change"= 15,943; total kernel lines for the resource: 127,619. (An intermediate invalidate attempt at 05:57:37,repl( StartingSyncT )toward the diskless peer, had no effect.)Manual recovery — disconnect, then connect with
--discard-my-dataon the Outdated node:UUID situation at handshake time: self current
45FB4CC9C2E63D32withbits:4096toward the peer; peer current45FB4CC9C2E63D33(differs only in the low bit) with a zeroed bitmap UUID slot for this node andbits:0— i.e. the peer finished the earlier resync (as target) and rotated, while this node was away; hence "Missed end of resync as sync-source".Expected vs actual
Expected: the connect resolves automatically — with the peer
UpToDateandPrimary, and the local disk Outdated, the local node should end up SyncTarget (which the manual--discard-my-dataproved is a trivial 16 MiB bitmap resync). Or, failing that, the state change should not be retried at full speed forever: the same refusal is recomputed ~13 times per second with no backoff, flooding the log (127k lines) and keeping the connection permanently inConnecting.Actual: livelock —
rule=sync-source-missed-finishinsists the Outdated node is the source, the source role is refused because the disk is Outdated (rv = -19), and the handshake repeats indefinitely. Only operator intervention (disconnect+connect --discard-my-data) recovers.Question
When
sync-source-missed-finishselects a sync source whose disk is Outdated and the peer is UpToDate, is falling back to the sync-target direction (the effect of--discard-my-data) safe/intended? If the current behaviour is deliberate, could the retry at least back off instead of looping at ~13 attempts/second?================================================================================
ORCHESTRATOR NOTES