Summary
When a diskful secondary rejoins the cluster ~2 minutes after a clean
shutdown, the sync handshake on the surviving primary decides
uuid_compare()=no-sync by rule=lost-quorum and clears the primary's
non-empty sync bitmap (bits:16), declaring the rejoining node
UpToDate without replicating the blocks written while it was away. The
replicas silently diverge while both report UpToDate. In our case a
workload was scheduled onto the stale replica minutes later, mounted it as
primary, and the missing ext4 metadata caused filesystem corruption
(multiply-claimed blocks), which then replicated back onto the good
replica.
We believe the lost-quorum no-sync rule is intended for the opposite
direction — a node that lost quorum and had its suspended writes discarded
should not force a resync of data it threw away. Here, the rejoining
secondary recorded a local quorum loss during its own shutdown (network
teardown while the device was still attached), while the surviving
primary held quorum via the diskless tiebreaker and kept writing
legitimate data. Clearing the survivor's bitmap discards real writes.
Environment
-
DRBD kernel module 9.3.3 (api:2/proto:118-124), GIT-hash
97da76040a6b31aaf9e12f1a167e77ca2b3cb43e, transports tcp/lb-tcp/rdma —
latest release at the time of writing
-
Built/loaded by LINSTOR satellite (Piraeus operator), LINSTOR controller
1.33.3
-
Ubuntu 24.04, kernel 6.8.0-139-generic
-
3-node k3s cluster; the affected resource has 2 diskful replicas
(node01, node02) + 1 diskless tiebreaker (node03)
-
Backing storage: LVM thin
-
Resource configuration (drbdsetup show, LINSTOR-generated defaults;
shared-secret redacted):
options {
on-no-data-accessible suspend-io;
quorum majority;
on-suspended-primary-outdated force-secondary;
}
# per-connection net options: rr-conflict retry-connect; verify-alg "crc32c";
# the connection to the diskless tiebreaker has "bitmap no" on volume 0
on-no-quorum is not set explicitly (so the suspend-io default applies,
consistent with the observed susp-io( no -> quorum ) log line).
Node IDs for readability below: node02 = peer-node-id 0, node03 = 1,
node01 = 2. All timestamps 2026-09-07, same NTP-synced clock domain.
Timeline
Context: rolling node maintenance. The resource pvc-c4770b79-…
(drbd1004, 5 GiB, ext4) hosts a continuously-writing workload
(SeaweedFS filer, leveldb).
-
15:13:42 — node02 demotes cleanly (workload evicted, clean
umount, role( Primary -> Secondary )).
-
15:13:57 — node01 auto-promotes and mounts; a pre-mount
fsck.ext4 -a passes (filesystem healthy at this point). Workload
writes continuously from here.
-
15:14:07–09 — node02 is shut down for maintenance.
k3s-killall.sh tears down the network namespace while the DRBD device
is still attached (secondary, unmounted). node02 logs:
drbd pvc-c4770b79-...: Disconnect because network namespace is exiting
drbd pvc-c4770b79-...: susp-io( no -> quorum )
drbd pvc-c4770b79-.../0 drbd1004: disk( UpToDate -> Consistent ) quorum( yes -> no )
drbd pvc-c4770b79-.../0 drbd1004: disk( Consistent -> UpToDate ) [lost-peer]
Node powers off at 15:14:09.
-
15:14:08 — node01 (primary, writing): PingAck did not arrive in time, then Would lose quorum, but using tiebreaker logic to keep.
Quorum retained via the diskless tiebreaker; writes continue and
accumulate in the bitmap towards node02.
-
15:15:20 — node02 boots; DRBD module reloaded.
-
15:15:58 — reconnect. Handshake on node01:
drbd pvc-c4770b79-.../0 drbd1004 k3s-node02: drbd_sync_handshake:
... self 4AA3D00411809698:4AA3D00411809699:A95E297FBC2654FC:2252CD43388F9E2E bits:16 flags:120
... peer 4AA3D00411809698:0000000000000000:A95E297FBC2654FC:2252CD43388F9E2E bits:0 flags:1020
... uuid_compare()=no-sync by rule=lost-quorum
... cleared bm UUID and bitmap 4AA3D00411809698:0000000000000000:A95E297FBC2654FC:2252CD43388F9E2E
... pdsk( DUnknown -> UpToDate ) repl( Off -> Established )
Expected: bitmap-based resync of the 16 out-of-sync blocks,
node01 → node02.
Actual: bitmap cleared, no resync, node02 immediately UpToDate.
Silent divergence from this moment.
-
15:16:24 — workload is rescheduled off node01; clean unmount and
demote on node01.
-
15:17:03 — workload is scheduled onto node02, which auto-promotes
and mounts the stale replica (pre-mount fsck -a did not catch it).
ext4 metadata written on node01 during 15:14–15:16 is missing, so ext4
hands out blocks that are already in use in the "real" state. These
writes replicate back to node01, poisoning both replicas identically.
-
15:20:03 — workload crashes; from here every pre-mount fsck fails
with UNEXPECTED INCONSISTENCY / "Duplicate or bad block in use! /
Multiply-claimed block(s)". Manual e2fsck -f -y later confirmed
multiply-claimed blocks across files whose mtimes straddle the
divergence window (some 15:13:59, some 15:17:05), i.e. a mix of the two
diverged states.
Throughout all of this, linstor resource list showed the resource as
healthy — both diskful replicas UpToDate, connections Ok.
Analysis / questions
- The current-UUIDs compare equal (
4AA3D00411809698 on both sides) even
though the primary wrote for ~2 minutes without the peer. Combined with
the peer's lost-quorum flag (flags:1020, exposed UUID
0000000000000000), this appears to land in the lost-quorum no-sync
branch — which then clears the survivor's non-empty bitmap.
- Shouldn't a no-sync decision be refused (or converted to a bitmap-based
resync) whenever self bits > 0? Whatever the UUID logic concludes, a
non-empty bitmap on the surviving side is positive evidence that the
peer is missing writes.
- Is it expected that a primary that keeps quorum via a diskless
tiebreaker does not rotate its current UUID when it loses a diskful
peer mid-write? If the UUID had been bumped, the handshake would
presumably have chosen a normal bitmap resync.
Reproduction sketch
- 3 nodes: A and B diskful, C diskless tiebreaker; quorum=majority.
- Mount the resource on A (primary) and write continuously.
- On B (secondary, attached): tear down networking in a way that lets
DRBD observe it while the device is still attached (we hit it via
k3s-killall.sh destroying the network namespace on shutdown), so B
records susp-io( no -> quorum ) / quorum( yes -> no ) before going
down.
- Keep writing on A (accumulate bitmap bits towards B).
- Boot B back within a couple of minutes.
- Observe the handshake on A:
no-sync by rule=lost-quorum with
self bits > 0, bitmap cleared, no resync.
Impact
Silent replica divergence with green status everywhere, followed by real
filesystem corruption of a production volume once the stale replica was
promoted. Recovered via manual e2fsck plus application-level repair.
Full journalctl excerpts from all three nodes for the whole window are
available on request, as is the LINSTOR resource definition.
Summary
When a diskful secondary rejoins the cluster ~2 minutes after a clean
shutdown, the sync handshake on the surviving primary decides
uuid_compare()=no-sync by rule=lost-quorumand clears the primary'snon-empty sync bitmap (
bits:16), declaring the rejoining nodeUpToDatewithout replicating the blocks written while it was away. Thereplicas silently diverge while both report
UpToDate. In our case aworkload was scheduled onto the stale replica minutes later, mounted it as
primary, and the missing ext4 metadata caused filesystem corruption
(multiply-claimed blocks), which then replicated back onto the good
replica.
We believe the
lost-quorumno-sync rule is intended for the oppositedirection — a node that lost quorum and had its suspended writes discarded
should not force a resync of data it threw away. Here, the rejoining
secondary recorded a local quorum loss during its own shutdown (network
teardown while the device was still attached), while the surviving
primary held quorum via the diskless tiebreaker and kept writing
legitimate data. Clearing the survivor's bitmap discards real writes.
Environment
DRBD kernel module 9.3.3 (api:2/proto:118-124), GIT-hash
97da76040a6b31aaf9e12f1a167e77ca2b3cb43e, transports tcp/lb-tcp/rdma —latest release at the time of writing
Built/loaded by LINSTOR satellite (Piraeus operator), LINSTOR controller
1.33.3
Ubuntu 24.04, kernel
6.8.0-139-generic3-node k3s cluster; the affected resource has 2 diskful replicas
(node01, node02) + 1 diskless tiebreaker (node03)
Backing storage: LVM thin
Resource configuration (
drbdsetup show, LINSTOR-generated defaults;shared-secret redacted):
on-no-quorumis not set explicitly (so the suspend-io default applies,consistent with the observed
susp-io( no -> quorum )log line).Node IDs for readability below: node02 = peer-node-id 0, node03 = 1,
node01 = 2. All timestamps 2026-09-07, same NTP-synced clock domain.
Timeline
Context: rolling node maintenance. The resource
pvc-c4770b79-…(
drbd1004, 5 GiB, ext4) hosts a continuously-writing workload(SeaweedFS filer, leveldb).
15:13:42 — node02 demotes cleanly (workload evicted, clean
umount,role( Primary -> Secondary )).15:13:57 — node01 auto-promotes and mounts; a pre-mount
fsck.ext4 -apasses (filesystem healthy at this point). Workloadwrites continuously from here.
15:14:07–09 — node02 is shut down for maintenance.
k3s-killall.shtears down the network namespace while the DRBD deviceis still attached (secondary, unmounted). node02 logs:
Node powers off at 15:14:09.
15:14:08 — node01 (primary, writing):
PingAck did not arrive in time, thenWould lose quorum, but using tiebreaker logic to keep.Quorum retained via the diskless tiebreaker; writes continue and
accumulate in the bitmap towards node02.
15:15:20 — node02 boots; DRBD module reloaded.
15:15:58 — reconnect. Handshake on node01:
Expected: bitmap-based resync of the 16 out-of-sync blocks,
node01 → node02.
Actual: bitmap cleared, no resync, node02 immediately
UpToDate.Silent divergence from this moment.
15:16:24 — workload is rescheduled off node01; clean unmount and
demote on node01.
15:17:03 — workload is scheduled onto node02, which auto-promotes
and mounts the stale replica (pre-mount
fsck -adid not catch it).ext4 metadata written on node01 during 15:14–15:16 is missing, so ext4
hands out blocks that are already in use in the "real" state. These
writes replicate back to node01, poisoning both replicas identically.
15:20:03 — workload crashes; from here every pre-mount fsck fails
with
UNEXPECTED INCONSISTENCY/ "Duplicate or bad block in use! /Multiply-claimed block(s)". Manual
e2fsck -f -ylater confirmedmultiply-claimed blocks across files whose mtimes straddle the
divergence window (some 15:13:59, some 15:17:05), i.e. a mix of the two
diverged states.
Throughout all of this,
linstor resource listshowed the resource ashealthy — both diskful replicas
UpToDate, connectionsOk.Analysis / questions
4AA3D00411809698on both sides) eventhough the primary wrote for ~2 minutes without the peer. Combined with
the peer's lost-quorum flag (
flags:1020, exposed UUID0000000000000000), this appears to land in thelost-quorumno-syncbranch — which then clears the survivor's non-empty bitmap.
resync) whenever
self bits > 0? Whatever the UUID logic concludes, anon-empty bitmap on the surviving side is positive evidence that the
peer is missing writes.
tiebreaker does not rotate its current UUID when it loses a diskful
peer mid-write? If the UUID had been bumped, the handshake would
presumably have chosen a normal bitmap resync.
Reproduction sketch
DRBD observe it while the device is still attached (we hit it via
k3s-killall.sh destroying the network namespace on shutdown), so B
records
susp-io( no -> quorum )/quorum( yes -> no )before goingdown.
no-sync by rule=lost-quorumwithself bits > 0, bitmap cleared, no resync.Impact
Silent replica divergence with green status everywhere, followed by real
filesystem corruption of a production volume once the stale replica was
promoted. Recovered via manual
e2fsckplus application-level repair.Full
journalctlexcerpts from all three nodes for the whole window areavailable on request, as is the LINSTOR resource definition.