Skip to content

Silent data divergence: sync handshake "no-sync by rule=lost-quorum" clears non-empty bitmap when diskful secondary rejoins after clean shutdown #145

Description

@magnusvin

Summary

When a diskful secondary rejoins the cluster ~2 minutes after a clean
shutdown, the sync handshake on the surviving primary decides
uuid_compare()=no-sync by rule=lost-quorum and clears the primary's
non-empty sync bitmap
(bits:16), declaring the rejoining node
UpToDate without replicating the blocks written while it was away. The
replicas silently diverge while both report UpToDate. In our case a
workload was scheduled onto the stale replica minutes later, mounted it as
primary, and the missing ext4 metadata caused filesystem corruption
(multiply-claimed blocks), which then replicated back onto the good
replica.

We believe the lost-quorum no-sync rule is intended for the opposite
direction — a node that lost quorum and had its suspended writes discarded
should not force a resync of data it threw away. Here, the rejoining
secondary
recorded a local quorum loss during its own shutdown (network
teardown while the device was still attached), while the surviving
primary held quorum via the diskless tiebreaker and kept writing
legitimate data
. Clearing the survivor's bitmap discards real writes.

Environment

  • DRBD kernel module 9.3.3 (api:2/proto:118-124), GIT-hash
    97da76040a6b31aaf9e12f1a167e77ca2b3cb43e, transports tcp/lb-tcp/rdma —
    latest release at the time of writing

  • Built/loaded by LINSTOR satellite (Piraeus operator), LINSTOR controller
    1.33.3

  • Ubuntu 24.04, kernel 6.8.0-139-generic

  • 3-node k3s cluster; the affected resource has 2 diskful replicas
    (node01, node02) + 1 diskless tiebreaker (node03)

  • Backing storage: LVM thin

  • Resource configuration (drbdsetup show, LINSTOR-generated defaults;
    shared-secret redacted):

    options {
        on-no-data-accessible   suspend-io;
        quorum                  majority;
        on-suspended-primary-outdated   force-secondary;
    }
    # per-connection net options: rr-conflict retry-connect; verify-alg "crc32c";
    # the connection to the diskless tiebreaker has "bitmap no" on volume 0
    

    on-no-quorum is not set explicitly (so the suspend-io default applies,
    consistent with the observed susp-io( no -> quorum ) log line).

Node IDs for readability below: node02 = peer-node-id 0, node03 = 1,
node01 = 2. All timestamps 2026-09-07, same NTP-synced clock domain.

Timeline

Context: rolling node maintenance. The resource pvc-c4770b79-…
(drbd1004, 5 GiB, ext4) hosts a continuously-writing workload
(SeaweedFS filer, leveldb).

  1. 15:13:42 — node02 demotes cleanly (workload evicted, clean
    umount, role( Primary -> Secondary )).

  2. 15:13:57 — node01 auto-promotes and mounts; a pre-mount
    fsck.ext4 -a passes (filesystem healthy at this point). Workload
    writes continuously from here.

  3. 15:14:07–09 — node02 is shut down for maintenance.
    k3s-killall.sh tears down the network namespace while the DRBD device
    is still attached (secondary, unmounted). node02 logs:

    drbd pvc-c4770b79-...: Disconnect because network namespace is exiting
    drbd pvc-c4770b79-...: susp-io( no -> quorum )
    drbd pvc-c4770b79-.../0 drbd1004: disk( UpToDate -> Consistent ) quorum( yes -> no )
    drbd pvc-c4770b79-.../0 drbd1004: disk( Consistent -> UpToDate ) [lost-peer]
    

    Node powers off at 15:14:09.

  4. 15:14:08 — node01 (primary, writing): PingAck did not arrive in time, then Would lose quorum, but using tiebreaker logic to keep.
    Quorum retained via the diskless tiebreaker; writes continue and
    accumulate in the bitmap towards node02.

  5. 15:15:20 — node02 boots; DRBD module reloaded.

  6. 15:15:58 — reconnect. Handshake on node01:

    drbd pvc-c4770b79-.../0 drbd1004 k3s-node02: drbd_sync_handshake:
    ... self 4AA3D00411809698:4AA3D00411809699:A95E297FBC2654FC:2252CD43388F9E2E bits:16 flags:120
    ... peer 4AA3D00411809698:0000000000000000:A95E297FBC2654FC:2252CD43388F9E2E bits:0 flags:1020
    ... uuid_compare()=no-sync by rule=lost-quorum
    ... cleared bm UUID and bitmap 4AA3D00411809698:0000000000000000:A95E297FBC2654FC:2252CD43388F9E2E
    ... pdsk( DUnknown -> UpToDate ) repl( Off -> Established )
    

    Expected: bitmap-based resync of the 16 out-of-sync blocks,
    node01 → node02.
    Actual: bitmap cleared, no resync, node02 immediately UpToDate.
    Silent divergence from this moment.

  7. 15:16:24 — workload is rescheduled off node01; clean unmount and
    demote on node01.

  8. 15:17:03 — workload is scheduled onto node02, which auto-promotes
    and mounts the stale replica (pre-mount fsck -a did not catch it).
    ext4 metadata written on node01 during 15:14–15:16 is missing, so ext4
    hands out blocks that are already in use in the "real" state. These
    writes replicate back to node01, poisoning both replicas identically.

  9. 15:20:03 — workload crashes; from here every pre-mount fsck fails
    with UNEXPECTED INCONSISTENCY / "Duplicate or bad block in use! /
    Multiply-claimed block(s)". Manual e2fsck -f -y later confirmed
    multiply-claimed blocks across files whose mtimes straddle the
    divergence window (some 15:13:59, some 15:17:05), i.e. a mix of the two
    diverged states.

Throughout all of this, linstor resource list showed the resource as
healthy — both diskful replicas UpToDate, connections Ok.

Analysis / questions

  • The current-UUIDs compare equal (4AA3D00411809698 on both sides) even
    though the primary wrote for ~2 minutes without the peer. Combined with
    the peer's lost-quorum flag (flags:1020, exposed UUID
    0000000000000000), this appears to land in the lost-quorum no-sync
    branch — which then clears the survivor's non-empty bitmap.
  • Shouldn't a no-sync decision be refused (or converted to a bitmap-based
    resync) whenever self bits > 0? Whatever the UUID logic concludes, a
    non-empty bitmap on the surviving side is positive evidence that the
    peer is missing writes.
  • Is it expected that a primary that keeps quorum via a diskless
    tiebreaker does not rotate its current UUID when it loses a diskful
    peer mid-write? If the UUID had been bumped, the handshake would
    presumably have chosen a normal bitmap resync.

Reproduction sketch

  1. 3 nodes: A and B diskful, C diskless tiebreaker; quorum=majority.
  2. Mount the resource on A (primary) and write continuously.
  3. On B (secondary, attached): tear down networking in a way that lets
    DRBD observe it while the device is still attached (we hit it via
    k3s-killall.sh destroying the network namespace on shutdown), so B
    records susp-io( no -> quorum ) / quorum( yes -> no ) before going
    down.
  4. Keep writing on A (accumulate bitmap bits towards B).
  5. Boot B back within a couple of minutes.
  6. Observe the handshake on A: no-sync by rule=lost-quorum with
    self bits > 0, bitmap cleared, no resync.

Impact

Silent replica divergence with green status everywhere, followed by real
filesystem corruption of a production volume once the stale replica was
promoted. Recovered via manual e2fsck plus application-level repair.

Full journalctl excerpts from all three nodes for the whole window are
available on request, as is the LINSTOR resource definition.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions