Skip to content

Sync drivers: per-peer zfs replication, source-only syncs, update_requires, and scaling fixes - #1137

Merged
cvaroqui merged 22 commits into
opensvc:mainfrom
cvaroqui:main
Sep 29, 2026
Merged

cvaroqui merged 22 commits into
opensvc:mainfrom
cvaroqui:main

Conversation

@cvaroqui

Copy link
Copy Markdown
Member

This PR reworks the sync resources after running a 3-node service with 30-second zfs, rsync and zfssnap syncs. That exposed several bugs, a few of which could overwrite the active node's data. CHANGELOG.md records the behaviour changes, and the book has a new "Data Replication" page.

Data safety

  • Syncs are sent only from the active node (d237460, 2e91ead)
    • A standby used to count as the source whenever none of its ip, fs, share, disk or container resources was down, which is always true when the service has none. In a test, both standbys each sent their own copy over the running node.
    • The check is now v2's: the reference resources (everything except app, sync and task) must be up in aggregate. A service with none of them isn't synced.
    • The scheduler now schedules a sync only on that node, so standbys no longer spawn and log a refused sync every period.
  • zfs sync receives with -u (7b187b2): zfs receive mounted the peers' copies, even with canmount=noauto. A standby's fs.zfs then read as up, which could make it pass the source check.
  • A zfs command failing on a peer is reported (94810ac): its error was lost, so a peer whose snapshot listing failed looked like one with no snapshots, and would have received a full copy with receive -F.
  • Datasets are no longer read as directories (3d02553): findmnt looked up the device name as a relative path. From /, where every action the daemon starts runs, a zfs dataset tank/fs mounted on /tank/fs was taken for that directory and read as unmounted.

sync.zfs: each peer has its own base snapshot (b9f92b4, 3ae2320, ba99ddc)

  • Per-peer bases: each run takes a snapshot named to the microsecond. Each peer is sent the changes since the newest snapshot it shares with the source, found by GUID. A peer that missed runs catches up on its own, and a failing peer no longer blocks the others.
  • Lag limits: a peer lagging past max_lag_age (24h) or holding more than max_lag_size (20%) on the source has its base released. It then needs om <path> instance full --rid <rid> --target <peer>, which the status warns about.
  • Safe first and diverged syncs: a peer that was never synced still gets a full copy automatically. A peer with no snapshot in common with the source is never overwritten without being asked.
  • Failover: a node that becomes the source forgets the peer state it kept from an earlier time as source.
  • Upgrade: the .sent/.tosend snapshots of older agents are used as bases on the first run.

update_requires

  • Rename (a3e63d4): sync_requires is now update_requires, and it gates update and full. sync_requires, sync_update_requires, sync_nodes_requires and sync_drp_requires are read as aliases.
  • Enforcement (cec2ea2): update_requires was never enforced before this series. It is now checked before a sync and holds the scheduled job while unmet.
  • Scheduler (6cdcb4c): it now evaluates job requirements from the last status after a daemon restart, and reports the first unmet condition; before, only the last condition checked counted.
  • Action names (a3e63d4): the actions are named update and full in logs, schedules and OPENSVC_ACTION. That's a breaking change for triggers that test sync_update.

Status

  • Sync in overall only (3f7b193): a sync resource counts in the instance's overall status only, as v2 did. An up sync counts as nothing, and a down sync as a warning. A standby receiving copies now reads down instead of warn.
  • Stale limit (a31c3b6): the limit was never computed from the schedule, because each driver's own Schedule field hid the one the check reads. With that fixed, a copy is stale once the sync due after the last one is half a period late.
  • zfssnap on standbys (a508b3a): the replicated snapshots are judged against the snapshot delay plus the replication delay. Stale messages name where their limit comes from.
  • R flag (a9687d8, b4c732f): syncs in progress write run files, so the resource and its instance show R, on every node, as tasks do. The daemon now creates missing run directories, so the first run of a task or sync is seen too.

Scaling

Measured on a 30-second schedule with 2 peers:

  • zfs: a run takes 0.46 s instead of 1.16 s, with one ssh login per peer instead of three.
  • rsync: a run takes 1.1 s instead of 2.4 s.

The changes behind that (194a467, 1c18330):

  • peers are synced in parallel;
  • each peer gets 3 last-sync posts instead of 3 × peers², and the responses are now closed;
  • the zfs driver uses one ssh connection per peer per run;
  • the state-file log lines moved to debug level, and the stats lines give the figures.

Other fixes

  • --target (7c02f0e) accepts node selector expressions, as the CHANGELOG already said, among the resource's configured peers. A selector matching none of them is an error; before, a node name selected no peer and the action still succeeded.
  • zfssnap (c6af66d) destroys old snapshots in the descendant datasets too when recursive.
  • Last-run file (b3a3eea): the file sent to peers is now the one the scheduler reads.
  • Last-sync write failures (2edf5a4) are reported after the sync, without skipping the local snapshot rotation.

…sync completes, rotating the local snapshots first
…g the base of a peer lagging past max_lag_age or max_lag_size
… flag a copy stale half a period after the sync due
…c records, and reuse one ssh connection per zfs peer
…s the resource and its instance running, as for a task
… is not read as the directory of the same name
… status when scheduling it, and report the first unmet one
…eplication delays, and say where the stale limit comes from
@cvaroqui
cvaroqui merged commit 7de366e into opensvc:main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant