Sync drivers: per-peer zfs replication, source-only syncs, update_requires, and scaling fixes - #1137
Merged
Merged
Conversation
…sync completes, rotating the local snapshots first
…he overall status, as v2 did
…ression, among the peers of the resource
…g the base of a peer lagging past max_lag_age or max_lag_size
… flag a copy stale half a period after the sync due
… the source since
…er holding no snapshot
…another within a second takes its own
…c records, and reuse one ssh connection per zfs peer
…st run of a task or sync is seen
…s the resource and its instance running, as for a task
… is not read as the directory of the same name
… status when scheduling it, and report the first unmet one
… a sync only when met
…oo when recursive
…ync_requires update_requires
…eplication delays, and say where the stale limit comes from
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR reworks the sync resources after running a 3-node service with 30-second zfs, rsync and zfssnap syncs. That exposed several bugs, a few of which could overwrite the active node's data. CHANGELOG.md records the behaviour changes, and the book has a new "Data Replication" page.
Data safety
-u(7b187b2):zfs receivemounted the peers' copies, even withcanmount=noauto. A standby's fs.zfs then read as up, which could make it pass the source check.receive -F.findmntlooked up the device name as a relative path. From/, where every action the daemon starts runs, a zfs datasettank/fsmounted on/tank/fswas taken for that directory and read as unmounted.sync.zfs: each peer has its own base snapshot (b9f92b4, 3ae2320, ba99ddc)
max_lag_age(24h) or holding more thanmax_lag_size(20%) on the source has its base released. It then needsom <path> instance full --rid <rid> --target <peer>, which the status warns about..sent/.tosendsnapshots of older agents are used as bases on the first run.update_requires
sync_requiresis nowupdate_requires, and it gatesupdateandfull.sync_requires,sync_update_requires,sync_nodes_requiresandsync_drp_requiresare read as aliases.update_requireswas never enforced before this series. It is now checked before a sync and holds the scheduled job while unmet.updateandfullin logs, schedules andOPENSVC_ACTION. That's a breaking change for triggers that testsync_update.Status
Schedulefield hid the one the check reads. With that fixed, a copy is stale once the sync due after the last one is half a period late.Scaling
Measured on a 30-second schedule with 2 peers:
The changes behind that (194a467, 1c18330):
Other fixes
--target(7c02f0e) accepts node selector expressions, as the CHANGELOG already said, among the resource's configured peers. A selector matching none of them is an error; before, a node name selected no peer and the action still succeeded.