Sync safety and RPO exposure, event-driven status refresh, narrower freeze adoption, and predictable listings - #1138
Merged
Merged
Conversation
…ted, and when a peer writes a state file of it
…objective, and mark those instances in om mon and the status tree
…ance, for wait_syncs_timeout at most, and schedule no sync during an orchestration
…ync on a node being drained
…ir type in a copy
…ning instead of waiting for them
… and wait for it as long as an orchestrated stop
… book does not read them as HTML tags
…ured environment, as docker exec does
… recovering it when it is removed Since the daemon announces a configuration file it wrote without waiting for the filesystem watcher, the watcher announced the same write again, 200ms later. On the node writing a configuration whose scope left it, the first announce ended the instance config manager, and the second started another one, which published the configuration for its peers a second time. A peer that had installed it on the first one and started its own manager took the second as a sign its file was foreign, and set the flag disabling the recovery of the file. The flag stayed, and a later removal of the file was not recovered from the peers, for some of the objects, depending on the timing. The announces of the daemon now go through configannounce, which records the modification time each file was announced at. The watcher consumes that record instead of announcing a write already announced: one announce hides one watcher event, at that modification time only, and a removal drops the record, so a file put back later is announced. A peer's configuration for the local node no longer disables the recovery of the local file either: it is fetched, not removed, and only a foreign configuration, local out of its scope or a peer's not for this node, ends in a removal the recovery must not undo.
…ode frozen since now, which froze the peers restarting with it The node monitor started with the node frozen since now, until discover scanned the frozen flag file, so the node was not seen unfrozen before the flag was known. That frozen date was published to the peers. A peer rejoining takes a node frozen after it went down for an operator freeze it missed, and freezes itself: a peer whose daemon restarted along with this one, as with "om daemon restart --node='*'" or a rolling deploy, froze on a freeze that never was, and the next nodes to rejoin froze on its freeze, real this time. The frozen state is now read from the flag file when the node monitor starts, as the daemon data does, so the node is neither seen unfrozen while frozen, nor frozen while it is not.
…ng back from down A daemon joining the cluster froze the node when any peer node had been frozen while it was down, and froze an ha instance when any peer instance had. A node frozen alone for its maintenance froze its peers as they rebooted, and a freeze the daemon took on its own, at the end of a rejoin grace period, spread from node to node through their restarts. OpenSVC v2 does the same, and its users are annoyed by the freezes nobody asked for. A frozen flag now records the scope of the freeze that raised it. The node monitor freezes the node for a freeze of the cluster with OSVC_FREEZE_SCOPE=cluster in the environment of the "node freeze" it forks, and the flag says "cluster". The instance monitor freezes the instance for a freeze of the object, and the flag says "object". Every other freeze leaves the flag empty, as the flags raised before this change are: a freeze of the node or of the instance alone. The node and instance statuses publish it as frozen_scope, "cluster" or "node", "object" or "instance", and a node coming back adopts only a freeze of the cluster, an instance only a freeze of the object, taken by a peer while it was down. A freeze adopted is one of the node or of the instance, and is not adopted again from there. The freeze of a node reached by a global expect logged that the node was not frozen: it says it is.
… the disk status is n/a, so a node standing by does not warn The registrations a device must have depend on whether the instance holds it, which the status of the resource the reservation belongs to tells: up, every path must be registered, down, none may be. A resource n/a, like a disk that is there whatever the state of the instance, tells nothing, and was taken as up, so every node standing by for the instance warned of "0/2 registrations". A reservation held with our key now witnesses that the instance holds the device, and every path must be registered. Not held, the registrations are logged and not judged, and the resource is down as the reservation is, as the reservation resource of v2 was on a node standing by.
…er, and flag nothing stopped for the stop ending it A provision without --leader and without --disable-rollback ended with a stop of the selection, run as the stop a user asks for: the instance was flagged stopped on purpose, so the daemon would not start it on its own, a resource provisioned with --rid was flagged stopped (X), so the daemon would not restart it, and a resource provisioned with --rid on a running instance was stopped there. The stop ending a provision is now a step of it: the action properties name the action a step belongs to, and a step raises neither the instance nor the resource stopped flag. It stops only the resources the provision found down, read from a status evaluated before it, and nothing on an instance found running. Running is judged on the resources provisioned already: a resource being added is down and would make a running instance read warn, while a node standing by with a disk up reads warn too. The rollback stack is not replayed instead, as v2 does: it holds the undo of provisioning steps too, which a provision that succeeded must keep.
…hanges, as a scsi reservation holder or a drbd peer does Some resource statuses read a state the actions of the peer instances change: the key holding a scsi persistent reservation, which a peer takes or drops as it starts or stops, the disk states of the drbd peers, which a peer changes as it takes its drbd resource down or up, the pair state and personality of a srdf device group, which a failover swaps. Nothing local tells, and the nodes standing by kept reading "reserved by <key>" after the holder had stopped, until the next scheduled status. A driver says so with the optional StatusDependsOnPeers interface, which disk.drbd and sync.symsrdfs implement, and the core answers it for every resource with scsireserv on, whatever its driver. The resource status publishes it as depends_on_peers. The instance monitor compares each peer instance status with the one before, and refreshes the local status 2s later when a resource depending on the peers had its peer counterpart change status, or the peer instance change availability. Only status values trigger it, not log lines nor timestamps, so two nodes depending on each other stop once their statuses settle. The refreshes asked by a peer event, this one and the state file a sync source writes, get a timer of their own: they were armed on the outdated timer, which every local status update arms again or stops, and a local update landing within the delay dropped them.
…reshes the node and peer events ask for A drbd resource changes state with no action of the node to tell: its peers connect once it is up, a resync ends, a peer takes its disk down. The status read at the end of an action saw the state of that moment, and a node bringing its drbd resource up kept reading its peers DUnknown until the next scheduled status. A resync length can not be guessed to read the status again at the right time. The drbdmon daemon component reads "drbdsetup events2", as mntmon reads the mount table, and publishes DrbdResourceUpdated for a drbd resource whose role, disk, peer disk, connection or replication state changed, once its changes settle. The instance monitor refreshes the status of an instance with a disk.drbd resource of that name. The monitor runs where the drbd event stream is usable: drbd 9, on every distribution down to el7. It disables itself, and says so once, where drbdsetup is not installed, where "drbdsetup events2 --now" fails, as with a drbd that does not know the stream, and where the stream ends right after it starts three times in a row. A node where the drbd kernel module is not loaded waits for it. The monitors report one change of the node from several sources at nearly the same time, a drbd promotion and the mount that follows it, and each asked for a status refresh at once. Every event-driven refresh, of a mount, an address, a drbd state, a peer instance change and a peer state file, now goes through one scheduler: the first event schedules a refresh a second later, which the events arriving meanwhile join, and the refresh is not due sooner than 5 seconds after the last status evaluation, whatever asked for it.
…he command chose an order, and count digit runs as numbers A listing assembled from maps, and from the answers of several nodes, came in a different order at each run: "om vol resource ls" listed the same resources shuffled from one call to the next. Only the exec, session and orchestration listings chose an order. A listing whose command sets no default sort, and whose caller asks for none, is now ordered on its columns, left to right, in json as in the table. A column this default order can not read, or that no item carries, is passed over, as the caller did not name it. The texts compare the way a reader counts: their digit runs compare as the numbers they write, so disk#2 comes before disk#10, and n2 before n10, whether the order is the default one or asked with --sort. The heartbeat listing, whose first columns are state icons, orders on its node, stream and peer, as its own sort did, now declared as its default sort, which also puts hb#2 before hb#10.
…us or a monitor yet The configuration of an instance is known before its status and its monitor are, as right after the daemon starts, or while its monitor is still starting. The resource listing read the resources of both without checking, and a nil monitor crashed the handler: "om vol resource ls" failed with an http2 INTERNAL_ERROR. An instance without a status has no resource to list yet and is passed over, and the monitor of a resource is added when there is one.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Syncs: safer stops and switches, visible recovery point objective
wait syncsmonitor state, forwait_syncs_timeoutat most (default 10m). Past it, the stop fails and the instance keeps running with its monitoring on.--interrupt-syncsonstopandswitch(interrupt_syncson the API) ends the running syncs instead, the rsync and ssh processes they started included.instance stopwaits for a sync holding the object lock, says so, and keeps to an explicit--waitlock.rpo_breached_at, explicit or derived from its schedule), and the instance publishes the earliest.om monmarks those instances with a redL, and the status tree tags themrpo-breached.outdated_at), and the daemon evaluates the status again then, and when a peer writes a state file of it.Status freshness: event-driven refreshes
depends_on_peers): a scsi persistent reservation,disk.drbd,sync.symsrdfs. When the peer counterpart changes, or the peer instance changes availability, the daemon refreshes the local status. Standby nodes no longer keep reading "reserved by " after the holder stopped.drbdmondaemon component: it readsdrbdsetup events2and publishesDrbdResourceUpdated, so connections coming up and resyncs ending refresh the status of the instances holding the drbd resource. It disables itself gracefully, logged once, where drbdsetup is missing or the event stream is unusable, and waits for the kernel module where it is not loaded yet. Checked on el7 (drbd 9.2.20) through Debian 13.status -r, and paces the refreshes to one per 5s after the last status evaluation.Freezes
om daemon restart --node='*'or a rolling deploy did. At startup, the node monitor published the node as frozen since then, and the peers adopted that freeze.frozen_scope:cluster/node,object/instance.Provision
X).Configuration replication
CLI and API
--sortnor the command chooses one, the rows are sorted on their columns, left to right, in json as in tables. Digit runs compare as numbers (disk#2beforedisk#10).om <path> container enterruns with the container's configured environment, asdocker execdoes.Behaviour changes (recorded in CHANGELOG.md)
wait_syncs_timeout, unless--interrupt-syncsis set.API and events
frozen_scope,outdated_at,rpo_breached_at;frozen_scope;depends_on_peers,outdated_at,rpo_breached_at.interrupt_syncson the stop and switch actions.wait_syncs_timeout.DrbdResourceUpdated,InstanceStateFileUpdated.Testing
Unit tests are added for the freeze scopes and adoption rules, the provision "running before" rule, drbd event parsing and the probe, event refresh coalescing and pacing, peer change detection, listing order and natural comparison, and config announce deduplication. Each change was also exercised on the dev2 3-node cluster: