Skip to content

Sync safety and RPO exposure, event-driven status refresh, narrower freeze adoption, and predictable listings - #1138

Merged
cvaroqui merged 19 commits into
opensvc:mainfrom
cvaroqui:main
Sep 29, 2026
Merged

cvaroqui merged 19 commits into
opensvc:mainfrom
cvaroqui:main

Conversation

@cvaroqui

Copy link
Copy Markdown
Member

Syncs: safer stops and switches, visible recovery point objective

  • A stop or a switch waits for the syncs running on the instance, in the new wait syncs monitor state, for wait_syncs_timeout at most (default 10m). Past it, the stop fails and the instance keeps running with its monitoring on.
  • --interrupt-syncs on stop and switch (interrupt_syncs on the API) ends the running syncs instead, the rsync and ssh processes they started included.
  • A local instance stop waits for a sync holding the object lock, says so, and keeps to an explicit --waitlock.
  • A drain, or a shutdown, ends the running syncs rather than waiting for them. No scheduled sync starts during an orchestration or on a node being drained.
  • Each sync resource publishes when its copy breaches its recovery point objective (rpo_breached_at, explicit or derived from its schedule), and the instance publishes the earliest. om mon marks those instances with a red L, and the status tree tags them rpo-breached.
  • A resource says when its status goes outdated (outdated_at), and the daemon evaluates the status again then, and when a peer writes a state file of it.

Status freshness: event-driven refreshes

  • A resource status that reads a state the peer instances change says so (depends_on_peers): a scsi persistent reservation, disk.drbd, sync.symsrdfs. When the peer counterpart changes, or the peer instance changes availability, the daemon refreshes the local status. Standby nodes no longer keep reading "reserved by " after the holder stopped.
  • New drbdmon daemon component: it reads drbdsetup events2 and publishes DrbdResourceUpdated, so connections coming up and resyncs ending refresh the status of the instances holding the drbd resource. It disables itself gracefully, logged once, where drbdsetup is missing or the event stream is unusable, and waits for the kernel module where it is not loaded yet. Checked on el7 (drbd 9.2.20) through Debian 13.
  • Every event-driven refresh (mount, ip, drbd, peer change, peer state file) goes through one scheduler. It coalesces the events of one change into one status -r, and paces the refreshes to one per 5s after the last status evaluation.
  • The scsi reservation check takes the reservation as the witness of a started instance when the disk status is n/a. Standby nodes no longer warn "0/2 registrations".

Freezes

  • A daemon restart no longer freezes the nodes restarting with it, as om daemon restart --node='*' or a rolling deploy did. At startup, the node monitor published the node as frozen since then, and the peers adopted that freeze.
  • A node coming back from down adopts only a freeze of the whole cluster or of a whole object, taken while it was down. A node or instance frozen alone, by a drain, at boot, by the end of a rejoin grace period, or by an adoption, stays the only one frozen. The node and instance statuses expose it as frozen_scope: cluster/node, object/instance.

Provision

  • Off the placement leader, a provision leaves the instance as it found it. On an instance found running, it stops nothing. On one found down, it stops only what it started. The stop ending it raises neither the instance nor the resource stopped flag (X).

Configuration replication

  • A configuration file the daemon writes is announced once. The watcher's second announce could leave a peer with a stale "disable recover" flag, and a later removal of the file was no longer recovered from the peers.

CLI and API

  • Listings come in a predictable order. When neither --sort nor the command chooses one, the rows are sorted on their columns, left to right, in json as in tables. Digit runs compare as numbers (disk#2 before disk#10).
  • om <path> container enter runs with the container's configured environment, as docker exec does.
  • The resource listing no longer crashes on an instance without a status or monitor yet.
  • The keyword texts put their placeholders in code spans, so the book no longer reads them as HTML tags.

Behaviour changes (recorded in CHANGELOG.md)

  • Stop and switch wait for running syncs, bounded by wait_syncs_timeout, unless --interrupt-syncs is set.
  • Freeze adoption at rejoin is narrowed to cluster and object freezes. Freeze flags written before the upgrade read as freezes of the node or instance alone.
  • A provision off the leader leaves the instance as found, and flags nothing stopped.

API and events

  • New fields:
    • instance status: frozen_scope, outdated_at, rpo_breached_at;
    • node status: frozen_scope;
    • resource status: depends_on_peers, outdated_at, rpo_breached_at.
  • New parameters: interrupt_syncs on the stop and switch actions.
  • New keyword: wait_syncs_timeout.
  • New events: DrbdResourceUpdated, InstanceStateFileUpdated.

Testing

Unit tests are added for the freeze scopes and adoption rules, the provision "running before" rule, drbd event parsing and the probe, event refresh coalescing and pacing, peer change detection, listing order and natural comparison, and config announce deduplication. Each change was also exercised on the dev2 3-node cluster:

  • carbonci scope test, 3 of 3 runs;
  • parallel daemon restarts;
  • freeze adoption with a node down;
  • scsi reservation cycles;
  • drbd peer down and up;
  • provision on running, standby and down instances.

…ted, and when a peer writes a state file of it
…objective, and mark those instances in om mon and the status tree
…ance, for wait_syncs_timeout at most, and schedule no sync during an orchestration
… and wait for it as long as an orchestrated stop
… recovering it when it is removed

Since the daemon announces a configuration file it wrote without waiting
for the filesystem watcher, the watcher announced the same write again,
200ms later. On the node writing a configuration whose scope left it, the
first announce ended the instance config manager, and the second started
another one, which published the configuration for its peers a second
time. A peer that had installed it on the first one and started its own
manager took the second as a sign its file was foreign, and set the flag
disabling the recovery of the file. The flag stayed, and a later removal
of the file was not recovered from the peers, for some of the objects,
depending on the timing.

The announces of the daemon now go through configannounce, which records
the modification time each file was announced at. The watcher consumes
that record instead of announcing a write already announced: one announce
hides one watcher event, at that modification time only, and a removal
drops the record, so a file put back later is announced.

A peer's configuration for the local node no longer disables the recovery
of the local file either: it is fetched, not removed, and only a foreign
configuration, local out of its scope or a peer's not for this node, ends
in a removal the recovery must not undo.
…ode frozen since now, which froze the peers restarting with it

The node monitor started with the node frozen since now, until discover
scanned the frozen flag file, so the node was not seen unfrozen before the
flag was known. That frozen date was published to the peers. A peer
rejoining takes a node frozen after it went down for an operator freeze it
missed, and freezes itself: a peer whose daemon restarted along with this
one, as with "om daemon restart --node='*'" or a rolling deploy, froze on
a freeze that never was, and the next nodes to rejoin froze on its freeze,
real this time.

The frozen state is now read from the flag file when the node monitor
starts, as the daemon data does, so the node is neither seen unfrozen
while frozen, nor frozen while it is not.
…ng back from down

A daemon joining the cluster froze the node when any peer node had been
frozen while it was down, and froze an ha instance when any peer instance
had. A node frozen alone for its maintenance froze its peers as they
rebooted, and a freeze the daemon took on its own, at the end of a rejoin
grace period, spread from node to node through their restarts. OpenSVC v2
does the same, and its users are annoyed by the freezes nobody asked for.

A frozen flag now records the scope of the freeze that raised it. The node
monitor freezes the node for a freeze of the cluster with
OSVC_FREEZE_SCOPE=cluster in the environment of the "node freeze" it forks,
and the flag says "cluster". The instance monitor freezes the instance for
a freeze of the object, and the flag says "object". Every other freeze
leaves the flag empty, as the flags raised before this change are: a
freeze of the node or of the instance alone.

The node and instance statuses publish it as frozen_scope, "cluster" or
"node", "object" or "instance", and a node coming back adopts only a freeze
of the cluster, an instance only a freeze of the object, taken by a peer
while it was down. A freeze adopted is one of the node or of the instance,
and is not adopted again from there.

The freeze of a node reached by a global expect logged that the node was
not frozen: it says it is.
… the disk status is n/a, so a node standing by does not warn

The registrations a device must have depend on whether the instance holds
it, which the status of the resource the reservation belongs to tells: up,
every path must be registered, down, none may be. A resource n/a, like a
disk that is there whatever the state of the instance, tells nothing, and
was taken as up, so every node standing by for the instance warned of
"0/2 registrations".

A reservation held with our key now witnesses that the instance holds the
device, and every path must be registered. Not held, the registrations
are logged and not judged, and the resource is down as the reservation
is, as the reservation resource of v2 was on a node standing by.
…er, and flag nothing stopped for the stop ending it

A provision without --leader and without --disable-rollback ended with a
stop of the selection, run as the stop a user asks for: the instance was
flagged stopped on purpose, so the daemon would not start it on its own,
a resource provisioned with --rid was flagged stopped (X), so the daemon
would not restart it, and a resource provisioned with --rid on a running
instance was stopped there.

The stop ending a provision is now a step of it: the action properties
name the action a step belongs to, and a step raises neither the instance
nor the resource stopped flag. It stops only the resources the provision
found down, read from a status evaluated before it, and nothing on an
instance found running. Running is judged on the resources provisioned
already: a resource being added is down and would make a running instance
read warn, while a node standing by with a disk up reads warn too.

The rollback stack is not replayed instead, as v2 does: it holds the undo
of provisioning steps too, which a provision that succeeded must keep.
…hanges, as a scsi reservation holder or a drbd peer does

Some resource statuses read a state the actions of the peer instances
change: the key holding a scsi persistent reservation, which a peer takes
or drops as it starts or stops, the disk states of the drbd peers, which
a peer changes as it takes its drbd resource down or up, the pair state and
personality of a srdf device group, which a failover swaps. Nothing local
tells, and the nodes standing by kept reading "reserved by <key>" after the
holder had stopped, until the next scheduled status.

A driver says so with the optional StatusDependsOnPeers interface, which
disk.drbd and sync.symsrdfs implement, and the core answers it for every
resource with scsireserv on, whatever its driver. The resource status
publishes it as depends_on_peers. The instance monitor compares each peer
instance status with the one before, and refreshes the local status 2s
later when a resource depending on the peers had its peer counterpart
change status, or the peer instance change availability. Only status
values trigger it, not log lines nor timestamps, so two nodes depending on
each other stop once their statuses settle.

The refreshes asked by a peer event, this one and the state file a sync
source writes, get a timer of their own: they were armed on the outdated
timer, which every local status update arms again or stops, and a local
update landing within the delay dropped them.
…reshes the node and peer events ask for

A drbd resource changes state with no action of the node to tell: its
peers connect once it is up, a resync ends, a peer takes its disk down. The
status read at the end of an action saw the state of that moment, and a
node bringing its drbd resource up kept reading its peers DUnknown until
the next scheduled status. A resync length can not be guessed to read the
status again at the right time.

The drbdmon daemon component reads "drbdsetup events2", as mntmon reads the
mount table, and publishes DrbdResourceUpdated for a drbd resource whose
role, disk, peer disk, connection or replication state changed, once its
changes settle. The instance monitor refreshes the status of an instance
with a disk.drbd resource of that name.

The monitor runs where the drbd event stream is usable: drbd 9, on every
distribution down to el7. It disables itself, and says so once, where
drbdsetup is not installed, where "drbdsetup events2 --now" fails, as with
a drbd that does not know the stream, and where the stream ends right
after it starts three times in a row. A node where the drbd kernel module
is not loaded waits for it.

The monitors report one change of the node from several sources at nearly
the same time, a drbd promotion and the mount that follows it, and each
asked for a status refresh at once. Every event-driven refresh, of a mount,
an address, a drbd state, a peer instance change and a peer state file,
now goes through one scheduler: the first event schedules a refresh a
second later, which the events arriving meanwhile join, and the refresh is
not due sooner than 5 seconds after the last status evaluation, whatever
asked for it.
…he command chose an order, and count digit runs as numbers

A listing assembled from maps, and from the answers of several nodes, came
in a different order at each run: "om vol resource ls" listed the same
resources shuffled from one call to the next. Only the exec, session and
orchestration listings chose an order.

A listing whose command sets no default sort, and whose caller asks for
none, is now ordered on its columns, left to right, in json as in the
table. A column this default order can not read, or that no item carries,
is passed over, as the caller did not name it.

The texts compare the way a reader counts: their digit runs compare as the
numbers they write, so disk#2 comes before disk#10, and n2 before n10,
whether the order is the default one or asked with --sort.

The heartbeat listing, whose first columns are state icons, orders on its
node, stream and peer, as its own sort did, now declared as its default
sort, which also puts hb#2 before hb#10.
…us or a monitor yet

The configuration of an instance is known before its status and its
monitor are, as right after the daemon starts, or while its monitor is
still starting. The resource listing read the resources of both without
checking, and a nil monitor crashed the handler: "om vol resource ls"
failed with an http2 INTERNAL_ERROR.

An instance without a status has no resource to list yet and is passed
over, and the monitor of a resource is added when there is one.
@cvaroqui
cvaroqui merged commit 71ede28 into opensvc:main Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant