Skip to content

Podman rootless - #1131

Merged
cvaroqui merged 40 commits into
opensvc:mainfrom
cvaroqui:podman-rootless
Sep 27, 2026
Merged

cvaroqui merged 40 commits into
opensvc:mainfrom
cvaroqui:podman-rootless

Conversation

@cvaroqui

Copy link
Copy Markdown
Member

No description provided.

"om <path> container logs" failed on every podman container with
'no container with name or ID "--tail" found': the name came before the
options, and podman stops reading options at the first argument that is
not one, so --tail and its value were read as two more container names.
The default of 100 lines passes --tail, so the command failed whatever
was asked of it.

Docker reads options wherever they are, which is how the order went
unnoticed. Put the name last, which both engines read.
A container run by an unprivileged user is placed by the systemd
instance of that user, in the subtree systemd delegates to them,
user@<uid>.service. A group om makes at the root of the hierarchy is
not above it, so nothing om writes there caps it.

A group can now be delegated: it keeps its ID, which orders the groups
and scopes a reset, and lives under the delegated subtree. The manager
keys the groups by where they live, so a group and its delegated copy
are made and capped on their own, and Ancestors returns the groups one
is nested in, for a delegated group to carry the cappings of its object
and namespace along.

A delegated group is made as the kernel's delegation model asks: the
directory and the files delegating it, cgroup.procs,
cgroup.subtree_control and cgroup.threads, are given to the user, so
their systemd can make the group of the container inside it, and the
files holding the cappings stay root's, so the user cannot lift them.
The subtree must exist, its absence meaning the systemd instance of the
user is not running, and is said so. A capping whose controller the
subtree was not delegated, like cpuset and io under a default
user@.service, is reported as not applied rather than failing the
others.
…less_user names

A podman container was always run by the agent, so by root, in root's
store. rootless_user names an unprivileged user to run it as instead:
every podman command of the container runs as that user, the container
lives in their store and their user namespace, and its processes are
mapped to their subordinate ids. rootless_group overrides their primary
group.

Every command, and not some: an engine keeps the containers of a user in
a store of that user, so a run as the user and an inspect as root see
two stores, and the container reads as down. The executor asks its
argser for a credential, the optional ExecutorCredentialer, and runs the
inspect, run, start, stop, pull, exec and logs commands as it, and
podman's own wait, with the home of the user as working directory and
the environment podman finds their store and runtime directory in.

What the node lacks for it is refused by name, before podman fails at a
boot with an error naming neither: the systemd instance of the user
must run without a login session, which "loginctl enable-linger"
makes it do, and the user needs subordinate ids in /etc/subuid and
/etc/subgid. The status warns, and the start refuses.

The pg keywords cap the container where podman places it: the groups of
the resource, of its object and of its namespace are made under the
subtree systemd delegates to the user, and --cgroup-parent names the
slice the systemd of the user then places the container in.

The resolver of the container is written under the runtime directory of
the user, which podman mounts it from as the user, rather than in the
var dir, which only root can enter.
The pg keywords of a docker or podman task capped nothing. A run is not
a start, so the action applied no group before it, and the engine placed
the container in a group made for it by name through --cgroup-parent,
uncapped: a task with pg_mem_limit = 64m ran with memory.max at max.

The oci task base now applies the groups before it starts the
container, through the driver rather than itself, so a driver placing
its groups elsewhere, like a rootless podman task, is the one asked.
…names

A podman task runs its command in a container of the podman container
driver, so it takes the rootless_user and rootless_group keywords and
hands them to that container: every podman command of the task runs as
the user, and the groups of the task are placed in the subtree systemd
delegates to them.

The status of the task says what keeps it from running rootless on the
node, as the status of a container does: the container driver exposes
RootlessIssues for it.
A rootless container runs its root as its user and its other ids as the
subordinate ids of that user, so a file its uid 101 reads has to be
owned, on the host, by an id only the container can say: the first
subordinate id of the user plus 100, on this node, whose /etc/subuid
may say otherwise than another's. An install could only name it by
hand, which breaks when the file changes or the container stops being
rootless.

{<rid>.uid.<id>} and {<rid>.gid.<id>} answer it, {<rid>.uid} and
{<rid>.gid} being the root of the container:

    install = /html/index.html from ./cfg/web key index.html mode 0640 user {container#1.uid.101} group {container#1.gid.101}

They are answered by the resource, through resource.IDMapper, as
{<rid>.capacity} is. A podman container or task maps its root to its
user, or to rootless_group for the gids, and its ids from 1 to the
subordinate ranges of the user, one after the other, as podman hands
them to newuidmap. A rootful one maps its ids to themselves, so an
install naming its owners by reference holds either way. A userns
mapping is refused: podman makes it when the container starts, and an
id computed before that would be a guess files get chowned by.

Only a resource id is taken for a resource: a key named uid, as
{env.uid}, is still the key.
"om <path> config show" split a key line at its first '#' or ';', and
drew the rest as an inline comment. The parser reads a comment only
after a space, so the '#' of a resource id is part of the value, and a
value naming one was drawn wrong: in "user {container#1.uid.101} group
{container#1.gid.101}", the first reference was left unclosed and not
drawn as one, and everything from its '#' was drawn as a single comment.

The value is now cut where the parser cuts it: at a '#' after a space,
or else at a ';' after a space.
The filesystem migration took any size keyword of a filesystem for the
om2 one, the size of a logical volume the filesystem made, and proposed
to turn it into a disk.lv resource. A quota-capped directory has a size
of its own in om3, and so does a tmpfs now, and a configuration holding
one would have had it moved to a volume nobody asked for.

om2 made a volume when the filesystem named a volume group, and carved
it at the size the filesystem gave. That pair, size and vg, is now the
condition. A filesystem with a size and no volume group made no volume
in om2 either, so it is left alone rather than read its group from its
device path.
…nt_opt to them

A resize grows a tmpfs by remounting it, and records the size it reached
in the size keyword of the resource. A tmpfs had none: its size was the
size option of mnt_opt, which a shm pool wrote at the size the volume
was made with, so the next mount of a resized volume brought it back to
that size.

fs.tmpfs takes size and mode keywords, which the mount is given in place
of the size and mode options of mnt_opt, the status saying those are
ignored. The size is a leaf, provisioning keyword like every size.

pool.shm takes a mode keyword, and writes the tmpfs of a volume with
size = {DEFAULT.size}, the size the volume is claimed with, where a
resize of the volume records, and mode = the pool's mode, or else the
mode= of its mnt_opt, 700 by default. Its other mount options are kept.

The configuration migration moves the size and mode options of a tmpfs
mnt_opt to the keywords: a volume's to {DEFAULT.size}, saying so when it
was mounted at another size than it is claimed with, a service's to its
number, node-scoped options to node-scoped keywords. A share of the
memory, size=50%, is kept where it is, and why is said.
An empty [node] section appeared in object configurations, as many as
five volumes of the lab: every filesystem resource carries prkey, whose
default is {node.node.prkey}, and building the resources dereferenced
it.

The dereference asked for the type of the node section in the
configuration of the object rather than the node's, and that read made
the section: the ini file makes a section, and a key, on the first
access to one it does not have, and Get, GetStrict, HasKey and Keys
went through that access. The next write of the configuration committed
what the read made, a resize recording the size it reached being the
one that did it here. Get of a missing key made the key too, which a
write would commit as an empty value.

The type of a node section is now read in the configuration of the
node, and the readers look sections and keys up without making them.
…the object

A configuration written through the api is acknowledged by the node that
wrote it, and reaches the other nodes of the object a moment later. A
caller going on to act on the instances had no way to know when that
moment was.

"om|ox <path> config update --wait [--time]", and the wait parameter of
PATCH /api/object/path/{namespace}/{kind}/{name}/config, hold the answer
until the configuration written has reached every live node of the
object: the nodes of the scope of the configuration written, a node it
adds included, whose data the daemon holds. A node whose heartbeats all
went stale is dropped from that data, fetches the configuration when it
comes back, and is not waited for.

A node has the configuration when the one it holds carries the
timestamp of the write, or a later one: a peer keeps the timestamp of
the configuration it fetches. The wait looks again on every event of an
instance configuration of the object, subscribed to before the write,
so it answers as the last node reports rather than at a poll.

A wait that expires is answered 408, naming the nodes the configuration
has not reached. The write is made all the same, and the answer carries
the OM-Last-Modified timestamp an action can require: writing again is
not the answer to it, as a += operation would apply twice.
…ng for the watcher

The filesystem watcher debounces the events of a file for 200ms, which it
has to for the writes it cannot tell apart, an editor saving in several
steps. A configuration written by the daemon paid it twice on its way to
the other nodes: on the node that wrote it, whose peers hear of it once
this node publishes the instance configuration it read back, and on each
peer, which installs the configuration it fetched by moving a file into
place, and reports it once it has read it back. A configuration update
waited for with --wait took some 750ms, where the transfers took 250ms.

The daemon now publishes ConfigFileUpdated as soon as it has written a
configuration file itself: an api update of a configuration, a
configuration file posted or put, a key of a cfg or sec object added,
changed or removed, and a configuration a peer installs. The watcher
still speaks, 200ms later, and finds the file already read at that
modification time, which does nothing.

An update waited for now answers in some 250ms.
A group an engine places a container in through systemd is a slice unit,
and systemd writes the cgroup files of a unit from its properties whenever
it starts it. A cap written in the files alone was put back to what the
properties said, no cap at all, the next time a container started in the
group: a rootful pg_cpu_quota was lost that way, and so were the caps of a
rootless task.

ApplyProc now also sets the caps as runtime properties of the slice, with
systemctl set-property --runtime, on the system manager, or on the manager
of the user a delegated group belongs to. Runtime properties are enough:
the actions make the groups again after a reboot, from the configuration.
… a rootless user

The groups of the namespace and of the node were copied in the tree of each
user running a rootless container, with their caps: each user tree was
given the whole budget again, which caps nothing the namespace consumes.

Only the groups of the object and its subsets are copied now. What a
namespace consumes wherever its containers run is for its claims to count.
… objects

A namespace configuration can now claim the compute of the cluster, the way
it claims the space of a pool or the addresses of a network:

	[claim#1]
	type = cpu
	limit = 800%
	default = 50%

	[claim#2]
	type = memory
	limit = 16g
	default = 512m

An object claims what its processes are capped to, on every instance it may
run at once: one started instance for a failover object, its flex target
for a flex one, counted on the nodes where they take the most, plus the
standby resources of the other instances. A process capped by nothing
makes the claim unbounded. The claims are published in the instance
configuration.

A container, or a podman, docker or oci task, of a namespace claiming a
default is given it when neither it, its subset nor its object says a cap.

A configuration write growing the claims of an object is weighed by the
node speaking for the cluster, through POST /api/compute/claim, with grants
covering the configuration propagation, and refused past the limit. An
instance start running more instances than the claims count is refused to
a non-root requester, and let through and logged for root.
Three pg keywords, on objects, resources and namespaces:

* pg_mem_high writes memory.high: past it the kernel reclaims and
  throttles the group rather than killing it. A memory claim still counts
  pg_mem_limit, the bound of the group.

* pg_pids_max writes pids.max, so a runaway fork loop fails in the group
  instead of exhausting the process table of the node. The pids
  controller is enabled on the groups, rootful and delegated.

* pg_cpu_burst writes cpu.max.burst, in the pg_cpu_quota notation, after
  the quota it cannot exceed.

The first two are also set as the MemoryHigh and TasksMax properties of
the slice. systemd has no property for the burst, and neither keeps nor
resets the file, so it is written at every apply. Each accepts "default",
and pg reset lifts them. memory.high and cpu.max.burst are of the unified
hierarchy only, and are said to be ignored on a v1 node.

The keyoprbac comment saying a pg keyword takes nothing from anybody is
corrected: raising a cap takes from a namespace claiming the compute, which
the claim check refuses past its limit.
Comment thread drivers/rescontainerocibase/credential.go
Comment thread drivers/rescontainerocibase/credential.go
Comment thread drivers/rescontainerpodman/rootless.go
Comment thread util/pg/linux.go
Comment thread util/pg/linux.go
Comment thread daemon/daemonapi/post_compute_claim.go
…out of it

The agent installs the keys of configs and secrets, and the directories of
the install keyword, as root in the head of a volume, which the containers
mounting the volume write in too. Every step followed links: a link a
container planted where a file is installed, or in place of a directory on
the way, led the write, the chmod and the chown of the agent onto a file of
the node. Pointed at /etc/shadow, the next install handed it to the
container's uid.

The install now runs its file operations through util/confined, a tree
rooted at the head with os.Root, which refuses a path, a link or a ".."
leading out of it, for each operation. A link where a key is installed is
removed and the file written in its place, and a link in place of a
directory is refused as one. The owner is changed with lchown, so a link is
never followed to its target.

An install naming no head, the daemon certificates and the key cache, writes
where the agent alone writes, and is left as it was.
…rant

The perm and dirperm keywords of a volume resource are converted with the
setuid and setgid bits, and the keys of configs and secrets are installed
with them. A namespace administrator could install a file of its choosing
setuid root, which runs as root for whoever executes it, on the node where
the volume is mounted or in a container mounting it.

A perm or dirperm value with either bit now needs the root grant, and so
does an install text naming such a mode. The mode words of an install text
are read as permission bits alone today, so the bits are not applied there,
but a mode that would be applied once the grammar honours them is not one to
let a user holding no root grant write now. The sticky bit is left open.
The rule refusing a host path in volume_mounts to a user holding no root
grant tested for a source beginning with an underscore, where a host path
begins with a slash, as the keyword documents it, as the driver reads it,
and as v2 read it. A namespace administrator could mount /etc, or any path
of the node, in a container, and the tests encoded the same underscore.

The rule now tests for the slash, and the tests name volume sources the way
a configuration does, a volume name, and host paths the way it does, a slash.
A volume_mounts source naming a volume was joined under the head of the
volume and handed to the engine as is. The containers mounting the volume
write in it, and a link one of them planted where the source is named, or
on the way to it, was followed by the engine when it mounted the source: a
link to / mounted the whole node in the next container started. A missing
source was made with a mkdir following the same links.

The source is now resolved when the container starts, refused unless it
stays in the head, and handed to the engine resolved. A missing source is
made in the head through util/confined, which follows no link out of it. A
volume with no head, whose sources were joined under /, has no mount source.

A link swapped in between the check and the mount of the engine is still
followed: only a mount through a file descriptor closes that, which the
engines do not offer.
… without following the user's links

The resolver of a rootless container is written under the runtime directory
of its user, which the user writes in too. The agent made the directories
on the way and chowned them to the user, then wrote the file, as root and
following links: a link the user planted in place of a directory, to /etc,
had the agent chown /etc to them, and a link in place of the file had it
write a file of the node.

The writes now run through util/confined, a tree rooted at the runtime
directory: a link leading out of it is refused, a link in place of a
directory is refused as one, the owner is changed with lchown, and a link in
place of the file is replaced. The executor hook is now the write itself,
WriteResolvConf, rather than the directory to write in, so the driver
running the engine as a user is the one writing where the user writes.
…e squatter writes

A namespace administrator could set rootless_user and rootless_group to any
account of the node, and the daemon ran every podman command of the
container as it: the namespace reached everything that account owns, the
rootless containers of other namespaces included, and a group like docker
or disk reached the node itself.

The namespace configuration now lists the accounts its objects may run
rootless containers as, with rootless_users and rootless_groups, which only
the squatter writes. An allowed user allows its primary group, and the root
account and group are never allowed. A write of a user holding no root grant
naming another account is refused, judged as the keywords evaluate, through
references and on every node of the object, and only where the write changes
the account. The ids are compared, and a node judging a write without the
account judges the name the list gives.

Root is not bound by the list, and a configuration does not say who wrote
it, so the container is not refused at runtime: its status warns that its
namespace does not allow its account. An account allowed in two namespaces
lets each reach the containers of the other: it is allowed, and warned
about, in the daemon log when a namespace configuration is written and in
the status of every container running as it.
The file mode converter read the leading digit of a 4 digit mode through a
table of its own: 2 as setuid, 4 as setgid, 6 as setgid and sticky. chmod,
and v2, which read these keywords with int(perm, 8), take 4 as setuid, 2 as
setgid and 1 as sticky, bits that add up. A tmpfs meant 2775 was mounted
4775.

The mode is now parsed as one octal number, and its 04000, 02000 and 01000
bits set the setuid, setgid and sticky flags os.FileMode keeps apart. It
feeds the perm and dirperm keywords of the data receivers, the mode of a
tmpfs and of the shm pool, and the perm of a raw disk: a value led by 2, 4,
5 or 6 now means what chmod says it means.
Comment thread util/confined/main.go
The test announcing a configuration file stopped its bus with Stop, which
drains the bus in a goroutine logging for a while after the test returned.
The next test of the package set the logger up meanwhile, and the race
detector of the PR tests reported the write of the time format against the
read of the drain.

The bus is now stopped by its context, and its goroutines waited for, so
nothing of it runs after the test.
…solves

The podman driver refused a rootless_user resolving to uid 0 but took any
group: rootless_group=root, or a user whose primary group is root, ran the
engine and the root of the container as gid 0, with the group rights of root
on the node. The driver now refuses gid 0 for every rootless container,
whoever wrote the configuration.

The namespace allow list judged an account the writing node does not
resolve by its name, and so let a listed "root" group through, which the
node running the container resolves to gid 0. The root group is refused
there too, by name, by gid, and by what the node resolves.
…ants on a speaker change, and count no @ALL cap

Two compute claims answered at once could both be weighed against the same
room before either was recorded, and the result of the recording was not
read: the namespace could be granted more than its limit. The weighing and
the recording now run under one lock, and a recording refused is answered
as a refusal.

The grants of the node speaking for the cluster live in memory, and a node
taking the speaking over answered compute claims without the grants its
peers had seen written, as pool and network claims did before their rebuild.
The rebuild now asks each peer what its own objects claim of the compute.

A cpu cap of @ALL was counted with the cpus of the node weighing the claim,
where the instance runs on another, which may have more. @ALL is not a
quantity a claim can count, so it bounds nothing in the claim: an object
capped by nothing else is unbounded, which a claimed namespace refuses.
…hat the agent opened

A volume_mounts source naming a volume was checked by the agent and handed
to the engine as a path in the volume, which the engine walked again when it
mounted it. The containers mounting the volume write in it, so a link one of
them swapped in between the check and the mount still led the engine out of
the volume: a link to / mounted the node in the next container.

The agent now opens the source through a tree confined to the head of the
volume, which follows no link out of it, and bind mounts what it opened,
through /proc/self/fd, onto a staging path under /run/opensvc-mnt, which
only root writes and a rootless engine traverses. The kernel follows that
link to the file opened, not to what the path leads to by then, and the
engine is handed the staging path, which holds no link to follow. It uses
bind mounts and openat alone, which the kernels of RHEL 7 have, where the
open_tree and move_mount api needs 5.2.

The staging mounts are made at every start, kept while the container exists
for an engine restarting it, and removed when it stops. A container an agent
without staging mounts started has none, and stops as before.

A kept container such an agent made mounts its sources by their path in the
volume, which its engine would walk again: its start is refused, with the
command removing it, since removing it loses what it wrote outside its
volumes. The status of a container mounting its sources that way, running or
not, warns about it, and says how to fix it.
A container kept after a stop, rm=false, survived an unprovision and a
purge: the driver unprovisioned nothing, and the container was left on the
node with what it wrote outside its volumes, for a start to find again.

The stop of an unprovision now removes a kept container, and its staging
mounts. The next start makes it again from the configuration. A resource
whose unprovision is disabled keeps its container.

It is done in the stop of the unprovision rather than as an unprovisioner,
which would have the container report a provisioned state where it reports
none.
…a restart

Changing a cap took a configuration update, then a pg update asked of each
node running an instance, once the configuration had landed there.

POST /api/object/path/{namespace}/{kind}/{name}/action/pg/set does both. It
writes the pg_* keywords of the sections it is given, DEFAULT, a subset or a
resource, through the same write as a configuration update, so the rbac
policy, the validation and the claim checks weigh it, and the caps hold
through a start, a failover and a reboot. It waits for the configuration to
reach every live node of the object, and asks each node running an instance
to apply the configuration of that write, which a node still holding an
older one refuses. It answers what each node did: applying, skipped as not
running, or failed.

The commands are the ones of the instances and of the resources:

	om <path> instance pg set [--rid <expr> | --subset <name>] --cpu-quota 80%
	om <path> container|app|task pg set [PATTERN]... --mem-limit 512m
	om <path> resource pg set RID... --pids-max 512

and the same of ox. A pattern naming no resource is refused, not a set of
nothing. The app, container and task command sets declare the subsystems
section their pg set is listed in.

It also sorts the imports of core/datarecv, left unsorted by the install
confinement.
…tarted

A trigger was run without the context of its action, so the timeout of the
action never reached it: a post provision script whose apt waited on a
connection the mirror had closed held the provision, and the orchestration
waiting on it, for hours past the one hour timeout.

The trigger now runs with the context of its action, in a process group of
its own, which command.WithProcessGroup kills as a whole on cancel: the
commands a trigger script started die with it, rather than outliving it
holding its output.

The trigger warnings also print their exit code, which a %s printed as
%!s(int=...).
… the object

The wait query parameter of a configuration update was not taken by the PUT
and POST of a whole configuration file, so a client writing one had no way
to know when it had landed on the nodes of the object.

Both now take it, with the meaning and the answers of the update: held
until every live node of the object holds the configuration written, or
answered 408 naming the nodes still behind, the write standing either way.
The file is written by the node the request reached, which need not be a
node of the object, so the nodes waited for are the ones of the
configuration written rather than the ones this node holds an instance
configuration of.

An invalid file was refused with a nil error for its reason. It now says
which keywords are wrong.
pg set wrote the caps of an object and applied them to its running
instances, and was offered as an instance and a resource command: it acted
on every node where its siblings act on the ones --node selects, and
answered a report of its own where they answer an exec.

The caps are now set by an orchestration, as a resize is:

	om <path> cap --set pg_mem_limit=4g --set container#1.pg_cpu_quota=80%
	POST /api/object/path/{namespace}/{kind}/{name}/action/cap?set=...

The set operations are the ones of a configuration update restricted to the
pg_* keywords of the object, a subset or a resource, with the = operator: a
cap holds one value, and a cap opens no keyword a configuration update does
not. The write goes through the configuration update, so the rbac policy,
the validation and the claim checks weigh it. The capped orchestration then
has every node wait for the configuration written and apply it where the
instance runs; an instance not running applies the caps at start, and a
failure is final on the instance it failed on. Without a set operation, the
caps the configuration holds are applied again.

The help of the command lists every pg_* keyword of the store with its
example and the first sentence of its text, so a keyword added later is
documented and completed without the command changing, and the first
sentences of the pg_* texts now say what each cap is.

instance pg set, the pg set of the driver groups and of the resources, and
the pg/set endpoint are removed. instance pg update and reset stay.

The options of a global expect were decoded as a generic value on the peers
receiving them, so the options of a resize never read as their type there,
and a peer did not wait for the configuration of the resize to land. The
options of resized and capped keep their type through the decoding and the
copy of the instance monitor.
…none is found

"config doc --kw pg_cpu_burst" printed nothing: a bare word reached the
keyword filter as a section with an empty option, and the filter had no
case for an option named without a section anyway, so it fell through.
The help of "cap" points at the keyword pages, which made the gap show.

A --kw value is now tried as the lookups doc.KeywordQueries returns, in
order, until one finds keywords: "<section>.<option>" as before, and a
bare word first as an option in any section, then as a section. With a
driver, a bare word is only an option. The five config doc commands, om
and ox, object, node and cluster, share this through doc.FindKeywords.

A value naming no keyword is an error, "keyword <kw> not found", rather
than an empty documentation that reads as a keyword with no text.

A keyword documented for no section in particular carries no rbac line:
the grant depends on the section, and the line shown was the one of the
empty section, which the policy reads as an unknown driver group
requiring root.
"cap --set pg_cpu_quota=50%@x" was written, validated and replicated,
and failed only when applied, on every node running an instance. The
validation never looked at a value it had no candidates for, and no
converter tells a cpu quota or a cpuset apart.

A keyword can now declare a Validate function, which the configuration
validation calls on the value of each line, raising an "unsupported
value" error when it refuses it. The value checked is the one of the
line, evaluated as the node its scope names, and not the one the key
evaluates to on the local node: a line scoped to a peer is what that
peer will use. A value naming something not built yet is left to be
checked when it is.

The pg_* keywords, of objects and of namespaces, declare the Valid*
functions of util/pg, which accept what the apply accepts:

  pg_cpus, pg_mems             a list of numbers and ranges, 0-2,4
  pg_cpu_quota, pg_cpu_burst   50%, 50%@ALL, 10%@2
  pg_cpu_shares, pg_*mem_*     a size
  pg_pids_max                  a count of processes, 1 or more
  pg_blkio_weight              1 to 10000
  pg_mem_swappiness            0 to 100
  pg_mem_oom_control           a flag

All of them accept "default", except pg_mem_swappiness and
pg_mem_oom_control, which are of the v1 hierarchy the reset is not
written for, and whose apply fails on it.

CPUQuota.Convert now parses through the parse the validation uses, and
so also refuses a negative percentage and a count of cpus below one,
which it used to write to cpu.max.
…eued

A cap refused because another orchestration was in progress had already
written its caps: the write came first, and the refusal after. The
caller was told the cap failed, and the next orchestration applied it.

The configuration update is now split in its checks,
prepareConfigUpdate, which runs the rbac policy, the validation and the
claim checks without writing, and its commit. The cap handler checks,
queues the orchestration, and commits only once it is accepted. The
orchestration names a configuration as recent as the request, less a
margin for the coarse clock files are stamped with, so the nodes wait for
the write that follows. A commit failing after the orchestration was
accepted aborts it, since it would wait for good. A cap changing nothing
writes nothing, and names the configuration in place.

A configuration the validation refuses is now answered 400, rather than
500, by the config update, resize and cap handlers, through the
ErrInvalidConfig the update returns.
…ainst

A configuration write is checked, by the rbac policy, the validation and
the claims, against the configuration read before it, and lands as a
whole file renamed over the one on disk. A file changed in between was
replaced by one checked against its predecessor: a keyword the other
write changed, as root changing a rootless_user, was put back without
the grant that change needs being asked of anyone. A cap holds its
write until its orchestration is queued, which made the window wide
enough to hit, and every api config writer had it, narrower.

A write now reads the file it is checked against, its base, before any
check reads it, and commits through configBase.commit, which takes a
daemon-wide mutex, compares the file with the base, and writes only if
they are the same. A file written, created or removed meanwhile refuses
the write with ErrConfigChanged, answered 409, for the caller to ask
again against the file as it now is. A cap whose orchestration was
already queued aborts it. An update changing nothing writes nothing.

This covers the config update, resize and cap handlers, and the PUT and
POST of an object configuration file and the PUT of the cluster one.
The PATCH of an object configuration declares its 409, and config
update and config migrate say what it holds.

The mutex is of the daemon: it does not serialize a write the daemon
api makes with one another process makes, as a local om or an action,
which are left the moment between the comparison and the rename.
A daemon api write compared the configuration file with the one it was
checked against, and renamed its own over it, under a mutex of the
daemon. The other writers took nothing: a local om, an action, and the
replication of a peer's file could land between the comparison and the
rename, and be replaced by a file checked against their predecessor.

Every installation of a configuration file is now made by xconfig,
under a flock(2) lock of the file: both write paths of xconfig.T, and
the replication, through xconfig.InstallFile. A write given a Base, the
file it was checked against, compares it with the file inside the lock,
and fails with ErrConfigChanged if another write landed. A write refused
keeps its base, so asking it again is refused again. The daemon api
hands its base to xconfig, and its mutex is gone.

The lock is lock.Exclusive, new in util/lock. It is held by an open
file description, so it excludes the goroutines of the daemon as it
excludes the other processes, where a fcntl lock belongs to the process
and lets every goroutine in. The kernel releases it when its holder
dies, so a crash leaves nothing to clean up. Its file is never removed:
removing it on release lets a waiter that opened it before and a
newcomer that creates it after both hold the lock. Solaris has no flock,
and falls back to a mutex of the process.

The lock files are kept out of the directory the daemon watches, under
var/lock/config, at the path of the configuration file in etc.

A local om write takes the lock and gives no base, so it can still
replace a daemon api write that landed after it read the file: setting a
base on every load is left for when the writers doing so through two
configurations are audited.
…h a client of no timeout

"cluster enroll --wait" failed after 30s, while the enrolled node was
still draining and restarting into the cluster:

  node qad12c23n2 was added to the cluster nodes but its heartbeat is
  not beating yet: context deadline exceeded (Client.Timeout or context
  cancellation while reading body)

The events were read with the client of the enroll request, whose
default timeout of 30s bounds the whole answer, the body read included,
so it cut the stream long before --timeout. cluster evict --wait,
cluster join, cluster leave and ox create --wait read theirs the same
way.

Each now reads its stream with a client of its own, with no timeout, and
bounded by the context of the command. The other requests keep the
default timeout.
…first of the nodename

A Debian node whose hosts file does not name it resolves its nodename,
through the myhostname nss module, to every address it has, link-local
ones first. The unicast tx bound every connection to the first one:

  set local ip to fe80::2023:24ff:fe17:111
  dial tcp: address [fe80::2023:24ff:fe17:111]:0: no suitable address found
  dial tcp [fe80::...:111]:0->[fe80::...:112%enp2s0]:10000: bind: invalid argument

so the peer never heard the node, and a join never completed.

The tx now keeps the usable addresses of the nodename, dropping the
link-local, loopback and unspecified ones: a peer knows the node by none
of these, and a link-local one can't be bound without its zone. It
resolves each peer, and dials its addresses a local address shares a
subnet with first, from that address, which is the one the kernel routes
the peer through and the one the peer rx expects the node from on that
network. It then dials the other peer addresses, from a local address of
their family, or from where the kernel routes them when none has it. The
routes are tried in turn within the dial timeout, so a peer of several
addresses keeps the fallback the dialer gave a name. An address
configured for the heartbeat stays the source of every connection.
Comment thread core/xconfig/install.go
// the write is checked. A file missing is a Base too: the write lands only if
// nobody created the file meanwhile.
func ReadBase(p string) (Base, error) {
data, err := os.ReadFile(p)

@aikido-pr-checks aikido-pr-checks Bot Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.

Show fix
Suggested change
data, err := os.ReadFile(p)
for _, seg := range strings.Split(filepath.ToSlash(p), "/") {
if seg == ".." {
return Base{}, fmt.Errorf("invalid file path")
}
}
data, err := os.ReadFile(p)

Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info

Comment thread util/lock/exclusive.go
// that opened it before the removal and a newcomer that creates it after both
// hold the lock, on two files of the same name.
func Exclusive(ctx context.Context, path string) (func(), error) {
f, err := os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0600)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.

Show fix

Remediation: Ignore this issue only after you've verified or sanitized the input going into this function.

Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info

Comment thread util/lock/exclusive.go
if err := os.MkdirAll(filepath.Dir(path), 0700); err != nil {
return nil, err
}
f, err = os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0600)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.

Show fix

Remediation: Ignore this issue only after you've verified or sanitized the input going into this function.

Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info

…not the first of the nodename"

This reverts commit e38f8d7.

The unicast heartbeat failed on the Debian QA nodes because their names
resolve inconsistently: no hosts file entry, and systemd-resolved
answering the single-label names over LLMNR with the address of whatever
link replied. Choosing the source address per peer fixed one side of it,
and the receivers still disagreed on the addresses they expect. The QA
environment is to resolve the node names consistently instead.
@cvaroqui
cvaroqui merged commit 47008f7 into opensvc:main Sep 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant