Podman rootless - #1131
Podman rootless#1131
Conversation
"om <path> container logs" failed on every podman container with 'no container with name or ID "--tail" found': the name came before the options, and podman stops reading options at the first argument that is not one, so --tail and its value were read as two more container names. The default of 100 lines passes --tail, so the command failed whatever was asked of it. Docker reads options wherever they are, which is how the order went unnoticed. Put the name last, which both engines read.
A container run by an unprivileged user is placed by the systemd instance of that user, in the subtree systemd delegates to them, user@<uid>.service. A group om makes at the root of the hierarchy is not above it, so nothing om writes there caps it. A group can now be delegated: it keeps its ID, which orders the groups and scopes a reset, and lives under the delegated subtree. The manager keys the groups by where they live, so a group and its delegated copy are made and capped on their own, and Ancestors returns the groups one is nested in, for a delegated group to carry the cappings of its object and namespace along. A delegated group is made as the kernel's delegation model asks: the directory and the files delegating it, cgroup.procs, cgroup.subtree_control and cgroup.threads, are given to the user, so their systemd can make the group of the container inside it, and the files holding the cappings stay root's, so the user cannot lift them. The subtree must exist, its absence meaning the systemd instance of the user is not running, and is said so. A capping whose controller the subtree was not delegated, like cpuset and io under a default user@.service, is reported as not applied rather than failing the others.
…less_user names A podman container was always run by the agent, so by root, in root's store. rootless_user names an unprivileged user to run it as instead: every podman command of the container runs as that user, the container lives in their store and their user namespace, and its processes are mapped to their subordinate ids. rootless_group overrides their primary group. Every command, and not some: an engine keeps the containers of a user in a store of that user, so a run as the user and an inspect as root see two stores, and the container reads as down. The executor asks its argser for a credential, the optional ExecutorCredentialer, and runs the inspect, run, start, stop, pull, exec and logs commands as it, and podman's own wait, with the home of the user as working directory and the environment podman finds their store and runtime directory in. What the node lacks for it is refused by name, before podman fails at a boot with an error naming neither: the systemd instance of the user must run without a login session, which "loginctl enable-linger" makes it do, and the user needs subordinate ids in /etc/subuid and /etc/subgid. The status warns, and the start refuses. The pg keywords cap the container where podman places it: the groups of the resource, of its object and of its namespace are made under the subtree systemd delegates to the user, and --cgroup-parent names the slice the systemd of the user then places the container in. The resolver of the container is written under the runtime directory of the user, which podman mounts it from as the user, rather than in the var dir, which only root can enter.
The pg keywords of a docker or podman task capped nothing. A run is not a start, so the action applied no group before it, and the engine placed the container in a group made for it by name through --cgroup-parent, uncapped: a task with pg_mem_limit = 64m ran with memory.max at max. The oci task base now applies the groups before it starts the container, through the driver rather than itself, so a driver placing its groups elsewhere, like a rootless podman task, is the one asked.
…names A podman task runs its command in a container of the podman container driver, so it takes the rootless_user and rootless_group keywords and hands them to that container: every podman command of the task runs as the user, and the groups of the task are placed in the subtree systemd delegates to them. The status of the task says what keeps it from running rootless on the node, as the status of a container does: the container driver exposes RootlessIssues for it.
A rootless container runs its root as its user and its other ids as the
subordinate ids of that user, so a file its uid 101 reads has to be
owned, on the host, by an id only the container can say: the first
subordinate id of the user plus 100, on this node, whose /etc/subuid
may say otherwise than another's. An install could only name it by
hand, which breaks when the file changes or the container stops being
rootless.
{<rid>.uid.<id>} and {<rid>.gid.<id>} answer it, {<rid>.uid} and
{<rid>.gid} being the root of the container:
install = /html/index.html from ./cfg/web key index.html mode 0640 user {container#1.uid.101} group {container#1.gid.101}
They are answered by the resource, through resource.IDMapper, as
{<rid>.capacity} is. A podman container or task maps its root to its
user, or to rootless_group for the gids, and its ids from 1 to the
subordinate ranges of the user, one after the other, as podman hands
them to newuidmap. A rootful one maps its ids to themselves, so an
install naming its owners by reference holds either way. A userns
mapping is refused: podman makes it when the container starts, and an
id computed before that would be a guess files get chowned by.
Only a resource id is taken for a resource: a key named uid, as
{env.uid}, is still the key.
"om <path> config show" split a key line at its first '#' or ';', and
drew the rest as an inline comment. The parser reads a comment only
after a space, so the '#' of a resource id is part of the value, and a
value naming one was drawn wrong: in "user {container#1.uid.101} group
{container#1.gid.101}", the first reference was left unclosed and not
drawn as one, and everything from its '#' was drawn as a single comment.
The value is now cut where the parser cuts it: at a '#' after a space,
or else at a ';' after a space.
The filesystem migration took any size keyword of a filesystem for the om2 one, the size of a logical volume the filesystem made, and proposed to turn it into a disk.lv resource. A quota-capped directory has a size of its own in om3, and so does a tmpfs now, and a configuration holding one would have had it moved to a volume nobody asked for. om2 made a volume when the filesystem named a volume group, and carved it at the size the filesystem gave. That pair, size and vg, is now the condition. A filesystem with a size and no volume group made no volume in om2 either, so it is left alone rather than read its group from its device path.
…nt_opt to them
A resize grows a tmpfs by remounting it, and records the size it reached
in the size keyword of the resource. A tmpfs had none: its size was the
size option of mnt_opt, which a shm pool wrote at the size the volume
was made with, so the next mount of a resized volume brought it back to
that size.
fs.tmpfs takes size and mode keywords, which the mount is given in place
of the size and mode options of mnt_opt, the status saying those are
ignored. The size is a leaf, provisioning keyword like every size.
pool.shm takes a mode keyword, and writes the tmpfs of a volume with
size = {DEFAULT.size}, the size the volume is claimed with, where a
resize of the volume records, and mode = the pool's mode, or else the
mode= of its mnt_opt, 700 by default. Its other mount options are kept.
The configuration migration moves the size and mode options of a tmpfs
mnt_opt to the keywords: a volume's to {DEFAULT.size}, saying so when it
was mounted at another size than it is claimed with, a service's to its
number, node-scoped options to node-scoped keywords. A share of the
memory, size=50%, is kept where it is, and why is said.
An empty [node] section appeared in object configurations, as many as
five volumes of the lab: every filesystem resource carries prkey, whose
default is {node.node.prkey}, and building the resources dereferenced
it.
The dereference asked for the type of the node section in the
configuration of the object rather than the node's, and that read made
the section: the ini file makes a section, and a key, on the first
access to one it does not have, and Get, GetStrict, HasKey and Keys
went through that access. The next write of the configuration committed
what the read made, a resize recording the size it reached being the
one that did it here. Get of a missing key made the key too, which a
write would commit as an empty value.
The type of a node section is now read in the configuration of the
node, and the readers look sections and keys up without making them.
…the object
A configuration written through the api is acknowledged by the node that
wrote it, and reaches the other nodes of the object a moment later. A
caller going on to act on the instances had no way to know when that
moment was.
"om|ox <path> config update --wait [--time]", and the wait parameter of
PATCH /api/object/path/{namespace}/{kind}/{name}/config, hold the answer
until the configuration written has reached every live node of the
object: the nodes of the scope of the configuration written, a node it
adds included, whose data the daemon holds. A node whose heartbeats all
went stale is dropped from that data, fetches the configuration when it
comes back, and is not waited for.
A node has the configuration when the one it holds carries the
timestamp of the write, or a later one: a peer keeps the timestamp of
the configuration it fetches. The wait looks again on every event of an
instance configuration of the object, subscribed to before the write,
so it answers as the last node reports rather than at a poll.
A wait that expires is answered 408, naming the nodes the configuration
has not reached. The write is made all the same, and the answer carries
the OM-Last-Modified timestamp an action can require: writing again is
not the answer to it, as a += operation would apply twice.
…ng for the watcher The filesystem watcher debounces the events of a file for 200ms, which it has to for the writes it cannot tell apart, an editor saving in several steps. A configuration written by the daemon paid it twice on its way to the other nodes: on the node that wrote it, whose peers hear of it once this node publishes the instance configuration it read back, and on each peer, which installs the configuration it fetched by moving a file into place, and reports it once it has read it back. A configuration update waited for with --wait took some 750ms, where the transfers took 250ms. The daemon now publishes ConfigFileUpdated as soon as it has written a configuration file itself: an api update of a configuration, a configuration file posted or put, a key of a cfg or sec object added, changed or removed, and a configuration a peer installs. The watcher still speaks, 200ms later, and finds the file already read at that modification time, which does nothing. An update waited for now answers in some 250ms.
A group an engine places a container in through systemd is a slice unit, and systemd writes the cgroup files of a unit from its properties whenever it starts it. A cap written in the files alone was put back to what the properties said, no cap at all, the next time a container started in the group: a rootful pg_cpu_quota was lost that way, and so were the caps of a rootless task. ApplyProc now also sets the caps as runtime properties of the slice, with systemctl set-property --runtime, on the system manager, or on the manager of the user a delegated group belongs to. Runtime properties are enough: the actions make the groups again after a reboot, from the configuration.
… a rootless user The groups of the namespace and of the node were copied in the tree of each user running a rootless container, with their caps: each user tree was given the whole budget again, which caps nothing the namespace consumes. Only the groups of the object and its subsets are copied now. What a namespace consumes wherever its containers run is for its claims to count.
… objects A namespace configuration can now claim the compute of the cluster, the way it claims the space of a pool or the addresses of a network: [claim#1] type = cpu limit = 800% default = 50% [claim#2] type = memory limit = 16g default = 512m An object claims what its processes are capped to, on every instance it may run at once: one started instance for a failover object, its flex target for a flex one, counted on the nodes where they take the most, plus the standby resources of the other instances. A process capped by nothing makes the claim unbounded. The claims are published in the instance configuration. A container, or a podman, docker or oci task, of a namespace claiming a default is given it when neither it, its subset nor its object says a cap. A configuration write growing the claims of an object is weighed by the node speaking for the cluster, through POST /api/compute/claim, with grants covering the configuration propagation, and refused past the limit. An instance start running more instances than the claims count is refused to a non-root requester, and let through and logged for root.
Three pg keywords, on objects, resources and namespaces: * pg_mem_high writes memory.high: past it the kernel reclaims and throttles the group rather than killing it. A memory claim still counts pg_mem_limit, the bound of the group. * pg_pids_max writes pids.max, so a runaway fork loop fails in the group instead of exhausting the process table of the node. The pids controller is enabled on the groups, rootful and delegated. * pg_cpu_burst writes cpu.max.burst, in the pg_cpu_quota notation, after the quota it cannot exceed. The first two are also set as the MemoryHigh and TasksMax properties of the slice. systemd has no property for the burst, and neither keeps nor resets the file, so it is written at every apply. Each accepts "default", and pg reset lifts them. memory.high and cpu.max.burst are of the unified hierarchy only, and are said to be ignored on a v1 node. The keyoprbac comment saying a pg keyword takes nothing from anybody is corrected: raising a cap takes from a namespace claiming the compute, which the claim check refuses past its limit.
…out of it The agent installs the keys of configs and secrets, and the directories of the install keyword, as root in the head of a volume, which the containers mounting the volume write in too. Every step followed links: a link a container planted where a file is installed, or in place of a directory on the way, led the write, the chmod and the chown of the agent onto a file of the node. Pointed at /etc/shadow, the next install handed it to the container's uid. The install now runs its file operations through util/confined, a tree rooted at the head with os.Root, which refuses a path, a link or a ".." leading out of it, for each operation. A link where a key is installed is removed and the file written in its place, and a link in place of a directory is refused as one. The owner is changed with lchown, so a link is never followed to its target. An install naming no head, the daemon certificates and the key cache, writes where the agent alone writes, and is left as it was.
…rant The perm and dirperm keywords of a volume resource are converted with the setuid and setgid bits, and the keys of configs and secrets are installed with them. A namespace administrator could install a file of its choosing setuid root, which runs as root for whoever executes it, on the node where the volume is mounted or in a container mounting it. A perm or dirperm value with either bit now needs the root grant, and so does an install text naming such a mode. The mode words of an install text are read as permission bits alone today, so the bits are not applied there, but a mode that would be applied once the grammar honours them is not one to let a user holding no root grant write now. The sticky bit is left open.
The rule refusing a host path in volume_mounts to a user holding no root grant tested for a source beginning with an underscore, where a host path begins with a slash, as the keyword documents it, as the driver reads it, and as v2 read it. A namespace administrator could mount /etc, or any path of the node, in a container, and the tests encoded the same underscore. The rule now tests for the slash, and the tests name volume sources the way a configuration does, a volume name, and host paths the way it does, a slash.
A volume_mounts source naming a volume was joined under the head of the volume and handed to the engine as is. The containers mounting the volume write in it, and a link one of them planted where the source is named, or on the way to it, was followed by the engine when it mounted the source: a link to / mounted the whole node in the next container started. A missing source was made with a mkdir following the same links. The source is now resolved when the container starts, refused unless it stays in the head, and handed to the engine resolved. A missing source is made in the head through util/confined, which follows no link out of it. A volume with no head, whose sources were joined under /, has no mount source. A link swapped in between the check and the mount of the engine is still followed: only a mount through a file descriptor closes that, which the engines do not offer.
… without following the user's links The resolver of a rootless container is written under the runtime directory of its user, which the user writes in too. The agent made the directories on the way and chowned them to the user, then wrote the file, as root and following links: a link the user planted in place of a directory, to /etc, had the agent chown /etc to them, and a link in place of the file had it write a file of the node. The writes now run through util/confined, a tree rooted at the runtime directory: a link leading out of it is refused, a link in place of a directory is refused as one, the owner is changed with lchown, and a link in place of the file is replaced. The executor hook is now the write itself, WriteResolvConf, rather than the directory to write in, so the driver running the engine as a user is the one writing where the user writes.
…e squatter writes A namespace administrator could set rootless_user and rootless_group to any account of the node, and the daemon ran every podman command of the container as it: the namespace reached everything that account owns, the rootless containers of other namespaces included, and a group like docker or disk reached the node itself. The namespace configuration now lists the accounts its objects may run rootless containers as, with rootless_users and rootless_groups, which only the squatter writes. An allowed user allows its primary group, and the root account and group are never allowed. A write of a user holding no root grant naming another account is refused, judged as the keywords evaluate, through references and on every node of the object, and only where the write changes the account. The ids are compared, and a node judging a write without the account judges the name the list gives. Root is not bound by the list, and a configuration does not say who wrote it, so the container is not refused at runtime: its status warns that its namespace does not allow its account. An account allowed in two namespaces lets each reach the containers of the other: it is allowed, and warned about, in the daemon log when a namespace configuration is written and in the status of every container running as it.
The file mode converter read the leading digit of a 4 digit mode through a table of its own: 2 as setuid, 4 as setgid, 6 as setgid and sticky. chmod, and v2, which read these keywords with int(perm, 8), take 4 as setuid, 2 as setgid and 1 as sticky, bits that add up. A tmpfs meant 2775 was mounted 4775. The mode is now parsed as one octal number, and its 04000, 02000 and 01000 bits set the setuid, setgid and sticky flags os.FileMode keeps apart. It feeds the perm and dirperm keywords of the data receivers, the mode of a tmpfs and of the shm pool, and the perm of a raw disk: a value led by 2, 4, 5 or 6 now means what chmod says it means.
The test announcing a configuration file stopped its bus with Stop, which drains the bus in a goroutine logging for a while after the test returned. The next test of the package set the logger up meanwhile, and the race detector of the PR tests reported the write of the time format against the read of the drain. The bus is now stopped by its context, and its goroutines waited for, so nothing of it runs after the test.
…solves The podman driver refused a rootless_user resolving to uid 0 but took any group: rootless_group=root, or a user whose primary group is root, ran the engine and the root of the container as gid 0, with the group rights of root on the node. The driver now refuses gid 0 for every rootless container, whoever wrote the configuration. The namespace allow list judged an account the writing node does not resolve by its name, and so let a listed "root" group through, which the node running the container resolves to gid 0. The root group is refused there too, by name, by gid, and by what the node resolves.
…ants on a speaker change, and count no @ALL cap Two compute claims answered at once could both be weighed against the same room before either was recorded, and the result of the recording was not read: the namespace could be granted more than its limit. The weighing and the recording now run under one lock, and a recording refused is answered as a refusal. The grants of the node speaking for the cluster live in memory, and a node taking the speaking over answered compute claims without the grants its peers had seen written, as pool and network claims did before their rebuild. The rebuild now asks each peer what its own objects claim of the compute. A cpu cap of @ALL was counted with the cpus of the node weighing the claim, where the instance runs on another, which may have more. @ALL is not a quantity a claim can count, so it bounds nothing in the claim: an object capped by nothing else is unbounded, which a claimed namespace refuses.
…hat the agent opened A volume_mounts source naming a volume was checked by the agent and handed to the engine as a path in the volume, which the engine walked again when it mounted it. The containers mounting the volume write in it, so a link one of them swapped in between the check and the mount still led the engine out of the volume: a link to / mounted the node in the next container. The agent now opens the source through a tree confined to the head of the volume, which follows no link out of it, and bind mounts what it opened, through /proc/self/fd, onto a staging path under /run/opensvc-mnt, which only root writes and a rootless engine traverses. The kernel follows that link to the file opened, not to what the path leads to by then, and the engine is handed the staging path, which holds no link to follow. It uses bind mounts and openat alone, which the kernels of RHEL 7 have, where the open_tree and move_mount api needs 5.2. The staging mounts are made at every start, kept while the container exists for an engine restarting it, and removed when it stops. A container an agent without staging mounts started has none, and stops as before. A kept container such an agent made mounts its sources by their path in the volume, which its engine would walk again: its start is refused, with the command removing it, since removing it loses what it wrote outside its volumes. The status of a container mounting its sources that way, running or not, warns about it, and says how to fix it.
A container kept after a stop, rm=false, survived an unprovision and a purge: the driver unprovisioned nothing, and the container was left on the node with what it wrote outside its volumes, for a start to find again. The stop of an unprovision now removes a kept container, and its staging mounts. The next start makes it again from the configuration. A resource whose unprovision is disabled keeps its container. It is done in the stop of the unprovision rather than as an unprovisioner, which would have the container report a provisioned state where it reports none.
…a restart
Changing a cap took a configuration update, then a pg update asked of each
node running an instance, once the configuration had landed there.
POST /api/object/path/{namespace}/{kind}/{name}/action/pg/set does both. It
writes the pg_* keywords of the sections it is given, DEFAULT, a subset or a
resource, through the same write as a configuration update, so the rbac
policy, the validation and the claim checks weigh it, and the caps hold
through a start, a failover and a reboot. It waits for the configuration to
reach every live node of the object, and asks each node running an instance
to apply the configuration of that write, which a node still holding an
older one refuses. It answers what each node did: applying, skipped as not
running, or failed.
The commands are the ones of the instances and of the resources:
om <path> instance pg set [--rid <expr> | --subset <name>] --cpu-quota 80%
om <path> container|app|task pg set [PATTERN]... --mem-limit 512m
om <path> resource pg set RID... --pids-max 512
and the same of ox. A pattern naming no resource is refused, not a set of
nothing. The app, container and task command sets declare the subsystems
section their pg set is listed in.
It also sorts the imports of core/datarecv, left unsorted by the install
confinement.
…tarted A trigger was run without the context of its action, so the timeout of the action never reached it: a post provision script whose apt waited on a connection the mirror had closed held the provision, and the orchestration waiting on it, for hours past the one hour timeout. The trigger now runs with the context of its action, in a process group of its own, which command.WithProcessGroup kills as a whole on cancel: the commands a trigger script started die with it, rather than outliving it holding its output. The trigger warnings also print their exit code, which a %s printed as %!s(int=...).
… the object The wait query parameter of a configuration update was not taken by the PUT and POST of a whole configuration file, so a client writing one had no way to know when it had landed on the nodes of the object. Both now take it, with the meaning and the answers of the update: held until every live node of the object holds the configuration written, or answered 408 naming the nodes still behind, the write standing either way. The file is written by the node the request reached, which need not be a node of the object, so the nodes waited for are the ones of the configuration written rather than the ones this node holds an instance configuration of. An invalid file was refused with a nil error for its reason. It now says which keywords are wrong.
pg set wrote the caps of an object and applied them to its running
instances, and was offered as an instance and a resource command: it acted
on every node where its siblings act on the ones --node selects, and
answered a report of its own where they answer an exec.
The caps are now set by an orchestration, as a resize is:
om <path> cap --set pg_mem_limit=4g --set container#1.pg_cpu_quota=80%
POST /api/object/path/{namespace}/{kind}/{name}/action/cap?set=...
The set operations are the ones of a configuration update restricted to the
pg_* keywords of the object, a subset or a resource, with the = operator: a
cap holds one value, and a cap opens no keyword a configuration update does
not. The write goes through the configuration update, so the rbac policy,
the validation and the claim checks weigh it. The capped orchestration then
has every node wait for the configuration written and apply it where the
instance runs; an instance not running applies the caps at start, and a
failure is final on the instance it failed on. Without a set operation, the
caps the configuration holds are applied again.
The help of the command lists every pg_* keyword of the store with its
example and the first sentence of its text, so a keyword added later is
documented and completed without the command changing, and the first
sentences of the pg_* texts now say what each cap is.
instance pg set, the pg set of the driver groups and of the resources, and
the pg/set endpoint are removed. instance pg update and reset stay.
The options of a global expect were decoded as a generic value on the peers
receiving them, so the options of a resize never read as their type there,
and a peer did not wait for the configuration of the resize to land. The
options of resized and capped keep their type through the decoding and the
copy of the instance monitor.
…none is found "config doc --kw pg_cpu_burst" printed nothing: a bare word reached the keyword filter as a section with an empty option, and the filter had no case for an option named without a section anyway, so it fell through. The help of "cap" points at the keyword pages, which made the gap show. A --kw value is now tried as the lookups doc.KeywordQueries returns, in order, until one finds keywords: "<section>.<option>" as before, and a bare word first as an option in any section, then as a section. With a driver, a bare word is only an option. The five config doc commands, om and ox, object, node and cluster, share this through doc.FindKeywords. A value naming no keyword is an error, "keyword <kw> not found", rather than an empty documentation that reads as a keyword with no text. A keyword documented for no section in particular carries no rbac line: the grant depends on the section, and the line shown was the one of the empty section, which the policy reads as an unknown driver group requiring root.
"cap --set pg_cpu_quota=50%@x" was written, validated and replicated, and failed only when applied, on every node running an instance. The validation never looked at a value it had no candidates for, and no converter tells a cpu quota or a cpuset apart. A keyword can now declare a Validate function, which the configuration validation calls on the value of each line, raising an "unsupported value" error when it refuses it. The value checked is the one of the line, evaluated as the node its scope names, and not the one the key evaluates to on the local node: a line scoped to a peer is what that peer will use. A value naming something not built yet is left to be checked when it is. The pg_* keywords, of objects and of namespaces, declare the Valid* functions of util/pg, which accept what the apply accepts: pg_cpus, pg_mems a list of numbers and ranges, 0-2,4 pg_cpu_quota, pg_cpu_burst 50%, 50%@ALL, 10%@2 pg_cpu_shares, pg_*mem_* a size pg_pids_max a count of processes, 1 or more pg_blkio_weight 1 to 10000 pg_mem_swappiness 0 to 100 pg_mem_oom_control a flag All of them accept "default", except pg_mem_swappiness and pg_mem_oom_control, which are of the v1 hierarchy the reset is not written for, and whose apply fails on it. CPUQuota.Convert now parses through the parse the validation uses, and so also refuses a negative percentage and a count of cpus below one, which it used to write to cpu.max.
…eued A cap refused because another orchestration was in progress had already written its caps: the write came first, and the refusal after. The caller was told the cap failed, and the next orchestration applied it. The configuration update is now split in its checks, prepareConfigUpdate, which runs the rbac policy, the validation and the claim checks without writing, and its commit. The cap handler checks, queues the orchestration, and commits only once it is accepted. The orchestration names a configuration as recent as the request, less a margin for the coarse clock files are stamped with, so the nodes wait for the write that follows. A commit failing after the orchestration was accepted aborts it, since it would wait for good. A cap changing nothing writes nothing, and names the configuration in place. A configuration the validation refuses is now answered 400, rather than 500, by the config update, resize and cap handlers, through the ErrInvalidConfig the update returns.
…ainst A configuration write is checked, by the rbac policy, the validation and the claims, against the configuration read before it, and lands as a whole file renamed over the one on disk. A file changed in between was replaced by one checked against its predecessor: a keyword the other write changed, as root changing a rootless_user, was put back without the grant that change needs being asked of anyone. A cap holds its write until its orchestration is queued, which made the window wide enough to hit, and every api config writer had it, narrower. A write now reads the file it is checked against, its base, before any check reads it, and commits through configBase.commit, which takes a daemon-wide mutex, compares the file with the base, and writes only if they are the same. A file written, created or removed meanwhile refuses the write with ErrConfigChanged, answered 409, for the caller to ask again against the file as it now is. A cap whose orchestration was already queued aborts it. An update changing nothing writes nothing. This covers the config update, resize and cap handlers, and the PUT and POST of an object configuration file and the PUT of the cluster one. The PATCH of an object configuration declares its 409, and config update and config migrate say what it holds. The mutex is of the daemon: it does not serialize a write the daemon api makes with one another process makes, as a local om or an action, which are left the moment between the comparison and the rename.
A daemon api write compared the configuration file with the one it was checked against, and renamed its own over it, under a mutex of the daemon. The other writers took nothing: a local om, an action, and the replication of a peer's file could land between the comparison and the rename, and be replaced by a file checked against their predecessor. Every installation of a configuration file is now made by xconfig, under a flock(2) lock of the file: both write paths of xconfig.T, and the replication, through xconfig.InstallFile. A write given a Base, the file it was checked against, compares it with the file inside the lock, and fails with ErrConfigChanged if another write landed. A write refused keeps its base, so asking it again is refused again. The daemon api hands its base to xconfig, and its mutex is gone. The lock is lock.Exclusive, new in util/lock. It is held by an open file description, so it excludes the goroutines of the daemon as it excludes the other processes, where a fcntl lock belongs to the process and lets every goroutine in. The kernel releases it when its holder dies, so a crash leaves nothing to clean up. Its file is never removed: removing it on release lets a waiter that opened it before and a newcomer that creates it after both hold the lock. Solaris has no flock, and falls back to a mutex of the process. The lock files are kept out of the directory the daemon watches, under var/lock/config, at the path of the configuration file in etc. A local om write takes the lock and gives no base, so it can still replace a daemon api write that landed after it read the file: setting a base on every load is left for when the writers doing so through two configurations are audited.
…h a client of no timeout "cluster enroll --wait" failed after 30s, while the enrolled node was still draining and restarting into the cluster: node qad12c23n2 was added to the cluster nodes but its heartbeat is not beating yet: context deadline exceeded (Client.Timeout or context cancellation while reading body) The events were read with the client of the enroll request, whose default timeout of 30s bounds the whole answer, the body read included, so it cut the stream long before --timeout. cluster evict --wait, cluster join, cluster leave and ox create --wait read theirs the same way. Each now reads its stream with a client of its own, with no timeout, and bounded by the context of the command. The other requests keep the default timeout.
…first of the nodename A Debian node whose hosts file does not name it resolves its nodename, through the myhostname nss module, to every address it has, link-local ones first. The unicast tx bound every connection to the first one: set local ip to fe80::2023:24ff:fe17:111 dial tcp: address [fe80::2023:24ff:fe17:111]:0: no suitable address found dial tcp [fe80::...:111]:0->[fe80::...:112%enp2s0]:10000: bind: invalid argument so the peer never heard the node, and a join never completed. The tx now keeps the usable addresses of the nodename, dropping the link-local, loopback and unspecified ones: a peer knows the node by none of these, and a link-local one can't be bound without its zone. It resolves each peer, and dials its addresses a local address shares a subnet with first, from that address, which is the one the kernel routes the peer through and the one the peer rx expects the node from on that network. It then dials the other peer addresses, from a local address of their family, or from where the kernel routes them when none has it. The routes are tried in turn within the dial timeout, so a peer of several addresses keeps the fallback the dialer gave a name. An address configured for the heartbeat stays the source of every connection.
| // the write is checked. A file missing is a Base too: the write lands only if | ||
| // nobody created the file meanwhile. | ||
| func ReadBase(p string) (Base, error) { | ||
| data, err := os.ReadFile(p) |
There was a problem hiding this comment.
Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.
Show fix
| data, err := os.ReadFile(p) | |
| for _, seg := range strings.Split(filepath.ToSlash(p), "/") { | |
| if seg == ".." { | |
| return Base{}, fmt.Errorf("invalid file path") | |
| } | |
| } | |
| data, err := os.ReadFile(p) |
Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info
| // that opened it before the removal and a newcomer that creates it after both | ||
| // hold the lock, on two files of the same name. | ||
| func Exclusive(ctx context.Context, path string) (func(), error) { | ||
| f, err := os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0600) |
There was a problem hiding this comment.
Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.
Show fix
Remediation: Ignore this issue only after you've verified or sanitized the input going into this function.
Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info
| if err := os.MkdirAll(filepath.Dir(path), 0700); err != nil { | ||
| return nil, err | ||
| } | ||
| f, err = os.OpenFile(path, os.O_CREATE|os.O_RDWR, 0600) |
There was a problem hiding this comment.
Potential file inclusion attack via reading file - high severity
If an attacker can control the input leading into the ReadFile function, they might be able to read sensitive files and launch further attacks with that information.
Show fix
Remediation: Ignore this issue only after you've verified or sanitized the input going into this function.
Reply @AikidoSec ignore: [REASON] to ignore this issue.
More info
…not the first of the nodename" This reverts commit e38f8d7. The unicast heartbeat failed on the Debian QA nodes because their names resolve inconsistently: no hosts file entry, and systemd-resolved answering the single-label names over LLMNR with the address of whatever link replied. Choosing the source address per peer fixed one side of it, and the receivers still disagreed on the addresses they expect. The QA environment is to resolve the node names consistently instead.
No description provided.