Latest upstream stable for el9. All four NVRs are already carried by the SO
prod repo, so no repo change is needed -- a so-repo-sync refresh is enough.
Tested on a managersearch and a heavynode (OL 9.8, 3.4.0), upgrading from
29.2.1/2.2.1 both by hand and through the state itself:
- The 29.8.1 RPM ships a byte-identical docker.service, so the full ExecStart
override in files/iptables-disabled.conf still resolves correctly and the
hand-written DOCKER/DOCKER-ISOLATION/DOCKER-USER chains came back
byte-identical on both nodes across upgrade, restart and reboot.
- update_holds re-pinned the versionlock from the old NVRs to the new ones
without intervention, so soup's path needs no change.
- docker-py 7.1.0 still creates sobridge and soauth (forced by removing both);
bridges keep their configured kernel names rather than br-<hash>.
- 29.6 changed how dynamic port allocation treats
net.ipv4.ip_local_reserved_ports; Strelka's 57314 is both published and
reserved, and docker-proxy still owns it with no bind errors.
- docker ps --format json gained a HealthStatus key. Additive, so so-status,
so-log-check and so-docker-prune all still parse it.
- containerd 2.3.6 ships the same config.toml, and it is %config(noreplace)
and unmodified on disk, so no .rpmnew and disabled_plugins=["cri"] survives.
- Zeek/Suricata/Strelka pipeline verified end-to-end with so-test: 111k packets
replayed, 0 capture loss, file extraction and ES ingest all landed.
The manifest unknown exclusion in so-log-check still fires on 29.8.1 -- it
comes from a tag lookup during the registry-to-registry image copy, not from
the 29.2.1 upgrade the old comment blamed -- so only the comment changes.
These accounts existed only for container UID mapping and filesystem
ownership, but user.present omitted shell:, so Salt fell through to the
platform useradd default and every one of them got /bin/bash. Pin them to
/sbin/nologin so none can be used as an interactive login or `su -` target.
socore keeps /bin/bash: `su socore -c '/usr/sbin/so-repo-sync'` in soup and
so-kernel-upgrade execs the account's passwd shell, and operator docs tell
users to su to socore. soqemussh keeps /bin/bash as an SSH login account.
elastic-agent, elastic-agent-pr and kafka are included alongside the accounts
named in the issue, being the same class with the same unset shell, so the
default is uniform.
Cron is unaffected: cronie runs jobs via the crontab SHELL (default /bin/sh),
not the passwd shell. suricata is the only account changed here that owns a
crontab, and somon has shipped as nologin with a working cron job already.
The zeek `runuser -l zeek` calls all run inside so-zeek via docker.run/exec,
so they resolve the shell from the image, not the host.
Verified on a 3.4.0 managersearch + sensor grid: highstate converges with the
shell as the only change and no failures, is idempotent on a second run, all
containers stay up, SOC still issues a Kratos login flow, and the suricata
surilogcompress cron job runs post-change ((suricata) CMD/CMDEND in
/var/log/cron) while `su - suricata` is now refused.
Closes#16256
so-telegraf mounted /var/run/docker.sock and joined the host docker group on
every node type. The :ro flag blocks write() to the inode, not connect() plus
HTTP over the socket, so any code execution inside the container could reach
POST /containers/create with Privileged:true and become root on the host.
No telegraf script used the socket; the only consumer was the native
[[inputs.docker]] plugin, and group_add 920 existed solely to feed it.
Container metrics now come from so-container-stats, a collector that runs on
the host from cron and writes influx line protocol to a file telegraf already
had mounted. This is the pattern so-status, so-raid-status and
so-elasticagent-status already use, so the privileged docker access stays on
the host side where root cron already ran it.
The collector runs as somon, a service account in the docker group with no
login shell and a locked password, rather than root. Docker group membership
is still root-equivalent on the host, so this is defense in depth rather than
a privilege boundary.
The telegraf scripts were root:939 mode 770, letting socore rewrite them for
code execution inside the container; they are now 750, which still allows the
read and execute telegraf needs. The container also ran with no group, giving
it gid 0, and now runs as 939:939. That alone would have broken
lasthighstate.sh, which reached /opt/so/log/salt only via the root group and
could not tell an unreadable file from a missing one, so it silently reported
a 56 year highstate age. Only the lasthighstate file is bind mounted now, and
the script tests readability instead of existence.
Everything inputs.docker collected beyond the five fields the shipped
dashboards query is available per stat under telegraf:container_stats,
annotated for SOC so an operator can enable it without editing files. Defaults
reproduce the previous output exactly. With every stat enabled the emitted
field set matches what inputs.docker wrote, verified by running the plugin
against the live socket and diffing: 53 fields, no type mismatches, no field
present on one side only.
Two deliberate differences: max_usage carries the real cgroup peak where the
daemon reports 0 on cgroup v2, and host-network containers emit no
docker_container_net row, matching inputs.docker. Docker label tags are not
restored, since nothing queries them and they cost significant cardinality.
so-status, the influxdb size cron, so-elasticagent-status, so-raid-status and
so-common-status-check truncated their output in place while telegraf read it,
so telegraf periodically saw an empty file and logged a parse error or emitted
empty values. They now write aside and rename. Measured on a live manager, the
old so-status cron left status.log empty for 215 of 10997 reads.
Tested on a fresh install, a converted grid and a 3.0 upgrade.
Adding a user failed with "Password does not meet the minimum requirements"
for passwords that were well over the eight character minimum.
so-user reads the password from stdin in updatePassword, but verifyEnvironment
runs first and calls kratosCurl, which ran docker exec with -i. That attaches
stdin and drains the pipe the caller sent the password on, so the later
read -rs saw EOF, password was empty, and expr length "" tripped the minimum
length check.
No kratosCurl or hydraCurl call sends a body on stdin; every one passes it as
a -d argument, so -i was never needed. Dropping it leaves the admin API
responses unchanged.
so-client had the same wrapper for the Hydra admin API. It does not currently
read stdin, but the flag drains its caller's pipe just the same, so it is
dropped there too.
sobridge carries most containers but held no annotation, so it was absent from
the config tree entirely. Its pillar entry is an empty map, so annotating it
adds a setting with no children for docker.networks.soauth.range to collide
with, and the tree still builds.
docker.map.jinja fills in sobridge's range and gateway at render time from
docker.range and docker.gateway, and resets the entry when it is not a mapping,
so nothing writes keys underneath it.
The sentence about sobridge is dropped from the soauth range description now
that sobridge documents itself.
SOC reported "Malformed config setting (docker.networks.soauth.range): Setting
name 'networks' conflicts with another similarly named setting" and refused to
render the config tree.
soc_docker.yaml annotated docker.networks itself, and every key in that block
was a scalar, so FlattenAnnotations registered docker.networks as a setting.
defaults.yaml holds a nested map there, so FlattenPillar separately registered
docker.networks.soauth.range, .gateway and .manager_only, and
HydrateAnnotations unions the two sets. The config tree gives each id segment
either a value or children, never both, so addToNode pushed docker.networks as
a childless leaf and then threw when docker.networks.soauth.range tried to
descend through it.
The annotations move down to the three leaves that actually carry values, so
docker.networks is a plain branch. This matches docker.containers, which is
nested the same way and has never carried annotations of its own.
docker.ulimits and the per-container networks list are unaffected: their pillar
values are lists rather than maps, so they stay single settings.
The description was written alongside the Kratos change and named only the
Kratos admin API, but so-hydra now sits on soauth too.
The trailing note about rebuilding the firewall is dropped, since the setting
is readonly and cannot be changed from SOC in the first place.