/opt/so/saltstack/default holds the source for every root-executed script --
/usr/sbin, the reactors, _runners/_modules/_beacons, the master engines and
salt-relay.sh -- plus every state the root master renders. SOC mounts
/opt/so/saltstack rw as uid 939, so root-owning /usr/sbin alone was not
enough: the next highstate would copy attacker-controlled bytes out of the
tree into the root-owned destination and run them.
SOC never writes under default/, it only reads it. Every SOC write targets
local/, which stays socore-owned, as does /opt/so/state. No mode is enforced
on default/ -- SOC reads that tree, and 750/640 would break its config load.
Also stops copy_new_files(), so-saltstack-update and setup from chowning the
tree back to socore, and replaces preserve: True in soup_scripts.sls, which
carried uid/gid in from the /tmp staging tree and would have undone the
ownership before the first post-soup highstate.
Salt installed the so-* scripts into /usr/sbin owned by unprivileged service
UIDs (939/socore, plus 930-960 per service) at mode 755, while root executes
those same files from cron, systemd and state cmd.run. Any file-write
primitive as one of those UIDs was therefore root.
Two mechanisms behind this are not visible in the diff:
file.recurse also manages the destination directory, so /usr/sbin itself was
chowned to whichever service UID ran last. A directory's owner may always
chmod it, so that UID could replace even the scripts already declared
user: root -- so-config-backup, so-suricata-eve-clean, so-nsm-mount-nvme.
usr_sbin_perms now pins the directory to root:root 555, the mode the
filesystem RPM ships.
Omitting user:/group: is a no-op on files that already exist, because
check_perms only chowns when a user is named. Explicit user: root is what
lets upgraded grids self-heal on the next highstate, and what makes a
revert chown back rather than silently do nothing.
setup runs so-minion -o=setup before /usr/sbin/so-common is installed, so
valid_ip4 was undefined, every MAINIP was rejected, and no minion pillar
was written. Pillar compile then failed for the new manager and setup gave
up waiting for the salt master. Use an inline IPv4 regex instead.
so-minion exported every line of the minion-controlled /opt/so/install.txt
into its root shell and wrote the values unescaped into a Jinja-rendered
pillar, allowing a rogue node to redefine PILLARFILE or run code on the
master at the next pillar compile. pcapspace also fed minion-returned
disk.usage output into bash arithmetic, which evaluates array subscripts.
- Parse install.txt against an allowlist of known keys; never export
- Validate MINION_ID before building pillar paths; make them readonly
- Validate node type, IP, interface, hostname, heap and core values
before any pillar is written; strip braces and control chars from
the free-text node description
- Require a numeric disk size before pcapspace arithmetic
- Refuse manager node types on add/addVM so a remote node cannot
rewrite the CA pillar; only setup may create them
The cleanup loop had no progress check or pass limit, so once /nsm was over
threshold with nothing left to reclaim it spun at full speed, writing 1.1 GB /
12.5M lines to sensor_clean.log in five hours. Stop when a pass removes
nothing, when a pass frees no space, or at MAX_PASSES, and drop the per-pass
"no old files" logging in favor of one actionable line.
Replace the pgrep guard with flock -n. A find|while read subshell inherits the
parent's argv, so pgrep -cf counted one instance as hundreds; it was also
check-then-act, which let cron stack up overlapping runs.
Paths now derive from SENSOR_DIR with env-overridable LOG/LOCK so the
over-threshold path can be tested against a scratch filesystem.
Upstream Elastic now ships binaries built for the x86-64-v3
micro-architecture level. This is not a Security Onion choice: nodes whose
CPUs predate x86-64-v3 can no longer run Elastic's own builds, so those
nodes break once they are upgraded.
Check for support before soup modifies anything, and require the operator
to type "override" to proceed when a node is unsupported or offline.
Runs after upgrade_check so a grid that is already current exits without
prompting. Targets only the roles that run a container built from the
so-elastic-agent image, using the role lists already maintained in
salt/reactor/pillar_push_map.yaml. Adds --skip-cpu-check to bypass the gate
for automation, and exit code 162 when the operator declines to override.
zeekctl's post-terminate archives the final logs in the background and returns
immediately unless StopWait is set, so the container exits and takes the archiving
with it, stranding unarchived logs in /nsm/zeek/spool/tmp on every restart. Docker's
default 10s grace is also too tight for the entrypoint's SIGTERM trap; overrunning it
means SIGKILL and crash directories on the next start.
Both are needed. StopWait alone gives the stop more work to do inside the same 10s
window, which was measured ending in SIGKILL with logs stranded in the spool.
disabled.sls used docker rm -f, which never delivers SIGTERM, so stop the container
before removing it.
Reported in discussion #16174.