Commit Graph
18801 Commits
Author SHA1 Message Date
Josh Patterson ba95b9bbc2 Address review feedback on so-push-drainer result tracking
Result checks walked dispatch records oldest-first with a cap of five
lookups per pass, counting records whose push was still running. Five
long-running pushes therefore used every slot on every 15s pass and newer,
finished pushes were not reported until one cleared. Check the least
recently checked records first and back off on pushes that are still
running (30s for the first two minutes, then age/4 up to 5 minutes),
recording checked_at in the dispatch record.

Catch any exception when writing a dispatch record so a failed write
cannot skip intent cleanup and re-dispatch the same intents every pass.
Log both output streams when no jid is found, and stop logging a traceback
when a record has already been removed.

Scope the test's salt mock to the drainer import. Run from the repo root,
'salt' resolves to this repo's salt/ directory as a namespace package, so
setdefault left it in place and test_load_push_cfg failed.

Verified on a 3.4.0 managersearch + sensor: a pushed highstate with soc
and telegraf pushes dispatched into it all reported success, with 25
result lookups across the three pushes instead of one per record per pass.
2026-10-02 10:35:30 -04:00
Josh Patterson 8de8ba811a Harden so-push-drainer result parsing
_orch_failures assumed every level of a jobs.lookup_jid result was a dict.
A list or string at the top level, in return.data, in a step's changes, or
in changes.ret raised AttributeError. Because result checks run before the
drain and a record is only removed after it is evaluated, one such record
would have failed every 15s pass and stopped all pushes until it was removed
by hand. Guard each shape, and evaluate each record under its own exception
handler so an unreadable result is logged and dropped instead of blocking
the drainer. Per-step parsing moves to _step_failures.

Search stdout as well as stderr for the async jid, in case salt-run logging
is routed to stdout.

Close the RotatingFileHandler in test_make_logger_adds_handler_once to
avoid a ResourceWarning on Python 3.12+.

Verified on a 3.4.0 standalone: real failed and successful orchestration
results parse as before, a record whose evaluation raises is logged and
removed while the next record still reports, and a replicated SOC change
to telegraf.output (and its revert) is pushed, rendered and logged as
succeeded.
2026-10-01 09:06:28 -04:00
Josh Patterson 9732e1c639 Trim tracebacks in push failure log lines
When an orchestration step raises, salt returns the full traceback as the
step comment, and the drainer wrote it verbatim, putting ~70 lines into
so-push-drainer.log per failure. Collapse comments to one line and, for
tracebacks, keep only the lead-in and the raised exception, e.g.
"apply_soc_1: An exception occurred in this state:
salt.exceptions.AuthenticationError: Authentication error occurred."

Seen on a standalone when a pushed highstate restarted salt-master while
two queued pushes were waiting: their orchestrations lost the master
connection and failed with AuthenticationError, although the minion
completed both state runs.
2026-09-30 16:18:37 -04:00
Josh Patterson 53f9ebcd46 FIX: queue auto-applied state runs instead of failing on conflict
orch.push_batch passed `kwarg: {queue: 2}` to salt.state, but in Salt
3006 queue is a top-level salt.state argument and salt.state always sets
the minion's queue kwarg from it (default False), so the kwarg block was
silently dropped and every pushed state ran with queue=False. The drainer
dispatches a separate async orchestration each 15s pass, so settings saved
more than ~15s apart overlap on the same minion and every run after the
first fails immediately with 'The function "state.sls" is running as PID
...'. The change then waits for the next scheduled highstate.

Seen on a 3.4.0 standalone: hydra.enabled, telegraf.output, and two soc
settings (including soc.config.licenseKey) were saved within 30s. The soc
state was dispatched while the telegraf state was still running and was
rejected, so the license key was not applied.

Use `queue: True`, as orch.deploy_newnode already does. An int is treated
as max_queue and still falls through to the conflict error once that many
state runs are active.

The failure was only visible in the master log, since the drainer
dispatches with --async and logged only "dispatch accepted". The drainer
now:
  - logs each dispatched action
  - parses the orchestration jid from salt-run's stderr (the only place
    --async reports it) and records it under /opt/so/state/push_dispatched
  - on later passes looks each jid up with jobs.lookup_jid and logs either
    "push succeeded" or an ERROR with the failed step, the per-minion
    failed states or rejection text, and the triggering paths
Lookups run outside the pending-intent lock since the reactors share it.

The beacon now logs each audit_settings row it emits and the reactor logs
the audit row id, so a single change can be traced from audit_settings to
its push result.

Adds so-push-drainer_test.py; the drainer is now held to the 100% coverage
requirement in python-test.

Verified on the standalone: a soc push dispatched while a 90s state run
was in progress queued behind it (queue=True in the job args), completed,
and the drainer logged "push succeeded" for its jid. The new result
parsing reports the original soc conflict and the hydra license failure
from the job cache.
2026-09-30 16:18:37 -04:00
coreyogburn 47d74f1ae1 Merge pull request #16270 from Security-Onion-Solutions/cogburn/automation
New Automation Fields
2026-09-30 11:10:51 -06:00
Corey Ogburn 855716846a New Automation Fields 2026-09-29 16:50:04 -06:00
Mike Reeves 8e35d70595 Merge pull request #16266 from Security-Onion-Solutions/mreeves/soai-context-1m
Raise SOAI Sonnet default small context limit to 1M
2026-09-29 12:19:20 -04:00
Jason Ertel e4625cfcae Merge pull request #16267 from Security-Onion-Solutions/jertel/wip
resolve startup errors
2026-09-29 12:07:26 -04:00
Jason Ertel eb803dce0e resolve startup errors 2026-09-29 12:02:48 -04:00
Mike Reeves a06f08217a Raise SOAI Sonnet default small context limit to 1M
Context is now flat-priced, so match contextLimitSmall to contextLimitLarge.
With equal limits the SOC assistant hides the increase-context toggle.
2026-09-29 11:42:29 -04:00
Josh Patterson 29d27cf255 Merge pull request #16261 from Security-Onion-Solutions/fix/telegraf-drop-docker-socket
FIX: remove the docker socket from so-telegraf
2026-09-29 10:46:54 -04:00
Jason Ertel 235a60e587 Merge pull request #16264 from Security-Onion-Solutions/jertel/wip
Metric alarms and more NTF annotations
2026-09-29 08:31:28 -04:00
Jason Ertel a8f7c46b0d Merge branch '3/dev' into jertel/wip 2026-09-28 13:47:19 -04:00
Jason Ertel 26d895ccb7 alarms and ntf 2026-09-28 13:47:16 -04:00
Josh Patterson b43efc458f Merge pull request #16263 from Security-Onion-Solutions/fix/service-account-nologin
FIX: use /sbin/nologin for service accounts
2026-09-28 13:05:26 -04:00
Josh Patterson 21222ff119 FIX: use /sbin/nologin for service accounts
These accounts existed only for container UID mapping and filesystem
ownership, but user.present omitted shell:, so Salt fell through to the
platform useradd default and every one of them got /bin/bash. Pin them to
/sbin/nologin so none can be used as an interactive login or `su -` target.

socore keeps /bin/bash: `su socore -c '/usr/sbin/so-repo-sync'` in soup and
so-kernel-upgrade execs the account's passwd shell, and operator docs tell
users to su to socore. soqemussh keeps /bin/bash as an SSH login account.

elastic-agent, elastic-agent-pr and kafka are included alongside the accounts
named in the issue, being the same class with the same unset shell, so the
default is uniform.

Cron is unaffected: cronie runs jobs via the crontab SHELL (default /bin/sh),
not the passwd shell. suricata is the only account changed here that owns a
crontab, and somon has shipped as nologin with a working cron job already.
The zeek `runuser -l zeek` calls all run inside so-zeek via docker.run/exec,
so they resolve the shell from the image, not the host.

Verified on a 3.4.0 managersearch + sensor grid: highstate converges with the
shell as the only change and no failures, is idempotent on a second run, all
containers stay up, SOC still issues a Kratos login flow, and the suricata
surilogcompress cron job runs post-change ((suricata) CMD/CMDEND in
/var/log/cron) while `su - suricata` is now refused.

Closes #16256
2026-09-25 09:23:09 -04:00
Mike Reeves 88fa7e7fb4 Merge pull request #16262 from Security-Onion-Solutions/TOoSmOotH-patch-4
Add openai_embeddings to the YAML configuration
2026-09-24 15:29:17 -04:00
Mike Reeves efe0581892 Add openai_embeddings to the YAML configuration 2026-09-24 15:27:51 -04:00
Josh Patterson 72f60fcaa9 FIX: remove the docker socket from so-telegraf
so-telegraf mounted /var/run/docker.sock and joined the host docker group on
every node type. The :ro flag blocks write() to the inode, not connect() plus
HTTP over the socket, so any code execution inside the container could reach
POST /containers/create with Privileged:true and become root on the host.

No telegraf script used the socket; the only consumer was the native
[[inputs.docker]] plugin, and group_add 920 existed solely to feed it.
Container metrics now come from so-container-stats, a collector that runs on
the host from cron and writes influx line protocol to a file telegraf already
had mounted. This is the pattern so-status, so-raid-status and
so-elasticagent-status already use, so the privileged docker access stays on
the host side where root cron already ran it.

The collector runs as somon, a service account in the docker group with no
login shell and a locked password, rather than root. Docker group membership
is still root-equivalent on the host, so this is defense in depth rather than
a privilege boundary.

The telegraf scripts were root:939 mode 770, letting socore rewrite them for
code execution inside the container; they are now 750, which still allows the
read and execute telegraf needs. The container also ran with no group, giving
it gid 0, and now runs as 939:939. That alone would have broken
lasthighstate.sh, which reached /opt/so/log/salt only via the root group and
could not tell an unreadable file from a missing one, so it silently reported
a 56 year highstate age. Only the lasthighstate file is bind mounted now, and
the script tests readability instead of existence.

Everything inputs.docker collected beyond the five fields the shipped
dashboards query is available per stat under telegraf:container_stats,
annotated for SOC so an operator can enable it without editing files. Defaults
reproduce the previous output exactly. With every stat enabled the emitted
field set matches what inputs.docker wrote, verified by running the plugin
against the live socket and diffing: 53 fields, no type mismatches, no field
present on one side only.

Two deliberate differences: max_usage carries the real cgroup peak where the
daemon reports 0 on cgroup v2, and host-network containers emit no
docker_container_net row, matching inputs.docker. Docker label tags are not
restored, since nothing queries them and they cost significant cardinality.

so-status, the influxdb size cron, so-elasticagent-status, so-raid-status and
so-common-status-check truncated their output in place while telegraf read it,
so telegraf periodically saw an empty file and logged a parse error or emitted
empty values. They now write aside and rename. Measured on a live manager, the
old so-status cron left status.log empty for 215 of 10997 reads.

Tested on a fresh install, a converted grid and a 3.0 upgrade.
2026-09-24 14:30:25 -04:00
Jorge Reyes 2684a5ca95 Merge pull request #16260 from Security-Onion-Solutions/reyesj2/16254
FIX: Fleet scripts failing when endpoints-initial is missing
2026-09-23 08:58:15 -05:00
reyesj2 bcee63bde5 remove policy precheck 2026-09-23 08:51:11 -05:00
reyesj2 6ce89eb323 jq -e exits 0 with empty input 2026-09-22 21:25:50 -05:00
reyesj2 ebab4b0d90 split between missing token and multiple enrollment tokens 2026-09-22 16:11:56 -05:00
reyesj2 db60c27da2 fail cleanly when endpoints-initial is missing or has multiple enrollment tokens 2026-09-22 12:24:23 -05:00
Jason Ertel f4defdfde0 Merge pull request #16258 from Security-Onion-Solutions/jertel/wip
upgrade whoislookup analyzer deps; notification annotations
2026-09-22 08:41:04 -04:00
Jason Ertel 36652e8f23 fixed malformed desc annotation 2026-09-22 08:30:44 -04:00
Jason Ertel 7bef194540 force transitive dep version 2026-09-22 08:28:44 -04:00
Jason Ertel e2bf2837fe notification annotations 2026-09-22 08:18:59 -04:00
Jason Ertel 06704dad22 upgrade whoislookup analyzer deps 2026-09-22 08:17:18 -04:00
Mike Reeves 47fe0758d0 Merge pull request #16255 from Security-Onion-Solutions/mreeves/fix-so-user-stdin-drain
Fix user creation failing with a valid password
2026-09-18 17:12:09 -04:00
Mike Reeves b71fd93f9d Stop docker exec from consuming so-user's piped password
Adding a user failed with "Password does not meet the minimum requirements"
for passwords that were well over the eight character minimum.

so-user reads the password from stdin in updatePassword, but verifyEnvironment
runs first and calls kratosCurl, which ran docker exec with -i. That attaches
stdin and drains the pipe the caller sent the password on, so the later
read -rs saw EOF, password was empty, and expr length "" tripped the minimum
length check.

No kratosCurl or hydraCurl call sends a body on stdin; every one passes it as
a -d argument, so -i was never needed. Dropping it leaves the admin API
responses unchanged.

so-client had the same wrapper for the Hydra admin API. It does not currently
read stdin, but the flag drains its caller's pipe just the same, so it is
dropped there too.
2026-09-18 17:04:17 -04:00
Mike Reeves d9eff9aa9e Merge pull request #16253 from Security-Onion-Solutions/mreeves/fix-docker-networks-setting-conflict
Fix config tree error from the docker.networks setting
2026-09-18 15:41:25 -04:00
Mike Reeves d8884dbd99 Document sobridge alongside the soauth network settings
sobridge carries most containers but held no annotation, so it was absent from
the config tree entirely. Its pillar entry is an empty map, so annotating it
adds a setting with no children for docker.networks.soauth.range to collide
with, and the tree still builds.

docker.map.jinja fills in sobridge's range and gateway at render time from
docker.range and docker.gateway, and resets the entry when it is not a mapping,
so nothing writes keys underneath it.

The sentence about sobridge is dropped from the soauth range description now
that sobridge documents itself.
2026-09-18 14:37:11 -04:00
Mike Reeves 6c0d4c15e8 Annotate the soauth network leaves instead of the networks map
SOC reported "Malformed config setting (docker.networks.soauth.range): Setting
name 'networks' conflicts with another similarly named setting" and refused to
render the config tree.

soc_docker.yaml annotated docker.networks itself, and every key in that block
was a scalar, so FlattenAnnotations registered docker.networks as a setting.
defaults.yaml holds a nested map there, so FlattenPillar separately registered
docker.networks.soauth.range, .gateway and .manager_only, and
HydrateAnnotations unions the two sets. The config tree gives each id segment
either a value or children, never both, so addToNode pushed docker.networks as
a childless leaf and then threw when docker.networks.soauth.range tried to
descend through it.

The annotations move down to the three leaves that actually carry values, so
docker.networks is a plain branch. This matches docker.containers, which is
nested the same way and has never carried annotations of its own.

docker.ulimits and the per-container networks list are unaffected: their pillar
values are lists rather than maps, so they stay single settings.
2026-09-18 13:42:41 -04:00
Jason Ertel 2e2f62f265 Merge pull request #16252 from Security-Onion-Solutions/jertel/wip
store in db
2026-09-18 12:47:32 -04:00
Jason Ertel b7a11a525c store in db 2026-09-18 12:41:35 -04:00
Mike Reeves 8eef95ea3e Merge pull request #16221 from Security-Onion-Solutions/mreeves/kratos-soauth-network
Isolate the Kratos admin API on a dedicated soauth docker network
2026-09-18 10:24:00 -04:00
Mike Reeves 65e261475d Mention Hydra in the soauth network description
The description was written alongside the Kratos change and named only the
Kratos admin API, but so-hydra now sits on soauth too.

The trailing note about rebuilding the firewall is dropped, since the setting
is readonly and cannot be changed from SOC in the first place.
2026-09-18 10:17:20 -04:00
Mike Reeves bbc28c88b7 Isolate the Hydra admin API on the soauth docker network
The Hydra admin API creates OAuth clients and introspects tokens without
authentication, relying on network isolation the same way Kratos does, but
so-hydra published 0.0.0.0:4445:4445 and sat on sobridge. With userland-proxy
left at its default of true, Docker ran a proxy listener on the host, so any
container reached the admin API through the manager IP or the bridge gateway
regardless of its docker network.

so-hydra now sits alone on soauth and no longer publishes 4445. 4444 is still
published, so the nginx proxy for /oauth2/token and the well-known endpoints is
unchanged. so-soc is already dual homed from the Kratos change and reaches the
admin API over soauth, so only its hostUrl moves. wait_for_hydra polls the
container address instead of the host port.

so-client reaches the admin API through docker exec, mirroring so-user. Exit
codes still propagate, so the --fail-with-body error handling is unchanged.

so-hydra is removed in up_to_3.4.0 alongside so-kratos and so-soc so the
highstate recreates it with the correct network membership. The auth network
range is already handled by set_soauth_range.
2026-09-17 16:58:49 -04:00
Jason Ertel edaacf79a7 Merge pull request #16251 from Security-Onion-Solutions/jertel/wip
notification annotations
2026-09-17 15:37:04 -04:00
Jason Ertel ecc643cd33 fix typo 2026-09-17 15:34:29 -04:00
Jason Ertel 08aaf7948e Merge branch '3/dev' into jertel/wip 2026-09-17 15:32:35 -04:00
Jason Ertel bb57545d08 new annotations for notifications 2026-09-17 15:32:21 -04:00
Josh Brower 0f7adbbecc Merge pull request #16247 from Security-Onion-Solutions/esql-fixes
Refactor for ESQL
2026-09-17 09:56:19 -04:00
Josh Patterson aeb4fe8f50 Merge pull request #16250 from Security-Onion-Solutions/fix/root-own-sbin-and-salt-tree
FIX: root-own /usr/sbin management scripts and the Salt default tree
2026-09-17 09:44:47 -04:00
defensivedepth f3aa39c5a4 Tweak name 2026-09-17 09:43:44 -04:00
Doug Burks 24077ba974 Merge pull request #16249 from Security-Onion-Solutions/dougburks/fix-eval-suricata-file-drops
FIX: suricata.fileinfo maps boolean gaps into long file.bytes.missing
2026-09-17 09:42:27 -04:00
Doug Burks b3567405f9 FIX: suricata.fileinfo maps boolean gaps into long file.bytes.missing 2026-09-17 07:57:44 -04:00
defensivedepth 1fc5bb7afa Refactor for ESQL 2026-09-17 07:56:42 -04:00
Josh Patterson 1f1d3ded41 FIX: root-own the Salt default tree so a SOC file-write cannot reach root code
/opt/so/saltstack/default holds the source for every root-executed script --
/usr/sbin, the reactors, _runners/_modules/_beacons, the master engines and
salt-relay.sh -- plus every state the root master renders. SOC mounts
/opt/so/saltstack rw as uid 939, so root-owning /usr/sbin alone was not
enough: the next highstate would copy attacker-controlled bytes out of the
tree into the root-owned destination and run them.

SOC never writes under default/, it only reads it. Every SOC write targets
local/, which stays socore-owned, as does /opt/so/state. No mode is enforced
on default/ -- SOC reads that tree, and 750/640 would break its config load.

Also stops copy_new_files(), so-saltstack-update and setup from chowning the
tree back to socore, and replaces preserve: True in soup_scripts.sls, which
carried uid/gid in from the /tmp staging tree and would have undone the
ownership before the first post-soup highstate.
2026-09-16 10:34:21 -04:00