Commit Graph
18834 Commits
Author SHA1 Message Date
defensivedepth 506918326d cleanup 2026-10-07 06:28:52 -04:00
defensivedepth 4244f72d95 Merge remote-tracking branch 'origin/3/dev' into esql-correlations 2026-10-06 06:52:53 -04:00
defensivedepth 06ffaf1373 Cleanup 2026-10-06 06:49:27 -04:00
defensivedepth ed58fb7429 Simplify check 2026-10-05 18:08:16 -04:00
Josh Patterson f7dbfba178 Merge pull request #16297 from Security-Onion-Solutions/revert-16274-fix/auto-apply-state-queue
Revert "Fix/auto apply state queue"
2026-10-05 17:54:16 -04:00
Josh Patterson 0a628bb7e7 Revert "Fix/auto apply state queue" 2026-10-05 17:43:08 -04:00
Josh Patterson bd6647e775 Merge pull request #16274 from Security-Onion-Solutions/fix/auto-apply-state-queue
Fix/auto apply state queue
2026-10-05 14:33:02 -04:00
Josh Brower da2c19188a Merge pull request #16287 from Security-Onion-Solutions/sigma-pipeline-dir
Move Sigma pipelines into a managed directory
2026-10-05 09:41:48 -04:00
defensivedepth 90b3d37be6 Clarify req 2026-10-05 08:14:13 -04:00
defensivedepth fcd2f67076 Merge remote-tracking branch 'origin/3/dev' into sigma-pipeline-dir 2026-10-05 07:44:53 -04:00
defensivedepth 50a44bda62 Merge branch 'sigma-pipeline-dir' into esql-correlations 2026-10-02 14:30:03 -04:00
defensivedepth a9cdd17694 Refactor sigma pipelines 2026-10-02 14:23:57 -04:00
defensivedepth f1954a9d86 Merge remote-tracking branch 'origin/3/dev' into esql-correlations 2026-10-02 13:49:30 -04:00
defensivedepth 3c62ef5e16 Cleanup 2026-10-02 13:47:03 -04:00
Josh Brower b32aaac290 Merge pull request #16278 from Security-Onion-Solutions/evtx-import-datastreams
EVTX Import cleanup
2026-10-02 13:30:57 -04:00
Josh Brower 1aee3f28dc Merge pull request #16279 from Security-Onion-Solutions/process-caseless-mappings
caseless for non-Defend sources
2026-10-02 13:30:45 -04:00
Josh Patterson 678cb0d5b2 Merge pull request #16286 from Security-Onion-Solutions/fix/docker-29.8.1
upgrade docker 29.8.1 and containerd 2.3.6
2026-10-02 12:05:04 -04:00
Josh Patterson 31c5190a1f add missing comma 2026-10-02 12:03:04 -04:00
Josh Patterson 4ce7a06abe upgrade docker to 29.8.1 and containerd.io to 2.3.6
Latest upstream stable for el9. All four NVRs are already carried by the SO
prod repo, so no repo change is needed -- a so-repo-sync refresh is enough.

Tested on a managersearch and a heavynode (OL 9.8, 3.4.0), upgrading from
29.2.1/2.2.1 both by hand and through the state itself:

- The 29.8.1 RPM ships a byte-identical docker.service, so the full ExecStart
  override in files/iptables-disabled.conf still resolves correctly and the
  hand-written DOCKER/DOCKER-ISOLATION/DOCKER-USER chains came back
  byte-identical on both nodes across upgrade, restart and reboot.
- update_holds re-pinned the versionlock from the old NVRs to the new ones
  without intervention, so soup's path needs no change.
- docker-py 7.1.0 still creates sobridge and soauth (forced by removing both);
  bridges keep their configured kernel names rather than br-<hash>.
- 29.6 changed how dynamic port allocation treats
  net.ipv4.ip_local_reserved_ports; Strelka's 57314 is both published and
  reserved, and docker-proxy still owns it with no bind errors.
- docker ps --format json gained a HealthStatus key. Additive, so so-status,
  so-log-check and so-docker-prune all still parse it.
- containerd 2.3.6 ships the same config.toml, and it is %config(noreplace)
  and unmodified on disk, so no .rpmnew and disabled_plugins=["cri"] survives.
- Zeek/Suricata/Strelka pipeline verified end-to-end with so-test: 111k packets
  replayed, 0 capture loss, file extraction and ES ingest all landed.

The manifest unknown exclusion in so-log-check still fires on 29.8.1 -- it
comes from a tag lookup during the registry-to-registry image copy, not from
the 29.2.1 upgrade the old comment blamed -- so only the comment changes.
2026-10-02 12:03:04 -04:00
Josh Patterson ba95b9bbc2 Address review feedback on so-push-drainer result tracking
Result checks walked dispatch records oldest-first with a cap of five
lookups per pass, counting records whose push was still running. Five
long-running pushes therefore used every slot on every 15s pass and newer,
finished pushes were not reported until one cleared. Check the least
recently checked records first and back off on pushes that are still
running (30s for the first two minutes, then age/4 up to 5 minutes),
recording checked_at in the dispatch record.

Catch any exception when writing a dispatch record so a failed write
cannot skip intent cleanup and re-dispatch the same intents every pass.
Log both output streams when no jid is found, and stop logging a traceback
when a record has already been removed.

Scope the test's salt mock to the drainer import. Run from the repo root,
'salt' resolves to this repo's salt/ directory as a namespace package, so
setdefault left it in place and test_load_push_cfg failed.

Verified on a 3.4.0 managersearch + sensor: a pushed highstate with soc
and telegraf pushes dispatched into it all reported success, with 25
result lookups across the three pushes instead of one per record per pass.
2026-10-02 10:35:30 -04:00
defensivedepth c9053dd564 Dont import correlation rules without esql 2026-10-02 07:41:10 -04:00
defensivedepth 43475452b3 set module 2026-10-01 19:17:52 -04:00
defensivedepth 99322cf26a Add additional mapping 2026-10-01 15:24:52 -04:00
coreyogburn 117548757f Merge pull request #16280 from Security-Onion-Solutions/cogburn/unified-automations
Unified Automations
2026-10-01 10:29:45 -06:00
Corey Ogburn 22bda63847 Unified Automations
Remove the template and mark automations as advanced, readonly, and stored in the DB.
2026-10-01 09:59:12 -06:00
defensivedepth 3d4f53b741 Additional ESQL tweaks 2026-10-01 11:32:42 -04:00
defensivedepth 2a4611df45 Add caseless mappings 2026-10-01 10:48:55 -04:00
defensivedepth 89f8bcd19f evtx-import fixup 2026-10-01 10:28:10 -04:00
Jason Ertel 563269cbac Merge pull request #16277 from Security-Onion-Solutions/jertel/wip
support empty yaml files
2026-10-01 10:10:55 -04:00
Jason Ertel 523c39d4f2 fix flake 2026-10-01 10:09:20 -04:00
Jason Ertel b4557e973c support empty yaml files 2026-10-01 10:03:39 -04:00
Jorge Reyes 0f53a7e0bc Merge pull request #16275 from Security-Onion-Solutions/reyesj2-521
review integration-defaults weird_integrations mappings, removed unus…
2026-10-01 08:36:57 -05:00
Josh Patterson 8de8ba811a Harden so-push-drainer result parsing
_orch_failures assumed every level of a jobs.lookup_jid result was a dict.
A list or string at the top level, in return.data, in a step's changes, or
in changes.ret raised AttributeError. Because result checks run before the
drain and a record is only removed after it is evaluated, one such record
would have failed every 15s pass and stopped all pushes until it was removed
by hand. Guard each shape, and evaluate each record under its own exception
handler so an unreadable result is logged and dropped instead of blocking
the drainer. Per-step parsing moves to _step_failures.

Search stdout as well as stderr for the async jid, in case salt-run logging
is routed to stdout.

Close the RotatingFileHandler in test_make_logger_adds_handler_once to
avoid a ResourceWarning on Python 3.12+.

Verified on a 3.4.0 standalone: real failed and successful orchestration
results parse as before, a record whose evaluation raises is logged and
removed while the next record still reports, and a replicated SOC change
to telegraf.output (and its revert) is pushed, rendered and logged as
succeeded.
2026-10-01 09:06:28 -04:00
reyesj2 d122ee7fea review integration-defaults weird_integrations mappings, removed unused, updated logstash integration naming 2026-09-30 16:49:19 -05:00
Josh Patterson 9732e1c639 Trim tracebacks in push failure log lines
When an orchestration step raises, salt returns the full traceback as the
step comment, and the drainer wrote it verbatim, putting ~70 lines into
so-push-drainer.log per failure. Collapse comments to one line and, for
tracebacks, keep only the lead-in and the raised exception, e.g.
"apply_soc_1: An exception occurred in this state:
salt.exceptions.AuthenticationError: Authentication error occurred."

Seen on a standalone when a pushed highstate restarted salt-master while
two queued pushes were waiting: their orchestrations lost the master
connection and failed with AuthenticationError, although the minion
completed both state runs.
2026-09-30 16:18:37 -04:00
Josh Patterson 53f9ebcd46 FIX: queue auto-applied state runs instead of failing on conflict
orch.push_batch passed `kwarg: {queue: 2}` to salt.state, but in Salt
3006 queue is a top-level salt.state argument and salt.state always sets
the minion's queue kwarg from it (default False), so the kwarg block was
silently dropped and every pushed state ran with queue=False. The drainer
dispatches a separate async orchestration each 15s pass, so settings saved
more than ~15s apart overlap on the same minion and every run after the
first fails immediately with 'The function "state.sls" is running as PID
...'. The change then waits for the next scheduled highstate.

Seen on a 3.4.0 standalone: hydra.enabled, telegraf.output, and two soc
settings (including soc.config.licenseKey) were saved within 30s. The soc
state was dispatched while the telegraf state was still running and was
rejected, so the license key was not applied.

Use `queue: True`, as orch.deploy_newnode already does. An int is treated
as max_queue and still falls through to the conflict error once that many
state runs are active.

The failure was only visible in the master log, since the drainer
dispatches with --async and logged only "dispatch accepted". The drainer
now:
  - logs each dispatched action
  - parses the orchestration jid from salt-run's stderr (the only place
    --async reports it) and records it under /opt/so/state/push_dispatched
  - on later passes looks each jid up with jobs.lookup_jid and logs either
    "push succeeded" or an ERROR with the failed step, the per-minion
    failed states or rejection text, and the triggering paths
Lookups run outside the pending-intent lock since the reactors share it.

The beacon now logs each audit_settings row it emits and the reactor logs
the audit row id, so a single change can be traced from audit_settings to
its push result.

Adds so-push-drainer_test.py; the drainer is now held to the 100% coverage
requirement in python-test.

Verified on the standalone: a soc push dispatched while a 90s state run
was in progress queued behind it (queue=True in the job args), completed,
and the drainer logged "push succeeded" for its jid. The new result
parsing reports the original soc conflict and the hydra license failure
from the job cache.
2026-09-30 16:18:37 -04:00
coreyogburn 47d74f1ae1 Merge pull request #16270 from Security-Onion-Solutions/cogburn/automation
New Automation Fields
2026-09-30 11:10:51 -06:00
Corey Ogburn 855716846a New Automation Fields 2026-09-29 16:50:04 -06:00
defensivedepth 98ffb6fa00 Initial Correlations support 2026-09-29 17:05:20 -04:00
Mike Reeves 8e35d70595 Merge pull request #16266 from Security-Onion-Solutions/mreeves/soai-context-1m
Raise SOAI Sonnet default small context limit to 1M
2026-09-29 12:19:20 -04:00
Jason Ertel e4625cfcae Merge pull request #16267 from Security-Onion-Solutions/jertel/wip
resolve startup errors
2026-09-29 12:07:26 -04:00
Jason Ertel eb803dce0e resolve startup errors 2026-09-29 12:02:48 -04:00
Mike Reeves a06f08217a Raise SOAI Sonnet default small context limit to 1M
Context is now flat-priced, so match contextLimitSmall to contextLimitLarge.
With equal limits the SOC assistant hides the increase-context toggle.
2026-09-29 11:42:29 -04:00
Josh Patterson 29d27cf255 Merge pull request #16261 from Security-Onion-Solutions/fix/telegraf-drop-docker-socket
FIX: remove the docker socket from so-telegraf
2026-09-29 10:46:54 -04:00
Jason Ertel 235a60e587 Merge pull request #16264 from Security-Onion-Solutions/jertel/wip
Metric alarms and more NTF annotations
2026-09-29 08:31:28 -04:00
Jason Ertel a8f7c46b0d Merge branch '3/dev' into jertel/wip 2026-09-28 13:47:19 -04:00
Jason Ertel 26d895ccb7 alarms and ntf 2026-09-28 13:47:16 -04:00
Josh Patterson b43efc458f Merge pull request #16263 from Security-Onion-Solutions/fix/service-account-nologin
FIX: use /sbin/nologin for service accounts
2026-09-28 13:05:26 -04:00
Josh Patterson 21222ff119 FIX: use /sbin/nologin for service accounts
These accounts existed only for container UID mapping and filesystem
ownership, but user.present omitted shell:, so Salt fell through to the
platform useradd default and every one of them got /bin/bash. Pin them to
/sbin/nologin so none can be used as an interactive login or `su -` target.

socore keeps /bin/bash: `su socore -c '/usr/sbin/so-repo-sync'` in soup and
so-kernel-upgrade execs the account's passwd shell, and operator docs tell
users to su to socore. soqemussh keeps /bin/bash as an SSH login account.

elastic-agent, elastic-agent-pr and kafka are included alongside the accounts
named in the issue, being the same class with the same unset shell, so the
default is uniform.

Cron is unaffected: cronie runs jobs via the crontab SHELL (default /bin/sh),
not the passwd shell. suricata is the only account changed here that owns a
crontab, and somon has shipped as nologin with a working cron job already.
The zeek `runuser -l zeek` calls all run inside so-zeek via docker.run/exec,
so they resolve the shell from the image, not the host.

Verified on a 3.4.0 managersearch + sensor grid: highstate converges with the
shell as the only change and no failures, is idempotent on a second run, all
containers stay up, SOC still issues a Kratos login flow, and the suricata
surilogcompress cron job runs post-change ((suricata) CMD/CMDEND in
/var/log/cron) while `su - suricata` is now refused.

Closes #16256
2026-09-25 09:23:09 -04:00
Mike Reeves 88fa7e7fb4 Merge pull request #16262 from Security-Onion-Solutions/TOoSmOotH-patch-4
Add openai_embeddings to the YAML configuration
2026-09-24 15:29:17 -04:00