Compare commits

...
Author SHA1 Message Date
Jorge Reyes c6e42131b2 Update version from 3.4.0 to 3.0.0-foxtrot 2026-10-06 13:45:22 -05:00
reyesj2 337dddf596 ES 9.4.8 2026-10-06 13:44:52 -05:00
Josh Patterson f7dbfba178 Merge pull request #16297 from Security-Onion-Solutions/revert-16274-fix/auto-apply-state-queue
Revert "Fix/auto apply state queue"
2026-10-05 17:54:16 -04:00
Josh Patterson 0a628bb7e7 Revert "Fix/auto apply state queue" 2026-10-05 17:43:08 -04:00
Josh Patterson bd6647e775 Merge pull request #16274 from Security-Onion-Solutions/fix/auto-apply-state-queue
Fix/auto apply state queue
2026-10-05 14:33:02 -04:00
Josh Brower da2c19188a Merge pull request #16287 from Security-Onion-Solutions/sigma-pipeline-dir
Move Sigma pipelines into a managed directory
2026-10-05 09:41:48 -04:00
Josh Patterson ba95b9bbc2 Address review feedback on so-push-drainer result tracking
Result checks walked dispatch records oldest-first with a cap of five
lookups per pass, counting records whose push was still running. Five
long-running pushes therefore used every slot on every 15s pass and newer,
finished pushes were not reported until one cleared. Check the least
recently checked records first and back off on pushes that are still
running (30s for the first two minutes, then age/4 up to 5 minutes),
recording checked_at in the dispatch record.

Catch any exception when writing a dispatch record so a failed write
cannot skip intent cleanup and re-dispatch the same intents every pass.
Log both output streams when no jid is found, and stop logging a traceback
when a record has already been removed.

Scope the test's salt mock to the drainer import. Run from the repo root,
'salt' resolves to this repo's salt/ directory as a namespace package, so
setdefault left it in place and test_load_push_cfg failed.

Verified on a 3.4.0 managersearch + sensor: a pushed highstate with soc
and telegraf pushes dispatched into it all reported success, with 25
result lookups across the three pushes instead of one per record per pass.
2026-10-02 10:35:30 -04:00
Josh Patterson 8de8ba811a Harden so-push-drainer result parsing
_orch_failures assumed every level of a jobs.lookup_jid result was a dict.
A list or string at the top level, in return.data, in a step's changes, or
in changes.ret raised AttributeError. Because result checks run before the
drain and a record is only removed after it is evaluated, one such record
would have failed every 15s pass and stopped all pushes until it was removed
by hand. Guard each shape, and evaluate each record under its own exception
handler so an unreadable result is logged and dropped instead of blocking
the drainer. Per-step parsing moves to _step_failures.

Search stdout as well as stderr for the async jid, in case salt-run logging
is routed to stdout.

Close the RotatingFileHandler in test_make_logger_adds_handler_once to
avoid a ResourceWarning on Python 3.12+.

Verified on a 3.4.0 standalone: real failed and successful orchestration
results parse as before, a record whose evaluation raises is logged and
removed while the next record still reports, and a replicated SOC change
to telegraf.output (and its revert) is pushed, rendered and logged as
succeeded.
2026-10-01 09:06:28 -04:00
Josh Patterson 9732e1c639 Trim tracebacks in push failure log lines
When an orchestration step raises, salt returns the full traceback as the
step comment, and the drainer wrote it verbatim, putting ~70 lines into
so-push-drainer.log per failure. Collapse comments to one line and, for
tracebacks, keep only the lead-in and the raised exception, e.g.
"apply_soc_1: An exception occurred in this state:
salt.exceptions.AuthenticationError: Authentication error occurred."

Seen on a standalone when a pushed highstate restarted salt-master while
two queued pushes were waiting: their orchestrations lost the master
connection and failed with AuthenticationError, although the minion
completed both state runs.
2026-09-30 16:18:37 -04:00
Josh Patterson 53f9ebcd46 FIX: queue auto-applied state runs instead of failing on conflict
orch.push_batch passed `kwarg: {queue: 2}` to salt.state, but in Salt
3006 queue is a top-level salt.state argument and salt.state always sets
the minion's queue kwarg from it (default False), so the kwarg block was
silently dropped and every pushed state ran with queue=False. The drainer
dispatches a separate async orchestration each 15s pass, so settings saved
more than ~15s apart overlap on the same minion and every run after the
first fails immediately with 'The function "state.sls" is running as PID
...'. The change then waits for the next scheduled highstate.

Seen on a 3.4.0 standalone: hydra.enabled, telegraf.output, and two soc
settings (including soc.config.licenseKey) were saved within 30s. The soc
state was dispatched while the telegraf state was still running and was
rejected, so the license key was not applied.

Use `queue: True`, as orch.deploy_newnode already does. An int is treated
as max_queue and still falls through to the conflict error once that many
state runs are active.

The failure was only visible in the master log, since the drainer
dispatches with --async and logged only "dispatch accepted". The drainer
now:
  - logs each dispatched action
  - parses the orchestration jid from salt-run's stderr (the only place
    --async reports it) and records it under /opt/so/state/push_dispatched
  - on later passes looks each jid up with jobs.lookup_jid and logs either
    "push succeeded" or an ERROR with the failed step, the per-minion
    failed states or rejection text, and the triggering paths
Lookups run outside the pending-intent lock since the reactors share it.

The beacon now logs each audit_settings row it emits and the reactor logs
the audit row id, so a single change can be traced from audit_settings to
its push result.

Adds so-push-drainer_test.py; the drainer is now held to the 100% coverage
requirement in python-test.

Verified on the standalone: a soc push dispatched while a 90s state run
was in progress queued behind it (queue=True in the job args), completed,
and the drainer logged "push succeeded" for its jid. The new result
parsing reports the original soc conflict and the hydra license failure
from the job cache.
2026-09-30 16:18:37 -04:00
4 changed files with 7 additions and 6 deletions

No files matched your search

+1 -1
View File
@@ -1 +1 @@
3.4.0
3.0.0-foxtrot
+1 -1
View File
@@ -1,7 +1,7 @@
elasticsearch:
enabled: false
esheap: '600m'
version: 9.4.5
version: 9.4.8
index_clean: true
data_retention_method: DLM
vm:
+1 -1
View File
@@ -22,7 +22,7 @@ kibana:
- default
- file
migrations:
discardCorruptObjects: "9.4.5"
discardCorruptObjects: "9.4.8"
telemetry:
enabled: False
xpack:
+4 -3
View File
@@ -1535,9 +1535,10 @@ verify_es_version_compatibility() {
["8.18.4"]="8.18.6 8.18.8 9.0.8"
["8.18.6"]="8.18.8 9.0.8"
["8.18.8"]="9.0.8"
["9.0.8"]="9.3.3 9.3.7 9.4.5"
["9.3.3"]="9.3.7 9.4.5"
["9.3.7"]="9.4.5"
["9.0.8"]="9.3.3 9.3.7 9.4.5 9.4.8"
["9.3.3"]="9.3.7 9.4.5 9.4.8"
["9.3.7"]="9.4.5 9.4.8"
["9.4.5"]="9.4.8"
)
# Elasticsearch MUST upgrade through these versions