mirror of
https://github.com/Security-Onion-Solutions/securityonion.git
synced 2026-10-03 04:54:43 +02:00
orch.push_batch passed `kwarg: {queue: 2}` to salt.state, but in Salt
3006 queue is a top-level salt.state argument and salt.state always sets
the minion's queue kwarg from it (default False), so the kwarg block was
silently dropped and every pushed state ran with queue=False. The drainer
dispatches a separate async orchestration each 15s pass, so settings saved
more than ~15s apart overlap on the same minion and every run after the
first fails immediately with 'The function "state.sls" is running as PID
...'. The change then waits for the next scheduled highstate.
Seen on a 3.4.0 standalone: hydra.enabled, telegraf.output, and two soc
settings (including soc.config.licenseKey) were saved within 30s. The soc
state was dispatched while the telegraf state was still running and was
rejected, so the license key was not applied.
Use `queue: True`, as orch.deploy_newnode already does. An int is treated
as max_queue and still falls through to the conflict error once that many
state runs are active.
The failure was only visible in the master log, since the drainer
dispatches with --async and logged only "dispatch accepted". The drainer
now:
- logs each dispatched action
- parses the orchestration jid from salt-run's stderr (the only place
--async reports it) and records it under /opt/so/state/push_dispatched
- on later passes looks each jid up with jobs.lookup_jid and logs either
"push succeeded" or an ERROR with the failed step, the per-minion
failed states or rejection text, and the triggering paths
Lookups run outside the pending-intent lock since the reactors share it.
The beacon now logs each audit_settings row it emits and the reactor logs
the audit row id, so a single change can be traced from audit_settings to
its push result.
Adds so-push-drainer_test.py; the drainer is now held to the 100% coverage
requirement in python-test.
Verified on the standalone: a soc push dispatched while a 90s state run
was in progress queued behind it (queue=True in the job args), completed,
and the drainer logged "push succeeded" for its jid. The new result
parsing reports the original soc conflict and the hydra license failure
from the job cache.