Compare commits

..
Author SHA1 Message Date
Jorge Reyes 97fddc0719 Merge pull request #16201 from Security-Onion-Solutions/reyesj2-patch-1
fix salt batching command
2026-08-28 16:32:08 -05:00
Jorge Reyes a5deee1444 fix salt batching command 2026-08-28 16:27:29 -05:00
Jorge Reyes 3585ccca79 Merge pull request #16190 from Security-Onion-Solutions/reyesj2/es945
UPGRADE: Elasticsearch 9.4.5
2026-08-28 16:18:13 -05:00
reyesj2 dd035beec4 include fleet state 2026-08-28 13:45:55 -05:00
reyesj2 0dbb7803ef Merge branch 'reyesj2/reworksoup' into reyesj2/es945 2026-08-28 13:28:53 -05:00
reyesj2 30deb00277 use correct version variable 2026-08-28 13:27:44 -05:00
reyesj2 192363bc2f Merge branch 'reyesj2/reworksoup' into reyesj2/es945 2026-08-28 12:19:20 -05:00
reyesj2 3d8f86883a after an ES upgrade run a final elasticsearch state to create/regenerate any needed addon index templates 2026-08-28 12:16:08 -05:00
Josh Patterson 96bef89ba9 Merge pull request #16200 from Security-Onion-Solutions/rotatehype
add log rotation for hypervisor logs
2026-08-28 09:54:59 -04:00
Josh Patterson a244640539 Merge remote-tracking branch 'origin/3/dev' into rotatehype
# Conflicts:
#	salt/logrotate/defaults.yaml
#	salt/logrotate/soc_logrotate.yaml
2026-08-28 09:26:36 -04:00
reyesj2 fdb975fdef Merge branch 'reyesj2/reworksoup' into reyesj2/es945 2026-08-27 22:12:44 -05:00
reyesj2 f8401bef37 exclude elasticsearch indexing error during upgrade for temporarily outdated policies 2026-08-27 21:17:27 -05:00
Jason Ertel ca96a15091 Merge pull request #16199 from Security-Onion-Solutions/jertel/wip
fix well-known paths
2026-08-27 16:31:27 -04:00
Jason Ertel 1bac9a218e fix well-known paths 2026-08-27 16:28:00 -04:00
reyesj2 4786d359fb exclude telegraf error during elasticsearch upgrade / master election 2026-08-27 14:39:46 -05:00
reyesj2 d771fbc444 upgrade integration policies directly after integration package upgrade 2026-08-27 14:18:39 -05:00
reyesj2 85ab4c69e5 rename 2026-08-27 14:17:29 -05:00
reyesj2 cb8e576d6b run elasticsearch state on remote minions when there is an ES upgrade. Prior to manager completing its first full highstate that includes kibana / elasticfleet 2026-08-27 12:52:59 -05:00
reyesj2 fae1754fec clean elasticsearch transform prior to elasticsearch integration package upgrade to prevent fleet automatic rollback 2026-08-27 12:50:32 -05:00
reyesj2 d33eb70af6 reverts 83aaa76 #15985 - allow full highstate on manager when locked 2026-08-27 12:12:38 -05:00
Jorge Reyes f45dcfdf73 Merge pull request #16195 from Security-Onion-Solutions/revert-16165-reyesj2-patch-stg
Revert "patch issue with fs.protected_symlinks"
2026-08-27 09:36:36 -05:00
Jorge Reyes 62da505ea7 Revert "patch issue with fs.protected_symlinks" 2026-08-27 09:21:57 -05:00
Josh Patterson 7e5b6f276f Merge pull request #16194 from Security-Onion-Solutions/fix/auto-state-apply-local-salt-files
Detect hand-placed local/salt files in Auto State Apply
2026-08-27 09:58:32 -04:00
Matthew Wright 376d29e376 Merge pull request #16191 from Security-Onion-Solutions/mwright/agent-studio-memory
Memory and Reconcile Persona Annotations
2026-08-27 09:34:32 -04:00
Josh Patterson 665772adb8 Merge remote-tracking branch 'origin/3/dev' into fix/auto-state-apply-local-salt-files 2026-08-26 15:16:21 -04:00
Josh Patterson 094b4d5e86 Detect hand-placed local/salt files in Auto State Apply
Auto State Apply fires on SOC config saves and on suricata/strelka rule
updates. Files a user creates or edits by hand under
/opt/so/saltstack/local/salt/ change no pillar, so nothing fired and the
change waited for the next scheduled highstate, now 120 minutes by default.
That gap is the 3.2 Known Issue in the docs.

Watch the directories the docs tell users to edit, and route them through
the push pipeline that already exists:

  zeek/policy                       -> zeek    (covers intel/ and custom/)
  zeek/zkg                          -> zeek
  elasticsearch/files/ingest        -> elasticsearch
  elasticsearch/roles               -> elasticsearch
  logstash/pipelines/config/custom  -> logstash

Tags are pillar_push_map.yaml app names, so the existing entries already
carry the right state and compound target, and no map entry changes.

Rename the beacon rules_beacon -> local_files_beacon. Rules are now one of
five kinds of file it watches, and the new name matches how its sibling
postgres_pillar_beacon is named: source, then what it watches.

Replace push_suricata.sls and push_strelka.sls with one push_files.sls bound
to salt/beacon/*/local_files_beacon/*, which looks the tag up in
pillar_push_map.yaml the same way push_pillar.sls does. The map's suricata
and strelka targets match the compounds those two reactors hardcoded, so
rule pushes are unchanged. The app comes from the event tag rather than the
payload because salt's beacon loop pops the beacon's tag key off the data.

Key watermarks by watched directory instead of by tag. zeek/policy and
zeek/zkg both emit the tag zeek, and a shared watermark would make them
overwrite each other's digest and emit on every poll.

Prune .git from the fingerprint walk. zkg packages must be git clones with a
clean working tree, so the watched tree carries full git metadata; walking it
every 15s is wasted work and git's own index and ref mtime churn would fire a
grid-wide zeek apply on its own. Placing or updating a package always touches
working-tree files too, so detection is unaffected.

The watch is an allowlist rather than the whole local salt tree because salt
writes into that tree itself: hypervisor/hosts/ is rewritten continuously by
virtual_node_manager.py and virtual_power_manager.py, libvirt/images/ holds
multi-GB qcow2 files, and elasticfleet/files/so_agent-installers/,
elasticsearch/files/users, ca/files/ and filebeat/files/ are all state-written.
Watching any of them would either self-retrigger or make the 15s poll walk
gigabytes.
2026-08-26 15:15:29 -04:00
Matthew Wright 3a3667996c make personas non-advanced and readonlyui 2026-08-26 15:08:17 -04:00
Mike Reeves dfa6f0b454 Merge pull request #16192 from Security-Onion-Solutions/TOoSmOotH/telegraf-exec-array-syntax
Use argv arrays for telegraf inputs.exec commands
2026-08-26 14:14:02 -04:00
Mike Reeves 9f6679c043 Use argv arrays for telegraf inputs.exec commands
Telegraf 1.39 deprecated bare string entries in inputs.exec commands and
will drop them in 1.45, warning on every start:

  W! DeprecationWarning: Value "/scripts/esindexsize.sh" for option
  "command" of plugin "inputs.exec" deprecated since version 1.39.0

Each entry is now a single-element argv array. Array form skips shell
parsing, which is fine here: every command is a bare script path from
telegraf's scripts list, no args or shell metacharacters.
2026-08-26 13:42:37 -04:00
Matthew Wright fb7d162de1 memory and reconcile persona annotations 2026-08-26 12:36:46 -04:00
reyesj2 0ee8aa8079 revert manager running two highstates 2026-08-26 09:11:39 -05:00
Mike Reeves a8785870af Merge pull request #16188 from Security-Onion-Solutions/TOoSmOotH/remove-stock-kernel
Remove the stock EL9 kernel once a node is running UEK8
2026-08-25 10:03:15 -04:00
Jason Ertel 00f948e4d2 Merge pull request #16189 from Security-Onion-Solutions/jertel/wip
add vector ext for agentic memory
2026-08-25 09:52:49 -04:00
Jason Ertel f6ab92fc24 add vector ext for agentic memory 2026-08-25 09:43:56 -04:00
Mike Reeves 5e9fd4a45b Remove the stock EL9 kernel once a node is running UEK8
The UEK8 rollout installs the new kernel and flips the boot default, but
leaves the stock EL9 (RHCK) packages behind: disk in /boot and a stale
GRUB entry on every upgraded node.

They cannot be removed in the same pass that installs UEK8. dnf's
protect_running_kernel refuses to erase the booted kernel-core, so the
removal has to wait until the node has rebooted onto 6.x. Waiting is the
safer sequencing anyway -- the node proves it comes up on UEK8 before its
fallback is deleted -- so this does not remove RHCK from the uek7 branch
either, where dnf would allow it.

so-kernel-upgrade grows a --cleanup mode that does only the removal and
no-ops (exit 0, with a log line) on a node not yet running UEK8. Its uek8
branch, which previously reported "nothing to do", now runs that cleanup
along with set_default_kernel_conf -- which also closes a gap where a node
that came up on UEK8 straight from a fresh install never had
DEFAULTKERNEL=kernel-uek-core written.

The common highstate calls --cleanup gated on the running kernel, so the
cleanup lands grid-wide as each node reboots: fresh installs reboot at the
end of setup, upgraded nodes whenever the admin schedules it. The rpm
check inside the script is the idempotency guard, so subsequent highstates
cost an rpm query rather than a dnf transaction, and the package list is
not duplicated into the state where it could drift.
2026-08-25 09:30:01 -04:00
coreyogburn e7f54b49c4 Merge pull request #16187 from Security-Onion-Solutions/cogburn/memory
Cogburn/memory
2026-08-24 16:14:57 -06:00
Corey Ogburn a127ef5714 Show Toggle in UI
Must specify bool fields with `forcedType: bool` in order for them to render as toggles in the UI.
2026-08-24 14:09:37 -06:00
Corey Ogburn fcb889a30c Specify Default Embed Model 2026-08-24 14:09:37 -06:00
Corey Ogburn 99e1d83358 Add Interval and Disable by Default
Added `memoryScanIntervalSeconds` with a default of 5 mins.

Opted to set `useMemoryScanner` to false so by default our user's sessions are not sent to the cloud before they have a chance to configure the new setting.
2026-08-24 14:09:36 -06:00
Corey Ogburn 60052e0910 Memory Defaults and Annotations 2026-08-24 14:09:36 -06:00
reyesj2 99c3c7f8aa ES 9.4.5 soup es compatibility update 2026-08-21 12:46:46 -05:00
reyesj2 cef1dcfcee add additional problematic indices / templates. Also only update index templates if the script patched any index / data stream 2026-08-21 12:16:41 -05:00
Josh Patterson cec3f7ed57 Merge pull request #16182 from Security-Onion-Solutions/feature/logstash-log-level
Expose Logstash log.level and log.format in SOC
2026-08-21 12:14:07 -04:00
Josh Patterson f566a8965d Merge branch '3/dev' into feature/logstash-log-level 2026-08-21 11:45:00 -04:00
Josh Patterson 905cc1c0dd Merge pull request #16180 from Security-Onion-Solutions/feature/logstash-pipeline-settings-discrete
FEATURE: Allow for tuning multiple Logstash pipelines in SOC
2026-08-21 11:33:57 -04:00
Josh Patterson 12744353fb Allow ten custom logstash pipelines instead of five 2026-08-21 10:21:12 -04:00
Josh Patterson 52fc0cb828 Expose Logstash log.level and log.format in SOC
The Logstash log level was hardcoded to info in log4j2.properties, and
logstash.yml carried no log.level key, so the only way to raise verbosity
for troubleshooting was to hand-edit a file that the next highstate
overwrites. Add log_x_level and log_x_format to logstash:config so both
render into logstash.yml, annotated as advanced per-node settings with
the value sets Logstash 9.3.7 accepts.

log4j2.properties gains jinja, so it moves to log4j2.properties.jinja and
is rendered by a discrete lslog4j2 state rather than the lsetcsync
recurse, which cannot rename. The recurse exclude_pat now matches both
names so it neither copies the template verbatim nor lets clean: True
delete the rendered file, matching how pipelines.yml is already handled.

The appender layout is selected at render time so log.format actually
changes the log output instead of being a dead setting, keeping the
existing file name so nothing downstream moves. rootLogger.level now
follows ls.log.level rather than claiming info regardless of the
configured level.
2026-08-21 09:47:09 -04:00
reyesj2 247d9cdb34 kibana spaces 9.4.5 2026-08-20 17:00:53 -05:00
reyesj2 088b761190 9.4.5 policy updates 2026-08-20 16:35:46 -05:00
Josh Patterson 5c3a69d742 Warn about two pipeline_settings combinations that stop a pipeline
Grid testing every permitted value on the manager pipeline surfaced two
combinations the UI allows that take the pipeline down, neither of which
the descriptions mentioned.

pipeline.ordered: true requires pipeline.workers: 1; with more workers the
pipeline fails to start with "enabling the 'pipeline.ordered' setting
requires the use of a single pipeline worker". Also correct the auto
wording: it only engages when workers is explicitly set to 1.

queue.max_bytes larger than the free space on /nsm/logstash fails queue
creation with "Unable to allocate N more bytes", rather than merely being
inadvisable.
2026-08-20 17:33:34 -04:00
Josh Patterson c1f256e630 Correct pipeline_settings annotations against Logstash 9.3.7
Widen the byte-size regex, which rejected values Logstash accepts and so
blocked the save in SOC: bare-letter units (1g, 512m, 64k), decimals
(1.5gb), whitespace before the unit, and a bare integer. Allow whitespace
in dead_letter_queue.retain.age (5 d). Both stay lowercase-only, matching
byte_value.rb and AbstractPipelineExt.parseToDuration.

Fix description gaps: queue.checkpoint.retry is a Windows/SAN workaround
Elastic does not otherwise recommend, batch metrics sampling is technical
preview, queue.checkpoint.interval is deprecated in 9.1, compression makes
a queue unreadable by Logstash before 9.2, flush_check_interval has a
1000ms floor, max_events counts unread events, and the path settings are
created by Logstash but reject symlinks. Note which settings apply only to
persisted queues or an enabled DLQ.

Drop the undocumented 'disabled' value from queue.compression.

Numeric fields stay stricter than NumericSetting, which has no validator
and would accept negatives, floats and NaN in event counts and intervals.
2026-08-20 16:35:31 -04:00
reyesj2 2dcc81ea7d run so-elasticsearch-systems-indices-patch script every highstate with no op if no unassigned replicas are found for known problematic indices 2026-08-20 14:30:53 -05:00
Josh Patterson 356da00395 Ignore malformed logstash pipeline_settings instead of failing the state
A non-mapping value under logstash:pipeline_settings:<pipeline> made
config.sls raise "'str object' has no attribute 'get'", which failed the
whole logstash.config render rather than just skipping the bad value.
pipelines.yml.jinja already guarded this; config.sls now does too, and
logs which pipeline was ignored.
2026-08-20 13:29:38 -04:00
Josh Brower d2ff29b7a8 Merge pull request #16175 from Security-Onion-Solutions/fix/defaultsigma
Update Sigma template
2026-08-20 09:45:17 -04:00
Josh Brower 7bdaf9338e Update Sigma template 2026-08-20 09:31:56 -04:00
reyesj2 35f545a858 find known problematic system indices and add missing auto_expand_replicas configuration to prevent yellow cluster 2026-08-19 16:30:38 -05:00
Josh Patterson dff3d76efd Expose Logstash 9.3.7 pipeline settings per pipeline in SOC
Add logstash:pipeline_settings carrying the 27 pipeline-scoped settings
Logstash 9.3.7 accepts, annotated individually per pipeline and rendered
into pipelines.yml. A blank setting inherits from logstash.yml. Restart
logstash when pipelines.yml changes, and add the missing managerhype
annotation.

Fixes #15090
2026-08-19 16:21:28 -04:00
Josh Patterson 6f3f58bd70 Merge pull request #16170 from Security-Onion-Solutions/fix/boot-highstate-marker
FIX: enable so-boot-highstate.service on non-manager nodes
2026-08-19 11:50:39 -04:00
Josh Patterson d62c53fc92 Merge remote-tracking branch 'origin/3/dev' into fix/boot-highstate-marker 2026-08-19 11:39:48 -04:00
Josh Patterson 2f2187f714 Write setup-complete marker on non-manager nodes
so-boot-highstate.service was never enabled outside managers: only the
manager branch of so-setup called mark_setup_complete, so the marker its
service.enabled gates on never existed on sensors, search nodes, receivers,
etc.

Move the marker state into salt.minion.boot_highstate as the sole owner
within a highstate. Non-managers never apply salt.minion during setup, so
reaching it means setup is done and the marker is unconditional -- this also
heals already-installed nodes. Managers keep the legacy startup_states gate,
since they do highstate mid-setup.

Also add the marker to setup.virt for salt-cloud guests (replacing the
startup_states line removed in fabecb82) and to so-setup's non-manager branch.
2026-08-19 11:39:47 -04:00
Jorge Reyes 6c37bc1f9b Merge pull request #16165 from Security-Onion-Solutions/reyesj2-patch-stg
patch issue with fs.protected_symlinks
2026-08-18 15:04:37 -05:00
reyesj2 c9a041ddb4 upstream sentinel_one_cloud_funnel integration patched data stream name
ref: https://github.com/elastic/integrations/commit/9f1513423ca26e6fded414c4f5a3a89409efb736
2026-08-17 22:08:19 -05:00
reyesj2 de3306e73c ES 9.4.5 2026-08-17 16:45:41 -05:00
reyesj2 4b74e2c320 allow for unavailable minions 2026-08-17 15:26:01 -05:00
reyesj2 3744c0bd6c fix issue with fs.protected_symlinks prior to checking for fleet health 2026-08-17 15:24:05 -05:00
Matthew Wright 563b9d7c3b Merge pull request #16158 from Security-Onion-Solutions/mwright/advanced-agent-studio
Agentic: Agent Studio Salt Annotations
2026-08-17 13:10:34 -04:00
Josh Patterson ec91f9b830 Merge pull request #16162 from Security-Onion-Solutions/fix/zeekctl-cron
Disable the Zeek stats log
2026-08-14 16:13:12 -04:00
Josh Patterson 7f3f99880f Disable the Zeek stats log
"zeekctl cron" writes node statistics to /nsm/zeek/logs/stats. The CPU and memory
half comes from a helper that shells out to top, which the Zeek container does not
include. The helper's "command not found" output is then parsed as process data, so
every cron run appended a line per node reading "bad output from top", which
so-log-check reports.

Nothing wrote that file before, since log_stats and update_http_stats only run from
"zeekctl cron". Set StatsLogEnable to 0 so neither runs, and mark it read only since
the CPU and memory statistics cannot work with this image. The interface counters it
also collects are not used anywhere in Security Onion, which tracks Zeek packet loss
separately through packetloss.log and Telegraf, so nothing is lost by turning this
off. Note in StatsLogExpireInterval that it does nothing while the stats log is off.
2026-08-14 16:04:48 -04:00
Josh Brower 3e7f508620 Merge pull request #16161 from Security-Onion-Solutions/fixtests
Add another pcap job fp
2026-08-14 13:59:19 -04:00
Josh Brower c4555a5514 Add another pcap job fp 2026-08-14 13:54:14 -04:00
Josh Brower d4d63fa60a Merge pull request #16160 from Security-Onion-Solutions/fixtests
Add fp check
2026-08-14 11:46:24 -04:00
Josh Brower dcb931b97c Update excluded errors in so-log-check script 2026-08-14 11:23:25 -04:00
Josh Brower 8e6b16bde0 Add fp check 2026-08-14 11:22:12 -04:00
Josh Patterson 63692aa1a0 Merge pull request #16159 from Security-Onion-Solutions/fix/zeekctl-cron
Run zeekctl cron so LogExpireInterval and the other expire settings take effect
2026-08-14 09:49:28 -04:00
Josh Patterson 2663ca87a2 Mark the mail-only zeekctl settings read only
MailTo, MailConnectionSummary and MailHostUpDown do nothing but send mail, and
the Zeek container has no mail program, so nothing they control can happen. Mark
them read only rather than offering knobs in SOC that cannot take effect.

MailConnectionSummary only gates the emailed copy; the connection summary is
generated and archived either way. MailTo also feeds Notice::mail_dest, but
Security Onion never enables the notice email action, so that half is inert too.
MailHostUpDown gates only the notification text - host status detection, the
plugin hook and the stored state all run regardless.

MinDiskSpace stays editable. It is not mail only: setting it to 0 skips the disk
space check entirely, and the warning it produces is not emailed but does appear
in the output of "zeekctl cron". Correct its description, and MailHostUpDown's,
which both claimed these settings have no visible effect.
2026-08-14 09:19:11 -04:00
Josh Patterson a337a3e4f6 Run zeekctl cron so the expire settings take effect
LogExpireInterval, StatsLogExpireInterval and CrashExpireInterval are only acted
on by "zeekctl cron", which nothing in the grid ran, so setting them in SOC did
nothing. Add so-zeek-cron and run it every 5 minutes, the interval upstream
recommends.

This also restarts a node that died unexpectedly and marks it crashed so a crash
report is written, which is what CrashExpireInterval then reaps.

The crontab runs as root because the script needs the docker socket; it drops to
the zeek user inside the container so the stats logs and zeekctl-config.sh it
writes stay owned by uid 937.

Annotate the five zeekctl settings that were previously undocumented. The regex
on LogExpireInterval matters: a bare number means days, and a value shorter than
LogRotationInterval raises ConfigurationError, which fails the zeekctl deploy in
the container entrypoint. Zeek then never starts while Salt still reports success
and the container still reports healthy. Excluding the min unit keeps that
unreachable at the default 3600 second rotation interval. MinDiskSpace and
MailHostUpDown only send mail and the image has no sendmail, so their
descriptions say they currently have no effect.
2026-08-13 16:18:24 -04:00
Josh Patterson 52791204e4 add logrotate for virtual_node_manager and so-salt-cloud 2026-05-19 13:40:19 -04:00
47 changed files with 1911 additions and 455 deletions
@@ -3,13 +3,14 @@
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# Custom salt beacon that watches the suricata/strelka rule directories for changes
# and emits a beacon event per changed directory. This replaces the stock salt
# `inotify` beacon, which leaks a kernel inotify instance every time the minion
# rebuilds the beacon loader's __context__ (orphaning the old pyinotify.Notifier
# without closing it) until fs.inotify.max_user_instances is exhausted and the
# beacon dies with EMFILE. Polling holds zero inotify instances, so the leak is
# impossible, and it keeps firing during state runs (no blackout).
# Custom salt beacon that watches hand-edited directories under
# /opt/so/saltstack/local/salt/ for changes and emits a beacon event per changed
# directory. This replaces the stock salt `inotify` beacon, which leaks a kernel
# inotify instance every time the minion rebuilds the beacon loader's __context__
# (orphaning the old pyinotify.Notifier without closing it) until
# fs.inotify.max_user_instances is exhausted and the beacon dies with EMFILE.
# Polling holds zero inotify instances, so the leak is impossible, and it keeps
# firing during state runs (no blackout).
#
# Detection is poll-based with a per-directory fingerprint persisted to
# WATERMARK_DIR: each pass walks the directory and hashes every file's
@@ -19,9 +20,10 @@
# up on the next one).
#
# Each emitted event carries the watched directory path under the configured tag
# (e.g. salt/beacon/<minion>/rules_beacon/suricata); the push_suricata / push_strelka
# reactors write a push intent, after which the existing so-push-drainer /
# orch.push_batch pipeline takes over unchanged.
# (e.g. salt/beacon/<minion>/local_files_beacon/zeek); the push_files reactor
# looks the tag up in salt/reactor/pillar_push_map.yaml and writes a push intent,
# after which the existing so-push-drainer / orch.push_batch pipeline takes over
# unchanged.
import hashlib
import logging
@@ -77,7 +79,9 @@ def _fingerprint(directory):
h = hashlib.sha1()
if os.path.isdir(directory):
entries = []
for root, _dirs, files in os.walk(directory):
for root, dirs, files in os.walk(directory):
# zkg packages are git clones; .git churn would fire a state apply on its own.
dirs[:] = [d for d in dirs if d != '.git']
for name in files:
full = os.path.join(root, name)
if _excluded(full):
@@ -94,20 +98,22 @@ def _fingerprint(directory):
return h.hexdigest()
def _watermark_file(tag):
return os.path.join(WATERMARK_DIR, 'rules_beacon_%s.hash' % tag)
def _watermark_file(tag, directory):
# Keyed by directory: zeek/policy and zeek/zkg share the tag `zeek`.
scope = hashlib.sha1(directory.encode('utf-8', 'surrogateescape')).hexdigest()[:12]
return os.path.join(WATERMARK_DIR, 'local_files_beacon_%s_%s.hash' % (tag, scope))
def _read_watermark(tag):
def _read_watermark(tag, directory):
try:
with open(_watermark_file(tag), 'r') as f:
with open(_watermark_file(tag, directory), 'r') as f:
return (f.read() or '').strip() or None
except IOError:
return None
def _write_watermark(tag, digest):
path = _watermark_file(tag)
def _write_watermark(tag, directory, digest):
path = _watermark_file(tag, directory)
try:
os.makedirs(WATERMARK_DIR, exist_ok=True)
tmp = path + '.tmp'
@@ -115,7 +121,7 @@ def _write_watermark(tag, digest):
f.write(digest)
os.rename(tmp, path)
except OSError:
log.exception('rules_beacon: failed to persist watermark to %s', path)
log.exception('local_files_beacon: failed to persist watermark to %s', path)
def beacon(config):
@@ -123,17 +129,17 @@ def beacon(config):
for directory, tag in _paths_from_config(config).items():
digest = _fingerprint(directory)
previous = _read_watermark(tag)
previous = _read_watermark(tag, directory)
# First run / missing watermark: seed the digest and emit nothing so a
# fresh host does not fire a spurious fleetwide push.
if previous is None:
_write_watermark(tag, digest)
_write_watermark(tag, directory, digest)
continue
if digest != previous:
_write_watermark(tag, digest)
_write_watermark(tag, directory, digest)
retval.append({'tag': tag, 'path': directory})
log.info('rules_beacon: change detected in %s, emitting %s', directory, tag)
log.info('local_files_beacon: change detected in %s, emitting %s', directory, tag)
return retval
+221
View File
@@ -0,0 +1,221 @@
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
import hashlib
import os
import shutil
import tempfile
import unittest
from unittest.mock import patch
import local_files_beacon
class TestRulesBeacon(unittest.TestCase):
def setUp(self):
# Isolate all on-disk state (watermarks and the dirs we fingerprint) in a
# throwaway tree, and point WATERMARK_DIR at it so the real read/write
# helpers run against actual files.
self.tmpdir = tempfile.mkdtemp()
self.state = os.path.join(self.tmpdir, 'state')
patcher = patch.object(local_files_beacon, 'WATERMARK_DIR', self.state)
patcher.start()
self.addCleanup(patcher.stop)
def tearDown(self):
shutil.rmtree(self.tmpdir, ignore_errors=True)
def _make_dir(self, name, files=None):
path = os.path.join(self.tmpdir, name)
os.makedirs(path, exist_ok=True)
for fname, content in (files or {}).items():
with open(os.path.join(path, fname), 'w') as f:
f.write(content)
return path
# -- trivial contract -------------------------------------------------
def test_virtual_returns_true(self):
self.assertTrue(local_files_beacon.__virtual__())
def test_validate_returns_valid(self):
self.assertEqual(local_files_beacon.validate({}), (True, 'valid'))
# -- _paths_from_config -----------------------------------------------
def test_paths_from_config_list_of_dicts(self):
config = [{'interval': 10}, {'paths': {'/a': 'suricata', '/b': 'strelka'}}]
self.assertEqual(
local_files_beacon._paths_from_config(config),
{'/a': 'suricata', '/b': 'strelka'},
)
def test_paths_from_config_plain_dict(self):
self.assertEqual(
local_files_beacon._paths_from_config({'paths': {'/a': 'suricata'}}),
{'/a': 'suricata'},
)
def test_paths_from_config_skips_non_dict_items(self):
self.assertEqual(local_files_beacon._paths_from_config(['bogus', 42]), {})
def test_paths_from_config_paths_not_a_dict(self):
self.assertEqual(local_files_beacon._paths_from_config({'paths': 'nope'}), {})
def test_paths_from_config_unexpected_type(self):
self.assertEqual(local_files_beacon._paths_from_config('nonsense'), {})
# -- _excluded --------------------------------------------------------
def test_excluded_matches_temp_and_editor_files(self):
for pathname in ('/rules/foo.swp', '/rules/foo~', '/rules/4913', '/rules/.#foo'):
self.assertTrue(local_files_beacon._excluded(pathname), pathname)
def test_excluded_allows_real_rule_files(self):
self.assertFalse(local_files_beacon._excluded('/rules/suricata.rules'))
# -- _fingerprint -----------------------------------------------------
def test_fingerprint_missing_dir_is_empty_tree_digest(self):
missing = os.path.join(self.tmpdir, 'does-not-exist')
self.assertEqual(local_files_beacon._fingerprint(missing), hashlib.sha1().hexdigest())
def test_fingerprint_changes_when_content_changes(self):
d = self._make_dir('rules', {'a.rules': 'alert'})
before = local_files_beacon._fingerprint(d)
with open(os.path.join(d, 'a.rules'), 'w') as f:
f.write('alert tcp any any -> any any') # different size
self.assertNotEqual(local_files_beacon._fingerprint(d), before)
def test_fingerprint_ignores_excluded_files(self):
d = self._make_dir('rules', {'a.rules': 'alert'})
before = local_files_beacon._fingerprint(d)
with open(os.path.join(d, 'a.rules.swp'), 'w') as f:
f.write('editor swap')
self.assertEqual(local_files_beacon._fingerprint(d), before)
def test_fingerprint_skips_unstatable_entries(self):
# A dangling symlink appears in os.walk's file list but os.stat raises
# OSError, exercising the except-continue path.
d = self._make_dir('rules', {'a.rules': 'alert'})
good = local_files_beacon._fingerprint(d)
os.symlink(os.path.join(d, 'missing-target'), os.path.join(d, 'broken.link'))
self.assertEqual(local_files_beacon._fingerprint(d), good)
def test_fingerprint_prunes_git_metadata(self):
# zkg packages are git clones, so the watched tree carries .git.
d = self._make_dir('zkg', {'pkg.zeek': 'print 1;'})
before = local_files_beacon._fingerprint(d)
git_dir = os.path.join(d, 'pkg', '.git', 'refs', 'heads')
os.makedirs(git_dir)
with open(os.path.join(git_dir, 'main'), 'w') as f:
f.write('0' * 40)
self.assertEqual(local_files_beacon._fingerprint(d), before)
def test_fingerprint_still_sees_worktree_next_to_git(self):
d = self._make_dir('zkg', {'pkg.zeek': 'print 1;'})
os.makedirs(os.path.join(d, 'pkg', '.git'))
before = local_files_beacon._fingerprint(d)
with open(os.path.join(d, 'pkg', 'scripts.zeek'), 'w') as f:
f.write('print 2;')
self.assertNotEqual(local_files_beacon._fingerprint(d), before)
# -- _read_watermark / _write_watermark -------------------------------
def test_watermark_round_trip(self):
local_files_beacon._write_watermark('suricata', '/rules/suricata', 'deadbeef')
self.assertEqual(
local_files_beacon._read_watermark('suricata', '/rules/suricata'), 'deadbeef')
def test_read_watermark_missing_returns_none(self):
self.assertIsNone(local_files_beacon._read_watermark('suricata', '/rules/suricata'))
def test_read_watermark_empty_file_returns_none(self):
os.makedirs(self.state, exist_ok=True)
with open(local_files_beacon._watermark_file('suricata', '/rules/suricata'), 'w') as f:
f.write('')
self.assertIsNone(local_files_beacon._read_watermark('suricata', '/rules/suricata'))
def test_write_watermark_swallows_oserror(self):
with patch.object(local_files_beacon.os, 'makedirs', side_effect=OSError):
local_files_beacon._write_watermark('suricata', '/rules/suricata', 'deadbeef')
self.assertIsNone(local_files_beacon._read_watermark('suricata', '/rules/suricata'))
def test_watermark_file_differs_per_directory_within_one_tag(self):
# zeek/policy and zeek/zkg share the tag 'zeek'.
self.assertNotEqual(
local_files_beacon._watermark_file('zeek', '/local/zeek/policy'),
local_files_beacon._watermark_file('zeek', '/local/zeek/zkg'),
)
def test_watermarks_are_independent_within_one_tag(self):
local_files_beacon._write_watermark('zeek', '/local/zeek/policy', 'policyhash')
local_files_beacon._write_watermark('zeek', '/local/zeek/zkg', 'zkghash')
self.assertEqual(
local_files_beacon._read_watermark('zeek', '/local/zeek/policy'), 'policyhash')
self.assertEqual(
local_files_beacon._read_watermark('zeek', '/local/zeek/zkg'), 'zkghash')
# -- beacon -----------------------------------------------------------
def _config(self, mapping):
return [{'paths': mapping}]
def test_beacon_seeds_first_run_and_emits_nothing(self):
with patch.object(local_files_beacon, '_fingerprint', return_value='hash1'), \
patch.object(local_files_beacon, '_read_watermark', return_value=None), \
patch.object(local_files_beacon, '_write_watermark') as mock_write:
result = local_files_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [])
mock_write.assert_called_once_with('suricata', '/rules/suricata', 'hash1')
def test_beacon_emits_on_change(self):
with patch.object(local_files_beacon, '_fingerprint', return_value='newhash'), \
patch.object(local_files_beacon, '_read_watermark', return_value='oldhash'), \
patch.object(local_files_beacon, '_write_watermark') as mock_write:
result = local_files_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [{'tag': 'suricata', 'path': '/rules/suricata'}])
mock_write.assert_called_once_with('suricata', '/rules/suricata', 'newhash')
def test_beacon_no_change_emits_nothing(self):
with patch.object(local_files_beacon, '_fingerprint', return_value='samehash'), \
patch.object(local_files_beacon, '_read_watermark', return_value='samehash'), \
patch.object(local_files_beacon, '_write_watermark') as mock_write:
result = local_files_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [])
mock_write.assert_not_called()
def test_beacon_end_to_end_with_real_files(self):
# Exercise the full stack (real fingerprint + real watermark files) across
# two poll passes: first seeds silently, second fires after a write.
d = self._make_dir('rules', {'a.rules': 'alert'})
config = self._config({d: 'suricata'})
self.assertEqual(local_files_beacon.beacon(config), []) # seed pass
self.assertEqual(local_files_beacon.beacon(config), []) # unchanged pass
with open(os.path.join(d, 'b.rules'), 'w') as f:
f.write('alert tcp any any -> any any')
self.assertEqual(local_files_beacon.beacon(config), [{'tag': 'suricata', 'path': d}])
def test_beacon_two_dirs_one_tag_do_not_flap(self):
# Tag-keyed watermarks would clobber each other and emit on every pass.
policy = self._make_dir('zeek/policy', {'intel.dat': '#fields\tindicator'})
zkg = self._make_dir('zeek/zkg', {'README': 'place packages here'})
config = self._config({policy: 'zeek', zkg: 'zeek'})
self.assertEqual(local_files_beacon.beacon(config), []) # seed pass
self.assertEqual(local_files_beacon.beacon(config), []) # idle
self.assertEqual(local_files_beacon.beacon(config), []) # still idle
with open(os.path.join(policy, 'intel.dat'), 'a') as f:
f.write('\nevil.com\tIntel::DOMAIN\tsource\n')
self.assertEqual(local_files_beacon.beacon(config), [{'tag': 'zeek', 'path': policy}])
self.assertEqual(local_files_beacon.beacon(config), []) # quiet again
if __name__ == '__main__':
unittest.main()
-172
View File
@@ -1,172 +0,0 @@
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
import hashlib
import os
import shutil
import tempfile
import unittest
from unittest.mock import patch
import rules_beacon
class TestRulesBeacon(unittest.TestCase):
def setUp(self):
# Isolate all on-disk state (watermarks and the dirs we fingerprint) in a
# throwaway tree, and point WATERMARK_DIR at it so the real read/write
# helpers run against actual files.
self.tmpdir = tempfile.mkdtemp()
self.state = os.path.join(self.tmpdir, 'state')
patcher = patch.object(rules_beacon, 'WATERMARK_DIR', self.state)
patcher.start()
self.addCleanup(patcher.stop)
def tearDown(self):
shutil.rmtree(self.tmpdir, ignore_errors=True)
def _make_dir(self, name, files=None):
path = os.path.join(self.tmpdir, name)
os.makedirs(path, exist_ok=True)
for fname, content in (files or {}).items():
with open(os.path.join(path, fname), 'w') as f:
f.write(content)
return path
# -- trivial contract -------------------------------------------------
def test_virtual_returns_true(self):
self.assertTrue(rules_beacon.__virtual__())
def test_validate_returns_valid(self):
self.assertEqual(rules_beacon.validate({}), (True, 'valid'))
# -- _paths_from_config -----------------------------------------------
def test_paths_from_config_list_of_dicts(self):
config = [{'interval': 10}, {'paths': {'/a': 'suricata', '/b': 'strelka'}}]
self.assertEqual(
rules_beacon._paths_from_config(config),
{'/a': 'suricata', '/b': 'strelka'},
)
def test_paths_from_config_plain_dict(self):
self.assertEqual(
rules_beacon._paths_from_config({'paths': {'/a': 'suricata'}}),
{'/a': 'suricata'},
)
def test_paths_from_config_skips_non_dict_items(self):
self.assertEqual(rules_beacon._paths_from_config(['bogus', 42]), {})
def test_paths_from_config_paths_not_a_dict(self):
self.assertEqual(rules_beacon._paths_from_config({'paths': 'nope'}), {})
def test_paths_from_config_unexpected_type(self):
self.assertEqual(rules_beacon._paths_from_config('nonsense'), {})
# -- _excluded --------------------------------------------------------
def test_excluded_matches_temp_and_editor_files(self):
for pathname in ('/rules/foo.swp', '/rules/foo~', '/rules/4913', '/rules/.#foo'):
self.assertTrue(rules_beacon._excluded(pathname), pathname)
def test_excluded_allows_real_rule_files(self):
self.assertFalse(rules_beacon._excluded('/rules/suricata.rules'))
# -- _fingerprint -----------------------------------------------------
def test_fingerprint_missing_dir_is_empty_tree_digest(self):
missing = os.path.join(self.tmpdir, 'does-not-exist')
self.assertEqual(rules_beacon._fingerprint(missing), hashlib.sha1().hexdigest())
def test_fingerprint_changes_when_content_changes(self):
d = self._make_dir('rules', {'a.rules': 'alert'})
before = rules_beacon._fingerprint(d)
with open(os.path.join(d, 'a.rules'), 'w') as f:
f.write('alert tcp any any -> any any') # different size
self.assertNotEqual(rules_beacon._fingerprint(d), before)
def test_fingerprint_ignores_excluded_files(self):
d = self._make_dir('rules', {'a.rules': 'alert'})
before = rules_beacon._fingerprint(d)
with open(os.path.join(d, 'a.rules.swp'), 'w') as f:
f.write('editor swap')
self.assertEqual(rules_beacon._fingerprint(d), before)
def test_fingerprint_skips_unstatable_entries(self):
# A dangling symlink appears in os.walk's file list but os.stat raises
# OSError, exercising the except-continue path.
d = self._make_dir('rules', {'a.rules': 'alert'})
good = rules_beacon._fingerprint(d)
os.symlink(os.path.join(d, 'missing-target'), os.path.join(d, 'broken.link'))
self.assertEqual(rules_beacon._fingerprint(d), good)
# -- _read_watermark / _write_watermark -------------------------------
def test_watermark_round_trip(self):
rules_beacon._write_watermark('suricata', 'deadbeef')
self.assertEqual(rules_beacon._read_watermark('suricata'), 'deadbeef')
def test_read_watermark_missing_returns_none(self):
self.assertIsNone(rules_beacon._read_watermark('suricata'))
def test_read_watermark_empty_file_returns_none(self):
os.makedirs(self.state, exist_ok=True)
with open(rules_beacon._watermark_file('suricata'), 'w') as f:
f.write('')
self.assertIsNone(rules_beacon._read_watermark('suricata'))
def test_write_watermark_swallows_oserror(self):
with patch.object(rules_beacon.os, 'makedirs', side_effect=OSError):
rules_beacon._write_watermark('suricata', 'deadbeef')
self.assertIsNone(rules_beacon._read_watermark('suricata'))
# -- beacon -----------------------------------------------------------
def _config(self, mapping):
return [{'paths': mapping}]
def test_beacon_seeds_first_run_and_emits_nothing(self):
with patch.object(rules_beacon, '_fingerprint', return_value='hash1'), \
patch.object(rules_beacon, '_read_watermark', return_value=None), \
patch.object(rules_beacon, '_write_watermark') as mock_write:
result = rules_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [])
mock_write.assert_called_once_with('suricata', 'hash1')
def test_beacon_emits_on_change(self):
with patch.object(rules_beacon, '_fingerprint', return_value='newhash'), \
patch.object(rules_beacon, '_read_watermark', return_value='oldhash'), \
patch.object(rules_beacon, '_write_watermark') as mock_write:
result = rules_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [{'tag': 'suricata', 'path': '/rules/suricata'}])
mock_write.assert_called_once_with('suricata', 'newhash')
def test_beacon_no_change_emits_nothing(self):
with patch.object(rules_beacon, '_fingerprint', return_value='samehash'), \
patch.object(rules_beacon, '_read_watermark', return_value='samehash'), \
patch.object(rules_beacon, '_write_watermark') as mock_write:
result = rules_beacon.beacon(self._config({'/rules/suricata': 'suricata'}))
self.assertEqual(result, [])
mock_write.assert_not_called()
def test_beacon_end_to_end_with_real_files(self):
# Exercise the full stack (real fingerprint + real watermark files) across
# two poll passes: first seeds silently, second fires after a write.
d = self._make_dir('rules', {'a.rules': 'alert'})
config = self._config({d: 'suricata'})
self.assertEqual(rules_beacon.beacon(config), []) # seed pass
self.assertEqual(rules_beacon.beacon(config), []) # unchanged pass
with open(os.path.join(d, 'b.rules'), 'w') as f:
f.write('alert tcp any any -> any any')
self.assertEqual(rules_beacon.beacon(config), [{'tag': 'suricata', 'path': d}])
if __name__ == '__main__':
unittest.main()
+14
View File
@@ -141,6 +141,20 @@ pin_nic_names:
- file: common_sbin
- file: statedir
# Once a node is actually running UEK8, the stock EL9 (RHCK) kernel packages are dead weight.
# They can't be removed any earlier -- dnf protects the running kernel -- so the cleanup waits
# for the reboot, which makes the highstate the natural place to catch it: fresh installs
# reboot at the end of setup, and upgraded nodes reboot whenever the admin schedules it.
# so-kernel-upgrade --cleanup checks rpm before touching dnf, so this costs an rpm query on
# every highstate after the first pass. The package list lives in the script only, so there
# is nothing here to drift out of sync with it.
remove_stock_kernel:
cmd.run:
- name: /usr/sbin/so-kernel-upgrade --cleanup
- onlyif: 'uname -r | grep -qE "^6\.[0-9]+.*uek"'
- require:
- file: common_sbin
common_sbin_jinja:
file.recurse:
- name: /usr/sbin
+81 -7
View File
@@ -5,10 +5,11 @@
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
#
# so-kernel-upgrade — install the UEK8 (6.x) kernel and make it the boot default.
# so-kernel-upgrade — install the UEK8 (6.x) kernel, make it the boot default, and once the
# node is running it, remove the stock EL9 kernel.
#
# Security Onion is moving off the EL9 stock kernel (RHCK, 5.14) and UEK7 (5.15) onto UEK8
# (6.x). Three things have to happen, and the tool has to drive each one:
# (6.x). Four things have to happen, and the tool has to drive each one:
#
# 1. Populate. The manager mirrors the UEK8 packages into /nsm/kernelrepo via so-repo-sync,
# and serves them to the grid over https://<manager>/kernelrepo. Until that sync runs the
@@ -26,10 +27,21 @@
# - From the stock EL9 kernel (RHCK, 5.14, no UEK) it is a flavor CROSS that is NOT
# auto-promoted, so the box keeps booting RHCK until grubby is told otherwise.
# This tool inspects the running kernel and only runs 'grubby --set-default' for RHCK.
# 4. Clean up. Once the node is actually RUNNING UEK8 the stock kernel packages are dead
# weight -- disk in /boot and a stale GRUB entry. They cannot come off any earlier:
# dnf's protect_running_kernel refuses to erase the booted kernel-core, so the removal
# has to wait for the reboot. Waiting is also the safer sequencing on its own terms --
# the node has proven it comes up on UEK8 before its fallback is deleted. That is why
# the removal does not happen in the uek7 branch either, where dnf would allow it.
#
# Every one of those failure modes is silent by default. This tool handles each case and fails
# loudly when it cannot, rather than reporting success while changing nothing.
#
# Invocation: with no arguments it drives the whole sequence for whatever kernel the node is
# on. With --cleanup it does the step 4 removal ONLY, and no-ops on a node that isn't running
# UEK8 yet -- that is the form the common highstate calls (remove_stock_kernel in
# salt/common/init.sls) so the cleanup lands grid-wide after each node reboots.
#
# Manager vs minion: only the manager owns /nsm/kernelrepo, so only the manager can populate
# it. If the repo is empty here, a manager runs so-repo-sync itself; a minion has no way to
# fix it and exits non-zero telling the admin to sync the manager first.
@@ -49,6 +61,11 @@ KERNEL_REPO_DIR="/nsm/kernelrepo"
REPOSYNC_CONF="/opt/so/conf/reposync/repodownload.conf"
GLOBAL_PILLAR="/opt/so/saltstack/local/pillar/global/soc_global.sls"
# Stock EL9 (RHCK) kernel packages, removed only once the node is running UEK8 (see step 4
# in the header). Left deliberately narrow: UEK7 kernel-uek builds age out on their own via
# installonly_limit=3, and kernel-devel/kernel-headers are not touched.
RHCK_PKGS="kernel kernel-core kernel-modules kernel-modules-core kernel-tools kernel-tools-libs"
log() { echo "[so-kernel-upgrade] $*"; }
die() { echo "[so-kernel-upgrade] ERROR: $*" >&2; exit 1; }
@@ -149,8 +166,13 @@ ensure_kernel_repo() {
}
reboot_notice() {
[ "$(uname -r)" = "$(basename "$1" | sed 's/^vmlinuz-//')" ] \
|| log "REBOOT REQUIRED to start using the UEK8 kernel (currently running $(uname -r))."
[ "$(uname -r)" = "$(basename "$1" | sed 's/^vmlinuz-//')" ] && return 0
log "REBOOT REQUIRED to start using the UEK8 kernel (currently running $(uname -r))."
# The stock kernel can't be removed until it stops being the running one, so say when
# that will happen rather than leaving the admin to wonder if it was missed.
[ -n "$(rhck_installed)" ] \
&& log "The stock EL9 kernel is left in place until then; it is removed by the next highstate after the reboot."
return 0
}
# Keep future kernel updates on the UEK line rather than falling back to RHCK. Oracle ships
@@ -162,6 +184,32 @@ set_default_kernel_conf() {
fi
}
# Which of RHCK_PKGS are actually installed, one per line. rpm -qa treats each argument as a
# name glob and prints only what it finds, so a package that was never installed (or is
# already gone) simply doesn't appear -- no "not installed" noise and no non-zero exit.
rhck_installed() {
rpm -qa $RHCK_PKGS 2>/dev/null
}
# Remove the stock EL9 kernel. Only ever called once the running kernel is UEK8. The rpm
# check above is the idempotency guard, so this is a cheap no-op on every highstate after
# the first one -- it costs an rpm query, not a dnf transaction.
remove_rhck() {
local installed; installed="$(rhck_installed)"
if [ -z "$installed" ]; then
log "no stock EL9 (RHCK) kernel packages installed; nothing to remove."
return 0
fi
log "running UEK8; removing the stock EL9 (RHCK) kernel packages:"
echo "$installed" | sed 's/^/[so-kernel-upgrade] /'
dnf -y remove $RHCK_PKGS || die "failed to remove the stock EL9 kernel packages"
installed="$(rhck_installed)"
[ -z "$installed" ] || die "dnf reported success but these remain: $(echo $installed)"
log "stock EL9 kernel packages removed."
}
# Make sure a UEK8 kernel is installed, leaving its boot entry in INSTALLED_UEK8. If one is
# already present we leave the repo alone -- it may be disabled or empty and we don't need it
# just to flip the boot default. Otherwise install the explicit NEVRA, not the bare package
@@ -184,12 +232,38 @@ ensure_uek8_installed() {
log "installed UEK8 kernel: $INSTALLED_UEK8"
}
# --cleanup does step 4 and nothing else. It exits 0 rather than failing on a node that
# isn't on UEK8 yet: the highstate gates on 'uname -r' before calling this, and a state that
# fails whenever that gate races would be worse than one that says what it's waiting for.
case "$1" in
"")
;;
--cleanup)
if [ "$(running_flavor)" != uek8 ]; then
log "not running a UEK8 kernel yet (currently $(uname -r)); leaving the stock EL9 kernel in place."
log "Run so-kernel-upgrade with no arguments to install UEK8, then reboot."
exit 0
fi
set_default_kernel_conf
remove_rhck
exit 0
;;
*)
echo "Usage: so-kernel-upgrade [--cleanup]" >&2
echo " (no arguments) install UEK8, make it the boot default, clean up once it's running" >&2
echo " --cleanup remove the stock EL9 kernel; no-op unless already running UEK8" >&2
exit 1
;;
esac
case "$(running_flavor)" in
uek8)
# Already on the 6.x UEK line. A plain 'dnf update' keeps this node current within the
# lineage and auto-promotes newer builds, so there is nothing for this tool to do.
log "already running a UEK8 kernel ($(uname -r)); nothing to do."
exit 0
# lineage and auto-promotes newer builds, so there is no install or grubby work left --
# only the step 4 cleanup, which this is the first point in the sequence that can run it.
log "already running a UEK8 kernel ($(uname -r)); no kernel install needed."
set_default_kernel_conf
remove_rhck
;;
uek7)
+5 -1
View File
@@ -134,6 +134,7 @@ if [[ $EXCLUDE_STARTUP_ERRORS == 'Y' ]]; then
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|Redis may have been restarted" # Redis likely restarted by salt
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|file already closed" # Go logging race condition during container restart
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|relation \"audit_settings\" does not exist" # salt checking for changes before SOC starts
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|Error in plugin: elasticsearch: Unable to retrieve master node information" # expected error while ES is upgrading/electing a master
fi
if [[ $EXCLUDE_FALSE_POSITIVE_ERRORS == 'Y' ]]; then
@@ -154,6 +155,8 @@ if [[ $EXCLUDE_FALSE_POSITIVE_ERRORS == 'Y' ]]; then
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|id.orig_h" # false positive (zeek test data)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|emerging-all.rules" # false positive (error in rulename)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|invalid query input" # false positive (Invalid user input in hunt query)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|no data available for the requested dates" # false positive (pcap cypress test submits a job with an empty time frame)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|no job processor" # false positive (same empty-time-frame job on import nodes, where no pcap processor runs)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|example" # false positive (example test data)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|status 200" # false positive (request successful, contained error string in content)
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|app_layer.error" # false positive (suricata 7) in stats.log e.g. app_layer.error.imap.parser | Total | 0
@@ -238,6 +241,7 @@ if [[ $EXCLUDE_KNOWN_ERRORS == 'Y' ]]; then
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|tcp 127.0.0.1:6791: bind: address already in use" # so-elastic-fleet agent restarting. Seen starting w/ 8.18.8 https://github.com/elastic/kibana/issues/201459
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|TransformTask\] \[logs-.*user so_kibana lacks the required permissions" # Known issue with integrations starting transform jobs that are explicitly not allowed to start as a system user
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|manifest unknown" # appears in so-dockerregistry log for so-tcpreplay following docker upgrade to 29.2.1-1
EXCLUDED_ERRORS="$EXCLUDED_ERRORS|Could not index event to Elasticsearch.*\"version\" => \"9.0.8\"" # Expected during Elastic upgrade temporarily, as policies referencing older pipelines are updated
fi
RESULT=0
@@ -302,4 +306,4 @@ else
echo -e "\nResult: One or more errors found"
fi
exit $RESULT
exit $RESULT
@@ -5,7 +5,7 @@
"package": {
"name": "endpoint",
"title": "Elastic Defend",
"version": "9.3.1",
"version": "9.4.1",
"requires_root": true
},
"enabled": true,
@@ -29,7 +29,7 @@
"\\.gz$"
],
"include_files": [],
"processors": "- dissect:\n tokenizer: \"/nsm/import/%{import.id}/evtx/%{import.file}\"\n field: \"log.file.path\"\n target_prefix: \"\"\n- decode_json_fields:\n fields: [\"message\"]\n target: \"\"\n- drop_fields:\n fields: [\"host\"]\n ignore_missing: true\n- add_fields:\n target: data_stream\n fields:\n type: logs\n dataset: system.security\n- add_fields:\n target: event\n fields:\n dataset: system.security\n module: system\n imported: true\n- add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.security-2.20.0\n- if:\n equals:\n winlog.channel: 'Microsoft-Windows-Sysmon/Operational'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: windows.sysmon_operational\n - add_fields:\n target: event\n fields:\n dataset: windows.sysmon_operational\n module: windows\n imported: true\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-windows.sysmon_operational-3.8.3\n- if:\n equals:\n winlog.channel: 'Application'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: system.application\n - add_fields:\n target: event\n fields:\n dataset: system.application\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.application-2.20.0\n- if:\n equals:\n winlog.channel: 'System'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: system.system\n - add_fields:\n target: event\n fields:\n dataset: system.system\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.system-2.20.0\n \n- if:\n equals:\n winlog.channel: 'Microsoft-Windows-PowerShell/Operational'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: windows.powershell_operational\n - add_fields:\n target: event\n fields:\n dataset: windows.powershell_operational\n module: windows\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-windows.powershell_operational-3.8.3\n- add_fields:\n target: data_stream\n fields:\n dataset: import",
"processors": "- dissect:\n tokenizer: \"/nsm/import/%{import.id}/evtx/%{import.file}\"\n field: \"log.file.path\"\n target_prefix: \"\"\n- decode_json_fields:\n fields: [\"message\"]\n target: \"\"\n- drop_fields:\n fields: [\"host\"]\n ignore_missing: true\n- add_fields:\n target: data_stream\n fields:\n type: logs\n dataset: system.security\n- add_fields:\n target: event\n fields:\n dataset: system.security\n module: system\n imported: true\n- add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.security-2.22.3\n- if:\n equals:\n winlog.channel: 'Microsoft-Windows-Sysmon/Operational'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: windows.sysmon_operational\n - add_fields:\n target: event\n fields:\n dataset: windows.sysmon_operational\n module: windows\n imported: true\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-windows.sysmon_operational-3.9.0\n- if:\n equals:\n winlog.channel: 'Application'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: system.application\n - add_fields:\n target: event\n fields:\n dataset: system.application\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.application-2.22.3\n- if:\n equals:\n winlog.channel: 'System'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: system.system\n - add_fields:\n target: event\n fields:\n dataset: system.system\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-system.system-2.22.3\n \n- if:\n equals:\n winlog.channel: 'Microsoft-Windows-PowerShell/Operational'\n then: \n - add_fields:\n target: data_stream\n fields:\n dataset: windows.powershell_operational\n - add_fields:\n target: event\n fields:\n dataset: windows.powershell_operational\n module: windows\n - add_fields:\n target: \"@metadata\"\n fields:\n pipeline: logs-windows.powershell_operational-3.9.0\n- add_fields:\n target: data_stream\n fields:\n dataset: import",
"tags": [
"import"
],
@@ -16,7 +16,6 @@
'awsfirehose.metrics': 'aws.cloudwatch',
'cribl.logs': 'cribl',
'cribl.metrics': 'cribl',
'sentinel_one_cloud_funnel.logins': 'sentinel_one_cloud_funnel.login',
'azure_application_insights.app_insights': 'azure.app_insights',
'azure_application_insights.app_state': 'azure.app_state',
'azure_billing.billing': 'azure.billing',
+19 -9
View File
@@ -68,6 +68,24 @@ so-elastic-fleet-package-upgrade:
- require:
- http: wait_for_so-kibana
# initial so-elasticsearch-templates run is earlier, but it can skip over templates that have component templates not yet installed to avoid elasticsearch rejecting the template.
so-elasticsearch-templates-after-fleet-packages:
cmd.run:
- name: /usr/sbin/so-elasticsearch-templates-load
- cwd: /opt/so
- unless: test -f /opt/so/state/estemplates.txt
- require:
- cmd: so-elastic-fleet-package-upgrade
so-elastic-fleet-integration-upgrade:
cmd.run:
- name: /usr/sbin/so-elastic-fleet-integration-upgrade
- retry:
attempts: 3
interval: 10
- require:
- cmd: so-elastic-fleet-package-upgrade
so-elastic-fleet-integrations:
cmd.run:
- name: /usr/sbin/so-elastic-fleet-integration-policy-load
@@ -86,21 +104,13 @@ so-elastic-agent-grid-upgrade:
- require:
- http: wait_for_so-kibana
so-elastic-fleet-integration-upgrade:
cmd.run:
- name: /usr/sbin/so-elastic-fleet-integration-upgrade
- retry:
attempts: 3
interval: 10
- require:
- http: wait_for_so-kibana
{# Optional integrations script doesn't need the retries like so-elastic-fleet-integration-upgrade which loads the default integrations #}
so-elastic-fleet-addon-integrations:
cmd.run:
- name: /usr/sbin/so-elastic-fleet-optional-integrations-load
- require:
- http: wait_for_so-kibana
- cmd: so-elasticsearch-templates-after-fleet-packages
{% if ELASTICFLEETMERGED.config.defend_filters.enable_auto_configuration %}
so-elastic-defend-manage-filters-file-watch:
@@ -10,6 +10,25 @@
PKG_LOAD_FAILURES=0
PKG_LOAD_FAILURES_NAMES=()
PKG_UPGRADED=0
cleanup_elasticsearch_fleet_transforms() {
local transforms transform_id attempt
if ! transforms=$(so-elasticsearch-query "_transform/logs-elasticsearch.index_pivot-default-*" --retry 1 --retry-delay 5); then
return 0
fi
while IFS= read -r transform_id; do
[ -n "$transform_id" ] || continue
for attempt in {1..3}; do
if so-elasticsearch-query "_transform/$transform_id?force=true" -XDELETE --fail --retry 1 --retry-delay 5; then
break
fi
sleep 5
done
done < <(jq -r '.transforms[]?.id' <<< "$transforms")
}
{%- for PACKAGE in SUPPORTED_PACKAGES %}
if INSTALLED_VERSION=$(elastic_fleet_package_version_check "{{ PACKAGE }}") && LATEST_VERSION=$(elastic_fleet_package_latest_version_check "{{ PACKAGE }}"); then
@@ -17,10 +36,25 @@ if INSTALLED_VERSION=$(elastic_fleet_package_version_check "{{ PACKAGE }}") && L
if [ "$INSTALLED_VERSION" == "$LATEST_VERSION" ]; then
echo "{{ PACKAGE }} integration version $INSTALLED_VERSION is already at the reported latest version $LATEST_VERSION, skipping upgrade."
else
{%- if PACKAGE == 'elasticsearch' %}
cleanup_elasticsearch_fleet_transforms
{%- endif %}
echo "Upgrading {{ PACKAGE }} package from $INSTALLED_VERSION to version $LATEST_VERSION..."
if ! elastic_fleet_package_install "{{ PACKAGE }}" "$LATEST_VERSION"; then
PKG_LOAD_FAILURES=$((PKG_LOAD_FAILURES + 1))
PKG_LOAD_FAILURES_NAMES+=("{{ PACKAGE }}")
# check that package has upgraded to the expected version after install command
elif ! LATEST_VERSION=$(elastic_fleet_package_latest_version_check "{{ PACKAGE }}"); then
echo "ERROR: Failed to get latest version information for integration {{ PACKAGE }} after upgrade attempt"
PKG_LOAD_FAILURES=$((PKG_LOAD_FAILURES + 1))
PKG_LOAD_FAILURES_NAMES+=("{{ PACKAGE }}")
elif INSTALLED_VERSION=$(elastic_fleet_package_version_check "{{ PACKAGE }}") && [ "$INSTALLED_VERSION" == "$LATEST_VERSION" ]; then
echo "{{ PACKAGE }} integration upgraded to version $LATEST_VERSION."
PKG_UPGRADED=$((PKG_UPGRADED + 1))
else
echo "ERROR: {{ PACKAGE }} integration still at ${INSTALLED_VERSION:-unknown}; expected $LATEST_VERSION"
PKG_LOAD_FAILURES=$((PKG_LOAD_FAILURES + 1))
PKG_LOAD_FAILURES_NAMES+=("{{ PACKAGE }}")
fi
fi
else
@@ -30,6 +64,11 @@ else
fi
{%- endfor %}
if [ $PKG_UPGRADED -gt 0 ]; then
echo "Elasticsearch template statefiles cleared after $PKG_UPGRADED package upgrade(s), so templates can reload."
rm -f /opt/so/state/estemplates.txt /opt/so/state/addon_estemplates.txt
fi
if [ $PKG_LOAD_FAILURES -gt 0 ]; then
echo "ERROR: Failed to upgrade $PKG_LOAD_FAILURES package(s):"
for PKG in "${PKG_LOAD_FAILURES_NAMES[@]}"; do
+7
View File
@@ -98,6 +98,13 @@ so-es-cluster-settings:
- docker_container: so-elasticsearch
- file: elasticsearch_sbin_jinja
- http: wait_for_so-elasticsearch
so-elasticsearch-system-indices-patch:
cmd.run:
- name: /usr/sbin/so-elasticsearch-system-indices-patch
- require:
- http: wait_for_so-elasticsearch
- file: so-elasticsearch-system-indices-patch-script
{% endif %}
# heavynodes will only load ILM policies for SO managed indices. (Indicies defined in elasticsearch/defaults.yaml)
+10
View File
@@ -42,6 +42,16 @@ elasticsearch_sbin:
- file_mode: 755
- exclude_pat:
- so-elasticsearch-pipelines # exclude this because we need to watch it for changes, we sync it in another state
- so-elasticsearch-system-indices-patch
- show_changes: False
so-elasticsearch-system-indices-patch-script:
file.managed:
- name: /usr/sbin/so-elasticsearch-system-indices-patch
- source: salt://elasticsearch/tools/sbin/so-elasticsearch-system-indices-patch
- user: 930
- group: 939
- mode: 755
- show_changes: False
elasticsearch_sbin_jinja:
+1 -1
View File
@@ -1,7 +1,7 @@
elasticsearch:
enabled: false
esheap: '600m'
version: 9.3.7
version: 9.4.5
index_clean: true
data_retention_method: DLM
vm:
@@ -0,0 +1,199 @@
#!/bin/bash
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
set -eo pipefail
SETTINGS='{"index":{"auto_expand_replicas":"0-1"}}'
KIBANA_PASSWORD=
INDEX_PATTERNS=(
'.entity_analytics.risk_score.lookup-*'
'.entity_analytics.watchlists.*'
'.entity_analytics.monitoring.users-*'
'.entity_analytics.entity-leads-*'
'.asset-criticality.asset-criticality-*'
'.workflows-executions'
'.workflows-step-executions'
'.entities.v2.latest.security_*'
'.entities.v2.history.security_*'
'risk-score.risk-score-latest-*'
)
DATA_STREAM_PATTERNS=(
'.entities.v2.updates.security_*'
'risk-score.risk-score-*'
'.rule-events'
'.alert-actions'
)
TEMPLATE_PATTERNS=(
'entities_v2_latest_security_default_index_template'
'entities_v2_history_security_default_index_template'
'.entities_v2_updates_security_default_index_template'
'.risk-score.risk-score-default-index-template'
'.rule-events'
'.alert-actions'
)
query_es() {
if so-elasticsearch-query "$@" --fail --retry 3 --retry-delay 5; then
return 0
fi
# retry failed attempts with so_kibana user (system managed indices reject so_elastic user)
local query_path="$1"
shift
if [[ -z "$KIBANA_PASSWORD" ]]; then
KIBANA_PASSWORD=$(salt-call pillar.get elasticsearch:auth:users:so_kibana_user:pass --out=newline_values_only)
fi
[[ -n "$KIBANA_PASSWORD" ]] || return 1
echo "Retrying ${query_path} as so_kibana." >&2
curl -K /opt/so/conf/elasticsearch/curl.config --user "so_kibana:${KIBANA_PASSWORD}" \
-s -k -L --fail --retry 3 --retry-delay 5 -H 'Content-Type: application/json' "https://localhost:9200/${query_path}" "$@"
}
# add auto_expand_replicas=0-1 to given index
set_auto_expand_replicas() {
local index="$1"
echo "Setting auto_expand_replicas to 0-1 on ${index}."
query_es "${index}/_settings" -XPUT -d "$SETTINGS" >/dev/null
}
# resolve index patterns and find each index with an unassigned replica
unassigned_replicas() {
local pattern="$1"
local resolved_indices response index
if ! resolved_indices=$(query_es "_resolve/index/${pattern}?expand_wildcards=all" 2>/dev/null); then
return 0
fi
while read -r index; do
if ! response=$(query_es "_cat/shards/${index}?format=json&h=index,prirep,state" 2>/dev/null); then
continue
fi
jq -r '.[]? | objects | select(.prirep == "r" and .state == "UNASSIGNED") | .index' <<<"$response"
done < <(jq -r '.indices[]?.name' <<<"$resolved_indices")
}
data_stream_indices() {
local pattern="$1"
local response
if ! response=$(query_es "_data_stream/${pattern}?expand_wildcards=all" 2>/dev/null); then
return 0
fi
jq -r '.data_streams[]?.indices[]?.index_name' <<<"$response"
}
update_system_indices() {
local pattern="$1"
local index
while read -r index; do
[[ -n "$index" ]] && set_auto_expand_replicas "$index"
done < <(unassigned_replicas "$pattern")
}
# update data stream backing indices with unassigned replicas
update_system_ds() {
local pattern="$1"
local index
while read -r index; do
while read -r unassigned_index; do
[[ -n "$unassigned_index" ]] && set_auto_expand_replicas "$unassigned_index"
done < <(unassigned_replicas "$index")
done < <(data_stream_indices "$pattern")
}
has_unassigned_replicas() {
local pattern="$1"
local index
index=$(unassigned_replicas "$pattern" | sed -n '1p')
[[ -n "$index" ]]
}
data_stream_has_unassigned_replicas() {
local pattern="$1"
local index
while read -r index; do
has_unassigned_replicas "$index" && return 0
done < <(data_stream_indices "$pattern")
return 1
}
needs_patch() {
local pattern
for pattern in "${INDEX_PATTERNS[@]}"; do
has_unassigned_replicas "$pattern" && return 0
done
for pattern in "${DATA_STREAM_PATTERNS[@]}"; do
data_stream_has_unassigned_replicas "$pattern" && return 0
done
return 1
}
# get index templates, update with auto_expand_replicas=0-1, and PUT back. Keeping mappings/settings/aliases in-place
update_system_templates() {
local pattern="$1"
local templates name response template auto_expand_replicas
if ! templates=$(query_es "_index_template/${pattern}" 2>/dev/null); then
return 0
fi
while read -r name; do
response=$(query_es "_index_template/${name}")
template=$(jq -c '.index_templates[0].index_template' <<<"$response")
auto_expand_replicas=$(jq -r '.template.settings["index.auto_expand_replicas"] // .template.settings.index.auto_expand_replicas // empty' <<<"$template")
[[ "$auto_expand_replicas" == "0-1" ]] && continue
template=$(jq '
if (.template.settings.index | type) == "object" then
.template.settings.index.auto_expand_replicas = "0-1"
else
.template.settings["index.auto_expand_replicas"] = "0-1"
end
| del(.created_date_millis, .modified_date_millis)
' <<<"$template")
echo "Setting auto_expand_replicas to 0-1 on index template ${name}."
query_es "_index_template/${name}" -XPUT -d "$template" >/dev/null
done < <(jq -r '.index_templates[]?.name' <<<"$templates")
}
if [[ "${1:-}" == "--check" ]]; then
needs_patch
exit $?
fi
if [[ $# -ne 0 ]]; then
echo "Usage: $0 [--check]" >&2
exit 1
fi
patched=false
for pattern in "${INDEX_PATTERNS[@]}"; do
if has_unassigned_replicas "$pattern"; then
update_system_indices "$pattern"
patched=true
fi
done
for pattern in "${DATA_STREAM_PATTERNS[@]}"; do
if data_stream_has_unassigned_replicas "$pattern"; then
update_system_ds "$pattern"
patched=true
fi
done
if [[ "$patched" == true ]]; then
for pattern in "${TEMPLATE_PATTERNS[@]}"; do
update_system_templates "$pattern"
done
fi
+1 -1
View File
@@ -22,7 +22,7 @@ kibana:
- default
- file
migrations:
discardCorruptObjects: "9.3.7"
discardCorruptObjects: "9.4.5"
telemetry:
enabled: False
xpack:
@@ -9,5 +9,5 @@ SESSIONCOOKIE=$(curl -K /opt/so/conf/elasticsearch/curl.config -c - -X GET http:
# Disable certain Features from showing up in the Kibana UI
echo
echo "Setting up default Kibana Space:"
curl -K /opt/so/conf/elasticsearch/curl.config -b "sid=$SESSIONCOOKIE" -L -X PUT "localhost:5601/api/spaces/space/default" -H 'kbn-xsrf: true' -H 'Content-Type: application/json' -d' {"id":"default","name":"Default","disabledFeatures":["ml","enterpriseSearch","logs","infrastructure","apm","uptime","monitoring","stackAlerts","actions","securitySolutionCasesV3","inventory","dataQuality","searchSynonyms","searchQueryRules","enterpriseSearchApplications","enterpriseSearchAnalytics","securitySolutionTimeline","securitySolutionNotes","securitySolutionRulesV1","entityManager","streams","cloudConnect","slo"]} ' >> /opt/so/log/kibana/misc.log
curl -K /opt/so/conf/elasticsearch/curl.config -b "sid=$SESSIONCOOKIE" -L -X PUT "localhost:5601/api/spaces/space/default" -H 'kbn-xsrf: true' -H 'Content-Type: application/json' -d' {"id":"default","name":"Default","disabledFeatures":["ml","enterpriseSearch","logs","infrastructure","apm","uptime","securitySolutionCasesV3","inventory","searchSynonyms","searchQueryRules","enterpriseSearchApplications","enterpriseSearchAnalytics","securitySolutionTimeline","securitySolutionNotes","securitySolutionRulesV4","securitySolutionAlertsV1","entityManager","slo","streams","anonymization","searchInferenceEndpoints","cloudConnect","queryActivity","automatic_import","stackAlerts","monitoring","dataQuality","actions"]} ' >> /opt/so/log/kibana/misc.log
echo
+20
View File
@@ -220,6 +220,26 @@ logrotate:
- extension .log
- dateext
- dateyesterday
/opt/so/log/salt/virtual_node_manager:
- daily
- rotate 14
- missingok
- copytruncate
- compress
- create
- extension .log
- dateext
- dateyesterday
/opt/so/log/salt/so-salt-cloud:
- daily
- rotate 14
- missingok
- copytruncate
- compress
- create
- extension .log
- dateext
- dateyesterday
/opt/so/log/salt/so-soup-grid-highstate:
- daily
- rotate 14
+14
View File
@@ -140,6 +140,20 @@ logrotate:
multiline: True
global: True
forcedType: "[]string"
"/opt/so/log/salt/virtual_node_manager":
description: List of logrotate options for this file.
title: /opt/so/log/salt/virtual_node_manager
advanced: True
multiline: True
global: True
forcedType: "[]string"
"/opt/so/log/salt/so-salt-cloud":
description: List of logrotate options for this file.
title: /opt/so/log/salt/so-salt-cloud
advanced: True
multiline: True
global: True
forcedType: "[]string"
"/opt/so/log/salt/so-soup-grid-highstate":
description: List of logrotate options for this file.
title: /opt/so/log/salt/so-soup-grid-highstate
+23 -3
View File
@@ -81,6 +81,14 @@ ls_custom_pipeline_conf_{{assigned_pipeline}}_{{pipeline}}:
{% for assigned_pipeline in ASSIGNED_PIPELINES %}
{# a blank per-pipeline setting falls back to the global logstash.yml value #}
{% set PARSED_OVERRIDES = LOGSTASH_MERGED.get('pipeline_settings', {}).get(assigned_pipeline, {}) %}
{% if PARSED_OVERRIDES is not mapping %}
{% do salt.log.warning('logstash: ignoring malformed pipeline_settings for pipeline ' ~ assigned_pipeline ~ '; expected a set of settings') %}
{% endif %}
{% set PIPELINE_OVERRIDES = PARSED_OVERRIDES if PARSED_OVERRIDES is mapping else {} %}
{% set THREADS = PIPELINE_OVERRIDES.get('pipeline_x_workers') or LOGSTASH_MERGED.config.pipeline_x_workers %}
{% set BATCH = PIPELINE_OVERRIDES.get('pipeline_x_batch_x_size') or LOGSTASH_MERGED.config.pipeline_x_batch_x_size %}
{% for CONFIGFILE in LOGSTASH_MERGED.defined_pipelines[assigned_pipeline] %}
ls_pipeline_{{assigned_pipeline}}_{{CONFIGFILE.split('.')[0] | replace("/","_") }}:
file.managed:
@@ -92,8 +100,8 @@ ls_pipeline_{{assigned_pipeline}}_{{CONFIGFILE.split('.')[0] | replace("/","_")
GLOBALS: {{ GLOBALS }}
ES_USER: "{{ salt['pillar.get']('elasticsearch:auth:users:so_elastic_user:user', '') }}"
ES_PASS: "{{ salt['pillar.get']('elasticsearch:auth:users:so_elastic_user:pass', '') }}"
THREADS: {{ LOGSTASH_MERGED.config.pipeline_x_workers }}
BATCH: {{ LOGSTASH_MERGED.config.pipeline_x_batch_x_size }}
THREADS: {{ THREADS }}
BATCH: {{ BATCH }}
{% else %}
- name: /opt/so/conf/logstash/pipelines/{{assigned_pipeline}}/{{CONFIGFILE.split('/')[1]}}
{% endif %}
@@ -125,6 +133,14 @@ lspipelinesyml:
- defaults:
ASSIGNED_PIPELINES: {{ ASSIGNED_PIPELINES }}
lslog4j2:
file.managed:
- name: /opt/so/conf/logstash/etc/log4j2.properties
- source: salt://logstash/etc/log4j2.properties.jinja
- template: jinja
- user: 931
- group: 939
lsetcsync:
file.recurse:
- name: /opt/so/conf/logstash/etc
@@ -133,7 +149,11 @@ lsetcsync:
- group: 939
- template: jinja
- clean: True
- exclude_pat: pipelines*
{#- both names are matched: the .jinja source so the recurse does not copy it verbatim,
and the rendered file so clean: True does not delete what lslog4j2 wrote #}
- exclude_pat:
- pipelines*
- log4j2.properties*
- defaults:
LOGSTASH_MERGED: {{ LOGSTASH_MERGED }}
+400
View File
@@ -42,6 +42,11 @@ logstash:
custom2: []
custom3: []
custom4: []
custom5: []
custom6: []
custom7: []
custom8: []
custom9: []
pipeline_config:
custom001: |-
filter {
@@ -60,10 +65,405 @@ logstash:
custom008: PLACEHOLDER
custom009: PLACEHOLDER
custom010: PLACEHOLDER
pipeline_settings:
fleet:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
manager:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
receiver:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
search:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom0:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom1:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom2:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom3:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom4:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom5:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom6:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom7:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom8:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
custom9:
pipeline_x_workers: ''
pipeline_x_batch_x_size: ''
pipeline_x_batch_x_delay: ''
pipeline_x_batch_x_metrics_x_sampling_mode: ''
pipeline_x_ordered: ''
pipeline_x_ecs_compatibility: ''
pipeline_x_reloadable: ''
queue_x_type: ''
queue_x_max_bytes: ''
queue_x_page_capacity: ''
queue_x_max_events: ''
queue_x_checkpoint_x_acks: ''
queue_x_checkpoint_x_writes: ''
queue_x_checkpoint_x_interval: ''
queue_x_checkpoint_x_retry: ''
queue_x_compression: ''
queue_x_drain: ''
dead_letter_queue_x_enable: ''
dead_letter_queue_x_max_bytes: ''
dead_letter_queue_x_flush_interval: ''
dead_letter_queue_x_flush_check_interval: ''
dead_letter_queue_x_storage_policy: ''
dead_letter_queue_x_retain_x_age: ''
path_x_queue: ''
path_x_dead_letter_queue: ''
config_x_debug: ''
config_x_support_escapes: ''
settings:
lsheap: 500m
config:
api_x_http_x_host: 0.0.0.0
log_x_level: info
log_x_format: plain
path_x_logs: /var/log/logstash
pipeline_x_workers: 1
pipeline_x_batch_x_size: 125
+2
View File
@@ -105,6 +105,8 @@ so-logstash:
{% endif %}
- watch:
- file: lsetcsync
- file: lslog4j2
- file: lspipelinesyml
- file: trusttheca
{% if GLOBALS.is_manager %}
- file: elasticsearch_cacerts
@@ -1,3 +1,4 @@
{%- from 'logstash/map.jinja' import LOGSTASH_MERGED -%}
status = error
name = LogstashPropertiesConfig
@@ -16,8 +17,14 @@ name = LogstashPropertiesConfig
appender.rolling.type = RollingFile
appender.rolling.name = rolling
appender.rolling.fileName = /var/log/logstash/logstash.log
{%- if LOGSTASH_MERGED.config.get('log_x_format', 'plain') == 'json' %}
appender.rolling.layout.type = JSONLayout
appender.rolling.layout.compact = true
appender.rolling.layout.eventEol = true
{%- else %}
appender.rolling.layout.type = PatternLayout
appender.rolling.layout.pattern = [%d{ISO8601}][%-5p][%-25c] %.10000m%n
{%- endif %}
appender.rolling.filePattern = /var/log/logstash/logstash-%d{yyyy-MM-dd}.log.gz
appender.rolling.policies.type = Policies
appender.rolling.policies.time.type = TimeBasedTriggeringPolicy
@@ -32,7 +39,5 @@ appender.rolling.strategy.action.condition.type = IfFileName
appender.rolling.strategy.action.condition.glob = *.gz
appender.rolling.strategy.action.condition.nested_condition.type = IfLastModified
appender.rolling.strategy.action.condition.nested_condition.age = 7D
rootLogger.level = info
rootLogger.level = ${sys:ls.log.level}
rootLogger.appenderRef.rolling.ref = rolling
#rootLogger.level = ${sys:ls.log.level}
#rootLogger.appenderRef.console.ref = ${sys:ls.log.format}_console
+13
View File
@@ -1,4 +1,17 @@
{%- from 'logstash/map.jinja' import LOGSTASH_MERGED %}
{%- set PIPELINE_SETTINGS = LOGSTASH_MERGED.get('pipeline_settings', {}) %}
{%- for assigned_pipeline in ASSIGNED_PIPELINES %}
- pipeline.id: {{ assigned_pipeline }}
path.config: "/usr/share/logstash/pipelines/{{ assigned_pipeline }}/"
{%- set extra = PIPELINE_SETTINGS.get(assigned_pipeline, {}) %}
{%- if extra is mapping %}
{#- values are emitted unquoted so yaml re-infers the type logstash expects:
4 as an integer, false as a boolean, 1024mb and auto as strings #}
{%- for key, value in extra | dictsort %}
{%- set rendered = key | replace('_x_', '.') %}
{%- if value not in ['', None] and rendered not in ['pipeline.id', 'path.config'] %}
{{ rendered }}: {{ value }}
{%- endif %}
{%- endfor %}
{%- endif %}
{% endfor -%}
+380
View File
@@ -16,6 +16,7 @@ logstash:
heavynode: *assigned_pipelines
searchnode: *assigned_pipelines
manager: *assigned_pipelines
managerhype: *assigned_pipelines
managersearch: *assigned_pipelines
fleet: *assigned_pipelines
defined_pipelines:
@@ -34,6 +35,11 @@ logstash:
custom2: *defined_pipelines
custom3: *defined_pipelines
custom4: *defined_pipelines
custom5: *defined_pipelines
custom6: *defined_pipelines
custom7: *defined_pipelines
custom8: *defined_pipelines
custom9: *defined_pipelines
pipeline_config:
custom001: &pipeline_config
description: Pipeline configuration for Logstash
@@ -51,6 +57,351 @@ logstash:
custom008: *pipeline_config
custom009: *pipeline_config
custom010: *pipeline_config
pipeline_settings:
manager: &pipeline_settings
pipeline_x_workers:
description: >-
Number of worker threads that run filters and outputs for this pipeline. May be set higher
than the CPU core count when outputs spend time waiting on I/O. Leave blank to use the value
from logstash.yml.
title: pipeline.workers
regex: '^$|^[1-9][0-9]*$'
regexFailureMessage: Must be blank, or a positive whole number.
advanced: True
global: False
helpLink: logstash
pipeline_x_batch_x_size:
description: >-
Maximum number of events an individual worker thread collects before running filters and
outputs. Larger batches are more efficient but increase heap use; total in-flight events is
workers multiplied by batch size. Leave blank to use the value from logstash.yml.
title: pipeline.batch.size
regex: '^$|^[1-9][0-9]*$'
regexFailureMessage: Must be blank, or a positive whole number.
advanced: True
global: False
helpLink: logstash
pipeline_x_batch_x_delay:
description: >-
Milliseconds a worker waits for the next event before running a batch that is not yet full.
Leave blank to use the value from logstash.yml.
title: pipeline.batch.delay
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
pipeline_x_batch_x_metrics_x_sampling_mode:
description: >-
Controls how often batch size metrics are collected for this pipeline, which helps tune
pipeline.batch.size to the batch sizes actually being processed. Fuller sampling consumes
additional heap. Elastic marks this setting as a technical preview that may change in a
future release. Leave blank to use the value from logstash.yml.
title: pipeline.batch.metrics.sampling_mode
options:
- ''
- 'disabled'
- 'minimal'
- 'full'
advanced: True
global: False
helpLink: logstash
pipeline_x_ordered:
description: >-
Whether event order is preserved through this pipeline. auto enables ordering only when
pipeline.workers is explicitly set to 1, and does nothing otherwise. Setting this to true
requires pipeline.workers to be 1 as well; with more workers this pipeline fails to start.
Leave blank to use the value from logstash.yml.
title: pipeline.ordered
options:
- ''
- 'auto'
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
pipeline_x_ecs_compatibility:
description: >-
Elastic Common Schema compatibility mode for plugins in this pipeline. Security Onion sets
this globally and it should rarely be changed per pipeline. Elastic considers values other
than disabled to be BETA, and they may produce unintended consequences when upgrading
Logstash. Leave blank to use the value from logstash.yml.
title: pipeline.ecs_compatibility
options:
- ''
- 'disabled'
- 'v1'
- 'v8'
advanced: True
global: False
helpLink: logstash
pipeline_x_reloadable:
description: >-
Whether this pipeline may be reloaded when its configuration changes. Leave blank to use the
value from logstash.yml.
title: pipeline.reloadable
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
queue_x_type:
description: >-
Queue backing this pipeline. persisted buffers events to disk under /nsm/logstash so they
survive a restart, at some throughput cost; memory does not. Leave blank to use the value
from logstash.yml.
title: queue.type
options:
- ''
- 'memory'
- 'persisted'
advanced: True
global: False
helpLink: logstash
queue_x_max_bytes:
description: >-
Total capacity of the persistent queue for this pipeline, in bytes. Only applies when
queue.type is persisted. The disk backing /nsm/logstash must have room for this much data or
the pipeline fails to start, reporting that it was unable to allocate the space. If both
queue.max_events and queue.max_bytes are set, whichever is reached first applies. Leave
blank to use the value from logstash.yml.
title: queue.max_bytes
regex: '^$|^[0-9]+$|^[0-9]+(\.[0-9]+)?\s*(b|kb?|mb?|gb?|tb?|pb?)$'
regexFailureMessage: Must be blank, or a size such as 512mb, 1gb, or 64k. Units are lowercase.
advanced: True
global: False
helpLink: logstash
queue_x_page_capacity:
description: >-
Size of the individual append-only page data files that make up the persistent queue for
this pipeline. Only applies when queue.type is persisted. Leave blank to use the value from
logstash.yml.
title: queue.page_capacity
regex: '^$|^[0-9]+$|^[0-9]+(\.[0-9]+)?\s*(b|kb?|mb?|gb?|tb?|pb?)$'
regexFailureMessage: Must be blank, or a size such as 512mb, 1gb, or 64k. Units are lowercase.
advanced: True
global: False
helpLink: logstash
queue_x_max_events:
description: >-
Maximum number of unread events in the persistent queue for this pipeline. 0 means
unlimited. Only applies when queue.type is persisted. Leave blank to use the value from
logstash.yml.
title: queue.max_events
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
queue_x_checkpoint_x_acks:
description: >-
Maximum number of acknowledged events before a checkpoint is forced. 0 means unlimited. Only
applies when queue.type is persisted. Leave blank to use the value from logstash.yml.
title: queue.checkpoint.acks
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
queue_x_checkpoint_x_writes:
description: >-
Maximum number of written events before a checkpoint is forced. Setting this to 1 gives
maximum durability at a severe performance cost. 0 means unlimited. Only applies when
queue.type is persisted. Leave blank to use the value from logstash.yml.
title: queue.checkpoint.writes
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
queue_x_checkpoint_x_interval:
description: >-
Milliseconds between forced checkpoints on the persistent queue head page. 0 eliminates
periodic checkpoints. Deprecated by Elastic as of Logstash 9.1. Only applies when queue.type
is persisted. Leave blank to use the value from logstash.yml.
title: queue.checkpoint.interval
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
queue_x_checkpoint_x_retry:
description: >-
When enabled, Logstash retries four times per attempted checkpoint write that fails; later
errors are not retried. Elastic describes this as a workaround for failed checkpoint writes
seen only on Windows and on filesystems with non-standard behaviour such as SANs, and does
not recommend enabling it otherwise. Only applies when queue.type is persisted. Leave blank
to use the value from logstash.yml.
title: queue.checkpoint.retry
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
queue_x_compression:
description: >-
Compression applied to persistent queue pages for this pipeline, trading CPU for disk: speed
favours the fastest operation, size the smallest files, and balanced sits between them. Once
compressed events have been written, that queue cannot be read by Logstash releases earlier
than 9.2. Only applies when queue.type is persisted. Leave blank to use the value from
logstash.yml.
title: queue.compression
options:
- ''
- 'none'
- 'speed'
- 'balanced'
- 'size'
advanced: True
global: False
helpLink: logstash
queue_x_drain:
description: >-
When enabled, Logstash waits for the persistent queue to drain before shutting down this
pipeline. Draining a large queue makes shutdown take considerably longer. Only applies when
queue.type is persisted. Leave blank to use the value from logstash.yml.
title: queue.drain
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_enable:
description: >-
Whether events this pipeline cannot process are written to a dead letter queue instead of
being dropped. Leave blank to use the value from logstash.yml.
title: dead_letter_queue.enable
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_max_bytes:
description: >-
Total capacity of the dead letter queue for this pipeline, in bytes. Only applies when
dead_letter_queue.enable is true. Leave blank to use the value from logstash.yml.
title: dead_letter_queue.max_bytes
regex: '^$|^[0-9]+$|^[0-9]+(\.[0-9]+)?\s*(b|kb?|mb?|gb?|tb?|pb?)$'
regexFailureMessage: Must be blank, or a size such as 512mb, 1gb, or 64k. Units are lowercase.
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_flush_interval:
description: >-
Milliseconds before an incomplete dead letter queue segment is flushed and made available to
the dead_letter_queue input. Lower values write more, smaller segment files; higher values
add latency before events can be read. Only applies when dead_letter_queue.enable is true.
Leave blank to use the value from logstash.yml.
title: dead_letter_queue.flush_interval
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_flush_check_interval:
description: >-
Milliseconds between checks for a stale dead letter queue segment needing a flush. Cannot be
set lower than 1000. Smaller values rotate segments sooner at the cost of CPU. Only applies
when dead_letter_queue.enable is true. Leave blank to use the value from logstash.yml.
title: dead_letter_queue.flush_check_interval
regex: '^$|^[0-9]+$'
regexFailureMessage: Must be blank, or a whole number.
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_storage_policy:
description: >-
Action taken when dead_letter_queue.max_bytes is reached: drop_newer stops accepting new
events, drop_older removes the oldest events to make room. Only applies when
dead_letter_queue.enable is true. Leave blank to use the value from logstash.yml.
title: dead_letter_queue.storage_policy
options:
- ''
- 'drop_newer'
- 'drop_older'
advanced: True
global: False
helpLink: logstash
dead_letter_queue_x_retain_x_age:
description: >-
How long an event is kept in the dead letter queue before Logstash removes it, such as 5d.
Units are d, h, m and s; there is no default unit, so one must be given. Only applies when
dead_letter_queue.enable is true. Leave blank to use the value from logstash.yml.
title: dead_letter_queue.retain.age
regex: '^$|^[0-9]+\s*[dhms]$'
regexFailureMessage: Must be blank, or a number followed by d, h, m, or s, such as 5d.
advanced: True
global: False
helpLink: logstash
path_x_queue:
description: >-
Directory inside the Logstash container holding the persistent queue for this pipeline. The
default lives under the /nsm/logstash bind mount; a path outside it will not survive a
container restart. Logstash creates the directory if it is missing, requires it to be
writable, and refuses to start if the path is a symlink. Only applies when queue.type is
persisted. Leave blank to use the value from logstash.yml.
title: path.queue
advanced: True
global: False
helpLink: logstash
path_x_dead_letter_queue:
description: >-
Directory inside the Logstash container holding the dead letter queue for this pipeline. The
default lives under the /nsm/logstash bind mount; a path outside it will not survive a
container restart. Logstash creates the directory if it is missing, requires it to be
writable, and refuses to start if the path is a symlink. Only applies when
dead_letter_queue.enable is true. Leave blank to use the value from logstash.yml.
title: path.dead_letter_queue
advanced: True
global: False
helpLink: logstash
config_x_debug:
description: >-
Whether the fully compiled configuration for this pipeline is written to the log. The output
may contain sensitive values from the pipeline configuration. Leave blank to use the value
from logstash.yml.
title: config.debug
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
config_x_support_escapes:
description: >-
Whether escape sequences such as \n and \t in this pipeline's quoted strings are
interpreted. Leave blank to use the value from logstash.yml.
title: config.support_escapes
options:
- ''
- 'true'
- 'false'
advanced: True
global: False
helpLink: logstash
fleet: *pipeline_settings
receiver: *pipeline_settings
search: *pipeline_settings
custom0: *pipeline_settings
custom1: *pipeline_settings
custom2: *pipeline_settings
custom3: *pipeline_settings
custom4: *pipeline_settings
custom5: *pipeline_settings
custom6: *pipeline_settings
custom7: *pipeline_settings
custom8: *pipeline_settings
custom9: *pipeline_settings
settings:
lsheap:
description: Heap size to use for logstash
@@ -62,6 +413,35 @@ logstash:
helpLink: logstash
readonly: True
advanced: True
log_x_level:
description: >-
Verbosity of the Logstash log at /opt/so/log/logstash/logstash.log. debug and trace produce
a very large volume of log data on a busy node and should be used only while troubleshooting;
the log rotates at 1GB and rotated files are deleted after 7 days. Setting this to debug is
also what makes the per-pipeline config.debug setting emit anything.
title: log.level
options:
- 'fatal'
- 'error'
- 'warn'
- 'info'
- 'debug'
- 'trace'
advanced: True
global: False
helpLink: logstash
log_x_format:
description: >-
Layout of the Logstash log. plain writes human readable lines; json writes one JSON object
per line, which is easier to parse but harder to read directly. The file name and location
do not change.
title: log.format
options:
- 'plain'
- 'json'
advanced: True
global: False
helpLink: logstash
path_x_logs:
description: Path inside the container to wrote logs.
helpLink: logstash
@@ -3,9 +3,16 @@ beacons:
postgres_pillar_beacon:
- interval: {{ AUTOAPPLY.drain_interval }}
- disable_during_state_run: False
rules_beacon:
local_files_beacon:
- interval: {{ AUTOAPPLY.drain_interval }}
- disable_during_state_run: False
# Tags are app names in salt/reactor/pillar_push_map.yaml.
# Allowlist on purpose: salt writes elsewhere under local/salt/ and would self-retrigger.
- paths:
/opt/so/saltstack/local/salt/suricata/rules: suricata
/opt/so/saltstack/local/salt/strelka/rules/compiled: strelka
/opt/so/saltstack/local/salt/zeek/policy: zeek
/opt/so/saltstack/local/salt/zeek/zkg: zeek
/opt/so/saltstack/local/salt/elasticsearch/files/ingest: elasticsearch
/opt/so/saltstack/local/salt/elasticsearch/roles: elasticsearch
/opt/so/saltstack/local/salt/logstash/pipelines/config/custom: logstash
+1 -1
View File
@@ -20,7 +20,7 @@ is older than debounce_seconds, this script:
with the deduped actions list passed as pillar kwargs
* deletes the contributed intent files on successful dispatch
Reactor sls files (push_suricata, push_strelka, push_pillar) write intents
Reactor sls files (push_files, push_pillar) write intents
but never dispatch directly
"""
+39 -10
View File
@@ -455,6 +455,19 @@ highstate() {
salt-call state.highstate -l info queue=True
}
upgrade_searchnode_elasticsearch() {
# Run the elasticsearch state across the true elastic cluster (non-heavy) with a retry attempt
# Excludes the manager, so that kibana & elasticfleet are not upgraded until searchnodes are upgraded.
echo "Getting ready to upgrade Elasticsearch across the grid. This may take a while..."
if salt -b 10% -C "I@elasticsearch:enabled and not G@role:so-heavynode and not G@id:${MINIONID}" state.apply elasticsearch queue=True; then
return 0
fi
echo "Initial elasticsearch state attempt had a problem; retrying in 30 seconds."
sleep 30
salt -b 10% -C "I@elasticsearch:enabled and not G@role:so-heavynode and not G@id:${MINIONID}" state.apply elasticsearch queue=True
}
push_grid_highstate() {
# Drive a batched, role-tiered highstate across the rest of the grid so remote minions
# pick up this upgrade now instead of waiting up to ~2.5 hours for their own scheduled
@@ -483,11 +496,10 @@ push_grid_highstate() {
masterlock() {
echo "Locking Salt Master"
mv -v $TOPFILE $BACKUPTOPFILE
# Render the real top file only for the host running soup; every other
# minion gets an empty top (no states) while the master is upgrading.
echo "{% if grains['id'] == '$MINIONID' %}" > $TOPFILE
cat $BACKUPTOPFILE >> $TOPFILE
echo "{% endif %}" >> $TOPFILE
echo "base:" > $TOPFILE
echo " $MINIONID:" >> $TOPFILE
echo " - ca" >> $TOPFILE
echo " - elasticsearch" >> $TOPFILE
}
masterunlock() {
@@ -1337,11 +1349,12 @@ verify_es_version_compatibility() {
local is_active_intermediate_upgrade=1
# supported upgrade paths for SO-ES versions
declare -A es_upgrade_map=(
["8.18.4"]="8.18.6 8.18.8 9.0.8"
["8.18.4"]="8.18.6 8.18.8 9.0.8"
["8.18.6"]="8.18.8 9.0.8"
["8.18.8"]="9.0.8"
["9.0.8"]="9.3.3 9.3.7"
["9.3.3"]="9.3.7"
["9.0.8"]="9.3.3 9.3.7 9.4.5"
["9.3.3"]="9.3.7 9.4.5"
["9.3.7"]="9.4.5"
)
# Elasticsearch MUST upgrade through these versions
@@ -2117,12 +2130,28 @@ main() {
# ensure the mine is updated and populated before highstates run, following the salt-master restart
update_salt_mine
# kick off a searchnode elasticsearch upgrade
set +e
if [[ "$es_version" != "$target_es_version" ]]; then
if salt-key -L accepted | grep -q "_searchnode$" 2>/dev/null; then
# only run if there is atleast 1 searchnode
upgrade_searchnode_elasticsearch
fi
fi
set -e
highstate
check_saltmaster_status
postupgrade_changes
[[ $is_airgap -eq 0 ]] && unmount_update
if [[ "$es_version" != "$target_es_version" ]]; then
# Run final elasticsearch / fleet state on manager to ensure addon index templates are created/regenerated and loaded
echo "Running final Elastic states at $(date +"%T.%6N"), after upgrade to $NEWVERSION"
salt-call state.apply elasticsearch,elasticfleet queue=True
fi
echo ""
echo "Upgrade to $NEWVERSION complete."
+1 -1
View File
@@ -260,7 +260,7 @@ http {
}
{% if 'api' in salt['pillar.get']('features', []) %}
location ~* (^/oauth2/token.*|^.well-known/jwks.json|^.well-known/openid-configuration) {
location ~* (^/oauth2/token.*|^/\.well-known/jwks.json|^/\.well-known/openid-configuration) {
limit_req zone=auth_throttle burst={{ NGINXMERGED.config.throttle_login_burst }} nodelay;
limit_req_status 429;
proxy_pass http://{{ GLOBALS.manager }}:4444;
+2
View File
@@ -29,6 +29,8 @@ psql -v ON_ERROR_STOP=1 --username "$POSTGRES_USER" --dbname "$POSTGRES_DB" <<-E
-- revoking CONNECT closes the soft edge entirely.
REVOKE CONNECT ON DATABASE "$POSTGRES_DB" FROM PUBLIC;
GRANT CONNECT ON DATABASE "$POSTGRES_DB" TO "$SO_POSTGRES_USER";
CREATE EXTENSION IF NOT EXISTS vector;
EOSQL
# Bootstrap the Telegraf metrics database. Per-minion roles + schemas are
+3
View File
@@ -1,3 +1,6 @@
# Read by push_pillar.sls (SOC config saves) and push_files.sls (local/salt file edits);
# both key on the app name. An app missing here waits for the next scheduled highstate.
#
# One pillar directory can map to multiple (state, tgt) actions.
# tgt is a raw salt compound expression. tgt_type is always "compound".
# Per-action `batch` / `batch_wait` override the orch defaults (25% / 15s).
+131
View File
@@ -0,0 +1,131 @@
#!py
# Reactor invoked by local_files_beacon when a watched directory under
# /opt/so/saltstack/local/salt/ changes. The beacon tag is an app name in
# pillar_push_map.yaml, so file changes and pillar changes route through the same
# table -- see salt/reactor/push_pillar.sls.
#
# The app comes from the event tag, not the payload: salt's beacon loop pops the
# beacon's 'tag' key off the data and appends it to the event tag instead (see
# salt/beacons/__init__.py). The reactor renderer sets both `tag` and `data` as
# module globals.
#
# Reactors never dispatch directly. The so-push-drainer schedule picks up ready
# intents, dedupes across pending files, and dispatches orch.push_batch.
import fcntl
import json
import logging
import os
import time
from salt.client import Caller
import yaml
LOG = logging.getLogger(__name__)
PENDING_DIR = '/opt/so/state/push_pending'
LOCK_FILE = os.path.join(PENDING_DIR, '.lock')
MAX_PATHS = 20
# The pillar_push_map.yaml is shipped via salt:// but the reactor runs on the
# master, which mounts the default saltstack tree at this path.
PUSH_MAP_PATH = '/opt/so/saltstack/default/salt/reactor/pillar_push_map.yaml'
_PUSH_MAP_CACHE = {'mtime': 0, 'data': None}
def _load_push_map():
try:
st = os.stat(PUSH_MAP_PATH)
except OSError:
LOG.warning('push_files: %s not found', PUSH_MAP_PATH)
return {}
if _PUSH_MAP_CACHE['mtime'] != st.st_mtime:
try:
with open(PUSH_MAP_PATH, 'r') as f:
_PUSH_MAP_CACHE['data'] = yaml.safe_load(f) or {}
except Exception:
LOG.exception('push_files: failed to load %s', PUSH_MAP_PATH)
_PUSH_MAP_CACHE['data'] = {}
_PUSH_MAP_CACHE['mtime'] = st.st_mtime
return _PUSH_MAP_CACHE['data'] or {}
def _push_enabled():
try:
caller = Caller()
return bool(caller.cmd('pillar.get', 'salt:auto_apply:enabled', True))
except Exception:
LOG.exception('push_files: pillar.get salt:auto_apply:enabled failed, assuming enabled')
return True
def _write_intent(key, actions, path):
now = time.time()
try:
os.makedirs(PENDING_DIR, exist_ok=True)
except OSError:
LOG.exception('push_files: cannot create %s', PENDING_DIR)
return
intent_path = os.path.join(PENDING_DIR, '{}.json'.format(key))
lock_fd = os.open(LOCK_FILE, os.O_CREAT | os.O_RDWR, 0o644)
try:
fcntl.flock(lock_fd, fcntl.LOCK_EX)
intent = {}
if os.path.exists(intent_path):
try:
with open(intent_path, 'r') as f:
intent = json.load(f)
except (IOError, ValueError):
intent = {}
intent.setdefault('first_touch', now)
intent['last_touch'] = now
intent['actions'] = actions
paths = intent.get('paths', [])
if path and path not in paths:
paths.append(path)
paths = paths[-MAX_PATHS:]
intent['paths'] = paths
tmp_path = intent_path + '.tmp'
with open(tmp_path, 'w') as f:
json.dump(intent, f)
os.rename(tmp_path, intent_path)
except Exception:
LOG.exception('push_files: failed to write intent %s', intent_path)
finally:
try:
fcntl.flock(lock_fd, fcntl.LOCK_UN)
finally:
os.close(lock_fd)
def run():
if not _push_enabled():
LOG.info('push_files: push disabled, skipping')
return {}
event = data.get('data', data) # noqa: F821 -- data provided by reactor
path = event.get('path', '')
app = tag.rsplit('/', 1)[-1].strip() # noqa: F821 -- tag provided by reactor
if not app:
LOG.debug('push_files: ignoring event with no app segment: tag=%s', tag) # noqa: F821
return {}
entry = _load_push_map().get(app)
if not entry:
LOG.warning(
'push_files: app "%s" is not in pillar_push_map.yaml; change will be '
'picked up at the next scheduled highstate (path=%s)',
app, path,
)
return {}
_write_intent('files_{}'.format(app), list(entry), path)
LOG.info('push_files: intent updated for %s (path=%s)', app, path)
return {}
-96
View File
@@ -1,96 +0,0 @@
#!py
# Reactor invoked by the rules_beacon poll beacon (salt/_beacons/rules_beacon.py) on rule
# file changes under /opt/so/saltstack/local/salt/strelka/rules/compiled/.
#
# Writes (or updates) a push intent at /opt/so/state/push_pending/rules_strelka.json
# and returns {}. The so-push-drainer schedule picks up ready intents, dedupes
# across pending files, and dispatches orch.push_batch. Reactors never dispatch
# directly
import fcntl
import json
import logging
import os
import time
from salt.client import Caller
LOG = logging.getLogger(__name__)
PENDING_DIR = '/opt/so/state/push_pending'
LOCK_FILE = os.path.join(PENDING_DIR, '.lock')
MAX_PATHS = 20
# Mirrors GLOBALS.sensor_roles in salt/vars/globals.map.jinja. Sensor-side
# strelka runs on exactly these four roles; so-import gets strelka.manager
# instead, which is not fired on pillar changes.
SENSOR_ROLES = ['so-eval', 'so-heavynode', 'so-sensor', 'so-standalone']
def _sensor_compound():
return ' or '.join('G@role:{}'.format(r) for r in SENSOR_ROLES)
def _push_enabled():
try:
caller = Caller()
return bool(caller.cmd('pillar.get', 'salt:auto_apply:enabled', True))
except Exception:
LOG.exception('push_strelka: pillar.get salt:auto_apply:enabled failed, assuming enabled')
return True
def _write_intent(key, actions, path):
now = time.time()
try:
os.makedirs(PENDING_DIR, exist_ok=True)
except OSError:
LOG.exception('push_strelka: cannot create %s', PENDING_DIR)
return
intent_path = os.path.join(PENDING_DIR, '{}.json'.format(key))
lock_fd = os.open(LOCK_FILE, os.O_CREAT | os.O_RDWR, 0o644)
try:
fcntl.flock(lock_fd, fcntl.LOCK_EX)
intent = {}
if os.path.exists(intent_path):
try:
with open(intent_path, 'r') as f:
intent = json.load(f)
except (IOError, ValueError):
intent = {}
intent.setdefault('first_touch', now)
intent['last_touch'] = now
intent['actions'] = actions
paths = intent.get('paths', [])
if path and path not in paths:
paths.append(path)
paths = paths[-MAX_PATHS:]
intent['paths'] = paths
tmp_path = intent_path + '.tmp'
with open(tmp_path, 'w') as f:
json.dump(intent, f)
os.rename(tmp_path, intent_path)
except Exception:
LOG.exception('push_strelka: failed to write intent %s', intent_path)
finally:
try:
fcntl.flock(lock_fd, fcntl.LOCK_UN)
finally:
os.close(lock_fd)
def run():
if not _push_enabled():
LOG.info('push_strelka: push disabled, skipping')
return {}
path = data.get('path', '') # noqa: F821 -- data provided by reactor
actions = [{'state': 'strelka', 'tgt': _sensor_compound()}]
_write_intent('rules_strelka', actions, path)
LOG.info('push_strelka: intent updated for path=%s', path)
return {}
-95
View File
@@ -1,95 +0,0 @@
#!py
# Reactor invoked by the rules_beacon poll beacon (salt/_beacons/rules_beacon.py) on rule
# file changes under /opt/so/saltstack/local/salt/suricata/rules/.
#
# Writes (or updates) a push intent at /opt/so/state/push_pending/rules_suricata.json
# and returns {}. The so-push-drainer schedule picks up ready intents, dedupes
# across pending files, and dispatches orch.push_batch. Reactors never dispatch
# directly
import fcntl
import json
import logging
import os
import time
from salt.client import Caller
LOG = logging.getLogger(__name__)
PENDING_DIR = '/opt/so/state/push_pending'
LOCK_FILE = os.path.join(PENDING_DIR, '.lock')
MAX_PATHS = 20
# Mirrors GLOBALS.sensor_roles in salt/vars/globals.map.jinja. Suricata also
# runs on so-import per salt/top.sls, so that role is appended below.
SENSOR_ROLES = ['so-eval', 'so-heavynode', 'so-sensor', 'so-standalone']
def _sensor_compound_plus_import():
return ' or '.join('G@role:{}'.format(r) for r in SENSOR_ROLES) + ' or G@role:so-import'
def _push_enabled():
try:
caller = Caller()
return bool(caller.cmd('pillar.get', 'salt:auto_apply:enabled', True))
except Exception:
LOG.exception('push_suricata: pillar.get salt:auto_apply:enabled failed, assuming enabled')
return True
def _write_intent(key, actions, path):
now = time.time()
try:
os.makedirs(PENDING_DIR, exist_ok=True)
except OSError:
LOG.exception('push_suricata: cannot create %s', PENDING_DIR)
return
intent_path = os.path.join(PENDING_DIR, '{}.json'.format(key))
lock_fd = os.open(LOCK_FILE, os.O_CREAT | os.O_RDWR, 0o644)
try:
fcntl.flock(lock_fd, fcntl.LOCK_EX)
intent = {}
if os.path.exists(intent_path):
try:
with open(intent_path, 'r') as f:
intent = json.load(f)
except (IOError, ValueError):
intent = {}
intent.setdefault('first_touch', now)
intent['last_touch'] = now
intent['actions'] = actions
paths = intent.get('paths', [])
if path and path not in paths:
paths.append(path)
paths = paths[-MAX_PATHS:]
intent['paths'] = paths
tmp_path = intent_path + '.tmp'
with open(tmp_path, 'w') as f:
json.dump(intent, f)
os.rename(tmp_path, intent_path)
except Exception:
LOG.exception('push_suricata: failed to write intent %s', intent_path)
finally:
try:
fcntl.flock(lock_fd, fcntl.LOCK_UN)
finally:
os.close(lock_fd)
def run():
if not _push_enabled():
LOG.info('push_suricata: push disabled, skipping')
return {}
path = data.get('path', '') # noqa: F821 -- data provided by reactor
actions = [{'state': 'suricata', 'tgt': _sensor_compound_plus_import()}]
_write_intent('rules_suricata', actions, path)
LOG.info('push_suricata: intent updated for path=%s', path)
return {}
+2 -4
View File
@@ -1,7 +1,5 @@
reactor:
- 'salt/beacon/*/rules_beacon/suricata':
- salt://reactor/push_suricata.sls
- 'salt/beacon/*/rules_beacon/strelka':
- salt://reactor/push_strelka.sls
- 'salt/beacon/*/local_files_beacon/*':
- salt://reactor/push_files.sls
- 'salt/beacon/*/postgres_pillar_beacon/audit_settings':
- salt://reactor/push_pillar.sls
+20 -2
View File
@@ -3,6 +3,8 @@
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
{% from 'vars/globals.map.jinja' import GLOBALS %}
# Manages /etc/systemd/system/so-boot-highstate.service, a Type=oneshot
# RemainAfterExit=yes unit that runs `salt-call state.highstate` exactly once
# per system boot. Replaces the legacy `startup_states: highstate` minion
@@ -19,9 +21,25 @@ so_boot_highstate_unit_file:
- onchanges_in:
- module: systemd_reload
# Non-managers never apply salt.minion during setup, so reaching this state means
# setup is finished and the marker is safe to write unconditionally. This also
# heals nodes installed before this fix, which have no marker and no legacy
# startup_states line to grep for. Managers do highstate mid-setup, so they only
# get the marker from the legacy upgrade signal; fresh installs get it from
# mark_setup_complete in setup/so-functions.
mark_setup_complete:
file.managed:
- name: /opt/so/state/setup-complete
- replace: false
- makedirs: True
{% if GLOBALS.is_manager %}
- onlyif: "grep -qx 'startup_states: highstate' /etc/salt/minion"
{% endif %}
- require_in:
- service: so_boot_highstate_service
# Only enable once setup is complete. Until then the gate file is missing and
# the unit's own ConditionPathExists would no-op it anyway -- this just keeps
# `systemctl is-enabled` honest for the sync_es_users gate.
# the unit's own ConditionPathExists would no-op it anyway.
so_boot_highstate_service:
service.enabled:
- name: so-boot-highstate.service
+4 -16
View File
@@ -87,27 +87,15 @@ set_log_levels:
# so-boot-highstate.service (managed in salt.minion.boot_highstate), which
# runs once per system boot only. Strip the line from /etc/salt/minion on
# upgrade; both the commented and uncommented forms historically existed.
# Ordered after mark_setup_complete (salt.minion.boot_highstate); the manager
# gate there greps for this line, so it must run before we delete it.
remove_startup_states:
file.line:
- name: /etc/salt/minion
- match: 'startup_states: highstate'
- mode: delete
# Upgrade-path bridge: systems that already passed setup under the old gate
# (`grep -x 'startup_states: highstate' /etc/salt/minion`) get a /opt/so/state/setup-complete
# marker so so-boot-highstate.service can be enabled and the so-user_sync cron
# in sync_es_users.sls keeps installing. Setup-in-progress systems instead get
# the marker from `mark_setup_complete` in setup/so-functions at the right
# moment. `replace: false` means we never overwrite a marker once written.
mark_setup_complete_for_upgrades:
file.managed:
- name: /opt/so/state/setup-complete
- replace: false
- makedirs: True
- onlyif: "grep -qx 'startup_states: highstate' /etc/salt/minion"
- require_in:
- file: remove_startup_states
- service: so_boot_highstate_service
- require:
- file: mark_setup_complete
{% endif %}
+9
View File
@@ -8,6 +8,15 @@ set_role_grain:
- name: role
- value: so-{{ grains.id.split("_") | last }}
# salt-cloud guests never run so-setup, so nothing else marks them setup-complete.
# Replaces the 'startup_states: highstate' line this state used to append. No
# GLOBALS import -- this runs before the guest's pillars exist.
mark_setup_complete_vm_guest:
file.managed:
- name: /opt/so/state/setup-complete
- replace: false
- makedirs: True
enable_salt_minion:
service.enabled:
- name: salt-minion
+25 -2
View File
@@ -1537,6 +1537,20 @@ soc:
Orchestrator: sonnet@SOAI
Investigator: gemma@SOAI
DetectionEngineer: gemma@SOAI
useMemory: true
useMemoryScanner: false
memoryScanIntervalSeconds: 300
memoryProximityThreshold: 0.8
messageProximityThreshold: 0.5
maxUserMemoriesToInclude: 5
maxGlobalMemoriesToInclude: 5
maxUserMemoriesToReconcile: 20
maxGlobalMemoriesToReconcile: 20
memoryModel: gemma@SOAI
embedModel: amazon.titan-embed-text-v2@SOAI
reconcileModel: gemma@SOAI
memoryPersona: ""
reconcilePersona: ""
onionconfig:
saltstackDir: /opt/so/saltstack
bypassEnabled: false
@@ -2671,7 +2685,7 @@ soc:
# The id (UUIDv4) is pregenerated and can safely be used.
# Click "Convert" to convert the Sigma rule to use Security Onion field mappings within an EQL query
#
# Rule Creation Guide: https://github.com/SigmaHQ/sigma/wiki/Rule-Creation-Guide
# Rule Creation Guide: https://github.com/SigmaHQ/sigma/wiki/Rule-Creation-High%E2%80%90Level-Guide
# Logsources: https://sigmahq.io/docs/basics/log-sources.html
title: 'A Short Capitalized Title With Less Than 50 Characters'
@@ -2683,7 +2697,7 @@ soc:
references:
- 'https://local.invalid'
author: '@SecurityOnion'
date: 'YYYY/MM/DD'
date: '[today]'
tags:
- detection.threat_hunting
- attack.technique_id
@@ -2727,5 +2741,14 @@ soc:
enabled: true
adapter: SOAI
charsPerTokenEstimate: 4
- id: amazon.titan-embed-text-v2
displayName: amazon.titan-embed-text-v2
origin: USA
contextLimitSmall: 8192
contextLimitLarge: 8192
lowBalanceColorAlert: 500000
enabled: true
adapter: SOAI
charsPerTokenEstimate: 4
+49
View File
@@ -845,6 +845,55 @@ soc:
DetectionEngineer:
description: This agent manages detections and their overrides, including tuning noisy rules and authoring rule content.
global: True
useMemory:
description: Enables the Memory system for OnionAI
global: True
forcedType: bool
useMemoryScanner:
description: Enables the memory scanner for automatic memory extraction from historical sessions.
global: True
forcedType: bool
memoryScanIntervalSeconds:
description: How long to wait in seconds between attempts to scan sessions for new memories.
global: True
memoryProximityThreshold:
description: Describes how close memories need to be on a floating point scale from 0.0 to 1.0 to be considered when reconciling new memories with old ones. This value is usually higher than messageProximityThreshold.
global: True
messageProximityThreshold:
description: Describes how close a memory needs to be to a user's message on a floating point scale from 0.0 to 1.0 to be included in the context. This value is usually lower than memoryProximityThreshold.
global: True
maxUserMemoriesToInclude:
description: Specify the max number of user-specific memories to include in the prompt when a user sends a message.
global: True
maxGlobalMemoriesToInclude:
description: Specify the max number of global memories to include in the prompt when a user sends a message.
global: True
maxUserMemoriesToReconcile:
description: When reconciling new user-specific memories with existing user-specific memories, this determines how many old memories may be considered.
global: True
maxGlobalMemoriesToReconcile:
description: When reconciling new global memories with existing global memories, this determines how many old memories may be considered.
global: True
memoryModel:
description: The model to use when extracting memories from sessions.
global: True
embedModel:
description: The model to use when embedding a memory as a vector. Note that only memories embedded using the same model may be compared and only memories created with the model specified here will be considered when informing an agent of existing memories.
global: True
advanced: True
reconcileModel:
description: The model to use when reconciling memories that contain nearly the same content.
global: True
memoryPersona:
description: Text appended to the built-in prompt of the memory extraction agent, managed from the Agent Studio. Use it to steer what is worth remembering.
global: True
readonlyUi: True
multiline: True
reconcilePersona:
description: Text appended to the built-in prompt of the memory reconciliation agent, managed from the Agent Studio. Use it to steer how new memories are merged with existing ones.
global: True
readonlyUi: True
multiline: True
client:
assistant:
enabled:
+3 -3
View File
@@ -335,7 +335,7 @@
{%- do TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]].remove('sostatus.sh') %}
[[inputs.exec]]
commands = [
"/scripts/sostatus.sh"
["/scripts/sostatus.sh"]
]
data_format = "influx"
timeout = "15s"
@@ -346,7 +346,7 @@
[[inputs.exec]]
commands = [
{%- for script in TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]] %}
"/scripts/{{script}}"{% if not loop.last %},{% endif %}
["/scripts/{{script}}"]{% if not loop.last %},{% endif %}
{%- endfor %}
]
data_format = "influx"
@@ -375,7 +375,7 @@
{%- if GLOBALS.is_manager or GLOBALS.role == 'so-heavynode' %}
[[ inputs.exec ]]
commands = [
"/scripts/esindexsize.sh"
["/scripts/esindexsize.sh"]
]
data_format = "influx"
interval = "1h"
+1 -1
View File
@@ -15,7 +15,7 @@ zeek:
MailHostUpDown: 0
LogRotationInterval: 3600
LogExpireInterval: 0
StatsLogEnable: 1
StatsLogEnable: 0
StatsLogExpireInterval: 0
StatusCmdShowAll: 0
CrashExpireInterval: 0
+5
View File
@@ -23,6 +23,11 @@ zeekpacketlosscron:
- identifier: zeekpacketlosscron
- user: root
zeekctlcron:
cron.absent:
- identifier: zeekctlcron
- user: root
{% else %}
{{sls}}_state_not_allowed:
+15
View File
@@ -87,6 +87,21 @@ zeekpacketlosscron:
- month: '*'
- dayweek: '*'
# LogExpireInterval, StatsLogExpireInterval and CrashExpireInterval are only acted on by
# 'zeekctl cron', so run it on the interval upstream recommends. This also restarts any
# node that died unexpectedly. Runs as root because the script needs the docker socket;
# it drops to the zeek user inside the container.
zeekctlcron:
cron.present:
- name: /usr/sbin/so-zeek-cron > /dev/null 2>&1
- identifier: zeekctlcron
- user: root
- minute: '*/5'
- hour: '*'
- daymonth: '*'
- month: '*'
- dayweek: '*'
{% else %}
{{sls}}_state_not_allowed:
+80
View File
@@ -58,6 +58,86 @@ zeek:
CompressLogs:
description: This setting enables compression of Zeek logs. If you are seeing packet loss at the top of the hour in Zeek or PCAP you might need to disable this by seting it to 0. This will use more disk space but save IO and CPU.
helpLink: zeek
LogExpireInterval:
description: >-
How long to keep rotated Zeek logs in /nsm/zeek/logs. A bare number means DAYS, so 7 means 7 days.
You may also give an explicit unit, such as "7 days" or "12 hr". Use 0 to keep logs forever.
This value must not be shorter than LogRotationInterval (3600 seconds by default), so the smallest
usable value is 1 hr - Zeek will fail to start if it is shorter. Expiry is applied by "zeekctl cron",
which runs every 5 minutes, and removes log files older than this based on their modification time.
regex: ^(0|[1-9][0-9]*( ?(day|hr)s?)?)$
regexFailureMessage: Enter 0, or a positive number optionally followed by "day" or "hr" (for example 7, "7 days", or "12 hr"). Minutes are not accepted because a log expire interval shorter than the log rotation interval prevents Zeek from starting.
helpLink: zeek
advanced: True
StatsLogEnable:
description: >-
Set to 1 to have "zeekctl cron" write node statistics to /nsm/zeek/logs/stats. This is
disabled because the CPU and memory portion depends on the "top" command, which the Zeek
container does not include, so every run records an error for each node instead. The
interface packet counters it also collects are not used anywhere in Security Onion, which
tracks Zeek packet loss separately through packetloss.log and Telegraf. It is read only
for that reason.
regex: ^[01]$
regexFailureMessage: You must enter 0 or 1.
helpLink: zeek
advanced: True
readonly: True
StatsLogExpireInterval:
description: >-
Number of days to keep entries in the Zeek stats log, or 0 to keep them forever.
Applied by "zeekctl cron", which runs every 5 minutes. This has no effect unless
StatsLogEnable is turned on, which it is not by default.
regex: ^[0-9]+$
regexFailureMessage: You must enter a whole number of days, or 0 to keep entries forever.
helpLink: zeek
advanced: True
CrashExpireInterval:
description: >-
Number of days to keep Zeek crash directories, or 0 to keep them forever.
Applied by "zeekctl cron", which runs every 5 minutes.
regex: ^[0-9]+$
regexFailureMessage: You must enter a whole number of days, or 0 to keep crash directories forever.
helpLink: zeek
advanced: True
MinDiskSpace:
description: >-
Percentage of free disk space below which ZeekControl reports a warning, or 0 to disable the check
entirely. The Zeek container does not include a mail program, so the warning is not emailed - it
appears in the output of "zeekctl cron" instead. This setting never deletes anything - cleanup based
on disk usage is handled separately by so-sensor-clean.
regex: ^([0-9]|[1-9][0-9]|100)$
regexFailureMessage: You must enter a percentage between 0 and 100.
helpLink: zeek
advanced: True
MailTo:
description: >-
Address that ZeekControl would send mail to, covering cron output and crash reports, and the address
Zeek's notice framework would use. The Zeek container does not include a mail program, and Security
Onion never enables the notice email action, so no mail is sent and this address is unused. It is
read only for that reason.
helpLink: zeek
advanced: True
readonly: True
MailConnectionSummary:
description: >-
Set to 1 to email the hourly connection summary. This only controls the emailed copy - the summary is
generated and archived with the other Zeek logs either way. The Zeek container does not include a mail
program, so no mail is sent and this setting has no effect. It is read only for that reason.
regex: ^[01]$
regexFailureMessage: You must enter 0 or 1.
helpLink: zeek
advanced: True
readonly: True
MailHostUpDown:
description: >-
Set to 1 to report when a Zeek node changes between the up and down states. The Zeek container does
not include a mail program, so this notification cannot be emailed. It is read only for that reason.
Host status detection still runs regardless of this setting - only the notification is affected.
regex: ^[01]$
regexFailureMessage: You must enter 0 or 1.
helpLink: zeek
advanced: True
readonly: True
policy:
custom:
filters:
+24
View File
@@ -0,0 +1,24 @@
#!/bin/bash
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# Run zeekctl's periodic maintenance tasks. This is what actually enforces
# LogExpireInterval, StatsLogExpireInterval and CrashExpireInterval - without a periodic
# 'zeekctl cron' those settings are inert no matter what they are set to.
# This also restarts any node that died unexpectedly, and marks it crashed so a crash report
# is produced. That is upstream's default cron behavior and it recovers a single node in
# place. The beacon in salt/_beacons/zeek.py is the only other recovery path, it is disabled
# by default (healthcheck:enabled), and it removes and recreates the whole container, so
# letting zeekctl handle a single dead worker avoids the heavier restart.
if ! docker ps --filter name=so-zeek --format '{{.Names}}' | grep -q '^so-zeek$'; then
exit 0
fi
# Run as the zeek user so the stats logs and zeekctl-config.sh this writes stay owned by
# uid 937 rather than root.
docker exec so-zeek runuser -l zeek -c '/opt/zeek/bin/zeekctl cron'
+1
View File
@@ -833,6 +833,7 @@ if ! [[ -f $install_opt_file ]]; then
check_sos_appliance
drop_install_options
hypervisor_local_states
mark_setup_complete
verify_setup
fi