Compare commits

...
Author SHA1 Message Date
Josh Patterson 72f60fcaa9 FIX: remove the docker socket from so-telegraf
so-telegraf mounted /var/run/docker.sock and joined the host docker group on
every node type. The :ro flag blocks write() to the inode, not connect() plus
HTTP over the socket, so any code execution inside the container could reach
POST /containers/create with Privileged:true and become root on the host.

No telegraf script used the socket; the only consumer was the native
[[inputs.docker]] plugin, and group_add 920 existed solely to feed it.
Container metrics now come from so-container-stats, a collector that runs on
the host from cron and writes influx line protocol to a file telegraf already
had mounted. This is the pattern so-status, so-raid-status and
so-elasticagent-status already use, so the privileged docker access stays on
the host side where root cron already ran it.

The collector runs as somon, a service account in the docker group with no
login shell and a locked password, rather than root. Docker group membership
is still root-equivalent on the host, so this is defense in depth rather than
a privilege boundary.

The telegraf scripts were root:939 mode 770, letting socore rewrite them for
code execution inside the container; they are now 750, which still allows the
read and execute telegraf needs. The container also ran with no group, giving
it gid 0, and now runs as 939:939. That alone would have broken
lasthighstate.sh, which reached /opt/so/log/salt only via the root group and
could not tell an unreadable file from a missing one, so it silently reported
a 56 year highstate age. Only the lasthighstate file is bind mounted now, and
the script tests readability instead of existence.

Everything inputs.docker collected beyond the five fields the shipped
dashboards query is available per stat under telegraf:container_stats,
annotated for SOC so an operator can enable it without editing files. Defaults
reproduce the previous output exactly. With every stat enabled the emitted
field set matches what inputs.docker wrote, verified by running the plugin
against the live socket and diffing: 53 fields, no type mismatches, no field
present on one side only.

Two deliberate differences: max_usage carries the real cgroup peak where the
daemon reports 0 on cgroup v2, and host-network containers emit no
docker_container_net row, matching inputs.docker. Docker label tags are not
restored, since nothing queries them and they cost significant cardinality.

so-status, the influxdb size cron, so-elasticagent-status, so-raid-status and
so-common-status-check truncated their output in place while telegraf read it,
so telegraf periodically saw an empty file and logged a parse error or emitted
empty values. They now write aside and rename. Measured on a live manager, the
old so-status cron left status.log empty for 215 of 10997 reads.

Tested on a fresh install, a converted grid and a 3.0 upgrade.
2026-09-24 14:30:25 -04:00
Jorge Reyes 2684a5ca95 Merge pull request #16260 from Security-Onion-Solutions/reyesj2/16254
FIX: Fleet scripts failing when endpoints-initial is missing
2026-09-23 08:58:15 -05:00
reyesj2 bcee63bde5 remove policy precheck 2026-09-23 08:51:11 -05:00
reyesj2 6ce89eb323 jq -e exits 0 with empty input 2026-09-22 21:25:50 -05:00
reyesj2 ebab4b0d90 split between missing token and multiple enrollment tokens 2026-09-22 16:11:56 -05:00
reyesj2 db60c27da2 fail cleanly when endpoints-initial is missing or has multiple enrollment tokens 2026-09-22 12:24:23 -05:00
Jason Ertel f4defdfde0 Merge pull request #16258 from Security-Onion-Solutions/jertel/wip
upgrade whoislookup analyzer deps; notification annotations
2026-09-22 08:41:04 -04:00
Jason Ertel 36652e8f23 fixed malformed desc annotation 2026-09-22 08:30:44 -04:00
Jason Ertel 7bef194540 force transitive dep version 2026-09-22 08:28:44 -04:00
Jason Ertel e2bf2837fe notification annotations 2026-09-22 08:18:59 -04:00
Jason Ertel 06704dad22 upgrade whoislookup analyzer deps 2026-09-22 08:17:18 -04:00
Mike Reeves 47fe0758d0 Merge pull request #16255 from Security-Onion-Solutions/mreeves/fix-so-user-stdin-drain
Fix user creation failing with a valid password
2026-09-18 17:12:09 -04:00
Mike Reeves b71fd93f9d Stop docker exec from consuming so-user's piped password
Adding a user failed with "Password does not meet the minimum requirements"
for passwords that were well over the eight character minimum.

so-user reads the password from stdin in updatePassword, but verifyEnvironment
runs first and calls kratosCurl, which ran docker exec with -i. That attaches
stdin and drains the pipe the caller sent the password on, so the later
read -rs saw EOF, password was empty, and expr length "" tripped the minimum
length check.

No kratosCurl or hydraCurl call sends a body on stdin; every one passes it as
a -d argument, so -i was never needed. Dropping it leaves the admin API
responses unchanged.

so-client had the same wrapper for the Hydra admin API. It does not currently
read stdin, but the flag drains its caller's pipe just the same, so it is
dropped there too.
2026-09-18 17:04:17 -04:00
Mike Reeves d9eff9aa9e Merge pull request #16253 from Security-Onion-Solutions/mreeves/fix-docker-networks-setting-conflict
Fix config tree error from the docker.networks setting
2026-09-18 15:41:25 -04:00
Mike Reeves d8884dbd99 Document sobridge alongside the soauth network settings
sobridge carries most containers but held no annotation, so it was absent from
the config tree entirely. Its pillar entry is an empty map, so annotating it
adds a setting with no children for docker.networks.soauth.range to collide
with, and the tree still builds.

docker.map.jinja fills in sobridge's range and gateway at render time from
docker.range and docker.gateway, and resets the entry when it is not a mapping,
so nothing writes keys underneath it.

The sentence about sobridge is dropped from the soauth range description now
that sobridge documents itself.
2026-09-18 14:37:11 -04:00
Mike Reeves 6c0d4c15e8 Annotate the soauth network leaves instead of the networks map
SOC reported "Malformed config setting (docker.networks.soauth.range): Setting
name 'networks' conflicts with another similarly named setting" and refused to
render the config tree.

soc_docker.yaml annotated docker.networks itself, and every key in that block
was a scalar, so FlattenAnnotations registered docker.networks as a setting.
defaults.yaml holds a nested map there, so FlattenPillar separately registered
docker.networks.soauth.range, .gateway and .manager_only, and
HydrateAnnotations unions the two sets. The config tree gives each id segment
either a value or children, never both, so addToNode pushed docker.networks as
a childless leaf and then threw when docker.networks.soauth.range tried to
descend through it.

The annotations move down to the three leaves that actually carry values, so
docker.networks is a plain branch. This matches docker.containers, which is
nested the same way and has never carried annotations of its own.

docker.ulimits and the per-container networks list are unaffected: their pillar
values are lists rather than maps, so they stay single settings.
2026-09-18 13:42:41 -04:00
Jason Ertel 2e2f62f265 Merge pull request #16252 from Security-Onion-Solutions/jertel/wip
store in db
2026-09-18 12:47:32 -04:00
Jason Ertel b7a11a525c store in db 2026-09-18 12:41:35 -04:00
Mike Reeves 8eef95ea3e Merge pull request #16221 from Security-Onion-Solutions/mreeves/kratos-soauth-network
Isolate the Kratos admin API on a dedicated soauth docker network
2026-09-18 10:24:00 -04:00
Mike Reeves 65e261475d Mention Hydra in the soauth network description
The description was written alongside the Kratos change and named only the
Kratos admin API, but so-hydra now sits on soauth too.

The trailing note about rebuilding the firewall is dropped, since the setting
is readonly and cannot be changed from SOC in the first place.
2026-09-18 10:17:20 -04:00
Mike Reeves bbc28c88b7 Isolate the Hydra admin API on the soauth docker network
The Hydra admin API creates OAuth clients and introspects tokens without
authentication, relying on network isolation the same way Kratos does, but
so-hydra published 0.0.0.0:4445:4445 and sat on sobridge. With userland-proxy
left at its default of true, Docker ran a proxy listener on the host, so any
container reached the admin API through the manager IP or the bridge gateway
regardless of its docker network.

so-hydra now sits alone on soauth and no longer publishes 4445. 4444 is still
published, so the nginx proxy for /oauth2/token and the well-known endpoints is
unchanged. so-soc is already dual homed from the Kratos change and reaches the
admin API over soauth, so only its hostUrl moves. wait_for_hydra polls the
container address instead of the host port.

so-client reaches the admin API through docker exec, mirroring so-user. Exit
codes still propagate, so the --fail-with-body error handling is unchanged.

so-hydra is removed in up_to_3.4.0 alongside so-kratos and so-soc so the
highstate recreates it with the correct network membership. The auth network
range is already handled by set_soauth_range.
2026-09-17 16:58:49 -04:00
Jason Ertel edaacf79a7 Merge pull request #16251 from Security-Onion-Solutions/jertel/wip
notification annotations
2026-09-17 15:37:04 -04:00
Jason Ertel ecc643cd33 fix typo 2026-09-17 15:34:29 -04:00
Jason Ertel 08aaf7948e Merge branch '3/dev' into jertel/wip 2026-09-17 15:32:35 -04:00
Jason Ertel bb57545d08 new annotations for notifications 2026-09-17 15:32:21 -04:00
Josh Brower 0f7adbbecc Merge pull request #16247 from Security-Onion-Solutions/esql-fixes
Refactor for ESQL
2026-09-17 09:56:19 -04:00
Josh Patterson aeb4fe8f50 Merge pull request #16250 from Security-Onion-Solutions/fix/root-own-sbin-and-salt-tree
FIX: root-own /usr/sbin management scripts and the Salt default tree
2026-09-17 09:44:47 -04:00
defensivedepth f3aa39c5a4 Tweak name 2026-09-17 09:43:44 -04:00
Doug Burks 24077ba974 Merge pull request #16249 from Security-Onion-Solutions/dougburks/fix-eval-suricata-file-drops
FIX: suricata.fileinfo maps boolean gaps into long file.bytes.missing
2026-09-17 09:42:27 -04:00
Doug Burks b3567405f9 FIX: suricata.fileinfo maps boolean gaps into long file.bytes.missing 2026-09-17 07:57:44 -04:00
defensivedepth 1fc5bb7afa Refactor for ESQL 2026-09-17 07:56:42 -04:00
Mike Reeves cf3a4ebc27 Isolate the Kratos admin API on a dedicated soauth docker network
The Kratos admin API is unauthenticated by design and relies on network
isolation, but so-kratos published 0.0.0.0:4434:4434 and daemon.json leaves
userland-proxy at its default of true. Docker therefore ran a proxy listener
on the host, and any container reached the admin API through the manager IP
or the bridge gateway regardless of its docker network.

so-kratos now sits alone on a new soauth network and no longer publishes
4434. 4433 is still published, so nginx is unchanged. so-soc is dual homed
and reaches the admin API over soauth, with publicHostUrl set explicitly.
so-user reaches it through docker exec, and wait_for_kratos polls the
container address instead of the host port.

firewall/iptables.jinja hardcoded sobridge in every generated DNAT, ACCEPT,
masquerade and isolation rule, so it is now driven by a docker:networks map
and a per-container networks list. Containers that do not declare one
default to sobridge, and .ip still resolves to the primary network's
address. soauth is only created on the roles that run so-kratos.

Setup already allows a custom docker range, so it now also prompts for the
auth network range and writes it to the docker pillar. Grids on the default
range need no decision during soup, since they get the same 172.17.2.0/24 a
fresh install would. Grids that set a custom range are prompted, defaulting
to the adjacent /24, and unattended upgrades take that default with a notice
rather than blocking.

so-kratos and so-soc are removed in up_to_3.4.0 so the highstate recreates
them with the correct network membership.
2026-09-09 16:09:00 -04:00
58 changed files with 2085 additions and 138 deletions
+25
View File
@@ -6,6 +6,9 @@ on:
- "salt/sensoroni/files/analyzers/**"
- "salt/manager/tools/sbin/**"
- "salt/_beacons/**"
- "salt/telegraf/tools/sbin_jinja/**"
- "salt/telegraf/defaults.yaml"
- "salt/telegraf/soc_telegraf.yaml"
jobs:
build:
@@ -34,3 +37,25 @@ jobs:
- name: Test with pytest
run: |
PYTHONPATH=${{ matrix.python-code-path }} pytest ${{ matrix.python-code-path }} --cov=${{ matrix.python-code-path }} --doctest-modules --cov-report=term --cov-fail-under=100 --cov-config=pytest.ini
telegraf-collector:
# so-container-stats is a jinja template rather than an importable module, so it gets its
# own job: the test renders it the way salt does, then drives it with a faked docker engine
# and cgroup tree. No container runtime is needed.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Set up Python
uses: actions/setup-python@v3
with:
python-version: "3.14"
- name: Install dependencies
run: |
python -m pip install --upgrade pip
python -m pip install flake8 pytest jinja2 pyyaml
- name: Lint with flake8
run: |
flake8 salt/telegraf/tools/sbin_jinja/so-container-stats_test.py --config=pytest.ini
- name: Test with pytest
run: |
pytest salt/telegraf/tools/sbin_jinja/so-container-stats_test.py -v
+3 -1
View File
@@ -217,9 +217,11 @@ sostatus_log:
- replace: False
# Install sostatus check cron. This is used to populate Grid.
# telegraf reads status.log on the same minute boundary this runs, so write aside and rename
# rather than truncating the file it is reading
so-status_check_cron:
cron.present:
- name: '/usr/sbin/so-status -j > /opt/so/log/sostatus/status.log 2>&1'
- name: '/usr/sbin/so-status -j > /opt/so/log/sostatus/status.log.tmp 2>&1; mv -f /opt/so/log/sostatus/status.log.tmp /opt/so/log/sostatus/status.log'
- identifier: so-status_check_cron
- user: root
- minute: '*/1'
+19 -6
View File
@@ -9,6 +9,7 @@ import sys
import subprocess
import os
import json
import tempfile
sys.path.append('/opt/saltstack/salt/lib/python3.10/site-packages/')
import salt.config
@@ -17,6 +18,21 @@ import salt.loader
__opts__ = salt.config.minion_config('/etc/salt/minion')
__grains__ = salt.loader.grains(__opts__)
def write_atomic(path, value):
# telegraf reads these files on its own schedule; replacing them by rename means it never
# reads a truncated file and reports an empty value as if it were real
directory = os.path.dirname(path)
handle, temp = tempfile.mkstemp(dir=directory)
try:
with os.fdopen(handle, 'w') as f:
f.write(str(value))
os.chmod(temp, 0o644)
os.replace(temp, path)
except Exception:
os.path.exists(temp) and os.unlink(temp)
raise
def check_needs_restarted():
osfam = __grains__['os_family']
val = '0'
@@ -34,8 +50,7 @@ def check_needs_restarted():
else:
fail("Unsupported OS")
with open(outfile, 'w') as f:
f.write(val)
write_atomic(outfile, val)
def check_for_fps():
feat = 'fps'
@@ -56,8 +71,7 @@ def check_for_fps():
# Unknown, so assume 0
fps = 0
with open('/opt/so/log/sostatus/fps_enabled', 'w') as f:
f.write(str(fps))
write_atomic('/opt/so/log/sostatus/fps_enabled', fps)
def check_for_lks():
feat = 'Lks'
@@ -80,8 +94,7 @@ def check_for_lks():
lks = 1
if lks:
break
with open('/opt/so/log/sostatus/lks_enabled', 'w') as f:
f.write(str(lks))
write_atomic('/opt/so/log/sostatus/lks_enabled', lks)
def fail(msg):
print(msg, file=sys.stderr)
+3 -1
View File
@@ -125,4 +125,6 @@ else
RAIDSTATUS=1
fi
echo "nsmraid=$RAIDSTATUS" > /opt/so/log/raid/status.log
# telegraf reads this file; write aside and rename so it never sees a half-written file
echo "nsmraid=$RAIDSTATUS" > /opt/so/log/raid/status.log.tmp
mv -f /opt/so/log/raid/status.log.tmp /opt/so/log/raid/status.log
+9 -2
View File
@@ -1,6 +1,12 @@
docker:
range: '172.17.1.0/24'
gateway: '172.17.1.1'
networks:
sobridge: {}
soauth:
range: '172.17.2.0/24'
gateway: '172.17.2.1'
manager_only: True
ulimits:
- name: nofile
soft: 1048576
@@ -58,18 +64,18 @@ docker:
ulimits: []
'so-kratos':
final_octet: 28
networks: ['soauth']
port_bindings:
- 0.0.0.0:4433:4433
- 0.0.0.0:4434:4434
custom_bind_mounts: []
extra_hosts: []
extra_env: []
ulimits: []
'so-hydra':
final_octet: 30
networks: ['soauth']
port_bindings:
- 0.0.0.0:4444:4444
- 0.0.0.0:4445:4445
custom_bind_mounts: []
extra_hosts: []
extra_env: []
@@ -128,6 +134,7 @@ docker:
ulimits: []
'so-soc':
final_octet: 34
networks: ['sobridge', 'soauth']
port_bindings:
- 0.0.0.0:9822:9822
custom_bind_mounts: []
+21 -3
View File
@@ -1,8 +1,26 @@
{% import_yaml 'docker/defaults.yaml' as DOCKERDEFAULTS %}
{% set DOCKERMERGED = salt['pillar.get']('docker', DOCKERDEFAULTS.docker, merge=True) %}
{% set RANGESPLIT = DOCKERMERGED.range.split('.') %}
{% set FIRSTTHREE = RANGESPLIT[0] ~ '.' ~ RANGESPLIT[1] ~ '.' ~ RANGESPLIT[2] ~ '.' %}
{% if DOCKERMERGED.networks.sobridge is not mapping %}
{% do DOCKERMERGED.networks.update({'sobridge': {}}) %}
{% endif %}
{% do DOCKERMERGED.networks['sobridge'].update({'range': DOCKERMERGED.range, 'gateway': DOCKERMERGED.gateway}) %}
{% for netname, net in DOCKERMERGED.networks.items() %}
{% set RANGESPLIT = net.range.split('.') %}
{% do net.update({'prefix': RANGESPLIT[0] ~ '.' ~ RANGESPLIT[1] ~ '.' ~ RANGESPLIT[2] ~ '.'}) %}
{% endfor %}
{% for container, vals in DOCKERMERGED.containers.items() %}
{% do DOCKERMERGED.containers[container].update({'ip': FIRSTTHREE ~ DOCKERMERGED.containers[container].final_octet}) %}
{% set CONTAINER_NETS = vals.get('networks', ['sobridge']) %}
{% set IPS = {} %}
{% for netname in CONTAINER_NETS %}
{% do IPS.update({netname: DOCKERMERGED.networks[netname].prefix ~ vals.final_octet}) %}
{% endfor %}
{% do DOCKERMERGED.containers[container].update({
'networks': CONTAINER_NETS,
'ips': IPS,
'network': CONTAINER_NETS[0],
'ip': IPS[CONTAINER_NETS[0]]
}) %}
{% endfor %}
+10 -6
View File
@@ -71,15 +71,19 @@ dockerreserveports:
- source: salt://common/files/99-reserved-ports.conf
- name: /etc/sysctl.d/99-reserved-ports.conf
sos_docker_net:
{% for NETNAME, NETWORK in DOCKERMERGED.networks.items() %}
{% if not NETWORK.get('manager_only') or GLOBALS.get('is_manager', False) %}
sos_docker_net_{{ NETNAME }}:
docker_network.present:
- name: sobridge
- subnet: {{ DOCKERMERGED.range }}
- gateway: {{ DOCKERMERGED.gateway }}
- name: {{ NETNAME }}
- subnet: {{ NETWORK.range }}
- gateway: {{ NETWORK.gateway }}
- options:
com.docker.network.bridge.name: 'sobridge'
com.docker.network.bridge.name: '{{ NETNAME }}'
com.docker.network.driver.mtu: '1500'
com.docker.network.bridge.enable_ip_masquerade: 'true'
com.docker.network.bridge.enable_icc: 'true'
com.docker.network.bridge.host_binding_ipv4: '0.0.0.0'
- unless: ip l | grep sobridge
- unless: ip l | grep {{ NETNAME }}
{% endif %}
{% endfor %}
+44
View File
@@ -7,6 +7,40 @@ docker:
description: Default docker IP range for containers.
helpLink: docker
advanced: True
networks:
sobridge:
description: |
The default docker network, carrying most containers. Its range and gateway are taken
from the docker.range and docker.gateway settings above rather than set here.
helpLink: docker
readonly: True
advanced: True
global: True
soauth:
range:
description: |
IP range for the soauth docker network, an isolated network for the authentication
services, so that the Kratos and Hydra admin APIs are only reachable from the
containers placed on it.
helpLink: docker
readonly: True
advanced: True
global: True
gateway:
description: Gateway for the soauth docker network.
helpLink: docker
readonly: True
advanced: True
global: True
manager_only:
description: |
Limits the soauth network to grid members running the authentication containers,
instead of creating it on every node.
helpLink: docker
readonly: True
advanced: True
global: True
forcedType: bool
ulimits:
description: |
Default ulimit settings applied to all containers via the Docker daemon. Each entry specifies a resource name (e.g. nofile, memlock, core, nproc) with soft and hard limits. Individual container ulimits override these defaults. Valid resource names include: cpu, fsize, data, stack, core, rss, nproc, nofile, memlock, as, locks, sigpending, msgqueue, nice, rtprio, rttime.
@@ -34,6 +68,16 @@ docker:
readonly: True
advanced: True
global: True
networks:
description: |
Docker networks this container is attached to. The first entry is the container's
primary network and determines the address its published ports are forwarded to.
Defaults to sobridge when unset.
helpLink: docker
readonly: True
advanced: True
global: True
forcedType: "[]string"
port_bindings:
description: List of port bindings for the container.
helpLink: docker
@@ -30,6 +30,56 @@ fleet_api() {
curl -sK /opt/so/conf/elasticsearch/curl.config -L "localhost:5601/api/fleet/${QUERYPATH}" "$@" --retry 3 --retry-delay 10 --fail 2>/dev/null
}
elastic_fleet_require_agent_policy() {
local AGENT_POLICY=$1
local POLICY_JSON
if ! POLICY_JSON=$(fleet_api "agent_policies/$AGENT_POLICY") || [ -z "$POLICY_JSON" ]; then
echo "Error: Agent policy '$AGENT_POLICY' was not found or is not visible in the current Kibana space." >&2
return 1
fi
if ! jq -e '.item.package_policies | type == "array"' <<<"$POLICY_JSON" >/dev/null 2>&1; then
echo "Error: Agent policy '$AGENT_POLICY' was not found or is not visible in the current Kibana space." >&2
return 1
fi
echo "$POLICY_JSON"
}
# Print the single active enrollment token for POLICY_ID.
# Exit 1: retryable (API failure, invalid response, no active token)
# Exit 2: multiple active tokens - Shouldn't get into this state without manual intervention
elastic_fleet_active_enrollment_token() {
local POLICY_ID=$1
local RESP TOKEN_COUNT API_KEY
if ! RESP=$(fleet_api "enrollment_api_keys?perPage=100" -H 'kbn-xsrf: true' -H 'Content-Type: application/json'); then
echo "Error: Failed to retrieve enrollment tokens for agent policy '$POLICY_ID'." >&2
return 1
fi
if ! jq -e '.list' <<<"$RESP" >/dev/null 2>&1; then
echo "Error: Invalid enrollment token response for agent policy '$POLICY_ID'." >&2
return 1
fi
TOKEN_COUNT=$(jq --arg pid "$POLICY_ID" '[.list[] | select(.policy_id == $pid and .active == true)] | length' <<<"$RESP")
if [ "${TOKEN_COUNT:-0}" -eq 0 ]; then
echo "Error: No active enrollment token found for agent policy '$POLICY_ID'." >&2
return 1
fi
if [ "$TOKEN_COUNT" -gt 1 ]; then
echo "Error: Found $TOKEN_COUNT active enrollment tokens for agent policy '$POLICY_ID'; expected exactly one." >&2
return 2
fi
API_KEY=$(jq -r --arg pid "$POLICY_ID" '.list[] | select(.policy_id == $pid and .active == true) | .api_key' <<<"$RESP")
echo "$API_KEY"
}
# Max number of concurrent Fleet write jobs (create/update). Override via env if needed.
MAX_FLEET_JOBS=${MAX_FLEET_JOBS:-10}
@@ -62,15 +112,7 @@ elastic_fleet_load_integrations_dir() {
i=0
# Fetch the agent policy a single time; we look up integration ids locally below.
if ! POLICY_JSON=$(fleet_api "agent_policies/$AGENT_POLICY"); then
echo "Error: Failed to retrieve agent policy '$AGENT_POLICY'."
rm -f "$FAIL_FILE"
rm -rf "$OUT_DIR"
return 1
fi
if ! jq -e '.item.package_policies' <<<"$POLICY_JSON" >/dev/null 2>&1; then
echo "Error: Invalid agent policy response for '$AGENT_POLICY'."
if ! POLICY_JSON=$(elastic_fleet_require_agent_policy "$AGENT_POLICY"); then
rm -f "$FAIL_FILE"
rm -rf "$OUT_DIR"
return 1
@@ -124,9 +166,15 @@ elastic_fleet_integration_check() {
JSON_STRING=$2
NAME=$(jq -r .name $JSON_STRING)
NAME=$(jq -r .name "$JSON_STRING")
INTEGRATION_ID=""
INTEGRATION_ID=$(/usr/sbin/so-elastic-fleet-agent-policy-view "$AGENT_POLICY" | jq -r '.item.package_policies[] | select(.name=="'"$NAME"'") | .id')
local POLICY_JSON
if ! POLICY_JSON=$(elastic_fleet_require_agent_policy "$AGENT_POLICY"); then
return 1
fi
INTEGRATION_ID=$(jq -r --arg n "$NAME" '.item.package_policies[]? | select(.name==$n) | .id' <<<"$POLICY_JSON")
}
@@ -148,7 +196,16 @@ elastic_fleet_integration_remove() {
NAME=$2
INTEGRATION_ID=$(/usr/sbin/so-elastic-fleet-agent-policy-view "$AGENT_POLICY" | jq -r '.item.package_policies[] | select(.name=="'"$NAME"'") | .id')
local POLICY_JSON
if ! POLICY_JSON=$(elastic_fleet_require_agent_policy "$AGENT_POLICY"); then
return 1
fi
INTEGRATION_ID=$(jq -r --arg n "$NAME" '.item.package_policies[]? | select(.name==$n) | .id' <<<"$POLICY_JSON")
if [ -z "$INTEGRATION_ID" ]; then
echo "Error: Integration '$NAME' was not found in agent policy '$AGENT_POLICY'." >&2
return 1
fi
JSON_STRING=$( jq -n \
--arg INTEGRATIONID "$INTEGRATION_ID" \
@@ -13,7 +13,10 @@ ERROR=false
for INTEGRATION in /opt/so/conf/elastic-fleet/integrations/elastic-defend/*.json
do
printf "\n\nInitial Endpoints Policy - Loading $INTEGRATION\n"
elastic_fleet_integration_check "endpoints-initial" "$INTEGRATION"
if ! elastic_fleet_integration_check "endpoints-initial" "$INTEGRATION"; then
ERROR=true
continue
fi
if [ -n "$INTEGRATION_ID" ]; then
printf "\n\nIntegration $NAME exists - Upgrading integration policy\n"
if ! elastic_fleet_integration_policy_upgrade "$INTEGRATION_ID"; then
@@ -7,20 +7,35 @@
. /usr/sbin/so-elastic-fleet-common
# Get all the fleet policies
json_output=$(curl -s -K /opt/so/conf/elasticsearch/curl.config -L -X GET "localhost:5601/api/fleet/agent_policies" -H 'kbn-xsrf: true')
if ! json_output=$(fleet_api "agent_policies" -H 'kbn-xsrf: true'); then
echo "Error: Failed to retrieve Fleet agent policies." >&2
exit 1
fi
if ! jq -e '.items' <<<"$json_output" >/dev/null 2>&1; then
echo "Error: Invalid Fleet agent policies response." >&2
exit 1
fi
# Extract the IDs that start with "FleetServer_"
POLICY=$(echo "$json_output" | jq -r '.items[] | select(.id | startswith("FleetServer_")) | .id')
POLICY=$(jq -r '.items[] | select(.id | startswith("FleetServer_")) | .id' <<<"$json_output")
# Iterate over each ID in the POLICY variable
for POLICYNAME in $POLICY; do
printf "\nUpdating Policy: $POLICYNAME\n"
# First get the Integration ID
INTEGRATION_ID=$(/usr/sbin/so-elastic-fleet-agent-policy-view "$POLICYNAME" | jq -r '.item.package_policies[] | select(.package.name == "fleet_server") | .id')
if ! POLICY_JSON=$(elastic_fleet_require_agent_policy "$POLICYNAME"); then
exit 1
fi
INTEGRATION_ID=$(jq -r '.item.package_policies[]? | select(.package.name == "fleet_server") | .id' <<<"$POLICY_JSON")
if [ -z "$INTEGRATION_ID" ]; then
echo "Error: fleet_server integration was not found in agent policy '$POLICYNAME'." >&2
exit 1
fi
# Modify the default integration policy to update the policy_id and an with the correct naming
UPDATED_INTEGRATION_POLICY=$(jq --arg policy_id "$POLICYNAME" --arg name "fleet_server-$POLICYNAME" '
UPDATED_INTEGRATION_POLICY=$(jq --arg policy_id "$POLICYNAME" --arg name "fleet_server-$POLICYNAME" '
.policy_id = $policy_id |
.name = $name' /opt/so/conf/elastic-fleet/integrations/fleet-server/fleet-server.json)
@@ -22,12 +22,19 @@ NUM_RUNNING=$(pgrep -cf "/bin/bash /sbin/so-elastic-agent-gen-installers")
for i in {1..30}
do
ENROLLMENTOKEN=$(curl -K /opt/so/conf/elasticsearch/curl.config -L "localhost:5601/api/fleet/enrollment_api_keys?perPage=100" -H 'kbn-xsrf: true' -H 'Content-Type: application/json' | jq .list | jq -r -c '.[] | select(.policy_id | contains("endpoints-initial")) | .api_key')
ENROLLMENTOKEN=$(elastic_fleet_active_enrollment_token "endpoints-initial")
TOKEN_RC=$?
if [ "$TOKEN_RC" -eq 2 ]; then
exit 1
fi
FLEETHOST=$(curl -K /opt/so/conf/elasticsearch/curl.config 'http://localhost:5601/api/fleet/fleet_server_hosts/grid-default' | jq -r '.item.host_urls[]' | paste -sd ',')
if [[ $FLEETHOST ]] && [[ $ENROLLMENTOKEN ]]; then break; else sleep 10; fi
if [[ -n "$FLEETHOST" ]] && [[ -n "$ENROLLMENTOKEN" ]]; then
break
fi
sleep 10
done
if [[ -z $FLEETHOST ]] || [[ -z $ENROLLMENTOKEN ]]; then
if [[ -z "$FLEETHOST" ]] || [[ -z "$ENROLLMENTOKEN" ]]; then
printf "\nFleet Host URL, Enrollment Token or Elastic Version empty - exiting..."
printf "\nFleet Host: $FLEETHOST, Enrollment Token: $ENROLLMENTOKEN\n"
exit 1
@@ -67,19 +74,25 @@ for GOOS in "${GOTARGETOS[@]}"; do
GOARCH="amd64"
if [[ $GOOS == 'darwin/arm64' ]]; then GOOS="darwin" && GOARCH="arm64"; fi
printf "\n\n### Generating $GOOS/$GOARCH Installer...\n"
docker run -e CGO_ENABLED=0 -e GOOS=$GOOS -e GOARCH=$GOARCH \
if ! docker run -e CGO_ENABLED=0 -e GOOS=$GOOS -e GOARCH=$GOARCH \
--mount type=bind,source=/etc/pki/tls/certs/,target=/workspace/files/cert/ \
--mount type=bind,source=/nsm/elastic-agent-workspace/,target=/workspace/files/elastic-agent/ \
--mount type=bind,source=/opt/so/saltstack/local/salt/elasticfleet/files/,target=/output/ \
{{ GLOBALS.registry_host }}:5000/{{ GLOBALS.image_repo }}/so-elastic-agent-builder:{{ GLOBALS.so_version }} go build -ldflags "-X main.fleetHostURLsList=$FLEETHOST -X main.enrollmentToken=$ENROLLMENTOKEN" -o /output/so-elastic-agent_${GOOS}_${GOARCH}
{{ GLOBALS.registry_host }}:5000/{{ GLOBALS.image_repo }}/so-elastic-agent-builder:{{ GLOBALS.so_version }} go build -ldflags "-X main.fleetHostURLsList=$FLEETHOST -X main.enrollmentToken=$ENROLLMENTOKEN" -o /output/so-elastic-agent_${GOOS}_${GOARCH}; then
printf "\n### ERROR: Failed to generate $GOOS/$GOARCH installer. Exiting...\n"
exit 1
fi
printf "\n### $GOOS/$GOARCH Installer Generated...\n"
done
printf "\n\n### Generating MSI...\n"
cp /opt/so/saltstack/local/salt/elasticfleet/files/so-elastic-agent_windows_amd64 /opt/so/saltstack/local/salt/elasticfleet/files/so-elastic-agent_windows_amd64.exe
docker run \
if ! docker run \
--mount type=bind,source=/opt/so/saltstack/local/salt/elasticfleet/files/,target=/output/ -w /output \
{{ GLOBALS.registry_host }}:5000/{{ GLOBALS.image_repo }}/so-elastic-agent-builder:{{ GLOBALS.so_version }} wixl -o so-elastic-agent_windows_amd64_msi --arch x64 /workspace/so-elastic-agent.wxs
{{ GLOBALS.registry_host }}:5000/{{ GLOBALS.image_repo }}/so-elastic-agent-builder:{{ GLOBALS.so_version }} wixl -o so-elastic-agent_windows_amd64_msi --arch x64 /workspace/so-elastic-agent.wxs; then
printf "\n### ERROR: Failed to generate MSI. Exiting...\n"
exit 1
fi
printf "\n### MSI Generated...\n"
# Verify installers were created
@@ -202,26 +202,9 @@ fi
### Finalization ###
# Query for Enrollment Tokens for default policies
if ENDPOINTSENROLLMENTOKEN_RAW=$(fleet_api "enrollment_api_keys" -H 'kbn-xsrf: true' -H 'Content-Type: application/json'); then
ENDPOINTSENROLLMENTOKEN=$(echo "$ENDPOINTSENROLLMENTOKEN_RAW" | jq .list | jq -r -c '.[] | select(.policy_id | contains("endpoints-initial")) | .api_key')
else
echo -e "\nFailed to query for Endpoints enrollment token"
exit 1
fi
if GRIDNODESENROLLMENTOKENGENERAL_RAW=$(fleet_api "enrollment_api_keys" -H 'kbn-xsrf: true' -H 'Content-Type: application/json'); then
GRIDNODESENROLLMENTOKENGENERAL=$(echo "$GRIDNODESENROLLMENTOKENGENERAL_RAW" | jq .list | jq -r -c '.[] | select(.policy_id | contains("so-grid-nodes_general")) | .api_key')
else
echo -e "\nFailed to query for Grid nodes - General enrollment token"
exit 1
fi
if GRIDNODESENROLLMENTOKENHEAVY_RAW=$(fleet_api "enrollment_api_keys" -H 'kbn-xsrf: true' -H 'Content-Type: application/json'); then
GRIDNODESENROLLMENTOKENHEAVY=$(echo "$GRIDNODESENROLLMENTOKENHEAVY_RAW" | jq .list | jq -r -c '.[] | select(.policy_id | contains("so-grid-nodes_heavy")) | .api_key')
else
echo -e "\nFailed to query for Grid nodes - Heavy enrollment token"
exit 1
fi
ENDPOINTSENROLLMENTOKEN=$(elastic_fleet_active_enrollment_token "endpoints-initial") || exit 1
GRIDNODESENROLLMENTOKENGENERAL=$(elastic_fleet_active_enrollment_token "so-grid-nodes_general") || exit 1
GRIDNODESENROLLMENTOKENHEAVY=$(elastic_fleet_active_enrollment_token "so-grid-nodes_heavy") || exit 1
# Store needed data in minion pillar
pillar_file=/opt/so/saltstack/local/pillar/minions/{{ GLOBALS.minion_id }}.sls
@@ -5,7 +5,8 @@
{ "rename": { "field": "message2.proto", "target_field": "network.transport", "ignore_missing": true } },
{ "rename": { "field": "message2.app_proto", "target_field": "network.protocol", "ignore_missing": true } },
{ "rename": { "field": "message2.fileinfo.filename", "target_field": "file.name", "ignore_missing": true } },
{ "rename": { "field": "message2.fileinfo.gaps", "target_field": "file.bytes.missing", "ignore_missing": true } },
{ "rename": { "field": "message2.fileinfo.gaps", "target_field": "suricata.fileinfo.gaps", "ignore_missing": true } },
{ "set": { "if": "ctx.suricata?.fileinfo?.gaps == false", "field": "file.bytes.missing", "value": 0 } },
{ "rename": { "field": "message2.fileinfo.magic", "target_field": "file.mime_type", "ignore_missing": true } },
{ "rename": { "field": "message2.fileinfo.md5", "target_field": "hash.md5", "ignore_missing": true } },
{ "rename": { "field": "message2.fileinfo.sha1", "target_field": "hash.sha1", "ignore_missing": true } },
+33 -14
View File
@@ -4,11 +4,19 @@
{%- set role = GLOBALS.role.split('-')[1] %}
{%- from 'firewall/containers.map.jinja' import NODE_CONTAINERS %}
{%- set NODE_NETWORKS = [] %}
{%- for NETNAME, NETWORK in DOCKERMERGED.networks.items() %}
{%- if not NETWORK.get('manager_only') or GLOBALS.get('is_manager', False) %}
{%- do NODE_NETWORKS.append(NETNAME) %}
{%- endif %}
{%- endfor %}
{%- set PR = [] %}
{%- set D1 = [] %}
{%- set D2 = [] %}
{%- for container in NODE_CONTAINERS %}
{%- set IP = DOCKERMERGED.containers[container].ip %}
{%- set BRIDGE = DOCKERMERGED.containers[container].network %}
{%- if DOCKERMERGED.containers[container].port_bindings is defined %}
{%- for binding in DOCKERMERGED.containers[container].port_bindings %}
{#- cant split int so we convert to string #}
@@ -35,11 +43,11 @@
{%- endif %}
{%- do PR.append("-A POSTROUTING -s " ~ DOCKERMERGED.containers[container].ip ~ "/32 -d " ~ DOCKERMERGED.containers[container].ip ~ "/32 -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ containerPort ~ " -j MASQUERADE") %}
{%- if bindip | length and bindip != '0.0.0.0' %}
{%- do D1.append("-A DOCKER -d " ~ bindip ~ "/32 ! -i sobridge -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ hostPort ~ " -j DNAT --to-destination " ~ DOCKERMERGED.containers[container].ip ~ ":" ~ containerPort) %}
{%- do D1.append("-A DOCKER -d " ~ bindip ~ "/32 ! -i " ~ BRIDGE ~ " -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ hostPort ~ " -j DNAT --to-destination " ~ DOCKERMERGED.containers[container].ip ~ ":" ~ containerPort) %}
{%- else %}
{%- do D1.append("-A DOCKER ! -i sobridge -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ hostPort ~ " -j DNAT --to-destination " ~ DOCKERMERGED.containers[container].ip ~ ":" ~ containerPort) %}
{%- do D1.append("-A DOCKER ! -i " ~ BRIDGE ~ " -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ hostPort ~ " -j DNAT --to-destination " ~ DOCKERMERGED.containers[container].ip ~ ":" ~ containerPort) %}
{%- endif %}
{%- do D2.append("-A DOCKER -d " ~ DOCKERMERGED.containers[container].ip ~ "/32 ! -i sobridge -o sobridge -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ containerPort ~ " -j ACCEPT") %}
{%- do D2.append("-A DOCKER -d " ~ DOCKERMERGED.containers[container].ip ~ "/32 ! -i " ~ BRIDGE ~ " -o " ~ BRIDGE ~ " -p " ~ proto ~ " -m " ~ proto ~ " --dport " ~ containerPort ~ " -j ACCEPT") %}
{%- endfor %}
{%- endif %}
{%- endfor %}
@@ -52,11 +60,15 @@
:DOCKER - [0:0]
-A PREROUTING -m addrtype --dst-type LOCAL -j DOCKER
-A OUTPUT ! -d 127.0.0.0/8 -m addrtype --dst-type LOCAL -j DOCKER
-A POSTROUTING -s {{DOCKERMERGED.range}} ! -o sobridge -j MASQUERADE
{%- for NETNAME in NODE_NETWORKS %}
-A POSTROUTING -s {{ DOCKERMERGED.networks[NETNAME].range }} ! -o {{ NETNAME }} -j MASQUERADE
{%- endfor %}
{%- for rule in PR %}
{{ rule }}
{%- endfor %}
-A DOCKER -i sobridge -j RETURN
{%- for NETNAME in NODE_NETWORKS %}
-A DOCKER -i {{ NETNAME }} -j RETURN
{%- endfor %}
{%- for rule in D1 %}
{{ rule }}
{%- endfor %}
@@ -97,10 +109,12 @@ COMMIT
{%- endif %}
-A FORWARD -j DOCKER-USER
-A FORWARD -j DOCKER-ISOLATION-STAGE-1
-A FORWARD -o sobridge -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A FORWARD -o sobridge -j DOCKER
-A FORWARD -i sobridge ! -o sobridge -j ACCEPT
-A FORWARD -i sobridge -o sobridge -j ACCEPT
{%- for NETNAME in NODE_NETWORKS %}
-A FORWARD -o {{ NETNAME }} -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A FORWARD -o {{ NETNAME }} -j DOCKER
-A FORWARD -i {{ NETNAME }} ! -o {{ NETNAME }} -j ACCEPT
-A FORWARD -i {{ NETNAME }} -o {{ NETNAME }} -j ACCEPT
{%- endfor %}
-A FORWARD -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A FORWARD -i lo -j ACCEPT
-A FORWARD -m conntrack --ctstate INVALID -j DROP
@@ -112,13 +126,18 @@ COMMIT
{%- for rule in D2 %}
{{ rule }}
{%- endfor %}
-A DOCKER-ISOLATION-STAGE-1 -i sobridge ! -o sobridge -j DOCKER-ISOLATION-STAGE-2
{% for NETNAME in NODE_NETWORKS %}
-A DOCKER-ISOLATION-STAGE-1 -i {{ NETNAME }} ! -o {{ NETNAME }} -j DOCKER-ISOLATION-STAGE-2
{%- endfor %}
-A DOCKER-ISOLATION-STAGE-1 -j RETURN
-A DOCKER-ISOLATION-STAGE-2 -o sobridge -j DROP
{%- for NETNAME in NODE_NETWORKS %}
-A DOCKER-ISOLATION-STAGE-2 -o {{ NETNAME }} -j DROP
{%- endfor %}
-A DOCKER-ISOLATION-STAGE-2 -j RETURN
-A DOCKER-USER ! -i sobridge -o sobridge -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-USER ! -i sobridge -o sobridge -j LOGGING
{%- for NETNAME in NODE_NETWORKS %}
-A DOCKER-USER ! -i {{ NETNAME }} -o {{ NETNAME }} -m conntrack --ctstate RELATED,ESTABLISHED -j ACCEPT
-A DOCKER-USER ! -i {{ NETNAME }} -o {{ NETNAME }} -j LOGGING
{%- endfor %}
-A DOCKER-USER -j RETURN
-A LOGGING -m limit --limit 2/min -j LOG --log-prefix "IPTables-dropped: "
-A LOGGING -j DROP
+6 -2
View File
@@ -4,8 +4,12 @@
{# add our ip to self #}
{% do FIREWALL_DEFAULT.firewall.hostgroups.self.append(GLOBALS.node_ip) %}
{# add dockernet range #}
{% do FIREWALL_DEFAULT.firewall.hostgroups.dockernet.append(DOCKERMERGED.range) %}
{# add dockernet ranges #}
{% for NETNAME, NETWORK in DOCKERMERGED.networks.items() %}
{% if not NETWORK.get('manager_only') or GLOBALS.get('is_manager', False) %}
{% do FIREWALL_DEFAULT.firewall.hostgroups.dockernet.append(NETWORK.range) %}
{% endif %}
{% endfor %}
{% if GLOBALS.role == 'so-idh' %}
{% from 'idh/opencanary_config.map.jinja' import IDH_PORTGROUPS %}
+3 -3
View File
@@ -26,8 +26,8 @@ so-hydra:
- hostname: hydra
- name: so-hydra
- networks:
- sobridge:
- ipv4_address: {{ DOCKERMERGED.containers['so-hydra'].ip }}
- soauth:
- ipv4_address: {{ DOCKERMERGED.containers['so-hydra'].ips['soauth'] }}
- binds:
- /opt/so/conf/hydra/:/hydra-conf:ro
- /opt/so/log/hydra/:/hydra-log:rw
@@ -73,7 +73,7 @@ delete_so-hydra_so-status.disabled:
wait_for_hydra:
http.wait_for_successful_query:
- name: 'http://{{ GLOBALS.manager }}:4444/health/alive'
- name: 'http://{{ DOCKERMERGED.containers['so-hydra'].ips['soauth'] }}:4444/health/alive'
- ssl: True
- verify_ssl: False
- status:
+3 -1
View File
@@ -94,9 +94,11 @@ metrics_link_file:
- docker_container: so-influxdb
# Install cron job to determine size of influxdb for telegraf
# telegraf reads this while the cron rewrites it, so write aside and rename rather than
# truncating in place. tgraflogdir recurses ownership, so the temp file is chowned to match
get_influxdb_size:
cron.present:
- name: 'du -s -k /nsm/influxdb | cut -f1 > /opt/so/log/telegraf/influxdb_size.log 2>&1'
- name: 'du -s -k /nsm/influxdb | cut -f1 > /opt/so/log/telegraf/influxdb_size.log.tmp 2>&1; chown 939:939 /opt/so/log/telegraf/influxdb_size.log.tmp; mv -f /opt/so/log/telegraf/influxdb_size.log.tmp /opt/so/log/telegraf/influxdb_size.log'
- identifier: get_influxdb_size
- user: root
- minute: '*/1'
+3 -3
View File
@@ -19,8 +19,8 @@ so-kratos:
- hostname: kratos
- name: so-kratos
- networks:
- sobridge:
- ipv4_address: {{ DOCKERMERGED.containers['so-kratos'].ip }}
- soauth:
- ipv4_address: {{ DOCKERMERGED.containers['so-kratos'].ips['soauth'] }}
- binds:
- /opt/so/conf/kratos/:/kratos-conf:ro
- /opt/so/log/kratos/:/kratos-log:rw
@@ -71,7 +71,7 @@ delete_so-kratos_so-status.disabled:
wait_for_kratos:
http.wait_for_successful_query:
- name: 'http://{{ GLOBALS.manager }}:4434/'
- name: 'http://{{ DOCKERMERGED.containers['so-kratos'].ips['soauth'] }}:4434/'
- ssl: True
- verify_ssl: False
- status:
+1 -1
View File
@@ -166,7 +166,7 @@ so-repo-sync:
so_fleetagent_status:
cron.present:
- name: /usr/sbin/so-elasticagent-status > /opt/so/log/agents/agentstatus.log 2>&1
- name: '/usr/sbin/so-elasticagent-status > /opt/so/log/agents/agentstatus.log.tmp 2>&1; mv -f /opt/so/log/agents/agentstatus.log.tmp /opt/so/log/agents/agentstatus.log'
- identifier: so_fleetagent_status
- user: root
- minute: '*/5'
+13 -8
View File
@@ -106,7 +106,8 @@ while [[ $# -gt 0 ]]; do
esac
done
hydraUrl=${HYDRA_URL:-http://127.0.0.1:4445}
hydraContainer=${HYDRA_CONTAINER:-so-hydra}
hydraUrl=${HYDRA_URL:-http://localhost:4445}
socRolesFile=${SOC_ROLES_FILE:-/opt/so/conf/soc/soc_clients_roles}
soUID=${SOCORE_UID:-939}
soGID=${SOCORE_GID:-939}
@@ -124,6 +125,10 @@ function fail() {
exit 1
}
function hydraCurl() {
docker exec "$hydraContainer" curl "$@"
}
function require() {
cmd=$1
which "$1" 2>&1 > /dev/null
@@ -133,8 +138,8 @@ function require() {
# Verify this environment is capable of running this script
function verifyEnvironment() {
require "jq"
require "curl"
response=$(curl -Ss -L ${hydraUrl}/health/alive)
require "docker"
response=$(hydraCurl -Ss -L ${hydraUrl}/health/alive)
[[ "$response" != '{"status":"ok"}' ]] && fail "Unable to communicate with Hydra; specify URL via HYDRA_URL environment variable"
}
@@ -164,7 +169,7 @@ function ensureRoleFileExists() {
}
function listClients() {
response=$(curl -Ss -L -f ${hydraUrl}/admin/clients)
response=$(hydraCurl -Ss -L -f ${hydraUrl}/admin/clients)
[[ $? != 0 ]] && fail "Unable to communicate with Hydra"
clientIds=$(echo "${response}" | jq -r ".[] | .client_id" | sort)
@@ -251,7 +256,7 @@ function createClient() {
EOF
)
response=$(curl -Ss -L --fail-with-body -X POST ${hydraUrl}/admin/clients -d "$body")
response=$(hydraCurl -Ss -L --fail-with-body -X POST ${hydraUrl}/admin/clients -d "$body")
if [[ $? != 0 ]]; then
error=$(echo $response | jq .error)
fail "Failed to submit request to Hydra: $error"
@@ -283,7 +288,7 @@ function update() {
EOF
)
response=$(curl -Ss -L --fail-with-body -X PATCH ${hydraUrl}/admin/clients/$id -d "$body")
response=$(hydraCurl -Ss -L --fail-with-body -X PATCH ${hydraUrl}/admin/clients/$id -d "$body")
if [[ $? != 0 ]]; then
error=$(echo $response | jq .error)
fail "Failed to submit request to Hydra: $error"
@@ -305,7 +310,7 @@ function generateSecret() {
EOF
)
response=$(curl -Ss -L --fail-with-body -X PATCH ${hydraUrl}/admin/clients/$id -d "$body")
response=$(hydraCurl -Ss -L --fail-with-body -X PATCH ${hydraUrl}/admin/clients/$id -d "$body")
if [[ $? != 0 ]]; then
error=$(echo $response | jq .error)
fail "Failed to submit request to Hydra: $error"
@@ -317,7 +322,7 @@ function deleteClient() {
[[ ${identityId} == "" ]] && fail "Client not found"
response=$(curl -Ss -XDELETE -L --fail-with-body "${hydraUrl}/admin/clients/$identityId")
response=$(hydraCurl -Ss -XDELETE -L --fail-with-body "${hydraUrl}/admin/clients/$identityId")
if [[ $? != 0 ]]; then
error=$(echo $response | jq .error)
fail "Failed to submit request to Hydra: $error"
+16 -11
View File
@@ -129,7 +129,8 @@ while [[ $# -gt 0 ]]; do
esac
done
kratosUrl=${KRATOS_URL:-http://127.0.0.1:4434/admin}
kratosContainer=${KRATOS_CONTAINER:-so-kratos}
kratosUrl=${KRATOS_URL:-http://localhost:4434/admin}
databasePath=${KRATOS_DB_PATH:-/nsm/kratos/db/db.sqlite}
databaseTimeout=${KRATOS_DB_TIMEOUT:-5000}
bcryptRounds=${BCRYPT_ROUNDS:-12}
@@ -154,6 +155,10 @@ function fail() {
exit 1
}
function kratosCurl() {
docker exec "$kratosContainer" curl "$@"
}
function require() {
cmd=$1
which "$1" 2>&1 > /dev/null
@@ -164,18 +169,18 @@ function require() {
function verifyEnvironment() {
require "htpasswd"
require "jq"
require "curl"
require "docker"
require "openssl"
require "sqlite3"
[[ ! -f $databasePath ]] && fail "Unable to find database file; specify path via KRATOS_DB_PATH environment variable"
response=$(curl -Ss -L ${kratosUrl}/)
response=$(kratosCurl -Ss -L ${kratosUrl}/)
[[ "$response" != "404 page not found" ]] && fail "Unable to communicate with Kratos; specify URL via KRATOS_URL environment variable"
}
function findIdByEmail() {
email=${1,,}
response=$(curl -Ss -L ${kratosUrl}/identities)
response=$(kratosCurl -Ss -L ${kratosUrl}/identities)
identityId=$(echo "${response}" | jq -r ".[] | select(.verifiable_addresses[0].value == \"$email\") | .id")
echo $identityId
}
@@ -416,7 +421,7 @@ function syncAll() {
}
function listUsers() {
response=$(curl -Ss -L ${kratosUrl}/identities)
response=$(kratosCurl -Ss -L ${kratosUrl}/identities)
[[ $? != 0 ]] && fail "Unable to communicate with Kratos"
users=$(echo "${response}" | jq -r ".[] | .verifiable_addresses[0].value" | sort)
@@ -495,7 +500,7 @@ function createUser() {
EOF
)
response=$(curl -Ss -L ${kratosUrl}/identities -d "$addUserJson")
response=$(kratosCurl -Ss -L ${kratosUrl}/identities -d "$addUserJson")
[[ $? != 0 ]] && fail "Unable to communicate with Kratos"
identityId=$(echo "${response}" | jq -r ".id")
@@ -518,7 +523,7 @@ function updateStatus() {
identityId=$(findIdByEmail "$email")
[[ ${identityId} == "" ]] && fail "User not found"
response=$(curl -Ss -L "${kratosUrl}/identities/$identityId")
response=$(kratosCurl -Ss -L "${kratosUrl}/identities/$identityId")
[[ $? != 0 ]] && fail "Unable to communicate with Kratos"
schemaId=$(echo "$response" | jq -r .schema_id)
@@ -531,7 +536,7 @@ function updateStatus() {
state="inactive"
fi
body="{ \"schema_id\": \"$schemaId\", \"state\": \"$state\", \"traits\": $traitBlock }"
response=$(curl -fSsL -XPUT -H "Content-Type: application/json" "${kratosUrl}/identities/$identityId" -d "$body")
response=$(kratosCurl -fSsL -XPUT -H "Content-Type: application/json" "${kratosUrl}/identities/$identityId" -d "$body")
[[ $? != 0 ]] && fail "Unable to update user"
}
@@ -550,7 +555,7 @@ function updateUserProfile() {
identityId=$(findIdByEmail "$email")
[[ ${identityId} == "" ]] && fail "User not found"
response=$(curl -Ss -L "${kratosUrl}/identities/$identityId")
response=$(kratosCurl -Ss -L "${kratosUrl}/identities/$identityId")
[[ $? != 0 ]] && fail "Unable to communicate with Kratos"
schemaId=$(echo "$response" | jq -r .schema_id)
@@ -559,7 +564,7 @@ function updateUserProfile() {
traitBlock="{\"email\":\"$email\",\"firstName\":\"$firstName\",\"lastName\":\"$lastName\",\"note\":\"$note\"}"
body="{ \"schema_id\": \"$schemaId\", \"state\": \"$state\", \"traits\": $traitBlock }"
response=$(curl -fSsL -XPUT -H "Content-Type: application/json" "${kratosUrl}/identities/$identityId" -d "$body")
response=$(kratosCurl -fSsL -XPUT -H "Content-Type: application/json" "${kratosUrl}/identities/$identityId" -d "$body")
[[ $? != 0 ]] && fail "Unable to update user"
}
@@ -569,7 +574,7 @@ function deleteUser() {
identityId=$(findIdByEmail "$email")
[[ ${identityId} == "" ]] && fail "User not found"
response=$(curl -Ss -XDELETE -L "${kratosUrl}/identities/$identityId")
response=$(kratosCurl -Ss -XDELETE -L "${kratosUrl}/identities/$identityId")
[[ $? != 0 ]] && fail "Unable to communicate with Kratos"
rolesTmpFile="${socRolesFile}.tmp"
+79
View File
@@ -28,6 +28,7 @@ INSTALLEDSALTVERSION=$(salt --versions-report | grep Salt: | awk '{print $2}')
# percentage like "25%"). Empty means so-soup-grid-highstate uses the salt:auto_apply:batch
# pillar default.
BATCHSIZE=
DEFAULT_DOCKER_RANGE='172.17.1.0/24'
SOUP_LOG=/root/soup.log
SOUP_DEBUG_LOG=/root/soup-debug.log
WHATWOULDYOUSAYYAHDOHERE=soup
@@ -605,6 +606,7 @@ preupgrade_changes() {
[[ "$INSTALLEDVERSION" == "3.0.0" ]] && up_to_3.1.0
[[ "$INSTALLEDVERSION" == "3.1.0" ]] && up_to_3.2.0
[[ "$INSTALLEDVERSION" == "3.2.0" ]] && up_to_3.3.0
[[ "$INSTALLEDVERSION" == "3.3.0" ]] && up_to_3.4.0
true
}
@@ -623,6 +625,7 @@ postupgrade_changes() {
[[ "$POSTVERSION" == "3.0.0" ]] && post_to_3.1.0
[[ "$POSTVERSION" == "3.1.0" ]] && post_to_3.2.0
[[ "$POSTVERSION" == "3.2.0" ]] && post_to_3.3.0
[[ "$POSTVERSION" == "3.3.0" ]] && post_to_3.4.0
# All applicable post-upgrade steps completed; clear the resume marker.
rm -f "$POSTVERSION_FILE"
true
@@ -1173,6 +1176,82 @@ post_to_3.3.0() {
}
### 3.3.0 End ###
### 3.4.0 Scripts ###
up_to_3.4.0() {
set_soauth_range
echo "Removing so-kratos, so-hydra and so-soc so they are recreated on the soauth network."
docker rm -f so-kratos so-hydra so-soc >> $SOUP_LOG 2>&1
INSTALLEDVERSION=3.4.0
}
set_soauth_range() {
local pillar_file=/opt/so/saltstack/local/pillar/docker/soc_docker.sls
local current_range suggested authnet authgw input
[[ -f "$pillar_file" ]] || return 0
current_range=$(so-yaml.py get -r "$pillar_file" docker.range 2>/dev/null) || return 0
# A default range gets the 172.17.2.0/24 from docker/defaults.yaml, same as a fresh
# install, so there is nothing to ask about.
[[ -n "$current_range" && "$current_range" != "$DEFAULT_DOCKER_RANGE" ]] || return 0
if so-yaml.py get -r "$pillar_file" docker.networks.soauth.range >/dev/null 2>&1; then
return 0
fi
suggested=$(echo "${current_range%%/*}" | awk -F'.' '{ printf "%s.%s.%s.%s", $1, $2, ($3 + 1) % 256, $4 }')
if [[ -z $UNATTENDED ]]; then
echo ""
echo "This grid uses a custom Docker range ($current_range). The authentication"
echo "services are moving to their own isolated network, which needs a second /24"
echo "that does not overlap it."
echo ""
while :; do
read -rp "Enter the network without the /24 suffix, or press Enter for ${suggested}: " input
[[ -z "$input" ]] && input="$suggested"
if valid_soauth_range "$input" "$current_range"; then
authnet="$input"
break
fi
echo "That range must be a valid IPv4 network, must not be within 172.17.0.0/24, and must not overlap ${current_range}."
done
else
if ! valid_soauth_range "$suggested" "$current_range"; then
FINAL_MESSAGE_QUEUE+=("WARNING: Unable to pick a range for the authentication network alongside $current_range. Set it manually before the next highstate:")
FINAL_MESSAGE_QUEUE+=(" - so-yaml.py add $pillar_file docker.networks.soauth.range <network>/24")
FINAL_MESSAGE_QUEUE+=(" - so-yaml.py add $pillar_file docker.networks.soauth.gateway <gateway>")
return 0
fi
authnet="$suggested"
FINAL_MESSAGE_QUEUE+=("NOTE: The authentication services moved to an isolated Docker network and were assigned ${authnet}/24.")
FINAL_MESSAGE_QUEUE+=(" - If that conflicts with your environment, update docker.networks.soauth in $pillar_file and run so-checkin.")
fi
authgw=$(echo "$authnet" | awk -F'.' '{print $1,$2,$3,1}' OFS='.')
echo "Assigning the authentication network the range ${authnet}/24."
so-yaml.py add "$pillar_file" docker.networks.soauth.range "${authnet}/24" >> $SOUP_LOG 2>&1
so-yaml.py add "$pillar_file" docker.networks.soauth.gateway "$authgw" >> $SOUP_LOG 2>&1
}
valid_soauth_range() {
local candidate=$1 docker_range=$2
valid_ip4 "$candidate" || return 1
[[ $candidate =~ ^172\.17\.0\. ]] && return 1
[[ "${candidate}/24" == "$docker_range" ]] && return 1
return 0
}
post_to_3.4.0() {
set_postversion 3.4.0
}
### 3.4.0 End ###
repo_sync() {
echo "Sync the local repo."
@@ -1,2 +1,3 @@
requests>=2.31.0
whoisit>=2.7.0
requests>=2.34.0
whoisit>=4.0.5
anyio>=4.15.1
+2
View File
@@ -14,6 +14,8 @@
{% do SOCDEFAULTS.soc.config.server.modules[module].update({'hostUrl': application_url}) %}
{% endfor %}
{% do SOCDEFAULTS.soc.config.server.modules.kratos.update({'publicHostUrl': 'http://' ~ DOCKERMERGED.containers['so-kratos'].ips['soauth'] ~ ':4433/'}) %}
{# add all grid heavy nodes to soc.server.modules.elastic.remoteHostUrls #}
{% for node_type, minions in salt['pillar.get']('elasticsearch:nodes', {}).items() %}
{% if node_type in ['heavynode'] %}
+5
View File
@@ -1380,6 +1380,7 @@ soc:
retryFailureMaxAttempts: 5
kratos:
hostUrl:
publicHostUrl:
hydra:
hostUrl:
elastalertengine:
@@ -1465,6 +1466,7 @@ soc:
- core
- emerging_threats_addon
useEsql: false
esqlCaseInsensitive: true
elastic:
hostUrl:
remoteHostUrls: []
@@ -1492,6 +1494,9 @@ soc:
org: Security Onion
bucket: telegraf/so_short_term
verifyCert: false
notification:
dismissedPruneDays: 30
enabled: false
playbook:
autoUpdateEnabled: true
playbookImportFrequencySeconds: 86400
+3 -1
View File
@@ -23,7 +23,9 @@ so-soc:
- name: so-soc
- networks:
- sobridge:
- ipv4_address: {{ DOCKERMERGED.containers['so-soc'].ip }}
- ipv4_address: {{ DOCKERMERGED.containers['so-soc'].ips['sobridge'] }}
- soauth:
- ipv4_address: {{ DOCKERMERGED.containers['so-soc'].ips['soauth'] }}
- binds:
- /nsm/rules:/nsm/rules:rw
- /opt/so/conf/strelka:/opt/sensoroni/yara:rw
+25
View File
@@ -1,6 +1,31 @@
name: Security Onion Baseline Pipeline
priority: 90
transformations:
# ES|QL scalar == returns null on multivalued fields; the
# backend reads this key and emits MV_INTERSECTS instead.
- id: declare_multivalue_fields
type: set_state
key: multivalue_fields
val:
- event.type
- event.action
- event.category
- tags
- process.args
- related.ip
- dns.resolved_ip
- id: esql_default_index
type: set_state
key: index
val: .ds-logs-*
- id: esql_source_metadata
type: set_state
key: metadata
val: "_id, _index, _source"
- id: esql_source_keep
type: set_state
key: keep
val: "_id, _index, _source"
- id: baseline_field_name_mapping
type: field_name_mapping
mapping:
+31
View File
@@ -155,6 +155,14 @@ soc:
description: Path to custom markdown templates for PDF report generation. All markdown files in this directory will be available as custom reports in the SOC Reports interface.
global: True
advanced: True
schedules:
title: Schedules
description: Schedules that are shared across the Security Onion product. Modify via one of the SOC Schedules view.
readonlyUi: True
global: True
forcedType: string
syntax: json
storage: db
subgrids:
title: Subordinate Grids
description: |
@@ -396,6 +404,11 @@ soc:
global: True
advanced: True
forcedType: bool
esqlCaseInsensitive:
description: "Match string values case-insensitively when converting Sigma rules. Applies to ES|QL only"
global: True
advanced: True
forcedType: bool
elastic:
index:
description: Comma-separated list of indices or index patterns (wildcard "*" supported) that SOC will search for records.
@@ -476,6 +489,24 @@ soc:
global: True
advanced: True
forcedType: bool
notification:
destinations:
title: Notification Destinations
description: JSON list of notifications. Modify via the SOC Notifications view.
readonlyUi: True
global: True
forcedType: string
syntax: json
storage: db
dismissedPruneDays:
title: Dismissed Retention Days
description: The number of days to retain dismissed notifications. When a notification is dismissed, it will be pruned after this many days. Only one user need dismiss a notification for it to be pruned.
forcedType: int
global: True
enabled:
description: Enables or disables the SOC notification module.
forcedType: bool
global: True
postgres:
host:
description: Hostname or IP address of the PostgreSQL server used by SOC. Defaults to the manager hostname.
+90 -10
View File
@@ -36,7 +36,7 @@ tgraf_sync_script_{{script}}:
- name: /opt/so/conf/telegraf/scripts/{{script}}
- user: root
- group: 939
- mode: 770
- mode: 750
- template: jinja
- source: salt://telegraf/scripts/{{script}}
- defaults:
@@ -49,7 +49,7 @@ tgraf_sync_script_esindexsize.sh:
- name: /opt/so/conf/telegraf/scripts/esindexsize.sh
- user: root
- group: 939
- mode: 770
- mode: 750
- source: salt://telegraf/scripts/esindexsize.sh
{# Copy conf/elasticsearch/curl.config for telegraf to use with esindexsize.sh #}
tgraf_sync_escurl_conf:
@@ -61,6 +61,80 @@ tgraf_sync_escurl_conf:
- source: salt://elasticsearch/curl.config
{% endif %}
# so-container-stats runs on the host as somon, a docker group member, so the container does
# not need the docker socket
somongroup:
group.present:
- name: somon
- gid: 961
# cron chdirs to $HOME before running a job, so home must exist
somon:
user.present:
- uid: 961
- gid: 961
- home: /opt/so/log/somon
- createhome: False
- shell: /sbin/nologin
- groups:
- docker
# renumbering an existing somon is a no-op on a fresh host and lets a host created
# before the id changed converge instead of failing the whole telegraf state
- allow_uid_change: True
- allow_gid_change: True
- require:
- group: somongroup
somonlogdir:
file.directory:
- name: /opt/so/log/somon
- user: 961
- group: 961
- mode: 755
# the lock file is not otherwise managed; recurse so a renumber rechowns it too
- recurse:
- user
- group
- require:
- user: somon
containers_log:
file.managed:
- name: /opt/so/log/somon/containers.log
- user: 961
- group: 961
- mode: 644
- replace: False
- require:
- file: somonlogdir
# telegraf reads on the same minute boundary the collector runs, and docker stats takes
# seconds, so write aside and rename rather than truncating the file telegraf is reading.
# ; not && so a failed run replaces the file instead of leaving stale metrics behind.
# flock -n keeps a run that outlives its minute from racing the next one over the same tmp
# file; the skipped run leaves a stale containers.log, which containers.sh discards by age
so-container-stats_cron:
cron.present:
- name: "flock -n /opt/so/log/somon/containers.lock -c '/usr/sbin/so-container-stats > /opt/so/log/somon/containers.log.tmp 2>&1; mv -f /opt/so/log/somon/containers.log.tmp /opt/so/log/somon/containers.log'"
- identifier: so-container-stats_cron
- user: somon
- minute: '*/1'
- hour: '*'
- daymonth: '*'
- month: '*'
- dayweek: '*'
- require:
- user: somon
# salt.lasthighstate touches this at order 9001, after the container starts; pre-create it so
# docker does not create a directory at the bind mount source
lasthighstate_placeholder:
file.managed:
- name: /opt/so/log/salt/lasthighstate
- mode: 644
- replace: False
- makedirs: True
telegraf_sbin:
file.recurse:
- name: /usr/sbin
@@ -69,14 +143,20 @@ telegraf_sbin:
- group: root
- file_mode: 755
#telegraf_sbin_jinja:
# file.recurse:
# - name: /usr/sbin
# - source: salt://telegraf/tools/sbin_jinja
# - user: 939
# - group: 939
# - file_mode: 755
# - template: jinja
# so-container-stats needs the per-stat toggles, so it renders instead of copying
tgraf_sbin_jinja:
file.recurse:
- name: /usr/sbin
- source: salt://telegraf/tools/sbin_jinja
- user: root
- group: root
- file_mode: 755
# the unit test lives beside the script; it must not ship or be rendered as jinja
- exclude_pat:
- "*_test.py"
- template: jinja
- defaults:
CONTAINER_STATS: {{ TELEGRAFMERGED.container_stats }}
tgrafconf:
file.managed:
+77
View File
@@ -10,6 +10,69 @@ telegraf:
flush_jitter: '0s'
debug: false
quiet: false
container_stats:
tags:
identity: False
engine:
n_containers: False
n_containers_running: False
n_containers_stopped: False
n_containers_paused: False
n_images: False
n_cpus: False
n_goroutines: False
n_used_file_descriptors: False
n_listener_events: False
memory_total: False
cpu:
usage_percent: True
usage_total: False
usage_in_usermode: False
usage_in_kernelmode: False
usage_system: False
throttling_periods: False
throttling_throttled_periods: False
throttling_throttled_time: False
container_id: False
mem:
usage_percent: True
usage: False
limit: False
max_usage: False
active_anon: False
active_file: False
inactive_anon: False
inactive_file: False
unevictable: False
pgfault: False
pgmajfault: False
container_id: False
net:
rx_bytes: True
rx_packets: False
rx_errors: False
rx_dropped: False
tx_bytes: False
tx_packets: False
tx_errors: False
tx_dropped: False
container_id: False
blkio:
io_service_bytes_recursive_read: False
io_service_bytes_recursive_write: False
container_id: False
status:
uptime_ns: True
oomkilled: True
pid: False
exitcode: False
restart_count: False
started_at: False
finished_at: False
container_id: False
health:
health_status: False
failing_streak: False
scripts:
eval:
- agentstatus.sh
@@ -19,6 +82,7 @@ telegraf:
- oldpcap.sh
- os.sh
- raid.sh
- containers.sh
- sostatus.sh
- suriloss.sh
- surirules.sh
@@ -34,6 +98,7 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
- suriloss.sh
- surirules.sh
@@ -47,6 +112,7 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
- features.sh
managerhype:
@@ -56,6 +122,7 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
- features.sh
managersearch:
@@ -66,12 +133,14 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
- features.sh
import:
- influxdbsize.sh
- lasthighstate.sh
- os.sh
- containers.sh
- sostatus.sh
sensor:
- checkfiles.sh
@@ -79,6 +148,7 @@ telegraf:
- oldpcap.sh
- os.sh
- raid.sh
- containers.sh
- sostatus.sh
- suriloss.sh
- surirules.sh
@@ -93,6 +163,7 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
- suriloss.sh
- surirules.sh
@@ -101,12 +172,14 @@ telegraf:
idh:
- lasthighstate.sh
- os.sh
- containers.sh
- sostatus.sh
searchnode:
- eps.sh
- lasthighstate.sh
- os.sh
- raid.sh
- containers.sh
- sostatus.sh
- features.sh
receiver:
@@ -115,16 +188,20 @@ telegraf:
- os.sh
- raid.sh
- redis.sh
- containers.sh
- sostatus.sh
fleet:
- lasthighstate.sh
- os.sh
- containers.sh
- sostatus.sh
hypervisor:
- lasthighstate.sh
- os.sh
- containers.sh
- sostatus.sh
desktop:
- lasthighstate.sh
- os.sh
- containers.sh
- sostatus.sh
+5
View File
@@ -13,6 +13,11 @@ so-telegraf:
docker_container.absent:
- force: True
so-container-stats_cron:
cron.absent:
- identifier: so-container-stats_cron
- user: somon
so-telegraf_so-status.disabled:
file.comment:
- name: /opt/so/conf/so-status/so-status.conf
+7 -4
View File
@@ -19,8 +19,7 @@ so-telegraf:
docker_container.running:
- image: {{ GLOBALS.registry_host }}:5000/{{ GLOBALS.image_repo }}/so-telegraf:{{ GLOBALS.so_version }}
- restart_policy: unless-stopped
- user: 939
- group_add: 939,920
- user: 939:939
- environment:
- HOST_ETC=/host/etc
- HOST_SYS=/host/sys
@@ -38,7 +37,6 @@ so-telegraf:
- /opt/so/conf/telegraf/etc/telegraf.conf:/etc/telegraf/telegraf.conf:ro
- /opt/so/conf/telegraf/node_config.json:/etc/telegraf/node_config.json:ro
- /var/run/utmp:/var/run/utmp:ro
- /var/run/docker.sock:/var/run/docker.sock:ro
- /:/host:ro
- /sys:/host/sys:ro
- /proc:/host/proc:ro
@@ -51,7 +49,8 @@ so-telegraf:
- /opt/so/log/suricata:/var/log/suricata:ro
- /opt/so/log/raid:/var/log/raid:ro
- /opt/so/log/sostatus:/var/log/sostatus:ro
- /opt/so/log/salt:/var/log/salt:ro
- /opt/so/log/somon:/var/log/somon:ro
- /opt/so/log/salt/lasthighstate:/var/log/salt/lasthighstate:ro
- /opt/so/log/agents:/var/log/agents:ro
{% if GLOBALS.is_manager or GLOBALS.role == 'so-heavynode' %}
- /opt/so/conf/telegraf/etc/escurl.config:/etc/telegraf/elasticsearch.config:ro
@@ -74,6 +73,8 @@ so-telegraf:
{% endfor %}
{% endif %}
- watch:
- file: tgraf_sbin_jinja
- file: lasthighstate_placeholder
- file: trusttheca
- x509: telegraf_crt
- x509: telegraf_key
@@ -83,6 +84,8 @@ so-telegraf:
- file: tgraf_sync_script_{{script}}
{% endfor %}
- require:
- file: lasthighstate_placeholder
- file: somonlogdir
- file: trusttheca
- x509: telegraf_crt
- x509: telegraf_key
+11 -7
View File
@@ -228,13 +228,6 @@
# ## bond interfaces.
# # bond_interfaces = ["bond0"]
# # Read metrics about docker containers
[[inputs.docker]]
# ## Docker Endpoint
# ## To use TCP, set endpoint = "tcp://[ip]:[port]"
# ## To use environment variables (ie, docker-machine), set endpoint = "ENV"
endpoint = "unix:///var/run/docker.sock"
#
# # Read stats from one or more Elasticsearch servers or clusters
{%- if GLOBALS.is_manager or GLOBALS.role == 'so-heavynode' %}
@@ -342,6 +335,17 @@
interval = "60s"
{%- endif %}
{%- if 'containers.sh' in TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]] %}
{%- do TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]].remove('containers.sh') %}
[[inputs.exec]]
commands = [
["/scripts/containers.sh"]
]
data_format = "influx"
timeout = "15s"
interval = "60s"
{%- endif %}
{%- if TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]] | length > 0 %}
[[inputs.exec]]
commands = [
+29
View File
@@ -0,0 +1,29 @@
#!/bin/bash
#
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# if this script isn't already running
if [[ ! "`pidof -x $(basename $0) -o %PPID`" ]]; then
CONTAINERSLOG=/var/log/somon/containers.log
# the collector rewrites this every minute; report nothing rather than repeating a stale
# file as if it were current, in case a run was skipped or the collector is wedged
MAXAGE=150
if [ -r "$CONTAINERSLOG" ]; then
AGE=$(( $(date +%s) - $(stat -c %Y "$CONTAINERSLOG") ))
if [ "$AGE" -le "$MAXAGE" ]; then
cat $CONTAINERSLOG
fi
fi
exit 0
fi
exit 0
+6 -4
View File
@@ -8,10 +8,12 @@
# if this script isn't already running
if [[ ! "`pidof -x $(basename $0) -o %PPID`" ]]; then
LAST_HIGHSTATE_END=$([ -e "/var/log/salt/lasthighstate" ] && date -r /var/log/salt/lasthighstate +%s || echo 0)
NOW=$(date +%s)
HIGHSTATE_AGE_SECONDS=$((NOW-LAST_HIGHSTATE_END))
echo "salt highstate_age_seconds=$HIGHSTATE_AGE_SECONDS"
if [ -r "/var/log/salt/lasthighstate" ]; then
LAST_HIGHSTATE_END=$(date -r /var/log/salt/lasthighstate +%s)
NOW=$(date +%s)
HIGHSTATE_AGE_SECONDS=$((NOW-LAST_HIGHSTATE_END))
echo "salt highstate_age_seconds=$HIGHSTATE_AGE_SECONDS"
fi
fi
+313
View File
@@ -53,6 +53,319 @@ telegraf:
forcedType: bool
advanced: True
helpLink: influxdb
container_stats:
tags:
identity:
description: Adds the container_image, container_version, engine_host and server_version tags to every container measurement, as the retired inputs.docker plugin did. Increases series cardinality. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
engine:
n_containers:
description: Total number of containers known to the Docker engine, running or not. Part of the engine-level docker measurement. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_containers_running:
description: Number of containers currently running. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_containers_stopped:
description: Number of containers currently stopped. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_containers_paused:
description: Number of containers currently paused. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_images:
description: Number of container images held by the Docker engine. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_cpus:
description: Number of CPUs the Docker engine reports for this host. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_goroutines:
description: Number of goroutines inside the Docker daemon. Diagnostic detail for the daemon itself. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_used_file_descriptors:
description: Number of file descriptors held open by the Docker daemon. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
n_listener_events:
description: Number of event listeners subscribed to the Docker daemon. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
memory_total:
description: Total physical memory the Docker engine reports for this host, in bytes. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
cpu:
usage_percent:
description: Percentage of host CPU consumed by the container. Required by the Container CPU Usage chart on the Security Onion Performance dashboard.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
usage_total:
description: Cumulative CPU time consumed by the container, in nanoseconds. Read from the cgroup. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
usage_in_usermode:
description: Cumulative CPU time consumed in user mode, in nanoseconds. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
usage_in_kernelmode:
description: Cumulative CPU time consumed in kernel mode, in nanoseconds. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
usage_system:
description: Cumulative host-wide CPU time, in nanoseconds. Used as the denominator when calculating container CPU percentage. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
throttling_periods:
description: Number of CPU enforcement periods the container has seen. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
throttling_throttled_periods:
description: Number of periods in which the container was throttled against its CPU limit. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
throttling_throttled_time:
description: Total time the container spent throttled, in nanoseconds. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
container_id: &containerid
description: Full 64 character container ID, emitted as a field on this measurement. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
mem:
usage_percent:
description: Percentage of its memory limit the container is using. Required by the Container Memory Usage chart on the Security Onion Performance dashboard.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
usage:
description: Container memory usage in bytes, excluding reclaimable page cache. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
limit:
description: Memory limit for the container in bytes. Reports total host memory when the container is unlimited. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
max_usage:
description: Peak memory usage for the container in bytes, read from the cgroup peak counter. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
active_anon:
description: Anonymous memory on the active LRU list, in bytes. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
active_file:
description: Page cache on the active LRU list, in bytes. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
inactive_anon:
description: Anonymous memory on the inactive LRU list, in bytes. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
inactive_file:
description: Page cache on the inactive LRU list, in bytes. This is the reclaimable cache subtracted from usage. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
unevictable:
description: Memory that cannot be reclaimed, in bytes. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
pgfault:
description: Cumulative number of page faults taken by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
pgmajfault:
description: Cumulative number of major page faults, those requiring disk access. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
container_id: *containerid
net:
rx_bytes:
description: Bytes received by the container across all interfaces except loopback. Required by the Container Traffic - Inbound chart on the Security Onion Performance dashboard.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
rx_packets:
description: Packets received by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
rx_errors:
description: Receive errors counted on the container interfaces. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
rx_dropped:
description: Received packets dropped by the container interfaces. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
tx_bytes:
description: Bytes transmitted by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
tx_packets:
description: Packets transmitted by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
tx_errors:
description: Transmit errors counted on the container interfaces. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
tx_dropped:
description: Transmitted packets dropped by the container interfaces. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
container_id: *containerid
blkio:
io_service_bytes_recursive_read:
description: Cumulative bytes read from block devices by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
io_service_bytes_recursive_write:
description: Cumulative bytes written to block devices by the container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
container_id: *containerid
status:
uptime_ns:
description: How long the container has been running, in nanoseconds. Required by the Container Uptime chart on the Security Onion Performance dashboard.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
oomkilled:
description: Whether the container was killed by the kernel out-of-memory handler. Required by the Most Recent Container Events table on the Security Onion Performance dashboard.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
pid:
description: Host process ID of the container main process. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
exitcode:
description: Exit code of the container main process. Meaningful once the container has stopped. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
restart_count:
description: Number of times the Docker engine has restarted this container. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
started_at:
description: Time the container last started, as a Unix timestamp in nanoseconds. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
finished_at:
description: Time the container last exited, as a Unix timestamp in nanoseconds. Absent for a container that has never exited. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
container_id: *containerid
health:
health_status:
description: Result of the container healthcheck, such as healthy, unhealthy or starting. Only emitted for containers that define a healthcheck. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
failing_streak:
description: Number of consecutive failed healthchecks. Only emitted for containers that define a healthcheck. Defaults to off.
forcedType: bool
global: True
advanced: True
helpLink: influxdb
scripts:
eval: &telegrafscripts
description: List of input.exec scripts to run for this node type. The script must be present in salt/telegraf/scripts.
@@ -0,0 +1,398 @@
#!/usr/bin/env python3
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# Runs from cron as somon, a docker group member, and emits influx line protocol for telegraf to
# read. This exists so so-telegraf does not need the docker socket; the container reads the output
# file instead.
# Measurement/tag/field names match telegraf's inputs.docker plugin because the InfluxDB
# dashboards query them directly. Which fields are emitted is set per stat in SOC; see
# telegraf.container_stats in defaults.yaml.
import json
SETTINGS = json.loads('''{{ CONTAINER_STATS | tojson }}''')
{% raw %}
import os
import re
import subprocess
import sys
from datetime import datetime, timezone
CGROUP_ROOT = '/sys/fs/cgroup'
NANOSEC = 10**9
# cgroup and docker report cpu time in microseconds; inputs.docker published nanoseconds
USEC_TO_NSEC = 1000
def want(group, field):
return bool(SETTINGS.get(group, {}).get(field))
def wants_any(group, fields):
return any(want(group, field) for field in fields)
def docker(args):
proc = subprocess.run(['docker'] + args, stdout=subprocess.PIPE, stderr=subprocess.DEVNULL, encoding='utf-8')
if proc.returncode != 0:
print('Container system error; unable to query docker', file=sys.stderr)
sys.exit(1)
return proc.stdout
def escape_tag(value):
return value.replace(',', '\\,').replace(' ', '\\ ').replace('=', '\\=')
def quote(value):
return '"{}"'.format(str(value).replace('\\', '\\\\').replace('"', '\\"'))
def integer(value):
return '{}i'.format(int(value))
def unsigned(value):
# inputs.docker published the cgroup and network counters as uint64; influx stores u and i
# as different field types, so matching it keeps historical series readable
return '{}u'.format(max(int(value), 0))
def read_text(path):
try:
with open(path) as handle:
return handle.read()
except OSError:
return ''
def read_pairs(path):
# cgroup files such as memory.stat and cpu.stat are "key value" per line
values = {}
for line in read_text(path).splitlines():
parts = line.split()
if len(parts) == 2:
try:
values[parts[0]] = int(parts[1])
except ValueError:
pass
return values
def read_value(path):
raw = read_text(path).strip()
try:
return int(raw)
except ValueError:
return None
def cgroup_path(pid):
# "0::/system.slice/docker-<id>.scope" on cgroup v2
for line in read_text('/proc/{}/cgroup'.format(pid)).splitlines():
parts = line.split(':', 2)
if len(parts) == 3 and parts[1] == '':
return CGROUP_ROOT + parts[2]
return None
def host_memory_total():
match = re.search(r'MemTotal:\s+(\d+) kB', read_text('/proc/meminfo'))
return int(match.group(1)) * 1024 if match else 0
def host_cpu_nanoseconds():
# inputs.docker's usage_system is the host-wide cpu time the daemon reads from /proc/stat
for line in read_text('/proc/stat').splitlines():
if line.startswith('cpu '):
ticks = sum(int(value) for value in line.split()[1:])
return int(ticks * NANOSEC / os.sysconf('SC_CLK_TCK'))
return 0
def parse_image(image):
# mirrors telegraf's internal/docker ParseImage so the tags match what inputs.docker emitted
domain = ''
remainder = image
if '/' in image:
head, _, tail = image.partition('/')
if '.' in head or ':' in head or head == 'localhost':
domain, remainder = head + '/', tail
if ':' in remainder:
name, _, version = remainder.rpartition(':')
return domain + name, version
return domain + remainder, 'unknown'
def to_nanoseconds(stamp):
# docker emits 9 fractional digits; fromisoformat takes at most 6 before python 3.11
stamp = stamp.rstrip('Z')
if '.' in stamp:
whole, _, frac = stamp.partition('.')
stamp = whole + '.' + frac[:6]
try:
parsed = datetime.fromisoformat(stamp).replace(tzinfo=timezone.utc)
except ValueError:
return None
if parsed.year <= 1:
return None
return int(parsed.timestamp() * NANOSEC)
def percent(value):
try:
return float(value.strip().rstrip('%'))
except ValueError:
return 0.0
def net_counters(pid):
# summed across interfaces except loopback, matching the daemon's per-container totals
rows = read_text('/proc/{}/net/dev'.format(pid)).splitlines()[2:]
if not rows:
return None
names = ['rx_bytes', 'rx_packets', 'rx_errors', 'rx_dropped']
totals = dict.fromkeys(names + ['tx_bytes', 'tx_packets', 'tx_errors', 'tx_dropped'], 0)
found = False
for row in rows:
iface, _, rest = row.partition(':')
if iface.strip() == 'lo':
continue
columns = rest.split()
if len(columns) < 12:
continue
found = True
for index, name in enumerate(names):
totals[name] += int(columns[index])
for index, name in enumerate(['tx_bytes', 'tx_packets', 'tx_errors', 'tx_dropped']):
totals[name] += int(columns[8 + index])
return totals if found else None
def blkio_counters(path):
# inputs.docker emitted the device=total row unconditionally, so a container that has done
# no block io reports zeros rather than dropping out of the measurement entirely
totals = {'io_service_bytes_recursive_read': 0, 'io_service_bytes_recursive_write': 0}
rows = read_text(path + '/io.stat').splitlines()
for row in rows:
for token in row.split()[1:]:
key, _, value = token.partition('=')
try:
if key == 'rbytes':
totals['io_service_bytes_recursive_read'] += int(value)
elif key == 'wbytes':
totals['io_service_bytes_recursive_write'] += int(value)
except ValueError:
pass
return totals
def collect_stats():
stats = {}
for line in docker(['stats', '--no-stream', '--format', '{{json .}}']).splitlines():
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
except json.JSONDecodeError:
continue
stats[entry.get('Name', '')] = entry
return stats
def collect_inspect():
ids = docker(['ps', '-aq']).split()
if not ids:
return []
try:
return json.loads(docker(['inspect'] + ids))
except json.JSONDecodeError:
return []
def collect_info():
try:
return json.loads(docker(['info', '--format', '{{json .}}']))
except json.JSONDecodeError:
return {}
class Emitter:
def __init__(self):
self.lines = []
def add(self, measurement, tags, fields):
if not fields:
return
tagset = ','.join('{}={}'.format(key, escape_tag(value)) for key, value in tags)
body = ','.join('{}={}'.format(key, fields[key]) for key in sorted(fields))
self.lines.append('{},{} {}'.format(measurement, tagset, body))
def main():
cpu_extra = ['usage_total', 'usage_in_usermode', 'usage_in_kernelmode',
'throttling_periods', 'throttling_throttled_periods', 'throttling_throttled_time']
mem_extra = ['usage', 'limit', 'max_usage', 'active_anon', 'active_file', 'inactive_anon',
'inactive_file', 'unevictable', 'pgfault', 'pgmajfault']
net_fields = ['rx_bytes', 'rx_packets', 'rx_errors', 'rx_dropped',
'tx_bytes', 'tx_packets', 'tx_errors', 'tx_dropped']
engine_fields = ['n_containers', 'n_containers_running', 'n_containers_stopped',
'n_containers_paused', 'n_images', 'n_cpus', 'n_goroutines',
'n_used_file_descriptors', 'n_listener_events']
identity = want('tags', 'identity')
need_stats = want('cpu', 'usage_percent')
need_cgroup_cpu = wants_any('cpu', cpu_extra)
need_cgroup_mem = wants_any('mem', mem_extra + ['usage_percent'])
need_blkio = wants_any('blkio', ['io_service_bytes_recursive_read',
'io_service_bytes_recursive_write', 'container_id'])
need_net = wants_any('net', net_fields + ['container_id'])
need_info = identity or wants_any('engine', engine_fields + ['memory_total'])
emitter = Emitter()
stats = collect_stats() if need_stats else {}
info = collect_info() if need_info else {}
engine_tags = [('engine_host', info.get('Name', '')),
('server_version', info.get('ServerVersion', ''))]
system_ns = host_cpu_nanoseconds() if want('cpu', 'usage_system') else 0
mem_total = host_memory_total() if wants_any('mem', ['limit', 'usage_percent']) else 0
if info:
counts = {'n_containers': 'Containers', 'n_containers_running': 'ContainersRunning',
'n_containers_stopped': 'ContainersStopped', 'n_containers_paused': 'ContainersPaused',
'n_images': 'Images', 'n_cpus': 'NCPU', 'n_goroutines': 'NGoroutines',
'n_used_file_descriptors': 'NFd', 'n_listener_events': 'NEventsListener'}
fields = {name: integer(info[key]) for name, key in counts.items()
if want('engine', name) and key in info}
emitter.add('docker', engine_tags, fields)
# inputs.docker published memory_total as its own point
if want('engine', 'memory_total') and 'MemTotal' in info:
emitter.add('docker', engine_tags, {'memory_total': integer(info['MemTotal'])})
for container in collect_inspect():
name = container.get('Name', '').lstrip('/')
if not name:
continue
state = container.get('State', {})
status = state.get('Status', 'unknown')
pid = state.get('Pid')
container_id = container.get('Id', '')
tags = [('container_name', name), ('container_status', status)]
if identity:
image, version = parse_image(container.get('Config', {}).get('Image', ''))
tags += [('container_image', image), ('container_version', version)] + engine_tags
status_fields = {}
if want('status', 'uptime_ns') or want('status', 'started_at') or want('status', 'finished_at'):
started = to_nanoseconds(state.get('StartedAt', ''))
finished = to_nanoseconds(state.get('FinishedAt', ''))
if started is not None:
if want('status', 'started_at'):
status_fields['started_at'] = integer(started)
if want('status', 'uptime_ns'):
end = finished if finished is not None and finished >= started else int(
datetime.now(timezone.utc).timestamp() * NANOSEC)
status_fields['uptime_ns'] = integer(end - started)
if finished is not None and want('status', 'finished_at'):
status_fields['finished_at'] = integer(finished)
if want('status', 'oomkilled'):
status_fields['oomkilled'] = 'true' if state.get('OOMKilled') else 'false'
if want('status', 'pid'):
status_fields['pid'] = integer(pid or 0)
if want('status', 'exitcode'):
status_fields['exitcode'] = integer(state.get('ExitCode', 0))
if want('status', 'restart_count'):
status_fields['restart_count'] = integer(container.get('RestartCount', 0))
if want('status', 'container_id'):
status_fields['container_id'] = quote(container_id)
emitter.add('docker_container_status', tags, status_fields)
health = state.get('Health')
if health:
health_fields = {}
if want('health', 'health_status'):
health_fields['health_status'] = quote(health.get('Status', ''))
if want('health', 'failing_streak'):
health_fields['failing_streak'] = integer(health.get('FailingStreak', 0))
emitter.add('docker_container_health', tags, health_fields)
if status != 'running' or not pid:
continue
path = cgroup_path(pid)
entry = stats.get(name, {})
cpu_fields = {}
if want('cpu', 'usage_percent'):
cpu_fields['usage_percent'] = percent(entry.get('CPUPerc', '0%'))
if want('cpu', 'usage_system'):
cpu_fields['usage_system'] = unsigned(system_ns)
if want('cpu', 'container_id'):
cpu_fields['container_id'] = quote(container_id)
if need_cgroup_cpu and path:
cpu = read_pairs(path + '/cpu.stat')
mapping = {'usage_total': 'usage_usec', 'usage_in_usermode': 'user_usec',
'usage_in_kernelmode': 'system_usec',
'throttling_periods': 'nr_periods',
'throttling_throttled_periods': 'nr_throttled',
'throttling_throttled_time': 'throttled_usec'}
for field, key in mapping.items():
if want('cpu', field) and key in cpu:
scale = 1 if field.startswith('throttling_') and field != 'throttling_throttled_time' else USEC_TO_NSEC
cpu_fields[field] = unsigned(cpu[key] * scale)
emitter.add('docker_container_cpu', tags + [('cpu', 'cpu-total')], cpu_fields)
mem_fields = {}
if want('mem', 'container_id'):
mem_fields['container_id'] = quote(container_id)
if need_cgroup_mem and path:
memory = read_pairs(path + '/memory.stat')
for field in ['active_anon', 'active_file', 'inactive_anon', 'inactive_file',
'unevictable', 'pgfault', 'pgmajfault']:
if want('mem', field) and field in memory:
mem_fields[field] = unsigned(memory[field])
current = read_value(path + '/memory.current')
# inputs.docker reports usage net of reclaimable page cache
usage = max(current - memory.get('inactive_file', 0), 0) if current is not None else None
raw_limit = read_value(path + '/memory.max')
limit = raw_limit if raw_limit is not None else mem_total
if want('mem', 'usage') and usage is not None:
mem_fields['usage'] = unsigned(usage)
if want('mem', 'limit'):
mem_fields['limit'] = unsigned(limit)
if want('mem', 'usage_percent') and usage is not None:
# same ratio inputs.docker computes, from the same two values
mem_fields['usage_percent'] = usage / limit * 100.0 if limit else 0.0
if want('mem', 'max_usage'):
peak = read_value(path + '/memory.peak')
if peak is not None:
mem_fields['max_usage'] = unsigned(peak)
emitter.add('docker_container_mem', tags, mem_fields)
# inputs.docker emitted nothing for host-network containers; its Networks map was empty
if need_net and container.get('HostConfig', {}).get('NetworkMode', '') != 'host':
counters = net_counters(pid)
if counters is not None:
net_out = {field: unsigned(counters[field]) for field in net_fields if want('net', field)}
if want('net', 'container_id'):
net_out['container_id'] = quote(container_id)
emitter.add('docker_container_net', tags + [('network', 'total')], net_out)
if need_blkio and path:
counters = blkio_counters(path)
if counters is not None:
blkio_out = {field: unsigned(value) for field, value in counters.items() if want('blkio', field)}
if want('blkio', 'container_id'):
blkio_out['container_id'] = quote(container_id)
emitter.add('docker_container_blkio', tags + [('device', 'total')], blkio_out)
print('\n'.join(emitter.lines))
if __name__ == '__main__':
main()
{% endraw %}
@@ -0,0 +1,635 @@
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# so-container-stats replaces telegraf's inputs.docker plugin, which was dropped along with the
# docker socket. The InfluxDB dashboards query its output by measurement, tag and field name, so
# these tests pin that contract: the shipped defaults, the per-stat toggles, the field types, and
# the value semantics copied from the plugin.
#
# The collector is a jinja template, so every test renders it the way salt does and imports the
# result. Docker and the cgroup filesystem are faked, so nothing here needs a container runtime.
import contextlib
import importlib.util
import io
import json
import os
import tempfile
import unittest
import jinja2
import yaml
HERE = os.path.dirname(os.path.abspath(__file__))
TEMPLATE = os.path.join(HERE, 'so-container-stats')
DEFAULTS = os.path.join(HERE, '..', '..', 'defaults.yaml')
# a container id is a 64 character hex string; the plugin published it in full
SOC_ID = 'a' * 64
TELEGRAF_ID = 'b' * 64
NGINX_ID = 'c' * 64
IDSTOOLS_ID = 'd' * 64
def shipped_defaults():
with open(DEFAULTS) as handle:
return yaml.safe_load(handle)['telegraf']['container_stats']
def all_enabled():
return {group: {field: True for field in fields} for group, fields in shipped_defaults().items()}
def render(settings):
"""Render the template as salt does, import it, and hand back the module."""
with open(TEMPLATE) as handle:
source = handle.read()
rendered = jinja2.Template(source, keep_trailing_newline=True).render(CONTAINER_STATS=settings)
path = os.path.join(tempfile.mkdtemp(), 'so_container_stats.py')
with open(path, 'w') as handle:
handle.write(rendered)
spec = importlib.util.spec_from_file_location('so_container_stats', path)
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
def split_escaped(text, sep):
"""Split on an unescaped, unquoted separator, the way influx line protocol is written."""
parts, current, escaped, quoted = [], '', False, False
for char in text:
if escaped:
current += char
escaped = False
elif char == '\\':
current += char
escaped = True
elif char == '"':
quoted = not quoted
current += char
elif char == sep and not quoted:
parts.append(current)
current = ''
else:
current += char
parts.append(current)
return parts
def split_on_space(line):
escaped = quoted = False
for index, char in enumerate(line):
if escaped:
escaped = False
elif char == '\\':
escaped = True
elif char == '"':
quoted = not quoted
elif char == ' ' and not quoted:
return line[:index], line[index + 1:]
return line, ''
def parse(output):
"""Parse line protocol into {(measurement, container): (tags, fields)}."""
points = {}
for line in output.splitlines():
line = line.strip()
if not line:
continue
head, fieldpart = split_on_space(line)
pieces = split_escaped(head, ',')
measurement, tags = pieces[0], {}
for piece in pieces[1:]:
key, _, value = piece.partition('=')
tags[key] = value
fields = {}
for piece in split_escaped(fieldpart, ','):
key, _, value = piece.partition('=')
fields[key] = value
key = (measurement, tags.get('container_name', ''))
if key in points:
# the engine measurement is published as two points, the way inputs.docker did
points[key][1].update(fields)
else:
points[key] = (tags, fields)
return points
def field_type(value):
if value.endswith('u'):
return 'unsigned'
if value.endswith('i'):
return 'integer'
if value.startswith('"'):
return 'string'
if value in ('true', 'false'):
return 'boolean'
return 'float'
def container(name, cid, pid, status='running', network='bridge', image='registry:5000/repo/img:3.4.0',
started='2026-09-24T10:00:00.123456789Z', finished='0001-01-01T00:00:00Z', health=None,
oomkilled=False, exitcode=0, restarts=0):
state = {'Status': status, 'Pid': pid, 'StartedAt': started, 'FinishedAt': finished,
'OOMKilled': oomkilled, 'ExitCode': exitcode}
if health is not None:
state['Health'] = health
return {'Id': cid, 'Name': '/' + name, 'State': state, 'RestartCount': restarts,
'Config': {'Image': image}, 'HostConfig': {'NetworkMode': network}}
class CollectorTestCase(unittest.TestCase):
"""Builds a fake docker engine and cgroup tree, then runs the collector against it."""
def setUp(self):
self.inspect = [
container('so-soc', SOC_ID, 1001),
container('so-telegraf', TELEGRAF_ID, 1002, network='host'),
container('so-nginx', NGINX_ID, 1003, health={'Status': 'healthy', 'FailingStreak': 0}),
]
self.stats = {
'so-soc': {'Name': 'so-soc', 'CPUPerc': '1.25%', 'MemPerc': '6.39%', 'NetIO': '1kB / 2kB'},
'so-telegraf': {'Name': 'so-telegraf', 'CPUPerc': '0.04%', 'MemPerc': '0.58%', 'NetIO': '0B / 0B'},
'so-nginx': {'Name': 'so-nginx', 'CPUPerc': '0.00%', 'MemPerc': '0.09%', 'NetIO': '3kB / 4kB'},
}
self.info = {'Name': 'sohost', 'ServerVersion': '29.2.1', 'Containers': 3, 'ContainersRunning': 3,
'ContainersStopped': 0, 'ContainersPaused': 0, 'Images': 9, 'NCPU': 8,
'NGoroutines': 42, 'NFd': 77, 'NEventsListener': 1, 'MemTotal': 16000000000}
# one cgroup per container, addressed through /proc/<pid>/cgroup exactly as the collector does
self.files = {
'/proc/meminfo': 'MemTotal: 15625000 kB\n',
'/proc/stat': 'cpu 100 200 300 400\ncpu0 1 2 3 4\n',
}
for pid in (1001, 1002, 1003):
self.files['/proc/%d/cgroup' % pid] = '0::/scope%d\n' % pid
self.cgroup(pid, 'cpu.stat',
'usage_usec 1000\nuser_usec 600\nsystem_usec 400\nnr_periods 5\nnr_throttled 2\nthrottled_usec 700\n')
self.cgroup(pid, 'memory.stat',
'active_anon 300\nactive_file 40\ninactive_anon 200\ninactive_file 50\nunevictable 0\npgfault 1234\npgmajfault 56\n')
self.cgroup(pid, 'memory.current', '1000\n')
self.cgroup(pid, 'memory.max', '4000\n')
self.cgroup(pid, 'memory.peak', '2500\n')
self.cgroup(pid, 'io.stat', '8:0 rbytes=100 wbytes=200 rios=1 wios=2\n252:0 rbytes=10 wbytes=20 rios=1 wios=1\n')
self.files['/proc/%d/net/dev' % pid] = (
'Inter-| Receive | Transmit\n'
' face |bytes packets errs drop fifo frame compressed multicast|bytes packets errs drop fifo colls carrier compressed\n'
' lo: 9999 99 9 9 0 0 0 0 9999 99 9 9 0 0 0 0\n'
' eth0: 1000 10 1 2 0 0 0 0 2000 20 3 4 0 0 0 0\n'
' eth1: 500 5 0 0 0 0 0 0 1000 10 0 0 0 0 0 0\n')
def cgroup(self, pid, name, contents):
self.files['/sys/fs/cgroup/scope%d/%s' % (pid, name)] = contents
def run_collector(self, settings=None, module=None):
module = module or render(settings if settings is not None else all_enabled())
def fake_docker(args):
if args[0] == 'stats':
return ''.join(json.dumps(entry) + '\n' for entry in self.stats.values())
if args[0] == 'ps':
return ' '.join(entry['Id'] for entry in self.inspect)
if args[0] == 'inspect':
return json.dumps(self.inspect)
if args[0] == 'info':
return json.dumps(self.info)
raise AssertionError('unexpected docker call: %s' % args)
module.docker = fake_docker
module.read_text = lambda path: self.files.get(path, '')
module.CGROUP_ROOT = '/sys/fs/cgroup'
buffer = io.StringIO()
with contextlib.redirect_stdout(buffer):
module.main()
self.output = buffer.getvalue()
return parse(self.output)
class TestShippedDefaults(CollectorTestCase):
def test_defaults_emit_only_the_dashboard_fields(self):
# the Security Onion Performance dashboard queries exactly these five
points = self.run_collector(shipped_defaults())
emitted = {(measurement, field) for (measurement, _), (_, fields) in points.items() for field in fields}
self.assertEqual(emitted, {
('docker_container_cpu', 'usage_percent'),
('docker_container_mem', 'usage_percent'),
('docker_container_net', 'rx_bytes'),
('docker_container_status', 'uptime_ns'),
('docker_container_status', 'oomkilled'),
})
def test_defaults_do_not_emit_the_opt_in_measurements(self):
points = self.run_collector(shipped_defaults())
measurements = {measurement for measurement, _ in points}
self.assertNotIn('docker', measurements)
self.assertNotIn('docker_container_blkio', measurements)
self.assertNotIn('docker_container_health', measurements)
def test_defaults_carry_the_tags_the_dashboard_filters_on(self):
points = self.run_collector(shipped_defaults())
tags, _ = points[('docker_container_cpu', 'so-soc')]
self.assertEqual(tags['container_status'], 'running')
self.assertEqual(tags['cpu'], 'cpu-total')
# identity tags are opt in, so they must be absent by default
self.assertNotIn('container_image', tags)
self.assertNotIn('engine_host', tags)
class TestToggles(CollectorTestCase):
def test_enabling_one_stat_adds_only_that_field(self):
settings = shipped_defaults()
settings['cpu']['usage_total'] = True
_, fields = self.run_collector(settings)[('docker_container_cpu', 'so-soc')]
self.assertIn('usage_total', fields)
self.assertNotIn('usage_in_usermode', fields)
def test_disabling_one_stat_leaves_its_neighbours(self):
settings = all_enabled()
settings['blkio']['io_service_bytes_recursive_read'] = False
_, fields = self.run_collector(settings)[('docker_container_blkio', 'so-soc')]
self.assertNotIn('io_service_bytes_recursive_read', fields)
self.assertIn('io_service_bytes_recursive_write', fields)
def test_identity_tags_are_added_when_enabled(self):
tags, _ = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
self.assertEqual(tags['container_image'], 'registry:5000/repo/img')
self.assertEqual(tags['container_version'], '3.4.0')
self.assertEqual(tags['engine_host'], 'sohost')
self.assertEqual(tags['server_version'], '29.2.1')
def test_everything_off_emits_nothing(self):
settings = {group: {field: False for field in fields} for group, fields in shipped_defaults().items()}
self.assertEqual(self.run_collector(settings), {})
class TestFieldTypes(CollectorTestCase):
"""inputs.docker wrote the cgroup and network counters as unsigned; influx treats u and i as
different field types, so a mismatch breaks queries spanning the change."""
def test_counter_fields_are_unsigned(self):
points = self.run_collector(all_enabled())
for measurement, field in (('docker_container_cpu', 'usage_total'),
('docker_container_cpu', 'throttling_periods'),
('docker_container_mem', 'usage'),
('docker_container_mem', 'limit'),
('docker_container_mem', 'pgfault'),
('docker_container_net', 'rx_bytes'),
('docker_container_blkio', 'io_service_bytes_recursive_read')):
_, fields = points[(measurement, 'so-soc')]
self.assertEqual(field_type(fields[field]), 'unsigned', '%s.%s' % (measurement, field))
def test_status_and_engine_fields_are_signed(self):
points = self.run_collector(all_enabled())
_, status = points[('docker_container_status', 'so-soc')]
for field in ('uptime_ns', 'pid', 'exitcode', 'restart_count', 'started_at'):
self.assertEqual(field_type(status[field]), 'integer', field)
_, engine = points[('docker', '')]
self.assertEqual(field_type(engine['n_containers']), 'integer')
def test_percentages_are_floats_and_ids_are_quoted_strings(self):
points = self.run_collector(all_enabled())
_, cpu = points[('docker_container_cpu', 'so-soc')]
self.assertEqual(field_type(cpu['usage_percent']), 'float')
self.assertEqual(field_type(cpu['container_id']), 'string')
self.assertEqual(cpu['container_id'], '"%s"' % SOC_ID)
_, status = points[('docker_container_status', 'so-soc')]
self.assertEqual(field_type(status['oomkilled']), 'boolean')
_, health = points[('docker_container_health', 'so-nginx')]
self.assertEqual(health['health_status'], '"healthy"')
class TestMemorySemantics(CollectorTestCase):
def test_usage_subtracts_reclaimable_page_cache(self):
# inputs.docker reports usage net of inactive_file: 1000 - 50
_, fields = self.run_collector(all_enabled())[('docker_container_mem', 'so-soc')]
self.assertEqual(fields['usage'], '950u')
def test_usage_percent_is_usage_over_limit(self):
_, fields = self.run_collector(all_enabled())[('docker_container_mem', 'so-soc')]
self.assertAlmostEqual(float(fields['usage_percent']), 950 / 4000 * 100.0)
def test_limit_falls_back_to_host_memory_when_unlimited(self):
for pid in (1001, 1002, 1003):
self.cgroup(pid, 'memory.max', 'max\n')
_, fields = self.run_collector(all_enabled())[('docker_container_mem', 'so-soc')]
self.assertEqual(fields['limit'], '%du' % (15625000 * 1024))
def test_max_usage_reports_the_cgroup_peak(self):
# the daemon reports 0 on cgroup v2, so this deliberately carries the real peak
_, fields = self.run_collector(all_enabled())[('docker_container_mem', 'so-soc')]
self.assertEqual(fields['max_usage'], '2500u')
class TestCpuSemantics(CollectorTestCase):
def test_microsecond_counters_are_published_as_nanoseconds(self):
_, fields = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
self.assertEqual(fields['usage_total'], '1000000u')
self.assertEqual(fields['usage_in_usermode'], '600000u')
self.assertEqual(fields['usage_in_kernelmode'], '400000u')
self.assertEqual(fields['throttling_throttled_time'], '700000u')
def test_throttling_counts_are_not_scaled(self):
_, fields = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
self.assertEqual(fields['throttling_periods'], '5u')
self.assertEqual(fields['throttling_throttled_periods'], '2u')
def test_usage_system_is_host_wide_cpu_time(self):
_, fields = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
self.assertEqual(field_type(fields['usage_system']), 'unsigned')
self.assertNotEqual(fields['usage_system'], '0u')
class TestNetwork(CollectorTestCase):
def test_counters_sum_interfaces_and_ignore_loopback(self):
_, fields = self.run_collector(all_enabled())[('docker_container_net', 'so-soc')]
self.assertEqual(fields['rx_bytes'], '1500u')
self.assertEqual(fields['rx_packets'], '15u')
self.assertEqual(fields['tx_bytes'], '3000u')
self.assertEqual(fields['rx_dropped'], '2u')
def test_host_network_containers_emit_no_row(self):
# inputs.docker skipped these: its Networks map is empty for --net=host
points = self.run_collector(all_enabled())
self.assertNotIn(('docker_container_net', 'so-telegraf'), points)
self.assertIn(('docker_container_net', 'so-soc'), points)
def test_total_tag_is_present(self):
tags, _ = self.run_collector(all_enabled())[('docker_container_net', 'so-soc')]
self.assertEqual(tags['network'], 'total')
class TestBlkio(CollectorTestCase):
def test_counters_sum_devices(self):
_, fields = self.run_collector(all_enabled())[('docker_container_blkio', 'so-soc')]
self.assertEqual(fields['io_service_bytes_recursive_read'], '110u')
self.assertEqual(fields['io_service_bytes_recursive_write'], '220u')
def test_container_with_no_io_reports_zero_rather_than_disappearing(self):
for pid in (1001, 1002, 1003):
self.cgroup(pid, 'io.stat', '')
tags, fields = self.run_collector(all_enabled())[('docker_container_blkio', 'so-soc')]
self.assertEqual(fields['io_service_bytes_recursive_read'], '0u')
self.assertEqual(tags['device'], 'total')
class TestStatus(CollectorTestCase):
def test_running_container_uptime_counts_from_start(self):
_, fields = self.run_collector(all_enabled())[('docker_container_status', 'so-soc')]
self.assertGreater(int(fields['uptime_ns'].rstrip('i')), 0)
self.assertNotIn('finished_at', fields)
def test_exited_container_reports_its_lifetime_and_finished_at(self):
self.inspect.append(container('so-idstools', IDSTOOLS_ID, 0, status='exited',
started='2026-09-24T10:00:00.000000000Z',
finished='2026-09-24T10:00:02.000000000Z', exitcode=3, restarts=1))
points = self.run_collector(all_enabled())
tags, fields = points[('docker_container_status', 'so-idstools')]
self.assertEqual(tags['container_status'], 'exited')
self.assertEqual(fields['uptime_ns'], '2000000000i')
self.assertEqual(fields['finished_at'], '1790244002000000000i')
self.assertEqual(fields['exitcode'], '3i')
self.assertEqual(fields['restart_count'], '1i')
# a stopped container has no live stats, so only the status row is emitted
self.assertNotIn(('docker_container_cpu', 'so-idstools'), points)
def test_health_is_emitted_only_for_containers_with_a_healthcheck(self):
points = self.run_collector(all_enabled())
self.assertIn(('docker_container_health', 'so-nginx'), points)
self.assertNotIn(('docker_container_health', 'so-soc'), points)
def test_oomkilled_is_reported(self):
self.inspect[0]['State']['OOMKilled'] = True
_, fields = self.run_collector(all_enabled())[('docker_container_status', 'so-soc')]
self.assertEqual(fields['oomkilled'], 'true')
class TestEngineMeasurement(CollectorTestCase):
def test_engine_counts_come_from_docker_info(self):
_, fields = self.run_collector(all_enabled())[('docker', '')]
self.assertEqual(fields['n_containers'], '3i')
self.assertEqual(fields['n_cpus'], '8i')
self.assertEqual(fields['n_used_file_descriptors'], '77i')
def test_engine_is_published_as_two_points(self):
# inputs.docker emitted memory_total on its own point, so keep that shape
self.run_collector(all_enabled())
engine = [line for line in self.output.splitlines() if line.startswith('docker,')]
self.assertEqual(len(engine), 2)
self.assertTrue(any('memory_total=' in line for line in engine))
self.assertTrue(any('n_containers=' in line for line in engine))
def test_engine_rows_are_tagged_with_host_and_version(self):
tags, _ = self.run_collector(all_enabled())[('docker', '')]
self.assertEqual(tags['engine_host'], 'sohost')
self.assertEqual(tags['server_version'], '29.2.1')
class TestLineProtocol(CollectorTestCase):
def test_tag_values_are_escaped(self):
self.inspect[0]['Name'] = '/odd name,with=chars'
self.stats['odd name,with=chars'] = self.stats.pop('so-soc')
self.stats['odd name,with=chars']['Name'] = 'odd name,with=chars'
output = self.run_collector(all_enabled())
self.assertIn(('docker_container_cpu', 'odd\\ name\\,with\\=chars'), output)
def test_image_without_a_tag_reports_version_unknown(self):
self.inspect[0]['Config']['Image'] = 'busybox'
tags, _ = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
self.assertEqual(tags['container_image'], 'busybox')
self.assertEqual(tags['container_version'], 'unknown')
def test_every_line_has_a_measurement_tagset_and_fieldset(self):
module = render(all_enabled())
def fake_docker(args):
if args[0] == 'stats':
return ''.join(json.dumps(entry) + '\n' for entry in self.stats.values())
if args[0] == 'ps':
return ' '.join(entry['Id'] for entry in self.inspect)
if args[0] == 'inspect':
return json.dumps(self.inspect)
return json.dumps(self.info)
module.docker = fake_docker
module.read_text = lambda path: self.files.get(path, '')
buffer = io.StringIO()
with contextlib.redirect_stdout(buffer):
module.main()
lines = [line for line in buffer.getvalue().splitlines() if line.strip()]
self.assertTrue(lines)
for line in lines:
head, fieldpart = split_on_space(line)
self.assertIn(',', head, line)
self.assertIn('=', fieldpart, line)
self.assertFalse(fieldpart.endswith(','), line)
# group -> (measurement, the container whose row carries it)
GROUP_TARGET = {
'engine': ('docker', ''),
'cpu': ('docker_container_cpu', 'so-soc'),
'mem': ('docker_container_mem', 'so-soc'),
'net': ('docker_container_net', 'so-soc'),
'blkio': ('docker_container_blkio', 'so-soc'),
'status': ('docker_container_status', 'so-soc'),
'health': ('docker_container_health', 'so-nginx'),
}
class TestEverySetting(CollectorTestCase):
"""Whatever is offered in defaults.yaml has to actually be collectable. These tests are driven
off that file, so a new setting that is never wired up fails here rather than shipping."""
def setUp(self):
super().setUp()
# an exited container so status.finished_at has a value to report
self.inspect.append(container('so-idstools', IDSTOOLS_ID, 0, status='exited',
started='2026-09-24T10:00:00.000000000Z',
finished='2026-09-24T10:00:02.000000000Z'))
def test_every_setting_emits_its_field_when_enabled(self):
points = self.run_collector(all_enabled())
for group, fields in shipped_defaults().items():
if group == 'tags':
continue
measurement, name = GROUP_TARGET[group]
if group == 'status':
# finished_at only exists for a container that has actually exited
_, exited = points[(measurement, 'so-idstools')]
self.assertIn('finished_at', exited)
_, emitted = points[(measurement, name)]
for field in fields:
if group == 'status' and field == 'finished_at':
continue
self.assertIn(field, emitted, '%s.%s is offered but never emitted' % (group, field))
def test_every_setting_is_individually_wired(self):
# enabling one stat on its own must produce exactly that field, proving each toggle is
# read rather than riding along with a neighbour
for group, fields in shipped_defaults().items():
if group == 'tags':
continue
measurement, name = GROUP_TARGET[group]
for field in fields:
settings = {other: {key: False for key in values} for other, values in shipped_defaults().items()}
settings[group][field] = True
target = 'so-idstools' if (group == 'status' and field == 'finished_at') else name
points = self.run_collector(settings)
self.assertIn((measurement, target), points, '%s.%s emitted no row' % (group, field))
_, emitted = points[(measurement, target)]
self.assertEqual(sorted(emitted), [field], '%s.%s did not emit itself alone' % (group, field))
def test_all_enabled_values_are_exact(self):
points = self.run_collector(all_enabled())
clock = os.sysconf('SC_CLK_TCK')
expected = {
('docker_container_cpu', 'so-soc'): {
'usage_percent': '1.25', 'usage_total': '1000000u', 'usage_in_usermode': '600000u',
'usage_in_kernelmode': '400000u', 'usage_system': '%du' % int(1000 * 10**9 / clock),
'throttling_periods': '5u', 'throttling_throttled_periods': '2u',
'throttling_throttled_time': '700000u', 'container_id': '"%s"' % SOC_ID,
},
('docker_container_mem', 'so-soc'): {
'usage': '950u', 'limit': '4000u', 'max_usage': '2500u', 'active_anon': '300u',
'active_file': '40u', 'inactive_anon': '200u', 'inactive_file': '50u',
'unevictable': '0u', 'pgfault': '1234u', 'pgmajfault': '56u',
'usage_percent': '23.75', 'container_id': '"%s"' % SOC_ID,
},
('docker_container_net', 'so-soc'): {
'rx_bytes': '1500u', 'rx_packets': '15u', 'rx_errors': '1u', 'rx_dropped': '2u',
'tx_bytes': '3000u', 'tx_packets': '30u', 'tx_errors': '3u', 'tx_dropped': '4u',
'container_id': '"%s"' % SOC_ID,
},
('docker_container_blkio', 'so-soc'): {
'io_service_bytes_recursive_read': '110u',
'io_service_bytes_recursive_write': '220u', 'container_id': '"%s"' % SOC_ID,
},
('docker', ''): {
'n_containers': '3i', 'n_containers_running': '3i', 'n_containers_stopped': '0i',
'n_containers_paused': '0i', 'n_images': '9i', 'n_cpus': '8i', 'n_goroutines': '42i',
'n_used_file_descriptors': '77i', 'n_listener_events': '1i',
'memory_total': '16000000000i',
},
('docker_container_health', 'so-nginx'): {
'health_status': '"healthy"', 'failing_streak': '0i',
},
}
for key, fields in expected.items():
_, emitted = points[key]
for field, value in fields.items():
self.assertEqual(emitted[field], value, '%s %s' % (key[0], field))
def test_status_values_are_exact(self):
# uptime is relative to now, so it is checked separately from the fixed fields
points = self.run_collector(all_enabled())
_, running = points[('docker_container_status', 'so-soc')]
self.assertEqual(running['pid'], '1001i')
self.assertEqual(running['exitcode'], '0i')
self.assertEqual(running['restart_count'], '0i')
self.assertEqual(running['oomkilled'], 'false')
self.assertEqual(running['container_id'], '"%s"' % SOC_ID)
self.assertEqual(int(running['started_at'].rstrip('i')) // 10**9, 1790244000)
self.assertGreater(int(running['uptime_ns'].rstrip('i')), 0)
_, exited = points[('docker_container_status', 'so-idstools')]
self.assertEqual(exited['uptime_ns'], '2000000000i')
self.assertEqual(exited['finished_at'], '1790244002000000000i')
def test_tags_identity_toggle_controls_the_identity_tags(self):
settings = shipped_defaults()
settings['tags']['identity'] = False
tags, _ = self.run_collector(settings)[('docker_container_cpu', 'so-soc')]
self.assertEqual(sorted(tags), ['container_name', 'container_status', 'cpu'])
settings['tags']['identity'] = True
tags, _ = self.run_collector(settings)[('docker_container_cpu', 'so-soc')]
self.assertEqual(sorted(tags), ['container_image', 'container_name', 'container_status',
'container_version', 'cpu', 'engine_host', 'server_version'])
def test_identity_tags_cover_every_documented_tag(self):
tags, _ = self.run_collector(all_enabled())[('docker_container_cpu', 'so-soc')]
for tag in ('container_image', 'container_version', 'engine_host', 'server_version'):
self.assertIn(tag, tags)
class TestTemplate(unittest.TestCase):
def test_template_renders_to_valid_python_for_the_shipped_defaults(self):
with open(TEMPLATE) as handle:
source = handle.read()
rendered = jinja2.Template(source, keep_trailing_newline=True).render(CONTAINER_STATS=shipped_defaults())
compile(rendered, 'so-container-stats', 'exec')
self.assertNotIn('{%', rendered)
# the docker format strings must survive rendering untouched
self.assertIn('{{json .}}', rendered)
def test_every_annotated_setting_exists_in_defaults(self):
# SOC reads both trees; an annotation without a default cannot be reverted in the UI
with open(os.path.join(HERE, '..', '..', 'soc_telegraf.yaml')) as handle:
annotated = yaml.safe_load(handle)['telegraf']['container_stats']
defaults = shipped_defaults()
for group, fields in annotated.items():
self.assertIn(group, defaults)
for field in fields:
self.assertIn(field, defaults[group], '%s.%s annotated but missing from defaults' % (group, field))
def test_every_default_setting_is_annotated_for_soc(self):
with open(os.path.join(HERE, '..', '..', 'soc_telegraf.yaml')) as handle:
annotated = yaml.safe_load(handle)['telegraf']['container_stats']
for group, fields in shipped_defaults().items():
self.assertIn(group, annotated)
for field in fields:
self.assertIn(field, annotated[group], '%s.%s missing a SOC annotation' % (group, field))
if __name__ == '__main__':
unittest.main()
+2 -2
View File
@@ -54,8 +54,8 @@
{%
do GLOBALS.update({
'application_urls': {
'hydra': 'http://' ~ GLOBALS.manager ~ ':4445/',
'kratos': 'http://' ~ GLOBALS.manager ~ ':4434/',
'hydra': 'http://' ~ DOCKERMERGED.containers['so-hydra'].ips['soauth'] ~ ':4445/',
'kratos': 'http://' ~ DOCKERMERGED.containers['so-kratos'].ips['soauth'] ~ ':4434/',
'elastic': 'https://' ~ GLOBALS.manager ~ ':9200/',
'influxdb': 'https://' ~ GLOBALS.manager ~ ':8086/'
}
+21
View File
@@ -276,9 +276,20 @@ collect_dockernet() {
whiptail_invalid_input
whiptail_dockernet_sosnet "$DOCKERNET"
done
whiptail_authnet_sosnet "$(adjacent_net "$DOCKERNET")"
while ! valid_ip4 "$AUTHNET" || [[ $AUTHNET =~ "172.17.0." ]] || [[ "$AUTHNET" == "$DOCKERNET" ]]; do
whiptail_invalid_input
whiptail_authnet_sosnet "$AUTHNET"
done
fi
}
adjacent_net() {
echo "$1" | awk -F'.' '{ printf "%s.%s.%s.%s", $1, $2, ($3 + 1) % 256, $4 }'
}
collect_gateway() {
whiptail_management_interface_gateway
@@ -1399,6 +1410,15 @@ docker_pillar() {
"docker:"\
" range: '$DOCKERNET/24'"\
" gateway: '$DOCKERGATEWAY'" > $docker_pillar_file
if [ ! -z "$AUTHNET" ]; then
AUTHGATEWAY=$(echo $AUTHNET | awk -F'.' '{print $1,$2,$3,1}' OFS='.')
printf '%s\n'\
" networks:"\
" soauth:"\
" range: '$AUTHNET/24'"\
" gateway: '$AUTHGATEWAY'" >> $docker_pillar_file
fi
fi
}
@@ -1591,6 +1611,7 @@ reserve_group_ids() {
logCmd "groupadd -g 949 elastic-agent"
logCmd "groupadd -g 947 elastic-fleet"
logCmd "groupadd -g 960 kafka"
logCmd "groupadd -g 961 somon"
}
reserve_ports() {
+13
View File
@@ -365,6 +365,18 @@ whiptail_dockernet_sosnet() {
}
whiptail_authnet_sosnet() {
[ -n "$TESTING" ] && return
AUTHNET=$(whiptail --title "$whiptail_title" --inputbox \
"\nEnter a second /24 size network range WITHOUT the /24 suffix. The authentication services are isolated on their own network so that the identity provider is not reachable from other containers. It must not overlap the range you just entered, and any range within 172.17.0.0/24 cannot be used." 13 65 "$1" 3>&1 1>&2 2>&3)
local exitstatus=$?
whiptail_check_exitstatus $exitstatus
}
whiptail_end_settings() {
[ -n "$TESTING" ] && return
@@ -427,6 +439,7 @@ whiptail_end_settings() {
[[ -n $WEBUSER ]] && __append_end_msg "Web User: $WEBUSER"
[[ -n $DOCKERNET ]] && __append_end_msg "Docker network: $DOCKERNET/24"
[[ -n $AUTHNET ]] && __append_end_msg "Authentication network: $AUTHNET/24"
if [[ ${#ntp_servers[@]} -gt 0 ]]; then
__append_end_msg "NTP Servers:"
for server in "${ntp_servers[@]}"; do