Files
securityonion/salt/telegraf/config.sls
Josh Patterson 72f60fcaa9 FIX: remove the docker socket from so-telegraf
so-telegraf mounted /var/run/docker.sock and joined the host docker group on
every node type. The :ro flag blocks write() to the inode, not connect() plus
HTTP over the socket, so any code execution inside the container could reach
POST /containers/create with Privileged:true and become root on the host.

No telegraf script used the socket; the only consumer was the native
[[inputs.docker]] plugin, and group_add 920 existed solely to feed it.
Container metrics now come from so-container-stats, a collector that runs on
the host from cron and writes influx line protocol to a file telegraf already
had mounted. This is the pattern so-status, so-raid-status and
so-elasticagent-status already use, so the privileged docker access stays on
the host side where root cron already ran it.

The collector runs as somon, a service account in the docker group with no
login shell and a locked password, rather than root. Docker group membership
is still root-equivalent on the host, so this is defense in depth rather than
a privilege boundary.

The telegraf scripts were root:939 mode 770, letting socore rewrite them for
code execution inside the container; they are now 750, which still allows the
read and execute telegraf needs. The container also ran with no group, giving
it gid 0, and now runs as 939:939. That alone would have broken
lasthighstate.sh, which reached /opt/so/log/salt only via the root group and
could not tell an unreadable file from a missing one, so it silently reported
a 56 year highstate age. Only the lasthighstate file is bind mounted now, and
the script tests readability instead of existence.

Everything inputs.docker collected beyond the five fields the shipped
dashboards query is available per stat under telegraf:container_stats,
annotated for SOC so an operator can enable it without editing files. Defaults
reproduce the previous output exactly. With every stat enabled the emitted
field set matches what inputs.docker wrote, verified by running the plugin
against the live socket and diffing: 53 fields, no type mismatches, no field
present on one side only.

Two deliberate differences: max_usage carries the real cgroup peak where the
daemon reports 0 on cgroup v2, and host-network containers emit no
docker_container_net row, matching inputs.docker. Docker label tags are not
restored, since nothing queries them and they cost significant cardinality.

so-status, the influxdb size cron, so-elasticagent-status, so-raid-status and
so-common-status-check truncated their output in place while telegraf read it,
so telegraf periodically saw an empty file and logged a parse error or emitted
empty values. They now write aside and rename. Measured on a live manager, the
old so-status cron left status.log empty for 215 of 10997 reads.

Tested on a fresh install, a converted grid and a 3.0 upgrade.
2026-09-24 14:30:25 -04:00

190 lines
5.5 KiB
YAML+Jinja

# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
{% from 'allowed_states.map.jinja' import allowed_states %}
{% if sls.split('.')[0] in allowed_states %}
{% from 'vars/globals.map.jinja' import GLOBALS %}
{% from 'telegraf/map.jinja' import TELEGRAFMERGED %}
{% from 'logstash/map.jinja' import LOGSTASH_MERGED %}
# add Telegraf to monitor all the things
tgraflogdir:
file.directory:
- name: /opt/so/log/telegraf
- makedirs: True
- user: 939
- group: 939
- recurse:
- user
- group
tgrafetcdir:
file.directory:
- name: /opt/so/conf/telegraf/etc
- makedirs: True
tgrafetsdir:
file.directory:
- name: /opt/so/conf/telegraf/scripts
- makedirs: True
{% for script in TELEGRAFMERGED.scripts[GLOBALS.role.split('-')[1]] %}
tgraf_sync_script_{{script}}:
file.managed:
- name: /opt/so/conf/telegraf/scripts/{{script}}
- user: root
- group: 939
- mode: 750
- template: jinja
- source: salt://telegraf/scripts/{{script}}
- defaults:
GLOBALS: {{ GLOBALS }}
{% endfor %}
{% if GLOBALS.is_manager or GLOBALS.role == 'so-heavynode' %}
tgraf_sync_script_esindexsize.sh:
file.managed:
- name: /opt/so/conf/telegraf/scripts/esindexsize.sh
- user: root
- group: 939
- mode: 750
- source: salt://telegraf/scripts/esindexsize.sh
{# Copy conf/elasticsearch/curl.config for telegraf to use with esindexsize.sh #}
tgraf_sync_escurl_conf:
file.managed:
- name: /opt/so/conf/telegraf/etc/escurl.config
- user: 939
- group: 939
- mode: 400
- source: salt://elasticsearch/curl.config
{% endif %}
# so-container-stats runs on the host as somon, a docker group member, so the container does
# not need the docker socket
somongroup:
group.present:
- name: somon
- gid: 961
# cron chdirs to $HOME before running a job, so home must exist
somon:
user.present:
- uid: 961
- gid: 961
- home: /opt/so/log/somon
- createhome: False
- shell: /sbin/nologin
- groups:
- docker
# renumbering an existing somon is a no-op on a fresh host and lets a host created
# before the id changed converge instead of failing the whole telegraf state
- allow_uid_change: True
- allow_gid_change: True
- require:
- group: somongroup
somonlogdir:
file.directory:
- name: /opt/so/log/somon
- user: 961
- group: 961
- mode: 755
# the lock file is not otherwise managed; recurse so a renumber rechowns it too
- recurse:
- user
- group
- require:
- user: somon
containers_log:
file.managed:
- name: /opt/so/log/somon/containers.log
- user: 961
- group: 961
- mode: 644
- replace: False
- require:
- file: somonlogdir
# telegraf reads on the same minute boundary the collector runs, and docker stats takes
# seconds, so write aside and rename rather than truncating the file telegraf is reading.
# ; not && so a failed run replaces the file instead of leaving stale metrics behind.
# flock -n keeps a run that outlives its minute from racing the next one over the same tmp
# file; the skipped run leaves a stale containers.log, which containers.sh discards by age
so-container-stats_cron:
cron.present:
- name: "flock -n /opt/so/log/somon/containers.lock -c '/usr/sbin/so-container-stats > /opt/so/log/somon/containers.log.tmp 2>&1; mv -f /opt/so/log/somon/containers.log.tmp /opt/so/log/somon/containers.log'"
- identifier: so-container-stats_cron
- user: somon
- minute: '*/1'
- hour: '*'
- daymonth: '*'
- month: '*'
- dayweek: '*'
- require:
- user: somon
# salt.lasthighstate touches this at order 9001, after the container starts; pre-create it so
# docker does not create a directory at the bind mount source
lasthighstate_placeholder:
file.managed:
- name: /opt/so/log/salt/lasthighstate
- mode: 644
- replace: False
- makedirs: True
telegraf_sbin:
file.recurse:
- name: /usr/sbin
- source: salt://telegraf/tools/sbin
- user: root
- group: root
- file_mode: 755
# so-container-stats needs the per-stat toggles, so it renders instead of copying
tgraf_sbin_jinja:
file.recurse:
- name: /usr/sbin
- source: salt://telegraf/tools/sbin_jinja
- user: root
- group: root
- file_mode: 755
# the unit test lives beside the script; it must not ship or be rendered as jinja
- exclude_pat:
- "*_test.py"
- template: jinja
- defaults:
CONTAINER_STATS: {{ TELEGRAFMERGED.container_stats }}
tgrafconf:
file.managed:
- name: /opt/so/conf/telegraf/etc/telegraf.conf
- user: 939
- group: 939
- mode: 660
- template: jinja
- source: salt://telegraf/etc/telegraf.conf
- show_changes: False
- defaults:
GLOBALS: {{ GLOBALS }}
TELEGRAFMERGED: {{ TELEGRAFMERGED }}
LOGSTASH_MERGED: {{ LOGSTASH_MERGED }}
# this file will be read by telegraf to send node details (management interface, monitor interface, etc)
# into influx
node_config:
file.managed:
- name: /opt/so/conf/telegraf/node_config.json
- source: salt://telegraf/node_config.json.jinja
- template: jinja
{% else %}
{{sls}}_state_not_allowed:
test.fail_without_changes:
- name: {{sls}}_state_not_allowed
{% endif %}