Compare commits

..
Author SHA1 Message Date
Mike Reeves 9ebf93cc26 Empty a default and create its partition in one transaction
Telegraf never stops writing. Clearing 50 defaults with separate
TRUNCATEs left the earliest ones refilled by the time maintenance tried
to attach today's child, which then failed on the default's constraint
and aborted the whole run. Doing both under one transaction makes the
concurrent inserts wait and land in the new partition.
2026-08-10 15:55:07 -04:00
Mike Reeves d69234146e Let partman premake forward across the gap a stall leaves
Retention drops every child once they all age out, partman refuses to
drop the last one, and with infinite_time_partitions off it will not
premake forward from a child that far in the past. The set is left with
one stale partition and no current one, so metrics land right back in
the default.

Set infinite_time_partitions on telegraf parents, in both the repair
script and the retention subcommand, and stop blaming the launcher when
pg_cron is not loaded at all.
2026-08-10 15:44:45 -04:00
Mike Reeves 706d46b395 Trim the comments on the Telegraf partition tooling
The rationale for the repair belongs in the commit history, not in a
40-line header on every script.
2026-08-10 15:08:14 -04:00
Mike Reeves fe4f7ad2f7 Repair a stalled grid in one step, and only when it needs it
so-telegraf-partition-repair cleared the backlog but left the cause in place:
pg_cron's launcher is still dead, so the grid re-stalls as soon as it walks off
the premade window. Operators on the preview release need something they can run
once, before they soup, that leaves Telegraf collecting again.

Replace it with so-telegraf-repair, which fixes both halves. The running release
already creates the pg_cron extension and registers telegraf-partman-maintenance
in so_telegraf; only the launcher is missing, because so_telegraf did not exist
when the postmaster started. Restarting so-postgres is therefore enough to get
the existing job firing, so this touches no configuration and duplicates none of
the postgres state's SQL -- group_role still migrates the job to the postgres
database on the next soup. It also reconciles premake to 7 and prefers
so_admin.telegraf_maintenance() when that state has already landed.

Exit status separates healthy (0) from needs-repair (1) from does-not-apply (2),
which is what soup now gates on. postupgrade_changes runs after the highstate,
so the database is already converted by then and the backlog is the only thing
left to detect. Truncating is destructive and most grids were never affected --
fresh installs in particular, since they have no Telegraf history at all -- so
soup asks first and skips silently rather than clearing defaults on every host.
2026-08-10 10:05:16 -04:00
Mike Reeves 668ab447a2 Exclude so_telegraf from the nightly Postgres backup
pg_dumpall dumped every database, and so_telegraf dominated the result: it
is the only database that grows with grid size and metric volume, while
everything else in the cluster is small and mostly static.

Nothing is lost by skipping it. The data is transient metrics on a 14-day
retention window, roles are globals so the per-minion telegraf logins are
still dumped, and the database itself is rebuilt after a restore without
operator action -- init-db.sh recreates it (run on every highstate by
postgres_bootstrap_soc_db, not just on a fresh volume), telegraf_users.sls
re-provisions the roles and schema, and Telegraf recreates its tables on
first write.
2026-08-07 15:02:55 -04:00
Mike Reeves a2a4d9314d Repair stalled Telegraf partitions from soup instead of every highstate
Truncating default partitions is destructive and premaking them is pg_cron's
job, so neither belongs in a state that runs on every checkin. Replace the two
telegraf_users states with so-telegraf-partition-repair, a standalone tool that
reports partition health and clears the backlog, and call it once from soup.

The script depends only on pg_partman, so it also runs against a grid that has
not yet picked up the new postgres state. It no-ops when nothing is stranded,
refuses to discard rows non-interactively without --yes, and reports when the
pg_cron job has never fired, which is the underlying cause rather than a
symptom the truncate addresses.

Hourly self-healing stays with so_admin.telegraf_maintenance() via pg_cron, so
a grid that never soups still recovers, just gradually and without discarding
in-retention metrics.
2026-08-06 12:28:40 -04:00
Mike Reeves 5d36d00dec Fix Telegraf metrics falling into pg_partman default partitions
pg_cron's launcher connects to cron.database_name at postmaster start and is
registered BGW_NEVER_RESTART. On a host upgraded onto an existing /nsm/postgres
volume, init-db.sh never runs, so so_telegraf does not exist when PostgreSQL
starts -- the launcher dies and never retries. Salt then creates the database,
the extension, and the schedule, all of which succeed, but no worker is left to
fire the job. partman.run_maintenance_proc() therefore never runs: partitions
stop being premade after create_parent's initial window and every metric lands
in <parent>_default. Retention never fires either.

That state is self-perpetuating. Once the default partition holds rows for a day
with no child, PostgreSQL cannot create that child at all -- attaching it would
violate the default partition's constraint -- so maintenance aborts on the first
parent it reaches. Fixing the scheduler alone does not recover a stalled grid.

Point cron.database_name at the always-present postgres database and register the
job with cron.schedule_in_database targeting so_telegraf, so the launcher no
longer depends on database creation order. group_role drops any registration left
behind in so_telegraf, and both halves are guarded on the live GUC so applying
postgres.telegraf_users before the postgresql.conf change has restarted the
container skips instead of failing.

Maintenance now runs so_admin.telegraf_maintenance(), which drains stranded rows
out of any default partition before calling partman: expired rows are deleted,
the rest are repartitioned. It runs from the state on every highstate as well as
hourly from pg_cron, so a grid whose worker is dead still recovers on its own.
The routines live in a postgres-owned schema so_telegraf has no rights on, since
pg_cron executes them as postgres.

Existing grids are recovered by a marker-guarded repair state that truncates the
non-empty defaults once per host. The backlog is mostly past retention already
and moving tens of GB just to delete most of it is not worth the WAL.

Also raise premake from 3 to 7, reconciled onto existing parents in the retention
subcommand, so an outage has a week of headroom before anything reaches a
default partition, and add a check subcommand reporting partition age, default
occupancy and last job status.
2026-08-05 17:39:39 -04:00
9 changed files with 474 additions and 31 deletions
+7 -18
View File
@@ -138,8 +138,6 @@ function getinstallinfo() {
log "ERROR" "Failed to source install variables"
return 1
fi
log "INFO" "Fetched install info for $MINION_ID (node type: ${NODETYPE:-unset})"
}
function pcapspace() {
@@ -485,7 +483,6 @@ function add_sensoroni_with_analyze_to_minion() {
# Sensor settings for the minion pillar
function add_sensor_to_minion() {
log "INFO" "Writing sensor configuration for $MINION_ID (interface: ${INTERFACE:-unset})"
{
echo "sensor:"
echo " interface: '$INTERFACE'"
@@ -512,8 +509,6 @@ function add_sensor_to_minion() {
log "ERROR" "Failed to add sensor configuration to $PILLARFILE"
return 1
fi
log "INFO" "Wrote sensor configuration for $MINION_ID"
}
function add_elastalert_to_minion() {
@@ -586,14 +581,11 @@ function add_telegraf_to_minion() {
# generates a password on first add and is a no-op on re-add so the cred
# is stable across repeated so-minion runs. postgres.telegraf_users on the
# manager creates/updates the DB role from the same pillar.
log "INFO" "Provisioning postgres telegraf credential for $MINION_ID"
so-telegraf-cred add "$MINION_ID"
local result=$?
if [ $result -ne 0 ]; then
log "ERROR" "Failed to provision postgres telegraf cred for $MINION_ID (exit code: $result)"
return 1
fi
log "INFO" "Provisioned postgres telegraf credential for $MINION_ID"
so-telegraf-cred add "$MINION_ID"
if [ $? -ne 0 ]; then
log "ERROR" "Failed to provision postgres telegraf cred for $MINION_ID"
return 1
fi
}
function add_influxdb_to_minion() {
@@ -1051,7 +1043,7 @@ function updateMineAndApplyStates() {
}
function setupMinionFiles() {
log "INFO" "Setting up minion files for $MINION_ID (pillar: $PILLARFILE)"
log "INFO" "Setting up minion files for $MINION_ID"
# Check to see if nodetype is set
if [ -z $NODETYPE ]; then
@@ -1077,10 +1069,7 @@ function setupMinionFiles() {
fi
# Create node-specific configuration
create$NODETYPE || {
log "ERROR" "Failed to create $NODETYPE configuration for $MINION_ID"
return 1
}
create$NODETYPE || return 1
# Ensure proper ownership after all content is written
ensure_socore_ownership || return 1
+21
View File
@@ -1013,9 +1013,30 @@ up_to_3.3.0() {
INSTALLEDVERSION=3.3.0
}
telegraf_repair() {
# Only grids whose Telegraf partitions stalled need this; --check exits 1
# when there is something to repair, so everyone else is left alone.
local repair=/usr/sbin/so-telegraf-repair
[[ -x "$repair" ]] || return 0
docker ps --format '{{.Names}}' | grep -qx so-postgres || return 0
echo "Checking Telegraf metric partitions."
local status=0
"$repair" --check >> "$SOUP_LOG" 2>&1 || status=$?
case "$status" in
0) echo " Telegraf partitions are healthy; nothing to repair." ;;
1) echo " Repairing stalled Telegraf partitions."
"$repair" --yes \
|| echo " warning: so-telegraf-repair failed; run it manually" >&2 ;;
*) echo " Skipping; Telegraf is not storing metrics in Postgres on this host." ;;
esac
}
post_to_3.3.0() {
# Recollate again since some internal DBs were excluded during 3.2.0 soup
recollate_postgres
telegraf_repair
}
### 3.3.0 End ###
+1 -1
View File
@@ -16,4 +16,4 @@ postgres:
logging_collector: 'off'
log_min_messages: 'warning'
shared_preload_libraries: pg_cron
cron.database_name: so_telegraf
cron.database_name: postgres
+1 -1
View File
@@ -83,7 +83,7 @@ postgres:
advanced: True
helpLink: postgres
cron.database_name:
description: Database pg_cron schedules jobs in. Must be so_telegraf so partman maintenance runs in the right database context.
description: Database pg_cron keeps its job metadata in. Must already exist when PostgreSQL starts, because pg_cron's launcher connects to it at startup and never retries if it is missing. The maintenance job itself targets so_telegraf.
global: True
advanced: True
helpLink: postgres
+6 -1
View File
@@ -47,7 +47,12 @@ trap 'rm -f "$TMPFILE"' EXIT
# Dump all databases and roles, compress. Write to a temp file so the final
# filename only ever appears for a complete, verified backup.
if ! docker exec so-postgres pg_dumpall -U postgres | gzip > "$TMPFILE"; then
#
# so_telegraf is excluded: it is transient metrics on a short retention window,
# it dominates the dump size, and it is rebuilt automatically after a restore --
# init-db.sh recreates the database and Telegraf recreates its tables on first
# write. Roles are globals, so the per-minion telegraf logins are still dumped.
if ! docker exec so-postgres pg_dumpall -U postgres --exclude-database=so_telegraf | gzip > "$TMPFILE"; then
log "ERROR: pg_dumpall/gzip failed; backup aborted"
exit 1
fi
+190 -7
View File
@@ -5,11 +5,18 @@ set -e
# Usage: so-telegraf-postgres <subcommand>
# create_db Ensure the so_telegraf database exists.
# group_role Provision the so_telegraf group role, telegraf/partman schemas,
# pg_partman, pg_cron, and the hourly partman maintenance job.
# pg_partman, the so_admin maintenance routines, and the hourly
# pg_cron maintenance job.
# user Create or update a per-minion login role granted to so_telegraf.
# Env: ROLE_USER, ROLE_PASS.
# retention Reconcile partman retention on telegraf parents.
# retention Reconcile partman retention and premake on telegraf parents.
# Env: RETENTION_DAYS.
# maintenance Drain default partitions and run partman maintenance.
# check Report partition health. Non-zero if any parent is unhealthy.
#
# A default partition holding rows for a day blocks creating that day's child,
# so maintenance drains defaults before calling partman. Use so-telegraf-repair
# on a grid already stuck in that state.
cmd="${1:?subcommand required}"
@@ -36,7 +43,6 @@ CREATE SCHEMA IF NOT EXISTS telegraf AUTHORIZATION so_telegraf;
GRANT USAGE, CREATE ON SCHEMA telegraf TO so_telegraf;
CREATE SCHEMA IF NOT EXISTS partman;
CREATE EXTENSION IF NOT EXISTS pg_partman SCHEMA partman;
CREATE EXTENSION IF NOT EXISTS pg_cron;
-- Telegraf (running as so_telegraf) calls partman.create_parent()
-- on first write of each metric, which needs USAGE on the partman
-- schema, EXECUTE on its functions/procedures, and write access to
@@ -51,12 +57,141 @@ ALTER DEFAULT PRIVILEGES IN SCHEMA partman
GRANT SELECT, INSERT, UPDATE, DELETE ON TABLES TO so_telegraf;
ALTER DEFAULT PRIVILEGES IN SCHEMA partman
GRANT USAGE, SELECT, UPDATE ON SEQUENCES TO so_telegraf;
-- Hourly partman maintenance. cron.schedule is idempotent by jobname.
SELECT cron.schedule(
-- pg_cron runs these as postgres, so they must not sit in a schema any
-- Telegraf role can create objects in.
CREATE SCHEMA IF NOT EXISTS so_admin AUTHORIZATION postgres;
REVOKE ALL ON SCHEMA so_admin FROM PUBLIC;
CREATE OR REPLACE PROCEDURE so_admin.telegraf_maintenance()
LANGUAGE plpgsql
AS $proc$
DECLARE
r record;
v_default text;
v_rows bigint;
BEGIN
-- No per-parent EXCEPTION handler: partition_data_proc commits internally,
-- and COMMIT is illegal while a subtransaction is active. A failing parent
-- aborts the run and the next pass retries.
FOR r IN
SELECT parent_table, retention
FROM partman.part_config
WHERE parent_table LIKE 'telegraf.%'
ORDER BY parent_table
LOOP
v_default := format('%I.%I',
split_part(r.parent_table, '.', 1),
split_part(r.parent_table, '.', 2) || '_default');
CONTINUE WHEN to_regclass(v_default) IS NULL;
EXECUTE format('SELECT count(*) FROM %s', v_default) INTO v_rows;
CONTINUE WHEN v_rows = 0;
RAISE WARNING 'so_admin.telegraf_maintenance: % rows stranded in %, draining',
v_rows, v_default;
-- Cheaper to delete expired rows than to repartition and then drop them.
IF r.retention IS NOT NULL THEN
EXECUTE format('DELETE FROM %s WHERE "time" < now() - %L::interval',
v_default, r.retention);
COMMIT;
END IF;
-- Bounded so a large backlog drains across several runs.
CALL partman.partition_data_proc(
p_parent_table := r.parent_table,
p_loop_count := 200,
p_source_table := v_default
);
COMMIT;
END LOOP;
CALL partman.run_maintenance_proc();
END;
$proc$;
CREATE OR REPLACE FUNCTION so_admin.telegraf_partition_status()
RETURNS TABLE (
parent_table text,
oldest_child date,
newest_child date,
days_ahead int,
retention text,
default_rows bigint,
default_size text
)
LANGUAGE plpgsql
AS $func$
DECLARE
r record;
v_default regclass;
BEGIN
FOR r IN
SELECT pc.parent_table AS pt, pc.retention AS ret
FROM partman.part_config pc
WHERE pc.parent_table LIKE 'telegraf.%'
ORDER BY pc.parent_table
LOOP
parent_table := r.pt;
retention := r.ret;
SELECT min(d), max(d) INTO oldest_child, newest_child
FROM (
SELECT to_date(substring(c.relname FROM '_p(\d{8})$'), 'YYYYMMDD') AS d
FROM pg_inherits i
JOIN pg_class c ON c.oid = i.inhrelid
WHERE i.inhparent = r.pt::regclass
AND pg_get_expr(c.relpartbound, c.oid) <> 'DEFAULT'
) s;
days_ahead := newest_child - current_date;
v_default := to_regclass(format('%I.%I',
split_part(r.pt, '.', 1),
split_part(r.pt, '.', 2) || '_default'));
IF v_default IS NULL THEN
default_rows := NULL;
default_size := NULL;
ELSE
EXECUTE format('SELECT count(*) FROM %s', v_default::text) INTO default_rows;
default_size := pg_size_pretty(pg_total_relation_size(v_default));
END IF;
RETURN NEXT;
END LOOP;
END;
$func$;
-- Drop the registration older releases left in so_telegraf.
SELECT CASE
WHEN current_setting('cron.database_name', true) IS DISTINCT FROM current_database()
AND EXISTS (SELECT 1 FROM pg_catalog.pg_extension WHERE extname = 'pg_cron')
THEN 'true' ELSE 'false'
END AS drop_stale_cron \gset
\if :drop_stale_cron
DROP EXTENSION pg_cron CASCADE;
\endif
EOSQL
# Guarded on the live GUC so applying this before the postgresql.conf change
# has restarted the container skips rather than failing.
docker exec -i so-postgres psql -v ON_ERROR_STOP=1 -U postgres -d postgres <<'EOSQL'
SELECT CASE WHEN current_setting('cron.database_name', true) = current_database()
THEN 'true' ELSE 'false' END AS cron_here \gset
\if :cron_here
CREATE EXTENSION IF NOT EXISTS pg_cron;
-- cron.schedule_in_database is idempotent by jobname.
SELECT cron.schedule_in_database(
'telegraf-partman-maintenance',
'17 * * * *',
'CALL partman.run_maintenance_proc()'
'CALL so_admin.telegraf_maintenance()',
'so_telegraf'
);
\else
\echo 'pg_cron metadata database is not `postgres` yet; skipping job registration.'
\endif
EOSQL
;;
@@ -90,6 +225,8 @@ EOSQL
: "${RETENTION_DAYS:?RETENTION_DAYS is required}"
# \gset + \if guards against a missing pg_partman without using a DO
# block (psql :var substitution doesn't reach into dollar-quoted code).
# premake is reconciled here because telegraf.conf only applies it to
# parents created from now on.
docker exec -i so-postgres psql \
-v ON_ERROR_STOP=1 \
-v retention_days="$RETENTION_DAYS" \
@@ -97,14 +234,60 @@ EOSQL
SELECT CASE WHEN EXISTS (SELECT 1 FROM pg_catalog.pg_extension WHERE extname = 'pg_partman')
THEN 'true' ELSE 'false' END AS has_partman \gset
\if :has_partman
-- infinite_time_partitions so a gap in metrics does not stop partman from
-- premaking forward, which is what leaves everything in the default.
UPDATE partman.part_config
SET retention = :'retention_days' || ' days',
retention_keep_table = false
retention_keep_table = false,
premake = 7,
infinite_time_partitions = true
WHERE parent_table LIKE 'telegraf.%';
\endif
EOSQL
;;
maintenance)
docker exec -i so-postgres psql -v ON_ERROR_STOP=1 -U postgres -d so_telegraf <<'EOSQL'
SELECT CASE WHEN to_regproc('so_admin.telegraf_maintenance') IS NOT NULL
THEN 'true' ELSE 'false' END AS has_proc \gset
\if :has_proc
CALL so_admin.telegraf_maintenance();
\else
\echo 'so_admin.telegraf_maintenance() is missing; run so-telegraf-postgres group_role first.'
\endif
EOSQL
;;
check)
docker exec -i so-postgres psql -U postgres -d so_telegraf <<'EOSQL'
\pset border 2
SELECT * FROM so_admin.telegraf_partition_status();
EOSQL
docker exec -i so-postgres psql -U postgres -d postgres <<'EOSQL'
\pset border 2
SELECT CASE WHEN to_regclass('cron.job_run_details') IS NOT NULL
THEN 'true' ELSE 'false' END AS has_cron \gset
\if :has_cron
SELECT d.status, d.return_message, d.start_time
FROM cron.job_run_details d
JOIN cron.job j ON j.jobid = d.jobid
WHERE j.jobname = 'telegraf-partman-maintenance'
ORDER BY d.start_time DESC
LIMIT 5;
\else
\echo 'pg_cron is not installed in this database.'
\endif
EOSQL
unhealthy=$(docker exec so-postgres psql -U postgres -d so_telegraf -tAc \
"SELECT count(*) FROM so_admin.telegraf_partition_status()
WHERE coalesce(default_rows, 0) > 0 OR coalesce(days_ahead, -1) < 1")
if [ "${unhealthy:-1}" != "0" ]; then
echo "so-telegraf-postgres check: $unhealthy telegraf parent(s) unhealthy" >&2
exit 1
fi
echo "so-telegraf-postgres check: all telegraf parents healthy"
;;
*)
echo "Unknown subcommand: $cmd" >&2
exit 1
+245
View File
@@ -0,0 +1,245 @@
#!/bin/bash
# Copyright Security Onion Solutions LLC and/or licensed to Security Onion Solutions LLC under one
# or more contributor license agreements. Licensed under the Elastic License 2.0 as shown at
# https://securityonion.net/license; you may not use this file except in compliance with the
# Elastic License 2.0.
# Put Telegraf metrics storage back in service on a grid where pg_partman
# maintenance stalled: raises premake, discards the rows stranded in default
# partitions, restarts so-postgres if pg_cron's launcher is dead, and runs
# maintenance once. Healthy grids are reported and left alone.
#
# Usage: so-telegraf-repair [--check] [--yes] [--no-restart]
# --check Report health and change nothing.
# --yes Skip the confirmation prompt (for soup and other automation).
# --no-restart Never restart so-postgres, even if pg_cron's launcher is dead.
#
# Exit status:
# 0 healthy, or repair completed
# 1 repair is needed (--check only)
# 2 cannot run here: so-postgres, so_telegraf or pg_partman is missing
set -e
# Matches p_premake in telegraf.conf's create_parent template.
PREMAKE=7
JOB_NAME=telegraf-partman-maintenance
CHECK_ONLY=false
ASSUME_YES=false
NO_RESTART=false
usage() { sed -n '/^# Usage:/,/^# 2 /p' "$0" | sed 's/^# \?//'; }
while [[ $# -gt 0 ]]; do
case "$1" in
--check|--dry-run) CHECK_ONLY=true ;;
--yes|-y) ASSUME_YES=true ;;
--no-restart) NO_RESTART=true ;;
-h|--help) usage; exit 0 ;;
*) echo "Unknown option: $1" >&2; usage >&2; exit 2 ;;
esac
shift
done
skip() { echo "$*"; exit 2; }
psql_tg() { docker exec -i so-postgres psql -U postgres -d so_telegraf "$@"; }
psql_pg() { docker exec -i so-postgres psql -U postgres -d postgres "$@"; }
# query_to_xml so the per-table row counts need no helper function installed.
REPORT="
WITH parents AS (
SELECT pc.parent_table,
pc.premake,
split_part(pc.parent_table, '.', 1) AS sch,
split_part(pc.parent_table, '.', 2) AS tbl
FROM partman.part_config pc
WHERE pc.parent_table LIKE 'telegraf.%'
), children AS (
SELECT p.parent_table,
max(to_date(substring(c.relname FROM '_p(\d{8})\$'), 'YYYYMMDD')) AS newest_child
FROM parents p
JOIN pg_class pt ON pt.oid = p.parent_table::regclass
JOIN pg_inherits i ON i.inhparent = pt.oid
JOIN pg_class c ON c.oid = i.inhrelid
WHERE pg_get_expr(c.relpartbound, c.oid) <> 'DEFAULT'
GROUP BY p.parent_table
), defaults AS (
SELECT p.parent_table,
p.premake,
format('%I.%I', p.sch, p.tbl || '_default') AS default_table,
to_regclass(format('%I.%I', p.sch, p.tbl || '_default')) AS default_oid
FROM parents p
)
SELECT d.parent_table,
c.newest_child,
(c.newest_child - current_date) AS days_ahead,
d.premake,
CASE WHEN d.default_oid IS NULL THEN NULL ELSE
(xpath('/row/cnt/text()',
query_to_xml(format('SELECT count(*) AS cnt FROM %s', d.default_table),
false, true, '')))[1]::text::bigint
END AS default_rows,
CASE WHEN d.default_oid IS NULL THEN NULL
ELSE pg_size_pretty(pg_total_relation_size(d.default_oid)) END AS default_size
FROM defaults d
LEFT JOIN children c ON c.parent_table = d.parent_table
ORDER BY 1
"
docker ps --format '{{.Names}}' | grep -qx so-postgres \
|| skip "so-postgres is not running; nothing to repair."
docker exec so-postgres psql -U postgres -tAc \
"SELECT 1 FROM pg_database WHERE datname='so_telegraf'" | grep -q 1 \
|| skip "The so_telegraf database does not exist; Telegraf is not writing to Postgres."
psql_tg -tAc "SELECT 1 FROM pg_extension WHERE extname='pg_partman'" | grep -q 1 \
|| skip "pg_partman is not installed in so_telegraf; nothing to repair."
parents=$(psql_tg -tAc \
"SELECT count(*) FROM partman.part_config WHERE parent_table LIKE 'telegraf.%'")
stranded=$(psql_tg -tAc "SELECT coalesce(sum(default_rows), 0) FROM ( $REPORT ) t")
behind=$(psql_tg -tAc \
"SELECT count(*) FROM ( $REPORT ) t WHERE coalesce(days_ahead, -1) < 1")
# premake < 7, or infinite_time_partitions off: without the latter partman
# refuses to premake forward across the gap the stall left behind.
misconfigured=$(psql_tg -tAc \
"SELECT count(*) FROM partman.part_config
WHERE parent_table LIKE 'telegraf.%'
AND (premake < $PREMAKE OR NOT infinite_time_partitions)")
# Both columns are matched because which one carries the launcher's name varies
# with the pg_cron version.
launcher=$(psql_pg -tAc \
"SELECT count(*) FROM pg_stat_activity
WHERE backend_type ILIKE '%pg_cron%' OR application_name ILIKE '%pg_cron%'")
# so_telegraf before the postgres state lands, postgres after.
cron_db=$(docker exec so-postgres psql -U postgres -tAc \
"SELECT current_setting('cron.database_name', true)" | tr -d '[:space:]')
last_run=never
if [[ -n "$cron_db" ]]; then
last_run=$(docker exec so-postgres psql -U postgres -d "$cron_db" -tAc \
"SELECT coalesce(max(d.start_time)::text, 'never')
FROM cron.job_run_details d JOIN cron.job j USING (jobid)
WHERE j.jobname = '$JOB_NAME'" 2>/dev/null | tr -d '[:space:]' || echo unknown)
[[ -n "$last_run" ]] || last_run=never
fi
# A grid that has never written a metric has nothing to recover, and an empty
# cron_db means pg_cron is not loaded at all, which no restart fixes.
restart_needed=false
[[ "$launcher" -eq 0 && "$parents" -gt 0 && -n "$cron_db" ]] && restart_needed=true
repair_needed=false
[[ "$stranded" -gt 0 ]] && repair_needed=true
[[ "$behind" -gt 0 ]] && repair_needed=true
[[ "$misconfigured" -gt 0 ]] && repair_needed=true
$restart_needed && repair_needed=true
echo "Telegraf partition status:"
psql_tg -c "$REPORT"
echo "Rows stranded in default partitions: $stranded"
echo "pg_cron metadata database: ${cron_db:-unset}"
echo "pg_cron launcher running: $([[ "$launcher" -gt 0 ]] && echo yes || echo no)"
echo "Last $JOB_NAME run: $last_run"
echo
if ! $repair_needed; then
echo "Telegraf partitions are healthy. Nothing to do."
exit 0
fi
if $CHECK_ONLY; then
echo "Repair is needed:"
[[ "$stranded" -gt 0 ]] && echo " * $stranded row(s) stranded in default partitions"
[[ "$behind" -gt 0 ]] && echo " * $behind parent(s) with no partition for the current window"
[[ "$misconfigured" -gt 0 ]] && echo " * $misconfigured parent(s) with stale partman settings"
$restart_needed && echo " * pg_cron's launcher is dead; maintenance is not running at all"
echo
echo "Re-run without --check to repair."
exit 1
fi
if [[ "$stranded" -gt 0 ]] && ! $ASSUME_YES; then
echo "This will permanently discard the $stranded stranded row(s) above."
$restart_needed && ! $NO_RESTART && \
echo "so-postgres will also be restarted, which briefly interrupts SOC."
[[ -t 0 ]] || { echo "Not a terminal; re-run with --yes to confirm." >&2; exit 2; }
read -r -p "Continue? [y/N] " answer
[[ "$answer" =~ ^[Yy]$ ]] || { echo "Aborted."; exit 0; }
fi
if [[ "$misconfigured" -gt 0 ]]; then
echo "Reconciling partman settings on $misconfigured parent(s)."
# GREATEST so an operator who raised premake further keeps their value.
psql_tg -v ON_ERROR_STOP=1 -c \
"UPDATE partman.part_config
SET premake = GREATEST(premake, $PREMAKE),
infinite_time_partitions = true
WHERE parent_table LIKE 'telegraf.%'"
fi
if [[ "$stranded" -gt 0 ]]; then
echo "Clearing default partitions."
# One transaction: Telegraf is still writing, so a default emptied without
# its partition in place is refilled before maintenance can attach one.
psql_tg -v ON_ERROR_STOP=1 <<'EOSQL'
DO $$
DECLARE
r record;
BEGIN
FOR r IN
SELECT pc.parent_table,
format('%I.%I', n.nspname, c.relname) AS default_table
FROM partman.part_config pc
JOIN pg_class p ON p.oid = pc.parent_table::regclass
JOIN pg_inherits i ON i.inhparent = p.oid
JOIN pg_class c ON c.oid = i.inhrelid
JOIN pg_namespace n ON n.oid = c.relnamespace
WHERE pc.parent_table LIKE 'telegraf.%'
AND pg_get_expr(c.relpartbound, c.oid) = 'DEFAULT'
LOOP
EXECUTE format('TRUNCATE TABLE %s', r.default_table);
PERFORM partman.create_partition_time(
r.parent_table, ARRAY[date_trunc('day', now())]::timestamptz[]);
END LOOP;
END
$$;
EOSQL
fi
if $restart_needed; then
if $NO_RESTART; then
echo "WARNING: pg_cron's launcher is dead and --no-restart was given."
echo " Maintenance will not run on its own until so-postgres is restarted."
else
echo "Restarting so-postgres to revive pg_cron's launcher."
docker restart so-postgres >/dev/null
for _ in $(seq 1 60); do
docker exec so-postgres pg_isready -U postgres -q 2>/dev/null && break
sleep 2
done
docker exec so-postgres pg_isready -U postgres -q \
|| { echo "so-postgres did not come back; check 'docker logs so-postgres'." >&2; exit 1; }
fi
fi
echo "Running partition maintenance."
# so_admin.telegraf_maintenance() only exists once the postgres state has landed.
psql_tg -v ON_ERROR_STOP=1 <<'EOSQL'
SELECT CASE WHEN to_regproc('so_admin.telegraf_maintenance') IS NOT NULL
THEN 'true' ELSE 'false' END AS has_proc \gset
\if :has_proc
CALL so_admin.telegraf_maintenance();
\else
CALL partman.run_maintenance_proc();
\endif
EOSQL
echo
echo "Telegraf partition status after repair:"
psql_tg -c "$REPORT"
echo "The $JOB_NAME job runs hourly at :17. Confirm it fired with:"
echo " so-telegraf-repair --check"
+1 -1
View File
@@ -122,7 +122,7 @@
create_templates = [
'''CREATE TABLE IF NOT EXISTS {{ .table }} ({{ .columns }}) PARTITION BY RANGE ("time")''',
'''ALTER TABLE {{ .table }} ALTER COLUMN "time" SET NOT NULL''',
'''SELECT partman.create_parent(p_parent_table := {{ printf "%s.%s" .table.Schema .table.Name | quoteLiteral }}, p_control := 'time', p_type := 'range', p_interval := '1 day', p_premake := 3) WHERE NOT EXISTS (SELECT 1 FROM partman.part_config WHERE parent_table = {{ printf "%s.%s" .table.Schema .table.Name | quoteLiteral }})'''
'''SELECT partman.create_parent(p_parent_table := {{ printf "%s.%s" .table.Schema .table.Name | quoteLiteral }}, p_control := 'time', p_type := 'range', p_interval := '1 day', p_premake := 7) WHERE NOT EXISTS (SELECT 1 FROM partman.part_config WHERE parent_table = {{ printf "%s.%s" .table.Schema .table.Name | quoteLiteral }})'''
]
tag_table_create_templates = [
'''CREATE TABLE IF NOT EXISTS {{ .table }} ({{ .columns }}, PRIMARY KEY (tag_id))'''
+2 -2
View File
@@ -971,8 +971,8 @@ docker_seed_registry() {
if [ -f /nsm/docker-registry/docker/registry.tar ]; then
logCmd "tar xvf /nsm/docker-registry/docker/registry.tar -C /nsm/docker-registry/docker"
logCmd "rm /nsm/docker-registry/docker/registry.tar"
elif [[ -d /nsm/docker-registry/docker/registry && ( -f /etc/SOCLOUD || "$is_airgap" == true ) ]]; then
echo "Using existing docker registry content"
elif [ -d /nsm/docker-registry/docker/registry ] && [ -f /etc/SOCLOUD ]; then
echo "Using existing docker registry content for cloud install"
else
if [ "$install_type" == 'IMPORT' ]; then
container_list 'so-import'