mirror of
https://github.com/Security-Onion-Solutions/securityonion.git
synced 2026-08-07 08:24:45 +02:00
pg_cron's launcher connects to cron.database_name at postmaster start and is registered BGW_NEVER_RESTART. On a host upgraded onto an existing /nsm/postgres volume, init-db.sh never runs, so so_telegraf does not exist when PostgreSQL starts -- the launcher dies and never retries. Salt then creates the database, the extension, and the schedule, all of which succeed, but no worker is left to fire the job. partman.run_maintenance_proc() therefore never runs: partitions stop being premade after create_parent's initial window and every metric lands in <parent>_default. Retention never fires either. That state is self-perpetuating. Once the default partition holds rows for a day with no child, PostgreSQL cannot create that child at all -- attaching it would violate the default partition's constraint -- so maintenance aborts on the first parent it reaches. Fixing the scheduler alone does not recover a stalled grid. Point cron.database_name at the always-present postgres database and register the job with cron.schedule_in_database targeting so_telegraf, so the launcher no longer depends on database creation order. group_role drops any registration left behind in so_telegraf, and both halves are guarded on the live GUC so applying postgres.telegraf_users before the postgresql.conf change has restarted the container skips instead of failing. Maintenance now runs so_admin.telegraf_maintenance(), which drains stranded rows out of any default partition before calling partman: expired rows are deleted, the rest are repartitioned. It runs from the state on every highstate as well as hourly from pg_cron, so a grid whose worker is dead still recovers on its own. The routines live in a postgres-owned schema so_telegraf has no rights on, since pg_cron executes them as postgres. Existing grids are recovered by a marker-guarded repair state that truncates the non-empty defaults once per host. The backlog is mostly past retention already and moving tens of GB just to delete most of it is not worth the WAL. Also raise premake from 3 to 7, reconciled onto existing parents in the retention subcommand, so an outage has a week of headroom before anything reaches a default partition, and add a check subcommand reporting partition age, default occupancy and last job status.
90 lines
3.4 KiB
YAML
90 lines
3.4 KiB
YAML
postgres:
|
|
enabled:
|
|
description: Whether the PostgreSQL database container is enabled on this grid. Backs the assistant store and the Telegraf metrics database.
|
|
forcedType: bool
|
|
readonly: True
|
|
helpLink: influxdb
|
|
telegraf:
|
|
retention_days:
|
|
description: Number of days of Telegraf metrics to keep in the so_telegraf database. Older partitions are dropped hourly by pg_partman.
|
|
forcedType: int
|
|
helpLink: postgres
|
|
config:
|
|
max_connections:
|
|
description: Maximum number of concurrent PostgreSQL connections.
|
|
forcedType: int
|
|
global: True
|
|
helpLink: postgres
|
|
shared_buffers:
|
|
description: Amount of memory PostgreSQL uses for shared buffers (e.g. 256MB, 1GB). Raising this improves read cache hit rate at the cost of system RAM.
|
|
global: True
|
|
helpLink: postgres
|
|
log_min_messages:
|
|
description: Minimum severity of server messages written to the PostgreSQL log.
|
|
options:
|
|
- debug1
|
|
- info
|
|
- notice
|
|
- warning
|
|
- error
|
|
- log
|
|
- fatal
|
|
global: True
|
|
helpLink: postgres
|
|
listen_addresses:
|
|
description: Interfaces PostgreSQL listens on. Must remain '*' so clients on the docker bridge network can connect.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
port:
|
|
description: TCP port PostgreSQL listens on inside the container. Firewall rules and container port mapping assume 5432.
|
|
forcedType: int
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
ssl:
|
|
description: Whether PostgreSQL accepts TLS connections. Must remain 'on' — pg_hba.conf requires hostssl for TCP.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
ssl_cert_file:
|
|
description: Path (inside the container) to the TLS server certificate. Salt-managed.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
ssl_key_file:
|
|
description: Path (inside the container) to the TLS server private key. Salt-managed.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
ssl_ca_file:
|
|
description: Path (inside the container) to the CA bundle PostgreSQL uses to verify client certificates. Salt-managed.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
hba_file:
|
|
description: Path (inside the container) to the pg_hba.conf authentication file. Salt-managed — edit salt/postgres/files/pg_hba.conf.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
log_destination:
|
|
description: Where PostgreSQL writes its server log. 'stderr' routes to the container log stream.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
logging_collector:
|
|
description: Whether to run a separate logging collector process. Disabled because the docker log stream already captures stderr.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
shared_preload_libraries:
|
|
description: Comma-separated list of extensions loaded at server start. Required for pg_cron which drives pg_partman maintenance — do not remove.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|
|
cron.database_name:
|
|
description: Database pg_cron keeps its job metadata in. Must already exist when PostgreSQL starts, because pg_cron's launcher connects to it at startup and never retries if it is missing. The maintenance job itself targets so_telegraf.
|
|
global: True
|
|
advanced: True
|
|
helpLink: postgres
|