Skip to content

Phase F foundation: OpenSearch security, independent alerting, snapshots - #11

Merged
ChiefGyk3D merged 3 commits into
masterfrom
claude/phase-f-foundation
Sep 2, 2026
Merged

ChiefGyk3D merged 3 commits into
masterfrom
claude/phase-f-foundation

Conversation

@ChiefGyk3D

Copy link
Copy Markdown
Owner

The first three foundation items from docs/game-plan.md Phase F, built as three commits so they can be tested and reverted independently. This branch is for live-machine test cycles — start with docs/phase-f-testing.md, which has the .env additions, the migration order for your live deployment, verification commands, and a rollback recipe.

F1 — The SIEM can now defend itself

  • OpenSearch security plugin enabled: transport + HTTP TLS with real certs (scripts/07-generate-opensearch-certs.sh), users/roles applied idempotently by scripts/08b-init-opensearch-security.sh (admin / scoped logstash writer / readonly for Grafana+exporter / kibanaserver)
  • Every client re-wired for https+auth: Logstash outputs (CA-verified), Grafana datasources (env-interpolated credentials), opensearch-dashboards (security plugin kept on), every script's curl, password rotation in change-passwords.sh
  • InfluxDB auth enabled with a self-healing bootstrap; Grafana + unifi-poller carry credentials
  • Anonymous requests to 9200 now fail — that's the point, and the testing doc verifies it explicitly

F2 note

Doppler (issue #4) intentionally not in this branch — the new OPENSEARCH_* env vars are designed to drop into the Doppler migration unchanged.

F3 — Alerting that survives Grafana dying

  • Alertmanager (loopback-only) + elasticsearch-exporter services; Prometheus rule_files with 8 health rules: cluster red/yellow, ingest-rate-zero (a silent pipeline pages instead of looking quiet), target down, container restart churn, filesystem filling, rule-eval failures
  • No credentials in prometheus.yml — the exporter holds them; the old dead /_prometheus scrape jobs (plugin never shipped in the stock image) are removed

F4 — A disk failure is no longer total evidence loss

  • path.repo on the warm tier, daily Snapshot Management policy (retain 14) over all six SIEM index patterns, and a documented restore drill in docs/backup-restore.md

Validated statically (all green)

compose renders without .env; promtool check config + check rules SUCCESS; shellcheck error-severity clean; 68 pytest tests (new YAML-invariant suite covering rules schema, route/receiver consistency, datasource https+auth, node security config); cert chain verified end-to-end locally (SANs, RFC2253 DNs match node configs, PKCS#8 keys).

Needs live-machine verification (the test cycles)

  • securityadmin.sh / hash.sh invocation against a real OpenSearch 2.19 container
  • Logstash opensearch-output TLS options (ssl, ssl_certificate_verification, cacert) accepted at boot (--config.test_and_exit runs in CI but the full handshake needs the cluster)
  • OpenSearch Dashboards env-var wiring (OPENSEARCH_USERNAME/PASSWORD/SSL_VERIFICATIONMODE) on the 2.19 image
  • Snapshot Management API shape (_plugins/_sm/policies) on 2.19
  • Live presence of the exporter/cadvisor metric names the alert rules reference
  • Grafana OpenSearch plugin behavior with basicAuth against the readonly role (may need a one-time permission widening)
  • amtool check-config and the Logstash config test run in CI (no docker daemon in the build environment)

Known follow-ups after F1 lands: docs/maintenance.md still shows pre-security example commands; Wazuh-indexer snapshots.

🤖 Generated with Claude Code

https://claude.ai/code/session_015XBQ5bbaaoKreyoWivRwVQ


Generated by Claude Code

- TLS everywhere on the main cluster: transport + HTTP, certs from new
  scripts/07-generate-opensearch-certs.sh (root CA + node/admin certs,
  RFC2253 DNs matched in node configs; guards against docker-created
  bind-mount directories)
- users/roles: admin, logstash writer (scoped to the six SIEM index
  patterns + templates), readonly (Grafana + exporter), kibanaserver;
  rendered + applied by scripts/08b-init-opensearch-security.sh
  (idempotent securityadmin flow, verifies auth-200 and anon-401)
- every client updated: Logstash outputs (https, logstash user, CA
  verification), Grafana datasources (readonly basicAuth via env
  interpolation), opensearch-dashboards (security plugin kept on,
  kibanaserver service user), all scripts' curl calls, ISM/smoketest/
  ingest-test, change-passwords rotation for the four new passwords
- prometheus: dead /_prometheus scrape jobs removed (plugin not
  shipped in the stock image) — metrics path moves to the exporter in
  the next commit; direct-scrape recipe kept as a comment
- InfluxDB auth enabled with self-healing bootstrap (CREATE USER
  exception on 1.8); Grafana + unifi-poller wired with credentials
- healthchecks accept 401 (auth working) and never carry the admin
  password into container env; deploy rsync protects server-side
  generated certs from --delete
- note: this commit's compose also carries mounts/services wiring for
  the alerting and snapshot commits that follow on this branch

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XBQ5bbaaoKreyoWivRwVQ
- alertmanager service (pinned, loopback-only API) with a route tree
  by severity and an n8n webhook receiver; tracked template + envsubst
  render script, default rendering works out of the box
- elasticsearch-exporter service (works against OpenSearch) as the
  sole authenticated metrics path into the secured cluster — no
  credentials in prometheus.yml
- prometheus rule_files + 8 SIEM health rules: cluster red/yellow,
  ingest-rate-zero (a silent pipeline must page, not look quiet),
  target down, container restart churn, filesystem filling,
  rule-evaluation failures (promtool: SUCCESS)
- CI now also runs promtool check rules and amtool check-config
- new tests/python/test_yaml_configs.py: rules schema/severity
  invariants, route/receiver consistency, datasource https+auth
  invariants, node security config invariants (68 tests total)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XBQ5bbaaoKreyoWivRwVQ
- path.repo on both nodes backed by /data/warm/snapshots; new
  scripts/10-snapshot-setup.sh registers the fs repository and a daily
  Snapshot Management policy (retain 14) over all six SIEM index
  patterns, via the secured endpoints
- docs/backup-restore.md: manual snapshots, the scheduled policy, a
  restore drill (restore-to-renamed-index, verify, delete), offsite
  copy guidance; Wazuh indexer snapshots flagged as follow-up
- docs/phase-f-testing.md: ordered per-cycle checklist for live-machine
  testing — .env additions, cert generation, live-deployment migration
  order (Influx admin before auth flip, securityadmin before client
  restarts), fresh-install bring-up, authenticated/anonymous curl
  verification, exporter metrics, forced test alert, snapshot+restore
  drill, and a full rollback recipe
- README service table + game-plan Phase F statuses updated

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015XBQ5bbaaoKreyoWivRwVQ
@ChiefGyk3D
ChiefGyk3D marked this pull request as ready for review September 2, 2026 22:42
@ChiefGyk3D
ChiefGyk3D merged commit f97f9d8 into master Sep 2, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants