Problem
nan-vm-operator holds leadership through a session advisory lock on a dedicated pgx.Conn (internal/db/advisorylock.go, AcquireLeaderLock). Nothing watches that connection after boot. If it dies (a CNPG switchover or failover, a pooler or TCP reset, a primary restart), Postgres releases the lock but the old pod keeps running its loops. A new pod, for example after a rollout or a crash-loop of the standby, can then acquire the lock and run at the same time.
Since helmcode/nan-vm-operator#43 (multi-host launch), each leader runs up to --dispatch-concurrency (16) workers plus the probe lane, the capacity publisher, the backup loop and the base-image loop. Two leaders means two writers on agents_runtime, the IP pool and the capacity tables. The per-row guards (pick + reserve critical section, in-flight set, finish-sequence stale-read guard) are in-process only and do not protect against a second process.
Proposal
- Watchdog goroutine: periodically
SELECT 1 on the lock connection, or check pg_locks for our AdvisoryLockKey held by our backend pid, every ~5s. On failure or loss, cancel the root context: workers drain, Run returns, and the process exits non-zero so Kubernetes restarts it and it re-competes for the lock.
- Optionally fence writes:
SetAssignedHost / ClaimIP could carry a leader epoch. That is a larger change.
- Test: kill the lock connection (
pg_terminate_backend) and assert the operator stops dispatching within the watchdog interval.
Context: review of nan-vm-operator#43 (concurrent dispatch). Before #43 the risk existed with one serial drain; with 16 workers its blast radius is larger.
Problem
nan-vm-operator holds leadership through a session advisory lock on a dedicated
pgx.Conn(internal/db/advisorylock.go,AcquireLeaderLock). Nothing watches that connection after boot. If it dies (a CNPG switchover or failover, a pooler or TCP reset, a primary restart), Postgres releases the lock but the old pod keeps running its loops. A new pod, for example after a rollout or a crash-loop of the standby, can then acquire the lock and run at the same time.Since helmcode/nan-vm-operator#43 (multi-host launch), each leader runs up to
--dispatch-concurrency(16) workers plus the probe lane, the capacity publisher, the backup loop and the base-image loop. Two leaders means two writers onagents_runtime, the IP pool and the capacity tables. The per-row guards (pick + reserve critical section, in-flight set, finish-sequence stale-read guard) are in-process only and do not protect against a second process.Proposal
SELECT 1on the lock connection, or checkpg_locksfor ourAdvisoryLockKeyheld by our backend pid, every ~5s. On failure or loss, cancel the root context: workers drain,Runreturns, and the process exits non-zero so Kubernetes restarts it and it re-competes for the lock.SetAssignedHost/ClaimIPcould carry a leader epoch. That is a larger change.pg_terminate_backend) and assert the operator stops dispatching within the watchdog interval.Context: review of nan-vm-operator#43 (concurrent dispatch). Before #43 the risk existed with one serial drain; with 16 workers its blast radius is larger.