Repository navigation
UN-4223 [FIX] Kill a stuck PG consumer child instead of restarting the whole pod - #2312
johnyrahul wants to merge 2 commits into
Conversation
…e whole pod One hung task (an LLM call that never returns) froze its child's heartbeat. The fleet probe reported the OLDEST child's age, so that single child failed liveness, and the container restart then waited out the full termination grace for the hung call, leaving the pod consuming nothing for up to ~2h. - The supervisor now SIGKILLs a child whose heartbeat stays frozen past the stuck-child cap and re-forks it. Its lease stops renewing, so the reaper redelivers the message (bounded by max_attempts). The cap defaults to HEALTH_STALE_SECONDS (the existing per-task bound) and is overridable via WORKER_PG_QUEUE_CONSUMER_STUCK_CHILD_SECONDS; with no health port it is off unless set explicitly. - /health now goes 503 only once at least half the children are stale. The oldest child's age and the stale-child count remain in the body. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
…ll metrics - Reseed a stuck-killed slot's heartbeat when it is reaped. In a two-child fleet one slot is the quorum, so the frozen age kept /health at 503 through the replacement's bootstrap and could still restart the pod. Crash exits are not reseeded, so crash-loop detection is unchanged. - Only kill children that have finished loading. With a cap shorter than the bootstrap, a child that had not polled yet could be killed before starting. - Warn at startup when STUCK_CHILD_SECONDS exceeds HEALTH_STALE_SECONDS, where a stuck child can fail the probe before it is killed. - Export pg_consumer_oldest_child_age_seconds and pg_consumer_stuck_child_kills_total, so a stuck child and each kill are visible while the half-of-fleet heartbeat stays healthy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
| def mark_stuck_killed(self, slot: int) -> None: | ||
| self._validate(slot) | ||
| self._stuck_killed.add(slot) | ||
| self._stuck_kill_count += 1 |
There was a problem hiding this comment.
Kill counter can overcount If a child exits just before the supervisor sends SIGKILL,
ProcessLookupError is suppressed but this line still increments pg_consumer_stuck_child_kills_total. The metric then reports a stuck-child kill that did not happen, making spontaneous exits look like supervisor interventions.
Prompt To Fix With AI
This is a comment left during a code review.
Path: workers/pg_queue_consumer/supervisor.py
Line: 400
Comment:
**Kill counter can overcount** If a child exits just before the supervisor sends SIGKILL, `ProcessLookupError` is suppressed but this line still increments `pg_consumer_stuck_child_kills_total`. The metric then reports a stuck-child kill that did not happen, making spontaneous exits look like supervisor interventions.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.
Unstract test resultsPer-group results
Critical paths
|



What
/healthprobe goes 503 only once at least half of the children are stale (previously: the single oldest child).Why
pg-executorchild hung inside an LLM call that never returned.freshness()reported the oldest child's age, so that one child turned/healthinto a 503. Liveness then restarted the container, all 20 children stopped consuming on SIGTERM, and the container waited outterminationGracePeriodSeconds(7260s) for the hung call. That meant about 2h of zero consumption per pod, and on three occasions both replicas were down at once.How
_kill_stuck_children, runs every monitor tick):LEASE_SECONDS, the reaper redelivers the message, andmax_attemptsbounds how often that repeats._join_childrentakes over.stuck_child_seconds_from_env):WORKER_PG_QUEUE_CONSUMER_HEALTH_STALE_SECONDS, the threshold the consumer already documents as the upper bound on one task and that used to restart the whole container. The bound is unchanged; it now costs one slot instead of the pod.WORKER_PG_QUEUE_CONSUMER_STUCK_CHILD_SECONDS, which must be finite and > 0.HEALTH_PORT(no probe to enforce the bound), nothing is killed unless the override is set.HEALTH_STALE_SECONDS, where a stuck child could fail the probe before it is killed.freshness()returnsquorum_age(), the age at leastceil(n × 0.5)children have reached: 10 of 20, 2 of 3, 1 of 2, and 1 of 1 (unchanged).infoverride is unchanged./healthJSON age key is nowquorum_child_seconds_since_poll. The body still carriesoldest_child_seconds_since_polland addsstale_children.Can this PR break any existing features. If yes, please list possible items. If no, please explain why. (PS: Admins do not merge the PR without this section filled)
HEALTH_STALE_SECONDSis now SIGKILLed and redelivered. Before, it triggered a full container restart, which also killed it after the grace period. So no task that used to finish is cut short. Every chart and compose consumer setsHEALTH_STALE_SECONDS >= VT.max_attemptsstops it./healthJSON: the age key was renamed. No consumer of the old key was found inunstractorunstract-cloud; it is still present in the body.pg_consumer_heartbeat_age_seconds: now reports the half-of-fleet age instead of the oldest child's.pg_consumer_oldest_child_age_seconds(the most-stale child) andpg_consumer_stuck_child_kills_total, so one stuck child and each kill stay visible while the half-of-fleet heartbeat looks healthy.Database Migrations
Env Config
WORKER_PG_QUEUE_CONSUMER_STUCK_CHILD_SECONDS(default:WORKER_PG_QUEUE_CONSUMER_HEALTH_STALE_SECONDSwhen a health port is set, else off).unstract-cloud: prodworkerPgExecutorhasHEALTH_STALE_SECONDS: 7260, so without an override a hung child is killed after about 2h. SettingWORKER_PG_QUEUE_CONSUMER_STUCK_CHILD_SECONDS: "3660"aligns it with the executor's 1h task and RPC limit.Relevant Docs
Related Issues or PRs
Dependencies Versions
Notes on Testing
workers/tests/test_pg_consumer_supervisor.py:/healthbody keys.test_pg_consumer_supervisor.pyandtest_pg_metrics.pypass: 109 tests.Screenshots
Checklist
I have read and understood the Contribution Guidelines.
🤖 Generated with Claude Code