Repository navigation
fix(loops): repair a run claimed but never closed after a mid-advance fault (#3316) - #3334
Merged
Merged
Conversation
… fault (#3316) advance_on_terminal claims run N (CAS runs_completed N-1 -> N) and then closes run N's row in a second write. A fault between the two (database is locked, a process kill) left the loop claimed with the run still `running`, and every redelivered terminal and reconcile_after_restart then lost the claim CAS, so the loop showed "running" forever. A delivery that loses the claim now recognises that state (loop non-terminal, runs_completed == N, run N still running) and closes the run itself. finalize_loop_run becomes a second CAS on status='running' and returns bool; only the caller that wins the close runs the tail, so a repair racing a live advance still dispatches run N+1 exactly once. No schema change. Red before the fix: test_m56_fault_after_claim_is_recoverable_on_restart (xfail marker removed). Red with the close CAS removed: test_m56_repair_that_loses_the_close_does_not_dispatch. Fixes #3316 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
loop_service.advance_on_terminalclaims run N on the loop row, then closes run N's row in a second write. A fault between the two (database is locked, a process kill) left the loop claimed with the run stillrunning. After that, every redelivered terminal andreconcile_after_restartlost the claim CAS, so the loop showed "running" forever.Fix (detect and repair):
runs_completed == Nand run N is stillrunning, it closes run N itself and carries on.finalize_loop_runis now a second CAS onstatus = 'running'and returnsbool. Only the caller that wins the close runs the rest of the advance, so a repair racing a live advance still dispatches run N+1 exactly once.Why not one transaction: it would need a new combined DB method, and it would leave the issue's test with no gap to inject the fault into. The repair also covers faults earlier in the close (the execution read, the gate lookup).
Tests
test_m56_fault_after_claim_is_recoverable_on_restart: xfail marker removed.test_m56_repair_that_loses_the_close_does_not_dispatch: a repair beaten to the close does not dispatch.finalize_loop_runin four existing test files updated to the new return contract.test_m61, bug: reconcile_after_restart racing a live loop advance runs the same iteration twice #3317, still strict).test_m56_fault…fails;status == 'running'condition removed →test_m56_repair…fails.Docs:
feature-flows/run-agent-loop.mddescribes the repair path.Fixes #3316
🤖 Generated with Claude Code