Repository navigation
test(kubernetes): retain bounded dependency failure diagnostics - #254
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and result
The Kubernetes fixture previously reported downstream Temporal connection errors and a generic wait timeout without retaining the external dependency's exit/OOM facts before teardown. One main-branch recovery run exposed this gap. This change records fixed lifecycle fields for the fixture's source and restore services, including stopped containers, before disconnecting their network and removing their state.
The probe has a nine-second command budget and 250ms maximum inherited-pipe drain, keeping it below ten seconds so it cannot prevent cleanup. Raw inspect/error/healthcheck contents and unknown health strings stay private. An intentionally stopped source service remains distinguishable from a failed restore service.
Scope and unresolved condition
This improves failure diagnosis; it does not claim to fix the original restore failure. Integration37337738012 lost its restored Temporal endpoint after initially connecting. The same
c56fc17actual-artifact candidate passed on Linux AMD64, and an unchanged macOS ARM64 source journey passed in 232.20s with no observed OOM/nonzero exit. Those passes do not establish the original root cause.No runtime, HTTP/OpenAPI, schema, SDK, timeout, resource or release-support change. This is an independently testable contributor diagnostic slice; no hosted credentials or new dependency are needed.
Acceptance and validation
exit_code=137,oom_killed=false, followed by removal of its containers, volumes, network and kind cluster. The injection passed: the normal 90-second wait failed as expected, lifecycle facts were retained, and all owned resources were removed. This deliberately induced failure is not evidence of the original failure's cause.WaitDelaynow releases those pipes, and the test cleans its owned child. Final race tests passed; both independent reviewers approved the fix with no remaining findings.f91e28db2ea5dfc2c97e0c19689a1afac2da5dc3with no remaining findings; all 12 required CI checks passed at that head, including the source Kubernetes journey. After merge, the final source must also pass manual actual-artifact candidate acceptance before publication. Alpha publication uses only the later reviewed merged revision; original failure evidence remains preserved.