Repository navigation
Bind guest reconnect authorization to the runner control connection - #114
Merged
Merged
Conversation
* feat(sessiond): authenticate guest reconnect before configuration refresh * fix(sessiond): refresh exec environments before reconnect readiness * docs: clarify configuration state after reconnect expiry * feat(driver): bound guest admission and bridge reconnect proofs (#116) * feat(driver): bound guest admission and bridge reconnect proofs * Require epoch-bound guest readiness acknowledgment (#117) * feat: require epoch-bound guest readiness acknowledgment * Integrate authenticated guest reconnect and guarded runner recovery (#118) * Wire authenticated guest handoff and guarded recovery admission * Fence delayed RPC delivery and reject ambiguous bootstrap preambles * Require exclusive state ownership for every microVM runner * Fence cold resume placement before launch and reconcile uncertain results * Preserve cold launch ownership across crashes and terminal races * Retain fresh and interrupted VM launches until process exit is proven * Use exact pending placement capacity in API explanations * Require exact instance and original process lifetime before VM teardown * Require host authority support to negotiate guest reconnect * fix: revalidate VM lifetime before shutdown escalation * feat: compose durable standalone guest recovery * fix(sessiond): launch the command supplied by microVM boot configuration * test(controld): await asynchronous placement dispatch before asserting * fix(driver): recognize original VM exit before reaping * fix(driver): require whole process exit before teardown * fix(microvm): wait for remaining recovered VMM threads * fix(microvm): resolve recovered cold resume configuration
jiashuoz
marked this pull request as ready for review
October 7, 2026 16:45
jiashuoz
added a commit
that referenced
this pull request
Oct 7, 2026
* Define bounded guest reconnect RPC messages and decoders * Bind guest reconnect authorization to the runner control connection (#114) * Authorize guest reconnect through one runner control connection * feat(sessiond): retain guest identity and authenticate reconnect (#115) * feat(sessiond): authenticate guest reconnect before configuration refresh * fix(sessiond): refresh exec environments before reconnect readiness * docs: clarify configuration state after reconnect expiry * feat(driver): bound guest admission and bridge reconnect proofs (#116) * feat(driver): bound guest admission and bridge reconnect proofs * Require epoch-bound guest readiness acknowledgment (#117) * feat: require epoch-bound guest readiness acknowledgment * Integrate authenticated guest reconnect and guarded runner recovery (#118) * Wire authenticated guest handoff and guarded recovery admission * Fence delayed RPC delivery and reject ambiguous bootstrap preambles * Require exclusive state ownership for every microVM runner * Fence cold resume placement before launch and reconcile uncertain results * Preserve cold launch ownership across crashes and terminal races * Retain fresh and interrupted VM launches until process exit is proven * Use exact pending placement capacity in API explanations * Require exact instance and original process lifetime before VM teardown * Require host authority support to negotiate guest reconnect * fix: revalidate VM lifetime before shutdown escalation * feat: compose durable standalone guest recovery * fix(sessiond): launch the command supplied by microVM boot configuration * test(controld): await asynchronous placement dispatch before asserting * fix(driver): recognize original VM exit before reaping * fix(driver): require whole process exit before teardown * fix(microvm): wait for remaining recovered VMM threads * fix(microvm): resolve recovered cold resume configuration
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Runner reconnect authorization now performs begin, guest proof exchange and acceptance through one captured agent connection. A control redial, local boot/placement replacement or deadline invalidates the attempt instead of moving its proof onto a replacement socket.
The optional driver host callback validates shared RPC responses, uses one five-second budget, limits work to one active attempt per session and 64 per runner, releases pending calls on failure, and returns only fixed errors with zero authority. The guest relay explicitly refuses begin/accept by method name alongside the existing high-bit origin guard. The design note documents the caller contract and remaining integration work.
This PR targets
feat/guest-reconnect-rpc(PR113). Merge that prerequisite, then rebase/retarget this branch to main. No capability is advertised or guest listener enabled. Guest key custody/enrollment, cold-resume placement ordering, current configuration delivery, bounded peer handling and relay takeover remain required before live qualification. Hosted integration additionally depends on Cloud PR146/147/148.Verification covers real agent WebSockets, signed round trip, wrong scope, strict decoding/refusals, guest-origin refusal, cancellation, late proof, admission limits, and replacement sockets/instances. A separately built callback consumer exercises success, wrong scope and revoked authority over the real runner control connection using a synthetic driver and proof provider; this is executable callback coverage, not a live VM continuity claim.
Local validation: full runner package race tests passed; focused reconnect race tests passed; build/vet passed; separately built execution probe passed. Local full verify encountered an existing latency cleanup timing failure (passed isolated) and a Docker registry timeout while obtaining the required test image. Full Linux CI is required before integration.
Final validation: Linux CI run36724284130 is fully green (full make verify, non-root jail ownership, CLI/client race checks, fleet syntax). Post-CI built callback execution passed. Independent and adversarial reviews both passed with no required findings. Both retained the documented pre-enablement requirements: cancellation-safe proof I/O, instance/epoch-bound configuration and relay installation, and cold-resume placement reconciliation.