Skip to content

feat: qualify B1 guest and runner recovery - #113

Merged
jiashuoz merged 2 commits into
mainfrom
feat/guest-reconnect-rpc
Oct 7, 2026
Merged

jiashuoz merged 2 commits into
mainfrom
feat/guest-reconnect-rpc

Conversation

@jiashuoz

@jiashuoz jiashuoz commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Surviving microVM guests can recover an authenticated control connection after relay or runner loss when guest recovery is enabled. Reconnect proves a process-memory key against current placement and accepted runner authority, applies current configuration, and fences the old relay before takeover.

Cold resume claims the resuming placement before launch and reconciles uncertain replies without creating a second placement. Recovery now waits through the interval between a VMM leader exiting and its last thread exiting, while retaining bounded waits and fresh signal authorization. If runner restart discarded the in-memory boot configuration, a negotiated versioned cold resume re-resolves it under the current pending claim and control connection before minting a new bootstrap token. Configuration and credentials stay out of durable metadata.

The standalone PostgreSQL dispatcher checks current owner/policy and separates guest bootstrap redemption from runner proof redemption. Guest startup honors boot-config argv. Durable launch ownership, process identity checks, terminal-generation fences and counted placements prevent premature reuse and duplicate reservations. Recovery remains opt-in via --microvm-guest-reconnect; volatile standalone stores do not negotiate it. Numeric-PID check/signal races remain documented; whole-process exit detection uses a process pidfd without claiming pidfd-based signaling.

Validation: full make verify; focused race and PostgreSQL/transport tests; full native non-root driver and runner race suites; reproduced-before/passing-after exit-wait and reconstructed-driver regressions. The synthetic KVM sequence passed cold suspend/resume, subsequent runner restart and guest deletion. Independent and adversarial reviews accepted the final code and paired Cloud pins.

The complete source-bound real-Claude B1 recovery run passed, including relay loss, gateway replacement, runner restart, the subsequent tool call, cold suspend/resume and another runner restart. All disposable resources were independently verified absent. The run qualified OSS df8ec14ed28d6f2a42ce6a114d815f9af52422db and Cloud 2b1b40eaed9f386339553724b8763cfb1574e7dc. No RAM snapshot or production rollout is included.

Includes the reviewed OSS B1 stack #113–#118, folded without changing the qualified runtime tree. Paired Cloud integration: https://github.com/tokencanopy/rainier-cloud/pull/149.

Qualification report: https://github.com/tokencanopy/rainier-cloud/blob/9673031/docs/experiments/microvm-b1-recovery-qualification.md

…114)

* Authorize guest reconnect through one runner control connection

* feat(sessiond): retain guest identity and authenticate reconnect (#115)

* feat(sessiond): authenticate guest reconnect before configuration refresh

* fix(sessiond): refresh exec environments before reconnect readiness

* docs: clarify configuration state after reconnect expiry

* feat(driver): bound guest admission and bridge reconnect proofs (#116)

* feat(driver): bound guest admission and bridge reconnect proofs

* Require epoch-bound guest readiness acknowledgment (#117)

* feat: require epoch-bound guest readiness acknowledgment

* Integrate authenticated guest reconnect and guarded runner recovery (#118)

* Wire authenticated guest handoff and guarded recovery admission

* Fence delayed RPC delivery and reject ambiguous bootstrap preambles

* Require exclusive state ownership for every microVM runner

* Fence cold resume placement before launch and reconcile uncertain results

* Preserve cold launch ownership across crashes and terminal races

* Retain fresh and interrupted VM launches until process exit is proven

* Use exact pending placement capacity in API explanations

* Require exact instance and original process lifetime before VM teardown

* Require host authority support to negotiate guest reconnect

* fix: revalidate VM lifetime before shutdown escalation

* feat: compose durable standalone guest recovery

* fix(sessiond): launch the command supplied by microVM boot configuration

* test(controld): await asynchronous placement dispatch before asserting

* fix(driver): recognize original VM exit before reaping

* fix(driver): require whole process exit before teardown

* fix(microvm): wait for remaining recovered VMM threads

* fix(microvm): resolve recovered cold resume configuration
@jiashuoz
jiashuoz marked this pull request as ready for review October 7, 2026 16:45
@jiashuoz
jiashuoz merged commit 4157d1c into main Oct 7, 2026
1 check passed
@jiashuoz
jiashuoz deleted the feat/guest-reconnect-rpc branch October 7, 2026 16:46
@jiashuoz jiashuoz changed the title Define strict guest reconnect RPC contract feat: qualify B1 guest and runner recovery Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant