Skip to content

feat: parallel rollout WAL files with master batch commits - #303

Merged
beinan merged 5 commits into
mainfrom
codex/rollout-staged-append
Oct 4, 2026
Merged

beinan merged 5 commits into
mainfrom
codex/rollout-staged-append

Conversation

@beinan

@beinan beinan commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Rollout WAL merges currently serialize each worker's complete delete/index/append cycle against the shared base table. Add an opt-in append-only path where workers prepare immutable files concurrently and the master publishes bounded groups in one transaction, with atomic MemWAL generation watermarks for recovery.

  • Enable selected owned rollout tables with ROLLOUT_APPEND_TARGETS. Generic/context merges keep their existing semantics. Default concurrency is 4 and each staged prefix retains a 64 MiB byte cap and the worker's shared slots/memory budget.
  • Dedicated catch-up Pods stage with their own CPU and shared memory budget; oversized speculative reads retry once alone after other readers release memory. The ordinary master handles metadata only.
  • The master commits completed groups without waiting indefinitely for slower workers, retries staging on another worker, and consumes actual read/encoding progress through NDJSON. Retired/unavailable workers' shard directories are discovered using metadata only.
  • Preserve the existing maintenance claim, manifest ownership checks, cancellation handling and version fence. Reject an unexpected commit version, mismatched worker dataset URI, changed schema/prefix, or partially stale result. Reopen atomic watermarks after uncertain commits instead of appending again.
  • Historical cutover prefixes use bounded ID-only lookups to exclude legacy append-without-drain rows. Updated legacy fallback preserves atomic watermarks after cutover, including when the feature is disabled and re-enabled.

This path relies on rollout's immutable, non-reused ID contract; it is not an upsert or cross-shard ingest-deduplication API. Workers must be deployed with compatible owned-merge configuration before enabling targets. Index coverage remains handled by normal maintenance. See docs/rollout-parallel-append.md for configuration, recovery and retention details.

Validation:

  • Local core regression: 274 passed, 13 ignored. Focused append tests exercise concurrent staging/single publication, live ingestion, failed drains/restart, migration and fallback, revocation, partial stale prefixes, byte limits and schema/version conflicts.
  • Real HTTP tests verify worker publication isolation and force two workers to reach a barrier concurrently. The etcd-backed dedicated executor test passes with no remote workers, two 2 MiB generations and a shared 3 MiB budget.
  • Clippy passes for core, master and server with all targets and warnings denied.
  • A local ARM64/debug/Lance 9 benchmark verifies all 512 final IDs/payloads: serial 1.141 s versus parallel 0.724 s (1.58x); base version increments fall from 11 to 2. This is a small, approximately 4 MiB local-filesystem workload, not a multi-Pod or Azure throughput result.
  • All 15 GitHub checks pass on af0dee7083574d45aaba895812c85d92e46aacb3: Rust tests, Python tests, coverage, clippy, three-platform test wheels and release packages, and distribution validation. This includes all 63 master etcd HA tests.

No production deployment or activation is included.

@beinan
beinan marked this pull request as ready for review October 4, 2026 04:10
@beinan
beinan merged commit 282c377 into main Oct 4, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant