Skip to content

ci: add opt-in paired race cache seed measurement - #456

Merged
XuPeng-SH merged 6 commits into
matrixorigin:mainfrom
XuPeng-SH:codex/race-seed-canary
Sep 20, 2026
Merged

XuPeng-SH merged 6 commits into
matrixorigin:mainfrom
XuPeng-SH:codex/race-seed-canary

Conversation

@XuPeng-SH

@XuPeng-SH XuPeng-SH commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Problem

UT cache seeding is default-off after #455. The Shanghai amd64-mo-shanghai-8c16g runner has no Docker daemon, and there is no controlled measurement of whether cache acquisition/import pays for itself over the complete race suite.

Change

Add a manual cold/warm A/B canary on the actual Shanghai runner class with immutable MatrixOne main SHA, CI checkout and builder digest. A disables seeding; B uses an opt-in daemonless registry transport that imports only the trusted producer's final cache COPY layers. Normal callers keep the existing Docker transport and seed remains default-off.

Each pair uses matched initial cache state, alternating AB/BA order, and records seed acquisition/import/cleanup plus clean/config/full race wall time, cgroup CPU/memory, disk and exact test/package outcomes. The comparison rejects identity drift, missing tests, partial imports and incomplete cleanup. Coverage is excluded.

Validation

  • Linux Python 3.12: 43 importer/cleanup/registry tests and 11 canary tests pass, including strict layer admission, compressed/uncompressed digest verification, malformed DEFLATE, cancellation, populated-cache preservation, forced EXDEV warm restore and exact outcome comparison.
  • actionlint 1.7.12 passes the reusable workflows, validation workflow and MatrixOne caller; whitespace checks pass.
  • Live task-owned Shanghai 8 CPU / 16 GiB probe, using producer 3ac87c30625fc082391e5384a4c94e50a2916cf6 and pinned image sha256:8137d222d3a3b29d119173639cea98931168ab7d886bf882894391f7804d5d88: state=seeded, producer ok, modules seeded, cleanup complete; 64,500 files / 15,819,961,352 bytes imported in 1,029.392s. A forced go test -race -count=1 smoke executed its body and passed.
  • No OOM/OOM kill, but memory reached about 16.4GB with 64 memory.max events. The diagnostic pod was deleted after collection.
  • Independent GPT-6 medium design and implementation review passed; all findings were addressed.

Limits

The live result proves capability, not speedup. Seeding consumed 1,029s of the 1,080s work budget and has little memory headroom. Full cold/warm MatrixOne race A/B has not run and this PR does not enable rollout. A thin MatrixOne manual entrypoint, pinned to the reviewed CI implementation commit, is required after this PR is merged so it can inherit the existing UT credentials.

See docs/race-seed-canary.md and docs/race-seed-validation.md for the measurement contract and evidence.

@XuPeng-SH XuPeng-SH left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed exact head 23eb3cc728e27db9d3e680bf78f69a44911a0bf5; self-authored PR, so COMMENT only.

[P2] Cold B admission must reject skipped module import.

scripts/race_seed_canary.py:321-326 validates overall seed status, producer health, cleanup, and imported build-file count, but ignores module_state. The importer can return state="seeded" with module_state="insufficient-space" after build-cache import succeeds but module import is skipped (actions/seed-go-caches/seed.py:489-515). B can therefore omit module acquisition/import while still being accepted as a valid comparison, making measured treatment and acquisition cost depend on available disk.

Require module_state="seeded" for cold B, allow only the expected preserved-populated state for warm B, and add regression coverage.

The 43 importer tests, 11 canary tests, and actionlint passed. Full Shanghai cold/warm A/B runtime evidence remains outstanding.

@XuPeng-SH XuPeng-SH left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-reviewed exact head 20a53a824c4526942ca5e570fb4f0c865f8c6492; self-authored PR, so COMMENT only.

PASS — the previous P2 is fixed at both sample admission and comparison: cold B now requires module_state=seeded, warm B requires preserved-populated, and regression coverage rejects missing, insufficient-space, and mismatched states. No additional actionable P0–P2 findings across the complete PR.

CI passed on an identical tree, including actionlint and 55 relevant tests. Evidence gaps remain for the full cold/warm race A/B and separately delivered MatrixOne caller integration; retain default-off rollout and do not infer a speedup or rollout conclusion yet.

@XuPeng-SH
XuPeng-SH merged commit c76c48c into matrixorigin:main Sep 20, 2026
1 check passed
@XuPeng-SH
XuPeng-SH deleted the codex/race-seed-canary branch September 20, 2026 03:27
XuPeng-SH added a commit to matrixorigin/matrixone that referenced this pull request Sep 20, 2026
## What

Add a manual-only MatrixOne workflow entrypoint for the paired race UT
cache-seeding experiment implemented by matrixorigin/CI#456.

The caller pins the reviewed CI implementation commit and forwards the
existing UT credentials with `secrets: inherit`. It is guarded to
`matrixorigin/matrixone` on `main`; ordinary PR CI and production seed
defaults are unchanged.

## Why a MatrixOne PR is needed

The reusable workflow lives in the CI repository, but the existing UT
credentials are available to MatrixOne workflows. Keeping the thin
caller here avoids copying credentials and lets a maintainer manually
select one immutable MatrixOne SHA, image digest, and one or three
paired repetitions.

## Validation

- actionlint v1.7.12 passes.
- CI importer/measurement suite: 43 importer tests and 11 canary tests
pass.
- A live task-owned Shanghai 8C16G probe imported 64,500 files / 15.82GB
with `state=seeded` and complete cleanup, then passed a forced `go test
-race -count=1` smoke.
- The import took 1,029.392s and approached the memory limit, so this
entrypoint is for measurement only; it does not claim speedup or enable
rollout.
- GPT-6 medium implementation review passed.

Depends on matrixorigin/CI#456. Merge CI#456 first, then this PR, then
dispatch the cold/warm A/B run.

Related #27076

---------

Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant