feat(xtest): shared benchmark preparation for local and CI runs - #626
Draft
dmihalcik-virtru wants to merge 1 commit into
Draft
dmihalcik-virtru wants to merge 1 commit into
dmihalcik-virtru wants to merge 1 commit into
Conversation
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: trueThanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
dmihalcik-virtru
force-pushed
the
DSPX-4372-s2-local-prep
branch
from
September 28, 2026 19:36
481b4da to
568f078
Compare
X-Test Failure Report |
dmihalcik-virtru
force-pushed
the
DSPX-4372-s2-local-prep
branch
from
September 28, 2026 19:54
568f078 to
d5e9aa2
Compare
One encoding of the benchmark's setup rules, reachable from a pytest run, a local shell and (in stage 5) the workflow. perf/config.py is a pure module -- no pytest import -- that owns BenchConfig and the request rules that were previously split between runner.BenchConfig, a lazy per-SDK select_arms that answered a typo with a skip, and a block of shell in xtest.yml. conftest validates every --bench option in pytest_configure, before payloads are built, so a bad threshold costs a second rather than several minutes of setup. otdf-sdk-mgr gains 'install benchmark': ordered request resolution that keeps alias->commit multiplicity, pr:N expansion, installed-path collision detection before any build, and a versioned manifest written on every exit path including the failing ones. The platform pin and the provisioning otdfctl pin are recorded apart from the measured arms, so the build under measurement no longer provisions the fixtures it is measured against. otdf-local starts the benchmark's platform-only shape correctly: ec and hybrid wrapping are set on the platform config rather than only on the km instances, a missing root key is generated once where every KAS reads it, and a KAS with no key fails loudly at start instead of reporting 'cipher: message authentication failed' at decrypt time. 'otdf-local env' exports the manifest pointer and OTDFCTL_HEADS as the JSON array conftest parses. It also finds and starts what otdf-sdk-mgr actually installs. Discovery looked only for a 'platform/' checkout beside xtest/, while 'install platform' creates xtest/platform/src/<ref>/ worktrees; several installed refs are an error naming OTDF_LOCAL_PLATFORM_DIR rather than an arbitrary pick. Those worktrees also lack the gitignored opentdf.yaml the manual setup makes by hand, so it is seeded once from the committed opentdf-dev.yaml -- before generation overwrites that file, since otherwise each run's output becomes the next run's input. Failed 'docker compose up' now reports what compose said instead of only that it failed. Refs DSPX-4372. Stage 2 of the split described in spec/DSPX-4372-t1.md.
dmihalcik-virtru
force-pushed
the
DSPX-4372-s2-local-prep
branch
from
September 29, 2026 13:13
d5e9aa2 to
500af5d
Compare
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Stage 2 of the six-stage split described in
spec/DSPX-4372-t1.md(added in #625). Replaces the corresponding part of #621.What this delivers on its own
A benchmark can be prepared and run locally by the same code path CI will use in stage 5, and the setup rules that govern it are encoded once.
xtest/perf/config.py— a pytest-free module that ownsBenchConfig(moved out ofperf/runner.py) and the request rules: SDK-name validation, the arm ceiling, duplicate rejection, finite positive budgets and timeouts,max_rounds >= min_rounds, and the1500 * K / 2budget default. Previously these lived in three disagreeing places —runner.BenchConfig.__post_init__, a lazy per-SDKselect_armsthat turned a typo into a silent skip, and a block of shell inxtest.yml.conftest.pytest_configurenow checks every--bench*option before session setup. A bad--bench-thresholdused to surface from thebench_configfixture, after payloads were built and CLIs installed.otdf-sdk-mgr install benchmark— resolves an ordered list of refs (first is the reference), expandspr:N, rejects installed-path collisions before any build, installs releases and source refs by the right method, and writes a versioned manifest toxtest/sdk/benchmark.installed.json.otdf-localbenchmark parity — the platform-only service shape the bench job wants actually works:ec_tdf_enabled/hybrid_tdf_enabledare set on the platform config rather than only on the km instances, a missingservices.kas.root_keyis generated once where the whole fleet reads it,otdf-local envexports the manifest pointer plusOTDFCTL_HEADS, and platform discovery finds thextest/platform/src/<ref>/worktreesotdf-sdk-mgractually creates (see below).Defects fixed while porting
conftest.load_otdfctlfalls back throughsdk/go/dist/main/otdfctl.shto a bare system binary, so "otdfctl" meant "whatever the go arm happens to be" — the build under measurement provisioned the fixtures it was measured against. The manifest now records a separateotdfctlpin and that alone is exported.OTDFCTL_HEADSis a JSON array, not a comma-separated list.conftest.pyreads it withjson.loadsandxtest.ymlwithfromJson. A comma-joined value is silently discarded and the run drops back onto the fallback above. A test now performs the exactjson.loadsconftestdoes.installation_method == "source"filter, reinstating the fallback for exactly the runs that asked not to have it.installed_tagis thedist/directory name for both methods.kas.pyaccepted an empty root key. The platform now generates one, and a KAS started without one fails atstart()with a message namingroot_keyinstead of surfacing ascipher: message authentication failedat decrypt time, several minutes later.cmd_tipstripssdk--/otdfctl--,normalize_versiononly v-prefixes, andrefs.ref_slugonly flattens slashes — sootdfctl/v0.24.0andv0.24.0both land indist/v0.24.0, undetected.dist_slugis now the single slug function for dist paths andcheck_collisionbuckets on it.check_collisionnow flags a shared directory only when the artifacts sit at different commits — two aliases at one commit produce one build, so sharing is fine and for measured arms it is the neutral outcome. The--otdfctlpin is resolved alongside the arms and joins the check when it shares theirdist/tree, so a pin at a different commit can no longer overwrite a build the run is measuring. An unresolvable pin is also now reported before the arms are compiled rather than after.installed_pathit had not installed. It recorded the directory if one happened to exist — routinely true from an earlier run — asserting a provenance the preparation never checked. It is empty now, unconditionally.otdf-local updiscarded Docker's reason for failing.DockerService.startcaptured compose's output and returned a bareFalse, so a run died with "✗ Failed to start Docker services" and nothing else; the only way to learn why was to rerun compose by hand. It now populates thestart_errorthe base class already defines, andupprints it. Found by hitting it: the real message waserror while creating mount source path ... chown ...: permission denied.statusis one ofsuccess/neutral/partial/failedandstatus_messageexplains whichever it is —neutralis not an error, so the field is not named for one.Dependencies and compatibility
install stable,install tip, andinstall releaseare untouched;install benchmarkis new.otdf-local envgains variables and drops none.--bench-baseline/--bench-candidatekeep their current meaning — they now fail fast on invalid input rather than late.BenchConfigmoved fromperf.runnertoperf.configand is re-imported inrunnerunder the same name, sofrom perf.runner import BenchConfigstill resolves.Validation
Re-run after rebasing onto #625 at
e4c7c8b0, its tip including the review corrections. The rebase was clean — this branch's diff against stage 1 is unchanged at 21 files, and its additions toperf/README.mdare a pure append to the section stage 1 rewrote.otdf-sdk-mgrotdf-local@pytest.mark.integration, see belowxtesttest_ls_no_services,test_ls_jsonandtest_clean_commandrequire a platform checkout to exist somewhere discoverable: without one_find_platform_dirraises andotdf-local lsexits 1. They pass in a workspace where a platform is installed — which is the earlier green run, and is what the discovery fix below is for — and they fail in one where none is. Checking outmainin this same workspace fails the same three identically, so nothing here regressed them; they are simply not runnable without a platform present.Eager validation checked against real pytest invocations:
Live preparation, actually run
install benchmarkwas run for real against this branch, not stubbed:It resolved and installed both arms (
v0.38.0as a release at308002c2,mainfrom source ate8a2f7f5), built the platform service frommain, installed the provisioning CLI separately atv0.38.1(99b51145), and wrotextest/sdk/benchmark.installed.jsonwith"status": "success". The source arm'sdist/main/.versionholdsref=main/sha=e8a2f7f5..., which is what the commit-reuse check reads.That run also exposed two of the defects listed above — the discarded Docker error, and a unit test that passed only because nothing had ever been installed in the workspace.
Platform discovery, fixed here
otdf-sdk-mgrinstalls the platform source toxtest/platform/src/<ref>/, butotdf_local.config.settings._find_platform_dirlooked only for aplatform/directorysibling to an ancestor of
xtest— i.e. the hand-clonedtests/platform/. The two layoutshad drifted apart on
main, soinstall benchmark --platform mainput a platform somewhereotdf-local upcould not see and the sequence inperf/README.mdworked only withOTDF_LOCAL_PLATFORM_DIRset by hand. Discovery now knows both: the sibling checkout stillwins where it exists, several installed refs are an error naming
OTDF_LOCAL_PLATFORM_DIRrather than an arbitrary pick, and the bare
platform.gitis excluded by shape rather thanby name.
Two further gaps surfaced on the way to a live run and are fixed with it:
opentdf.yaml. It is gitignored in the platform repoand made by hand in the manual setup (
cp opentdf-dev.yaml opentdf.yaml); nothing ininstall platformcreates it, soupdied with a bareFileNotFoundErrorfromload_yaml. It is now seeded once from the committedopentdf-dev.yaml— and the seedinghas to happen before the first generation, because
opentdf-dev.yamlis also wheregeneration writes; without a separate pristine copy each run's output becomes the next
run's input and the golden keyring entries accumulate a copy per run.
keys/are not generated by any documented step, so Keycloak's bindmounts resolve to empty directories Docker then fails to chown. Run the platform's own
.github/scripts/init-temp-keys.shonce. This one is left as an environment step, notcode: it belongs to the platform checkout, not to
otdf-local.Live measured run
The full sequence from
perf/README.mdwas executed end to end against real services — nostubs, no
OTDF_LOCAL_PLATFORM_DIRoverride:7 cells executed, 0 skipped (
"skipped": {}in the JSON), 10 rounds and 2 warmup roundseach, every one stopping on
max_roundsrather than the budget:go-encrypt-1MiB-controlgo-encrypt-1KiBgo-encrypt-1MiBgo-encrypt-32MiBgo-decrypt-1KiBgo-decrypt-1MiBgo-decrypt-32MiB13 improvements, 0 regressions,
trustworthy: true,noise_floor.tripped: false. Theenvironment exported by
otdf-local envcarriedXT_TMP_DIR,BENCH_INSTALLATION_MANIFESTand
OTDFCTL_HEADS='["v0.38.1"]'— the JSON-array formconftestparses — so the runprovisioned with the pinned CLI rather than with a build it was measuring.
The run also exercises #625's new code path.
go-encrypt-1MiB/wallcame backratio 0.493, ci_high 0.578, p_value 1.0, p_adjusted 1.0, p_value_faster 0.00098, p_adjusted_faster 0.00107 → IMPROVED, so both tails are computed and BH-adjustedindependently on a real measurement rather than only in the unit tests.
It does not, though, discriminate the old rule from the new one, and an earlier
version of this paragraph claimed it did.
main's IMPROVED clause wasp_adjusted > 1 - alphaon the slower tail;p_adjustedis1.0here and1.0 > 0.95holds, somainwould have called this cell IMPROVED too. That iswhat the fix predicts. BH never lowers a p-value, so reading an adjusted
upper-tail p as evidence of improvement gets easier as a run grows: the defect
shows up as false IMPROVED verdicts on null data under multiplicity, not as
changed verdicts on a genuine 2x speedup.
test_adjusted_slower_tail_is_not_read_as_evidence_of_improvementin #625 is what covers the discriminating case.