Repository navigation
test(system): chain tests on a rolled-out fleet, with a per-chain go/no-go report (abilityai/trinity-enterprise#794) - #3331
Conversation
Live test matrix against a local dev instance (2026-10-07)I ran
Defects the matrix found, now fixed in
What J19 left on the kept agent, read back from the platform: the agent holds 🤖 Generated with Claude Code |
|
Regression details (head_sha: `a8bae72ee49c1b6db48c57ded4adc05d34b9862e`)Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (4)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (4)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Seed Backend unit-suite regression diffPer-XML totals
❌ New failures introduced by HEAD (4)Tests failing under HEAD that did not fail under BASE in any seed:
Legend: [F] = assertion failure, [E] = collection or fixture error. Reproduce locally: |
|
merge-train: ejected from this batch — rides the next train once fixed. The Tier 1 failures are this PR's own:
The body also has no closing keyword — add |
…no-go report (trinity-enterprise#794) The 1.0 system test's chains (#783): each checks that a change in one part of a rolled-out system shows up in another, on records the system wrote — never on an agent's own report. - tests/system_chains/: a report plugin and five chains (J15–J19 in the journey catalog). Collected only under TRINITY_CHAIN_TESTS=1, so they never ride along with run-full.sh or the per-PR journey lane. - The report (chain-report.json + .md) gives each chain one verdict: passed / failed at a named step / partial / not run. Only passed is evidence; a skip, a step that could not run, or a test that recorded no step is never reported passed (pinned by test_ent794_chain_report.py). - J19 (objective → person) runs on any stack with Docker: canon objective, number recorded with the agent's own key, role card and operator read show it against the target, then stale past 2x cadence. The brief step is recorded not run — the platform puts no objective number in a brief — so it reports partial. Passed every platform step on a live stack. - J18 (library change → every holder) needs a github.com skills repo the run can push to (CHAIN_SKILLS_REPO / CHAIN_SKILLS_GITHUB_TOKEN), else not run. Fleet re-inject is off by default; the chain turns it on for the run, says so, and restores it. - J15 / J16 / J17 are skeletons reporting not run, blocked by ent#793 (the reference company) and ent#804 (change propagation / fact routing). - scripts/system/run_chains.sh (exit 0 only when every chain passed) and a manual system-chains workflow that boots a ref and uploads the report. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a visual check (ent#794)
Found by an eight-case matrix against a live dev instance:
- A chain that died before its fixture ran (the suite-wide login failing)
was missing from the report, and an empty report exited 3 ("not yet")
instead of failed. Such a chain is now reported failed from its marker,
naming the setup error (pytest's crash line, not the traceback footer);
an empty report says nothing ran and exits 1.
- J18's repo check raised on a network error and read as a failure; an
unreachable repo is a precondition, so it is now not run (with retries).
- run_chains.sh: 0 every chain passed, 1 a chain failed or none reported,
3 only partial / not run — a real failure is distinguishable.
- CHAIN_KEEP_AGENTS=1 leaves the agents a chain created in place, listed
in the report as kept agents, so they can be opened in the UI.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- system_chains/ is an opt-in NON_TIER_DIRS entry in run-full.sh, so the api tier never collects it and the #2888 directory guard owns it. - The chain doc moves beside its harnesses (tests/system_chains/README.md), with a pointer row in STRATEGY.md; docs/testing/ stays the #2339 live set. - test_ent794_chain_report scopes its sys.modules registration to the load via MonkeyPatch.context, so nothing leaks into the session (#762). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
a8bae72 to
549f222
Compare
vybe
left a comment
There was a problem hiding this comment.
/validate-pr (Lane C, because of the new workflow): rebased onto dev, with the registry.json conflict resolved by keeping both sides. Fixed the merge-train ejection causes in 549f222: system_chains/ is now an opt-in NON_TIER_DIRS entry; the chain doc moved to tests/system_chains/README.md, with a pointer in STRATEGY.md, so docs/testing/ stays the #2339 live set; and the sys.modules registration is scoped through MonkeyPatch.context. All checks green, including regression diff and the sys.modules lint. The workflow is dispatch-only with contents: read, and inputs.ref is used only in with:. Security greps came back clean.
Summary
These are the 1.0 system test's chains (ent#783). Each chain checks that a change in one part of a rolled-out system shows up in another part. A chain passes only on records the system wrote: platform rows, files the platform wrote into a container, and the platform's own responses. It never passes on an agent's report that it is done.
devCHAIN_SKILLS_REPO,CHAIN_SKILLS_GITHUB_TOKEN); otherwise not runHow it works
tests/system_chains/tests/journeys/catalog.yaml.TRINITY_CHAIN_TESTS=1, so they never run insiderun-full.shor the per-PR journey lane.The report (
chain-report.jsonandchain-report.md) gives each chain one verdict:Only passed counts as evidence. A skip, a step that couldn't run, or a test that recorded no step is never reported as passed. The summary is a go only when every chain passed.
scripts/system/run_chains.shruns the chains against a target and exits 0 only when every chain passed..github/workflows/system-chains.ymlis manual-dispatch: it boots a ref, runs the chains and uploads the report. GitHub registers manual dispatch only from the default branch, so this can't be dispatched until the next release cut carries it tomain. The exception is recorded in the bug(ci): issue promotion has not run since #2769 — the pull_request_target trigger is inert until it reaches main #2814 parity guard; until then, userun_chains.sh.Docs:
docs/testing/SYSTEM_CHAINS.md, and a regenerateddocs/testing/JOURNEYS.md.Findings from building it
skills_library_auto_reinject_enabled). With it off, a library change reaches an agent that holds the skill only when that agent restarts. The chain turns it on for the run, says so in the report, and restores the old value afterwards.runninga moment before its own server accepts file writes (503). The chains wait for the server the way any client must.Test Plan
tests/unit/test_ent794_chain_report.pyhas 9 tests on the verdict rules. As a mutation check, I made partial report as passed; 2 of them go red.test_2338, with the count bumped to 19), Python-version parity, workflow trigger parity, and root test placement (100 passed).scripts/system/run_chains.sh -k j19against a dev stack.Refs abilityai/trinity-enterprise#794. This is partial: the issue stays open for J15–J17, which depend on ent#793 and ent#804, and for J18's repo.
🤖 Generated with Claude Code