Skip to content

test(system): chain tests on a rolled-out fleet, with a per-chain go/no-go report (abilityai/trinity-enterprise#794) - #3331

Merged
vybe merged 3 commits into
devfrom
feature/ent794-chain-tests
Oct 7, 2026
Merged

vybe merged 3 commits into
devfrom
feature/ent794-chain-tests

Conversation

@dolho

@dolho dolho commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Summary

These are the 1.0 system test's chains (ent#783). Each chain checks that a change in one part of a rolled-out system shows up in another part. A chain passes only on records the system wrote: platform rows, files the platform wrote into a container, and the platform's own responses. It never passes on an agent's report that it is done.

Chain On dev
J19 An objective reaches the person Runs on any stack with Docker. All six platform steps passed on a live stack. The brief step is recorded as not run, so the chain reports partial (see below).
J18 A library change reaches every holder Runs when given a test skills repo (CHAIN_SKILLS_REPO, CHAIN_SKILLS_GITHUB_TOKEN); otherwise not run
J15 Roll-out · J16 Canon fact · J17 Fact routing Skeletons that report not run, blocked by ent#793 (the reference company) and ent#804 (change propagation / fact routing)

How it works

  • tests/system_chains/

    • A report plugin plus the five chains, which are J15–J19 in tests/journeys/catalog.yaml.
    • They are collected only when TRINITY_CHAIN_TESTS=1, so they never run inside run-full.sh or the per-PR journey lane.
    • The journey tier's create, poll and key helpers are reused rather than duplicated.
  • The report (chain-report.json and chain-report.md) gives each chain one verdict:

    • passed;
    • failed, at a named step;
    • partial;
    • not run, with a reason.

    Only passed counts as evidence. A skip, a step that couldn't run, or a test that recorded no step is never reported as passed. The summary is a go only when every chain passed.

  • scripts/system/run_chains.sh runs the chains against a target and exits 0 only when every chain passed.

  • .github/workflows/system-chains.yml is manual-dispatch: it boots a ref, runs the chains and uploads the report. GitHub registers manual dispatch only from the default branch, so this can't be dispatched until the next release cut carries it to main. The exception is recorded in the bug(ci): issue promotion has not run since #2769 — the pull_request_target trigger is inert until it reaches main #2814 parity guard; until then, use run_chains.sh.

  • Docs: docs/testing/SYSTEM_CHAINS.md, and a regenerated docs/testing/JOURNEYS.md.

Findings from building it

  1. J19 brief step. The platform puts no objective number into a brief, so any number in a brief is written by the model. The chain doesn't accept that, so it reports this step as not run (verdict partial) until the platform carries the number. Follow-up to file.
  2. J18 and fleet re-inject. Re-injecting agents after a library sync is off by default (skills_library_auto_reinject_enabled). With it off, a library change reaches an agent that holds the skill only when that agent restarts. The chain turns it on for the run, says so in the report, and restores the old value afterwards.
  3. Soft-deleted agents keep their seat. J19's first real runs showed this: an agent that holds a seat (ent#811, abilityai/trinity-enterprise#818) still holds it after it's deleted, because the seat-holder row is only cleaned up at the hard purge. A new agent then can't take that seat (409 naming the deleted agent). I'm reporting it on fix(security): sandbox voice panel HTML in iframe (CRITICAL XSS) #818. The chain now uses a fresh seat id per run, so it doesn't depend on earlier runs cleaning up.
  4. Readiness race. An agent reports running a moment before its own server accepts file writes (503). The chains wait for the server the way any client must.

Test Plan

  • tests/unit/test_ent794_chain_report.py has 9 tests on the verdict rules. As a mutation check, I made partial report as passed; 2 of them go red.
  • Guards pass: catalog (test_2338, with the count bumped to 19), Python-version parity, workflow trigger parity, and root test placement (100 passed).
  • Live run: scripts/system/run_chains.sh -k j19 against a dev stack.
    • J19 is partial: 6 of 6 platform steps passed and the brief step was not run. Exit code 3, as designed.
    • The full run reports J15/J16/J17/J18 as not run, each with its reason.
  • J18 needs a test skills repo before it can run (see above).

Refs abilityai/trinity-enterprise#794. This is partial: the issue stays open for J15–J17, which depend on ent#793 and ent#804, and for J18's repo.

🤖 Generated with Claude Code

@dolho

dolho commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

Live test matrix against a local dev instance (2026-10-07)

I ran scripts/system/run_chains.sh against a running dev stack. The stack was v1.0.0-aws.2, with the #3271 assignments and seat code loaded.

# Case Expected Result
A Collection gate off (TRINITY_CHAIN_TESTS unset) nothing collected ✅ 0 collected
B Runner without a password exit 2 with a message ✅
C Full run J19 partial at the brief step; J15–J18 not run with reasons; exit 3; nothing left behind; the re-inject setting unchanged ✅
D J19 with an impossible stale deadline failed at "when recording stops, the number reads stale on the card"; exit 1; agent cleaned up ✅ after fix
E J18 with an unreachable repo and a bad token not run, with the reason ✅ after fix
F Wrong admin password all 5 chains reported failed with "setup failed before the chain started: … 401"; exit 1 ✅ after fix
G A second run straight after (fresh seat and objective ids per run) partial again; nothing left behind ✅
H CHAIN_KEEP_AGENTS=1 agent kept and listed in the report as kept agents ✅

Defects the matrix found, now fixed in 4fbfd5125

  • F: when the suite-wide login failed, every chain errored before its fixture ran. The chains were missing from the report, and the empty report exited 3 ("not yet"). A chain that dies that early is now reported failed, built from its marker, and the reason is pytest's crash line. A report with no chains exits 1.
  • E: a network error during J18's repo precondition check read as a failure. An unreachable repo now means the chain is not run; the check retries three times first.
  • D: a real failure and a partial or not-run chain both exited 3. The exit codes are now 0 (all passed), 1 (a chain failed, or no chain was reported) and 3 (only partial or not run).

What J19 left on the kept agent, read back from the platform: the agent holds chain-pipeline-owner-* (seat_source: holds). The objective "Reach 100 qualified pipeline this quarter" is owned, and chain_pipeline reads 42 / 100, stale: true.

🤖 Generated with Claude Code

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown

⚠️ Nightly unit-suite found regressions when this PR is merged into dev (3 of 3 seeds).

Regression details (head_sha: `a8bae72ee49c1b6db48c57ded4adc05d34b9862e`)

Seed 12345

Backend unit-suite regression diff

Per-XML totals

Side Path Total Pass Fail Error Skip
base junit-base-pr3331-12345.xml 24211 24093 0 0 118
head junit-head-pr3331-12345.xml 24223 24101 4 0 118

❌ New failures introduced by HEAD (4)

Tests failing under HEAD that did not fail under BASE in any seed:

  • [F] test_2080_harness_contract::test_every_test_directory_is_covered_by_a_tier
  • [F] test_2080_harness_contract::test_the_directory_guard_passes_on_the_real_tree
  • [F] test_2339_testing_docs_consolidated::test_docs_testing_top_level_is_exactly_the_live_set
  • [F] test_lint_sys_modules::test_committed_baseline_matches_current_repo_state

Legend: [F] = assertion failure, [E] = collection or fixture error.
Identity = (classname, name, kind); union taken across all input XMLs.


Seed 67890

Backend unit-suite regression diff

Per-XML totals

Side Path Total Pass Fail Error Skip
base junit-base-pr3331-67890.xml 24211 24093 0 0 118
head junit-head-pr3331-67890.xml 24223 24101 4 0 118

❌ New failures introduced by HEAD (4)

Tests failing under HEAD that did not fail under BASE in any seed:

  • [F] test_2080_harness_contract::test_every_test_directory_is_covered_by_a_tier
  • [F] test_2080_harness_contract::test_the_directory_guard_passes_on_the_real_tree
  • [F] test_2339_testing_docs_consolidated::test_docs_testing_top_level_is_exactly_the_live_set
  • [F] test_lint_sys_modules::test_committed_baseline_matches_current_repo_state

Legend: [F] = assertion failure, [E] = collection or fixture error.
Identity = (classname, name, kind); union taken across all input XMLs.


Seed 99999

Backend unit-suite regression diff

Per-XML totals

Side Path Total Pass Fail Error Skip
base junit-base-pr3331-99999.xml 24211 24093 0 0 118
head junit-head-pr3331-99999.xml 24223 24101 4 0 118

❌ New failures introduced by HEAD (4)

Tests failing under HEAD that did not fail under BASE in any seed:

  • [F] test_2080_harness_contract::test_every_test_directory_is_covered_by_a_tier
  • [F] test_2080_harness_contract::test_the_directory_guard_passes_on_the_real_tree
  • [F] test_2339_testing_docs_consolidated::test_docs_testing_top_level_is_exactly_the_live_set
  • [F] test_lint_sys_modules::test_committed_baseline_matches_current_repo_state

Legend: [F] = assertion failure, [E] = collection or fixture error.
Identity = (classname, name, kind); union taken across all input XMLs.

Reproduce locally: git merge dev && ( cd tests && python -m pytest unit/ -m "not slow" -p randomly --randomly-seed=12345 )

@vybe

vybe commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

merge-train: ejected from this batch — rides the next train once fixed. The Tier 1 failures are this PR's own:

  1. lint (sys.modules pollution check): tests/unit/test_ent794_chain_report.py:22 — bare sys.modules[...] assignment; use monkeypatch.setitem.
  2. regression diff — 4 new failures: test_2080_harness_contract::test_every_test_directory_is_covered_by_a_tier and ::test_the_directory_guard_passes_on_the_real_tree (the new tests/system_chains/ needs a tier), test_2339_testing_docs_consolidated::test_docs_testing_top_level_is_exactly_the_live_set (new docs/testing/SYSTEM_CHAINS.md), test_lint_sys_modules::test_committed_baseline_matches_current_repo_state.

The body also has no closing keyword — add Fixes abilityai/trinity-enterprise#<N> so the issue promotes on merge.

@vybe vybe added the status-needs-fix PR has an unaddressed review/validation finding; cleared by the author's next push (#2815) label Oct 7, 2026
dolho and others added 2 commits October 7, 2026 19:32
…no-go report (trinity-enterprise#794)

The 1.0 system test's chains (#783): each checks that a change in one part
of a rolled-out system shows up in another, on records the system wrote —
never on an agent's own report.

- tests/system_chains/: a report plugin and five chains (J15–J19 in the
  journey catalog). Collected only under TRINITY_CHAIN_TESTS=1, so they
  never ride along with run-full.sh or the per-PR journey lane.
- The report (chain-report.json + .md) gives each chain one verdict:
  passed / failed at a named step / partial / not run. Only passed is
  evidence; a skip, a step that could not run, or a test that recorded no
  step is never reported passed (pinned by test_ent794_chain_report.py).
- J19 (objective → person) runs on any stack with Docker: canon objective,
  number recorded with the agent's own key, role card and operator read show
  it against the target, then stale past 2x cadence. The brief step is
  recorded not run — the platform puts no objective number in a brief — so
  it reports partial. Passed every platform step on a live stack.
- J18 (library change → every holder) needs a github.com skills repo the
  run can push to (CHAIN_SKILLS_REPO / CHAIN_SKILLS_GITHUB_TOKEN), else not
  run. Fleet re-inject is off by default; the chain turns it on for the run,
  says so, and restores it.
- J15 / J16 / J17 are skeletons reporting not run, blocked by ent#793 (the
  reference company) and ent#804 (change propagation / fact routing).
- scripts/system/run_chains.sh (exit 0 only when every chain passed) and a
  manual system-chains workflow that boots a ref and uploads the report.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…a visual check (ent#794)

Found by an eight-case matrix against a live dev instance:

- A chain that died before its fixture ran (the suite-wide login failing)
  was missing from the report, and an empty report exited 3 ("not yet")
  instead of failed. Such a chain is now reported failed from its marker,
  naming the setup error (pytest's crash line, not the traceback footer);
  an empty report says nothing ran and exits 1.
- J18's repo check raised on a network error and read as a failure; an
  unreachable repo is a precondition, so it is now not run (with retries).
- run_chains.sh: 0 every chain passed, 1 a chain failed or none reported,
  3 only partial / not run — a real failure is distinguishable.
- CHAIN_KEEP_AGENTS=1 leaves the agents a chain created in place, listed
  in the report as kept agents, so they can be opened in the UI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- system_chains/ is an opt-in NON_TIER_DIRS entry in run-full.sh, so the
  api tier never collects it and the #2888 directory guard owns it.
- The chain doc moves beside its harnesses (tests/system_chains/README.md),
  with a pointer row in STRATEGY.md; docs/testing/ stays the #2339 live set.
- test_ent794_chain_report scopes its sys.modules registration to the load
  via MonkeyPatch.context, so nothing leaks into the session (#762).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe
vybe force-pushed the feature/ent794-chain-tests branch from a8bae72 to 549f222 Compare October 7, 2026 18:33

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

/validate-pr (Lane C, because of the new workflow): rebased onto dev, with the registry.json conflict resolved by keeping both sides. Fixed the merge-train ejection causes in 549f222: system_chains/ is now an opt-in NON_TIER_DIRS entry; the chain doc moved to tests/system_chains/README.md, with a pointer in STRATEGY.md, so docs/testing/ stays the #2339 live set; and the sys.modules registration is scoped through MonkeyPatch.context. All checks green, including regression diff and the sys.modules lint. The workflow is dispatch-only with contents: read, and inputs.ref is used only in with:. Security greps came back clean.

@vybe
vybe merged commit db4c1e9 into dev Oct 7, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

status-needs-fix PR has an unaddressed review/validation finding; cleared by the author's next push (#2815)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants