Skip to content

docs: prompt to commit experiment artifacts; close out failure-mode A9 - #122

Merged
codexceed merged 2 commits into
mainfrom
docs/commit-experiment-artifacts
Aug 24, 2026
Merged

codexceed merged 2 commits into
mainfrom
docs/commit-experiment-artifacts

Conversation

@codexceed

Copy link
Copy Markdown
Owner

Description

Adds a rule requiring agent-driven experiments to end with an artifact-commit prompt, and closes out failure-mode A9 — marking it fixed, recording its validation, and documenting that its raw run files are gone. Docs only; no code paths touched.

Motivation

While auditing evidence for the queued provider-reliability post, the six eval result files it cites (evals/results/2026062*.jsonl, plus the eval_run_* journal trees) turned out to be absent. They were never committed — not gitignored, not deleted in any commit, simply never staged — and the local copies did not survive a machine migration.

The findings survived anyway, but only by luck of discipline: every load-bearing number had been written into ADR-0028, ADR-0029, the A9 catalog entry, and the R5 validation note at the time. Had that not happened, the post would be unwritable. increment-4-plan.md (the oracle-gaming build plan, also untracked) was lost the same way and did not leave that kind of trace.

The common cause is that durability depended on remembering to git add at the end of a long session — precisely when it is least likely to happen. Since experiments here are agent-driven, the fix belongs in the agent's instructions rather than in the eval runner.

Separately, A9 still read "pending merge + eval validation" two months after both had happened, so the catalog was misreporting the status of a closed defect.

Changes

Artifact-commit rule (CLAUDE.md, mirrored as principle 9 in evals/CLAUDE.md):

  • When the agent executes an experiment, the run is not finished when the numbers land. Enumerate every artifact produced and prompt the user to confirm committing them before reporting completion. Never auto-commit; never leave silently untracked.
  • The prompt must name paths, file counts and total size, so it is a decision rather than a reflex.
  • Artifacts are split into two classes because their risk differs. Results (*.jsonl, .summary.json, generated plots) are small and are the citable receipt, so they are offered by default. Journal trees under --no-cleanup are bulky and carry full trajectories, so they are asked for separately, with size, and scanned first — failure mode D1 is a secret reaching log and context, which makes reflexive commit the wrong default there.
  • A decline is recorded in the summary, so a lost artifact is a logged choice rather than an accident. An artifact too large to commit must be replaced with something citable (a summary, a digest, or the numbers transcribed into the research note), because evidence that cannot be cited breaks Ship the receipts later.

Failure-mode A9 (docs/research/failure-modes.md):

  • Marked ✅ fixed. ADR-0028 (R1–R4, feat(harness): transport-layer retry + request timeout for model calls (ADR-0028) #87) and ADR-0029 (R5, feat(model): R5 async streaming model calls — idle-timeout, mid-call cancellation, streaming fallback (ADR-0029) #89) both merged 2026-06-21, validated by 2026-06-21-eval-r5-postmerge-validation.md: 80 rows at concurrency 8 — the load case that produced the incident — with 0 transport_error, 0 iterations==0, and 525/525 decisions tagged transport="native_stream" with 0 fallbacks.
  • Records ADR-0029's correction to ADR-0028's framing: httpx has only per-operation timeouts and no total, so R1's "240s request timeout" was always a per-read bound — which is why a legitimate 358s generation slipped past it.
  • Keeps the status honest rather than simply green: what is fixed is the layering (a transport failure no longer routes through the model parse-retry). The recovery path has never fired on a real hang — proven correct offline and non-destructive to long legitimate work, but not proven to catch a live stall. Fault injection is still the next step.
  • New Artifacts bullet recording that the raw run files this entry cites are lost, and that the numbers survive only because they were written down at the time.

Testing

Docs-only change; no code paths touched, so no test suite applies.

  • Audited every load-bearing claim in the provider-reliability writing kit against committed sources to establish what survives; all but two (a 647s/22.4k-token streamed cell and a per-chunk micro-benchmark) are attested in ADR-0028, ADR-0029, A9, or the R5 validation note.
  • Confirmed the June artifacts were never tracked (git log --all -- 'evals/results/20260620*' is empty, no delete commit, git check-ignore reports not ignored) before describing them as lost rather than removed.
  • Verified the R5 note's 525/525 figure supersedes the kit's 491/491 at larger scale, and used the stronger one.
  • Checked evals/CLAUDE.md renumbering: the inserted principle 9 pushed the LLM-judge principle to 10, and no other doc references the principle count.
  • Grepped for downstream references to A9 being open; the remaining 🔧 markers (B2, C1, C3) are unrelated entries.

Reviewer notes

  • ADR-0028's Raw: eval_run_20260620T142752Z/ pointer is now a dead reference, and I deliberately left it alone — it is an accepted ADR and a record of what was true when written, and CLAUDE.md says supersede rather than edit those. The loss is recorded once, in the catalog, which is the living doc.
  • No ADR accompanies the artifact rule: it is workflow policy, not an architecture decision with rejected alternatives. The journals-in-git question does have a real trade-off and could justify one later if the default is ever revisited.

Sarthak Joshi added 2 commits August 24, 2026 12:20
…periments

The 2026-06 transport-resilience eval artifacts (evals/results/2026062*)
were never committed and did not survive a machine migration. The
findings survived only because the numbers had been written into
ADR-0028/0029 and a research note; the raw JSONL a write-up would cite
under 'ship the receipts' is gone.

Adds a rule to the root CLAUDE.md: when the agent executes an
experiment, the run is not done when the numbers land — enumerate the
artifacts and prompt the user to confirm committing them. Never
auto-commit, never leave silently untracked. The prompt has to name
paths, counts and sizes so it is a decision rather than a reflex.

Splits artifacts into two classes because their risk differs. Results
are small and are the citable receipt, so they are offered by default.
Journal trees are bulky and carry full trajectories, so they are asked
for separately and scanned first — failure mode D1 is a secret reaching
log and context, which makes reflexive commit the wrong default there.

A decline is recorded in the summary so a lost artifact is a choice, not
an accident, and an artifact too large to commit has to be replaced with
something citable.

Mirrored as principle 9 in evals/CLAUDE.md, where make eval runs live.
…rtifacts are lost

A9 still read 'fix implemented (ADR-0028, PR #87) — pending merge + eval
validation'. Both happened on 2026-06-21: #87 and #89 merged, and
2026-06-21-eval-r5-postmerge-validation.md is the validation — 80 rows
at concurrency 8, the load case that produced the incident, with 0
transport_error, 0 iterations==0, and 525/525 decisions streamed with 0
fallbacks.

Records ADR-0029's correction to ADR-0028's framing: httpx has only
per-operation timeouts and no total, so the '240s request timeout' was
always a per-read bound, which is why a legitimate 358s generation
slipped past it.

Keeps the status honest rather than simply green. What is fixed is the
layering — a transport failure no longer routes through the model
parse-retry. The recovery path has never fired on a real hang, so it is
proven correct offline and non-destructive to long legitimate work, but
not proven to catch a live stall. Fault injection is still the next step.

Also records that the raw run files this entry cites were never
committed and are lost to a machine migration; the numbers survive only
because they were written here and into the ADRs at the time. That is
the loss the new artifact-commit rule exists to prevent.
@codexceed

Copy link
Copy Markdown
Owner Author

Reviewer quiz — artifact-commit rule + A9 closeout

4 questions on this change (answers inside)

1. The rule splits artifacts into two classes with different defaults. Which class is not offered for commit by default, and which catalogued failure mode is the reason?

2. A9 is now marked ✅, but the entry deliberately withholds a stronger claim. What is fixed, and what remains unproven?

3. Fill in the blank: ADR-0029 corrected ADR-0028's framing by establishing that httpx has only ______________ timeouts and no ______________ — which is why a legitimate 358s generation slipped past a "240s" bound.

4. Why does the PR leave ADR-0028's now-dead Raw: artifact pointer unedited — (a) it is out of scope, (b) it is an accepted ADR and a record of what was true when written, (c) the file may come back?


Answers

1. Journal trees (eval_run_<stamp>/**/journal.jsonl) — asked for separately, with size, and scanned first. The reason is D1: a secret reaching log and context. Trajectories carry whatever the agent read, so reflexive commit is the wrong default.

2. Fixed: the layering — a transport failure no longer routes through the model parse-retry. Unproven: the recovery path. No hang has occurred since the fix, so zero retries have fired; it is proven correct offline and non-destructive to long legitimate work, but not proven to catch a live stall.

3. …per-operation (connect/read/write/pool) timeouts and no total. R1's "240s request timeout" was always a per-read bound — 240s of silence, not 240s of elapsed time.

4. (b) — it is an accepted ADR and a record of what was true when written; CLAUDE.md says supersede rather than edit those. The loss is recorded once in the catalog, which is the living doc.

@codexceed codexceed self-assigned this Aug 24, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b1f67c5850

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread CLAUDE.md

| Class | Examples | Default |
| --- | --- | --- |
| **Results** — small, structured, already the citable receipt | `evals/results/*.jsonl` + `.summary.json`, generated plots under `docs/research/assets/` | **Offer to commit.** These are what *Ship the receipts* points a reader at. |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make approved result artifacts addable

When the user accepts this offer, the documented result and journal paths are still ignored by .gitignore (*.jsonl, evals/results/*, and eval_run*/), so ordinary git status will hide them and git add <path> refuses to stage them. As git add -h documents, only --force allows adding otherwise ignored files. Update the ignore policy when an artifact is approved, or explicitly require force-adding the enumerated paths; otherwise the new durability workflow can report acceptance without preserving its receipts.

AGENTS.md reference: AGENTS.md:L34-L43

Useful? React with 👍 / 👎.

@codexceed
codexceed merged commit cd8b04a into main Aug 24, 2026
2 checks passed
@codexceed
codexceed deleted the docs/commit-experiment-artifacts branch August 24, 2026 11:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant