Skip to content

Hawkins judgment-record anchor, fresh user-testing battery, and the persona x modality grid - #23

Merged
mzargham merged 60 commits into
mainfrom
user-testing/browser-pass1
Oct 1, 2026
Merged

mzargham merged 60 commits into
mainfrom
user-testing/browser-pass1

Conversation

@mzargham

@mzargham mzargham commented Oct 1, 2026

Copy link
Copy Markdown
Member

Summary

Three bodies of work, in sequence on this branch:

  • Fresh single-modality user-testing battery (DL-074 through DL-083): every chapter (1-10) re-checkpointed end to end. All ten PASS. Surfaced a recurring seam-narration question across judgment-record notebooks (DL-075), escalated and tracked through the rest of the battery.
  • Hawkins judgment-record anchor (DL-084, closes DL-075): ReviewRecord gains subject_ref, the tutorial's own analog of Hawkins' Assurance Claim Point, cross-checked against a new ReviewRecordRef model-side metadata tag. Every judgment record across Ch2/3/4/6/8/10 and the exercises/ track now carries a real, checkable anchor. Grounded in a direct read of the Hawkins et al. 2011 primary source, not just prior curated extracts.
  • Large-scale user testing: the persona x modality grid (DL-085): a 10-cell grid (persona x where-a-learner-encounters-the-repo: GitHub repo, GitHub Pages, local clone reading chapters, local clone doing exercises), synthesized by the ACE. Zero blocking content defects. Fixed a real, three-times-confirmed stale-worktree-provisioning bug in the grid-dispatch methodology. Two exercise-scaffolding findings (first-run convention, carry-forward) were escalated to Z and resolved: Ch8's exercise now gates like Ch9/10 instead of mixing conventions; carry-forward between exercises is now documented (docs/setup.md) rather than silently assumed, with Ch6's own two required snapshots called out explicitly where they're created.

Also included: Apache-2.0 LICENSE, README/setup.md corrections found during manual browser testing, a CI/CD deploy-readiness plan and a large-scale user-testing design doc (both specs, not yet executed further), and a new docs/contributor.md section explaining the repo's own multi-agent harness to a first-time contributor.

Test plan

  • uv run pytest tests/ glossary/tests/ -v — 423 passed, 7 deselected (full suite, every commit)
  • uv run python scripts/check_construction.py --check
  • Every retrofitted judgment-record notebook re-executed end to end via nbconvert, not assumed from a builder's own report
  • exercises/ch08/exercise.ipynb's new gate verified to fire with a named diagnostic via nbconvert
  • Manual browser sweep of all 58 rendered pages, search, dark/light mode, mobile viewport
  • Every grid-cell finding either independently re-verified against the real branch (not assumed from a subagent's own report) or explicitly escalated to Z and resolved

🤖 Generated with Claude Code

DL-073. The Pass 3 battery's own execution checklist and report format
hardcoded cell indices (Cell 0, 5, 6), assuming every notebook has exactly
the 7-cell minimum skeleton -- already false for construction-zone
notebooks, judgment-record notebooks (17-27 cells), and Chapter 10's own
distributed-seam design. Now identifies each step by content type, matching
toaster-recipe's own already-established rule. Also gives the Novice
persona a concrete rule for "one sentence" (count terminal periods, not
semicolons), closing a recorded false-positive pattern (Ch1/7/8/9).

tutorial-style-guide/SKILL.md carries the identical cell-index assumption
independently -- flagged, not fixed here (different skill, own DL gate).
…iewRecordRef)

Grounded in a direct read of the primary source (Hawkins et al. 2011, the
authors' own self-archived copy) and a live probe against OpenSysML v0.9.0
confirming SysML v2's own metadata/about mechanism (SS7.27.2) works for this.
Gives ReviewRecord a real, bidirectional anchor to the model it previously
lacked, and retires the recurring judgment-record seam question (DL-075) by
giving it something concrete to narrate instead of a choice between readings.
Decomposes the approved spec into 14 tasks: schema + query helper (TDD),
the ReviewRecordRef construct introduced once in Ch2, per-chapter retrofit
of every original-authoring record (Ch2-4,6,8,10) with real subject_ref
values and matching model tags, Ch9/Ch10 reconstructions carried forward,
every negative control fixed so it keeps demonstrating exactly its own
one intended validation error, skill updates, glossary edge, DL-075
closure, and a full verification sweep. Grounded in a fresh empirical
probe (exact OpenSysML v0.9.0 JSON shapes for MetadataUsage/
annotatedElement/identifier, confirmed live, not assumed from the spec's
earlier probe) and a full survey of all 13 ReviewRecord-constructing
notebooks in the repo today.
…-record anchor

Implements Task 2 of docs/superpowers/plans/2026-10-01-hawkins-judgment-record-anchor-plan.md.
Given a loaded model, finds every ReviewRecordRef metadata tag and the one
subject it names, independent of any notebook's own Python objects. Shapes
confirmed against a live OpenSysML v0.9.0 probe (MetadataUsage.type and
.annotatedElement are both list-valued; the identifier attribute's literal
value is reached via ownedMember -> declaredName=="identifier" -> value ->
LiteralString.value).

Built, tested and independently reviewed (different model) in isolated
worktrees; integrated here via cherry-pick rather than branch merge, since
both worktrees were provisioned from a stale base commit missing this
branch's own recent history -- merging would have pulled that divergence
in. Reviewer's one cosmetic finding (an import placed mid-file instead of
at the top) fixed during integration, using the file's own existing
query.<fn>() call convention instead of a bare import.
…k-execution sweep

Task 1's independent reviewer (different model than the builder, per this
repo's own review discipline) found two real gaps in the original plan
that neither the spec nor Tasks 1-14 covered:

1. exercises/ch03,04,06,08,09,10 each mirror a real chapter judgment
   record (e.g. AS-C08-EX mirrors AS-C08, same subject) but were never in
   scope -- a learner who correctly completes one of these exercises by
   mirroring its chapter notebook would hit an unexplained subject_ref
   error nothing in the exercise instructions mentions. New Task 15 fixes
   this the same mechanical way as Tasks 4-9: reuse each chapter's own
   already-decided subject_ref value.

2. Nothing in the plan actually executes a notebook end to end -- pytest
   and this repo's current CI both skip notebook execution entirely
   (still a placeholder pending WP-8) -- so a retrofit step silently
   missed in Tasks 3-9 or 15 would go undetected by every check the plan
   originally specified. New Task 14 Step 6 runs every judgment notebook
   through nbconvert and treats any subject_ref-shaped AssertionError as
   a sign a specific retrofit step was missed.

Both gaps are the kind of real, evidence-backed finding this repo's own
review discipline (author and reviewer on different models) exists to
catch before it ships, not after.
… Point

Add a gl:refines edge from this tutorial's own subject_ref/ReviewRecordRef
convention to the existing Hawkins Assurance Claim Point definition. Includes
a gl:gloss (required by `glossary check`'s tutorial-definition rule since the
full gl:text exceeds 240 characters) and the resulting docs/glossary.md
regeneration via `glossary render`, per the tutorial-glossary skill's
"change definitions only through the graph, then check, then render" rule.
… faithful gloss

The reviewer (different model than the builder) correctly failed the original
commit on two points: it set gl:status gl:confirmed and gl:confirmedBy "Z" on
an edge Z has not actually seen (tutorial-glossary SKILL.md rule 1: agents
propose, only Z confirms; check is supposed to reject an agent name as
confirmedBy but didn't catch this because "Z" is also the name a real
confirmation uses), and it placed the edge as def-tutorial--subject-ref in
hawkins.ttl rather than def-tutorial--assurance-claim-point in tutorial.ttl,
where every other src-tutorial refinement edge actually lives (SKILL.md's own
"Edge:" instructions name the convention this violated).

Rebuilt the edge through glossary.graph.save_graph (the canonical, byte-
deterministic writer the skill requires, not a hand-written Turtle block) with
gl:status gl:proposed and no gl:confirmedBy, in tutorial.ttl under the correct
id. Also tightened the gloss itself per the reviewer's finding: the first
version said the tutorial's construct IS an Assurance Claim Point rather than
an analog of one, and dropped the "without importing GSN's argument-graph
apparatus" caveat that says what the narrowing gives up -- both restored.

docs/glossary.md regenerates back to showing only Hawkins' own confirmed sense
of the term, which is correct: a proposed edge never appears in a rendered
gloss (SKILL.md rule 1), so the public page stays accurate until Z actually
reviews and confirms this one. `glossary check` now exits 0 (was failing with
a stale-docs-page error after the revert); `lookup` shows the edge tagged
[proposed]; full suite still 423 passed, 7 deselected.

Left open, flagged for the user rather than fixed here (blocked by this
session's own auto-mode permission policy on git history rewrites): two
earlier commits on this branch (ca1b460, and the original 04820d7 this
commit fixes forward from) carry a Co-Authored-By trailer, which violates
this repo's own stated no-trailer convention. Both are unpushed -- nothing
has gone to origin/pass1/harness-alignment -- so a message-only cleanup
remains safe for the user to do by hand whenever they next touch this
branch's history.
…m claim

The reviewer (different model than the builder) passed the construct itself
but found two real content gaps worth fixing before Tasks 4-9 copy this
notebook's pattern as their own template:

1. cell-05's transition sentence said "the next cell turns to the judgment
   side", written before the builder's own resolution of a plan gap inserted
   a new construction-zone group (the ReviewRecordRef definition) ahead of
   it. Reworded to describe what actually comes next. Also dropped a now-
   redundant clause from cell-06 ("...and the model element it's an
   assumption about") that duplicated what subject_ref already establishes.

2. The seam cell claims the Python record and the model's own tag "agree",
   but nothing on the page actually showed the tag resolved in the loaded
   model -- the same printed Validation errors: [] would appear whether the
   tag existed or not (confirmed by the reviewer's own probe: cutting the
   tag out of the model and re-validating still returns []). Added one line
   to the validation cell printing the resolved tag itself
   (get_review_record_refs(model), filtered to this record's identifier)
   immediately before the validation-errors line, so the seam's claim is
   something the reader has actually just watched, not merely asserted.

Re-executed the notebook end to end (clean) and re-ran
scripts/check_construction.py --check and the full test suite (423 passed,
7 deselected, both unchanged).
…icting its own prose and the actual subject (ApplyHeat)

Caught by Task 5's own builder, who followed the correct prose instruction and verified empirically (a fresh OpenSysML v0.9.0 probe: metadata about an action def, not a part usage, loads cleanly and populates annotatedElement correctly) rather than the plan's own wrong code block.
…ng transitions

The reviewer (different model than the builder) failed this task on one
binding-rule violation and two narration gaps, all in cell transitions,
not in the records, model or validation logic (those all passed clean):

1. 03-stopping-judgment.ipynb had two code cells adjacent with no markdown
   between them (TOASTER_INCREMENT assembly, then the claim cell) --
   toaster-recipe's own pacing rule is explicit that this never happens.
   Inserted one markdown cell between them.

2. 02-second-level.ipynb's claim cells for both AC-C06 and AS-C06 followed
   markdown that only looked backward at the just-printed tag fragment,
   with nothing announcing the claim was next -- every other retrofitted
   notebook has this forward bridge. Appended one sentence to each of the
   two existing cells rather than inserting new ones, since both already
   sat immediately before their claim cell.

3. Two minor staleness fixes: 02 cell-05's construction-zone group list
   didn't include "anchor" (Task 5 already made the equivalent fix in
   Ch4); 03's closing cell still said "this chapter's two prior
   notebooks" after this notebook's own anchor-tag addition made it a
   third contributor to ch06-cumulative.sysml.

4. One wording nit in 03's seam cell (two sentences both opening with
   `validate_record`) tightened per the reviewer's own suggested fix.

Re-executed both notebooks end to end (clean), confirmed zero adjacent
code-cell pairs remain in either (checked programmatically, not just by
re-reading), and re-ran scripts/check_construction.py --check and the
full test suite (423 passed, 7 deselected, both unchanged).
The reviewer (different model than the builder) failed this task on one
real omission: the plan's Task 7 Step 3 explicitly required rewriting the
seam cell to narrate the bridged tag/validate_record connection (the
DL-075 resolution language every other retrofitted notebook already
carries), but the builder's diff never touched it -- it still read only
the pre-existing "claim/proof/counterevidence" sentence, with nothing
interpreting the new "Model tag:" line printed just above it.

Prepended the standard bridged-connection sentence, kept the original
sentence as a second one (it's still true and still worth saying). Also
tightened a bolted-on-sounding wording nit the reviewer flagged in a
nearby cell per their own suggested rewording.

Re-executed the notebook end to end (clean) and re-ran
scripts/check_construction.py --check and the full test suite (423
passed, 7 deselected, both unchanged).
…ting, now propagated into evidence.py, query.py and toaster-recipe/SKILL.md

Caught by Task 11's own builder while cross-checking its edit against already-merged files. The corruption traces back to the plan document itself (written by me), not to any builder's own work -- every builder that copied the plan's text verbatim reproduced it faithfully.
…ut exemption uniqueness

The reviewer (different model than the builder) failed the new AI-C10
exemption-narration cell on two factual errors:

1. It said AI-C10's premises list has "four real findings" -- it has five
   (the notebook-01 traceability-graph finding, AS-C06, AS-C08, AI-C06 and
   AC-C10, confirmed by reading the actual list and by the notebook's own
   earlier cell, which already says "the three ledger records plus AC-C10").

2. It said AI-C10 is "the one record in this tutorial where [the
   exemption] applies" -- false even within this same notebook:
   AI-C10-DRAFT (the negative control four cells earlier) has a non-empty
   premises list and uses the exact same exemption; AI-C04 (Chapter 4)
   does too. Reworded to say AI-C10 is the one record that actually uses
   the exemption rather than carrying a real anchor anyway, which is what
   the notebook's own earlier chapters in fact show.

Also fixed a smaller attribution issue the reviewer flagged (crediting
the exemption to "Hawkins SS3.1's own child-claims-supporting-a-parent
case" rather than correctly as this tutorial's own narrowing rule for the
pure cross-record synthesis case -- Hawkins SS3.1 justifies the
asserted_inference *kind*, not this specific exemption), and updated a
now-stale cell listing what validate_record() checks (it predated Task 1
and never mentioned subject_ref), matching the wording precedent already
established in Task 8's own fix to chapters/ch09-coverage-sufficiency/
02-evidence-completeness.ipynb.

Re-executed the notebook end to end (clean) and re-ran
scripts/check_construction.py --check and the full test suite (423
passed, 7 deselected, both unchanged).
…s Task 9's

The reviewer (different model than the builder) failed this skill edit on
the exact same overclaim Task 9's own review caught and fixed in the
notebook (commit f7dcbb3): "AI-C10 is the one record in this tutorial
that uses this exemption" is false even within this tutorial --
AI-C10-DRAFT (chapters/ch10-traceability-signoff/03-engineering-signoff.ipynb
cell 4) is a negative control with a non-empty premises list and no
subject_ref, and it relies on the very same exemption to validate to
exactly one error.

The builder copied this sentence verbatim from the plan's own Task 10
text before Task 9's review fix existed to correct against -- the error
originates in the plan, not in any builder's own judgment. Reworded to
match the corrected notebook wording exactly: AI-C10 is the one
*original* record that actually relies on the exemption rather than
carrying a real anchor anyway; negative controls rely on it too.

Re-ran the full test suite and glossary check (423 passed, 7 deselected;
0 glossary errors, both unchanged).
…t match reality

Task 11's own review (toaster-recipe) flagged that three sources describe
the judgment-record construction zone's first group three different ways:
toaster-recipe says the anchor group builds the tag fragment before the
claim; the shipped notebooks (Ch2, Ch8, etc.) do exactly that, in two
separate cells; but toaster-review-protocol's own diagram -- unchanged
since before this whole body of work started -- still showed subject_ref
folded into the same group as claim/model_ref, no tag fragment at all,
and a final validate_record(record) call with no model= argument.

Rewrote the diagram to match what's actually shipped: a new first group
building the subject_ref value and the ReviewRecordRef tag fragment
(mirroring a model-increment cell's own fragment-then-assemble pattern),
a second group assembling and printing TOASTER_INCREMENT, then the claim
group, then the original four Hawkins-taxonomy groups unchanged. Updated
the final assembly cell to pass model=model and print the resolved
Model tag line, matching every retrofitted notebook. Added one sentence
distinguishing original-authoring notebooks (seven groups, a real new
tag) from Python-only reconstructions (five groups, no new SysML, but
still exercising the cross-representation check against an
already-committed tag).

Re-ran the full test suite and glossary check (423 passed, 7 deselected;
0 glossary errors, both unchanged).
… missed negative control

Task 15's own builder implemented the plan's literal subject_ref table
(copied from the ToasterDemo chapter retrofits), then flagged that this
was semantically backwards: every exercise notebook runs its own
parallel CoffeeDemo exercise, not an extension of ToasterDemo, and every
record's own model_ref field already correctly names a CoffeeDemo
element one line away from the ToasterDemo subject_ref I'd just told it
to add. Since no exercise notebook passes model= to validate_record,
this had zero effect on any test or assertion -- but it's a real,
confusing, wrong-domain reference a learner would see.

Swapped every subject_ref to its already-present CoffeeDemo analog
(reading each record's own model_ref as the source of truth for what
that analog is): ApplyHeat->ApplyWater, heatGenerationReq->brewReq,
ResistanceCoil->WaterMover, HeatingAssembly::heatGen->
BrewAssembly::mover, deliveredEnergyBoundedBySupply->
deliveredMassBoundedBySupply, EnergyConservationReq->
massConservationReq (matching AC-C10-EX's own model_ref hint,
"CoffeeDemo::massConservationReq"), timely->tempCheck (ch03, matching
its own model_ref placeholder's hint).

Also fixed the builder's second finding: exercises/ch09/exercise.ipynb
has an empty-identifier negative control (cell 35) the plan's own
AST-based record inventory missed entirely (an empty identifier string
structurally can't match a "-EX"-suffixed search). It had gained an
unintended second validation error under the new required-ness rule,
same as every other negative control fixed throughout this plan -- added
subject_ref="CoffeeDemo::deliveredMassBoundedBySupply" so it again
demonstrates exactly one error.

Every touched file re-validated with nbformat; full test suite re-run
(423 passed, 7 deselected, unchanged -- exercises aren't in the pytest
suite, so this confirms no collateral breakage elsewhere).
… the definition

The reviewer (different model than both the original builder and the
orchestrator's own follow-up fix) caught that AC-C10-EX's subject_ref
("CoffeeDemo::massConservationReq", lowercase) named a requirement
*usage*, while the real chapter's AC-C10 is explicitly about the
requirement *definition* (chapters/ch10-traceability-signoff/
01-traceability-graph.ipynb cells 37/39/49 all anchor on
EnergyConservationReq, the definition, not energyConservationReq, the
usage -- and the exercise's own cell 23 asks the learner to confirm the
tie is "attributed to your own new requirement's DEFINITION (not its
usage)", so the subject_ref I'd written contradicted the exercise's own
instructions). Fixed to "CoffeeDemo::MassConservationReq" (capitalized,
the definition form), matching the real chapter's own choice.

Also re-saved the four other files my previous fix touched
(ensure_ascii=False) to stop unicode characters like em-dashes and
degree signs from round-tripping as \uXXXX escapes, which the reviewer
flagged as unnecessary diff noise -- the content is unchanged, only the
on-disk byte representation.

Re-validated all six touched notebooks with nbformat and re-ran the full
test suite (423 passed, 7 deselected, unchanged).
… DL-075

Consolidated decision-log entry for the whole Hawkins judgment-record
anchor body of work (Tasks 1-15 of the implementation plan), covering
the schema change, the SysML construct, the query helper, every chapter
and exercise retrofit, both skill updates, and the glossary edge. Closes
DL-075 (the recurring judgment-record seam escalation, nine occurrences
across the fresh user-testing battery) by making its own recommended
reading concrete and checkable rather than ruling among its three
options -- an addendum on DL-075 itself cross-references this entry.

Records, rather than silently drops, every real gap found and
deliberately left for a later pass during this session's own review
cycles: a narrow cross-representation asymmetry in validate_record for
an exempt-but-tagged record (affects no record that exists today); a
few documentation cross-references not yet updated (Ch2's index.md/
conclusion.md, a cumulative-model header comment); one general-vs-
judgment-record seam-cell wording question (one sentence vs. two) left
unreconciled across two reviewers who both accepted it; a pre-existing,
unrelated file-path error in toaster-review-protocol's own worked
example; a pre-existing, repo-wide section-symbol typo predating this
session, deliberately left untouched as out of scope; and two commits
on this branch carrying a Co-Authored-By trailer, which could not be
cleaned up because the needed history rewrite was blocked by this
session's own auto-mode git-safety policy -- flagged for the user to
fix by hand, since nothing on this branch has been pushed.
Found by building and navigating the rendered MyST book (not just reading
source): three real, user-visible defects.

1. Chapter 8's page title ("Constraint Checking") didn't match its own
   navigation title in myst.yml ("Checking and Revision"), and the
   sidebar's wording is the more accurate one -- it covers notebook 03's
   revision/staleness content too, which "Constraint Checking" alone
   omits. Fixed the page's own H1 to match, in both
   chapters/ch08-checking/index.md and docs/index.md's own chapter
   outline table, which had the same stale wording as a third copy.

2. Chapter 9's page title used a dash ("Chapter 9 - Coverage and
   Sufficiency") where every other chapter, including myst.yml's own TOC
   entry for Chapter 9 itself, uses a colon. Normalized for consistency.

3. The home page's "Case studies" link pointed at docs/case-studies/ (a
   bare directory with no index page and no myst.yml entry), which
   404s. Added the one real file there
   (2026-09-30-energy-conservation-requirement-tie.md) to myst.yml's
   project TOC and pointed the link at it directly.

Verified live in the rendered book (hot-reloaded dev server, not just
grep): all three now resolve and read correctly. Full test suite and
check_construction.py --check both still pass.
Three README inaccuracies found by direct comparison against the repo:
- Attribution credited only Douglas Part 3; docs/references.md already
  documents that both Part 3 and Part 4 are the tutorial's ground truth
  (Part 4 is cited directly in Chapter 2's requirement-def notebook).
- The "build the site" command (npx mystmd build --execute) doesn't match
  the command this repo actually uses and documents elsewhere
  (.claude/launch.json, docs/setup.md: npx mystmd start --execute, a dev
  server, not a one-shot static build).
- The src/toaster module list was missing two real modules
  (conformance.py, modelcheck.py -- the latter is Chapter 8's own Z3
  wrapper, not a minor omission).

Two docs/setup.md inaccuracies, found by actually running an exercise
notebook rather than trusting the prose:
- "apply the same construct... to a different part of the toaster" is
  wrong -- every exercise builds a parallel coffee-maker model
  (CoffeeDemo), not a different part of the same toaster. Verified
  directly against exercises/ch01/exercise.ipynb.
- The sysml-toolkit paragraph said "no chapter currently uses this; it
  becomes relevant once Chapter 8 is re-derived to need it" -- stale from
  before Chapter 8's own re-derivation. Chapter 8 already uses it
  directly (toaster.modelcheck.verify_holds, confirmed by reading the
  real source) to prove deliveredEnergyBoundedBySupply.

Verified live in the rendered book (hot-reloaded) and against the full
test suite (423 passed, 7 deselected, unchanged).
README, myst.yml, and every rendered page footer already declare
Apache-2.0, but no LICENSE file was committed to back that claim --
GitHub's license detector and badge need the actual text at the repo
root, not just a reference to it. Standard, unmodified Apache License
2.0 text (the canonical form every GitHub license-detector match is
built against); no copyright-year placeholder filled in, since that
template line in the Appendix is guidance for individual source-file
headers, not something to customize in the root LICENSE file itself.
Two documents, both drafts for review before any implementation:

1. docs/superpowers/plans/2026-10-01-ci-cd-deploy-readiness-plan.md --
   an 8-task implementation plan turning the CI workflow's deploy
   placeholder into the real seven-step pipeline docs/contributor.md
   already specifies, ending with Z's own deliberate flip of if: false
   (never bundled into the same change as the pipeline work itself).
   Grounded in what's already documented and what's confirmed missing
   by direct inspection (check_construction.py/check_conformance.py
   exist and work but aren't wired into CI; no provenance-manifest
   generator exists at all; no /toaster base-path config exists yet).

2. docs/superpowers/specs/2026-10-01-large-scale-user-testing-design.md
   -- extends the existing user-testing skill's persona battery (which
   today only tests one modality, a local chapter read-through) across
   a persona x modality grid (GitHub repo, GitHub Pages, local clone
   reading chapters, local clone doing exercises), scoped to nine cells
   for a first run. The interpretive synthesis layer is the existing
   ACE, not a new role -- its synthesis protocol gains one new step,
   interviewing a specific sub-agent by name (SendMessage resuming its
   own transcript) when a structured finding is ambiguous, before
   triaging and ruling or escalating, the same way this session's own
   reviewer agents probed builder self-reports all night rather than
   trusting them at face value.

Full test suite unaffected (423 passed, 7 deselected).
…rktree dispatch, void false alarms, escalate exercise-scaffolding to Z

- decisions/log.md: DL-085, the full synthesis -- zero blocking content defects
  across the grid; four subject_ref/ReviewRecordRef claims voided as false
  alarms traced to a hardcoded cd path in user-testing/SKILL.md; model_ref and
  the requirement_coverage() feature-chain limitation ruled, not gaps; two
  exercise-scaffolding clusters (first-run convention, carry-forward) escalated
  to Z with briefs; five minor/cosmetic items batched for Z's direction.
- .claude/skills/user-testing/SKILL.md: fix the hardcoded cd that caused the
  one false NEEDS-FIX (mixed-path read across a worktree and the main checkout).
- docs/superpowers/specs/2026-10-01-large-scale-user-testing-design.md: fix
  the nine/ten cell arithmetic error; stop telling dispatchers to rely on the
  Agent tool's own isolation:worktree mechanism, which the orchestrator role
  already bans for exactly this reason; record this run's confirmed findings
  in Open items.
- docs/superpowers/plans/2026-10-01-ci-cd-deploy-readiness-plan.md: Task 5's
  static-build link checker must also cover exercise links, which fall through
  to a dev-only Jupyter server today and are unverified on the static build.
- src/toaster/query.py: docstring recording that requirement_coverage() keys
  on a satisfy relationship's top-level subject only.
… DL-085)

README: point uv/mystmd newcomers to docs/setup.md's fuller explanation; note
that AGENTS.md/CLAUDE.md/DEFERRED.md are the project's own working contract,
not learner material, with a pointer to docs/contributor.md.

docs/contributor.md: new section explaining the multi-agent harness (agent
roles, skills, decisions/) that builds and reviews this tutorial's content,
so a contributor knows where to start for extending, clarifying, or
reviewing didactic content -- per Z's own framing of what this page should do.
…y-forward

Brief A (gate ch08 only): exercises/ch08/exercise.ipynb's first cell now
asserts model.ok immediately, matching ch09/ch10's own gating pattern,
instead of printing and continuing into an uncaught AssertionError two
cells later. Verified by nbconvert: the gate now fires at cell 1 with a
message naming the actual problem.

Brief B (document, don't collapse): investigated the 'two Ch6 states'
finding before touching anything -- it's deliberate, not an oversight.
exercises/ch06/exercise.ipynb hashes AS-C06-EX (the mechanism-selection
judgment) against the pre-Impeller state and AI-C06-EX (the stopping
judgment) against the post-Impeller state, encoding DL-065's own
judgment-precedes-construction ordering. Collapsing to one state would
erase that. Instead: docs/setup.md gains a 'Keep your model between
chapters' section, and exercises/ch06/exercise.ipynb's own two snapshot
cells each gain a comment flagging exactly what to save and why.
@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

@mzargham
mzargham merged commit 9e269ad into main Oct 1, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant