Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
60 commits
Select commit Hold shift + click to select a range
de74112
Fix user-testing skill: content-type cell IDs, Novice one-sentence rule
mzargham Oct 1, 2026
06af2ef
Record DL-074: fresh user-testing checkpoint PASS for Chapter 1
mzargham Oct 1, 2026
9f01a76
Record DL-075: fresh user-testing checkpoint PASS for Chapter 2, esca…
mzargham Oct 1, 2026
eeb99ba
Record DL-076: fresh user-testing checkpoint PASS for Chapter 7
mzargham Oct 1, 2026
ef9e03d
Record DL-077: fresh user-testing checkpoint PASS for Chapter 10
mzargham Oct 1, 2026
2fbb1d5
Record DL-078: fresh user-testing checkpoint PASS for Chapter 6
mzargham Oct 1, 2026
7fe6bd3
Regenerate ch05 interconnection figure (stale allocate-end label from…
mzargham Oct 1, 2026
fae49b4
Record DL-079 and DL-080: fresh user-testing checkpoint PASS for Chap…
mzargham Oct 1, 2026
8bcd529
Record DL-081: fresh user-testing checkpoint PASS for Chapter 8
mzargham Oct 1, 2026
cdb96b9
Record DL-082: fresh user-testing checkpoint PASS for Chapter 3
mzargham Oct 1, 2026
2800f07
Record DL-083: fresh user-testing checkpoint PASS for Chapter 9 -- fr…
mzargham Oct 1, 2026
abfd35c
Spec: anchor Hawkins judgment records to the model (subject_ref + Rev…
mzargham Oct 1, 2026
41ef33f
Plan: implementation plan for Hawkins judgment-record anchor design
mzargham Oct 1, 2026
2931ae2
Add get_review_record_refs: model-to-Python direction of the judgment…
mzargham Oct 1, 2026
345f179
Ignore agent-managed git worktrees under .claude/worktrees/
mzargham Oct 1, 2026
d0eb376
Add subject_ref to ReviewRecord: the tutorial's Assurance Claim Point…
mzargham Oct 1, 2026
8e6649a
Fix test_simulate.py fixture for the new subject_ref required-ness rule
mzargham Oct 1, 2026
ac8b339
Lint cleanup: sort imports, drop unused hash_content, use dict litera…
mzargham Oct 1, 2026
3ca701a
Plan amendment: add Task 15 (exercises/ retrofit) and Task 14 noteboo…
mzargham Oct 1, 2026
1dcd8f9
Glossary: tie subject_ref/ReviewRecordRef to Hawkins' Assurance Claim…
mzargham Oct 1, 2026
b4d3ac6
Fix Task 12 per independent review: proposed status, correct file/id,…
mzargham Oct 1, 2026
2691c61
Introduce ReviewRecordRef metadata def; tag AC-001's real subject (no…
mzargham Oct 1, 2026
3e5ab8a
Fix Task 3 per independent review: accurate transitions, grounded sea…
mzargham Oct 1, 2026
e777710
Retrofit AC-C03/AS-C03 with subject_ref and ReviewRecordRef tags (bot…
mzargham Oct 1, 2026
3649d0b
Retrofit AI-C04 with subject_ref and ReviewRecordRef tag (about Apply…
mzargham Oct 1, 2026
03c293b
Fix plan typo: Task 5's SysML snippet said 'about applyHeat', contrad…
mzargham Oct 1, 2026
65fe98d
Retrofit AC-C06/AS-C06/AI-C06 with subject_ref and tags; fix AI-BAD's…
mzargham Oct 1, 2026
2992a0b
Fix Task 6 per independent review: pacing-rule violation, stale/missi…
mzargham Oct 1, 2026
8686a5d
Retrofit AS-C08 with subject_ref and tag; fix revision-flow's negativ…
mzargham Oct 1, 2026
ecc180c
Fix Task 7 per independent review: the seam cell was never rewritten
mzargham Oct 1, 2026
15f15e6
Carry subject_ref forward in Ch9 reconstructions; fix negative contro…
mzargham Oct 1, 2026
f6a71eb
Retrofit AC-C10 with subject_ref and tag; carry ledger values forward…
mzargham Oct 1, 2026
3a34ca4
toaster-recipe: judgment-record construction zone narrates subject_re…
mzargham Oct 1, 2026
5e5aa18
Fix a section-symbol mojibake (SS -> §) introduced by my own plan-wri…
mzargham Oct 1, 2026
bb2368a
toaster-review-protocol: document subject_ref as this tutorial's ACP …
mzargham Oct 1, 2026
5e796a6
Fix Task 9 per independent review: wrong premise count, overclaim abo…
mzargham Oct 1, 2026
080d898
Fix Task 10 per independent review: same AI-C10 exemption overclaim a…
mzargham Oct 1, 2026
c715dc9
Fix Task 11 reviewer's open question: construction-zone diagram didn'…
mzargham Oct 1, 2026
6f045f3
Retrofit exercises/ with subject_ref, mirroring each chapter's own value
mzargham Oct 1, 2026
58fa10d
Fix Task 15's builder-flagged findings: wrong-domain subject_ref, one…
mzargham Oct 1, 2026
6effc04
Fix Task 15 per independent review: AC-C10-EX pointed at a usage, not…
mzargham Oct 1, 2026
7a3e135
Record DL-084: Hawkins judgment records anchored to the model; closes…
mzargham Oct 1, 2026
7068805
Browser testing: fix title mismatches and a broken home-page link
mzargham Oct 1, 2026
0a63325
Bring README and setup.md in line with the repo's actual current state
mzargham Oct 1, 2026
1ab9750
Add Apache-2.0 LICENSE file
mzargham Oct 1, 2026
6d404b1
Plan and design: CI/CD deploy-readiness, large-scale user-testing grid
mzargham Oct 1, 2026
4746680
User testing M1-novice: GitHub repo modality, Novice persona evaluation
mzargham Oct 1, 2026
0b70292
User testing M2-practitioner: GitHub Pages modality, SE Practitioner …
mzargham Oct 1, 2026
5e20be9
User testing: M2-novice browser-rendered book, Chapters 1-2, Novice p…
mzargham Oct 1, 2026
9e44df7
User testing: M2-returning browser-rendered book, Chapter 10, Returni…
mzargham Oct 1, 2026
df31fa5
User testing M3-novice: Ch2 execution report with structured findings
mzargham Oct 1, 2026
0c96551
User testing M3-practitioner: SE Practitioner persona evaluation of C…
mzargham Oct 1, 2026
69863b7
User testing M3-returning: Ch10 traceability/signoff, Returning Learn…
mzargham Oct 1, 2026
3948385
User testing M4-novice: local clone exercise, Novice persona, Chapter 1
mzargham Oct 1, 2026
2102056
User testing M4-practitioner: ch08 exercise cold run confirms uncaugh…
mzargham Oct 1, 2026
62dcd5e
User testing M4-returning: local clone exercise modality, Returning L…
mzargham Oct 1, 2026
fb336ba
ACE synthesis of the 10-cell user-testing grid (DL-085): fix stale-wo…
mzargham Oct 1, 2026
7dcd6f2
Contextualize root harness files for a first-time visitor (M1-novice,…
mzargham Oct 1, 2026
1803739
Implement Z's two DL-085 rulings: gate ch08's exercise, document carr…
mzargham Oct 1, 2026
75c8b04
DL-085: record Z's decisions and what was implemented
mzargham Oct 1, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 11 additions & 1 deletion .claude/skills/toaster-recipe/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -126,6 +126,13 @@ below" — never a sentence built around naming the categories themselves. `user
simulated-learner checklist judges this behaviorally (does removing any label still leave the
connection legible?), which is exactly the test a lint rule can't run.

**Judgment-record notebooks specifically:** the seam cell narrates one bridged connection — the
tagged SysML text (the subject plus its `ReviewRecordRef` usage), the tool that loads it and
cross-checks both the Python record and the model tag, and a result showing they agree
(`validate_record` plus `get_review_record_refs`) — not a choice between the model's own
construct/tool/result triad and the record's own fields/`validate_record`/result triad
(`decisions/log.md` DL-075, retired by this convention).

## Chapter index.md — 6-element recipe

1. Purpose — engineering question and model state after completing the chapter
Expand Down Expand Up @@ -160,7 +167,10 @@ someone happens to reread it.
A notebook that builds a `ReviewRecord` (an `asserted_context`, `asserted_inference` or
`asserted_solution` judgment) uses `toaster-review-protocol`'s own construction-zone pattern for
it, not one dense call: name each group of fields, narrate what it's for, print it, then assemble.
The size limits below are relaxed for this content (see that skill for the exact grouping and why).
The first group now names the record's own subject (`subject_ref`) alongside its `claim`, and
builds the matching `ReviewRecordRef` metadata-tag fragment the same way a model-increment cell
builds any other named fragment (see toaster-review-protocol's own subject_ref section). The size limits below
are relaxed for this content (see that skill for the exact grouping and why).

## Size limits (A6 review criteria)

Expand Down
77 changes: 65 additions & 12 deletions .claude/skills/toaster-review-protocol/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,11 +13,45 @@ description: Hawkins et al. 2011 judgment record fields, three ACP types, two ev

## Three judgment sites and ACP kinds

| Site | Kind | Hawkins ref |
|---|---|---|
| Assumption or context used for a claim | `asserted_context` | §3.2 |
| Child claims supporting a parent | `asserted_inference` | §3.1 |
| Evidence supporting a conclusion | `asserted_solution` | §3.3 |
| Site | Kind | Hawkins ref | `subject_ref` |
|---|---|---|---|
| Assumption or context used for a claim | `asserted_context` | §3.2 | required |
| Child claims supporting a parent | `asserted_inference` | §3.1 | required unless `premises` is non-empty |
| Evidence supporting a conclusion | `asserted_solution` | §3.3 | required |

## `subject_ref`: this tutorial's narrowed Assurance Claim Point

Hawkins' own Assurance Claim Point (ACP) is never free-floating: every confidence argument is
anchored to one specific, located assertion in the argument (Hawkins 2011, Sec. 3, p. 8 —
`glid:def-hawkins--assurance-claim-point`). `subject_ref` is this tutorial's own narrowed,
single-element analog: the one qualified name the record's `claim` is directly about, checkable
both from Python (`validate_record(record, model=model)` resolves it via `model.find()`) and from
the model's own side, via a real SysML metadata tag (SysML v2 formal/2026-03-02 §7.27.2,
MetadataDefinition):

```sysml
metadata def ReviewRecordRef {
attribute identifier : ScalarValues::String;
}

metadata ac001Tag : ReviewRecordRef about nominal {
identifier = "AC-001";
}
```

(the committed text in `models/ch02-cumulative.sysml`, built in
`chapters/ch02-requirements/03-judgment-context.ipynb`.)

The `about` clause binds the usage's inherited `annotatedElement` feature to the named subject — a
real, queryable model relationship, not a string a reader has to trust. `subject_ref` is required
for `asserted_context` and `asserted_solution`; for `asserted_inference` it may stay empty only
when `premises` is non-empty (the pure cross-record synthesis case — `AI-C10` is the one original
record that actually relies on this exemption rather than carrying a real anchor anyway; negative
controls such as `AI-C10-DRAFT` rely on it too). `src/toaster/query.py`'s `get_review_record_refs()` is the
model-to-Python direction: given a loaded model, it finds every `ReviewRecordRef` tag and what it's
about, independent of any notebook's own Python objects. `validate_record` cross-checks both
directions automatically whenever a `model` is passed and a tag already exists for that record's
own `identifier`.

## ReviewRecord required fields

Expand All @@ -28,6 +62,7 @@ record = ReviewRecord(
identifier="RR-001",
kind="asserted_solution",
claim="DeliveredEnergy >= 50000 J at nominal operating conditions",
subject_ref="ToasterDemo::HeatGenerator::deliveredEnergy",
model_ref="models/ch07-snapshot.sysml",
content_hash=hash_content(open("models/ch07-snapshot.sysml").read()),
scope="nominal operating envelope: P=800W, t=120s, eta=0.7",
Expand Down Expand Up @@ -64,7 +99,18 @@ group of fields, narrate what it's for, print it, then assemble. Group by the qu
Hawkins' taxonomy is answering, not by the dataclass's field order:

```
[markdown] narration: what is being claimed, and about what
[markdown] narration: what is being claimed, and what, specifically, it is about
[code] subject_ref = "ToasterDemo::..."
SOME_TAG = """\
metadata someTag : ReviewRecordRef about <subject> {
identifier = "..."
}
"""
print(SOME_TAG)
[markdown] narration: this fragment is the same text now committed in the cumulative model
[code] TOASTER_INCREMENT = SOME_TAG # or assembled with any other new fragment this notebook adds
print(TOASTER_INCREMENT)
[markdown] narration: the claim itself comes next
[code] claim = "..."
model_ref = "..."
[markdown] narration: what standard the claim is checked against (appropriateness)
Expand All @@ -82,21 +128,28 @@ Hawkins' taxonomy is answering, not by the dataclass's field order:
[code] counterevidence = "..."
residual_uncertainties = "..."
[markdown] narration: assembling the record from the named parts above
[code] record = ReviewRecord(identifier=..., kind=..., claim=claim, model_ref=model_ref,
content_hash=hash_content(source), scope=scope, criteria=criteria,
[code] record = ReviewRecord(identifier=..., kind=..., claim=claim, subject_ref=subject_ref,
model_ref=model_ref, content_hash=hash_content(source), scope=scope, criteria=criteria,
premises=premises, assumption_refs=assumption_refs,
evidence_refs=evidence_refs, rationale=rationale,
counterevidence=counterevidence,
residual_uncertainties=residual_uncertainties,
disposition="pending", dependency_freshness="current",
engineering_conclusion=..., record_kind="worked_example")
errors = validate_record(record)
errors = validate_record(record, model=model)
tag = next((t for t in get_review_record_refs(model) if t["identifier"] == record.identifier), None)
print(f"Model tag: {tag}")
print(f"Validation errors: {errors}")
```

Five groups, five narration cells, matching the model-fragment construction zone's pacing rule (no
two code cells adjacent). Each `print`ed group is the record's own reflection, the same role a
printed `TOASTER_INCREMENT` plays for a model fragment.
Seven groups when the notebook is introducing a new tag (the two anchor groups above, then the five
Hawkins-taxonomy groups); five when it is a Python-only reconstruction that cites an already-tagged
identifier from an earlier chapter (no new SysML, so no anchor groups, but `model=model` and the
`Model tag` lookup still run, exercising the cross-representation check against the already-committed
tag). Every code cell is still followed by a markdown cell narrating what's next (no two code cells
adjacent). Each printed group is its own reflection, the same role a printed `TOASTER_INCREMENT`
plays for a model fragment — and `TOASTER_INCREMENT` here really is the Hawkins record's own model-
side anchor, assembled and loaded the same way any other chapter's model increment is.

**Size limit:** `toaster-recipe`'s ≤600 words / ≤50 lines budget is sized for a notebook whose main
content is one model construct. A notebook whose construct is a judgment record may exceed it — the
Expand Down
59 changes: 44 additions & 15 deletions .claude/skills/user-testing/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,36 +27,65 @@ One agent per persona. No two agents with identical persona in one checkpoint ru

## Execution checklist (`simulated-learner` must run these in order)

Identify each step below by what the cell *does*, not by a fixed position: a construction-zone
notebook may have several fragment cells before assembly, a judgment-record notebook can run
17-27 real cells, and some chapters carry the seam across several cells' prose rather than one
dedicated cell (Chapter 10's own distributed-seam design is a real, valid instance of this, not a
gap) — the same "by content type, not cell index" rule `toaster-recipe` already states for its own
review. If you cannot find a cell matching a step below, say so explicitly rather than guessing
which numbered cell it must be.

For each sub-notebook in the assigned chapter(s):

1. **Read index.md** — does it orient you? Note any undefined terms or missing prerequisites.
2. **Cell 0** — read the concept statement. Is it exactly one sentence? Does it state what you will learn?
3. **Cell 1** — read the context paragraph. Does it locate this notebook in the arc? Is there a link to the prior notebook where needed?
4. **Cell 2 (execute)** — run the model-loading code. Record: `model.ok`, any diagnostic output.
5. **Cell 3 (execute)** — run the negative control. Record: `bad.ok` (must be False), printed diagnostic message.
6. **Cell 4 (execute)** — run the demonstration. Record: output produced; note if it matches what cell 0 promised.
7. **Cell 5** — read the Tall seam. AGENTS.md 1.10 binds that learner content **never names** Tall or "the three worlds" (the `tall-named` lint rule, `glossary/lint_rules.toml`, DL-028, already enforces the never-name half in CI). Your job is the half a lint rule cannot judge: does the cell **address the seam in behavior** — is it clear, without naming the lens, that the SysML text, the tool that loads and runs it, and the rendered/printed result are three distinct things the reader has just seen connect? Record which of the three you could each point to concretely from what the cell actually showed, and whether a reader who had not been told there were "three worlds" would still notice the seam.
8. **Cell 6** — read the exercise pointer. Is it one sentence? Does it describe what the exercise asks?
9. **Read conclusion.md** — three paragraphs (what was built / what this establishes / what comes next) plus exercise reference?
2. **The concept-statement cell** — read it. Is it exactly one sentence? (A single sentence may
contain a semicolon joining two independent clauses and still be one sentence — count terminal
periods, not semicolons or conjunctions, before judging this a failure.) Does it state what you
will learn?
3. **The context cell** — read it. Does it locate this notebook in the arc? Is there a link to the
prior notebook where needed?
4. **The model-increment cell(s) (execute)** — run the model-loading code. Record: `model.ok`, any
diagnostic output.
5. **The negative-control cell (execute)** — run it. Record: `bad.ok` (must be False), printed
diagnostic message.
6. **The demonstration cell(s) (execute)** — run them. Record: output produced; note if it matches
what the concept-statement cell promised.
7. **The seam cell(s)** — read them. AGENTS.md 1.10 binds that learner content **never names** Tall
or "the three worlds" (the `tall-named` lint rule, `glossary/lint_rules.toml`, DL-028, already
enforces the never-name half in CI). Your job is the half a lint rule cannot judge: does the
content **address the seam in behavior** — is it clear, without naming the lens, that the SysML
text, the tool that loads and runs it, and the rendered/printed result are three distinct things
the reader has just seen connect? Record which of the three you could each point to concretely
from what the notebook actually showed, and whether a reader who had not been told there were
"three worlds" would still notice the seam.
8. **The exercise-pointer cell** — read it. Is it one sentence? Does it describe what the exercise
asks?
9. **Read conclusion.md** — three paragraphs (what was built / what this establishes / what comes
next) plus exercise reference?

Execution command:

```sh
cd /Users/z/Documents/GitHub/toaster
cd <the notebook's own directory in YOUR worktree, never the main checkout>
uv run python - <<'EOF'
[paste cell code here]
EOF
```

A worktree-isolated cell that runs from the main checkout's path instead of its own worktree reads
some files (e.g. notebook text) from the wrong branch state while reading others (e.g. model files)
from its own worktree — a mixed-path read that produces false findings. Confirmed as the cause of a
false NEEDS-FIX verdict in the first grid run (`decisions/log.md` DL-085).

## Report format

```
LEARNER [ID] — [Persona] — Ch[N]

EXECUTION RESULTS:
- nb[N] cell2: ok=[True/False] | [diagnostic if any]
- nb[N] cell3: neg_ok=[True/False] | diagnostic: [message]
- nb[N] cell4: output=[one-line summary]
- nb[N] model-increment: ok=[True/False] | [diagnostic if any]
- nb[N] negative-control: neg_ok=[True/False] | diagnostic: [message]
- nb[N] demonstration: output=[one-line summary]
[repeat for each notebook]

NARRATIVE OBSERVATIONS (top 3, each quoting exact text):
Expand All @@ -65,9 +94,9 @@ NARRATIVE OBSERVATIONS (top 3, each quoting exact text):
3. "[exact quote]" — [learner reaction in one sentence]

STRUCTURAL CHECKS:
- Cell 0 one sentence: [yes/no]
- Cell 5 addresses the seam without naming it: [yes/no] — [which of the three you could point to; if no, what's missing]
- Cell 6 one sentence: [yes/no]
- Concept-statement cell is one sentence: [yes/no]
- Seam addressed in behavior without naming it: [yes/no] — [which of the three you could point to; if no, what's missing]
- Exercise-pointer cell is one sentence: [yes/no]
- conclusion.md three paragraphs + exercise reference: [yes/no]

OVERALL: [PASS/NEEDS-FIX] — one sentence.
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -45,3 +45,6 @@ glossary/sources/local/*
# reproducible across machines: see chapters/ch08-checking/02-violation-witness.ipynb)
chapters/*/companion-check-scratch/
exercises/*/companion-check-scratch/

# Agent-managed git worktrees (builder/reviewer isolation during plan execution)
.claude/worktrees/
Loading
Loading