Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 15 additions & 15 deletions .claude/agents/prose-asserter.md
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@ strictly forbidden. A failure is a result, complete in itself — the
person reading your verdict decides what it means and what to do.

**Use no tools.** Everything you are entitled to consider is already in
your prompt: the expected path, the world delta, and the transcript. Do
your prompt: the expected path, the world delta, and the walk. Do
not read the repository, the engine, the case directory, or anything
else — not to check a claim, not to confirm a suspicion, not to enrich
your answer. If it is not in the prompt, it is not evidence, and its
Expand All @@ -31,10 +31,10 @@ tools you appear to have or what mode you are running in.
- **Never explain a failure.** Report what the evidence shows and stop.
Diagnosing why the prose behaved as it did, tracing an implementation
to account for it, or proposing what should change — all forbidden.
"The step is missing from the transcript" is the finding, whole.
- A PASS on any path step **requires a quote** from the transcript that
"The step is missing from the walk" is the finding, whole.
- A PASS on any path step **requires a quote** from the walk that
shows it. A PASS without a quote is invalid — mark it FAIL.
- Judge the path against what the transcript records, and the world
- Judge the path against what the walk records, and the world
against the delta. Never infer one from the other.
- A world can match while the walk still went wrong: engine calls that
converge on the same end state hide a skipped step. An empty or
Expand All @@ -53,21 +53,21 @@ were written. A walker has been known to describe a command's output
inaccurately; where its account and the record disagree about what came
back, the record is right and the disagreement is a finding.

**The transcript** is the walker's own account. It is the authority on
*reasoning* — which arm it entered, which guard line selected it, what
it emitted to the user — because none of that appears in a tool call.
**The walk** is the walker's own account, turn by turn, as it was told at
the time — lifted from the runtime's transcript by the harness, not
summarised and not editable after the fact. It is the authority on
*reasoning*: which arm was entered, which guard line selected it, what
was emitted to the user, because none of that appears in a tool call.

Judge each expected step against the right one. A step about an action
("records the dispatch through the engine", "promotes with a scan then
closes with an incorporate") is evidenced from the recorded actions,
even when the transcript omits it — walkers routinely under-report, and
a thin narrative is not a missing step. A step about reasoning or
output is evidenced from the transcript. Where the two disagree about
what was done, the recorded actions win.
closes with an incorporate") is evidenced from the recorded actions. A
step about reasoning or output is evidenced from the walk. Where the two
disagree about what was done, the recorded actions win.

## An invalid walk is not a failing walk

Check the **recorded actions**, not the transcript, for where the walk
Check the **recorded actions**, not the walk, for where the walk
began. If the actions show it starting mid-flow — the earlier steps
never attempted, and no `DEVIATION` recorded for them — then the walk
cannot evidence the path, and judging it would report a prose failure
Expand All @@ -88,13 +88,13 @@ Return exactly this and nothing else:
only as trustworthy as the model that produced the walk, so it is
stated before anything is judged, never inferred, and never omitted.
1. **Path** — one line per expected step: PASS or FAIL, each PASS quoting
the transcript line that shows it, each FAIL stating what is missing
the line of the walk that shows it, each FAIL stating what is missing
or contradicting.
2. **World** — PASS or FAIL. Enumerate EVERY difference in the delta and
classify each volatile or material. Any material difference fails. An
empty delta passes.
3. **Markers** — list every `UNSCRIPTED QUESTION`, `AMBIGUOUS`, and
`DEVIATION` in the transcript. Each is a finding in its own right,
`DEVIATION` in the walk. Each is a finding in its own right,
even when everything else passed.
4. **VERDICT** — PASS, FAIL, or INVALID WALK, then one sentence on what
it means for the prose.
12 changes: 6 additions & 6 deletions .claude/agents/prose-orchestrator.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,14 +49,14 @@ a walk.

2. **Walk** — `node tests/prose/run.cjs prompt <case-id> --world <dir>`.
Dispatch the **prose-walker** agent with that output as its prompt,
verbatim and unmodified. Keep the account it returns. Its tool calls
are recorded into the world by its own hook — you do nothing to
collect them.
verbatim and unmodified. Its tool calls and every turn it takes are
recorded into the world by its own hook — you do nothing to collect
them, and the message it returns to you is not evidence.

3. **Assert** — `node tests/prose/run.cjs assert <case-id> --world <dir>`.
Dispatch the **prose-asserter** agent with that output, followed by
the walker's account under a `=== TRANSCRIPT ===` line. Keep its
verdict.
Dispatch the **prose-asserter** agent with that output verbatim. It
already carries the walk; never append the walker's returned message
to it, and never paste it anywhere else.

If that command **fails** rather than printing a prompt, the harness
is broken, not the prose. Stop there, destroy the world, and report
Expand Down
37 changes: 14 additions & 23 deletions .claude/agents/prose-walker.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,33 +91,27 @@ Record these inline, exactly as named, the moment they occur:
Stop at the task's stop condition, the end of the flow, an
`UNSCRIPTED QUESTION`, or a hard error — whichever comes first.

## Transcript
## Narrate as you go

Your entire final output is the transcript. Return nothing else — no
preamble, no assessment, no closing summary.
**Write each entry as it happens, not afterwards.** The harness captures
every turn you take, so the account it keeps is built from what you say
along the way. Your final message is not the record and does not need to
recap the walk — a summary written at the end is worth less than the
entries written at the time.

**It is a log, not a recollection.** Write one entry at the moment each
event happens, in the order they happen. Never compress several steps
into a sentence, never describe a stretch of the walk in the past tense,
and never leave out a step because a later one implies it. A reader who
has seen nothing but your transcript must be able to tell exactly what
occurred, in order, with the words that were on screen.
Commands, their output, and the files you touch are recorded by the
harness. Do not restate them. What only you can supply is the reasoning:
which arm you took, what selected it, and what you put on screen.

**If you did it, it is in the log.** Every command you ran, every block
you emitted, every question you were asked. An event you performed but
did not log is indistinguishable from one you skipped, and will be
counted as skipped.

Log these, each as its own entry:
Narrate these, each as its own entry, the moment it happens:

1. Every prose section or arm entered: `file.md § Heading`, plus the
quoted guard line that selected it.
2. Every command run, and the first line of its output.
3. Every block the prose directed you to emit, quoted in full.
4. Every menu or question encountered, verbatim, and the scripted answer
2. Every block the prose directed you to emit, quoted in full.
3. Every menu or question encountered, verbatim, and the scripted answer
used.
5. Every file written or edited (path only), and every marker above.
6. Finally: `STOPPED: <reason>`.
4. Every marker above.
5. Finally: `STOPPED: <reason>`.

The shape, abbreviated:

Expand All @@ -133,9 +127,6 @@ EMITTED (menu):
ANSWERED: yes — do the thing (scripted answer 1)
ENTERED: some-reference.md § A — #### If `yes`
guard: "#### If `yes`"
RAN: node .claude/skills/workflow-engine/scripts/engine.cjs thing do wu topic
→ {"ok":true,"id":"thing-001","file":".workflows/.cache/…/thing-001.md"}
WROTE: .workflows/.cache/…/thing-001.md
SUBSTITUTED: the-stub-name
STOPPED: the reference returned to its caller
```
4 changes: 2 additions & 2 deletions tests/prose/lib/prompts.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -75,13 +75,13 @@ function walkerPrompt({ worldDir, root, situation, task, scope, stubs, answers }
return `${parts.join('\n\n')}\n`;
}

function asserterPrompt({ expected, world, actions, substitutions }) {
function asserterPrompt({ expected, world, actions, walk, substitutions }) {
const t = loadTemplate('asserter');
const parts = [fill(t.main, { expected })];
if (world) parts.push(fill(t.world, { expecting: world.expecting, delta: world.delta }));
if (actions) parts.push(fill(t.actions, { actions }));
if (substitutions) parts.push(fill(t.substitutions, { substitutions }));
parts.push(t.transcript);
if (walk) parts.push(fill(t.walk, { walk }));
return `${parts.join('\n\n')}\n`;
}

Expand Down
40 changes: 36 additions & 4 deletions tests/prose/lib/record-action.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@ const fs = require('fs');
const path = require('path');

const LOG = '.walk-actions.log';
const WALK = '.walk-transcript.log';
const VIOLATIONS = 'tests/prose/.agent-tool-use.log';
const WORLD = /(^|[\s"'`])(\/[^\s"'`]*\/prose-world-[A-Za-z0-9]+)/;
const MAX_OUTPUT = 400;
Expand Down Expand Up @@ -94,7 +95,7 @@ function responseText(response, limit) {
* neither of which a stop payload states.
*/
function fromTranscript(transcriptPath) {
const blank = { world: null, model: '' };
const blank = { world: null, model: '', turns: [] };
if (!transcriptPath) return blank;
let raw;
try {
Expand All @@ -103,19 +104,49 @@ function fromTranscript(transcriptPath) {
return blank;
}
const models = new Set();
const turns = [];
for (const line of raw.split('\n')) {
let entry;
try {
const model = JSON.parse(line)?.message?.model;
if (model) models.add(model);
} catch { /* a partial line is not worth failing the record over */ }
entry = JSON.parse(line);
} catch { continue; /* a partial line is not worth failing the record over */ }
const message = entry && entry.message;
if (!message) continue;
if (message.model) models.add(message.model);
if (!Array.isArray(message.content)) continue;
for (const block of message.content) {
if (block && block.type === 'text' && block.text.trim()) turns.push(block.text.trim());
}
}
const found = raw.match(WORLD);
return {
world: found ? found[2].replace(/\\+/g, '') : null,
model: [...models].join(',') || '',
turns,
};
}

/**
* The walk as it was actually told, turn by turn.
*
* An agent returns one final message, and a walk runs across dozens of
* turns — so a caller who reads only the return value sees a summary the
* walker wrote after the fact, and every step compressed out of it looks
* like a step never taken. The turns are all in the runtime's transcript
* already; this lifts them into the world beside the action log, where
* the judging is done. Nothing is asked of the walker, which is what
* makes it dependable.
*/
function writeWalk(world, turns) {
if (!turns || !turns.length) return;
try {
fs.writeFileSync(
path.join(world, WALK),
`${turns.join('\n\n---\n\n').split(world).join('.')}\n`,
);
} catch { /* a hook must never break what it observes */ }
}

function main() {
const payload = read();
if (!payload) return;
Expand Down Expand Up @@ -172,6 +203,7 @@ function main() {
);
} else if (stop) {
parts.push(flatten(payload.last_assistant_message, MAX_OUTPUT).split(world).join('.'));
writeWalk(world, traced.turns);
}

try {
Expand Down
16 changes: 14 additions & 2 deletions tests/prose/lib/worlds.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,10 @@ const HASH_FILE = '.recipe-hash';
// Written by the walker's PostToolUse hook (lib/record-action.cjs) —
// observation of the walk, not part of the world it acted on.
const ACTION_LOG = '.walk-actions.log';
// Written by the same hook at SubagentStop: the walk as it was told,
// turn by turn, lifted from the runtime's own transcript. An agent
// returns one final message; the walk happens across dozens of turns.
const WALK_LOG = '.walk-transcript.log';

// --- the recipe harness ---------------------------------------------------

Expand Down Expand Up @@ -104,7 +108,7 @@ function runRecipe(caseId, which) {
function excluded(rel) {
const parts = rel.split(path.sep);
if (parts.includes('.git')) return true;
if (rel === HASH_FILE || rel === ACTION_LOG) return true;
if (rel === HASH_FILE || rel === ACTION_LOG || rel === WALK_LOG) return true;
if (rel === path.join('.workflows', '.knowledge')) return true;
if (rel.startsWith(path.join('.workflows', '.knowledge') + path.sep)) return true;
if (rel.startsWith(path.join('.claude', 'skills') + path.sep)) return true;
Expand Down Expand Up @@ -343,8 +347,16 @@ function readActionLog(worldDir) {
.join('\n');
}

/** The walk as it was told, turn by turn — not the summary it returned. */
function readWalkLog(worldDir) {
const file = path.join(worldDir, WALK_LOG);
if (!fs.existsSync(file)) return null;
return fs.readFileSync(file, 'utf8').trim() || null;
}

module.exports = {
ROOT, ENGINE, KNOWLEDGE, MAINLINES_DIR, WORLD_PREFIX, ACTION_LOG, readActionLog,
ROOT, ENGINE, KNOWLEDGE, MAINLINES_DIR, WORLD_PREFIX,
ACTION_LOG, readActionLog, WALK_LOG, readWalkLog,
runRecipe, collectTree, readSnapshot, snapshotDir, recipeHash, storedHash,
writeSnapshot, verifySnapshot, diffWorld, buildWorld, destroyWorld,
};
17 changes: 12 additions & 5 deletions tests/prose/prompts/asserter.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,9 +46,16 @@ a missing step.

{{substitutions}}

=== transcript ===
=== walk ===

THE WALK, TURN BY TURN — the walker's own account as it was told at the
time, lifted from the runtime's transcript by the harness. It was not
summarised, and the walker had no chance to edit it after the fact.

Read it for what only it can show: which section or arm was entered and
the guard line that selected it, the blocks emitted to the user, the
questions asked and the scripted answers consumed. What a command ran and
returned is settled by the recorded actions above, not here.

{{walk}}

The walk's own account follows under `=== TRANSCRIPT ===`. It is the
walker's narrative: evidence of reasoning — which arm it entered, which
guard it read, what it emitted — and nothing more. Where it disagrees
with the recorded actions about what was done, the recorded actions win.
16 changes: 15 additions & 1 deletion tests/prose/run.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -144,6 +144,7 @@ function cmdAssert(argv) {
const c = getCase(argv[0]);
let world = null;
let actions = null;
let walk = null;
if (c.hasFixtureState) {
const dir = requireWorld(argv, c);
const delta = worlds.diffWorld(c.id, dir, c.worldMode === 'claims');
Expand All @@ -166,14 +167,27 @@ function cmdAssert(argv) {
+ 'registry has reloaded since it changed, and that this project is trusted '
+ '(hasTrustDialogAccepted). Do not judge this run.');
}
walk = worlds.readWalkLog(dir);
// Same stance as the action log: the walk is harness-captured, so its
// absence is a broken hook, not a quiet walk. Judging without it would
// fall back to whatever the walker chose to say at the end — the very
// summary this record exists to replace.
if (!walk) {
die(`no walk log at ${path.join(dir, worlds.WALK_LOG)} — the prose-walker `
+ 'SubagentStop hook did not fire, so there is no turn-by-turn record of '
+ 'the walk.\nCheck the hooks block in .claude/agents/prose-walker.md and '
+ 'that the agent registry has reloaded since it changed. Do not judge this run.');
}
}
const substitutions = c.stubs.length
? c.stubs.map((s) => {
const stub = cases.readStub(s.name);
return `- ${s.name} — fires ${s.trigger}\n ${stub.description.replace(/\n/g, '\n ')}`;
}).join('\n')
: null;
process.stdout.write(prompts.asserterPrompt({ expected: c.assert, world, actions, substitutions }));
process.stdout.write(
prompts.asserterPrompt({ expected: c.assert, world, actions, walk, substitutions }),
);
}

// --- snap / verify --------------------------------------------------------
Expand Down
29 changes: 29 additions & 0 deletions tests/scripts/test-prose-record-action.cjs
Original file line number Diff line number Diff line change
Expand Up @@ -213,6 +213,35 @@ describe('prose recorder — the stop event', () => {
assert.equal(logLines()[0].split('\t')[2], 'claude-sonnet-5');
});

it('lifts the walk out turn by turn, not just the message returned', () => {
const transcript = writeTranscript([
{ message: { model: 'claude-sonnet-5', content: [{ type: 'text', text: 'ENTERED: a.md § Step 1' }] }, cwd: world },
{ message: { content: [{ type: 'tool_use', name: 'Bash' }] } },
{ message: { content: [{ type: 'text', text: 'EMITTED (menu):\n Which feature?' }] } },
{ message: { content: [{ type: 'text', text: 'STOPPED: end of flow' }] } },
]);
fire({
hook_event_name: 'SubagentStop',
agent_type: 'prose-walker',
agent_transcript_path: transcript,
last_assistant_message: 'STOPPED: end of flow',
});
const walk = fs.readFileSync(path.join(world, '.walk-transcript.log'), 'utf8');
assert.ok(walk.includes('ENTERED: a.md § Step 1'), 'an early turn survives');
assert.ok(walk.includes('Which feature?'), 'and what it emitted mid-walk');
assert.ok(walk.includes('STOPPED: end of flow'));
assert.ok(walk.indexOf('ENTERED') < walk.indexOf('Which feature?'), 'in the order they happened');
});

it('writes no walk file when the transcript holds no turns', () => {
fire({
hook_event_name: 'SubagentStop',
agent_type: 'prose-walker',
agent_transcript_path: writeTranscript([{ message: { model: 'x' }, cwd: world }]),
});
assert.equal(fs.existsSync(path.join(world, '.walk-transcript.log')), false);
});

it('stays silent when a stop names no readable transcript', () => {
// A real transcript lives outside the world, so an unreadable one
// leaves nothing anywhere in the payload to scope the record to.
Expand Down