Skip to content

docs(upstream): answer "What a task costs on Opus 5.5" with routing, observability and pointer updates - #5827

Merged
kyle-sexton merged 23 commits into
mainfrom
feat/opus-5-5-task-cost-digest
Oct 2, 2026
Merged

kyle-sexton merged 23 commits into
mainfrom
feat/opus-5-5-task-cost-digest

Conversation

@kyle-sexton

Copy link
Copy Markdown
Contributor

Closes #5763

Summary

This PR applies the claude.dev post What a task costs on Opus 5.5 to the repo. The changes come from a verified digest of the post and a full planning interview. The repo acts on the official docs each item rests on, not on the post itself. docs/upstream/opus-5-5-task-cost.md maps every item in the post to its docs section and records the decision made for it.

Fix

  • Upstream record, docs/upstream/opus-5-5-task-cost.md: maps every item in the post to a docs section. The post's figures appear only in the record and are labelled vendor-reported.
  • Cost-claims rule, .claude/rules/cost-claims.md: guidance links the costs and pricing docs and states no prices or per-task figures.
  • Pointers:
    • audit-instructions, audit-pass and its doctor handoff name /doctor prompt-audit and defer to it where the two overlap. The 12 overlapping criteria rows stay, because the two checks complement each other (Q42).
    • prompt-caching.md and draft-goal-condition link costs#why-usage-climbs-in-a-long-session instead of repeating it.
  • Per-phase Sonnet routing, Q20:
    • Plans gain a per-phase Model column.
    • A phase marked sonnet runs on the new implementation:scoped-implementer (sonnet, effort medium), spawned with an explicit model: sonnet. If the provider has no Sonnet alias, it falls back to implementation:implementer.
    • Unmarked or complex phases stay on implementation:implementer (opus, high).
    • A sync test keeps the two agents' shared worker contract identical.
    • docs/plugin-philosophy.md, the loop-lane README and the loop-lane prompts state the split.
  • Observability: harness-ops:observability compare compares two sessions' tokens by type, split by model and effort, and reconciles them against the cost metric (TDD, session-compare.test.sh).
  • Research gate: outcome-gate row 7 now passes MEDIUM claims that are listed as gaps (Q6).
  • audit-instructions I17-a: points at the live model-config#extended-thinking section instead of listing models (Q32). I17-b is unchanged.
  • Docs queue: adds "Spending your effort" and "Prompt caching is everything" as correlate-only entries (Q34).
  • Instruction-placement: the eval-fixture exclusion moves into the shared discovery code, so the render-index and glob checks agree.
  • No-restatement rule: lines this branch added to skill and agent bodies now follow the links-only rule from feat(playbooks): apply the Sonnet 5.5 prompting guide across the marketplace #5767.

Two decisions differ from the docs or from another session, and both are logged in the deviation record:

  • Unmarked phases stay on Opus, while the Claude Code costs page says Sonnet handles most coding. The platform model guide says most workloads start with Opus 5.5, and this PR follows that.
  • scoped-implementer runs at medium effort, not high. This follows the Sonnet 5.5 effort guidance for well-specified agentic coding. The "Spending Your Effort" session's user adopted this design (its Q59).

Verification

Results on b1218703b, with origin/main merged:

  • check-changelog-parity.sh: --check, --check-bump origin/main, --check-order and --check-preserved origin/main each pass.
  • validate-plugin-contracts.test.sh: 138 passed, 0 failed. It includes the gate that every shipped agent sets model:.
  • agent-contract-sync.test.sh: 5 passed, 0 failed. It checks that implementer and scoped-implementer carry the same contract, skills and tools.
  • session-compare.test.sh: 36 passed, 0 failed.
  • plugins/discovery/scripts/contract.test.sh: passes, including the 20000-byte re-attach slice.
  • audit-instructions scanner tests: instruction-scan 140, emit-findings 184 and finding-ids 49, all passing.
  • check-skill: passes for implementation, harness-config, harness-ops and planning. audit-instructions SKILL.md is 499 lines and audit-pass is 497, under the 500 cap.
  • affected-tests.sh (after the first main merge): every shell suite passes, and the 8 delegated Python suites pass (989 passed, 1 skipped). check-script-contract.test.sh could not run locally because htmlhint is not installed.
  • Docs anchors: every pointer added or changed was fetched on 2026-10-02 and resolves.
  • implement-dispatch eval id 13 covers the three routes, including the explicit model: sonnet. The route has not run live yet, because the installed plugin predates scoped-implementer.

Known gaps that this branch doesn't cause:

  • The glob gate fails on main's .claude/rules/mod-authoring.md, whose plugins/*/types/** glob matches no files.
  • One machine-health Pester test needs Windows.

The plan and the deviation log are on #5763. The removed docs/topics/ slice is still in history at 1dc4ccceb.

Related

🤖 Generated with Claude Code

kyle-sexton and others added 19 commits October 1, 2026 21:10
Add docs/upstream/opus-5-5-task-cost.md: every post item mapped to the
official docs section it rests on, the post's figures labelled
vendor-reported, 12 post-vs-docs conflicts with recheck triggers, and the
stated deviations. Keep CLAUDE_CODE_SUBAGENT_MODEL_FORCE declined in
claude-code.md and note that the plain variable does not move built-in
Plan or Explore. Commit the approved plan and design for #5763.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Add a compare action that reads two sessions from the local OTEL store:
tokens by type (cache writes as their own type) split by model and
effort, and api_request event cost reconciled against the
claude_code.cost.usage metric, flagging an event undercount. A store
with non-delta metrics exits 2 instead of reconciling. cc-otel.sql is
unchanged; effort and temporality are read from raw attributes.

Refs #5763

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lementer

Add a Model column (sonnet | opus | frontier) to the plan template's
per-phase routing table and a scoped-implementer agent (sonnet, medium
effort). implement-dispatch spawns it with an explicit model: sonnet for
sonnet rows, falls back to implementer when a non-Anthropic provider
would resolve the sonnet alias to an older model, and keeps implementer
for unrouted, opus, unknown and frontier rows. Both agents carry the
worker contract inline; a sync test fails when the copies diverge. The
work skill's fix worker and the loop-lane tier roles say the same.

Refs #5763

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gaps

Row 7 now says only accepted claims must be HIGH confidence; a MEDIUM or
LOW claim listed in the Gaps section is not accepted, so an artifact of
honestly listed vendor-only claims passes. The verifier, discipline 11
and the phases context say the same. Four wording trims keep the
effort-ceiling sentence inside the re-attach slice the contract test
asserts.

Refs #5763

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Link the Claude Code prompt-caching page and the costs doc's "Why usage
climbs in a long session" from the prompt-caching chapter and
draft-goal-condition; name /doctor prompt-audit for Claude Code
configuration beside /claude-api prompt-audit in audit-instructions,
audit-pass and the bundled-claude-api reference. Add a path-scoped
cost-claims rule: cost claims point at the costs and pricing docs, and
only docs/upstream records list vendor figures. Instruction-placement
discovery skips eval fixture trees, so the AGENTS.md rules index and the
glob gate no longer pick up fixture rules.

Refs #5763

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…onnet 5.5 effort guidance

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-philosophy and the loop-lane profile

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sts as correlates

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ded-thinking section

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he task-cost record

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…records

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kyle-sexton
kyle-sexton marked this pull request as ready for review October 2, 2026 16:34
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T16:38:56.716350Z 6b66e32 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 42s —— View job


Review complete

Scoped to the diff against origin/main (57 files, +1329/-109) per /review:code-review's criteria: correctness and alignment with CLAUDE.md/repo conventions, high-signal only.

Findings

1. Stale agent-count enumeration in docs/plugin-philosophy.md (Model tiers section)

docs/plugin-philosophy.md#L1379 states:

Four pin effort: medium: planning plan-reviewer …; implementation phase-verifier …; and review ecosystem-specialist and discovery explorer …

This PR adds plugins/implementation/agents/scoped-implementer.md with effort: medium (and model: sonnet), but does not update this sentence. Counting the shipped agents now (grep -n '^effort:' plugins/*/agents/*.md) gives five agents pinned at medium: plan-reviewer, phase-verifier, ecosystem-specialist, explorer, and the new scoped-implementer — not four. The enumeration is a closed, named list (the kind this doc elsewhere treats as load-bearing, e.g. the "Eleven pin effort: high" sentence two lines above, which is still accurate), so a reader relying on it to audit effort pins across the repo will silently miss the new agent. No test or script enforces this count, so it won't be caught except by inspection.

Fix: update the sentence to "Five pin effort: medium" and add implementation scoped-implementer to the named list.

Fix this →

What checked out

  • scoped-implementer.md / implementer.md: the shared `` blocks are byte-identical, and agent-contract-sync.test.sh verifies this plus the `skills:`/`tools:` frontmatter stay in sync, including negative cases that confirm the check can actually fail.
  • implement-dispatch/SKILL.md's new per-phase Model-column routing (sonnet → scoped-implementer with explicit model: sonnet, provider-variable guard against untrusted sonnet aliases on Bedrock/Vertex/Foundry/etc., frontier/security-surface → implementer at the frontier alias, everything else → implementer) is consistent with plan-template.md's routing-table eligibility rules and with the new eval case (id 13), which exercises all three routes plus the "orchestrator's standing model: opus" override case.
  • plugins/instruction-placement/scripts/lib/discover.sh's new IP_EVAL_FIXTURE_ERE exclusion is applied in both ip_discover_rules and ip_discover_nested_instructions, and is exercised by new fixture-tree cases in discover.test.sh, glob-tools.test.sh, and render-index.test.sh (rules/files inside the fixture excluded, same-named files outside it still discovered).
  • session-compare.sh/session-compare.sql: session-id values are validated against ^[A-Za-z0-9._-]+$ before being interpolated into SQL literals (the injection guard the script's own comment documents), temporality/missing-row/non-delta-metric failure paths are distinct and each covered by a test case, and the cost-reconciliation arithmetic (metric_usd - events_usd, the match/events short/events exceed metric classification) matches the SQL's column order.
  • Doc-pointer updates (I17 in criteria.md, prompt-caching.md, draft-goal-condition/SKILL.md, the doctor-handoff.md addition) all carry a pointer + as-of date + recheck trigger per .claude/rules/skill-bodies-state-current-rules.md, and none restate the linked section's content.

No other correctness issues found in the reviewed diff. The author-claimed verification results in the PR description (test pass counts, anchor checks, etc.) are reported as claimed, not independently re-run here.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 41s —— View job


Security review complete

No security issues found.

This PR is almost entirely documentation and agent/skill prompt content (routing tables, pointer updates, changelogs, plan templates) plus two new shell/SQL additions. I reviewed the security-relevant surface specifically:

  • plugins/harness-ops/skills/observability/scripts/session-compare.sh (blob) interpolates session ids into DuckDB SQL string literals, but validates each id against ^[A-Za-z0-9._-]+$ before use, which closes the injection path (quotes, semicolons, etc. are rejected). The store path is escaped via sql_path() (doubles ') before being embedded in a SQL literal too.
  • plugins/harness-ops/skills/observability/otel/session-compare.sql (blob) consumes the ids via DuckDB getvariable() rather than re-interpolating strings, which is the safer pattern.
  • plugins/instruction-placement/scripts/lib/discover.sh fixture-exclusion change uses a fixed ERE constant, no injection surface.
  • New implementation:scoped-implementer agent (the Sonnet routing tier) carries the same tool cage as implementer (enforced by the new agent-contract-sync.test.sh), explicitly excludes security-surface work from downward routing, and the orchestrator-side routing logic in implement-dispatch/SKILL.md preserves the existing rule that a per-invocation model can only route upward, never undercut the frontmatter binding for source-editing work.
  • Instruction-surface check: none of the diff's edits to CLAUDE.md/AGENTS.md/skill or agent bodies delete or weaken a standing guardrail — they replace restated facts with pointers to live docs (per the repo's no-restatement rule), and the removed security-relevant exception lists (e.g. I17's thinking no-effect set) are now read live from the linked docs section rather than dropped.
  • No changes touch GitHub Actions workflows, secrets handling, auth, or input from untrusted external sources.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 6b66e32eea

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/harness-ops/skills/observability/otel/session-compare.sql Outdated
…enter's effort pin

A session with token points but no claude_code.cost.usage point counted its
token points as metric points, so it reconciled as "events exceed metric"
instead of exiting 2. Count cost points only, with a regression case.

plugin-philosophy's medium-effort list now names scoped-implementer (five).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 43s —— View job


Security review complete

No security issues found.

Reviewed the diff against origin/main at cc1b2f2 (57 files, +1339/-111), scoped to the skill's criteria: logic/trust-boundary/injection findings static analysis misses, plus instruction-surface deletions checked against the exception register.

What I looked at:

  • plugins/harness-ops/skills/observability/scripts/session-compare.sh: session ids are validated against ^[A-Za-z0-9._-]+$ before being interpolated into DuckDB SQL string literals (L38), closing the injection path, and the store path is single-quote-escaped via sql_path() (L46-L50) before the same treatment. The action is read-only (no store writes).
  • plugins/harness-ops/skills/observability/otel/session-compare.sql: ids flow through getvariable(), not re-interpolated strings. The latest commit (cc1b2f2) already fixes the one substantive issue Codex flagged here (the cost-point count was previously over all metric types, not just claude_code.cost.usage); the metric CTE now filters WHERE metric_name = 'claude_code.cost.usage' before counting — that was a correctness/report-integrity bug, not an exploitable security path, and it's resolved.
  • New implementation:scoped-implementer agent (the Sonnet routing tier, plugins/implementation/agents/scoped-implementer.md): carries the identical tool cage as implementer (enforced by the new agent-contract-sync.test.sh, which also verifies the check can fail via four perturbation cases). Routing text and the new eval (implement-dispatch/evals/evals.json id 13) both confirm security-surface work is excluded from downward routing and stays on implementer at the frontier alias regardless of what a plan's routing table says — the per-invocation model parameter can only route a phase upward, never undercut the frontmatter binding for source-editing work.
  • Instruction-surface check: the diff narrows/replaces several restated passages with pointers (I17 in criteria.md, prompt-caching.md, draft-goal-condition/SKILL.md, loop-lane docs, the capability-tier table) per the repo's no-restatement rule. None of these removes an enforced guardrail — the "every security-surface work class, always → frontier" rule in docs/conventions/loop-lane/README.md's capability-tier table is preserved unchanged, and the stricter sonnet-eligibility list in plan-template.md explicitly excludes security-surface work. I checked docs/conventions/instruction-exception-register/README.md, which is present in this checkout, and found no protected-class rule weakened by this diff.
  • plugins/instruction-placement/scripts/lib/discover.sh: the new eval-fixture exclusion uses a fixed ERE constant shared across both discovery functions and the render-index/glob checks; no injection surface (no untrusted input reaches it).
  • No changes touch GitHub Actions workflows, secrets handling, auth, or input from untrusted external sources.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 10s —— View job


Review complete

Ran /review:code-review against origin/main...cc1b2f2 (57 files, +1339/-111).

Findings

None. No inline comments posted.

Both issues raised by the prior reviews on commit 6b66e32 are fixed in the latest commit cc1b2f2 ("fix: count only cost points in session-compare and list scoped-implementer's effort pin"):

  • Codex's P2 on session-compare.sql: the metric CTE now filters to claude_code.cost.usage rows before count(*), so a session with token points but no cost point correctly fails the missing-data check instead of reconciling as events exceed metric. A regression case (s-t) was added to session-compare.test.sh.
  • Claude's finding on docs/plugin-philosophy.md: the medium-effort enumeration now says "Five pin effort: medium" and names implementation scoped-implementer.

I independently re-checked the routing logic this PR introduces, since it's the highest-risk surface:

  • implement-dispatch/SKILL.md's per-phase Model-column dispatch (sonnet → scoped-implementer with explicit model: sonnet, provider-alias guard, frontier/security-surface → implementer at the frontier alias, else implementer) is consistent with plan-template.md's routing-table rules and observability/SKILL.md's new compare action wiring.
  • agent-contract-sync.test.sh correctly diffs the `` blocks and skills:/`tools:` frontmatter between `implementer.md` and `scoped-implementer.md`, and its four self-test perturbations (edited contract line, dropped end marker, tools drift, skills drift) each exercise a real failure path.
  • instruction-placement/scripts/lib/discover.sh's new IP_EVAL_FIXTURE_ERE is declared once at module scope and consumed by both ip_discover_rules and ip_discover_nested_instructions; the forward reference inside ip_discover_rules (defined above the variable) is safe since the assignment runs at source time before either function is called.
  • session-compare.sh's awk reconciliation (column indices, events short/events exceed metric/match classification, the $3 == 0 missing-data check) matches the SQL's output column order after the fix.

I did not re-run the test suites or check-skill/linting locally — this sandbox denies Bash execution of scripts (only trivial commands like pwd/git diff were permitted; bash <script> and even git fetch required approval that wasn't available). The verification results in the PR description are author-claimed and not independently re-run here. Local origin/main was already current (a287fa4), so the diff above is accurate.

kyle-sexton added a commit that referenced this pull request Oct 2, 2026
…y turn cost (#5834)

No related issue: follow-ups found while landing #5827. The two findings
that need a design decision are filed as #5832 and #5833.

## Summary

- `harness-config:audit-instructions` said that `/doctor prompt-audit`'s
write posture was undocumented. The upstream memory page now documents
it.
- `fleet:reach`'s relay probe table stated a per-turn dollar figure. It
now links the official cost-tracking docs instead.

## Fix

- `plugins/harness-config/skills/audit-instructions/SKILL.md`, `doctor`
Boundary section:
- The mutation gate states our decision: applying the proposed edits is
the person's call, and the skill never chains into `/doctor`.
- The availability line notes that the subcommand needs the bundled
`claude-api` skill to be on. SKILL.md stays at 499 lines.
- `reference/native-doctor.md`: the "undocumented" row is replaced by a
links-only record. It points to
<https://code.claude.com/docs/en/memory#audit-your-instruction-files>,
carries the as-of date 2026-10-02, and gives a recheck trigger.
- `plugins/fleet/skills/reach/reference/relay.md`: the Cost row keeps
what the cost went to (SessionStart hooks and context loading) and links
<https://code.claude.com/docs/en/costs#track-your-costs>.
- Version bumps with CHANGELOG entries: harness-config to 1.3.2 (#5827
holds 1.3.1) and fleet to 0.2.2.

## Verification

- `scripts/check-changelog-parity.sh --check`, `--check-bump
origin/main`, `--check-order` and `--check-preserved origin/main` all
pass.
- `scripts/check-changed-skills.sh origin/main`: 2 skills, 0 errors. The
one warning is pre-existing, in the untouched `reference/criteria.md`.
- markdownlint-cli2 reports 0 issues on the 5 touched markdown files.
`typos` is clean.
- `run-plugin-tests.sh` passes on the 6 audit-instructions suites plus
`check-doc-citations` and `check-structure`.

## Related

- #5827: the session that found these.
- #5832: the mod-authoring rule glob matches no file.
- #5833: the discovery research gate passes an artifact with zero
accepted claims.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Renumber this branch's plugin versions above main's: harness-config 1.3.3,
harness-ops 2.3.0, instruction-placement 0.19.1, knowledge 0.15.2,
planning 0.61.0, work-items 0.44.2. This branch's CHANGELOG entry stays first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 54s —— View job


Code review complete — no new findings

Skip-gate hit: this head already has a successful review from this lane that still applies.

8982b2b is a clean two-parent merge (cc1b2f2 + 0059cf3, no conflicts) of origin/main into this branch. I confirmed the PR's own diff is unchanged since the last review at cc1b2f2:

  • git diff origin/main..HEAD now shows 61 files (vs. the 57 in this PR's own change set) — the 4 extra (plugins/miro/...) are origin/main's own newer commit (73015ad, not yet merged into this branch) and are not part of this PR.
  • For every file this PR actually touches that also appears in git diff cc1b2f2..8982b2b (e.g. docs/plugin-philosophy.md, plugins/harness-config/skills/audit-instructions/SKILL.md, plugins/implementation/README.md), the content landing on this branch is main's own concurrent edits merging in cleanly — not a change to this PR's authored content. I diffed origin/main..HEAD directly for these files and confirmed they still carry exactly this PR's intended edits (e.g. docs/plugin-philosophy.md still reads "Five pin effort: medium" and lists implementation scoped-implementer, matching the fix already applied in cc1b2f2).

That cc1b2f2 review (and the one before it) already covered this PR's full diff and came back clean — both prior findings (the session-compare.sql cost-point count and the stale "Four pin effort: medium" enumeration) were fixed in that commit, and no PR-authored content has changed since. No new inline comments to post.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 47s —— View job


Security review complete

No security issues found.

HEAD 8982b2b is a merge of origin/main into this branch. I diffed cc1b2f2 (the commit the two prior security-review passes on this PR already covered in full, both "no security issues") against 8982b2b and confirmed the PR's own content is unchanged — the only delta in any file this PR touches is one unrelated convention-registry row in docs/plugin-philosophy.md, pulled in from main (plugin-option-naming), not authored by this PR. The rest of the merge brought in main-only files (e.g. .claude-plugin/marketplace.json version bumps) that this PR does not otherwise touch.

Re-verified the security-relevant surface still holds at this HEAD:

No changes touch GitHub Actions workflows, secrets handling, auth, or untrusted external input.

Main took harness-ops 2.3.0, so this branch's entry moves to 2.4.0 and stays first.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 45s —— View job


Security review complete

No security issues found.

Skip-gate hit: this head already has a successful security review that still applies.

1c328f2 is another merge of origin/main into this branch (on top of 8982b2b, which the two prior security-review passes on this PR already covered in full as clean). I diffed 8982b2b..1c328f2 directly:

plugins/harness-ops/.claude-plugin/plugin.json     |   2 +-
plugins/harness-ops/CHANGELOG.md                   |  46 +++-
.../skills/inventory/scripts/inventory.py          |  49 +++-
.../skills/inventory/scripts/js/parser_helper.cjs  | 255 ++++++++++++++++++++-
.../skills/inventory/scripts/parser_reader.py      |  62 +++--
.../skills/inventory/scripts/test_parser_reader.py | 132 +++++++++++
.../inventory/scripts/test_reader_findings.py      | 166 ++++++++++++--
plugins/miro/.claude-plugin/plugin.json            |   2 +-
plugins/miro/CHANGELOG.md                          |  10 +
plugins/miro/server/src/launcher.test.ts           |   9 +
plugins/miro/server/src/launcher.ts                |   2 +-

All 11 files are either entirely outside this PR's own change set (harness-ops/skills/inventory/**, all of plugins/miro/** — these belong to main's own commit 496e995, PR #5828's --reader=parser work, and are not authored by this PR) or, for the two files this PR does touch (plugins/harness-ops/.claude-plugin/plugin.json, plugins/harness-ops/CHANGELOG.md), the delta is only a version renumber (2.3.0 → 2.4.0) to make room for main's concurrently-landed 2.3.0 entry — this PR's own CHANGELOG entry is unchanged in content, just shifted down and re-numbered.

Re-confirmed the security-relevant surface from the prior passes is untouched at this HEAD (git diff 8982b2b..1c328f2 shows no diff on any of these):

  • plugins/harness-ops/skills/observability/scripts/session-compare.sh: session ids still validated against ^[A-Za-z0-9._-]+$ before SQL interpolation; sql_path() (L46) still escapes the store path.
  • plugins/harness-ops/skills/observability/otel/session-compare.sql: the cost-point count still filters to metric_name = 'claude_code.cost.usage' (the fix that resolved Codex's P2 is intact).
  • plugins/implementation/agents/scoped-implementer.md: still carries the identical tool cage as implementer.md, enforced by agent-contract-sync.test.sh.
  • .claude/rules/cost-claims.md content at this HEAD is byte-identical to the version reviewed at 8982b2b.
  • Instruction-surface edits (I17 pointer, prompt-caching.md, draft-goal-condition, loop-lane docs) remain pointer-only replacements; no enforced guardrail is dropped, consistent with the prior passes. docs/conventions/instruction-exception-register/README.md is present in this checkout and was consulted.

No changes touch GitHub Actions workflows, secrets handling, auth, or untrusted external input.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 19s —— View job


Code review complete — skip gate hit, no new findings

Skip gate: this head already has a successful review from this lane that still applies.

1c328f2 is two further merges of origin/main on top of 8982b2b, which this lane already
reviewed in full (clean, both earlier findings fixed). I diffed every file this PR actually
touches between 8982b2b and 1c328f2 and confirmed none of this PR's own authored content
changed:

  • The only deltas are in plugins/harness-ops/.claude-plugin/plugin.json (version bump
    2.3.0 → 2.4.0) and plugins/harness-ops/CHANGELOG.md (a new ## [2.3.0] section describing
    the unrelated inventory.py --reader=parser writes/mutations work) — both are origin/main's
    own concurrent release content from feat(harness-ops): take the inventory's writes and mutations from the parser under --reader=parser #5828, stacked above/below this PR's entries, not edits to
    lines this PR authored.
  • Every other file this PR touches (docs/plugin-philosophy.md,
    plugins/harness-config/skills/audit-instructions/SKILL.md, the implementation/planning/
    knowledge/instruction-placement/work-items/playbooks CHANGELOGs, etc.) is byte-identical
    between 8982b2b and 1c328f2.

Re-confirmed the two fixes from the prior review passes (at cc1b2f2) are still intact at
1c328f2:

No inline comments posted — there is no new or changed content on this head to anchor findings
to. Note: this sandbox's working tree had .claude/rules/cost-claims.md deleted on disk prior to
my session starting (pre-existing local artifact, not a commit on this branch); I did not have
approval to run the git checkout -- to restore it, so I reviewed that file's committed content
via git show instead — it is unaffected and unchanged in the PR's actual history.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 49s —— View job


Code review complete — skip gate hit, no new findings

Skip gate: this head already has a successful review from this lane that still applies.

1571d54 merges origin/main (bringing in main's 223f37c, "retire the provenance shim, ban rename shims, and drop stale run output") on top of 1c328f2, which this lane already reviewed in full and found clean (both earlier findings — the session-compare.sql cost-point count and the stale "Four pin effort: medium" enumeration — were fixed back at cc1b2f2).

I diffed 1c328f2..1571d54 directly, scoped to every file/directory this PR touches (AGENTS.md, docs/conventions/loop-lane/README.md, docs/upstream/claude-code.md, docs/upstream/opus-5-5-task-cost.md, plugins/discovery, plugins/harness-config, plugins/harness-ops, plugins/implementation, plugins/instruction-placement, plugins/knowledge, plugins/planning, plugins/playbooks, plugins/work-items, prompts/loops, .claude/rules/cost-claims.md): zero lines of diff. The 34 files touched by 223f37c (plugins/provenance/**, plugins/attribution/**, docs/hook-migration-audit.md, docs/setup-contract-campaign-follow-ups.md, etc.) are entirely outside this PR's own change set.

Re-confirmed the two previously-fixed items are still intact at this HEAD:

No inline comments posted — there is no new or changed content on this head to anchor findings to.

Note: this sandbox's working tree has .claude/rules/cost-claims.md deleted on disk (a pre-existing local artifact from before my session started, not a commit on this branch). I reviewed its committed content via git show 1571d54:.claude/rules/cost-claims.md instead — unaffected and unchanged from the version prior passes reviewed.

@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 51s —— View job


Security review complete

No security issues found.

Skip-gate hit: this head already has a successful security review that still applies.

1571d54 is a further merge of origin/main on top of 1c328f2, which this lane already reviewed in full (clean — both earlier findings from Codex/code-review were fixed, and no security issues were found across three prior passes on this branch). I diffed 1c328f2..1571d54 directly:

 .claude-plugin/marketplace.json                     |  11 -
 .claude/ai-slop.json                                |   5 +-
 .claude/settings.json                               |   1 -
 docs/architecture/landscape.json                    |   4 +-
 docs/catalog.md                                     |   1 -
 docs/hook-migration-audit.md                        | 156 ---
 docs/migration-playbook.md                          |  14 +-
 docs/setup-contract-campaign-follow-ups.md          |  41 -
 plugins/attribution/.claude-plugin/plugin.json      |   2 +-
 plugins/attribution/CHANGELOG.md                    |  11 +
 plugins/attribution/README.md                       |   6 -
 plugins/attribution/skills/audit/scripts/lib.sh     |  19 +-
 plugins/attribution/skills/audit/scripts/list-corpus.test.sh | 34 -
 plugins/attribution/skills/setup/SKILL.md           |  13 +-
 plugins/provenance/**                               | (removed entirely)
 scripts/cheatsheet-config.mjs                       |   1 -
 scripts/em-dash-purged-paths.txt                    |   1 -
 scripts/skill-leaf-name-registry.txt                |   6 +-
 (plus a removed research/evidence/decision-log directory tree)

Every one of these files belongs to origin/main's own commit 223f37c ("chore!: retire the provenance shim, ban rename shims, and drop stale run output", #5790), merged into this branch — not to this PR's own change set. None of the 57 files this PR actually touches (listed in the PR's file list: plugins/harness-ops/skills/observability/**, plugins/implementation/agents/scoped-implementer.md, docs/upstream/opus-5-5-task-cost.md, .claude/rules/cost-claims.md, etc.) appear in this diff, so this PR's authored content is unchanged since the clean review at 1c328f2.

Re-confirmed the security-relevant surface from the prior passes is still intact at this HEAD:

Note: this sandbox's working tree shows .claude/rules/cost-claims.md as locally deleted (git status → D .claude/rules/cost-claims.md), matching a prior review's observation — this is a pre-existing local artifact in this checkout, not a commit on this branch. git show 1571d54:.claude/rules/cost-claims.md confirms the file is intact and unchanged in the branch's actual history.

No changes touch GitHub Actions workflows, secrets handling, auth, or untrusted external input.

@kyle-sexton
kyle-sexton merged commit 9b3a077 into main Oct 2, 2026
20 checks passed
@kyle-sexton
kyle-sexton deleted the feat/opus-5-5-task-cost-digest branch October 2, 2026 17:44
kyle-sexton added a commit that referenced this pull request Oct 3, 2026
…5895)

No related issue: adopts the claude.dev eval-design and hillclimbing
post across six plugins; follow-ups are #5878 and #5879.

## Summary

Applies what this repository takes from the eval-design and hillclimbing
post, with every eval score taken only from a run the new validity gate
calls VALID. One commit per plugin:

- **evals 0.4.0**: noise report, run-validity gate, judge calibration,
offline grader tests, routing by repository kind, a 31-case suite with
labelled samples for every `llm` grader, and six settings.
- **skill-quality 0.26.0**: hard-case reasons in the evals schema and
trigger-noise reporting.
- **discovery 0.26.0**: a first-party content claim with one publisher
passes the research gate as `HIGH (single source)`, with the flag kept
visible.
- **harness-config 1.3.4**: the `audit-instructions` records for the
bundled `claude-api` and `doctor` skills are links-only.
- **context-budget 0.7.4**: skill pruning routes to `/skill-doctor`,
with its boundary recorded.
- **playbooks 0.17.2**: skill-authoring says how to carry an eval
failure back into a skill, and allows an old-patterns names table.

## Fix

- `plugin-eval` now runs three scripts before a score counts.
`run-validity.py` prints VALID or INVALID from the result and the kept
traces. `noise-report.py` prints intervals and a paired
with-versus-without delta. `calibrate-judge.py` scores each `llm` judge
against labelled samples and fails it below 90% agreement.
- The suite grew from 3 to 30 cases. Each hard case says why it is hard,
and four controls must not invoke an evals skill.
- Records that point at Anthropic pages state our decision, then a
pointer, an as-of date and a recheck trigger. They no longer restate the
upstream text.
- `.claude/rules/cost-claims.md` (added by #5827) gains a second
exception. A skill that prices its own runs before spending may state
run costs it measured on its own suite, each dated, with the version,
the setup and a recheck trigger. `plugin-eval`'s sizing anchors are that
case. **Reviewer: please confirm this fits the rule's intent.** Its
author's session has closed.
- The evals options follow the plugin option naming convention:
noun-phrase boolean titles, and `options` pickers for the three
fixed-set strings.

## Verification

- Before the rebase onto main (job 11, 72507f118): **VALID**. With the
plugin 0.99 (95% CI 0.97 to 1.00), without 0.31 (0.16 to 0.47), delta
**+0.68 (+0.52 to +0.83)**. The three remaining misses are judge errors
on calibrated graders (follow-up #5879).
- Every `llm` grader is at or above 90% agreement with its labels.
- The earlier tip (276c5fd), full suite, 3 runs, sonnet judge:
**VALID**. With the plugin 1.00 (95% CI 0.99 to 1.00), without 0.29
(0.14 to 0.44), delta **+0.71 (+0.56 to +0.85)**; 30 of 31 cases at
1.00. The one miss is the recurring judge error in #5879. An earlier run
the same day was INVALID: every no-plugin run in one case hit a
sandbox-probe denial, and one with-plugin run did not fire the skill.
Neither recurred.
- Compressing `plugin-eval/SKILL.md` for cost (two `/claude-api
hillclimb` iterations) was **not adopted**. The second, run beside that
same-day control, was VALID and cut agent cost per request 11.1%. It
dropped two explanatory clauses that graders depend on, though, and two
cases regressed (with-arm 0.97 against 0.99). A version with both
clauses restored is follow-up #5956.
- This tip (50c2769), after the respelling and review fixes below: the
13 `llm` graders whose samples were respelled recalibrate at 100%
agreement. Full suite, 3 runs, sonnet judge: **VALID**. With the plugin
0.98 (95% CI 0.94 to 1.00), without 0.28 (0.13 to 0.44), delta **+0.69
(+0.54 to +0.84)**; 30 of 31 cases at 1.00. The one miss is
`measurable-criterion` (0.33): in two of three runs the reply set a
target without stating its basis. These commits changed only that case's
calibration samples, which no run reads, so the miss is model variance,
not a regression. Later merges from main change nothing under
`plugins/evals`.
- Wording changes after job 11: option titles and descriptions, three
links-only rewordings, one caveat clause, two line merges in
`plugin-eval/SKILL.md`. None changes a rubric, a sample or a routing
rule.
- A security review of this diff found two issues in scripts it adds,
both fixed with tests in 6ca9f6a. `calibrate-judge.py build` could
write outside `--out` through a crafted `case.yaml` grader name, and
`render-review.py` Markdown cells left link and image syntax live.
Neither script runs inside an eval, so the eval result above still
applies.
- After the ready flip, CI's typos (en-us) and machine-specific-paths
checks failed on this PR's eval files. 8c792eb respells them to
American English, renames the `expected-behaviour` grader to
`expected-behavior`, and swaps `/home/dev/` example paths for `/srv/`.
No label, expected value or rubric criterion changed. A fresh-context
verifier confirmed that.
- Review fixes. 8055109: `calibrate-judge.py` no longer counts a run
it could not show reproduced the labeled sample toward agreement (Codex
P1, security review), and the `md_cell` docstring now states the
bare-URL contract. 1479903: `measure-invocation.sh compare` builds its
interval from paired per-probe deltas (Codex P2). Both have tests.
- Local checks on the merged head (after merging main): changelog parity
(check, order, bump), `validate-plugin-contracts.mjs`,
`check-changed-skills.sh origin/main`, `render-index.sh check`,
`overlap.py generate --check`, `check-spoke-plugin-root.sh --check`,
`validate-cases.py` (0 FAIL), the affected shell suites, and every
Python test under `plugins/evals`. The only failures are outside this
diff: htmlhint is not installed locally, and two hook timing tests
(guardrails, pr-body-linkage-gate) went over their 8000 ms budget under
load.

## Related

- #5878: the eval sandbox reads only a skill's hub `SKILL.md`.
- #5879: calibration gaps for the two judges that missed in job 11, and
the must-fail samples the agent will not reproduce.
- #5956: the v2 `plugin-eval/SKILL.md` compression with the two clauses
restored.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Oct 3, 2026
…review, session-flow and telemetry (#5920)

No related issue: first of two PRs carrying this repo's share of the
decisions from digesting the "Spending your effort" post
(https://claude.dev/blog/spending-your-effort); the second PR follows
once this one merges.

## Summary

Make effort an explicit, recorded choice across lanes, review fan-out,
session-flow advice and telemetry, and add a check that flags when the
upstream effort guidance changes.

- **harness-ops 3.0.0 (breaking):** the lane launcher refuses a lane
whose config names no `effort`, always passes `--effort`, and warns once
per run when `CLAUDE_CODE_EFFORT_LEVEL` may override lane levels and
agent pins. Migration: add `lanes[].effort` to every lane in
`lanes.json`. The session event log records the effort level (`n/a` on
events that never carry one, `unset` when missing) and an allowlist of
documented non-content hook input fields, read only from the payload's
top level.
- **harness-config:** new `check-effort-pins.sh` drift check (audit
`effort-pins` scope, also run by the changelog skill) hashes
model-config's effort sections against a committed baseline and flags
pins in agents, skills, lane configs and Workflow scripts without
editing them. unhobble records the session effort.
- **knowledge:** `check-html-rows.py`, a standing gate for quotes taken
from a page's HTML, with a synthetic negative-control suite.
- **review:** fanout's generic slice and extract agents run at an
explicit effort; named agents keep their pins; the report shows each
leaf as `label@level` (the level requested).
- **session-flow:** workflow gives per-stage effort advice by reading
model-config's live table and never sets effort; continue-in-background
picks the resumed task's level from that table and passes `--effort`.
- **songwriting:** object-writer pinned at `medium`, provisional until
an eval compares levels.
- **work-items, source-control:** the work-loop and babysit-loop state
blocks record the effort a lane actually ran at.

## Fix

One commit per area, a links-only conformance pass (upstream specifics
are pointed at, never restated, per
`.claude/rules/skill-bodies-state-current-rules.md`), and the version
bumps and CHANGELOG entries as the last commit.

## Verification

- Each phase was checked by a fresh-context phase verifier against the
plan's acceptance criteria (all PASS after fixes).
- Suites: `check-effort-pins.test.sh` (67 checks),
`lane-launcher.test.sh` (232 cases), `session-event-log.test.sh` (139),
`check-html-rows.test.sh` (18), plus the reader suites for the event
log; spawn ratchets unchanged.
- `scripts/validate-plugins.sh`, `scripts/check-changelog-parity.sh
--check` and `--check-bump origin/main`, and
`scripts/check-changed-skills.sh origin/main` pass.
- `scripts/run-plugin-tests.sh`: every suite for the changed plugins
passes. Two failures outside this branch's files:
`plugins/code-tidying/scripts/evals-fixtures.test.sh` (fails the same
way on main, missing tooling) and a wall-clock budget in
`plugins/guardrails/hooks/block-root-delete-target.test.sh` that overran
under a machine load average of ~30; this branch changes nothing under
`plugins/guardrails`.

## Related

- #5767 and
#5827
(merged peer work this builds on)
- #5682
(replay-sweep scope approved alongside this work)

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
kyle-sexton added a commit that referenced this pull request Oct 3, 2026
… OTEL columns, with the upstream record (#5992)

No related issue: second of two PRs carrying this repo's share of the
decisions from digesting the "Spending your effort" post
(https://claude.dev/blog/spending-your-effort); the first was
[#5920](#5920).

## Summary

Finish the effort work that had to wait for #5920 and its two sibling
PRs
([#5767](#5767),
[#5827](#5827))
to merge, and record every decision the post drove.

- **planning, session-flow:** the interview's session-config
recommendation names an implement level and a verify level, each from
model-config's effort table, neither below `medium`, and no level when
the page cannot be read. Orchestrate's imperative 7 keeps code-changing,
verifying and likely-edge-case work off the lower effort.
- **harness-ops:** an opt-in `session_event_log_content` (default off)
adds the payload's top-level content strings to event-log rows, cut at
64 KB with `<key>_truncated` or a row-level `content_truncated` marker.
Rows over 4000 bytes now append under an exclusive-create lock: bash
writes long lines in 4 KB pieces, so concurrent long appends interleaved
(15 to 18 of 33 parallel 64 KB appends corrupt before, 0 after). Short
rows keep the single unlocked write. The OTEL store exposes every
attribute the monitoring page documents as a typed column, effort
included; cold macros union a typed stub with older Parquet by name. A
new `CC_OTEL_COLD_KEEP_CONTENT` (default keep, `=0` scrubs) covers
content columns beyond user prompts, which
`CC_OTEL_COLD_KEEP_USER_PROMPTS` still governs; `session-compare`
reports a missing level as `unset`.
- **harness-config:** the audit checklist runs the effort-pin drift
check.
- **implementation:** the upward-only binding now covers Workflow
`effort` and `model` for the implementer and phase verifier. Pins stay
as #5885 set them (implementer `medium`, phase-verifier `high`).
- **knowledge:** map-corpus's effort gotcha points at the Workflow probe
record in `docs/plugin-philosophy.md`.
- **playbooks:** prompt-caching's effort note points at Claude Code's
own cache page.
- **docs:** `plugin-philosophy.md`'s low-effort sweep rule also excludes
likely-edge-case work;
`docs/upstream/claude-dev-spending-your-effort-blog.md` records each
decision and where it landed.

## Fix

One commit per area, then version bumps (harness-config 1.6.0,
harness-ops 3.3.0, implementation 0.21.0, knowledge 0.18.0, planning
0.63.0, playbooks 0.17.5, session-flow 0.47.0, each with a CHANGELOG
entry), with the upstream record last. No breaking change.

Deliberate departures from the plan: the implementer keeps `effort:
medium` as
[#5885](#5885)
set it (the plan predates it); the event-log lock and
`content_truncated` marker were added after a fresh verifier found
interleaved long rows and a silently dropped cut field.

## Verification

- Two fresh-context phase verifiers checked every phase against the
plan's acceptance criteria on the committed diff: 23/27 then 22/22, with
the one failure fixed and two criteria superseded (#5885) or stale
(three newer workflow-only agents leave effort unpinned on purpose).
- `scripts/run-plugin-tests.sh --jobs 4`: all suites pass except
`plugins/code-tidying/scripts/evals-fixtures.test.sh`, which fails the
same way on main (missing tooling); this PR does not touch code-tidying.
- Suites touched here: `session-event-log.test.sh` 168/168,
`cc-otel.test.sh` 22/22, `prune-otel-store.test.sh` 355 passed,
`session-compare.test.sh` 38/38, `claude-observability.test.sh` 73/73,
`check-effort-pins.test.sh` 69/69, `interview-defenses.test.sh` 172/172,
`agent-contract-sync.test.sh` 5/5; spawn ratchets unchanged.
- `measure-hook-log-budget.sh --samples 5`: 0 corrupt lines at 4, 16 and
64 KB on two runs (15 to 18 of 33 at 64 KB before the lock).
- `scripts/validate-plugins.sh` and `scripts/check-changelog-parity.sh
--check-bump origin/main` pass.

## Related

-
[#5920](#5920):
first PR of this pair.
-
[#5767](#5767),
[#5827](#5827):
sibling PRs this one waited on.
-
[#5885](#5885):
set the implementation agents' pins this PR keeps.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

docs(upstream): answer 'What a task costs on Opus 5.5'

1 participant