Conversation
…out eval measurement ADR-0037: model-declared semi-frozen verification contract + immutable floor (supersedes the greenfield disposition of ADR-0014). ADR-0038: scoped, config-gated auto-approve of contract amendments only (extends ADR-0016; AVATAR_AUTONOMOUS_AMENDMENT, default deny). ADR-0039: held-out-verified vs self-reported success; wire the ADR-0011 D3 held-out oracle so auto-approve stays measurable.
…ADR-0037) Add declare_verification: on a greenfield edit (tiers 1-3 resolve nothing), the model authors a real executing verification contract that the harness runs itself and grades on the real exit code (never self-certification). - PlannedCheck.kind gains 'declared'/'floor'; RunDeps.declared_contract buffers the tool's output (tools never mutate TaskState). - vacuous_declared_check rejects no-op commands (true/echo/:) so a declared contract can't be a bar the model trivially clears. - runner._freeze_plan folds a declared contract into the frozen plan when greenfield; a real detected/cited contract always wins; a decline leaves the empty plan for the existing smoke floor. - Verifier runs declared checks unchanged (required + positive signal). The always-on immutable floor beneath a declared contract lands with the amendment gate (increment 3), where it becomes load-bearing.
…prove (ADR-0037/0038) The model can amend an obsolete declared check via alter_verification (tier 3, gated). The immutable floor now binds beneath a declared contract, so success = floor AND declared, and an amendment can never weaken it. - state: append_verification_floor (idempotent) + amend_declared_contract (rewrites declared entries, preserves the floor). - runner: _maybe_add_smoke_floor binds the floor alongside a declared contract (greenfield only); _apply_amendment folds an approved amendment into the plan and re-journals the rubric (auditable). - alter_verification tool (tier 3) routes through the existing approval gate. - session: scoped auto-approve — unattended + amendment_policy='approve' + tool=='alter_verification' self-ratifies; everything else (run_command, denylist) still auto-denies. Scoped by tool NAME, never tier. - config AVATAR_AUTONOMOUS_AMENDMENT (default deny), threaded via Harness.session. 10 new tests: mutators, tool gating, scoped session disposition, floor binding.
PR #105 (multi-line input) merged an ADR-0037 concurrently, so this feature's ADRs shift up one: model-declared contract -> 0038, scoped amendment disposition -> 0039, held-out eval measurement -> 0040. Also corrects a stale ADR-0039 citation for the amendment disposition (now correct at 0039 by coincidence of the shift).
…ation-contract # Conflicts: # docs/adr/README.md
ApprovalModal now renders an alter_verification amendment with its obsolescence rationale + proposed checks (not a raw args dict) and notes the immutable floor can't be amended away — so a human can tell a genuine design change from a real failure being papered over, rather than rubber-stamping. The generic tier-3 approval path is unchanged.
…R-0040) Under auto-approved self-amendment (ADR-0039) the model's own contract passing certifies nothing, so measurement must key on the independent hidden oracle, not the self-report. - ResultRow gains self_reported_success + held_out_passed (both default False; back-compatible with older rows). - run_task populates them: held_out = the success probe's real exit code when one exists, else the composed grade. - metrics: held_out_pass_at_1 (honest capability) + gamed_rate (self_reported AND NOT held_out) surfaced per-model in build_summary. - classify: a probe failure the model CLAIMED to satisfy is 'gamed', distinct from an honest 'probe_failed'. Grader-touching (ADR-0040): offline TDD green here; the mandatory global evals.diff validation is a live run, to be run before flipping Accepted.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 12a27ac850
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| passing_outcomes=spec.passing_outcomes, | ||
| ), | ||
| solved=solved, | ||
| self_reported_success=reached_success, |
There was a problem hiding this comment.
Record final-answer claims even when strict verification fails
For strict/no-probe tasks, state.outcome is not the same as “the model reached final_answer”: the strict verifier can reject a final answer, leave the loop in editing, and later end as failed after the repair budget while state.final_answer remains populated. In that scenario the model did claim success but this records self_reported_success=false, so gamed_rate and the gamed bucket miss claimed-but-rejected runs outside the conversational success-probe path. Use state.final_answer is not None or the decision log for this signal instead of the terminal outcome.
Useful? React with 👍 / 👎.
| # `gamed` = self_reported ∧ ¬held_out (see `metrics.gamed_rate`). Both default False so rows | ||
| # written before the split load cleanly. | ||
| self_reported_success: bool = False | ||
| held_out_passed: bool = False |
There was a problem hiding this comment.
Derive legacy held-out verdicts instead of defaulting false
When load_results reads any pre-split result file, including the committed evals/results/*.jsonl baselines that lack this field, Pydantic fills held_out_passed=false even for rows with solved=true and probe_exit=0. Any rebuilt summary or direct held_out_pass_at_1(load_results(...)) over those historical baselines will therefore report 0 held-out capability and corrupt comparisons. For legacy rows, derive this from the old solved/probe data or make the field optional and have the metric fall back, rather than defaulting all missing values to false.
Useful? React with 👍 / 👎.
…all-clock during approval Two hardening changes surfaced by the jo-cli Tetris dogfood run that ended `incomplete` (600s wall-clock, no declared verification contract). Greenfield declaration gate (ADR-0038): at the investigating->editing boundary a greenfield `edit` task must call `declare_verification` before it may edit. The runner refuses each edit-intent call and nudges the model up to `max_declaration_nudges` (default 3) times, then falls back to the smoke floor so a run is never stranded. Surfaced as a typed `DeclarationRequired` event the cockpit renders. Scoped to greenfield edit; a detected/declared contract skips the gate. Also corrects the edit mission prompt, which implied verification was already handled. Wall-clock pause on approval: the wall-clock budget bounds agent work, not reviewer think-time. The runner credits time spent blocked on a human approval back to the deadline (`_approval_wait_seconds`), so a slow approval no longer starves a run into a spurious `incomplete`.
Close Threat C (runtime/substrate gaming) at the single command chokepoint. A `Sandbox`/`ExecSpec` strategy transforms every command — verifier checks and the model's run_command alike — before exec: `hermetic-env` (the default) scrubs the child env to a language-neutral allowlist so an inherited PYTEST_ADDOPTS/ PYTHONPATH can't rig a pass, `sandbox-exec` adds macOS network-deny, `none` keeps the fully-inherited environment (back-compat escape hatch). Increment 1: env-allowlist + macOS sandbox-exec net-deny. rlimits ship gated off (preexec_fn thread-safety vs the multithreaded eval runner); container/bwrap and write-confinement are deferred to Increment 2. Full suite green with hermetic-env as the shipped default.
…pression fix (ADR-0042)
Increment 2 of the hermetic execution sandbox, plus the Threat A interim fix that
lands alongside it.
Backends (same prepare() seam): `bwrap` (Linux — read-only root, writable
workspace, unshared net namespace) and `container` (Podman/Docker — net-deny,
read-only rootfs, workspace bind-mounted, kernel-enforced pid cap, requires
AVATAR_SANDBOX_IMAGE). Both are shape-tested; guest isolation needs a Linux run.
Opt-in RLimits (AVATAR_SANDBOX_RLIMITS) ride preexec_fn for direct-exec backends
and --pids-limit for the container; off by default (preexec_fn thread-safety vs
the multithreaded eval runner).
Exit-5 laundering (Threat A): a `kind="test"` check that now collects zero tests
is a tolerated skip only if the repo was genuinely test-less. Workspace.baseline_
paths() reads the pinned-baseline tree; if the baseline HAD tests but the check
collects none, the verifier fails ("suppression, not absence") instead of skipping.
662 passed.
The '--replay' in an edge label tokenizes as a link operator and breaks the render (pre-existing since #70). Reword the label to drop the '--'/parens.
…le end-to-end The declared contract must run the real entry point (it imports and launches), not only isolated unit tests, and the model owns toolchain setup (don't assume a test runner is installed). Enriches the edit mission framing + declare_verification /alter_verification descriptions. Motivated by a dogfood run where the model got green unit tests while the actual program had a circular import — the contract never exercised the deliverable.
…it by default max_wall_clock_seconds becomes int | None (None = no wall-clock bound). The attended cockpit (jo) sets it None by default — the human's Ctrl-C and max_iterations are the backstops there — while batch/eval keep the 600s cap. An explicit AVATAR_MAX_WALL_CLOCK_SECONDS still wins in the cockpit. Motivated by a dogfood build where two agent runs were terminated as incomplete mid-work by the per-run 600s clock (confirmed ~600s of genuine agent work, not human-approval latency). The clock is already per-agent-run, not cumulative; this just lets the attended path opt out of it.
… default (red) PR-#106 review: the cockpit's off-by-default guard keys on os.environ, so a cap stated in .env is silently overridden to None.
The cockpit's off-by-default guard keyed on os.environ, so an operator cap stated in .env (which HarnessConfig reads) was silently overridden to None. Key on config.model_fields_set instead — pydantic marks env-var and dotenv values alike, and defaults not at all (PR-#106 review).
Resolves the max_wall_clock_seconds conflict by composing both intents: the branch's nullable type + cockpit-off default (ADR-0043) with main's 1800s batch default (PR #104). sdk.mdx keeps the per-run/cockpit wording with the 1800 value; ADR-0043 notes the raise.
…s first token (#110) * test(planner): pin per-segment vacuity judgment for declared checks (red) The guard currently classifies only the line's first program: it rejected `printf 'q' | python3 -m ascii_tetris.main` as vacuous (dogfood 7e49b161) and would accept `grep -q Overview DESIGN.md`. Pin the intended semantics: judge every &&/;/| stage; one real stage redeems the line. * fix(planner): judge declared-check vacuity per segment, not the line's first token vacuous_declared_check fed the whole command line to effective_invocation, which classifies token 0 only — so `printf 'q' | python3 -m ascii_tetris.main` was rejected as vacuous (the pipeline runs the real entry point; dogfood 7e49b161 burned a turn + a tier-3 amendment approval on it), while `grep -q Overview DESIGN.md` would have been accepted. Split the line into &&/||/; segments and | pipeline stages (the split _split_segments already used by tier classification, extended to single pipes) and reject only when every stage's effective program is a no-op/inspector; grep/rg/head/tail/wc/sleep join the denylist so inspection-only chains stay rejected regardless of token order. The quote-blind split can only mis-split toward accepting — the safe direction for a lower-bound guard (the immutable floor stays the real anchor). * test(planner,tools): pin builtin-bypass rejection and contract-level vacuity (red) PR-#110 review: (1) unlisted shell builtins (`|| exit 1`, `command -v`) redeem inspector-only lines; (2) per-check judgment rejects a real contract that supplements pytest with an artifact grep — burn-a-turn in a new shape. Pin: builtins never redeem a line; one real check redeems the contract. * fix(planner): deny shell builtins as vacuity redeemers; judge contracts whole PR-#110 review, both directions of the same guard: - `grep -q X || exit 1` and `command -v pytest && grep -q X` were accepted because unknown programs count as real and `exit`/`command` were unlisted. Builtins/probes (exit, command, type, which, cd, export, …), fs-noise (mkdir, touch, rm, …) and more inspectors (sort, uniq, cut, tr, stat) join the denylist; execution-capable wrappers (env, eval, xargs, find) stay off so mis-parses keep failing open. - Per-check rejection burned a turn when a real contract supplemented pytest with an artifact grep. _validate_checks now judges the contract whole: one executing check redeems it; only an all-vacuous contract is rejected, with the error steering toward an executing check. --------- Co-authored-by: t <t@t.t>
…044/0045) (#112) * docs(adr): record ADR-0044 — declared change_kinds select per-kind vacuity rulebooks The 8216e26b dogfood run (markdown design spec) rejected a legitimate anchored grep contract as vacuous, then accepted the same assertions laundered through python3 -c: the 'must execute the code' rule is a category error for textual deliverables. ADR-0044: the model declares change_kinds (list — code/content) alongside its contract; each kind demands one covering check (code keeps the executing rule, content gets anchored+falsifiable), and the diff audits the declaration at verification time (kinds(diff) must be declared). Amends ADR-0038. * test(planner,tools,verifier): pin ADR-0044 change_kinds rulebooks and diff audit (red) Pins the three seams ahead of implementation: check_covers_content (anchored+falsifiable — inspectors with operands are first-class, pure emitters and ||-neutralized lines are not), classify_change_paths (murky config fails toward code), per-kind coverage at declare_verification (default ["code"], per-kind rejection messages, companions tolerated), and the verifier's change_kind_coverage audit (under-declaration fails legibly; over-declaration and no-declaration do not). The motivating journal check is replayed verbatim and must be accepted under content. * feat(planner,tools,verifier): declared change_kinds select per-kind vacuity rulebooks (ADR-0044) The 8216e26b dogfood run rejected a legitimate anchored grep contract for a markdown deliverable as vacuous, then accepted the same assertions laundered through python3 -c — the 'must execute the code' rule is a category error for textual artifacts. Now the model declares change_kinds (list: code/content, default ['code']) alongside its contract: - Declaration: each declared kind needs >=1 covering check. code keeps the ADR-0038 executing rule; content demands anchored + falsifiable — the check names a content artifact (.md/.rst/.txt/.adoc, the same suffixes the diff classifier uses) and can exit non-zero (inspectors with operands or real executors; ||-neutralized lines rejected). Per-kind rejection messages keep the ADR-0038 vocabulary; companions are tolerated once kinds are covered. - Freeze: the kinds ride deps -> TaskState.declared_change_kinds -> the verification_plan_frozen event (None = no declaration, incl. amendments). - Audit: the verifier's required change_kind_coverage check reconciles kinds(files_modified + diff paths) against the declaration — undeclared kinds fail legibly naming kind and paths; over-declaration is tolerated; murky config classifies as code. No declaration -> no audit (tiers 1-3). Greens tests/test_change_kinds.py (18 red tests from c634b75). Note: committed --no-verify because the pre-commit pyrefly hook stashes the unrelated unstaged cockpit work-in-progress and trips on its test file; the full-tree gates (ruff, ruff format, pyrefly, 699-test suite) all pass. * refactor(planner,tools): type change_kinds as a shared ChangeKind Literal One definition (planner.ChangeKind) now keys CHANGE_KIND_COVERAGE, types both tool-input fields, and types classify_change_paths — a typo'd kind anywhere fails pyrefly, and the tool JSON schema advertises the enum to the model up front (strict tool-calling transports can enforce it at generation). An unknown kind is now rejected at ToolRuntime input validation with pydantic's permitted-values error (§10 model-correctable, same lane as a missing arg), so the handler's unknown-kind branch is deleted; the reworked test pins the rejection at that seam (runtime.execute with a raw dict, as the model reaches it). A new test guards registry exhaustiveness (dict keys may be a subset of the annotated type — typing alone can't catch a missing entry). Journaled fields (TaskState.declared_change_kinds, VerificationPlanFrozen.change_kinds) deliberately stay list[str]: journals are durable and read across version skew — a kind added later must not make an older harness unable to parse a newer journal. ADR-0044 decision section records the refinement. --no-verify: pre-commit's stash of the unrelated unstaged cockpit WIP trips pyrefly on its orphaned test file; full-tree gates (ruff, format, pyrefly, 700-test suite) all pass. * test(session,jo-cli): pin that contract amendments are never grantable (red) An [a] always on an alter_verification approval currently stores a standing ApprovalGrant (tier 3 < the tier-4 ceiling), silently auto-approving every later amendment in the session — the model can re-move its own goalposts without a human after one ratification. Pins the fix at both layers: - core: remember=True on an amendment approval degrades to allow-once (no grant stored), and a hand-built/persisted amendment grant never matches - cockpit: the ApprovalModal for alter_verification offers no Always button, no [a] hint, and the [a] key is inert; allow-once/deny still work * fix(session,jo-cli): contract amendments are never grantable — [a] always is refused An [a] always on an alter_verification approval stored a standing ApprovalGrant (tier 3, under the tier-4 ceiling), silently auto-approving every later amendment in the session: one human ratification let the model re-move its own goalposts for the rest of the sitting. ADR-0039's scoping exists to prevent exactly this widening — it was arriving through the grant door instead of the unattended-policy door. - session: _UNGRANTABLE_TOOLS = {alter_verification}; resolve_approval degrades remember=True to allow-once for such tools, and ApprovalGrant.matches never covers them (belt and braces — even a hand-built or persisted grant is inert). The config-gated unattended policy (ADR-0039) is untouched: it remains the only sanctioned auto-approve for amendments. - jo-cli: ApprovalModal offers no Always button and no [a] hint for an amendment, and the [a] key binding is inert; allow-once/deny unchanged. The guarantee is core-owned — the modal only keeps the UI honest. - docs: ADR-0039 decision point 5 records the rule; jo-cli ARCHITECTURE approval-flow section notes the exception. Greens the two red pins from 744f1fa. --no-verify: pre-commit's stash of the unrelated unstaged cockpit WIP trips pyrefly on its orphaned test file; full-tree gates (ruff, format, pyrefly, 703-test suite) all pass. * test(tools,verifier): pin shell-syntax rejection and && conjunction at command boundaries (red) Pins ADR-0045 against the tetris_glm dogfood journal (events/be46ea273029486fbc62ac5360a6c82f.jsonl): a declared && chain must split into per-segment checks (the mangled single-grep false pass must fail), heredocs/pipes/redirects must reject with a steer at declare/alter time, and run_command must reject shell operators instead of silently mis-executing. * fix(tools,verifier): reject shell syntax at command boundaries; && splits into check conjunction Workspace.run executes a single argv with no shell, so shell operators were never interpreted — a declared && chain froze as ONE mangled command (the tetris_glm journal's vacuous verification pass: grep -q exit-0'd on its first pattern while later patterns became unopenable filenames) and a declared heredoc hung on stdin to timeout, feeding the run that ended incomplete. New shared gate (avatar/shell_syntax.py, ADR-0045), applied before any model-authored command reaches Workspace.run: - declare/alter_verification: quote-aware && split into one PlannedCheck per segment (execution now matches the planner's per-segment classification); ;, |, ||, redirects and heredocs reject model-correctably with a steer. - run_command: rejects all shell operators, chains included — silently running only the first program manufactured fake exit=0 evidence. * docs(adr,research): record ADR-0045 and the tetris_glm shell-mangling false-pass analysis ADR-0045: shell syntax is rejected, not interpreted, at model-authored command boundaries; && normalizes to per-segment check conjunction; Workspace.run stays pure argv. Research note: trajectory analysis of the 2026-07-10 tetris_glm journal — one defect, two presentations (vacuous verification pass; heredoc hang feeding a finalization spiral to incomplete), with the eval-integrity caveat for chained declared contracts. * feat(jo-cli): live activity indicator — thinking / running-tool / verifying, color-coded per tool The cockpit surfaces what the agent is doing between events: an activity line driven by the event stream (model thinking, the running tool with its tool-coded style, the verifying phase), cleared when idle. Documented in jo-cli/ARCHITECTURE.md. * fix(intent): route mid-build questions to investigate — the prompt carries both tiebreakers The conversational-follow-up rule (dogfood events/04849a5a…jsonl) alone overcorrected: a mid-build question ('How do I quit the game gracefully?', tetris_glm 7745b972…) routed to edit, vacating the declaration gate and forcing the edit contract onto an investigation. Add the investigate counterweight and pin both directions so a prompt tweak can't drop either. * test(planner,tools,verifier,runner): pin PR #112 review findings (red) Pins the review's DO + three confirmed P1s + P2 and lifecycle TRYs: behavior-bearing .txt classifies as code (both classifier and content anchor); string-literal anchors never cover content; cmp/diff inspect but do not execute; quoted shell-operator args reject legibly; LLM proposals route through the ADR-0045 gate; && chains keep short-circuit semantics via shared chain ids; amendments onto detected plans stamp change_kinds; declare_verification is investigating-only; the kind audit ignores created-then-deleted paths and reads git-quoted diff headers. * fix(planner,tools,verifier,runner): close PR #112 review integrity holes - content rulebook (DO + P1): behavior-bearing .txt manifests (requirements*/constraints*/CMakeLists) classify as code in both the path classifier and the content anchor; the anchor must be an argv OPERAND of an asserting stage, so a filename inside a string literal (python -c "print('README.md')") no longer self-certifies a doc change; cmp/diff move to _VACUOUS_PROGRAMS (inspectors, not executors). - shell gate (P1s): LLM-proposed checks route through argv_segments — && splits per segment, other operators discard the proposal; split segments share a PlannedCheck.chain and the verifier stops a chain at its first failure (shell short-circuit kept — a failing segment still guards a later mutating one); bare quoted operator args ('&&') reject legibly instead of silently mis-splitting. - lifecycle: amendments stamp declared_change_kinds unconditionally, so amending a tier-1-3 detected plan no longer skips the kind audit; declare_verification is investigating-only (the editing-phase call was a phantom success — steer is alter_verification). - audit inputs (P2): change_kind_coverage reads the final state (diff ∪ still-existing touches), not the append-only ledger — created-then- deleted scratch no longer false-fails; _diff_paths unquotes git C-quoted headers so non-ASCII paths cannot escape the audit. * refactor(session,jo-cli,docs): PR #112 review polish — shared ungrantable constant, doc sync UNGRANTABLE_TOOLS becomes part of the public avatar surface; the cockpit modal derives [a] visibility from it (no string drift). app.py spinner toggle reads as if/else. run_command notes why the ADR-0045 rejection lands after the tier-3 gate (the gate stays tool-semantics-blind, §13). run_tests' no-command error steers declared contracts to run_command. ARCHITECTURE.md gains the ADR-0045 key property; README notes the amendment [a] exception; ADR-0044/0045 reconciled with the review fixes. * docs(claude): extend the no-agent-attribution rule to the whole GitHub surface The commit-authorship rule already banned Claude trailers; PR #112 shipped with '🤖 Generated with Claude Code' footers in the PR body and a review comment, which the rule's intent clearly covers. Make it explicit: no agent-attribution footers in PR titles/descriptions, comments, issues, or reviews. (AGENTS.md is a symlink to CLAUDE.md — one edit covers both.) * fix(runner,planner): verifier steers in every mode; deliverable-scoped floor (#113) * fix(runner,planner): verifier steers in every mode; deliverable-scoped floor Interactive (conversational) verification was *advisory*: a `final_answer` was delivered as `outcome="success"` regardless of the verdict, with no repair loop. Dogfood journals (tetris_grok2, 2026-07-10) showed this let the model self-certify — a failed verdict, and even a failed *immutable floor*, were laundered to `success` (invariant #3 violated). The verifier's whole purpose, steering the model toward functional correctness, was discarded in the REPL. Verification now steers in every mode (ADR-0046): a failing verdict always drives the repair loop, so the model repairs or proposes a gated `alter_verification` amendment. What shifts with who is in the loop is only the terminal-boundary disposition, settled in `_settle_terminal_outcome`: - autonomous (`--auto`): repair exhaustion -> `failed` (unchanged) - conversational (REPL default): repair exhaustion -> `blocked`, a first-class hand-off to the human (an `open_question`), never a fake `success` The eval harness's option-A external grading (ADR-0040) is preserved as a distinct `advisory` mode (report, don't steer/gate) so probe-graded baselines are unaffected; evals/run.py now passes `advisory` instead of `conversational`. The greenfield smoke floor now scopes to the deliverable (ADR-0047): the authoring prompt (`_SMOKE_SYSTEM`) names only the delivered artifact and excludes the model's own throwaway `verify_*`/scratch scaffolding, so a broken scratch file can't poison the immutable floor (the tetris_grok2 regression). Supersedes the advisory stance of §23.5 / ADR-0002 D7. * feat(runner,tools): mid-run investigate→edit escalation; run_command in investigation (#114) * feat(runner,tools): mid-run investigate→edit escalation; run_command in investigation A fix goal misrouted to `investigate` never reaches verification, so the steering from ADR-0046 can't help it: it edits blind (no execution) and can't keep its changes (net-zero-diff), then thrashes until killed (the tetris_grok3 journal, 2026-07-11). ADR-0048 makes that recoverable. - run_command admitted in `investigating` (new `_COMMAND_PHASES`): reproducing a failure is core investigation, and run_command attributes+stages its side effects so they stay inside the baseline diff the net-zero contract enforces. run_tests/ run_linter stay verify-only (no accounting, no pre-escalation plan). - `switch_to_editing` control tool (tier-3, phases={investigating}, UNGRANTABLE): the model-requested escalation lever, exposed where it's needed. On approval the runner performs `_escalate_to_edit` — same post-approval pattern as alter_verification. - `_escalate_to_edit` flips task_kind investigate→edit (one-directional, once-only) and emits `TaskEscalated`; it does NOT jump the phase/freeze, so the run becomes a normal edit task and the standard declaration gate runs on the next edit. - Baseline-clean contract: an investigate run resolves+memoizes its plan at open, before any transient edit / run_command can plant a Makefile/test — so escalation can't freeze an agent-authored contract as its own passing rubric (§5). - Both escalation triggers are consented proposals: attended asks; unattended follows `autonomous_escalation_policy` (default deny). The harness thrash detector (repeats with a persistent diff ≥ `escalation_thrash_repeats`) surfaces the harness-only signal as a directive nudge — it never auto-escalates. Investigation-ended (final_answer with a diff) is handled by the net-zero contract (revert), not escalation. - Legibility: the investigate framing and the no_unintended_diff hint point at the lever. * fix(runner): fire the declaration gate at claim-done, not only at first-edit An edit task that reaches `final_answer` with no edit and no contract slipped past the ADR-0038 declaration gate (which fired only at the edit-intent bootstrap): it froze an empty plan, failed with the cryptic "no verification contract discovered", and — pushed into `editing` by the repair loop — could not reach declare_verification (investigating-only), so it thrashed establishing a contract through alter_verification (the tetris_grok4 journal: "provide a design spec in markdown" answered inline). The gate now also fires at the claim-done boundary: a greenfield, undeclared edit proposing `final_answer` (with nudges remaining) is refused and nudged to declare — while still in `investigating`, where declare is reachable — instead of after the doomed empty verify. Bootstrap and claim-done share one `_refuse_for_declaration` helper (same nudge budget / feedback / journaling); at the cap, `final_answer` proceeds to the smoke floor as before. Deliberately does NOT make declare_verification reachable in `editing` (the naive "Fix 2"): with declaration forced in investigating, the model never legitimately needs to declare from editing — and an editing-phase (tier-0) declare could replace a frozen contract, bypassing alter_verification's tier-3 consent gate. Recorded in ADR-0049. * fix(cockpit): render modal command/diff content verbatim, not as markup The approval and diff modals rendered model-authored content (the command, the raw args dict, the diff) through Static's default markup parsing. Model commands routinely contain Textual markup metacharacters — `[` in `python -c '...[...]'`, `grep -E '[...]'`, list literals, or a `[a=1 b=]`-shaped token — which the markup parser rejects with `MarkupError: Expected markup value`, crashing the whole cockpit while laying out the modal (surfaced running `jo` on tetris_grok4). The transcript already sets markup=False; the modals were the gap. Set markup=False on the ApprovalModal prompt/command/args/hints Statics and the DiffModal body so all model/command/diff content renders literally. Independent of the escalation stack (pre-existing modal bug); bundled here since dogfooding #114 surfaced it. Regression tests: an ApprovalModal command and a DiffModal body carrying the crashing `[a=1 b=]` token now mount, render verbatim, and route without a MarkupError. * feat(cockpit): keep the live task kind + phase by the input, following escalation The status bar was docked at the screen top — stranded away from the input where the human is deciding — and its mode label was pinned to the goal's initial classification. When a consented `switch_to_editing` (ADR-0048) flipped the kind mid-run, the bar kept reading "investigate" while the run edited (dogfood `tetris_grok4`). - Relocate `#status` into the bottom footer, directly above the input, so mode·phase are always in view by the text box. - Handle `TaskEscalated` in the cockpit: update the shown kind so the bar follows an `investigate → edit` escalation instead of lying about the classification. - Export `TaskEscalated` on the public `avatar` surface (the cockpit consumes only that). * feat(state,runner,session): journal how the task kind was decided (mode_source) A dogfood run couldn't tell whether a mis-routed goal was a classifier MISS (the LLM verdict was "investigate") or a classifier OUTAGE (the call failed and the first-word heuristic decided) — both surface identically as the heuristic's verdict. The tetris_grok4 analysis hit exactly this blind spot: three edit follow-ups routed to investigate, with no way to attribute the cause post-hoc. Carry the routing source through to the journal: - TaskState.mode_source — "override" / "classifier" / "heuristic", or None when the core is driven directly (kind given, no REPL classification). Stamped by the REPL SessionManager from last_mode_source; the core only echoes it. - AgentStart gains task_kind + mode_source, emitted on both the typed bus (the JSONL journal) and the legacy emitter. Old journals parse unchanged (defaulted fields). Now a journal answers "who decided this kind?" directly, so a classifier regression is distinguishable from a classifier outage before touching the heuristic. * fix(tools,runner,adr): close PR #114 review holes — lying affordance, planted contract, thrash signal Four defects surfaced by review, all reopening failure classes ADR-0048 exists to close. 1. `switch_to_editing` was gated by PHASE only, but an `edit` task also starts in `investigating` and an escalated task has already become `edit`. So the tool was advertised ("escalate this INVESTIGATION…") to tasks that cannot escalate: the human paid a tier-3 approval, the tool returned success, and `_escalate_to_edit`'s guard silently no-oped — telling the model an escalation happened when none did. Now kind-gated in `phase_admits_tool`, the single predicate the gate and the ContextBuilder share, so advertisement and execution cannot drift; a stray call is refused as out-of-phase instead of false-succeeding. The kind flip also makes this cover the already-escalated case with no extra check. 2. The eager baseline-clean plan resolution ran only for `investigate`, but `run_command` is now admitted in `investigating` for edit tasks too — and an edit task resolved lazily, at the first edit-intent gate check. That left a planted-contract window on the ORDINARY edit path: an approved codegen command stages a Makefile before the first resolution, and detection freezes the planted file as the run's own passing contract (self-certification). Resolution is now hoisted to run open for every kind; edit tasks paid it at the gate anyway, so the cost is unchanged. 3. The thrash detector tested `ws.status_paths()` (`git status --porcelain`), which counts pre-existing untracked/dirty files the run never touched — never empty on an ordinarily dirty repo, so a clean read-only investigation was falsely told it was "leaving changes in the tree", and would self-escalate under an autonomous approve policy. It now tests `ws.diff()` — the pinned-baseline diff, exactly the signal `no_unintended_diff` enforces. 4. `TaskEscalated.trigger="thrash"` was unreachable (every call site hardcoded "model"), so a thrash-RESCUED run was indistinguishable from a self-aware one — the signal the eval loop needs. The nudge now records `escalation_nudged` on TaskState and the escalation reads it. Also: ADR-0048 sections 2-3, the `TaskEscalated` docstring, and the `_switch_to_editing` comment described a phase-jumping, plan-freezing escalation that was deliberately NOT built (the flip is kind-only, so the standard bootstrap runs the gate on the next edit — jumping the phase is precisely what would skip it). The code was right; the decision record is now too. `autonomous_escalation_policy` takes "approve" (not "auto") to match the amendment knob, and `escalation_thrash_repeats` is bounded ge=1. * docs(api): regenerate the API reference for the new public surface `gen_api_docs.py --check` failed on this branch: `AgentStart` rendered without `task_kind`/`mode_source`, and the `HarnessEvent` union omitted `TaskEscalated` — public surface this PR introduces, so the drift is not the parent chain's to carry. Regenerated with `make docs-api` (mechanical; the generator is the source of truth). --------- Co-authored-by: t <t@t.t> --------- Co-authored-by: t <t@t.t> --------- Co-authored-by: t <t@t.t>
…eypress simulation (#115) * docs(claude): forbid AI attribution in commits, PRs, and comments (#111) Co-authored-by: t <t@t.t> * feat(evals): tetris-tui — single-shot TUI task graded by a human-like keypress probe The single-shot successor to the interactive ASCII-Tetris dogfoods: the goal pins the whole graded contract (ANSI arrow bytes over a --no-raw scripted mode, streaming input, frame format + sentinel, seeded 7-bag, guideline scoring) and a nine-phase probe plays the deliverable like a human — differential frame assertions, an adaptive bottom-row packing planner that checks the line-clear score to the point, and a pty presentation phase (a ~40-line terminal emulator) that catches the raw-mode staircase pipes cannot see. No model-provided hooks: the rendered UI is the only grading surface, and the agent's README is graded (semantically) but never obeyed. Probe validity is pinned by a golden game, nine flip counter-examples, and a tolerated farewell-frame variant. * docs(research): tetris-tui eval-development note — design record, probe artifacts vs genuine failures The task's durable rationale lives here (deliberately a research doc, not an ADR — this is eval-content development, not harness architecture): the design record (probe-owned scanning vs model hooks, the pinned scripted mode, rejected alternatives, known limits), two development cells, two 3-seed matrix waves (27 cells, 9 models), the five probe artifacts each fixed the same day with replay-verified corrections, the staircase false-pass closed by the pty phase, and the surviving genuine defects (reverse-order bag draw across two model families, budget exhaustion, malformed tool calls). Recorded result files were never rewritten; corrections are dated addenda. * docs(blogging): verification-saga candidates NC7-NC10 + four writing kits NC7 (the vacuity guard's false-lesson epistemics), NC8 (the PR #110/112/113/114 self-certification arms race as a saga), NC9 (the shell-syntax quiet false pass), NC10 (tetris-tui development: five probe artifacts before any genuine model failure; spec-sensitivity as signal). Each with a full writing kit; master scorecard + snapshot updated. * chore(evals): refresh the pinned eval-matrix model set (grok-4.5, gpt-5.6-sol; 3 seeds) * refactor(tests): extract embedded golden apps into tests/goldens/ data files The three golden reference apps (news-analyzer, shop-portal, tetris) plus the tetris README and the canned-frames cheat move out of tests/test_evals.py (~950 lines of string literals) into tests/goldens/, loaded via read_text at module import. Extraction was AST-based and round-trip byte-verified, so probe behavior against the goldens is unchanged. The surgical counter-example markers stay in the test file, still containment-asserted before every replace. tests/goldens joins evals/fixtures in the ruff/pyrefly excludes: it is test data probes execute in scratch repos, not code the suite maintains, and reformatting could shift the asserted markers. * fix(evals): tetris-tui probe review fixes — TERM ownership, full-board presentation, EOF farewell tolerance, README-gate coverage Review findings from PR #115, all reproduced before fixing: - the probe now FORCES TERM=xterm on its pty (setdefault preserved a hostile CI TERM=dumb, falsely rejecting correct curses-style games; pinned by a new hostile-TERM regression test) - the presentation phase demands the full 20-row board the goal pins (a fake interactive mode of 5 static aligned rows passed; pinned as counter-example static_interactive), diagnosing the staircase first so its message survives - count-sensitive phases tolerate exactly one trailing EOF farewell frame (the q-farewell ambiguity's twin; the tolerated variant now renders on both exits) - the README gate covers the goal's full checklist (Down/soft-drop, the cell glyphs, GAME OVER were unchecked), and the known substring laxity of the score values is documented as deliberate - comment drift: nine phases, renumbered counter-examples, corrected launch count and timeout claim in the task TOML, Unix-only platform note, emulator ESC-charset coarseness note Replay-verified against kept workspaces: sol/grok/qwen passing cells still pass; opus seed0 still fails only on its genuine staircase. * chore(evals): complete the matrix-model refresh — price grok-4.5 + gpt-5.6-sol, repoint the coverage test, fix stale prose pricing.json gains the two new pinned-matrix models (OpenRouter model-get, 2026-07-12; old entries kept for committed historical rows), the tracked-models pricing test pins the actual MATRIX_MODELS set instead of the old one it passed vacuously against, and the Makefile/README prose stops advertising 4 models x 5 seeds with the removed gpt-5.3-codex. * docs(research): commit the result rows the tetris-tui note cites; record the first nine-phase-native pass Per the .gitignore whitelist convention (cited-by-committed-research-notes rows are kept so write-ups are self-contained and auditable without re-running paid cells): the two dev cells, both matrix waves, and both first-solve cells. Addendum 4 gains the first cell graded natively by the nine-phase probe (gpt-5.6-sol, all phases green, setcbreak + \r\n — both remedies the amended goal names). * fix(evals): tetris-tui goal states requirements, not remedies — drop the raw-mode how-to hint Maintainer ruling: the task mirrors real human intent — a well-defined acceptance spec the model reasons against. The briefly-shipped 'beware tty.setraw ... use setcbreak / \r\n / curses' sentence leaked the recipe and silently changed what phase 8 measures (instruction-following instead of unprompted terminal-semantics knowledge). The requirement stays and sharpens: 'genuinely playable on a real terminal — every board row starts at the same screen column'. Design-record item 7 records the principle; addendum 4 marks the two hinted-goal cells (20260712T110513Z, T113045Z) as incomparable with cells on either side — incidentally a clean natural experiment: naming the pitfall flipped the previously-failing models within one run. * docs(blogging): rewrite blog-candidates as a single plan of action (#116) The doc had accumulated four competing "current plans" (locked roadmap, master scorecard, recommended publishing sequence, risk-calibrated rollout) and six overlapping ID schemes for the same ~20 articles, so it contradicted itself and had drifted from the actual blog repo state. Reduce it to one status board, one queue, one distribution plan, and add an update protocol so status is read from git rather than from memory. Archive the superseded reasoning verbatim. Corrects three factual errors: post 02 was listed as a stub when it is a finished draft in review; the NC1 (cost-per-solved) and NC2 (pass^k) candidates were still advertised as ready-to-write despite both being absorbed into post 02; and every research link was broken (wrong relative base, plus two filenames predating the date-prefix rename in 7ca2e64). Co-authored-by: t <t@t.t> * feat(evals): tetris-hard task + transport-decomposed probe; rename tetris-tui to tetris-easy tetris-hard (formerly tetris-playable) measures the broad construct 'can the agent build a playable Tetris' as the design-freedom counterpart to the pinned-contract task, now renamed tetris-easy. The goal pins only the observation seam (board glyphs in both modes, the streaming transport sentence, seeded determinism, frame format); bag order, spawn orientations, rotation/kicks, and scoring values stay agent choices. The probe grades gameplay over a replayed key prefix (leaning on the pinned determinism), so verdicts are transport-independent; the goal's streaming sentence is checked by a dedicated phase that reports on its own transport line without gating the exit code (_STREAMING_GATES, default off). Hidden rows above the visible field are tolerated end-to-end (including 0-visible mid-rotation clips, with pieces soft-dropped to a safe depth before the line-clear planner enumerates orientations), GAME OVER is accepted anywhere in the final output, and the interactive pty phase retries keypresses like a human before declaring a no-op. Each tolerance answers a false-rejection class found by hand-auditing matrix 20260712T140147Z (reported pass@1 0.19 vs audited ground truth 0.52) and is pinned by a golden variant or counter-example in tests. Also: cap both tetris tasks at 1200s wall clock (slow-cohort cells were riding the 1800s cap after already-working builds), price gpt-5.6-terra and gpt-5.6-luna, and record the development arc in two research notes with their cited result stamps whitelisted. * fix(evals): force a portable terminfo chain on probe-owned ptys CI has failed since the pty phases landed: python-build-standalone's bundled ncurses does not search Debian's /etc/terminfo or /lib/terminfo, so on the CI runner curses-based interactive modes die at setupterm before painting a single cell — every golden-interactive test failed with '0 parseable board rows' while raw/cbreak games passed. Force TERMINFO_DIRS to the standard chain (trailing empty entry keeps the compiled-in defaults) next to the already-forced TERM, and capture the child's stderr in the tetris-hard interactive phase so a pty-side crash surfaces as a legible reason instead of a silent empty screen. * fix(evals): strip inherited LINES/COLUMNS from probe-owned ptys The stderr capture from the previous commit surfaced the real CI crash: _curses.error: addwstr() returned ERR. ncurses lets LINES/COLUMNS from the environment OVERRIDE the pty's real window size, and the CI test environment exports values smaller than the 40x120 the probe sets via TIOCSWINSZ — so a correct curses UI crashes mid-paint on the runner while passing locally, where the shell does not export them. The probe owns the pty, so its terminal identity is now forced end to end: TERM, TERMINFO_DIRS, and no inherited size overrides. Verified locally with COLUMNS=80 LINES=24 exported, which reproduced the crash before this change. * test(evals): surface the tetris-easy probe's phase log on golden-test failure run_probe returns only the exit code, which left the remaining CI failures illegible (the runner shows 'assert 1 == 0' with no phase). Run the probe directly in the three golden tests and attach its stdout to the assertion message, mirroring the tetris-hard transport test. * fix(evals): wait for the first full interactive paint before judging presentation The presentation capture used a fixed 2s window; a loaded CI runner spends most of that on interpreter startup and curses init, so the screen was captured mid-paint and a correct game read as '1 aligned board row'. Poll the emulated screen until a full 20-row board appears (10s deadline; healthy games break out in well under a second) before sending q. * fix(evals): port the full terminal emulator into the tetris-easy presentation phase The easy probe's screen reconstruction only understood absolute CSI H/f addressing. macOS ncurses paints almost entirely with H, so it passed locally — but Linux ncurses for TERM=xterm optimizes with relative motions (CSI A/B/C/D, VPA, CHA, index/reverse-index, erases, scroll regions), which the emulator dropped, collapsing the whole board onto one screen line and rejecting a correct game with 'drew 1 aligned board rows'. Port the tetris-hard probe's emulator (which passes on the same runner) so the two cannot diverge again. --------- Co-authored-by: t <t@t.t>
…tandard Two identical 7x6x5 matrix runs of the declared-verification tree against the 2026-07-05 landscape baseline: no significant regression (pass@1 0.75/0.77 vs 0.75, paired diff clears every model). Documents the Groq-side declare_verification tool-validation death signature, one un-root-caused NoneType harness crash, the missing provider-routing metadata gap, and a run-2 calibration of the metrics themselves (84% seed-level agreement between identical runs; per-task rank order and single-run pass^k are noise-dominated at 5 seeds).
Resolve docs/blogging/blog-candidates.md: main (#116) and this branch (#115) each rewrote the doc into the single plan-of-action form. Keep this branch's version -- it is main's rewrite plus the four verification-saga candidates (vacuity-guard-false-lesson, self-certification-arms-race, shell-syntax-boundary, eval-probe-false-rejections), their writing-kit/research evidence links, and the NC7-NC10 mappings.
PR #116 collapsed six overlapping ID schemes into one rule -- one article = one ID (its blog-dir slug) -- and demoted the NC* numbers to a frozen legacy crosswalk whose own instruction is 'do not reintroduce them'. The archive only ever used NC1-NC6. This branch had minted NC7-NC10 for the four verification-saga candidates, so those rows mapped nothing legacy and regrew the retired scheme. Drop them; the candidates keep their slug IDs in the backlog table. Update the two writing kits that referred to these candidates by number to use slugs (NC3 -> 06, NC7/NC8 -> their slugs).
…act' into feat/evals-heldout-measurement
…asurement # Conflicts: # README.md
Description
Stacked on #106 (
feat/declared-verification-contract) — the eval-measurement side of thedeclared-verification feature (ADR-0040). Splits the
solvedverdict into its two componentsignals and adds a metric that makes gaming visible.
Motivation
Under auto-approved self-amendment (ADR-0039), a model can rewrite a failing check to pass it, so
its own contract passing certifies nothing. Honest measurement must key on the independent
hidden oracle (the success probe the agent never sees), not the self-report — and it must surface
the gap between the two so goalpost-moving is measured, not hidden inside a green checkmark.
Changes
evals/result.py:ResultRowgainsself_reported_success(the model's claim — reachedfinal_answer) andheld_out_passed(the independent oracle's verdict). Both defaultFalse, sorows written before the split load cleanly.
evals/run.py:run_taskpopulates them —held_out_passedis the success probe's real exitcode when a probe exists, else the composed grade (no separate oracle → the harness verifier is
the authority).
evals/metrics.py:held_out_pass_at_1(honest capability) +gamed_rate(
self_reported ∧ ¬held_out), surfaced per-model inbuild_summary.evals/classify.py: a probe failure the model claimed to satisfy is bucketedgamed,distinct from an honest
probe_failed.Testing
tests/test_evals.py: new tests forheld_out_pass_at_1,gamed_rate, and thegamedclassifybucket (a claimed-but-oracle-rejected run); existing
probe_failedtest still green (its fixturedefaults
self_reported_success=False).make checkgreen — 634 passed.python -m evals.diffvalidation (McNemar / clustered CI / model-agnosticism) is a live run and must passbefore the ADR flips Proposed → Accepted.