Skip to content

phase3+gate2: B2 baseline, gated evaluation, schema v1.1, docs restructure - #15

Merged
Atik203 merged 2 commits into
devfrom
atik
Sep 16, 2026
Merged

Atik203 merged 2 commits into
devfrom
atik

Conversation

@Atik203

@Atik203 Atik203 commented Sep 16, 2026

Copy link
Copy Markdown
Owner

What this PR does

Brings atik up to date with the merged docs cleanup and implements the rest of Phase 3 plus all of Phase 4 (Gate 1 & Gate 2 evidence), the schema v1.1 unfreeze, and the docs restructure.

Phase 3 — B2 ToolGate baseline (Gate 1)

  • Real Hoare pre/post contracts: InjecAgent evaluated universe 79/79, MCPTox 65 contracts (8.1% distinct / 37.2% availability-weighted; the 736-tool tail + every poisoned registration stays no_contract).
  • World-state with directories + best-effort seeding from the trusted request; coverage frozen in configs/b2_coverage.json (scripts/freeze_b2_coverage.py).
  • Divergence notes and the blueprint Sec 10 fallback validation are documented in docs/experiments/03_b2_toolgate.md.

Phase 4 — Gate build (Gate 2)

  • AgentLoop is a real ReAct loop that executes only through GateMiddleware.execute(); trace provenance (model/embedding metadata + run-time contract) is logged per call.
  • Gated benchmark runner: --gate none|ours|toolgate on harness/run_injecagent.py, harness/run_mcptox.py, and the scripts/run_*_ours.py pilots; the ordering invariant is enforced structurally (policy built before attacker content).
  • Integration run (100 InjecAgent + 100 MCPTox × 3 conditions): InjecAgent ASR-valid 9/100 → 0/100 with 0 false-positive blocks; MCPTox attack-influenced 22 → 10 (12 blocked); p95 ≤ 23 ms on CPU.
  • Threshold sweep / ASR–FPR Pareto from cached traces (eval/sweep.py::sweep_cases, scripts/sweep_gated.py): knee at τ = 0.75 (pilot FPR 4%).
  • Error taxonomy (scripts/error_taxonomy.py) + rule fixes (path scope, destructive-file verbs, shell-operator guard) and the system_change schema v1.1 unfreeze (30-request LLM re-validation).

Docs

  • blueprint.md moved to docs/blueprint.md (indexed ToC/anchors, key-content callouts; no design changes).
  • New docs/README.md and docs/experiments/README.md hubs; 10 short notes merged into 4 indexed reports (E1–E10) with datasets, artifacts, reproduce commands, and supervisor Q&A.

Verification

  • pytest -q → 130 passed (data-dependent tests skip without benchmark clones).
  • ruff check clean on all new/changed files (pre-existing repo-wide findings untouched).
  • 42 local doc links verified; all cross-references updated (README, AGENTS, setup, roadmap, CHANGELOG, configs, scripts).
  • Roadmap checkboxes + CHANGELOG updated in the same PR (per CONTRIBUTING).

Notes for reviewers

  • docs/experiments/README.md is the entry point: gate status matrix, E1–E10 index, artifact map, and the old→new docs merge map.
  • The MCPTox residual (~half of attack-influenced calls) is registration-trust, documented as a scope boundary rather than hidden.
  • B2's MCPTox coverage is low by construction (manual per-tool contracts vs an 801-tool attacker-registered toolset); numbers and limitations are frozen in configs/b2_coverage.json and the E7/E8 reports.

…category

- intent_schema v1.1: system_change required, fail-closed default disallow; 9 few-shots; prompt authorization examples; heuristic fallback
- rules: token-exact system category (create/update/delete/deletion/disable/grant/unlock/schedule/deploy/manage/...), delete verbs moved from file category; Manager-suffix false positive fixed
- 30-request LLM re-validation (docs/experiments/parser_spotcheck_v2.md); pilot FPR stays 4%
- post-unfreeze runs (tau=0.75): InjecAgent 0/100 succ (last miss vetoed), MCPTox attack-influenced 10 (12 blocked); docs updated (errors/integration/sweep), 130 tests
… blueprint into docs/

- blueprint.md -> docs/blueprint.md with ToC, section/component anchors and key-content callouts (no design changes)
- new docs/README.md (docs hub) and docs/experiments/README.md (status matrix, E1-E10 index, datasets, models, artifact map, reproduce commands, merge map, cross-cutting supervisor Q&A)
- 10 experiment notes merged into 4 reports: 01 gate-0 foundation (E1-E4), 02 parser/schema (E5-E6), 03 B2 ToolGate (E7), 04 gate-2 gated eval (E8-E10); each with at-a-glance, goal, setup, results, artifacts, limitations, supervisor Q&A
- references updated (README, AGENTS, setup, roadmap, CHANGELOG, issue template, configs schema comment, build_pilot_set docstring); 42 local links verified; 130 tests green
Copilot AI lite review requested due to automatic review settings September 16, 2026 19:39

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@Atik203
Atik203 merged commit 320af0c into dev Sep 16, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants