diff --git a/.claude/commands/agent-runtime-kit/upgrade.md b/.claude/commands/agent-runtime-kit/upgrade.md index 1783ae2..e5ea7b2 100644 --- a/.claude/commands/agent-runtime-kit/upgrade.md +++ b/.claude/commands/agent-runtime-kit/upgrade.md @@ -1,216 +1,20 @@ --- name: "Agent Runtime Kit: Upgrade" -description: Run the local agent-runtime-kit SDK evolution workflow and produce a safe upgrade PR +description: Inspect and safely upgrade agent-runtime-kit vendor SDK dependencies category: Workflow tags: [agent-runtime-kit, sdk-evolution, upgrade, workflow] --- -Run the local SDK evolution agent for `agent-runtime-kit`. +# Agent Runtime Kit SDK Upgrade -Use this when the user asks to upgrade, refresh, or evolve agent-runtime-kit -against current Claude Agent SDK, OpenAI Codex SDK, Codex CLI binary, or Google -Antigravity SDK releases. +Read and follow the canonical repository workflow in +`.codex/skills/agent-runtime-kit-upgrade/SKILL.md`. That file owns candidate +discovery, no-cooloff resolution, report-first gates, local upgrade +authorization, verification, and PR publication rules; do not duplicate or +weaken them here. -Default runtime for this Claude command: `claude-agent-sdk`. - -## Guardrails - -- All AI reasoning, planning, implementation decisions, structured output, and - review MUST go through `python -m examples.sdk_evolution_agent` and therefore - through agent-runtime-kit runtime primitives. -- Do not call Anthropic, OpenAI, Google, Bedrock, Vertex, or other model APIs - directly from this command. -- Local tools are allowed for deterministic work: Git, `gh`, `uv`, Python, - package metadata fetching, filesystem inspection, SDK introspection, tests, - and report inspection. -- Do not scrape unsupported credentials. Use only supported SDK auth surfaces. -- Do not auto-merge. Do not publish a release. Open or update a draft PR only - when requested or when the user explicitly asks for an upgrade PR. -- Do not run from a dirty or divergent checkout. Create a fresh worktree from - `origin/main` unless the user explicitly gives a different base. - -## Inputs - -The user may specify: - -- runtime: `claude-agent-sdk`, `codex-agent-sdk`, or `antigravity-agent-sdk` -- package subset, otherwise inspect all supported packages -- branch name, otherwise derive one from the current timestamp -- whether to create a draft PR -- whether implementation is allowed - -If unspecified, inspect all packages: - -- `claude-agent-sdk` -- `openai-codex` -- `openai-codex-cli-bin` -- `google-antigravity` - -## Preflight - -1. Announce the active checkout and the new worktree path. -2. Fetch current remote state: - - ```bash - git fetch origin --prune - ``` - -3. Create a new worktree from the chosen base: - - ```bash - git worktree add -b "sdk-evolution-upgrade-$(date +%Y%m%d-%H%M%S)" \ - /tmp/ark-sdk-evolution-upgrade origin/main - ``` - -4. In the new worktree, verify local tooling: - - ```bash - gh auth status - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked python -m examples.sdk_evolution_agent --help - ``` - -5. Resolve the runtime that will run the AI-backed stages. Use - `claude-agent-sdk` unless the user explicitly selected another runtime. - Change the runtime and uv extra together: - - `claude-agent-sdk` -> `--extra claude` - - `codex-agent-sdk` -> `--extra codex` - - `antigravity-agent-sdk` -> `--extra antigravity` - -6. Verify provider auth through supported mechanisms only: - - Claude: Anthropic API key, Claude Code auth, or Claude Code provider - environment/settings such as Bedrock or Vertex modes. Use the AWS SDK - credential chain for Bedrock and Google Application Default Credentials for - Vertex; do not read credential files yourself. - - Codex: local Codex auth/config. The SDK evolution runner injects - `CODEX_HOME=~/.codex_agent_runtime_sdk` for Codex-backed stages and mirrors - the normal `~/.codex/auth.json` cache into that isolated home before use. - - Antigravity: `GEMINI_API_KEY`, `GOOGLE_API_KEY`, or Google Application - Default Credentials with project/location environment variables. - -7. If the selected runtime is `codex-agent-sdk`, prepare and verify fresh auth - for the dedicated SDK Codex home before running the evolution agent: - - ```bash - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex - ``` - - The helper creates `~/.codex_agent_runtime_sdk`, removes uv freshness cutoff - variables, mirrors `~/.codex/auth.json` into that isolated home when the - normal cache is newer, and checks `codex login status` against that exact - home. If it exits non-zero, STOP before running `examples.sdk_evolution_agent` - and refresh the normal Codex login cache: - - ```bash - uv run --locked --extra codex codex login --device-auth - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex - ``` - -## Report-Only Evidence Pass - -Run a report-only pass first. Explicitly bypass freshness cutoffs because fresh -upstream SDK releases are the point of this workflow: - -The runner removes environment cutoffs and passes -`--exclude-newer-package =false` for every monitored vendor package. -The project declares the same package-scoped exemptions, while every other -dependency keeps the repository's normal eight-day delay. - -```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra claude python -m examples.sdk_evolution_agent \ - --runtime claude-agent-sdk \ - --refresh-preview \ - --inspect-candidates \ - --package claude-agent-sdk \ - --package openai-codex \ - --package openai-codex-cli-bin \ - --package google-antigravity -``` - -`--inspect-candidates` is explicit consent to install and import a missing or -drifted locked baseline and resolver-selected candidates in credential-scrubbed -temporary environments. - -If the user chose another runtime, replace both the `--runtime` value and the -matching uv extra using the mapping above. Do not add direct model calls. - -Inspect the newest `reports/sdk-evolution//` directory and summarize: - -- `evidence.json` -- `release_notes.json` -- `api_diffs.json` -- `behavior_probes.json` -- `behavior_diffs.json` -- `behavior_summary.json` -- `current_state.json` -- `direction_analysis.json` -- `architecture_decision.json` -- `review.json` -- `report.md` - -Stop before implementation if any of these are true: - -- required candidate API diffs are missing, -- required release-note evidence could not be collected, -- `behavior_summary.json` is missing, malformed, has an unknown status, or - reports `fail` / `incomplete`, -- `architecture_decision.json` has `manual_design_required: true`, -- the reviewer rejects the evidence or design, -- recursive self-adaptation is required and the report does not include a safe - migration plan for the agent's own use of `AgentTask`, `RuntimeRegistry`, - adapters, output schemas, event sinks, permission profiles, or - `AgentResult`. - -`pass` means complete unchanged evidence; `changed` means complete -non-breaking evidence; `incomplete` means required observations could not be -proved; and `fail` means a required contract failed or a breaking diff was -observed. - -## Implementation Pass - -Only run implementation when the report-only pass supports it and the user wants -an upgrade branch or PR: - -```bash -BRANCH="sdk-evolution-upgrade-$(date +%Y%m%d-%H%M%S)" - -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra claude python -m examples.sdk_evolution_agent \ - --runtime claude-agent-sdk \ - --refresh-preview \ - --inspect-candidates \ - --implementation-enabled \ - --create-branch \ - --branch-name "$BRANCH" \ - --draft-pr \ - --pr-base main \ - --commit-message "Run SDK evolution update" \ - --pr-title "Run SDK evolution update across vendor packages" \ - --package claude-agent-sdk \ - --package openai-codex \ - --package openai-codex-cli-bin \ - --package google-antigravity -``` - -If the user chose another runtime, replace both the `--runtime` value and the -matching uv extra using the mapping above. - -## Verification - -After implementation, run or verify: - -```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv lock --check -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked ruff check . -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked mypy -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked pytest -``` - -If a draft PR was created, watch CI until it finishes or clearly report that it -is still running. Include the PR URL, report path, changed SDK versions, -`behavior_summary.json` status and reasons, architecture decision, reviewer -result, test results, uncertainty, and manual review checklist in the final -response. +For this Claude command, use `claude-agent-sdk` with `--extra claude` unless the +user selected another runtime. Preserve every other authorization boundary from +the canonical workflow: an explicit SDK upgrade request permits the gated local +dependency update, while branch, commit, push, PR, merge, and release actions +remain separately authorized. diff --git a/.codex/skills/agent-runtime-kit-upgrade/SKILL.md b/.codex/skills/agent-runtime-kit-upgrade/SKILL.md index da01d5e..bd803a7 100644 --- a/.codex/skills/agent-runtime-kit-upgrade/SKILL.md +++ b/.codex/skills/agent-runtime-kit-upgrade/SKILL.md @@ -1,97 +1,97 @@ --- name: agent-runtime-kit-upgrade -description: Run the local agent-runtime-kit SDK evolution workflow to inspect and safely upgrade Claude Agent SDK, OpenAI Codex SDK, Codex CLI binary, and Google Antigravity SDK dependencies. Use when the user asks to upgrade, refresh, evolve, or create a PR for agent-runtime-kit against current upstream agent SDK releases. +description: Inspect and safely upgrade agent-runtime-kit's Claude Agent SDK, OpenAI Codex SDK and coupled CLI binary, and Google Antigravity SDK dependencies. Use for SDK freshness checks, upgrades, compatibility repairs, or upgrade PRs in agent-runtime-kit. --- -# Agent Runtime Kit Upgrade - -Run the local SDK evolution agent for `agent-runtime-kit`. - -Default runtime for this Codex skill: `codex-agent-sdk`. - -## Non-Negotiables - -- Route all AI reasoning, planning, implementation decisions, structured output, - and review through `python -m examples.sdk_evolution_agent`. -- Do not call OpenAI, Anthropic, Google, Bedrock, Vertex, or other model APIs - directly from this skill. -- Use local tools only for deterministic work: Git, `gh`, `uv`, Python, package - metadata fetching, filesystem inspection, SDK introspection, tests, and report - inspection. -- Do not scrape unsupported credentials. -- Do not auto-merge or publish releases. -- Do not run from a dirty or divergent checkout. Create a fresh worktree from - `origin/main` unless the user explicitly gives another base. - -## Preflight - -1. Announce the active checkout and the new worktree path. -2. Fetch remote state and create a fresh worktree: - - ```bash - git fetch origin --prune - git worktree add -b "sdk-evolution-upgrade-$(date +%Y%m%d-%H%M%S)" \ - /tmp/ark-sdk-evolution-upgrade origin/main - ``` - -3. In the new worktree, verify local tooling: - - ```bash - gh auth status - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked python -m examples.sdk_evolution_agent --help - ``` - -4. Resolve the runtime that will run the AI-backed stages. Use - `codex-agent-sdk` unless the user explicitly selected another runtime. - Change the runtime and uv extra together: - - `claude-agent-sdk` -> `--extra claude` - - `codex-agent-sdk` -> `--extra codex` - - `antigravity-agent-sdk` -> `--extra antigravity` - -5. Use supported provider auth only: - - Claude: Anthropic API key, Claude Code auth, or Claude Code provider - settings such as Bedrock or Vertex. Bedrock must use the AWS SDK credential - chain; Vertex must use Google Application Default Credentials or supported - environment variables. - - Codex: supported local Codex login cache. The runner injects - `CODEX_HOME=~/.codex_agent_runtime_sdk` for Codex-backed stages and mirrors - the normal `~/.codex/auth.json` cache into that isolated home before use. - - Antigravity: `GEMINI_API_KEY`, `GOOGLE_API_KEY`, or Google Application - Default Credentials with a project and optional location variables. - -6. If the selected runtime is `codex-agent-sdk`, prepare and verify fresh auth - for the dedicated SDK Codex home before running the evolution agent: - - ```bash - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex - ``` - - The helper creates `~/.codex_agent_runtime_sdk`, removes uv freshness cutoff - variables, mirrors `~/.codex/auth.json` into that isolated home when the - normal cache is newer, and checks `codex login status` against that exact - home. If it exits non-zero, STOP before running `examples.sdk_evolution_agent` - and refresh the normal Codex login cache: - - ```bash - uv run --locked --extra codex codex login --device-auth - env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ - uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex - ``` - -## Report-Only First - -Always run a report-only pass before implementation. Explicitly bypass uv -freshness cutoffs: - -The runner removes environment cutoffs and passes -`--exclude-newer-package =false` for every monitored vendor package. -The project declares the same package-scoped exemptions, while every other -dependency keeps the repository's normal eight-day delay. +# Agent Runtime Kit SDK Upgrade + +Use the repository-owned evolution runner for evidence collection, candidate +inspection, compatibility gates, dependency updates, and verification. + +## Authorization + +- Always run a report-only pass before changing dependencies. +- An explicit request to upgrade, refresh, or update the SDKs authorizes the + gated local `pyproject.toml`, `uv.lock`, and compatibility-manifest update + after the report passes. +- A bare skill invocation or freshness question is report-only. +- Creating a branch, commit, push, or pull request requires an explicit request + for that publication action. Never auto-merge or publish a release. + +## Invariants + +- Discover each monitored package's latest upstream release independently of + the repository's current constraints. A constraint that excludes a newer + release is a result to investigate, not proof that the repository is current. +- Keep these facts separate in the report: upstream latest, compatible + candidate, current resolver result, prospective resolver result, and applied + lock version. +- `openai-codex-cli-bin` is coupled to the exact version required by the + latest published `openai-codex`. Record a newer standalone CLI release as a + staged artifact, not a blocked candidate. It may be fingerprinted in a + credential-scrubbed disposable environment for implementation-history + analysis, but never select it into the project unless the SDK requires it. +- Direction-of-travel means trends in actual upstream implementations. Inspect + recent release files and Python definitions even when no dependency update is + pending. Never turn the direction report into upgrade, hold, resolver, or + lockfile advice; that belongs to the architecture and implementation stages. +- Run UV without freshness delays. The runner removes cutoff environment + variables and uses `uv lock --exclude-newer false`; do not reintroduce the + retired repository cooloff or package-specific exemptions. +- Candidate API snapshots and behavior probes must cover every exact + `(package, baseline, candidate)` transition. Field presence alone is not a + semantic probe: construct the provider configuration and exercise the + adapter-owned contract. +- Never promote or create a PR unless the update produced a non-empty diff, the + lock contains every inspected compatible candidate exactly, and all + verification commands passed. + +## Checkout Preflight + +Fetch and compare the current checkout with the requested base. Work directly +only when it is the intended clean checkout. Preserve dirty or unrelated work. +When isolation is needed, use a unique path rather than the old shared `/tmp` +path: ```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ +git fetch origin --prune +scratch="$(mktemp -d "${TMPDIR:-/tmp}/ark-sdk-evolution.XXXXXX")" +worktree="$scratch/worktree" +branch="sdk-evolution-upgrade-$(date +%Y%m%d-%H%M%S)-$$" +git worktree add -b "$branch" "$worktree" origin/main +``` + +Announce the selected checkout and base. Do not delete a worktree that contains +uncommitted work. If repository instructions mention a GSD command that is not +actually installed, record the investigation and plan in the available task +plan, then continue; an unavailable wrapper is not a reason to abandon an +explicitly requested repair. + +## Runtime Preflight + +Use `codex-agent-sdk` unless the user selected another runtime. Match runtime +and optional extra: + +- `claude-agent-sdk` -> `--extra claude` +- `codex-agent-sdk` -> `--extra codex` +- `antigravity-agent-sdk` -> `--extra antigravity` + +Use supported authentication only. For Codex, prepare the dedicated SDK home: + +```bash +env -u UV_EXCLUDE_NEWER \ + uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex +``` + +If authentication fails, stop and report that concrete blocker. Do not scrape +or reconstruct credentials. + +## Report-Only Pass + +Run all four monitored packages unless the user explicitly narrowed scope: + +```bash +env -u UV_EXCLUDE_NEWER \ uv run --locked --extra codex python -m examples.sdk_evolution_agent \ --runtime codex-agent-sdk \ --refresh-preview \ @@ -102,83 +102,82 @@ env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ --package google-antigravity ``` -`--inspect-candidates` is explicit consent to install and import a missing or -drifted locked baseline and resolver-selected candidates in credential-scrubbed -temporary environments. - -If the user explicitly chooses another runtime, replace both the `--runtime` -value and the matching uv extra using the mapping above. Codex-backed runs -should use the runner's built-in `gpt-5.5` and -`reasoning_effort=xhigh` policy; do not implement model selection outside the -runner. - -Inspect the newest `reports/sdk-evolution//` directory: - -- `evidence.json` -- `release_notes.json` -- `api_diffs.json` -- `behavior_probes.json` -- `behavior_diffs.json` -- `behavior_summary.json` -- `current_state.json` -- `direction_analysis.json` -- `architecture_decision.json` -- `review.json` -- `report.md` - -Stop before implementation when candidate API diffs are missing, required -release-note evidence is missing, `behavior_summary.json` is missing, malformed, -has an unknown status, or reports `fail` / `incomplete`, -`manual_design_required` is true, the reviewer rejects the evidence or design, -or recursive self-adaptation lacks a safe migration plan. `pass` means complete -unchanged evidence; `changed` means complete non-breaking evidence; -`incomplete` means required observations could not be proved; and `fail` means -a required contract failed or a breaking diff was observed. - -Recursive self-adaptation means the upgrade affects the runner's own use of -`AgentTask`, `RuntimeRegistry`, adapters, output schemas, event sinks, -permission profiles, typed unsupported-feature errors, or `AgentResult`. The -upgrade must update those usages, tests, and docs in the same scoped change, or -stop for manual design review. - -## Implementation And PR - -Only run implementation when the report-only pass supports it and the user wants -an upgrade branch or PR: +`--inspect-candidates` is explicit consent to install and import exact locked baselines and candidate +versions in credential-scrubbed temporary environments. The runner reuses each +package/version environment for its API snapshot and behavior probe, retries +transient metadata/install failures, and prints progress while it works. -```bash -BRANCH="sdk-evolution-upgrade-$(date +%Y%m%d-%H%M%S)" +Read the newest `reports/sdk-evolution//report.md` plus its raw JSON +artifacts, especially `implementation_diffs.json` and `behavior_summary.json`. +The implementation-trend section must describe observed source/artifact changes +and explicit binary limitations, never dependency action. Treat any of these as a hard +implementation blocker: + +- missing or failed no-cooloff prospective resolver preview; +- upstream candidate absent from the exact API-diff inventory; +- missing required release evidence; +- behavior status `fail` or `incomplete`, malformed/contradictory evidence, or + a missing exact-version comparison; +- `manual_design_required`, reviewer rejection, or an unsafe architecture + decision; +- recursive runtime impact without an explicit adaptation and rerun plan. + +`pass` means complete unchanged behavior evidence. `changed` means complete, +non-breaking behavior evidence. Neither status permits ignoring snapshot import +errors or a failed resolver preview. + +## Apply a Requested Local Upgrade -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ +When the user explicitly requested an upgrade and the report-only pass is green, +rerun with `--implementation-enabled` and without PR flags: + +```bash +env -u UV_EXCLUDE_NEWER \ uv run --locked --extra codex python -m examples.sdk_evolution_agent \ --runtime codex-agent-sdk \ --refresh-preview \ --inspect-candidates \ --implementation-enabled \ - --create-branch \ - --branch-name "$BRANCH" \ - --draft-pr \ - --pr-base main \ - --commit-message "Run SDK evolution update" \ - --pr-title "Run SDK evolution update across vendor packages" \ --package claude-agent-sdk \ --package openai-codex \ --package openai-codex-cli-bin \ --package google-antigravity ``` -## Verification +The deterministic implementation lane widens only excluding upper bounds, +updates the lock with `--exclude-newer false`, verifies exact target versions, +updates `src/agent_runtime_kit/compatibility.py` to the exact tested SDK and +coupled-runtime versions, and rolls all three artifacts back on resolver +mismatch or test failure. If a candidate requires adapter source changes, the +runner must block; repair the evidenced contract in the scoped coding workflow, +add a regression test, and rerun the report before applying dependencies. + +## Pull Request Mode + +Only when the user requested a PR, add a unique branch and the PR flags to the +green implementation pass: + +```bash +--create-branch --branch-name "$branch" --draft-pr --pr-base main +``` + +The runner may stage only its reported `changed_paths`. Verify the exact pushed +head, PR URL, changed files, and CI state. Do not claim a PR exists merely +because `--draft-pr` was requested; the runner skips publication when the +change is empty, unapplied, or unverified. + +## Final Verification -Run or verify: +Verify the deciding live surface again: ```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv lock --check -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked ruff check . -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked mypy -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE uv run --locked pytest +env -u UV_EXCLUDE_NEWER uv lock --check +env -u UV_EXCLUDE_NEWER uv run --locked ruff check . +env -u UV_EXCLUDE_NEWER uv run --locked mypy +env -u UV_EXCLUDE_NEWER uv run --locked pytest ``` -If a draft PR was created, watch CI until it finishes or clearly report that it -is still running. Final output should include the PR URL, report path, changed -SDK versions, `behavior_summary.json` status and reasons, architecture decision, -reviewer result, test results, uncertainty, and manual review checklist. +Report the baseline and applied versions, candidate classifications, behavior +status, exact verification results, changed files, report path, and any +remaining blocker or uncertainty. For PR work, also report the verified URL, +head SHA, and current checks. diff --git a/AGENTS.md b/AGENTS.md index 80dc9ea..7336c8e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -57,10 +57,10 @@ capabilities needed for real work. |------------|---------|---------|-----------------| | Python | >=3.10 | Package runtime | Claude Agent SDK, Codex SDK, and Google Antigravity all currently advertise Python 3.10+ compatibility. | | Pydantic | 2.13.4 current; use >=2.12 | Public request/result validation where useful | Mature typed validation without forcing callers into framework-specific models. | -| claude-agent-sdk | 0.2.96 current | Claude runtime adapter | Official Agent SDK for Claude Code-style local agent execution. | -| openai-codex | 0.1.0b3 current | Codex runtime adapter | Official Python SDK for Codex app-server integration. | -| openai-codex-cli-bin | 0.136.0 current | Codex runtime dependency | Pinned Codex CLI runtime used by the Python SDK package. | -| google-antigravity | 0.1.2 current | Antigravity runtime adapter | Official Google Antigravity Python SDK for local agent harness integration. | +| claude-agent-sdk | 0.2.148 current | Claude runtime adapter | Official Agent SDK for Claude Code-style local agent execution. | +| openai-codex | 0.147.0 current | Codex runtime adapter | Official Python SDK for Codex app-server integration. | +| openai-codex-cli-bin | 0.147.0 current compatible | Codex runtime dependency | Exact runtime dependency selected by openai-codex 0.147.0; standalone 0.149.0 is not independently usable in this lane. | +| google-antigravity | 0.1.15 current | Antigravity runtime adapter | Official Google Antigravity Python SDK for local agent harness integration. | | anyio or asyncio | stdlib plus optional anyio | Async runtime compatibility | Vendor SDKs are async; the public API should be async-first. | | OpenTelemetry API | 1.42.1 current | Optional event/trace integration | Mestre already normalizes agent events into span-event-shaped payloads; community users will expect observability hooks. | @@ -125,10 +125,10 @@ capabilities needed for real work. | Package | Current Version Checked | Python | Notes | |---------|-------------------------|--------|-------| -| claude-agent-sdk | 0.2.96 | >=3.10 | Newer than Mestre's pinned 0.2.91; tests must detect option-surface drift. | -| openai-codex | 0.1.0b3 | >=3.10 | Beta package; isolate Codex SDK API drift behind adapter boundaries. | -| openai-codex-cli-bin | 0.136.0 | >=3.10 | Runtime dependency for Codex SDK. | -| google-antigravity | 0.1.2 | >=3.10 | Includes compiled runtime wheels; install from PyPI rather than source checkout. | +| claude-agent-sdk | 0.2.148 | >=3.10 | Adapter-contract tests must detect option-surface drift. | +| openai-codex | 0.147.0 | >=3.10 | Pre-1.0 package; isolate Codex SDK API drift behind adapter boundaries. | +| openai-codex-cli-bin | 0.147.0 | >=3.10 | Exact runtime dependency for openai-codex 0.147.0. | +| google-antigravity | 0.1.15 | >=3.10 | Includes compiled runtime wheels; install from PyPI rather than source checkout. | | google-genai | 2.8.0 | >=3.10 | Not a core runtime adapter dependency unless Antigravity or future Google direct paths need it. | ## Sources @@ -163,14 +163,16 @@ Architecture not yet mapped. Follow existing patterns found in the codebase. ## Project Skills -No project skills found. Add skills to any of: `.claude/skills/`, `.agents/skills/`, `.cursor/skills/`, `.github/skills/`, or `.codex/skills/` with a `SKILL.md` index file. +- `.codex/skills/agent-runtime-kit-upgrade/SKILL.md` owns the report-first, + no-cooloff vendor SDK discovery and upgrade workflow. Use it for SDK + freshness checks, compatibility repairs, local upgrades, and upgrade PRs. ## GSD Workflow Enforcement -Before using Edit, Write, or other file-changing tools, start work through a GSD command so planning artifacts and execution context stay in sync. +Before using Edit, Write, or other file-changing tools, start work through a GSD command when a GSD entry point is actually installed so planning artifacts and execution context stay in sync. Use these entry points: @@ -178,7 +180,14 @@ Use these entry points: - `/gsd-debug` for investigation and bug fixing - `/gsd-execute-phase` for planned phase work -Do not make direct repo edits outside a GSD workflow unless the user explicitly asks to bypass it. +If none of these entry points is installed, do not fabricate or wait indefinitely +for a nonexistent wrapper. For an explicitly authorized change, record the plan +with the available task-planning mechanism, make the scoped edit, and validate it +with the repository's normal tests. Report that fallback in the handoff. + +Do not make direct repo edits outside an available GSD workflow unless the user +explicitly asks to bypass it or the unavailable-entry-point fallback above +applies. diff --git a/CHANGELOG.md b/CHANGELOG.md index 518a23b..a0fd38e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,55 @@ All notable changes to this project are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## 0.5.1 - 2026-08-29 + +### Added + +- SDK evolution reports now include recent-release implementation evidence: + shipped-file fingerprints, Python-definition diffs, opaque-runtime limits, + and implementation trends grounded in observed upstream artifacts. +- Exact baseline/candidate behavior probes now exercise adapter-owned semantic + contracts in reusable, credential-scrubbed disposable environments. +- Evolution artifacts now preserve independent upstream discovery, current and + prospective resolver results, exact candidate classifications, implementation + diffs, and stronger current-state provenance. + +### Changed + +- Updated tested vendor runtimes from Claude Agent SDK `0.2.106` to `0.2.148`, + OpenAI Codex SDK `0.1.0b3` to `0.147.0`, its coupled CLI binary `0.137.0a4` + to `0.147.0`, and Google Antigravity SDK `0.1.4` to `0.1.15`. +- Raised the validated Codex SDK range to `<0.148`; the standalone Codex CLI + `0.149.0` remains an inspected artifact rather than an install candidate + because Codex SDK `0.147.0` requires CLI `0.147.0` exactly. +- SDK freshness checks now bypass repository and ambient UV release cooloffs so + upstream discovery, prospective resolution, and applied locks describe the + same current package state. +- Direction-of-travel analysis now reports trends in actual upstream + implementations instead of emitting dependency upgrade or hold advice. +- The repository-owned upgrade skill is the canonical operator workflow for + report-first discovery, gated implementation, and separately authorized PR + publication. + +### Fixed + +- Candidate discovery no longer treats the current dependency bounds as proof + that no newer upstream release exists; excluded releases receive an + independent prospective resolver check. +- Missing optional SDKs, failed imports, skipped probes, malformed evidence, and + mismatched baseline/candidate versions can no longer produce a false-green + compatibility result or bypass implementation gates. +- Dependency application now updates project constraints, `uv.lock`, and the + compatibility manifest atomically, verifies exact inspected versions, and + restores all three artifacts if resolution or validation fails. +- Antigravity now maps short public session IDs deterministically to + provider-safe UUID conversation IDs while preserving the caller-visible + session identifier. +- Codex CLI binary snapshots use the correct import surface, and publication is + skipped for empty, unapplied, rolled-back, or unverified changes. +- Candidate installation and release-note collection now retry bounded transient + failures and report progress instead of appearing stalled. + ## 0.5.0 - 2026-07-11 ### Added diff --git a/docs/sdk-evolution-agent-design.md b/docs/sdk-evolution-agent-design.md index b5f20d9..f952913 100644 --- a/docs/sdk-evolution-agent-design.md +++ b/docs/sdk-evolution-agent-design.md @@ -14,7 +14,8 @@ or manual design stop. The SDK evolution agent should answer these questions for every run: - What package versions are installed, locked, and available upstream? -- Which packages does the resolver actually want to update? +- Which newer upstream releases exist, which are excluded by current bounds, + and which exact versions are prospectively resolvable together? - What changed in public API shape? - What changed in documented behavior or product direction? - Which adapter behavior contracts still pass on the candidate versions? @@ -75,18 +76,22 @@ Step responsibilities: `pyproject.toml`, `uv.lock`, installed distributions, package metadata, configured source hints, local environment facts, and supported auth availability. This produces raw facts, not recommendations. -- **Resolve update candidates**: Run the targeted resolver preview with - freshness cutoffs removed. This step decides which packages are real update - candidates for the run. It should use resolver output rather than only PyPI - `latest` metadata, especially for prerelease packages. -- **Inspect current and candidate APIs**: Load API snapshot and diff artifacts - from the last update run, then focus new inspection on packages that the - resolver selected for update or packages whose evidence is missing, stale, or - incompatible with the current evidence schema. This step owns API snapshot and - API diff artifacts. If the evidence signature changes, the agent may need to - refresh the current-state snapshot or gather more current-state data before - comparing candidates. If an update candidate has no candidate API diff, the - run should not proceed to implementation. +- **Resolve update candidates**: Build the candidate inventory from independent + upstream package metadata before consulting current constraints. Run both the + current targeted resolver preview and, when an upper bound excludes a newer + release, a no-mutation prospective preview in a temporary manifest with only + the excluding upper bound widened. Resolver output classifies feasibility; it + must not define away candidates hidden by the current manifest. +- **Inspect current and candidate APIs**: Treat `uv.lock` as the baseline and + inspect every exact compatible candidate. Reuse one credential-scrubbed + package/version environment for its snapshot and behavior probe. If an update + candidate has no exact candidate API diff, the run must not proceed to + implementation. +- **Inspect recent implementations**: When explicit candidate inspection is + enabled, inspect the three most recent published releases independently of + whether an update is pending. Fingerprint shipped files and Python AST + definitions, compare adjacent releases, and preserve opaque-artifact + limitations. This evidence is longitudinal analysis, not resolver input. - **Collect changelog and release-note evidence**: Fetch or read official changelogs, release pages, docs changelogs, repository releases, and package metadata links. This step records what changed according to the vendor and @@ -100,10 +105,12 @@ Step responsibilities: snapshots, API diffs, release-note evidence, behavior probe results, source references, and uncertainty into a compact bundle for the AI stages. This step should preserve provenance so later reasoning can be traced back to evidence. -- **Direction analysis through agent-runtime-kit**: Ask a runtime, via - `AgentTask`, to infer direction-of-travel themes from the evidence. This step - identifies whether changes look isolated or part of a broader SDK direction, - but it does not own the concrete implementation plan. +- **Implementation-trend analysis through agent-runtime-kit**: Ask a runtime, + via `AgentTask`, to infer implementation patterns from exact recent-release + file and definition diffs, corroborated by API, behavior, and release-note + evidence. This step identifies whether changes look isolated or part of a + broader SDK implementation direction. It must not recommend upgrade, hold, + resolver, or lockfile actions; it does not own the implementation plan. - **Architecture decision and update plan through agent-runtime-kit**: Ask a runtime, via `AgentTask`, to turn direction analysis into the concrete plan: adapter-only, test-only, docs-only, capability metadata change, @@ -123,26 +130,26 @@ Step responsibilities: report with the evidence bundle, analysis, decision, reviewer output, uncertainty, blocked reasons, and the exact manual review questions. This is a valid end state, not a failed run. -- **Apply safe implementation**: Apply only the changes allowed by the accepted - architecture decision and deterministic gates. This may include lockfile - updates, adapter changes, tests, docs, examples, compatibility shims, or report - changes. It must not implement changes that were classified as - `manual_design_required`. +- **Apply safe implementation**: The built-in deterministic implementation lane + widens only excluding dependency upper bounds, refreshes the lock to the exact + inspected compatible versions, and rolls both files back on mismatch or + verification failure. Adapter, public API, test, or documentation changes are + scoped coding work triggered by a blocked report; they require regression + evidence and a green rerun before the dependency lane proceeds. - **Run verification**: Run the verification commands required by the architecture decision. At minimum, this should cover formatting/linting, typing, unit tests, lock checks, report generation checks, and any available live smoke needed for the affected runtime behavior. - **Promote updated state to current baseline**: After implementation and - verification pass, save the updated lock/package/API/release-note/probe state - as the new current-state baseline for the next run. This promotion should be - explicit, atomic, and tied to the verified commit or workspace state. Failed, - blocked, or manual-design-required runs must not replace the current baseline. + verification pass, record the updated lock/package/API/release-note/probe state + in `current_state.json`, tied to the verified workspace state. Failed, blocked, + or manual-design-required runs are never marked promoted. - **Write report and optional draft PR**: Write the final local report with evidence, decisions, implementation summary, baseline-promotion result, test results, uncertainty, and manual checklist. If explicitly configured and authenticated, create or update a draft PR. This step must never auto-merge. -Every box before direction analysis is deterministic. AI stages may interpret +Every box before implementation-trend analysis is deterministic. AI stages may interpret evidence, but they should not invent evidence that was not collected. ## Operating Modes @@ -176,7 +183,7 @@ uv run --locked --extra antigravity python -m examples.sdk_evolution_agent \ Before a Codex-backed run, prepare the dedicated SDK evolution auth home: ```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ +env -u UV_EXCLUDE_NEWER \ uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex ``` @@ -241,30 +248,24 @@ The agent checks: - `pyproject.toml` dependency declarations. - `uv.lock` versions. - Installed distributions in the local environment. -- PyPI metadata and recent releases. -- `uv lock --dry-run -P ...` output with freshness cutoffs removed. +- Independent PyPI latest and recent-release metadata. +- Current constrained and prospective temporary-manifest + `uv lock --dry-run --exclude-newer false -P ...` output. -`uv lock --dry-run` is the source of truth for update candidates when it is -available. PyPI `latest` metadata is useful context, but it can be misleading -for prerelease packages. For example, a locked prerelease can be newer than the -stable value reported by package metadata. +Upstream metadata is the source of truth for whether a newer direct SDK release +exists. Prospective resolution is the source of truth for the compatible set +that can be applied. For `openai-codex-cli-bin`, the compatible candidate is the +exact version required by the Codex SDK candidate; a newer standalone CLI +release remains visible but is classified as blocked by that coupling. ### 2. API Shape Evidence -The agent should treat the lockfile as the current SDK baseline. If the active -Python environment has drifted from `uv.lock`, the agent inspects the locked -baseline in an isolated virtualenv instead of using the installed package. API -inspection artifacts are reusable evidence from the last update run when their -schema, lockfile version, and artifact hashes still match. A normal run starts -by loading the prior `api_snapshots/` and `api_diffs.json` artifacts, then -inspects only the packages that need fresh facts: - -- packages selected by the resolver for update, -- packages whose prior artifacts are missing, -- packages whose prior artifacts were produced by an older evidence schema, -- packages whose current locked or installed version no longer matches the - artifact baseline, -- packages needed to answer a specific adapter-compatibility question. +The agent treats the lockfile as the current SDK baseline. If the active Python +environment has drifted from `uv.lock`, it inspects the locked baseline in an +isolated virtualenv instead of using the installed package. Each run captures +fresh evidence for every selected package and every independently discovered +compatible candidate. A package/version environment is cached only within the +run, so its API snapshot and behavior probe reuse one exact install. For importable packages, snapshots record: @@ -283,15 +284,9 @@ This catches obvious adapter risks: API shape is necessary but insufficient. It does not prove behavior. -After a successful implementation, the candidate API snapshots and diffs that -were verified must be promoted to the current-state baseline. That ensures the -next run compares new upstream candidates against the SDK state that was -actually accepted, not against stale pre-update artifacts. - -If the evidence schema changes, promotion should include a schema refresh of the -current package state even when the package version did not change. Otherwise -future runs may compare candidate evidence against artifacts that no longer mean -the same thing. +After a successful implementation, `uv.lock` becomes the next run's +authoritative baseline and `current_state.json` records the evidence that was +accepted. ### 3. Changelog and Release-Note Evidence @@ -344,7 +339,7 @@ Behavior probes should cover these contracts: | Contract | Why API diffs are not enough | Example probe | | --- | --- | --- | -| Request construction | Constructor signatures can stay stable while fields change meaning. | Assert adapter builds expected SDK options/config objects. | +| Request construction | Signatures can stay stable while validators or field meaning changes. | Construct the exact SDK options/config object with adapter-owned values; Antigravity includes provider-compatible mapped conversation IDs. | | Permission mapping | Permission mode names can stay present while policy behavior changes. | Strict/default/permissive tests for each adapter. | | Sandbox and workspace semantics | Behavior can shift across SDK or CLI layers without a Python signature change. | Codex sandbox enum and run argument contract tests, plus smoke where possible. | | Streaming and event order | New message types may not break imports but may be dropped. | Feed fake vendor messages and assert emitted event order. | @@ -371,7 +366,7 @@ Each probe result should include: - skipped reason when optional credentials are missing. `behavior_diffs.json` compares observed locked-baseline probes against -candidate-version probes for resolver-selected updates. Breaking candidate +candidate-version probes for independently inventoried compatible updates. Breaking candidate probe changes block implementation deterministically before any local lock update. @@ -411,12 +406,19 @@ sequenceDiagram The AI stages should receive compacted, source-referenced evidence. They should not be asked to inspect the filesystem directly during report-only analysis. +The direction stage receives implementation diffs as primary evidence. Its +structured package status is `observed`, `opaque-runtime`, `no-transition`, or +`unavailable`; package freshness and resolver status are not implementation +trends. A deterministic postcondition replaces operational advice if a runtime +returns it in a trend field. ## Decision Gates The agent should fail closed. Implementation is blocked when: -- the resolver reports an update but candidate API diffs are missing, +- independent discovery finds a candidate but the prospective no-cooloff + resolver preview is missing or failed, +- an exact candidate API diff is missing, - release notes exist but were not collected, - release notes are unavailable and the API or behavior evidence is ambiguous, - behavior probes fail, @@ -534,6 +536,8 @@ evidence.json release_notes.json api_snapshots/ api_diffs.json +implementation_snapshots/ +implementation_diffs.json behavior_probes.json behavior_diffs.json behavior_summary.json @@ -553,7 +557,7 @@ report.md - API diff count and affected packages, - behavior probe status, - current-state baseline promotion status, -- direction-of-travel themes, +- upstream implementation patterns across exact release intervals, - architecture decision, - reviewer status, - implementation result, @@ -568,7 +572,7 @@ artifact-aware. It should record: - commit SHA or explicit dirty-worktree marker, - lockfile hash, - package names and accepted current versions, -- paths or content hashes for current API snapshots, +- paths or content hashes for current API and implementation snapshots, - paths or content hashes for release-note evidence, - paths or content hashes for behavior probe results, - a path or content hash for the deterministic behavior summary, @@ -589,9 +593,11 @@ Promotion rules should be conservative: Changelogs are incomplete. They often omit small behavior changes and may lag package releases. -API snapshots are shallow. Python introspection can miss behavior encoded in -runtime binaries, generated models, callbacks, subprocesses, environment -variables, or remote services. +Public API snapshots are shallow. Implementation-history fingerprints add +file-level and Python-definition evidence, but hashes still cannot explain +behavior encoded in runtime binaries, generated models, callbacks, subprocesses, +environment variables, or remote services. Opaque artifacts must remain an +explicit limitation rather than being reverse-inferred from version numbers. Live probes are environment-sensitive. They prove that one local credential and runtime setup worked at one time. They do not replace unit or contract probes. @@ -657,10 +663,16 @@ The example implements the deterministic evidence artifacts described above: package versions, artifact hashes, and promotion status. The implementation path is gated by deterministic checks before the local -lockfile update runs. Missing candidate API diffs, unavailable required +manifest/lock update runs. A missing or failed prospective resolver preview, +missing candidate API diffs, unavailable required release-note evidence, behavior summaries that are failed, incomplete, missing, malformed, or unknown, reviewer rejection, `manual_design_required`, and unresolved recursive self-adaptation all block implementation. When -implementation is allowed, the example applies the resolver-selected SDK lock -update locally, runs verification, writes the report artifacts, commits them, -pushes the branch, and opens a draft PR when configured. +implementation is allowed, the example widens only excluding direct dependency +bounds, applies the exact prospectively selected SDK set with freshness cutoffs +disabled, updates the compatibility manifest to the exact tested SDK and +coupled-runtime versions, verifies the resolved versions, and runs lint, typing, +tests, and lock checks. It restores the project manifest, lockfile, and +compatibility manifest on mismatch or verification failure. Only explicitly +requested draft-PR mode may stage the reported changed paths, commit, push, and +open a PR, and only after a verified non-empty update. diff --git a/docs/sdk-evolution-agent.md b/docs/sdk-evolution-agent.md index 1533cd2..15789a4 100644 --- a/docs/sdk-evolution-agent.md +++ b/docs/sdk-evolution-agent.md @@ -42,7 +42,7 @@ directory is created with private permissions before the Codex runtime starts. Run the auth preflight before real Codex-backed SDK evolution runs: ```bash -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ +env -u UV_EXCLUDE_NEWER \ uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex ``` @@ -53,13 +53,14 @@ normal Codex login cache and rerun the helper: ```bash uv run --locked --extra codex codex login --device-auth -env -u UV_EXCLUDE_NEWER -u UV_EXCLUDE_NEWER_PACKAGE \ +env -u UV_EXCLUDE_NEWER \ uv run --locked --extra codex python -m examples.sdk_evolution_agent.auth ensure-codex ``` Codex-backed SDK evolution runs explicitly choose `gpt-5.5` with `reasoning_effort=xhigh` for the AI stages that analyze direction, decide the -update plan, implement allowed changes, and review the result. This model policy +update plan, and review the result. Dependency application and verification are +deterministic. This model policy is applied only to `codex-agent-sdk`; Claude and Antigravity runs keep their provider-native model selection because `gpt-5.5` is not a valid model override for those adapters. @@ -86,6 +87,8 @@ Each run writes a timestamped directory under `reports/sdk-evolution/` with: - `release_notes.json` - `api_snapshots/` - `api_diffs.json` +- `implementation_snapshots/` +- `implementation_diffs.json` - `behavior_probes.json` - `behavior_diffs.json` - `behavior_summary.json` @@ -98,7 +101,7 @@ Each run writes a timestamped directory under `reports/sdk-evolution/` with: - `report.md` The report separates deterministic facts from runtime-generated analysis and -calls out uncertainty, release-note coverage, API diffs, behavior diffs, +calls out uncertainty, release-note coverage, API diffs, implementation trends, behavior diffs, baseline promotion, recursive self-adaptation impact, implementation status, test results, reviewer output, and manual review items. @@ -112,53 +115,84 @@ installed distributions, then compares it with upstream package metadata for: - `openai-codex-cli-bin` - `google-antigravity` -When `--refresh-preview` is used, the targeted `uv lock --dry-run -P ...` -preview removes freshness cutoff environment variables, including -`UV_EXCLUDE_NEWER`, and passes `--exclude-newer-package =false` for -each monitored SDK/runtime package. The project declares the same package-scoped -exemptions, so approved lock updates and ordinary locked CI agree while all other -dependencies retain the repository's eight-day delay. +When `--refresh-preview` is used, the runner removes freshness-cutoff +environment variables and executes `uv lock --dry-run --exclude-newer false -P +...`. The repository no longer declares the retired eight-day cooloff or +package-specific exemptions, so candidate research, applied locks, and ordinary +locked CI use one policy. ## Candidate API Inspection -The command treats `uv.lock` as the current baseline. If the locked SDK is +The command treats `uv.lock` as the current baseline and PyPI package metadata +as an independent candidate source. If the locked SDK is missing from the active `.venv` or the installed version differs, the agent can inspect the locked baseline in a temporary isolated virtualenv instead of trusting the missing or drifted environment. -When a refresh preview is available, package update candidates come from the -resolver's `uv lock --dry-run -P ...` output, not only from PyPI's `latest` -metadata. With `--inspect-candidates`, the agent installs each missing or -drifted locked baseline and each resolver update candidate in a temporary -isolated virtualenv — with a credential-scrubbed environment (throwaway `HOME`, -`PATH` only) — and writes API snapshots plus an `api_diffs.json` entry, and runs -the behavior probes against the candidate the same way. This avoids false -downgrade diffs for packages whose locked -prerelease is newer than PyPI's stable latest field. Candidate inspection is -opt-in because it executes freshly downloaded upstream code; without the flag, -candidates are recorded as explicit `skip` entries rather than silently -missing evidence. - -If `uv lock --dry-run -P ...` reports an SDK update but the run cannot produce a -candidate-version API diff for that package, implementation is blocked and the +The current constrained preview is evidence, not the candidate inventory. When +an upstream release is excluded by an upper bound, the runner copies the +manifest and lock into a temporary workspace, widens only that excluding bound, +and performs a prospective no-cooloff preview. The original checkout is not +mutated during research. The Codex CLI candidate is the exact runtime selected +by the latest published Codex SDK. A newer standalone CLI wheel is recorded as +`sdk-coupled-no-update`: it is a release-staging artifact, not a blocked or +SDK-usable candidate. It may be inspected in a disposable environment for +implementation-history evidence, but it is never selected into the project +unless a published Codex SDK requires it. + +With `--inspect-candidates`, the agent installs each missing or drifted locked +baseline, every exact compatible candidate, and the three most recent published +releases in credential-scrubbed temporary virtualenvs (throwaway `HOME`, `PATH` +only). One package/version environment is reused wherever API, behavior, and +implementation-history inspection overlap, and transient metadata/install +failures are retried. Candidate inspection writes the exact compatibility +transition to `api_diffs.json` and executes the adapter contract against that +same version. Independently, recent-release inspection fingerprints shipped +files and Python AST definitions, compares adjacent releases, and writes the +result to `implementation_diffs.json`. Source is never copied into the report; +only paths, counts, sizes, and hashes are retained. Candidate inspection is +opt-in because it executes freshly downloaded upstream code; historical +inspection uses the same explicit consent. Without the flag, missing trend +evidence is reported as `no-transition` rather than being replaced with +dependency advice. + +## Upstream Implementation Trends + +`direction_analysis.json` is an implementation-trend report. It answers what +changed across actual recent SDK implementations: modules and files, Python +definitions, source size, runtime artifacts, public API, observed adapter +behavior, capabilities, and deprecations. It does not recommend upgrading, +holding, resolving, or changing the lockfile; those decisions belong to +`architecture_decision.json` and the deterministic implementation gates. + +The direction stage receives `implementation_diffs.json` as its primary +evidence. A deterministic postcondition replaces any release-operation advice +that leaks into an implementation-trend field. Opaque executables are reported +as opaque: artifact changes can be proved by hashes, but their internal design +cannot be inferred from a wheel. + +If the independent candidate inventory contains an SDK update but the run cannot +produce an exact candidate-version API diff for that package, implementation is blocked and the architecture decision is marked `manual_design_required`. An empty added / removed / changed diff is valid; a missing diff object is not. Behavior probes intentionally separate observed SDK surface churn from adapter -contract breakage. `behavior_probes.json` records fields and parameters seen in -current and candidate packages, while `behavior_diffs.json` compares the -required adapter contract. `behavior_summary.json` records the deterministic +contract breakage. `behavior_probes.json` records fields and parameters and +constructs provider configurations with adapter-owned values; this catches +validation changes that field-presence inspection misses. `behavior_diffs.json` +compares the required adapter contract. `behavior_summary.json` records the deterministic assessment, counts, and reasons used by the implementation gate. Its status is `pass` for complete unchanged evidence, `changed` for complete non-breaking changes, `incomplete` for probe errors, skips, malformed records, or missing exact-version comparisons, and `fail` for a failed required contract or a -breaking diff. A package with no resolver-selected update needs only a valid +breaking diff. A package with no compatible candidate update needs only a valid current baseline at the locked version (or the observed installed version when no lock entry exists); it does not need a candidate probe. An ambient SDK that has drifted away from the lock therefore produces `incomplete`, not `pass`. Optional field changes remain visible in the report and API diffs without being treated as a contract failure. -Resolver transitions are parsed once as exact `(package, from, to)` triples and +Candidate transitions are built once as exact `(package, from, to)` triples from +the independent package inventory plus successful prospective resolution, then shared by snapshot collection, behavior assessment, and implementation gates. An API diff or behavior comparison for a different package or version does not satisfy the expected transition. Snapshot import and execution errors are also @@ -184,7 +218,8 @@ Implementation is still blocked when: - the architecture decision sets `manual_design_required`, - the reviewer rejects the evidence or design, -- a resolver-selected update lacks a candidate API diff, +- the no-cooloff prospective resolver preview is missing or failed, +- an independently discovered compatible update lacks an exact candidate API diff, - required release-note evidence could not be collected, - `behavior_summary.json` is missing, malformed, has an unknown status, reports `fail` / `incomplete`, is internally inconsistent, or uses the wrong @@ -216,9 +251,14 @@ uv run --locked --extra claude python -m examples.sdk_evolution_agent \ --draft-pr ``` -When `--draft-pr` is set, the agent stages `uv.lock` and the run report -directory, commits them with `--commit-message`, pushes the branch, and opens a -draft PR with `gh`. It never auto-merges. +When `--draft-pr` is set, publication occurs only after an applied, verified, +non-empty update. The agent stages only the `changed_paths` reported by the +implementation (normally `pyproject.toml`, `uv.lock`, and +`src/agent_runtime_kit/compatibility.py`), commits them with `--commit-message`, +pushes the branch, and opens a draft PR with `gh`. The gitignored report is +embedded in the PR body rather than blindly staged. A requested PR is skipped +when the update was blocked, empty, rolled back, or failed verification. It +never auto-merges. The command uses local Git and `gh` authentication. It never auto-merges, auto-publishes, or scrapes unsupported credentials. diff --git a/examples/sdk_evolution_agent/behavior.py b/examples/sdk_evolution_agent/behavior.py index 11d9d2f..a11e856 100644 --- a/examples/sdk_evolution_agent/behavior.py +++ b/examples/sdk_evolution_agent/behavior.py @@ -15,7 +15,12 @@ from pathlib import Path from typing import Any, cast -from examples.sdk_evolution_agent.collectors import ResolverTransition, parse_refresh_transitions +from agent_runtime_kit.adapters.antigravity import _provider_conversation_id +from examples.sdk_evolution_agent.collectors import ResolverTransition, candidate_transitions +from examples.sdk_evolution_agent.inspection import ( + CandidatePreparationError, + prepare_cached_candidate_environment, +) from examples.sdk_evolution_agent.models import BehaviorDiff, BehaviorProbeResult from examples.sdk_evolution_agent.snapshots import DEFAULT_MODULES, isolated_env @@ -33,6 +38,8 @@ "required_start_params", "run_params", "start_params", + "construction_cases", + "construction_failures", } ) _SENSITIVE_PATH_START_RE = re.compile( @@ -113,41 +120,69 @@ def probe_candidate_in_venv( step = "virtual-environment-creation" try: - with tempfile.TemporaryDirectory(prefix="ark-sdk-behavior-") as directory: - venv = Path(directory) / ".venv" - # Scrub the environment for every subprocess that touches freshly - # downloaded upstream code (same scrub as the API snapshots): a - # throwaway HOME and only PATH, so a malicious or buggy candidate - # package cannot read the caller's credentials/config. - env = isolated_env(Path(directory)) - subprocess.run( - (python, "-m", "venv", str(venv)), - check=True, - text=True, - capture_output=True, - timeout=timeout, - env=env, - ) - bin_dir = "Scripts" if sys.platform == "win32" else "bin" - venv_python = venv / bin_dir / "python" - step = "package-installation" - subprocess.run( - (str(venv_python), "-m", "pip", "install", f"{package}=={version}"), - check=True, - text=True, - capture_output=True, - timeout=timeout, - env=env, - ) + cached = prepare_cached_candidate_environment( + package, + version, + python=python, + timeout=timeout, + runner=subprocess.run, + env_factory=isolated_env, + ) + if cached is not None: step = "probe-execution" - completed = subprocess.run( - (str(venv_python), "-c", _PROBE_SCRIPT, package, version, scope), - check=True, - text=True, - capture_output=True, + completed = _run_probe_script( + cached.python, + package=package, + version=version, + scope=scope, timeout=timeout, - env=env, + env=cached.env, ) + else: + with tempfile.TemporaryDirectory(prefix="ark-sdk-behavior-") as directory: + venv = Path(directory) / ".venv" + # Scrub the environment for every subprocess that touches freshly + # downloaded upstream code (same scrub as the API snapshots): a + # throwaway HOME and only PATH, so a malicious or buggy candidate + # package cannot read the caller's credentials/config. + env = isolated_env(Path(directory)) + subprocess.run( + (python, "-m", "venv", str(venv)), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=env, + ) + bin_dir = "Scripts" if sys.platform == "win32" else "bin" + venv_python = venv / bin_dir / "python" + step = "package-installation" + subprocess.run( + (str(venv_python), "-m", "pip", "install", f"{package}=={version}"), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=env, + ) + step = "probe-execution" + completed = _run_probe_script( + venv_python, + package=package, + version=version, + scope=scope, + timeout=timeout, + env=env, + ) + except CandidatePreparationError as exc: + cause = exc.cause + if isinstance(cause, subprocess.TimeoutExpired): + detail = f"timed out after {cause.timeout}s" + elif isinstance(cause, subprocess.CalledProcessError): + detail = _bounded_text(cause.stderr or cause.stdout or str(cause)) + else: + detail = _bounded_text(cause) + return (_probe_execution_failure(package, version, scope, exc.step, detail),) except subprocess.TimeoutExpired as exc: return ( _probe_execution_failure( @@ -187,6 +222,33 @@ def probe_candidate_in_venv( ) +def _run_probe_script( + python: Path, + *, + package: str, + version: str, + scope: str, + timeout: int, + env: Mapping[str, str], +) -> subprocess.CompletedProcess[str]: + return subprocess.run( + ( + str(python), + "-c", + _PROBE_SCRIPT, + package, + version, + scope, + json.dumps(_probe_inputs(package), sort_keys=True), + ), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=dict(env), + ) + + def diff_behavior_results(results: Sequence[BehaviorProbeResult]) -> tuple[BehaviorDiff, ...]: """Compare current and candidate behavior probes for each package/probe.""" @@ -285,7 +347,7 @@ def behavior_expectations_from_evidence(evidence: Mapping[str, object]) -> dict[ issues.append(f"deterministic evidence package {index} must have a non-empty name") continue packages.append(package) - expectations = build_behavior_expectations(packages, parse_refresh_transitions(evidence)) + expectations = build_behavior_expectations(packages, candidate_transitions(evidence)) expectations["expectation_issues"] = issues return expectations @@ -751,7 +813,11 @@ def _is_probe_execution_error(result: BehaviorProbeResult) -> bool: def _has_contract_failure_evidence(result: BehaviorProbeResult) -> bool: missing = result.details.get("missing") - return bool(missing) and _is_sequence_payload(missing) + construction_failures = result.details.get("construction_failures") + return bool( + (bool(missing) and _is_sequence_payload(missing)) + or (bool(construction_failures) and _is_sequence_payload(construction_failures)) + ) def _probe_package( @@ -928,21 +994,55 @@ def _probe_antigravity(*, version: str | None, scope: str) -> BehaviorProbeResul "mcp_servers", } missing = sorted(expected - fields) + construction_cases = _probe_inputs(package)["conversation_ids"] + construction_failures = _antigravity_construction_failures(config_cls, construction_cases) + failed = bool(missing or construction_failures) return BehaviorProbeResult( package=package, version=version, scope=scope, probe="adapter-contract", - status="fail" if missing else "pass", + status="fail" if failed else "pass", summary=( - "Antigravity LocalAgentConfig exposes required adapter fields." - if not missing - else "Antigravity LocalAgentConfig is missing required adapter fields." + "Antigravity LocalAgentConfig exposes and constructs the required adapter contract." + if not failed + else "Antigravity LocalAgentConfig rejected the required adapter contract." ), - details={"fields": sorted(fields), "required_fields": sorted(expected), "missing": missing}, + details={ + "fields": sorted(fields), + "required_fields": sorted(expected), + "missing": missing, + "construction_cases": sorted(construction_cases), + "construction_failures": construction_failures, + }, ) +def _probe_inputs(package: str) -> dict[str, list[str]]: + if package != "google-antigravity": + return {} + mapped = _provider_conversation_id("conv-77") + return { + "conversation_ids": [ + mapped or "", + "a" * 32, + ] + } + + +def _antigravity_construction_failures( + config_cls: Any, + conversation_ids: Sequence[str], +) -> list[str]: + failures: list[str] = [] + for index, conversation_id in enumerate(conversation_ids): + try: + config_cls(conversation_id=conversation_id) + except Exception as exc: + failures.append(f"case-{index}:{type(exc).__name__}:{_bounded_text(exc, limit=160)}") + return failures + + def _fields(cls: Any) -> set[str]: if hasattr(cls, "model_fields"): return set(cls.model_fields) @@ -971,6 +1071,14 @@ def _contract_details(result: BehaviorProbeResult) -> dict[str, Any]: contract["required_start_params"] = _normalized_contract_field( details.get("required_start_params") ) + if "construction_cases" in details: + contract["construction_cases"] = _normalized_contract_field( + details.get("construction_cases") + ) + if "construction_failures" in details: + contract["construction_failures"] = _normalized_contract_field( + details.get("construction_failures") + ) return contract @@ -1109,6 +1217,7 @@ def _string_or_none(value: object) -> str | None: from pathlib import Path package, version, scope = sys.argv[1:4] + probe_inputs = json.loads(sys.argv[4]) if len(sys.argv) > 4 else {} sensitive_path_start_re = re.compile( r"(?P^|[^A-Za-z0-9_.-])" r"(?P(?:[A-Za-z][A-Za-z0-9+.-]*://|[A-Za-z]:[\\/]|\\\\|/))" @@ -1245,7 +1354,8 @@ def result(probe, status, summary, details): config_module = importlib.import_module( "google.antigravity.connections.local.local_connection_config" ) - config_fields = fields(getattr(config_module, "LocalAgentConfig")) + config_cls = getattr(config_module, "LocalAgentConfig") + config_fields = fields(config_cls) expected = { "model", "api_key", "vertex", "project", "location", "system_instructions", "capabilities", "policies", "workspaces", @@ -1253,15 +1363,28 @@ def result(probe, status, summary, details): "mcp_servers", } missing = sorted(expected - config_fields) + conversation_ids = probe_inputs.get("conversation_ids", []) + construction_failures = [] + for index, conversation_id in enumerate(conversation_ids): + try: + config_cls(conversation_id=conversation_id) + except Exception as exc: + construction_failures.append( + f"case-{index}:{type(exc).__name__}:{str(exc)[:160]}" + ) + contract_failed = bool(missing or construction_failures) payload = [result( "adapter-contract", - "fail" if missing else "pass", - "Antigravity LocalAgentConfig exposes required adapter fields." if not missing - else "Antigravity LocalAgentConfig is missing required adapter fields.", + "fail" if contract_failed else "pass", + "Antigravity LocalAgentConfig exposes and constructs the required adapter contract." + if not contract_failed + else "Antigravity LocalAgentConfig rejected the required adapter contract.", { "fields": sorted(config_fields), "required_fields": sorted(expected), "missing": missing, + "construction_cases": sorted(conversation_ids), + "construction_failures": construction_failures, }, )] else: diff --git a/examples/sdk_evolution_agent/cli.py b/examples/sdk_evolution_agent/cli.py index 77cc121..ca03478 100644 --- a/examples/sdk_evolution_agent/cli.py +++ b/examples/sdk_evolution_agent/cli.py @@ -7,19 +7,25 @@ from pathlib import Path from typing import Any +from packaging.version import InvalidVersion, Version + from agent_runtime_kit import AgentRuntime, RuntimeRegistry from examples.sdk_evolution_agent.behavior import collect_behavior_evidence from examples.sdk_evolution_agent.collectors import ( CommandRunner, PypiClient, + candidate_transitions, + candidate_update_versions, collect_evidence, - parse_refresh_transitions, - refresh_update_versions, + read_uv_lock_versions, run_lock_update, run_verification_commands, + update_compatibility_manifest, + widen_project_dependency_bounds, ) from examples.sdk_evolution_agent.current_state import build_current_state from examples.sdk_evolution_agent.events import JsonlEventSink +from examples.sdk_evolution_agent.inspection import candidate_environment_cache from examples.sdk_evolution_agent.models import ( DEFAULT_PACKAGES, RunContext, @@ -35,8 +41,9 @@ stage_paths, ) from examples.sdk_evolution_agent.release_notes import collect_release_notes -from examples.sdk_evolution_agent.report import write_run_report +from examples.sdk_evolution_agent.report import write_json, write_run_report from examples.sdk_evolution_agent.snapshots import ( + diff_implementation_snapshot_groups, diff_snapshot_groups, snapshot_candidate_in_venv, snapshot_current_api, @@ -177,6 +184,7 @@ async def run_agent( raise RuntimeError( f"failed to create branch {options.branch_name}: {branch_result.stderr}" ) + _progress("collecting independent upstream and resolver evidence") evidence = collect_evidence( options.workspace, packages=options.packages, @@ -184,26 +192,48 @@ async def run_agent( pypi_client=pypi_client, command_runner=command_runner, ) - transitions = parse_refresh_transitions(evidence) - update_versions = refresh_update_versions(evidence) - snapshots = _collect_snapshots(evidence, inspect_candidates=options.inspect_candidates) - api_diffs = [to_jsonable(diff) for diff in diff_snapshot_groups(snapshots)] - release_notes = [ - to_jsonable(item) - for item in collect_release_notes(evidence.get("packages", []), update_versions) - ] - behavior = to_jsonable( - collect_behavior_evidence( - evidence.get("packages", []), - update_versions, - inspect_candidates=options.inspect_candidates, - expected_transitions=transitions, + transitions = candidate_transitions(evidence) + update_versions = candidate_update_versions(evidence) + _progress( + "candidate inventory complete: " + + ( + ", ".join(f"{name}=={version}" for name, version in sorted(update_versions.items())) + or "no newer SDK releases" ) ) + with candidate_environment_cache(): + _progress("collecting current and candidate API snapshots") + snapshots = _collect_snapshots(evidence, inspect_candidates=options.inspect_candidates) + api_diffs = [to_jsonable(diff) for diff in diff_snapshot_groups(snapshots)] + _progress("collecting recent upstream implementation history") + implementation_snapshots = _collect_implementation_snapshots( + evidence, + compatibility_snapshots=snapshots, + inspect_candidates=options.inspect_candidates, + ) + implementation_diffs = [ + to_jsonable(diff) + for diff in diff_implementation_snapshot_groups(implementation_snapshots) + ] + _progress("collecting release notes and executable behavior probes") + release_notes = [ + to_jsonable(item) + for item in collect_release_notes(evidence.get("packages", []), update_versions) + ] + behavior = to_jsonable( + collect_behavior_evidence( + evidence.get("packages", []), + update_versions, + inspect_candidates=options.inspect_candidates, + expected_transitions=transitions, + ) + ) + _progress("running direction, architecture, and review stages") direction, architecture, review = await run_analysis_pipeline( selected_runtime, evidence=evidence, api_diffs=api_diffs, + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, context=RunContext( @@ -231,6 +261,7 @@ async def run_agent( config["event_log_path"] = str(context.event_log_path) if options.implementation_enabled and implementation.get("allowed"): + _progress("applying inspected dependency candidates") implementation = _run_local_sdk_update( options, update_versions=update_versions, @@ -251,6 +282,10 @@ async def run_agent( evidence=evidence, snapshots=[to_jsonable(snapshot) for snapshot in snapshots], api_diffs=api_diffs, + implementation_snapshots=[ + to_jsonable(snapshot) for snapshot in implementation_snapshots + ], + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, current_state=current_state, @@ -263,9 +298,7 @@ async def run_agent( context, promoted=promoted, status=( - "promoted" - if promoted - else str(implementation.get("blocked_reason") or "skipped") + "promoted" if promoted else str(implementation.get("blocked_reason") or "skipped") ), implementation=implementation, ) @@ -275,6 +308,10 @@ async def run_agent( evidence=evidence, snapshots=[to_jsonable(snapshot) for snapshot in snapshots], api_diffs=api_diffs, + implementation_snapshots=[ + to_jsonable(snapshot) for snapshot in implementation_snapshots + ], + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, current_state=current_state, @@ -284,10 +321,12 @@ async def run_agent( review=review, ) - if options.draft_pr: + if _should_create_pr(options.draft_pr, implementation): + _progress("creating verified draft pull request") git_results = _create_autonomous_pr( options.workspace, report_path=report_path, + changed_paths=tuple(implementation.get("changed_paths") or ()), options=options, command_runner=command_runner, ) @@ -298,6 +337,10 @@ async def run_agent( evidence=evidence, snapshots=[to_jsonable(snapshot) for snapshot in snapshots], api_diffs=api_diffs, + implementation_snapshots=[ + to_jsonable(snapshot) for snapshot in implementation_snapshots + ], + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, current_state=current_state, @@ -312,6 +355,42 @@ async def run_agent( options=options, command_runner=command_runner, ) + elif options.draft_pr: + implementation["pr_skipped_reason"] = ( + "draft PR requires an applied, verified, non-empty change" + ) + report_path = _write_full_report( + context, + config=config, + evidence=evidence, + snapshots=[to_jsonable(snapshot) for snapshot in snapshots], + api_diffs=api_diffs, + implementation_snapshots=[ + to_jsonable(snapshot) for snapshot in implementation_snapshots + ], + implementation_diffs=implementation_diffs, + release_notes=release_notes, + behavior=behavior, + current_state=current_state, + direction=direction, + architecture=architecture, + implementation=implementation, + review=review, + ) + # The report renders promotion state, so the report written immediately + # before build_current_state is necessarily provisional. Recompute the + # manifest after the final report write and persist only current_state; + # rewriting the report here would recreate the hash cycle. + current_state = build_current_state( + context, + promoted=promoted, + status=( + "promoted" if promoted else str(implementation.get("blocked_reason") or "skipped") + ), + implementation=implementation, + ) + write_json(context.report_root / "current_state.json", current_state) + _progress(f"report ready: {report_path}") return report_path finally: if close_owned_runtime: @@ -338,35 +417,100 @@ def _run_local_sdk_update( return { **implementation, "applied": False, - "blocked_reason": "no resolver-selected SDK updates", + "changed_paths": [], + "blocked_reason": "no newer upstream SDK candidates", } - update_result = run_lock_update( - options.workspace, - packages, - command_runner=command_runner, - ) + pyproject = options.workspace / "pyproject.toml" + lockfile = options.workspace / "uv.lock" + compatibility_manifest = options.workspace / "src" / "agent_runtime_kit" / "compatibility.py" + originals = { + path: path.read_bytes() if path.exists() else None + for path in (pyproject, lockfile, compatibility_manifest) + } results = list(implementation.get("verification_results") or []) - results.append(to_jsonable(update_result)) - applied = update_result.returncode == 0 changes = list(implementation.get("changes") or []) - if applied: - changes.append("Updated uv.lock for resolver-selected SDK packages: " + ", ".join(packages)) - verification_commands = tuple(DEFAULT_VERIFICATION_COMMANDS) + try: + bound_changes = widen_project_dependency_bounds(pyproject, update_versions) + changes.extend(f"Widened dependency bound: {change}" for change in bound_changes) + update_result = run_lock_update( + options.workspace, + packages, + command_runner=command_runner, + ) + results.append(to_jsonable(update_result)) + if update_result.returncode != 0: + raise RuntimeError(update_result.stderr or update_result.stdout or "uv lock failed") + + locked_versions = read_uv_lock_versions(lockfile) + mismatches = { + package: (version, locked_versions.get(package)) + for package, version in update_versions.items() + if locked_versions.get(package) != version + } + if mismatches: + detail = ", ".join( + f"{package}: expected {expected}, resolved {actual or ''}" + for package, (expected, actual) in sorted(mismatches.items()) + ) + raise RuntimeError(f"lockfile did not contain inspected candidate versions ({detail})") + + manifest_changes = update_compatibility_manifest( + compatibility_manifest, + pyproject, + update_versions, + ) + changes.extend(f"Updated compatibility manifest: {change}" for change in manifest_changes) + + changed_paths = [ + str(path.relative_to(options.workspace)) + for path, original in originals.items() + if (path.read_bytes() if path.exists() else None) != original + ] + if not changed_paths: + raise RuntimeError("resolver reported success but produced no workspace changes") + changes.append("Updated inspected SDK packages: " + ", ".join(packages)) verification_results = run_verification_commands( options.workspace, - verification_commands, + tuple(DEFAULT_VERIFICATION_COMMANDS), command_runner=command_runner, ) results.extend(to_jsonable(verification_results)) + failed = [result for result in verification_results if result.returncode != 0] + if failed: + failed_result = failed[0] + command = " ".join(failed_result.command) + raise RuntimeError( + f"verification failed: {command} (exit {failed_result.returncode}); " + "inspect verification_results for full output" + ) + except (OSError, RuntimeError, ValueError) as exc: + _restore_paths(originals) + return { + **implementation, + "applied": False, + "changes": changes, + "changed_paths": [], + "verification_results": results, + "blocked_reason": str(exc), + } return { **implementation, - "applied": applied, + "applied": True, "changes": changes, + "changed_paths": changed_paths, "verification_results": results, - "blocked_reason": "" if applied else update_result.stderr or update_result.stdout, + "blocked_reason": "", } +def _restore_paths(originals: dict[Path, bytes | None]) -> None: + for path, content in originals.items(): + if content is None: + path.unlink(missing_ok=True) + else: + path.write_bytes(content) + + def _write_full_report( context: RunContext, *, @@ -374,6 +518,8 @@ def _write_full_report( evidence: dict[str, Any], snapshots: list[dict[str, Any]], api_diffs: list[dict[str, Any]], + implementation_snapshots: list[dict[str, Any]], + implementation_diffs: list[dict[str, Any]], release_notes: list[dict[str, Any]], behavior: dict[str, Any], current_state: dict[str, Any], @@ -389,6 +535,8 @@ def _write_full_report( evidence=evidence, snapshots=snapshots, api_diffs=api_diffs, + implementation_snapshots=implementation_snapshots, + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, current_state=current_state, @@ -410,48 +558,64 @@ def _verification_passed(implementation: dict[str, Any]) -> bool: ) +def _should_create_pr(draft_pr: bool, implementation: dict[str, Any]) -> bool: + """Publish only a verified change that was actually applied.""" + + changed_paths = implementation.get("changed_paths") + return bool( + draft_pr + and implementation.get("allowed") + and implementation.get("applied") + and isinstance(changed_paths, list | tuple) + and changed_paths + and _verification_passed(implementation) + ) + + def _create_autonomous_pr( root: Path, *, report_path: Path, + changed_paths: tuple[str, ...], options: RunOptions, command_runner: CommandRunner | None, ) -> list[dict[str, Any]]: branch_name = options.branch_name or _current_branch(root, command_runner=command_runner) body = build_draft_pr_body(report_path.read_text(encoding="utf-8")) - # Only stage the lockfile. The report lives under the gitignored default - # report dir, so `git add`-ing it returned rc=1 and previously broke the - # --draft-pr flow; its content is already embedded in the PR body above. - paths = ("uv.lock",) - results = [ - to_jsonable(stage_paths(root, paths, command_runner=command_runner)), - to_jsonable( - commit_staged( - root, - message=options.commit_message, - command_runner=command_runner, - ) - ), - ] + results: list[dict[str, Any]] = [] + staged = stage_paths(root, changed_paths, command_runner=command_runner) + results.append(to_jsonable(staged)) + _require_command_success(staged, "stage SDK update") + committed = commit_staged( + root, + message=options.commit_message, + command_runner=command_runner, + ) + results.append(to_jsonable(committed)) + _require_command_success(committed, "commit SDK update") if branch_name: - results.append( - to_jsonable(push_branch(root, branch_name=branch_name, command_runner=command_runner)) - ) - results.append( - to_jsonable( - create_draft_pr( - root, - title=options.pr_title, - body=body, - base=options.pr_base, - head=branch_name, - command_runner=command_runner, - ) - ) + pushed = push_branch(root, branch_name=branch_name, command_runner=command_runner) + results.append(to_jsonable(pushed)) + _require_command_success(pushed, "push SDK update branch") + created = create_draft_pr( + root, + title=options.pr_title, + body=body, + base=options.pr_base, + head=branch_name, + command_runner=command_runner, ) + results.append(to_jsonable(created)) + _require_command_success(created, "create draft PR") return results +def _require_command_success(result: Any, action: str) -> None: + if result.returncode != 0: + detail = result.stderr or result.stdout or f"exit code {result.returncode}" + raise RuntimeError(f"failed to {action}: {detail}") + + def _commit_final_autonomous_pr_report( root: Path, *, @@ -483,9 +647,7 @@ def _commit_final_autonomous_pr_report( raise RuntimeError(f"failed to commit final autonomous PR report: {detail}") -def _report_dir_committable( - root: Path, path: str, *, command_runner: CommandRunner | None -) -> bool: +def _report_dir_committable(root: Path, path: str, *, command_runner: CommandRunner | None) -> bool: runner = command_runner if runner is None: from examples.sdk_evolution_agent.collectors import run_command @@ -520,15 +682,13 @@ def _collect_snapshots(evidence: dict[str, Any], *, inspect_candidates: bool = F # off, every snapshot uses the already-installed version via snapshot_current_api, # which imports nothing new. snapshots = [] - update_versions = refresh_update_versions(evidence) - refresh_preview_seen = evidence.get("refresh_preview") is not None + update_versions = candidate_update_versions(evidence) for package in evidence.get("packages", []): if not isinstance(package, dict): continue name = str(package.get("name")) locked = package.get("locked_version") installed = package.get("installed_version") - baseline = locked or installed if inspect_candidates and locked and locked != installed: snapshots.append(snapshot_candidate_in_venv(name, str(locked))) else: @@ -536,16 +696,74 @@ def _collect_snapshots(evidence: dict[str, Any], *, inspect_candidates: bool = F if not inspect_candidates: continue candidate = update_versions.get(name) - if candidate is None and not refresh_preview_seen: - latest = package.get("latest_version") - if latest and latest != baseline: - candidate = str(latest) if candidate: snapshots.append(snapshot_candidate_in_venv(name, candidate)) return snapshots +def _collect_implementation_snapshots( + evidence: dict[str, Any], + *, + compatibility_snapshots: list[Any], + inspect_candidates: bool = False, +) -> list[Any]: + """Inspect adjacent recent releases for implementation-trend evidence.""" + + by_exact_version = { + (snapshot.package, str(snapshot.version)): snapshot + for snapshot in compatibility_snapshots + if snapshot.version and not snapshot.import_error + } + history: list[Any] = [] + for package in evidence.get("packages", []): + if not isinstance(package, dict): + continue + name = str(package.get("name") or "") + if not name: + continue + recent = [str(version) for version in package.get("recent_versions", []) if version] + if not inspect_candidates: + matching = [ + snapshot for snapshot in compatibility_snapshots if snapshot.package == name + ] + history.extend(matching[:1]) + continue + versions = _ordered_versions(recent) + if len(versions) < 2: + versions = _ordered_versions( + [ + *recent, + str(package.get("locked_version") or ""), + str(package.get("candidate_version") or ""), + ] + ) + for version in versions: + snapshot = by_exact_version.get((name, version)) + if snapshot is None: + snapshot = snapshot_candidate_in_venv(name, version) + if not snapshot.import_error: + by_exact_version[(name, version)] = snapshot + history.append(snapshot) + return history + + +def _ordered_versions(versions: list[str]) -> list[str]: + unique = {version for version in versions if version} + + def key(version: str) -> tuple[int, Version | str]: + try: + return (1, Version(version)) + except InvalidVersion: + return (0, version) + + return sorted(unique, key=key) + + def _refresh_update_versions(evidence: dict[str, Any]) -> dict[str, str]: - """Compatibility wrapper around the shared exact transition parser.""" + """Compatibility wrapper around the independent candidate inventory.""" + + return candidate_update_versions(evidence) + - return refresh_update_versions(evidence) +def _progress(message: str) -> None: + print(f"[sdk-evolution] {message}", flush=True) diff --git a/examples/sdk_evolution_agent/collectors.py b/examples/sdk_evolution_agent/collectors.py index 6d888ae..ddf7fa8 100644 --- a/examples/sdk_evolution_agent/collectors.py +++ b/examples/sdk_evolution_agent/collectors.py @@ -7,13 +7,26 @@ import os import re import shlex +import shutil import subprocess +import tempfile +import time import urllib.request from collections.abc import Callable, Mapping, Sequence -from dataclasses import dataclass +from dataclasses import dataclass, replace from pathlib import Path from typing import Any +from packaging.requirements import InvalidRequirement, Requirement +from packaging.specifiers import InvalidSpecifier, Specifier +from packaging.utils import canonicalize_name +from packaging.version import InvalidVersion, Version + +try: + import tomllib +except ModuleNotFoundError: # pragma: no cover - exercised on Python 3.10 + import tomli as tomllib + from examples.sdk_evolution_agent.models import ( DEFAULT_PACKAGES, CommandResult, @@ -23,7 +36,7 @@ ) FRESHNESS_CUTOFF_ENV_VARS = ("UV_EXCLUDE_NEWER",) -SDK_EVOLUTION_EXCLUDE_NEWER_PACKAGE = "false" +SDK_EVOLUTION_EXCLUDE_NEWER = "false" _REFRESH_TRANSITION_RE = re.compile( r"^[ \t]*Update[ \t]+(?P[A-Za-z0-9_.-]+)[ \t]+" @@ -107,8 +120,19 @@ def collect_evidence( ], } if include_refresh_preview: - preview = run_refresh_preview(root, packages, command_runner=command_runner) - evidence["refresh_preview"] = to_jsonable(preview) + current_preview = run_refresh_preview(root, packages, command_runner=command_runner) + evidence["refresh_preview"] = to_jsonable(current_preview) + evidence["packages"] = _annotate_candidate_states(evidence) + if _needs_prospective_preview(evidence): + prospective = run_prospective_refresh_preview( + root, + packages, + candidate_update_versions(evidence), + command_runner=command_runner, + ) + evidence["current_refresh_preview"] = evidence["refresh_preview"] + evidence["refresh_preview"] = to_jsonable(prospective) + evidence["packages"] = _annotate_candidate_states(evidence) return evidence @@ -117,15 +141,47 @@ def read_pyproject_dependency_specs(path: Path) -> dict[str, str]: if not path.exists(): return {} - text = path.read_text(encoding="utf-8") + data = tomllib.loads(path.read_text(encoding="utf-8")) + monitored = {canonicalize_name(package): package for package in DEFAULT_PACKAGES} specs: dict[str, str] = {} - for package in DEFAULT_PACKAGES: - match = re.search(rf"{re.escape(package)}[^\"'\],]*", text) - if match: - specs[package] = match.group(0) + for raw in _project_requirement_strings(data, include_constraints=True): + try: + requirement = Requirement(raw) + except InvalidRequirement: + continue + package = monitored.get(canonicalize_name(requirement.name)) + if package is not None: + specs.setdefault(package, raw) return specs +def _project_requirement_strings( + data: Mapping[str, Any], + *, + include_constraints: bool, +) -> tuple[str, ...]: + """Return actual project requirement strings, never comments or keyword mentions.""" + + requirements: list[str] = [] + project = data.get("project") + if isinstance(project, Mapping): + direct = project.get("dependencies") + if isinstance(direct, list): + requirements.extend(item for item in direct if isinstance(item, str)) + optional = project.get("optional-dependencies") + if isinstance(optional, Mapping): + for group in optional.values(): + if isinstance(group, list): + requirements.extend(item for item in group if isinstance(item, str)) + if include_constraints: + tool = data.get("tool") + uv = tool.get("uv") if isinstance(tool, Mapping) else None + constraints = uv.get("constraint-dependencies") if isinstance(uv, Mapping) else None + if isinstance(constraints, list): + requirements.extend(item for item in constraints if isinstance(item, str)) + return tuple(requirements) + + def read_uv_lock_versions(path: Path) -> dict[str, str]: """Read package versions from uv.lock without depending on a TOML parser.""" @@ -168,6 +224,7 @@ def detect_package_versions( """Detect local and upstream version state for vendor SDK packages.""" states: list[PackageVersionState] = [] + metadata_by_package: dict[str, Mapping[str, Any]] = {} for package in packages: sources = list(PACKAGE_SOURCE_HINTS.get(package, ())) latest_version: str | None = None @@ -175,6 +232,7 @@ def detect_package_versions( unavailable_reason = "" try: metadata = pypi_client(package) + metadata_by_package[package] = metadata latest_version = str(metadata.get("info", {}).get("version") or "") if not latest_version: latest_version = None @@ -198,28 +256,89 @@ def detect_package_versions( note=unavailable_reason, ) ) + installed = installed_version(package) + baseline = locked_versions.get(package) or installed + candidate = latest_version if _is_newer_version(latest_version, baseline) else None states.append( PackageVersionState( name=package, pyproject_spec=pyproject_specs.get(package), locked_version=locked_versions.get(package), - installed_version=installed_version(package), + installed_version=installed, latest_version=latest_version, + candidate_version=candidate, + candidate_status=( + "upstream-newer-unresolved" + if candidate is not None + else "metadata-unavailable" + if latest_version is None + else "current" + ), recent_versions=recent_versions, sources=tuple(sources), unavailable_reason=unavailable_reason, ) ) + sdk_selected_version, sdk_requirement = _exact_runtime_dependency( + metadata_by_package.get("openai-codex"), + dependency="openai-codex-cli-bin", + ) + if sdk_selected_version is not None: + states = [ + replace( + state, + sdk_selected_version=sdk_selected_version, + sdk_requirement=sdk_requirement, + ) + if state.name == "openai-codex-cli-bin" + else state + for state in states + ] return tuple(states) +def _exact_runtime_dependency( + metadata: Mapping[str, Any] | None, + *, + dependency: str, +) -> tuple[str | None, str | None]: + """Return an exact dependency pin advertised by the latest SDK metadata.""" + + if not isinstance(metadata, Mapping): + return None, None + info = metadata.get("info") + requires_dist = info.get("requires_dist") if isinstance(info, Mapping) else None + if not isinstance(requires_dist, Sequence) or isinstance(requires_dist, str | bytes): + return None, None + for raw in requires_dist: + if not isinstance(raw, str): + continue + try: + requirement = Requirement(raw) + except InvalidRequirement: + continue + if canonicalize_name(requirement.name) != canonicalize_name(dependency): + continue + for specifier in requirement.specifier: + if specifier.operator == "==" and "*" not in specifier.version: + return specifier.version, raw + return None, None + + def fetch_pypi_metadata(package: str) -> Mapping[str, Any]: """Fetch PyPI JSON metadata for one package.""" url = f"https://pypi.org/pypi/{package}/json" request = urllib.request.Request(url, headers={"User-Agent": "agent-runtime-kit-sdk-evolution"}) - with urllib.request.urlopen(request, timeout=20) as response: - return json.loads(response.read().decode("utf-8")) + for attempt in range(3): + try: + with urllib.request.urlopen(request, timeout=20) as response: + return json.loads(response.read().decode("utf-8")) + except (OSError, TimeoutError): + if attempt == 2: + raise + time.sleep(0.25 * (attempt + 1)) + raise AssertionError("unreachable") def installed_version(package: str) -> str | None: @@ -259,16 +378,9 @@ def cutoff_free_env(env: Mapping[str, str] | None = None) -> tuple[dict[str, str def build_refresh_preview_command(packages: Sequence[str]) -> tuple[str, ...]: """Build the targeted uv refresh preview command.""" - command = ["uv", "lock", "--dry-run"] + command = ["uv", "lock", "--dry-run", "--exclude-newer", SDK_EVOLUTION_EXCLUDE_NEWER] for package in packages: command.extend(("-P", package)) - for package in packages: - command.extend( - ( - "--exclude-newer-package", - f"{package}={SDK_EVOLUTION_EXCLUDE_NEWER_PACKAGE}", - ) - ) return tuple(command) @@ -293,6 +405,33 @@ def run_refresh_preview( ) +def run_prospective_refresh_preview( + root: Path, + packages: Sequence[str], + candidates: Mapping[str, str], + *, + command_runner: CommandRunner | None = None, +) -> CommandResult: + """Preview candidates beyond current upper bounds without mutating the checkout.""" + + command_runner = command_runner or run_command + with tempfile.TemporaryDirectory(prefix="ark-sdk-prospective-") as directory: + workspace = Path(directory) + for name in ("pyproject.toml", "uv.lock"): + source = root / name + if source.exists(): + shutil.copy2(source, workspace / name) + pyproject = workspace / "pyproject.toml" + if not pyproject.exists(): + return CommandResult( + command=build_refresh_preview_command(packages), + returncode=2, + stderr="prospective preview requires pyproject.toml", + ) + widen_project_dependency_bounds(pyproject, candidates) + return run_refresh_preview(workspace, packages, command_runner=command_runner) + + def parse_refresh_transitions(evidence: Mapping[str, Any]) -> tuple[ResolverTransition, ...]: """Parse exact resolver transitions from refresh-preview stdout and stderr.""" @@ -320,6 +459,400 @@ def refresh_update_versions(evidence: Mapping[str, Any]) -> dict[str, str]: } +def candidate_update_versions(evidence: Mapping[str, Any]) -> dict[str, str]: + """Return every newer upstream candidate, independent of current resolver caps.""" + + resolver_targets = refresh_update_versions(evidence) + updates: dict[str, str] = {} + packages = evidence.get("packages") + if not isinstance(packages, Sequence) or isinstance(packages, str | bytes): + return updates + for package in packages: + if not isinstance(package, Mapping): + continue + name = str(package.get("name") or "") + if not name: + continue + baseline = _string_or_none(package.get("locked_version")) or _string_or_none( + package.get("installed_version") + ) + candidate_key_present = "candidate_version" in package + candidate = _string_or_none(package.get("candidate_version")) + latest = _string_or_none(package.get("latest_version")) + if candidate is None and not candidate_key_present and _is_newer_version(latest, baseline): + candidate = latest + if candidate is None: + candidate = resolver_targets.get(name) + if candidate is not None and candidate != baseline: + updates[name] = candidate + return updates + + +def candidate_transitions(evidence: Mapping[str, Any]) -> tuple[ResolverTransition, ...]: + """Build the complete baseline-to-upstream candidate set for safety gates.""" + + updates = candidate_update_versions(evidence) + transitions: list[ResolverTransition] = [] + packages = evidence.get("packages") + if isinstance(packages, Sequence) and not isinstance(packages, str | bytes): + for package in packages: + if not isinstance(package, Mapping): + continue + name = str(package.get("name") or "") + candidate = updates.get(name) + baseline = _string_or_none(package.get("locked_version")) or _string_or_none( + package.get("installed_version") + ) + if name and candidate and baseline: + transitions.append(ResolverTransition(name, baseline, candidate)) + # Resolver transitions are supplemental evidence for older/minimal bundles + # that lack package metadata. They never replace independently discovered + # candidates when the package inventory is present. + covered = {transition.package for transition in transitions} + transitions.extend( + transition + for transition in parse_refresh_transitions(evidence) + if transition.package not in covered + ) + return tuple(sorted(set(transitions))) + + +def _annotate_candidate_states(evidence: Mapping[str, Any]) -> list[dict[str, Any]]: + """Classify upstream freshness separately from current resolver selection.""" + + resolver_targets = refresh_update_versions(evidence) + current_preview = evidence.get("current_refresh_preview") + current_targets = ( + refresh_update_versions({"refresh_preview": current_preview}) + if isinstance(current_preview, Mapping) + else resolver_targets + ) + raw_packages = evidence.get("packages") + if not isinstance(raw_packages, Sequence) or isinstance(raw_packages, str | bytes): + return [] + packages: list[dict[str, Any]] = [] + for raw in raw_packages: + if not isinstance(raw, Mapping): + continue + package = dict(raw) + name = str(package.get("name") or "") + baseline = _string_or_none(package.get("locked_version")) or _string_or_none( + package.get("installed_version") + ) + latest = _string_or_none(package.get("latest_version")) + resolver_target = resolver_targets.get(name) + current_target = current_targets.get(name) + sdk_selected_version = _string_or_none(package.get("sdk_selected_version")) + sdk_requirement = _string_or_none(package.get("sdk_requirement")) + candidate = ( + (resolver_target or sdk_selected_version) + if name == "openai-codex-cli-bin" + and (resolver_target or sdk_selected_version) is not None + and (resolver_target or sdk_selected_version) != baseline + else latest + if name != "openai-codex-cli-bin" and _is_newer_version(latest, baseline) + else None + ) + status, reason = _candidate_status( + name=name, + baseline=baseline, + candidate=candidate, + latest=latest, + resolver_target=resolver_target, + current_target=current_target, + pyproject_spec=_string_or_none(package.get("pyproject_spec")), + metadata_available=latest is not None, + sdk_selected_version=sdk_selected_version, + sdk_requirement=sdk_requirement, + ) + package.update( + candidate_version=candidate, + resolver_target_version=resolver_target, + candidate_status=status, + candidate_reason=reason, + ) + packages.append(package) + return packages + + +def _candidate_status( + *, + name: str, + baseline: str | None, + candidate: str | None, + latest: str | None, + resolver_target: str | None, + current_target: str | None, + pyproject_spec: str | None, + metadata_available: bool, + sdk_selected_version: str | None, + sdk_requirement: str | None, +) -> tuple[str, str]: + if candidate is None: + if not metadata_available: + return "metadata-unavailable", "upstream package metadata was unavailable" + if name == "openai-codex-cli-bin" and _is_newer_version(latest, baseline): + selected = sdk_selected_version or baseline or "an earlier runtime" + requirement = f" ({sdk_requirement})" if sdk_requirement else "" + return ( + "sdk-coupled-no-update", + f"latest published openai-codex selects CLI {selected}{requirement}; " + f"standalone CLI {latest} is a staged runtime artifact, not an " + "SDK-usable update", + ) + return "current", "no newer upstream release was observed" + if current_target == candidate: + if name == "openai-codex-cli-bin" and latest and latest != candidate: + return ( + "sdk-coupled-candidate", + f"current resolution selects the CLI {candidate} required by openai-codex; " + f"standalone CLI latest is staged at {latest}", + ) + return "resolver-selected", "current project constraints select the upstream candidate" + if resolver_target == candidate: + if name == "openai-codex-cli-bin" and latest and latest != candidate: + return ( + "prospective-coupled-candidate", + f"Codex SDK candidate selects CLI {candidate}; standalone CLI latest is {latest}", + ) + if pyproject_spec and not _requirement_allows(pyproject_spec, candidate): + return ( + "blocked-by-project-constraint", + f"{pyproject_spec} excludes candidate {candidate}; " + "prospective resolution admits it", + ) + return ( + "prospective-resolver-selected", + "prospective resolution selected the candidate after relaxing excluding upper bounds", + ) + if resolver_target: + return ( + "resolver-selected-nonlatest", + f"resolver selected {resolver_target}, while upstream latest is {candidate}", + ) + if name == "openai-codex-cli-bin" and sdk_selected_version == candidate: + latest_note = f"; standalone CLI latest is {latest}" if latest != candidate else "" + return ( + "sdk-coupled-candidate", + f"latest published openai-codex selects CLI {candidate}{latest_note}", + ) + if pyproject_spec and not _requirement_allows(pyproject_spec, candidate): + return ( + "blocked-by-project-constraint", + f"{pyproject_spec} excludes upstream candidate {candidate}", + ) + if name == "openai-codex-cli-bin": + return ( + "sdk-coupling-unresolved", + "the published Codex SDK dependency and resolver did not select the same CLI version", + ) + return ( + "not-selected-by-resolver", + f"upstream candidate {candidate} is newer than baseline {baseline or ''}", + ) + + +def _needs_prospective_preview(evidence: Mapping[str, Any]) -> bool: + packages = evidence.get("packages") + if not isinstance(packages, Sequence) or isinstance(packages, str | bytes): + return False + return any( + isinstance(package, Mapping) + and package.get("candidate_version") + and package.get("candidate_status") != "resolver-selected" + for package in packages + ) + + +def widen_project_dependency_bounds( + path: Path, + candidates: Mapping[str, str], +) -> tuple[str, ...]: + """Relax excluding upper bounds just enough to admit inspected candidates.""" + + text = path.read_text(encoding="utf-8") + data = tomllib.loads(text) + monitored = {canonicalize_name(name): version for name, version in candidates.items()} + changes: list[str] = [] + replacements: dict[str, str] = {} + for raw in _project_requirement_strings(data, include_constraints=False): + try: + requirement = Requirement(raw) + except InvalidRequirement: + continue + candidate = monitored.get(canonicalize_name(requirement.name)) + if candidate is None or requirement.specifier.contains(candidate, prereleases=True): + continue + widened = _widen_requirement_upper_bound(raw, candidate) + replacements[raw] = widened + change = f"{raw} -> {widened}" + if change not in changes: + changes.append(change) + for old, new in replacements.items(): + before = text + text = text.replace(f'"{old}"', f'"{new}"') + text = text.replace(f"'{old}'", f"'{new}'") + if text == before: + raise ValueError(f"could not locate parsed requirement {old!r} in {path}") + tomllib.loads(text) + if changes: + path.write_text(text, encoding="utf-8") + return tuple(changes) + + +def update_compatibility_manifest( + path: Path, + pyproject_path: Path, + candidates: Mapping[str, str], +) -> tuple[str, ...]: + """Keep committed tested-version metadata aligned with an SDK lock refresh.""" + + if not path.exists(): + raise ValueError(f"compatibility manifest is missing: {path}") + text = path.read_text(encoding="utf-8") + original = text + project_specs = read_pyproject_dependency_specs(pyproject_path) + changes: list[str] = [] + + for package, candidate in sorted(candidates.items()): + entry_span = _compatibility_entry_span(text, package) + if entry_span is None: + text, count = re.subn( + rf'(PackageVersion\(package="{re.escape(package)}", version=")' + r'[^\"]+("\))', + rf"\g<1>{candidate}\g<2>", + text, + ) + if count != 1: + raise ValueError( + f"expected one compatibility runtime dependency for {package}, found {count}" + ) + changes.append(f"{package} tested runtime dependency -> {candidate}") + continue + + start, end = entry_span + block = text[start:end] + block, tested_count = re.subn( + r'(?m)^( tested_version=")[^\"]+(",)$', + rf"\g<1>{candidate}\g<2>", + block, + ) + if tested_count != 1: + raise ValueError( + f"expected one tested_version field for {package}, found {tested_count}" + ) + + raw_requirement = project_specs.get(package) + if raw_requirement is None: + raise ValueError(f"project requirement is missing for {package}") + specifier = _requirement_specifier_text(raw_requirement, package) + block, specifier_count = re.subn( + r'(?m)^( version_specifier=")[^\"]+(",)$', + rf"\g<1>{specifier}\g<2>", + block, + ) + if specifier_count != 1: + raise ValueError( + f"expected one version_specifier field for {package}, found {specifier_count}" + ) + text = text[:start] + block + text[end:] + changes.append(f"{package} tested version -> {candidate} ({specifier})") + + compile(text, str(path), "exec") + if text != original: + path.write_text(text, encoding="utf-8") + return tuple(changes) + + +def _compatibility_entry_span(text: str, package: str) -> tuple[int, int] | None: + package_line = re.compile(rf'(?m)^ package="{re.escape(package)}",$') + matches = tuple(package_line.finditer(text)) + if not matches: + return None + if len(matches) != 1: + raise ValueError(f"expected one compatibility entry for {package}, found {len(matches)}") + start = text.rfind(" RuntimeCompatibility(\n", 0, matches[0].start()) + if start < 0: + raise ValueError(f"could not find compatibility entry start for {package}") + next_entry = text.find(" RuntimeCompatibility(\n", matches[0].end()) + end = next_entry if next_entry >= 0 else text.find("\n)\n", matches[0].end()) + if end < 0: + raise ValueError(f"could not find compatibility entry end for {package}") + return start, end + + +def _requirement_specifier_text(raw: str, package: str) -> str: + requirement = Requirement(raw) + if canonicalize_name(requirement.name) != canonicalize_name(package): + raise ValueError(f"requirement {raw!r} does not describe {package}") + match = re.match(rf"^\s*{re.escape(requirement.name)}\s*", raw, re.IGNORECASE) + if match is None: + raise ValueError(f"could not parse requirement spelling from {raw!r}") + specifier = raw[match.end() :].split(";", 1)[0].strip() + if not specifier or str(requirement.specifier) == "": + raise ValueError(f"requirement {raw!r} has no version specifier") + return specifier + + +def _widen_requirement_upper_bound(raw: str, candidate: str) -> str: + next_minor = _next_minor_upper_bound(candidate) + changed = False + + def replace_upper(match: re.Match[str]) -> str: + nonlocal changed + token = match.group(0) + try: + allows = Specifier(token).contains(candidate, prereleases=True) + except InvalidSpecifier: + return token + if allows: + return token + changed = True + return f"<{next_minor}" + + widened = re.sub(r"(?=!~])(?:<=|<)\s*[^,;\s]+", replace_upper, raw) + if not changed: + raise ValueError( + f"{raw!r} excludes {candidate}, but no safely widenable upper bound was found" + ) + if not Requirement(widened).specifier.contains(candidate, prereleases=True): + raise ValueError(f"widened requirement {widened!r} still excludes {candidate}") + return widened + + +def _next_minor_upper_bound(version: str) -> str: + parsed = Version(version) + release = parsed.release + if len(release) < 2: + return str(release[0] + 1) + return f"{release[0]}.{release[1] + 1}" + + +def _requirement_allows(raw: str, version: str) -> bool: + try: + return Requirement(raw).specifier.contains(version, prereleases=True) + except (InvalidRequirement, InvalidVersion): + return False + + +def _is_newer_version(candidate: str | None, baseline: str | None) -> bool: + if candidate is None: + return False + if baseline is None: + return True + try: + return Version(candidate) > Version(baseline) + except InvalidVersion: + return _version_key(candidate) > _version_key(baseline) + + +def _string_or_none(value: object) -> str | None: + if value is None: + return None + text = str(value) + return text or None + + def run_lock_update( root: Path, packages: Sequence[str], @@ -330,16 +863,9 @@ def run_lock_update( command_runner = command_runner or run_command env, removed = cutoff_free_env() - command = ["uv", "lock"] + command = ["uv", "lock", "--exclude-newer", SDK_EVOLUTION_EXCLUDE_NEWER] for package in packages: command.extend(("-P", package)) - for package in packages: - command.extend( - ( - "--exclude-newer-package", - f"{package}={SDK_EVOLUTION_EXCLUDE_NEWER_PACKAGE}", - ) - ) result = command_runner(tuple(command), cwd=root, env=env) return CommandResult( command=result.command, @@ -359,9 +885,10 @@ def run_verification_commands( """Run verification commands requested by an architecture decision.""" command_runner = command_runner or run_command + env, _removed = cutoff_free_env() results: list[CommandResult] = [] for command in commands: - results.append(command_runner(tuple(shlex.split(command)), cwd=root)) + results.append(command_runner(tuple(shlex.split(command)), cwd=root, env=env)) return tuple(results) @@ -370,19 +897,29 @@ def run_command( *, cwd: Path | None = None, env: Mapping[str, str] | None = None, - timeout: int = 120, + timeout: int = 900, ) -> CommandResult: """Run a local command and capture output.""" - completed = subprocess.run( - tuple(command), - cwd=cwd, - env=dict(env) if env is not None else None, - text=True, - capture_output=True, - timeout=timeout, - check=False, - ) + try: + completed = subprocess.run( + tuple(command), + cwd=cwd, + env=dict(env) if env is not None else None, + text=True, + capture_output=True, + timeout=timeout, + check=False, + ) + except subprocess.TimeoutExpired as exc: + return CommandResult( + command=tuple(command), + returncode=124, + stdout=_decode_timeout_output(exc.stdout), + stderr=( + f"command timed out after {timeout}s: {_decode_timeout_output(exc.stderr)}" + ).strip(), + ) return CommandResult( command=tuple(command), returncode=completed.returncode, @@ -391,6 +928,12 @@ def run_command( ) +def _decode_timeout_output(value: str | bytes | None) -> str: + if value is None: + return "" + return value.decode(errors="replace") if isinstance(value, bytes) else value + + def _version_key(version: str) -> tuple[Any, ...]: parts: list[Any] = [] for part in re.split(r"[.\-+_]", version): diff --git a/examples/sdk_evolution_agent/current_state.py b/examples/sdk_evolution_agent/current_state.py index 7c91ae1..189395e 100644 --- a/examples/sdk_evolution_agent/current_state.py +++ b/examples/sdk_evolution_agent/current_state.py @@ -46,6 +46,7 @@ def _artifact_refs(report_root: Path, *, workspace: Path) -> dict[str, dict[str, "evidence.json", "release_notes.json", "api_diffs.json", + "implementation_diffs.json", "behavior_probes.json", "behavior_diffs.json", "behavior_summary.json", @@ -70,6 +71,13 @@ def _artifact_refs(report_root: Path, *, workspace: Path) -> dict[str, dict[str, "path": _portable_path(path, workspace=workspace), "sha256": _sha256(path), } + implementation_snapshots_dir = report_root / "implementation_snapshots" + if implementation_snapshots_dir.exists(): + for path in sorted(implementation_snapshots_dir.glob("*.json")): + refs[f"implementation_snapshots/{path.name}"] = { + "path": _portable_path(path, workspace=workspace), + "sha256": _sha256(path), + } return refs diff --git a/examples/sdk_evolution_agent/inspection.py b/examples/sdk_evolution_agent/inspection.py new file mode 100644 index 0000000..c275605 --- /dev/null +++ b/examples/sdk_evolution_agent/inspection.py @@ -0,0 +1,129 @@ +"""Shared, credential-scrubbed candidate environments for one evolution run.""" + +from __future__ import annotations + +import contextlib +import contextvars +import hashlib +import subprocess +import sys +import tempfile +import time +from collections.abc import Callable, Iterator, Mapping, Sequence +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any + + +@dataclass(frozen=True) +class CandidateEnvironment: + """A prepared isolated interpreter containing one exact package version.""" + + python: Path + env: Mapping[str, str] + + +@dataclass +class _CandidateEnvironmentCache: + root: Path + environments: dict[tuple[str, str, str], CandidateEnvironment] = field(default_factory=dict) + + +class CandidatePreparationError(RuntimeError): + """A candidate environment could not be created or populated.""" + + def __init__(self, step: str, cause: BaseException) -> None: + super().__init__(str(cause)) + self.step = step + self.cause = cause + + +_ACTIVE_CACHE: contextvars.ContextVar[_CandidateEnvironmentCache | None] = contextvars.ContextVar( + "sdk_evolution_candidate_environment_cache", default=None +) + + +@contextlib.contextmanager +def candidate_environment_cache() -> Iterator[None]: + """Reuse each package/version install across snapshots and behavior probes.""" + + with tempfile.TemporaryDirectory(prefix="ark-sdk-inspection-") as directory: + token = _ACTIVE_CACHE.set(_CandidateEnvironmentCache(Path(directory))) + try: + yield + finally: + _ACTIVE_CACHE.reset(token) + + +def prepare_cached_candidate_environment( + package: str, + version: str, + *, + python: str, + timeout: int, + runner: Callable[..., Any], + env_factory: Callable[[Path], Mapping[str, str]], +) -> CandidateEnvironment | None: + """Prepare or return a cached environment; return ``None`` outside a run cache.""" + + cache = _ACTIVE_CACHE.get() + if cache is None: + return None + key = (package, version, python) + existing = cache.environments.get(key) + if existing is not None: + _progress(f"reusing isolated {package}=={version} environment") + return existing + + digest = hashlib.sha256("\0".join(key).encode()).hexdigest()[:16] + workspace = cache.root / digest + workspace.mkdir(parents=True, exist_ok=False) + venv = workspace / ".venv" + env = dict(env_factory(workspace)) + _progress(f"creating isolated environment for {package}=={version}") + try: + runner( + (python, "-m", "venv", str(venv)), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=env, + ) + except (OSError, subprocess.CalledProcessError, subprocess.TimeoutExpired) as exc: + raise CandidatePreparationError("virtual-environment-creation", exc) from exc + + bin_dir = "Scripts" if sys.platform == "win32" else "bin" + venv_python = venv / bin_dir / "python" + install_command: Sequence[str] = ( + str(venv_python), + "-m", + "pip", + "install", + f"{package}=={version}", + ) + for attempt in range(3): + try: + _progress(f"installing {package}=={version} (attempt {attempt + 1}/3)") + runner( + install_command, + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=env, + ) + break + except (OSError, subprocess.CalledProcessError, subprocess.TimeoutExpired) as exc: + if attempt == 2 or isinstance(exc, OSError): + raise CandidatePreparationError("package-installation", exc) from exc + time.sleep(0.25 * (attempt + 1)) + + prepared = CandidateEnvironment(python=venv_python, env=env) + cache.environments[key] = prepared + _progress(f"isolated environment ready for {package}=={version}") + return prepared + + +def _progress(message: str) -> None: + print(f"[sdk-evolution] {message}", flush=True) diff --git a/examples/sdk_evolution_agent/models.py b/examples/sdk_evolution_agent/models.py index 9ce477e..7eb0a92 100644 --- a/examples/sdk_evolution_agent/models.py +++ b/examples/sdk_evolution_agent/models.py @@ -58,6 +58,12 @@ class PackageVersionState: locked_version: str | None = None installed_version: str | None = None latest_version: str | None = None + candidate_version: str | None = None + resolver_target_version: str | None = None + candidate_status: str = "unknown" + candidate_reason: str = "" + sdk_selected_version: str | None = None + sdk_requirement: str | None = None recent_versions: tuple[str, ...] = () sources: tuple[SourceRef, ...] = () unavailable_reason: str = "" @@ -73,14 +79,39 @@ class ApiMember: module: str = "" +@dataclass(frozen=True) +class ImplementationFile: + """Content fingerprint for one file shipped inside an SDK implementation.""" + + path: str + kind: str + sha256: str + size: int + line_count: int = 0 + + +@dataclass(frozen=True) +class ImplementationDefinition: + """AST fingerprint for one Python definition in an SDK implementation.""" + + name: str + kind: str + path: str + sha256: str + + @dataclass(frozen=True) class ApiSnapshot: - """Public API snapshot for an inspected package version.""" + """Public API and implementation snapshot for an inspected package version.""" package: str version: str | None module: str members: tuple[ApiMember, ...] = () + implementation_files: tuple[ImplementationFile, ...] = () + implementation_definitions: tuple[ImplementationDefinition, ...] = () + implementation_status: str = "unavailable" + implementation_note: str = "" import_error: str | None = None source: str = "current-environment" @@ -97,6 +128,30 @@ class ApiDiff: changed: tuple[str, ...] = () +@dataclass(frozen=True) +class ImplementationDiff: + """Observed implementation changes between two exact SDK releases.""" + + package: str + from_version: str | None + to_version: str | None + status: str + python_files_before: int = 0 + python_files_after: int = 0 + source_lines_before: int = 0 + source_lines_after: int = 0 + files_added: tuple[str, ...] = () + files_removed: tuple[str, ...] = () + files_changed: tuple[str, ...] = () + definitions_added: tuple[str, ...] = () + definitions_removed: tuple[str, ...] = () + definitions_changed: tuple[str, ...] = () + opaque_artifacts_added: tuple[str, ...] = () + opaque_artifacts_removed: tuple[str, ...] = () + opaque_artifacts_changed: tuple[str, ...] = () + limitations: tuple[str, ...] = () + + @dataclass(frozen=True) class ReleaseNoteEvidence: """Release-note evidence collected for one package interval.""" diff --git a/examples/sdk_evolution_agent/release_notes.py b/examples/sdk_evolution_agent/release_notes.py index 57f5b63..1567313 100644 --- a/examples/sdk_evolution_agent/release_notes.py +++ b/examples/sdk_evolution_agent/release_notes.py @@ -6,6 +6,7 @@ import json import os import re +import time import urllib.parse import urllib.request from collections.abc import Callable, Mapping, Sequence @@ -109,7 +110,7 @@ def collect_release_notes( to_version=None, status="not-needed", sources=RELEASE_NOTE_SOURCES.get(name, ()), - unavailable_reason="no resolver-selected update", + unavailable_reason="no newer compatible candidate", ) ) continue @@ -124,6 +125,8 @@ def collect_release_notes( source_results.append(source) continue checked_urls.append(source.url) + if use_default_fetcher: + _progress(f"checking {name} release evidence: {source.label}") try: text = _fetch_source_text( source, @@ -165,6 +168,8 @@ def collect_release_notes( continue checked_urls.append(linked_url) label = f"{source.label} matching release note" + if use_default_fetcher: + _progress(f"checking {name} release evidence: {label}") try: linked_text = fetcher(linked_url) except Exception as exc: @@ -233,8 +238,7 @@ def fetch_url_text(url: str) -> str: """Fetch a release-note source as text.""" request = urllib.request.Request(url, headers={"User-Agent": "agent-runtime-kit-sdk-evolution"}) - with urllib.request.urlopen(request, timeout=20) as response: - raw = response.read() + raw = _urlopen_bytes(request) if raw.startswith(b"\x1f\x8b"): raw = gzip.decompress(raw) return raw.decode("utf-8", errors="replace") @@ -296,8 +300,7 @@ def _fetch_github_discussions_index(url: str, *, token: str) -> str: }, method="POST", ) - with urllib.request.urlopen(request, timeout=20) as response: - response_payload = json.loads(response.read().decode("utf-8", errors="replace")) + response_payload = json.loads(_urlopen_bytes(request).decode("utf-8", errors="replace")) if response_payload.get("errors"): raise RuntimeError(str(response_payload["errors"])) return _format_github_discussions_index(response_payload, category_slug=category_slug) @@ -432,3 +435,19 @@ def _string_or_none(value: object) -> str | None: return None text = str(value) return text or None + + +def _urlopen_bytes(request: urllib.request.Request) -> bytes: + for attempt in range(3): + try: + with urllib.request.urlopen(request, timeout=20) as response: + return response.read() + except (OSError, TimeoutError): + if attempt == 2: + raise + time.sleep(0.25 * (attempt + 1)) + raise AssertionError("unreachable") + + +def _progress(message: str) -> None: + print(f"[sdk-evolution] {message}", flush=True) diff --git a/examples/sdk_evolution_agent/report.py b/examples/sdk_evolution_agent/report.py index 9318295..c23f7eb 100644 --- a/examples/sdk_evolution_agent/report.py +++ b/examples/sdk_evolution_agent/report.py @@ -30,6 +30,8 @@ def write_run_report( evidence: dict[str, Any], snapshots: list[dict[str, Any]], api_diffs: list[dict[str, Any]], + implementation_snapshots: list[dict[str, Any]] | None = None, + implementation_diffs: list[dict[str, Any]] | None = None, release_notes: list[dict[str, Any]], behavior: dict[str, Any], current_state: dict[str, Any], @@ -41,11 +43,14 @@ def write_run_report( ) -> Path: """Write all run artifacts and return report.md.""" + implementation_snapshots = implementation_snapshots or [] + implementation_diffs = implementation_diffs or [] context.report_root.mkdir(parents=True, exist_ok=True) write_json(context.report_root / "config.json", config) write_json(context.report_root / "evidence.json", evidence) write_json(context.report_root / "release_notes.json", release_notes) write_json(context.report_root / "api_diffs.json", api_diffs) + write_json(context.report_root / "implementation_diffs.json", implementation_diffs) write_json(context.report_root / "behavior_probes.json", behavior.get("results", [])) write_json(context.report_root / "behavior_diffs.json", behavior.get("diffs", [])) write_json( @@ -65,6 +70,15 @@ def write_run_report( for index, snapshot in enumerate(snapshots, start=1): package = str(snapshot.get("package", "snapshot")).replace("/", "-") write_json(snapshots_dir / f"{index:02d}-{package}.json", snapshot) + implementation_snapshots_dir = context.report_root / "implementation_snapshots" + implementation_snapshots_dir.mkdir(exist_ok=True) + for index, snapshot in enumerate(implementation_snapshots, start=1): + package = str(snapshot.get("package", "snapshot")).replace("/", "-") + version = str(snapshot.get("version") or "unknown").replace("/", "-") + write_json( + implementation_snapshots_dir / f"{index:02d}-{package}-{version}.json", + snapshot, + ) if pr_body is not None: (context.report_root / "draft_pr_body.md").write_text(pr_body, encoding="utf-8") report_path = context.report_root / "report.md" @@ -74,6 +88,8 @@ def write_run_report( evidence=evidence, snapshots=snapshots, api_diffs=api_diffs, + implementation_snapshots=implementation_snapshots, + implementation_diffs=implementation_diffs, release_notes=release_notes, behavior=behavior, current_state=current_state, @@ -93,6 +109,8 @@ def render_markdown_report( evidence: dict[str, Any], snapshots: list[dict[str, Any]], api_diffs: list[dict[str, Any]], + implementation_snapshots: list[dict[str, Any]] | None = None, + implementation_diffs: list[dict[str, Any]] | None = None, release_notes: list[dict[str, Any]], behavior: dict[str, Any], current_state: dict[str, Any], @@ -103,17 +121,25 @@ def render_markdown_report( ) -> str: """Render the human-readable local report.""" + implementation_snapshots = implementation_snapshots or [] + implementation_diffs = implementation_diffs or [] packages = evidence.get("packages", []) package_lines = [] for package in packages: if not isinstance(package, dict): continue package_lines.append( - "- {name}: locked={locked} installed={installed} latest={latest}".format( + ( + "- {name}: locked={locked} installed={installed} upstream_latest={latest} " + "sdk_selected={sdk_selected} candidate={candidate} status={status}" + ).format( name=package.get("name"), locked=package.get("locked_version"), installed=package.get("installed_version"), latest=package.get("latest_version"), + sdk_selected=package.get("sdk_selected_version"), + candidate=package.get("candidate_version"), + status=package.get("candidate_status"), ) ) manual = architecture.get("manual_design_required") @@ -148,12 +174,48 @@ def render_markdown_report( ) for snapshot in snapshot_errors ] + implementation_snapshot_errors = [ + snapshot + for snapshot in implementation_snapshots + if isinstance(snapshot, dict) and snapshot.get("import_error") + ] + implementation_snapshot_error_lines = [ + "- {package}@{version}: {error}".format( + package=snapshot.get("package"), + version=snapshot.get("version"), + error=_one_line(snapshot.get("import_error")), + ) + for snapshot in implementation_snapshot_errors + ] behavior_reason_lines = [ f"- {_one_line(reason)}" for reason in behavior_reasons if isinstance(reason, str) and reason ] promotion = current_state.get("promotion", {}) if isinstance(current_state, dict) else {} + implementation_diff_lines = [ + ( + "- {package} {from_version} -> {to_version}: {status}; " + "Python files +{files_added}/-{files_removed}/~{files_changed}; " + "definitions +{definitions_added}/-{definitions_removed}/~{definitions_changed}; " + "source lines {lines_before} -> {lines_after}" + ).format( + package=item.get("package"), + from_version=item.get("from_version"), + to_version=item.get("to_version"), + status=item.get("status"), + files_added=len(item.get("files_added", [])), + files_removed=len(item.get("files_removed", [])), + files_changed=len(item.get("files_changed", [])), + definitions_added=len(item.get("definitions_added", [])), + definitions_removed=len(item.get("definitions_removed", [])), + definitions_changed=len(item.get("definitions_changed", [])), + lines_before=item.get("source_lines_before", 0), + lines_after=item.get("source_lines_after", 0), + ) + for item in implementation_diffs + if isinstance(item, dict) + ] return "\n".join( [ "# SDK Evolution Agent Report", @@ -179,6 +241,15 @@ def render_markdown_report( "", f"- Diff count: `{len(api_diffs)}`", "", + "## Upstream Implementation History", + "", + f"- Status: `{'incomplete' if implementation_snapshot_errors else 'pass'}`", + f"- Release snapshots: `{len(implementation_snapshots)}`", + f"- Adjacent release transitions: `{len(implementation_diffs)}`", + f"- Snapshot errors: `{len(implementation_snapshot_errors)}`", + *implementation_snapshot_error_lines, + *(implementation_diff_lines or ["- No implementation transitions inspected."]), + "", "## Release Notes", "", *(release_lines or ["- No SDK update release-note evidence required."]), @@ -196,7 +267,7 @@ def render_markdown_report( f"- Diff count: `{len(behavior_diffs)}`", *(behavior_reason_lines or ["- No behavior evidence issues recorded."]), "", - "## Direction Of Travel", + "## Upstream Implementation Trends", "", "```json", json.dumps(direction, indent=2, sort_keys=True, default=str), diff --git a/examples/sdk_evolution_agent/schemas.py b/examples/sdk_evolution_agent/schemas.py index 1c73b83..f176ff4 100644 --- a/examples/sdk_evolution_agent/schemas.py +++ b/examples/sdk_evolution_agent/schemas.py @@ -21,11 +21,25 @@ class SchemaValidationError(ValueError): "type": "array", "items": { "type": "object", - "required": ["name", "direction", "evidence"], + "required": [ + "name", + "evidence_status", + "implementation_trend", + "observed_transitions", + "evidence", + ], "additionalProperties": False, "properties": { "name": {"type": "string"}, - "direction": {"type": "string"}, + "evidence_status": { + "type": "string", + "enum": ["observed", "opaque-runtime", "no-transition", "unavailable"], + }, + "implementation_trend": {"type": "string"}, + "observed_transitions": { + "type": "array", + "items": {"type": "string"}, + }, "evidence": {"type": "array", "items": {"type": "string"}}, }, }, @@ -34,11 +48,13 @@ class SchemaValidationError(ValueError): "type": "array", "items": { "type": "object", - "required": ["name", "summary"], + "required": ["name", "implementation_pattern", "packages", "evidence"], "additionalProperties": False, "properties": { "name": {"type": "string"}, - "summary": {"type": "string"}, + "implementation_pattern": {"type": "string"}, + "packages": {"type": "array", "items": {"type": "string"}}, + "evidence": {"type": "array", "items": {"type": "string"}}, }, }, }, diff --git a/examples/sdk_evolution_agent/snapshots.py b/examples/sdk_evolution_agent/snapshots.py index dbca816..6760db4 100644 --- a/examples/sdk_evolution_agent/snapshots.py +++ b/examples/sdk_evolution_agent/snapshots.py @@ -2,6 +2,8 @@ from __future__ import annotations +import ast +import hashlib import importlib import inspect import json @@ -10,11 +12,22 @@ import sys import tempfile import textwrap -from collections.abc import Sequence +from collections.abc import Mapping, Sequence from pathlib import Path from typing import Any -from examples.sdk_evolution_agent.models import ApiDiff, ApiMember, ApiSnapshot +from examples.sdk_evolution_agent.inspection import ( + CandidatePreparationError, + prepare_cached_candidate_environment, +) +from examples.sdk_evolution_agent.models import ( + ApiDiff, + ApiMember, + ApiSnapshot, + ImplementationDefinition, + ImplementationDiff, + ImplementationFile, +) DEFAULT_MODULES = { "claude-agent-sdk": "claude_agent_sdk", @@ -25,7 +38,7 @@ def snapshot_current_api(package: str, *, version: str | None = None) -> ApiSnapshot: - """Capture public API for a package importable in the current environment.""" + """Capture public API and implementation fingerprints in the current environment.""" module_name = DEFAULT_MODULES.get(package, package.replace("-", "_")) try: @@ -49,11 +62,16 @@ def snapshot_current_api(package: str, *, version: str | None = None) -> ApiSnap module=str(getattr(value, "__module__", "")), ) ) + implementation = _inspect_implementation(module, package=package, module_name=module_name) return ApiSnapshot( package=package, version=version or str(getattr(module, "__version__", "") or ""), module=module_name, members=tuple(sorted(members, key=lambda item: item.name)), + implementation_files=implementation[0], + implementation_definitions=implementation[1], + implementation_status=implementation[2], + implementation_note=implementation[3], ) @@ -97,6 +115,137 @@ def diff_snapshot_groups(snapshots: Sequence[ApiSnapshot]) -> tuple[ApiDiff, ... return tuple(diffs) +def diff_implementation_snapshots( + before: ApiSnapshot, + after: ApiSnapshot, +) -> ImplementationDiff: + """Diff file and Python-definition fingerprints for two exact releases.""" + + if ( + before.import_error + or after.import_error + or before.implementation_status == "unavailable" + or after.implementation_status == "unavailable" + ): + errors = tuple( + detail + for detail in ( + f"{before.version}: {before.import_error}" if before.import_error else "", + f"{after.version}: {after.import_error}" if after.import_error else "", + ( + f"{before.version}: " + f"{before.implementation_note or 'implementation unavailable'}" + if before.implementation_status == "unavailable" + else "" + ), + ( + f"{after.version}: {after.implementation_note or 'implementation unavailable'}" + if after.implementation_status == "unavailable" + else "" + ), + ) + if detail + ) + return ImplementationDiff( + package=before.package, + from_version=before.version, + to_version=after.version, + status="unavailable", + limitations=errors or ("implementation snapshot was unavailable",), + ) + + before_files = {item.path: item for item in before.implementation_files} + after_files = {item.path: item for item in after.implementation_files} + before_python = { + path: item for path, item in before_files.items() if item.kind in {"python", "stub"} + } + after_python = { + path: item for path, item in after_files.items() if item.kind in {"python", "stub"} + } + before_opaque = { + path: item for path, item in before_files.items() if item.kind not in {"python", "stub"} + } + after_opaque = { + path: item for path, item in after_files.items() if item.kind not in {"python", "stub"} + } + before_definitions = {item.name: item for item in before.implementation_definitions} + after_definitions = {item.name: item for item in after.implementation_definitions} + opaque_artifacts_added = tuple(sorted(set(after_opaque) - set(before_opaque))) + opaque_artifacts_removed = tuple(sorted(set(before_opaque) - set(after_opaque))) + opaque_artifacts_changed = tuple( + sorted( + path + for path in set(before_opaque) & set(after_opaque) + if before_opaque[path].sha256 != after_opaque[path].sha256 + ) + ) + + status = "observed" + limitations: tuple[str, ...] = () + if before.package == "openai-codex-cli-bin": + status = "opaque-runtime" + limitations = ( + "Python packaging code is inspectable, but the bundled Codex executable " + "is opaque in wheel artifacts.", + ) + elif not before_python and not after_python: + status = "opaque-runtime" if before_opaque or after_opaque else "unavailable" + limitations = ("No inspectable Python implementation source was found.",) + elif opaque_artifacts_added or opaque_artifacts_removed or opaque_artifacts_changed: + limitations = ( + "Opaque or native artifacts changed; fingerprints prove artifact drift, not " + "their internal implementation design.", + ) + + return ImplementationDiff( + package=before.package, + from_version=before.version, + to_version=after.version, + status=status, + python_files_before=len(before_python), + python_files_after=len(after_python), + source_lines_before=sum(item.line_count for item in before_python.values()), + source_lines_after=sum(item.line_count for item in after_python.values()), + files_added=tuple(sorted(set(after_python) - set(before_python))), + files_removed=tuple(sorted(set(before_python) - set(after_python))), + files_changed=tuple( + sorted( + path + for path in set(before_python) & set(after_python) + if before_python[path].sha256 != after_python[path].sha256 + ) + ), + definitions_added=tuple(sorted(set(after_definitions) - set(before_definitions))), + definitions_removed=tuple(sorted(set(before_definitions) - set(after_definitions))), + definitions_changed=tuple( + sorted( + name + for name in set(before_definitions) & set(after_definitions) + if before_definitions[name].sha256 != after_definitions[name].sha256 + ) + ), + opaque_artifacts_added=opaque_artifacts_added, + opaque_artifacts_removed=opaque_artifacts_removed, + opaque_artifacts_changed=opaque_artifacts_changed, + limitations=limitations, + ) + + +def diff_implementation_snapshot_groups( + snapshots: Sequence[ApiSnapshot], +) -> tuple[ImplementationDiff, ...]: + """Diff adjacent, already-ordered release snapshots for every package.""" + + diffs: list[ImplementationDiff] = [] + grouped: dict[str, list[ApiSnapshot]] = {} + for snapshot in snapshots: + grouped.setdefault(snapshot.package, []).append(snapshot) + for group in grouped.values(): + for index in range(1, len(group)): + diffs.append(diff_implementation_snapshots(group[index - 1], group[index])) + return tuple(diffs) + + def isolated_env(home: Path) -> dict[str, str]: """A minimal environment for candidate subprocesses: PATH + a throwaway HOME. @@ -130,45 +279,68 @@ def snapshot_candidate_in_venv( module_name = DEFAULT_MODULES.get(package, package.replace("-", "_")) step = "virtual environment creation" try: - with tempfile.TemporaryDirectory(prefix="ark-sdk-snapshot-") as directory: - venv = Path(directory) / ".venv" - # Scrub the environment for every subprocess that touches freshly downloaded - # upstream code: give it a throwaway HOME and only PATH, so a malicious or - # buggy candidate package cannot read the caller's credentials/config. - env = isolated_env(Path(directory)) - subprocess.run( - (python, "-m", "venv", str(venv)), - check=True, - timeout=timeout, - env=env, - ) - bin_dir = "Scripts" if sys.platform == "win32" else "bin" - venv_python = venv / bin_dir / "python" - step = "package installation" - subprocess.run( - (str(venv_python), "-m", "pip", "install", f"{package}=={version}"), - check=True, - text=True, - capture_output=True, - timeout=timeout, - env=env, - ) + cached = prepare_cached_candidate_environment( + package, + version, + python=python, + timeout=timeout, + runner=subprocess.run, + env_factory=isolated_env, + ) + if cached is not None: step = "snapshot execution" - completed = subprocess.run( - ( - str(venv_python), - "-c", - _SNAPSHOT_SCRIPT, - package, - version, - module_name, - ), - check=True, - text=True, - capture_output=True, + completed = _run_snapshot_script( + cached.python, + package=package, + version=version, + module_name=module_name, timeout=timeout, - env=env, + env=cached.env, ) + else: + with tempfile.TemporaryDirectory(prefix="ark-sdk-snapshot-") as directory: + venv = Path(directory) / ".venv" + # Scrub the environment for every subprocess that touches freshly downloaded + # upstream code: give it a throwaway HOME and only PATH, so a malicious or + # buggy candidate package cannot read the caller's credentials/config. + env = isolated_env(Path(directory)) + subprocess.run( + (python, "-m", "venv", str(venv)), + check=True, + timeout=timeout, + env=env, + ) + bin_dir = "Scripts" if sys.platform == "win32" else "bin" + venv_python = venv / bin_dir / "python" + step = "package installation" + subprocess.run( + (str(venv_python), "-m", "pip", "install", f"{package}=={version}"), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=env, + ) + step = "snapshot execution" + completed = _run_snapshot_script( + venv_python, + package=package, + version=version, + module_name=module_name, + timeout=timeout, + env=env, + ) + except CandidatePreparationError as exc: + step = exc.step.replace("-", " ") + cause = exc.cause + if isinstance(cause, subprocess.TimeoutExpired): + detail = f"{step} timed out after {cause.timeout}s" + elif isinstance(cause, subprocess.CalledProcessError): + output = _bounded_failure_detail(cause.stderr or cause.stdout or str(cause)) + detail = f"{step} failed: {output}" + else: + detail = f"{step} failed: {_bounded_failure_detail(cause)}" + return _failed_isolated_snapshot(package, version, module_name, detail) except subprocess.TimeoutExpired as exc: return _failed_isolated_snapshot( package, @@ -199,6 +371,15 @@ def snapshot_candidate_in_venv( version=raw["version"], module=raw["module"], members=tuple(ApiMember(**item) for item in raw.get("members", ())), + implementation_files=tuple( + ImplementationFile(**item) for item in raw.get("implementation_files", ()) + ), + implementation_definitions=tuple( + ImplementationDefinition(**item) + for item in raw.get("implementation_definitions", ()) + ), + implementation_status=str(raw.get("implementation_status") or "unavailable"), + implementation_note=str(raw.get("implementation_note") or ""), import_error=raw.get("import_error"), source="isolated-venv", ) @@ -210,6 +391,25 @@ def snapshot_candidate_in_venv( return _failed_isolated_snapshot(package, version, module_name, detail) +def _run_snapshot_script( + python: Path, + *, + package: str, + version: str, + module_name: str, + timeout: int, + env: Mapping[str, str], +) -> subprocess.CompletedProcess[str]: + return subprocess.run( + (str(python), "-c", _SNAPSHOT_SCRIPT, package, version, module_name), + check=True, + text=True, + capture_output=True, + timeout=timeout, + env=dict(env), + ) + + def _failed_isolated_snapshot( package: str, version: str, @@ -250,12 +450,276 @@ def _signature(value: Any) -> str: return "" +def _inspect_implementation( + module: Any, + *, + package: str, + module_name: str, +) -> tuple[ + tuple[ImplementationFile, ...], + tuple[ImplementationDefinition, ...], + str, + str, +]: + roots = _module_roots(module) + if not roots: + return (), (), "unavailable", "imported module has no inspectable filesystem path" + + files: list[ImplementationFile] = [] + definitions: list[ImplementationDefinition] = [] + seen_paths: set[str] = set() + for root_index, (root, single_file) in enumerate(roots, start=1): + candidates = (root,) if single_file else tuple(sorted(root.rglob("*"))) + for path in candidates: + if not path.is_file() or path.is_symlink() or "__pycache__" in path.parts: + continue + if path.suffix in {".pyc", ".pyo"}: + continue + relative = path.name if single_file else str(path.relative_to(root)) + if len(roots) > 1: + relative = f"root-{root_index}/{relative}" + relative = f"{module_name.replace('.', '/')}/{relative}" + if relative in seen_paths: + continue + seen_paths.add(relative) + try: + content = path.read_bytes() + except OSError: + continue + kind = _implementation_file_kind(path, content) + line_count = ( + content.count(b"\n") + int(bool(content) and not content.endswith(b"\n")) + if kind in {"python", "stub"} + else 0 + ) + files.append( + ImplementationFile( + path=relative, + kind=kind, + sha256=hashlib.sha256(content).hexdigest(), + size=len(content), + line_count=int(line_count), + ) + ) + if kind in {"python", "stub"}: + definitions.extend(_definition_fingerprints(content, path=relative)) + + status = "opaque-runtime" if package == "openai-codex-cli-bin" else "observed" + note = ( + "bundled executable internals are opaque; wrapper files and artifact hashes are observed" + if status == "opaque-runtime" + else "" + ) + if not files: + status = "unavailable" + note = "module path contained no readable implementation files" + return ( + tuple(sorted(files, key=lambda item: item.path)), + tuple(sorted(definitions, key=lambda item: item.name)), + status, + note, + ) + + +def _module_roots(module: Any) -> tuple[tuple[Path, bool], ...]: + roots: list[tuple[Path, bool]] = [] + module_path = getattr(module, "__path__", None) + if module_path is not None: + for value in module_path: + path = Path(str(value)).resolve() + if path.exists(): + roots.append((path, False)) + if not roots: + module_file = getattr(module, "__file__", None) + if module_file: + path = Path(str(module_file)).resolve() + if path.is_file(): + roots.append((path, True)) + return tuple(roots) + + +def _implementation_file_kind(path: Path, content: bytes) -> str: + if path.suffix == ".py": + return "python" + if path.suffix == ".pyi": + return "stub" + if path.suffix.lower() in {".so", ".dylib", ".dll", ".pyd"}: + return "native-extension" + if content.startswith((b"\x7fELF", b"\xcf\xfa\xed\xfe", b"\xca\xfe\xba\xbe", b"MZ")): + return "executable" + return "opaque-artifact" + + +def _definition_fingerprints( + content: bytes, + *, + path: str, +) -> tuple[ImplementationDefinition, ...]: + try: + tree = ast.parse(content.decode("utf-8"), filename=path) + except (SyntaxError, UnicodeDecodeError): + return () + definitions: list[ImplementationDefinition] = [] + + def visit(node: ast.AST, parents: tuple[str, ...] = ()) -> None: + for child in ast.iter_child_nodes(node): + if isinstance(child, (ast.ClassDef, ast.FunctionDef, ast.AsyncFunctionDef)): + name = ".".join((*parents, child.name)) + kind = ( + "class" + if isinstance(child, ast.ClassDef) + else "async-function" + if isinstance(child, ast.AsyncFunctionDef) + else "function" + ) + fingerprint = hashlib.sha256( + ast.dump(child, include_attributes=False).encode("utf-8") + ).hexdigest() + definitions.append( + ImplementationDefinition( + name=f"{path}:{name}", + kind=kind, + path=path, + sha256=fingerprint, + ) + ) + visit(child, (*parents, child.name)) + else: + visit(child, parents) + + visit(tree) + return tuple(definitions) + + _SNAPSHOT_SCRIPT = textwrap.dedent( """ + import ast + import hashlib import importlib import inspect import json import sys + from pathlib import Path + + def module_roots(module): + roots = [] + module_path = getattr(module, "__path__", None) + if module_path is not None: + for value in module_path: + path = Path(str(value)).resolve() + if path.exists(): + roots.append((path, False)) + if not roots: + module_file = getattr(module, "__file__", None) + if module_file: + path = Path(str(module_file)).resolve() + if path.is_file(): + roots.append((path, True)) + return roots + + def file_kind(path, content): + if path.suffix == ".py": + return "python" + if path.suffix == ".pyi": + return "stub" + if path.suffix.lower() in {".so", ".dylib", ".dll", ".pyd"}: + return "native-extension" + if content.startswith( + (b"\\x7fELF", b"\\xcf\\xfa\\xed\\xfe", b"\\xca\\xfe\\xba\\xbe", b"MZ") + ): + return "executable" + return "opaque-artifact" + + def definition_fingerprints(content, path): + try: + tree = ast.parse(content.decode("utf-8"), filename=path) + except (SyntaxError, UnicodeDecodeError): + return [] + definitions = [] + + def visit(node, parents=()): + for child in ast.iter_child_nodes(node): + if isinstance(child, (ast.ClassDef, ast.FunctionDef, ast.AsyncFunctionDef)): + name = ".".join((*parents, child.name)) + if isinstance(child, ast.ClassDef): + kind = "class" + elif isinstance(child, ast.AsyncFunctionDef): + kind = "async-function" + else: + kind = "function" + fingerprint = hashlib.sha256( + ast.dump(child, include_attributes=False).encode("utf-8") + ).hexdigest() + definitions.append({ + "name": f"{path}:{name}", + "kind": kind, + "path": path, + "sha256": fingerprint, + }) + visit(child, (*parents, child.name)) + else: + visit(child, parents) + + visit(tree) + return definitions + + def inspect_implementation(module, package, module_name): + roots = module_roots(module) + if not roots: + return [], [], "unavailable", "imported module has no inspectable filesystem path" + files = [] + definitions = [] + seen_paths = set() + for root_index, (root, single_file) in enumerate(roots, start=1): + candidates = (root,) if single_file else tuple(sorted(root.rglob("*"))) + for path in candidates: + if not path.is_file() or path.is_symlink() or "__pycache__" in path.parts: + continue + if path.suffix in {".pyc", ".pyo"}: + continue + relative = path.name if single_file else str(path.relative_to(root)) + if len(roots) > 1: + relative = f"root-{root_index}/{relative}" + relative = f"{module_name.replace('.', '/')}/{relative}" + if relative in seen_paths: + continue + seen_paths.add(relative) + try: + content = path.read_bytes() + except OSError: + continue + kind = file_kind(path, content) + line_count = ( + content.count(b"\\n") + + int(bool(content) and not content.endswith(b"\\n")) + if kind in {"python", "stub"} + else 0 + ) + files.append({ + "path": relative, + "kind": kind, + "sha256": hashlib.sha256(content).hexdigest(), + "size": len(content), + "line_count": int(line_count), + }) + if kind in {"python", "stub"}: + definitions.extend(definition_fingerprints(content, relative)) + status = "opaque-runtime" if package == "openai-codex-cli-bin" else "observed" + note = ( + "bundled executable internals are opaque; wrapper files and artifact " + "hashes are observed" + if status == "opaque-runtime" + else "" + ) + if not files: + status = "unavailable" + note = "module path contained no readable implementation files" + return ( + sorted(files, key=lambda item: item["path"]), + sorted(definitions, key=lambda item: item["name"]), + status, + note, + ) package, version, module_name = sys.argv[1:4] try: @@ -282,11 +746,21 @@ def _signature(value: Any) -> str: "signature": signature, "module": str(getattr(value, "__module__", "")), }) + ( + implementation_files, + implementation_definitions, + implementation_status, + implementation_note, + ) = inspect_implementation(module, package, module_name) payload = { "package": package, "version": version, "module": module_name, "members": sorted(members, key=lambda item: item["name"]), + "implementation_files": implementation_files, + "implementation_definitions": implementation_definitions, + "implementation_status": implementation_status, + "implementation_note": implementation_note, "import_error": None, } except Exception as exc: @@ -295,6 +769,10 @@ def _signature(value: Any) -> str: "version": version, "module": module_name, "members": [], + "implementation_files": [], + "implementation_definitions": [], + "implementation_status": "unavailable", + "implementation_note": "implementation inspection failed", "import_error": str(exc), } print(json.dumps(payload, sort_keys=True)) diff --git a/examples/sdk_evolution_agent/stages.py b/examples/sdk_evolution_agent/stages.py index 703b20e..016191c 100644 --- a/examples/sdk_evolution_agent/stages.py +++ b/examples/sdk_evolution_agent/stages.py @@ -3,6 +3,7 @@ from __future__ import annotations import json +import re from collections.abc import Mapping, Sequence from pathlib import Path from typing import Any @@ -33,7 +34,7 @@ assess_behavior_payload, behavior_expectations_from_evidence, ) -from examples.sdk_evolution_agent.collectors import parse_refresh_transitions +from examples.sdk_evolution_agent.collectors import candidate_transitions, candidate_update_versions from examples.sdk_evolution_agent.models import ( RUNTIME_CONTRACT_SYMBOLS, ApiDiff, @@ -210,6 +211,7 @@ async def run_analysis_pipeline( *, evidence: Mapping[str, Any], api_diffs: Sequence[Mapping[str, Any]], + implementation_diffs: Sequence[Mapping[str, Any]], release_notes: Sequence[Mapping[str, Any]], behavior: Mapping[str, Any], context: RunContext, @@ -219,9 +221,14 @@ async def run_analysis_pipeline( stage_payload = { "evidence": evidence, "api_diffs": list(api_diffs), + "implementation_diffs": _compact_stage_value( + list(implementation_diffs), + list_limit=12, + ), "release_notes": list(release_notes), "behavior": behavior, } + _progress("starting direction analysis") direction = await run_stage( runtime, stage="direction-analysis", @@ -229,13 +236,24 @@ async def run_analysis_pipeline( schema=DIRECTION_ANALYSIS_SCHEMA, context=context, ) + _progress("direction analysis complete") + direction = with_implementation_direction_guard( + direction, + evidence=evidence, + implementation_diffs=implementation_diffs, + ) direction = _compact_stage_output(direction) + _progress("starting architecture decision") architecture = await run_stage( runtime, stage="architecture-decision", payload={ "evidence": evidence, "api_diffs": list(api_diffs), + "implementation_diffs": _compact_stage_value( + list(implementation_diffs), + list_limit=12, + ), "release_notes": list(release_notes), "behavior": behavior, "direction_analysis": direction, @@ -243,18 +261,30 @@ async def run_analysis_pipeline( schema=ARCHITECTURE_DECISION_SCHEMA, context=context, ) - architecture = with_recursive_impact(architecture, api_diffs) + _progress("architecture decision complete; applying deterministic gates") + architecture = with_recursive_impact( + architecture, + api_diffs, + evidence=evidence, + runtime=context.runtime, + ) + architecture = with_resolver_preview_guard(architecture, evidence) architecture = with_candidate_api_diff_guard(architecture, evidence, api_diffs) architecture = with_release_note_guard(architecture, release_notes) architecture = with_behavior_probe_guard(architecture, evidence, behavior) architecture = with_manual_design_gate(architecture) architecture = _compact_stage_output(architecture) + _progress("starting independent review") review = await run_stage( runtime, stage="review", payload={ "evidence": evidence, "api_diffs": list(api_diffs), + "implementation_diffs": _compact_stage_value( + list(implementation_diffs), + list_limit=12, + ), "release_notes": list(release_notes), "behavior": behavior, "direction_analysis": direction, @@ -263,6 +293,7 @@ async def run_analysis_pipeline( schema=REVIEWER_OUTPUT_SCHEMA, context=context, ) + _progress("independent review complete") return direction, architecture, review @@ -275,7 +306,7 @@ async def maybe_run_implementation( review: Mapping[str, Any], context: RunContext, ) -> dict[str, Any]: - """Run implementation only if decision gates permit it.""" + """Evaluate whether the deterministic compatible-dependency update may run.""" gate = evaluate_implementation_gate( architecture, @@ -294,6 +325,7 @@ async def maybe_run_implementation( return { "applied": False, "allowed": True, + "mode": "compatible-dependency-refresh", "changes": [], "verification_results": [], "blocked_reason": "", @@ -351,29 +383,87 @@ def detects_recursive_impact(api_diffs: Sequence[Mapping[str, Any] | ApiDiff]) - def with_recursive_impact( architecture: Mapping[str, Any], api_diffs: Sequence[Mapping[str, Any] | ApiDiff], + *, + evidence: Mapping[str, Any] | None = None, + runtime: str | None = None, ) -> dict[str, Any]: """Ensure recursive runtime-contract impacts are explicit.""" result = dict(architecture) - if not detects_recursive_impact(api_diffs): + candidate_packages = set(candidate_update_versions(evidence or {})) + runtime_packages = _runtime_dependency_packages(runtime) + runtime_dependency_impact = bool(candidate_packages & runtime_packages) + if not detects_recursive_impact(api_diffs) and not runtime_dependency_impact: return result result["recursive_self_adaptation_impact"] = True - result.setdefault( - "self_adaptation_plan", - [ + if not result.get("self_adaptation_plan"): + result["self_adaptation_plan"] = [ "Update examples/sdk_evolution_agent runtime usage, schemas, tests, and docs " "in the same scoped change.", - ], + "Rerun the evolution report through the updated runtime after lockfile verification.", + ] + findings = list(result.get("findings") or []) + evidence_items = ["api_diffs"] if detects_recursive_impact(api_diffs) else [] + evidence_items.extend( + f"active runtime dependency candidate: {package}" + for package in sorted(candidate_packages & runtime_packages) + ) + findings.append( + { + "classification": "recursive-runtime-validation", + "summary": ( + "An active analysis-runtime dependency will change and requires " + "post-update verification through that runtime." + ), + "evidence": evidence_items, + } ) + result["findings"] = findings + return result + + +def _runtime_dependency_packages(runtime: str | None) -> set[str]: + normalized = str(runtime or "").lower().replace("_", "-") + if "codex" in normalized: + return {"openai-codex", "openai-codex-cli-bin"} + if "claude" in normalized: + return {"claude-agent-sdk"} + if "antigravity" in normalized: + return {"google-antigravity"} + return set() + + +def with_resolver_preview_guard( + architecture: Mapping[str, Any], + evidence: Mapping[str, Any], +) -> dict[str, Any]: + """Require a successful prospective resolution before applying candidates.""" + + if not candidate_update_versions(evidence): + return dict(architecture) + preview = evidence.get("refresh_preview") + if isinstance(preview, Mapping) and preview.get("returncode") == 0: + return dict(architecture) + + result = dict(architecture) + result["safe_to_implement"] = False + result["manual_design_required"] = True findings = list(result.get("findings") or []) findings.append( { "classification": "manual-design-required", - "summary": "Runtime contract changes affect the SDK evolution agent itself.", - "evidence": ["api_diffs"], + "summary": "SDK candidates were not proven resolvable after relaxing excluding bounds.", + "evidence": [ + "refresh_preview is missing" + if not isinstance(preview, Mapping) + else f"refresh_preview exited {preview.get('returncode')!r}" + ], } ) result["findings"] = findings + uncertainty = list(result.get("uncertainty") or []) + uncertainty.append("A successful no-cooloff prospective resolver preview is required.") + result["uncertainty"] = uncertainty return result @@ -384,7 +474,7 @@ def with_candidate_api_diff_guard( ) -> dict[str, Any]: """Block SDK update implementation when candidate API evidence is missing.""" - transitions = parse_refresh_transitions(evidence) + transitions = candidate_transitions(evidence) if not transitions: return dict(architecture) observed = { @@ -590,6 +680,227 @@ def _compact_stage_value(value: Any, *, string_limit: int = 800, list_limit: int return value +_OPERATIONAL_DIRECTION_RE = re.compile( + r"\b(?:upgrade|hold|keep|resolver|lockfile|candidate|adapter[- ]contract|" + r"safe to implement)\b", + re.IGNORECASE, +) + + +def with_implementation_direction_guard( + direction: Mapping[str, Any], + *, + evidence: Mapping[str, Any], + implementation_diffs: Sequence[Mapping[str, Any]], +) -> dict[str, Any]: + """Anchor direction output to implementation evidence, never release operations.""" + + grouped: dict[str, list[Mapping[str, Any]]] = {} + for item in implementation_diffs: + package = str(item.get("package") or "") + if package: + grouped.setdefault(package, []).append(item) + runtime_packages = { + str(item.get("name") or ""): item + for item in direction.get("packages", []) + if isinstance(item, Mapping) and item.get("name") + } + package_names = [ + str(item.get("name")) + for item in evidence.get("packages", []) + if isinstance(item, Mapping) and item.get("name") + ] + packages: list[dict[str, Any]] = [] + replaced_operation_advice = False + for package in package_names: + diffs = grouped.get(package, []) + source = runtime_packages.get(package, {}) + status = _implementation_evidence_status(package, diffs) + trend = str(source.get("implementation_trend") or "").strip() + if not trend or _OPERATIONAL_DIRECTION_RE.search(trend): + replaced_operation_advice = replaced_operation_advice or bool(trend) + trend = _deterministic_implementation_trend(package, status, diffs) + packages.append( + { + "name": package, + "evidence_status": status, + "implementation_trend": trend, + "observed_transitions": [ + f"{item.get('from_version')} -> {item.get('to_version')}" + for item in diffs + if item.get("from_version") and item.get("to_version") + ], + "evidence": _implementation_evidence_lines(diffs), + } + ) + + themes = [ + dict(item) + for item in direction.get("themes", []) + if isinstance(item, Mapping) + and item.get("name") + and item.get("implementation_pattern") + and not _OPERATIONAL_DIRECTION_RE.search( + f"{item.get('name', '')} {item.get('implementation_pattern', '')}" + ) + ] + if not themes: + observed = [item["name"] for item in packages if item["evidence_status"] != "no-transition"] + themes = [ + { + "name": "Inspected implementation evolution", + "implementation_pattern": ( + "Recent release transitions are summarized from source and artifact " + "fingerprints; opaque runtime internals remain explicitly out of scope." + if observed + else "No adjacent release implementations were inspected in this run." + ), + "packages": observed, + "evidence": [ + f"{len(implementation_diffs)} adjacent release transition(s) inspected." + ], + } + ] + + uncertainty = [ + str(item) + for item in direction.get("uncertainty", []) + if isinstance(item, str) and not _OPERATIONAL_DIRECTION_RE.search(item) + ] + for package in packages: + if package["evidence_status"] == "unavailable": + uncertainty.append( + f"{package['name']}: implementation snapshots were unavailable for at least " + "one observed interval." + ) + elif package["evidence_status"] == "opaque-runtime": + uncertainty.append( + f"{package['name']}: bundled executable internals cannot be inferred from " + "wheel artifact fingerprints." + ) + for item in implementation_diffs: + package = str(item.get("package") or "unknown package") + for limitation in item.get("limitations", []): + if isinstance(limitation, str) and limitation: + uncertainty.append(f"{package}: {limitation}") + if replaced_operation_advice: + uncertainty.append( + "The runtime returned release-operation advice in the trend field; it was replaced " + "with deterministic implementation evidence." + ) + return { + "packages": packages, + "themes": themes, + "uncertainty": list(dict.fromkeys(uncertainty)), + } + + +def _implementation_evidence_status( + package: str, + diffs: Sequence[Mapping[str, Any]], +) -> str: + if not diffs: + return "no-transition" + statuses = {str(item.get("status") or "unavailable") for item in diffs} + if package == "openai-codex-cli-bin" or "opaque-runtime" in statuses: + return "opaque-runtime" + if statuses == {"unavailable"}: + return "unavailable" + return "observed" + + +def _deterministic_implementation_trend( + package: str, + status: str, + diffs: Sequence[Mapping[str, Any]], +) -> str: + if status == "no-transition": + return "No adjacent release implementations were inspected in this run." + if status == "unavailable": + return "The requested implementation history could not be inspected reliably." + file_changes = sum( + len(item.get(field, [])) + for item in diffs + for field in ("files_added", "files_removed", "files_changed") + ) + definition_changes = sum( + len(item.get(field, [])) + for item in diffs + for field in ("definitions_added", "definitions_removed", "definitions_changed") + ) + if status == "opaque-runtime": + opaque_changes = sum( + len(item.get(field, [])) + for item in diffs + for field in ( + "opaque_artifacts_added", + "opaque_artifacts_removed", + "opaque_artifacts_changed", + ) + ) + return ( + f"Across {len(diffs)} inspected release transition(s), the Python wrapper changed " + f"in {file_changes} file event(s) and {definition_changes} definition event(s); " + f"{opaque_changes} opaque artifact event(s) were observed, but bundled runtime " + "internals are not inspectable from wheels." + ) + if file_changes == 0 and definition_changes == 0: + return ( + f"Across {len(diffs)} inspected release transition(s), the inspectable Python " + "implementation was unchanged at file and definition level." + ) + return ( + f"Across {len(diffs)} inspected release transition(s), {file_changes} Python file " + f"event(s) and {definition_changes} definition event(s) were observed in {package}." + ) + + +def _implementation_evidence_lines( + diffs: Sequence[Mapping[str, Any]], +) -> list[str]: + if not diffs: + return ["No adjacent implementation release snapshots were available."] + lines: list[str] = [] + for item in diffs: + files = ( + len(item.get("files_added", [])), + len(item.get("files_removed", [])), + len(item.get("files_changed", [])), + ) + definitions = ( + len(item.get("definitions_added", [])), + len(item.get("definitions_removed", [])), + len(item.get("definitions_changed", [])), + ) + opaque = ( + len(item.get("opaque_artifacts_added", [])), + len(item.get("opaque_artifacts_removed", [])), + len(item.get("opaque_artifacts_changed", [])), + ) + lines.append( + f"{item.get('from_version')} -> {item.get('to_version')}: " + f"Python files +{files[0]}/-{files[1]}/~{files[2]}, " + f"definitions +{definitions[0]}/-{definitions[1]}/~{definitions[2]}, " + f"opaque artifacts +{opaque[0]}/-{opaque[1]}/~{opaque[2]}, " + f"source lines {item.get('source_lines_before', 0)} -> " + f"{item.get('source_lines_after', 0)}." + ) + notable = [ + *item.get("files_added", []), + *item.get("files_removed", []), + *item.get("files_changed", []), + *item.get("definitions_added", []), + *item.get("definitions_removed", []), + *item.get("definitions_changed", []), + *item.get("opaque_artifacts_added", []), + *item.get("opaque_artifacts_removed", []), + *item.get("opaque_artifacts_changed", []), + ] + if notable: + lines.append("Changed implementation elements: " + ", ".join(notable[:6])) + return lines + + def _stage_system_prompt(stage: str, schema: JsonSchema) -> str: prompt = ( "You are running inside the local SDK evolution agent. " @@ -604,6 +915,20 @@ def _stage_system_prompt(stage: str, schema: JsonSchema) -> str: f"Current stage: {stage}. " f"Output schema: {json.dumps(schema, sort_keys=True)}" ) + if stage == "direction-analysis": + prompt += ( + " Direction means upstream implementation trajectory, not dependency action. " + "Treat implementation_diffs as the primary evidence and API diffs, behavior " + "probes, and release notes only as corroboration. Describe observed changes in " + "modules, files, Python definitions, runtime architecture, capabilities, behavior, " + "and deprecations across exact release intervals. Never recommend upgrading, " + "holding, keeping, resolving, or changing a lockfile; architecture-decision owns " + "release feasibility and action. A package without an inspected interval must say " + "no-transition, not current or hold. For an opaque binary runtime, distinguish " + "observable wrapper/artifact changes from unknowable executable internals. Themes " + "must describe implementation patterns supported by implementation-diff evidence, " + "not package freshness, resolver state, or reporting mode." + ) if stage in {"architecture-decision", "review"}: prompt += ( " Deterministic gate policy: candidate API diffs prove API shape drift, " @@ -617,8 +942,18 @@ def _stage_system_prompt(stage: str, schema: JsonSchema) -> str: "that the removed symbols are used. Failed, incomplete, malformed, or " "internally inconsistent behavior evidence; missing exact candidate API " "transitions; unavailable required release-note evidence; " - "reviewer-identified unsupported vendor behavior, or recursive " - "runtime-contract impact remain hard blockers. Release-note status found " + "or reviewer-identified unsupported vendor behavior remain hard blockers. " + "A current direct-dependency upper bound that excludes a newer SDK is not " + "a manual-design blocker when independent discovery found it, prospective " + "resolution selects it, and its exact API and behavior evidence passes; the " + "implementation lane is explicitly designed to widen that excluding bound. " + "Likewise, an active analysis-runtime dependency candidate is recursive " + "impact but not by itself a blocker when exact candidate probes pass and the " + "self-adaptation plan requires a post-update rerun through that runtime. " + "A newer standalone Codex CLI is not required evidence when the Codex SDK " + "candidate pins and the resolver selects a different exact CLI candidate. " + "Recursive impact without such a plan, or with failed/incomplete candidate " + "evidence, remains a hard blocker. Release-note status found " "is direct release-note evidence. Status no-matching-version is source " "coverage with explicit uncertainty, not unavailable evidence." ) @@ -627,6 +962,10 @@ def _stage_system_prompt(stage: str, schema: JsonSchema) -> str: return prompt +def _progress(message: str) -> None: + print(f"[sdk-evolution] {message}", flush=True) + + def _stage_permissions(runtime: AgentRuntime, *, write_enabled: bool) -> PermissionProfile: permissions = PermissionProfile( mode=PermissionMode.CAUTIOUS if write_enabled else PermissionMode.STRICT, @@ -665,8 +1004,13 @@ def _fixture_payload(stage: str, task: AgentTask) -> dict[str, Any]: packages = [ { "name": package.get("name"), - "direction": "unknown", - "evidence": ["deterministic package metadata"], + "evidence_status": "no-transition", + "implementation_trend": ( + "Fixture runtime defers interpretation to deterministic implementation " + "fingerprints." + ), + "observed_transitions": [], + "evidence": ["deterministic implementation diff inventory"], } for package in source.get("evidence", {}).get("packages", []) if isinstance(package, dict) @@ -675,11 +1019,16 @@ def _fixture_payload(stage: str, task: AgentTask) -> dict[str, Any]: "packages": packages, "themes": [ { - "name": "runtime SDK evolution", - "summary": "Fixture runtime records evidence for human or real-runtime review.", + "name": "runtime implementation evolution", + "implementation_pattern": ( + "Fixture runtime records implementation evidence for human or " + "real-runtime review." + ), + "packages": [item["name"] for item in packages], + "evidence": ["implementation diff inventory"], } ], - "uncertainty": ["fake runtime cannot infer real upstream product direction"], + "uncertainty": ["fake runtime cannot infer upstream implementation intent"], } if stage == "architecture-decision": return { diff --git a/pyproject.toml b/pyproject.toml index 70bb681..c6eb9f7 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "hatchling.build" [project] name = "agent-runtime-kit" -version = "0.5.0" +version = "0.5.1" description = "One typed runtime API for Claude, Codex, and Antigravity agent SDKs." readme = "README.md" requires-python = ">=3.10" @@ -48,12 +48,12 @@ dependencies = ["jsonschema>=4.18,<5"] claude = ["claude-agent-sdk>=0.2.87,<0.3"] # 0.1.0b3 floor: earlier betas pin a codex binary with no manylinux wheels, # making the extra uninstallable on glibc Linux. -codex = ["openai-codex>=0.1.0b3,<0.145"] +codex = ["openai-codex>=0.1.0b3,<0.148"] antigravity = ["google-antigravity>=0.1.2,<0.2"] all = [ "claude-agent-sdk>=0.2.87,<0.3", "google-antigravity>=0.1.2,<0.2", - "openai-codex>=0.1.0b3,<0.145", + "openai-codex>=0.1.0b3,<0.148", ] [project.urls] @@ -65,6 +65,7 @@ Issues = "https://github.com/ebarti/agent-runtime-kit/issues" dev = [ "build>=1.2", "mypy>=1.8", + "packaging>=24,<27", "pydantic>=2,<3", "pytest>=8.0", "pytest-asyncio>=0.23", @@ -75,13 +76,6 @@ dev = [ ] [tool.uv] -# Supply-chain freshness delay: never resolve releases younger than 8 days, so a -# freshly compromised upload cannot reach the lockfile before the ecosystem has -# had time to notice. Committed here (not just a developer env var) so CI and -# every contributor resolve under the same policy and `uv sync --locked` agrees -# everywhere. Monitored vendor SDK/runtime packages are exempt so the audited -# evolution workflow and ordinary locked CI share one current dependency state. -exclude-newer = "8 days" # openai-codex-cli-bin ships per-platform binary wheels; require the platforms we # develop and run CI on so resolution never locks a version missing one of them. required-environments = [ @@ -92,12 +86,6 @@ required-environments = [ # marker here opts that one package into pre-release resolution. constraint-dependencies = ["openai-codex-cli-bin>=0.134.0a1"] -[tool.uv.exclude-newer-package] -claude-agent-sdk = false -google-antigravity = false -openai-codex = false -openai-codex-cli-bin = false - [tool.hatch.build.targets.wheel] packages = ["src/agent_runtime_kit"] diff --git a/src/agent_runtime_kit/adapters/antigravity.py b/src/agent_runtime_kit/adapters/antigravity.py index e470f7e..ea42c06 100644 --- a/src/agent_runtime_kit/adapters/antigravity.py +++ b/src/agent_runtime_kit/adapters/antigravity.py @@ -8,6 +8,7 @@ import logging import os import re +import uuid from collections.abc import Mapping from dataclasses import dataclass from pathlib import Path @@ -63,6 +64,7 @@ logger = logging.getLogger(__name__) _MCP_SERVER_NAME_PATTERN = re.compile(r"^[a-zA-Z0-9_-]+$") +_ANTIGRAVITY_CONVERSATION_NAMESPACE = uuid.UUID("7e47bc2d-a8c4-4d2b-9c27-3aeef32142ba") class AntigravityAgentRuntime: @@ -97,9 +99,7 @@ def __init__( reuse_process: bool = False, ) -> None: self._default_model = default_model - self._supported_models = validate_model_configuration( - default_model, supported_models - ) + self._supported_models = validate_model_configuration(default_model, supported_models) self._api_key = api_key self._vertex = vertex self._project = project @@ -170,9 +170,7 @@ async def check_readiness(self) -> RuntimeReadiness: metadata={"failure": "adc-probe", "error_type": type(exc).__name__}, ) project = ( - self._project - or _env_first("GOOGLE_CLOUD_PROJECT", "GCLOUD_PROJECT") - or adc.project + self._project or _env_first("GOOGLE_CLOUD_PROJECT", "GCLOUD_PROJECT") or adc.project ) if not adc.credentials_available or not project: return self._missing_credentials_readiness(availability) @@ -187,9 +185,7 @@ async def check_readiness(self) -> RuntimeReadiness: }, ) - def _missing_credentials_readiness( - self, availability: RuntimeAvailability - ) -> RuntimeReadiness: + def _missing_credentials_readiness(self, availability: RuntimeAvailability) -> RuntimeReadiness: return RuntimeReadiness.not_ready( self.kind, reason=AvailabilityReason.MISSING_CREDENTIALS, @@ -351,7 +347,7 @@ def _build_config( "capabilities": capabilities, "policies": policies, "workspaces": _workspaces(task), - "conversation_id": _conversation_id(task), + "conversation_id": _provider_conversation_id(_conversation_id(task)), "save_dir": str(self._runtime_dir("antigravity-sessions")), "app_data_dir": str(self._runtime_dir("antigravity-app-data")), "response_schema": dict(schema) if schema is not None else None, @@ -405,15 +401,18 @@ async def _invoke( sdk=sdk, config=config, ) - structured_output, usage_metadata, session_id, stop_reason = ( - await self._chat_agent( - task, - agent=agent, - sdk=sdk, - text_parts=text_parts, - tool_calls=tool_calls, - wants_structured=schema is not None, - ) + ( + structured_output, + usage_metadata, + session_id, + stop_reason, + ) = await self._chat_agent( + task, + agent=agent, + sdk=sdk, + text_parts=text_parts, + tool_calls=tool_calls, + wants_structured=schema is not None, ) except BaseException: # Evict the reused agent on any non-normal exit — including @@ -431,20 +430,19 @@ async def _invoke( raise else: async with sdk.agent_cls(config) as agent: - structured_output, usage_metadata, session_id, stop_reason = ( - await self._chat_agent( - task, - agent=agent, - sdk=sdk, - text_parts=text_parts, - tool_calls=tool_calls, - wants_structured=schema is not None, - ) + structured_output, usage_metadata, session_id, stop_reason = await self._chat_agent( + task, + agent=agent, + sdk=sdk, + text_parts=text_parts, + tool_calls=tool_calls, + wants_structured=schema is not None, ) - # Fall back to the caller's conversation id when the SDK does not echo one, - # so a resumed task always returns a usable session_id (matches Claude). - session_id = session_id or _conversation_id(task) + # Keep the public runtime session id stable even when the provider-facing + # value had to be expanded to satisfy Antigravity's minimum-length rule. + # A provider-created id is still returned when the caller supplied none. + session_id = _conversation_id(task) or session_id process_metadata = ( self._process_reuse_metadata(process_reused) if self._reuse_process else None @@ -856,9 +854,8 @@ def _google_adc_project() -> str | None: def _is_missing_google_credentials(exc: Exception) -> bool: - return ( - type(exc).__name__ == "DefaultCredentialsError" - and type(exc).__module__.startswith("google.auth") + return type(exc).__name__ == "DefaultCredentialsError" and type(exc).__module__.startswith( + "google.auth" ) @@ -1004,6 +1001,20 @@ def _conversation_id(task: AgentTask) -> str | None: return task.session_id +def _provider_conversation_id(conversation_id: str | None) -> str | None: + """Map short public session ids to stable Antigravity-compatible ids. + + Antigravity validates ``conversation_id`` at config construction time and + currently requires at least 32 characters. The public runtime contract does + not impose that vendor-specific restriction, so short ids are deterministically + namespaced while already-compatible ids pass through unchanged. + """ + + if conversation_id is None or len(conversation_id) >= 32: + return conversation_id + return str(uuid.uuid5(_ANTIGRAVITY_CONVERSATION_NAMESPACE, conversation_id)) + + def _agent_key(task: AgentTask, config: Any) -> tuple[Any, ...]: conversation_id = _conversation_id(task) if conversation_id: @@ -1036,9 +1047,7 @@ def _usage_from(value: Any) -> Usage: else None ), output_tokens=( - output_tokens + thoughts - if output_tokens is not None and thoughts is not None - else None + output_tokens + thoughts if output_tokens is not None and thoughts is not None else None ), cache_read_tokens=cache_read, total_tokens=total, diff --git a/src/agent_runtime_kit/compatibility.py b/src/agent_runtime_kit/compatibility.py index a5ff706..b3b097e 100644 --- a/src/agent_runtime_kit/compatibility.py +++ b/src/agent_runtime_kit/compatibility.py @@ -66,17 +66,17 @@ def __post_init__(self) -> None: package="claude-agent-sdk", module="claude_agent_sdk", version_specifier=">=0.2.87,<0.3", - tested_version="0.2.128", + tested_version="0.2.148", ), RuntimeCompatibility( kind=AgentRuntimeKind.CODEX_AGENT_SDK, extra="codex", package="openai-codex", module="openai_codex", - version_specifier=">=0.1.0b3,<0.145", - tested_version="0.144.4", + version_specifier=">=0.1.0b3,<0.148", + tested_version="0.147.0", tested_runtime_dependencies=( - PackageVersion(package="openai-codex-cli-bin", version="0.144.4"), + PackageVersion(package="openai-codex-cli-bin", version="0.147.0"), ), ), RuntimeCompatibility( @@ -85,7 +85,7 @@ def __post_init__(self) -> None: package="google-antigravity", module="google.antigravity", version_specifier=">=0.1.2,<0.2", - tested_version="0.1.8", + tested_version="0.1.15", ), ) diff --git a/tests/test_antigravity_adapter.py b/tests/test_antigravity_adapter.py index 010b8d8..43515e5 100644 --- a/tests/test_antigravity_adapter.py +++ b/tests/test_antigravity_adapter.py @@ -19,6 +19,7 @@ ) from agent_runtime_kit._errors import UnsupportedTaskInputError from agent_runtime_kit.adapters import AntigravityAgentRuntime +from agent_runtime_kit.adapters.antigravity import _provider_conversation_id from agent_runtime_kit.testing import RecordingEventSink @@ -389,7 +390,7 @@ class NarrowConfig: def __init__( self, *, - model: str | None = None, + model: str | None = None, api_key: str | None = None, capabilities: Any = None, policies: Any = None, @@ -480,7 +481,7 @@ class NoWorkspacesConfig: def __init__( self, *, - model: str | None = None, + model: str | None = None, api_key: str | None = None, capabilities: Any = None, policies: Any = None, @@ -593,6 +594,53 @@ async def chat(self, prompt: str) -> NoIdResponse: assert result.session_id == "conv-77" +def test_antigravity_short_public_session_ids_map_to_stable_provider_ids() -> None: + first = _provider_conversation_id("conv-77") + second = _provider_conversation_id("conv-77") + + assert first == second + assert first != "conv-77" + assert first is not None + assert len(first) >= 32 + assert _provider_conversation_id(None) is None + assert _provider_conversation_id("a" * 32) == "a" * 32 + + +@pytest.mark.asyncio +async def test_antigravity_maps_provider_id_but_preserves_public_session_id( + tmp_path: Path, +) -> None: + seen: dict[str, Any] = {} + + class NoIdAgent: + def __init__(self, config: FakeConfig) -> None: + seen.update(config.kwargs) + self.conversation_id = None + + async def __aenter__(self) -> NoIdAgent: + return self + + async def __aexit__(self, *args: object) -> None: + return None + + async def chat(self, prompt: str) -> FakeResponse: + return FakeResponse(prompt, _chunks) + + runtime = AntigravityAgentRuntime( + api_key="key", + data_dir=tmp_path, + agent_cls=NoIdAgent, + config_cls=FakeConfig, + types_module=FakeTypes, + policy_module=FakePolicy, + ) + + result = await runtime.run(AgentTask(goal="x", session_id="conv-77")) + + assert seen["conversation_id"] == _provider_conversation_id("conv-77") + assert result.session_id == "conv-77" + + @pytest.mark.asyncio async def test_antigravity_max_tokens_stop_reason_fails(tmp_path: Path) -> None: class TruncatedResponse: @@ -1076,9 +1124,7 @@ async def test_antigravity_strict_honors_read_only_allow_list(tmp_path: Path) -> await runtime.run( AgentTask( goal="task", - permissions=PermissionProfile( - mode=PermissionMode.STRICT, allowed_tools=("view_file",) - ), + permissions=PermissionProfile(mode=PermissionMode.STRICT, allowed_tools=("view_file",)), ) ) @@ -1251,9 +1297,7 @@ async def test_antigravity_rejects_network(tmp_path: Path) -> None: runtime = make_runtime(data_dir=tmp_path) with pytest.raises(UnsupportedTaskInputError): - await runtime.run( - AgentTask(goal="task", permissions=PermissionProfile(network=True)) - ) + await runtime.run(AgentTask(goal="task", permissions=PermissionProfile(network=True))) @pytest.mark.asyncio diff --git a/tests/test_sdk_evolution_agent.py b/tests/test_sdk_evolution_agent.py index 230178a..4e3a30b 100644 --- a/tests/test_sdk_evolution_agent.py +++ b/tests/test_sdk_evolution_agent.py @@ -1,5 +1,6 @@ from __future__ import annotations +import hashlib import json import os import stat @@ -38,22 +39,38 @@ probe_current_package, summarize_behavior, ) -from examples.sdk_evolution_agent.cli import RunOptions, _collect_snapshots, parse_args, run_agent +from examples.sdk_evolution_agent.cli import ( + RunOptions, + _collect_implementation_snapshots, + _collect_snapshots, + _run_local_sdk_update, + _should_create_pr, + _verification_passed, + parse_args, + run_agent, +) from examples.sdk_evolution_agent.collectors import ( ResolverTransition, build_refresh_preview_command, + candidate_transitions, + candidate_update_versions, collect_evidence, cutoff_free_env, parse_refresh_transitions, + read_pyproject_dependency_specs, run_lock_update, run_refresh_preview, + widen_project_dependency_bounds, ) from examples.sdk_evolution_agent.current_state import build_current_state +from examples.sdk_evolution_agent.inspection import candidate_environment_cache from examples.sdk_evolution_agent.models import ( ApiSnapshot, BehaviorDiff, BehaviorProbeResult, CommandResult, + ImplementationDefinition, + ImplementationFile, RunContext, SourceRef, to_jsonable, @@ -72,6 +89,7 @@ ) from examples.sdk_evolution_agent.snapshots import ( DEFAULT_MODULES, + diff_implementation_snapshots, diff_snapshots, snapshot_candidate_in_venv, snapshot_current_api, @@ -88,9 +106,11 @@ run_stage, with_behavior_probe_guard, with_candidate_api_diff_guard, + with_implementation_direction_guard, with_manual_design_gate, with_recursive_impact, with_release_note_guard, + with_resolver_preview_guard, ) @@ -142,11 +162,16 @@ def runner( assert seen["command"] == build_refresh_preview_command( ("claude-agent-sdk", "google-antigravity") ) - assert seen["command"][-4:] == ( - "--exclude-newer-package", - "claude-agent-sdk=false", - "--exclude-newer-package", - "google-antigravity=false", + assert seen["command"] == ( + "uv", + "lock", + "--dry-run", + "--exclude-newer", + "false", + "-P", + "claude-agent-sdk", + "-P", + "google-antigravity", ) assert "UV_EXCLUDE_NEWER" not in seen["env"] assert result.removed_env == ("UV_EXCLUDE_NEWER",) @@ -220,19 +245,103 @@ def runner( assert seen["command"] == ( "uv", "lock", + "--exclude-newer", + "false", "-P", "claude-agent-sdk", "-P", "google-antigravity", - "--exclude-newer-package", - "claude-agent-sdk=false", - "--exclude-newer-package", - "google-antigravity=false", ) assert "UV_EXCLUDE_NEWER" not in seen["env"] assert result.removed_env == ("UV_EXCLUDE_NEWER",) +def test_pyproject_specs_are_structural_and_ignore_mentions_in_comments( + tmp_path: Path, +) -> None: + pyproject = tmp_path / "pyproject.toml" + pyproject.write_text( + """ +[project] +name = "fixture" +version = "0.0.0" +keywords = ["openai-codex", "google-antigravity"] + +[project.optional-dependencies] +codex = ["openai-codex>=0.1.0b3,<0.145"] +all = ["openai-codex>=0.1.0b3,<0.145", "google-antigravity>=0.1.2,<0.2"] + +[tool.uv] +# openai-codex-cli-bin ships platform wheels; this comment is not a requirement. +constraint-dependencies = ["openai-codex-cli-bin>=0.134.0a1"] +""", + encoding="utf-8", + ) + + assert read_pyproject_dependency_specs(pyproject) == { + "openai-codex": "openai-codex>=0.1.0b3,<0.145", + "openai-codex-cli-bin": "openai-codex-cli-bin>=0.134.0a1", + "google-antigravity": "google-antigravity>=0.1.2,<0.2", + } + + +def test_candidate_inventory_is_independent_of_constrained_resolver_output() -> None: + evidence = { + "packages": [ + { + "name": "openai-codex", + "locked_version": "0.144.4", + "installed_version": "0.144.4", + "latest_version": "0.147.0", + "candidate_version": "0.147.0", + "candidate_status": "blocked-by-project-constraint", + } + ], + "refresh_preview": {"returncode": 0, "stdout": "", "stderr": ""}, + } + + assert parse_refresh_transitions(evidence) == () + assert candidate_update_versions(evidence) == {"openai-codex": "0.147.0"} + assert candidate_transitions(evidence) == ( + ResolverTransition("openai-codex", "0.144.4", "0.147.0"), + ) + + +def test_widen_project_dependency_bounds_only_relaxes_excluding_upper_bound( + tmp_path: Path, +) -> None: + pyproject = tmp_path / "pyproject.toml" + pyproject.write_text( + """ +[project] +name = "fixture" +version = "0.0.0" + +[project.optional-dependencies] +codex = ["openai-codex>=0.1.0b3,<0.145"] +claude = ["claude-agent-sdk>=0.2.87,<0.3"] +all = [ + "openai-codex>=0.1.0b3,<0.145", + "claude-agent-sdk>=0.2.87,<0.3", +] +""", + encoding="utf-8", + ) + + changes = widen_project_dependency_bounds( + pyproject, + { + "openai-codex": "0.147.0", + "claude-agent-sdk": "0.2.145", + }, + ) + + text = pyproject.read_text(encoding="utf-8") + assert text.count("openai-codex>=0.1.0b3,<0.148") == 2 + assert "claude-agent-sdk>=0.2.87,<0.3" in text + assert changes == ("openai-codex>=0.1.0b3,<0.145 -> openai-codex>=0.1.0b3,<0.148",) + + def test_collect_evidence_records_versions_and_sources(tmp_path: Path) -> None: (tmp_path / "pyproject.toml").write_text( """ @@ -266,6 +375,145 @@ def test_collect_evidence_records_versions_and_sources(tmp_path: Path) -> None: assert evidence["adapter_sources"] +def test_collect_evidence_prospectively_resolves_candidates_hidden_by_caps( + tmp_path: Path, +) -> None: + (tmp_path / "pyproject.toml").write_text( + """ +[project] +name = "fixture" +version = "0.0.0" + +[project.optional-dependencies] +codex = ["openai-codex>=0.1.0b3,<0.145"] +""", + encoding="utf-8", + ) + (tmp_path / "uv.lock").write_text( + """ +[[package]] +name = "openai-codex" +version = "0.144.4" + +[[package]] +name = "openai-codex-cli-bin" +version = "0.144.4" +""", + encoding="utf-8", + ) + + metadata = { + "openai-codex": {"info": {"version": "0.147.0"}, "releases": {}}, + "openai-codex-cli-bin": {"info": {"version": "0.149.0"}, "releases": {}}, + } + preview_roots: list[Path] = [] + + def runner( + command: tuple[str, ...], + *, + cwd: Path | None = None, + env: dict[str, str] | None = None, + ) -> CommandResult: + del env + assert cwd is not None + preview_roots.append(cwd) + if cwd == tmp_path: + return CommandResult(command=command, returncode=0) + assert "openai-codex>=0.1.0b3,<0.148" in (cwd / "pyproject.toml").read_text() + return CommandResult( + command=command, + returncode=0, + stderr=( + "Update openai-codex v0.144.4 -> v0.147.0\n" + "Update openai-codex-cli-bin v0.144.4 -> v0.147.0\n" + ), + ) + + evidence = collect_evidence( + tmp_path, + packages=("openai-codex", "openai-codex-cli-bin"), + include_refresh_preview=True, + pypi_client=lambda package: metadata[package], + command_runner=runner, + ) + + assert len(preview_roots) == 2 + assert candidate_update_versions(evidence) == { + "openai-codex": "0.147.0", + "openai-codex-cli-bin": "0.147.0", + } + packages = {item["name"]: item for item in evidence["packages"]} + assert packages["openai-codex"]["candidate_status"] == "blocked-by-project-constraint" + assert packages["openai-codex-cli-bin"]["candidate_status"] == "prospective-coupled-candidate" + assert packages["openai-codex-cli-bin"]["latest_version"] == "0.149.0" + + +def test_codex_cli_staged_artifact_is_not_reported_as_a_blocked_candidate( + tmp_path: Path, +) -> None: + (tmp_path / "pyproject.toml").write_text( + """ +[project] +name = "fixture" +version = "0.0.0" + +[project.optional-dependencies] +codex = ["openai-codex>=0.1.0b3,<0.148"] + +[tool.uv] +constraint-dependencies = ["openai-codex-cli-bin>=0.134.0a1"] +""", + encoding="utf-8", + ) + (tmp_path / "uv.lock").write_text( + """ +[[package]] +name = "openai-codex" +version = "0.147.0" + +[[package]] +name = "openai-codex-cli-bin" +version = "0.147.0" +""", + encoding="utf-8", + ) + metadata = { + "openai-codex": { + "info": { + "version": "0.147.0", + "requires_dist": ["openai-codex-cli-bin==0.147.0"], + }, + "releases": {"0.147.0": [{}]}, + }, + "openai-codex-cli-bin": { + "info": {"version": "0.149.0"}, + "releases": {"0.147.0": [{}], "0.149.0": [{}]}, + }, + } + + evidence = collect_evidence( + tmp_path, + packages=("openai-codex", "openai-codex-cli-bin"), + include_refresh_preview=True, + pypi_client=lambda package: metadata[package], + command_runner=lambda command, **kwargs: CommandResult( + command=command, + returncode=0, + stderr="No lockfile changes detected\n", + ), + ) + + packages = {item["name"]: item for item in evidence["packages"]} + cli = packages["openai-codex-cli-bin"] + assert cli["latest_version"] == "0.149.0" + assert cli["sdk_selected_version"] == "0.147.0" + assert cli["candidate_version"] is None + assert cli["candidate_status"] == "sdk-coupled-no-update" + assert "staged runtime artifact" in cli["candidate_reason"] + assert "blocked" not in cli["candidate_status"] + assert candidate_update_versions(evidence) == {} + + def test_release_notes_collects_matching_update_source() -> None: notes = collect_release_notes( [ @@ -848,6 +1096,65 @@ def fake_run(args: Any, **kwargs: Any) -> Any: assert env.get("HOME") != os.environ.get("HOME") +def test_candidate_snapshot_and_behavior_reuse_one_isolated_install( + monkeypatch: pytest.MonkeyPatch, +) -> None: + calls: list[tuple[str, ...]] = [] + + def fake_run(args: Any, **kwargs: Any) -> Any: + del kwargs + command = tuple(args) + calls.append(command) + if len(command) > 2 and command[1] == "-c" and len(command) == 6: + payload = ( + '{"package":"claude-agent-sdk","version":"9.9.9",' + '"module":"claude_agent_sdk","members":[],"import_error":null}' + ) + elif len(command) > 2 and command[1] == "-c": + payload = ( + '[{"package":"claude-agent-sdk","version":"9.9.9",' + '"scope":"candidate","probe":"adapter-contract","status":"pass",' + '"summary":"ok","details":{"missing":[]}}]' + ) + else: + payload = "" + return types.SimpleNamespace( + stdout=payload, + stderr="", + returncode=0, + ) + + monkeypatch.setattr("examples.sdk_evolution_agent.snapshots.subprocess.run", fake_run) + + with candidate_environment_cache(): + snapshot = snapshot_candidate_in_venv("claude-agent-sdk", "9.9.9") + (probe,) = probe_candidate_in_venv("claude-agent-sdk", "9.9.9") + + install_calls = [call for call in calls if "pip" in call and "install" in call] + assert len(install_calls) == 1 + assert len(calls) == 4 # venv, install, snapshot, cached-environment behavior probe + assert snapshot.import_error is None + assert probe.status == "pass" + + +def test_antigravity_probe_uses_constructible_provider_session_ids() -> None: + class LengthCheckingConfig: + def __init__(self, *, conversation_id: str) -> None: + if len(conversation_id) < 32: + raise ValueError("conversation id must have at least 32 characters") + + inputs = behavior_module._probe_inputs("google-antigravity")["conversation_ids"] + + assert "conv-77" not in inputs + assert behavior_module._antigravity_construction_failures(LengthCheckingConfig, inputs) == [] + failures = behavior_module._antigravity_construction_failures( + LengthCheckingConfig, + ["conv-77"], + ) + assert len(failures) == 1 + assert "ValueError" in failures[0] + + @pytest.mark.parametrize( ("package", "failure_call", "error", "failure_step", "probe"), [ @@ -976,8 +1283,7 @@ def test_codex_cli_exception_redaction_discards_entire_path_suffix( ) -> None: detail = behavior_module._safe_exception_detail( ValueError( - f"REDACTION_SENTINEL at {private_path} trailing {secret_suffix} " - + ("x" * 1_000) + f"REDACTION_SENTINEL at {private_path} trailing {secret_suffix} " + ("x" * 1_000) ) ) @@ -1012,15 +1318,11 @@ def test_codex_cli_binary_probe_contains_each_failure( def metadata_version(package: str) -> str: assert package == "openai-codex-cli-bin" if failure == "metadata": - raise RuntimeError( - f"METADATA_SENTINEL at {metadata_path} " + ("m" * 1_000) - ) + raise RuntimeError(f"METADATA_SENTINEL at {metadata_path} " + ("m" * 1_000)) return "1.2.3" def helper_error() -> Path: - raise ValueError( - f"HELPER_SENTINEL at {unc_path} " + ("h" * 1_000) - ) + raise ValueError(f"HELPER_SENTINEL at {unc_path} " + ("h" * 1_000)) module = types.SimpleNamespace(bundled_codex_path=lambda: tmp_path / "codex") if failure == "helper-missing": @@ -1036,8 +1338,7 @@ def import_module(name: str) -> Any: assert name == "codex_cli_bin" if failure == "module": raise ModuleNotFoundError( - f"MODULE_SENTINEL: No module named codex_cli_bin at {drive_path} " - + ("i" * 1_000) + f"MODULE_SENTINEL: No module named codex_cli_bin at {drive_path} " + ("i" * 1_000) ) return module @@ -1084,24 +1385,18 @@ def test_codex_cli_binary_probe_redacts_path_conversion_and_file_check_errors( expected_summary: str, secret_suffix: str, ) -> None: - conversion_url = ( - "https://example.invalid/private folder/CONVERSION_SECRET_SUFFIX" - ) + conversion_url = "https://example.invalid/private folder/CONVERSION_SECRET_SUFFIX" file_check_path = tmp_path / "private file check" / "FILE_CHECK_SECRET_SUFFIX" class InvalidPath: def __fspath__(self) -> str: - raise ValueError( - f"CONVERSION_SENTINEL at {conversion_url} " + ("c" * 1_000) - ) + raise ValueError(f"CONVERSION_SENTINEL at {conversion_url} " + ("c" * 1_000)) if failure == "file-check": def fail_file_check(self: Path) -> bool: del self - raise OSError( - f"FILE_CHECK_SENTINEL at {file_check_path} " + ("f" * 1_000) - ) + raise OSError(f"FILE_CHECK_SENTINEL at {file_check_path} " + ("f" * 1_000)) monkeypatch.setattr(Path, "is_file", fail_file_check) @@ -1200,37 +1495,25 @@ def test_embedded_codex_cli_failures_keep_binary_probe_label( expected_summary: str, secret_suffix: str | None, ) -> None: - metadata_path = ( - tmp_path / "private metadata" / "EMBEDDED_METADATA_SECRET_SUFFIX" - ) + metadata_path = tmp_path / "private metadata" / "EMBEDDED_METADATA_SECRET_SUFFIX" drive_path = r"D:\Program Files\Private\EMBEDDED_MODULE_SECRET_SUFFIX" unc_path = r"\\candidate-host\Private Share\EMBEDDED_HELPER_SECRET_SUFFIX" - conversion_url = ( - "https://example.invalid/private folder/EMBEDDED_CONVERSION_SECRET_SUFFIX" - ) - file_check_path = ( - tmp_path / "private file check" / "EMBEDDED_FILE_CHECK_SECRET_SUFFIX" - ) + conversion_url = "https://example.invalid/private folder/EMBEDDED_CONVERSION_SECRET_SUFFIX" + file_check_path = tmp_path / "private file check" / "EMBEDDED_FILE_CHECK_SECRET_SUFFIX" def fail_metadata(package: str) -> str: if failure == "metadata": raise RuntimeError( - f"EMBEDDED_METADATA_SENTINEL for {package} at {metadata_path} " - + ("m" * 1_000) + f"EMBEDDED_METADATA_SENTINEL for {package} at {metadata_path} " + ("m" * 1_000) ) return "1.2.3" def helper_error() -> Path: - raise ValueError( - f"EMBEDDED_HELPER_SENTINEL at {unc_path} " + ("h" * 1_000) - ) + raise ValueError(f"EMBEDDED_HELPER_SENTINEL at {unc_path} " + ("h" * 1_000)) class InvalidPath: def __fspath__(self) -> str: - raise ValueError( - f"EMBEDDED_CONVERSION_SENTINEL at {conversion_url} " - + ("c" * 1_000) - ) + raise ValueError(f"EMBEDDED_CONVERSION_SENTINEL at {conversion_url} " + ("c" * 1_000)) def bundled_path() -> object: if failure == "helper-error": @@ -1243,19 +1526,14 @@ def bundled_path() -> object: def fail_file_check(self: Path) -> bool: del self - raise OSError( - f"EMBEDDED_FILE_CHECK_SENTINEL at {file_check_path} " - + ("f" * 1_000) - ) + raise OSError(f"EMBEDDED_FILE_CHECK_SENTINEL at {file_check_path} " + ("f" * 1_000)) monkeypatch.setattr(Path, "is_file", fail_file_check) def import_module(name: str) -> Any: assert name == "codex_cli_bin" if failure == "module": - raise ModuleNotFoundError( - f"EMBEDDED_MODULE_SENTINEL at {drive_path} " + ("i" * 1_000) - ) + raise ModuleNotFoundError(f"EMBEDDED_MODULE_SENTINEL at {drive_path} " + ("i" * 1_000)) return module helper = bundled_path @@ -1407,9 +1685,7 @@ def test_behavior_summary_marks_errors_skips_missing_and_malformed_as_incomplete def test_behavior_summary_contains_malformed_nested_contract_details() -> None: transition = ResolverTransition("claude-agent-sdk", "1.0.0", "2.0.0") - baseline = _probe( - "claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []} - ) + baseline = _probe("claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []}) malformed_candidate = BehaviorProbeResult( package="claude-agent-sdk", version="2.0.0", @@ -1436,9 +1712,7 @@ def test_behavior_summary_contains_malformed_nested_contract_details() -> None: def test_behavior_summary_treats_missing_fields_plus_error_as_contract_failure() -> None: transition = ResolverTransition("claude-agent-sdk", "1.0.0", "2.0.0") - baseline = _probe( - "claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []} - ) + baseline = _probe("claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []}) candidate = _probe( "claude-agent-sdk", "2.0.0", @@ -1545,9 +1819,7 @@ def test_behavior_summary_status_precedence_is_fail_then_incomplete_then_changed def test_behavior_probe_guard_blocks_complete_breaking_candidate_payload() -> None: - baseline = _probe( - "google-antigravity", "1.0.0", "current-baseline", "pass", {"missing": []} - ) + baseline = _probe("google-antigravity", "1.0.0", "current-baseline", "pass", {"missing": []}) candidate = _probe( "google-antigravity", "2.0.0", @@ -1610,9 +1882,7 @@ def test_behavior_probe_guard_allows_valid_pass_and_changed_evidence() -> None: ) assert ( - with_behavior_probe_guard(architecture, _sdk_evidence(), pass_payload)[ - "safe_to_implement" - ] + with_behavior_probe_guard(architecture, _sdk_evidence(), pass_payload)["safe_to_implement"] is True ) assert ( @@ -1642,9 +1912,7 @@ def test_behavior_probe_guard_blocks_self_declared_empty_expectations() -> None: assert guarded["safe_to_implement"] is False assert guarded["manual_design_required"] is True - assert "contradicts deterministic evidence" in " ".join( - guarded["findings"][-1]["evidence"] - ) + assert "contradicts deterministic evidence" in " ".join(guarded["findings"][-1]["evidence"]) def test_behavior_probe_guard_blocks_failed_incomplete_and_invalid_evidence() -> None: @@ -1815,9 +2083,7 @@ def test_report_exposes_snapshot_and_behavior_evidence_failures() -> None: def test_report_persists_and_renders_recomputed_behavior_summary(tmp_path: Path) -> None: - baseline = _probe( - "claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []} - ) + baseline = _probe("claude-agent-sdk", "1.0.0", "current-baseline", "pass", {"missing": []}) transition = ResolverTransition("claude-agent-sdk", "1.0.0", "2.0.0") behavior = _behavior_payload( [ @@ -1833,9 +2099,9 @@ def test_report_persists_and_renders_recomputed_behavior_summary(tmp_path: Path) expected_packages=["claude-agent-sdk"], expected_transitions=[transition], ) - behavior["summary"] = _behavior_payload( - [baseline], expected_packages=["claude-agent-sdk"] - )["summary"] + behavior["summary"] = _behavior_payload([baseline], expected_packages=["claude-agent-sdk"])[ + "summary" + ] assert behavior["summary"]["status"] == "pass" report_root = tmp_path / "reports" / "run-1" @@ -1867,9 +2133,7 @@ def test_report_persists_and_renders_recomputed_behavior_summary(tmp_path: Path) review={}, ) - persisted = json.loads( - (report_root / "behavior_summary.json").read_text(encoding="utf-8") - ) + persisted = json.loads((report_root / "behavior_summary.json").read_text(encoding="utf-8")) report = report_path.read_text(encoding="utf-8") assert persisted["status"] == "fail" assert persisted["contract_failure_count"] == 1 @@ -1968,6 +2232,100 @@ def run_new(value: str, *, verbose: bool = False) -> str: assert diff.changed == ("run",) +def test_snapshot_and_diff_actual_python_implementation( + tmp_path: Path, + monkeypatch: pytest.MonkeyPatch, +) -> None: + package = tmp_path / "trend_sdk" + package.mkdir() + implementation = package / "__init__.py" + implementation.write_text( + "def run(value: str) -> str:\n return value\n", + encoding="utf-8", + ) + monkeypatch.syspath_prepend(str(tmp_path)) + before = snapshot_current_api("trend-sdk", version="1.0.0") + implementation.write_text( + ( + "def run(value: str) -> str:\n" + " return value.upper()\n\n" + "def stream(value: str):\n" + " yield value\n" + ), + encoding="utf-8", + ) + after = snapshot_current_api("trend-sdk", version="1.1.0") + + diff = diff_implementation_snapshots(before, after) + + assert before.implementation_status == "observed" + assert before.implementation_files[0].path == "trend_sdk/__init__.py" + assert diff.status == "observed" + assert diff.files_changed == ("trend_sdk/__init__.py",) + assert diff.definitions_added == ("trend_sdk/__init__.py:stream",) + assert diff.definitions_changed == ("trend_sdk/__init__.py:run",) + assert diff.source_lines_after > diff.source_lines_before + + +def test_implementation_diff_marks_cli_binary_internals_opaque() -> None: + wrapper_before = ImplementationFile("codex_cli_bin/__init__.py", "python", "a", 10, 1) + wrapper_after = ImplementationFile("codex_cli_bin/__init__.py", "python", "b", 11, 1) + binary_before = ImplementationFile("codex_cli_bin/codex", "executable", "c", 100) + binary_after = ImplementationFile("codex_cli_bin/codex", "executable", "d", 120) + before = ApiSnapshot( + package="openai-codex-cli-bin", + version="0.147.0", + module="codex_cli_bin", + implementation_files=(wrapper_before, binary_before), + implementation_definitions=( + ImplementationDefinition("codex_cli_bin/__init__.py:path", "function", "x", "a"), + ), + implementation_status="opaque-runtime", + ) + after = ApiSnapshot( + package="openai-codex-cli-bin", + version="0.149.0", + module="codex_cli_bin", + implementation_files=(wrapper_after, binary_after), + implementation_definitions=( + ImplementationDefinition("codex_cli_bin/__init__.py:path", "function", "x", "b"), + ), + implementation_status="opaque-runtime", + ) + + diff = diff_implementation_snapshots(before, after) + + assert diff.status == "opaque-runtime" + assert diff.files_changed == ("codex_cli_bin/__init__.py",) + assert diff.opaque_artifacts_changed == ("codex_cli_bin/codex",) + assert "opaque" in diff.limitations[0] + + +def test_implementation_diff_does_not_treat_an_unavailable_snapshot_as_added_source() -> None: + before = ApiSnapshot( + package="claude-agent-sdk", + version="0.2.146", + module="claude_agent_sdk", + implementation_status="unavailable", + implementation_note="wheel contained no readable source", + ) + after = ApiSnapshot( + package="claude-agent-sdk", + version="0.2.147", + module="claude_agent_sdk", + implementation_files=( + ImplementationFile("claude_agent_sdk/__init__.py", "python", "a", 10, 1), + ), + implementation_status="observed", + ) + + diff = diff_implementation_snapshots(before, after) + + assert diff.status == "unavailable" + assert diff.files_added == () + assert "no readable source" in diff.limitations[0] + + @pytest.mark.parametrize( ("failure_call", "detail"), [ @@ -2190,6 +2548,56 @@ def candidate_snapshot(package: str, version: str) -> ApiSnapshot: ] +def test_collect_snapshots_inspects_newer_upstream_even_when_resolver_is_constrained( + monkeypatch: pytest.MonkeyPatch, +) -> None: + calls: list[tuple[str, str, str | None]] = [] + + def current_snapshot(package: str, *, version: str | None = None) -> ApiSnapshot: + calls.append(("current", package, version)) + return ApiSnapshot(package=package, version=version, module="openai_codex") + + def candidate_snapshot(package: str, version: str) -> ApiSnapshot: + calls.append(("candidate", package, version)) + return ApiSnapshot( + package=package, + version=version, + module="openai_codex", + source="isolated-venv", + ) + + monkeypatch.setattr( + "examples.sdk_evolution_agent.cli.snapshot_current_api", + current_snapshot, + ) + monkeypatch.setattr( + "examples.sdk_evolution_agent.cli.snapshot_candidate_in_venv", + candidate_snapshot, + ) + + _collect_snapshots( + { + "packages": [ + { + "name": "openai-codex", + "locked_version": "0.144.4", + "installed_version": "0.144.4", + "latest_version": "0.147.0", + "candidate_version": "0.147.0", + "candidate_status": "blocked-by-project-constraint", + } + ], + "refresh_preview": {"returncode": 0, "stdout": "", "stderr": ""}, + }, + inspect_candidates=True, + ) + + assert calls == [ + ("current", "openai-codex", "0.144.4"), + ("candidate", "openai-codex", "0.147.0"), + ] + + def test_collect_snapshots_uses_locked_baseline_when_environment_drifted( monkeypatch: pytest.MonkeyPatch, ) -> None: @@ -2380,6 +2788,55 @@ def isolated_snapshot(package: str, version: str) -> ApiSnapshot: assert snapshots[0].import_error == "not installed" +def test_implementation_history_inspects_recent_releases_when_project_is_current( + monkeypatch: pytest.MonkeyPatch, +) -> None: + calls: list[tuple[str, str]] = [] + current = ApiSnapshot( + package="claude-agent-sdk", + version="0.2.147", + module="claude_agent_sdk", + implementation_status="observed", + ) + + def isolated_snapshot(package: str, version: str) -> ApiSnapshot: + calls.append((package, version)) + return ApiSnapshot( + package=package, + version=version, + module="claude_agent_sdk", + implementation_status="observed", + source="isolated-venv", + ) + + monkeypatch.setattr( + "examples.sdk_evolution_agent.cli.snapshot_candidate_in_venv", + isolated_snapshot, + ) + + snapshots = _collect_implementation_snapshots( + { + "packages": [ + { + "name": "claude-agent-sdk", + "locked_version": "0.2.147", + "latest_version": "0.2.147", + "candidate_version": None, + "recent_versions": ["0.2.147", "0.2.146", "0.2.145"], + } + ] + }, + compatibility_snapshots=[current], + inspect_candidates=True, + ) + + assert [snapshot.version for snapshot in snapshots] == ["0.2.145", "0.2.146", "0.2.147"] + assert calls == [ + ("claude-agent-sdk", "0.2.145"), + ("claude-agent-sdk", "0.2.146"), + ] + + def test_candidate_api_diff_guard_blocks_missing_update_diff() -> None: guarded = with_candidate_api_diff_guard( { @@ -2406,6 +2863,64 @@ def test_candidate_api_diff_guard_blocks_missing_update_diff() -> None: ) +@pytest.mark.parametrize( + "preview", + [None, {"returncode": 1, "stderr": "resolution failed"}], +) +def test_resolver_preview_guard_blocks_unproven_candidate_resolution( + preview: dict[str, Any] | None, +) -> None: + evidence: dict[str, Any] = { + "packages": [ + { + "name": "openai-codex", + "locked_version": "0.144.4", + "latest_version": "0.147.0", + "candidate_version": "0.147.0", + } + ] + } + if preview is not None: + evidence["refresh_preview"] = preview + + guarded = with_resolver_preview_guard( + { + "findings": [], + "safe_to_implement": True, + "manual_design_required": False, + "uncertainty": [], + }, + evidence, + ) + + assert guarded["safe_to_implement"] is False + assert guarded["manual_design_required"] is True + assert "resolvable" in guarded["findings"][-1]["summary"] + + +def test_resolver_preview_guard_accepts_successful_prospective_resolution() -> None: + architecture = { + "findings": [], + "safe_to_implement": True, + "manual_design_required": False, + } + guarded = with_resolver_preview_guard( + architecture, + { + "packages": [ + { + "name": "openai-codex", + "locked_version": "0.144.4", + "candidate_version": "0.147.0", + } + ], + "refresh_preview": {"returncode": 0}, + }, + ) + + assert guarded == architecture + + @pytest.mark.parametrize( "api_diff", [ @@ -2494,6 +3009,70 @@ def test_schema_validation_rejects_missing_required_field() -> None: validate_mapping({"packages": [], "themes": []}, DIRECTION_ANALYSIS_SCHEMA, name="stage") +def test_direction_guard_replaces_upgrade_advice_with_implementation_evidence() -> None: + guarded = with_implementation_direction_guard( + { + "packages": [ + { + "name": "claude-agent-sdk", + "evidence_status": "observed", + "implementation_trend": "Hold the current baseline; no upgrade is needed.", + "observed_transitions": [], + "evidence": ["The lockfile is current."], + } + ], + "themes": [ + { + "name": "No Upgrade Action", + "implementation_pattern": "Keep the lockfile unchanged.", + "packages": ["claude-agent-sdk"], + "evidence": ["Resolver found no changes."], + }, + { + "name": "Adapter Contracts Stable Across Candidate Check", + "implementation_pattern": "The adapter contract passed for the candidate.", + "packages": ["claude-agent-sdk"], + "evidence": ["Behavior probes passed."], + } + ], + "uncertainty": [], + }, + evidence={"packages": [{"name": "claude-agent-sdk"}]}, + implementation_diffs=[ + { + "package": "claude-agent-sdk", + "from_version": "0.2.145", + "to_version": "0.2.146", + "status": "observed", + "source_lines_before": 100, + "source_lines_after": 104, + "files_added": [], + "files_removed": [], + "files_changed": ["claude_agent_sdk/_version.py"], + "definitions_added": [], + "definitions_removed": [], + "definitions_changed": [], + "opaque_artifacts_added": [], + "opaque_artifacts_removed": [], + "opaque_artifacts_changed": ["claude_agent_sdk/_bundled/claude"], + } + ], + ) + + package = guarded["packages"][0] + assert package["evidence_status"] == "observed" + assert package["observed_transitions"] == ["0.2.145 -> 0.2.146"] + assert "hold" not in package["implementation_trend"].lower() + assert "upgrade" not in package["implementation_trend"].lower() + assert package["evidence"][0] == ( + "0.2.145 -> 0.2.146: Python files +0/-0/~1, definitions +0/-0/~0, " + "opaque artifacts +0/-0/~1, source lines 100 -> 104." + ) + assert "claude_agent_sdk/_version.py" in package["evidence"][1] + assert guarded["themes"][0]["name"] == "Inspected implementation evolution" + assert any("replaced" in item for item in guarded["uncertainty"]) + + @pytest.mark.asyncio async def test_stage_execution_uses_agent_task_runtime_primitives(tmp_path: Path) -> None: runtime = RecordingRuntime() @@ -2523,6 +3102,8 @@ async def test_stage_execution_uses_agent_task_runtime_primitives(tmp_path: Path assert runtime.task.working_directory == tmp_path assert runtime.task.permissions.filesystem is FilesystemAccess.READ_ONLY assert runtime.task.metadata["stage"] == "direction-analysis" + assert "upstream implementation trajectory" in runtime.task.system + assert "Never recommend upgrading" in runtime.task.system assert "model" not in runtime.task.metadata assert "reasoning_effort" not in runtime.task.metadata @@ -2668,6 +3249,248 @@ def test_reviewer_approved_status_allows_implementation() -> None: assert gate.allowed is True +@pytest.mark.parametrize( + ("implementation", "expected"), + [ + ({"allowed": False, "applied": False, "verification_results": []}, False), + ({"allowed": True, "applied": False, "verification_results": []}, False), + ( + { + "allowed": True, + "applied": True, + "changed_paths": [], + "verification_results": [{"returncode": 0}], + }, + False, + ), + ( + { + "allowed": True, + "applied": True, + "changed_paths": ["uv.lock"], + "verification_results": [{"returncode": 1}], + }, + False, + ), + ( + { + "allowed": True, + "applied": True, + "changed_paths": ["pyproject.toml", "uv.lock"], + "verification_results": [{"returncode": 0}], + }, + True, + ), + ], +) +def test_draft_pr_requires_applied_verified_nonempty_change( + implementation: dict[str, Any], + expected: bool, +) -> None: + assert _should_create_pr(True, implementation) is expected + assert _should_create_pr(False, implementation) is False + + +def test_local_sdk_update_rolls_back_when_lock_misses_inspected_candidate( + tmp_path: Path, +) -> None: + pyproject = tmp_path / "pyproject.toml" + lockfile = tmp_path / "uv.lock" + pyproject.write_text( + '[project.optional-dependencies]\ncodex = ["openai-codex>=0.1.0b3,<0.145"]\n', + encoding="utf-8", + ) + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.144.4"\n', + encoding="utf-8", + ) + compatibility = _write_compatibility_fixture(tmp_path) + original_pyproject = pyproject.read_bytes() + original_lock = lockfile.read_bytes() + original_compatibility = compatibility.read_bytes() + + def runner( + command: tuple[str, ...], + *, + cwd: Path | None = None, + env: dict[str, str] | None = None, + ) -> CommandResult: + del env + assert cwd == tmp_path + if command[:3] == ("uv", "lock", "--exclude-newer"): + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.146.0"\n', + encoding="utf-8", + ) + return CommandResult(command=command, returncode=0) + + result = _run_local_sdk_update( + RunOptions(workspace=tmp_path, runtime="fake"), + update_versions={"openai-codex": "0.147.0"}, + implementation={"allowed": True, "changes": [], "verification_results": []}, + command_runner=runner, + ) + + assert result["applied"] is False + assert result["changed_paths"] == [] + assert "expected 0.147.0, resolved 0.146.0" in result["blocked_reason"] + assert pyproject.read_bytes() == original_pyproject + assert lockfile.read_bytes() == original_lock + assert compatibility.read_bytes() == original_compatibility + + +def test_local_sdk_update_widens_bound_and_verifies_exact_candidate( + tmp_path: Path, +) -> None: + pyproject = tmp_path / "pyproject.toml" + lockfile = tmp_path / "uv.lock" + pyproject.write_text( + '[project.optional-dependencies]\ncodex = ["openai-codex>=0.1.0b3,<0.145"]\n', + encoding="utf-8", + ) + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.144.4"\n' + '[[package]]\nname = "openai-codex-cli-bin"\nversion = "0.144.4"\n', + encoding="utf-8", + ) + compatibility = _write_compatibility_fixture(tmp_path) + commands: list[tuple[str, ...]] = [] + + def runner( + command: tuple[str, ...], + *, + cwd: Path | None = None, + env: dict[str, str] | None = None, + ) -> CommandResult: + del env + assert cwd == tmp_path + commands.append(command) + if command[:3] == ("uv", "lock", "--exclude-newer"): + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.147.0"\n' + '[[package]]\nname = "openai-codex-cli-bin"\nversion = "0.147.0"\n', + encoding="utf-8", + ) + return CommandResult(command=command, returncode=0) + + result = _run_local_sdk_update( + RunOptions(workspace=tmp_path, runtime="fake"), + update_versions={ + "openai-codex": "0.147.0", + "openai-codex-cli-bin": "0.147.0", + }, + implementation={"allowed": True, "changes": [], "verification_results": []}, + command_runner=runner, + ) + + assert result["applied"] is True + assert result["changed_paths"] == [ + "pyproject.toml", + "uv.lock", + "src/agent_runtime_kit/compatibility.py", + ] + assert "openai-codex>=0.1.0b3,<0.148" in pyproject.read_text() + compatibility_text = compatibility.read_text(encoding="utf-8") + assert 'version_specifier=">=0.1.0b3,<0.148"' in compatibility_text + assert 'tested_version="0.147.0"' in compatibility_text + assert 'PackageVersion(package="openai-codex-cli-bin", version="0.147.0")' in ( + compatibility_text + ) + assert any(command[:3] == ("uv", "run", "--locked") for command in commands) + assert _verification_passed(result) is True + + +def test_local_sdk_update_rolls_back_manifest_after_verification_failure( + tmp_path: Path, +) -> None: + pyproject = tmp_path / "pyproject.toml" + lockfile = tmp_path / "uv.lock" + pyproject.write_text( + '[project.optional-dependencies]\ncodex = ["openai-codex>=0.1.0b3,<0.145"]\n', + encoding="utf-8", + ) + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.144.4"\n', + encoding="utf-8", + ) + compatibility = _write_compatibility_fixture(tmp_path) + originals = { + pyproject: pyproject.read_bytes(), + lockfile: lockfile.read_bytes(), + compatibility: compatibility.read_bytes(), + } + + def runner( + command: tuple[str, ...], + *, + cwd: Path | None = None, + env: dict[str, str] | None = None, + ) -> CommandResult: + del env + assert cwd == tmp_path + if command[:3] == ("uv", "lock", "--exclude-newer"): + lockfile.write_text( + '[[package]]\nname = "openai-codex"\nversion = "0.147.0"\n', + encoding="utf-8", + ) + return CommandResult( + command=command, + returncode=1 if command == ("uv", "run", "--locked", "pytest") else 0, + stdout="one failed" if command[-1:] == ("pytest",) else "", + ) + + result = _run_local_sdk_update( + RunOptions(workspace=tmp_path, runtime="fake"), + update_versions={"openai-codex": "0.147.0"}, + implementation={"allowed": True, "changes": [], "verification_results": []}, + command_runner=runner, + ) + + assert result["applied"] is False + assert result["changed_paths"] == [] + assert result["blocked_reason"] == ( + "verification failed: uv run --locked pytest (exit 1); " + "inspect verification_results for full output" + ) + for path, original in originals.items(): + assert path.read_bytes() == original + + +def _write_compatibility_fixture( + root: Path, + *, + package: str = "openai-codex", + specifier: str = ">=0.1.0b3,<0.145", + tested_version: str = "0.144.4", + runtime_dependency: tuple[str, str] | None = ( + "openai-codex-cli-bin", + "0.144.4", + ), +) -> Path: + path = root / "src" / "agent_runtime_kit" / "compatibility.py" + path.parent.mkdir(parents=True) + lines = [ + "COMPATIBILITY_MANIFEST = (", + " RuntimeCompatibility(", + f' package="{package}",', + f' version_specifier="{specifier}",', + f' tested_version="{tested_version}",', + ] + if runtime_dependency is not None: + dependency_package, dependency_version = runtime_dependency + lines.extend( + ( + " tested_runtime_dependencies=(", + " PackageVersion(" + f'package="{dependency_package}", version="{dependency_version}"),', + " ),", + ) + ) + lines.extend((" ),", ")", "")) + path.write_text("\n".join(lines), encoding="utf-8") + return path + + @pytest.mark.asyncio async def test_run_agent_report_only_generates_artifacts( tmp_path: Path, @@ -2716,6 +3539,8 @@ async def test_run_agent_report_only_generates_artifacts( assert (report_path.parent / "evidence.json").exists() assert (report_path.parent / "release_notes.json").exists() assert (report_path.parent / "api_diffs.json").exists() + assert (report_path.parent / "implementation_diffs.json").exists() + assert (report_path.parent / "implementation_snapshots").is_dir() assert (report_path.parent / "behavior_probes.json").exists() assert (report_path.parent / "behavior_diffs.json").exists() assert (report_path.parent / "behavior_summary.json").exists() @@ -2729,6 +3554,14 @@ async def test_run_agent_report_only_generates_artifacts( encoding="utf-8" ) assert "Recursive self-adaptation impact" in report_path.read_text(encoding="utf-8") + assert "## Upstream Implementation Trends" in report_path.read_text(encoding="utf-8") + assert "## Direction Of Travel" not in report_path.read_text(encoding="utf-8") + current_state = json.loads( + (report_path.parent / "current_state.json").read_text(encoding="utf-8") + ) + assert current_state["artifacts"]["report.md"]["sha256"] == hashlib.sha256( + report_path.read_bytes() + ).hexdigest() @pytest.mark.asyncio @@ -2797,7 +3630,7 @@ def runner( cwd: Path | None = None, env: dict[str, str] | None = None, ) -> CommandResult: - del cwd, env + del env commands.append(command) # check-ignore exits 0 => path IS ignored (the default report dir). return CommandResult(command=command, returncode=0) @@ -2829,7 +3662,7 @@ def runner( cwd: Path | None = None, env: dict[str, str] | None = None, ) -> CommandResult: - del cwd, env + del env commands.append(command) # check-ignore exits 128 when it cannot judge the path (outside the repo). if command[:2] == ("git", "check-ignore"): @@ -2873,6 +3706,13 @@ async def test_run_agent_autonomous_pr_path( """, encoding="utf-8", ) + _write_compatibility_fixture( + tmp_path, + package="claude-agent-sdk", + specifier=">=0.2", + tested_version="0.2.1", + runtime_dependency=None, + ) monkeypatch.setattr( "examples.sdk_evolution_agent.cli.snapshot_current_api", lambda package, *, version=None: ApiSnapshot( @@ -2927,7 +3767,7 @@ def runner( cwd: Path | None = None, env: dict[str, str] | None = None, ) -> CommandResult: - del cwd, env + del env commands.append(command) if command[:3] == ("uv", "lock", "--dry-run"): return CommandResult( @@ -2936,6 +3776,11 @@ def runner( stderr="Update claude-agent-sdk v0.2.1 -> v0.3.0\n", ) if command[:2] == ("uv", "lock"): + assert cwd is not None + (cwd / "uv.lock").write_text( + '\n[[package]]\nname = "claude-agent-sdk"\nversion = "0.3.0"\n', + encoding="utf-8", + ) return CommandResult(command=command, returncode=0, stdout="updated") if command[:2] == ("git", "check-ignore"): # Report dir is tracked in this scenario -> not ignored (exit 1). @@ -2966,10 +3811,10 @@ def runner( assert ( "uv", "lock", + "--exclude-newer", + "false", "-P", "claude-agent-sdk", - "--exclude-newer-package", - "claude-agent-sdk=false", ) in commands assert any(command[:3] == ("git", "commit", "-m") for command in commands) assert any(command[:4] == ("gh", "pr", "create", "--draft") for command in commands) @@ -3290,7 +4135,18 @@ async def run(self, task: AgentTask) -> AgentResult: self.task = task stage = task.metadata["stage"] if stage == "direction-analysis": - payload = {"packages": [], "themes": [], "uncertainty": []} + payload = { + "packages": [], + "themes": [ + { + "name": "Implementation evidence", + "implementation_pattern": "Interpret exact implementation diffs.", + "packages": [], + "evidence": [], + } + ], + "uncertainty": [], + } elif stage == "architecture-decision": payload = { "findings": [], @@ -3394,8 +4250,7 @@ def _sdk_evidence( evidence["refresh_preview"] = { "stdout": "", "stderr": ( - f"Update {package} v{locked_version or installed_version} " - f"-> v{candidate_version}\n" + f"Update {package} v{locked_version or installed_version} -> v{candidate_version}\n" ), } return evidence diff --git a/tests/test_sdk_evolution_docs.py b/tests/test_sdk_evolution_docs.py index e3c7ab8..f05d3f6 100644 --- a/tests/test_sdk_evolution_docs.py +++ b/tests/test_sdk_evolution_docs.py @@ -10,8 +10,8 @@ CLAUDE_RUNBOOK = Path(".claude/commands/agent-runtime-kit/upgrade.md") PUBLIC_GUIDE = Path("docs/sdk-evolution-agent.md") DESIGN_GUIDE = Path("docs/sdk-evolution-agent-design.md") -AUTHORITATIVE_DOCS = (CODEX_RUNBOOK, CLAUDE_RUNBOOK, PUBLIC_GUIDE, DESIGN_GUIDE) -RUNBOOKS = (CODEX_RUNBOOK, CLAUDE_RUNBOOK) +AUTHORITATIVE_DOCS = (CODEX_RUNBOOK, PUBLIC_GUIDE, DESIGN_GUIDE) +RUNBOOKS = (CODEX_RUNBOOK,) PUBLIC_DOCS = (PUBLIC_GUIDE, DESIGN_GUIDE) RUNTIME_EXTRAS = { "claude-agent-sdk": "claude", @@ -51,10 +51,7 @@ def test_public_guides_cover_every_runtime_extra_mapping() -> None: def test_runbooks_include_report_and_implementation_commands() -> None: - defaults = { - CODEX_RUNBOOK: "codex-agent-sdk", - CLAUDE_RUNBOOK: "claude-agent-sdk", - } + defaults = {CODEX_RUNBOOK: "codex-agent-sdk"} for path, runtime in defaults.items(): commands = [ command @@ -99,13 +96,21 @@ def test_behavior_summary_is_in_artifacts_and_runbook_handoffs() -> None: for path in RUNBOOKS: text = _read(path) normalized = " ".join(text.split()) - assert "`behavior_summary.json` status and reasons" in text - assert "a missing or drifted locked baseline" in normalized + assert "behavior_summary.json" in text assert "credential-scrubbed" in text - for blocker in ("missing", "malformed", "unknown status", "`fail`", "`incomplete`"): + assert "exact" in normalized + for blocker in ("missing", "failed", "`fail`", "`incomplete`"): assert blocker in text +def test_claude_command_delegates_to_the_canonical_workflow() -> None: + text = _read(CLAUDE_RUNBOOK) + + assert ".codex/skills/agent-runtime-kit-upgrade/SKILL.md" in text + assert "`claude-agent-sdk` with `--extra claude`" in text + assert "publication" in text + + def test_codex_auth_commands_use_the_locked_codex_extra() -> None: for path in AUTHORITATIVE_DOCS: for command in _bash_commands(path): diff --git a/uv.lock b/uv.lock index 9916bfe..9b5ed09 100644 --- a/uv.lock +++ b/uv.lock @@ -11,16 +11,6 @@ required-markers = [ "platform_machine == 'x86_64' and sys_platform == 'linux'", ] -[options] -exclude-newer = "2026-07-20T15:29:02.807093Z" -exclude-newer-span = "P8D" - -[options.exclude-newer-package] -openai-codex-cli-bin = false -google-antigravity = false -openai-codex = false -claude-agent-sdk = false - [manifest] constraints = [{ name = "openai-codex-cli-bin", specifier = ">=0.134.0a1" }] @@ -35,7 +25,7 @@ wheels = [ [[package]] name = "agent-runtime-kit" -version = "0.5.0" +version = "0.5.1" source = { editable = "." } dependencies = [ { name = "jsonschema" }, @@ -61,6 +51,7 @@ codex = [ dev = [ { name = "build" }, { name = "mypy" }, + { name = "packaging" }, { name = "pydantic" }, { name = "pytest" }, { name = "pytest-asyncio" }, @@ -77,8 +68,8 @@ requires-dist = [ { name = "google-antigravity", marker = "extra == 'all'", specifier = ">=0.1.2,<0.2" }, { name = "google-antigravity", marker = "extra == 'antigravity'", specifier = ">=0.1.2,<0.2" }, { name = "jsonschema", specifier = ">=4.18,<5" }, - { name = "openai-codex", marker = "extra == 'all'", specifier = ">=0.1.0b3,<0.145" }, - { name = "openai-codex", marker = "extra == 'codex'", specifier = ">=0.1.0b3,<0.145" }, + { name = "openai-codex", marker = "extra == 'all'", specifier = ">=0.1.0b3,<0.148" }, + { name = "openai-codex", marker = "extra == 'codex'", specifier = ">=0.1.0b3,<0.148" }, ] provides-extras = ["claude", "codex", "antigravity", "all"] @@ -86,6 +77,7 @@ provides-extras = ["claude", "codex", "antigravity", "all"] dev = [ { name = "build", specifier = ">=1.2" }, { name = "mypy", specifier = ">=1.8" }, + { name = "packaging", specifier = ">=24,<27" }, { name = "pydantic", specifier = ">=2,<3" }, { name = "pytest", specifier = ">=8.0" }, { name = "pytest-asyncio", specifier = ">=0.23" }, @@ -390,21 +382,22 @@ wheels = [ [[package]] name = "claude-agent-sdk" -version = "0.2.128" +version = "0.2.148" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "anyio" }, + { name = "jsonschema" }, { name = "mcp" }, { name = "sniffio" }, { name = "typing-extensions", marker = "python_full_version < '3.11'" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/a7/e8/3a9622b31f9ee22274e13a620e5e75ac38454d14538391b2fe3bc7eb76dc/claude_agent_sdk-0.2.128.tar.gz", hash = "sha256:2ac7b2b3bc56ae9037fd284c8690d3dafab9493ecd28d8974bba79a418e1b800", size = 309369, upload-time = "2026-07-25T01:48:25.884Z" } +sdist = { url = "https://files.pythonhosted.org/packages/7e/f4/fb7f81b31e44f69af9f53b93c32e2047f3ca15cd6311cfe72bfdf420f1f5/claude_agent_sdk-0.2.148.tar.gz", hash = "sha256:45c9972fa72dc6006745239c6838cc3cea7a9b42d6927c89f7bdb6996a1983b0", size = 344717, upload-time = "2026-08-28T18:33:51.812Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/45/ff/03613a38a84285cd85f114fc26d4431f545d2b3b344eb8f69f9b0f18f3c6/claude_agent_sdk-0.2.128-py3-none-macosx_11_0_arm64.whl", hash = "sha256:2e47ee95be68cb07612fd5288a40f3307763da5a5adf66d6e03ea49dc0495c9c", size = 75183825, upload-time = "2026-07-25T01:48:30.45Z" }, - { url = "https://files.pythonhosted.org/packages/a3/df/2adbd3077f1a1cada39cd1508990d4008a8ac9ec74c801b24b1b302348b0/claude_agent_sdk-0.2.128-py3-none-macosx_11_0_x86_64.whl", hash = "sha256:55e8fa28918af5620f391692efe80527ad9f2294cec14dadeea93c5161dbefc4", size = 80305667, upload-time = "2026-07-25T01:48:35.091Z" }, - { url = "https://files.pythonhosted.org/packages/cc/56/776a41af67a53d794b420dcd2f1075b585ed3725faa9827164f4a2e945dd/claude_agent_sdk-0.2.128-py3-none-manylinux_2_17_aarch64.whl", hash = "sha256:6c10cc1c5403b2b0b3d8deb79ad5271a8e49b9db6460193599b91f49f16b4cf7", size = 84617823, upload-time = "2026-07-25T01:48:39.297Z" }, - { url = "https://files.pythonhosted.org/packages/17/1c/37044abbddf2141c4eeb6cc8e15a486491549a27283029d73db2c4f3c3b7/claude_agent_sdk-0.2.128-py3-none-manylinux_2_17_x86_64.whl", hash = "sha256:cc4a0f20337e227fe16f00edf7cbe109349707adb440fe51ca6c44f9f8d7d21a", size = 85663462, upload-time = "2026-07-25T01:48:43.982Z" }, - { url = "https://files.pythonhosted.org/packages/a3/c9/ffbd8113080f87a4cd7bbfc8b00a974ea7189e3f666b5d7a2903dec7b74f/claude_agent_sdk-0.2.128-py3-none-win_amd64.whl", hash = "sha256:37b87c8e75daa2a6dc74da6d788e0981a2e5e3a6be80cfa4c287abce2bf70e0b", size = 85566882, upload-time = "2026-07-25T01:48:48.635Z" }, + { url = "https://files.pythonhosted.org/packages/b8/f5/f4affb9f882bcf61422ea48491effb6ee2de28bfbdafd0ec90d1b959dad2/claude_agent_sdk-0.2.148-py3-none-macosx_11_0_arm64.whl", hash = "sha256:026e2b54bb3f3909712ff462b8469e2b1fc5471cfde796abc23f32d30b2c8ba1", size = 83882833, upload-time = "2026-08-28T18:33:56.238Z" }, + { url = "https://files.pythonhosted.org/packages/af/cf/e1b11dd9fac877c8b44f022cf895f54f23854b916e09b32360e71209904c/claude_agent_sdk-0.2.148-py3-none-macosx_11_0_x86_64.whl", hash = "sha256:c8cca1f4e4eb4e0c9f44ccfca6f651cf35e8f3914e483632ae8867c5619818a0", size = 88314490, upload-time = "2026-08-28T18:34:01.284Z" }, + { url = "https://files.pythonhosted.org/packages/34/d4/53de4727576fa6b5a261414d420895f4eb5f0369e232d403f1096205709e/claude_agent_sdk-0.2.148-py3-none-manylinux_2_17_aarch64.whl", hash = "sha256:3fbbea7cc30099d14f894e12fef68b0ee42fa9982567df94e59499b6aff55959", size = 94014630, upload-time = "2026-08-28T18:34:06.238Z" }, + { url = "https://files.pythonhosted.org/packages/d8/e9/2ea469751ed4df75a911c4bf4dcad1ff5b71ccc9aa1a4aa108583a2a66b5/claude_agent_sdk-0.2.148-py3-none-manylinux_2_17_x86_64.whl", hash = "sha256:221efbc0f583c80258189ab93e4f0d6b6a61ad61335de1fdf3e2331ad423dbeb", size = 94210928, upload-time = "2026-08-28T18:34:11.592Z" }, + { url = "https://files.pythonhosted.org/packages/c3/ea/86ccd7a34446c37174e53f03c389b25f168c7f4fd4f00f3e130baae9c902/claude_agent_sdk-0.2.148-py3-none-win_amd64.whl", hash = "sha256:f155252940fd77f1bcc707836924ca3a9691d19bbc464914ae2b7ccf17f0410f", size = 96875177, upload-time = "2026-08-28T18:34:16.378Z" }, ] [[package]] @@ -602,7 +595,7 @@ name = "exceptiongroup" version = "1.3.1" source = { registry = "https://pypi.org/simple" } dependencies = [ - { name = "typing-extensions", marker = "python_full_version < '3.11'" }, + { name = "typing-extensions" }, ] sdist = { url = "https://files.pythonhosted.org/packages/50/79/66800aadf48771f6b62f7eb014e352e5d06856655206165d775e675a02c9/exceptiongroup-1.3.1.tar.gz", hash = "sha256:8b412432c6055b0b7d14c310000ae93352ed6754f70fa8f7c34141f91c4e3219", size = 30371, upload-time = "2025-11-21T23:01:54.787Z" } wheels = [ @@ -611,7 +604,7 @@ wheels = [ [[package]] name = "google-antigravity" -version = "0.1.8" +version = "0.1.15" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "absl-py" }, @@ -623,11 +616,11 @@ dependencies = [ { name = "websockets" }, ] wheels = [ - { url = "https://files.pythonhosted.org/packages/0b/5f/0f5d6b210faddb4e6d2b0c25072ae0b2430294f5dec0b4eabecc196ea32b/google_antigravity-0.1.8-py3-none-macosx_11_0_arm64.whl", hash = "sha256:be5a4853ae7cb7bff4fbba44073cbed96e1fdcdeebd9602b324bf3abe5614ab2", size = 31829862, upload-time = "2026-07-23T19:20:02.908Z" }, - { url = "https://files.pythonhosted.org/packages/a7/cb/2d74bfc57e9f7f579a97f8b445599639ac4a101bcf44bab2a9c58c1f634f/google_antigravity-0.1.8-py3-none-manylinux_2_17_aarch64.whl", hash = "sha256:06773b250668993cf59e3a90ba00f1d6fc52483612d9463aa3e71e4a2a30a20b", size = 37289461, upload-time = "2026-07-23T19:20:06.571Z" }, - { url = "https://files.pythonhosted.org/packages/93/54/9972ecf8e0b5e3d12067550462078be9babf8dc0f686e80ffe70daef0cad/google_antigravity-0.1.8-py3-none-manylinux_2_17_x86_64.whl", hash = "sha256:bd912cc0ac74fef8026c227500b7de06b3749a78f0148b99d9033e8c952787e0", size = 39396935, upload-time = "2026-07-23T19:20:09.493Z" }, - { url = "https://files.pythonhosted.org/packages/bb/ee/52f63973142aa00b722d7d841b60991798352564538f8533c9eb27bb9546/google_antigravity-0.1.8-py3-none-win_amd64.whl", hash = "sha256:edc1690c3116d1da31b033951d62a56a005f73d13486ced1d11564c0b98ca244", size = 36678881, upload-time = "2026-07-23T19:20:13.11Z" }, - { url = "https://files.pythonhosted.org/packages/a1/f6/503b6efb48903cfca252fb22077c15849da302e3d3f4d7a7b7cd2438e146/google_antigravity-0.1.8-py3-none-win_arm64.whl", hash = "sha256:999bdc9ed8d413a6abff96b15105e72b7c596fcd7fce41448f6613ee0947a8b4", size = 33138761, upload-time = "2026-07-23T19:20:16.811Z" }, + { url = "https://files.pythonhosted.org/packages/c9/95/142eee8db8be5a553f9fd18f762336f585138f5fe8ce85cb5892c7ec8ee0/google_antigravity-0.1.15-py3-none-macosx_11_0_arm64.whl", hash = "sha256:3e281ab5ee2017e6c57ffd33053ed51300fa9bbc08f81e728d5d53a1243217ac", size = 32916793, upload-time = "2026-08-26T19:08:08.899Z" }, + { url = "https://files.pythonhosted.org/packages/7c/8d/f844f97b5f0cd4b432491032be93f9a90a6a4c14387df03c107d119eee6a/google_antigravity-0.1.15-py3-none-manylinux_2_17_aarch64.musllinux_1_1_aarch64.whl", hash = "sha256:e7c3c39df2ab3bbe5a2fdfa7d2ee7c8958bd6bb40a6b97aa43ab18c5d36e717c", size = 35736612, upload-time = "2026-08-26T19:08:16.85Z" }, + { url = "https://files.pythonhosted.org/packages/29/2c/3837d346fb5caa454db32eb785934d37ce7313cc319b90edca9866a279d3/google_antigravity-0.1.15-py3-none-manylinux_2_17_x86_64.musllinux_1_1_x86_64.whl", hash = "sha256:ee28bc56c895f21623b8bb7b147f91ae791fc5c2b03b76bcea62a998f8ed068d", size = 39623445, upload-time = "2026-08-26T19:08:21.005Z" }, + { url = "https://files.pythonhosted.org/packages/40/4e/b998ebd8b565a15b7d9193b33f385191ae72f372272fc835cb9875a0813e/google_antigravity-0.1.15-py3-none-win_amd64.whl", hash = "sha256:a51106034129825f87a1eb0d7ef78a2efa9605ba289fa4850f242aa037767b20", size = 38008309, upload-time = "2026-08-26T19:08:25.706Z" }, + { url = "https://files.pythonhosted.org/packages/8e/5d/135860fc10ea92f612889f10ab13b6e1c563573442edb678651150f5447c/google_antigravity-0.1.15-py3-none-win_arm64.whl", hash = "sha256:47dcf6282d4cf2b4f0a36575598adb6e9ee50fd7ca6ed40f7034d371152e8732", size = 34293602, upload-time = "2026-08-26T19:08:30.278Z" }, ] [[package]] @@ -729,7 +722,7 @@ name = "importlib-metadata" version = "9.0.0" source = { registry = "https://pypi.org/simple" } dependencies = [ - { name = "zipp", marker = "python_full_version < '3.11'" }, + { name = "zipp" }, ] sdist = { url = "https://files.pythonhosted.org/packages/a9/01/15bb152d77b21318514a96f43af312635eb2500c96b55398d020c93d86ea/importlib_metadata-9.0.0.tar.gz", hash = "sha256:a4f57ab599e6a2e3016d7595cfd72eb4661a5106e787a95bcc90c7105b831efc", size = 56405, upload-time = "2026-03-20T06:42:56.999Z" } wheels = [ @@ -953,30 +946,30 @@ wheels = [ [[package]] name = "openai-codex" -version = "0.144.4" +version = "0.147.0" source = { registry = "https://pypi.org/simple" } dependencies = [ { name = "openai-codex-cli-bin" }, { name = "pydantic" }, ] -sdist = { url = "https://files.pythonhosted.org/packages/b9/7d/4b999b0ddd05c22d83cc2e879c95e1e92e3b7808b58ac561c4ea91fca1c4/openai_codex-0.144.4.tar.gz", hash = "sha256:91c63a7cb213441569f130e593386b34657ab9e726ae88af255f0ecb8de08ea5", size = 68324, upload-time = "2026-07-17T23:42:17.013Z" } +sdist = { url = "https://files.pythonhosted.org/packages/2b/ac/beb56a8a5a37b90e40e418a29449f6c345f45cc326d8b021ca23aae12e8d/openai_codex-0.147.0.tar.gz", hash = "sha256:e0ab4ea3ac44585a98a80df11e220519fec4641d508d830d2c1d656ce471e2ed", size = 73470, upload-time = "2026-08-18T06:13:11.314Z" } wheels = [ - { url = "https://files.pythonhosted.org/packages/17/35/5d2a13e38d91278019f18757ac10426d2649b7ec6f031818c792f64559e9/openai_codex-0.144.4-py3-none-any.whl", hash = "sha256:de1513a6e94b9a8d7728a3b74298bc1469428ade10ba0ef2d5db47dd1cb606f5", size = 76244, upload-time = "2026-07-17T23:42:15.658Z" }, + { url = "https://files.pythonhosted.org/packages/3f/14/1a36ddc96152160768793689cd0924ffa78ce10e1bbb817fb4b006a96479/openai_codex-0.147.0-py3-none-any.whl", hash = "sha256:ab2e0b3a41dba5a62be8561397cf3e7913afb53b5372ad881002a6f0b77e6a0a", size = 81344, upload-time = "2026-08-18T06:13:09.897Z" }, ] [[package]] name = "openai-codex-cli-bin" -version = "0.144.4" +version = "0.147.0" source = { registry = "https://pypi.org/simple" } wheels = [ - { url = "https://files.pythonhosted.org/packages/79/30/7e457c007a32aa7333a78438a9f532504c3b77a1ea57e7808855712c2c0f/openai_codex_cli_bin-0.144.4-py3-none-macosx_10_9_x86_64.whl", hash = "sha256:4d587d152d5f0aa25f21ded4d08aac5150da48639b29fc3e57b2a8b413d06903", size = 126851183, upload-time = "2026-07-15T00:14:00.171Z" }, - { url = "https://files.pythonhosted.org/packages/65/eb/64c180514a2cc3e2500e486813f5a8d7f7e349342e9bebd78d99ddd9791a/openai_codex_cli_bin-0.144.4-py3-none-macosx_11_0_arm64.whl", hash = "sha256:05db505a9c7f020f58b70837a94e00d32a50086986c267bcc44ea97b573d4a05", size = 116474758, upload-time = "2026-07-15T00:14:08.177Z" }, - { url = "https://files.pythonhosted.org/packages/85/34/921ab692c7ed140941a91e39b84c6deafddb45072793f3a5f6dcbf86f59a/openai_codex_cli_bin-0.144.4-py3-none-manylinux_2_17_aarch64.whl", hash = "sha256:cd4bb31b8a477a3adba22139c8c493693fa5292ba21e86eef08ace3e96c41294", size = 119437596, upload-time = "2026-07-15T00:14:17.418Z" }, - { url = "https://files.pythonhosted.org/packages/25/62/39e630cf8b7b2e2444a5a6669235a6cd88004869a25bb7f3f629ed35638c/openai_codex_cli_bin-0.144.4-py3-none-manylinux_2_17_x86_64.whl", hash = "sha256:4106229c38f37245c3eea8d904426afd45b8bf19831765715647f5402f757c06", size = 128817757, upload-time = "2026-07-15T00:14:30.539Z" }, - { url = "https://files.pythonhosted.org/packages/23/24/318f91a95baaff845307e529d138d42a7202e6bd0e626556058eab3dead5/openai_codex_cli_bin-0.144.4-py3-none-musllinux_1_1_aarch64.whl", hash = "sha256:d2d3fada11731938e3d3e3660819d5324b7f92e825a0ed5aede63100b36ffc9a", size = 119437594, upload-time = "2026-07-15T00:14:39.327Z" }, - { url = "https://files.pythonhosted.org/packages/2e/a0/65e6fd3a6fba52639937801d9ce517b0c48026458723bb9f08eb07a1dd35/openai_codex_cli_bin-0.144.4-py3-none-musllinux_1_1_x86_64.whl", hash = "sha256:70fb62ed7755e332dd8b00ed48e22ee16df10f4a977335039e237cd330490df3", size = 128817755, upload-time = "2026-07-15T00:14:47.849Z" }, - { url = "https://files.pythonhosted.org/packages/e3/8a/3ab5fc97352e8f780d472c8dd6d7504d6a670cd7dede106d461ac1171999/openai_codex_cli_bin-0.144.4-py3-none-win_amd64.whl", hash = "sha256:56e142974467332f1f669b89f2636e06b3ada09413512abb2f24c33b4a4f59fc", size = 140595092, upload-time = "2026-07-15T00:14:58.484Z" }, - { url = "https://files.pythonhosted.org/packages/70/1a/3aa52ab8f89e0f596b6eea835e3f3494697da51917d3231f5f48f92e2bcb/openai_codex_cli_bin-0.144.4-py3-none-win_arm64.whl", hash = "sha256:2ad01058db7181323ae7c6217e560702e38c4f107595b99ab1d0abc9f184890a", size = 130665105, upload-time = "2026-07-15T00:15:08.861Z" }, + { url = "https://files.pythonhosted.org/packages/46/85/3302e5265f35941a24f614573e1a27e6f218d7fa71e4dd3b1353d1726c5a/openai_codex_cli_bin-0.147.0-py3-none-macosx_10_9_x86_64.whl", hash = "sha256:19c3a72a0eac6706bb5023088ab7eb31d8cfc06bf7860250b5c5b90f4505489b", size = 116948927, upload-time = "2026-08-18T06:05:53.479Z" }, + { url = "https://files.pythonhosted.org/packages/b8/0c/6474ef2f854ecf217550d06667a8ac988b5656c448b680e5b5ece5ae5152/openai_codex_cli_bin-0.147.0-py3-none-macosx_11_0_arm64.whl", hash = "sha256:b851943fffc48aa7c5c130b6a34be09964833d2785546eda96d749427c6e24f2", size = 107573930, upload-time = "2026-08-18T06:05:57.701Z" }, + { url = "https://files.pythonhosted.org/packages/59/4b/0efce457a301271301b9229b36b3f53ca94d5f4db1d605c86e0da9833565/openai_codex_cli_bin-0.147.0-py3-none-manylinux_2_17_aarch64.whl", hash = "sha256:aab3e27ce07cd7bc7a78708efa40375b28e89f586851c495ef55aa6cb6806d11", size = 111220520, upload-time = "2026-08-18T06:06:01.636Z" }, + { url = "https://files.pythonhosted.org/packages/f2/e4/5e4fb0f61ca90c2bb42a8823e08fe05957c54591d1e161c56a93d27d565d/openai_codex_cli_bin-0.147.0-py3-none-manylinux_2_17_x86_64.whl", hash = "sha256:cb3907d633cda87c4b68b47ff4979e6054c02a08c41b15aa2e9a689558a61074", size = 119921665, upload-time = "2026-08-18T06:06:05.666Z" }, + { url = "https://files.pythonhosted.org/packages/f9/83/a52e32fc63bde10996e38cb4e1175c402c578ee7c7ca79b96abee5219127/openai_codex_cli_bin-0.147.0-py3-none-musllinux_1_1_aarch64.whl", hash = "sha256:ea0bfdee98164fc7d589eea5623aff06c8f8f2ba3f9bc550297429db45ea68f7", size = 111220518, upload-time = "2026-08-18T06:06:09.538Z" }, + { url = "https://files.pythonhosted.org/packages/e9/b6/c159da0a1ec136fde6d2742ec05347e311f000dfe19024ed7d708cf60aa5/openai_codex_cli_bin-0.147.0-py3-none-musllinux_1_1_x86_64.whl", hash = "sha256:a9be23f46326494b7ebf4bd98cbe985542f5ad076355f3b0fb636a984df56816", size = 119921665, upload-time = "2026-08-18T06:06:15.061Z" }, + { url = "https://files.pythonhosted.org/packages/54/38/1bb47e08b9c523b673358e4f5597fc99699e855ec63bd02fdc157f1679c9/openai_codex_cli_bin-0.147.0-py3-none-win_amd64.whl", hash = "sha256:1a403ff803ae27e078189a0fd24687f32d43c46f152e50cdbe3eddc9e302f697", size = 127586208, upload-time = "2026-08-18T06:06:19.395Z" }, + { url = "https://files.pythonhosted.org/packages/d2/1e/a44a8da140ef080aebac6da6517e8fe6ac1aad49436c8b1f9b87e735b813/openai_codex_cli_bin-0.147.0-py3-none-win_arm64.whl", hash = "sha256:be0c8b9e34b067151d0964a24e9f4b8b48ea516448a08ef2636f357ebdc877b7", size = 117863410, upload-time = "2026-08-18T06:06:23.554Z" }, ] [[package]]