diff --git a/.github/PULL_REQUEST_TEMPLATE.md b/.github/PULL_REQUEST_TEMPLATE.md index 34c39220..89438761 100644 --- a/.github/PULL_REQUEST_TEMPLATE.md +++ b/.github/PULL_REQUEST_TEMPLATE.md @@ -18,7 +18,7 @@ ## Checklist -- [ ] I read [`CONTRIBUTING.md`](../CONTRIBUTING.md) and (for subsystem work) the local `CLAUDE.md`. +- [ ] I read [`CONTRIBUTING.md`](../CONTRIBUTING.md) and (for subsystem work) the local `AGENTS.md`. - [ ] `pre-commit run --all-files` passes. - [ ] `pytest tests/` passes; I added/updated tests for new behavior. - [ ] Bug fixes cite `file:line` in the description. diff --git a/.github/dependabot.yml b/.github/dependabot.yml index 3767128a..7aded83c 100644 --- a/.github/dependabot.yml +++ b/.github/dependabot.yml @@ -1,5 +1,5 @@ version: 2 -# All updates open against `dev` — the integration branch (see CLAUDE.md §6); +# All updates open against `dev` — the integration branch (see AGENTS.md §6); # `main` only receives promotion PRs. updates: - package-ecosystem: "pip" diff --git a/.planning/PROJECT.md b/.planning/PROJECT.md index 55c1a9d8..7c41cb17 100644 --- a/.planning/PROJECT.md +++ b/.planning/PROJECT.md @@ -166,6 +166,6 @@ This document evolves at phase transitions and milestone boundaries. This document and its siblings are *agent-maintained*. They succeed when: - A fresh AI agent dropped into the repo can answer *"what is jvagent for?"* by reading this file alone. -- A fresh AI agent dropped into a subsystem (`jvagent/core/`, `jvagent/action/orchestrator/`, etc.) can do correct local work by reading the local `CLAUDE.md` alone. +- A fresh AI agent dropped into a subsystem (`jvagent/core/`, `jvagent/action/orchestrator/`, etc.) can do correct local work by reading the local `AGENTS.md` alone. - Every claim about runtime behavior in [`SPEC.md`](SPEC.md) cites a file:line in the source. - Every load-bearing design decision is captured in an [`adr/`](adr/). diff --git a/.planning/README.md b/.planning/README.md index 4beac0b5..5253429a 100644 --- a/.planning/README.md +++ b/.planning/README.md @@ -1,7 +1,7 @@ # .planning — agent-facing design docs Reference material for **AI agents and human contributors** working on jvagent. -The repo-root [`CLAUDE.md`](../CLAUDE.md) is the agent entry point; this folder +The repo-root [`AGENTS.md`](../AGENTS.md) is the agent entry point; this folder holds the deeper normative specs, reference guides, runbooks, and decision records it links into. User-facing onboarding lives in the root [`README.md`](../README.md). diff --git a/.planning/adr/0033-identity-and-locking-substrate.md b/.planning/adr/0033-identity-and-locking-substrate.md index 004ad738..73cf20d9 100644 --- a/.planning/adr/0033-identity-and-locking-substrate.md +++ b/.planning/adr/0033-identity-and-locking-substrate.md @@ -2,7 +2,7 @@ - **Status:** Accepted — implemented 2026-07-17 (identity index #86, boot dedupe #87, contextvar hygiene #91, per-session conversation lock #99, turn-lock lease renewal #100, distributed bootstrap lease #101). Remaining: cross-worker upsert-by-identity for User/Conversation (in-process races locked; cross-process now rides the lease infra — small follow-up). - **Date:** 2026-07-16 -- **Supersedes / amends:** contracts in `jvagent/core/CLAUDE.md` §3, `jvagent/memory/CLAUDE.md` §3/§7 (the "compound index rejects on save" claim), and the singleton-enforcement narrative in `jvagent/action/actions.py`. +- **Supersedes / amends:** contracts in `jvagent/core/AGENTS.md` §3, `jvagent/memory/AGENTS.md` §3/§7 (the "compound index rejects on save" claim), and the singleton-enforcement narrative in `jvagent/action/actions.py`. - **Related:** review [`.planning/reviews/2026-07-16-core-review.md`](../reviews/2026-07-16-core-review.md); ADR-0003 (interaction pruning), ADR-0020 (public auth / session tokens). ## Context @@ -54,7 +54,7 @@ Run a lightweight identity-reconcile (dedupe by identity tuple, using raw record - **Positive:** the duplicate-singleton class (C1–C6) and the lost-update / double-claim class (C7–C9, H18) are closed at the substrate. Mongo deployments stop silently losing actions. The default `json` adapter gets a real single-writer guarantee via the bootstrap lease. - **Costs:** raw-record queries are slightly more code than `find_one`; a distributed lease adds a dependency for multi-replica correctness (optional, with a documented single-writer fallback). Boot-time reconcile adds bounded startup work. -- **Contract changes:** update `core/CLAUDE.md` and `memory/CLAUDE.md` to state that uniqueness is enforced by the application layer (upsert-by-identity + lease), and remove the false "compound index rejects on save" claim for the default adapter. +- **Contract changes:** update `core/AGENTS.md` and `memory/AGENTS.md` to state that uniqueness is enforced by the application layer (upsert-by-identity + lease), and remove the false "compound index rejects on save" claim for the default adapter. - **Test debt:** requires the currently-absent concurrency suite (duplicate create, lease expiry, contextvar reentrancy, double-claim) — treated as acceptance criteria, not follow-up. ## Alternatives considered diff --git a/.planning/archive/executive-build-prompt.md b/.planning/archive/executive-build-prompt.md index d813d664..feacda6f 100644 --- a/.planning/archive/executive-build-prompt.md +++ b/.planning/archive/executive-build-prompt.md @@ -9,7 +9,7 @@ > [`../docs/ORCHESTRATOR.md`](../../docs/ORCHESTRATOR.md). This file is retained only to > preserve the original design intent, mirroring how ADR-0010 itself is kept. > -> Paste this to Claude Code (or any coding agent) at the repo root of `jvagent`. It is self-contained but assumes the repo's `CLAUDE.md`, `.planning/SPEC.md`, and `.planning/adr/0010-executive-centers-architecture.md` are present. +> Paste this to Claude Code (or any coding agent) at the repo root of `jvagent`. It is self-contained but assumes the repo's `AGENTS.md`, `.planning/SPEC.md`, and `.planning/adr/0010-executive-centers-architecture.md` are present. --- @@ -17,7 +17,7 @@ Implement the **Executive + Centers** deployment pattern specified in `.planning/adr/0010-executive-centers-architecture.md`. It is a new, additive pattern that ships **alongside** the Rails pattern — no forced migration, no harness subsumption. ADR-0010 is the **source of truth**; if anything below conflicts with it, the ADR wins and you flag the conflict. -Read first, in order: `CLAUDE.md`, `.planning/adr/0010-executive-centers-architecture.md`, `.planning/SPEC.md` (§3.3–§3.4, §11). +Read first, in order: `AGENTS.md`, `.planning/adr/0010-executive-centers-architecture.md`, `.planning/SPEC.md` (§3.3–§3.4, §11). ## Mental model (so you make correct judgment calls) @@ -63,7 +63,7 @@ A brain. One **Executive** (prefrontal cortex, light model) engages trivial conv **M8 — Scaffolder profile + example.** `executive` profile in `jvagent/scaffold/builtin_profiles/`; `examples/jvagent_app/agents/jvagent/executive_agent/`. *Tests:* scaffold smoke; `jvagent ... validate` passes; a second orchestrator at `-200` rejected. *Exit:* `jvagent app create --profile executive` yields a working agent. -**M9 — Observability + docs + parity.** Activation-trace events on `Interaction`; `docs/ORCHESTRATOR.md`; `PATTERNS.md` + `GLOSSARY.md` + top-level `CLAUDE.md` updated; a smoke harness (`tests/action/executive/smoke_executive.py`) running the 6-utterance suite, archived under `baselines/`. *Exit:* a turn is fully traceable from one log query; smoke runs clean; pattern documented as a peer. +**M9 — Observability + docs + parity.** Activation-trace events on `Interaction`; `docs/ORCHESTRATOR.md`; `PATTERNS.md` + `GLOSSARY.md` + top-level `AGENTS.md` updated; a smoke harness (`tests/action/executive/smoke_executive.py`) running the 6-utterance suite, archived under `baselines/`. *Exit:* a turn is fully traceable from one log query; smoke runs clean; pattern documented as a peer. ## Testing & verification diff --git a/.planning/config.json b/.planning/config.json index 4464854a..2db242d8 100644 --- a/.planning/config.json +++ b/.planning/config.json @@ -52,5 +52,5 @@ }, "project_code": "jvagent", "agent_skills": {}, - "claude_md_path": "./CLAUDE.md" + "claude_md_path": "./AGENTS.md" } diff --git a/.planning/plans/2026-08-03-skill-only-tools.md b/.planning/plans/2026-08-03-skill-only-tools.md index 39ecc628..214edcb6 100644 --- a/.planning/plans/2026-08-03-skill-only-tools.md +++ b/.planning/plans/2026-08-03-skill-only-tools.md @@ -26,7 +26,7 @@ Why a new module rather than more lines in `orchestrator_interact_action.py`: that file is already 3388 lines, and the repo's established pattern is one concern per module (`access.py`, `egress.py`, `catalog.py`, `continuation.py`). The gate is a closed, testable unit with no orchestrator dependencies. -**Conventions for every task:** the commit gate in [CLAUDE.md](../../CLAUDE.md) §6 is mandatory — `pre-commit run --all-files` **and** the affected pytest slice must pass before each commit. Do not use `--no-verify`. Commits are authored as the repo user with no Claude co-author trailer. +**Conventions for every task:** the commit gate in [AGENTS.md](../../AGENTS.md) §6 is mandatory — `pre-commit run --all-files` **and** the affected pytest slice must pass before each commit. Do not use `--no-verify`. Commits are authored as the repo user with no Claude co-author trailer. --- diff --git a/.planning/reference/action-authoring.md b/.planning/reference/action-authoring.md index e632a832..df83b299 100644 --- a/.planning/reference/action-authoring.md +++ b/.planning/reference/action-authoring.md @@ -472,6 +472,15 @@ async def on_register(self): ## 13. Common pitfalls +Actions inherit jvspatial's protected entity setter. Declare persisted state with `attribute(...)` and runtime-only underscore state with Pydantic `PrivateAttr` on the class. Do this for clients, registries, caches, and per-turn flags set after initialization. Declared private methods may be replaced on an instance in tests; undeclared `_name` assignments raise `AttributeError`. Keep test doubles on real IDs when exercising JsonDB, which now rejects malformed entity IDs. + +```python +from pydantic import PrivateAttr + +class MyAction(Action): + _registry: dict[str, object] = PrivateAttr(default_factory=dict) +``` + | Mistake | Fix | |---|---| | Forgetting `from . import endpoints` in `__init__.py` | Endpoints don't register. Add the import. | @@ -479,6 +488,7 @@ async def on_register(self): | Top-level `InteractAction` not routing to children | Child actions never execute. Call `await visitor.visit(child)` from `execute()`. | | Long sync work inside `execute()` | Blocks the response. Use `run_in_background=True` for non-critical work, or enqueue a `PROACTIVE` task via `TaskStore.enqueue_proactive` / `TaskMonitor`. | | Mutating `self.metadata` directly | Lost on next load — `metadata` is rebuilt from `info.yaml`. Use `attribute(...)` fields for persistent state. | +| Assigning `self._registry` without declaring it | jvspatial rejects the undeclared name. Add a typed Pydantic `PrivateAttr` on the class. | | Swallowing exceptions in lifecycle hooks | Errors go silent. Let them propagate — the framework's `enable()`/`disable()` wrappers log them. | | Hard-coding API keys | Use `attribute(default="")` + agent.yaml `${ENV_VAR}` indirection. | | Skipping `info.yaml` | Loader skips the action package. Always ship one. | diff --git a/.planning/reference/jvspatial-integration.md b/.planning/reference/jvspatial-integration.md index 3cd7a3b2..13f50a4c 100644 --- a/.planning/reference/jvspatial-integration.md +++ b/.planning/reference/jvspatial-integration.md @@ -7,7 +7,7 @@ ## 1. Where jvspatial lives - **Source**: `/Users/eldonmarks/Briefcase/dev/jv/jvspatial` (sibling directory in this workspace). -- **Pip install**: declared in [`pyproject.toml`](../../pyproject.toml) as `jvspatial==0.0.21`. +- **Pip install**: [`pyproject.toml`](../../pyproject.toml) and both requirements files temporarily pin the exact patched jvspatial Git commit. Replace all three refs with the published patched version before publishing jvagent; a fresh `pip install -e '.[test]'` must resolve the reviewed commit in `jvspatial`'s `direct_url.json`. - **Own docs**: jvspatial has its own [`README.md`](../../../jvspatial/README.md) and [`SPEC.md`](../../../jvspatial/SPEC.md). Treat those as authoritative for anything below. --- @@ -20,8 +20,8 @@ Object ── persistence-capable Pydantic-style base └── Node ── graph node, has edges + visitor support ├── Edge ── relationship between two Nodes - └── Walker ── traverses a graph, visits Nodes - └── Root ── singleton Node, anchor for everything + └── Root ── singleton Node, anchor for everything +Walker ── separate traversal class; visits Nodes ``` - `Object` (`jvspatial/core/entities/object.py:19`) — base persistence-capable class with id, entity type, graph context. All entity types inherit. Pydantic-aware. @@ -30,6 +30,8 @@ Object ── persistence-capable Pydantic-style base - `Walker` (`jvspatial/core/entities/walker.py:83`) — traversal agent. Visit queue + trail. Built-in protection: `max_steps=10000`, `max_visits_per_node=100`, `max_execution_time=300s`, `max_queue_size=1000`. - `Root` (`jvspatial/core/entities/root.py:11`) — singleton; id fixed at `"n.Root.root"`. Created once. +`Object.__setattr__` rejects undeclared fields, including new underscore names. Use `attribute(...)` for persisted state and Pydantic `PrivateAttr` for action-local runtime state (for example a registry, client, or cached dispatch state). A private method already declared on the class can be replaced on an instance for a test double; this does not allow an arbitrary new `_helper`. Do not bypass the setter with `object.__setattr__`. + ### 2.2 Walker API (read by every InteractAction author) | Method | Signature | Purpose | @@ -77,6 +79,8 @@ async def my_handler(...): ... - `roles=[...]` restricts to those roles. - Endpoints inside an action package's `endpoints.py` are auto-discovered at action register time. +With auth enabled, jvspatial caps register, login, forgot-password, and reset-password at five requests per 60 seconds per IP even when its global limiter is off. Test hosts or applications with an equivalent trusted limiter may explicitly set `RateLimitConfig(auth_entrypoint_rate_limit_enabled=False)`; do not assume the global limiter flag covers those routes. Authenticated mounted ASGI and raw FastAPI routes may lack jvspatial endpoint metadata and remain reachable after authentication; built-in `/status`, `/logs`, and `/graph` routes still require admin roles. + ### 2.5 Persistence jvspatial supports five backends usable from jvagent, selected via env vars: @@ -100,7 +104,7 @@ await node.save() # Required after property mutatio results = await MyNode.find({"context.x": 1}) # Mongo-style query ``` -`Object.count()` does not exist — use `len(await Entity.find(query))` (jvspatial SPEC.md). For high-cardinality counts, design indexes accordingly. +`Object.count(query)` is available; use it instead of loading every matching entity just to count them. ### 2.6 Context diff --git a/.planning/reference/memory-and-pruning.md b/.planning/reference/memory-and-pruning.md index 86b45775..c9d8a109 100644 --- a/.planning/reference/memory-and-pruning.md +++ b/.planning/reference/memory-and-pruning.md @@ -1,6 +1,6 @@ # Memory & Pruning -> Deep dive on `User` / `Conversation` / `Interaction` lifecycle and the rolling-window pruning mechanism. Companion to [`SPEC.md`](../SPEC.md) §5, [`jvagent/memory/CLAUDE.md`](../../jvagent/memory/CLAUDE.md), [`adr/0003-interaction-limit-pruning.md`](../adr/0003-interaction-limit-pruning.md). +> Deep dive on `User` / `Conversation` / `Interaction` lifecycle and the rolling-window pruning mechanism. Companion to [`SPEC.md`](../SPEC.md) §5, [`jvagent/memory/AGENTS.md`](../../jvagent/memory/AGENTS.md), [`adr/0003-interaction-limit-pruning.md`](../adr/0003-interaction-limit-pruning.md). --- diff --git a/.planning/reference/observability.md b/.planning/reference/observability.md index 4cc782d3..12c1c5c2 100644 --- a/.planning/reference/observability.md +++ b/.planning/reference/observability.md @@ -119,4 +119,4 @@ If you add one of these, document it here and in [`/docs/logging.md`](../../docs - [`/docs/logging.md`](../../docs/logging.md) (734 lines) — comprehensive logging architecture - [`/docs/error-logging.md`](../../docs/error-logging.md) (401 lines) — error rollup mechanics - [`/docs/interaction-logging.md`](../../docs/interaction-logging.md) (275 lines) — turn-level events + INTERACTION level -- Local subsystem guide: [`/jvagent/logging/CLAUDE.md`](../../jvagent/logging/CLAUDE.md) +- Local subsystem guide: [`/jvagent/logging/AGENTS.md`](../../jvagent/logging/AGENTS.md) diff --git a/.planning/reviews/2026-07-16-core-review.md b/.planning/reviews/2026-07-16-core-review.md index ab989227..a37f79cd 100644 --- a/.planning/reviews/2026-07-16-core-review.md +++ b/.planning/reviews/2026-07-16-core-review.md @@ -151,7 +151,7 @@ Lower-severity items (equal-timestamp chain infinite loop `interaction.py:741-75 ### Phase 5 — Test debt (blocks regressions in all of the above) - **Concurrency suite** (currently zero): duplicate User/Conversation/action creation; lock-lease expiry mid-turn; contextvar reentrancy; proactive double-claim; concurrent turns on one conversation. -- **Pruning regression suite**: the two test files `memory/CLAUDE.md` references **do not exist**; `_prune_old_interactions` (cap, chain rewiring, `last_interaction_id`, limit-sync) is essentially untested. +- **Pruning regression suite**: the two test files `memory/AGENTS.md` references **do not exist**; `_prune_old_interactions` (cap, chain rewiring, `last_interaction_id`, limit-sync) is essentially untested. - **Interact auth**: `log`-mode no-leak, streaming-path identity guard parity, rate-limiter spoofing/isolation. - **Repair multi-tick resume**, **loop-lifecycle** (scheduler actually fires after `run_server`), **reply endpoint authz**, **`get_model_action` fallback**, deregistration cleanup paths. @@ -160,6 +160,6 @@ Lower-severity items (equal-timestamp chain infinite loop `interaction.py:741-75 ## 3. Suggested sequencing for review - **Ship now as isolated security patches:** Phase 0 items 1–4 (rate limiter, reply authz, reason leak, log-mode token). Each is small and independently testable. -- **One design ADR** for Phase 1+2 (identity + locking substrate) — it changes contracts in `core/CLAUDE.md` and `memory/CLAUDE.md` and supersedes the "compound index rejects on save" claim, which is false on the default adapter. +- **One design ADR** for Phase 1+2 (identity + locking substrate) — it changes contracts in `core/AGENTS.md` and `memory/AGENTS.md` and supersedes the "compound index rejects on save" claim, which is false on the default adapter. - **Phase 3/4** as per-subsystem PRs behind the substrate work. - Treat **Phase 5** as acceptance criteria, not a follow-up: the high-severity band is entirely untested today. diff --git a/.planning/reviews/2026-09-01-full-code-review.md b/.planning/reviews/2026-09-01-full-code-review.md index 4a8f3723..3ae05419 100644 --- a/.planning/reviews/2026-09-01-full-code-review.md +++ b/.planning/reviews/2026-09-01-full-code-review.md @@ -27,7 +27,7 @@ The codebase is well-engineered where invariants are mechanically checkable (asy | **pytest** (`pytest tests/ -p no:randomly`) | **PASS** (exit 0) | 3,558 tests collected | | **mypy dev venv** (`mypy jvagent/`) | **FAIL** | **270 errors in 48 files** (local mypy 1.20.1 with jvspatial installed) | -**Mypy gap:** CLAUDE.md asserts the pre-commit hook and bare `mypy jvagent/` are aligned; they are not. Inside the isolated pre-commit env every jvspatial type is `Any`, so boundary errors like `"Object" has no attribute "name"` (`cli/agent_commands.py:248`) and `AuthConfig` keyword mismatches (`cli/server_config.py:418`) are invisible to the enforced gate. Representative errors live precisely at the jvagent↔jvspatial boundary — the area most likely to hide real runtime bugs. +**Mypy gap:** AGENTS.md asserts the pre-commit hook and bare `mypy jvagent/` are aligned; they are not. Inside the isolated pre-commit env every jvspatial type is `Any`, so boundary errors like `"Object" has no attribute "name"` (`cli/agent_commands.py:248`) and `AuthConfig` keyword mismatches (`cli/server_config.py:418`) are invisible to the enforced gate. Representative errors live precisely at the jvagent↔jvspatial boundary — the area most likely to hide real runtime bugs. --- @@ -280,7 +280,7 @@ Memory user-listing returns HTTP 200 with `total=0` on backend outage; `_send_to Pre-commit passes; dev venv reports 270 errors. Hook lacks jvspatial/pydantic/httpx stubs, so boundary types are unchecked — precisely where integration bugs live. -**Fix theme:** Add jvspatial (and key deps) to hook `additional_dependencies`, or pin a shared stub package; align CLAUDE.md claim with reality. +**Fix theme:** Add jvspatial (and key deps) to hook `additional_dependencies`, or pin a shared stub package; align AGENTS.md claim with reality. --- @@ -387,7 +387,7 @@ Most HIGH and MEDIUM items from the 2026-09-01 review are now addressed. Remaini - Prior core review: [2026-07-16-core-review.md](2026-07-16-core-review.md) (H19 = C3) - Prior once-over: [2026-07-17-once-over.md](2026-07-17-once-over.md) -- Agent guide: [CLAUDE.md](../../CLAUDE.md) +- Agent guide: [AGENTS.md](../../AGENTS.md) - Orchestrator design: [docs/ORCHESTRATOR.md](../../docs/ORCHESTRATOR.md) - Thin harness: [docs/thin-harness.md](../../docs/thin-harness.md) diff --git a/.planning/runbooks/local-dev.md b/.planning/runbooks/local-dev.md index 58b8f82c..761d3138 100644 --- a/.planning/runbooks/local-dev.md +++ b/.planning/runbooks/local-dev.md @@ -24,6 +24,8 @@ pip install -e ".[dev]" pre-commit install # one-time hook setup (pre-commit + pre-push) ``` +On the security compatibility branch, the dependency files temporarily pin a specific jvspatial Git commit. Do not replace it with an editable sibling checkout or an older PyPI version when validating this branch: a fresh install must use the pinned commit. After the patched jvspatial release, replace the Git ref in `pyproject.toml`, `requirements.txt`, and `requirements-all.txt` together, then rerun the full install and tests. + `pre-commit run --all-files` checks **tracked** files only; stage new files (`git add -A`) before running it or they are silently skipped. The installed commit-time hook covers staged files, so keep the hooks installed @@ -209,7 +211,7 @@ jvagent . |---|---|---| | `KeyError: 'JVAGENT_ADMIN_PASSWORD'` | `.env` not loaded or var missing | Confirm `.env` is at app root; `cat .env \| grep JVAGENT_ADMIN_PASSWORD` | | `RuntimeError: attached to different loop` | jvspatial entity cached from a prior event loop | Restart the process; this often resolves on warm path | -| `Unknown argument: --foo` | Flag not recognized by `cli/main.py` | `jvagent --help` or read [`jvagent/cli/CLAUDE.md`](../../jvagent/cli/CLAUDE.md) | +| `Unknown argument: --foo` | Flag not recognized by `cli/main.py` | `jvagent --help` or read [`jvagent/cli/AGENTS.md`](../../jvagent/cli/AGENTS.md) | | `--source and --merge require --update` | Misused flags | Add `--update` or drop the modifier | | `--purge` exits with "only allowed in development mode" | `JVSPATIAL_ENVIRONMENT` not `development` | `export JVSPATIAL_ENVIRONMENT=development` | | `action package not found: namespace/foo` | dir layout mismatch with `info.yaml` name | Match `info.yaml:package.name` to dir path | @@ -220,7 +222,7 @@ jvagent . ## 13. Reading list (after this runbook) -- [`/CLAUDE.md`](../../CLAUDE.md) — agent guide +- [`/AGENTS.md`](../../AGENTS.md) — agent guide - [`.planning/architecture.md`](../architecture.md) — diagrams - [`/docs/ORCHESTRATOR.md`](../../docs/ORCHESTRATOR.md) — Executive pattern deep dive - [`/docs/scaffolding.md`](../../docs/scaffolding.md) — `jvagent app create` reference diff --git a/.planning/specs/2026-06-13-replyaction-sole-egress-design.md b/.planning/specs/2026-06-13-replyaction-sole-egress-design.md index ec764466..65aea854 100644 --- a/.planning/specs/2026-06-13-replyaction-sole-egress-design.md +++ b/.planning/specs/2026-06-13-replyaction-sole-egress-design.md @@ -80,11 +80,11 @@ Today ReplyAction applies the directives/params already on the interaction in it `scaffold/builtin_profiles/{minimal,orchestrator,research}.yaml` and `examples/jvagent_app/.../agent.yaml` — replace PersonaAction with ReplyAction in the action set; update READMEs/architecture docs. ### 5.7 Delete PersonaAction -Remove `jvagent/action/persona/` (action, endpoints, prompt_builder, prompts, info.yaml, README) and its dedicated tests (`test_persona_*.py`). Update `tests/CLAUDE.md`, `action/CLAUDE.md`, `interact/CLAUDE.md` egress decision trees to ReplyAction. +Remove `jvagent/action/persona/` (action, endpoints, prompt_builder, prompts, info.yaml, README) and its dedicated tests (`test_persona_*.py`). Update `tests/AGENTS.md`, `action/AGENTS.md`, `interact/AGENTS.md` egress decision trees to ReplyAction. ## 6. Files touched -Modify: `reply/reply_action.py` (history subsume, gather, N=1 publish), `reply/history.py` (new helper), `memory/interaction.py` (standalone flag), `base.py` (`get_responder`, docstrings), `interact/base.py` (respond collapse), `handoff_interact_action.py`, `parameters.py`, `model/context.py`, `core/agent.py`, `core/profiling.py`, scaffold profiles, examples, CLAUDE.md egress trees. +Modify: `reply/reply_action.py` (history subsume, gather, N=1 publish), `reply/history.py` (new helper), `memory/interaction.py` (standalone flag), `base.py` (`get_responder`, docstrings), `interact/base.py` (respond collapse), `handoff_interact_action.py`, `parameters.py`, `model/context.py`, `core/agent.py`, `core/profiling.py`, scaffold profiles, examples, AGENTS.md egress trees. Delete: `jvagent/action/persona/**`, `tests/action/test_persona_*.py`. ADR: refine ADR-0024 / new ADR-0025 "ReplyAction is the single output contract." @@ -95,7 +95,7 @@ ADR: refine ADR-0024 / new ADR-0025 "ReplyAction is the single output contract." - **C — `get_responder` → ReplyAction only**; remove PersonaAction fallback + isinstance. - **D — migrate direct consumers** (handoff history, parameters/model/agent/profiling references). - **E — scaffold profiles + examples → ReplyAction**. -- **F — delete `persona/` + dedicated tests**; update CLAUDE.md egress trees. +- **F — delete `persona/` + dedicated tests**; update AGENTS.md egress trees. - **G — full verification** + ADR. Each phase keeps the suite green (excluding the unrelated `web_fetch` dep gap). Phases A-C make ReplyAction sufficient before any deletion. diff --git a/AGENTS.md b/AGENTS.md index 004fdedb..759c0777 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,5 +1,250 @@ -# AGENTS.md +# AGENTS.md — jvagent Master Agent Guide -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents (Codex CLI, Gemini CLI, etc.). +> This file is the entry point for **AI agents** (Claude Code, Codex CLI, Gemini CLI, etc.) working on jvagent. Human contributors should start with [`README.md`](README.md). Both audiences are welcome here, but agent-targeted reference docs live under [`.planning/`](.planning/) and scoped `AGENTS.md` files live in the relevant source directories. -## Imported Claude Cowork project instructions +Keep agent guidance in this file or the nearest scoped `AGENTS.md`. Do not create `CLAUDE.md` shims. + +--- + +## 1. What jvagent is (60-second version) + +A modular AI-agent platform built on [jvspatial](.planning/reference/jvspatial-integration.md)'s object-spatial graph framework. + +**jvspatial security compatibility:** keep `pyproject.toml`, `requirements.txt`, and `requirements-all.txt` on the same published jvspatial version (`0.1.0`), and verify a fresh install when updating the pin. jvspatial rejects undeclared instance attributes on entity subclasses: declare persisted fields with `attribute(...)` and runtime-only underscore state with Pydantic `PrivateAttr`. See [the integration reference](.planning/reference/jvspatial-integration.md) and [action authoring](.planning/reference/action-authoring.md). + +- An *app* declares one or more *agents* in YAML. +- Each agent owns a graph of *actions* (plugins) plus a per-user memory subgraph (`User → Conversation → Interaction`). +- Incoming traffic at `POST /agents/{id}/interact` becomes an `Interaction`; an `InteractWalker` visits the agent's `InteractAction`s in weight order; the **Orchestrator** action (weight `-200`) runs the whole turn in one `execute()`: a deterministic **continuation check** (resume an active flow from the conversation `TaskStore`), then a bounded **think-act-observe loop** over a unified tool surface. Routing = tool selection; turn-lock = an active flow that hasn't returned `COMPLETE`/`YIELD`. +- Production-shaped: namespaced plugins, lifecycle hooks, response bus with channel adapters, rolling-window memory pruning, separate logs DB. + +Use cases: turn-based chatbots, channel adapters (WhatsApp / Messenger / email / web), long-running autonomous agents. + +--- + +## 2. Where to find things (the only table you need) + +| You want to... | Read | +|---|---| +| **Navigate the design docs** | [`.planning/README.md`](.planning/README.md) (folder index) | +| **Get the big picture** | [`.planning/PROJECT.md`](.planning/PROJECT.md) | +| **v2.0 Harness Excellence** | [`.planning/ROADMAP.md`](.planning/ROADMAP.md) · [`.planning/REQUIREMENTS.md`](.planning/REQUIREMENTS.md) · [`docs/HARNESS_EXCELLENCE_PLAN.md`](docs/HARNESS_EXCELLENCE_PLAN.md) | +| **Look up normative semantics** (invariants, contracts) | [`.planning/SPEC.md`](.planning/SPEC.md) | +| **Choose a deployment pattern** (Orchestrator) | [`.planning/PATTERNS.md`](.planning/PATTERNS.md) | +| **See diagrams** (boot, interact, executive, pruning) | [`.planning/architecture.md`](.planning/architecture.md) | +| **Define a term** | [`.planning/GLOSSARY.md`](.planning/GLOSSARY.md) | +| **Build a new action** | [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md) | +| **Thin harness principle** (platform-wide) | [`docs/thin-harness.md`](docs/thin-harness.md) | +| **Build / extend an interview skill** | [`docs/thin-harness.md`](docs/thin-harness.md) + [`jvagent/action/interview/docs/thin-harness.md`](jvagent/action/interview/docs/thin-harness.md) (profile) + [`jvagent/action/interview/AGENTS.md`](jvagent/action/interview/AGENTS.md) | +| **See every existing action** | [`.planning/reference/actions-catalog.md`](.planning/reference/actions-catalog.md) | +| **Understand the jvspatial dependency** | [`.planning/reference/jvspatial-integration.md`](.planning/reference/jvspatial-integration.md) | +| **Understand memory pruning** | [`.planning/reference/memory-and-pruning.md`](.planning/reference/memory-and-pruning.md) | +| **Tune / query logging** | [`.planning/reference/observability.md`](.planning/reference/observability.md) + [`docs/logging.md`](docs/logging.md) | +| **Find a config key** | [`.planning/reference/configuration-keys.md`](.planning/reference/configuration-keys.md) + [`docs/environment-keys-reference.md`](docs/environment-keys-reference.md) | +| **Run on PostgreSQL** | [`docs/postgres.md`](docs/postgres.md) | +| **Understand the Orchestrator pattern** | [`docs/ORCHESTRATOR.md`](docs/ORCHESTRATOR.md) + ADRs [0012](.planning/adr/0012-skill-executive-architecture.md) (architecture), [0013](.planning/adr/0013-togglable-deterministic-turn-lock.md) (turn-lock), [0014](.planning/adr/0014-identity-on-agent-replyaction-egress.md) (identity/egress), [0015](.planning/adr/0015-skill-executive-configuration-surface.md) (config surface), [0016](.planning/adr/0016-model-gearing-light-heavy.md) (model gearing), [0017](.planning/adr/0017-two-skill-specs-code-execution-substrate.md) (skill specs + code execution), [0018](.planning/adr/0018-lean-tool-surfacing.md) (lean surfacing), [0019](.planning/adr/0019-orchestrator-resumable-plan.md) (resumable plan), [0041](.planning/adr/0041-gearing-and-cost-policy-in-core.md) (gearing/cost in core), [0042](.planning/adr/0042-session-context-ground-truth.md) (session clock/channel), [0048](.planning/adr/0048-parallel-tool-dispatch.md) (parallel tool dispatch), [0044](.planning/adr/0044-native-tool-calling-protocol.md) (native tool-calling protocol) | +| **Document conversational test scenarios (CUCS)** | [`.planning/reference/conversation-use-cases.md`](.planning/reference/conversation-use-cases.md) + [ADR-0027](.planning/adr/0027-conversation-use-case-spec.md) | +| **Run jvagent locally** | [`.planning/runbooks/local-dev.md`](.planning/runbooks/local-dev.md) | +| **Add a new action end-to-end** | [`.planning/runbooks/add-action.md`](.planning/runbooks/add-action.md) | +| **Canary the LiteLLM transport on one deployment** | [`.planning/runbooks/litellm-canary.md`](.planning/runbooks/litellm-canary.md) + [`docs/language-models.md`](docs/language-models.md) (migration guide) | +| **Send a proactive (agent-initiated) message** | [`docs/proactive-messages.md`](docs/proactive-messages.md) | +| **Embed the customer-facing popup chat** | [`docs/jvmessenger.md`](docs/jvmessenger.md) + [ADR-0035](.planning/adr/0035-embeddable-chat-messenger.md) (source: [`jvmessenger/`](jvmessenger/), served by `jvagent messenger` from [`jvagent/messenger/`](jvagent/messenger/)) | +| **See design rationale** | [`.planning/adr/`](.planning/adr/) | +| **User-facing onboarding** | [`README.md`](README.md) | + +--- + +## 3. Graph hierarchy (memorize this) + +``` +Root → App → Agents → Agent ─┬─ Actions → Action(s) → [InteractAction subclass] + └─ Memory → User → Conversation → Interaction* +``` + +- Top-level `InteractAction`s are visited by `InteractWalker` in ascending `weight` order. +- Sub-`InteractAction`s connected as children require explicit `visitor.visit(child)` from the parent's `execute()`. +- `Interaction`s are bidirectionally chained after the second one is added. +- `Conversation.interaction_limit` controls rolling-window pruning; `0` disables. + +Source anchors: +- App: [`jvagent/core/app.py:21`](jvagent/core/app.py) +- Agent: [`jvagent/core/agent.py:30`](jvagent/core/agent.py) +- Action base: [`jvagent/action/base.py:49`](jvagent/action/base.py) +- InteractAction: [`jvagent/action/interact/base.py:27`](jvagent/action/interact/base.py) +- InteractWalker: `jvagent/action/interact/interact_walker.py:47` +- Orchestrator: [`jvagent/action/orchestrator/orchestrator_interact_action.py`](jvagent/action/orchestrator/orchestrator_interact_action.py) + supporting modules ([`continuation.py`](jvagent/action/orchestrator/continuation.py), [`tools.py`](jvagent/action/orchestrator/tools.py), [`core_tools.py`](jvagent/action/orchestrator/core_tools.py), [`catalog.py`](jvagent/action/orchestrator/catalog.py), [`skills.py`](jvagent/action/orchestrator/skills.py)) +- Conversation + pruning: `jvagent/memory/conversation.py:235` (`add_interaction`) + `:490` (`_prune_old_interactions`) + +--- + +## 4. Per-subsystem guides (drop into each one before editing) + +When working inside a subdirectory, read its local `AGENTS.md` first — it's stricter and more local than this file. + +| Subdir | Local guide | +|---|---| +| `jvagent/core/` | [`jvagent/core/AGENTS.md`](jvagent/core/AGENTS.md) | +| `jvagent/memory/` | [`jvagent/memory/AGENTS.md`](jvagent/memory/AGENTS.md) | +| `jvagent/action/` | [`jvagent/action/AGENTS.md`](jvagent/action/AGENTS.md) | +| `jvagent/action/interact/` | [`jvagent/action/interact/AGENTS.md`](jvagent/action/interact/AGENTS.md) | +| `jvagent/cli/` | [`jvagent/cli/AGENTS.md`](jvagent/cli/AGENTS.md) | +| `jvagent/logging/` | [`jvagent/logging/AGENTS.md`](jvagent/logging/AGENTS.md) | +| `tests/` | [`tests/AGENTS.md`](tests/AGENTS.md) | + +Each local guide is ≤ 150 lines and self-contained for that directory. + +--- + +## 5. Development commands + +```bash +# Install +pip install -e ".[dev]" + +# Run the server (defaults to ./examples/jvagent_app or arg path) +jvagent # uses cwd +jvagent examples/jvagent_app # explicit app root +jvagent path/to/app --debug # verbose +jvagent path/to/app --update # apply merge YAML sync +jvagent path/to/app --update --source # destructive YAML sync +jvagent path/to/app --serverless # serverless single-worker + +# Subcommands +jvagent path/to/app bootstrap # bootstrap graph without starting server +jvagent path/to/app status # diagnostic snapshot +jvagent path/to/app validate # validate app.yaml + agents +jvagent bundle path/to/app # generate Dockerfile + +# Chat UI (bundled jvchat, served on its own port — see docs/jvchat.md) +jvagent chat # serve the bundled UI at http://127.0.0.1:3000 +jvagent chat --url https://my-agent # point the UI at a remote agent + +# Scaffolding +jvagent app create --yes --dir ./my_app --app-id my_app --title "My App" \ + --author "You" --agent jvagent/main_bot@minimal --profile minimal + +# Tests +pytest tests/ # all +pytest tests/action/orchestrator/ -v # one slice +pre-commit run --all-files # full lint pass + +# Lint / type +black jvagent/ +isort jvagent/ --profile black +flake8 jvagent/ --config=.flake8 +# Bare `mypy jvagent/` reads [tool.mypy] in pyproject.toml (aligned with the +# pre-commit mypy hook: strict_optional off, check_untyped_defs off, +# ignore_missing_imports, follow_imports=silent). That is the enforced gate — +# not full strict typing. +mypy jvagent/ +``` + +Full CLI reference in [`jvagent/cli/AGENTS.md`](jvagent/cli/AGENTS.md) and [`docs/scaffolding.md`](docs/scaffolding.md). + +--- + +## 6. Conventions to follow + +### Commit gate (MANDATORY — run before every commit) + +Before **any** `git commit`, both of these MUST pass — no exceptions, including +docs-touching commits (pre-commit still runs trailing-whitespace / YAML / secret +hooks): + +```bash +pre-commit run --all-files # black, isort, flake8, mypy, secrets, etc. +pytest tests/ # or the affected slice(s) at minimum +``` + +- If a hook **reformats** files (black/isort), re-stage and re-run until the run is + clean — a hook that "modifies files" is a FAILURE until a re-run passes with no + changes. Never commit with a hook still reporting changes. +- If `pytest` has any failure, fix it or explicitly call it out; do not commit over + red tests. +- `pre-commit install` also installs a pre-push pytest hook (full suite on push); + the manual run above is still required before every commit. +- **Stage first, then run.** `pre-commit run --all-files` only checks files git + tracks — an untracked new file is skipped and reports nothing, so `git add -A` + BEFORE the run (PRs #176/#177 went red in CI on exactly this: new test files + passed locally untracked, CI's black reformatted them). With the git hooks + installed (`pre-commit install`) the commit-time hook covers staged files + regardless; check `.git/hooks/pre-commit` exists in your clone. +- Do not use `git commit --no-verify` to bypass the gate. +- Applies on every branch, including hotfix/docs/chore branches. + +### When editing source +- **Type-annotate new and touched code.** Prefer annotations for readability and + IDE support; the enforced mypy gate matches pre-commit / `pyproject.toml` + (`strict_optional = false`, `check_untyped_defs = false`) — it does **not** + require full strict typing across the tree. +- **Use `attribute(...)` for all persisted Node fields.** Plain class attributes are not persisted. +- **Add a test slice** in `tests/action/{name}/` or `tests/{subsystem}/` for any new behavior. +- **Run the commit gate above** (`pre-commit run --all-files` + `pytest`) before every commit and before claiming a change is done. +- **Cite file:line** in commit messages and PR descriptions when fixing bugs — `core/app.py:124` beats "fixed the App singleton". + +### When editing docs +- **Reference, don't duplicate.** New docs link to the existing `docs/*.md` rather than rewriting them. +- **File:line refs for every claim** about runtime behavior. +- **Update [`.planning/GLOSSARY.md`](.planning/GLOSSARY.md)** when introducing a new term used in 2+ places. +- **ADRs are immutable** once accepted. To change a decision, write a new ADR that supersedes the old one. + +### When adding a feature +- **Read [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md)** first if it's a new Action. +- **Stay within the action's directory** — cross-cutting changes should be unusual. +- **Honor lifecycle hooks**: `on_register`, `on_enable`, `on_startup`, `on_disable`, `on_deregister`. +- **Default to `run_in_background=True`** for analytics, model updates, follow-ups — anything not required for the user-facing response. + +--- + +## 7. Configuration resolution (precedence) + +1. CLI flag (`--update`, `--source`, `--merge`, `--debug`, `--serverless`) +2. Environment variable (resolved via `jvspatial.env.env`) +3. `app.yaml` at the app root +4. `agent.yaml` under `agents/` +5. Action `attribute(default=...)` + +`Model HTTP retries`: `BaseModelAction` / `LanguageModelAction` expose `max_retries`, `retry_initial_delay`, `retry_max_delay`, `retry_backoff_multiplier`, `retry_jitter`, `retry_on_status_codes`. Tune per-action in `agent.yaml`. See [`docs/language-models.md`](docs/language-models.md). + +--- + +## 8. Common traps + +| Trap | What goes wrong | Fix | +|---|---|---| +| Forgetting `from . import endpoints` in `__init__.py` | HTTP routes don't register | Add the import | +| Mutating a `protected=True` field with `=` assignment | Silently dropped on some paths | Use `object.__setattr__` + `save()` (see `set_app_update_mode` at [`app.py:596`](jvagent/core/app.py)) | +| Top-level `InteractAction` not routing to children | Children never execute | Call `await visitor.visit(child)` in `execute()` | +| Setting `Agent.interaction_limit` very low after long history | Latency spike on next append | Pruning is capped per-call by `JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL` (default 100); see [`adr/0003`](.planning/adr/0003-interaction-limit-pruning.md) | +| Caching jvspatial objects across event loops | `RuntimeError: attached to different loop` on serverless warm starts | Use the per-loop lock pattern from [`app.py:97-117`](jvagent/core/app.py) | +| Using `count()` on a jvspatial entity | Method may not exist | `len(await Entity.find(query))` | +| Long blocking work in `InteractAction.execute()` | Slow user-facing response | Use `run_in_background=True` or enqueue a `PROACTIVE` task (`TaskMonitor`) | +| Creating new App nodes | Singleton violation | Always use `await App.get()` | +| Fattening the harness (prep steering, extractors, auto-store, orchestrator special-casing) | Server fights the model; regressions to pre-refactor behavior | Follow [docs/thin-harness.md](docs/thin-harness.md); for interviews also [interview profile](jvagent/action/interview/docs/thin-harness.md); extend SOP + skill extensions instead | + +--- + +## 9. Roadmap and in-flight work + +- **v2.0 Harness Excellence** (active): [`.planning/ROADMAP.md`](.planning/ROADMAP.md), [`.planning/REQUIREMENTS.md`](.planning/REQUIREMENTS.md), [`docs/HARNESS_EXCELLENCE_PLAN.md`](docs/HARNESS_EXCELLENCE_PLAN.md). +- Orchestrator design: [`.planning/adr/0012-skill-executive-architecture.md`](.planning/adr/0012-skill-executive-architecture.md). +- v1 history: [`.planning/archive/EXECUTIVE-ROADMAP.md`](.planning/archive/EXECUTIVE-ROADMAP.md). +- ADRs: [`.planning/adr/`](.planning/adr/). + +--- + +## 10. Out of scope for jvagent itself + +- Database adapter internals → jvspatial. +- Auth / JWT / HTTP wire format → jvspatial. +- Model-provider API quirks → individual `LanguageModelAction` subclasses, not the core. +- Frontend chat UI internals → `jvchat/` reference client (built bundle is served by `jvagent chat`; the Python side is just a static server — see [`docs/jvchat.md`](docs/jvchat.md)). + +--- + +## 11. If you only read 3 files... + +1. [`.planning/SPEC.md`](.planning/SPEC.md) — what jvagent guarantees. +2. The local `AGENTS.md` for the subsystem you're touching. +3. [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md) — if you're adding behavior. + +Everything else is reachable from those. diff --git a/CHANGELOG.md b/CHANGELOG.md index d647b168..a6c7b671 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,6 +8,8 @@ and this project adheres to [PEP 440](https://peps.python.org/pep-0440/) / ## [Unreleased] +## [0.1.8rc19] - 2026-09-27 + ### Added - **Unified capability discovery (`find_capability`, ADR-0055).** Primary lean discovery meta-tool ranks matching **skills** then **tools** in one observation, with `use_skill` / `load_tool` next-step cues. `find_tool` / `find_skill` remain aliases. Loop protocol, lean partial-list hint, and unknown-tool bounce steer to `find_capability` first so domain SOPs activate instead of find_tool thrash. @@ -35,6 +37,8 @@ and this project adheres to [PEP 440](https://peps.python.org/pep-0440/) / ### Changed +- **Agent guide migration.** Root and scoped `AGENTS.md` files now contain the former `CLAUDE.md` guidance. The `CLAUDE.md` files are removed; contributor and tool references point to `AGENTS.md`. + - **Artifact handler notify and PageIndex webhook import.** Async ingest reverse-indexes the jvforge `job_id` (including nested 202 bodies). `register_job` fails the ingest if the index cannot be saved. Notify requires the minted `api_key` before any job lookup, reloads the action, scans sibling `ArtifactHandlerInteractAction` nodes only after that key matches, and returns **503** on an unknown `job_id` so jvforge retries. Warm Lambda evicts the process entity cache before `Action.get`; raw `find()` `context.jvforge_job_index` is source of truth so a stale empty cached node cannot 503 a job Dynamo already has. The notify route is action-loaded only and still requires a trusted jvforge `/v1/artifacts/{job_id}` URL. Import, ready-answer generate, and WhatsApp/Messenger send run in-request via sequential `jvspatial.create_task` (Shape B); a returned schedule is awaited so Lambda cannot freeze it. A failed send leaves the reverse-index job and returns **503**. PageIndex graph import uses the same await. PageIndex LLM completions use `POST /api/pageindex/interact/webhook/{agent_id}` (old retrieval path 404s). Lambda `.docx`/Office saves sniff MIME via `file`/`file-libs` + `python-magic` in `Dockerfile.base` instead of falling through to `octet-stream`. Failures stay `logger.error` / `logger.exception`. `check_ingest_status` still prefers PageIndex, then polls jvforge and pull-imports `webhook_failed` / `completed` artifacts. - **ResponseBus now enforces the single-egress latch.** The first delivered @@ -92,6 +96,10 @@ and this project adheres to [PEP 440](https://peps.python.org/pep-0440/) / ### Fixed +- **Wire prompt-contract fixture.** Direct graph bootstrap now applies jvagent's UTF-8 persistence default, so the wire tests measure the same prompt text as the normal CLI startup instead of jvspatial's ASCII-folding default. + +- **jvspatial 0.1.0 security compatibility.** Runtime action and walker state is declared with Pydantic `PrivateAttr`; instance patches of declared private helpers continue to work, and stale interview mocks for a removed helper are gone. JsonDB test doubles use a missing conversation ID instead of an invalid `MagicMock` ID. Fresh installs require the published `jvspatial==0.1.0` release. + - **Atomic per-interaction egress claim across ResponseBus instances.** A process-wide `InteractionEgressRecord` (keyed by `interaction_id`) is the single durable latch for user delivery and `message_type=final`. Rematerialized diff --git a/CLAUDE.md b/CLAUDE.md deleted file mode 100644 index 6d405492..00000000 --- a/CLAUDE.md +++ /dev/null @@ -1,248 +0,0 @@ -# CLAUDE.md — jvagent Master Agent Guide - -> This file is the entry point for **AI agents** (Claude Code, Codex CLI, Gemini CLI, etc.) working on jvagent. Human contributors should start with [`README.md`](README.md). Both audiences are welcome here, but agent-targeted reference docs live under [`.planning/`](.planning/) and per-subsystem `CLAUDE.md` files are scattered through the source tree. -> -> **AGENTS.md** at the repo root is a one-line pointer to this file. - ---- - -## 1. What jvagent is (60-second version) - -A modular AI-agent platform built on [jvspatial](.planning/reference/jvspatial-integration.md)'s object-spatial graph framework. - -- An *app* declares one or more *agents* in YAML. -- Each agent owns a graph of *actions* (plugins) plus a per-user memory subgraph (`User → Conversation → Interaction`). -- Incoming traffic at `POST /agents/{id}/interact` becomes an `Interaction`; an `InteractWalker` visits the agent's `InteractAction`s in weight order; the **Orchestrator** action (weight `-200`) runs the whole turn in one `execute()`: a deterministic **continuation check** (resume an active flow from the conversation `TaskStore`), then a bounded **think-act-observe loop** over a unified tool surface. Routing = tool selection; turn-lock = an active flow that hasn't returned `COMPLETE`/`YIELD`. -- Production-shaped: namespaced plugins, lifecycle hooks, response bus with channel adapters, rolling-window memory pruning, separate logs DB. - -Use cases: turn-based chatbots, channel adapters (WhatsApp / Messenger / email / web), long-running autonomous agents. - ---- - -## 2. Where to find things (the only table you need) - -| You want to... | Read | -|---|---| -| **Navigate the design docs** | [`.planning/README.md`](.planning/README.md) (folder index) | -| **Get the big picture** | [`.planning/PROJECT.md`](.planning/PROJECT.md) | -| **v2.0 Harness Excellence** | [`.planning/ROADMAP.md`](.planning/ROADMAP.md) · [`.planning/REQUIREMENTS.md`](.planning/REQUIREMENTS.md) · [`docs/HARNESS_EXCELLENCE_PLAN.md`](docs/HARNESS_EXCELLENCE_PLAN.md) | -| **Look up normative semantics** (invariants, contracts) | [`.planning/SPEC.md`](.planning/SPEC.md) | -| **Choose a deployment pattern** (Orchestrator) | [`.planning/PATTERNS.md`](.planning/PATTERNS.md) | -| **See diagrams** (boot, interact, executive, pruning) | [`.planning/architecture.md`](.planning/architecture.md) | -| **Define a term** | [`.planning/GLOSSARY.md`](.planning/GLOSSARY.md) | -| **Build a new action** | [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md) | -| **Thin harness principle** (platform-wide) | [`docs/thin-harness.md`](docs/thin-harness.md) | -| **Build / extend an interview skill** | [`docs/thin-harness.md`](docs/thin-harness.md) + [`jvagent/action/interview/docs/thin-harness.md`](jvagent/action/interview/docs/thin-harness.md) (profile) + [`jvagent/action/interview/CLAUDE.md`](jvagent/action/interview/CLAUDE.md) | -| **See every existing action** | [`.planning/reference/actions-catalog.md`](.planning/reference/actions-catalog.md) | -| **Understand the jvspatial dependency** | [`.planning/reference/jvspatial-integration.md`](.planning/reference/jvspatial-integration.md) | -| **Understand memory pruning** | [`.planning/reference/memory-and-pruning.md`](.planning/reference/memory-and-pruning.md) | -| **Tune / query logging** | [`.planning/reference/observability.md`](.planning/reference/observability.md) + [`docs/logging.md`](docs/logging.md) | -| **Find a config key** | [`.planning/reference/configuration-keys.md`](.planning/reference/configuration-keys.md) + [`docs/environment-keys-reference.md`](docs/environment-keys-reference.md) | -| **Run on PostgreSQL** | [`docs/postgres.md`](docs/postgres.md) | -| **Understand the Orchestrator pattern** | [`docs/ORCHESTRATOR.md`](docs/ORCHESTRATOR.md) + ADRs [0012](.planning/adr/0012-skill-executive-architecture.md) (architecture), [0013](.planning/adr/0013-togglable-deterministic-turn-lock.md) (turn-lock), [0014](.planning/adr/0014-identity-on-agent-replyaction-egress.md) (identity/egress), [0015](.planning/adr/0015-skill-executive-configuration-surface.md) (config surface), [0016](.planning/adr/0016-model-gearing-light-heavy.md) (model gearing), [0017](.planning/adr/0017-two-skill-specs-code-execution-substrate.md) (skill specs + code execution), [0018](.planning/adr/0018-lean-tool-surfacing.md) (lean surfacing), [0019](.planning/adr/0019-orchestrator-resumable-plan.md) (resumable plan), [0041](.planning/adr/0041-gearing-and-cost-policy-in-core.md) (gearing/cost in core), [0042](.planning/adr/0042-session-context-ground-truth.md) (session clock/channel), [0048](.planning/adr/0048-parallel-tool-dispatch.md) (parallel tool dispatch), [0044](.planning/adr/0044-native-tool-calling-protocol.md) (native tool-calling protocol) | -| **Document conversational test scenarios (CUCS)** | [`.planning/reference/conversation-use-cases.md`](.planning/reference/conversation-use-cases.md) + [ADR-0027](.planning/adr/0027-conversation-use-case-spec.md) | -| **Run jvagent locally** | [`.planning/runbooks/local-dev.md`](.planning/runbooks/local-dev.md) | -| **Add a new action end-to-end** | [`.planning/runbooks/add-action.md`](.planning/runbooks/add-action.md) | -| **Canary the LiteLLM transport on one deployment** | [`.planning/runbooks/litellm-canary.md`](.planning/runbooks/litellm-canary.md) + [`docs/language-models.md`](docs/language-models.md) (migration guide) | -| **Send a proactive (agent-initiated) message** | [`docs/proactive-messages.md`](docs/proactive-messages.md) | -| **Embed the customer-facing popup chat** | [`docs/jvmessenger.md`](docs/jvmessenger.md) + [ADR-0035](.planning/adr/0035-embeddable-chat-messenger.md) (source: [`jvmessenger/`](jvmessenger/), served by `jvagent messenger` from [`jvagent/messenger/`](jvagent/messenger/)) | -| **See design rationale** | [`.planning/adr/`](.planning/adr/) | -| **User-facing onboarding** | [`README.md`](README.md) | - ---- - -## 3. Graph hierarchy (memorize this) - -``` -Root → App → Agents → Agent ─┬─ Actions → Action(s) → [InteractAction subclass] - └─ Memory → User → Conversation → Interaction* -``` - -- Top-level `InteractAction`s are visited by `InteractWalker` in ascending `weight` order. -- Sub-`InteractAction`s connected as children require explicit `visitor.visit(child)` from the parent's `execute()`. -- `Interaction`s are bidirectionally chained after the second one is added. -- `Conversation.interaction_limit` controls rolling-window pruning; `0` disables. - -Source anchors: -- App: [`jvagent/core/app.py:21`](jvagent/core/app.py) -- Agent: [`jvagent/core/agent.py:30`](jvagent/core/agent.py) -- Action base: [`jvagent/action/base.py:49`](jvagent/action/base.py) -- InteractAction: [`jvagent/action/interact/base.py:27`](jvagent/action/interact/base.py) -- InteractWalker: `jvagent/action/interact/interact_walker.py:47` -- Orchestrator: [`jvagent/action/orchestrator/orchestrator_interact_action.py`](jvagent/action/orchestrator/orchestrator_interact_action.py) + supporting modules ([`continuation.py`](jvagent/action/orchestrator/continuation.py), [`tools.py`](jvagent/action/orchestrator/tools.py), [`core_tools.py`](jvagent/action/orchestrator/core_tools.py), [`catalog.py`](jvagent/action/orchestrator/catalog.py), [`skills.py`](jvagent/action/orchestrator/skills.py)) -- Conversation + pruning: `jvagent/memory/conversation.py:235` (`add_interaction`) + `:490` (`_prune_old_interactions`) - ---- - -## 4. Per-subsystem guides (drop into each one before editing) - -When working inside a subdirectory, read its local `CLAUDE.md` first — it's stricter and more local than this file. - -| Subdir | Local guide | -|---|---| -| `jvagent/core/` | [`jvagent/core/CLAUDE.md`](jvagent/core/CLAUDE.md) | -| `jvagent/memory/` | [`jvagent/memory/CLAUDE.md`](jvagent/memory/CLAUDE.md) | -| `jvagent/action/` | [`jvagent/action/CLAUDE.md`](jvagent/action/CLAUDE.md) | -| `jvagent/action/interact/` | [`jvagent/action/interact/CLAUDE.md`](jvagent/action/interact/CLAUDE.md) | -| `jvagent/cli/` | [`jvagent/cli/CLAUDE.md`](jvagent/cli/CLAUDE.md) | -| `jvagent/logging/` | [`jvagent/logging/CLAUDE.md`](jvagent/logging/CLAUDE.md) | -| `tests/` | [`tests/CLAUDE.md`](tests/CLAUDE.md) | - -Each local guide is ≤ 150 lines and self-contained for that directory. - ---- - -## 5. Development commands - -```bash -# Install -pip install -e ".[dev]" - -# Run the server (defaults to ./examples/jvagent_app or arg path) -jvagent # uses cwd -jvagent examples/jvagent_app # explicit app root -jvagent path/to/app --debug # verbose -jvagent path/to/app --update # apply merge YAML sync -jvagent path/to/app --update --source # destructive YAML sync -jvagent path/to/app --serverless # serverless single-worker - -# Subcommands -jvagent path/to/app bootstrap # bootstrap graph without starting server -jvagent path/to/app status # diagnostic snapshot -jvagent path/to/app validate # validate app.yaml + agents -jvagent bundle path/to/app # generate Dockerfile - -# Chat UI (bundled jvchat, served on its own port — see docs/jvchat.md) -jvagent chat # serve the bundled UI at http://127.0.0.1:3000 -jvagent chat --url https://my-agent # point the UI at a remote agent - -# Scaffolding -jvagent app create --yes --dir ./my_app --app-id my_app --title "My App" \ - --author "You" --agent jvagent/main_bot@minimal --profile minimal - -# Tests -pytest tests/ # all -pytest tests/action/orchestrator/ -v # one slice -pre-commit run --all-files # full lint pass - -# Lint / type -black jvagent/ -isort jvagent/ --profile black -flake8 jvagent/ --config=.flake8 -# Bare `mypy jvagent/` reads [tool.mypy] in pyproject.toml (aligned with the -# pre-commit mypy hook: strict_optional off, check_untyped_defs off, -# ignore_missing_imports, follow_imports=silent). That is the enforced gate — -# not full strict typing. -mypy jvagent/ -``` - -Full CLI reference in [`jvagent/cli/CLAUDE.md`](jvagent/cli/CLAUDE.md) and [`docs/scaffolding.md`](docs/scaffolding.md). - ---- - -## 6. Conventions to follow - -### Commit gate (MANDATORY — run before every commit) - -Before **any** `git commit`, both of these MUST pass — no exceptions, including -docs-touching commits (pre-commit still runs trailing-whitespace / YAML / secret -hooks): - -```bash -pre-commit run --all-files # black, isort, flake8, mypy, secrets, etc. -pytest tests/ # or the affected slice(s) at minimum -``` - -- If a hook **reformats** files (black/isort), re-stage and re-run until the run is - clean — a hook that "modifies files" is a FAILURE until a re-run passes with no - changes. Never commit with a hook still reporting changes. -- If `pytest` has any failure, fix it or explicitly call it out; do not commit over - red tests. -- `pre-commit install` also installs a pre-push pytest hook (full suite on push); - the manual run above is still required before every commit. -- **Stage first, then run.** `pre-commit run --all-files` only checks files git - tracks — an untracked new file is skipped and reports nothing, so `git add -A` - BEFORE the run (PRs #176/#177 went red in CI on exactly this: new test files - passed locally untracked, CI's black reformatted them). With the git hooks - installed (`pre-commit install`) the commit-time hook covers staged files - regardless; check `.git/hooks/pre-commit` exists in your clone. -- Do not use `git commit --no-verify` to bypass the gate. -- Applies on every branch, including hotfix/docs/chore branches. - -### When editing source -- **Type-annotate new and touched code.** Prefer annotations for readability and - IDE support; the enforced mypy gate matches pre-commit / `pyproject.toml` - (`strict_optional = false`, `check_untyped_defs = false`) — it does **not** - require full strict typing across the tree. -- **Use `attribute(...)` for all persisted Node fields.** Plain class attributes are not persisted. -- **Add a test slice** in `tests/action/{name}/` or `tests/{subsystem}/` for any new behavior. -- **Run the commit gate above** (`pre-commit run --all-files` + `pytest`) before every commit and before claiming a change is done. -- **Cite file:line** in commit messages and PR descriptions when fixing bugs — `core/app.py:124` beats "fixed the App singleton". - -### When editing docs -- **Reference, don't duplicate.** New docs link to the existing `docs/*.md` rather than rewriting them. -- **File:line refs for every claim** about runtime behavior. -- **Update [`.planning/GLOSSARY.md`](.planning/GLOSSARY.md)** when introducing a new term used in 2+ places. -- **ADRs are immutable** once accepted. To change a decision, write a new ADR that supersedes the old one. - -### When adding a feature -- **Read [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md)** first if it's a new Action. -- **Stay within the action's directory** — cross-cutting changes should be unusual. -- **Honor lifecycle hooks**: `on_register`, `on_enable`, `on_startup`, `on_disable`, `on_deregister`. -- **Default to `run_in_background=True`** for analytics, model updates, follow-ups — anything not required for the user-facing response. - ---- - -## 7. Configuration resolution (precedence) - -1. CLI flag (`--update`, `--source`, `--merge`, `--debug`, `--serverless`) -2. Environment variable (resolved via `jvspatial.env.env`) -3. `app.yaml` at the app root -4. `agent.yaml` under `agents/` -5. Action `attribute(default=...)` - -`Model HTTP retries`: `BaseModelAction` / `LanguageModelAction` expose `max_retries`, `retry_initial_delay`, `retry_max_delay`, `retry_backoff_multiplier`, `retry_jitter`, `retry_on_status_codes`. Tune per-action in `agent.yaml`. See [`docs/language-models.md`](docs/language-models.md). - ---- - -## 8. Common traps - -| Trap | What goes wrong | Fix | -|---|---|---| -| Forgetting `from . import endpoints` in `__init__.py` | HTTP routes don't register | Add the import | -| Mutating a `protected=True` field with `=` assignment | Silently dropped on some paths | Use `object.__setattr__` + `save()` (see `set_app_update_mode` at [`app.py:596`](jvagent/core/app.py)) | -| Top-level `InteractAction` not routing to children | Children never execute | Call `await visitor.visit(child)` in `execute()` | -| Setting `Agent.interaction_limit` very low after long history | Latency spike on next append | Pruning is capped per-call by `JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL` (default 100); see [`adr/0003`](.planning/adr/0003-interaction-limit-pruning.md) | -| Caching jvspatial objects across event loops | `RuntimeError: attached to different loop` on serverless warm starts | Use the per-loop lock pattern from [`app.py:97-117`](jvagent/core/app.py) | -| Using `count()` on a jvspatial entity | Method may not exist | `len(await Entity.find(query))` | -| Long blocking work in `InteractAction.execute()` | Slow user-facing response | Use `run_in_background=True` or enqueue a `PROACTIVE` task (`TaskMonitor`) | -| Creating new App nodes | Singleton violation | Always use `await App.get()` | -| Fattening the harness (prep steering, extractors, auto-store, orchestrator special-casing) | Server fights the model; regressions to pre-refactor behavior | Follow [docs/thin-harness.md](docs/thin-harness.md); for interviews also [interview profile](jvagent/action/interview/docs/thin-harness.md); extend SOP + skill extensions instead | - ---- - -## 9. Roadmap and in-flight work - -- **v2.0 Harness Excellence** (active): [`.planning/ROADMAP.md`](.planning/ROADMAP.md), [`.planning/REQUIREMENTS.md`](.planning/REQUIREMENTS.md), [`docs/HARNESS_EXCELLENCE_PLAN.md`](docs/HARNESS_EXCELLENCE_PLAN.md). -- Orchestrator design: [`.planning/adr/0012-skill-executive-architecture.md`](.planning/adr/0012-skill-executive-architecture.md). -- v1 history: [`.planning/archive/EXECUTIVE-ROADMAP.md`](.planning/archive/EXECUTIVE-ROADMAP.md). -- ADRs: [`.planning/adr/`](.planning/adr/). - ---- - -## 10. Out of scope for jvagent itself - -- Database adapter internals → jvspatial. -- Auth / JWT / HTTP wire format → jvspatial. -- Model-provider API quirks → individual `LanguageModelAction` subclasses, not the core. -- Frontend chat UI internals → `jvchat/` reference client (built bundle is served by `jvagent chat`; the Python side is just a static server — see [`docs/jvchat.md`](docs/jvchat.md)). - ---- - -## 11. If you only read 3 files... - -1. [`.planning/SPEC.md`](.planning/SPEC.md) — what jvagent guarantees. -2. The local `CLAUDE.md` for the subsystem you're touching. -3. [`.planning/reference/action-authoring.md`](.planning/reference/action-authoring.md) — if you're adding behavior. - -Everything else is reachable from those. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 022f751c..00e4e038 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -2,8 +2,8 @@ Thanks for your interest in contributing! This guide covers the practical workflow. For the architecture and where things live, start with -[`CLAUDE.md`](CLAUDE.md) (the agent/contributor map) and the per-subsystem -`CLAUDE.md` files. +[`AGENTS.md`](AGENTS.md) (the agent/contributor map) and the per-subsystem +`AGENTS.md` files. ## Code of Conduct @@ -46,7 +46,7 @@ no changes — a "files were modified by this hook" result is a failure. Do not jvagent validate examples/jvagent_app # app YAML stays valid ``` -Conventions (see [`CLAUDE.md` §6](CLAUDE.md)): +Conventions (see [`AGENTS.md` §6](AGENTS.md)): - **Type-annotate everything** — Pydantic and jvspatial rely on it. - **Use `attribute(...)`** for all persisted Node fields (plain class diff --git a/README.md b/README.md index 31939b60..16d06b16 100644 --- a/README.md +++ b/README.md @@ -190,7 +190,7 @@ The graph is the source of truth. Top-level `InteractAction`s are visited by the ### Actions -An **Action** is a namespaced plugin (`namespace/action_name`) declared by an `info.yaml`. **Persisted fields use `attribute(...)`** so they live on the graph; plain class attributes do not persist. Actions expose tools via `get_tools()`, discover peers with `get_action()`, and honor lifecycle hooks. **InteractActions** additionally participate in a turn and can be dispatched as tools by the Orchestrator. To build one, start with the [action authoring contract](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/action-authoring.md). +An **Action** is a namespaced plugin (`namespace/action_name`) declared by an `info.yaml`. **Persisted fields use `attribute(...)`** so they live on the graph; runtime-only underscore fields use Pydantic `PrivateAttr` because jvspatial rejects undeclared entity attributes. Actions expose tools via `get_tools()`, discover peers with `get_action()`, and honor lifecycle hooks. **InteractActions** additionally participate in a turn and can be dispatched as tools by the Orchestrator. To build one, start with the [action authoring contract](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/action-authoring.md). ### The turn @@ -275,13 +275,13 @@ jvagent resolves configuration by precedence (highest first): ### For AI agents & contributors -Agent-facing design docs live under [`.planning/`](https://github.com/TrueSelph/jvagent/blob/main/.planning/README.md); the root [`CLAUDE.md`](https://github.com/TrueSelph/jvagent/blob/main/CLAUDE.md) is the entry point (also surfaced as [`AGENTS.md`](https://github.com/TrueSelph/jvagent/blob/main/AGENTS.md)). +Agent-facing design docs live under [`.planning/`](.planning/README.md); the root [`AGENTS.md`](AGENTS.md) is the entry point. - [Project vision](https://github.com/TrueSelph/jvagent/blob/main/.planning/PROJECT.md) · [SPEC](https://github.com/TrueSelph/jvagent/blob/main/.planning/SPEC.md) · [Patterns](https://github.com/TrueSelph/jvagent/blob/main/.planning/PATTERNS.md) · [Architecture diagrams](https://github.com/TrueSelph/jvagent/blob/main/.planning/architecture.md) · [Glossary](https://github.com/TrueSelph/jvagent/blob/main/.planning/GLOSSARY.md) - [Action authoring](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/action-authoring.md) · [Memory & pruning](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/memory-and-pruning.md) · [Observability](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/observability.md) · [jvspatial integration](https://github.com/TrueSelph/jvagent/blob/main/.planning/reference/jvspatial-integration.md) - [Decision records (ADRs)](https://github.com/TrueSelph/jvagent/tree/main/.planning/adr/) · [Specs](https://github.com/TrueSelph/jvagent/tree/main/.planning/specs/) · [Plans](https://github.com/TrueSelph/jvagent/tree/main/.planning/plans/) - Runbooks: [local dev](https://github.com/TrueSelph/jvagent/blob/main/.planning/runbooks/local-dev.md) · [add an action](https://github.com/TrueSelph/jvagent/blob/main/.planning/runbooks/add-action.md) -- Per-subsystem guides: [`core`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/core/CLAUDE.md) · [`memory`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/memory/CLAUDE.md) · [`action`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/action/CLAUDE.md) · [`interact`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/action/interact/CLAUDE.md) · [`cli`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/cli/CLAUDE.md) · [`logging`](https://github.com/TrueSelph/jvagent/blob/main/jvagent/logging/CLAUDE.md) · [`tests`](https://github.com/TrueSelph/jvagent/blob/main/tests/CLAUDE.md) +- Per-subsystem guides: [`core`](jvagent/core/AGENTS.md) · [`memory`](jvagent/memory/AGENTS.md) · [`action`](jvagent/action/AGENTS.md) · [`interact`](jvagent/action/interact/AGENTS.md) · [`cli`](jvagent/cli/AGENTS.md) · [`logging`](jvagent/logging/AGENTS.md) · [`tests`](tests/AGENTS.md) - [Changelog](https://github.com/TrueSelph/jvagent/blob/main/CHANGELOG.md) ## Authors & maintainers diff --git a/docs/postgres.md b/docs/postgres.md index 35872087..70f62208 100644 --- a/docs/postgres.md +++ b/docs/postgres.md @@ -93,7 +93,7 @@ Two defects had to be fixed upstream, both shipped in 0.0.16 ([TrueSelph/jvspati File-backed adapters never notice this. It is the same class of bug as the per-event-loop lock pattern jvagent already uses at [`core/app.py:100-124`](../jvagent/core/app.py). -Per [`CLAUDE.md`](../CLAUDE.md) §4 and [`adr/0006`](../.planning/adr/0006-jvspatial-dependency.md), database adapter behavior is jvspatial's to own — if Postgres misbehaves, fix it there rather than working around it in jvagent. +Per [`AGENTS.md`](../AGENTS.md) §4 and [`adr/0006`](../.planning/adr/0006-jvspatial-dependency.md), database adapter behavior is jvspatial's to own — if Postgres misbehaves, fix it there rather than working around it in jvagent. --- diff --git a/docs/thin-harness.md b/docs/thin-harness.md index 75fbe04c..f813fbc4 100644 --- a/docs/thin-harness.md +++ b/docs/thin-harness.md @@ -91,5 +91,5 @@ Subsystem profiles should list concrete test files that guard their contract. ## See also - [`.planning/reference/action-authoring.md`](../.planning/reference/action-authoring.md) — building new actions on this contract -- [`jvagent/action/CLAUDE.md`](../jvagent/action/CLAUDE.md) — action subsystem guide +- [`jvagent/action/AGENTS.md`](../jvagent/action/AGENTS.md) — action subsystem guide - [`jvagent/skills/README.md`](../jvagent/skills/README.md) — skill placement and specs diff --git a/jvagent/action/AGENTS.md b/jvagent/action/AGENTS.md index b4156a6e..61099384 100644 --- a/jvagent/action/AGENTS.md +++ b/jvagent/action/AGENTS.md @@ -1 +1,146 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/action/ — Agent Guide + +> Local guide for the action plugin layer. Cross-link: root [`/AGENTS.md`](../../AGENTS.md), [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md), [`/.planning/reference/actions-catalog.md`](../../.planning/reference/actions-catalog.md). + +--- + +## 1. What this directory owns + +The plugin-loadable extension surface of jvagent: + +- **`Action` base** ([`base.py:49`](base.py)) — Node subclass with lifecycle hooks, attribute config, endpoint registration, child-cascade delete, tool exposure to the Orchestrator's tool surface. +- **`InteractAction`** ([`interact/base.py:27`](interact/base.py)) — see `interact/AGENTS.md`. +- **Specialized bases**: `BaseModelAction`, `LanguageModelAction`, `BaseWebSearchAction`, `BaseSTTAction`, `BaseTTSAction`, `VectorStore`. +- **Concrete plugins** organized by topic: language models, response/bus, Orchestrator, memory-related, channel adapters, productivity integrations, tasks. Catalog in [`/.planning/reference/actions-catalog.md`](../../.planning/reference/actions-catalog.md). +- **Loader/registry** in `loader/`. +- **Plugin contracts** in `plugin_contracts.py`. + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `base.py:49` | `Action` base class (canonical) | +| `base.py:259` | `get_tools()` — default discovers `@tool`-decorated methods via `collect_tools` | +| `jvagent/tooling/tool_decorator.py` | `@tool` decorator + `collect_tools` — preferred way to publish tools | +| `jvagent/tooling/signature_schema.py` | signature → JSON Schema deriver used by `@tool` | +| `base.py:180` | `get_capabilities()` — for ReplyAction prompt aggregation | +| `base.py:300` | `delete(cascade=True)` — walks outgoing edges and cascade-deletes children | +| `base.py:331-403` | Lifecycle hook contracts (`on_register`, `on_reload`, `post_register`, `on_startup`, `on_enable`, `on_disable`, `on_deregister`); `pulse`/`healthcheck` at `670`/`678` | +| `base.py:469-587` | Endpoint discovery + unregistration (relies on `/actions/{action_id}/` path prefix) | +| `base.py:588-667` | Module unload safety (skips core + shared modules) | +| `base.py:834-981` | `get_action()` / `get_action_by_base_class()` / `get_model_action()` — cross-action lookup | +| `base.py:982-1058` | Package metadata accessors (namespace, version, type) | +| `base.py:1145-1230` | File storage helpers (action-scoped paths) | +| `interact/base.py:27` | `InteractAction` (see `interact/AGENTS.md`) | +| `actions.py` | `Actions` manager node | +| `endpoints.py` | Top-level action HTTP routes (~9 routes) | +| `loader/` | Action loader, registry, plugin discovery | +| `plugin_contracts.py` | Plugin protocol definitions | +| `streaming.py` | Streaming response helpers | + +--- + +## 3. Contracts (don't break) + +1. **`Action` subclasses MUST set `archetype` in `info.yaml`** to match the Python class name. The loader uses it. +2. **Action endpoints MUST live under `/actions/{action_id}/...`** ([`base.py:490`](base.py)). Deregister scans this prefix; non-conforming endpoints leak after `on_deregister`. +3. **`get_action()` is `O(1)`; `get_action_by_base_class()` is `O(n)`.** Don't use the latter in hot paths. +4. **Lifecycle hooks MUST not swallow exceptions** ([`base.py:694`](base.py)). The framework's `enable()`/`disable()`/`reload()` wrappers log errors automatically with the action context — silencing them hides bugs. +5. **`Action.metadata` is owned by the loader.** Mutations to it are not persisted across restarts. Use `attribute(...)` fields for persistent state. + Declare runtime-only underscore state with Pydantic `PrivateAttr`; jvspatial rejects undeclared instance attributes. +6. **Child Nodes attached via outgoing edges are cascade-deleted** when the action is deleted ([`base.py:300`](base.py)). Always connect via `await self.connect(child, direction="out")`. +7. **`is_singleton` default is `True`** ([`base.py:296`](base.py)). Override `config.singleton: false` in `info.yaml` if multiple instances per agent are allowed. +8. **Thin harness** — Actions expose capabilities via `get_tools()`; they must not classify user intent, inject prep steering, auto-store extracted values, or inline multi-step workflows. Put judgment in skill SOPs and domain logic in skill extensions. See [`docs/thin-harness.md`](../../docs/thin-harness.md). +9. **Prefer the `@tool` decorator over hand-built `Tool()`** — decorate an `async def` method with `@tool` ([`jvagent/tooling/tool_decorator.py`](../tooling/tool_decorator.py)) and the base `get_tools()` auto-publishes it. Name = `{action_name}__{method}` (override with `@tool(name=...)`); description = method docstring's first paragraph; JSON Schema = signature (use `Annotated[T, "desc"]` for per-arg docs). Hand-built `Tool()` and `get_tools()` overrides still work; only override when a tool can't be a decorated method (combine with `collect_tools(self)`). + +--- + +## 4. The four-file pattern + +Every action package MUST contain: + +``` +{namespace}/{action_name}/ +├── __init__.py # exports class + imports endpoints (for @endpoint registration) +├── {action_name}.py # Action subclass +├── endpoints.py # @endpoint-decorated routes +└── info.yaml # package metadata +``` + +Skeleton and full templates: [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md). + +--- + +## 5. Cross-action lookup table + +| Need | Method | Cost | +|---|---|---| +| One specific class | `await self.get_action(MyActionClass)` | O(1) (cached index) | +| Class by name string | `await self.get_action("MyActionClass")` | O(1) | +| Any subclass of a base | `await self.get_action_by_base_class(Base)` | O(n) isinstance scan | +| Any LM provider | `await self.get_model_action(required=True)` | O(1) if `model_action_type` set, else falls back to base scan | +| The agent | `await self.get_agent()` | Cached | +| The App | `await self.get_app()` | Cached singleton | + +--- + +## 6. Tests + +- `tests/action/{name}/` per-action tests. +- `tests/action/test_action_loader.py` — plugin loading. +- `tests/action/test_action_endpoints.py` — endpoint discovery. +- `tests/action/test_plugin_system.py` — plugin contracts. +- `tests/test_tool_schema_audit.py` — tool schema sanity. + +```bash +pytest tests/action/ -v +``` + +--- + +## 7. Adding a new Action + +The detailed walkthrough lives at [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md) and the runbook at [`/.planning/runbooks/add-action.md`](../../.planning/runbooks/add-action.md). Short version: + +1. Pick base class (Action / InteractAction / specialized). +2. Choose namespace (`jvagent/` for core, otherwise `contrib/` or `custom/`). +3. Create the 4-file directory. +4. Define `attribute(...)` fields; implement lifecycle hooks + `execute` (if InteractAction). Publish tools by decorating methods with `@tool` (preferred over overriding `get_tools()`). +5. Wire endpoints under `/actions/{action_id}/...`. +6. Add tests under `tests/action/{name}/`. +7. Update [`actions-catalog.md`](../../.planning/reference/actions-catalog.md). + +--- + +## 8. Traps specific to action/ + +| Trap | Fix | +|---|---| +| Missing `from . import endpoints` in `__init__.py` | Add it; otherwise routes don't register | +| Class name doesn't match `archetype` in info.yaml | Loader silently skips the package | +| Heavy work in `__init__` | Use lifecycle hooks (`on_register`, `on_enable`) instead | +| Forgetting `@compound_index` when adding a queried field | Slow queries at scale | +| Writing to `self.metadata` | Not persisted; use `attribute(...)` | +| Endpoints not under `/actions/{action_id}/` | Deregister leaks them | +| Recursive `await self.get_action(MyAction)` calls | OK (cache returns same instance) but expensive isinstance walks aren't | + +--- + +## 9. Subdirectory pointers + +| Subdir | Local guide | +|---|---| +| `interact/` | [`interact/AGENTS.md`](interact/AGENTS.md) | +| `interview/` | [`interview/AGENTS.md`](interview/AGENTS.md) | +| `orchestrator/` | (see [`/.planning/adr/0012-skill-executive-architecture.md`](../../.planning/adr/0012-skill-executive-architecture.md)) | +| All other action dirs | Per-package `info.yaml` + class docstring | + +--- + +## 10. Out of scope here + +- Walker mechanics: see `interact/AGENTS.md`. +- Executive prompt/loop specifics: see [`/docs/ORCHESTRATOR.md`](../../docs/ORCHESTRATOR.md). +- Memory graph: see `jvagent/memory/AGENTS.md`. diff --git a/jvagent/action/CLAUDE.md b/jvagent/action/CLAUDE.md deleted file mode 100644 index e76cb13d..00000000 --- a/jvagent/action/CLAUDE.md +++ /dev/null @@ -1,145 +0,0 @@ -# jvagent/action/ — Agent Guide - -> Local guide for the action plugin layer. Cross-link: root [`/CLAUDE.md`](../../CLAUDE.md), [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md), [`/.planning/reference/actions-catalog.md`](../../.planning/reference/actions-catalog.md). - ---- - -## 1. What this directory owns - -The plugin-loadable extension surface of jvagent: - -- **`Action` base** ([`base.py:49`](base.py)) — Node subclass with lifecycle hooks, attribute config, endpoint registration, child-cascade delete, tool exposure to the Orchestrator's tool surface. -- **`InteractAction`** ([`interact/base.py:27`](interact/base.py)) — see `interact/CLAUDE.md`. -- **Specialized bases**: `BaseModelAction`, `LanguageModelAction`, `BaseWebSearchAction`, `BaseSTTAction`, `BaseTTSAction`, `VectorStore`. -- **Concrete plugins** organized by topic: language models, response/bus, Orchestrator, memory-related, channel adapters, productivity integrations, tasks. Catalog in [`/.planning/reference/actions-catalog.md`](../../.planning/reference/actions-catalog.md). -- **Loader/registry** in `loader/`. -- **Plugin contracts** in `plugin_contracts.py`. - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `base.py:49` | `Action` base class (canonical) | -| `base.py:259` | `get_tools()` — default discovers `@tool`-decorated methods via `collect_tools` | -| `jvagent/tooling/tool_decorator.py` | `@tool` decorator + `collect_tools` — preferred way to publish tools | -| `jvagent/tooling/signature_schema.py` | signature → JSON Schema deriver used by `@tool` | -| `base.py:180` | `get_capabilities()` — for ReplyAction prompt aggregation | -| `base.py:300` | `delete(cascade=True)` — walks outgoing edges and cascade-deletes children | -| `base.py:331-403` | Lifecycle hook contracts (`on_register`, `on_reload`, `post_register`, `on_startup`, `on_enable`, `on_disable`, `on_deregister`); `pulse`/`healthcheck` at `670`/`678` | -| `base.py:469-587` | Endpoint discovery + unregistration (relies on `/actions/{action_id}/` path prefix) | -| `base.py:588-667` | Module unload safety (skips core + shared modules) | -| `base.py:834-981` | `get_action()` / `get_action_by_base_class()` / `get_model_action()` — cross-action lookup | -| `base.py:982-1058` | Package metadata accessors (namespace, version, type) | -| `base.py:1145-1230` | File storage helpers (action-scoped paths) | -| `interact/base.py:27` | `InteractAction` (see `interact/CLAUDE.md`) | -| `actions.py` | `Actions` manager node | -| `endpoints.py` | Top-level action HTTP routes (~9 routes) | -| `loader/` | Action loader, registry, plugin discovery | -| `plugin_contracts.py` | Plugin protocol definitions | -| `streaming.py` | Streaming response helpers | - ---- - -## 3. Contracts (don't break) - -1. **`Action` subclasses MUST set `archetype` in `info.yaml`** to match the Python class name. The loader uses it. -2. **Action endpoints MUST live under `/actions/{action_id}/...`** ([`base.py:490`](base.py)). Deregister scans this prefix; non-conforming endpoints leak after `on_deregister`. -3. **`get_action()` is `O(1)`; `get_action_by_base_class()` is `O(n)`.** Don't use the latter in hot paths. -4. **Lifecycle hooks MUST not swallow exceptions** ([`base.py:694`](base.py)). The framework's `enable()`/`disable()`/`reload()` wrappers log errors automatically with the action context — silencing them hides bugs. -5. **`Action.metadata` is owned by the loader.** Mutations to it are not persisted across restarts. Use `attribute(...)` fields for persistent state. -6. **Child Nodes attached via outgoing edges are cascade-deleted** when the action is deleted ([`base.py:300`](base.py)). Always connect via `await self.connect(child, direction="out")`. -7. **`is_singleton` default is `True`** ([`base.py:296`](base.py)). Override `config.singleton: false` in `info.yaml` if multiple instances per agent are allowed. -8. **Thin harness** — Actions expose capabilities via `get_tools()`; they must not classify user intent, inject prep steering, auto-store extracted values, or inline multi-step workflows. Put judgment in skill SOPs and domain logic in skill extensions. See [`docs/thin-harness.md`](../../docs/thin-harness.md). -9. **Prefer the `@tool` decorator over hand-built `Tool()`** — decorate an `async def` method with `@tool` ([`jvagent/tooling/tool_decorator.py`](../tooling/tool_decorator.py)) and the base `get_tools()` auto-publishes it. Name = `{action_name}__{method}` (override with `@tool(name=...)`); description = method docstring's first paragraph; JSON Schema = signature (use `Annotated[T, "desc"]` for per-arg docs). Hand-built `Tool()` and `get_tools()` overrides still work; only override when a tool can't be a decorated method (combine with `collect_tools(self)`). - ---- - -## 4. The four-file pattern - -Every action package MUST contain: - -``` -{namespace}/{action_name}/ -├── __init__.py # exports class + imports endpoints (for @endpoint registration) -├── {action_name}.py # Action subclass -├── endpoints.py # @endpoint-decorated routes -└── info.yaml # package metadata -``` - -Skeleton and full templates: [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md). - ---- - -## 5. Cross-action lookup table - -| Need | Method | Cost | -|---|---|---| -| One specific class | `await self.get_action(MyActionClass)` | O(1) (cached index) | -| Class by name string | `await self.get_action("MyActionClass")` | O(1) | -| Any subclass of a base | `await self.get_action_by_base_class(Base)` | O(n) isinstance scan | -| Any LM provider | `await self.get_model_action(required=True)` | O(1) if `model_action_type` set, else falls back to base scan | -| The agent | `await self.get_agent()` | Cached | -| The App | `await self.get_app()` | Cached singleton | - ---- - -## 6. Tests - -- `tests/action/{name}/` per-action tests. -- `tests/action/test_action_loader.py` — plugin loading. -- `tests/action/test_action_endpoints.py` — endpoint discovery. -- `tests/action/test_plugin_system.py` — plugin contracts. -- `tests/test_tool_schema_audit.py` — tool schema sanity. - -```bash -pytest tests/action/ -v -``` - ---- - -## 7. Adding a new Action - -The detailed walkthrough lives at [`/.planning/reference/action-authoring.md`](../../.planning/reference/action-authoring.md) and the runbook at [`/.planning/runbooks/add-action.md`](../../.planning/runbooks/add-action.md). Short version: - -1. Pick base class (Action / InteractAction / specialized). -2. Choose namespace (`jvagent/` for core, otherwise `contrib/` or `custom/`). -3. Create the 4-file directory. -4. Define `attribute(...)` fields; implement lifecycle hooks + `execute` (if InteractAction). Publish tools by decorating methods with `@tool` (preferred over overriding `get_tools()`). -5. Wire endpoints under `/actions/{action_id}/...`. -6. Add tests under `tests/action/{name}/`. -7. Update [`actions-catalog.md`](../../.planning/reference/actions-catalog.md). - ---- - -## 8. Traps specific to action/ - -| Trap | Fix | -|---|---| -| Missing `from . import endpoints` in `__init__.py` | Add it; otherwise routes don't register | -| Class name doesn't match `archetype` in info.yaml | Loader silently skips the package | -| Heavy work in `__init__` | Use lifecycle hooks (`on_register`, `on_enable`) instead | -| Forgetting `@compound_index` when adding a queried field | Slow queries at scale | -| Writing to `self.metadata` | Not persisted; use `attribute(...)` | -| Endpoints not under `/actions/{action_id}/` | Deregister leaks them | -| Recursive `await self.get_action(MyAction)` calls | OK (cache returns same instance) but expensive isinstance walks aren't | - ---- - -## 9. Subdirectory pointers - -| Subdir | Local guide | -|---|---| -| `interact/` | [`interact/CLAUDE.md`](interact/CLAUDE.md) | -| `interview/` | [`interview/CLAUDE.md`](interview/CLAUDE.md) | -| `orchestrator/` | (see [`/.planning/adr/0012-skill-executive-architecture.md`](../../.planning/adr/0012-skill-executive-architecture.md)) | -| All other action dirs | Per-package `info.yaml` + class docstring | - ---- - -## 10. Out of scope here - -- Walker mechanics: see `interact/CLAUDE.md`. -- Executive prompt/loop specifics: see [`/docs/ORCHESTRATOR.md`](../../docs/ORCHESTRATOR.md). -- Memory graph: see `jvagent/memory/CLAUDE.md`. diff --git a/jvagent/action/base.py b/jvagent/action/base.py index 54ea68c0..74f3a227 100644 --- a/jvagent/action/base.py +++ b/jvagent/action/base.py @@ -25,6 +25,7 @@ from jvspatial.core import Node from jvspatial.core.annotations import attribute, compound_index +from pydantic import PrivateAttr if TYPE_CHECKING: from jvagent.action.manifest import Manifest @@ -131,6 +132,8 @@ class MyAction(Action): reachable via outgoing edges are cascade-deleted. """ + _property_override_keys: set[str] = PrivateAttr(default_factory=set) + # Core Attributes agent_id: str = attribute( indexed=True, diff --git a/jvagent/action/interact/AGENTS.md b/jvagent/action/interact/AGENTS.md index b4156a6e..b7407623 100644 --- a/jvagent/action/interact/AGENTS.md +++ b/jvagent/action/interact/AGENTS.md @@ -1 +1,139 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/action/interact/ — Agent Guide + +> Local guide for the interaction subsystem. Cross-link: [`/.planning/SPEC.md`](../../../.planning/SPEC.md) §3, [`/.planning/architecture.md`](../../../.planning/architecture.md) §3. + +--- + +## 1. What this directory owns + +The HTTP-facing interaction pipeline: + +- `InteractWalker` — the jvspatial Walker subclass that drives `/interact` traffic. +- `InteractAction` base — contract for executable actions in the pipeline. +- `endpoints.py` — `POST /agents/{id}/interact` and supporting routes. +- Walker payload bootstrap (User / Conversation / Interaction resolution). +- Background-action queueing and post-response execution. +- Access control enforcement per visit. + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `base.py:27` | `InteractAction` abstract base | +| `base.py:59` | `weight` attribute (top-tier ordering) | +| `base.py:71` | `always_execute` flag | +| `base.py:81` | `run_in_background` flag | +| `base.py:92` | `anchors: List[str]` for routing | +| `../base.py:170` | `parameters: List[Dict]` — behavioural params live on the `Action` base | +| `base.py:117-132` | `get_anchors(conversation)` — dynamic anchors override | +| `base.py:249-291` | `execute(visitor)` — abstract contract | +| `base.py:293-374` | `publish()` — direct write to response bus | +| `base.py:376-405` | `publish_thought()` — thought-category emit | +| `base.py:407-543` | `respond()` — generate via the agent egress responder (ReplyAction) | +| `interact_walker.py:36-45` | `InteractionInitResult` dataclass | +| `interact_walker.py:47-150` | `InteractWalker` core (state + properties) | +| `interact_walker.py:221` | `enforce_interact_action_access()` — access control | +| `interact_walker.py:267-414` | `_bootstrap_interaction()` — User / Conversation / Interaction resolve | +| `interact_walker.py:635` | `on_interact_action()` — per-visit callback | +| `response_builder.py:178` | `build_interact_response` (SSE helpers imported from `action/response/streaming`) | +| `webhook_pipeline.py:53` | `run_background_actions()` post-response runner (re-exported via `endpoints.py`) | +| `endpoints.py:208` | `/interact` endpoint decorator (handler `interact_endpoint` at :296) | + +--- + +## 3. Contracts (don't break) + +1. **`InteractAction.execute()` is the only entry point** the walker calls. Don't add side channels. +2. **Top-level actions order by `weight` ascending.** Sub-actions (connected to other InteractActions) do not — they run in graph order, only when the parent calls `await visitor.visit(child)`. +3. **`run_in_background=True` defers execution.** The walker queues the action; `_run_background_actions(walker)` fires it after response is sent. Each background action is isolated in try/except — failures must not propagate. +4. **`always_execute=True` bypasses routing exclusion** but does NOT bypass access control. The order is: access check → routing → execute. +5. **`execute()` is called inside a walker `visiting()` context** — `visitor.here` is set to the action node. Don't break that contract by mutating the walker queue from inside without using `visitor.visit()` / `visitor.prepend()`. +6. **`publish()` requires `visitor.response_bus` + `visitor.session_id`** ([`base.py:337-344`](base.py)). Early-return with a warning if either is missing. +7. **`publish()` with `stream=None` defaults to `visitor.stream`** ([`base.py:320`](base.py)). For non-streaming channels (WhatsApp), this is set False on the walker. +8. **Walker-revisit capability**: `visitor.prepend([self])` exists as a walker primitive — an action MAY enqueue itself (or another node) to be visited again. No shipped pattern currently relies on multi-visit turns; the Orchestrator runs its whole turn in a single `execute()` call (no revisit), carrying loop state locally inside `_run_loop`. + +--- + +## 4. Response emission decision tree + +``` +Generated text from a model? +└─ Use respond() — goes through ReplyAction (the single responder: gathers queued + directives + parameters, sources history, renders one reply in the Agent identity) + +Pre-built string (canned, system message, summary)? +├─ Visible to user → publish(content, stream=False) +└─ Internal trace (reasoning, plan)? → publish_thought(content, thought_type="reasoning") + +Need direct write without compose or history? +└─ publish() with explicit channel/metadata +``` + +--- + +## 5. Background action contract + +```python +class MyInteractAction(InteractAction): + run_in_background: bool = attribute(default=True) + + async def execute(self, visitor): + # This runs AFTER the user-facing response is sent. + # visitor.interaction is closed by the time we run. + # Failures here are caught — they don't impact the user response. + ... +``` + +Implementation lives in `webhook_pipeline.py:53` (`run_background_actions`). Each background action is wrapped: + +```python +try: + await action.execute(walker) +except Exception: + logger.error(...) +``` + +Use for: analytics, model updates, follow-up emails, scheduled task creation. + +--- + +## 6. Tests + +- `tests/action/interact/` — walker + bootstrap unit tests. +- `tests/action/gating/` — access control + always_execute tests. +- `tests/action/test_interact_walker.py` — end-to-end visit semantics. + +```bash +pytest tests/action/interact/ tests/action/gating/ -v +``` + +--- + +## 7. Traps specific to interact/ + +| Trap | Fix | +|---|---| +| Top-level `InteractAction` with children but no `visitor.visit(child)` call in `execute()` | Children never run. Explicitly route. | +| Calling `publish()` with `stream=True` on a non-streaming channel | Adapter mishandles. Pass `stream=False` or let visitor.stream propagate. | +| `await visitor.visit(self)` for re-visit | Cycle risk; walker may trip `max_visits_per_node=100`. If you must re-enqueue, use `visitor.prepend([self])` and persist state explicitly (no shipped pattern needs this — the Executive avoids re-visits entirely). | +| Setting `run_in_background=True` on an action that emits the user response | Response never reaches the client. Background = post-response only. | +| Long sleeps in `execute()` | Blocks the walker; latency spike. Use background or enqueue via `TaskMonitor` / `queue_task`. | +| Reading `visitor.interaction` in a background action | It's closed/saved by then — read-only, don't mutate. | +| Forgetting to call `await visitor.add_directives(...)` before `respond()` | Directives won't reach ReplyAction. | + +--- + +## 8. Don't touch from outside interact/ + +- `InteractWalker._bootstrap_interaction()` semantics — they're tightly coupled to memory/. +- The order of pre-visit access control checks — bypassing creates security holes. +- Background-action try/except wrapping — without it, one failure cascades. + +--- + +## 9. Out of scope here + +- Memory graph mutation: see `jvagent/memory/AGENTS.md`. +- Channel adapters: see `jvagent/action/response/`. diff --git a/jvagent/action/interact/CLAUDE.md b/jvagent/action/interact/CLAUDE.md deleted file mode 100644 index 30d09198..00000000 --- a/jvagent/action/interact/CLAUDE.md +++ /dev/null @@ -1,139 +0,0 @@ -# jvagent/action/interact/ — Agent Guide - -> Local guide for the interaction subsystem. Cross-link: [`/.planning/SPEC.md`](../../../.planning/SPEC.md) §3, [`/.planning/architecture.md`](../../../.planning/architecture.md) §3. - ---- - -## 1. What this directory owns - -The HTTP-facing interaction pipeline: - -- `InteractWalker` — the jvspatial Walker subclass that drives `/interact` traffic. -- `InteractAction` base — contract for executable actions in the pipeline. -- `endpoints.py` — `POST /agents/{id}/interact` and supporting routes. -- Walker payload bootstrap (User / Conversation / Interaction resolution). -- Background-action queueing and post-response execution. -- Access control enforcement per visit. - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `base.py:27` | `InteractAction` abstract base | -| `base.py:59` | `weight` attribute (top-tier ordering) | -| `base.py:71` | `always_execute` flag | -| `base.py:81` | `run_in_background` flag | -| `base.py:92` | `anchors: List[str]` for routing | -| `../base.py:170` | `parameters: List[Dict]` — behavioural params live on the `Action` base | -| `base.py:117-132` | `get_anchors(conversation)` — dynamic anchors override | -| `base.py:249-291` | `execute(visitor)` — abstract contract | -| `base.py:293-374` | `publish()` — direct write to response bus | -| `base.py:376-405` | `publish_thought()` — thought-category emit | -| `base.py:407-543` | `respond()` — generate via the agent egress responder (ReplyAction) | -| `interact_walker.py:36-45` | `InteractionInitResult` dataclass | -| `interact_walker.py:47-150` | `InteractWalker` core (state + properties) | -| `interact_walker.py:221` | `enforce_interact_action_access()` — access control | -| `interact_walker.py:267-414` | `_bootstrap_interaction()` — User / Conversation / Interaction resolve | -| `interact_walker.py:635` | `on_interact_action()` — per-visit callback | -| `response_builder.py:178` | `build_interact_response` (SSE helpers imported from `action/response/streaming`) | -| `webhook_pipeline.py:53` | `run_background_actions()` post-response runner (re-exported via `endpoints.py`) | -| `endpoints.py:208` | `/interact` endpoint decorator (handler `interact_endpoint` at :296) | - ---- - -## 3. Contracts (don't break) - -1. **`InteractAction.execute()` is the only entry point** the walker calls. Don't add side channels. -2. **Top-level actions order by `weight` ascending.** Sub-actions (connected to other InteractActions) do not — they run in graph order, only when the parent calls `await visitor.visit(child)`. -3. **`run_in_background=True` defers execution.** The walker queues the action; `_run_background_actions(walker)` fires it after response is sent. Each background action is isolated in try/except — failures must not propagate. -4. **`always_execute=True` bypasses routing exclusion** but does NOT bypass access control. The order is: access check → routing → execute. -5. **`execute()` is called inside a walker `visiting()` context** — `visitor.here` is set to the action node. Don't break that contract by mutating the walker queue from inside without using `visitor.visit()` / `visitor.prepend()`. -6. **`publish()` requires `visitor.response_bus` + `visitor.session_id`** ([`base.py:337-344`](base.py)). Early-return with a warning if either is missing. -7. **`publish()` with `stream=None` defaults to `visitor.stream`** ([`base.py:320`](base.py)). For non-streaming channels (WhatsApp), this is set False on the walker. -8. **Walker-revisit capability**: `visitor.prepend([self])` exists as a walker primitive — an action MAY enqueue itself (or another node) to be visited again. No shipped pattern currently relies on multi-visit turns; the Orchestrator runs its whole turn in a single `execute()` call (no revisit), carrying loop state locally inside `_run_loop`. - ---- - -## 4. Response emission decision tree - -``` -Generated text from a model? -└─ Use respond() — goes through ReplyAction (the single responder: gathers queued - directives + parameters, sources history, renders one reply in the Agent identity) - -Pre-built string (canned, system message, summary)? -├─ Visible to user → publish(content, stream=False) -└─ Internal trace (reasoning, plan)? → publish_thought(content, thought_type="reasoning") - -Need direct write without compose or history? -└─ publish() with explicit channel/metadata -``` - ---- - -## 5. Background action contract - -```python -class MyInteractAction(InteractAction): - run_in_background: bool = attribute(default=True) - - async def execute(self, visitor): - # This runs AFTER the user-facing response is sent. - # visitor.interaction is closed by the time we run. - # Failures here are caught — they don't impact the user response. - ... -``` - -Implementation lives in `webhook_pipeline.py:53` (`run_background_actions`). Each background action is wrapped: - -```python -try: - await action.execute(walker) -except Exception: - logger.error(...) -``` - -Use for: analytics, model updates, follow-up emails, scheduled task creation. - ---- - -## 6. Tests - -- `tests/action/interact/` — walker + bootstrap unit tests. -- `tests/action/gating/` — access control + always_execute tests. -- `tests/action/test_interact_walker.py` — end-to-end visit semantics. - -```bash -pytest tests/action/interact/ tests/action/gating/ -v -``` - ---- - -## 7. Traps specific to interact/ - -| Trap | Fix | -|---|---| -| Top-level `InteractAction` with children but no `visitor.visit(child)` call in `execute()` | Children never run. Explicitly route. | -| Calling `publish()` with `stream=True` on a non-streaming channel | Adapter mishandles. Pass `stream=False` or let visitor.stream propagate. | -| `await visitor.visit(self)` for re-visit | Cycle risk; walker may trip `max_visits_per_node=100`. If you must re-enqueue, use `visitor.prepend([self])` and persist state explicitly (no shipped pattern needs this — the Executive avoids re-visits entirely). | -| Setting `run_in_background=True` on an action that emits the user response | Response never reaches the client. Background = post-response only. | -| Long sleeps in `execute()` | Blocks the walker; latency spike. Use background or enqueue via `TaskMonitor` / `queue_task`. | -| Reading `visitor.interaction` in a background action | It's closed/saved by then — read-only, don't mutate. | -| Forgetting to call `await visitor.add_directives(...)` before `respond()` | Directives won't reach ReplyAction. | - ---- - -## 8. Don't touch from outside interact/ - -- `InteractWalker._bootstrap_interaction()` semantics — they're tightly coupled to memory/. -- The order of pre-visit access control checks — bypassing creates security holes. -- Background-action try/except wrapping — without it, one failure cascades. - ---- - -## 9. Out of scope here - -- Memory graph mutation: see `jvagent/memory/CLAUDE.md`. -- Channel adapters: see `jvagent/action/response/`. diff --git a/jvagent/action/interview/AGENTS.md b/jvagent/action/interview/AGENTS.md index b4156a6e..6e47a29c 100644 --- a/jvagent/action/interview/AGENTS.md +++ b/jvagent/action/interview/AGENTS.md @@ -1 +1,119 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# interview — Agent Guide + +--- + +## What this is + +`InterviewAction` is a pure `Action` (not `InteractAction`) that registers eight fixed `interview__*` tools plus per-skill custom tools. The **orchestrator LLM** reads each skill's `SKILL.md` procedure and drives multi-turn interviews by calling tools — the action manages session state, validation, hooks, and task tracking; it does **not** classify user intent or choose the next question itself. + +Session state lives in `conversation.context["interview"]` as a lightweight `InterviewSession` dataclass (field values, skipped fields, status, scratch `context` dict). + +**Design contract:** always build and extend interviews with the **[thin harness principle](../../../docs/thin-harness.md)** (platform) and **[interview profile](docs/thin-harness.md)** — thick SOP + skill extensions, thin server steering. + +--- + +## Foundation vs skill extensions + +**`interview/` is a reusable foundation.** It must stay domain-agnostic: no signup/training phrases, no per-skill field names, no business validators hardcoded in `interview_action.py`. + +| Layer | Location (per consuming app) | Owns | +|-------|------------------------------|------| +| **Foundation** | `jvagent/action/interview/` | `interview__*` tools, session lifecycle, hook dispatch, validator *invocation*, runtime-ready turn-lock gate, generic pipeline | +| **Base SOP** | `SKILL.md` (action root) | Inherited via `extends: action:jvagent/interview` | +| **Spec** | `skills//SKILL.md` frontmatter `interview:` | Fields, order, branches, validator names, pre/post processors, handlers, skill_tools | +| **Procedure** | `skills//SKILL.md` body | Custom behavioral rules only (base composed via `extends`) | +| **Implementation** | `skills//scripts/custom_tools.py` | Validators, pre/post tools, completion handlers, custom LLM tools | + +When fixing behavior for one skill (e.g. `validate_full_name`, training slot matching), change the **skill extension** — not the foundation — unless the bug is in generic plumbing (chaining, validator dispatch, hook dispatch, session persistence). + +**Terminal cleanup:** `complete`, `cancel`, and `interview_complete` validators call `clear_interview_context()` — wipes `conversation.context` except platform keys (`new_user`) and any `retain_context_keys` returned by the completion handler or validator. Do not persist interview scratch in `conversation.context` unless opting in via `retain_context_keys`. + +--- + +## File map + +``` +interview/ +├── SKILL.md # Base SOP (extends target) +├── interview_action.py # Action shell: discovery, turn-lock hooks, skill activation +├── spec.py # Frontmatter parsing: FieldDef / ForEachDef / InterviewSpec / registry +├── for_each.py # Per-item subpart expansion state + iteration helpers +├── session.py # InterviewSession + conversation persistence +├── flow.py # Branch evaluation, path walk, prune +├── hooks.py # custom_tools.py loader; the ctx interface (HookExecutionContext); +│ # call_hook; validator dispatch; internal directive framing +├── directive_compose.py # Internal: merge hook directives into tool-response envelopes +├── validators.py # Built-in validators +├── engine.py # The 8 tool handlers + activation + skill-tool dispatch +├── tools.py # Tool definitions binding to engine +├── tasks.py # interview SKILL-task lifecycle +├── procedure.py # SOP composition +├── _validate_contract.py # Skill frontmatter ↔ custom_tools.py validation +├── info.yaml +├── README.md / AGENTS.md +├── examples/ # Reference skill packages (not auto-discovered) +└── docs/ # How-to guides + skill_custom_instructions.md +``` + +--- + +## Creating a new interview skill (minimum steps) + +1. Copy [`examples/example_interview/`](examples/example_interview/) → `agents///skills//`. For **per-item subparts**, also read [`examples/example_for_each_interview/`](examples/example_for_each_interview/). +2. Align `name` in folder and `SKILL.md` frontmatter. +3. Implement every `function:` referenced in frontmatter `interview:` inside `scripts/custom_tools.py`. +4. Write `SKILL.md` custom instructions only; set `extends: action:jvagent/interview` (see `docs/skill_custom_instructions.md`). +5. Set `extends: action:jvagent/interview` and `requires-actions: [InterviewAction]`. Add custom LLM tools to additive `allowed-tools` only. +6. Register skill in agent `orchestrator.skills:`. +7. Enable `InterviewAction` in agent actions. + +See [README.md](README.md) and [docs/extending.md](docs/extending.md) for validators, hooks, review/reset/completion handlers. + +--- + +## Key invariants + +Full tables: **[interview profile](docs/thin-harness.md)** (+ [platform](../../../docs/thin-harness.md)). Summary: + +1. **Thin harness** — no server intent classification, no prep observations, no activation auto-store, no merge-inlined next/review responses, no `extractors` in frontmatter. +2. **Hook functions are not LLM tools** — only entries in frontmatter `interview.skill_tools` become `{skill}__{name}` tools. Reset uses `handlers.reset` (invoked via `interview__reset()`). +3. **Model owns extraction and chaining** — `interview__set_fields` + base SOP; read `ok` before advancing; post-processors do not run when `ok: false`. +4. **`response_directive` beats `next_field`** when they conflict — one action per turn. +5. **Review before complete** — always call `interview__review()` before `interview__complete()` unless review sets `terminate: true` or `confirm: auto` chains complete. +6. **Contract discovery** — `InterviewRegistry` scans dirs from `Action.resolve_skill_scan_dirs()` (app `skills/` + action-bundled paths). Author interview skills under `agents/.../skills//` (ADR-0023). Reference packages live under `examples/` (not discovered). +7. **Never reuse stale field values** from older chat turns unless the user repeats them in the latest message. +8. **Domain logic in skills** — validators, processors, handlers, and branching in `custom_tools.py` + `interview:` frontmatter; never in foundation code. + +### `for_each` subpart invariants + +9. **`ctx.field_def.key` not literals in processors.** Post-processors and validators receive `ctx.field_def` — use `ctx.field_def.key` to identify the current field instead of a bare string. Hard-coding `"my_field_name"` in a post_processor couples the hook to one field name and breaks silently on rename. +10. **`ctx.value` is `None` in post_processors.** The value was stored before the hook runs. Read it with `ctx.session.get_value(ctx.field_def.key)`. +11. **`ctx.get_for_each_records(parent_key)` in handlers.** Review and complete handlers must call `ctx.get_for_each_records("parent_key")` (or `ctx.get_for_each_records(some_key_variable)`) instead of `ctx.session.context["for_each"]["parent_key"]["records"]`. The internal path is framework-private. `ctx.get_for_each_records()` returns `[]` gracefully on skip. +12. **Wipe-before-validate protection.** The engine wipes existing for_each expansion ONLY after the new parent value passes validation. A failed re-submission leaves the old expansion intact so the children remain reachable. Do not rely on for_each state being wiped on every parent re-submission. +13. **Item dict shape.** `ctx.expand_for_each(items=[...])` items are plain Python values or dicts with optional `"id"` and `"label"` keys. Primitive values (`str`, `int`) are used as both id and label. Dicts that do not have `"id"` fall back to the list index as id — prefer explicit `{"id": ..., "label": ...}` for readability. + +--- + +## Tests + +```bash +pytest tests/action/interview/ -v +``` + +--- + +## Read next + +| Doc | Topic | +|-----|-------| +| [docs/thin-harness.md](../../../docs/thin-harness.md) | **Thin harness** (jvagent-wide principle) | +| [docs/thin-harness.md](docs/thin-harness.md) | **Interview profile** (subsystem invariants) | +| [docs/frontmatter-schema.md](docs/frontmatter-schema.md) | Canonical `interview:` YAML schema | +| [README.md](README.md) | Reading paths, tool envelope, live skill patterns | +| [docs/multi-turn-flow.md](docs/multi-turn-flow.md) | Turn-by-turn lifecycle, turn-lock, session states | +| [docs/extending.md](docs/extending.md) | Validators, processors, handlers, skill tools, **`for_each`** | +| [docs/troubleshooting.md](docs/troubleshooting.md) | Common failures and fixes | +| [examples/example_interview/](examples/example_interview/) | Reference implementation | +| [examples/example_for_each_interview/](examples/example_for_each_interview/) | **`for_each` subparts** reference | +| [CUCS witness scenarios](examples/example_account_gating/use-cases/) | Domain-neutral conversation use cases | +| [`.planning/reference/conversation-use-cases.md`](../../../.planning/reference/conversation-use-cases.md) | Conversation Use Case Specification | diff --git a/jvagent/action/interview/CLAUDE.md b/jvagent/action/interview/CLAUDE.md deleted file mode 100644 index 0513db06..00000000 --- a/jvagent/action/interview/CLAUDE.md +++ /dev/null @@ -1,119 +0,0 @@ -# interview — Agent Guide - ---- - -## What this is - -`InterviewAction` is a pure `Action` (not `InteractAction`) that registers eight fixed `interview__*` tools plus per-skill custom tools. The **orchestrator LLM** reads each skill's `SKILL.md` procedure and drives multi-turn interviews by calling tools — the action manages session state, validation, hooks, and task tracking; it does **not** classify user intent or choose the next question itself. - -Session state lives in `conversation.context["interview"]` as a lightweight `InterviewSession` dataclass (field values, skipped fields, status, scratch `context` dict). - -**Design contract:** always build and extend interviews with the **[thin harness principle](../../../docs/thin-harness.md)** (platform) and **[interview profile](docs/thin-harness.md)** — thick SOP + skill extensions, thin server steering. - ---- - -## Foundation vs skill extensions - -**`interview/` is a reusable foundation.** It must stay domain-agnostic: no signup/training phrases, no per-skill field names, no business validators hardcoded in `interview_action.py`. - -| Layer | Location (per consuming app) | Owns | -|-------|------------------------------|------| -| **Foundation** | `jvagent/action/interview/` | `interview__*` tools, session lifecycle, hook dispatch, validator *invocation*, runtime-ready turn-lock gate, generic pipeline | -| **Base SOP** | `SKILL.md` (action root) | Inherited via `extends: action:jvagent/interview` | -| **Spec** | `skills//SKILL.md` frontmatter `interview:` | Fields, order, branches, validator names, pre/post processors, handlers, skill_tools | -| **Procedure** | `skills//SKILL.md` body | Custom behavioral rules only (base composed via `extends`) | -| **Implementation** | `skills//scripts/custom_tools.py` | Validators, pre/post tools, completion handlers, custom LLM tools | - -When fixing behavior for one skill (e.g. `validate_full_name`, training slot matching), change the **skill extension** — not the foundation — unless the bug is in generic plumbing (chaining, validator dispatch, hook dispatch, session persistence). - -**Terminal cleanup:** `complete`, `cancel`, and `interview_complete` validators call `clear_interview_context()` — wipes `conversation.context` except platform keys (`new_user`) and any `retain_context_keys` returned by the completion handler or validator. Do not persist interview scratch in `conversation.context` unless opting in via `retain_context_keys`. - ---- - -## File map - -``` -interview/ -├── SKILL.md # Base SOP (extends target) -├── interview_action.py # Action shell: discovery, turn-lock hooks, skill activation -├── spec.py # Frontmatter parsing: FieldDef / ForEachDef / InterviewSpec / registry -├── for_each.py # Per-item subpart expansion state + iteration helpers -├── session.py # InterviewSession + conversation persistence -├── flow.py # Branch evaluation, path walk, prune -├── hooks.py # custom_tools.py loader; the ctx interface (HookExecutionContext); -│ # call_hook; validator dispatch; internal directive framing -├── directive_compose.py # Internal: merge hook directives into tool-response envelopes -├── validators.py # Built-in validators -├── engine.py # The 8 tool handlers + activation + skill-tool dispatch -├── tools.py # Tool definitions binding to engine -├── tasks.py # interview SKILL-task lifecycle -├── procedure.py # SOP composition -├── _validate_contract.py # Skill frontmatter ↔ custom_tools.py validation -├── info.yaml -├── README.md / CLAUDE.md / AGENTS.md -├── examples/ # Reference skill packages (not auto-discovered) -└── docs/ # How-to guides + skill_custom_instructions.md -``` - ---- - -## Creating a new interview skill (minimum steps) - -1. Copy [`examples/example_interview/`](examples/example_interview/) → `agents///skills//`. For **per-item subparts**, also read [`examples/example_for_each_interview/`](examples/example_for_each_interview/). -2. Align `name` in folder and `SKILL.md` frontmatter. -3. Implement every `function:` referenced in frontmatter `interview:` inside `scripts/custom_tools.py`. -4. Write `SKILL.md` custom instructions only; set `extends: action:jvagent/interview` (see `docs/skill_custom_instructions.md`). -5. Set `extends: action:jvagent/interview` and `requires-actions: [InterviewAction]`. Add custom LLM tools to additive `allowed-tools` only. -6. Register skill in agent `orchestrator.skills:`. -7. Enable `InterviewAction` in agent actions. - -See [README.md](README.md) and [docs/extending.md](docs/extending.md) for validators, hooks, review/reset/completion handlers. - ---- - -## Key invariants - -Full tables: **[interview profile](docs/thin-harness.md)** (+ [platform](../../../docs/thin-harness.md)). Summary: - -1. **Thin harness** — no server intent classification, no prep observations, no activation auto-store, no merge-inlined next/review responses, no `extractors` in frontmatter. -2. **Hook functions are not LLM tools** — only entries in frontmatter `interview.skill_tools` become `{skill}__{name}` tools. Reset uses `handlers.reset` (invoked via `interview__reset()`). -3. **Model owns extraction and chaining** — `interview__set_fields` + base SOP; read `ok` before advancing; post-processors do not run when `ok: false`. -4. **`response_directive` beats `next_field`** when they conflict — one action per turn. -5. **Review before complete** — always call `interview__review()` before `interview__complete()` unless review sets `terminate: true` or `confirm: auto` chains complete. -6. **Contract discovery** — `InterviewRegistry` scans dirs from `Action.resolve_skill_scan_dirs()` (app `skills/` + action-bundled paths). Author interview skills under `agents/.../skills//` (ADR-0023). Reference packages live under `examples/` (not discovered). -7. **Never reuse stale field values** from older chat turns unless the user repeats them in the latest message. -8. **Domain logic in skills** — validators, processors, handlers, and branching in `custom_tools.py` + `interview:` frontmatter; never in foundation code. - -### `for_each` subpart invariants - -9. **`ctx.field_def.key` not literals in processors.** Post-processors and validators receive `ctx.field_def` — use `ctx.field_def.key` to identify the current field instead of a bare string. Hard-coding `"my_field_name"` in a post_processor couples the hook to one field name and breaks silently on rename. -10. **`ctx.value` is `None` in post_processors.** The value was stored before the hook runs. Read it with `ctx.session.get_value(ctx.field_def.key)`. -11. **`ctx.get_for_each_records(parent_key)` in handlers.** Review and complete handlers must call `ctx.get_for_each_records("parent_key")` (or `ctx.get_for_each_records(some_key_variable)`) instead of `ctx.session.context["for_each"]["parent_key"]["records"]`. The internal path is framework-private. `ctx.get_for_each_records()` returns `[]` gracefully on skip. -12. **Wipe-before-validate protection.** The engine wipes existing for_each expansion ONLY after the new parent value passes validation. A failed re-submission leaves the old expansion intact so the children remain reachable. Do not rely on for_each state being wiped on every parent re-submission. -13. **Item dict shape.** `ctx.expand_for_each(items=[...])` items are plain Python values or dicts with optional `"id"` and `"label"` keys. Primitive values (`str`, `int`) are used as both id and label. Dicts that do not have `"id"` fall back to the list index as id — prefer explicit `{"id": ..., "label": ...}` for readability. - ---- - -## Tests - -```bash -pytest tests/action/interview/ -v -``` - ---- - -## Read next - -| Doc | Topic | -|-----|-------| -| [docs/thin-harness.md](../../../docs/thin-harness.md) | **Thin harness** (jvagent-wide principle) | -| [docs/thin-harness.md](docs/thin-harness.md) | **Interview profile** (subsystem invariants) | -| [docs/frontmatter-schema.md](docs/frontmatter-schema.md) | Canonical `interview:` YAML schema | -| [README.md](README.md) | Reading paths, tool envelope, live skill patterns | -| [docs/multi-turn-flow.md](docs/multi-turn-flow.md) | Turn-by-turn lifecycle, turn-lock, session states | -| [docs/extending.md](docs/extending.md) | Validators, processors, handlers, skill tools, **`for_each`** | -| [docs/troubleshooting.md](docs/troubleshooting.md) | Common failures and fixes | -| [examples/example_interview/](examples/example_interview/) | Reference implementation | -| [examples/example_for_each_interview/](examples/example_for_each_interview/) | **`for_each` subparts** reference | -| [CUCS witness scenarios](examples/example_account_gating/use-cases/) | Domain-neutral conversation use cases | -| [`.planning/reference/conversation-use-cases.md`](../../../.planning/reference/conversation-use-cases.md) | Conversation Use Case Specification | diff --git a/jvagent/action/interview/README.md b/jvagent/action/interview/README.md index c4865b0c..8e5edc1c 100644 --- a/jvagent/action/interview/README.md +++ b/jvagent/action/interview/README.md @@ -4,7 +4,7 @@ LLM-driven interview framework for structured data collection. The orchestrator **Custom interview skills** are two-file packages under `agents///skills//` ([ADR-0023 placement standard](../../.planning/adr/0023-skill-placement-standard.md)). Copy [`examples/example_interview/`](examples/example_interview/) as a template, set `extends: action:jvagent/interview`, `requires-actions: [InterviewAction]`, and `task-lock: true` for turn-lock. -**Agent entry point:** [CLAUDE.md](CLAUDE.md) +**Agent entry point:** [AGENTS.md](AGENTS.md) ## Documentation @@ -14,7 +14,7 @@ LLM-driven interview framework for structured data collection. The orchestrator |----------|------------| | **Building a new interview skill** | [Quick start](#quick-start) → [docs/extending.md](docs/extending.md) → [examples/example_interview/](examples/example_interview/) | | **Per-item subpart questions (`for_each`)** | [docs/frontmatter-schema.md](docs/frontmatter-schema.md#per-item-subparts-fieldsfor_each) → [docs/extending.md](docs/extending.md#per-item-subparts-for_each) → [examples/example_for_each_interview/](examples/example_for_each_interview/) | -| **AI agent editing this package** | [CLAUDE.md](CLAUDE.md) | +| **AI agent editing this package** | [AGENTS.md](AGENTS.md) | | **Debugging a stuck turn** | [docs/troubleshooting.md](docs/troubleshooting.md) → [docs/multi-turn-flow.md](docs/multi-turn-flow.md) | | **Authoring `SKILL.md` body only** | [docs/skill_custom_instructions.md](docs/skill_custom_instructions.md) (base SOP: [SKILL.md](SKILL.md)) | diff --git a/jvagent/action/interview/__init__.py b/jvagent/action/interview/__init__.py index 547be2c8..1b5cfa16 100644 --- a/jvagent/action/interview/__init__.py +++ b/jvagent/action/interview/__init__.py @@ -6,7 +6,7 @@ package has no ``skills/`` subdir. Reference templates are under ``examples/`` (not discovered). Declare ``extends: action:jvagent/interview``. -Documentation: ``README.md``, ``CLAUDE.md``, ``docs/``. +Documentation: ``README.md``, ``AGENTS.md``, ``docs/``. """ from .interview_action import InterviewAction diff --git a/jvagent/action/interview/interview_action.py b/jvagent/action/interview/interview_action.py index 17819e42..49961a20 100644 --- a/jvagent/action/interview/interview_action.py +++ b/jvagent/action/interview/interview_action.py @@ -7,6 +7,7 @@ from typing import Any, Callable, Dict, List, Optional, Tuple from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.base import Action from jvagent.action.parameters import SCOPE_ORCHESTRATION, SCOPE_RESPONSE @@ -111,9 +112,7 @@ class InterviewAction(Action): def get_capabilities(self) -> List[str]: return [self.description] if self.description else [] - def __init__(self, **kwargs): - super().__init__(**kwargs) - self._registry = InterviewRegistry() + _registry: InterviewRegistry = PrivateAttr(default_factory=InterviewRegistry) # -- discovery ---------------------------------------------------------- diff --git a/jvagent/action/leadgen/leadgen_action.py b/jvagent/action/leadgen/leadgen_action.py index ccb85474..6b4aac4e 100644 --- a/jvagent/action/leadgen/leadgen_action.py +++ b/jvagent/action/leadgen/leadgen_action.py @@ -7,6 +7,7 @@ from typing import Any, Dict, List from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.base import Action @@ -64,9 +65,7 @@ class LeadGenAction(Action): ), ) - def __init__(self, **kwargs): - super().__init__(**kwargs) - self._registry = LeadGenRegistry() + _registry: LeadGenRegistry = PrivateAttr(default_factory=LeadGenRegistry) async def on_register(self): await super().on_register() diff --git a/jvagent/action/mcp/mcp_action.py b/jvagent/action/mcp/mcp_action.py index 9fba2f28..fbec5ec4 100644 --- a/jvagent/action/mcp/mcp_action.py +++ b/jvagent/action/mcp/mcp_action.py @@ -8,6 +8,7 @@ from typing import Any, Dict, List, Optional, Tuple from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.base import Action from jvagent.action.mcp.client import MCPClientWrapper @@ -309,9 +310,7 @@ class MCPAction(Action): description="Optional override for sandbox root; defaults to jvspatial files root (JVSPATIAL_FILES_ROOT_PATH).", ) - def __init__(self, **kwargs: Any) -> None: - super().__init__(**kwargs) - self._servers_by_name: Dict[str, _ServerEntry] = {} + _servers_by_name: Dict[str, _ServerEntry] = PrivateAttr(default_factory=dict) def _strip_trailing_path_arg(self, args: List[str]) -> List[str]: """Delegate to :func:`strip_trailing_path_arg` (keeps e.g. ``@scope/pkg``).""" diff --git a/jvagent/action/model/base.py b/jvagent/action/model/base.py index 69c3c9df..1cdad530 100644 --- a/jvagent/action/model/base.py +++ b/jvagent/action/model/base.py @@ -15,6 +15,7 @@ import httpx from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.base import Action @@ -107,6 +108,9 @@ class BaseModelAction(Action, ABC): total_duration: Cumulative query duration in seconds """ + _deadline_coherence_warned: bool = PrivateAttr(default=False) + _last_result: Any = PrivateAttr(default=None) + # Common configuration attributes. API credentials are resolved # exclusively from environment variables via ``api_key_from_context()``; # the legacy ``api_key`` attribute is no longer accepted on Model actions. diff --git a/jvagent/action/model/embedding/base.py b/jvagent/action/model/embedding/base.py index f5c05ed2..097d70f9 100644 --- a/jvagent/action/model/embedding/base.py +++ b/jvagent/action/model/embedding/base.py @@ -8,6 +8,7 @@ from typing import List, Optional from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.model.base import BaseModelAction @@ -31,6 +32,9 @@ class EmbeddingModelAction(BaseModelAction, ABC): >>> print(f"Embedding dimensions: {len(vector)}") """ + _calling_action_name: Optional[str] = PrivateAttr(default=None) + _recorded_usage_tokens: int = PrivateAttr(default=0) + embedding_dimensions: int = attribute( default=0, description="Expected embedding dimensions (0 = auto-detect)", ge=0 ) diff --git a/jvagent/action/orchestrator/orchestrator_interact_action.py b/jvagent/action/orchestrator/orchestrator_interact_action.py index 811244ab..d5e10230 100644 --- a/jvagent/action/orchestrator/orchestrator_interact_action.py +++ b/jvagent/action/orchestrator/orchestrator_interact_action.py @@ -42,6 +42,7 @@ ) from jvspatial.core.annotations import attribute +from pydantic import PrivateAttr from jvagent.action.interact.base import InteractAction from jvagent.action.interact.utils.uploads import DEFAULT_UPLOAD_KEYS @@ -294,6 +295,9 @@ class OrchestratorInteractAction( ): """The sole pattern orchestrator (ADR-0012), weight ``-200``.""" + _actions_enum_failed: bool = PrivateAttr(default=False) + _max_tokens_fallback_warned: bool = PrivateAttr(default=False) + weight: int = attribute( default=-200, description="Pattern-orchestrator slot at -200 (the sole turn orchestrator).", diff --git a/jvagent/action/pageindex/document_walker.py b/jvagent/action/pageindex/document_walker.py index f9d8c3ef..3824d558 100644 --- a/jvagent/action/pageindex/document_walker.py +++ b/jvagent/action/pageindex/document_walker.py @@ -9,6 +9,7 @@ from typing import Any, Dict, List, Optional from jvspatial.core import Walker, on_visit +from pydantic import PrivateAttr from .models import ( DocumentContentEdge, @@ -33,6 +34,13 @@ class DocumentWalker(Walker): Stops traversal when report size reaches limit (early termination). """ + _query: str = PrivateAttr(default="") + _query_lower: str = PrivateAttr(default="") + _query_regex: Optional[re.Pattern] = PrivateAttr(default=None) + _limit: Optional[int] = PrivateAttr(default=None) + _only_enabled: bool = PrivateAttr(default=True) + _include: Optional[List[str]] = PrivateAttr(default=None) + def __init__( self, query: str = "", diff --git a/jvagent/action/stt_action/deepgram/deepgram.py b/jvagent/action/stt_action/deepgram/deepgram.py index 87ff2225..d2234f13 100644 --- a/jvagent/action/stt_action/deepgram/deepgram.py +++ b/jvagent/action/stt_action/deepgram/deepgram.py @@ -9,6 +9,7 @@ from deepgram.core.api_error import ApiError from jvspatial.core.annotations import attribute from jvspatial.env import env +from pydantic import PrivateAttr from jvagent.action.stt_action.base import BaseSTTAction @@ -36,6 +37,9 @@ class DeepgramSTTAction(BaseSTTAction): """Speech-to-text action using the Deepgram API.""" + _deepgram_client: Optional[AsyncDeepgramClient] = PrivateAttr(default=None) + _deepgram_client_key: Optional[str] = PrivateAttr(default=None) + model: str = attribute( default="nova-2", description="Model to use (enhanced, nova, base, nova-2)", diff --git a/jvagent/action/whatsapp_voice/whatsapp_voice_action.py b/jvagent/action/whatsapp_voice/whatsapp_voice_action.py index ff22b4a9..e29a516c 100644 --- a/jvagent/action/whatsapp_voice/whatsapp_voice_action.py +++ b/jvagent/action/whatsapp_voice/whatsapp_voice_action.py @@ -7,6 +7,7 @@ from jvspatial.core.annotations import attribute from jvspatial.env import env +from pydantic import PrivateAttr from jvagent.action.base import Action from jvagent.action.interact.webhook_pipeline import get_conversation_with_lock @@ -30,6 +31,9 @@ class WhatsAppVoiceAction(Action): to handle realtime audio and bridge utterances to the jvagent Orchestrator. """ + _active_calls: Dict[str, str] = PrivateAttr(default_factory=dict) + _jvvoice: Optional[JvvoiceClient] = PrivateAttr(default=None) + jvvoice_base_url: str = attribute( default="", description="jvvoice connector API base URL; when empty, JVVOICE_BASE_URL env is used", diff --git a/jvagent/cli/AGENTS.md b/jvagent/cli/AGENTS.md index b4156a6e..69c0b93b 100644 --- a/jvagent/cli/AGENTS.md +++ b/jvagent/cli/AGENTS.md @@ -1 +1,143 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/cli/ — Agent Guide + +> Local guide for the CLI and server bootstrap. Cross-link: root [`/AGENTS.md`](../../AGENTS.md), [`/.planning/runbooks/local-dev.md`](../../.planning/runbooks/local-dev.md), [`/docs/scaffolding.md`](../../docs/scaffolding.md). + +--- + +## 1. What this directory owns + +- The `jvagent` and `python -m jvagent` entry points. +- Argument parsing, app-root extraction, flag handling. +- Subcommand dispatch: `status`, `agent`, `action`, `skill`, `bootstrap`, `bundle`, `app`, `validate`, `stress-seed`. +- Server bootstrap from `app.yaml` (`create_server_from_config`). +- Graph bootstrap (`bootstrap_application_graph`). + +It does **not** own: the actual graph node definitions (that's `core/`), HTTP server runtime (that's jvspatial). + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `__init__.py` | Exports `main` | +| `main.py:118-244` | `main()` — top-level entry; parses args, dispatches | +| `main.py:58-115` | `_first_app_root_path()` — extracts app root from arg list | +| `main.py:142-149` | `--serverless` flag handling | +| `main.py:152-159` | `--debug` flag | +| `main.py:161-179` | `--update` / `--source` / `--merge` parsing (mutually exclusive rules) | +| `main.py:181-192` | `--purge` (dev-only) | +| `main.py:205-226` | Subcommand dispatch | +| `main.py:239-244` | Default → `run_server()` | +| `commands.py` | Backward-compat re-export facade (re-exports the handlers below) | +| `server.py` | `load_app_env`, `purge_app_data`, `run_server`, `show_status`, `bootstrap_only`, `StartupLogCounter` | +| `agent_commands.py` | `handle_agent_command`, `handle_action_command`, `handle_bundle_command`, `list_agents`, `uninstall_agent` | +| `skill_commands.py` | `handle_skill_command` | +| `validate.py` | `run_validate`, `print_usage` | +| `chat.py` | `handle_chat_command` (serves the jvchat UI) | +| `server_config.py:59-180` | `create_server_from_config()` — reads app.yaml + env, builds jvspatial Server | +| `server_config.py` | `_set_db_env_from_config`, `_import_core_endpoint_modules` | +| `bootstrap.py` | `bootstrap_application_graph()` — orchestrates App + Agents + Actions creation | +| `app_commands.py` | `jvagent app create / profile new / ...` subcommands | +| `__main__.py:5-6` | `python -m jvagent` → `cli.main:main()` | + +--- + +## 3. CLI shape (memorize) + +``` +jvagent [app_root] [SUBCOMMAND] [FLAGS] + +DEFAULT (no subcommand) → start HTTP server + +SUBCOMMANDS: + status diagnostic snapshot of an app's graph + agent ... agent CRUD (add, list, enable, disable, delete) + action ... action CRUD on an agent + skill ... skill management + bootstrap bootstrap graph then exit (no server) + bundle emit Dockerfile + deployment artifacts + app create|profile scaffold a new app or profile + validate validate app.yaml + agents + chat serve the bundled jvchat UI on its own port (jvagent/webui) + stress-seed generate synthetic graph for testing + +FLAGS (apply to default + bootstrap): + --debug verbose logging + --update merge-mode YAML sync + --update --source destructive YAML sync (DESTRUCTIVE) + --update --merge explicit merge (== --update alone) + --purge wipe DB (dev only; requires JVSPATIAL_ENVIRONMENT=development) + --serverless single-worker, SERVERLESS_MODE=true +``` + +`--source` and `--merge` REQUIRE `--update`. They are mutually exclusive ([`main.py:167-172`](main.py)). + +--- + +## 4. Contracts (don't break) + +1. **App root resolution** ([`main.py:58-115`](main.py)) — strips path tokens, keeps subcommands + flags + stress-seed args. Don't break this for `jvagent ./myapp agent add ...`. +2. **`--purge` MUST require dev mode** ([`main.py:185-190`](main.py)). Never relax. +3. **`bootstrap_only` and `run_server` reset `App.update_mode` to `run`** after a successful sync. Don't remove this; otherwise one-shot ops persist across restarts. +4. **`load_app_env(app_root)` runs first** ([`main.py:130`](main.py)). Anything reading env vars before this is incorrect. +5. **`set_app_root(app_root)` must run before any node lookup** ([`main.py:134`](main.py)). Cache/config keys depend on it. +6. **`_set_db_env_from_config(app_root)` ([`main.py:150`](main.py))** translates app.yaml DB stanza into env vars jvspatial reads. Must precede any DB call. + +--- + +## 5. Subcommand handler conventions + +Each `handle_*_command` (in `agent_commands.py` / `skill_commands.py`): + +- Receives `args` (post-subcommand) and `app_root`. +- Uses `asyncio.run(...)` for async operations. +- Exits with non-zero on errors via `sys.exit(N)`. +- Prints user-facing output to stdout/stderr (no return values). + +When adding a subcommand: + +1. Add the name to `DISPATCH` ([`main.py:42-55`](main.py)). +2. Add a dispatch branch in the `if args[0] in DISPATCH:` block ([`main.py:205-226`](main.py)). +3. Add the handler in the relevant sibling module (`agent_commands.py`, `server.py`, `skill_commands.py`, `validate.py`); re-export it from `commands.py` for back-compat. +4. Document in [`/.planning/runbooks/`](../../.planning/runbooks/) if non-trivial. + +--- + +## 6. Tests + +- `tests/cli/` — argparse/dispatch tests. +- `tests/test_env_load.py` — `load_app_env` precedence. +- `tests/scaffold/` — `app create` flow. + +```bash +pytest tests/cli/ tests/test_env_load.py -v +``` + +--- + +## 7. Traps specific to cli/ + +| Trap | Fix | +|---|---| +| Treating a `STRESS_FLAG_NAMES` value as a path | `_first_app_root_path()` handles this — don't bypass it | +| Forgetting to strip a flag from `args` after parsing | The default-server path will see it and error "Unknown argument" | +| Calling jvspatial before `set_app_root()` + `load_app_env()` | Wrong DB / paths | +| Running `--update --source` with no app root | App root defaults to `cwd`; destructive op on the wrong dir. Always pass explicit path for source mode. | +| Mixing `--purge` with `--source` | `--purge` deletes the DB, then `--source` rebuilds — works, but slow. Use `--update --source` alone to overwrite. | + +--- + +## 8. Don't touch from outside cli/ + +- The argument-parsing semantics — many runbooks and docs encode them. +- The dispatch table — subcommand name is part of the public CLI contract. +- The `update_mode` reset behavior — its absence makes cold starts unpredictable. + +--- + +## 9. Out of scope here + +- Server runtime (uvicorn + FastAPI internals): jvspatial. +- Endpoint registration: handled by jvspatial server based on imported endpoint modules. +- Graph repair: see `jvagent/core/graph_repair*.py` and `core/AGENTS.md`. diff --git a/jvagent/cli/CLAUDE.md b/jvagent/cli/CLAUDE.md deleted file mode 100644 index 1e5de722..00000000 --- a/jvagent/cli/CLAUDE.md +++ /dev/null @@ -1,143 +0,0 @@ -# jvagent/cli/ — Agent Guide - -> Local guide for the CLI and server bootstrap. Cross-link: root [`/CLAUDE.md`](../../CLAUDE.md), [`/.planning/runbooks/local-dev.md`](../../.planning/runbooks/local-dev.md), [`/docs/scaffolding.md`](../../docs/scaffolding.md). - ---- - -## 1. What this directory owns - -- The `jvagent` and `python -m jvagent` entry points. -- Argument parsing, app-root extraction, flag handling. -- Subcommand dispatch: `status`, `agent`, `action`, `skill`, `bootstrap`, `bundle`, `app`, `validate`, `stress-seed`. -- Server bootstrap from `app.yaml` (`create_server_from_config`). -- Graph bootstrap (`bootstrap_application_graph`). - -It does **not** own: the actual graph node definitions (that's `core/`), HTTP server runtime (that's jvspatial). - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `__init__.py` | Exports `main` | -| `main.py:118-244` | `main()` — top-level entry; parses args, dispatches | -| `main.py:58-115` | `_first_app_root_path()` — extracts app root from arg list | -| `main.py:142-149` | `--serverless` flag handling | -| `main.py:152-159` | `--debug` flag | -| `main.py:161-179` | `--update` / `--source` / `--merge` parsing (mutually exclusive rules) | -| `main.py:181-192` | `--purge` (dev-only) | -| `main.py:205-226` | Subcommand dispatch | -| `main.py:239-244` | Default → `run_server()` | -| `commands.py` | Backward-compat re-export facade (re-exports the handlers below) | -| `server.py` | `load_app_env`, `purge_app_data`, `run_server`, `show_status`, `bootstrap_only`, `StartupLogCounter` | -| `agent_commands.py` | `handle_agent_command`, `handle_action_command`, `handle_bundle_command`, `list_agents`, `uninstall_agent` | -| `skill_commands.py` | `handle_skill_command` | -| `validate.py` | `run_validate`, `print_usage` | -| `chat.py` | `handle_chat_command` (serves the jvchat UI) | -| `server_config.py:59-180` | `create_server_from_config()` — reads app.yaml + env, builds jvspatial Server | -| `server_config.py` | `_set_db_env_from_config`, `_import_core_endpoint_modules` | -| `bootstrap.py` | `bootstrap_application_graph()` — orchestrates App + Agents + Actions creation | -| `app_commands.py` | `jvagent app create / profile new / ...` subcommands | -| `__main__.py:5-6` | `python -m jvagent` → `cli.main:main()` | - ---- - -## 3. CLI shape (memorize) - -``` -jvagent [app_root] [SUBCOMMAND] [FLAGS] - -DEFAULT (no subcommand) → start HTTP server - -SUBCOMMANDS: - status diagnostic snapshot of an app's graph - agent ... agent CRUD (add, list, enable, disable, delete) - action ... action CRUD on an agent - skill ... skill management - bootstrap bootstrap graph then exit (no server) - bundle emit Dockerfile + deployment artifacts - app create|profile scaffold a new app or profile - validate validate app.yaml + agents - chat serve the bundled jvchat UI on its own port (jvagent/webui) - stress-seed generate synthetic graph for testing - -FLAGS (apply to default + bootstrap): - --debug verbose logging - --update merge-mode YAML sync - --update --source destructive YAML sync (DESTRUCTIVE) - --update --merge explicit merge (== --update alone) - --purge wipe DB (dev only; requires JVSPATIAL_ENVIRONMENT=development) - --serverless single-worker, SERVERLESS_MODE=true -``` - -`--source` and `--merge` REQUIRE `--update`. They are mutually exclusive ([`main.py:167-172`](main.py)). - ---- - -## 4. Contracts (don't break) - -1. **App root resolution** ([`main.py:58-115`](main.py)) — strips path tokens, keeps subcommands + flags + stress-seed args. Don't break this for `jvagent ./myapp agent add ...`. -2. **`--purge` MUST require dev mode** ([`main.py:185-190`](main.py)). Never relax. -3. **`bootstrap_only` and `run_server` reset `App.update_mode` to `run`** after a successful sync. Don't remove this; otherwise one-shot ops persist across restarts. -4. **`load_app_env(app_root)` runs first** ([`main.py:130`](main.py)). Anything reading env vars before this is incorrect. -5. **`set_app_root(app_root)` must run before any node lookup** ([`main.py:134`](main.py)). Cache/config keys depend on it. -6. **`_set_db_env_from_config(app_root)` ([`main.py:150`](main.py))** translates app.yaml DB stanza into env vars jvspatial reads. Must precede any DB call. - ---- - -## 5. Subcommand handler conventions - -Each `handle_*_command` (in `agent_commands.py` / `skill_commands.py`): - -- Receives `args` (post-subcommand) and `app_root`. -- Uses `asyncio.run(...)` for async operations. -- Exits with non-zero on errors via `sys.exit(N)`. -- Prints user-facing output to stdout/stderr (no return values). - -When adding a subcommand: - -1. Add the name to `DISPATCH` ([`main.py:42-55`](main.py)). -2. Add a dispatch branch in the `if args[0] in DISPATCH:` block ([`main.py:205-226`](main.py)). -3. Add the handler in the relevant sibling module (`agent_commands.py`, `server.py`, `skill_commands.py`, `validate.py`); re-export it from `commands.py` for back-compat. -4. Document in [`/.planning/runbooks/`](../../.planning/runbooks/) if non-trivial. - ---- - -## 6. Tests - -- `tests/cli/` — argparse/dispatch tests. -- `tests/test_env_load.py` — `load_app_env` precedence. -- `tests/scaffold/` — `app create` flow. - -```bash -pytest tests/cli/ tests/test_env_load.py -v -``` - ---- - -## 7. Traps specific to cli/ - -| Trap | Fix | -|---|---| -| Treating a `STRESS_FLAG_NAMES` value as a path | `_first_app_root_path()` handles this — don't bypass it | -| Forgetting to strip a flag from `args` after parsing | The default-server path will see it and error "Unknown argument" | -| Calling jvspatial before `set_app_root()` + `load_app_env()` | Wrong DB / paths | -| Running `--update --source` with no app root | App root defaults to `cwd`; destructive op on the wrong dir. Always pass explicit path for source mode. | -| Mixing `--purge` with `--source` | `--purge` deletes the DB, then `--source` rebuilds — works, but slow. Use `--update --source` alone to overwrite. | - ---- - -## 8. Don't touch from outside cli/ - -- The argument-parsing semantics — many runbooks and docs encode them. -- The dispatch table — subcommand name is part of the public CLI contract. -- The `update_mode` reset behavior — its absence makes cold starts unpredictable. - ---- - -## 9. Out of scope here - -- Server runtime (uvicorn + FastAPI internals): jvspatial. -- Endpoint registration: handled by jvspatial server based on imported endpoint modules. -- Graph repair: see `jvagent/core/graph_repair*.py` and `core/CLAUDE.md`. diff --git a/jvagent/core/AGENTS.md b/jvagent/core/AGENTS.md index b4156a6e..17aaafd7 100644 --- a/jvagent/core/AGENTS.md +++ b/jvagent/core/AGENTS.md @@ -1 +1,123 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/core/ — Agent Guide + +> Local guide for editing the core subsystem. Cross-link: root [`/AGENTS.md`](../../AGENTS.md), spec [`/.planning/SPEC.md`](../../.planning/SPEC.md). + +--- + +## 1. What this directory owns + +The graph-level skeleton of an app: + +- `App` (singleton root node) and `Agent` / `Agents` nodes. +- YAML loaders that translate `app.yaml` and `agent.yaml` into graph state. +- Config resolution (env → app.yaml → defaults). +- Bootstrap / update-mode handling (`run` / `merge` / `source`). +- Graph repair (stale node cleanup, reconciliation). +- Caching, profiling, observability primitives. +- Core HTTP endpoints (under `core/endpoints/`). + +It does **not** own: per-user state (that's `memory/`), action plugins (that's `action/`), HTTP server bootstrap (that's `cli/`). + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `app.py:21` | `App` node — singleton, file storage, timezone, update_mode | +| `app.py:90-124` | Per-event-loop lock pattern (use this verbatim for any node singleton) | +| `app.py:596` | `set_app_update_mode` — example of mutating a `protected=True` field | +| `agent.py:30` | `Agent` node + `Agent.get(agent_id)` cached fetch | +| `agent.py:108-118` | Cache-aware fetch via `cache_manager.get_agent()` | +| `agent.py:244` | `Agent.get_memory()` — resolves attached `Memory` node | +| `agent.py:256` | `Agent.get_response_bus()` — lazy per-agent `ResponseBus` singleton | +| `agent.py:271-358` | `Agent.send_proactive_message()` — programmatic, response-only message send; resolves User/Conversation, creates empty-utterance Interaction, publishes via `ResponseBus` (auto-records + auto-dispatches to channel adapter). See [`docs/proactive-messages.md`](../../docs/proactive-messages.md). | +| `agents.py:17` | `Agents` branchpoint node | +| `app_loader.py` | `app.yaml → App` translation | +| `agent_loader.py` | `agent.yaml → Agent + Action graph` | +| `app_yaml_validator.py` | Schema-level validation of `app.yaml` | +| `agent_yaml_validator.py` | Same, for `agent.yaml` | +| `config.py:60-150` | `ConfigKey` / `ConfigSchema` precedence resolution | +| `env_resolver.py` | Expands `${ENV_VAR}` placeholders in config | +| `cache.py` | Per-agent action-type index; `cache_manager` for `Agent.get()` | +| `bootstrap_logger.py` | Startup logging context manager | +| `bootstrap_update_mode.py` | `--update` / `--merge` / `--source` handling | +| `graph_repair*.py` + `repair_phases/` | Reconciliation jobs (stale node cleanup, orphan repair) | +| `graph_repair_job.py` | Top-level repair orchestration | +| `endpoints/` | Core HTTP routes (`agents.py`, `app.py`, `conversation.py`, `status.py`, `graph_repair.py`) | +| `app_context.py` | `set_app_root()` / `get_app_root()` — used by CLI on startup | +| `profiling.py` | `profile_enabled`, `profiled_request` decorators | +| `jvspatial_compat.py` | Version compatibility shims | + +--- + +## 3. Contracts (don't break) + +1. **`App` is a singleton.** Always go through `await App.get()` — never construct it directly. +2. **`App._cached_app` and per-loop lock** ([`app.py:90-124`](app.py)) must be preserved exactly. Serverless warm-starts depend on it. +3. **`Agent.save()` invalidates cache** ([`agent.py:408`](agent.py)). If you bypass `.save()`, also invalidate manually via `invalidate_agent_cache(agent.id)`. +4. **`App.update_mode` resets to `run`** after a successful bootstrap (in `cli/server.py:run_server` / `bootstrap_only`). Don't let one-shot merge/source operations persist. +5. **`protected=True` fields require `object.__setattr__` + `await save()`** (see [`set_app_update_mode`](app.py)). Plain assignment is dropped silently in bulk overwrite paths. +6. **App.now() timezone semantics** ([`app.py:251`](app.py)): + - Returns `datetime` if `fmt` is `None`, else a formatted `str`. + - If `App.timezone` (IANA name, e.g. `America/New_York`) is set, the + returned `datetime` is **timezone-aware** in that zone. + - If `App.timezone` is unset or invalid, returns `datetime.now()` — + **naïve local time** (NOT UTC). Be careful with arithmetic against + timezone-aware datetimes. + - `app_now_aware_utc(app)` ([`app.py:577`](app.py)) normalizes both + branches into a `datetime` aware in UTC. Use it whenever you need to + compare or subtract `App.now()` against another timestamp. + +--- + +## 4. Adding to this directory + +| If you're adding... | Read first | +|---|---| +| A new App-level config key | `config.py` (`ConfigKey`, `ConfigSchema` pattern) | +| A new App field | Match the `attribute(...)` style from `app.py`. Add to `app_yaml_validator.py`. | +| A graph repair phase | `repair_phases/` — phases are isolated and registered via `repair_state.py`. | +| A core HTTP route | `endpoints/`. Use `@endpoint("/api/...")` with `auth=True, roles=["admin"]` unless public. | +| A bootstrap hook | `cli/bootstrap.py` is the orchestrator; add intermediate steps in `core/startup.py`. | +| Programmatic proactive send (agent → user, no inbound) | Call `await agent.send_proactive_message(user_id=..., content=..., channel=...)`. See [`docs/proactive-messages.md`](../../docs/proactive-messages.md). Do NOT publish to `ResponseBus` directly from outside the walker pipeline — use this method so the bound Interaction is created and the response is recorded. | + +--- + +## 5. Tests + +`tests/core/` mirrors this layout. Run a slice: + +```bash +pytest tests/core/ -v +``` + +For bootstrap-flow regressions: `tests/test_stress_seed_graph.py` exercises a full graph build. + +--- + +## 6. Traps specific to core/ + +| Trap | Fix | +|---|---| +| Reading `app.file_storage_provider` before `App` is loaded | `await App.get()` first; it may be `None` during cold init. | +| Calling `Agent.get()` with kwargs and an agent_id | The kwargs path bypasses cache. Pass only `agent_id` for cached lookups ([`agent.py:89`](agent.py)). | +| Forgetting to `await save()` after mutating a config key on the App node | Persists nothing. Add the `save()`. | +| Manually constructing event-loop locks | Use the dict-keyed-by-`id(loop)` pattern from [`app.py:90-124`](app.py). | +| Skipping `agent_yaml_validator` after editing `agent_loader.py` | YAML schema drift. Update both. | + +--- + +## 7. Don't touch from outside core/ + +- `Agent` / `Agents` / `App` class definitions — they're the canonical Node shape. +- `cache.py` invalidation contract — if you change cache keys, find every call site of `invalidate_agent_cache`. +- `graph_repair_*.py` — repair phases run on every cold start; bugs here cause user-visible boot failures. + +--- + +## 8. Out of scope here + +- Per-user data (User/Conversation/Interaction): see `jvagent/memory/`. +- Action lifecycle (`on_register`, `on_enable`, etc.): see `jvagent/action/AGENTS.md`. +- HTTP server start: see `jvagent/cli/AGENTS.md`. diff --git a/jvagent/core/CLAUDE.md b/jvagent/core/CLAUDE.md deleted file mode 100644 index 4b8ccfec..00000000 --- a/jvagent/core/CLAUDE.md +++ /dev/null @@ -1,123 +0,0 @@ -# jvagent/core/ — Agent Guide - -> Local guide for editing the core subsystem. Cross-link: root [`/CLAUDE.md`](../../CLAUDE.md), spec [`/.planning/SPEC.md`](../../.planning/SPEC.md). - ---- - -## 1. What this directory owns - -The graph-level skeleton of an app: - -- `App` (singleton root node) and `Agent` / `Agents` nodes. -- YAML loaders that translate `app.yaml` and `agent.yaml` into graph state. -- Config resolution (env → app.yaml → defaults). -- Bootstrap / update-mode handling (`run` / `merge` / `source`). -- Graph repair (stale node cleanup, reconciliation). -- Caching, profiling, observability primitives. -- Core HTTP endpoints (under `core/endpoints/`). - -It does **not** own: per-user state (that's `memory/`), action plugins (that's `action/`), HTTP server bootstrap (that's `cli/`). - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `app.py:21` | `App` node — singleton, file storage, timezone, update_mode | -| `app.py:90-124` | Per-event-loop lock pattern (use this verbatim for any node singleton) | -| `app.py:596` | `set_app_update_mode` — example of mutating a `protected=True` field | -| `agent.py:30` | `Agent` node + `Agent.get(agent_id)` cached fetch | -| `agent.py:108-118` | Cache-aware fetch via `cache_manager.get_agent()` | -| `agent.py:244` | `Agent.get_memory()` — resolves attached `Memory` node | -| `agent.py:256` | `Agent.get_response_bus()` — lazy per-agent `ResponseBus` singleton | -| `agent.py:271-358` | `Agent.send_proactive_message()` — programmatic, response-only message send; resolves User/Conversation, creates empty-utterance Interaction, publishes via `ResponseBus` (auto-records + auto-dispatches to channel adapter). See [`docs/proactive-messages.md`](../../docs/proactive-messages.md). | -| `agents.py:17` | `Agents` branchpoint node | -| `app_loader.py` | `app.yaml → App` translation | -| `agent_loader.py` | `agent.yaml → Agent + Action graph` | -| `app_yaml_validator.py` | Schema-level validation of `app.yaml` | -| `agent_yaml_validator.py` | Same, for `agent.yaml` | -| `config.py:60-150` | `ConfigKey` / `ConfigSchema` precedence resolution | -| `env_resolver.py` | Expands `${ENV_VAR}` placeholders in config | -| `cache.py` | Per-agent action-type index; `cache_manager` for `Agent.get()` | -| `bootstrap_logger.py` | Startup logging context manager | -| `bootstrap_update_mode.py` | `--update` / `--merge` / `--source` handling | -| `graph_repair*.py` + `repair_phases/` | Reconciliation jobs (stale node cleanup, orphan repair) | -| `graph_repair_job.py` | Top-level repair orchestration | -| `endpoints/` | Core HTTP routes (`agents.py`, `app.py`, `conversation.py`, `status.py`, `graph_repair.py`) | -| `app_context.py` | `set_app_root()` / `get_app_root()` — used by CLI on startup | -| `profiling.py` | `profile_enabled`, `profiled_request` decorators | -| `jvspatial_compat.py` | Version compatibility shims | - ---- - -## 3. Contracts (don't break) - -1. **`App` is a singleton.** Always go through `await App.get()` — never construct it directly. -2. **`App._cached_app` and per-loop lock** ([`app.py:90-124`](app.py)) must be preserved exactly. Serverless warm-starts depend on it. -3. **`Agent.save()` invalidates cache** ([`agent.py:408`](agent.py)). If you bypass `.save()`, also invalidate manually via `invalidate_agent_cache(agent.id)`. -4. **`App.update_mode` resets to `run`** after a successful bootstrap (in `cli/server.py:run_server` / `bootstrap_only`). Don't let one-shot merge/source operations persist. -5. **`protected=True` fields require `object.__setattr__` + `await save()`** (see [`set_app_update_mode`](app.py)). Plain assignment is dropped silently in bulk overwrite paths. -6. **App.now() timezone semantics** ([`app.py:251`](app.py)): - - Returns `datetime` if `fmt` is `None`, else a formatted `str`. - - If `App.timezone` (IANA name, e.g. `America/New_York`) is set, the - returned `datetime` is **timezone-aware** in that zone. - - If `App.timezone` is unset or invalid, returns `datetime.now()` — - **naïve local time** (NOT UTC). Be careful with arithmetic against - timezone-aware datetimes. - - `app_now_aware_utc(app)` ([`app.py:577`](app.py)) normalizes both - branches into a `datetime` aware in UTC. Use it whenever you need to - compare or subtract `App.now()` against another timestamp. - ---- - -## 4. Adding to this directory - -| If you're adding... | Read first | -|---|---| -| A new App-level config key | `config.py` (`ConfigKey`, `ConfigSchema` pattern) | -| A new App field | Match the `attribute(...)` style from `app.py`. Add to `app_yaml_validator.py`. | -| A graph repair phase | `repair_phases/` — phases are isolated and registered via `repair_state.py`. | -| A core HTTP route | `endpoints/`. Use `@endpoint("/api/...")` with `auth=True, roles=["admin"]` unless public. | -| A bootstrap hook | `cli/bootstrap.py` is the orchestrator; add intermediate steps in `core/startup.py`. | -| Programmatic proactive send (agent → user, no inbound) | Call `await agent.send_proactive_message(user_id=..., content=..., channel=...)`. See [`docs/proactive-messages.md`](../../docs/proactive-messages.md). Do NOT publish to `ResponseBus` directly from outside the walker pipeline — use this method so the bound Interaction is created and the response is recorded. | - ---- - -## 5. Tests - -`tests/core/` mirrors this layout. Run a slice: - -```bash -pytest tests/core/ -v -``` - -For bootstrap-flow regressions: `tests/test_stress_seed_graph.py` exercises a full graph build. - ---- - -## 6. Traps specific to core/ - -| Trap | Fix | -|---|---| -| Reading `app.file_storage_provider` before `App` is loaded | `await App.get()` first; it may be `None` during cold init. | -| Calling `Agent.get()` with kwargs and an agent_id | The kwargs path bypasses cache. Pass only `agent_id` for cached lookups ([`agent.py:89`](agent.py)). | -| Forgetting to `await save()` after mutating a config key on the App node | Persists nothing. Add the `save()`. | -| Manually constructing event-loop locks | Use the dict-keyed-by-`id(loop)` pattern from [`app.py:90-124`](app.py). | -| Skipping `agent_yaml_validator` after editing `agent_loader.py` | YAML schema drift. Update both. | - ---- - -## 7. Don't touch from outside core/ - -- `Agent` / `Agents` / `App` class definitions — they're the canonical Node shape. -- `cache.py` invalidation contract — if you change cache keys, find every call site of `invalidate_agent_cache`. -- `graph_repair_*.py` — repair phases run on every cold start; bugs here cause user-visible boot failures. - ---- - -## 8. Out of scope here - -- Per-user data (User/Conversation/Interaction): see `jvagent/memory/`. -- Action lifecycle (`on_register`, `on_enable`, etc.): see `jvagent/action/CLAUDE.md`. -- HTTP server start: see `jvagent/cli/CLAUDE.md`. diff --git a/jvagent/logging/AGENTS.md b/jvagent/logging/AGENTS.md index b4156a6e..cfb801d2 100644 --- a/jvagent/logging/AGENTS.md +++ b/jvagent/logging/AGENTS.md @@ -1 +1,96 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/logging/ — Agent Guide + +> Local guide for the logging subsystem. Cross-link: [`/.planning/reference/observability.md`](../../.planning/reference/observability.md), [`/docs/logging.md`](../../docs/logging.md), [`/docs/error-logging.md`](../../docs/error-logging.md), [`/docs/interaction-logging.md`](../../docs/interaction-logging.md). + +--- + +## 1. What this directory owns + +- The custom `INTERACTION` log level registration. +- The `GET /logs/agents/{agent_id}` query endpoint. +- Integration with jvspatial's logging service (separate `logs` database). + +It does **not** own: per-interaction observability metrics (those live on `Interaction.observability_metrics` — see `memory/AGENTS.md`), the actual log storage (jvspatial), or the bootstrap-startup log counter (that's in `core/`). + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `__init__.py` | Module init | +| `service.py` | Registers the `INTERACTION` log level + jvspatial logging service bridge | +| `endpoints.py:21-113` | `GET /logs/agents/{agent_id}` — filter by agent_id, time range, query | + +--- + +## 3. Logging tiers + +| Tier | Where | Storage | +|---|---|---| +| Standard Python `logging` | `logger = logging.getLogger(__name__)` everywhere | stderr (configurable via jvspatial `configure_standard_logging`) | +| `INTERACTION` level events | Auto-recorded during walker traversal | `logs` DB (separate from main jvspatial DB) | +| Per-interaction observability metrics | `Interaction.observability_metrics` (dict) | Main DB on the Interaction node | +| Per-interaction usage tally | `Interaction.usage` | Main DB on the Interaction node | +| Per-interaction action/event/directive lists | `Interaction.actions`, `Interaction.events`, `Interaction.directives` | Main DB | +| Bootstrap warnings / errors | `core/bootstrap_logger.py:BootstrapLogger` context manager | stderr + `core/startup.py` counter | +| `DBLogHandler` errors | Auto-routed by handler | `logs` DB | + +--- + +## 4. Contracts + +1. **Logs DB is separate from main DB.** Always: `get_logging_service(database_name="logs")`. Don't accidentally use the main jvspatial context. +2. **`INTERACTION` level must be registered before any module that emits at that level imports.** `service.py` handles this; don't reorder. +3. **`Interaction.observability_metrics` is the per-interaction aggregator.** Other code SHOULD merge events into it rather than emit standalone log records, so they're recoverable per interaction. +4. **Log retention is enforced via `App.log_retention_days`** ([`core/app.py:65`](../core/app.py)) — default 60 days. `TaskMonitor.tick` calls [`logging/retention.py`](retention.py) `purge_logs_past_retention`. Set to **`0` to disable** purge (unbounded growth — document intentionally). + +--- + +## 5. Adding to this directory + +| If you're adding... | Where | +|---|---| +| A new query/filter on logs | `endpoints.py` — add a query param + filter clause | +| A new log level | `service.py` — register before imports | +| A new observability field on Interaction | `memory/interaction.py` + agree on the schema in `observability_metrics` | +| A separate logging DB | Don't. Use the existing `logs` DB or extend jvspatial. | + +--- + +## 6. Tests + +- Add tests under `tests/logging/` (create if missing). +- `tests/unit/` and `tests/integration/` may have related coverage. + +```bash +pytest tests/logging/ -v # if exists +``` + +--- + +## 7. Traps specific to logging/ + +| Trap | Fix | +|---|---| +| Forgetting `preserve_handler_class_names=["DBLogHandler", "StartupLogCounter"]` when reconfiguring logging | Custom handlers drop; logs go to stderr only. See `cli/main.py:34`. | +| Querying logs DB with main jvspatial context | Misses. Use `get_logging_service(database_name="logs")`. | +| Emitting massive JSON blobs in log messages | DB pressure; storage growth. Use `observability_metrics` for structured per-interaction data. | +| Setting `App.log_retention_days = 0` | Disables purge (logs grow forever). Use a positive value in production. | +| Logging inside `_run_background_actions` without try/except | Failures propagate. The wrapper already catches; don't double-handle. | + +--- + +## 8. Don't touch from outside logging/ + +- The `INTERACTION` level integer value — third-party log consumers depend on it. +- The `logs` DB name string — endpoint contract. +- `DBLogHandler` class name — listed in `preserve_handler_class_names` at boot. + +--- + +## 9. Out of scope here + +- jvspatial's logging service internals. +- Per-action error reporting policy (each action handles its own try/except + logger.error). +- Tracing / OpenTelemetry — not currently integrated. diff --git a/jvagent/logging/CLAUDE.md b/jvagent/logging/CLAUDE.md deleted file mode 100644 index 369723de..00000000 --- a/jvagent/logging/CLAUDE.md +++ /dev/null @@ -1,96 +0,0 @@ -# jvagent/logging/ — Agent Guide - -> Local guide for the logging subsystem. Cross-link: [`/.planning/reference/observability.md`](../../.planning/reference/observability.md), [`/docs/logging.md`](../../docs/logging.md), [`/docs/error-logging.md`](../../docs/error-logging.md), [`/docs/interaction-logging.md`](../../docs/interaction-logging.md). - ---- - -## 1. What this directory owns - -- The custom `INTERACTION` log level registration. -- The `GET /logs/agents/{agent_id}` query endpoint. -- Integration with jvspatial's logging service (separate `logs` database). - -It does **not** own: per-interaction observability metrics (those live on `Interaction.observability_metrics` — see `memory/CLAUDE.md`), the actual log storage (jvspatial), or the bootstrap-startup log counter (that's in `core/`). - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `__init__.py` | Module init | -| `service.py` | Registers the `INTERACTION` log level + jvspatial logging service bridge | -| `endpoints.py:21-113` | `GET /logs/agents/{agent_id}` — filter by agent_id, time range, query | - ---- - -## 3. Logging tiers - -| Tier | Where | Storage | -|---|---|---| -| Standard Python `logging` | `logger = logging.getLogger(__name__)` everywhere | stderr (configurable via jvspatial `configure_standard_logging`) | -| `INTERACTION` level events | Auto-recorded during walker traversal | `logs` DB (separate from main jvspatial DB) | -| Per-interaction observability metrics | `Interaction.observability_metrics` (dict) | Main DB on the Interaction node | -| Per-interaction usage tally | `Interaction.usage` | Main DB on the Interaction node | -| Per-interaction action/event/directive lists | `Interaction.actions`, `Interaction.events`, `Interaction.directives` | Main DB | -| Bootstrap warnings / errors | `core/bootstrap_logger.py:BootstrapLogger` context manager | stderr + `core/startup.py` counter | -| `DBLogHandler` errors | Auto-routed by handler | `logs` DB | - ---- - -## 4. Contracts - -1. **Logs DB is separate from main DB.** Always: `get_logging_service(database_name="logs")`. Don't accidentally use the main jvspatial context. -2. **`INTERACTION` level must be registered before any module that emits at that level imports.** `service.py` handles this; don't reorder. -3. **`Interaction.observability_metrics` is the per-interaction aggregator.** Other code SHOULD merge events into it rather than emit standalone log records, so they're recoverable per interaction. -4. **Log retention is enforced via `App.log_retention_days`** ([`core/app.py:65`](../core/app.py)) — default 60 days. `TaskMonitor.tick` calls [`logging/retention.py`](retention.py) `purge_logs_past_retention`. Set to **`0` to disable** purge (unbounded growth — document intentionally). - ---- - -## 5. Adding to this directory - -| If you're adding... | Where | -|---|---| -| A new query/filter on logs | `endpoints.py` — add a query param + filter clause | -| A new log level | `service.py` — register before imports | -| A new observability field on Interaction | `memory/interaction.py` + agree on the schema in `observability_metrics` | -| A separate logging DB | Don't. Use the existing `logs` DB or extend jvspatial. | - ---- - -## 6. Tests - -- Add tests under `tests/logging/` (create if missing). -- `tests/unit/` and `tests/integration/` may have related coverage. - -```bash -pytest tests/logging/ -v # if exists -``` - ---- - -## 7. Traps specific to logging/ - -| Trap | Fix | -|---|---| -| Forgetting `preserve_handler_class_names=["DBLogHandler", "StartupLogCounter"]` when reconfiguring logging | Custom handlers drop; logs go to stderr only. See `cli/main.py:34`. | -| Querying logs DB with main jvspatial context | Misses. Use `get_logging_service(database_name="logs")`. | -| Emitting massive JSON blobs in log messages | DB pressure; storage growth. Use `observability_metrics` for structured per-interaction data. | -| Setting `App.log_retention_days = 0` | Disables purge (logs grow forever). Use a positive value in production. | -| Logging inside `_run_background_actions` without try/except | Failures propagate. The wrapper already catches; don't double-handle. | - ---- - -## 8. Don't touch from outside logging/ - -- The `INTERACTION` level integer value — third-party log consumers depend on it. -- The `logs` DB name string — endpoint contract. -- `DBLogHandler` class name — listed in `preserve_handler_class_names` at boot. - ---- - -## 9. Out of scope here - -- jvspatial's logging service internals. -- Per-action error reporting policy (each action handles its own try/except + logger.error). -- Tracing / OpenTelemetry — not currently integrated. diff --git a/jvagent/memory/AGENTS.md b/jvagent/memory/AGENTS.md index b4156a6e..fd7398a3 100644 --- a/jvagent/memory/AGENTS.md +++ b/jvagent/memory/AGENTS.md @@ -1 +1,130 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# jvagent/memory/ — Agent Guide + +> Local guide for the memory subsystem. Cross-link: [`/AGENTS.md`](../../AGENTS.md), [`/.planning/reference/memory-and-pruning.md`](../../.planning/reference/memory-and-pruning.md), [`/.planning/adr/0003-interaction-limit-pruning.md`](../../.planning/adr/0003-interaction-limit-pruning.md). + +--- + +## 1. What this directory owns + +The per-agent, per-user state graph: + +``` +Memory (per-agent) → User (per memory_id+user_id) → Conversation (per session_id) → Interaction* +``` + +Plus the locking, pruning, and persistence helpers that keep this graph consistent under concurrency. + +--- + +## 2. Key files + +| File | Purpose | +|---|---| +| `manager.py:18` | `Memory` node + `Memory.get_user()` (locks → unlocked fetch) | +| `manager.py:60-100` | User lookup; lock-then-fetch | +| `user.py:25` | `User` node — `memory_id`, `user_id`, `memory`, `memory_tags` | +| `user.py:16-24` | Compound unique index on `(memory_id, user_id)` — DO NOT DROP | +| `conversation.py:39` | `Conversation` node | +| `conversation.py:199-238` | `add_interaction()` — public wrapper with conversation_mutation_lock | +| `conversation.py:240-295` | `_add_interaction_unlocked()` — chain edges + trigger prune | +| `conversation.py:297-367` | `_prune_old_interactions()` — bounded-work rolling window | +| `interaction.py:47` | `Interaction` node — utterance, response, actions, directives, events, parameters | +| `lock_manager.py` | Per-`(memory_id, user_id)` async lock to prevent duplicate User creation | +| `distributed_conversation_lock.py` | Cross-process conversation lock (when configured) | +| `task_store.py` | Task node CRUD on Conversation/Interaction | +| `evidence_log.py` | Evidence / citation logging for memory | +| `services/` | Memory-related service helpers | +| `endpoints.py` | HTTP routes for user/conversation/interaction queries | +| `README.md` | Existing user-facing memory notes — keep, don't duplicate here | + +--- + +## 3. Contracts (don't break) + +1. **`User` is unique per `(memory_id, user_id)`.** The compound index at `user.py:16-24` enforces this — never drop it. Concurrent creates MUST go through `lock_manager`. +2. **First `Interaction` connects to `Conversation` with `direction="out"`** ([`conversation.py:272`](conversation.py)). Subsequent ones connect to the previous `Interaction` with `direction="both"` ([`conversation.py:270`](conversation.py)). +3. **Pruning never removes the last `Interaction`** ([`conversation.py:333-336`](conversation.py)). If `next_interaction` is `None`, stop. +4. **Pruning is bounded per call** by `JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL` (default 100, [`conversation.py:317-323`](conversation.py)). The remainder happens on subsequent appends or via `Memory.apply_interaction_limit_pruning_for_connected_users`. +5. **`Agent.interaction_limit = 0` disables pruning entirely.** Code paths MUST early-return when limit is `0`. +6. **`Conversation.interaction_count` and `last_interaction_at`** are written together with the edge insert in `add_interaction()` ([`conversation.py:272-277`](conversation.py)). Don't update one without the other. +7. **Durable user state** lives in `User.memory` + `User.memory_tags`. One-time read migration copies legacy `context.user_model` into `memory["user_model"]` on user load. + +--- + +## 3a. Canonical memory fields (audit) + +| Field | Scope | Shape | Use for | +|---|---|---|---| +| `User.memory` | Cross-session, per user | `dict[str, Any]` — markdown-keyed flat map | Orchestrator `memory_set` / `memory_get` / `memory_append` / `memory_search`; durable user facts and preferences | +| `User.memory_tags` | Per user | `dict[str, list[str]]` | Tag metadata keyed by `User.memory` keys | +| `Conversation.context` | Per session | `dict[str, Any]` | Ephemeral turn state: routing buffers (`deferred_fragments`), flow flags, interview session handles — not long-term user memory | +| `Conversation.memory` | Per session | `dict[str, str]` | Session-scoped markdown map (same tool surface as user memory but scoped to one conversation) | +| `Conversation.memory_tags` | Per session | `dict[str, list[str]]` | Tags for `Conversation.memory` keys | + +**Rule of thumb:** durable cross-session state → `User.memory`; session-only scratch → `Conversation.context` or `Conversation.memory`. + +--- + +## 4. Pruning math (memorize) + +``` +to_remove = interaction_count - interaction_limit +max_prune = min(to_remove, JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL) +walk from = first interaction +stop when = removed == max_prune OR next_interaction is None +``` + +After pruning, `last_interaction_id` is verified — if stale, rebuilt by traversal ([`conversation.py:354-364`](conversation.py)). + +--- + +## 5. Adding to this directory + +| If you're adding... | Where | +|---|---| +| A new field on User/Conversation/Interaction | Add via `attribute(...)`. Update `endpoints.py` query shape if external. | +| A new pruning strategy | Modify `_prune_old_interactions`; preserve bounded-work + last-interaction invariants. Add a regression test in `tests/test_comprehensive_pruning.py`. | +| Cross-conversation memory query | Use `Memory.get_user()` + walk; do not bypass locks. | +| Distributed locking | Extend `distributed_conversation_lock.py`. | + +--- + +## 6. Tests + +- `tests/memory/` — unit tests for User/Conversation/Interaction CRUD. +- `tests/test_comprehensive_pruning.py` — full pruning regression suite. +- `tests/test_interview_path_pruning_and_convergence.py` — branching + pruning interaction. +- `tests/test_pruning_fix.py` — specific pruning bug regression. + +```bash +pytest tests/memory/ tests/test_comprehensive_pruning.py -v +``` + +--- + +## 7. Traps specific to memory/ + +| Trap | Fix | +|---|---| +| Creating a User without acquiring the lock | Duplicate rows; compound index rejects on save | Use `Memory.get_user()` which locks first | +| Editing `interaction_count` directly | Drift from actual edge count | Always mutate via `add_interaction` / `_prune_old_interactions` | +| Adding fields without `attribute()` | Not persisted | Use `attribute(...)` | +| Spawning a Walker over a long Conversation without limits | Walker `max_steps` (10000) trips | Either bound the traversal or paginate manually | +| Deleting an Interaction mid-chain manually | Leaves dangling bidirectional edges | Use the pruning routine or rewire both sides | +| Legacy `context.user_model` on load | One-time migration | Copied into `User.memory["user_model"]` automatically | + +--- + +## 8. Don't touch from outside memory/ + +- The compound index at `user.py:16-24` — call sites elsewhere assume `(memory_id, user_id)` uniqueness. +- `_prune_old_interactions` invariants — there are subtle correctness tests guarding these. +- `lock_manager` — bypassing produces duplicate Users that the DB then rejects, causing intermittent failures. + +--- + +## 9. Out of scope here + +- Action plugins: see `jvagent/action/AGENTS.md`. +- The Executive's Skills center memory tools wrap this layer. +- PageIndex (document index): see `jvagent/action/pageindex/`. diff --git a/jvagent/memory/CLAUDE.md b/jvagent/memory/CLAUDE.md deleted file mode 100644 index 84ec2779..00000000 --- a/jvagent/memory/CLAUDE.md +++ /dev/null @@ -1,130 +0,0 @@ -# jvagent/memory/ — Agent Guide - -> Local guide for the memory subsystem. Cross-link: [`/CLAUDE.md`](../../CLAUDE.md), [`/.planning/reference/memory-and-pruning.md`](../../.planning/reference/memory-and-pruning.md), [`/.planning/adr/0003-interaction-limit-pruning.md`](../../.planning/adr/0003-interaction-limit-pruning.md). - ---- - -## 1. What this directory owns - -The per-agent, per-user state graph: - -``` -Memory (per-agent) → User (per memory_id+user_id) → Conversation (per session_id) → Interaction* -``` - -Plus the locking, pruning, and persistence helpers that keep this graph consistent under concurrency. - ---- - -## 2. Key files - -| File | Purpose | -|---|---| -| `manager.py:18` | `Memory` node + `Memory.get_user()` (locks → unlocked fetch) | -| `manager.py:60-100` | User lookup; lock-then-fetch | -| `user.py:25` | `User` node — `memory_id`, `user_id`, `memory`, `memory_tags` | -| `user.py:16-24` | Compound unique index on `(memory_id, user_id)` — DO NOT DROP | -| `conversation.py:39` | `Conversation` node | -| `conversation.py:199-238` | `add_interaction()` — public wrapper with conversation_mutation_lock | -| `conversation.py:240-295` | `_add_interaction_unlocked()` — chain edges + trigger prune | -| `conversation.py:297-367` | `_prune_old_interactions()` — bounded-work rolling window | -| `interaction.py:47` | `Interaction` node — utterance, response, actions, directives, events, parameters | -| `lock_manager.py` | Per-`(memory_id, user_id)` async lock to prevent duplicate User creation | -| `distributed_conversation_lock.py` | Cross-process conversation lock (when configured) | -| `task_store.py` | Task node CRUD on Conversation/Interaction | -| `evidence_log.py` | Evidence / citation logging for memory | -| `services/` | Memory-related service helpers | -| `endpoints.py` | HTTP routes for user/conversation/interaction queries | -| `README.md` | Existing user-facing memory notes — keep, don't duplicate here | - ---- - -## 3. Contracts (don't break) - -1. **`User` is unique per `(memory_id, user_id)`.** The compound index at `user.py:16-24` enforces this — never drop it. Concurrent creates MUST go through `lock_manager`. -2. **First `Interaction` connects to `Conversation` with `direction="out"`** ([`conversation.py:272`](conversation.py)). Subsequent ones connect to the previous `Interaction` with `direction="both"` ([`conversation.py:270`](conversation.py)). -3. **Pruning never removes the last `Interaction`** ([`conversation.py:333-336`](conversation.py)). If `next_interaction` is `None`, stop. -4. **Pruning is bounded per call** by `JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL` (default 100, [`conversation.py:317-323`](conversation.py)). The remainder happens on subsequent appends or via `Memory.apply_interaction_limit_pruning_for_connected_users`. -5. **`Agent.interaction_limit = 0` disables pruning entirely.** Code paths MUST early-return when limit is `0`. -6. **`Conversation.interaction_count` and `last_interaction_at`** are written together with the edge insert in `add_interaction()` ([`conversation.py:272-277`](conversation.py)). Don't update one without the other. -7. **Durable user state** lives in `User.memory` + `User.memory_tags`. One-time read migration copies legacy `context.user_model` into `memory["user_model"]` on user load. - ---- - -## 3a. Canonical memory fields (audit) - -| Field | Scope | Shape | Use for | -|---|---|---|---| -| `User.memory` | Cross-session, per user | `dict[str, Any]` — markdown-keyed flat map | Orchestrator `memory_set` / `memory_get` / `memory_append` / `memory_search`; durable user facts and preferences | -| `User.memory_tags` | Per user | `dict[str, list[str]]` | Tag metadata keyed by `User.memory` keys | -| `Conversation.context` | Per session | `dict[str, Any]` | Ephemeral turn state: routing buffers (`deferred_fragments`), flow flags, interview session handles — not long-term user memory | -| `Conversation.memory` | Per session | `dict[str, str]` | Session-scoped markdown map (same tool surface as user memory but scoped to one conversation) | -| `Conversation.memory_tags` | Per session | `dict[str, list[str]]` | Tags for `Conversation.memory` keys | - -**Rule of thumb:** durable cross-session state → `User.memory`; session-only scratch → `Conversation.context` or `Conversation.memory`. - ---- - -## 4. Pruning math (memorize) - -``` -to_remove = interaction_count - interaction_limit -max_prune = min(to_remove, JVAGENT_MAX_INTERACTIONS_PRUNED_PER_CALL) -walk from = first interaction -stop when = removed == max_prune OR next_interaction is None -``` - -After pruning, `last_interaction_id` is verified — if stale, rebuilt by traversal ([`conversation.py:354-364`](conversation.py)). - ---- - -## 5. Adding to this directory - -| If you're adding... | Where | -|---|---| -| A new field on User/Conversation/Interaction | Add via `attribute(...)`. Update `endpoints.py` query shape if external. | -| A new pruning strategy | Modify `_prune_old_interactions`; preserve bounded-work + last-interaction invariants. Add a regression test in `tests/test_comprehensive_pruning.py`. | -| Cross-conversation memory query | Use `Memory.get_user()` + walk; do not bypass locks. | -| Distributed locking | Extend `distributed_conversation_lock.py`. | - ---- - -## 6. Tests - -- `tests/memory/` — unit tests for User/Conversation/Interaction CRUD. -- `tests/test_comprehensive_pruning.py` — full pruning regression suite. -- `tests/test_interview_path_pruning_and_convergence.py` — branching + pruning interaction. -- `tests/test_pruning_fix.py` — specific pruning bug regression. - -```bash -pytest tests/memory/ tests/test_comprehensive_pruning.py -v -``` - ---- - -## 7. Traps specific to memory/ - -| Trap | Fix | -|---|---| -| Creating a User without acquiring the lock | Duplicate rows; compound index rejects on save | Use `Memory.get_user()` which locks first | -| Editing `interaction_count` directly | Drift from actual edge count | Always mutate via `add_interaction` / `_prune_old_interactions` | -| Adding fields without `attribute()` | Not persisted | Use `attribute(...)` | -| Spawning a Walker over a long Conversation without limits | Walker `max_steps` (10000) trips | Either bound the traversal or paginate manually | -| Deleting an Interaction mid-chain manually | Leaves dangling bidirectional edges | Use the pruning routine or rewire both sides | -| Legacy `context.user_model` on load | One-time migration | Copied into `User.memory["user_model"]` automatically | - ---- - -## 8. Don't touch from outside memory/ - -- The compound index at `user.py:16-24` — call sites elsewhere assume `(memory_id, user_id)` uniqueness. -- `_prune_old_interactions` invariants — there are subtle correctness tests guarding these. -- `lock_manager` — bypassing produces duplicate Users that the DB then rejects, causing intermittent failures. - ---- - -## 9. Out of scope here - -- Action plugins: see `jvagent/action/CLAUDE.md`. -- The Executive's Skills center memory tools wrap this layer. -- PageIndex (document index): see `jvagent/action/pageindex/`. diff --git a/jvagent/version.py b/jvagent/version.py index 6b9e2a64..086f2037 100644 --- a/jvagent/version.py +++ b/jvagent/version.py @@ -1,3 +1,3 @@ """Version information for jvagent package.""" -__version__ = "0.1.8rc16" +__version__ = "0.1.8rc19" diff --git a/pyproject.toml b/pyproject.toml index b8a1fc65..e9203606 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -30,8 +30,8 @@ classifiers = [ dependencies = [ "aiohttp>=3.9.0", - # CI records the resolved jvspatial version after install (see .github/workflows/test-jvagent.yaml). - "jvspatial==0.0.21", + "jvspatial==0.1.0", + "pydantic[email]>=2.13.5", "python-dotenv>=1.0.0", "pyyaml>=6.0.0", "httpx>=0.27.0", diff --git a/requirements-all.txt b/requirements-all.txt index a6a767ff..b449925d 100644 --- a/requirements-all.txt +++ b/requirements-all.txt @@ -6,7 +6,7 @@ # Must stay in sync with [project] dependencies in pyproject.toml — # enforced by tests/test_requirements_sync.py. aiohttp>=3.9.0 -jvspatial==0.0.21 +jvspatial==0.1.0 python-dotenv>=1.0.0 pyyaml>=6.0.0 httpx>=0.27.0 diff --git a/requirements.txt b/requirements.txt index 8fc60f38..c370707b 100644 --- a/requirements.txt +++ b/requirements.txt @@ -3,7 +3,8 @@ # Test-only deps (incl. Docling for PageIndex): pyproject.toml # [project.optional-dependencies] test — install with: pip install -e ".[test]" aiohttp>=3.9.0 -jvspatial==0.0.21 +jvspatial==0.1.0 +pydantic[email]>=2.13.5 python-dotenv>=1.0.0 pyyaml>=6.0.0 httpx>=0.27.0 diff --git a/tests/AGENTS.md b/tests/AGENTS.md index b4156a6e..495f24cd 100644 --- a/tests/AGENTS.md +++ b/tests/AGENTS.md @@ -1 +1,159 @@ -See [CLAUDE.md](CLAUDE.md) — same agent guide, alternate filename for non-Claude AI agents. +# tests/ — Agent Guide + +> Local guide for the test suite. Cross-link: root [`/AGENTS.md`](../AGENTS.md), [`/.planning/runbooks/local-dev.md`](../.planning/runbooks/local-dev.md). + +--- + +## 1. Layout (mirrors `jvagent/`) + +``` +tests/ +├── conftest.py # session-level fixtures +├── harness/ # ADR-0054 contracts + HarnessRuntime (HP-02…12) +├── conformance/ # host-neutral harness suite (marker: harness_conformance) +├── action/ # per-action unit tests +│ ├── orchestrator/ # Orchestrator loop +│ ├── interact/ # walker bootstrap + visit semantics +│ ├── interview/ # branching, convergence, pruning +│ ├── mcp/ +│ ├── model/ +│ ├── pageindex/ +│ ├── postiz_action/ +│ ├── response/ +│ ├── task_creation_interact_action/ +│ ├── task_monitor/ +│ ├── whatsapp/ +│ ├── google/ +│ ├── facebook_action/ +│ ├── email_action/ +│ ├── access_control/ +│ ├── test_action_loader.py +│ ├── test_action_endpoints.py +│ ├── test_plugin_system.py +│ ├── test_no_persona_imports.py, test_reply*.py +│ ├── test_secrets.py +│ ├── test_vision*.py +│ └── ... +├── core/ # framework-level tests +├── memory/ # memory subsystem tests +├── cli/ # CLI argparse + dispatch +├── scaffold/ # `jvagent app create` flow +├── bundle/ # bundle/Dockerfile generation +├── skills/ # skill discovery + dispatch +├── integration/ # end-to-end flows +├── wire/ # WIRE-CONTRACT tier — see §3.1 +├── unit/ # cross-cutting unit tests +├── test_comprehensive_pruning.py +├── test_pruning_fix.py +├── test_interview_branch_cache.py +├── test_interview_path_pruning_and_convergence.py +├── test_stress_seed_graph.py +├── test_tool_schema_audit.py +├── test_embed.py +└── test_env_load.py +``` + +--- + +## 2. Running a slice + +```bash +pytest tests/ # everything (slow) +pytest tests/action/orchestrator/ -v # Orchestrator +pytest -k pruning # by keyword +pytest --lf # last-failed only +pytest -x # stop on first failure +pytest -n auto # parallel (pytest-xdist if installed) +``` + +--- + +## 3. Fixtures + conventions + +- `pytest-asyncio` is configured for `async def` tests. +- `conftest.py` provides DB context fixtures. Don't re-create one per test. +- Mock external HTTP via `pytest-httpx` or `respx`. +- For walker-level tests, construct an `InteractWalker` directly; see `tests/action/interact/` for patterns. +- For full integration, see `tests/integration/`. +- For stress / synthetic data: `tests/test_stress_seed_graph.py` + the `stress-seed` CLI subcommand. + +### 3.1 The wire tier (`tests/wire/`) + +Everything else in this suite constructs actions in memory. `tests/wire/` does +not: it bootstraps a real app graph from YAML, loads the action **back out of +the database**, and asserts on the exact prompt a tick would send. Only the +model is stubbed. + +Use it when the thing that can break is *wiring* rather than logic: + +| Symptom it catches | Why unit tests miss it | +|---|---| +| A stale persisted attribute beating a new code default | The constructed object has the new default | +| A render site that stopped calling its helper | The helper's own test still passes | +| A rule that renders twice, or not at all | Neither is visible without the assembled prompt | +| Interpreter-order dependence reaching the model | Needs two processes with different `PYTHONHASHSEED` | + +```bash +pytest tests/wire/ -q # ~5s, boots a real graph per test +``` + +Conventions: +- **Assert on captured wire text, never on a helper's return value.** A helper + can be correct while the code that calls it is not — that is the failure mode + this tier exists for. +- The `wire` fixture is function-scoped on purpose. jvspatial objects bind to + the event loop that created them; a session-scoped graph fails across + pytest-asyncio's per-test loops. +- Cross-process checks go through `tests/wire/_render_once.py` as a subprocess + so the caller controls `PYTHONHASHSEED`. +- **Mutation-check new tests here.** Break the behaviour on purpose and confirm + the test goes red before trusting it. + +--- + +## 4. When you add a feature + +Add at least one test slice: + +| Touched | Add tests at | +|---|---| +| `core/` | `tests/core/` | +| `memory/` | `tests/memory/` + regression in `tests/test_comprehensive_pruning.py` if it affects pruning | +| `action/{name}/` | `tests/action/{name}/` | +| `action/interact/` | `tests/action/interact/` + `tests/action/access_control/` if access control changes | +| `action/orchestrator/` | `tests/action/orchestrator/` | +| Prompt assembly, parameter rendering, persisted config | `tests/wire/` as well — logic tests do not see wiring | +| `cli/` | `tests/cli/` | +| Tool schemas | check `tests/test_tool_schema_audit.py` still passes; add cases | + +For pure-doc PRs, no tests are required, but `pre-commit run --all-files` still runs. + +--- + +## 5. Contracts + +1. **Tests must not depend on a running MongoDB unless explicitly marked.** Use the JSON backend (`JVSPATIAL_DB_TYPE=json` in `conftest.py`). +2. **Tests must clean up after themselves.** Use the fixture-managed DB context; do not write to the production `jvdb/`. +3. **No real network calls.** Mock the HTTP layer. +4. **Async tests must use `@pytest.mark.asyncio` and `async def`.** +5. **Test names start with `test_`**; helper files are `_helpers.py` or `conftest.py`. + +--- + +## 6. Traps specific to tests/ + +| Trap | Fix | +|---|---| +| Tests pass alone but fail together | DB context leaks between tests. Use the per-test fixture from `conftest.py`. | +| Walker tests time out | Default `max_execution_time=300` — set lower in test setup if needed. | +| Mocking jvspatial entities directly | Brittle. Mock at the HTTP / model boundary instead. | +| Hard-coding action IDs in tests | Use the fixture that creates the action and returns its ID. | +| Skipping the commit gate | `pre-commit run --all-files` + `pytest` must pass before **every** commit (root [`AGENTS.md` §6](../AGENTS.md)). Never `--no-verify`. | +| Asserting on log strings | Use `caplog` fixture, not string match on stderr. | + +--- + +## 7. Don't touch from outside tests/ + +- `conftest.py` — shared fixtures with order constraints. +- Stress-seed scenarios — they drive `tests/test_stress_seed_graph.py` and a CLI subcommand simultaneously. diff --git a/tests/CLAUDE.md b/tests/CLAUDE.md deleted file mode 100644 index cc2f6e44..00000000 --- a/tests/CLAUDE.md +++ /dev/null @@ -1,159 +0,0 @@ -# tests/ — Agent Guide - -> Local guide for the test suite. Cross-link: root [`/CLAUDE.md`](../CLAUDE.md), [`/.planning/runbooks/local-dev.md`](../.planning/runbooks/local-dev.md). - ---- - -## 1. Layout (mirrors `jvagent/`) - -``` -tests/ -├── conftest.py # session-level fixtures -├── harness/ # ADR-0054 contracts + HarnessRuntime (HP-02…12) -├── conformance/ # host-neutral harness suite (marker: harness_conformance) -├── action/ # per-action unit tests -│ ├── orchestrator/ # Orchestrator loop -│ ├── interact/ # walker bootstrap + visit semantics -│ ├── interview/ # branching, convergence, pruning -│ ├── mcp/ -│ ├── model/ -│ ├── pageindex/ -│ ├── postiz_action/ -│ ├── response/ -│ ├── task_creation_interact_action/ -│ ├── task_monitor/ -│ ├── whatsapp/ -│ ├── google/ -│ ├── facebook_action/ -│ ├── email_action/ -│ ├── access_control/ -│ ├── test_action_loader.py -│ ├── test_action_endpoints.py -│ ├── test_plugin_system.py -│ ├── test_no_persona_imports.py, test_reply*.py -│ ├── test_secrets.py -│ ├── test_vision*.py -│ └── ... -├── core/ # framework-level tests -├── memory/ # memory subsystem tests -├── cli/ # CLI argparse + dispatch -├── scaffold/ # `jvagent app create` flow -├── bundle/ # bundle/Dockerfile generation -├── skills/ # skill discovery + dispatch -├── integration/ # end-to-end flows -├── wire/ # WIRE-CONTRACT tier — see §3.1 -├── unit/ # cross-cutting unit tests -├── test_comprehensive_pruning.py -├── test_pruning_fix.py -├── test_interview_branch_cache.py -├── test_interview_path_pruning_and_convergence.py -├── test_stress_seed_graph.py -├── test_tool_schema_audit.py -├── test_embed.py -└── test_env_load.py -``` - ---- - -## 2. Running a slice - -```bash -pytest tests/ # everything (slow) -pytest tests/action/orchestrator/ -v # Orchestrator -pytest -k pruning # by keyword -pytest --lf # last-failed only -pytest -x # stop on first failure -pytest -n auto # parallel (pytest-xdist if installed) -``` - ---- - -## 3. Fixtures + conventions - -- `pytest-asyncio` is configured for `async def` tests. -- `conftest.py` provides DB context fixtures. Don't re-create one per test. -- Mock external HTTP via `pytest-httpx` or `respx`. -- For walker-level tests, construct an `InteractWalker` directly; see `tests/action/interact/` for patterns. -- For full integration, see `tests/integration/`. -- For stress / synthetic data: `tests/test_stress_seed_graph.py` + the `stress-seed` CLI subcommand. - -### 3.1 The wire tier (`tests/wire/`) - -Everything else in this suite constructs actions in memory. `tests/wire/` does -not: it bootstraps a real app graph from YAML, loads the action **back out of -the database**, and asserts on the exact prompt a tick would send. Only the -model is stubbed. - -Use it when the thing that can break is *wiring* rather than logic: - -| Symptom it catches | Why unit tests miss it | -|---|---| -| A stale persisted attribute beating a new code default | The constructed object has the new default | -| A render site that stopped calling its helper | The helper's own test still passes | -| A rule that renders twice, or not at all | Neither is visible without the assembled prompt | -| Interpreter-order dependence reaching the model | Needs two processes with different `PYTHONHASHSEED` | - -```bash -pytest tests/wire/ -q # ~5s, boots a real graph per test -``` - -Conventions: -- **Assert on captured wire text, never on a helper's return value.** A helper - can be correct while the code that calls it is not — that is the failure mode - this tier exists for. -- The `wire` fixture is function-scoped on purpose. jvspatial objects bind to - the event loop that created them; a session-scoped graph fails across - pytest-asyncio's per-test loops. -- Cross-process checks go through `tests/wire/_render_once.py` as a subprocess - so the caller controls `PYTHONHASHSEED`. -- **Mutation-check new tests here.** Break the behaviour on purpose and confirm - the test goes red before trusting it. - ---- - -## 4. When you add a feature - -Add at least one test slice: - -| Touched | Add tests at | -|---|---| -| `core/` | `tests/core/` | -| `memory/` | `tests/memory/` + regression in `tests/test_comprehensive_pruning.py` if it affects pruning | -| `action/{name}/` | `tests/action/{name}/` | -| `action/interact/` | `tests/action/interact/` + `tests/action/access_control/` if access control changes | -| `action/orchestrator/` | `tests/action/orchestrator/` | -| Prompt assembly, parameter rendering, persisted config | `tests/wire/` as well — logic tests do not see wiring | -| `cli/` | `tests/cli/` | -| Tool schemas | check `tests/test_tool_schema_audit.py` still passes; add cases | - -For pure-doc PRs, no tests are required, but `pre-commit run --all-files` still runs. - ---- - -## 5. Contracts - -1. **Tests must not depend on a running MongoDB unless explicitly marked.** Use the JSON backend (`JVSPATIAL_DB_TYPE=json` in `conftest.py`). -2. **Tests must clean up after themselves.** Use the fixture-managed DB context; do not write to the production `jvdb/`. -3. **No real network calls.** Mock the HTTP layer. -4. **Async tests must use `@pytest.mark.asyncio` and `async def`.** -5. **Test names start with `test_`**; helper files are `_helpers.py` or `conftest.py`. - ---- - -## 6. Traps specific to tests/ - -| Trap | Fix | -|---|---| -| Tests pass alone but fail together | DB context leaks between tests. Use the per-test fixture from `conftest.py`. | -| Walker tests time out | Default `max_execution_time=300` — set lower in test setup if needed. | -| Mocking jvspatial entities directly | Brittle. Mock at the HTTP / model boundary instead. | -| Hard-coding action IDs in tests | Use the fixture that creates the action and returns its ID. | -| Skipping the commit gate | `pre-commit run --all-files` + `pytest` must pass before **every** commit (root [`CLAUDE.md` §6](../CLAUDE.md)). Never `--no-verify`. | -| Asserting on log strings | Use `caplog` fixture, not string match on stderr. | - ---- - -## 7. Don't touch from outside tests/ - -- `conftest.py` — shared fixtures with order constraints. -- Stress-seed scenarios — they drive `tests/test_stress_seed_graph.py` and a CLI subcommand simultaneously. diff --git a/tests/action/interact/test_endpoints.py b/tests/action/interact/test_endpoints.py index 033ff0fc..c69773dc 100644 --- a/tests/action/interact/test_endpoints.py +++ b/tests/action/interact/test_endpoints.py @@ -242,6 +242,7 @@ async def test_build_interact_response_development_mode(self): interaction.events = [] interaction.observability_metrics = [] interaction.streamed = False + interaction.conversation_id = None report = [{"test": "report"}] @@ -370,6 +371,7 @@ async def test_build_interact_response_no_report(self): interaction.events = [] interaction.observability_metrics = [] interaction.streamed = False + interaction.conversation_id = None response = await build_interact_response( user_id="usr_123", diff --git a/tests/action/interview/test_awaiting_fields.py b/tests/action/interview/test_awaiting_fields.py index a3a0fc55..b1613c84 100644 --- a/tests/action/interview/test_awaiting_fields.py +++ b/tests/action/interview/test_awaiting_fields.py @@ -35,7 +35,6 @@ async def test_activation_includes_awaiting_fields_not_field_definitions(signup_ visitor = SimpleNamespace(conversation=conv) action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() result = json.loads( await action._handle_start("signup_interview", visitor, user_message="sign up") @@ -60,7 +59,6 @@ async def test_activation_includes_awaiting_fields_not_field_definitions(signup_ async def test_on_skill_activate_includes_awaiting_fields(signup_action): action, _spec = signup_action action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() action._get_conversation = AsyncMock(return_value=None) note = await action.on_skill_activate("signup_interview", user_message="sign up") diff --git a/tests/action/interview/test_idempotent_resubmit.py b/tests/action/interview/test_idempotent_resubmit.py index 5d3c3888..26fceca1 100644 --- a/tests/action/interview/test_idempotent_resubmit.py +++ b/tests/action/interview/test_idempotent_resubmit.py @@ -37,7 +37,6 @@ async def _start(action): conv.save = AsyncMock() visitor = SimpleNamespace(conversation=conv, utterance="My name is Eldon Marks") action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() await action._handle_start( "signup_interview", visitor, user_message="My name is Eldon Marks" ) diff --git a/tests/action/interview/test_interview_set_field_validation.py b/tests/action/interview/test_interview_set_field_validation.py index f089a3b5..91cc41c2 100644 --- a/tests/action/interview/test_interview_set_field_validation.py +++ b/tests/action/interview/test_interview_set_field_validation.py @@ -155,7 +155,6 @@ async def _passthrough_post_processors(*args, **kwargs): async def test_init_does_not_auto_store_tracking_from_user_message(pre_alert_action): action, contract = pre_alert_action action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() action._get_conversation = AsyncMock(return_value=None) result = json.loads( @@ -174,7 +173,6 @@ async def test_init_does_not_auto_store_tracking_from_user_message(pre_alert_act async def test_init_without_extractable_data_asks_first_question(pre_alert_action): action, contract = pre_alert_action action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() action._get_conversation = AsyncMock(return_value=None) result = json.loads( @@ -226,7 +224,6 @@ async def test_next_field_falls_back_to_phone_question_off_whatsapp( async def test_init_does_not_auto_store_phone_from_user_message(onboarding_action): action, contract = onboarding_action action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() action._get_conversation = AsyncMock(return_value=None) result = json.loads( diff --git a/tests/action/interview/test_interview_skill_activate.py b/tests/action/interview/test_interview_skill_activate.py index 53df7d41..7bb29dc2 100644 --- a/tests/action/interview/test_interview_skill_activate.py +++ b/tests/action/interview/test_interview_skill_activate.py @@ -30,7 +30,6 @@ def _interview_action_with_contracts() -> InterviewAction: _SKILLS_DIR / "pre_alert_interview" ) action._get_conversation = AsyncMock(return_value=None) - action._ensure_active_task = AsyncMock() return action diff --git a/tests/action/interview/test_path_regression_remediation.py b/tests/action/interview/test_path_regression_remediation.py index aec1debf..58e45cef 100644 --- a/tests/action/interview/test_path_regression_remediation.py +++ b/tests/action/interview/test_path_regression_remediation.py @@ -105,7 +105,6 @@ async def test_activation_awaiting_fields_only_first_field(signup_action): visitor = SimpleNamespace(conversation=conv) action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() result = json.loads( await action._handle_start("signup_interview", visitor, user_message=_OPENING) @@ -125,7 +124,6 @@ async def test_activation_opening_extracts_user_name(signup_action): visitor = SimpleNamespace(conversation=conv, utterance=_OPENING) action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() await action._handle_start("signup_interview", visitor, user_message=_OPENING) diff --git a/tests/action/interview/test_review_gate.py b/tests/action/interview/test_review_gate.py index dfcd5925..1bdc7c0a 100644 --- a/tests/action/interview/test_review_gate.py +++ b/tests/action/interview/test_review_gate.py @@ -28,7 +28,6 @@ def signup_action(): action = InterviewAction(metadata={"agent_dir": str(ORCHESTRATOR_AGENT_DIR)}) spec = load_interview_spec_from_skill(SIGNUP_INTERVIEW_SKILL_DIR) action._registry._specs[spec.name] = spec - action._ensure_active_task = AsyncMock() action._close_task = AsyncMock() return action, spec @@ -43,7 +42,6 @@ async def _persist(session, _visitor=None): await save_session(conversation, session) action._save_session = _persist - action._ensure_active_task = AsyncMock() await action.on_skill_activate( "signup_interview", visitor, user_message=visitor.utterance ) diff --git a/tests/action/interview/test_signup_activation_inline.py b/tests/action/interview/test_signup_activation_inline.py index 34c1783f..65e2b1d7 100644 --- a/tests/action/interview/test_signup_activation_inline.py +++ b/tests/action/interview/test_signup_activation_inline.py @@ -32,7 +32,6 @@ def signup_action(): async def test_on_skill_activate_notes_skill_procedure(signup_action): action, _spec = signup_action action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() action._get_conversation = AsyncMock(return_value=None) note = await action.on_skill_activate( @@ -58,7 +57,6 @@ async def test_activation_set_fields_then_model_chains_next_field(signup_action) visitor = SimpleNamespace(conversation=conv, utterance=_OPENING) action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() await action._handle_start("signup_interview", visitor, user_message=_OPENING) @@ -82,7 +80,6 @@ async def test_set_field_idempotent_when_field_already_stored(signup_action): visitor = SimpleNamespace(conversation=conv, utterance=_OPENING) action._save_session = AsyncMock() - action._ensure_active_task = AsyncMock() await action._handle_start("signup_interview", visitor, user_message=_OPENING) diff --git a/tests/action/interview/test_signup_golden_path.py b/tests/action/interview/test_signup_golden_path.py index 42e5eed8..b90d4aaf 100644 --- a/tests/action/interview/test_signup_golden_path.py +++ b/tests/action/interview/test_signup_golden_path.py @@ -24,7 +24,6 @@ def signup_action(): action = InterviewAction(metadata={"agent_dir": str(ORCHESTRATOR_AGENT_DIR)}) spec = load_interview_spec_from_skill(SIGNUP_INTERVIEW_SKILL_DIR) action._registry._specs[spec.name] = spec - action._ensure_active_task = AsyncMock() action._close_task = AsyncMock() return action, spec @@ -45,7 +44,6 @@ async def _persist(session, _visitor=None): await save_session(conversation, session) action._save_session = _persist - action._ensure_active_task = AsyncMock() await action.on_skill_activate( "signup_interview", diff --git a/tests/action/task_trigger/test_task_trigger_schema.py b/tests/action/task_trigger/test_task_trigger_schema.py index aac9e4fe..17ff0ba9 100644 --- a/tests/action/task_trigger/test_task_trigger_schema.py +++ b/tests/action/task_trigger/test_task_trigger_schema.py @@ -37,6 +37,7 @@ def get_tasks(status=None, owner_action=None): async def test_triggers_proactive_task_from_spec_v2(monkeypatch): conversation = MagicMock() + conversation.id = None # In-memory fixture: do not refresh from JsonDB. conversation.tasks = [] conversation.save = AsyncMock() diff --git a/tests/memory/test_task_store_proactive.py b/tests/memory/test_task_store_proactive.py index 04fd3fd7..e0c7a2b9 100644 --- a/tests/memory/test_task_store_proactive.py +++ b/tests/memory/test_task_store_proactive.py @@ -12,6 +12,7 @@ def _make_store(): conv = MagicMock() + conv.id = None # In-memory fixture: no persisted conversation to refresh. conv.tasks = [] conv.save = AsyncMock() return TaskStore(conv), conv diff --git a/tests/wire/conftest.py b/tests/wire/conftest.py index e9c611ee..bb463bb4 100644 --- a/tests/wire/conftest.py +++ b/tests/wire/conftest.py @@ -14,7 +14,7 @@ Each test bootstraps its own app (~1s). That is deliberate: jvspatial objects bind to the event loop that created them, and a session-scoped graph shared across pytest-asyncio's per-test loops produces "attached to different loop" — -see the trap table in the root CLAUDE.md. +see the trap table in the root AGENTS.md. """ from __future__ import annotations @@ -30,6 +30,9 @@ async def wire(tmp_path, monkeypatch): from jvagent.core.app_context import clear_app_root, set_app_root monkeypatch.setenv("JVSPATIAL_ENABLE_DEFERRED_SAVES", "false") + # The wire fixture bootstraps directly, bypassing the CLI's runtime default. + # Preserve the UTF-8 prompt text that production jvagent persists. + monkeypatch.setenv("JVSPATIAL_TEXT_NORMALIZATION_ENABLED", "false") app_root = write_app(tmp_path) set_app_root(app_root) try: