diff --git a/PAA.md b/PAA.md index 1c9de3d..686fbfe 100644 --- a/PAA.md +++ b/PAA.md @@ -18,6 +18,7 @@ where the implementation deliberately stops. | `AutonomyEvent`, `EventStore` | Autonomy event | Uses the contract's fourteen-field event representation and an implementation protocol with explicit ordering, transaction, uniqueness, and append-only semantics. | | `SqliteEventStore` | Evidence/event log substrate | Supplies the default append-only SQLite store, including storage-level update/delete rejection. | | `store_evidence`, `verify_evidence` | Evidence record binding | Content-addresses exact evidence bytes with SHA-256 and fails closed on missing or changed bytes. | +| `OperatingRecord`, `SqliteOperatingRecordStore` | Operating accounting | Stores usage, prices and provenance separately from evidence files and autonomy events, retrievable by subject; no transition rule reads it. | | `import_events` | Archive replay | Imports already validated contract-shaped events without regenerating identifiers or timestamps. The legacy conformance capture proves field and projection continuity across extraction. | | `paa-contracts` conformance suite | Published contract | Checks schema vocabulary, declarations, event histories, evidence addressing, invalid semantic cases, and the pre-cutover capture against the same packaged corpus. | @@ -80,9 +81,26 @@ implementation. ### Worker identity Events record an actor selected explicitly, from a configured environment -variable, or from the OS login. The current contract has no durable worker -identity or attestation field, so an evidence window cannot prove which worker -produced every verdict. That requires a future contract revision. +variable, or from the OS login. This actor is distinct from the optional +evidence-record `worker {id, version, configuration_ref}` and the evaluator-side +`producer`. Worker attribution is consumer-asserted identity, not attestation. +The exact evidence bytes, including worker attribution, are retained. Current +transition logic does not inspect worker identity or operating cost. + +### Future declaration rules — draft only, not implemented + +A later contract revision may declare a single-configuration evidence window: +every qualifying verdict would need explicit matching worker identity and a +configuration reference. Missing attribution would not establish compliance; +evidence from another configuration would not silently count toward the window. + +A separate configuration-change rule may require restoring the declared safer +position before operating a changed configuration. It would need to define +which identity changes trigger it, the fallback when already at the safest +position, window reset behavior, and atomicity with the consumer's configuration +switch. Such a rule must be separately versioned and approved. This revision +adds no declaration fields or automatic demotion, and does not claim to enforce +"no inherited autonomy." Cost would remain outside transition eligibility. ### Governed-effect atomicity with the default store diff --git a/README.md b/README.md index 008957e..ab4a44e 100644 --- a/README.md +++ b/README.md @@ -11,12 +11,14 @@ PAA is implementation-neutral by construction — it describes *what* a governed - **Position resolution** — current autonomy position is never stored. It is folded fresh from the declaration's `initial_position` plus the latest exact-scope `position_changed` event. - **Evidence binding** — every motion binds to the exact bytes of its evidence artifact by SHA-256, re-verified at approval. Tamper or loss is a fail-closed error. - **An append-only event store** — append-only enforced by the storage layer, not by convention. +- **Operating-record storage** — a separate optional append-only store preserves subject-linked usage, prices, worker configurations, and source attribution. It has no role in resolving authority. ## What it does not do - **Produce evaluator verdicts.** The runtime governs; consumers evaluate. Which evaluators exist and what code produces each verdict is consumer domain data, supplied as a registry. - **Evaluate promotion rules.** Thresholds and windows are *declared*, not machine-evaluated. Approval is an operator judgment. -- **Carry worker identity.** The contract has no worker-identity field yet, so evidence windows cannot prove which worker produced them. Tracked for a later contract cycle. +- **Enforce worker-configuration policies.** Evidence records can carry optional worker attribution; evidence storage preserves those bytes. Single-configuration windows and configuration-change demotion remain future declaration rules, not implemented behavior. +- **Calculate or govern by cost.** Consumers produce prices, retain constituent usage, distinguish actuals from estimates, assess coverage, and reconcile overlapping summaries. Neither operating records nor worker attribution changes transition eligibility in this revision. ## Install @@ -122,7 +124,7 @@ The conformance suite runs against the published contract artifacts rather than fixtures of its own, so that "passes the published conformance suite" is a claim about the contract and not about this repo's idea of it. -Those artifacts come from `paa-contracts` — the four normative schemas, the +Those artifacts come from `paa-contracts` — the five normative schemas, the positive fixture corpus, and the invalid-case tables. It is the other package this repo publishes, and a workspace member here, so the extra is all it takes: @@ -186,6 +188,61 @@ release would carry the corpus. ## Status -`0.3.0`, tracking the `paa-task/0.2.1-draft` and `paa-autonomy-event/0.1.0-draft` schema families. Package and spec versions drift independently — the schema families a release targets are stated here and asserted by the conformance suite, not inferred from the package version. +`0.4.0` adds optional operating-record storage against `paa-operating-record/0.1.0-draft` and preserves optional worker attribution in `paa-evidence-record/0.2.0-draft`. Existing `paa-evidence-record/0.1.0-draft` records without worker attribution remain valid. No evidence files or event histories require migration. Task and event families remain `paa-task/0.2.1-draft` and `paa-autonomy-event/0.1.0-draft`. Package and schema-family versions are independent; this PR prepares release versions but does not publish packages. -`0.3.0` is a breaking change to the declaration access layer, and it is the change that made the sentence above true. `0.2.0` claimed the `paa-task/0.2.1-draft` family while implementing an older evaluator identity — a single `oracle` field where the contract has `evaluation_basis` and `epistemic_status` — and a `position_policy` requiring all four positions at fixed modes, where the contract admits any non-empty subset with per-evaluator placement overrides. It could not load a single published declaration. Building the conformance suite is what surfaced that; `PaaEvaluator`, `ProducerRegistration`, and `PaaPositionPolicy` changed shape to fix it. Nothing was published at `0.2.0`, so no consumer is stranded. +`0.3.0` introduced a breaking change to the declaration access layer to align it with the published task family. `0.2.0` claimed the `paa-task/0.2.1-draft` family while implementing an older evaluator identity — a single `oracle` field where the contract has `evaluation_basis` and `epistemic_status` — and a `position_policy` requiring all four positions at fixed modes, where the contract admits any non-empty subset with per-evaluator placement overrides. It could not load a single published declaration. Building the conformance suite is what surfaced that; `PaaEvaluator`, `ProducerRegistration`, and `PaaPositionPolicy` changed shape to fix it. Nothing was published at `0.2.0`, so no consumer is stranded. + +## Optional operating records + +Evidence remains content-addressed files. Autonomy events remain in `EventStore`. +`SqliteOperatingRecordStore` adds an independent `operating_records` table and +connection; it may use the same database path or a separate database. It does +not extend the `EventStore` protocol or participate in motion transactions. + +```python +from pathlib import Path +from paa_runtime import SqliteOperatingRecordStore, decode_operating_record + +record = decode_operating_record(Path("operating-record.json").read_bytes()) +accounting = SqliteOperatingRecordStore(Path("paa_runtime.db")) +try: + accounting.append(record) + records = accounting.get_by_subject(record["subject"]) +finally: + accounting.close() +``` + +An append validates structure and commits one record atomically. Reusing a +record ID is an error, including an identical retry of an already committed +write; read by subject to reconcile an uncertain write before retrying. Reads +return detached records in insertion order. Subject kind and ID both match +exactly; each result retains task, declaration version, scope, and worker +configuration. Multiple tasks, attempts, and summaries may share a subject. +No update/delete API exists; SQLite triggers reject updates, deletes, and +replacement inserts. As with the event store, this is not protection against +an administrator dropping triggers or rewriting the database. + +`usage` keys are open, nonnegative quantities; recommended names include +`input_tokens`, `output_tokens`, `cached_tokens`, and `llm_calls`. Null usage, +individual quantities, or price explicitly means unavailable. Zero is a real +measurement. A price requires `currency`, `amount`, and opaque `basis` together. +Optional components carry additional costs **not already included** in the +base price. Omitted components make no completeness claim. Component kinds and +currencies are consumer-defined; the runtime neither sums nor converts them. + +Source references must identify the attributed work and attempts, including +failures and superseded retries, and preserve constituent quantities, +model/rate identities, pricing bases, and measured/estimated coverage. Pipeline +summaries are allowed; a catalog reference and aggregate tokens alone cannot +reprice mixed-model work. Do not sum pipeline totals and their constituent +task records as independent charges. Multiple verdicts about one output do +not create more charges. The runtime validates references structurally, not +their external contents or accounting accuracy. + +Readers compute effective cost for one configuration, population, and window: +attributed costs of **all** attempts divided by distinct accepted outcomes in +that same population. Acceptance rules belong to the task/consumer. Zero +accepted outcomes means undefined effective cost, not zero; missing prices +cannot support an unqualified total-cost claim. Effective cost is not a stored +field, and operating records cannot grant authority or offset a behavioral +failure. diff --git a/conformance/_corpus.py b/conformance/_corpus.py index 564b680..fb2b9e0 100644 --- a/conformance/_corpus.py +++ b/conformance/_corpus.py @@ -22,6 +22,7 @@ #: Which schema governs each case table's documents. SCHEMA_FOR_KIND: dict[str, str] = { + "operating": "paa-operating-record", "task": "paa-task", "evidence": "paa-evidence-record", "decision": "paa-decision-artifact", diff --git a/conformance/test_corpus_integrity.py b/conformance/test_corpus_integrity.py index 4585279..392372a 100644 --- a/conformance/test_corpus_integrity.py +++ b/conformance/test_corpus_integrity.py @@ -51,6 +51,7 @@ ] PUBLISHED_FIXTURES = [ + *(("operating", p) for p in contracts.operating_record_paths()), *(("task", p) for p in contracts.task_declaration_paths()), *(("event", p) for p in contracts.autonomy_event_paths()), *(("evidence", p) for p in contracts.evidence_record_paths()), @@ -70,17 +71,17 @@ class TestTheCorpusIsWhatItClaims: guard against that. """ - def test_the_case_tables_carry_ninety_five_cases(self) -> None: - assert len(ALL_CASES) == 95 + def test_the_case_table_count_is_pinned(self) -> None: + assert len(ALL_CASES) == 147 - def test_fifty_of_them_are_structural(self) -> None: - assert len(STRUCTURAL_CASES) == 50 + def test_the_structural_case_count_is_pinned(self) -> None: + assert len(STRUCTURAL_CASES) == 102 def test_fifteen_of_them_are_pinned(self) -> None: assert len(PINNED_CASES) == 15 def test_every_published_fixture_is_discoverable(self) -> None: - assert len(PUBLISHED_FIXTURES) == 17 + assert len(PUBLISHED_FIXTURES) == 22 class TestFormatAssertionIsLive: diff --git a/conformance/test_evidence_integrity.py b/conformance/test_evidence_integrity.py index 0fdf29b..1b7524c 100644 --- a/conformance/test_evidence_integrity.py +++ b/conformance/test_evidence_integrity.py @@ -41,7 +41,7 @@ def _sha_from_path(path: Path) -> str: class TestCorpusIsPresent: def test_evidence_records_are_discoverable(self) -> None: - assert len(contracts.evidence_record_paths()) == 3 + assert len(contracts.evidence_record_paths()) == 4 def test_decision_artifacts_are_discoverable(self) -> None: assert len(contracts.decision_artifact_paths()) == 5 diff --git a/conformance/test_operating_records.py b/conformance/test_operating_records.py new file mode 100644 index 0000000..12513e1 --- /dev/null +++ b/conformance/test_operating_records.py @@ -0,0 +1,94 @@ +"""Published operating records round-trip without becoming policy inputs.""" + +from __future__ import annotations + +import json +from pathlib import Path + +import paa_contracts as contracts +import pytest +from jsonschema import Draft7Validator + +from conformance._corpus import case_documents, violations +from paa_runtime import RuntimeConfig, SqliteEventStore, approve, propose, show +from paa_runtime.operating import ( + CURRENT_OPERATING_SCHEMA, + OperatingRecordError, + decode_operating_record, +) +from paa_runtime.operating_store import SqliteOperatingRecordStore + + +def test_operating_schema_stamp_matches_contract() -> None: + assert contracts.schema_version("paa-operating-record") == CURRENT_OPERATING_SCHEMA + + +@pytest.mark.parametrize("schema_id", ["paa-operating-record", "paa-evidence-record"]) +def test_revised_contract_is_a_valid_draft7_schema(schema_id: contracts.SchemaId) -> None: + Draft7Validator.check_schema(contracts.load_schema(schema_id)) + + +def test_worker_definitions_agree() -> None: + operating = contracts.load_schema("paa-operating-record")["definitions"]["worker"] + evidence = dict(contracts.load_schema("paa-evidence-record")["definitions"]["worker"]) + evidence.pop("description") + assert operating == evidence + + +@pytest.mark.parametrize("path", contracts.operating_record_paths(), ids=lambda path: path.name) +def test_operating_fixture_round_trips(path: Path, tmp_path: Path) -> None: + record = decode_operating_record(path.read_bytes()) + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + assert store.get_by_subject(record["subject"]) == (json.loads(path.read_bytes()),) + finally: + store.close() + + +@pytest.mark.parametrize("case", contracts.invalid_cases("operating"), ids=lambda case: case["id"]) +def test_decoder_rejects_published_invalid_case(case: contracts.InvalidCase) -> None: + _, mutated = case_documents("operating", case) + with pytest.raises(OperatingRecordError): + decode_operating_record(json.dumps(mutated).encode()) + + +def test_current_evidence_without_worker_remains_valid() -> None: + record = json.loads(contracts.evidence_record_paths()[0].read_bytes()) + record["record_schema"] = "paa-evidence-record/0.2.0-draft" + record.pop("worker", None) + assert violations("evidence", record) == () + + +def test_operating_records_do_not_change_motion_outcomes( + runtime_config: RuntimeConfig, tmp_path: Path, +) -> None: + """Same database, deliberately changed cost/worker; approval still independent.""" + events = SqliteEventStore(runtime_config.db_path) + operating = SqliteOperatingRecordStore(runtime_config.db_path) + evidence = tmp_path / "report.json" + evidence.write_bytes(b'{"operator_report": "illustrative"}') + try: + motion = propose( + events, runtime_config, task="outbound_content_publish", scope="publish:farcaster", + to_position="hotl", evidence_path=evidence, actor="test:operator", + ) + before = show( + events, runtime_config, task="outbound_content_publish", scope="publish:farcaster", + ) + for index, path in enumerate(contracts.operating_record_paths()): + record = decode_operating_record(path.read_bytes()) + record.update(task="outbound_content_publish", scope="publish:farcaster") + record["worker"]["configuration_ref"] = f"cfg:changed-{index}" + operating.append(record) + assert show( + events, runtime_config, task="outbound_content_publish", scope="publish:farcaster", + ) == before + approve( + events, runtime_config, motion_id=motion.motion_id, + actor="test:operator", reason="independent operator approval", + ) + assert events.get_autonomy_events()[-1].to_position == "hotl" + finally: + operating.close() + events.close() diff --git a/examples/runtime-conformance/evidence-records/evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json b/examples/runtime-conformance/evidence-records/evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json new file mode 100644 index 0000000..cd2f3ab --- /dev/null +++ b/examples/runtime-conformance/evidence-records/evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json @@ -0,0 +1 @@ +{"boundary":{"input_ref":"app://inbound_candidate/1001","output_ref":"app://surfaced_event/1001"},"declaration_version":1,"evaluator":{"authority":"advisory","epistemic_status":"proxy","evaluation_basis":{"kind":"rubric","ref":"response_quality_rubric"},"property":"response_quality","target":"output","technique":"llm_judge","version":"1"},"payload":{"model":"critic-v1","rubric_id":"response_quality_rubric_v1","score":0.86},"payload_schema":"https://paa.dev/payload-schemas/response-quality-llm.schema.json","producer":{"id":"llm-critic","version":"1"},"record_id":"evrec-inbound-case-1001-worker","record_schema":"paa-evidence-record/0.2.0-draft","scope":null,"source_references":["critiques:5501"],"subject":{"id":"case-1001","kind":"case"},"task":"inbound_reply_surfacing","timestamps":{"completed_at":"2026-01-05T12:00:02.000Z","recorded_at":"2026-01-05T12:00:02.500Z","started_at":"2026-01-05T12:00:00.000Z"},"verdict":{"reason_codes":["tone_within_bounds","no_hallucinated_claim"],"value":"approve"},"worker":{"id":"reply-drafter","version":"response-v7","configuration_ref":"cfg:reply-drafter-v7"}} diff --git a/examples/runtime-conformance/invalid/evidence-cases.json b/examples/runtime-conformance/invalid/evidence-cases.json index f8803c0..bbfc653 100644 --- a/examples/runtime-conformance/invalid/evidence-cases.json +++ b/examples/runtime-conformance/invalid/evidence-cases.json @@ -2,18 +2,31 @@ { "id": "evidence_missing_producer", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "remove", "path": "/producer" }], + "mutations": [ + { + "kind": "remove", + "path": "/producer" + } + ], "expected": { "stage": "structural", "code": "required", "path": "", - "params": { "missingProperty": "producer" } + "params": { + "missingProperty": "producer" + } } }, { "id": "evidence_invalid_started_at", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "set", "path": "/timestamps/started_at", "value": "not-a-date" }], + "mutations": [ + { + "kind": "set", + "path": "/timestamps/started_at", + "value": "not-a-date" + } + ], "expected": { "stage": "structural", "code": "format", @@ -23,7 +36,13 @@ { "id": "evidence_invalid_task_identifier", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "set", "path": "/task", "value": "Inbound Reply" }], + "mutations": [ + { + "kind": "set", + "path": "/task", + "value": "Inbound Reply" + } + ], "expected": { "stage": "structural", "code": "pattern", @@ -33,7 +52,13 @@ { "id": "evidence_undeclared_scope", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "set", "path": "/scope", "value": "publish:bluesky" }], + "mutations": [ + { + "kind": "set", + "path": "/scope", + "value": "publish:bluesky" + } + ], "expected": { "stage": "evidence_semantic", "code": "evidence.undeclared_scope", @@ -43,7 +68,13 @@ { "id": "evidence_unregistered_declaration_version", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "set", "path": "/declaration_version", "value": 99 }], + "mutations": [ + { + "kind": "set", + "path": "/declaration_version", + "value": 99 + } + ], "expected": { "stage": "evidence_semantic", "code": "evidence.unregistered_declaration_version", @@ -53,11 +84,158 @@ { "id": "evidence_evaluator_identity_mismatch", "base": "evidence/paa/cd85082f0ef66db18183acd3e350668078d31cbb6ff68683a9efcb552157662c/evidence.json", - "mutations": [{ "kind": "set", "path": "/evaluator/version", "value": "99" }], + "mutations": [ + { + "kind": "set", + "path": "/evaluator/version", + "value": "99" + } + ], "expected": { "stage": "evidence_semantic", "code": "evidence.evaluator_identity_mismatch", "path": "/evaluator" } + }, + { + "id": "evidence_worker_missing_id", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/id" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "evidence_worker_empty_id", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/id", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/id" + } + }, + { + "id": "evidence_worker_missing_version", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/version" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "evidence_worker_empty_version", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/version", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/version" + } + }, + { + "id": "evidence_worker_missing_configuration_ref", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/configuration_ref" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "evidence_worker_empty_configuration_ref", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/configuration_ref", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/configuration_ref" + } + }, + { + "id": "evidence_worker_null", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/worker", + "value": null + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/worker" + } + }, + { + "id": "evidence_worker_unknown_field", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/model", + "value": "opaque-model" + } + ], + "expected": { + "stage": "structural", + "code": "additionalProperties", + "path": "/worker" + } + }, + { + "id": "evidence_worker_old_version", + "base": "evidence/paa/421d1e7672aa1fa96ff2426a4b242d711a25cf23f27160accabcca25abc515fa/evidence.json", + "mutations": [ + { + "kind": "set", + "path": "/record_schema", + "value": "paa-evidence-record/0.1.0-draft" + } + ], + "expected": { + "stage": "structural", + "code": "not", + "path": "" + } } ] diff --git a/examples/runtime-conformance/invalid/operating-cases.json b/examples/runtime-conformance/invalid/operating-cases.json new file mode 100644 index 0000000..e36e304 --- /dev/null +++ b/examples/runtime-conformance/invalid/operating-cases.json @@ -0,0 +1,707 @@ +[ + { + "id": "operating_missing_record_schema", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/record_schema" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_record_id", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/record_id" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_task", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/task" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_declaration_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/declaration_version" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_scope", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/scope" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_subject", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/subject" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_worker", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_usage", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/usage" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_price", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/price" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_timestamps", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/timestamps" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_missing_source_references", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/source_references" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "" + } + }, + { + "id": "operating_worker_missing_id", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/id" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "operating_worker_empty_id", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/id", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/id" + } + }, + { + "id": "operating_worker_missing_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/version" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "operating_worker_empty_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/version", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/version" + } + }, + { + "id": "operating_worker_missing_configuration_ref", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/worker/configuration_ref" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/worker" + } + }, + { + "id": "operating_worker_empty_configuration_ref", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/worker/configuration_ref", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/worker/configuration_ref" + } + }, + { + "id": "operating_price_missing_basis", + "base": "task-priced.json", + "mutations": [ + { + "kind": "remove", + "path": "/price/basis" + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/price" + } + }, + { + "id": "operating_price_empty_basis", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/price/basis", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/price/basis" + } + }, + { + "id": "operating_negative_price", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/price/amount", + "value": -1 + } + ], + "expected": { + "stage": "structural", + "code": "minimum", + "path": "/price/amount" + } + }, + { + "id": "operating_unavailable_amount", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/price/amount", + "value": null + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/price/amount" + } + }, + { + "id": "operating_invalid_usage", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/usage/input_tokens", + "value": "many" + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/usage/input_tokens" + } + }, + { + "id": "operating_negative_usage", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/usage/input_tokens", + "value": -1 + } + ], + "expected": { + "stage": "structural", + "code": "minimum", + "path": "/usage/input_tokens" + } + }, + { + "id": "operating_empty_usage", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/usage", + "value": {} + } + ], + "expected": { + "stage": "structural", + "code": "minProperties", + "path": "/usage" + } + }, + { + "id": "operating_boolean_usage", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/usage/input_tokens", + "value": true + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/usage/input_tokens" + } + }, + { + "id": "operating_empty_usage_key", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/usage", + "value": { + "": 1 + } + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/usage" + } + }, + { + "id": "operating_unknown_field", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/effective_cost", + "value": 0.1 + } + ], + "expected": { + "stage": "structural", + "code": "additionalProperties", + "path": "" + } + }, + { + "id": "operating_invalid_subject", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/subject/kind", + "value": "pipeline" + } + ], + "expected": { + "stage": "structural", + "code": "enum", + "path": "/subject/kind" + } + }, + { + "id": "operating_invalid_task", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/task", + "value": "Bad Task" + } + ], + "expected": { + "stage": "structural", + "code": "pattern", + "path": "/task" + } + }, + { + "id": "operating_empty_scope", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/scope", + "value": "" + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/scope" + } + }, + { + "id": "operating_invalid_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/declaration_version", + "value": 0 + } + ], + "expected": { + "stage": "structural", + "code": "minimum", + "path": "/declaration_version" + } + }, + { + "id": "operating_fractional_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/declaration_version", + "value": 1.5 + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/declaration_version" + } + }, + { + "id": "operating_boolean_version", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/declaration_version", + "value": true + } + ], + "expected": { + "stage": "structural", + "code": "type", + "path": "/declaration_version" + } + }, + { + "id": "operating_unsupported_schema", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/record_schema", + "value": "paa-operating-record/99" + } + ], + "expected": { + "stage": "structural", + "code": "const", + "path": "/record_schema" + } + }, + { + "id": "operating_invalid_timestamp", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/timestamps/started_at", + "value": "not-a-date" + } + ], + "expected": { + "stage": "structural", + "code": "format", + "path": "/timestamps/started_at" + } + }, + { + "id": "operating_missing_timezone", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/timestamps/recorded_at", + "value": "2026-01-05T12:00:02" + } + ], + "expected": { + "stage": "structural", + "code": "format", + "path": "/timestamps/recorded_at" + } + }, + { + "id": "operating_invalid_calendar_date", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/timestamps/completed_at", + "value": "2026-02-30T12:00:02Z" + } + ], + "expected": { + "stage": "structural", + "code": "format", + "path": "/timestamps/completed_at" + } + }, + { + "id": "operating_empty_references", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/source_references", + "value": [] + } + ], + "expected": { + "stage": "structural", + "code": "minItems", + "path": "/source_references" + } + }, + { + "id": "operating_duplicate_references", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/source_references", + "value": [ + "a", + "a" + ] + } + ], + "expected": { + "stage": "structural", + "code": "uniqueItems", + "path": "/source_references" + } + }, + { + "id": "operating_empty_reference", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/source_references", + "value": [ + "" + ] + } + ], + "expected": { + "stage": "structural", + "code": "minLength", + "path": "/source_references/0" + } + }, + { + "id": "operating_component_missing_basis", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/components", + "value": [ + { + "kind": "review", + "quantity": 1, + "unit": "minute", + "price": { + "currency": "USD", + "amount": 1 + } + } + ] + } + ], + "expected": { + "stage": "structural", + "code": "required", + "path": "/components/0/price" + } + }, + { + "id": "operating_component_negative_quantity", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/components", + "value": [ + { + "kind": "review", + "quantity": -1, + "unit": "minute", + "price": null + } + ] + } + ], + "expected": { + "stage": "structural", + "code": "minimum", + "path": "/components/0/quantity" + } + }, + { + "id": "operating_component_unknown_field", + "base": "task-priced.json", + "mutations": [ + { + "kind": "set", + "path": "/components", + "value": [ + { + "kind": "review", + "quantity": 1, + "unit": "minute", + "price": null, + "extra": 1 + } + ] + } + ], + "expected": { + "stage": "structural", + "code": "additionalProperties", + "path": "/components/0" + } + } +] diff --git a/examples/runtime-conformance/operating-records/failed-unavailable.json b/examples/runtime-conformance/operating-records/failed-unavailable.json new file mode 100644 index 0000000..a9b18ba --- /dev/null +++ b/examples/runtime-conformance/operating-records/failed-unavailable.json @@ -0,0 +1,26 @@ +{ + "record_schema": "paa-operating-record/0.1.0-draft", + "record_id": "op-case-1001-failed-attempt-2", + "task": "inbound_reply_surfacing", + "declaration_version": 1, + "scope": null, + "subject": { + "kind": "case", + "id": "case-1001" + }, + "worker": { + "id": "reply-drafter", + "version": "response-v7", + "configuration_ref": "cfg:reply-drafter-v7" + }, + "usage": null, + "price": null, + "timestamps": { + "started_at": "2026-01-05T12:00:00.000Z", + "completed_at": "2026-01-05T12:00:02.000Z", + "recorded_at": "2026-01-05T12:00:02.500Z" + }, + "source_references": [ + "example://usage/case-1001/failed-attempt-2" + ] +} diff --git a/examples/runtime-conformance/operating-records/pipeline-summary.json b/examples/runtime-conformance/operating-records/pipeline-summary.json new file mode 100644 index 0000000..9de6ecb --- /dev/null +++ b/examples/runtime-conformance/operating-records/pipeline-summary.json @@ -0,0 +1,54 @@ +{ + "record_schema": "paa-operating-record/0.1.0-draft", + "record_id": "op-run-2001-summary", + "task": "inbound_reply_surfacing", + "declaration_version": 1, + "scope": null, + "subject": { + "kind": "run", + "id": "run-2001" + }, + "worker": { + "id": "reply-pipeline", + "version": "v3", + "configuration_ref": "cfg:all-models-prompts-routing-evaluators-v3" + }, + "usage": { + "input_tokens": 2000, + "output_tokens": 500, + "tool_seconds": null + }, + "price": { + "currency": "USD", + "amount": 0.032, + "basis": "example-catalog:v1" + }, + "timestamps": { + "started_at": "2026-01-05T12:00:00.000Z", + "completed_at": "2026-01-05T12:00:02.000Z", + "recorded_at": "2026-01-05T12:00:02.500Z" + }, + "source_references": [ + "example://usage/run-2001/draft/attempt-1", + "example://usage/run-2001/critic/attempt-1", + "example://review/run-2001" + ], + "components": [ + { + "kind": "human_review", + "quantity": 3.5, + "unit": "minutes", + "price": { + "currency": "USD", + "amount": 4.08, + "basis": "example-review-rate:2026q3" + } + }, + { + "kind": "tool_execution", + "quantity": null, + "unit": "seconds", + "price": null + } + ] +} diff --git a/examples/runtime-conformance/operating-records/task-priced.json b/examples/runtime-conformance/operating-records/task-priced.json new file mode 100644 index 0000000..892d6ee --- /dev/null +++ b/examples/runtime-conformance/operating-records/task-priced.json @@ -0,0 +1,35 @@ +{ + "record_schema": "paa-operating-record/0.1.0-draft", + "record_id": "op-case-1001-attempt-1", + "task": "inbound_reply_surfacing", + "declaration_version": 1, + "scope": null, + "subject": { + "kind": "case", + "id": "case-1001" + }, + "worker": { + "id": "reply-drafter", + "version": "response-v7", + "configuration_ref": "cfg:reply-drafter-v7" + }, + "usage": { + "input_tokens": 1840, + "output_tokens": 412, + "cached_tokens": 0, + "llm_calls": 2 + }, + "price": { + "currency": "USD", + "amount": 0.0117, + "basis": "example-catalog:v1" + }, + "timestamps": { + "started_at": "2026-01-05T12:00:00.000Z", + "completed_at": "2026-01-05T12:00:02.000Z", + "recorded_at": "2026-01-05T12:00:02.500Z" + }, + "source_references": [ + "example://usage/case-1001/attempt-1" + ] +} diff --git a/examples/runtime-conformance/operating-records/zero-price.json b/examples/runtime-conformance/operating-records/zero-price.json new file mode 100644 index 0000000..9659bb1 --- /dev/null +++ b/examples/runtime-conformance/operating-records/zero-price.json @@ -0,0 +1,35 @@ +{ + "record_schema": "paa-operating-record/0.1.0-draft", + "record_id": "op-case-1002-zero", + "task": "inbound_reply_surfacing", + "declaration_version": 1, + "scope": null, + "subject": { + "kind": "case", + "id": "case-1002" + }, + "worker": { + "id": "reply-drafter", + "version": "response-v7", + "configuration_ref": "cfg:reply-drafter-v7" + }, + "usage": { + "input_tokens": 0, + "output_tokens": 0, + "custom_measure": 0.5 + }, + "price": { + "currency": "USD", + "amount": 0, + "basis": "example-free-tier:v1" + }, + "timestamps": { + "started_at": "2026-01-05T12:00:00.000Z", + "completed_at": "2026-01-05T12:00:02.000Z", + "recorded_at": "2026-01-05T12:00:02.500Z" + }, + "source_references": [ + "example://usage/case-1002/attempt-1" + ], + "components": [] +} diff --git a/package.json b/package.json index b3eafef..7769dd5 100644 --- a/package.json +++ b/package.json @@ -1,7 +1,7 @@ { "name": "@paadev/paa-contracts", - "version": "0.1.0", - "description": "Published contract artifacts of the Progressive Autonomy Architecture: four normative JSON Schemas, the positive fixture corpus every implementation is checked against, and the table-driven invalid-case matrices.", + "version": "0.2.0", + "description": "Published contract artifacts of the Progressive Autonomy Architecture: five normative JSON Schemas, the positive fixture corpus every implementation is checked against, and the table-driven invalid-case matrices.", "license": "MIT", "homepage": "https://www.paa.dev", "repository": { diff --git a/packages/paa-contracts/README.md b/packages/paa-contracts/README.md index d046dfc..949d521 100644 --- a/packages/paa-contracts/README.md +++ b/packages/paa-contracts/README.md @@ -1,6 +1,6 @@ # paa-contracts -The published contract artifacts of the [Progressive Autonomy Architecture](https://www.paa.dev): four normative JSON Schemas, the positive fixture corpus every implementation is checked against, and the table-driven invalid-case matrices. +The published contract artifacts of the [Progressive Autonomy Architecture](https://www.paa.dev): five normative JSON Schemas, the positive fixture corpus every implementation is checked against, and the table-driven invalid-case matrices. No runtime logic, no dependencies. This package is data and honest paths to it. @@ -85,9 +85,9 @@ The consequence: a wheel can only be built from a full checkout of this repo. Th | Path | What | |---|---| -| `schemas/` | `paa-task`, `paa-evidence-record`, `paa-decision-artifact`, `paa-autonomy-event` | +| `schemas/` | `paa-task`, `paa-evidence-record`, `paa-decision-artifact`, `paa-autonomy-event`, `paa-operating-record` | | `examples/paa-tasks/` | four valid declarations + 63 invalid cases | -| `examples/runtime-conformance/` | evidence records, decision artifacts, autonomy-event sequences, payload companion schemas, 32 invalid cases | +| `examples/runtime-conformance/` | evidence and operating records, decision artifacts, autonomy-event sequences, payload companion schemas, 84 invalid cases | | `examples/runtime-conformance/invalid/fixtures/tampered-evidence/` | a deliberately byte-mismatched artifact, so tamper detection has something real to fail on | ## Versioning @@ -100,10 +100,38 @@ contracts.schema_version("paa-autonomy-event") # 'paa-autonomy-event/0.1.0-dra ## Development +### Contract changes in 0.2.0 + +`paa-evidence-record/0.2.0-draft` admits optional `worker` attribution. When +present, all of `id`, `version`, and opaque `configuration_ref` are required +and nonempty. `producer` continues to identify the evaluator side, not the +worker. The schema also accepts existing `0.1.0-draft` records without worker +attribution; adding worker attribution requires the new record stamp. Existing +content-addressed fixtures are retained byte-for-byte. + +`paa-operating-record/0.1.0-draft` is a new optional artifact. Its task, version, +scope, subject and worker identify the accounting boundary, with constituent +attempt attribution carried by durable `source_references`. Usage and prices +may be explicitly null when unavailable. Usage keys and component kinds are +open vocabulary. Component prices require the same currency/amount/basis +shape as the base price, and must not repeat costs already included there. +Measured/estimated coverage and repricing detail must remain in referenced +source artifacts; schema conformance does not verify those external claims. + +`operating_record_paths()` discovers four illustrative positive fixtures; +`invalid_cases("operating")` provides 43 structural negative cases. The +pipeline example's references stand for constituent usage sources, not actual +production measurements. Consumers must resolve their own durable sources, +reconcile overlapping summaries, retain failed-attempt charges, and compute +any cost-per-accepted-outcome metric over matching populations. Neither cost +nor worker identity is consumed by runtime transition rules in this revision. + +### Commands + Released from this repo, where the schemas and fixtures live alongside the reference implementation that is measured against them. ```bash -uv run pytest # includes registry parity against contract-registry.mjs +uv run pytest # includes registry parity against schemas on disk uv run ruff check src tests uv run mypy src/paa_contracts uv build # must run from a full repo checkout diff --git a/packages/paa-contracts/pyproject.toml b/packages/paa-contracts/pyproject.toml index c28caca..e93db6e 100644 --- a/packages/paa-contracts/pyproject.toml +++ b/packages/paa-contracts/pyproject.toml @@ -1,7 +1,7 @@ [project] name = "paa-contracts" -version = "0.1.0" -description = "Published Progressive Autonomy Architecture contract artifacts: the four normative JSON Schemas, the positive fixture corpus, and the invalid-case tables" +version = "0.2.0" +description = "Published Progressive Autonomy Architecture contract artifacts: the five normative JSON Schemas, the positive fixture corpus, and the invalid-case tables" readme = "README.md" requires-python = ">=3.12" license = "MIT" diff --git a/packages/paa-contracts/src/paa_contracts/__init__.py b/packages/paa-contracts/src/paa_contracts/__init__.py index 2a85291..2a4bd48 100644 --- a/packages/paa-contracts/src/paa_contracts/__init__.py +++ b/packages/paa-contracts/src/paa_contracts/__init__.py @@ -1,6 +1,6 @@ """Published Progressive Autonomy Architecture contract artifacts. -This package is the contract side of the PAA split: the four normative JSON +This package is the contract side of the PAA split: the five normative JSON Schemas, the positive fixture corpus every implementation is checked against, and the table-driven invalid-case matrices. It contains no runtime logic and has no dependencies — it is data, plus honest paths to that data. @@ -25,7 +25,7 @@ from pathlib import Path from typing import Any, Literal, NotRequired, TypedDict -__version__ = "0.1.0" +__version__ = "0.2.0" class ContractsUnavailableError(RuntimeError): @@ -38,6 +38,7 @@ class ContractsUnavailableError(RuntimeError): SchemaId = Literal[ + "paa-operating-record", "paa-task", "paa-evidence-record", "paa-decision-artifact", @@ -50,17 +51,18 @@ class ContractsUnavailableError(RuntimeError): #: in other languages keep their own list and check it the same way, against #: the artifacts rather than against this one. SCHEMA_IDS: tuple[SchemaId, ...] = ( + "paa-operating-record", "paa-task", "paa-evidence-record", "paa-decision-artifact", "paa-autonomy-event", ) -CaseKind = Literal["task", "evidence", "decision", "event"] +CaseKind = Literal["task", "evidence", "decision", "event", "operating"] -#: The invalid-case tables. Every case in all four shares one shape, which is +#: The invalid-case tables. Every case shares one shape, which is #: what lets a single accessor serve all of them. -CASE_KINDS: tuple[CaseKind, ...] = ("task", "evidence", "decision", "event") +CASE_KINDS: tuple[CaseKind, ...] = ("task", "evidence", "decision", "event", "operating") # One edit applied to a positive fixture to produce an invalid document. @@ -203,6 +205,7 @@ def _resolve_data_root() -> tuple[Path, Literal["packaged", "worktree"]]: ) _CASE_TABLES: Mapping[CaseKind, Path] = { + "operating": RUNTIME_FIXTURES_ROOT / "invalid" / "operating-cases.json", "task": TASK_FIXTURES_ROOT / "invalid" / "cases.json", "evidence": RUNTIME_FIXTURES_ROOT / "invalid" / "evidence-cases.json", "decision": RUNTIME_FIXTURES_ROOT / "invalid" / "decision-cases.json", @@ -213,6 +216,7 @@ def _resolve_data_root() -> tuple[Path, Literal["packaged", "worktree"]]: #: plain filenames; evidence and decision bases are ``evidence/paa//…`` #: refs relative to their content-addressed root. _CASE_BASE_ROOTS: Mapping[CaseKind, Path] = { + "operating": RUNTIME_FIXTURES_ROOT / "operating-records", "task": TASK_FIXTURES_ROOT, "evidence": RUNTIME_FIXTURES_ROOT / "evidence-records", "decision": RUNTIME_FIXTURES_ROOT / "decision-artifacts", @@ -276,6 +280,11 @@ def autonomy_event_paths() -> tuple[Path, ...]: return _sorted_files(RUNTIME_FIXTURES_ROOT / "autonomy-events", ".json") +def operating_record_paths() -> tuple[Path, ...]: + """Valid operating records, including summaries and unavailable costs.""" + return _sorted_files(RUNTIME_FIXTURES_ROOT / "operating-records", ".json") + + def evidence_record_paths() -> tuple[Path, ...]: """The valid content-addressed evidence records.""" return _content_addressed_files(RUNTIME_FIXTURES_ROOT / "evidence-records") @@ -370,6 +379,7 @@ def resolve_case_base(kind: CaseKind, case: InvalidCase) -> Path: "case_stages", "decision_artifact_paths", "evidence_record_paths", + "operating_record_paths", "invalid_cases", "load_schema", "payload_schema_paths", diff --git a/packages/paa-contracts/tests/test_contracts.py b/packages/paa-contracts/tests/test_contracts.py index 6dc7e99..b13270a 100644 --- a/packages/paa-contracts/tests/test_contracts.py +++ b/packages/paa-contracts/tests/test_contracts.py @@ -52,8 +52,8 @@ def test_every_published_root_resolves(self, root: Path) -> None: class TestSchemas: - def test_all_four_contracts_are_present(self) -> None: - assert len(contracts.SCHEMA_IDS) == 4 + def test_all_five_contracts_are_present(self) -> None: + assert len(contracts.SCHEMA_IDS) == 5 @pytest.mark.parametrize("schema_id", contracts.SCHEMA_IDS) def test_schema_file_exists(self, schema_id: str) -> None: @@ -81,7 +81,8 @@ class TestPositiveFixtures: [ (contracts.task_declaration_paths, 4), (contracts.autonomy_event_paths, 5), - (contracts.evidence_record_paths, 3), + (contracts.evidence_record_paths, 4), + (contracts.operating_record_paths, 4), (contracts.decision_artifact_paths, 5), (contracts.payload_schema_paths, 2), ], @@ -98,6 +99,7 @@ def test_corpus_size_is_pinned(self, accessor: object, expected_count: int) -> N contracts.task_declaration_paths, contracts.autonomy_event_paths, contracts.evidence_record_paths, + contracts.operating_record_paths, contracts.decision_artifact_paths, contracts.payload_schema_paths, ], @@ -120,7 +122,7 @@ def test_task_fixtures_are_yaml_not_the_invalid_case_table(self) -> None: class TestInvalidCases: @pytest.mark.parametrize( ("kind", "expected_count"), - [("task", 63), ("evidence", 6), ("decision", 11), ("event", 15)], + [("task", 63), ("evidence", 15), ("decision", 11), ("event", 15), ("operating", 43)], ) def test_case_table_size_is_pinned(self, kind: str, expected_count: int) -> None: assert len(contracts.invalid_cases(kind)) == expected_count # type: ignore[arg-type] @@ -153,6 +155,7 @@ def test_case_ids_are_unique_within_a_table(self, kind: str) -> None: ("evidence", ("evidence_semantic", "structural")), ("decision", ("decision_semantic", "structural")), ("event", ("event_semantic", "structural")), + ("operating", ("structural",)), ], ) def test_stage_vocabulary_is_pinned(self, kind: str, expected_stages: tuple[str, ...]) -> None: diff --git a/packages/paa-contracts/uv.lock b/packages/paa-contracts/uv.lock index dc2b8ca..b32a159 100644 --- a/packages/paa-contracts/uv.lock +++ b/packages/paa-contracts/uv.lock @@ -245,7 +245,7 @@ wheels = [ [[package]] name = "paa-contracts" -version = "0.1.0" +version = "0.2.0" source = { editable = "." } [package.dev-dependencies] diff --git a/pyproject.toml b/pyproject.toml index 36317ea..42fc790 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "paa-runtime" -version = "0.3.0" +version = "0.4.0" description = "Reference implementation of the Progressive Autonomy Architecture control plane: task declarations, motion lifecycle, position resolution, content-addressed evidence, and an append-only autonomy event store" readme = "README.md" requires-python = ">=3.12" @@ -34,7 +34,7 @@ Documentation = "https://www.paa.dev/reference/schema" Repository = "https://github.com/RankOneLabs/paa" [project.optional-dependencies] -# Conformance runs against the published contract artifacts: the four schemas, +# Conformance runs against the published contract artifacts: the five schemas, # the positive fixture corpus, and the invalid-case tables, all carried by # paa-contracts. # @@ -60,7 +60,7 @@ Repository = "https://github.com/RankOneLabs/paa" # rfc3987, which is GPL, and this repo is MIT. conformance = [ "jsonschema[format-nongpl]>=4.23", - "paa-contracts", + "paa-contracts>=0.2.0", ] # One repo, two published packages. The contract is data with no dependencies diff --git a/schemas/paa-evidence-record.schema.json b/schemas/paa-evidence-record.schema.json index 93db612..ed6a896 100644 --- a/schemas/paa-evidence-record.schema.json +++ b/schemas/paa-evidence-record.schema.json @@ -3,7 +3,7 @@ "$id": "https://paa.dev/paa-evidence-record.schema.json", "title": "PAA Evidence Record", "description": "Universal runtime envelope for one evaluator's verdict on one subject, produced against a PAA task declaration.", - "x-paa-schema-version": "paa-evidence-record/0.1.0-draft", + "x-paa-schema-version": "paa-evidence-record/0.2.0-draft", "type": "object", "required": [ "record_schema", @@ -23,6 +23,17 @@ ], "additionalProperties": false, "definitions": { + "worker": { + "type": "object", + "description": "Identity of the worker that performed the judged work, distinct from the evaluator-side producer. configuration_ref identifies the complete configuration; PAA does not interpret it.", + "required": ["id", "version", "configuration_ref"], + "additionalProperties": false, + "properties": { + "id": { "type": "string", "minLength": 1 }, + "version": { "type": "string", "minLength": 1 }, + "configuration_ref": { "type": "string", "minLength": 1 } + } + }, "evaluator_target": { "type": "string", "enum": ["input", "process", "output", "outcome"] @@ -56,7 +67,7 @@ "properties": { "record_schema": { "type": "string", - "const": "paa-evidence-record/0.1.0-draft" + "enum": ["paa-evidence-record/0.1.0-draft", "paa-evidence-record/0.2.0-draft"] }, "record_id": { "type": "string", @@ -134,6 +145,7 @@ } } }, + "worker": { "$ref": "#/definitions/worker" }, "producer": { "type": "object", "required": ["id", "version"], @@ -169,5 +181,11 @@ "type": "object", "description": "Task-specific evidence detail. Typed by the referenced payload_schema." } - } + }, + "allOf": [ + { + "if": { "properties": { "record_schema": { "const": "paa-evidence-record/0.1.0-draft" } } }, + "then": { "not": { "required": ["worker"] } } + } + ] } diff --git a/schemas/paa-operating-record.schema.json b/schemas/paa-operating-record.schema.json new file mode 100644 index 0000000..c802ce6 --- /dev/null +++ b/schemas/paa-operating-record.schema.json @@ -0,0 +1,90 @@ +{ + "$schema": "http://json-schema.org/draft-07/schema#", + "$id": "https://paa.dev/paa-operating-record.schema.json", + "title": "PAA Operating Record", + "description": "Optional subject-linked accounting artifact, independent of evaluator verdict count and autonomy transitions. Task or pipeline summaries must retain source attribution to constituent attempts, usage, model/rate identities and pricing bases. Readers must not add a summary and its constituents as separate charges.", + "x-paa-schema-version": "paa-operating-record/0.1.0-draft", + "type": "object", + "required": ["record_schema", "record_id", "task", "declaration_version", "scope", "subject", "worker", "usage", "price", "timestamps", "source_references"], + "additionalProperties": false, + "definitions": { + "worker": { + "type": "object", + "required": ["id", "version", "configuration_ref"], + "additionalProperties": false, + "properties": { + "id": { "type": "string", "minLength": 1 }, + "version": { "type": "string", "minLength": 1 }, + "configuration_ref": { "type": "string", "minLength": 1 } + } + }, + "price": { + "type": ["object", "null"], + "description": "Null means unavailable, never zero. A non-null price declares the currency and opaque rate/catalog basis; provenance must distinguish measured amounts from estimates and state coverage.", + "required": ["currency", "amount", "basis"], + "additionalProperties": false, + "properties": { + "currency": { "type": "string", "minLength": 1 }, + "amount": { "type": "number", "minimum": 0 }, + "basis": { "type": "string", "minLength": 1 } + } + } + }, + "properties": { + "record_schema": { "const": "paa-operating-record/0.1.0-draft" }, + "record_id": { "type": "string", "minLength": 1 }, + "task": { "type": "string", "pattern": "^[a-z][a-z0-9_]*$", "minLength": 1 }, + "declaration_version": { "type": "integer", "minimum": 1 }, + "scope": { "type": ["string", "null"], "minLength": 1 }, + "subject": { + "type": "object", + "required": ["kind", "id"], + "additionalProperties": false, + "properties": { + "kind": { "type": "string", "enum": ["case", "run"] }, + "id": { "type": "string", "minLength": 1 } + } + }, + "worker": { "$ref": "#/definitions/worker" }, + "usage": { + "type": ["object", "null"], + "description": "Consumer-defined nonnegative quantities. input_tokens, output_tokens, cached_tokens and llm_calls are recommended, not an exhaustive vocabulary. Null quantities (or null usage) are explicitly unavailable. Omitted measures make no coverage claim. Aggregate tokens alone cannot reprice mixed-model work.", + "minProperties": 1, + "propertyNames": { "minLength": 1 }, + "additionalProperties": { "type": ["number", "null"], "minimum": 0 } + }, + "price": { "$ref": "#/definitions/price" }, + "components": { + "type": "array", + "description": "Additional costs NOT already included in price. Component kinds have consumer-defined meaning. An omitted array makes no completeness claim.", + "items": { + "type": "object", + "required": ["kind", "quantity", "unit", "price"], + "additionalProperties": false, + "properties": { + "kind": { "type": "string", "minLength": 1 }, + "quantity": { "type": ["number", "null"], "minimum": 0 }, + "unit": { "type": "string", "minLength": 1 }, + "price": { "$ref": "#/definitions/price" } + } + } + }, + "timestamps": { + "type": "object", + "required": ["started_at", "completed_at", "recorded_at"], + "additionalProperties": false, + "properties": { + "started_at": { "type": "string", "format": "date-time" }, + "completed_at": { "type": "string", "format": "date-time" }, + "recorded_at": { "type": "string", "format": "date-time" } + } + }, + "source_references": { + "type": "array", + "description": "Durable references identifying covered work, scopes, configurations, individual attempts (including failed and superseded retries), quantities, model/rate identities, pricing bases and measurement coverage. Multiple verdicts do not create new charges. PAA preserves these references but cannot verify external provenance or deduplicate overlapping summaries.", + "minItems": 1, + "uniqueItems": true, + "items": { "type": "string", "minLength": 1 } + } + } +} diff --git a/src/paa_runtime/__init__.py b/src/paa_runtime/__init__.py index a5a4bc0..06abbbe 100644 --- a/src/paa_runtime/__init__.py +++ b/src/paa_runtime/__init__.py @@ -75,6 +75,22 @@ store_evidence, verify_evidence, ) +from paa_runtime.operating import ( + CURRENT_OPERATING_SCHEMA, + OperatingComponent, + OperatingPrice, + OperatingRecord, + OperatingRecordError, + RecordSubject, + RecordTimestamps, + WorkerIdentity, + decode_operating_record, +) +from paa_runtime.operating_store import ( + OperatingRecordStore, + OperatingStoreError, + SqliteOperatingRecordStore, +) from paa_runtime.replay import import_events from paa_runtime.service import ( Motion, @@ -101,13 +117,25 @@ from paa_runtime.sqlite_store import SqliteEventStore from paa_runtime.store import AutonomyEvent, EventStore -__version__ = "0.3.0" +__version__ = "0.4.0" __all__ = [ # Construction surface "RuntimeConfig", "SqliteEventStore", "EventStore", + "SqliteOperatingRecordStore", + "OperatingRecordStore", + "OperatingStoreError", + "OperatingRecord", + "OperatingRecordError", + "OperatingComponent", + "OperatingPrice", + "RecordSubject", + "RecordTimestamps", + "WorkerIdentity", + "CURRENT_OPERATING_SCHEMA", + "decode_operating_record", # Lifecycle API "propose", "approve", diff --git a/src/paa_runtime/operating.py b/src/paa_runtime/operating.py new file mode 100644 index 0000000..65d7f4e --- /dev/null +++ b/src/paa_runtime/operating.py @@ -0,0 +1,219 @@ +"""Named wire types and an IO decoder for paa-operating-record/0.1.0-draft. + +These types mirror schemas/paa-operating-record.schema.json. WorkerIdentity +also mirrors the optional evidence worker definition. Neither is a policy +input: the motion service does not import this module. +""" + +from __future__ import annotations + +import json +import math +import re +from datetime import datetime +from typing import Literal, Never, NotRequired, TypedDict, cast + +CURRENT_OPERATING_SCHEMA = "paa-operating-record/0.1.0-draft" + + +class WorkerIdentity(TypedDict): + """The worker role, its version, and opaque complete configuration handle.""" + + id: str + version: str + configuration_ref: str + + +class RecordSubject(TypedDict): + """The same subject identity used by evidence records.""" + + kind: Literal["case", "run"] + id: str + + +class OperatingPrice(TypedDict): + """A priced amount with a consumer-owned currency and rate/catalog basis.""" + + currency: str + amount: int | float + basis: str + + +class OperatingComponent(TypedDict): + """Additional cost not already included in the record's base price.""" + + kind: str + quantity: int | float | None + unit: str + price: OperatingPrice | None + + +class RecordTimestamps(TypedDict): + """RFC 3339 timestamps for the attributed work and its recording.""" + + started_at: str + completed_at: str + recorded_at: str + + +class OperatingRecord(TypedDict): + """Wire representation; unavailable measurements are None, not zero. + + Usage keys and component kinds are open consumer vocabulary. Source + references preserve attribution; this envelope cannot verify external + provenance, price accuracy, coverage, or overlap between summaries. + """ + + record_schema: Literal["paa-operating-record/0.1.0-draft"] + record_id: str + task: str + declaration_version: int + scope: str | None + subject: RecordSubject + worker: WorkerIdentity + usage: dict[str, int | float | None] | None + price: OperatingPrice | None + timestamps: RecordTimestamps + source_references: list[str] + components: NotRequired[list[OperatingComponent]] + + +class OperatingRecordError(ValueError): + """Malformed record at the JSON IO boundary, with operation and field path.""" + + +def _fail(path: str, detail: str) -> Never: + raise OperatingRecordError(f"decode operating record at {path}: {detail}") + + +def _object( + value: object, path: str, required: set[str], optional: frozenset[str] = frozenset(), +) -> dict[str, object]: + # Dynamic JSON is unknown until every member has been validated below. + if not isinstance(value, dict) or not all(isinstance(key, str) for key in value): + _fail(path, "expected an object") + result = cast(dict[str, object], value) + if required - result.keys() or result.keys() - required - optional: + _fail(path, "missing required or unknown fields") + return result + + +def _text(value: object, path: str) -> str: + if not isinstance(value, str) or not value: + _fail(path, "expected a nonempty string") + return value + + +def _quantity(value: object, path: str) -> None: + if value is None: + return + if isinstance(value, bool) or not isinstance(value, (int, float)): + _fail(path, "expected a nonnegative finite number or null") + if value < 0 or (isinstance(value, float) and not math.isfinite(value)): + _fail(path, "expected a nonnegative finite number or null") + + +def _price(value: object, path: str) -> None: + if value is None: + return + price = _object(value, path, {"currency", "amount", "basis"}) + _text(price["currency"], path + "/currency") + _text(price["basis"], path + "/basis") + if price["amount"] is None: + _fail(path + "/amount", "use null price when the amount is unavailable") + _quantity(price["amount"], path + "/amount") + + +def _timestamp(value: object, path: str) -> None: + stamp = _text(value, path) + pattern = ( + r"[0-9]{4}-[0-9]{2}-[0-9]{2}[Tt](?:[01][0-9]|2[0-3]):[0-5][0-9]:[0-5][0-9]" + r"(?:\.[0-9]+)?(?:[Zz]|[+-](?:[01][0-9]|2[0-3]):[0-5][0-9])" + ) + if re.fullmatch(pattern, stamp) is None: + _fail(path, "expected an RFC 3339 date-time with timezone") + try: + datetime.fromisoformat(stamp.upper().replace("Z", "+00:00")) + except ValueError as error: + _fail(path, str(error)) + + +def decode_operating_record(data: bytes) -> OperatingRecord: + """Decode and structurally check JSON at the store/import boundary. + + No schema package is needed in production. Conformance tests run the + published positive and negative corpus through both this decoder and + the JSON Schema. Semantic attribution remains the consumer's duty. + """ + try: + return _decode_operating_record(data) + except OperatingRecordError: + raise + except (ValueError, TypeError, OverflowError, RecursionError) as error: + raise OperatingRecordError(f"decode operating record: {error}") from error + + +def _decode_operating_record(data: bytes) -> OperatingRecord: + document: object = json.loads(data) + record = _object(document, "/", { + "record_schema", "record_id", "task", "declaration_version", "scope", + "subject", "worker", "usage", "price", "timestamps", "source_references", + }, frozenset({"components"})) + if record["record_schema"] != CURRENT_OPERATING_SCHEMA: + _fail("/record_schema", "unsupported schema version") + _text(record["record_id"], "/record_id") + if re.fullmatch(r"[a-z][a-z0-9_]*", _text(record["task"], "/task")) is None: + _fail("/task", "invalid task identifier") + version = record["declaration_version"] + _quantity(version, "/declaration_version") + if not isinstance(version, (int, float)) or version < 1 or int(version) != version: + _fail("/declaration_version", "expected a positive integer") + record["declaration_version"] = int(version) + if record["scope"] is not None: + _text(record["scope"], "/scope") + subject = _object(record["subject"], "/subject", {"kind", "id"}) + if subject["kind"] not in ("case", "run"): + _fail("/subject/kind", "expected case or run") + _text(subject["id"], "/subject/id") + worker = _object(record["worker"], "/worker", {"id", "version", "configuration_ref"}) + for key, value in worker.items(): + _text(value, "/worker/" + key) + usage = record["usage"] + if usage is not None: + if not isinstance(usage, dict) or not usage: + _fail("/usage", "expected a nonempty quantity object or null") + for key, value in cast(dict[str, object], usage).items(): + _text(key, "/usage") + _quantity(value, "/usage/" + key) + _price(record["price"], "/price") + if "components" in record: + components = record["components"] + if not isinstance(components, list): + _fail("/components", "expected an array") + for index, value in enumerate(cast(list[object], components)): + path = f"/components/{index}" + component = _object(value, path, {"kind", "quantity", "unit", "price"}) + _text(component["kind"], path + "/kind") + _text(component["unit"], path + "/unit") + _quantity(component["quantity"], path + "/quantity") + _price(component["price"], path + "/price") + timestamps = _object(record["timestamps"], "/timestamps", { + "started_at", "completed_at", "recorded_at", + }) + for key, value in timestamps.items(): + _timestamp(value, "/timestamps/" + key) + references = record["source_references"] + if not isinstance(references, list) or not references: + _fail("/source_references", "expected a nonempty array") + strings = [_text(value, "/source_references") for value in cast(list[object], references)] + if len(set(strings)) != len(strings): + _fail("/source_references", "duplicate source references") + # The sole cast to the wire type follows validation of every field. + return cast(OperatingRecord, record) + + +__all__ = [ + "CURRENT_OPERATING_SCHEMA", "WorkerIdentity", "RecordSubject", "OperatingPrice", + "OperatingComponent", "RecordTimestamps", "OperatingRecord", "OperatingRecordError", + "decode_operating_record", +] diff --git a/src/paa_runtime/operating_store.py b/src/paa_runtime/operating_store.py new file mode 100644 index 0000000..7bcc5aa --- /dev/null +++ b/src/paa_runtime/operating_store.py @@ -0,0 +1,130 @@ +"""Independent append-only operating records; no dependency on authority state. + +The SQLite implementation may use the event store's database path or a separate +database. It does not alter EventStore, evidence files, or the motion service. +One append is one transaction. Record IDs are unique; subjects are not, since +attempts and summaries may share a subject without representing new outcomes. +""" + +from __future__ import annotations + +import json +import sqlite3 +from pathlib import Path +from typing import Protocol + +from paa_runtime.operating import OperatingRecord, RecordSubject, decode_operating_record + + +class OperatingStoreError(RuntimeError): + """Storage failure with operation and record/subject context.""" + + +class OperatingRecordStore(Protocol): + """Append records and retrieve all records for an exact subject. + + Append rejects reused record IDs, including identical duplicates. No + update/delete API exists and mutation must be blocked in storage. + Retrieval returns insertion order and preserves all fields, including + source references, missing measurements, and optional components. + Consumers must reconcile overlapping records before aggregating costs. + """ + + def append(self, record: OperatingRecord) -> None: + """Atomically append one structurally valid record.""" + ... + + def get_by_subject(self, subject: RecordSubject) -> tuple[OperatingRecord, ...]: + """All records with this exact subject kind and ID, in insertion order.""" + ... + + +_SCHEMA_DDL = """ +CREATE TABLE IF NOT EXISTS operating_records ( + sequence INTEGER PRIMARY KEY AUTOINCREMENT, + record_id TEXT NOT NULL UNIQUE, + subject_kind TEXT NOT NULL CHECK(subject_kind IN ('case', 'run')), + subject_id TEXT NOT NULL, + document TEXT NOT NULL +); +CREATE INDEX IF NOT EXISTS operating_records_subject_idx + ON operating_records(subject_kind, subject_id, sequence); +CREATE TRIGGER IF NOT EXISTS operating_records_no_replace +BEFORE INSERT ON operating_records +WHEN EXISTS ( + SELECT 1 FROM operating_records + WHERE record_id = NEW.record_id OR sequence = NEW.sequence +) +BEGIN + SELECT RAISE(ABORT, 'operating_records record already exists'); +END; +CREATE TRIGGER IF NOT EXISTS operating_records_no_update +BEFORE UPDATE ON operating_records +BEGIN + SELECT RAISE(ABORT, 'operating_records is append-only'); +END; +CREATE TRIGGER IF NOT EXISTS operating_records_no_delete +BEFORE DELETE ON operating_records +BEGIN + SELECT RAISE(ABORT, 'operating_records is append-only'); +END; +""" + + +class SqliteOperatingRecordStore: + """Owns a connection and a sibling table, never an autonomy event stream.""" + + def __init__(self, db_path: Path) -> None: + try: + self._conn = sqlite3.connect(db_path) + except sqlite3.Error as error: + raise OperatingStoreError(f"open operating store {db_path}: {error}") from error + try: + # Enables deletion triggers for REPLACE as well as explicit DELETE. + self._conn.execute("PRAGMA recursive_triggers=ON") + self._conn.execute("PRAGMA journal_mode=WAL") + self._conn.execute("PRAGMA synchronous=NORMAL") + self._conn.executescript(_SCHEMA_DDL) + self._conn.commit() + except sqlite3.Error as error: + self._conn.close() + raise OperatingStoreError(f"initialize operating store {db_path}: {error}") from error + + def close(self) -> None: + """Release this store's connection.""" + self._conn.close() + + def append(self, record: OperatingRecord) -> None: + """Validate at IO, then insert atomically; duplicate IDs fail closed.""" + try: + document = json.dumps(record, allow_nan=False) + except (ValueError, TypeError, RecursionError) as error: + raise OperatingStoreError(f"serialize operating record: {error}") from error + checked = decode_operating_record(document.encode("utf-8")) + try: + with self._conn: + self._conn.execute( + """INSERT INTO operating_records + (record_id, subject_kind, subject_id, document) VALUES (?, ?, ?, ?)""", + (checked["record_id"], checked["subject"]["kind"], + checked["subject"]["id"], document), + ) + except sqlite3.Error as error: + raise OperatingStoreError( + f"append operating record {checked['record_id']!r}: {error}" + ) from error + + def get_by_subject(self, subject: RecordSubject) -> tuple[OperatingRecord, ...]: + """Return detached, validated records; no costing or policy decisions.""" + try: + rows = self._conn.execute( + """SELECT document FROM operating_records + WHERE subject_kind = ? AND subject_id = ? ORDER BY sequence""", + (subject["kind"], subject["id"]), + ).fetchall() + except sqlite3.Error as error: + raise OperatingStoreError(f"read operating subject {subject!r}: {error}") from error + return tuple(decode_operating_record(row[0].encode("utf-8")) for row in rows) + + +__all__ = ["OperatingRecordStore", "SqliteOperatingRecordStore", "OperatingStoreError"] diff --git a/tests/test_operating_store.py b/tests/test_operating_store.py new file mode 100644 index 0000000..fdeee9a --- /dev/null +++ b/tests/test_operating_store.py @@ -0,0 +1,201 @@ +"""Operating accounting is append-only, lossless, and independent of authority.""" + +from __future__ import annotations + +import copy +import json +import sqlite3 +from pathlib import Path +from typing import Literal + +import pytest + +from paa_runtime.operating import OperatingRecord, OperatingRecordError, decode_operating_record +from paa_runtime.operating_store import OperatingStoreError, SqliteOperatingRecordStore + + +@pytest.fixture +def record() -> OperatingRecord: + """Minimal illustrative task attempt, shaped by the operating contract.""" + return { + "record_schema": "paa-operating-record/0.1.0-draft", + "record_id": "attempt-1", "task": "reply_draft", "declaration_version": 1, + "scope": None, "subject": {"kind": "run", "id": "run-1"}, + "worker": {"id": "drafter", "version": "v1", "configuration_ref": "cfg:1"}, + "usage": {"input_tokens": 100, "output_tokens": None}, + "price": {"currency": "USD", "amount": 0.001, "basis": "catalog:v1"}, + "timestamps": { + "started_at": "2026-01-01T00:00:00Z", "completed_at": "2026-01-01T00:00:01Z", + "recorded_at": "2026-01-01T00:00:02Z", + }, + "source_references": ["app://usage/run-1/attempt-1"], + } + + +def test_record_survives_reopening(tmp_path: Path, record: OperatingRecord) -> None: + path = tmp_path / "records.db" + store = SqliteOperatingRecordStore(path) + store.append(record) + store.close() + reopened = SqliteOperatingRecordStore(path) + try: + assert reopened.get_by_subject(record["subject"]) == (record,) + finally: + reopened.close() + + +def test_attempts_share_subject_without_overwriting( + tmp_path: Path, record: OperatingRecord, +) -> None: + retry = copy.deepcopy(record) + retry.update(record_id="attempt-2", price=None, usage=None) + retry["source_references"] = ["app://usage/run-1/failed-attempt-2"] + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + store.append(retry) + assert store.get_by_subject(record["subject"]) == (record, retry) + finally: + store.close() + + +def test_subject_lookup_matches_kind_and_id(tmp_path: Path, record: OperatingRecord) -> None: + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + assert store.get_by_subject({"kind": "case", "id": "run-1"}) == () + finally: + store.close() + + +def test_subject_lookup_excludes_other_ids(tmp_path: Path, record: OperatingRecord) -> None: + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + assert store.get_by_subject({"kind": "run", "id": "run-2"}) == () + finally: + store.close() + + +def test_same_subject_retains_task_scope_and_configuration( + tmp_path: Path, record: OperatingRecord, +) -> None: + other = copy.deepcopy(record) + other.update(record_id="other-task", task="reply_review", scope="review:public") + other["worker"]["configuration_ref"] = "cfg:review" + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + store.append(other) + assert store.get_by_subject(record["subject"]) == (record, other) + finally: + store.close() + + +def test_reused_record_id_fails_and_original_remains( + tmp_path: Path, record: OperatingRecord, +) -> None: + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + with pytest.raises(OperatingStoreError, match="append operating record 'attempt-1'"): + store.append(record) + assert store.get_by_subject(record["subject"]) == (record,) + finally: + store.close() + + +@pytest.mark.parametrize("statement", [ + "UPDATE operating_records SET document = '{}'", + "DELETE FROM operating_records", + "INSERT OR REPLACE INTO operating_records " + "(record_id, subject_kind, subject_id, document) VALUES ('attempt-1', 'run', 'r', '{}')", + "INSERT OR REPLACE INTO operating_records " + "(sequence, record_id, subject_kind, subject_id, document) " + "VALUES (1, 'other', 'run', 'r', '{}')", +]) +def test_storage_rejects_mutation_from_another_connection( + tmp_path: Path, record: OperatingRecord, statement: str, +) -> None: + path = tmp_path / "records.db" + store = SqliteOperatingRecordStore(path) + store.append(record) + connection = sqlite3.connect(path) + try: + with pytest.raises(sqlite3.IntegrityError): + connection.execute(statement) + connection.rollback() + assert store.get_by_subject(record["subject"]) == (record,) + finally: + connection.close() + store.close() + + +def test_caller_mutation_cannot_change_stored_record( + tmp_path: Path, record: OperatingRecord, +) -> None: + expected = copy.deepcopy(record) + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + store.append(record) + record["worker"]["configuration_ref"] = "cfg:changed" + read = store.get_by_subject(record["subject"]) + read[0]["source_references"].append("new-reference") + assert store.get_by_subject(record["subject"]) == (expected,) + finally: + store.close() + + +def test_invalid_record_is_not_partially_written(tmp_path: Path, record: OperatingRecord) -> None: + record["worker"]["configuration_ref"] = "" + store = SqliteOperatingRecordStore(tmp_path / "records.db") + try: + with pytest.raises(OperatingRecordError, match="/worker/configuration_ref"): + store.append(record) + assert store.get_by_subject(record["subject"]) == () + finally: + store.close() + + +@pytest.mark.parametrize("value", [float("nan"), float("inf"), -float("inf")]) +def test_non_json_numbers_are_rejected(record: OperatingRecord, value: float) -> None: + record["usage"] = {"tokens": value} + with pytest.raises(OperatingRecordError, match="finite"): + decode_operating_record(json.dumps(record).encode()) + + +def test_malformed_json_fails_closed() -> None: + with pytest.raises(OperatingRecordError, match="decode operating record"): + decode_operating_record(b"not-json") + + +@pytest.mark.parametrize("stamp", [ + "2026-01-01T00:00:00+01:99", "2026-01-01T24:00:00Z", "2026-02-30T00:00:00Z", +]) +@pytest.mark.parametrize("field", ["started_at", "completed_at", "recorded_at"]) +def test_invalid_timestamp_reports_its_field_path( + record: OperatingRecord, stamp: str, + field: Literal["started_at", "completed_at", "recorded_at"], +) -> None: + record["timestamps"][field] = stamp + with pytest.raises(OperatingRecordError, match=f"/timestamps/{field}"): + decode_operating_record(json.dumps(record).encode()) + + +def test_database_open_failure_has_operation_context(tmp_path: Path) -> None: + with pytest.raises(OperatingStoreError, match="open operating store"): + SqliteOperatingRecordStore(tmp_path / "missing-directory" / "records.db") + + +def test_multiple_connections_observe_committed_records( + tmp_path: Path, record: OperatingRecord, +) -> None: + path = tmp_path / "records.db" + first = SqliteOperatingRecordStore(path) + second = SqliteOperatingRecordStore(path) + try: + first.append(record) + assert second.get_by_subject(record["subject"]) == (record,) + finally: + second.close() + first.close() diff --git a/uv.lock b/uv.lock index b6970ee..40c60e9 100644 --- a/uv.lock +++ b/uv.lock @@ -361,7 +361,7 @@ wheels = [ [[package]] name = "paa-contracts" -version = "0.1.0" +version = "0.2.0" source = { editable = "packages/paa-contracts" } [package.dev-dependencies] @@ -382,7 +382,7 @@ dev = [ [[package]] name = "paa-runtime" -version = "0.3.0" +version = "0.4.0" source = { editable = "." } dependencies = [ { name = "pyyaml" },