diff --git a/docs/SURFACE.md b/docs/SURFACE.md index 6261f456c..29c6887f3 100644 --- a/docs/SURFACE.md +++ b/docs/SURFACE.md @@ -246,7 +246,8 @@ No process runs between events: the handler wakes, executes to its next await, p one, while a worker that dies stops renewing and the lease expires. The run's wallclock budget is separate and stops *new* work, draining whatever is already running rather than cancelling a step mid-flight. A wrapper step and - a native step therefore get the same duration. + a native step therefore get the same duration by default. Agent steps may + declare the per-step `timeout` described below, enforced on both paths. `flows check` resolves the binary (a path is relative to the declaring flow or project config; a bare name resolves via `PATH`) and caches each resolved @@ -347,7 +348,8 @@ authoring-time narrowing, not a kernel guarantee. ### Per-agent permissions in TypeScript -Supported `f.agent` calls accept an optional `permissions` declaration: +Supported `f.agent` calls accept an optional `timeout` (see Agent step timeouts) +and an optional `permissions` declaration: ```ts const draft = await f.agent("writer", { @@ -641,6 +643,68 @@ the kernel kills the command's process group and journals `completionReason: tim `f.run` refuses with code `lease_exceeded`. The override applies only to that invocation; calls without options retain the default. +### Agent step timeouts + +`f.agent(name, { task, timeout?: string | number })` accepts the same duration +syntax as `f.run`: numeric milliseconds or strings with `ms`, `s`, or `m` +(including `'1.5s'`). Omitting `timeout` keeps the existing unlimited CLI +execution duration. The authoring/`flows check` ceiling is **60 minutes** +(3600000 ms), inclusive: measured repair work reached 44m09s, so 60m allows +about 35% headroom and leaves room in a 2h flow for publishing. This is a +compiler/SDK validation limit; the kernel validates positive milliseconds +fitting in `i64` and carries the declaration to the worker. + +```ts +const repair = await f.agent('repair', { task: 'Repair failing checks.', timeout: '45m' }); +if (repair.completionReason === 'timeout') { + await f.run('git push'); // publish what the agent already committed +} +f.done('success'); +``` + +**A declared timeout resolves; it does not throw.** `AgentResult` includes +`completionReason: 'success' | 'timeout'`. Handle the timeout by branching on +that field, not with `catch`. The worker stops the CLI process group at the +execution deadline, using the existing graceful stop and forced-kill escalation, +then journals `step.completed` with `completionReason: 'timeout'`. Settlement +includes process-stop confirmation, so it can occur shortly after the deadline. +Failure to confirm the stop remains a failure, not a recoverable timeout. +The agent child run remains failed (`step_failed`) in status/dashboard views; +the authored flow may continue and succeed. Other failures still throw. +A predicate `.gate(result => ...)` still runs on this result; returning false +lets an author fail the flow on timeout. Named data gates require successful +producer output; a timeout that prevents such a gate from running still fails +the operation, rather than bypassing the declared check. + +On timeout, `summary` contains journaled timeout evidence and `artifacts` is +`[]`. Native CLI partial output is retained in transport failure evidence; +native artifact paths remain under `trajectory_tail.transcript.artifacts.paths` +in the journal (a bounded list). Wrappers discard partial stdout on timeout +and provide the deadline message; they have no transcript artifact list. A +wrapper result envelope emitted before the deadline still supplies its usage. +No workspace reset occurs: committed work and uncommitted edits remain, and +the following step sees that potentially dirty tree even with +`recoveryMode: 'reset'`. Unlike crash/lease recovery in RFC Appendix A, timeout +settles the step with no successor attempt. Resume replays the recorded timeout +and does not execute that agent again. Incurred spend remains charged: +reported usage is priced as usual, and usage never reported before the deadline +is journaled as dollar-unmetered rather than as a measured $0. + +`timeout` with `maxIterations > 1` is refused, so semantic iterations cannot +multiply the limit. Transport recovery before a timeout can start a new CLI +execution with its own timer; this is an execution deadline, not a total flow +budget. `transport: 'relay'` with a timeout is refused because the local worker +cannot stop the remote process. Authored option errors use the existing +`agent_cli_unresolved` refusal, before child admission. + +Declarative YAML/JSON uses `timeoutMs: 2700000` on agent steps: **integer +milliseconds only**, no duration strings. It lowers to journal spec +`timeout_ms`. This enforced duration is separate from +`requirements.expectedDurationMs`, which is only a placement estimate. +Use matching SDK, daemon, and worker versions: older daemons refuse the new +field, while older workers do not enforce it. The native and wrapper paths +both honor the declaration; no default wrapper duration cap is introduced. + ### Repair before failure: `onNonZero: 'record'` A deterministic step is gated on its exit code by default: a nonzero exit fails @@ -781,7 +845,9 @@ vocabulary: recorded before author code can reach the operation, so a `catch` cannot hide it. `done("step_failed")` does not change that: it declares a verdict about checks the body ran and read for itself, and is not a way to continue past a - step that failed. + step that failed. A declared agent timeout is the explicit exception: + it journals a failed step but resolves with `completionReason: 'timeout'`, + allowing the flow to continue (see Agent step timeouts). 3. **Finish your derived work before `done()`.** If a handler chained onto a step is still in flight when the body returns, the run is refused with `unsettled_derived_work` rather than recorded as a success nobody can prove. diff --git a/evidence/agent-timeout/README.md b/evidence/agent-timeout/README.md new file mode 100644 index 000000000..452aa9e29 --- /dev/null +++ b/evidence/agent-timeout/README.md @@ -0,0 +1,65 @@ +# Agent timeout verification + +Commands ran from the repository root unless a `cd` is shown. The linked files +contain captured stdout/stderr, including failures; they are not paraphrases. +Dependencies were installed in surface and SDK, the SDK's surface dependency was +linked to this checkout's built surface, and `ops/cargo.sh` installed its private +toolchain. No workflow files or gate workflows were changed. + +| Literal command | Captured output | +| --- | --- | +| `npm run build --prefix packages/sdk` | [build-sdk.txt](build-sdk.txt) | +| `sh ops/cargo.sh build --manifest-path kernel/Cargo.toml` | [build-kernel.txt](build-kernel.txt) | +| `npm run typecheck --prefix packages/sdk` | [typecheck.txt](typecheck.txt) | +| `npm run typecheck:tests --prefix packages/sdk` | [test-typecheck.txt](test-typecheck.txt) | +| `sh ops/cargo.sh test --manifest-path kernel/Cargo.toml -p relayflowd-core` | [kernel.txt](kernel.txt) | +| `npm run generate --prefix packages/schema` | [schema-generate.txt](schema-generate.txt) | +| `npm test --prefix packages/schema` | [schema-tests.txt](schema-tests.txt) | + +Final SDK regression command, with [captured output](final-regressions.txt): + +```sh +cd packages/sdk && npx vitest run tests/agent-timeout.test.ts tests/agent-timeout-outcome.test.ts tests/agent-timeout-worker.test.ts tests/agent-timeout-live.test.ts tests/step-lease.test.ts tests/spec-parity.test.ts tests/verb-field-lint.test.ts tests/wrapper-execution-duration.test.ts tests/stop-process-group.test.ts tests/worker-cli.test.ts tests/authored-node-result.test.ts tests/authored-retried-child-resume-live.test.ts tests/authored-agent-artifacts.test.ts +``` + +Earlier broader command, with [captured failures](regression-tests.txt): + +```sh +cd packages/sdk && npx vitest run tests/agent-timeout.test.ts tests/agent-timeout-worker.test.ts tests/agent-timeout-live.test.ts tests/step-lease.test.ts tests/spec-parity.test.ts tests/verb-field-lint.test.ts tests/wrapper-execution-duration.test.ts tests/stop-process-group.test.ts tests/worker-cli.test.ts tests/authored-node-result.test.ts tests/live-kernel.test.ts tests/authored-retried-child-resume-live.test.ts tests/authored-agent-artifacts.test.ts +``` + +That run found a stale per-verb descriptor expectation (updated to admit agent +`timeoutMs`) and nine live-kernel failures in this checkout. The checkout inherits +`type: commonjs` from `/home/daytona/package.json`; its extensionless ESM fixtures +did not emit their wrapper handshake. A daemon-discovery test also needed an +explicit binary path. + +The live-kernel suite was rerun in `/tmp/agent-timeout-baseline`, an isolated +worktree **overlaid with the changed source, new files, and current SDK dist**. +Despite the directory name, this is the implemented code, not a baseline test. +Its SDK dependencies link to the installed dependencies in the working checkout. +No fixtures were rewritten to change their behavior. The analyzer opt-out flag +allows the existing suite to report unavailable live Claude access if needed; +the captured output records whether the analyzer actually executed. + +```sh +cd /tmp/agent-timeout-baseline/packages/sdk && RELAYFLOWD_BIN=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 ./node_modules/.bin/vitest run tests/live-kernel.test.ts +``` + +[Captured isolated live-kernel output](isolated-live-kernel.txt). + +The timeout-specific runtime and kill/resume tests ran in the original checkout +as part of the final SDK regression command. They do not depend on the isolated +checkout or on live provider access. + +The remaining real Claude analyzer failure was also run against the original +branch head (`c88c3d0`) in `/tmp/agent-timeout-original`, with its own original +surface and SDK rebuilt. It used the same daemon executable. The original +source produced the same execution-gate failure; this is not a green live +analyzer acceptance claim. + +```sh +cd /tmp/agent-timeout-original/packages/sdk && RELAYFLOWD_BIN=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd ./node_modules/.bin/vitest run tests/live-kernel.test.ts -t 'hn-monitor analyze-story reaches done through the real Claude analyzer CLI' +``` + +[Captured original-head analyzer output](original-analyzer.txt). diff --git a/evidence/agent-timeout/build-kernel.txt b/evidence/agent-timeout/build-kernel.txt new file mode 100644 index 000000000..fac58b91f --- /dev/null +++ b/evidence/agent-timeout/build-kernel.txt @@ -0,0 +1,4 @@ + Compiling relayflowd-core v0.1.0 (/home/daytona/.relayflow-v2-supervisor/durable/repository/kernel/relayflowd-core) + Compiling relayflowd-journal v0.1.0 (/home/daytona/.relayflow-v2-supervisor/durable/repository/kernel/relayflowd-journal) + Compiling relayflowd v0.1.0 (/home/daytona/.relayflow-v2-supervisor/durable/repository/kernel/relayflowd) + Finished `dev` profile [unoptimized + debuginfo] target(s) in 3.00s diff --git a/evidence/agent-timeout/build-sdk.txt b/evidence/agent-timeout/build-sdk.txt new file mode 100644 index 000000000..419855a15 --- /dev/null +++ b/evidence/agent-timeout/build-sdk.txt @@ -0,0 +1,4 @@ + +> @relayflows/sdk@2.0.39 build +> tsc && node scripts/make-cli-executable.mjs + diff --git a/evidence/agent-timeout/final-regressions.txt b/evidence/agent-timeout/final-regressions.txt new file mode 100644 index 000000000..9005887f2 --- /dev/null +++ b/evidence/agent-timeout/final-regressions.txt @@ -0,0 +1,67 @@ + + RUN v2.1.9 /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk + + ✓ tests/authored-node-result.test.ts (45 tests) 35ms + ✓ tests/verb-field-lint.test.ts (99 tests) 384ms + ✓ tests/wrapper-execution-duration.test.ts (7 tests) 10858ms + ✓ keeps the handshake deadline independent of the removed execution deadline 10062ms + ✓ still lets a lease abort stop an unlimited wrapper before it produces output 511ms + ✓ tests/stop-process-group.test.ts (11 tests) 14607ms + ✓ every stop reaches the process group, not just the direct child > exits the run after an execution-timeout stop 1003ms + ✓ every stop reaches the process group, not just the direct child > exits the run after a protocol terminate stop 653ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after a protocol terminate stop 1584ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after an execution-timeout stop 1944ms + ✓ every stop reaches the process group, not just the direct child > holds the loop open long enough for the escalation to run 1104ms + ✓ every stop reaches the process group, not just the direct child > bounds forced-stop confirmation when a group remains unprovable 1009ms + ✓ a wrapper that exits with no execution deadline still drains > reports the wrapper result and reaps a grandchild holding its pipes 797ms + ✓ a wrapper that exits with no execution deadline still drains > reaps a SIGTERM-deaf grandchild holding its pipes 1886ms + ✓ a wrapper that exits with no execution deadline still drains > settles on its own deadline when an escaped holder withholds close 4289ms + ✓ tests/authored-agent-artifacts.test.ts (5 tests) 356ms + ✓ tests/spec-parity.test.ts (46 tests) 538ms + ✓ tests/authored-retried-child-resume-live.test.ts (2 tests) 1499ms + ✓ an authored step whose child run retried an attempt > resumes after a kill mid-step and reads the terminal completion, not the crashed one 1361ms + ✓ tests/agent-timeout-live.test.ts (3 tests) 4486ms + ✓ journals timeout, stops the process, runs a predicate gate and publishes under a budget header 1231ms + ✓ replays timeout after killing the root and daemon, without executing the agent again 2078ms + ✓ lets an author reject timeout through a predicate gate 1176ms + ✓ tests/agent-timeout-worker.test.ts (8 tests) 1981ms + ✓ stops claude at its declared limit, preserving work and stopping descendants 523ms + ✓ stops codex at its declared limit, preserving work and stopping descendants 519ms + ✓ stops wrapper at its declared limit, preserving work and stopping descendants 540ms + ✓ tests/agent-timeout-outcome.test.ts (6 tests) 12ms + ✓ tests/agent-timeout.test.ts (23 tests) 38ms + ✓ tests/worker-cli.test.ts (25 tests) 30887ms + ✓ registered CLI model defaults > passes the same priced Claude default to the real provider invocation 532ms + ✓ registered CLI model defaults > uses the explicitly supplied agent environment for the provider subprocess 581ms + ✓ registered CLI model defaults > dispatches canonical generic bytes with the preflight-proved adapter identity 571ms + ✓ direct transport lifecycle evidence > classifies only the exact Codex stdin lifecycle signature as retryable 583ms + ✓ direct transport lifecycle evidence > records a signal close separately from an ordinary nonzero exit 1000ms + ✓ direct transport lifecycle evidence > records a spawn error code without treating a missing executable as transient 405ms + ✓ direct transport lifecycle evidence > journals classified lifecycle evidence and reports crashed instead of generic worker_error 554ms + ✓ step discovery environment > names the run, step, attempt and an absolute data dir for a direct agent spawn 624ms + ✓ step discovery environment > exports none of the four without a data dir, even when the worker inherited them 499ms + ✓ wrapper discovery environment > sets the four names from the dispatch and still refuses ambient values and other secrets 530ms + ✓ wrapper discovery environment > exports none of the four to a wrapper without a data dir, even when the worker inherited them 521ms + ✓ custom wrapper execution identity > passes an explicit safe environment at identification and execution 577ms + ✓ custom wrapper execution identity > refuses a wrapper symlink retarget before delivering private values 505ms + ✓ custom wrapper execution identity > bounds wrapper execution after acknowledgement 557ms + ✓ custom wrapper execution identity > bounds captured wrapper output 553ms + ✓ custom wrapper execution identity > refuses a duplicate execute protocol frame 552ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a conforming wrapper leaks a stdio pipe to a background helper 1867ms + ✓ custom wrapper execution bounds are reader-owned > resolves when the leaked helper inherits stderr only 1849ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a wrapper leaks a stdio pipe and exits before identifying 3464ms + ✓ custom wrapper execution bounds are reader-owned > journals a completionReason at the default bound when a wrapper leaks a stdio pipe 11584ms + ✓ custom wrapper execution bounds are reader-owned > accepts an execute token and an over-8KiB payload flushed in one write 506ms + ✓ custom wrapper execution bounds are reader-owned > accepts the same over-8KiB payload whether or not it coalesces with the execute token 1499ms + ✓ custom wrapper execution bounds are reader-owned > still bounds an un-terminated handshake buffer and names the bound 507ms + ✓ delivers the journaled memory pack to the real wrapper and excludes its charge from completion usage 466ms + ✓ tests/step-lease.test.ts (36 tests) 66522ms + ✓ f.run leases against the live kernel > enforces 10000 ms for 'sleep 5; printf ok' 5084ms + ✓ f.run leases against the live kernel > enforces 40000 ms for 'sleep 31; printf ok' 31086ms + ✓ f.run leases against the live kernel > enforces 30000 ms for 'sleep 31; printf ok' 30128ms + + Test Files 13 passed (13) + Tests 316 passed (316) + Start at 20:46:54 + Duration 83.60s (transform 1.76s, setup 138ms, collect 6.15s, tests 132.20s, environment 2ms, prepare 666ms) + diff --git a/evidence/agent-timeout/isolated-live-kernel.txt b/evidence/agent-timeout/isolated-live-kernel.txt new file mode 100644 index 000000000..07a5c64a6 --- /dev/null +++ b/evidence/agent-timeout/isolated-live-kernel.txt @@ -0,0 +1,73 @@ + + RUN v2.1.9 /tmp/agent-timeout-baseline/packages/sdk + +stdout | tests/live-kernel.test.ts +LIVE_KERNEL relayflowd=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd +LIVE_KERNEL flows=/tmp/agent-timeout-baseline/packages/sdk/dist/cli.js + +stdout | tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +LIVE_ANALYZER ready: claude -p --model claude-haiku-4-5-20251001 round-trip OK + +stdout | tests/live-kernel.test.ts > surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once +LIVE_KERNEL kill -9 pid=23748 run=01M41RDPD3E4N5M9G1G9MFVMHV while step=two state=Running + + ❯ tests/live-kernel.test.ts (32 tests | 1 failed) 63296ms + ✓ built flows CLI against live relayflowd > twenty-six-step reuses 25 durable completions after editing the failed final step 2390ms + ✓ built flows CLI against live relayflowd > runs rung (a), parks rung (b), and keeps JSON report-shaped 2885ms + ✓ built flows CLI against live relayflowd > allows a deterministic run to exceed the bounded request timeout 32583ms + ✓ built flows CLI against live relayflowd > follows a live worker dispatch through flows run 609ms + ✓ built flows CLI against live relayflowd > runs an agent CLI end to end through the SDK worker 629ms + ✓ built flows CLI against live relayflowd > f.agent lowers to a real agent step and dispatches through a live worker 660ms + ✓ built flows CLI against live relayflowd > can always get a parked run to a late-attaching worker 5586ms + ✓ built flows CLI against live relayflowd > reports a real manual-recovery NeedsHuman state as parked 494ms + ✓ built flows CLI against live relayflowd > runs hn-monitor analyze-story end-to-end via a stub agent CLI (gate 2 clause 2 demo) 611ms + ✓ built flows CLI against live relayflowd > hn-monitor analyze-story FAILS verification when the CLI omits required schema fields 607ms + ✓ built flows CLI against live relayflowd > agent step preserves the CliResult wrapper as output when the CLI emits non-JSON text 561ms + ✓ built flows CLI against live relayflowd > AgentWorker exposes wake_context to the CLI via RELAYFLOW_WAKE_CONTEXT env var (real analyzer prerequisite) 640ms + ✓ built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_WAKE_CONTEXT UNSET when the run has no wake_context (undefined-vs-null pin) 594ms + ✓ built flows CLI against live relayflowd > AgentWorker passes a declared model to an identified wrapper as RELAYFLOW_MODEL 656ms + ✓ built flows CLI against live relayflowd > AgentWorker refuses a nonconforming journal-submitted wrapper before exposing RELAYFLOW_MODEL 510ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw claude adapter with its real model flag 1250ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw codex adapter with its real model flag 595ms + ✓ built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_MODEL UNSET when the step declares no model 561ms + × built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI 4897ms + → expected { …(3) } to match object { Object (gate, verdict) } +(1 matching property omitted from actual) + ✓ built flows CLI against live relayflowd > preflights before journaling and names an unreachable socket 897ms + ✓ built flows CLI against live relayflowd > starts exactly one daemon when two runs race for one empty data dir 531ms + ✓ subprocess_gate output capture against live relayflowd > journals the gate command output and reports both tails 1055ms + ✓ surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once 980ms + ✓ a relayflow can be scheduled: tick source against live relayflowd > a tick spawns a real run whose step reports the SCHEDULED instant 631ms + ✓ a relayflow can be scheduled: tick source against live relayflowd > TWO ticks for ONE scheduled instant produce exactly ONE run 1135ms + +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +AssertionError: expected { …(3) } to match object { Object (gate, verdict) } +(1 matching property omitted from actual) + +- Expected ++ Received + + Object { +- "gate": "json_schema", +- "verdict": "pass", ++ "gate": "execution", ++ "verdict": "fail", + } + + ❯ tests/live-kernel.test.ts:1317:49 + 1315| // the promoted output, not the test re-deriving it: the spec is + 1316| // the unmodified canonical one, so this record is the gate. + 1317| expect(stepCompleted!.payload.verification).toMatchObject({ + | ^ + 1318| gate: 'json_schema', + 1319| verdict: 'pass', + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯ + + Test Files 1 failed (1) + Tests 1 failed | 31 passed (32) + Start at 20:47:49 + Duration 65.28s (transform 1.10s, setup 17ms, collect 1.79s, tests 63.30s, environment 0ms, prepare 46ms) + diff --git a/evidence/agent-timeout/kernel.txt b/evidence/agent-timeout/kernel.txt new file mode 100644 index 000000000..ef6a14383 --- /dev/null +++ b/evidence/agent-timeout/kernel.txt @@ -0,0 +1,160 @@ + Compiling syn v3.0.4 + Compiling zerovec-derive v0.11.6 + Compiling displaydoc v0.2.7 + Compiling serde_derive v1.0.229 + Compiling ref-cast-impl v1.0.27 + Compiling ref-cast v1.0.27 + Compiling thiserror-impl v2.0.20 + Compiling zerotrie v0.2.5 + Compiling zerovec v0.11.8 + Compiling thiserror v2.0.20 + Compiling tinystr v0.8.4 + Compiling potential_utf v0.1.6 + Compiling icu_collections v2.3.0 + Compiling icu_locale_core v2.3.0 + Compiling serde v1.0.229 + Compiling icu_provider v2.3.1 + Compiling icu_properties v2.3.0 + Compiling icu_normalizer v2.3.0 + Compiling fluent-uri v0.3.2 + Compiling ahash v0.8.12 + Compiling email_address v0.2.9 + Compiling ulid v1.2.1 + Compiling referencing v0.33.0 + Compiling idna_adapter v1.2.2 + Compiling idna v1.1.0 + Compiling jsonschema v0.33.0 + Compiling relayflowd-core v0.1.0 (/home/daytona/.relayflow-v2-supervisor/durable/repository/kernel/relayflowd-core) + Finished `test` profile [unoptimized + debuginfo] target(s) in 8.93s + Running unittests src/lib.rs (/home/daytona/.relayflows-toolchain/target/2962130851/debug/deps/relayflowd_core-ad6b37c3a492c16b) + +running 86 tests +test channel::tests::malformed_payloads_and_invalid_new_channel_appends_leave_state_unchanged ... ok +test channel::tests::send_retry_is_stable_and_conflicting_content_is_rejected ... ok +test channel::tests::forged_deliveries_and_acknowledgements_fail_closed ... ok +test clock::tests::simulated_clock_is_explicitly_advanced ... ok +test channel::tests::delivery_replay_and_independent_acknowledged_offsets ... ok +test entry::attempt_started_tests::a_pre_field_attempt_started_still_reads ... ok +test entry::attempt_started_tests::max_transport_retries_is_always_journaled ... ok +test entry::completion_reason_tests::all_covers_every_serialized_label ... ok +test entry::completion_reason_tests::every_journal_label_matches_serialized ... ok +test entry::reported_cost_tests::step_completed_without_reported_cost_round_trips_unchanged ... ok +test journal::tests::memory_journal_assigns_sequences_and_rolls_epochs ... ok +test machine::parallel_tests::every_declared_mutable_surface_participates_in_conflict_selection ... ok +test machine::parallel_tests::external_ancestor_and_descendant_paths_conflict_but_siblings_do_not ... ok +test machine::parallel_tests::crash_resume_preserves_each_parallel_lease_exactly_once ... ok +test machine::parallel_tests::machine_starts_every_runnable_step_in_authored_order ... ok +test machine::parallel_tests::failed_run_drains_open_siblings_before_terminal_entry ... ok +test machine::parallel_tests::overlapping_agent_surfaces_are_serialized_in_authored_order ... ok +test machine::parallel_tests::workspace_ancestor_and_descendant_paths_conflict_but_siblings_do_not ... ok +test machine::recovery_tests::manual_recovery_parks_a_worker_reported_transport_loss_instead_of_redispatching ... ok +test machine::parallel_tests::parallel_lanes_do_not_cross_the_dependency_barrier_early ... ok +test machine::parallel_tests::disjoint_agent_lanes_merge_pins_in_either_completion_order ... ok +test machine::recovery_tests::recovery_journals_the_wait_human_a_torn_manual_park_never_wrote ... ok +test machine::recovery_tests::reset_recovery_still_retries_a_worker_reported_transport_loss ... ok +test machine::tests::a_retryable_worker_reported_reason_is_not_rewritten_by_the_retry_branch ... ok +test machine::tests::cancel_request_closes_the_active_lease_before_the_terminal_fact ... ok +test machine::tests::all_backing_off_steps_return_timers ... ok +test machine::tests::crashed_attempt_does_not_consume_an_iteration ... ok +test machine::tests::deterministic_lease_rejects_invalid_and_foreign_fields ... ok +test machine::tests::classified_transport_retry_is_bounded_and_preserves_pin_and_idempotency ... ok +test machine::tests::deterministic_lease_override_and_default_are_journaled ... ok +test machine::tests::durable_cancel_request_outranks_crash_recovery ... ok +test machine::tests::every_reason_label_matches_its_serialized_form ... ok +test machine::tests::each_failed_attempt_journals_its_own_output_beside_an_identical_verdict ... ok +test machine::tests::failed_deterministic_completion_preserves_exit_code_and_stderr ... ok +test machine::tests::machine_starts_runnable_step_with_stable_effect_key ... ok +test machine::tests::inspect_recovery_injects_the_dirty_pin_completion_reason_and_tail ... ok +test machine::tests::ordinary_worker_error_is_not_retried_by_either_budget ... ok +test machine::tests::every_failed_run_terminates_with_declared_completion_reasons ... ok +test machine::tests::manual_recovery_parks_needs_human_and_never_redispatches ... ok +test machine::tests::repeated_cancel_request_is_idempotent ... ok +test machine::tests::successful_memo_is_never_scheduled_again ... ok +test machine::tests::the_exhaustion_label_replaces_verification_failed_over_identical_evidence ... ok +test machine::tests::reset_recovery_dispatches_the_original_pinned_revision ... ok +test machine::tests::refused_dispatch_is_re_elected_within_max_iterations_only ... ok +test machine::tests::verification_failure_schedules_a_durable_retry ... ok +test memory::tests::caps_compare_exact_decimals_and_each_token_dimension ... ok +test machine::tests::worker_reported_failure_without_detail_still_records_a_verification ... ok +test retry::tests::jitter_is_repeatable_and_bounded ... ok +test schema::tests::a_property_named_ref_is_not_a_reference ... ok +test schema::tests::in_document_uri_references_resolve_to_the_node_they_name ... ok +test schema::tests::references_the_bound_leaves_opaque_are_refused_by_the_engine ... ok +test schema::tests::refusal_names_the_cycle_it_found ... ok +test schema::tests::shared_declarations_and_boolean_schemas_are_validated ... ok +test spec::tests::a_misspelled_step_level_key_is_a_parse_error ... ok +test spec::tests::a_misspelled_verification_gate_key_is_a_parse_error_not_a_dropped_gate ... ok +test spec::tests::agent_cwd_is_carried_and_must_be_run_root_relative ... ok +test spec::tests::agent_without_cwd_hashes_as_before ... ok +test spec::tests::cli_identity_is_internal_to_provider_steps ... ok +test spec::tests::cycles_are_rejected ... ok +test spec::tests::external_surface_paths_must_have_one_canonical_spelling ... ok +test spec::tests::on_non_zero_parses_the_two_declared_policies_and_refuses_anything_else ... ok +test spec::tests::preflight_data_is_fail_closed ... ok +test schema::tests::every_accepted_corpus_schema_is_accepted ... ok +test schema::tests::every_refused_corpus_schema_compiles_but_is_refused_by_the_bound ... ok +test spec::tests::spec_version_is_semver_and_gated ... ok +test spec::tests::the_default_policy_is_absent_from_the_serialized_boundary_spec ... ok +test spec::tests::the_full_ladder_parses_in_the_one_dialect ... ok +test spec::tests::unknown_root_and_nested_fields_are_rejected ... ok +test spec::tests::workspace_mounts_and_worktrees_must_have_one_canonical_spelling ... ok +test spec::tests::zero_agent_flow_is_valid ... ok +test state::budget::tests::adds_costs_exactly_beyond_machine_decimal_precision ... ok +test state::budget::tests::overflow_and_malformed_cost_leave_total_unchanged ... ok +test state::tests::a_completion_that_omits_a_surface_does_not_drop_it_from_the_pin_chain ... ok +test state::tests::budget_decimal_strings_add_without_floats ... ok +test state::tests::end_pin_chain_is_enforced_and_a_broken_chain_is_a_hard_error ... ok +test state::tests::journal_replays_data_gate_verdict_without_rerunning_completed_code ... ok +test verify::tests::a_recorded_nonzero_exit_passes_its_gate_and_names_the_code ... ok +test verify::tests::an_unbounded_schema_in_a_journal_fails_its_gate_instead_of_aborting ... ok +test verify::tests::deterministic_output_requires_successful_exit_and_content ... ok +test verify::tests::json_schema_is_a_control_gate ... ok +test verify::tests::recording_an_exit_code_does_not_relax_a_declared_content_gate ... ok +test verify::tests::recording_does_not_absorb_a_missing_or_sentinel_exit_status ... ok +test verify::tests::the_default_policy_still_fails_a_nonzero_exit ... ok +test spec::tests::sdk_boundary_accepts_a_valid_10_000_step_reverse_chain ... ok +test spec::tests::sdk_boundary_rejects_a_10_000_step_cycle_with_a_typed_error ... ok +test schema::tests::deeply_nested_schemas_do_not_overflow_the_checker ... ok + +test result: ok. 86 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.71s + + Running tests/memoization.rs (/home/daytona/.relayflows-toolchain/target/2962130851/debug/deps/memoization-8b67c7279372dfe2) + +running 4 tests +test distinct_large_kernel_integers_do_not_alias_through_float_rounding ... ok +test match_reuses_output_with_provenance_and_zero_cost_without_dispatch ... ok +test changed_spec_or_input_dispatches_and_legacy_or_failed_records_miss ... ok +test canonical_corpus_agrees_with_typescript_and_key_permutations ... ok + +test result: ok. 4 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.01s + + Running tests/spec_parity.rs (/home/daytona/.relayflows-toolchain/target/2962130851/debug/deps/spec_parity-1a3112c1ec34bebe) + +running 18 tests +test agent_cwd_is_not_accepted_on_other_verbs_or_under_another_name ... ok +test a_non_string_agent_cwd_fails_closed_rather_than_defaulting ... ok +test agent_timeout_bounds_match_the_kernel_corpus ... ok +test agent_timeout_has_identical_canonical_bytes_and_hash ... ok +test an_undeclared_agent_cwd_is_not_serialized ... ok +test agent_cwd_declaration_acceptance_matches_the_sdk_corpus ... ok +test memory_declaration_acceptance_matches_the_sdk_corpus ... ok +test placement_requirements_have_identical_canonical_bytes_and_hash ... ok +test agent_working_directories_have_identical_canonical_bytes_and_hash ... ok +test step_memory_has_identical_canonical_bytes_and_hash ... ok +test placement_declaration_acceptance_matches_the_sdk_corpus ... ok +test the_kernel_parses_the_deterministic_rung_and_stamps_the_same_hash ... ok +test the_kernel_parses_the_event_triggered_spec_and_stamps_the_same_hash ... ok +test the_kernel_parses_the_rung_c_agent_spec_and_stamps_the_same_hash ... ok +test the_kernel_round_trips_declared_agent_transports_and_rejects_unknown_values ... ok +test the_kernel_parses_the_sdk_compiled_spec_and_stamps_the_same_hash ... ok +test the_kernel_parses_the_advisory_repair_spec_and_stamps_the_same_hash ... ok +test the_kernel_parses_the_rung_b_spec_and_stamps_the_same_hash ... ok + +test result: ok. 18 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.03s + + Doc-tests relayflowd_core + +running 0 tests + +test result: ok. 0 passed; 0 failed; 0 ignored; 0 measured; 0 filtered out; finished in 0.00s + diff --git a/evidence/agent-timeout/original-analyzer.txt b/evidence/agent-timeout/original-analyzer.txt new file mode 100644 index 000000000..533542eba --- /dev/null +++ b/evidence/agent-timeout/original-analyzer.txt @@ -0,0 +1,46 @@ + + RUN v2.1.9 /tmp/agent-timeout-original/packages/sdk + +stdout | tests/live-kernel.test.ts +LIVE_KERNEL relayflowd=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd +LIVE_KERNEL flows=/tmp/agent-timeout-original/packages/sdk/dist/cli.js + +stdout | tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +LIVE_ANALYZER ready: claude -p --model claude-haiku-4-5-20251001 round-trip OK + + ❯ tests/live-kernel.test.ts (32 tests | 1 failed | 31 skipped) 4356ms + × built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI 4355ms + → expected { …(3) } to match object { Object (gate, verdict) } +(1 matching property omitted from actual) + +⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +AssertionError: expected { …(3) } to match object { Object (gate, verdict) } +(1 matching property omitted from actual) + +- Expected ++ Received + + Object { +- "gate": "json_schema", +- "verdict": "pass", ++ "gate": "execution", ++ "verdict": "fail", + } + + ❯ tests/live-kernel.test.ts:1317:49 + 1315| // the promoted output, not the test re-deriving it: the spec is + 1316| // the unmodified canonical one, so this record is the gate. + 1317| expect(stepCompleted!.payload.verification).toMatchObject({ + | ^ + 1318| gate: 'json_schema', + 1319| verdict: 'pass', + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/1]⎯ + + Test Files 1 failed (1) + Tests 1 failed | 31 skipped (32) + Start at 20:51:14 + Duration 6.08s (transform 1000ms, setup 16ms, collect 1.54s, tests 4.36s, environment 0ms, prepare 43ms) + diff --git a/evidence/agent-timeout/regression-tests.txt b/evidence/agent-timeout/regression-tests.txt new file mode 100644 index 000000000..6596b7357 --- /dev/null +++ b/evidence/agent-timeout/regression-tests.txt @@ -0,0 +1,329 @@ + + RUN v2.1.9 /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk + +stdout | tests/live-kernel.test.ts +LIVE_KERNEL relayflowd=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd +LIVE_KERNEL flows=/home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk/dist/cli.js + + ✓ tests/stop-process-group.test.ts (11 tests) 14744ms + ✓ every stop reaches the process group, not just the direct child > exits the run after an execution-timeout stop 990ms + ✓ every stop reaches the process group, not just the direct child > exits the run after a protocol terminate stop 547ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after a protocol terminate stop 1544ms + ✓ every stop reaches the process group, not just the direct child > kills a SIGTERM-deaf grandchild after an execution-timeout stop 1928ms + ✓ every stop reaches the process group, not just the direct child > holds the loop open long enough for the escalation to run 1095ms + ✓ every stop reaches the process group, not just the direct child > terminate() forces a group that outlives SIGTERM 589ms + ✓ every stop reaches the process group, not just the direct child > bounds forced-stop confirmation when a group remains unprovable 1008ms + ✓ a wrapper that exits with no execution deadline still drains > reports the wrapper result and reaps a grandchild holding its pipes 778ms + ✓ a wrapper that exits with no execution deadline still drains > reaps a SIGTERM-deaf grandchild holding its pipes 1869ms + ✓ a wrapper that exits with no execution deadline still drains > settles on its own deadline when an escaped holder withholds close 4319ms + ✓ tests/authored-node-result.test.ts (45 tests) 22ms + ❯ tests/verb-field-lint.test.ts (99 tests | 1 failed) 345ms + × closed per-verb step fields > pins the per-verb descriptor and generates every foreign-field pair from it 7ms + → expected { …(3) } to deeply equal { …(3) } + ✓ tests/wrapper-execution-duration.test.ts (7 tests) 10873ms + ✓ keeps the handshake deadline independent of the removed execution deadline 10060ms + ✓ still lets a lease abort stop an unlimited wrapper before it produces output 513ms + ✓ tests/authored-agent-artifacts.test.ts (5 tests) 355ms + ✓ tests/spec-parity.test.ts (46 tests) 500ms + ✓ tests/worker-cli.test.ts (25 tests) 31000ms + ✓ registered CLI model defaults > passes the same priced Claude default to the real provider invocation 538ms + ✓ registered CLI model defaults > uses the explicitly supplied agent environment for the provider subprocess 535ms + ✓ registered CLI model defaults > dispatches canonical generic bytes with the preflight-proved adapter identity 495ms + ✓ direct transport lifecycle evidence > classifies only the exact Codex stdin lifecycle signature as retryable 493ms + ✓ direct transport lifecycle evidence > records a signal close separately from an ordinary nonzero exit 1031ms + ✓ direct transport lifecycle evidence > records a spawn error code without treating a missing executable as transient 411ms + ✓ direct transport lifecycle evidence > journals classified lifecycle evidence and reports crashed instead of generic worker_error 668ms + ✓ step discovery environment > names the run, step, attempt and an absolute data dir for a direct agent spawn 616ms + ✓ step discovery environment > exports none of the four without a data dir, even when the worker inherited them 500ms + ✓ wrapper discovery environment > sets the four names from the dispatch and still refuses ambient values and other secrets 630ms + ✓ wrapper discovery environment > exports none of the four to a wrapper without a data dir, even when the worker inherited them 520ms + ✓ custom wrapper execution identity > passes an explicit safe environment at identification and execution 513ms + ✓ custom wrapper execution identity > refuses a wrapper symlink retarget before delivering private values 469ms + ✓ custom wrapper execution identity > bounds wrapper execution after acknowledgement 562ms + ✓ custom wrapper execution identity > bounds captured wrapper output 469ms + ✓ custom wrapper execution identity > refuses a duplicate execute protocol frame 545ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a conforming wrapper leaks a stdio pipe to a background helper 1778ms + ✓ custom wrapper execution bounds are reader-owned > resolves when the leaked helper inherits stderr only 1824ms + ✓ custom wrapper execution bounds are reader-owned > resolves when a wrapper leaks a stdio pipe and exits before identifying 3563ms + ✓ custom wrapper execution bounds are reader-owned > journals a completionReason at the default bound when a wrapper leaks a stdio pipe 11522ms + ✓ custom wrapper execution bounds are reader-owned > accepts an execute token and an over-8KiB payload flushed in one write 540ms + ✓ custom wrapper execution bounds are reader-owned > accepts the same over-8KiB payload whether or not it coalesces with the execute token 1720ms + ✓ custom wrapper execution bounds are reader-owned > still bounds an un-terminated handshake buffer and names the bound 513ms + ✓ delivers the journaled memory pack to the real wrapper and excludes its charge from completion usage 543ms + ✓ tests/authored-retried-child-resume-live.test.ts (2 tests) 1564ms + ✓ an authored step whose child run retried an attempt > resumes after a kill mid-step and reads the terminal completion, not the crashed one 1419ms + ✓ tests/agent-timeout-live.test.ts (3 tests) 4485ms + ✓ journals timeout, stops the process, runs a predicate gate and publishes under a budget header 1261ms + ✓ replays timeout after killing the root and daemon, without executing the agent again 1969ms + ✓ lets an author reject timeout through a predicate gate 1255ms + ✓ tests/agent-timeout-worker.test.ts (8 tests) 1975ms + ✓ stops claude at its declared limit, preserving work and stopping descendants 523ms + ✓ stops codex at its declared limit, preserving work and stopping descendants 516ms + ✓ stops wrapper at its declared limit, preserving work and stopping descendants 536ms + ✓ tests/agent-timeout.test.ts (23 tests) 37ms +stdout | tests/live-kernel.test.ts > surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once +LIVE_KERNEL kill -9 pid=18853 run=01M41R3AJWBNG7DZZCDAVFFQDC while step=two state=Running + + ❯ tests/live-kernel.test.ts (32 tests | 9 failed) 57453ms + ✓ built flows CLI against live relayflowd > twenty-six-step reuses 25 durable completions after editing the failed final step 2651ms + ✓ built flows CLI against live relayflowd > runs rung (a), parks rung (b), and keeps JSON report-shaped 3629ms + ✓ built flows CLI against live relayflowd > allows a deterministic run to exceed the bounded request timeout 32485ms + ✓ built flows CLI against live relayflowd > follows a live worker dispatch through flows run 636ms + ✓ built flows CLI against live relayflowd > runs an agent CLI end to end through the SDK worker 516ms + ✓ built flows CLI against live relayflowd > f.agent lowers to a real agent step and dispatches through a live worker 738ms + ✓ built flows CLI against live relayflowd > can always get a parked run to a late-attaching worker 5585ms + ✓ built flows CLI against live relayflowd > reports a real manual-recovery NeedsHuman state as parked 463ms + × built flows CLI against live relayflowd > runs hn-monitor analyze-story end-to-end via a stub agent CLI (gate 2 clause 2 demo) 572ms + → expected { …(12) } to match object { output: { …(3) }, …(1) } +(22 matching properties omitted from actual) + × built flows CLI against live relayflowd > hn-monitor analyze-story FAILS verification when the CLI omits required schema fields 655ms + → expected { …(12) } to match object { …(3) } +(21 matching properties omitted from actual) + × built flows CLI against live relayflowd > agent step preserves the CliResult wrapper as output when the CLI emits non-JSON text 605ms + → expected null not to be null + × built flows CLI against live relayflowd > AgentWorker exposes wake_context to the CLI via RELAYFLOW_WAKE_CONTEXT env var (real analyzer prerequisite) 718ms + → Cannot read properties of null (reading 'story_title') + × built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_WAKE_CONTEXT UNSET when the run has no wake_context (undefined-vs-null pin) 562ms + → Cannot read properties of null (reading 'env_present') + ✓ built flows CLI against live relayflowd > AgentWorker passes a declared model to an identified wrapper as RELAYFLOW_MODEL 594ms + ✓ built flows CLI against live relayflowd > AgentWorker refuses a nonconforming journal-submitted wrapper before exposing RELAYFLOW_MODEL 566ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw claude adapter with its real model flag 517ms + ✓ built flows CLI against live relayflowd > AgentWorker executes the raw codex adapter with its real model flag 580ms + × built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_MODEL UNSET when the step declares no model 642ms + → Cannot read properties of null (reading 'story_title') + × built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI 31ms + → LIVE_ANALYZER_UNAVAILABLE: "/home/daytona/.relayflow-v2-supervisor/durable/repository/testdata/preflight/analyze-story-claude-cli" does not identify as relayflows-agent-cli-v1 — failing because gate-2 acceptance requires the real analyzer to execute. Set RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 only if this run is not gate evidence. + ✓ built flows CLI against live relayflowd > preflights before journaling and names an unreachable socket 873ms + × built flows CLI against live relayflowd > starts exactly one daemon when two runs race for one empty data dir 505ms + → WARNING [unprovable_effects] Step "greet" command "echo" resolves, but its effects cannot be proven before execution. +WARNING [unprovable_effects] Step "shout" command "echo" resolves, but its effects cannot be proven before execution. +WARNING [editor_schema_missing] For editor validation, add this first line: # yaml-language-server: $schema=https://schema.relayflows.dev/v0.1/flows.schema.json +REFUSED [relayflowd_not_found] No relayflowd binary could be found. Install the runtime package for this host (@relayflows/runtime-linux-x64), or set RELAYFLOWD_BIN to a relayflowd executable. Tried: /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk/dist/relayflowd. +: expected 2 to be +0 // Object.is equality + ✓ subprocess_gate output capture against live relayflowd > journals the gate command output and reports both tails 1039ms + ✓ surface resume after a real daemon kill > resumes a three-step run with each successful completion exactly once 944ms + × a relayflow can be scheduled: tick source against live relayflowd > a tick spawns a real run whose step reports the SCHEDULED instant 610ms + → expected null to deeply equal { schedule_id: 'heartbeat-1m', …(3) } + ✓ tests/step-lease.test.ts (36 tests) 66623ms + ✓ f.run leases against the live kernel > enforces 10000 ms for 'sleep 5; printf ok' 5089ms + ✓ f.run leases against the live kernel > enforces 40000 ms for 'sleep 31; printf ok' 31170ms + ✓ f.run leases against the live kernel > enforces 30000 ms for 'sleep 31; printf ok' 30125ms + +⎯⎯⎯⎯⎯⎯ Failed Tests 10 ⎯⎯⎯⎯⎯⎯⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > runs hn-monitor analyze-story end-to-end via a stub agent CLI (gate 2 clause 2 demo) +AssertionError: expected { …(12) } to match object { output: { …(3) }, …(1) } +(22 matching properties omitted from actual) + +- Expected ++ Received + + Object { +- "output": Object { +- "reasoning": "stub agent runtime — deterministic output for gate-2 clause-2 demo", +- "relevance_score": 5, +- "story_title": "stub", +- }, ++ "output": null, + "verification": Object { +- "gate": "json_schema", +- "verdict": "pass", ++ "gate": "execution", ++ "verdict": "fail", + }, + } + + ❯ tests/live-kernel.test.ts:657:36 + 655| && (entry as { step_id?: string }).step_id === 'analyze-story', + 656| ) as { payload: { output: unknown; verification: unknown } } | und… + 657| expect(stepCompleted?.payload).toMatchObject({ + | ^ + 658| output: { + 659| story_title: 'stub', + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story FAILS verification when the CLI omits required schema fields +AssertionError: expected { …(12) } to match object { …(3) } +(21 matching properties omitted from actual) + +- Expected ++ Received + + Object { +- "completionReason": "retries_exhausted", ++ "completionReason": "worker_error", + "output": null, + "verification": Object { +- "gate": "json_schema", ++ "gate": "execution", + "verdict": "fail", + }, + } + + ❯ tests/live-kernel.test.ts:752:36 + 750| // its verification record names the json_schema rejection. The re… + 751| // parsed value is nulled before the completion is persisted. + 752| expect(stepCompleted?.payload).toMatchObject({ + | ^ + 753| completionReason: 'retries_exhausted', + 754| output: null, + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > agent step preserves the CliResult wrapper as output when the CLI emits non-JSON text +AssertionError: expected null not to be null + ❯ tests/live-kernel.test.ts:823:24 + 821| // here (parseJsonOutput returned null on non-JSON stdout) and + 822| // these assertions would all fail. + 823| expect(output).not.toBeNull(); + | ^ + 824| expect(output.exit_code).toBe(0); + 825| expect(output.stdout_tail).toContain('looked at the story'); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[3/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > AgentWorker exposes wake_context to the CLI via RELAYFLOW_WAKE_CONTEXT env var (real analyzer prerequisite) +TypeError: Cannot read properties of null (reading 'story_title') + ❯ tests/live-kernel.test.ts:891:42 + 889| ) as { payload: { output: { story_title: string; reasoning: string… + 890| expect(stepCompleted).toBeDefined(); + 891| expect(stepCompleted!.payload.output.story_title).toBe(`echoed:${s… + | ^ + 892| expect(stepCompleted!.payload.output.reasoning).toContain(String(s… + 893| + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[4/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_WAKE_CONTEXT UNSET when the run has no wake_context (undefined-vs-null pin) +TypeError: Cannot read properties of null (reading 'env_present') + ❯ tests/live-kernel.test.ts:958:38 + 956| ) as { payload: { output: { env_present: boolean } } } | undefined; + 957| expect(completed).toBeDefined(); + 958| expect(completed!.payload.output.env_present).toBe(false); + | ^ + 959| + 960| delete process.env.RELAYFLOW_WAKE_CONTEXT; + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[5/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > AgentWorker leaves RELAYFLOW_MODEL UNSET when the step declares no model +TypeError: Cannot read properties of null (reading 'story_title') + ❯ tests/live-kernel.test.ts:1194:38 + 1192| expect(completed).toBeDefined(); + 1193| // UNSET, not EMPTY and not the leaked parent value. + 1194| expect(completed!.payload.output.story_title).toBe('model:UNSET'); + | ^ + 1195| + 1196| delete process.env.RELAYFLOW_MODEL; + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[6/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > hn-monitor analyze-story reaches done through the real Claude analyzer CLI +Error: LIVE_ANALYZER_UNAVAILABLE: "/home/daytona/.relayflow-v2-supervisor/durable/repository/testdata/preflight/analyze-story-claude-cli" does not identify as relayflows-agent-cli-v1 — failing because gate-2 acceptance requires the real analyzer to execute. Set RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 only if this run is not gate evidence. + ❯ tests/live-kernel.test.ts:1223:15 + 1221| const notice = `LIVE_ANALYZER_UNAVAILABLE: ${readiness.detail}`; + 1222| if (process.env['RELAYFLOWS_ALLOW_ANALYZER_SKIP'] !== '1') { + 1223| throw new Error( + | ^ + 1224| `${notice} — failing because gate-2 acceptance requires the … + 1225| + 'Set RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 only if this run is … + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[7/10]⎯ + + FAIL tests/live-kernel.test.ts > built flows CLI against live relayflowd > starts exactly one daemon when two runs race for one empty data dir +AssertionError: WARNING [unprovable_effects] Step "greet" command "echo" resolves, but its effects cannot be proven before execution. +WARNING [unprovable_effects] Step "shout" command "echo" resolves, but its effects cannot be proven before execution. +WARNING [editor_schema_missing] For editor validation, add this first line: # yaml-language-server: $schema=https://schema.relayflows.dev/v0.1/flows.schema.json +REFUSED [relayflowd_not_found] No relayflowd binary could be found. Install the runtime package for this host (@relayflows/runtime-linux-x64), or set RELAYFLOWD_BIN to a relayflowd executable. Tried: /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk/dist/relayflowd. +: expected 2 to be +0 // Object.is equality + +- Expected ++ Received + +- 0 ++ 2 + + ❯ tests/live-kernel.test.ts:1388:40 + 1386| ]); + 1387| + 1388| expect(first.status, first.stderr).toBe(0); + | ^ + 1389| expect(second.status, second.stderr).toBe(0); + 1390| expect(first.stdout).toContain('completionReason: success'); + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[8/10]⎯ + + FAIL tests/live-kernel.test.ts > a relayflow can be scheduled: tick source against live relayflowd > a tick spawns a real run whose step reports the SCHEDULED instant +AssertionError: expected null to deeply equal { schedule_id: 'heartbeat-1m', …(3) } + +- Expected: +Object { + "lag_ms": 43000, + "schedule_id": "heartbeat-1m", + "scheduled_for_ms": 1764000000000, + "slot": 29400000, +} + ++ Received: +null + + ❯ tests/live-kernel.test.ts:1739:39 + 1737| // The bound: the run reports the grid instant and its own lag, so… + 1738| // backfilled run can tell it is running for a slot from the past. + 1739| expect(completed!.payload.output).toEqual({ + | ^ + 1740| schedule_id: 'heartbeat-1m', + 1741| slot: 29_400_000, + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[9/10]⎯ + + FAIL tests/verb-field-lint.test.ts > closed per-verb step fields > pins the per-verb descriptor and generates every foreign-field pair from it +AssertionError: expected { …(3) } to deeply equal { …(3) } + +- Expected ++ Received + + Object { + "agent": Array [ ++ "timeoutMs", + "instruction", + "agent", + "cli", + "model", + "cwd", + "transport", + "surfaces", + "recoveryMode", + "permissions", + "output", + ], + "deterministic": Array [ + "command", + "timeoutMs", + "lease_ms", + "onNonZero", + ], + "llm": Array [ + "prompt", + "model", + "cli", + "output", + ], + } + + ❯ tests/verb-field-lint.test.ts:205:33 + 203| 'requirements', + 204| ]); + 205| expect(STEP_FIELDS_BY_TYPE).toEqual({ + | ^ + 206| deterministic: ['command', 'timeoutMs', 'lease_ms', 'onNonZero'], + 207| llm: ['prompt', 'model', 'cli', 'output'], + +⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[10/10]⎯ + + Test Files 2 failed | 11 passed (13) + Tests 10 failed | 332 passed (342) + Start at 20:42:14 + Duration 99.23s (transform 1.66s, setup 81ms, collect 6.14s, tests 189.98s, environment 2ms, prepare 526ms) + diff --git a/evidence/agent-timeout/schema-generate.txt b/evidence/agent-timeout/schema-generate.txt new file mode 100644 index 000000000..29985e5ba --- /dev/null +++ b/evidence/agent-timeout/schema-generate.txt @@ -0,0 +1,5 @@ + +> @relayflows/schema@0.1.0 generate +> node ../../scripts/generate-json-schema.mjs + +Generated packages/schema/flows.schema.json (74 definitions) diff --git a/evidence/agent-timeout/schema-tests.txt b/evidence/agent-timeout/schema-tests.txt new file mode 100644 index 000000000..c30769989 --- /dev/null +++ b/evidence/agent-timeout/schema-tests.txt @@ -0,0 +1,99 @@ + +> @relayflows/schema@0.1.0 test +> bun test tests + +bun test v1.3.6 (d530ed99) + +tests/parity.test.ts: +(pass) flows check fixture parity: advisory-repair.flow.yaml [72.00ms] +(pass) flows check fixture parity: agent-cwd.flow.yaml [11.00ms] +(pass) flows check fixture parity: agent-timeout.flow.yaml [6.00ms] +(pass) flows check fixture parity: backlog-picker.flow.yaml [10.00ms] +(pass) flows check fixture parity: budget-guarded.flow.yaml [5.00ms] +(pass) flows check fixture parity: dir-watcher.flow.yaml [29.00ms] +(pass) flows check fixture parity: hello-agent.flow.yaml [12.00ms] +(pass) flows check fixture parity: hello-deterministic.flow.yaml [6.00ms] +(pass) flows check fixture parity: hello-ladder.flow.yaml [40.00ms] +(pass) flows check fixture parity: hello-llm.flow.yaml [26.00ms] +(pass) flows check fixture parity: hn-monitor.flow.yaml [59.00ms] +(pass) flows check fixture parity: json-schema-invalid.flow.yaml [12.00ms] +(pass) flows check fixture parity: step-memory.flow.yaml [8.00ms] +(pass) flows check fixture parity: step-placement.flow.yaml [4.00ms] +(pass) flows check fixture parity: tick-heartbeat.flow.yaml [56.00ms] +(pass) structural parity: unknown root key [2.00ms] +(pass) structural parity: unsupported version +(pass) structural parity: no steps +(pass) structural parity: step typo +(pass) structural parity: empty command +(pass) structural parity: positive timeout +(pass) structural parity: fractional retry [1.00ms] +(pass) structural parity: negative transport retry +(pass) structural parity: fractional transport retry +(pass) structural parity: zero transport retry +(pass) structural parity: u32 max transport retry +(pass) structural parity: transport retry above u32 +(pass) structural parity: wrong step field [1.00ms] +(pass) structural parity: nonzero exit gate +(pass) structural parity: legacy zero exit gate +(pass) structural parity: boolean schema [2.00ms] +(pass) structural parity: nested invalid schema [12.00ms] +(pass) structural parity: bad memory budget [1.00ms] +(pass) structural parity: unsafe memory budget +(pass) structural parity: empty memory query +(pass) structural parity: unsafe duration +(pass) structural parity: bad money +(pass) structural parity: bad input index +(pass) structural parity: blank input name +(pass) structural parity: trigger silence budget [1.00ms] +(pass) structural parity: workspace: single grant +(pass) structural parity: tools.fs: single grant +(pass) structural parity: workspace: grant array +(pass) structural parity: tools.fs: grant array +(pass) structural parity: workspace: non-string item +(pass) structural parity: tools.fs: non-string item +(pass) structural parity: workspace: nested array +(pass) structural parity: tools.fs: nested array +(pass) structural parity: llm: output object [16.00ms] +(pass) structural parity: llm: boolean output [1.00ms] +(pass) structural parity: llm: output and verification [7.00ms] +(pass) structural parity: llm: exit gate [1.00ms] +(pass) structural parity: llm: trimmed model +(pass) structural parity: llm: control in model +(pass) structural parity: agent: output object [15.00ms] +(pass) structural parity: agent: boolean output [1.00ms] +(pass) structural parity: agent: output and verification [17.00ms] +(pass) structural parity: agent: exit gate [1.00ms] +(pass) structural parity: agent: trimmed model +(pass) structural parity: agent: control in model +(pass) step examples compile and validate [1.00ms] +(pass) generated schema satisfies the existing bounded-reference rule [9.00ms] +(pass) semantic checks remain explicit runtime responsibilities +(pass) embedded dialect http://json-schema.org/draft-04/schema# [17.00ms] +(pass) embedded dialect http://json-schema.org/draft-06/schema# [15.00ms] +(pass) embedded dialect http://json-schema.org/draft-07/schema# [15.00ms] +(pass) embedded dialect https://json-schema.org/draft/2019-09/schema [28.00ms] +(pass) embedded dialect https://json-schema.org/draft/2020-12/schema [29.00ms] +(pass) header hint is warning-only, first-line aware, and never edits input [18.00ms] +(pass) canonical surface parity: "repo" +(pass) canonical surface parity: "/repo/src" +(pass) canonical surface parity: "pr://github/example" +(pass) canonical surface parity: "/" +(pass) canonical surface parity: "pr://" +(pass) canonical surface parity: "" +(pass) canonical surface parity: " repo" +(pass) canonical surface parity: "repo " +(pass) canonical surface parity: "repo//src" +(pass) canonical surface parity: "repo/../src" +(pass) canonical surface parity: "repo/." +(pass) canonical surface parity: ":/bad//path" +(pass) named declarations and selected input paths use authoring shapes [43.00ms] + +tests/smoke.test.ts: +(pass) all exported spec type nodes have documented definitions [24.00ms] +(pass) regeneration is byte-stable and committed schema has not drifted [485.00ms] +(pass) npm tarball contains only data and documentation with no runtime dependencies [304.00ms] + + 85 pass + 0 fail + 4019 expect() calls +Ran 85 tests across 2 files. [2.31s] diff --git a/evidence/agent-timeout/test-typecheck.txt b/evidence/agent-timeout/test-typecheck.txt new file mode 100644 index 000000000..6ae9c3503 --- /dev/null +++ b/evidence/agent-timeout/test-typecheck.txt @@ -0,0 +1,4 @@ + +> @relayflows/sdk@2.0.39 typecheck:tests +> tsc -p tsconfig.tests.json + diff --git a/evidence/agent-timeout/typecheck.txt b/evidence/agent-timeout/typecheck.txt new file mode 100644 index 000000000..ee46739ea --- /dev/null +++ b/evidence/agent-timeout/typecheck.txt @@ -0,0 +1,4 @@ + +> @relayflows/sdk@2.0.39 typecheck +> tsc --noEmit && tsc -p tsconfig.type-tests.json + diff --git a/examples/babysitter/tests/platform.test.ts b/examples/babysitter/tests/platform.test.ts index c440706ef..cece0d07b 100644 --- a/examples/babysitter/tests/platform.test.ts +++ b/examples/babysitter/tests/platform.test.ts @@ -8,7 +8,7 @@ for (const [name, dependency] of Object.entries({ durableCrossRunNotification: 'Durable per-head notification receipt and atomic delivery claim', journaledCiObservation: 'Journal-backed script-memory write/recall across runs (memory.learn currently refuses)', enforcedAgentWriteScope: 'Enforced readonly agent workspace and credential scopes (gate 8 / #442)', - selectiveAgentExitRetry: 'Typed agent exit-code/retry policy in authored Step; AgentResult only has summary/artifacts', + selectiveAgentExitRetry: 'Typed agent exit-code/retry policy in authored Step; AgentResult has summary/artifacts and success/timeout completionReason', durableSubscriptionLiveness: 'Durable cross-run wake records for the liveness sweep; a subscription that stops firing is currently indistinguishable from a quiet one', })) { test(name, { todo: dependency }, () => assert.equal(capabilities[name as keyof typeof capabilities], true)); diff --git a/kernel/relayflowd-core/src/spec.rs b/kernel/relayflowd-core/src/spec.rs index 41bbce492..f3b80fea5 100644 --- a/kernel/relayflowd-core/src/spec.rs +++ b/kernel/relayflowd-core/src/spec.rs @@ -156,6 +156,13 @@ impl RunSpec { step.id ))); } + if let StepKind::Agent { timeout_ms: Some(ms), .. } = &step.kind + && (*ms == 0 || *ms > i64::MAX as u64) + { + return Err(SpecError::Malformed(format!( + "step {}: timeout_ms must be positive and fit in i64", step.id + ))); + } let cli = match &step.kind { StepKind::Llm { cli, .. } | StepKind::Agent { cli, .. } => cli, StepKind::Deterministic { .. } => &None, @@ -350,6 +357,7 @@ const STEP_COMMON_FIELDS: &[&str] = &[ const STEP_DETERMINISTIC_FIELDS: &[&str] = &["command", "timeout_ms", "lease_ms", "on_non_zero"]; const STEP_LLM_FIELDS: &[&str] = &["prompt", "model", "cli", "cli_identity"]; const STEP_AGENT_FIELDS: &[&str] = &[ + "timeout_ms", "instruction", "cli", "cli_identity", @@ -508,6 +516,9 @@ pub enum StepKind { }, Agent { instruction: String, + /// CLI execution deadline, enforced by the attached worker. + #[serde(default, skip_serializing_if = "Option::is_none")] + timeout_ms: Option, #[serde(default, skip_serializing_if = "Option::is_none")] cli: Option, /// Model the declared CLI should use. The kernel never calls a model diff --git a/kernel/relayflowd-core/tests/spec_parity.rs b/kernel/relayflowd-core/tests/spec_parity.rs index c1f6bb6c0..54be5a21f 100644 --- a/kernel/relayflowd-core/tests/spec_parity.rs +++ b/kernel/relayflowd-core/tests/spec_parity.rs @@ -305,3 +305,26 @@ fn agent_cwd_is_not_accepted_on_other_verbs_or_under_another_name() { ); } } + +#[test] +fn agent_timeout_has_identical_canonical_bytes_and_hash() { + assert_parity( + include_str!("../../../testdata/agent-timeout.spec.canonical.json"), + include_str!("../../../testdata/agent-timeout.spec.sha256"), + ); +} + +#[test] +fn agent_timeout_bounds_match_the_kernel_corpus() { + let cases: Vec = serde_json::from_str(include_str!("../../../testdata/agent-timeout-cases.json")).unwrap(); + for case in cases { + let value = serde_json::json!({"steps":[{ + "id":"a", "type":"agent", "instruction":"repair", "timeout_ms":case["timeoutMs"] + }]}); + assert_eq!(RunSpec::parse(&value).and_then(|s| s.validate()).is_ok(), + case["kernelValid"].as_bool().unwrap(), "{}", case["name"]); + } + let value = serde_json::json!({"steps":[{"id":"a", "type":"agent", "instruction":"repair"}]}); + let parsed = RunSpec::parse(&value).unwrap(); + assert!(serde_json::to_value(parsed).unwrap()["steps"][0].get("timeout_ms").is_none()); +} diff --git a/packages/schema/flows.schema.json b/packages/schema/flows.schema.json index 0fc803e8e..9ec59e97b 100644 --- a/packages/schema/flows.schema.json +++ b/packages/schema/flows.schema.json @@ -1389,6 +1389,13 @@ "minimum": 0, "maximum": 4294967295 }, + "timeoutMs": { + "title": "timeoutMs", + "description": "timeoutMs in the Relayflows spec.", + "type": "integer", + "minimum": 1, + "maximum": 3600000 + }, "instruction": { "title": "instruction", "description": "instruction in the Relayflows spec.", @@ -2437,6 +2444,11 @@ "type": "string", "const": "agent" }, + "timeout_ms": { + "title": "timeout_ms", + "description": "timeout_ms in the Relayflows spec.", + "type": "number" + }, "instruction": { "title": "instruction", "description": "instruction in the Relayflows spec.", diff --git a/packages/sdk/src/authored-node-runner.ts b/packages/sdk/src/authored-node-runner.ts index 5d5e88999..cf2c8f075 100644 --- a/packages/sdk/src/authored-node-runner.ts +++ b/packages/sdk/src/authored-node-runner.ts @@ -360,15 +360,15 @@ export async function verifyAuthoredNodeResult( await journal.hello('flows-authored-result-verifier'); for (const [index, claimed] of [...ordered, ...gates].entries()) { if (!claimed || typeof claimed.id !== 'string' || typeof claimed.runId !== 'string' - || (index < ordered.length && ordinal(claimed.id) !== index+1) || claimed.completionReason !== 'success' + || (index < ordered.length && ordinal(claimed.id) !== index+1) || !['success', 'timeout'].includes(claimed.completionReason) || runs.has(claimed.runId)) invalid(`claim ${claimed?.id}`); runs.add(claimed.runId); const state = await journal.runGet(claimed.runId); - if (state.run_id !== claimed.runId || state.status !== 'completed' + if (state.run_id !== claimed.runId || state.status !== (claimed.completionReason === 'timeout' ? 'failed' : 'completed') || state.steps[claimed.id]?.state !== 'done') invalid(`state ${claimed.id} ${state.status} ${state.steps[claimed.id]?.state}`); const entries: Array<{ seq: number; entry_type: string; step_id?: string; payload?: { - completionReason?: string; spec?: { name?: string; steps?: Array<{id?:string;type?:string;command?:string}> }; + completionReason?: string; spec?: { name?: string; steps?: Array<{id?:string;type?:string;command?:string;timeout_ms?:number}> }; }; }> = []; let fromSeq = 1; @@ -397,10 +397,13 @@ export async function verifyAuthoredNodeResult( const completed = settlingCompletion(entries, claimed.id); const gateCompleted = lowered.length === 2 ? settlingCompletion(entries, `${claimed.id}.gate`) : undefined; const terminalFacts = entries.filter(entry => entry.entry_type === 'run.completed'); + const timedOut = claimed.completionReason === 'timeout'; + if (timedOut && (claimed === terminal || lowered.length !== 1 || step?.type !== 'agent' + || !Number.isSafeInteger(step.timeout_ms) || step.timeout_ms! <= 0)) invalid(`timeout declaration ${claimed.id}`); if (spec?.name !== `${metadata.flowName}/${claimed.id}` || !specShape - || step?.id !== claimed.id || completed?.payload?.completionReason !== 'success' + || step?.id !== claimed.id || completed?.payload?.completionReason !== claimed.completionReason || (lowered.length === 2 && gateCompleted?.payload?.completionReason !== 'success') - || terminalFacts.length !== 1 || terminalFacts[0]?.payload?.completionReason !== 'success') invalid(`evidence ${claimed.id} spec=${spec?.name} step=${step?.id} completed=${completed?.payload?.completionReason}`); + || terminalFacts.length !== 1 || terminalFacts[0]?.payload?.completionReason !== (timedOut ? 'step_failed' : 'success')) invalid(`evidence ${claimed.id} spec=${spec?.name} step=${step?.id} completed=${completed?.payload?.completionReason}`); if (claimed === terminal) { // The claimed verdict must match the marker the journal actually // recorded, so an IPC frame cannot claim `success` over a run whose diff --git a/packages/sdk/src/authored-worker-step.ts b/packages/sdk/src/authored-worker-step.ts index 3897fed99..0f163dfc8 100644 --- a/packages/sdk/src/authored-worker-step.ts +++ b/packages/sdk/src/authored-worker-step.ts @@ -2,7 +2,7 @@ import { dirname, resolve, isAbsolute, relative, sep } from 'node:path'; import type { AuthoredBudget } from './authored-budget.js'; import { parseBudget } from './budget.js'; import type { AgentOptions, AgentResult, LlmOptions, NamedGate } from '@relayflows/surface'; -import { compileSpec, toKernelSpec } from './compile.js'; +import { compileSpec, toKernelSpec, parseAgentStepTimeout } from './compile.js'; import { isTransportRetries } from './validate.js'; import { authoredPreflight } from './authored-preflight.js'; import { classifyOutcome, type RunLifecycleOptions, type RunReport } from './cli/run.js'; @@ -19,6 +19,7 @@ import { alsoRecord, recordAuthoredChild } from './authored-step-index.js'; import type { StepFailedDetails } from './failure-kinds.js'; import { authoredWorkerSlots, type AuthoredWorkerSlots } from './worker-slots.js'; +const AGENT_TIMEOUT = Symbol('declared agent timeout'); const WORKSPACE_PERMISSION_ANNOTATION = /:\s*(readonly|readwrite)\s*$/i; /** Agent and LLM calls share the exact declarative preflight and lease wait. */ @@ -118,6 +119,7 @@ export function authoredWorkerRunner( const diagnostic = execution.report.diagnostics.at(-1); const details = stepDetails(diagnostic); const reason = execution.report.completionReason; + let indexFailure = ''; let message = diagnostic?.message ?? `flow "${definition.name}" step "${id}" did not complete successfully ` + `(status: ${execution.report.status ?? 'unknown'})`; @@ -128,13 +130,21 @@ export function authoredWorkerRunner( journalSteps.push(Object.freeze({ id, runId: outcome.run_id, completionReason: details.completionReason, ...stepEdges?.(id), })); - message += await alsoRecord(journal, rootRunId, { + indexFailure = await alsoRecord(journal, rootRunId, { step: id, runId: outcome.run_id, state: 'completed', completionReason: details.completionReason, ...(details.stepId === undefined ? {} : { kernelStep: details.stepId }), ...stepEdges?.(id), }); } + message += indexFailure; + // Resolve inside consume so budget accounting does not latch a failure. + // A failed index append must still fail closed, even for a tolerated timeout. + if (indexFailure === '' && reason === 'step_failed' && execution.report.status === 'failed' + && spec.steps.length === 1 && step.type === 'agent' && step.timeoutMs !== undefined + && details?.stepId === id && details.stepType === 'agent' && details.completionReason === 'timeout') { + return { [AGENT_TIMEOUT]: details.stderrTail ?? message }; + } throw new AuthoredFlowExecutionError( 'step_failed', message, isSurfaceCompletionReason(reason) ? reason : undefined, @@ -244,11 +254,21 @@ export function authoredWorkerRunner( if (cwdTransport !== undefined) { throw new AuthoredFlowExecutionError('agent_cli_unresolved', `f.agent options.${cwdTransport}.`); } + let timeoutMs: number | undefined; + if (options.timeout !== undefined) { + try { timeoutMs = parseAgentStepTimeout(options.timeout); } + catch (error) { throw new AuthoredFlowExecutionError('agent_cli_unresolved', String(error)); } + if (options.transport === 'relay' || (options.maxIterations ?? 1) > 1) { + throw new AuthoredFlowExecutionError('agent_cli_unresolved', + 'f.agent timeout requires direct transport and maxIterations 1.'); + } + } const permissions = options.permissions; const permissionsSnapshot = permissions === undefined ? undefined : snapshotJsonValue(permissions, 'f.agent options.permissions') as unknown as PermissionsSpec; const output = await run({ id, type: 'agent', instruction: options.task, + ...(timeoutMs === undefined ? {} : { timeoutMs }), ...(permissionsSnapshot === undefined ? {} : { permissions: permissionsSnapshot }), ...(localAgentStream === undefined ? {} : { surfaces: { streams: [{ stream: localAgentStream }] } }), ...(options.workspace === undefined ? {} : { surfaces: { workspace: [{ surface: options.workspace }] } }), @@ -264,6 +284,9 @@ export function authoredWorkerRunner( if (typeof output !== 'object' || output === null || Array.isArray(output)) { throw new AuthoredFlowExecutionError('journal_protocol_violation', `step "${id}" produced a non-object output`); } + if (AGENT_TIMEOUT in output && timeoutMs !== undefined) { + return { completionReason: 'timeout', summary: String(output[AGENT_TIMEOUT]), artifacts: [] }; + } const stdout = 'stdout_tail' in output ? output.stdout_tail : undefined; // The worker that ran the CLI measured the artifacts and journaled them // in the step's output; read that fact back rather than re-scanning a @@ -271,7 +294,7 @@ export function authoredWorkerRunner( const journaled = 'artifacts' in output ? output.artifacts : undefined; const artifacts = Array.isArray(journaled) && journaled.every(entry => typeof entry === 'string') ? [...journaled] : []; - return { summary: typeof stdout === 'string' ? stdout : JSON.stringify(output), artifacts }; + return { completionReason: 'success', summary: typeof stdout === 'string' ? stdout : JSON.stringify(output), artifacts }; }, async llm(id: string, prompt: string, options?: LlmOptions, verification?: NamedGate): Promise { if (options !== undefined && (typeof options !== 'object' || options === null || options.output === undefined)) { diff --git a/packages/sdk/src/compile.ts b/packages/sdk/src/compile.ts index cca74c527..6d84cd7b2 100644 --- a/packages/sdk/src/compile.ts +++ b/packages/sdk/src/compile.ts @@ -74,7 +74,12 @@ export class CompileError extends Error { * survives float-precision (`1.1 * 1000 = 1100.0000000000002`) rather than * being rejected by `isSafeInteger`. */ -export function parseStepTimeout(timeout: unknown): number { +export const AGENT_STEP_TIMEOUT_MAX_MS = 60 * 60_000; +export function parseAgentStepTimeout(timeout: unknown): number { + return parseStepTimeout(timeout, true); +} + +export function parseStepTimeout(timeout: unknown, agent = false): number { const unitToMs: Record<'ms' | 's' | 'm', number> = { ms: 1, s: 1000, m: 60_000 }; let milliseconds: number; if (typeof timeout === 'number') { @@ -96,10 +101,10 @@ export function parseStepTimeout(timeout: unknown): number { milliseconds = NaN; } if (!Number.isSafeInteger(milliseconds) || milliseconds <= 0) { - throw new CompileError(['f.run timeout must be a positive whole number of milliseconds or a duration such as "10s" or "5m".'], 'timeout_invalid'); + throw new CompileError([`${agent ? 'f.agent' : 'f.run'} timeout must be a positive whole number of milliseconds or a duration such as "10s" or "5m".`], 'timeout_invalid'); } - if (milliseconds > 15 * 60_000) { - throw new CompileError(['f.run timeout exceeds the maximum of 15 minutes declared in SURFACE.md.'], 'lease_exceeded'); + if (milliseconds > (agent ? AGENT_STEP_TIMEOUT_MAX_MS : 15 * 60_000)) { + throw new CompileError([`${agent ? 'f.agent' : 'f.run'} timeout exceeds the maximum of ${agent ? 60 : 15} minutes declared in SURFACE.md.`], 'lease_exceeded'); } return milliseconds; } @@ -211,7 +216,7 @@ function compileStep(step: StepSpec): StepSpec { type: 'deterministic', command: s.command, verification: verification ?? { type: 'exit_code' as const }, - // #138: `timeoutMs` is deterministic-only — worker-backed verbs own + // Timeout fields stay per-verb — worker-backed verbs own // their dispatch timeout. It must be spread HERE and nowhere in `base`. ...(s.timeoutMs !== undefined ? { timeoutMs: s.timeoutMs } : {}), ...(s.lease_ms !== undefined ? { lease_ms: parseStepTimeout(s.lease_ms) } : {}), @@ -242,6 +247,7 @@ function compileStep(step: StepSpec): StepSpec { ...base, type: 'agent', instruction: s.instruction, + ...(s.timeoutMs === undefined ? {} : { timeoutMs: s.timeoutMs }), ...(verification !== undefined ? { verification: outputGate(verification, s.id) } : {}), ...(s.agent !== undefined ? { agent: s.agent } : {}), ...(s.cli !== undefined ? { cli: s.cli } : {}), @@ -473,7 +479,7 @@ function kernelStepToAuthoring(value: unknown, at: string): unknown { : type === 'llm' ? ['prompt', 'model', 'cli', 'cli_identity'] as const : type === 'agent' - ? ['instruction', 'cli', 'model', 'recovery_mode', 'surfaces', 'permissions', 'cwd', 'transport', 'cli_identity'] as const + ? ['timeout_ms', 'instruction', 'cli', 'model', 'recovery_mode', 'surfaces', 'permissions', 'cwd', 'transport', 'cli_identity'] as const : []; assertKernelKeys(step, [...commonKeys, ...typeKeys], at); const transportRetries = step['retry'] === undefined @@ -508,6 +514,7 @@ function kernelStepToAuthoring(value: unknown, at: string): unknown { return { ...common, instruction: step['instruction'], + ...(step['timeout_ms'] === undefined ? {} : { timeoutMs: step['timeout_ms'] }), ...(step['recovery_mode'] !== undefined ? { recoveryMode: step['recovery_mode'] } : {}), ...copyDefined(step, ['cli', 'model', 'surfaces', 'cwd', 'transport']), ...(step['permissions'] !== undefined @@ -680,6 +687,7 @@ function toKernelStep(step: StepSpec, cliIdentity?: string): ResolvedKernelStepS ...common, type: 'agent', instruction: step.instruction, + ...(step.timeoutMs === undefined ? {} : { timeout_ms: step.timeoutMs }), ...(step.cli !== undefined ? { cli: step.cli } : {}), ...(step.model !== undefined ? { model: step.model } : {}), ...(step.cwd !== undefined ? { cwd: step.cwd } : {}), diff --git a/packages/sdk/src/spec.ts b/packages/sdk/src/spec.ts index 6cbf2d06a..51a0db928 100644 --- a/packages/sdk/src/spec.ts +++ b/packages/sdk/src/spec.ts @@ -290,6 +290,7 @@ export interface LlmStepSpec extends BaseStepSpec { */ export interface AgentStepSpec extends BaseStepSpec { type: 'agent'; + timeoutMs?: number; instruction: string; verification?: OutputVerificationSpec; /** Named authoring declaration selected from `FlowSpec.agents`. Compiled away. */ @@ -513,6 +514,7 @@ export interface KernelPermissionsSpec { export interface KernelAgentStep extends KernelStepCommon { type: 'agent'; + timeout_ms?: number; instruction: string; cli?: string; model?: string; diff --git a/packages/sdk/src/step-fields.ts b/packages/sdk/src/step-fields.ts index 64239a8d6..79a1f3f79 100644 --- a/packages/sdk/src/step-fields.ts +++ b/packages/sdk/src/step-fields.ts @@ -38,5 +38,5 @@ export const STEP_COMMON_FIELDS = [ export const STEP_FIELDS_BY_TYPE = { deterministic: ['command', 'timeoutMs', 'lease_ms', 'onNonZero'], llm: ['prompt', 'model', 'cli', 'output'], - agent: ['instruction', 'agent', 'cli', 'model', 'cwd', 'transport', 'surfaces', 'recoveryMode', 'permissions', 'output'], + agent: ['timeoutMs', 'instruction', 'agent', 'cli', 'model', 'cwd', 'transport', 'surfaces', 'recoveryMode', 'permissions', 'output'], } as const satisfies Record; diff --git a/packages/sdk/src/validate.ts b/packages/sdk/src/validate.ts index 8be7c7ade..c68de0232 100644 --- a/packages/sdk/src/validate.ts +++ b/packages/sdk/src/validate.ts @@ -490,6 +490,13 @@ class Validator { } private validateAgent(st: AgentStepSpec, at: string): void { + if (st.timeoutMs !== undefined) { + if (!isPosInt(st.timeoutMs) || st.timeoutMs > 60 * 60_000) { + this.fail(`${at}.timeoutMs: expected a positive integer at most 3600000 (60 minutes)`); + } + if (st.transport === 'relay') this.fail(`${at}.timeoutMs: unsupported with relay transport`); + if ((st.maxIterations ?? 1) > 1) this.fail(`${at}.timeoutMs: requires maxIterations 1`); + } if (!isNonEmptyString(st.instruction)) { this.fail(`${at}.instruction: expected a non-empty string`); } diff --git a/packages/sdk/src/worker-cli.ts b/packages/sdk/src/worker-cli.ts index bd0eeb99e..1f29e778c 100644 --- a/packages/sdk/src/worker-cli.ts +++ b/packages/sdk/src/worker-cli.ts @@ -121,6 +121,8 @@ export async function runAgentCli( relayContext?: AgentRelayContext, processEnvironment: NodeJS.ProcessEnv = process.env, cliIdentity?: string, + // Trailing parameter preserves existing positional callers. + stepTimeoutMs?: number, ): Promise { signal?.throwIfAborted(); if (signal !== undefined && process.platform === 'win32') { @@ -167,17 +169,21 @@ export async function runAgentCli( // discovery names are set from this dispatch, exactly as for a direct spawn. const wrapperEnv = wrapperEnvironment(processEnvironment); applyStepEnvironment(wrapperEnv, sidechannel); - return requirePricedUsage(decodeWrapperResult(await runWrapperSession( + const session = await runWrapperSession( cli, instruction, wakeContext, effectiveModel, wrapperEnv, - wrapperLimits, + stepTimeoutMs === undefined ? wrapperLimits : { ...wrapperLimits, executionTimeoutMs: stepTimeoutMs }, signal, cwd, argv0, - )), effectiveModel); + ); + const decoded = decodeWrapperResult(session); + // A wrapper that hit its deadline after reporting usage is still charged + // for it; what it wrote is not the step's output. + return requirePricedUsage(session.transport?.cause === 'timeout' ? { ...decoded, stdout_tail: '' } : decoded, effectiveModel); } const env: NodeJS.ProcessEnv = { ...processEnvironment }; @@ -188,6 +194,8 @@ export async function runAgentCli( applyStepEnvironment(env, sidechannel); const invocation = mode === 'llm' ? llmExecution(kind, instruction, effectiveModel) : agentExecution(kind, instruction, effectiveModel); + if (stepTimeoutMs !== undefined) invocation.timeoutMs = stepTimeoutMs; + if (wakeContext !== undefined) { try { env[WAKE_CONTEXT_ENV] = JSON.stringify(wakeContext); diff --git a/packages/sdk/src/worker-spend.ts b/packages/sdk/src/worker-spend.ts index a3c9a9fd5..fb9b42628 100644 --- a/packages/sdk/src/worker-spend.ts +++ b/packages/sdk/src/worker-spend.ts @@ -35,7 +35,10 @@ function unmeteredUsage(input: number | undefined, output: number | undefined): */ export function workerSpend(result: WorkerCliResult, model?: string): { result: WorkerCliResult; usage: StepUsage | undefined } { try { - const usage = pricedUsage(model, result.tokens_input, result.tokens_output) + // Usage that was never reported is unknown, not a measured $0, even for a + // priced model (a step stopped before its CLI reported, for one). + const reported = result.tokens_input !== undefined && result.tokens_output !== undefined; + const usage = (reported ? pricedUsage(model, result.tokens_input, result.tokens_output) : undefined) ?? unmeteredUsage(result.tokens_input, result.tokens_output); return { result, usage }; } diff --git a/packages/sdk/src/worker.ts b/packages/sdk/src/worker.ts index 9b631cace..4afb28317 100644 --- a/packages/sdk/src/worker.ts +++ b/packages/sdk/src/worker.ts @@ -156,7 +156,7 @@ export class AgentWorker extends EventEmitter { spec.transport === 'relay' ? 'relay' : 'direct', { runId: dispatch.run_id, stepId: dispatch.step_id, idempotencyKey: dispatch.idempotency_key, dataDir: this.options.dataDir, resultSchema: spec.verification?.json_schema }, - this.options.environment, spec.cli_identity) + this.options.environment, spec.cli_identity, spec.timeout_ms) : Promise.resolve({ exit_code: null, stdout_tail: '', stderr_tail: 'agent step has no declared CLI' })); const { result, usage } = workerSpend(completed, effectiveModel); const cost = reportedCost(completed, effectiveModel); diff --git a/packages/sdk/src/wrapper-session.ts b/packages/sdk/src/wrapper-session.ts index e94fe989c..35805b39f 100644 --- a/packages/sdk/src/wrapper-session.ts +++ b/packages/sdk/src/wrapper-session.ts @@ -1,3 +1,4 @@ +import { transportEvidence, type CliTransportEvidence } from './cli-transport-evidence.js'; import { spawn, type ChildProcessWithoutNullStreams } from 'node:child_process'; import { childStop, ownsProcessGroup } from './child-stop.js'; import { reapOnExit } from './agent-reaper.js'; @@ -29,6 +30,7 @@ export interface WrapperSessionLimits { } export interface WrapperSessionResult { + transport?: CliTransportEvidence; exit_code: number | null; stdout_tail: string; stderr_tail: string; @@ -157,7 +159,7 @@ async function executePinnedWrapper( let exitCode: number | null = null; /** The handshake deadline, then the execution deadline if there is one. */ let lifecycleTimer: NodeJS.Timeout | undefined; - /** The post-exit drain grace, armed only when execution is unlimited. */ + /** The post-exit drain grace, armed once the wrapper has exited and acknowledged. */ let drainTimer: NodeJS.Timeout | undefined; const finish = (result: WrapperSessionResult): void => { @@ -231,7 +233,7 @@ async function executePinnedWrapper( }; signal?.addEventListener('abort', onAbort, { once: true }); if (signal?.aborted) { onAbort(); return; } - const terminate = (message: string): void => { + const terminate = (message: string, transport?: WrapperSessionResult['transport'], captured = ''): void => { if (protocolError !== undefined) return; protocolError = message; if (lifecycleTimer !== undefined) clearTimeout(lifecycleTimer); @@ -247,7 +249,7 @@ async function executePinnedWrapper( // The shared stop owns both the liveness bound and settlement. Its // callback preserves this refusal once group death is proved, or // replaces it with a fail-closed confirmation error when it is not. - finishAfterStop(failure(message), 'terminate'); + finishAfterStop({ ...failure(message), stdout_tail: captured, ...(transport === undefined ? {} : { transport }) }, 'terminate'); }; const startExecutionTimer = (): void => { if (lifecycleTimer !== undefined) clearTimeout(lifecycleTimer); @@ -258,9 +260,29 @@ async function executePinnedWrapper( // and what bounds the work from here is the lease abort — the same thing // that bounds a `claude` or `codex` step. if (limits.executionTimeoutMs <= 0) return; - lifecycleTimer = setTimeout(() => terminate( + lifecycleTimer = setTimeout(() => { + lifecycleTimer = undefined; + // The deadline bounds a wrapper that is still running. One that has + // already exited did not time out, whatever it left holding stdio: + // its settle belongs to the post-exit drain, which reports the + // wrapper's own exit. Re-arming here is idempotent. + if (childExited) { + startDrainGrace(); + return; + } + onExecutionDeadline(); + }, limits.executionTimeoutMs); + }; + const onExecutionDeadline = (): void => { + terminate( `CLI ${JSON.stringify(cli)} execution timed out after ${limits.executionTimeoutMs}ms.`, - ), limits.executionTimeoutMs); + transportEvidence({ phase: 'timeout', cause: 'timeout', retryable: false, exitCode: null, signal: null, + stderr: `CLI execution timed out after ${limits.executionTimeoutMs}ms.` }, env), + // What the wrapper wrote before its deadline, so a result envelope it + // already emitted still reports its usage. Only the caller decides + // what of it, if anything, is output. + stdout.join('') + executionPending, + ); }; const exceedsOutputLimit = (additionalBytes: number): boolean => { const pendingBytes = Buffer.byteLength(executionPending); @@ -382,13 +404,13 @@ async function executePinnedWrapper( finishOnChildExit(result); }; /** - * Arm the post-exit drain, for an unlimited session that has both seen the - * wrapper exit and consumed its acknowledgement. A session with an explicit - * execution deadline keeps its existing behaviour: that deadline is what - * bounds a withheld `'close'` there, and nothing about it changes. + * Arm the post-exit drain, for a session that has both seen the wrapper + * exit and consumed its acknowledgement. This holds with or without an + * execution deadline: the deadline bounds a wrapper still running, and a + * wrapper that exited — even with a descendant holding stdio — settles as + * its own exit here rather than waiting out the deadline as a timeout. */ const startDrainGrace = (): void => { - if (limits.executionTimeoutMs > 0) return; if (!childExited || phase !== 'execute') return; if (settled || drainTimer !== undefined || protocolError !== undefined) return; drainTimer = setTimeout(drainExpired, DRAIN_AFTER_EXIT_MS); diff --git a/packages/sdk/tests/agent-timeout-live.test.ts b/packages/sdk/tests/agent-timeout-live.test.ts new file mode 100644 index 000000000..83045a40f --- /dev/null +++ b/packages/sdk/tests/agent-timeout-live.test.ts @@ -0,0 +1,106 @@ +import { existsSync, readFileSync, readdirSync, writeFileSync } from 'node:fs'; +import { join, resolve } from 'node:path'; +import { afterEach, expect, it } from 'vitest'; +import { AUTHORED_STEP_STREAM } from '../src/authored-step-index.js'; +import { JournalClient } from '../src/journal-client.js'; +import { socketPathFor } from '../src/daemon-connection.js'; +import { chainFixture } from './flow-chain-fixture.js'; + +const cleanups: Array<() => Promise> = []; +afterEach(async () => { for (const close of cleanups.splice(0).reverse()) await close(); }); + +function fixture(pause: boolean, gate = true) { + const f = chainFixture(); + cleanups.push(() => f.close()); + writeFileSync(f.wrapper, `#!/usr/bin/env node +import { receiveWrapperRequest } from ${JSON.stringify(resolve('../../testdata/preflight/wrapper-session.mjs'))}; +import { appendFileSync, writeFileSync } from 'node:fs'; +if (process.argv[2] === 'auth') process.exit(0); +const request = await receiveWrapperRequest(); +if (request?.instruction.includes('TIMEOUT_TEST')) { + appendFileSync(${JSON.stringify(f.calls)}, 'executed\\n'); + writeFileSync(${JSON.stringify(join(f.root, 'work.txt'))}, 'keep this work'); + writeFileSync(${JSON.stringify(join(f.root, 'agent.pid'))}, String(process.pid)); + // A complete, priced result envelope, then a process that never exits. + process.stdout.write(JSON.stringify({ protocol: 'relayflows-agent-cli-v1-result', output: 'partial', + usage: { input_tokens: 100, output_tokens: 20 } })); + setInterval(() => {}, 1000); +} else if (request) process.stdout.write('probe ok'); +`); + writeFileSync(join(f.root, 'flows.json'), JSON.stringify({ cli: f.wrapper, models: ['claude-opus-5'] })); + writeFileSync(f.flowPath, `import { flow } from '@relayflows/surface'; +export default flow('timeout-test', { budget: { wallclock: '2m' } }, async f => { + const result = await f.agent('repair', { task: 'TIMEOUT_TEST', timeout: '300ms', model: 'claude-opus-5' }) + .gate(r => r.completionReason === 'timeout' && ${gate}); + if (result.completionReason === 'timeout') { + await f.run(${JSON.stringify(pause ? 'test -f paused || { touch paused; sleep 120; }; cat work.txt > published' : 'cat work.txt > published')}); + } + f.done('success'); +}); +`); + return f; +} + +async function evidence(f: ReturnType) { + const root = readdirSync(join(f.data, 'runs')).filter(x => x.endsWith('.sqlite3')).sort()[0]!.slice(0, -8); + const client = new JournalClient(socketPathFor(f.data)); + await client.connect(); await client.hello('agent-timeout-test'); + try { + const records = (await client.streamRead(root, AUTHORED_STEP_STREAM, 0, 1000)).messages + .map((m: any) => m.message ?? m) as Array>; + const child = records.find(r => r.state === 'completed' && r.completionReason === 'timeout')!; + expect(child).toBeDefined(); + const entries = (await client.journalRead(child.runId as string, 1, 1000)).entries as any[]; + expect(entries.filter(e => e.entry_type === 'step.completed').map(e => e.payload.completionReason)).toEqual(['timeout']); + // The usage the wrapper reported before its deadline is charged, priced: + // claude-opus-5 at 100 * $5/M + 20 * $25/M = $0.001000. + expect(entries.find(e => e.entry_type === 'step.completed').payload.budget) + .toEqual({ tokens_in: 100, tokens_out: 20, dollars: '0.001000' }); + expect(entries.find(e => e.entry_type === 'run.spawned').payload.spec.steps[0].timeout_ms).toBe(300); + expect((await client.runGet(child.runId as string)).status).toBe('failed'); + } finally { client.close(); } + return root; +} + +it('journals timeout, stops the process, runs a predicate gate and publishes under a budget header', async () => { + const f = fixture(false); + const run = f.invoke('run', '--local-agent', '--no-observer-link', '--json', '--data-dir', f.data, f.flowPath, '--input', '{}'); + expect(run.status, run.stdout + run.stderr).toBe(0); + expect(JSON.parse(run.stdout).completionReason).toBe('success'); + expect(readFileSync(join(f.root, 'published'), 'utf8')).toBe('keep this work'); + expect(readFileSync(f.calls, 'utf8')).toBe('executed\n'); + expect(() => process.kill(Number(readFileSync(join(f.root, 'agent.pid'), 'utf8')), 0)).toThrow(); + await evidence(f); +}, 60000); + +it('replays timeout after killing the root and daemon, without executing the agent again', async () => { + const f = fixture(true); + const run = f.invokeAsync('run', '--local-agent', '--no-observer-link', '--json', '--data-dir', f.data, f.flowPath, '--input', '{}'); + let output = ''; + run.stdout?.on('data', chunk => { output += chunk; }); + run.stderr?.on('data', chunk => { output += chunk; }); + const deadline = Date.now() + 30000; + while (!existsSync(join(f.root, 'paused'))) { + if (Date.now() > deadline || run.exitCode !== null) throw new Error(`did not reach publish: ${output}`); + await new Promise(resolve => setTimeout(resolve, 25)); + } + const root = await evidence(f); + const exited = new Promise(resolve => run.once('exit', () => resolve())); + run.kill('SIGKILL'); await exited; + const { pid } = JSON.parse(readFileSync(join(f.data, 'connection.json'), 'utf8')); + process.kill(pid, 'SIGKILL'); + const resumed = f.invoke('resume', '--local-agent', '--no-observer-link', '--json', '--data-dir', f.data, root); + expect(resumed.status, resumed.stdout + resumed.stderr).toBe(0); + expect(JSON.parse(resumed.stdout).completionReason).toBe('success'); + expect(readFileSync(f.calls, 'utf8')).toBe('executed\n'); + expect(readFileSync(join(f.root, 'published'), 'utf8')).toBe('keep this work'); + await evidence(f); +}, 90000); + +it('lets an author reject timeout through a predicate gate', async () => { + const f = fixture(false, false); + const run = f.invoke('run', '--local-agent', '--no-observer-link', '--json', '--data-dir', f.data, f.flowPath, '--input', '{}'); + expect(run.status, run.stdout + run.stderr).not.toBe(0); + expect(existsSync(join(f.root, 'published'))).toBe(false); + expect(readFileSync(f.calls, 'utf8')).toBe('executed\n'); +}, 60000); diff --git a/packages/sdk/tests/agent-timeout-outcome.test.ts b/packages/sdk/tests/agent-timeout-outcome.test.ts new file mode 100644 index 000000000..b0a62fca5 --- /dev/null +++ b/packages/sdk/tests/agent-timeout-outcome.test.ts @@ -0,0 +1,53 @@ +import { describe, expect, it, vi } from 'vitest'; +import { authoredWorkerRunner } from '../src/authored-worker-step.js'; +import { JournalClient } from '../src/journal-client.js'; +import { compileSpec } from '../src/compile.js'; + +const mocks = vi.hoisted(() => ({ classify: vi.fn() })); +vi.mock('../src/cli/run.js', () => ({ classifyOutcome: mocks.classify })); +vi.mock('../src/authored-preflight.js', () => ({ + authoredPreflight: () => async (spec: unknown) => ({ report: { ok: true }, flow: compileSpec(spec) }), +})); + +function setup(reason = 'timeout') { + const journal = new JournalClient('/unused'); + vi.spyOn(journal, 'runStart').mockResolvedValue({ run_id: 'child', status: 'failed', completed_steps: 0 }); + const append = vi.spyOn(journal, 'streamAppend').mockResolvedValue({ offset: 0 } as never); + mocks.classify.mockResolvedValue({ exitCode: 1, report: { status: 'failed', completionReason: 'step_failed', + diagnostics: [{ stepId: 'a', stepType: 'agent', completionReason: reason, stderrTail: 'deadline' }] } }); + const runner = authoredWorkerRunner({ name: 'test' }, journal, '/flow.ts', [], {}, + undefined, undefined, undefined, 'root'); + return { runner, append, journal }; +} + +describe('recoverable agent timeout boundaries', () => { + it('refuses an undeclared timeout and other failures', async () => { + await expect(setup().runner.agent('a', { task: 'repair' })).rejects.toMatchObject({ code: 'step_failed' }); + await expect(setup('worker_error').runner.agent('a', { task: 'repair', timeout: '1s' })) + .rejects.toMatchObject({ code: 'step_failed' }); + }); + it('does not bypass a lowered named gate that never ran', async () => { + await expect(setup().runner.agent('a', { task: 'repair', timeout: '1s' }, + { type: 'artifact_exists', path: 'work.txt' })).rejects.toMatchObject({ code: 'step_failed' }); + }); + it('fails closed if the completed timeout cannot be indexed', async () => { + const { runner, append } = setup(); + append.mockResolvedValueOnce({ offset: 0 } as never).mockRejectedValueOnce(new Error('disk full')); + await expect(runner.agent('a', { task: 'repair', timeout: '1s' })).rejects.toThrow('disk full'); + }); + it('lowers a duration and resolves the journaled timeout exactly once', async () => { + const { runner, append, journal } = setup(); + await expect(runner.agent('a', { task: 'repair', timeout: '45m' })) + .resolves.toEqual({ summary: 'deadline', artifacts: [], completionReason: 'timeout' }); + expect(journal.runStart).toHaveBeenCalledWith(expect.objectContaining({ + steps: [expect.objectContaining({ timeout_ms: 2700000 })], + }), undefined, expect.any(String)); + expect(append).toHaveBeenCalledTimes(2); + }); + it.each([{ transport: 'relay' as const }, { maxIterations: 2 }])('refuses %j before admission', async options => { + const { runner, journal } = setup(); + await expect(runner.agent('a', { task: 'repair', timeout: '1s', ...options })) + .rejects.toMatchObject({ code: 'agent_cli_unresolved' }); + expect(journal.runStart).not.toHaveBeenCalled(); + }); +}); diff --git a/packages/sdk/tests/agent-timeout-worker.test.ts b/packages/sdk/tests/agent-timeout-worker.test.ts new file mode 100644 index 000000000..5258eeafc --- /dev/null +++ b/packages/sdk/tests/agent-timeout-worker.test.ts @@ -0,0 +1,86 @@ +import { chmodSync, mkdtempSync, readFileSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { afterEach, expect, it } from 'vitest'; +import { runAgentCli } from '../src/worker-cli.js'; +import { agentCompletionReason } from '../src/cli-transport-evidence.js'; +import { workerSpend } from '../src/worker-spend.js'; + +const roots: string[] = []; +afterEach(() => { for (const root of roots.splice(0)) rmSync(root, { recursive: true, force: true }); }); +function cli(name: string, source: string) { + const root = mkdtempSync(join(tmpdir(), 'agent-timeout-worker-')); roots.push(root); + const path = join(root, name); + writeFileSync(path, `#!/usr/bin/env node\n${source}`); chmodSync(path, 0o755); + return { root, path }; +} + +it.each(['claude', 'codex', 'wrapper'])('stops %s at its declared limit, preserving work and stopping descendants', async name => { + const source = ` +const { writeFileSync } = require('node:fs'); +const { spawn } = require('node:child_process'); +function work() { + writeFileSync('work.txt', 'preserved'); + writeFileSync('pid', String(process.pid)); + spawn(process.execPath, ['-e', "require('node:fs').writeFileSync('descendant', String(process.pid)); setInterval(() => {}, 1000)"], { stdio: 'inherit' }); + setInterval(() => {}, 1000); +} +${name === 'wrapper' ? `process.stdout.write('relayflows-agent-cli-v1\\n'); +process.stdin.resume(); process.stdin.on('end', () => { + process.stdout.write('relayflows-agent-cli-v1-execute\\n'); work(); +});` : 'process.stdin.resume(); work();'} +`; + const f = cli(name, source); + const result = await runAgentCli(f.path, 'repair', undefined, undefined, undefined, + new AbortController().signal, 'agent', undefined, f.root, 'direct', undefined, process.env, undefined, 500); + expect(agentCompletionReason(result)).toBe('timeout'); + expect(result.transport).toMatchObject({ phase: 'timeout', cause: 'timeout', retryable: false }); + expect(readFileSync(join(f.root, 'work.txt'), 'utf8')).toBe('preserved'); + for (const name of ['pid', 'descendant']) { + const pid = Number(readFileSync(join(f.root, name), 'utf8')); + // Linux may retain a dead descendant as a zombie until its init reaps it. + try { process.kill(pid, 0); expect(readFileSync(`/proc/${pid}/stat`, 'utf8')).toMatch(/\) Z /); } + catch (error) { expect((error as NodeJS.ErrnoException).code).toBe('ESRCH'); } + } +}, 15000); + +it('keeps timeout evidence when a priced wrapper has no usage', async () => { + const f = cli('wrapper', `process.stdout.write('relayflows-agent-cli-v1\\n'); +process.stdin.resume(); process.stdin.on('end', () => { +process.stdout.write('relayflows-agent-cli-v1-execute\\n'); setInterval(() => {}, 1000); +});`); + const result = await runAgentCli(f.path, 'repair', undefined, 'claude-sonnet-4-6', undefined, + new AbortController().signal, 'agent', undefined, f.root, 'direct', undefined, process.env, undefined, 50); + expect(agentCompletionReason(result)).toBe('timeout'); + expect(result.stderr_tail).toMatch(/usage/i); + // Never measured, so never journaled as a measured $0. + expect(workerSpend(result, 'claude-sonnet-4-6').usage).toEqual({ tokens_in: 0, tokens_out: 0, dollars_unmetered: true }); +}); + +it('keeps the usage a wrapper reported before it hung past its deadline', async () => { + const envelope = JSON.stringify({ protocol: 'relayflows-agent-cli-v1-result', output: 'partial', usage: { input_tokens: 100, output_tokens: 20 } }); + const f = cli('wrapper', `process.stdout.write('relayflows-agent-cli-v1\\n'); +process.stdin.resume(); process.stdin.on('end', () => { +process.stdout.write('relayflows-agent-cli-v1-execute\\n' + ${JSON.stringify(envelope)}); setInterval(() => {}, 1000); +});`); + const result = await runAgentCli(f.path, 'repair', undefined, 'claude-opus-5', undefined, + new AbortController().signal, 'agent', undefined, f.root, 'direct', undefined, process.env, undefined, 200); + expect(agentCompletionReason(result)).toBe('timeout'); + expect(result).toMatchObject({ tokens_input: 100, tokens_output: 20, stdout_tail: '' }); + // claude-opus-5: 100 * $5/M + 20 * $25/M = $0.001000. + expect(workerSpend(result, 'claude-opus-5').usage).toEqual({ tokens_in: 100, tokens_out: 20, dollars: '0.001000' }); +}); + +it.each([ + ['handshake timeout', 'setInterval(() => {}, 1000);'], + ['wrong identity', "process.stdout.write('not-a-wrapper\\n'); setInterval(() => {}, 1000);"], + ['duplicate execute', "process.stdout.write('relayflows-agent-cli-v1\\n'); process.stdin.resume(); process.stdin.on('end', () => process.stdout.write('relayflows-agent-cli-v1-execute\\nrelayflows-agent-cli-v1-execute\\n'));"], + ['output limit', "process.stdout.write('relayflows-agent-cli-v1\\n'); process.stdin.resume(); process.stdin.on('end', () => process.stdout.write('relayflows-agent-cli-v1-execute\\n' + 'x'.repeat(1000)));"], +])('keeps %s a worker error despite a declared execution timeout', async (_, source) => { + const f = cli('wrapper', source); + const result = await runAgentCli(f.path, 'repair', undefined, undefined, + { handshakeTimeoutMs: 200, maxOutputBytes: 128 }, new AbortController().signal, + 'agent', undefined, f.root, 'direct', undefined, process.env, undefined, 1000); + expect(agentCompletionReason(result)).toBe('worker_error'); + expect(result.transport).toBeUndefined(); +}); diff --git a/packages/sdk/tests/agent-timeout.test.ts b/packages/sdk/tests/agent-timeout.test.ts new file mode 100644 index 000000000..f2b2bf473 --- /dev/null +++ b/packages/sdk/tests/agent-timeout.test.ts @@ -0,0 +1,36 @@ +import { readFileSync } from 'node:fs'; +import { flow } from '@relayflows/surface'; +import { describe, expect, it } from 'vitest'; +import { executeAuthoredFlow } from '../src/authored-flow-executor.js'; +import { compileSpec, kernelToAuthoring, parseAgentStepTimeout, toKernelSpec } from '../src/compile.js'; +import { JournalClient } from '../src/journal-client.js'; +import { validateSpec } from '../src/validate.js'; + +const spec = (fields = {}) => ({ version: '0.1.0', steps: [{ id: 'a', type: 'agent', instruction: 'repair', ...fields }] }); +describe('agent timeout declaration', () => { + it.each([['45m', 2700000], ['90s', 90000], [2700000, 2700000], ['1.5s', 1500], ['60m', 3600000]])( + 'parses %s and preserves the kernel round trip', (value, ms) => { + expect(parseAgentStepTimeout(value)).toBe(ms); + const compiled = compileSpec(spec({ timeoutMs: ms })); + const kernel = toKernelSpec(compiled); + expect(kernel.steps[0]).toHaveProperty('timeout_ms', ms); + expect(kernelToAuthoring(kernel)).toEqual(compiled); + }); + it.each([0, -1, NaN, Infinity, '5', 'forever', '61m', 3600001])('refuses %s before admission', async timeout => { + expect(() => parseAgentStepTimeout(timeout)).toThrow(); + await expect(executeAuthoredFlow(flow('invalid', async f => { + await f.agent('repair', { task: 'repair', timeout }); + f.done('success'); + }), new JournalClient('/never-connect'))).rejects.toMatchObject({ code: 'agent_cli_unresolved' }); + }); + it('keeps omission absent', () => { + expect(toKernelSpec(compileSpec(spec())).steps[0]).not.toHaveProperty('timeout_ms'); + }); + const cases = JSON.parse(readFileSync(new URL('../../../testdata/agent-timeout-cases.json', import.meta.url), 'utf8')); + it.each(cases)('$name', ({ timeoutMs, valid }: { timeoutMs: unknown; valid: boolean }) => { + expect(validateSpec(spec({ timeoutMs })).ok).toBe(valid); + }); + it.each([{ transport: 'relay' }, { maxIterations: 2 }])('refuses an unenforceable bound: %j', fields => { + expect(validateSpec(spec({ timeoutMs: 1000, ...fields })).ok).toBe(false); + }); +}); diff --git a/packages/sdk/tests/authored-node-result.test.ts b/packages/sdk/tests/authored-node-result.test.ts index 9f5345f53..c74c806da 100644 --- a/packages/sdk/tests/authored-node-result.test.ts +++ b/packages/sdk/tests/authored-node-result.test.ts @@ -37,6 +37,23 @@ beforeEach(()=>{ })); }); describe('authored IPC result durable verification',()=>{ + it('accepts only a declared, durably settled agent timeout', async () => { + const claimed = result(); + claimed.journalSteps[0]!.completionReason = 'timeout'; + const step = { id: 'run-1', type: 'agent', timeout_ms: 1000 }; + records.set('child-1', [ + { entry_type: 'run.spawned', payload: { spec: { name: 'example/run-1', steps: [step] } } }, + { entry_type: 'step.completed', step_id: 'run-1', payload: { completionReason: 'timeout', disposition: 'step_done' } }, + { entry_type: 'run.completed', payload: { completionReason: 'step_failed' } }, + ]); + mocks.runGet.mockImplementation(async (run_id: string) => ({ run_id, + status: run_id === 'child-1' ? 'failed' : 'completed', + steps: { [run_id === 'child-1' ? 'run-1' : 'complete-2']: { state: 'done' } }, + })); + await expect(verifyAuthoredNodeResult(claimed, metadata, 'root', 'socket')).resolves.toBeUndefined(); + delete (step as { timeout_ms?: number }).timeout_ms; + await expect(verifyAuthoredNodeResult(claimed, metadata, 'root', 'socket')).rejects.toThrow('no matching durable completion'); + }); it('accepts a predicate-gated flow: the `.gate` child is verified but does not consume an ordinal',async()=>{ const claimed:AuthoredFlowExecutionResult={rootRunId:'root',name:'example',completionReason:'success',journalSteps:[ {id:'run-1',runId:'child-1',completionReason:'success'}, diff --git a/packages/sdk/tests/spec-parity.test.ts b/packages/sdk/tests/spec-parity.test.ts index c3f407b07..00a4779fe 100644 --- a/packages/sdk/tests/spec-parity.test.ts +++ b/packages/sdk/tests/spec-parity.test.ts @@ -27,7 +27,7 @@ function fixture(name: string): string { } describe('spec parity: one dialect at the SDK<->kernel boundary', () => { - for (const name of ['hello-deterministic', 'hello-ladder', 'hello-llm', 'hello-agent', 'step-memory', 'step-placement', 'agent-cwd']) { + for (const name of ['hello-deterministic', 'hello-ladder', 'hello-llm', 'hello-agent', 'step-memory', 'step-placement', 'agent-cwd', 'agent-timeout']) { it(`compiles ${name} to the pinned canonical JSON`, () => { const yaml = fixture(`${name}.flow.yaml`); const canonical = compileYamlToCanonicalJson(yaml); diff --git a/packages/sdk/tests/verb-field-lint.test.ts b/packages/sdk/tests/verb-field-lint.test.ts index 53b4f00a8..88bb79a25 100644 --- a/packages/sdk/tests/verb-field-lint.test.ts +++ b/packages/sdk/tests/verb-field-lint.test.ts @@ -183,7 +183,7 @@ describe('closed per-verb step fields', () => { 'workspace', 'tools', ]); expect(AGENT_DECLARATION_FIELDS).toEqual(['cli', 'model']); - // `timeoutMs` is deliberately absent: it is a deterministic-only authoring + // `timeoutMs` is deliberately absent: it is a per-verb authoring // field, not a common one. Pinned so the move cannot be silently undone. expect(STEP_COMMON_FIELDS).toEqual([ 'id', @@ -205,14 +205,13 @@ describe('closed per-verb step fields', () => { expect(STEP_FIELDS_BY_TYPE).toEqual({ deterministic: ['command', 'timeoutMs', 'lease_ms', 'onNonZero'], llm: ['prompt', 'model', 'cli', 'output'], - agent: ['instruction', 'agent', 'cli', 'model', 'cwd', 'transport', 'surfaces', 'recoveryMode', 'permissions', 'output'], + agent: ['timeoutMs', 'instruction', 'agent', 'cli', 'model', 'cwd', 'transport', 'surfaces', 'recoveryMode', 'permissions', 'output'], }); expect(CROSS_VERB_STEP_FIELDS.map(({ label }) => label).sort()).toEqual([ 'agent foreign command', 'agent foreign lease_ms', 'agent foreign onNonZero', 'agent foreign prompt', - 'agent foreign timeoutMs', 'deterministic foreign agent', 'deterministic foreign cli', 'deterministic foreign cwd', diff --git a/packages/sdk/tests/worker-cli.test.ts b/packages/sdk/tests/worker-cli.test.ts index 7c2c3fefc..51f0be836 100644 --- a/packages/sdk/tests/worker-cli.test.ts +++ b/packages/sdk/tests/worker-cli.test.ts @@ -14,6 +14,7 @@ import { join } from 'node:path'; import { afterEach, describe, expect, it } from 'vitest'; import type { JournalClient } from '../src/journal-client.js'; import type { Pins } from '../src/protocol.js'; +import { agentCompletionReason } from '../src/cli-transport-evidence.js'; import { AgentWorker } from '../src/worker.js'; import { cliInvocationArgv0, runAgentCli } from '../src/worker-cli.js'; @@ -450,6 +451,7 @@ process.stdin.on('end', () => { expect(result.exit_code).toBeNull(); expect(result.stderr_tail).toMatch(/identity changed/i); + expect(agentCompletionReason(result)).toBe('worker_error'); expect(realpathSync(declared)).toBe(realpathSync(replacement)); expect(existsSync(requestEvidence)).toBe(false); expect(existsSync(replacementEvidence)).toBe(false); @@ -531,6 +533,10 @@ describe('custom wrapper execution bounds are reader-owned', () => { * there the wrapper is still alive when the deadline fires, so SIGTERM * closes its own pipes and 'close' arrives. That test proves the timer * FIRES. This one proves the bound HOLDS. + * + * The wrapper exited 0 before its deadline, so it did not time out: the + * post-exit drain settles it as its own exit even with a deadline armed, + * and the descendant holding the pipe does not turn it into a timeout. */ function leakyWrapperSource(inherit: 'inherit' | ['ignore', 'inherit', 'ignore'], holdMs: number): string { return ` @@ -562,8 +568,9 @@ process.stdin.on('end', () => { }); const elapsed = Date.now() - started; - expect(result.exit_code).toBeNull(); - expect(result.stderr_tail).toMatch(/timed out after 300ms/i); + expect(result.transport?.cause).not.toBe('timeout'); + expect(result.exit_code).toBe(0); + expect(result.stdout_tail).toBe('{"ok":true}\n'); expect(elapsed).toBeLessThan(5_000); }, 20_000); @@ -583,8 +590,9 @@ process.stdin.on('end', () => { }); const elapsed = Date.now() - started; - expect(result.exit_code).toBeNull(); - expect(result.stderr_tail).toMatch(/timed out after 300ms/i); + expect(result.transport?.cause).not.toBe('timeout'); + expect(result.exit_code).toBe(0); + expect(result.stdout_tail).toBe('{"ok":true}\n'); expect(elapsed).toBeLessThan(5_000); }, 20_000); diff --git a/packages/sdk/tests/wrapper-exit-drain.test.ts b/packages/sdk/tests/wrapper-exit-drain.test.ts index a111367be..333a442fa 100644 --- a/packages/sdk/tests/wrapper-exit-drain.test.ts +++ b/packages/sdk/tests/wrapper-exit-drain.test.ts @@ -2,7 +2,7 @@ import { chmodSync, existsSync, mkdtempSync, rmSync, writeFileSync } from 'node: import { tmpdir } from 'node:os'; import { join } from 'node:path'; import { afterEach, expect, it, vi } from 'vitest'; -import { runAgentCli } from '../src/worker-cli.js'; +import { agentCompletionReason, runAgentCli } from '../src/worker-cli.js'; /** * What the 300 s execution deadline was really doing for a session that had @@ -54,9 +54,9 @@ function makeDirectory(): string { */ function pipeHolder( stdio: 'inherit' | readonly string[], - { deaf = false, detached = true }: { deaf?: boolean; detached?: boolean } = {}, + { deaf = false, detached = true, holdMs = HOLD_MS }: { deaf?: boolean; detached?: boolean; holdMs?: number } = {}, ): string { - const script = `${deaf ? 'process.on(\'SIGTERM\', () => {});' : ''}setTimeout(() => {}, ${HOLD_MS});`; + const script = `${deaf ? 'process.on(\'SIGTERM\', () => {});' : ''}setTimeout(() => {}, ${holdMs});`; return ` const { spawn } = require('node:child_process'); spawn(process.execPath, ['-e', ${JSON.stringify(script)}], { @@ -105,6 +105,37 @@ process.exit(0); * `['ignore', 'inherit', 'ignore']` — that is stdout — so this path had no * coverage under its own name. */ +/** + * An agent `timeout` arms an execution deadline on every wrapper session. That + * deadline bounds a wrapper still RUNNING; it says nothing about one that has + * already exited. A clean exit whose descendant keeps the pipe open must still + * settle from the drain as the wrapper's own exit, not wait out the deadline + * and be journaled as a timeout. + */ +it('settles an exited wrapper from the drain even when an execution deadline is armed', async () => { + const directory = makeDirectory(); + const wrapper = makeWrapper(directory, 'drain-with-deadline', ` +process.stdout.write('{"ok":"drained-before-deadline"}\\n'); +`, `${pipeHolder('inherit', { holdMs: 15_000 })} +process.exit(0); +`); + const stepTimeoutMs = 8_000; + + const started = Date.now(); + const result = await runAgentCli( + wrapper, 'instruction', undefined, undefined, undefined, undefined, 'agent', undefined, directory, + undefined, undefined, undefined, undefined, stepTimeoutMs, + ); + const elapsed = Date.now() - started; + + expect(result.transport?.cause).not.toBe('timeout'); + expect(agentCompletionReason(result)).not.toBe('timeout'); + expect(result.stderr_tail).toBe(''); + expect(result.exit_code).toBe(0); + expect(result.stdout_tail).toBe('{"ok":"drained-before-deadline"}\n'); + expect(elapsed).toBeLessThan(SETTLE_BOUND_MS); +}, 30_000); + it('settles a wrapper that exits leaving only stderr held open', async () => { const directory = makeDirectory(); const wrapper = makeWrapper(directory, 'drain-stderr', ` diff --git a/packages/sdk/type-tests/step-fields.ts b/packages/sdk/type-tests/step-fields.ts index f3588bcf6..9088e8d95 100644 --- a/packages/sdk/type-tests/step-fields.ts +++ b/packages/sdk/type-tests/step-fields.ts @@ -33,7 +33,6 @@ const agent = { id: 'agent', type: 'agent', instruction: 'act', - // @ts-expect-error timeoutMs is a deterministic-step-only authoring field. timeoutMs: 1_000, } satisfies AgentStepSpec; diff --git a/packages/surface/src/context.ts b/packages/surface/src/context.ts index 7c9489157..d5d392615 100644 --- a/packages/surface/src/context.ts +++ b/packages/surface/src/context.ts @@ -7,6 +7,7 @@ import type { Activity, ActivityOptions } from "./activity.js"; import type { TriggerSource } from "./triggers.js"; export interface AgentResult { + completionReason: 'success' | 'timeout'; summary: string; /** * Files the agent created or changed under its working directory, @@ -28,6 +29,9 @@ export interface PermissionsSpec { } export interface AgentOptions { + /** CLI time limit: milliseconds or ms/s/m duration, at most 60m. + * Resolves with completionReason: timeout. No default; direct transport only, maxIterations 1. */ + timeout?: string | number; task: string; workspace?: string; /** Validated declaration only; not currently enforced (gate 8 / #442). */ diff --git a/scripts/schema-constraints.mjs b/scripts/schema-constraints.mjs index c8383908e..307af3bb1 100644 --- a/scripts/schema-constraints.mjs +++ b/scripts/schema-constraints.mjs @@ -31,6 +31,7 @@ export function applyConstraints(defs, version) { } property('KernelRetryPolicy', 'max_transport_retries', u32); property('DeterministicStepSpec', 'timeoutMs', positive); + property('AgentStepSpec', 'timeoutMs', { ...positive, maximum: 3600000 }); for (const field of ['maxTokensIn', 'maxTokensOut']) property('BudgetSpec', field, integer); property('BudgetSpec', 'maxDollars', decimal); property('MemorySpec', 'query', { pattern: '\\S' }); diff --git a/testdata/agent-timeout-cases.json b/testdata/agent-timeout-cases.json new file mode 100644 index 000000000..32b9524ee --- /dev/null +++ b/testdata/agent-timeout-cases.json @@ -0,0 +1,9 @@ +[ + {"name":"one millisecond", "timeoutMs":1, "valid":true, "kernelValid":true}, + {"name":"60 minute ceiling", "timeoutMs":3600000, "valid":true, "kernelValid":true}, + {"name":"over authoring ceiling", "timeoutMs":3600001, "valid":false, "kernelValid":true}, + {"name":"zero", "timeoutMs":0, "valid":false, "kernelValid":false}, + {"name":"negative", "timeoutMs":-1, "valid":false, "kernelValid":false}, + {"name":"fractional", "timeoutMs":1.5, "valid":false, "kernelValid":false}, + {"name":"YAML duration strings are not milliseconds", "timeoutMs":"45m", "valid":false, "kernelValid":false} +] diff --git a/testdata/agent-timeout.flow.yaml b/testdata/agent-timeout.flow.yaml new file mode 100644 index 000000000..20dc9326e --- /dev/null +++ b/testdata/agent-timeout.flow.yaml @@ -0,0 +1,11 @@ +version: '0.1.0' +name: agent-timeout +steps: + - id: repair + type: agent + instruction: Repair the workspace. + timeoutMs: 2700000 + - id: review + type: agent + dependsOn: [repair] + instruction: Review the workspace without a deadline. diff --git a/testdata/agent-timeout.spec.canonical.json b/testdata/agent-timeout.spec.canonical.json new file mode 100644 index 000000000..5f1e8e9e6 --- /dev/null +++ b/testdata/agent-timeout.spec.canonical.json @@ -0,0 +1 @@ +{"name":"agent-timeout","steps":[{"depends_on":[],"id":"repair","instruction":"Repair the workspace.","max_iterations":1,"recovery_mode":"reset","retry":{"initial_backoff_ms":100,"jitter_percent":20,"max_backoff_ms":60000,"multiplier":2},"timeout_ms":2700000,"type":"agent","verification":{}},{"depends_on":["repair"],"id":"review","instruction":"Review the workspace without a deadline.","max_iterations":1,"recovery_mode":"reset","retry":{"initial_backoff_ms":100,"jitter_percent":20,"max_backoff_ms":60000,"multiplier":2},"type":"agent","verification":{}}],"version":"0.1.0"} diff --git a/testdata/agent-timeout.spec.sha256 b/testdata/agent-timeout.spec.sha256 new file mode 100644 index 000000000..ca4ffa275 --- /dev/null +++ b/testdata/agent-timeout.spec.sha256 @@ -0,0 +1 @@ +9046a1076346f394d60439da9798758eb334a696e62d405a33eed4f1ba8cc9ab