Skip to content

fix(flows): stop long agent steps at their FLOW_TIME allowance (#138) - #139

Merged
AgentRelayBot merged 2 commits into
mainfrom
arc-138-agent-timeouts
Oct 4, 2026
Merged

AgentRelayBot merged 2 commits into
mainfrom
arc-138-agent-timeouts

Conversation

@AgentRelayBot

@AgentRelayBot AgentRelayBot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #138

What

FLOW_TIME decided whether a long agent step could start, but nothing stopped it at its allowance (a check-repair agent ran 38m / $54). The generated cloud flow now passes each allowance as a hard f.agent timeout (relayflows 2.0.40, AgentWorkforce/flows#606), and handles completionReason === "timeout" explicitly. A timeout is never treated as success.

Agent step Limit On timeout
check-repair-N repairMinutes (45m) Counts as a failed repair attempt. It re-checks, then stops repairing (a second repair of the same failures would most likely run out too). If the checks still fail, the existing base-check, draft and report path runs.
adversary-N reviewMinutes (20m) Treated as an unresolved review. review.clean is not read, even if the stopped reviewer left one. Same path as an unresolved review: fix round (traditional, round 1), or the PR goes to draft and the report is posted. That report now says the review was stopped at its limit (review_timeout=yes).
fixer fixerMinutes (45m) Its work is kept. The flow checks, pushes and re-reviews it as usual.
check-discovery new discoveryMinutes (15m; measured 6m and 9m24s) Discards any half-written .relayflow/check.sh and falls back to the ecosystem default.

Each branch prints a clear console.error line.

Left unbounded, with the reason in a comment: planner, plan-reviewer, implementer, prototypes and comparator. They are the mandatory path that produces the change. A hard stop there leaves nothing worth publishing, and no measured run overran on them. The header's wallclock still bounds them.

Budget arithmetic

Each allowance is ≤ agentLimitMaxMinutes (60, the runtime ceiling). The repairStartMinutes / reviewStartMinutes / fixRoundStartMinutes arithmetic is unchanged, because it already assumed each step ran its full allowance. Now the runtime enforces that. The budget sim now models the runtime: an agent that would run past its timeout is charged exactly the limit and resolves with completionReason: "timeout". The edge-of-every-guard sweep runs repair, review and fixer at 10× their allowances. It asserts that each one timed out and was charged exactly its limit, and that publishing still fits (more than 100 timed-out steps checked).

Local target (item 4)

Only the cloud target emits timeout. The local kit pins RELAYFLOWS_VERSION 2.0.26, and that version's AgentOptions has no timeout: a cloud source fails tsc there with TS2353 'timeout' does not exist in type 'AgentOptions'. The local flow keeps the same timedOut() branches. They are typed against unknown, so they compile against 2.0.26's AgentResult, which has no completionReason, and they never fire. A test fails once the pin reaches 2.0.40, as a reminder to turn the limits on locally.

Checked by hand: every generated cloud source typechecks against @relayflows/surface@2.0.40, and every local source typechecks against @relayflows/surface@2.0.26 (tsc --strict, exit 0).

Tests (red → green)

  • flow-workflows.test.ts:
    • Per workflow, each f.agent call carries exactly its FLOW_TIME limit (or none).
    • Each timeout branch is present.
    • The local source states no limits.
    • FLOW_REVIEW_BLOCKED_COMMAND with review_timeout=yes gives the timeout note in the posted report.
  • flow-budget.test.ts: behaviour tests for repair (fails re-check, and passes re-check), adversary (with a stale review.clean), first-review timeout leading to a fix round, fixer, discovery and the local target, plus the extended sweep.
  • flow-onboarding.test.ts: updated for the review_timeout=no; prefix.

Before the change, 18 of the new tests failed. Now they all pass, and tsc --noEmit is clean. Full web npx vitest run locally: 378 of 379 pass. The one failure is flow-agent-settings.test.ts › gives every generated preset agent an explicit supported CLI/model pair ("Test timed out in 5000ms" under full-suite load). It fails the same way on clean main and passes when run alone.

Do not redeploy: the orchestrator rolls this onto the live Gardens after merge.

🤖 Generated with Claude Code


Note

Medium Risk
Changes production cloud flow orchestration for repairs, reviews, and check discovery; behavior is well-tested but mis-handling timeouts could draft PRs or skip repairs incorrectly.

Overview
Cloud-generated factory flows now pass hard f.agent timeouts (relayflows 2.0.40) for optional long steps—check-discovery, check-repair, adversary reviews, and the fixer—using new FLOW_TIME allowances including 15m discovery and a 60m runtime ceiling. Steps that hit the limit resolve with completionReason: "timeout" via a shared timedOut() helper; the flow never treats that as success.

Repair: a timed-out repair is one failed attempt (re-check, no second repair), then the existing draft/report path if checks still fail. Review: timeouts are always unresolved (no review.clean, stale review.md cleared between rounds); blocked PR text can include review_timeout=yes. Discovery: drops a partial .relayflow/check.sh and falls back to the ecosystem default. Fixer: keeps partial work and continues check/push/re-review. Planner, implementer, and similar mandatory agents stay unbounded on cloud; local generated flows omit timeout until the kit pins relayflows ≥ 2.0.40.

Budget simulation and workflow tests were extended to charge agents at their limits and assert each timeout branch.

Reviewed by Cursor Bugbot for commit 2d639df. Bugbot is set up for automated code reviews on this repo. Configure here.


Summary by cubic

Stops long agent steps at their FLOW_TIME allowance as hard f.agent timeouts (relayflows 2.0.40), so a step that once started only when its allowance fit but then ran on past it (a check-repair agent ran 38m / $54) is now stopped there. A timeout is handled explicitly per step and never treated as success.

  • A timed-out check-repair counts as a failed repair attempt: its committed work is re-checked, no further repair is tried, and remaining failures take the existing draft-and-report path.
  • A timed-out adversary review is unresolved and never clean, even if it left a review.clean. Each reviewer now starts from a cleared review.md, so a stopped one can never pass off a previous round's findings; the posted report says the review was stopped at its limit.
  • A timed-out fixer's work is kept and checked, pushed, and re-reviewed as usual.
  • A timed-out check-discovery discards any half-written check.sh and falls back to the ecosystem default.
  • Planner, plan-reviewer, implementer, prototypes, and comparator stay unbounded: a hard stop there would leave nothing worth publishing.

Only the cloud target states limits; the local kit pins RELAYFLOWS_VERSION 2.0.26, which has no agent timeout, so its flow keeps the same branches that never fire. The budget sim now charges a step that would run past its limit exactly the limit and resolves it with completionReason: "timeout", and a sweep asserts such steps still leave room to publish. Tests also flag any non-literal agent timeout so the no-limit checks cannot pass over one.

Written for commit 2d639df. Summary will update on new commits.

Review in cubic

Pass the FLOW_TIME allowances as hard f.agent timeouts (relayflows
2.0.40, AgentWorkforce/flows#606) on check-repair, the adversary reviews,
the fixer and check-discovery, and handle completionReason "timeout"
explicitly on each: a repair is a failed attempt (re-check, no further
repair), a review is unresolved (never clean), a fixer's work is kept
and checked, and a discovery falls back to the ecosystem default.

Only the cloud target states limits; the local kit's pinned 2.0.26
refuses the option. The budget sweep now runs the long agents past
their limits and charges each exactly its limit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 27242e9e-7fb5-448f-a698-8fa8e9644610
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 93067f83-c6f3-4b3d-add7-c8c4302ee720
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 4 files

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread web/lib/flow-workflows.ts
Comment thread web/lib/test/flow-workflows.test.ts Outdated
@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployed!

Environment URL
Web https://c2f91951-agentrelay-web.agent-workforce.workers.dev

This is a Cloudflare Workers preview version of this PR's build.

…nt timeouts in tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 27242e9e-7fb5-448f-a698-8fa8e9644610
@AgentRelayBot
AgentRelayBot merged commit 9652513 into main Oct 4, 2026
5 checks passed
@AgentRelayBot
AgentRelayBot deleted the arc-138-agent-timeouts branch October 4, 2026 09:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Garden agents need a hard stop: pass FLOW_TIME allowances as f.agent timeouts (relayflows 2.0.40)

1 participant