Skip to content

relayflowd run.get times out under CPU load and kills the run: a flow's own test suite can destroy 30 steps of journaled work - #613

Draft
agent-relay-code[bot] wants to merge 1 commit into
mainfrom
relayflow/flows-software-garden-9c0d555d
Draft

agent-relay-code[bot] wants to merge 1 commit into
mainfrom
relayflow/flows-software-garden-9c0d555d

Conversation

@agent-relay-code

@agent-relay-code agent-relay-code Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

A flow's CPU-heavy step could turn a delayed run.get into a terminal root failure. This change reduces read pressure, retries read timeouts, and leaves the root resumable when the daemon still cannot answer.

  • Worker completion waits use journal watch pushes, with snapshots every two seconds for live lease deadlines. Scoped watch sessions close on completion, cancellation, or failure; no new daemon verb is required.
  • Opted-in clients serialize read-only requests and retry timeouts with fresh IDs, escalating bounds, jitter, and a 300-second total budget including queue time. An unconditional lazy reader session avoids blocking behind deterministic commands. Writes and interactive clients retain their existing behavior.
  • Read interruptions skip root terminalization. Both run and resume reports retain the root ID, status: running, and the resume command. A fresh probe distinguishes daemon_unresponsive from daemon_unreachable, without claiming that a failed probe proves the daemon is dead.
  • Heartbeat timeouts become lease loss, so the kernel owns recovery. The standalone CLI's Node child preserves read interruptions across its error frame.

Retry opt-ins: runDirectFlow, declarative executeCheckedFlow, resumeFlow, and authored-node-entry. Worker peers and utility clients keep single-shot bounded requests.

The retry policy, completion wait, and diagnostic classifier live in separate small modules. The existing large transport file retains the typed verb wrappers; cli/run.ts shrinks. No Rust execution logic, journal format, workflow files, or gate definitions change.

Design notes: watches alone cannot unblock a snapshot or child-journal read queued behind a different child's deterministic command, so the unconditional reader remains necessary. The protocol has no run.unwatch; closing each scoped watch connection releases its server watcher. readChildJournal normally reads once per completed child; adopted unfinished children can reread after settling, and retry-backoff inspection can reread. This change does not add a journal cursor cache or daemon snapshot cache.

Validation details and remaining failures are recorded below. “Parked” here means execution is left resumable without a fabricated terminal journal fact; it is not a journaled human park and does not assert journal integrity without reading it.

Validation:

  • Captured commands and literal output: 20 focused files / 210 tests passed, plus the separately added real-daemon resume test. That test suspends the daemon during an actual child read, observes daemon_unresponsive without worker_error terminalization, then resumes the root and checks saved\nafter\n exactly once. The CPU-load case saturates the daemon's CPU allocation (one affinity-pinned core on Linux) while reads remain in flight.
  • Typecheck command and captured result: npx tsc --noEmit and npx tsc -p tsconfig.tests.json, exit 0 with no output.
  • Four mutations were reverted individually, tested failing, restored byte-for-byte, and tested passing: reader routing, two-second cadence, heartbeat lease classification, and root parking. Both outputs and matching source SHA256 values are pasted in the verification transcript. Reproduction script.

The full suite is not green. The final production build was tested with the explicit daemon binary after building the local Surface package; the additional resume test above was added afterward and run separately. Complete full-suite output:

$ cd packages/sdk && RELAYFLOWD_BIN=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd npm test
 Test Files  6 failed | 248 passed | 1 skipped (255)
      Tests  77 failed | 3774 passed | 22 skipped (3873)
     Errors  1 error

Failing files, copied from that output:

 ❯ tests/live-kernel.test.ts (32 tests | 8 failed) 57015ms
 ❯ tests/hosted-extension-isolation.test.ts (22 tests | 14 failed) 350ms
 ❯ tests/authored-node-runtime.test.ts (16 tests | 16 skipped) 12ms
 ❯ tests/hosted-extension-protocol.test.ts (27 tests | 8 failed) 40086ms
 ❯ tests/canonical-software-factory.test.ts (25 tests | 25 failed) 152ms
 ❯ tests/stuck-run-triage.test.ts (22 tests | 22 failed) 74ms

These include unavailable bubblewrap, the standalone Bun prerequisite, hosted-extension timeouts, Surface flow-identity refusals, and live wrapper/analyzer failures. They remain unresolved here; standalone Bun integration is consequently unverified. One live stub's handshake failure also reproduces using unmodified SDK commit 3c58ee16d10a9e2400db980f5bbafac84e437f20 against the same fixture path: current probe/output, current output, baseline probe, baseline output. This baseline comparison covers that specific failure, not all six suites.

Checks

The checks fail on the base commit too, so these failures were not introduced by this change: they come from the repository itself or from the environment the checks ran in. This pull request is a draft until someone looks.

What ran (.relayflow/check.sh)
#!/bin/sh
set -e
# Fresh-machine equivalent of this repo's CI gates, in CI's own order.
#
# Mirrors, in this order:
#   .github/workflows/cloud-runtime-artifact.yml  (kernel + SDK, the main gate)
#   .github/workflows/schema-publish.yml          (its pull_request `validate` job)
#   .github/workflows/surface-package.yml         (scripts/surface-package-gate.sh)
#   .github/workflows/review-swarm-wrapper-guard.yml (the two parity checks only)
#
# Deliberately NOT run here, because this machine cannot:
#   - publish.yml's publish/pages jobs: NPM_TOKEN / OIDC / GitHub Pages
#     deployment. Its one secret-free check, scripts/publish.test.mjs, is below.
#   - review-swarm.yml and swarm-wrapper-guard.sh: need GH_TOKEN plus a real
#     pull_request_target event, and the swarm needs model credentials.
#   - actions/upload-artifact steps: no artifact store, and nothing checks them.
#   - tests/live-kernel.test.ts's live-analyzer case: needs a `claude` binary and
#     model access. RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 below is exactly what CI
#     sets, and means this run is not gate-2 acceptance evidence.
#   - packages/ts-plugin's `npm test`: no CI workflow runs it today.

ROOT=$(CDPATH= cd -- "$(dirname -- "$0")/.." && pwd -P)
cd "$ROOT"

for tool in node npm bun; do
  command -v "$tool" >/dev/null 2>&1 || {
    echo "check: $tool is required (CI provisions node 22 and bun 1.4.0)" >&2
    exit 127
  }
done

# ---------------------------------------------------------------- toolchains

# CI gets cargo from dtolnay/rust-toolchain: a default toolchain in the standard
# rustup home, driven by plain `cargo`. Reproduce that rather than ops/cargo.sh —
# the wrapper redirects RUSTUP_HOME to a private tree, which leaves a rustup shim
# on PATH with no toolchain to choose ("rustup could not choose a version of
# cargo to run"). That is why the kernel CI steps call cargo directly too.
if ! command -v cargo >/dev/null 2>&1; then
  [ -x "$HOME/.cargo/bin/cargo" ] || curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs \
    | sh -s -- -y --default-toolchain stable --profile minimal --no-modify-path
  PATH="$HOME/.cargo/bin:$PATH"
  export PATH
fi
cargo --version

# Hosted-extension sandbox tests exercise a real bubblewrap namespace command.
# Best-effort: the cases that need it are `it.skipIf(!existsSync('/usr/bin/bwrap'))`,
# so a machine without it still gets a meaningful (smaller) run, and a package
# install is not worth failing the whole check over.
if [ ! -x /usr/bin/bwrap ] && command -v sudo >/dev/null 2>&1 && command -v apt-get >/dev/null 2>&1; then
  sudo apt-get update \
    && sudo apt-get install --yes --no-install-recommends bubblewrap \
    || echo "check: bubblewrap unavailable; sandbox cases will skip themselves" >&2
fi
if [ -x /usr/bin/bwrap ] && sysctl kernel.apparmor_restrict_unprivileged_userns >/dev/null 2>&1; then
  # Ubuntu 24.04 blocks unprivileged user namespaces through AppArmor by default.
  sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0 || true
fi

# ------------------------------------------------------------------ installs

# --ignore-scripts throughout: each package's build is an explicit step below,
# so nothing is skipped, and npm's prepare cannot run before its devDependencies
# exist.
( cd packages/surface && bun install --frozen-lockfile --ignore-scripts && bun run build )
npm ci --prefix packages/sdk --ignore-scripts

# `npm ci` pulls @relayflows/surface from the REGISTRY, not the copy just built.
# Any SDK code importing a symbol that has not shipped yet then fails typecheck.
# `./` is load-bearing: without it npm reads the path as a GitHub shorthand.
npm install ./packages/surface --prefix packages/sdk --no-save --ignore-scripts

# workflows/*.flow.ts import @relayflows/surface from the repo root, where a
# fresh checkout has no node_modules at all; workflow-driving tests cannot load
# their subject without this link.
mkdir -p node_modules/@relayflows
ln -sfn ../../packages/sdk/node_modules/@relayflows/surface node_modules/@relayflows/surface
node -e "console.log(require.resolve('@relayflows/surface'))"

# ------------------------------------------------------- builds and codegen

( cd kernel && cargo build --locked --release -p relayflowd )

# Re-assert the executable bit on the preflight CLI fixtures (test:prep's other
# half). They are committed 100755, but an EACCES deep inside a preflight test is
# opaque enough to be worth one line.
[ ! -d testdata/preflight ] || find testdata/preflight -name '*-cli' -type f -exec chmod +x {} +

# packages/sdk/dist is required: several test files fail at collection without
# it, so a bare `vitest run` is not enough.
npm run typecheck --prefix packages/sdk
npm run build --prefix packages/sdk
npm run typecheck:tests --prefix packages/sdk

# --------------------------------------------------------------------- tests

node --test scripts/cloud-artifact.test.mjs
node --test scripts/publish.test.mjs

( cd kernel && cargo test --workspace )

# The whole SDK suite, not named files: naming files is how ~22 of ~26 test files
# went uncovered. RELAYFLOWD_BIN points at the release binary built above instead
# of compiling a second debug copy through ops/cargo.sh.
( cd packages/sdk \
  && RELAYFLOWS_ALLOW_ANALYZER_SKIP=1 \
     RELAYFLOWD_BIN="$ROOT/kernel/target/release/relayflowd" \
     ./node_modules/.bin/vitest run )

# Committed schema must be regenerable and the generator deterministic. This
# rewrites packages/schema/flows.schema.json in place, exactly as CI does.
node scripts/generate-json-schema.mjs
git diff --exit-code -- packages/schema/flows.schema.json
cp packages/schema/flows.schema.json /tmp/flows.schema.first.json
node scripts/generate-json-schema.mjs
diff -q /tmp/flows.schema.first.json packages/schema/flows.schema.json
( cd packages/schema && bun run test )

# Lens parity between the pre-swarm runner and the cloud swarm spec: pure file
# comparisons, no GitHub context needed. PRESWARM_ALLOW_MISSING_CLI=1 is what the
# guard workflow sets, because no runner has every lens CLI installed.
sh ops/preswarm-check/lens-parity-check.sh
PRESWARM_ALLOW_MISSING_CLI=1 sh ops/preswarm-check/lens-cli-parity-check.sh

# Last on purpose: this gate repacks the surface and installs the TARBALL over
# packages/sdk/node_modules/@relayflows/surface, so running it earlier would
# change the tree the SDK suite above is meant to test. Needs bash (set -o pipefail).
bash scripts/surface-package-gate.sh

# Standalone CLI and cloud artifact, with the verifier and the exact-archive
# smoke. CI uploads the archive afterwards; that upload is the omitted part.
mkdir -p dist/cloud-artifact-input
node scripts/build-standalone-cli.mjs bun-linux-x64 dist/cloud-artifact-input/flows
node scripts/cloud-artifact.mjs build \
  --relayflowd kernel/target/release/relayflowd \
  --flows-executable dist/cloud-artifact-input/flows \
  --output-dir dist/cloud-artifact \
  --source-commit "$(git rev-parse HEAD)"
archive="$(find dist/cloud-artifact -name '*.tar.gz' -type f -print -quit)"
node scripts/cloud-artifact.mjs verify \
  --archive "$archive" \
  --sha256 "$(awk '{print $1}' "$archive.sha256")"
rm -rf dist/cloud-artifact-smoke
mkdir -p dist/cloud-artifact-smoke
tar -xzf "$archive" -C dist/cloud-artifact-smoke
dist/cloud-artifact-smoke/bin/relayflowd --help
dist/cloud-artifact-smoke/bin/flows check --json testdata/hello-deterministic.flow.yaml
Output on this branch (last 80 lines)
 FAIL  tests/live-kernel.test.ts > a relayflow can be scheduled: tick source against live relayflowd > a tick spawns a real run whose step reports the SCHEDULED instant
AssertionError: expected null to deeply equal { schedule_id: 'heartbeat-1m', …(3) }

- Expected: 
Object {
  "lag_ms": 43000,
  "schedule_id": "heartbeat-1m",
  "scheduled_for_ms": 1764000000000,
  "slot": 29400000,
}

+ Received: 
null

 ❯ tests/live-kernel.test.ts:1739:39
    1737|     // The bound: the run reports the grid instant and its own lag, so…
    1738|     // backfilled run can tell it is running for a slot from the past.
    1739|     expect(completed!.payload.output).toEqual({
       |                                       ^
    1740|       schedule_id: 'heartbeat-1m',
    1741|       slot: 29_400_000,

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[22/32]⎯

 FAIL  tests/software-garden-babysitter-composition.test.ts > canonical Software Garden + Babysitter composition > propagates capability denial without a retry or fallback
AssertionError: expected Error: Hosted extension sandbox exited wi… { code: '…' } to be Error: live babysit label is absent // Object.is equality

- Expected
+ Received

- [Error: live babysit label is absent]
+ [Error: Hosted extension sandbox exited without a valid completion (exit 1): bwrap: loopback: Failed RTM_NEWADDR: Operation not permitted
+ ]

 ❯ tests/software-garden-babysitter-composition.test.ts:139:7
    137|       const refusal = new Error('live babysit label is absent');
    138|       let calls = 0;
    139|       await expect(runHostedSoftwareGardenBabysitter({
       |       ^
    140|         flowPath: installed.flowPath,
    141|         dispatch: dispatch(),

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[23/32]⎯

⎯⎯⎯⎯⎯⎯ Unhandled Errors ⎯⎯⎯⎯⎯⎯

Vitest caught 1 unhandled error during the test run.
This might cause false positive tests. Resolve unhandled errors to make sure your tests are not affected.

⎯⎯⎯⎯ Unhandled Rejection ⎯⎯⎯⎯⎯
Error: Hosted extension sandbox exited without a valid completion (exit 1): bwrap: loopback: Failed RTM_NEWADDR: Operation not permitted

 ❯ refuse src/hosted-extension-protocol.ts:135:21
    133|       : new PluginError('plugin_unsupported', 'Hosted capability rejec…
    134|     const refuse = (message: string) => {
    135|       const error = new PluginError('plugin_unsupported', message);
       |                     ^
    136|       CHILD_PROCESS_KILL(child, 'SIGKILL');
    137|       if (capabilityState === 'pending') {
 ❯ ChildProcess.<anonymous> src/hosted-extension-protocol.ts:234:21
 ❯ ChildProcess.emit node:events:520:22
 ❯ maybeClose node:internal/child_process:1084:16
 ❯ Socket.<anonymous> node:internal/child_process:456:11
 ❯ Socket.emit node:events:508:20
 ❯ Pipe.<anonymous> node:net:347:12

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯
Serialized Error: { code: 'plugin_unsupported' }
This error originated in "tests/hosted-extension-protocol.test.ts" test file. It doesn't mean the error was thrown inside the file itself, but while it was running.
The latest test that might've caused the error is "rejects two forged calls after the authoritative first outcome settles". It might mean one of the following:
- The error was thrown, while Vitest was running this test.
- If the error occurred after the test had been completed, this was the last documented test before it was thrown.
⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯

 Test Files  6 failed | 249 passed | 1 skipped (256)
      Tests  31 failed | 3823 passed | 20 skipped (3874)
     Errors  1 error
   Start at  11:01:00
   Duration  614.66s (transform 4.47s, setup 1.15s, collect 71.72s, tests 1636.42s, environment 33ms, prepare 10.79s)

Output on the base commit (last 80 lines)
cargo 1.99.0 (5f94df478 2026-08-27)
sysctl: permission denied on key "kernel.apparmor_restrict_unprivileged_userns"
/tmp/relayflow-recipe.TWt90j: 66: cd: can't cd to packages/surface

Fixes #522


Summary by cubic

Prevents a CPU-heavy step from destroying journaled work when a delayed run.get times out under load. Flow-execution runs now retry read-only requests within a bounded budget and, if the daemon still cannot answer, stay resumable instead of being terminalized.

Bug Fixes

  • Completion waits use run.watch pushes with lease-deadline snapshots every two seconds; scoped watch sessions close on completion, cancellation, or failure.
  • Opt-in clients (runDirectFlow, declarative executeCheckedFlow, resumeFlow, and authored-node-entry) serialize read-only requests and retry with fresh IDs, escalating bounds, jitter, and a 300-second total budget; writes and interactive clients keep their existing single-shot behavior.
  • Read interruptions skip root terminalization: both run and resume reports keep the root ID, status: running, and the resume command, and a fresh probe distinguishes daemon_unresponsive from daemon_unreachable.
  • Heartbeat timeouts now count as lease loss, so the kernel owns recovery; the standalone CLI's Node child preserves read interruptions across its error frame.

Validation

  • Adds a live test that suspends the daemon during a real child read, observes daemon_unresponsive without worker_error terminalization, resumes the root, and checks the journaled effect runs exactly once.
  • The full suite is not green; the 77 failures (bubblewrap, standalone Bun, hosted-extension timeouts, Surface flow-identity refusals) reproduce on the base commit and are unrelated to this change.

Written for commit a50a8de. Summary will update on new commits.

Review in cubic

@coderabbitai

coderabbitai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Important

Review skipped

Bot user detected.

To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: 4ff99450-e026-4eef-a2db-cfdd363232f9

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@agent-relay-code

Copy link
Copy Markdown
Contributor Author

Relayflow: the adversarial review did not pass. This branch is not approved: the flow stopped here and did not mark it ready to merge.

Review of PR #613

Changes requested. Reviewed head a50a8dee0074b71c867298abf8dd1ec5f641d3d6 against parent 3c58ee16d10a9e2400db980f5bbafac84e437f20. No production code or existing tests changed by this review. review.clean is deliberately absent.

Findings

P1 — A disconnected read session still terminalizes the authored root

Location: packages/sdk/src/journal-client.ts:214–219, with propagation through journal-read-policy.ts:64 and authored-root.ts:372–402.

The new dedicated reader can disconnect independently while the primary and root worker sessions remain healthy. requestOnce then rejects with an ordinary connection error; the retry policy retries only JournalRequestTimeoutError, and the root's new parking branch recognizes only timeout/unresponsive errors. Consequently a transient read transport failure reaches terminalizeRootFailure, which records terminal worker_error and exhausts the root's retries. The cached reader also remains unusable after disconnect.

The reproduction uses a real loopback socket for reads and the existing authored-root fixture for journal writes: the first run.get times out, the second closes its socket. The observed result is journal client: connection closed and a root completion with reason: worker_error. This is a transport interruption, not a body failure, and still violates the requirement to preserve unreadable runs for resume.

Carry a typed read-transport interruption across the SDK and Node bridge; reconnect retryable reads within their remaining budget, and route exhausted transport interruptions through the resumable report. Keep protocol refusals and write failures distinct. Add coverage for a reader disconnect after a timeout while the root worker connection remains available.

P2 — Cancellation cannot interrupt the new long-running snapshot wait

Location: packages/sdk/src/cli/running-step.ts:66–72.

Once the two-second poll starts await client.runGet(runId), cancellation only rejects the completed promise, which this await no longer observes. A read can consume the full new 300-second budget after cancellation or root lease loss before cleanup closes the watch. If the read eventually returns a completed snapshot, this function can even return normally despite the abort. The added cancellation test covers only pending watch registration, so it misses this branch.

The reproduction starts the snapshot read, aborts the signal, advances another 100 ms, and observes that the wait is still unfinished. Make read cancellation settle and drain the underlying request promptly, preserving authored promise-scope accounting, and propagate the abort before interpreting the snapshot. Add cancellation coverage after the first snapshot has begun.

PR comments and scope

Read the complete PR diff, source call paths, added/changed tests, and the submitted verification artifacts. The PR context and inline comments are captured in evidence/pr-613-review/pr-context.json and inline-comments.json using:

gh pr view 613 --json url,headRefOid,baseRefOid,comments,reviews
gh api --paginate repos/AgentWorkforce/flows/pulls/613/comments

At review time the only conversation comment was CodeRabbit's “Review skipped / Bot user detected”; reviews and inline comments were both empty. No substantive reviewer findings were available to reconcile.

Verification performed for this review

The following are independent executions at the reviewed head. The failing probes use the existing authored-root fixture, a real loopback read socket for the disconnect case, and simulated time for cancellation. They do not claim a real-daemon disconnect reproduction or mutation verification. The probe script creates and removes a temporary test file without modifying existing tests.

Reproduce both findings:

python3 evidence/pr-613-review/probe.py

Captured output (also evidence/pr-613-review/regressions.log):

$ cd packages/sdk && npx vitest run tests/review-read-disconnect.test.ts -t review: --maxWorkers=1 --minWorkers=1

 RUN  v2.1.9 /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk

stdout | tests/review-read-disconnect.test.ts > review: a reader disconnect after a timeout must not terminalize the root
{"attempts":2,"error":"journal client: connection closed","completions":[{"attempt":1,"reason":"worker_error"}]}

stdout | tests/review-read-disconnect.test.ts > review: cancellation must interrupt a pending snapshot read
{"canceled":true,"finished":false}

 ❯ tests/review-read-disconnect.test.ts (28 tests | 2 failed | 26 skipped) 39ms
   × review: a reader disconnect after a timeout must not terminalize the root 35ms
     → expected [ { attempt: 1, …(1) } ] to deeply equal []
   × review: cancellation must interrupt a pending snapshot read 3ms
     → expected false to be true // Object.is equality

⎯⎯⎯⎯⎯⎯⎯ Failed Tests 2 ⎯⎯⎯⎯⎯⎯⎯

 FAIL  tests/review-read-disconnect.test.ts > review: a reader disconnect after a timeout must not terminalize the root
AssertionError: expected [ { attempt: 1, …(1) } ] to deeply equal []

- Expected
+ Received

- Array []
+ Array [
+   Object {
+     "attempt": 1,
+     "reason": "worker_error",
+   },
+ ]

 ❯ tests/review-read-disconnect.test.ts:637:38
    635|       { dataDir: '/unused', admissionKey: 'review-disconnect' }).catch…
    636|     console.log(JSON.stringify({ attempts, error: error.message, compl…
    637|     expect(journal.peer.completions).toEqual([]);
       |                                      ^
    638|   } finally {
    639|     transport.close();

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/2]⎯

 FAIL  tests/review-read-disconnect.test.ts > review: cancellation must interrupt a pending snapshot read
AssertionError: expected false to be true // Object.is equality

- Expected
+ Received

- true
+ false

 ❯ tests/review-read-disconnect.test.ts:666:22
    664|     await vi.advanceTimersByTimeAsync(100);
    665|     console.log(JSON.stringify({ canceled: controller.signal.aborted, …
    666|     expect(finished).toBe(true);
       |                      ^
    667|   } finally {
    668|     release({ steps: { step: { state: 'completed' } } });

⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[2/2]⎯

 Test Files  1 failed (1)
      Tests  2 failed | 26 skipped (28)
   Start at  11:13:59
   Duration  1.67s (transform 869ms, setup 15ms, collect 1.44s, tests 39ms, environment 0ms, prepare 46ms)

exit=1

Focused existing regression suite:

cd packages/sdk && RELAYFLOWD_BIN=/home/daytona/.relayflows-toolchain/target/2962130851/debug/relayflowd npx vitest run tests/journal-client-read-timeout.test.ts tests/running-step-watch.test.ts tests/heartbeat-timeout.test.ts tests/run-daemon-unresponsive.test.ts tests/authored-root.test.ts tests/run-read-load-live.test.ts tests/read-timeout-resume-live.test.ts --maxWorkers=1 --minWorkers=1

 RUN  v2.1.9 /home/daytona/.relayflow-v2-supervisor/durable/repository/packages/sdk

(node:108318) [FLOWS_ROOT_LEASE_LOST] Warning: authored root run_id=root-run attempt=1: lease_conflict: attempt has no active worker lease. Waiting for the kernel to retry it.
(Use `node --trace-warnings ...` to show where the warning was created)
 ✓ tests/authored-root.test.ts (26 tests) 411ms
 ✓ tests/journal-client-read-timeout.test.ts (13 tests) 1111ms
   ✓ a recovered read timeout does not become an authored callback failure 364ms
 ✓ tests/run-read-load-live.test.ts (2 tests) 2696ms
   ✓ completes a CPU-saturating deterministic flow with reads in flight and preserves its journal 2083ms
   ✓ drains read and watch promises before an authored flow completes 612ms
 ✓ tests/read-timeout-resume-live.test.ts (1 test) 1222ms
   ✓ parks an unreadable authored root and resumes without repeating its journaled effect 1221ms
 ✓ tests/running-step-watch.test.ts (2 tests) 2121ms
   ✓ uses pushes for completion with lease-cadence reads and releases its watcher 2117ms
 ✓ tests/run-daemon-unresponsive.test.ts (2 tests) 4ms
 ✓ tests/heartbeat-timeout.test.ts (1 test) 16ms

 Test Files  7 passed (7)
      Tests  47 passed (47)
   Start at  11:12:52
   Duration  11.59s (transform 942ms, setup 47ms, collect 2.86s, tests 7.58s, environment 1ms, prepare 313ms)

exit=0

Type checking:

cd packages/sdk && npm run typecheck

> @relayflows/sdk@2.0.42 typecheck
> tsc --noEmit && tsc -p tsconfig.type-tests.json

exit=0

The PR also includes a full-suite run with failures in evidence/run-read-timeout/full-suite-final.log. That is author-supplied evidence, not an independent full-suite run from this review; it must not be described as green. This review did not rerun the full suite or establish the baseline for all of its failures. macOS CPU starvation was not reproduced in this Linux environment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

relayflowd run.get times out under CPU load and kills the run: a flow's own test suite can destroy 30 steps of journaled work

0 participants