Skip to content

fix(sdk): de-flake #611 (word_count_bounds EPIPE, queued-agents timing) - #612

Merged
khaliqgant merged 2 commits into
mainfrom
fix/611-sdk-flakes
Oct 4, 2026
Merged

khaliqgant merged 2 commits into
mainfrom
fix/611-sdk-flakes

Conversation

@AgentRelayBot

@AgentRelayBot AgentRelayBot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Closes #611. Fixes the two SDK tests that have been failing almost every linux-x64-artifact run, which is blocking v2.0.41 and #610.

(B) named-gate-diagnostics › "reports output that is not a count"

The real errno is EPIPE, not ETXTBSY. The full CI message (run 37196319563, attempt 2) is:
word_count_bounds: could not run wc -w: EPIPE

The stub wc (printf 'not a number') exits without reading stdin. If it exits before spawnSync writes input, the write fails with EPIPE. The lowering then hit if(result.error) fail('could not run …') before it checked the status, signal and stdout, which spawnSync still fills in (verified: {error: EPIPE, status: 0, stdout: 'not a number'}).

This is also a product bug. Any wc that stops reading early (it crashed, was killed, or prints its error and exits) was reported as "could not run", and what it actually did was hidden.

Fix (src/named-gate-lowering.ts): EPIPE no longer counts as "could not run", and it never affects the verdict. The verdict comes from the signal, the exit status and the output, as before. Whether the input write lost the race with wc's exit is timing, so it must not change the result.

Repro: new tests send an input of about 100 KB to a stub that exits without reading stdin. spawnSync's stdin is a socketpair. On macOS its buffer is small, so the stub always exits first: against main's lowering, 3 of 3 fail with exactly CI's could not run wc -…, and with the fix 20 of 20 pass. On Linux the socketpair buffer holds more than one env string can carry (128 KB), so the race is not forced there. The tests check the right verdict whichever way the race goes, and Linux's evidence is CI's original EPIPE failure. CI attempt 1 on this PR showed this: a test that relied on observing EPIPE failed on Linux and was removed together with the rule it tested.

The ETXTBSY hypothesis did not hold. Vitest 2 uses the forks pool and these tests are fully synchronous, and the errno says otherwise. So the write-then-exec helpers are left unchanged.

(A) authored-parallel-agents › "never starts queued agents once the body has failed"

Each f.agent runs its preflight before it asks for a worker slot (authored-worker-step.ts: check(authoring), then slots.run(admit)), and that preflight probes the wrapper's auth. The test failed the body on a fixed 100 ms timer. When preflight took longer than 100 ms, the body failed while no agent held the slot. Teardown then refused all three agents, which is the correct product behaviour, so no session ran and spans.jsonl was never created (ENOENT).

This was the test's timing assumption, not a product bug. Fix: the wrapper marks when a session has received its request, and the body fails only once that has happened. The assertion is unchanged: exactly one agent ran and the two queued agents were refused.

Repro:

  • Deterministic (wrapper auth probe slowed to 300 ms): the old test fails with exactly CI's ENOENT … spans.jsonl; the new test passes.
  • Stress (macOS, 14 CPU burners, 10 iterations): old test failed 10 of 10 with ENOENT; new test passed 10 of 10.

No blanket retries

Neither fix retries a test. (B) changes how the product classifies the result, and (A) replaces a fixed delay with the event the test actually depends on.

🤖 Generated with Claude Code


Summary by cubic

Fixes the two SDK tests that have failed nearly every linux-x64-artifact run, blocking v2.0.41 and #610.

Bug Fixes

  • word_count_bounds no longer reports EPIPE as "could not run" when wc exits before reading all of its input; the verdict now comes from the signal, exit status, and output. A numeric count from a wc that stopped reading early is refused because it does not cover the selected text.
  • The queued-agents test fails its body once a session has actually started instead of on a fixed 100ms delay, which could fire before any preflight finished and cause spans.jsonl to never be written.

Written for commit 67d18b2. Summary will update on new commits.

Review in cubic


Note

Low Risk
Targeted gate error classification and test synchronization; behavior change is stricter rejection of partial wc counts, which fixes incorrect diagnostics rather than loosening verification.

Overview
Fixes flaky CI around #611 by correcting word_count_bounds diagnostics and stabilizing a parallel-agents test.

For word_count_bounds, when wc exits before consuming all stdin, spawnSync can surface EPIPE on the input write. That case is no longer treated as “could not run wc”; the gate still evaluates signal, exit status, and stdout (and stderr suffix). If wc returns a numeric count but did not read the full selected text, the gate fails with a message that the count does not cover the input—instead of masking the real outcome behind EPIPE.

The authored-parallel-agents test “never starts queued agents once the body has failed” no longer fails the flow body on a fixed 100ms timer (which could fire before preflight finished and any agent held a slot). It waits until a session has actually started (via a started marker in the test wrapper), then fails the body; expectations stay the same (one span, queued agents refused).

New named-gate-diagnostics cases use ~100KB input so wc always exits before the pipe buffer fills, making EPIPE/partial-read behavior deterministic.

Reviewed by Cursor Bugbot for commit 67d18b2. Bugbot is set up for automated code reviews on this repo. Configure here.

…tion in the queued-agents test

(B) named-gate-diagnostics "reports output that is not a count": CI's full
message was "could not run wc -w: EPIPE", not ETXTBSY. A wc that exits
before reading all of `input` makes spawnSync's write fail with EPIPE,
and the lowering reported that as "could not run" before looking at the
status, signal and output spawnSync still returns. Product fix: EPIPE is
classified by what wc did; a numeric count from a wc that stopped
reading early is refused because it does not cover the selected text.
Deterministic regressions use an input larger than the pipe buffer.

(A) authored-parallel-agents "never starts queued agents once the body
has failed": each f.agent runs its preflight (which probes the wrapper)
before it asks for a worker slot. The body failed on a fixed 100ms
timer; when preflight outlasted it, teardown correctly refused all
three agents and spans.jsonl was never written (ENOENT). The body now
fails once a session has actually started.

Closes #611

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 5b3c6dad-4537-4bf9-a327-02331aa48dc1
@AgentRelayBot

Copy link
Copy Markdown
Contributor Author

bugbot run

@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration
  • Configuration used: Organization UI
  • Review profile: CHILL
  • Plan: Advanced
  • Run ID: fa088e93-a220-414a-b0f0-f1459f4111a9
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@cursor cursor Bot left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale Bugbot comment from a previous run.

Linux CI showed spawnSync's stdin socketpair holds a ~100 KB input, so
the race is not forced there; refusing a count only when EPIPE was
observed would make the verdict depend on that race. Drop the refusal
and its test; the remaining tests hold whichever way the race goes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

Session-Id: 5b3c6dad-4537-4bf9-a327-02331aa48dc1
@AgentRelayBot

Copy link
Copy Markdown
Contributor Author

bugbot run

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

✅ Bugbot reviewed your changes and found no new issues!

Comment @cursor review or bugbot run to trigger another review on this PR

Reviewed by Cursor Bugbot for commit cd1f567. Configure here.

@AgentRelayBot
AgentRelayBot marked this pull request as ready for review October 4, 2026 12:49

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 3 files

Re-trigger cubic

@khaliqgant
khaliqgant merged commit 60e61b4 into main Oct 4, 2026
10 checks passed
@khaliqgant
khaliqgant deleted the fix/611-sdk-flakes branch October 4, 2026 15:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Two flaky SDK tests keep failing linux-x64-artifact: authored-parallel-agents spans.jsonl ENOENT, named-gate wc ETXTBSY

2 participants