Skip to content

fix: long agent-team sessions no longer slow down (TUI dev-React leak, backend busy polls) - #957

Merged
ericleepi314 merged 1 commit into
mainfrom
fix/team-session-slowdown
Sep 28, 2026
Merged

ericleepi314 merged 1 commit into
mainfrom
fix/team-session-slowdown

Conversation

@ericleepi314

Copy link
Copy Markdown
Collaborator

Summary

A clawcodex session running an agent team became very slow after a few hours. The session was still live, so the investigation measured it directly rather than guessing. It found two defects that both got worse the longer the session ran.

1. The TUI shipped React's development build and leaked memory on every render (root cause)

Symptom

  • The TUI process was at 3.1 GB.
  • An earlier TUI from the same kind of session left a memory diagnostic: 3.23 GB of heap after 6.6 h, against a 4.5 GB V8 limit.
  • Almost all of it was old-space: large_object_space held only 37 MB. That points to millions of small retained objects rather than a few big strings.
  • At that size, V8 runs frequent full garbage collections that pause the UI.

Cause

  • ui-tui/scripts/build.mjs never defined process.env.NODE_ENV, so esbuild kept React's runtime switch.
  • The launcher runs node dist/entry.js with NODE_ENV unset, so every user loaded the development builds of react, react-reconciler and scheduler.
  • React 19.2's development performance tracks call performance.measure() with a detail.devtools payload whenever a component re-renders.
  • Node keeps every entry in its global performance timeline until clearMeasures() is called, and nothing calls it.

Proof

  • The real build ran in a PTY against a fake agent-server running one long turn, with the post-gc() heap read over the inspector.
  • A heap-snapshot diff of a spinner-only run shows the new objects retained via (GC roots) → observerCallback → measureEntryBuffer → PerformanceMeasure → detail.devtools.properties, at 4,506 new measures per minute.

Heap after forced GC:

workload before after
spinner only +2.4 MB / 30 s, unbounded +0.6 MB / 30 s, bounded (see below)
+ 20 teammate progress frames/s +5.4 MB / 30 s +0.6 MB / 30 s
+ 2 tool calls/s (4 KB results) +13 MB / 30 s +1.6 MB / 30 s (transcript content)

What's left after the fix is Ink's Output.charCache. It holds at most 16,384 lines and clears when it overflows, so memory follows a sawtooth: about 31 minutes per cycle for the spinner alone, faster under streaming. It peaks around 46 MB for spinner-sized lines, or more for full-width transcript lines. That cap existed before this PR.

Fix

  • A build-time define: { 'process.env.NODE_ENV': '"production"' }.
    • The bundle now contains no development reconciler, no performance.measure( and no runtime NODE_ENV check.
    • Bundle size drops from 5.12 MB to 4.34 MB.
  • The bun run src/entry.tsx fallback, used when there's no built bundle, gets the same define: bun run --define, verified to load the production build.
  • NODE_ENV is deliberately not set at runtime. The agent-server and the user's Bash commands inherit the TUI's environment, and npm install skips devDependencies under NODE_ENV=production.
  • The other NODE_ENV readers in the bundle only gate development warnings and hooks:
    • Our ink sources compare it against development and test. Those checks were already off at runtime.
    • nanostores' !== 'production' dev hooks were on and are now off. Nothing changes, because cleanStores isn't bundled.
  • Production React reports its own errors as short codes in Ink's error screen. Component names and stacks survive, because the bundle isn't minified.

2. An idle team kept the backend busy, and the cost grew with the session

Symptom

  • The backend used 10% CPU while idle, about 8 CPU-hours in total.
  • ps -M put nearly all of it on three threads: 4.5 h, 1.7 h and 1.5 h.

Cause: two 50 ms loops re-read files that rarely change.

  • TeamRuntime._poll → sweep_mailboxes re-parsed every message each inbox had ever held, then sliced [offset:].
  • Each idle teammate re-read and parsed the whole task board through claim_next_task.

Fix: skip a file that hasn't changed since it was last read to the end.

  • sweep_mailboxes takes an optional consumed map of inbox versions (inode, size and mtime, from the new file_version()).
  • The teammate loop remembers the board version from its last empty claim.
  • Versions are taken before the read, so a write that races the read shows up on the next tick. Latency stays at 50 ms.
  • A failed read is never recorded as consumed. read_mailbox(strict=True) raises on an unreadable inbox instead of returning [], including when an exists() pre-check would have hidden the error, and the sweep retries.
  • An idle teammate still re-reads the board every 2 s, because a same-size rewrite can repeat the inode, size and mtime on filesystems with coarse mtimes.

CPU at the real 50 ms cadence, on Python 3.10 (the live interpreter), against a copy of the live team's inboxes and board:

loop before after
team mailbox sweep 2.46 ms per sweep (4.2% of a core) 0.42 ms (0.7%)
idle teammate wake 1.05 ms (2.0%) 0.35 ms (0.67%)

The remaining floor is the 50 ms wake itself (0.54% per teammate). The Claude Code reference polls every 500 ms.

Not a clawcodex defect

The lead's model steps went from about 15 s to about 50 s, but seconds per 1k output tokens stayed flat and the cache hit rate stayed at 95–97%. The steps got slower because each one produced more output (median 363 → 843 tokens) and read a context of up to 520K tokens. /compact helps with that.

Test plan

  • vitest buildProductionReact.test.ts runs the real build into a temp dir. It asserts the production reconciler is present, and that the development reconciler, performance.measure( and process.env.NODE_ENV are absent.
  • pytest, test_mailbox_poller.py:
    • an unchanged inbox is never re-read;
    • a line appended during a read is delivered on the next sweep;
    • an inbox whose open() failed once is retried on the next sweep;
    • an inbox whose exists() check failed once is retried on the next sweep.
  • pytest, test_team_runtime_e2e.py:
    • an idle teammate claims once across ten wakes, then still claims a new task;
    • a task created while a claim is in flight is still claimed;
    • with no board file, every wake claims;
    • a board whose version never changes is re-checked.
  • pytest, test_tui_launcher.py: the bun fallback carries the define.
  • Mutation-checked all 10 guards: the define, plus 9 backend and launcher guards. Each test fails when its guard is reverted.
  • The team and mailbox suites pass 5 runs in a row.
  • Full pytest: 10,822 passed. The one failure, test_headless_keeps_the_damped_wording…, is pre-existing: it fails the same way on main.
  • Full ui-tui vitest: 1,922 passed. The 7 failures also fail on main, and this PR changes no source file they exercise.
  • Critic review loop.

Follow-ups (not in this PR)

  • ui-tui has no CI job, so the new vitest guard only runs locally.
  • The launcher runs node dist/entry.js without the --max-old-space-size=8192 from the shebang in entry.tsx, because the build strips the shebang. memoryMonitor.ts derives its thresholds from the real ~4.5 GB limit; only its comments still say 8 GB.
  • The lead's session shows five near-total prompt-cache misses within a minute of the previous call. The worst was 492K uncached tokens. That isn't cache expiry, so the cause is worth checking.
  • Existing installs pick up the TUI fix the next time the installer runs npm run build in ui-tui.

🤖 Generated with Claude Code

…, backend busy polls)

A session running an agent team grew slower over a few hours. Measured on
the live session:
- The TUI was at 3.1 GB. An earlier TUI from the same kind of session hit
  3.2 GB against V8's 4.5 GB limit.
- The idle backend burned about 8 CPU-hours.

TUI (root cause): scripts/build.mjs never defined NODE_ENV, and the launcher
runs `node dist/entry.js` with NODE_ENV unset, so every user loaded React's
development build. Its performance tracks call performance.measure() as
components re-render, and Node keeps every entry in its global timeline.
A heap-snapshot diff traced the growth to measureEntryBuffer: about 200 MB/h
from the spinner alone, roughly three times that with agent progress.
- Fold NODE_ENV to "production" at build time. Setting it at runtime would
  leak into the agent-server and user commands.
- The bun fallback gets the same define.

Backend: two 50 ms loops re-read files that rarely change.
- The team mailbox sweep re-parsed every message each inbox had ever held.
- Each idle teammate re-read the whole task board.
Both now skip a file whose (inode, size, mtime) version is unchanged since
they last read it to the end. The version is taken before the read, so a
racing write is picked up next tick. A failed read is never marked
consumed. An idle teammate still re-reads the board every 2 s, because
coarse-mtime filesystems can repeat a version.

At the real cadence on Python 3.10:
- sweep: 2.46 -> 0.42 ms
- idle wake: 1.05 -> 0.35 ms

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

Test Results

     5 files   1 021 suites   21m 51s ⏱️
16 092 tests 16 070 ✅ 22 💤 0 ❌
32 154 runs  32 083 ✅ 71 💤 0 ❌

Results for commit 843ddb8.

@ericleepi314
ericleepi314 merged commit 07567fa into main Sep 28, 2026
8 checks passed
@ericleepi314
ericleepi314 deleted the fix/team-session-slowdown branch September 28, 2026 03:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant