Skip to content

perf(ci): markdown-format, hook-utils and cloud-bootstrap suites add 275 s p50 to test-windows #5990

Description

@kyle-sexton

Three more test-windows suites (markdown-format hook, shared lib tests, cloud-bootstrap accounting) spend 275 s p50 combined, mostly flat per-case spawn cost on Git Bash.

Measurements

Step p50 p95 min max Detail
Test the markdown-format hook 119.0 s 133.4 s 80 145 167 PASS + 4 SKIPPED in 114.4 s, about 0.68 s/case; 86 hook-invocation call sites; 3,049 lines
Run shared lib tests (hook-utils) 85.0 s 92.1 s 62 96 597 cases in 82.7 s, about 0.14 s/case
Test cloud-bootstrap plugin accounting 71.0 s 78.0 s 47 81 79 assertions over 13 scenarios in 70.0 s, 2.7 to 4.5 s per scenario

Linux spawn proxy (strace -f): markdown-format 7,317 forks and 11,687 execs; cloud-bootstrap 2,283 forks and 8,665 execs; hook-utils not collected (stalled under strace in the timing cases). Windows timings from run 37090684624 and 40 runs 37075818834..37091316717.

Cause

  • markdown-format: no dominant case. Largest gaps are the SessionStart kill-switch-off probe at 11.2 s and kill-switch-unset at 5.9 s, then many 3 to 5 s cases. The C1 fd1-leak case measures a 715 ms base three times against an 8 s sink sleep; the sleep is not on the critical path (delta 6 ms).
  • hook-utils: about 20 s is two buffer_stdin timing cases that exercise idle and stall windows by design (sliced stall path 12.25 s, pre-4.1 delimiter-read fallback 8.12 s). The rest is flat spawn cost.
  • cloud-bootstrap: each scenario extracts the plugin block, builds fixture git repos and a stub claude CLI, and runs the block. Flat cost per scenario.

Proposed fix

Put these suites in separate jobs (the 4-job layout in #5830 puts markdown-format and hook-utils together; a finer split would separate them). Do not run them concurrently inside one job: hook-utils and markdown-format assert timing and scaling (for example bash_parse_segments: 8x the command costs 7x and the buffer_stdin idle bounds), and the hook-utils buffering cases stalled under strace, which shows they are sensitive to a slowed host. Per-suite spawn reduction is a further option: for cloud-bootstrap, build the fixture repos once and reuse them across scenarios; for markdown-format, batch the 86 hook invocations where cases allow. Measure per-case forks first.

Accuracy guard

Separate jobs change no test. For fixture sharing in cloud-bootstrap, scenarios must stay isolated (confirm no scenario mutates shared state), the 13 scenarios and 79 assertions must all still run, and the PASS list must match before and after. The by-design timing cases in hook-utils stay unchanged.

Related

#5886 (CI lane work), #5830 (13-minute PR run), #3716 (local test wall clock on Git Bash), #5958 (where it was observed). #5935 does not touch these suites.

🤖 Generated with Claude Code

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    needs-triageNot yet classified. Floor until a type and one priority tier are set.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions