Skip to content

feat(performance): make monitoring always available and explain slowdowns #944

Description

@Juliusolsson05

Status (quality-loop audit, 2026-09-25)

Landed: Monitoring shipped in #949, #955, #958 and #986; monitorCoordinator.start() runs unconditionally.

Remaining: The qualification run in docs/superpowers/plans/2026-09-12-performance-monitoring-qualification.md: closed-monitor packaged CPU, utility RSS, startup p95, 8 h workload, 24 h disk growth, signed arm64 and Intel.

Acceptance: A qualification run with those numbers attached. Owner call: does it block a release, or get tracked on its own? (needs-owner)

Original report

Agent Code needs performance monitoring available to every user without an environment flag, reachable through Settings and an ordinary command. It should explain perceived slowness, preserve useful recent evidence, and distinguish application overhead from agent/provider activity.

The planning audit at main 115e26f found an existing opt-in PerformanceService and renderer client, a visible-pane-only process panel, a separate system-memory poller, and always-on AppRunJournal/freeze/heap watchdogs. These are useful foundations, but collection ownership, retention, and user-facing explanations are split. #767 already documents diagnostic overhead; simply enabling all existing tracing for everyone would not satisfy this feature.

Intended behavior:

  • A lightweight baseline runs for all users, including when the dashboard is closed.
  • Settings → Performance and a normal Performance Monitor command open the same application-wide view.
  • Show responsiveness, CPU/memory, all managed agent processes (including detached agents), costly operations, recent incidents, sample age, and missing coverage.
  • Explain slowdowns using measured evidence and confidence, with drill-down to a timeline. CPU activity alone is not an error or proof of a cause.
  • Preserve bounded local history across restart, and allow explicit content-minimized report export. Central uploads are an unresolved product option, not implied by always-on monitoring.
  • Offer time-limited advanced profiling separately, with clear overhead and artifact-content information.
  • Consolidate existing samplers and use bounded queues, isolated storage/aggregation, and self-observation. Production monitoring must not trigger automatic stop-the-world heap snapshots.

Acceptance:

  • Define and measure idle/streaming overhead, hot-path cost, resident memory, disk growth, and multi-window scaling before enabling the new baseline by default.
  • Handle sleep/resume, renderer reload/crash, process-ID reuse, slow/full disk, collector failure, unsupported metrics, and prolonged heavy agent output without unbounded work or false healthy states.
  • Maintain current incident correlation and crash evidence while migrating existing diagnostics; exact phase boundaries and retained exceptions belong in the implementation plan.
  • Add meaningful behavioral/fault/packaged-app verification and reproducible overhead benchmarks. Budget numbers in the design are provisional targets until measured.
  • Every implementation PR receives two independent Agent Code orchestration reviews and disposition of valid feedback.

Related: #767 (diagnostic overhead), #103 (general optimization), #327/#365 (memory investigations), #370 (crash evidence), #243 (dictation reliability). Product-usage analytics #210 is a separate concern.

Status: detailed architecture and rollout plan committed and pushed on feat/performance-monitoring (first commit 984a4ed4, refined at 32a928b6). Production monitoring has not been changed; no implementation PR is open yet.

Implementation plan covers shared sampler ownership, independent baseline/verbose configuration, bounded transport and retention, process attribution, operation/incident explanations, Settings/commands, advanced profiling, privacy, failure handling, provisional overhead budgets, and six reviewable implementation stages.

Planning validation: current-source and existing-issue audit, official Electron/Node/browser API checks, source-reference validation, and git diff whitespace checks. Runtime tests and performance benchmarks have not been run for this documentation-only pass. The initial budgets must be measured before production rollout.

Implementation progress — September 12

Implementation of all six stages is authorized. Stage 1 is PR #949 (bounded contracts, queues, latency rollups and synthetic producer benchmark); two independent Agent Code orchestration reviewers found malformed histogram inputs, now fixed with regression coverage. Stage 2 (#950) implements the shared baseline sampler, bounded utility-process collector, renderer lifecycle/transport credit, and metadata-only automatic heap pressure. Thirteen focused collector/renderer tests pass; full type-check/build verification is running. User-facing UI, full process attribution, incidents, history/export, profiling and soak qualification remain in the planned follow-on stages.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    class:C7-new-featureBug cluster of a recently shipped featureneeds-ownerOwner decision neededtype:featureNew capability (out of quality-loop fix scope)

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions