From efbe85c0dc1ca8636ffcb4b5ff3816c413e8d5f4 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Thu, 1 Oct 2026 14:12:46 -0700 Subject: [PATCH 01/13] Split build-and-test into phase slices The skill becomes a core (rules that bind in every phase, a step map, steps 1-3) plus one reference per phase: build.md (step 0a-0c), drive.md (step 0d-0f and the drive rules), evidence.md and funding.md. Rule ids are unchanged. The skill-read gate delivers the build slice with the build scripts and the drive and evidence slices with capture-buy-quote.sh and with every maestro drive. The funding slice arrives with a value-moving log-attempt.sh call. build-and-test-slices.test.py holds the core and each slice under 20,000 characters and checks that every pre-split rule id is placed once. --- .cursor/skills/build-and-test/SKILL.md | 146 ++--------- .../build-and-test/maestro/buy-quote.yaml | 2 +- .../skills/build-and-test/references/build.md | 75 ++++++ .../skills/build-and-test/references/drive.md | 68 +++++ .../build-and-test/references/evidence.md | 12 + .../build-and-test/references/funding.md | 12 + .../hooks/require-playbook-before-drive.sh | 41 ++- .../hooks/require-skill-read-for-scripts.sh | 29 +- .../hooks/tests/build-and-test-slices.test.py | 248 ++++++++++++++++++ .../build-and-test-pre-split.SKILL.md | 178 +++++++++++++ 10 files changed, 685 insertions(+), 126 deletions(-) create mode 100644 .cursor/skills/build-and-test/references/build.md create mode 100644 .cursor/skills/build-and-test/references/drive.md create mode 100644 .cursor/skills/build-and-test/references/evidence.md create mode 100644 .cursor/skills/build-and-test/references/funding.md create mode 100644 agent-watcher/hooks/tests/build-and-test-slices.test.py create mode 100644 agent-watcher/hooks/tests/fixtures/build-and-test-pre-split.SKILL.md diff --git a/.cursor/skills/build-and-test/SKILL.md b/.cursor/skills/build-and-test/SKILL.md index 6192754b..80554eb0 100644 --- a/.cursor/skills/build-and-test/SKILL.md +++ b/.cursor/skills/build-and-test/SKILL.md @@ -9,126 +9,46 @@ metadata: Verify the active repo builds cleanly before /one-shot marks a task complete. Returns a clear PASS/FAIL signal the caller can include in the Asana summary or use to gate the watch loop. - + Inspect the current working directory to decide what to run: -1. If `package.json` `name` is `edge-react-gui` → iOS UI test (maestro) path (step 0). Check this first. +1. If `package.json` `name` is `edge-react-gui` → iOS UI test path (step 0: `references/build.md`, then `references/drive.md`). Check this first. 2. Else if the repo is an EdgeApp gui DEPENDENCY (per `gui-dependency-integration`) → run its own checks (the TS/Node path below) AND the gui integration test. A dep change is NOT done until it runs in the app. 3. Else if `package.json` exists and a `tsconfig.json` exists → Node + TypeScript path (step 1). 4. Else if `package.json` exists with a `test` script but no tsconfig → Node path (step 2). 5. If `Cargo.toml` exists → not implemented yet, fall through to placeholder. 6. Otherwise → placeholder mode (step 3). -On FAIL, surface the exact command, exit code, and last 30 lines of output. Do not try to fix anything inside this skill — the caller decides whether to amend or block. -This skill does NOT edit source code, commit, push, or change Asana state by default — verification + results only. The only exceptions are the explicitly scoped rules below: `testids-over-coordinates` (testID additions as a separate commit IN the task's single gui PR — never a separate testID-only PR), `gui-dependency-integration` (gui-side changes a dep actually needs, committed on the gui branch), and the LOCAL-ONLY, never-committed corePlugins edits — `single-asset-plugin-trim` (currency trim) and `force-swap-provider-locally` (force a swap provider). -Scoped exception to `no-mutation`, test-infrastructure only. TESTIDS FIRST, COORDINATES LAST. When a maestro flow needs to drive an element that has no stable selector (text match fails and no `testID` exists), the DEFAULT action is to ADD the missing `testID` prop to that component in the gui worktree and drive via it — do NOT struggle with coordinate taps. A `testID` is a JS-only prop: Metro reload picks it up in seconds (no native rebuild), so adding one is cheaper than even a single round of coordinate trial-and-error, and it de-brittles the suite for every future run. Coordinate taps are permitted ONLY for surfaces you cannot edit (system dialogs, native pickers, third-party views that don't forward `testID`) or when a reload would destroy unrecoverable in-flight app state — and any coordinate tap that survives into the PROOF flow must be called out in the run report with why a testID was not possible. Commit discipline: commit the testID additions as a SEPARATE commit, distinct from any feature commit; change ONLY `testID` props, never component logic; update the maestro selector(s) to use them. MESSAGE NAMES THE SURFACE: subject `test: add testIDs to ` (e.g. `test: add testIDs to ExchangeScene swap pills`), with every added id listed in the body — never a generic subject like `test: add missing testIDs for maestro selectors`: identical subjects across runs make these commits indistinguishable when a human cherry-picks between branches. WHERE THE COMMIT LANDS — always THE TASK'S SINGLE GUI PR, the same one whose test surfaced the need; NEVER a separate testID-only PR. A task has at most ONE gui PR: for a gui-feature task the testID commit rides that feature PR; for a DEP-repo task whose maestro test drives the gui, the testIDs go in the task's ONE gui integration PR (and if the testIDs are the only gui change, that PR IS the task's gui PR — a first-class PR, not a throwaway), on the SAME `/` gui branch the run already provisioned. Do NOT cut a second gui branch/PR for the testIDs when the task already has (or will have) a gui PR — that splits one task's gui work across two PRs, which is the mistake this forbids. If no selector was missing, this rule is a no-op. -OPTIMIZATION (optional, LOCAL-ONLY — never committed). When the task targets a SINGLE asset and the maestro test needs to drive that asset's wallet, you MAY temporarily comment out the unrelated currency plugins in the gui worktree's `src/util/corePlugins.ts` (the `currencyPlugins` map) — keeping the plugin(s) the task needs — to cut app load/init time (fewer plugins to spin up). This is a test-harness speedup ONLY: it must NEVER land in a commit or PR. Revert it before any commit, or rely on it living only in the throwaway test build; if you commit after trimming, verify `git status`/`git diff` does NOT include `corePlugins.ts`. Skip entirely for multi-asset tasks or tasks that don't drive a wallet. -To FORCE a specific swap provider for a test (so the engine routes through it instead of a competitor), edit the gui worktree's `src/util/corePlugins.ts` `swapPlugins` map and set every OTHER provider to `false`, leaving only the target's `*_INIT` truthy — LOCAL-ONLY, same corePlugins surface and same never-committed discipline as `single-asset-plugin-trim` (revert before any commit; verify `git status`/`git diff` excludes `corePlugins.ts`, or rely on the throwaway build). Do NOT force a provider by toggling the in-app **Settings → Exchange Settings**: that state is ACCOUNT-SYNCED, so on a shared roster account it thrashes against every parallel session and persists to the next run — parallel-UNSAFE and forbidden as the forcing lever. Use Exchange Settings only to READ/diagnose why a provider is absent, never to set routing. (`Preferred`/`preferPluginId` also do not pin — engine reverts to best-rate ~60s — so the local hard disable is the reliable lever.) See sim-testing-playbook "Feature-enablement check". -Deterministic operations (sim selection, RN build, capture loop) MUST run via the companion scripts under `~/.cursor/skills/build-and-test/scripts/`. Do not inline their logic as raw bash blocks in this SKILL.md or in agent reasoning. -START the sim-testing phase with ONE call: `~/.cursor/skills/build-and-test/scripts/slot-preflight.sh` (defaults to `$AGENT_SIM_UDID`/`$AGENT_METRO_PORT`; pass `--repo ` when cwd is not the gui repo). It boots the sim if needed and answers, deterministically, the questions runs kept re-deriving at high friction cost: is Metro mine or squatted, is the app installed, does the installed native side match the worktree (`.agent-native-build-stamp` vs `ios/Podfile.lock`), are node_modules present. OBEY its final `PLAN:` line — `ready` (drive now, no build), `js-only`/`install`, or `full-rebuild` — and run the exact `INVOKE:` command it prints (verbatim; no flag-guessing), then its `WAIT:` command when it prints one. Do NOT re-derive any of its checks manually, and do NOT start a second Metro when it reports one running. -When verification needs RUNTIME state from the running app — why a check evaluates false, the actual value of a variable, which code path executed (e.g. the Swap/Maya "investigate outage" kind of task) — use the `/debugger` skill (`~/.cursor/skills/debugger/SKILL.md`), do NOT hand-roll a CDP/WebSocket attach. It sets a `file:line` breakpoint over Metro's Hermes inspector and reports the call stack + locals. It is already slot-aware: `check-metro.sh` and `cdp-attach.js` default to `$AGENT_METRO_PORT`, so in a parallel slot it targets THIS session's Metro (base 8181) with no port flags. Static questions (where is X defined) stay grep/read — `/debugger` is only for live runtime state. -A critical-path wait (build, Metro bundle, screenshot, app-ready — anything you cannot proceed without) MUST be a single BLOCKING call inside the CURRENT turn. NEVER end your turn and hand the wait to a backgrounded shell expecting "the background task will re-invoke me when it finishes." That makes your own forward progress depend on an external re-invoke, and when the wait can't complete you idle forever with no one driving — the failure that wedged the BitcoinDepot and piratechain runs. This is the same disease one-shot's `never-self-respawn` already forbids: *"any wait is a single blocking call in THIS process."* Concretely: -- **Do the wait, get a result, react — all in this turn.** Foreground it. The harness's background-completion → re-invoke is for genuinely parallel/optional work, NOT for a step the next step depends on. -- **Bound every wait with `timeout `** so it ALWAYS terminates (success OR timeout) and control returns to you to react. (`timeout` IS available — macOS ships no `timeout`/`gtimeout`, so it's provided on PATH by the portable shim `~/.cursor/skills/timeout.sh`; `timeout 180 ` just works.) An unbounded `until grep ; do sleep 5; done` / `while ! ; do sleep; done` hangs forever the moment the marker never appears (wrong logfile, wrong marker, build died). A timed-out wait is a real FAIL/retry to handle now — never a reason to spawn another waiter. -- **iOS builds run detached, then wait in chunks.** A cold build outlives the Bash tool's 600s cap (the call is killed mid-build), so never run a full build as one foreground call or under a hand-rolled nohup/poll loop: use step 0c's two commands and re-run `ios-rn-build-wait.sh` in this turn while it exits 7. Wait exits 0/1/2 are the build's own result; 3 = stalled (no log output for 10 min) and already killed: read the printed log tail, fix, start a fresh `--detach`; 4 = no build to wait on: start one. +On FAIL, surface the exact command, exit code, and last 30 lines of output. Do not try to fix anything inside this skill: the caller decides whether to amend or block. +This skill does NOT edit source code, commit, push, or change Asana state by default: verification + results only. The only exceptions are the explicitly scoped rules: `testids-over-coordinates` (testID additions as a separate commit IN the task's single gui PR, never a separate testID-only PR), `gui-dependency-integration` (gui-side changes a dep actually needs, committed on the gui branch), and the LOCAL-ONLY, never-committed corePlugins edits: `single-asset-plugin-trim` (currency trim) and `force-swap-provider-locally` (force a swap provider). +Deterministic operations (sim selection, RN build, capture loop) MUST run via the companion scripts under `~/.cursor/skills/build-and-test/scripts/`. Do not inline their logic as raw bash blocks in this skill's files or in agent reasoning. +A critical-path wait (build, Metro bundle, screenshot, app-ready: anything you cannot proceed without) MUST be a single BLOCKING call inside the CURRENT turn. NEVER end your turn and hand the wait to a backgrounded shell expecting the background task to re-invoke you when it finishes: your forward progress then depends on an external re-invoke, and when the wait cannot complete you idle forever with no one driving. Same contract as one-shot's `never-self-respawn` and `pr-watch-bounded-poll`; an outside watchdog is NOT the safety net. +- **Do the wait, get a result, react, all in this turn.** Foreground it. Background-completion re-invoke is for parallel/optional work, NOT for a step the next step depends on. +- **Bound every wait with `timeout `** so it ALWAYS terminates (success OR timeout) and control returns to you. (`timeout` is on PATH via the portable shim `~/.cursor/skills/timeout.sh`; macOS ships none.) An unbounded `until grep ; do sleep 5; done` hangs forever when the marker never appears. A timed-out wait is a real FAIL/retry to handle now, never a reason to spawn another waiter. +- **iOS builds run detached, then wait in chunks.** A cold build outlives the Bash tool's 600s cap, so never run a full build as one foreground call or under a hand-rolled nohup/poll loop: use step 0c's two commands (`references/build.md`, which also lists the wait's exit codes) and re-run `ios-rn-build-wait.sh` in this turn while it exits 7. - **Any other long compile** (gradle, a hand-run xcodebuild) needs a stall check inside its bounded wait: log mtime frozen with no live compiler children means HUNG now; kill, diagnose, retry instead of waiting out the timeout. Use `capture-buy-quote.sh` (bounded retry cycles) for app capture. -- **Detect readiness against the resource you actually started, not a guessed log line:** `timeout`-bounded `curl` against the Metro you launched on its REAL port (`/status`, then the `index.bundle` URL) — never `grep` a logfile whose name/marker you assumed (the bug here: Metro logged to `gui-metro2.log` but the waiter grepped `gui-metro.log`). -Mirrors one-shot's `never-self-respawn` and `pr-watch-bounded-poll`. Recovery by an outside watchdog is explicitly NOT the safety net — the agent must not hang in the first place. -Never assume a repo's package manager — repos migrate between npm and yarn (edge-react-gui is currently yarn-locked; package-lock.json was removed upstream). All install/run/pack operations go through the shared dispatcher `~/.cursor/skills/pm.sh`, which detects the lockfile (`package-lock.json`→npm, `yarn.lock`→yarn, both/neither→npm). Companion scripts in this skill already dispatch through it; do not hand-write `npm ...`/`yarn ...` against a repo without checking `pm.sh detect`. -A change to an EdgeApp gui DEPENDENCY is NOT fully tested until it runs in the app — its own `tsc`/jest passing is necessary but NOT sufficient. Gui dependencies = the Edge-owned repos `edge-react-gui` consumes: `edge-core-js`, `edge-currency-accountbased`, `edge-currency-plugins`, `edge-exchange-plugins`, `edge-login-ui-rn`, `edge-currency-monero`, `react-native-piratechain`, `react-native-zcash`, `react-native-zano`. When the repo under test is one of these, after its own checks you MUST also run the gui integration test, autonomously (NO prompting): -1. **Co-located gui worktree:** ensure one exists — create via `~/.config/agent-watcher/setup-task-workspace.sh --task-gid --repo edge-react-gui` if absent (sibling of the dep worktree under `~/git/.agent-worktrees//`, so updot can find it). -2. **Link the MODIFIED dep into the app — the mechanism, and whether you flip any `DEBUG_*` flag, is YOUR per-task call** (depends on what the task changed and how you want to verify it; it is NOT a fixed per-dep rule). Run repo scripts with each repo's package manager (lockfile: `yarn.lock`→yarn, `package-lock.json`→npm; **yarn is being phased out — check, don't assume**). The toolbox: - - **`updot` — bakes the built dep into the gui's `node_modules`.** Works for ANY dep, no dev-server, no runtime race → the safe default for headless/automated runs. ` updot ` then the gui's `prepare` (npm form: `npm run updot -- && npm run prepare`; add `prepare.ios` for native-module deps), then rebuild. The dep's `DEBUG_*` flag stays FALSE (you baked it in). - - **`DEBUG_` flag + the dep's live webpack dev-server — webview-plugin deps only** (`DEBUG_ACCOUNTBASED`:8082, `DEBUG_EXCHANGES`:8083, `DEBUG_CURRENCY_PLUGINS`:8084, `DEBUG_PLUGINS`:8101 — these ports are HARDCODED in each dep package's `debugUri` and are HOST-GLOBAL). Set the flag TRUE in the gui's `env.json` AND run the dep's `yarn start`/`npm start` (webpack serve) backgrounded for the test; the webview loads the local bundle live (sim reaches host localhost), no gui rebuild. Pick this when live iteration helps; if it flakes (dev-server unreachable, ATS/cleartext, recompile race) fall back to updot. - - **Parallel-slot port rule (this bit the Swap/Maya run):** a `DEBUG_` dev-server port is a SINGLE-OCCUPANT host resource — only ONE slot can serve a given dep at a time. A second concurrent session needing the SAME dep MUST use updot instead. Before starting the dev-server, check the port is free: `lsof -nP -iTCP: -sTCP:LISTEN`; if another slot holds it, use updot. Your slot's Metro runs on `$AGENT_METRO_PORT` (base **8181**, i.e. 8181/8182/8183…), deliberately OUTSIDE the 808x DEBUG range so Metro never collides with a dev-server — do NOT pass a `--port` that drags Metro back into 808x. When in doubt in a parallel slot, prefer updot: it has no shared port and is collision-free by construction. - - **`DEBUG_EXCHANGES` crash-loop trap (Swap/Maya):** the gui's `allowDebugging` flag (which permits the cleartext localhost load) is OR-gated on `DEBUG_ACCOUNTBASED || DEBUG_CORE || DEBUG_CURRENCY_PLUGINS || DEBUG_PLUGINS` — **`DEBUG_EXCHANGES` is NOT in that set**, so enabling it ALONE crash-loops the app. Co-enable one that IS (e.g. `DEBUG_ACCOUNTBASED`); note that drags in its 8082 dev-server, so plan ports per the rule above. Also: swap/exchange plugin code runs in **edge-core-js's webview context, not the Metro bundle** — serve patched dep code via the dev-server (or `updot`-bake it); do NOT sync patched `lib/` into `node_modules` expecting Metro to bundle it. - - **`edge-core-js`: prefer `updot`, avoid `DEBUG_CORE`.** `DEBUG_CORE` loads the WHOLE core from hardcoded `http://localhost:8080/` (`edge-core-js/.../react-native-webview.tsx`: `source={debug ? 'http://localhost:8080/' : null}`) — races init, cleartext/ATS-sensitive, and any hiccup takes the entire app down (the long-standing "DEBUG_CORE is buggy"). updot is reliable for core. - Only link the dep(s) THIS task modifies; leave every other dep's `DEBUG_*` at its env.json default. Keep flags consistent with what you actually linked — a `DEBUG_*` left true with no dev-server running will break that dep. -3. **Login:** the test account auto-logs-in via the `YOLO_*` env knobs (set by workspace init to the roster's `agent` account from `~/.config/edge-secrets/test-accounts.json`, consumed in `LoginScene.tsx` — pinned by `setup-task-workspace.sh` on every worktree's env.json copy). Keep them set so the maestro run reaches the logged-in app; when the change is to `edge-login-ui-rn` specifically, these are the lever for exercising the login flow — adjust only if the change requires driving the login UI differently. -4. **Make the gui-side changes the feature NEEDS to run, then run the gui maestro path (step 0)** against that build. A dep change almost always needs gui-side wiring to actually function — plugin init / apiKey, provider/plugin registration, imports, config. Those gui changes are PART OF THE WORK, not optional: complete ALL of them (and commit on the gui worktree's branch) so the app is fully runnable with the feature, autonomously, do NOT prompt. Do not stop at "the dep compiles" or "it links" — if the feature doesn't load/run in the app yet, the implementation is NOT done. -PASS requires the maestro app test to pass with the dep change linked AND the actual feature exercised to its terminal success (`test-drives-the-real-action`). A dep whose unit checks pass but that isn't fully wired into a runnable app, or that runs but whose real action was never executed, is a FAIL. -SCOPE DOES NOT EXEMPT THE TEST (the Houdini-prototype rationalization, 2026-06-11): a task that scopes its deliverable to the dep repo, calls itself a prototype, or explicitly defers PRODUCTION gui integration to follow-up work still gets THIS integration test. The wiring in steps 1-4 is TEST SCAFFOLDING in the task's gui WORKTREE (plugin registration, env.json keys, dep linking) — it is not an "unrequested production change"; nothing lands in the gui repo unless the task asks for it. Likewise "the plugin is unvetted prototype code, a real swap through it is irreversible" is NOT a blocker: vetting it with a small sanctioned-roster swap is exactly what this test exists to do (see one-shot `yolo-true-blockers` carve-out). -PLATFORM: default to iOS. Provision the iOS sim, run the iOS maestro flow, and credit `iOS Sim` UNLESS the task EXPLICITLY calls out Android (task title/description says Android, the task is tagged Android, or the change is under `android/` only). For an Android-called-out task, run the ANDROID path instead of (or in addition to) iOS: `./gradlew :app:assembleDebug` from the gui worktree's `android/` is the build verification, and a successful APK is the terminal-success signal for a BUILD-ONLY fix (GitHub `pr-checks.yml` does NOT build Android, so these regressions are invisible to GitHub checks — the local assembleDebug is what catches them). Credit `Android Sim` and log the attempt via `log-attempt.sh --category test-drive --result success|failed:`. The Android build needs gitignored secrets the node_modules clone does not carry (`android/app/google-services.json`, `EdgeApiKey.java`, `android/app/src/main/assets/edge-core/plugin-bundle.js`, a generated `android/local.properties` with `sdk.dir`) — `setup-task-workspace.sh` copies them; and the Android SDK (`ANDROID_HOME`/`ANDROID_SDK_ROOT`) must be present in the env. Run gradle with `--no-daemon` (or a per-slot `GRADLE_USER_HOME`) for parallel-safety; assembleDebug is CPU/RAM-heavy, so do not run many concurrently. A genuine in-app Android drive (AVD + maestro) is a larger path; the build-only check closes the regression gap for build/native fixes. If a task touches BOTH platforms, exercise iOS and credit both. -DEFAULT to physically exercising the change in the running app on the sim. Almost ANY task can be tested in-app — a swap, a send, a settings toggle, an onboarding/account-creation flow, a specific wallet action, a bug repro. `tsc`/jest/build passing is NECESSARY BUT NOT SUFFICIENT: a change is not verified until you have driven the actual changed behavior in the app via maestro and seen the expected result — to its TERMINAL success, not a precursor (see `test-drives-the-real-action` for the exact bar: execute the real action, e.g. an actual swap, not just a quote). Do NOT skip the sim test because static analysis "looks right", because the diff is small, or because authoring a flow is effort (Rango shipped a swap-plugin change with NO in-app test — that is the failure this rule forbids). Specifically: before setting `blocked = Yes` with reason "can't verify / no defensible default" on a bug, repro, or investigation task, you MUST first attempt the most-specific RUNTIME REPRO you can construct — build the relevant flavor (e.g. `ENABLE_MAESTRO_BUILD=true` for test-server flows) and drive the precise maestro flow. "I can only trace it statically" is NOT a blocker. Block only if the repro is genuinely un-runnable here (missing creds/KYC/datastore the slot can't provide). For FUNDS specifically: the ONLY funds blocker is an OBSERVED TRUE LOSS — an attempted swap/send that failed AND lost principal. Fees/slippage NEVER count as loss (budgeted at $15 equivalent per run, per the playbook), and blocked-ness is established by ATTEMPTING, never predicted. -When the asset or feature under test genuinely CANNOT be driven in the sim — a default-disabled plugin that crashes the debug build on enable (e.g. `BOTANIX_INIT`), a date-gated change that only activates after a future date, or a non-GUI-dependency surface — the bar is NOT "skip the in-app test and declare `verified: not-run`". Do BOTH: (1) drive a PROXY that exercises the SAME mechanism your change routes through and capture proof of it reaching its terminal state (e.g. for a keys-only create-wallet exclusion, a hardcoded-enabled keys-only asset like `bitcoinsv` hits the identical exclusion path — see the sim-testing playbook); and (2) unit-test the gate/branch your change adds (the condition that routes the target asset) so the logic is covered even though the asset itself cannot run. Record both in the Testing section and set `verify_blockers: [precondition]` (asset un-runnable here) with the proxy drive + unit test as the evidence. That combination is a sanctioned PASS; `verified: not-run` with an empty `verify_blockers` is not. When the un-drivable surface is VISUAL, the combination is not sufficient on its own — add the hack-verified frame per `hack-verify-visual-changes`. -LOG every value-moving action and every test-drive/repro the moment it resolves, via `~/.config/agent-watcher/log-attempt.sh --gid --action "" --result success|failed:|loss:|blocked: --category swap|send|sweep|test-drive|repro`. This attempt-log (`$XDG_STATE_HOME/agent-watcher/attempts/.jsonl`) is the AUTHORITATIVE, agent-location-independent record of what the run actually attempted — the concession-validation gate reads it to tell a real wall (`loss:`/`failed:`/`blocked:` after an attempt) from a predicted one, on BOTH a formal `--blocked yes` AND a silent DOWNGRADE-finalize (completing or opening a PR without reaching the prescribed in-app success), and the eval reads it as ground truth for testing-depth instead of trusting transcript narration. RESULT semantics: `success` = reached terminal success; `failed:` = attempted, no success, principal safe (fees only); `loss:` = attempted, FAILED, principal unrecoverable (the ONLY funds condition that legitimizes a block); `blocked:` = attempted up to a precondition the slot genuinely cannot satisfy (real provider halt, geo-block confirmed by attempt). This matters most after the tester becomes its own agent: the drive will live in the tester's context, NOT this transcript, so the log is the only place a concession gate or the eval can see it — write it on EVERY attempt now so the contract is already in force when that split lands. -The test harness is YOURS to build — its absence is NEVER a blocker. When driving the real behavior needs scaffolding that does not exist yet, CREATE it locally and uncommitted: author a new maestro `.yaml` flow (per `maestro-flows-are-shortcuts` — expected, not exceptional), add a missing `testID` (`testids-over-coordinates`), trim unrelated plugins (`single-asset-plugin-trim`), disable a crashing module (the piratechain local-disable), or HARD-CODE the inputs the code path reads — fixtures, seed data, info-server/remote-config payloads, feature-flag state, a forced provider. "No maestro flow exists for this", "the data comes from a remote server I don't control", "there's no fixture", "the feature isn't enabled by default" are NOT blockers and NOT reasons to stop at static analysis — they are scaffolding to BUILD. KEY DISTINCTION: hard-code the INPUTS to REACH and exercise the real logic, never fake the OUTPUT to fabricate a pass. Injecting a disable-map into the store so the REAL `isSpendBrandDisabled` filter runs against controlled data is correct; hard-coding "this brand is hidden" to skip the filter is not — the changed code path must actually execute. LOCAL-ONLY discipline (same as `single-asset-plugin-trim`): this scaffolding is throwaway — it must NEVER land in a commit/PR. Revert it before any commit (verify `git status`/`git diff` is clean of it), or rely on it living only in the disposable test build. If you find yourself writing `blocked = Yes` or "could not test because doesn't exist", stop: build the scaffolding and drive the test. -Before the sim-test phase, READ `~/.cursor/skills/build-and-test/references/sim-testing-playbook.md` — it is short and holds the working knowledge (funding floors, account roster/switching, feature-enablement gotchas, investigation order) that otherwise gets re-learned every run. Then COMPOSE, don't re-derive: parameterized subflows live in `~/.cursor/skills/build-and-test/maestro/common/` (`login-if-needed`, `dismiss-startup-modals`, `select-swap-pair`, `confirm-slider` — the slider is SOLVED there; never re-derive the gesture). Copy the subflows you need next to your task flow and `runFlow` them; author NEW task-specific `.yaml` liberally for what the task actually changed (expected, not exceptional), keeping task flows LOCAL (`.syncignore`d from the agent repo; never committed to the gui repo — its `maestro/` is the heavyweight verification suite, reference-only for selectors). What DOES get committed to the gui: missing `testID`s, per `testids-over-coordinates` (add them the moment a selector is missing — never grind coordinates on an editable component). EXPLORATION vs PROOF: for exploring screens/selectors use the **maestro MCP tools** (persistent driver — no ~2-min `maestro test` startup per probe). CAUTION — the maestro MCP daemon binds to one device and DRIFTS. `maestro-mcp-wrapper.sh` start-pins it (`maestro --device "$AGENT_SIM_UDID" mcp`), but the start-pin does NOT survive slot-sim relaunch cycling: the daemon rebinds to a different booted pool sim mid-run, so it then drives/observes ANOTHER slot's sim (the confirmed cause of a red-springboard "proof" of the wrong device). Therefore: (1) the MCP is for EXPLORATION ONLY — re-verify its bound device (list_devices + an inspect matching your app's expected state) at the START OF EVERY observation block, not once, because it can rebind between blocks; (2) ALL PROOF evidence comes from the maestro CLI — canonical slot invocation: `maestro --device "$AGENT_SIM_UDID" --driver-host-port $((AGENT_METRO_PORT + 1000)) test ` (the per-slot driver port keeps parallel slots' iOS drivers off each other's ports; hook-enforced; the MCP daemon uses +2000 via its wrapper) — + `simctl io` against that SAME UDID — NEVER an MCP screenshot and NEVER a parallel `simctl io` against a different device than the CLI drove. If a verify shows the MCP on the wrong device, drop to CLI for the rest of the run. For the repeatable PROOF run compose ONE yaml flow and run it once via `capture-buy-quote.sh --flow ` (CLI-driven, auto-pins device + driver port from the slot env; that run produces the PR evidence screenshots). When a run teaches you something durable, PROPOSE it as a `[playbook]`-tagged bullet in your run report's Dev Notes & Gotchas section — do NOT edit the playbook directly. A NEW reusable drive sequence (not knowledge, a FLOW) is proposed the same way with a `[flow]` tag: name, params, one-line purpose, and the FULL yaml EMBEDDED in the report as a fenced block (worktrees are pruned on retention — a path reference dies with the worktree; the report attachment is the durable copy). A change to an EXISTING library flow (genericize, new param, split into subflows) uses `[flow-update]`: name the flow, the change, and the compatibility argument — new params MUST default to current behavior so existing callers are unaffected, and a rename/split must say so explicitly (callers get grepped at promotion). Promotion is NEVER done by task runs: proposals wait in reports for the eval's manually-triggered flow-consolidation pass. The playbook is operator-curated: entries in it may be trusted without re-verification precisely because every one was reviewed and promoted by the operator; agent-written proposals wait in reports until promoted. SCOPE a `[playbook]` proposal TIGHTLY — it earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE — if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), do NOT propose it as a recipe; propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything already covered by an existing rule. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately — miscategorizing them bloats the playbook and dilutes the load-bearing lines. -The test is COMPLETE only when the ACTUAL end-to-end user action the task is about has EXECUTED successfully in the app and you've captured proof of its terminal success state — NOT a precursor or a partial step. The repeated failure is stopping SHORT: Rango declared a swap-plugin change tested at the QUOTE; a quote is NOT an executed swap. Per-action bar: a SWAP is done at the executed-swap success scene (e.g. "Congratulations" / order submitted), NEVER at the quote; a SEND at the broadcast/confirmation screen, not an address entered; a feature at its real user-visible outcome, not "it builds" / "the plugin loaded". (Exception: when the task's deliverable IS the precursor — e.g. the buy-quote smoke test exists to render a quote — then that's the bar. Identify the actual user-facing outcome and drive to IT.) -ALL prerequisites to reach that terminal state are MANDATORY and not skippable — and the FIRST, most-skipped one is finishing the IMPLEMENTATION itself: complete EVERY code change across ALL required repos to make the feature integrated and actually RUNNABLE in the sim — the core/dep change AND the gui-side wiring it needs (plugin init / apiKey, provider/plugin registration, imports, config). A partial implementation that never makes the app fully runnable is the upstream failure here: agents do part of the work and stop before the feature even loads, let alone executes. Do NOT stop at "core change written" or "it compiles" — wire it all the way into the gui so the app RUNS with the feature. THEN: link the modified dep into core+gui (`gui-dependency-integration`), build, fund or switch accounts (`funded-test-accounts`), force the provider (`force-swap-provider-locally` — the local corePlugins hack, NOT in-app Exchange Settings), and execute. "The task isn't complete until you can execute a successful swap in the sim, so all parts are necessary including core/gui." Drive through every step EAGERLY; do not declare done, and do not block, until the real action has actually run — unless you hit a genuine precondition the slot truly cannot satisfy (real funds/KYC/finality), in which case capture what you have and `blocked = Yes` with the specific precondition. -CEILING (so you don't over-grind once you're actually there): the bar is the IN-APP success state, not EXTERNAL finality — once the success scene shows and you've captured proof, you are DONE; do NOT then wait for on-chain settlement / full balance sync / provider-side completion (minutes-to-never, out of scope). Run each maestro flow as a SINGLE bounded in-turn call (`timeout maestro test `), never backgrounded-and-polled (re-pays the ~2-min driver startup; it's the `blocking-in-turn-waits` footgun); bound every `extendedWaitUntil`. -A change with a USER-VISIBLE surface (a new row, badge, spinner, empty state, error copy, layout shift) is NEVER finalized on logic evidence alone. When the natural trigger cannot be reproduced on the sim — the window is too fast, the state needs a cold engine, the condition depends on a remote/timing precondition the slot cannot hold — you still owe the PIXELS: HACK-VERIFY it. Force the state with a TEMPORARY UNCOMMITTED edit in the worktree (hard-code the branch condition true, stub the slow call, pin the state field), drive to the screen, capture the frame, then REVERT the hack and prove the tree is clean (`git status --porcelain` empty for that file) before any commit. This is `build-the-test-harness` applied to rendering: you are forcing the INPUT state so the REAL component renders, never faking the output pixels — a mocked-up image or a design comp is not evidence. Unit tests that assert the element renders are complementary, not a substitute: they prove the condition wires up, not that the thing looks right on the device. LABEL THE RESULT so the provenance is never lost: save the frame as `/tmp/agent-proof---HACKED-.png` (the literal token `HACKED`), which makes `pr-attach-screenshots.sh` caption it 🪓 HACK-FORCED and banner the comment — the attach step requires `--hack-note` with a one-line description of the exact hack (pr-create `attach-test-evidence`), so write that line when you make the hack and reuse it; state in the report's Testing section which frames were hack-forced, what the hack was, and that it was reverted. A RE-TEST after any code change lands at a new head sha, which retires the previous frames: carry forward the ones the change left true and retire the ones it invalidated (pr-create `evidence-dispositions`), so a run never leaves pixels standing that its own fix made false. What a hack-verified frame does and does not buy: it PROVES the rendering (layout, copy, colors, no reflow); it does NOT prove the trigger fires in production, so the trigger still needs its own evidence (a unit test, a log line, a code-path argument) and stays named in `verify_blockers`. -The milestone screenshots you capture per `test-drives-the-real-action` are PR EVIDENCE, not throwaways — they get attached to the PR (by `/pr-create`'s `attach-test-evidence`). Capture them deliberately: MULTIPLE if needed to actually illustrate the change (typically: the feature enabled/visible, the action in progress, the terminal success state — e.g. provider-in-settings → quote → confirm → success scene). Save as `/tmp/agent-proof---.png` where `NN` orders them and `slug` is a short human-readable description (`01-provider-enabled`, `02-sonic-quote`, `03-swap-success`) — the slug becomes the image caption on the PR, so write it for a reviewer, not for yourself. One screenshot is enough only when the change is fully visible in a single frame. A frame captured under a forced state carries the `HACKED` token in its filename per `hack-verify-visual-changes` — never strip it to make the evidence look stronger. -VERIFY THE PIXELS before attaching: Read each captured PNG and confirm it actually renders the scene its slug claims. A springboard/home screen, a red error screen, a lock screen, or a blank frame is NOT proof of anything — recapture, never attach. A FAILED read is also a fail: if your own Read of the image errors or returns "could not be processed / removed", you have NOT pixel-verified it — never label such a screenshot "pixel-verified", never attach it, and never credit `Simulator` on it; recapture until a Read SUCCEEDS and shows the claimed terminal-success scene. "I drove the flow so the proof must be right" is precisely the trap (the maestro/MCP driver can observe a different device than simctl photographed); the successful read of a non-error scene is the only thing that substantiates the credit. CAPTURE FROM THE DEVICE YOU DROVE: when the maestro MCP drove the flow, capture via the MCP's own take_screenshot (it renders the daemon's bound device); when the maestro CLI drove it, `simctl io` against the SAME `--device` UDID. A parallel `simctl io` against a different device than the one maestro actually drove photographs the wrong simulator and produces sincere-but-false evidence. -For tasks needing a FUNDED asset (swaps/sends/sync-observation), the test-account ROSTER is the LOCAL-ONLY file `~/.config/edge-secrets/test-accounts.json`: roles `agent` (the agents' own account and **the default YOLO login**; 2FA ON, password + OTP key in its `credsFile`, so its 2FA is never a user-only-credential wall), `primary` (heavily funded but cluttered with leftover assets), `qa-a`, `qa-b` (region California/USA), each with username, PIN, and notes. Refer to accounts BY ROLE in anything synced, committed, or posted; usernames and PINs never leave that file. This roster is EXHAUSTIVE and the scope of any account search: the sim also contains many junk/leftover accounts — do NOT trawl beyond the roster. **Test on the `agent` account and acquire assets by SWAPPING on it**: (1) the agent account already holds the asset → use it; (2) otherwise swap into it on the agent account from a wallet it holds (creating the destination wallet per `create-missing-destination-wallet`), driven to the success scene like any tested swap; if no provider quotes the pair, enable every provider for the requote, worktree-locally only: undo any `force-swap-provider-locally` edit, and set any `*_INIT` that the worktree env.json has as `false` to `{}`. NEVER send funds from another roster account to the agent account: its balance grows only through swaps it executes. Size each acquiring swap per the minimum-viable-amounts rule (binding floor + 10-20% buffer); what it buys stays in the agent account for later runs. (3) Only when the agent account holds nothing that clears any provider floor for a route to the asset, run the test on the roster account that holds the asset and say so in the run report; "no test account holds X" is NOT a valid conclusion until each ROSTER account was actually checked. Move a test off the agent account otherwise ONLY when it needs another account's own state (qa-b's region, a per-account setting or history the task names). (`YOLO_*` auto-login starts every run on the `agent` account and re-asserts it on each relaunch.) **Switch accounts by EDITING the worktree's `env.json`** (`YOLO_USERNAME`/`YOLO_PIN` → target roster account, then `simctl terminate` + `launch`; auto-login does the rest) — NOT by driving the in-app account dropdown, which is slow and fumble-prone. Never brute-force a PIN prompt (exponential lockout) — look it up in the roster. -If you ALREADY hold a funded, provider-supported pair — both wallets present and the source funded above the floor (e.g. funded BTC + an existing FTM wallet → BTC→FTM on SideShift) — that swap is EXECUTABLE NOW and you MUST drive it to terminal success per `test-drives-the-real-action` (quote → confirm slider → broadcast/success scene). Do NOT abandon a held executable pair for a slower-settling or riskier alternative (a slower new-wallet path, a different target asset, etc.) — pick the executable pair you have and finish it. A slow off-chain settle is NOT a reason to abandon: the bar is the in-app success scene, not on-chain finality (`test-drives-the-real-action` CEILING). NEVER report `blocked = Yes` with reason "blocked on funding / no executable pair" while any roster account holds a funded, provider-supported wallet you have not driven to completion — that conclusion is invalid until you have actually attempted the held pair to the confirm slider. Sanctioned majors (BTC/ETH/USDC and similar, supported by nearly every provider) are spendable funding sources; a $2k+ funded roster account is never "blocked on funding". Falling back to direct-API / boot-init verification (sim-testing-playbook Fabric-SIGABRT entry) is permitted ONLY after a genuine funded attempt on a held executable pair is interrupted by the build crash — not as a first resort and not while an untried funded pair exists. SIZE every value-moving action (swap, send, sweep) per the playbook's minimum-viable-amounts rule: discover the binding floor first (provider pair minimum / dust limit), then move the smallest amount that clears it with a 10-20% buffer — never a round default like $20, and never repeat a successful value-moving action for extra evidence. -When a test SENDS to another wallet (a same-asset wallet on another roster account, a second wallet on the same account, a swap/transfer/sweep target) and that destination wallet does not exist, CREATE IT and continue: in whichever roster account the test needs it, via Add Wallet (or Add / Edit Tokens for a token). A receive-side wallet needs no funds, so the roster-search-first step of `funded-test-accounts` does not apply to it. "No destination wallet", "the receiving account has no X wallet", or "would need to create a wallet" is NEVER a blocker, a concession, or a reason to switch to a weaker verification; the completion judge and blocker validator deny it on sight. Wallet creation is a supported test path (playbook). -In a watcher slot (`$AGENT_SIM_UDID` set), resolve your simulator ONLY via `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"` — NEVER by `--runtime`/`--device`. Raw `simctl ... booted` is hook-blocked in slot sessions (multiple sims boot concurrently; `booted` is ambiguous) — pass `$AGENT_SIM_UDID` explicitly to every ad-hoc simctl call, and select that device in the maestro MCP before driving. By-name resolution targets the SHARED MASTER sim ("iPhone 16 Pro Max"); running builds/maestro on the master pollutes the golden image every clone is cut from. `select-ios-sim.sh` now refuses by-name in slot mode (override: `--allow-master`). Your slot clone DOES carry the Edge app + the logged-in test account (APFS copy-on-write from the master) once booted — note `get_app_container` returns NOTHING on a SHUT/never-booted clone, a FALSE negative; boot first (the scripts do) and trust the clone. Do NOT trigger a fresh rebuild on that false negative (it wastes minutes AND wipes the cloned login state → the onboarding screen). DISTINCT and LEGITIMATE: `ios-rn-build.sh` will itself force a rebuild when the cached app's NATIVE build drifted from the worktree (its `ios/Podfile.lock` stamp no longer matches — e.g. a slot image baked before develop's reanimated-4/new-arch migration, which otherwise throws at render on current-develop JS). That is the script self-healing a stale image, NOT the false-negative footgun — let it rebuild (a reinstall keeps the data container, so login survives); do not pass `--skip-install` to dodge it. A JS-only change leaves Podfile.lock identical and still hits the fast path. If `$AGENT_SIM_UDID` is set but `select-ios-sim --accept-udid` HARD-FAILS ("not found") — you are a RESUMED session whose slot sim was recycled after completion — you cannot fix it in-process (no self-respawn). Report it and set `blocked = Yes` noting the operator must re-provision via `~/.config/agent-watcher/resume-task.sh --task-gid ` (allocates a fresh slot+sim+port and relaunches you with working env). Do NOT fall back to the master or a by-name sim. +- **Detect readiness against the resource you actually started, not a guessed log line:** `timeout`-bounded `curl` against the Metro you launched on its REAL port (`/status`, then the `index.bundle` URL); never `grep` a logfile whose name or marker you assumed. +Never assume a repo's package manager: repos migrate between npm and yarn. All install/run/pack operations go through the shared dispatcher `~/.cursor/skills/pm.sh`, which detects the lockfile (`package-lock.json`→npm, `yarn.lock`→yarn, both/neither→npm). Companion scripts in this skill already dispatch through it; do not hand-write `npm ...`/`yarn ...` against a repo without checking `pm.sh detect`. +PLATFORM: default to iOS. Provision the iOS sim, run the iOS flow, and credit `iOS Sim` UNLESS the task EXPLICITLY calls out Android (task title/description says Android, the task is tagged Android, or the change is under `android/` only). For an Android-called-out task, run the ANDROID path instead of (or in addition to) iOS: `./gradlew :app:assembleDebug` from the gui worktree's `android/` is the build verification, and a successful APK is the terminal-success signal for a BUILD-ONLY fix (GitHub `pr-checks.yml` does NOT build Android, so the local assembleDebug is the only check that catches these regressions). Credit `Android Sim` and log the attempt via `log-attempt.sh --category test-drive --result success|failed:`. The Android build needs gitignored secrets the node_modules clone does not carry (`android/app/google-services.json`, `EdgeApiKey.java`, `android/app/src/main/assets/edge-core/plugin-bundle.js`, a generated `android/local.properties` with `sdk.dir`); `setup-task-workspace.sh` copies them, and the Android SDK (`ANDROID_HOME`/`ANDROID_SDK_ROOT`) must be present in the env. Run gradle with `--no-daemon` (or a per-slot `GRADLE_USER_HOME`) for parallel-safety; assembleDebug is CPU/RAM-heavy, so do not run many concurrently. A genuine in-app Android drive (AVD + maestro) is a larger path; the build-only check closes the regression gap for build/native fixes. If a task touches BOTH platforms, exercise iOS and credit both. +DEFAULT to physically exercising the change in the running app on the sim. Almost ANY task can be tested in-app: a swap, a send, a settings toggle, an onboarding/account-creation flow, a specific wallet action, a bug repro. `tsc`/jest/build passing is NECESSARY BUT NOT SUFFICIENT: a change is not verified until you have driven the actual changed behavior in the app and seen the expected result, to its TERMINAL success, not a precursor (`test-drives-the-real-action` sets the bar). Do NOT skip the sim test because static analysis "looks right", because the diff is small, or because authoring a flow is effort. Before setting `blocked = Yes` with reason "can't verify / no defensible default" on a bug, repro, or investigation task, you MUST first attempt the most-specific RUNTIME REPRO you can construct: build the relevant flavor (e.g. `ENABLE_MAESTRO_BUILD=true` for test-server flows) and drive the precise flow. "I can only trace it statically" is NOT a blocker. Block only if the repro is genuinely un-runnable here (missing creds/KYC/datastore the slot can't provide). For FUNDS specifically: the ONLY funds blocker is an OBSERVED TRUE LOSS, an attempted swap/send that failed AND lost principal. Fees/slippage NEVER count as loss (budgeted at $15 equivalent per run, per the playbook), and blocked-ness is established by ATTEMPTING, never predicted. +The test is COMPLETE only when the ACTUAL end-to-end user action the task is about has EXECUTED successfully in the app and you've captured proof of its terminal success state, NOT a precursor or a partial step. Per-action bar: a SWAP is done at the executed-swap success scene (e.g. "Congratulations" / order submitted), NEVER at the quote; a SEND at the broadcast/confirmation screen, not an address entered; a feature at its real user-visible outcome, not "it builds" / "the plugin loaded". (Exception: when the task's deliverable IS the precursor, e.g. the buy-quote smoke test exists to render a quote, that is the bar. Identify the actual user-facing outcome and drive to IT.) +ALL prerequisites to reach that terminal state are MANDATORY and not skippable, and the first one is finishing the IMPLEMENTATION itself: complete EVERY code change across ALL required repos so the feature is integrated and actually RUNNABLE in the sim: the core/dep change AND the gui-side wiring it needs (plugin init / apiKey, provider/plugin registration, imports, config). Do NOT stop at "core change written" or "it compiles". THEN: link the modified dep into core+gui (`gui-dependency-integration`), build, fund or switch accounts (`funded-test-accounts`), force the provider (`force-swap-provider-locally`, the local corePlugins edit, NOT in-app Exchange Settings), and execute. Drive through every step EAGERLY; do not declare done, and do not block, until the real action has actually run, unless you hit a genuine precondition the slot truly cannot satisfy (real funds/KYC/finality), in which case capture what you have and `blocked = Yes` with the specific precondition. +CEILING: the bar is the IN-APP success state, not EXTERNAL finality. Once the success scene shows and you've captured proof, you are DONE; do NOT then wait for on-chain settlement, full balance sync, or provider-side completion. Run each flow as a SINGLE bounded in-turn call (`timeout maestro test `), never backgrounded-and-polled (`blocking-in-turn-waits`); bound every `extendedWaitUntil`. +LOG every value-moving action and every test-drive/repro the moment it resolves, via `~/.config/agent-watcher/log-attempt.sh --gid --action "" --result success|failed:|loss:|blocked: --category swap|send|sweep|test-drive|repro`. This attempt-log (`$XDG_STATE_HOME/agent-watcher/attempts/.jsonl`) is the AUTHORITATIVE record of what the run actually attempted: the concession-validation gate reads it to tell a real wall (`loss:`/`failed:`/`blocked:` after an attempt) from a predicted one, on BOTH a formal `--blocked yes` AND a silent DOWNGRADE-finalize (completing or opening a PR without reaching the prescribed in-app success), and the eval reads it as ground truth for testing depth instead of trusting transcript narration. RESULT semantics: `success` = reached terminal success; `failed:` = attempted, no success, principal safe (fees only); `loss:` = attempted, FAILED, principal unrecoverable (the ONLY funds condition that legitimizes a block); `blocked:` = attempted up to a precondition the slot genuinely cannot satisfy (real provider halt, geo-block confirmed by attempt). +The test harness is YOURS to build; its absence is NEVER a blocker. When driving the real behavior needs scaffolding that does not exist yet, CREATE it locally and uncommitted: author a new `.yaml` flow (per `maestro-flows-are-shortcuts`, expected, not exceptional), add a missing `testID` (`testids-over-coordinates`), trim unrelated plugins (`single-asset-plugin-trim`), disable a crashing module, or HARD-CODE the inputs the code path reads: fixtures, seed data, info-server/remote-config payloads, feature-flag state, a forced provider. "No flow exists for this", "the data comes from a remote server I don't control", "there's no fixture", "the feature isn't enabled by default" are NOT blockers and NOT reasons to stop at static analysis: they are scaffolding to BUILD. KEY DISTINCTION: hard-code the INPUTS to REACH and exercise the real logic, never fake the OUTPUT to fabricate a pass. Injecting a disable-map into the store so the REAL `isSpendBrandDisabled` filter runs against controlled data is correct; hard-coding "this brand is hidden" to skip the filter is not: the changed code path must actually execute. LOCAL-ONLY discipline (same as `single-asset-plugin-trim`): this scaffolding must NEVER land in a commit/PR. Revert it before any commit (verify `git status`/`git diff` is clean of it), or rely on it living only in the disposable test build. If you find yourself writing `blocked = Yes` or "could not test because doesn't exist", stop: build the scaffolding and drive the test. - + -A real on-simulator UI test that logs into the pre-provisioned test account, navigates to the Buy tab, requests a $500 quote, and captures a proof screenshot. PASS requires the screenshot to actually render the resolved quote. +| Phase | Reference (under `~/.cursor/skills/build-and-test/`) | Rules it owns | Delivered by the gate at | +|---|---|---|---| +| Step 0a-0c: preflight, sim selection, RN build, linking a gui dependency into the app | `references/build.md` | `preflight-before-build-decisions`, `slot-sim-is-the-clone`, `gui-dependency-integration` | first `slot-preflight.sh`, `select-ios-sim.sh`, `ios-rn-build.sh` or `ios-rn-build-wait.sh` call | +| Step 0d-0f: driving the app | `references/drive.md` | `maestro-flows-are-shortcuts`, `testids-over-coordinates`, `single-asset-plugin-trim`, `force-swap-provider-locally`, `runtime-inspection-via-debugger`, `spaced-pin-taps`, `no-hideKeyboard`, `no-hierarchy-polling-on-buy` | first maestro drive or `capture-buy-quote.sh` call | +| Proof frames, forced visual states, un-runnable assets | `references/evidence.md` | `proof-screenshots-for-pr`, `hack-verify-visual-changes`, `unrunnable-asset-proxy-verification` | first maestro drive or `capture-buy-quote.sh` call | +| Any test that needs a funded asset or moves value (swap, send, sweep) | `references/funding.md` | `funded-test-accounts`, `executable-pair-must-complete`, `create-missing-destination-wallet` | first `log-attempt.sh --category swap\|send\|sweep` (a backstop: read it yourself BEFORE choosing an account or a pair) | +| Steps 1-3: Node/TypeScript, Node, placeholder | this file | n/a | n/a | -**Parallel-session env contract:** when the agent-watcher spawns this session as one of several parallel slots, it exports `$AGENT_SIM_UDID` (the slot's cloned simulator) and `$AGENT_METRO_PORT` (the slot's Metro port) into the shell. The scripts below honor them automatically — `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"` skips name/runtime resolution and trusts the clone, and `ios-rn-build.sh` falls back to `$AGENT_SIM_UDID` / `$AGENT_METRO_PORT` when `--udid` / `--port` are not passed (forwarding a non-8081 port to `react-native run-ios`). On a manual run with neither var set, behavior is unchanged: resolve the iOS 18 sim by name and use Metro 8081. +Sim working knowledge (funding floors, roster switching, feature-enablement gotchas, investigation order) is in `references/sim-testing-playbook.md`; `maestro-flows-are-shortcuts` owns when to read it. -### 0a. Prerequisites (check, install if missing) - -- `xcrun -version` → Xcode CLT -- `maestro --version` → install with `curl -Ls "https://get.maestro.mobile.dev" | bash`, then add `$HOME/.maestro/bin` to PATH. maestro needs JDK 11+; Temurin 17 works. - -### 0b. Resolve + boot the simulator - -There can be multiple "iPhone 16 Pro Max" devices across runtimes. **Only the iOS 18 device holds the test accounts** (the `funded-test-accounts` roster; default login: the `agent` account). The iOS 26.x device does NOT. - -```bash -UDID=$(~/.cursor/skills/build-and-test/scripts/select-ios-sim.sh \ - --runtime "iOS 18" --device "iPhone 16 Pro Max" --boot) -``` - -If the script exits 2 (ambiguous), narrow `--runtime` (e.g. `"iOS 18.6"`). - -### 0c. Build + install + launch the app - -```bash -~/.cursor/skills/build-and-test/scripts/ios-rn-build.sh \ - --udid "$UDID" --bundle-id co.edgesecure.app --detach -~/.cursor/skills/build-and-test/scripts/ios-rn-build-wait.sh --udid "$UDID" # re-run while it exits 7 -``` - -Skips the full RN build when the app is already installed (cached path: seconds; a fresh build is usually just a few minutes — the Hermes prebuilt is prefetched). Pass `--force-rebuild` to always rebuild. - -### 0d. Run the maestro capture - -```bash -~/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh -``` - -Drives `maestro/buy-quote-input.yaml` (login → Buy → $500), then captures via an external simctl screenshot burst — keeping the last frame taken while the app was alive. Retries up to 5 cycles. Writes `/tmp/agent-mvp-buy-quote-screenshot.png` on success. - -### 0e. PASS / FAIL contract - -On capture-buy-quote.sh exit 0, the screenshot must visibly show **USD 500**, a non-empty **Amount BTC**, and the **`1 BTC = USD`** line. Emit: - -``` -build-and-test: PASS (iOS maestro — Buy $500 quote) -screenshot: /tmp/agent-mvp-buy-quote-screenshot.png -``` - -On exit nonzero, emit FAIL with the last 30 lines of the script's output: - -``` -build-and-test: FAIL — Buy $500 quote not captured - -``` - -Return success exit only on PASS. - -### 0f. Critical gotchas baked into the flow (do not "fix" them) - -Edge's RN keypad drops digits tapped too fast → wrong PIN → exponential lockout (465s → 914s → …). Each PIN digit tap in `buy-quote-input.yaml` uses `waitToSettleTimeoutMs`. Never speed it up. If a run logs "Invalid PIN: Account locked for N seconds", wait — do NOT tap. -On this debug build, `hideKeyboard` reliably triggers an RN Fabric text-measure SIGABRT. The flow leaves the keyboard up. Do not add `hideKeyboard` steps. -`assertVisible`/`extendedWaitUntil` traverse the a11y hierarchy on a poll loop, provoking the same Fabric crash on the Buy scene. The flow stops polling once the amount is entered; the capture script uses external simctl screenshots (no hierarchy traversal). - - + Run, in order: @@ -166,13 +86,3 @@ build-and-test: placeholder mode — no commands executed (repo shape not auto-d ``` Return success. - - -Re-run `select-ios-sim.sh` with a more specific `--runtime` (e.g. `"iOS 18.6"`). If still ambiguous, surface the list to the caller and set `blocked = Yes` on the Asana task with the candidate UDIDs and ask which to use. -Run `xcrun simctl shutdown all && xcrun simctl erase ` is destructive — do NOT run it. Set `blocked = Yes` with the boot error. -Re-run step 0b. If it fails twice, set `blocked = Yes`. -Acceptable in --yolo. Watch loop should NOT timeout the iteration during a known cold-build window — that's handled by /one-shot's `iOS prep budget` policy. -Emit FAIL with the maestro tail. Do NOT set `blocked = Yes` unless the failure mode is clearly a true-blocker (e.g. simulator died entirely, app uninstalled). A normal capture exhaustion is a real test FAIL the caller (watch loop) should react to. -Set `blocked = Yes` with the install error and a note about JDK requirement. -Set `blocked = Yes` — the test relies on the roster accounts (default: the `agent` account) being present on the sim image. Re-provisioning is a human step. - diff --git a/.cursor/skills/build-and-test/maestro/buy-quote.yaml b/.cursor/skills/build-and-test/maestro/buy-quote.yaml index 6fcc6ac5..47e00bac 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote.yaml @@ -3,7 +3,7 @@ # Target: the iPhone 16 Pro Max / iOS 18 simulator that holds the pre-provisioned # roster account `qa-b` (see ~/.config/edge-secrets/test-accounts.json: region # California/USA). The flow quotes into a BTC wallet, so confirm the logged-in account has one. The Edge -# app must already be installed (build + install steps are in build-and-test/SKILL.md). +# app must already be installed (build + install steps are in build-and-test/references/build.md). # # Run: maestro test ~/.cursor/skills/build-and-test/maestro/buy-quote.yaml # diff --git a/.cursor/skills/build-and-test/references/build.md b/.cursor/skills/build-and-test/references/build.md new file mode 100644 index 00000000..72e363d5 --- /dev/null +++ b/.cursor/skills/build-and-test/references/build.md @@ -0,0 +1,75 @@ +Governs step 0a-0c of `/build-and-test` (preflight, simulator selection, RN build, linking a gui dependency into the app); the core step map points here. + + + +| Script | Purpose | Exit codes | +|--------|---------|------------| +| `slot-preflight.sh` | Boot the slot sim and print the build plan (`PLAN:`, `INVOKE:`, `WAIT:` lines) | 0 = plan printed | +| `select-ios-sim.sh` | Resolve (and with `--boot`, boot) the simulator; prints the UDID | 2 = ambiguous device/runtime | +| `ios-rn-build.sh` | Build, install and launch the app; `--detach` returns at once | 2 = sim not booted | +| `ios-rn-build-wait.sh` | Wait on a detached build in bounded chunks | 0/1/2 = the build's own result; 7 = still running, re-run it; 3 = stalled (no log output for 10 min) and already killed: read the printed log tail, fix, start a fresh `--detach`; 4 = no build to wait on: start one | + + + +START the sim-testing phase with ONE call: `~/.cursor/skills/build-and-test/scripts/slot-preflight.sh` (defaults to `$AGENT_SIM_UDID`/`$AGENT_METRO_PORT`; pass `--repo ` when cwd is not the gui repo). It boots the sim if needed and answers deterministically: is Metro mine or squatted, is the app installed, does the installed native side match the worktree (`.agent-native-build-stamp` vs `ios/Podfile.lock`), are node_modules present. OBEY its final `PLAN:` line (`ready` = drive now, no build; `js-only`/`install`; or `full-rebuild`) and run the exact `INVOKE:` command it prints (verbatim; no flag-guessing), then its `WAIT:` command when it prints one. Do NOT re-derive any of its checks manually, and do NOT start a second Metro when it reports one running. +In a watcher slot (`$AGENT_SIM_UDID` set), resolve your simulator ONLY via `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"`, NEVER by `--runtime`/`--device`. Raw `simctl ... booted` is hook-blocked in slot sessions (several sims boot concurrently, so `booted` is ambiguous): pass `$AGENT_SIM_UDID` explicitly to every ad-hoc simctl call, and select that device in the maestro MCP before driving. By-name resolution targets the SHARED MASTER sim ("iPhone 16 Pro Max"); builds or drives on the master pollute the golden image every clone is cut from. `select-ios-sim.sh` refuses by-name in slot mode (override: `--allow-master`). +- **Trust the clone.** Once booted it carries the Edge app and the logged-in test account (APFS copy-on-write from the master). `get_app_container` returns NOTHING on a SHUT/never-booted clone, a FALSE negative: boot first (the scripts do). Do NOT trigger a fresh rebuild on that false negative: it wastes minutes AND wipes the cloned login state. +- **Let the script's own rebuild run.** `ios-rn-build.sh` forces a rebuild when the cached app's NATIVE build drifted from the worktree (its `ios/Podfile.lock` stamp no longer matches). A reinstall keeps the data container, so login survives; do not pass `--skip-install` to dodge it. A JS-only change leaves Podfile.lock identical and takes the fast path. +- **Recycled sim.** If `$AGENT_SIM_UDID` is set but `select-ios-sim --accept-udid` HARD-FAILS ("not found"), you are a RESUMED session whose slot sim was recycled; you cannot fix it in-process (no self-respawn). Report it and set `blocked = Yes`, noting the operator must re-provision via `~/.config/agent-watcher/resume-task.sh --task-gid `. Do NOT fall back to the master or a by-name sim. +A change to an EdgeApp gui DEPENDENCY is NOT fully tested until it runs in the app; its own `tsc`/jest passing is necessary but NOT sufficient. Gui dependencies = the Edge-owned repos `edge-react-gui` consumes: `edge-core-js`, `edge-currency-accountbased`, `edge-currency-plugins`, `edge-exchange-plugins`, `edge-login-ui-rn`, `edge-currency-monero`, `react-native-piratechain`, `react-native-zcash`, `react-native-zano`. When the repo under test is one of these, after its own checks you MUST also run the gui integration test, autonomously (NO prompting): +1. **Co-located gui worktree:** ensure one exists; create via `~/.config/agent-watcher/setup-task-workspace.sh --task-gid --repo edge-react-gui` if absent (sibling of the dep worktree under `~/git/.agent-worktrees//`, so updot can find it). +2. **Link the MODIFIED dep into the app. The mechanism, and whether you flip any `DEBUG_*` flag, is YOUR per-task call** (it depends on what the task changed and how you want to verify it). Run repo scripts with each repo's package manager (`lockfile-driven-pm`). The toolbox: + - **`updot`: bakes the built dep into the gui's `node_modules`.** Works for ANY dep, no dev-server, no runtime race: the safe default for headless/automated runs. ` updot ` then the gui's `prepare` (npm form: `npm run updot -- && npm run prepare`; add `prepare.ios` for native-module deps), then rebuild. The dep's `DEBUG_*` flag stays FALSE (you baked it in). + - **`DEBUG_` flag + the dep's live webpack dev-server, webview-plugin deps only** (`DEBUG_ACCOUNTBASED`:8082, `DEBUG_EXCHANGES`:8083, `DEBUG_CURRENCY_PLUGINS`:8084, `DEBUG_PLUGINS`:8101; these ports are HARDCODED in each dep package's `debugUri` and are HOST-GLOBAL). Set the flag TRUE in the gui's `env.json` AND run the dep's `yarn start`/`npm start` (webpack serve) backgrounded for the test; the webview loads the local bundle live (sim reaches host localhost), no gui rebuild. Pick this when live iteration helps; if it flakes (dev-server unreachable, ATS/cleartext, recompile race) fall back to updot. + - **Parallel-slot port rule:** a `DEBUG_` dev-server port is a SINGLE-OCCUPANT host resource; only ONE slot can serve a given dep at a time. Before starting the dev-server, check the port is free: `lsof -nP -iTCP: -sTCP:LISTEN`; if another slot holds it, use updot. Your slot's Metro runs on `$AGENT_METRO_PORT` (base **8181**), deliberately OUTSIDE the 808x DEBUG range; do NOT pass a `--port` that drags Metro back into 808x. When in doubt in a parallel slot, prefer updot: it has no shared port. + - **`DEBUG_EXCHANGES` crash-loop trap:** the gui's `allowDebugging` flag (which permits the cleartext localhost load) is OR-gated on `DEBUG_ACCOUNTBASED || DEBUG_CORE || DEBUG_CURRENCY_PLUGINS || DEBUG_PLUGINS`. **`DEBUG_EXCHANGES` is NOT in that set**, so enabling it ALONE crash-loops the app. Co-enable one that IS (e.g. `DEBUG_ACCOUNTBASED`); that drags in its 8082 dev-server, so plan ports per the rule above. Swap/exchange plugin code runs in **edge-core-js's webview context, not the Metro bundle**: serve patched dep code via the dev-server (or `updot`-bake it); do NOT sync patched `lib/` into `node_modules` expecting Metro to bundle it. + - **`edge-core-js`: prefer `updot`, avoid `DEBUG_CORE`.** `DEBUG_CORE` loads the WHOLE core from hardcoded `http://localhost:8080/`: it races init, is cleartext/ATS-sensitive, and any hiccup takes the entire app down. + Only link the dep(s) THIS task modifies; leave every other dep's `DEBUG_*` at its env.json default. Keep flags consistent with what you actually linked: a `DEBUG_*` left true with no dev-server running breaks that dep. +3. **Login:** the test account auto-logs-in via the `YOLO_*` env knobs (set by workspace init to the roster's `agent` account from `~/.config/edge-secrets/test-accounts.json`, consumed in `LoginScene.tsx`, pinned by `setup-task-workspace.sh` on every worktree's env.json copy). Keep them set so the run reaches the logged-in app; when the change is to `edge-login-ui-rn` specifically, these are the lever for exercising the login flow: adjust only if the change requires driving the login UI differently. +4. **Make the gui-side changes the feature NEEDS to run, then run the gui path (step 0)** against that build. A dep change almost always needs gui-side wiring to function: plugin init / apiKey, provider/plugin registration, imports, config. Those gui changes are PART OF THE WORK: complete ALL of them (and commit on the gui worktree's branch) so the app is fully runnable with the feature, autonomously, do NOT prompt. If the feature doesn't load/run in the app yet, the implementation is NOT done. +PASS requires the app test to pass with the dep change linked AND the actual feature exercised to its terminal success (`test-drives-the-real-action`). A dep whose unit checks pass but that isn't fully wired into a runnable app, or that runs but whose real action was never executed, is a FAIL. +SCOPE DOES NOT EXEMPT THE TEST: a task that scopes its deliverable to the dep repo, calls itself a prototype, or explicitly defers PRODUCTION gui integration to follow-up work still gets THIS integration test. The wiring in steps 1-4 is TEST SCAFFOLDING in the task's gui WORKTREE (plugin registration, env.json keys, dep linking), not an unrequested production change; nothing lands in the gui repo unless the task asks for it. Likewise "the plugin is unvetted prototype code, a real swap through it is irreversible" is NOT a blocker: vetting it with a small sanctioned-roster swap is what this test exists to do (see one-shot `yolo-true-blockers` carve-out). + + + + +A real on-simulator UI test that logs into the pre-provisioned test account, navigates to the Buy tab, requests a $500 quote, and captures a proof screenshot. Steps 0d-0f (the drive and the PASS/FAIL contract) are in `references/drive.md`. + +**Parallel-session env contract:** when the agent-watcher spawns this session as one of several parallel slots, it exports `$AGENT_SIM_UDID` (the slot's cloned simulator) and `$AGENT_METRO_PORT` (the slot's Metro port). The scripts honor them automatically: `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"` skips name/runtime resolution and trusts the clone, and `ios-rn-build.sh` falls back to `$AGENT_SIM_UDID` / `$AGENT_METRO_PORT` when `--udid` / `--port` are not passed (forwarding a non-8081 port to `react-native run-ios`). In a slot, `preflight-before-build-decisions` replaces 0b-0c: run the preflight and obey its plan. On a manual run with neither var set: resolve the iOS 18 sim by name and use Metro 8081, as below. + +### 0a. Prerequisites (check, install if missing) + +- `xcrun -version` → Xcode CLT +- `maestro --version` → install with `curl -Ls "https://get.maestro.mobile.dev" | bash`, then add `$HOME/.maestro/bin` to PATH. maestro needs JDK 11+; Temurin 17 works. + +### 0b. Resolve + boot the simulator + +There can be multiple "iPhone 16 Pro Max" devices across runtimes. **Only the iOS 18 device holds the test accounts** (the `funded-test-accounts` roster; default login: the `agent` account). The iOS 26.x device does NOT. + +```bash +UDID=$(~/.cursor/skills/build-and-test/scripts/select-ios-sim.sh \ + --runtime "iOS 18" --device "iPhone 16 Pro Max" --boot) +``` + +If the script exits 2 (ambiguous), narrow `--runtime` (e.g. `"iOS 18.6"`). + +### 0c. Build + install + launch the app + +```bash +~/.cursor/skills/build-and-test/scripts/ios-rn-build.sh \ + --udid "$UDID" --bundle-id co.edgesecure.app --detach +~/.cursor/skills/build-and-test/scripts/ios-rn-build-wait.sh --udid "$UDID" # re-run while it exits 7 +``` + +Skips the full RN build when the app is already installed (cached path: seconds; a fresh build is usually a few minutes, the Hermes prebuilt is prefetched). Pass `--force-rebuild` to always rebuild. + + + + +Re-run `select-ios-sim.sh` with a more specific `--runtime` (e.g. `"iOS 18.6"`). If still ambiguous, surface the list to the caller and set `blocked = Yes` on the Asana task with the candidate UDIDs and ask which to use. +`xcrun simctl shutdown all && xcrun simctl erase ` is destructive: do NOT run it. Set `blocked = Yes` with the boot error. +Re-run step 0b. If it fails twice, set `blocked = Yes`. +Acceptable in --yolo. The watch loop should NOT timeout the iteration during a known cold-build window; /one-shot's `iOS prep budget` policy handles that. +Set `blocked = Yes` with the install error and a note about the JDK requirement. +Set `blocked = Yes`: the test relies on the roster accounts (default: the `agent` account) being present on the sim image. Re-provisioning is a human step. + diff --git a/.cursor/skills/build-and-test/references/drive.md b/.cursor/skills/build-and-test/references/drive.md new file mode 100644 index 00000000..3563bd1e --- /dev/null +++ b/.cursor/skills/build-and-test/references/drive.md @@ -0,0 +1,68 @@ +Governs step 0d-0f of `/build-and-test` (driving the app on the sim: the flow library, selectors, the corePlugins levers, runtime inspection); the core step map points here. Proof-frame rules are in `references/evidence.md`, funding rules in `references/funding.md`. + + + +| Script | Purpose | Exit codes | +|--------|---------|------------| +| `capture-buy-quote.sh [--flow ]` | Run one flow on the maestro CLI (device and driver port pinned from the slot env), then capture via an external simctl screenshot burst; up to 5 retry cycles | 0 = captured; nonzero = FAIL (step 0e) | + + + +Before the sim-test phase, READ `~/.cursor/skills/build-and-test/references/sim-testing-playbook.md`: it holds the working knowledge (funding floors, account roster/switching, feature-enablement gotchas, investigation order) that otherwise gets re-learned every run. +- **Compose, don't re-derive.** Parameterized subflows live in `~/.cursor/skills/build-and-test/maestro/common/` (`login-if-needed`, `dismiss-startup-modals`, `select-swap-pair`, `confirm-slider`; the slider gesture is SOLVED there, never re-derive it). Copy the subflows you need next to your task flow and `runFlow` them; author NEW task-specific `.yaml` liberally for what the task actually changed. Task flows stay LOCAL (`.syncignore`d from the agent repo; never committed to the gui repo, whose `maestro/` is the heavyweight verification suite, reference-only for selectors). What DOES get committed to the gui: missing `testID`s, per `testids-over-coordinates`. +- **Exploration on the MCP, proof on the CLI.** For exploring screens/selectors use the maestro MCP tools (persistent driver, no ~2-min `maestro test` startup per probe). The MCP daemon binds to one device and DRIFTS: `maestro-mcp-wrapper.sh` start-pins it, but the pin does not survive slot-sim relaunch cycling, and the daemon can rebind to another slot's sim mid-run. So: (1) re-verify its bound device (list_devices + an inspect matching your app's expected state) at the START OF EVERY observation block; (2) ALL PROOF evidence comes from the maestro CLI, canonical slot invocation `maestro --device "$AGENT_SIM_UDID" --driver-host-port $((AGENT_METRO_PORT + 1000)) test ` (the per-slot driver port keeps parallel slots' iOS drivers apart; hook-enforced; the MCP daemon uses +2000 via its wrapper), plus `simctl io` against that SAME UDID. NEVER an MCP screenshot, and NEVER a `simctl io` against a different device than the CLI drove. If a verify shows the MCP on the wrong device, drop to CLI for the rest of the run. +- **One proof flow.** For the repeatable PROOF run compose ONE yaml flow and run it once via `capture-buy-quote.sh --flow `; that run produces the PR evidence screenshots. +- **Proposals, never direct edits.** The playbook and the flow library are operator-curated: entries may be trusted without re-verification because the operator reviewed and promoted each one. Task runs never promote; proposals wait in run reports for the eval's manually-triggered flow-consolidation pass. In the report's Dev Notes & Gotchas section: + - `[playbook]`: durable knowledge, one bullet. + - `[flow]`: a NEW reusable drive sequence: name, params, one-line purpose, and the FULL yaml EMBEDDED as a fenced block (worktrees are pruned on retention; the report attachment is the durable copy). + - `[flow-update]`: a change to an EXISTING library flow (genericize, new param, split into subflows): name the flow, the change, and the compatibility argument. New params MUST default to current behavior so existing callers are unaffected, and a rename/split must say so explicitly (callers get grepped at promotion). +- **Scope a `[playbook]` proposal tightly.** It earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE: if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING instead. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything an existing rule already covers. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately. +Scoped exception to `no-mutation`, test-infrastructure only. TESTIDS FIRST, COORDINATES LAST. When a flow needs to drive an element that has no stable selector (text match fails and no `testID` exists), ADD the missing `testID` prop to that component in the gui worktree and drive via it. A `testID` is a JS-only prop: Metro reload picks it up in seconds (no native rebuild), so adding one is cheaper than a single round of coordinate trial-and-error, and it de-brittles the suite for every future run. Coordinate taps are permitted ONLY for surfaces you cannot edit (system dialogs, native pickers, third-party views that don't forward `testID`) or when a reload would destroy unrecoverable in-flight app state, and any coordinate tap that survives into the PROOF flow must be called out in the run report with why a testID was not possible. +- **Commit discipline:** commit the testID additions as a SEPARATE commit, distinct from any feature commit; change ONLY `testID` props, never component logic; update the flow selector(s) to use them. +- **Message names the surface:** subject `test: add testIDs to ` (e.g. `test: add testIDs to ExchangeScene swap pills`), with every added id listed in the body. Never a generic subject: identical subjects across runs make these commits indistinguishable when a human cherry-picks between branches. +- **Where the commit lands:** always THE TASK'S SINGLE GUI PR, the one whose test surfaced the need; NEVER a separate testID-only PR. A task has at most ONE gui PR: for a gui-feature task the testID commit rides that feature PR; for a DEP-repo task whose test drives the gui, the testIDs go in the task's ONE gui integration PR (if the testIDs are the only gui change, that PR IS the task's gui PR), on the SAME `/` gui branch the run already provisioned. +If no selector was missing, this rule is a no-op. +OPTIMIZATION (optional, LOCAL-ONLY, never committed). When the task targets a SINGLE asset and the test needs to drive that asset's wallet, you MAY temporarily comment out the unrelated currency plugins in the gui worktree's `src/util/corePlugins.ts` (the `currencyPlugins` map), keeping the plugin(s) the task needs, to cut app load/init time. This is a test-harness speedup ONLY: it must NEVER land in a commit or PR. Revert it before any commit, or rely on it living only in the throwaway test build; if you commit after trimming, verify `git status`/`git diff` does NOT include `corePlugins.ts`. Skip entirely for multi-asset tasks or tasks that don't drive a wallet. +To FORCE a specific swap provider for a test (so the engine routes through it instead of a competitor), edit the gui worktree's `src/util/corePlugins.ts` `swapPlugins` map and set every OTHER provider to `false`, leaving only the target's `*_INIT` truthy. LOCAL-ONLY, same never-committed discipline as `single-asset-plugin-trim`. Do NOT force a provider by toggling the in-app **Settings → Exchange Settings**: that state is ACCOUNT-SYNCED, so on a shared roster account it thrashes against every parallel session and persists to the next run. Use Exchange Settings only to READ/diagnose why a provider is absent, never to set routing. (`Preferred`/`preferPluginId` also do not pin: the engine reverts to best-rate in about 60s.) See sim-testing-playbook "Feature-enablement check". +When verification needs RUNTIME state from the running app (why a check evaluates false, the actual value of a variable, which code path executed), use the `/debugger` skill (`~/.cursor/skills/debugger/SKILL.md`); do NOT hand-roll a CDP/WebSocket attach. It sets a `file:line` breakpoint over Metro's Hermes inspector and reports the call stack + locals, and it is slot-aware: `check-metro.sh` and `cdp-attach.js` default to `$AGENT_METRO_PORT`. Static questions (where is X defined) stay grep/read. + + + + +### 0d. Run the maestro capture + +```bash +~/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh +``` + +Drives `maestro/buy-quote-input.yaml` (login → Buy → $500), then captures via an external simctl screenshot burst, keeping the last frame taken while the app was alive. Retries up to 5 cycles. Writes `/tmp/agent-mvp-buy-quote-screenshot.png` on success. + +### 0e. PASS / FAIL contract + +On capture-buy-quote.sh exit 0, the screenshot must visibly show **USD 500**, a non-empty **Amount BTC**, and the **`1 BTC = USD`** line. Emit: + +``` +build-and-test: PASS (iOS maestro — Buy $500 quote) +screenshot: /tmp/agent-mvp-buy-quote-screenshot.png +``` + +On exit nonzero, emit FAIL with the last 30 lines of the script's output: + +``` +build-and-test: FAIL — Buy $500 quote not captured + +``` + +Return success exit only on PASS. + +### 0f. Critical gotchas baked into the flow (do not "fix" them) + +Edge's RN keypad drops digits tapped too fast → wrong PIN → exponential lockout (465s → 914s → …). Each PIN digit tap in `buy-quote-input.yaml` uses `waitToSettleTimeoutMs`. Never speed it up. If a run logs "Invalid PIN: Account locked for N seconds", wait; do NOT tap. +On this debug build, `hideKeyboard` reliably triggers an RN Fabric text-measure SIGABRT. The flow leaves the keyboard up. Do not add `hideKeyboard` steps. +`assertVisible`/`extendedWaitUntil` traverse the a11y hierarchy on a poll loop, provoking the same Fabric crash on the Buy scene. The flow stops polling once the amount is entered; the capture script uses external simctl screenshots (no hierarchy traversal). + + + + +Emit FAIL with the maestro tail. Do NOT set `blocked = Yes` unless the failure mode is clearly a true-blocker (e.g. simulator died entirely, app uninstalled). A normal capture exhaustion is a real test FAIL the caller (watch loop) should react to. + diff --git a/.cursor/skills/build-and-test/references/evidence.md b/.cursor/skills/build-and-test/references/evidence.md new file mode 100644 index 00000000..77852c39 --- /dev/null +++ b/.cursor/skills/build-and-test/references/evidence.md @@ -0,0 +1,12 @@ +Governs the proof a `/build-and-test` drive produces: PR screenshots, forced visual states, and the bar for an asset that cannot run on the sim; the core step map points here. + + +The milestone screenshots you capture per `test-drives-the-real-action` are PR EVIDENCE: they get attached to the PR (by `/pr-create`'s `attach-test-evidence`). Capture MULTIPLE if needed to illustrate the change (typically: the feature enabled/visible, the action in progress, the terminal success state, e.g. provider-in-settings → quote → confirm → success scene). Save as `/tmp/agent-proof---.png` where `NN` orders them and `slug` is a short human-readable description (`01-provider-enabled`, `02-sonic-quote`, `03-swap-success`); the slug becomes the image caption on the PR, so write it for a reviewer. One screenshot is enough only when the change is fully visible in a single frame. A frame captured under a forced state carries the `HACKED` token in its filename per `hack-verify-visual-changes`; never strip it. +- **Verify the pixels before attaching.** Read each captured PNG and confirm it renders the scene its slug claims. A springboard/home screen, a red error screen, a lock screen, or a blank frame is NOT proof: recapture, never attach. A FAILED read is also a fail: if your Read of the image errors or returns "could not be processed / removed", you have NOT pixel-verified it; never label it "pixel-verified", never attach it, and never credit `Simulator` on it. Recapture until a Read SUCCEEDS and shows the claimed terminal-success scene. Having driven the flow does not substitute for the read. +- **Capture from the device you drove.** When the maestro MCP drove the flow, capture via the MCP's own take_screenshot (it renders the daemon's bound device); when the maestro CLI drove it, `simctl io` against the SAME `--device` UDID. A `simctl io` against a different device than the one maestro drove photographs the wrong simulator. +A change with a USER-VISIBLE surface (a new row, badge, spinner, empty state, error copy, layout shift) is NEVER finalized on logic evidence alone. When the natural trigger cannot be reproduced on the sim (the window is too fast, the state needs a cold engine, the condition depends on a remote/timing precondition the slot cannot hold), you still owe the PIXELS: HACK-VERIFY it. Force the state with a TEMPORARY UNCOMMITTED edit in the worktree (hard-code the branch condition true, stub the slow call, pin the state field), drive to the screen, capture the frame, then REVERT the hack and prove the tree is clean (`git status --porcelain` empty for that file) before any commit. This is `build-the-test-harness` applied to rendering: force the INPUT state so the REAL component renders, never fake the output pixels; a mocked-up image or a design comp is not evidence. Unit tests that assert the element renders are complementary, not a substitute. +- **Label the result.** Save the frame as `/tmp/agent-proof---HACKED-.png` (the literal token `HACKED`), which makes `pr-attach-screenshots.sh` caption it 🪓 HACK-FORCED and banner the comment. The attach step requires `--hack-note` with a one-line description of the exact hack (pr-create `attach-test-evidence`): write that line when you make the hack and reuse it. State in the report's Testing section which frames were hack-forced, what the hack was, and that it was reverted. +- **Re-test retires frames.** A re-test after any code change lands at a new head sha, which retires the previous frames: carry forward the ones the change left true and retire the ones it invalidated (pr-create `evidence-dispositions`). +- **What the frame buys.** It PROVES the rendering (layout, copy, colors, no reflow); it does NOT prove the trigger fires in production, so the trigger still needs its own evidence (a unit test, a log line, a code-path argument) and stays named in `verify_blockers`. +When the asset or feature under test genuinely CANNOT be driven in the sim (a default-disabled plugin that crashes the debug build on enable, e.g. `BOTANIX_INIT`; a date-gated change that only activates after a future date; a non-GUI-dependency surface), the bar is NOT "skip the in-app test and declare `verified: not-run`". Do BOTH: (1) drive a PROXY that exercises the SAME mechanism your change routes through and capture proof of it reaching its terminal state (e.g. for a keys-only create-wallet exclusion, a hardcoded-enabled keys-only asset like `bitcoinsv` hits the identical exclusion path; see the sim-testing playbook); and (2) unit-test the gate/branch your change adds (the condition that routes the target asset) so the logic is covered even though the asset itself cannot run. Record both in the Testing section and set `verify_blockers: [precondition]` (asset un-runnable here) with the proxy drive + unit test as the evidence. That combination is a sanctioned PASS; `verified: not-run` with an empty `verify_blockers` is not. When the un-drivable surface is VISUAL, add the hack-verified frame per `hack-verify-visual-changes`. + diff --git a/.cursor/skills/build-and-test/references/funding.md b/.cursor/skills/build-and-test/references/funding.md new file mode 100644 index 00000000..3964087b --- /dev/null +++ b/.cursor/skills/build-and-test/references/funding.md @@ -0,0 +1,12 @@ +Governs any `/build-and-test` drive that needs a funded asset or moves value (swap, send, sweep); the core step map points here. Read it BEFORE choosing an account or a pair. + + +For tasks needing a FUNDED asset (swaps/sends/sync-observation), the test-account ROSTER is the LOCAL-ONLY file `~/.config/edge-secrets/test-accounts.json`: roles `agent` (the agents' own account and **the default YOLO login**; 2FA ON, password + OTP key in its `credsFile`, so its 2FA is never a user-only-credential wall), `primary` (heavily funded but cluttered with leftover assets), `qa-a`, `qa-b` (region California/USA), each with username, PIN, and notes. Refer to accounts BY ROLE in anything synced, committed, or posted; usernames and PINs never leave that file. The roster is EXHAUSTIVE and the scope of any account search: the sim also contains many junk/leftover accounts; do NOT trawl beyond the roster. +**Test on the `agent` account and acquire assets by SWAPPING on it:** +1. The agent account already holds the asset → use it. +2. Otherwise swap into it on the agent account from a wallet it holds (creating the destination wallet per `create-missing-destination-wallet`), driven to the success scene like any tested swap. If no provider quotes the pair, enable every provider for the requote, worktree-locally only: undo any `force-swap-provider-locally` edit, and set any `*_INIT` that the worktree env.json has as `false` to `{}`. NEVER send funds from another roster account to the agent account: its balance grows only through swaps it executes. Size each acquiring swap per the minimum-viable-amounts rule (binding floor + 10-20% buffer); what it buys stays in the agent account for later runs. +3. Only when the agent account holds nothing that clears any provider floor for a route to the asset, run the test on the roster account that holds the asset and say so in the run report. "No test account holds X" is NOT a valid conclusion until each ROSTER account was actually checked. +Move a test off the agent account otherwise ONLY when it needs another account's own state (qa-b's region, a per-account setting or history the task names). `YOLO_*` auto-login starts every run on the `agent` account and re-asserts it on each relaunch. **Switch accounts by EDITING the worktree's `env.json`** (`YOLO_USERNAME`/`YOLO_PIN` → target roster account, then `simctl terminate` + `launch`; auto-login does the rest), NOT by driving the in-app account dropdown. Never brute-force a PIN prompt (exponential lockout): look it up in the roster. +If you ALREADY hold a funded, provider-supported pair (both wallets present and the source funded above the floor, e.g. funded BTC + an existing FTM wallet → BTC→FTM on SideShift), that swap is EXECUTABLE NOW and you MUST drive it to terminal success per `test-drives-the-real-action` (quote → confirm slider → broadcast/success scene). Do NOT abandon a held executable pair for a slower-settling or riskier alternative (a new-wallet path, a different target asset): pick the executable pair you have and finish it. A slow off-chain settle is NOT a reason to abandon: the bar is the in-app success scene, not on-chain finality. NEVER report `blocked = Yes` with reason "blocked on funding / no executable pair" while any roster account holds a funded, provider-supported wallet you have not driven to completion; that conclusion is invalid until you have attempted the held pair to the confirm slider. Sanctioned majors (BTC/ETH/USDC and similar, supported by nearly every provider) are spendable funding sources; a $2k+ funded roster account is never "blocked on funding". Falling back to direct-API / boot-init verification (sim-testing-playbook Fabric-SIGABRT entry) is permitted ONLY after a genuine funded attempt on a held executable pair is interrupted by the build crash, not as a first resort and not while an untried funded pair exists. SIZE every value-moving action (swap, send, sweep) per the playbook's minimum-viable-amounts rule: discover the binding floor first (provider pair minimum / dust limit), then move the smallest amount that clears it with a 10-20% buffer; never a round default like $20, and never repeat a successful value-moving action for extra evidence. +When a test SENDS to another wallet (a same-asset wallet on another roster account, a second wallet on the same account, a swap/transfer/sweep target) and that destination wallet does not exist, CREATE IT and continue: in whichever roster account the test needs it, via Add Wallet (or Add / Edit Tokens for a token). A receive-side wallet needs no funds, so the roster-search-first step of `funded-test-accounts` does not apply to it. "No destination wallet", "the receiving account has no X wallet", or "would need to create a wallet" is NEVER a blocker, a concession, or a reason to switch to a weaker verification; the completion judge and blocker validator deny it on sight. Wallet creation is a supported test path (playbook). + diff --git a/agent-watcher/hooks/require-playbook-before-drive.sh b/agent-watcher/hooks/require-playbook-before-drive.sh index bac74cf1..3ed4ea1a 100755 --- a/agent-watcher/hooks/require-playbook-before-drive.sh +++ b/agent-watcher/hooks/require-playbook-before-drive.sh @@ -17,12 +17,22 @@ # a loop only occurs if the agent refuses the read. The deny message carries the # flow-library index (the old nudge's payload) so composition guidance still # arrives at the drive moment. +# +# PHASE SLICES. /build-and-test is a core plus per-phase references, and a +# maestro drive is the first call of its drive phase, which calls no companion +# script on the MCP and bare-CLI paths. So this hook also requires the drive and +# evidence slices (units `build-and-test:drive`, `build-and-test:evidence`) and +# delivers the missing ones in the deny, through lib/skill-read-gate.sh: same +# markers, same transcript credit, same one round trip as +# require-skill-read-for-scripts.sh, which covers capture-buy-quote.sh. A deny +# that owes both the playbook and slices says so once. set -euo pipefail [ -n "${AGENT_TASK_GID:-}" ] || exit 0 MARKER="/tmp/agent-playbook-read-$AGENT_TASK_GID" INPUT=$(cat) +LIB="$HOME/.config/agent-watcher/hooks/lib" TOOL=$(printf '%s' "$INPUT" | jq -r '.tool_name // empty' 2>/dev/null || true) IS_DRIVE=0 case "$TOOL" in @@ -45,7 +55,6 @@ case "$TOOL" in # hierarchy subcommand, or capture-buy-quote.sh / maestro-mcp-wrapper.sh # executed. `maestro --version`, `ls .../maestro`, and grep/cat of maestro # paths are not drives. - LIB="$HOME/.config/agent-watcher/hooks/lib" [ -f "$LIB/maestro-cmd.sh" ] || exit 0 . "$LIB/maestro-cmd.sh" if [ -n "$(maestro_cmd_segments "$CMD" "$CMD_M" 2>/dev/null || true)" ]; then @@ -55,12 +64,25 @@ case "$TOOL" in esac [ "$IS_DRIVE" = 1 ] || exit 0 +# Slices this drive still owes (empty when the lib is unavailable: fail open). +SLICES="" +if [ -f "$LIB/skill-read-gate.sh" ]; then + export SKILL_READ_KEY="$AGENT_TASK_GID" + . "$LIB/skill-read-gate.sh" + SLICES=$(skill_read_missing build-and-test:drive build-and-test:evidence) + if [ -n "$SLICES" ]; then + TRANSCRIPT=$(printf '%s' "$INPUT" | jq -r '.transcript_path // empty' 2>/dev/null || true) + skill_read_credit_from_transcript "$TRANSCRIPT" $SLICES + SLICES=$(skill_read_missing $SLICES) + fi +fi + # Playbook already read: pass, but once per run inject the working-set check at # the FIRST post-read drive. Salience delivery, not availability — the playbook # is force-read (below) yet its working-set bullet was read-and-missed on an # asset task (HOOD, 2026-08-12). Fires only when corePlugins.ts is untouched; # a trimmed worktree or a non-gui repo sees nothing. Never blocks. -if [ -f "$MARKER" ]; then +if [ -f "$MARKER" ] && [ -z "$SLICES" ]; then NUDGE_FLAG="/tmp/agent-coreplugins-nudge-$AGENT_TASK_GID" [ -f "$NUDGE_FLAG" ] && exit 0 : > "$NUDGE_FLAG" @@ -73,6 +95,14 @@ if [ -f "$MARKER" ]; then exit 0 fi +if [ -f "$MARKER" ]; then + { + echo "BLOCKED: this maestro drive is the first of the /build-and-test drive phase, whose rules are not yet in this session's context. They are delivered below; re-run the same call." + skill_read_deliver $SLICES + } >&2 + exit 2 +fi + PLAYBOOK="$HOME/.cursor/skills/build-and-test/references/sim-testing-playbook.md" cat >&2 <&2 +fi exit 2 diff --git a/agent-watcher/hooks/require-skill-read-for-scripts.sh b/agent-watcher/hooks/require-skill-read-for-scripts.sh index 0179d33d..5be5ad23 100755 --- a/agent-watcher/hooks/require-skill-read-for-scripts.sh +++ b/agent-watcher/hooks/require-skill-read-for-scripts.sh @@ -22,8 +22,11 @@ # it. Without this the split would silently relax every moved rule, which is # how a gate erodes into a formality. The map below is the whole list; a script # with no entry requires nothing from it and stays quiet. /pr-land took the -# same split (2026-09-30, units `pr-land:`); its entries apply in every -# session, since it runs outside orch too. +# same split (2026-09-30, units `pr-land:`), and so did /build-and-test +# (2026-10-01, units `build-and-test:`); their entries apply in every +# session, since both run outside orch too. /build-and-test's drive and evidence +# slices also arrive on every maestro drive, which calls no companion script: +# require-playbook-before-drive.sh owns that vector. # # Markers come from mark-skill-read.sh and # inject-run-context.sh; on the would-block path the transcript is scanned for @@ -47,7 +50,7 @@ # Scope: every session. Orch runs (AGENT_TASK_GID set) key markers by the gid # and get the whole map; any other session keys them sess- (what # mark-skill-read.sh writes there) and gets only the skill-directory rule and -# the /pr-land slices, because the remaining shared-script entries are /one-shot +# the /pr-land and /build-and-test slices, because the remaining shared-script entries are /one-shot # phase slices and orch-only intake/completion steps. No session_id: no-op. Exit 0 allow, exit 2 # block. set -uo pipefail @@ -141,6 +144,10 @@ need '(pr-land-prepare|changelog-union-merge)\.sh([[:space:]]|$)' pr-land:prepar need '(pr-land-automerge|pr-merge-watch|pr-land-merge|force-land-rationale)\.sh([[:space:]]|$)' pr-land:merge need '(pr-land-publish|npm-publish-web|npm-auth-wait|upgrade-dep)\.sh([[:space:]]|$)' pr-land:publish need '(pr-bot-findings-sweep|pr-land-extract-asana-task)\.sh([[:space:]]|$)' pr-land:post-merge +# /build-and-test phase slices, every session. All four build scripts and the +# capture script live in build-and-test/scripts, so the core is already required. +need '(slot-preflight|select-ios-sim|ios-rn-build|ios-rn-build-wait)\.sh([[:space:]]|$)' build-and-test:build +need 'capture-buy-quote\.sh([[:space:]]|$)' build-and-test:drive build-and-test:evidence [ "$ORCH" = 1 ] || return 0 # Shared top-level scripts with one governing skill, and the /one-shot phase @@ -161,8 +168,9 @@ need 'asana-get-context\.sh([[:space:]]|$)' $INTAKE_UNITS need 'setup-task-workspace\.sh([[:space:]]|$)' one-shot:implementation need 'lint-commit\.sh([[:space:]]|$)' im one-shot:implementation need 'set-tested\.sh([[:space:]]|$)' one-shot:testing -# build-and-test's drive scripts: the ones that build or drive the app on the -# sim (select-ios-sim.sh / slot-preflight.sh only pick and check a slot). +# build-and-test's scripts that build or drive the app on the sim start the +# /one-shot testing phase (select-ios-sim.sh / slot-preflight.sh only pick and +# check a slot, so they pull build-and-test:build above and nothing here). need '(capture-buy-quote|ios-rn-build|ios-rn-build-wait)\.sh([[:space:]]|$)' one-shot:testing need 'asana-review-field\.sh([[:space:]]|$)' one-shot:review need 'pr-create\.sh([[:space:]]|$)' one-shot:pr @@ -183,6 +191,16 @@ case "$SEG_M" in case "$CMD_M" in *agent-run-report*) SEG_NEEDED="$SEG_NEEDED one-shot:report" ;; esac ;; esac +# log-attempt.sh for a value-moving category: the funding slice, as a backstop. +# The log call follows the action, so this cannot put the rules ahead of the +# account or pair choice (the core step map tells the agent to read the slice +# first); it guarantees they are in context before the run concludes anything +# about funding. The category is often quoted, so it is read from the raw +# segment, which the mention-stripped view blanks. +if [ -n "$(invocations 'log-attempt\.sh([[:space:]]|$)')" ] \ + && printf '%s' "$SEG_RAW" | grep -qE -- "--category[[:space:]=]+[\"']?(swap|send|sweep)([\"'[:space:]]|\$)"; then + SEG_NEEDED="$SEG_NEEDED build-and-test:funding" +fi # update-status.sh: --blocked is the blocked completion whatever status rides # with it; a plain Complete is the finalize gate. US_TAIL=$(invocations 'update-status\.sh([[:space:]]|$)') @@ -201,6 +219,7 @@ NEEDED=""; BLAME="" while read -r B64M B64RAW HD; do [ -n "${B64M:-}" ] || continue SEG_M=$(printf '%s' "$B64M" | base64 -d 2>/dev/null) || continue + SEG_RAW=$(printf '%s' "$B64RAW" | base64 -d 2>/dev/null) || SEG_RAW="$SEG_M" segment_units for u in $SEG_NEEDED; do NEEDED="$NEEDED $u" diff --git a/agent-watcher/hooks/tests/build-and-test-slices.test.py b/agent-watcher/hooks/tests/build-and-test-slices.test.py new file mode 100644 index 00000000..b20eb788 --- /dev/null +++ b/agent-watcher/hooks/tests/build-and-test-slices.test.py @@ -0,0 +1,248 @@ +#!/usr/bin/env python3 +"""Contract test for the /build-and-test core + phase references split and its gates. + +Run: python3 ~/.config/agent-watcher/hooks/tests/build-and-test-slices.test.py + BUILD_AND_TEST_DIR= python3 .../build-and-test-slices.test.py # staged tree + +Same budget as pr-land-slices.test.py: a re-attached skill body is truncated at +20,000 characters after compaction, so the core and every rule-bearing slice +stay under it. The pinned pre-split file +(fixtures/build-and-test-pre-split.SKILL.md) is the baseline for rule ids only: +every id it carried must still exist exactly once. Step 0 is split across +build.md (0a-0c) and drive.md (0d-0f), so step ids are not checked for +uniqueness. sim-testing-playbook.md and throwaway-accounts.md are working +knowledge, not rule slices; the playbook has its own read gate. + +Gate half, two hooks: + require-skill-read-for-scripts.sh build / drive+evidence on the companion + scripts (every session), funding on a value-moving log-attempt.sh (orch). + require-playbook-before-drive.sh drive+evidence on a maestro drive (orch), + alone or in the same deny as the playbook block. +""" +import glob +import json +import os +import re +import subprocess +import sys + +LIMIT = 20000 +HERE = os.path.dirname(os.path.abspath(__file__)) +HOOKS = os.path.dirname(HERE) +GATE = os.path.join(HOOKS, 'require-skill-read-for-scripts.sh') +DRIVE_GATE = os.path.join(HOOKS, 'require-playbook-before-drive.sh') +BT = os.environ.get('BUILD_AND_TEST_DIR', os.path.expanduser('~/.cursor/skills/build-and-test')) +REFS = os.path.join(BT, 'references') +SCRIPTS = os.path.join(BT, 'scripts') +FIXTURE = os.path.join(HERE, 'fixtures', 'build-and-test-pre-split.SKILL.md') +SLICES = ['build', 'drive', 'evidence', 'funding'] +GID = 'btslicetest' +SID = 'btslicetest-session' +PLAYBOOK_MARKER = '/tmp/agent-playbook-read-' + GID +NUDGE_FLAG = '/tmp/agent-coreplugins-nudge-' + GID + +RULE_RE = re.compile(r'', re.S) +fails = [] + + +def check(cond, msg): + if not cond: + fails.append(msg) + + +def read(p): + return open(p, encoding='utf-8').read() + + +pre = read(FIXTURE) +core = read(os.path.join(BT, 'SKILL.md')) +texts = [('SKILL.md', core)] + [(s + '.md', read(os.path.join(REFS, s + '.md'))) for s in SLICES] + +# --- 1. sizes ---------------------------------------------------------------- +for name, text in texts: + check(len(text) < LIMIT, '%s is %d chars, must be under %d' % (name, len(text), LIMIT)) + +# --- 2. rule ids: every baseline id placed exactly once ----------------------- +seen = {} +for name, text in texts: + for rid in RULE_RE.findall(text): + check(rid not in seen, 'rule %s appears in both %s and %s' % (rid, seen.get(rid), name)) + seen[rid] = name +for rid in RULE_RE.findall(pre): + check(rid in seen, 'rule %s was lost in the split' % rid) + +# --- 3. step map names every slice, and names each rule in the slice that holds it +step_map = core[core.find('')] +for s in SLICES: + check('references/%s.md' % s in step_map, 'core step map never names references/%s.md' % s) +for row in step_map.split('\n'): + m = re.search(r'`references/([a-z-]+)\.md`', row) + if not m or not row.startswith('|'): + continue + for rid in re.findall(r'`([a-z][a-zA-Z-]+)`', row.split('|')[3]): + check(seen.get(rid) == m.group(1) + '.md', + 'step map lists %s under %s.md, found in %s' % (rid, m.group(1), seen.get(rid))) + +# --- 4. gate map entries resolve to real scripts -------------------------------- +NEED_RE = re.compile(r"^need\s+'([^']+)'\s+(.*?)\s*$", re.M) +mapped = [(m.group(1), m.group(2).split()) for m in NEED_RE.finditer(read(GATE)) + if any(u.startswith('build-and-test:') for u in m.group(2).split())] +check(mapped, 'no build-and-test: entries in %s' % GATE) +by_script = {} +for script_re, units in mapped: + for n in re.sub(r'\\\.sh.*$', '', script_re).strip('()').split('|'): + check(os.path.exists(os.path.join(SCRIPTS, n + '.sh')), + 'gate entry %r names %s.sh, not in %s' % (script_re, n, SCRIPTS)) + by_script[n] = [u.split(':', 1)[1] for u in units if u.startswith('build-and-test:')] + + +# --- helpers ------------------------------------------------------------------- +def clear(): + for key in (GID, 'sess-' + SID): + for f in glob.glob('/tmp/agent-skill-read-%s-*' % key): + os.remove(f) + for f in (PLAYBOOK_MARKER, NUDGE_FLAG): + if os.path.exists(f): + os.remove(f) + + +def mark(unit, orch=True): + open('/tmp/agent-skill-read-%s-%s' % (GID if orch else 'sess-' + SID, unit), 'w').close() + + +def run(gate, payload, orch): + payload = dict(payload, session_id=SID) + env = {k: v for k, v in os.environ.items() + if k not in ('AGENT_TASK_GID', 'SKILL_READ_KEY', 'AGENT_DELIVERABLE')} + if orch: + env['AGENT_TASK_GID'] = GID + p = subprocess.run(['bash', gate], input=json.dumps(payload), + capture_output=True, text=True, env=env) + return p.returncode, p.stderr + + +def bash(cmd): + return {'tool_name': 'Bash', 'tool_input': {'command': cmd}} + + +def delivered(err, sl): + return ('/build-and-test %s phase contract' % sl in err + and read(os.path.join(REFS, sl + '.md')).strip() in err) + + +def premark(orch, *units): + """Satisfy the units a call needs that this test is not about.""" + for u in ('build-and-test', 'one-shot:testing') + units: + mark(u, orch) + + +# --- 5. companion scripts deliver their slices, orch and plain ------------------ +for orch in (True, False): + label = 'orch' if orch else 'plain session' + for script, slices in sorted(by_script.items()): + clear() + premark(orch) + cmd = '~/.cursor/skills/build-and-test/scripts/%s.sh' % script + rc, err = run(GATE, bash(cmd), orch) + check(rc == 2, '%s: %s.sh with no slice: rc=%d, want 2' % (label, script, rc)) + for sl in SLICES: + check(delivered(err, sl) == (sl in slices), + '%s: %s.sh delivery of %s slice: got %s, want %s' + % (label, script, sl, delivered(err, sl), sl in slices)) + rc, err = run(GATE, bash(cmd), orch) + check(rc == 0, '%s: retry of %s.sh: rc=%d, want 0\n%s' % (label, script, rc, err[:300])) + # --help executes no step. + clear() + premark(orch) + rc, _ = run(GATE, bash('~/.cursor/skills/build-and-test/scripts/ios-rn-build.sh --help'), orch) + check(rc == 0, '%s: ios-rn-build.sh --help was gated (rc=%d)' % (label, rc)) + # A command that only mentions a script path executes nothing. + rc, _ = run(GATE, bash('grep -n sim ~/.cursor/skills/build-and-test/scripts/select-ios-sim.sh'), orch) + check(rc == 0, '%s: a grep of select-ios-sim.sh was gated (rc=%d)' % (label, rc)) + +# --- 6. funding backstop: value-moving log-attempt.sh, orch only ---------------- +LA = '~/.config/agent-watcher/log-attempt.sh --gid 1 --action "x" --result success --category %s' +for cat, want in [('swap', True), ('send', True), ('sweep', True), ('"swap"', True), + ('test-drive', False), ('repro', False)]: + clear() + rc, err = run(GATE, bash(LA % cat), True) + check((rc == 2 and delivered(err, 'funding')) == want, + 'orch: log-attempt --category %s: rc=%d funding delivered=%s, want gated=%s' + % (cat, rc, delivered(err, 'funding'), want)) + if want: + rc, _ = run(GATE, bash(LA % cat), True) + check(rc == 0, 'orch: retry of log-attempt --category %s: rc=%d, want 0' % (cat, rc)) +clear() +rc, err = run(GATE, bash(LA % 'swap'), False) +check(rc == 0, 'plain session: log-attempt --category swap was gated (rc=%d)' % rc) +clear() +rc, err = run(GATE, bash('echo "log-attempt.sh --category swap"'), True) +check(rc == 0, 'orch: an echo quoting log-attempt.sh was gated (rc=%d)' % rc) + +# --- 7. maestro drives: drive + evidence through the playbook gate -------------- +CLI = bash('maestro --device ABC --driver-host-port 9182 test /tmp/flow.yaml') +MCP = {'tool_name': 'mcp__maestro__run', 'tool_input': {'yaml': '- tapOn: x'}} +for name, payload in (('maestro CLI', CLI), ('maestro MCP', MCP)): + # Playbook read, slices owed: deny delivers both slices, no playbook block. + clear() + open(PLAYBOOK_MARKER, 'w').close() + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 2 and delivered(err, 'drive') and delivered(err, 'evidence'), + '%s: playbook read, slices owed: rc=%d, want 2 with both slices' % (name, rc)) + check('no maestro drive before the sim-testing playbook' not in err, + '%s: playbook block shown although the playbook marker exists' % name) + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 0, '%s: retry after slice delivery: rc=%d, want 0\n%s' % (name, rc, err[:300])) + + # Nothing read: one deny carries the playbook block and both slices; the + # retry then owes only the playbook. + clear() + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 2 and 'no maestro drive before the sim-testing playbook' in err + and delivered(err, 'drive') and delivered(err, 'evidence'), + '%s: nothing read: rc=%d, want 2 with playbook block and both slices' % (name, rc)) + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 2 and 'phase contract' not in err, + '%s: second deny re-delivered slices or passed (rc=%d)' % (name, rc)) + open(PLAYBOOK_MARKER, 'w').close() + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 0, '%s: playbook and slices satisfied: rc=%d, want 0' % (name, rc)) + + # Slices read, playbook not: the playbook block alone. + clear() + mark('build-and-test:drive') + mark('build-and-test:evidence') + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 2 and 'phase contract' not in err, + '%s: slices read, playbook owed: rc=%d, want 2 without slices' % (name, rc)) + + # Plain session: the hook is orch-only. + clear() + rc, err = run(DRIVE_GATE, payload, False) + check(rc == 0, '%s: plain session was gated (rc=%d)' % (name, rc)) + +# Non-drives pass with nothing read. +clear() +for name, payload in [ + ('maestro --version', bash('maestro --version')), + ('ls of the maestro dir', bash('ls ~/.cursor/skills/build-and-test/maestro')), + ('unrelated command', bash('git status')), + ('MCP inspect_screen', {'tool_name': 'mcp__maestro__inspect_screen', 'tool_input': {}}), + ('MCP take_screenshot', {'tool_name': 'mcp__maestro__take_screenshot', 'tool_input': {}}), + ('MCP list_devices', {'tool_name': 'mcp__maestro__list_devices', 'tool_input': {}})]: + rc, err = run(DRIVE_GATE, payload, True) + check(rc == 0, '%s was gated as a drive (rc=%d)' % (name, rc)) +clear() + +# --- report ------------------------------------------------------------------ +for name, text in texts: + print(' %-14s %6d chars' % (name, len(text))) +print('rules: %d baseline, %d placed; script gates: %s' + % (len(RULE_RE.findall(pre)), len(seen), + ', '.join('%s->%s' % (k, '+'.join(v)) for k, v in sorted(by_script.items())))) +if fails: + print('\nFAIL (%d)' % len(fails)) + for f in fails: + print(' - ' + f) + sys.exit(1) +print('\nPASS') diff --git a/agent-watcher/hooks/tests/fixtures/build-and-test-pre-split.SKILL.md b/agent-watcher/hooks/tests/fixtures/build-and-test-pre-split.SKILL.md new file mode 100644 index 00000000..6192754b --- /dev/null +++ b/agent-watcher/hooks/tests/fixtures/build-and-test-pre-split.SKILL.md @@ -0,0 +1,178 @@ +--- +name: build-and-test +description: Run build and test verification for the active repo. Detects edge-react-gui and runs a real iOS UI test via maestro (Buy $500 quote with proof screenshot); detects Node/TypeScript repos and runs `tsc --noEmit` + smoke checks; falls back to a placeholder ack for unknown repo shapes. Use during the Testing phase of /one-shot. +metadata: + author: j0ntz +--- + + +Verify the active repo builds cleanly before /one-shot marks a task complete. Returns a clear PASS/FAIL signal the caller can include in the Asana summary or use to gate the watch loop. + + + +Inspect the current working directory to decide what to run: +1. If `package.json` `name` is `edge-react-gui` → iOS UI test (maestro) path (step 0). Check this first. +2. Else if the repo is an EdgeApp gui DEPENDENCY (per `gui-dependency-integration`) → run its own checks (the TS/Node path below) AND the gui integration test. A dep change is NOT done until it runs in the app. +3. Else if `package.json` exists and a `tsconfig.json` exists → Node + TypeScript path (step 1). +4. Else if `package.json` exists with a `test` script but no tsconfig → Node path (step 2). +5. If `Cargo.toml` exists → not implemented yet, fall through to placeholder. +6. Otherwise → placeholder mode (step 3). +On FAIL, surface the exact command, exit code, and last 30 lines of output. Do not try to fix anything inside this skill — the caller decides whether to amend or block. +This skill does NOT edit source code, commit, push, or change Asana state by default — verification + results only. The only exceptions are the explicitly scoped rules below: `testids-over-coordinates` (testID additions as a separate commit IN the task's single gui PR — never a separate testID-only PR), `gui-dependency-integration` (gui-side changes a dep actually needs, committed on the gui branch), and the LOCAL-ONLY, never-committed corePlugins edits — `single-asset-plugin-trim` (currency trim) and `force-swap-provider-locally` (force a swap provider). +Scoped exception to `no-mutation`, test-infrastructure only. TESTIDS FIRST, COORDINATES LAST. When a maestro flow needs to drive an element that has no stable selector (text match fails and no `testID` exists), the DEFAULT action is to ADD the missing `testID` prop to that component in the gui worktree and drive via it — do NOT struggle with coordinate taps. A `testID` is a JS-only prop: Metro reload picks it up in seconds (no native rebuild), so adding one is cheaper than even a single round of coordinate trial-and-error, and it de-brittles the suite for every future run. Coordinate taps are permitted ONLY for surfaces you cannot edit (system dialogs, native pickers, third-party views that don't forward `testID`) or when a reload would destroy unrecoverable in-flight app state — and any coordinate tap that survives into the PROOF flow must be called out in the run report with why a testID was not possible. Commit discipline: commit the testID additions as a SEPARATE commit, distinct from any feature commit; change ONLY `testID` props, never component logic; update the maestro selector(s) to use them. MESSAGE NAMES THE SURFACE: subject `test: add testIDs to ` (e.g. `test: add testIDs to ExchangeScene swap pills`), with every added id listed in the body — never a generic subject like `test: add missing testIDs for maestro selectors`: identical subjects across runs make these commits indistinguishable when a human cherry-picks between branches. WHERE THE COMMIT LANDS — always THE TASK'S SINGLE GUI PR, the same one whose test surfaced the need; NEVER a separate testID-only PR. A task has at most ONE gui PR: for a gui-feature task the testID commit rides that feature PR; for a DEP-repo task whose maestro test drives the gui, the testIDs go in the task's ONE gui integration PR (and if the testIDs are the only gui change, that PR IS the task's gui PR — a first-class PR, not a throwaway), on the SAME `/` gui branch the run already provisioned. Do NOT cut a second gui branch/PR for the testIDs when the task already has (or will have) a gui PR — that splits one task's gui work across two PRs, which is the mistake this forbids. If no selector was missing, this rule is a no-op. +OPTIMIZATION (optional, LOCAL-ONLY — never committed). When the task targets a SINGLE asset and the maestro test needs to drive that asset's wallet, you MAY temporarily comment out the unrelated currency plugins in the gui worktree's `src/util/corePlugins.ts` (the `currencyPlugins` map) — keeping the plugin(s) the task needs — to cut app load/init time (fewer plugins to spin up). This is a test-harness speedup ONLY: it must NEVER land in a commit or PR. Revert it before any commit, or rely on it living only in the throwaway test build; if you commit after trimming, verify `git status`/`git diff` does NOT include `corePlugins.ts`. Skip entirely for multi-asset tasks or tasks that don't drive a wallet. +To FORCE a specific swap provider for a test (so the engine routes through it instead of a competitor), edit the gui worktree's `src/util/corePlugins.ts` `swapPlugins` map and set every OTHER provider to `false`, leaving only the target's `*_INIT` truthy — LOCAL-ONLY, same corePlugins surface and same never-committed discipline as `single-asset-plugin-trim` (revert before any commit; verify `git status`/`git diff` excludes `corePlugins.ts`, or rely on the throwaway build). Do NOT force a provider by toggling the in-app **Settings → Exchange Settings**: that state is ACCOUNT-SYNCED, so on a shared roster account it thrashes against every parallel session and persists to the next run — parallel-UNSAFE and forbidden as the forcing lever. Use Exchange Settings only to READ/diagnose why a provider is absent, never to set routing. (`Preferred`/`preferPluginId` also do not pin — engine reverts to best-rate ~60s — so the local hard disable is the reliable lever.) See sim-testing-playbook "Feature-enablement check". +Deterministic operations (sim selection, RN build, capture loop) MUST run via the companion scripts under `~/.cursor/skills/build-and-test/scripts/`. Do not inline their logic as raw bash blocks in this SKILL.md or in agent reasoning. +START the sim-testing phase with ONE call: `~/.cursor/skills/build-and-test/scripts/slot-preflight.sh` (defaults to `$AGENT_SIM_UDID`/`$AGENT_METRO_PORT`; pass `--repo ` when cwd is not the gui repo). It boots the sim if needed and answers, deterministically, the questions runs kept re-deriving at high friction cost: is Metro mine or squatted, is the app installed, does the installed native side match the worktree (`.agent-native-build-stamp` vs `ios/Podfile.lock`), are node_modules present. OBEY its final `PLAN:` line — `ready` (drive now, no build), `js-only`/`install`, or `full-rebuild` — and run the exact `INVOKE:` command it prints (verbatim; no flag-guessing), then its `WAIT:` command when it prints one. Do NOT re-derive any of its checks manually, and do NOT start a second Metro when it reports one running. +When verification needs RUNTIME state from the running app — why a check evaluates false, the actual value of a variable, which code path executed (e.g. the Swap/Maya "investigate outage" kind of task) — use the `/debugger` skill (`~/.cursor/skills/debugger/SKILL.md`), do NOT hand-roll a CDP/WebSocket attach. It sets a `file:line` breakpoint over Metro's Hermes inspector and reports the call stack + locals. It is already slot-aware: `check-metro.sh` and `cdp-attach.js` default to `$AGENT_METRO_PORT`, so in a parallel slot it targets THIS session's Metro (base 8181) with no port flags. Static questions (where is X defined) stay grep/read — `/debugger` is only for live runtime state. +A critical-path wait (build, Metro bundle, screenshot, app-ready — anything you cannot proceed without) MUST be a single BLOCKING call inside the CURRENT turn. NEVER end your turn and hand the wait to a backgrounded shell expecting "the background task will re-invoke me when it finishes." That makes your own forward progress depend on an external re-invoke, and when the wait can't complete you idle forever with no one driving — the failure that wedged the BitcoinDepot and piratechain runs. This is the same disease one-shot's `never-self-respawn` already forbids: *"any wait is a single blocking call in THIS process."* Concretely: +- **Do the wait, get a result, react — all in this turn.** Foreground it. The harness's background-completion → re-invoke is for genuinely parallel/optional work, NOT for a step the next step depends on. +- **Bound every wait with `timeout `** so it ALWAYS terminates (success OR timeout) and control returns to you to react. (`timeout` IS available — macOS ships no `timeout`/`gtimeout`, so it's provided on PATH by the portable shim `~/.cursor/skills/timeout.sh`; `timeout 180 ` just works.) An unbounded `until grep ; do sleep 5; done` / `while ! ; do sleep; done` hangs forever the moment the marker never appears (wrong logfile, wrong marker, build died). A timed-out wait is a real FAIL/retry to handle now — never a reason to spawn another waiter. +- **iOS builds run detached, then wait in chunks.** A cold build outlives the Bash tool's 600s cap (the call is killed mid-build), so never run a full build as one foreground call or under a hand-rolled nohup/poll loop: use step 0c's two commands and re-run `ios-rn-build-wait.sh` in this turn while it exits 7. Wait exits 0/1/2 are the build's own result; 3 = stalled (no log output for 10 min) and already killed: read the printed log tail, fix, start a fresh `--detach`; 4 = no build to wait on: start one. +- **Any other long compile** (gradle, a hand-run xcodebuild) needs a stall check inside its bounded wait: log mtime frozen with no live compiler children means HUNG now; kill, diagnose, retry instead of waiting out the timeout. Use `capture-buy-quote.sh` (bounded retry cycles) for app capture. +- **Detect readiness against the resource you actually started, not a guessed log line:** `timeout`-bounded `curl` against the Metro you launched on its REAL port (`/status`, then the `index.bundle` URL) — never `grep` a logfile whose name/marker you assumed (the bug here: Metro logged to `gui-metro2.log` but the waiter grepped `gui-metro.log`). +Mirrors one-shot's `never-self-respawn` and `pr-watch-bounded-poll`. Recovery by an outside watchdog is explicitly NOT the safety net — the agent must not hang in the first place. +Never assume a repo's package manager — repos migrate between npm and yarn (edge-react-gui is currently yarn-locked; package-lock.json was removed upstream). All install/run/pack operations go through the shared dispatcher `~/.cursor/skills/pm.sh`, which detects the lockfile (`package-lock.json`→npm, `yarn.lock`→yarn, both/neither→npm). Companion scripts in this skill already dispatch through it; do not hand-write `npm ...`/`yarn ...` against a repo without checking `pm.sh detect`. +A change to an EdgeApp gui DEPENDENCY is NOT fully tested until it runs in the app — its own `tsc`/jest passing is necessary but NOT sufficient. Gui dependencies = the Edge-owned repos `edge-react-gui` consumes: `edge-core-js`, `edge-currency-accountbased`, `edge-currency-plugins`, `edge-exchange-plugins`, `edge-login-ui-rn`, `edge-currency-monero`, `react-native-piratechain`, `react-native-zcash`, `react-native-zano`. When the repo under test is one of these, after its own checks you MUST also run the gui integration test, autonomously (NO prompting): +1. **Co-located gui worktree:** ensure one exists — create via `~/.config/agent-watcher/setup-task-workspace.sh --task-gid --repo edge-react-gui` if absent (sibling of the dep worktree under `~/git/.agent-worktrees//`, so updot can find it). +2. **Link the MODIFIED dep into the app — the mechanism, and whether you flip any `DEBUG_*` flag, is YOUR per-task call** (depends on what the task changed and how you want to verify it; it is NOT a fixed per-dep rule). Run repo scripts with each repo's package manager (lockfile: `yarn.lock`→yarn, `package-lock.json`→npm; **yarn is being phased out — check, don't assume**). The toolbox: + - **`updot` — bakes the built dep into the gui's `node_modules`.** Works for ANY dep, no dev-server, no runtime race → the safe default for headless/automated runs. ` updot ` then the gui's `prepare` (npm form: `npm run updot -- && npm run prepare`; add `prepare.ios` for native-module deps), then rebuild. The dep's `DEBUG_*` flag stays FALSE (you baked it in). + - **`DEBUG_` flag + the dep's live webpack dev-server — webview-plugin deps only** (`DEBUG_ACCOUNTBASED`:8082, `DEBUG_EXCHANGES`:8083, `DEBUG_CURRENCY_PLUGINS`:8084, `DEBUG_PLUGINS`:8101 — these ports are HARDCODED in each dep package's `debugUri` and are HOST-GLOBAL). Set the flag TRUE in the gui's `env.json` AND run the dep's `yarn start`/`npm start` (webpack serve) backgrounded for the test; the webview loads the local bundle live (sim reaches host localhost), no gui rebuild. Pick this when live iteration helps; if it flakes (dev-server unreachable, ATS/cleartext, recompile race) fall back to updot. + - **Parallel-slot port rule (this bit the Swap/Maya run):** a `DEBUG_` dev-server port is a SINGLE-OCCUPANT host resource — only ONE slot can serve a given dep at a time. A second concurrent session needing the SAME dep MUST use updot instead. Before starting the dev-server, check the port is free: `lsof -nP -iTCP: -sTCP:LISTEN`; if another slot holds it, use updot. Your slot's Metro runs on `$AGENT_METRO_PORT` (base **8181**, i.e. 8181/8182/8183…), deliberately OUTSIDE the 808x DEBUG range so Metro never collides with a dev-server — do NOT pass a `--port` that drags Metro back into 808x. When in doubt in a parallel slot, prefer updot: it has no shared port and is collision-free by construction. + - **`DEBUG_EXCHANGES` crash-loop trap (Swap/Maya):** the gui's `allowDebugging` flag (which permits the cleartext localhost load) is OR-gated on `DEBUG_ACCOUNTBASED || DEBUG_CORE || DEBUG_CURRENCY_PLUGINS || DEBUG_PLUGINS` — **`DEBUG_EXCHANGES` is NOT in that set**, so enabling it ALONE crash-loops the app. Co-enable one that IS (e.g. `DEBUG_ACCOUNTBASED`); note that drags in its 8082 dev-server, so plan ports per the rule above. Also: swap/exchange plugin code runs in **edge-core-js's webview context, not the Metro bundle** — serve patched dep code via the dev-server (or `updot`-bake it); do NOT sync patched `lib/` into `node_modules` expecting Metro to bundle it. + - **`edge-core-js`: prefer `updot`, avoid `DEBUG_CORE`.** `DEBUG_CORE` loads the WHOLE core from hardcoded `http://localhost:8080/` (`edge-core-js/.../react-native-webview.tsx`: `source={debug ? 'http://localhost:8080/' : null}`) — races init, cleartext/ATS-sensitive, and any hiccup takes the entire app down (the long-standing "DEBUG_CORE is buggy"). updot is reliable for core. + Only link the dep(s) THIS task modifies; leave every other dep's `DEBUG_*` at its env.json default. Keep flags consistent with what you actually linked — a `DEBUG_*` left true with no dev-server running will break that dep. +3. **Login:** the test account auto-logs-in via the `YOLO_*` env knobs (set by workspace init to the roster's `agent` account from `~/.config/edge-secrets/test-accounts.json`, consumed in `LoginScene.tsx` — pinned by `setup-task-workspace.sh` on every worktree's env.json copy). Keep them set so the maestro run reaches the logged-in app; when the change is to `edge-login-ui-rn` specifically, these are the lever for exercising the login flow — adjust only if the change requires driving the login UI differently. +4. **Make the gui-side changes the feature NEEDS to run, then run the gui maestro path (step 0)** against that build. A dep change almost always needs gui-side wiring to actually function — plugin init / apiKey, provider/plugin registration, imports, config. Those gui changes are PART OF THE WORK, not optional: complete ALL of them (and commit on the gui worktree's branch) so the app is fully runnable with the feature, autonomously, do NOT prompt. Do not stop at "the dep compiles" or "it links" — if the feature doesn't load/run in the app yet, the implementation is NOT done. +PASS requires the maestro app test to pass with the dep change linked AND the actual feature exercised to its terminal success (`test-drives-the-real-action`). A dep whose unit checks pass but that isn't fully wired into a runnable app, or that runs but whose real action was never executed, is a FAIL. +SCOPE DOES NOT EXEMPT THE TEST (the Houdini-prototype rationalization, 2026-06-11): a task that scopes its deliverable to the dep repo, calls itself a prototype, or explicitly defers PRODUCTION gui integration to follow-up work still gets THIS integration test. The wiring in steps 1-4 is TEST SCAFFOLDING in the task's gui WORKTREE (plugin registration, env.json keys, dep linking) — it is not an "unrequested production change"; nothing lands in the gui repo unless the task asks for it. Likewise "the plugin is unvetted prototype code, a real swap through it is irreversible" is NOT a blocker: vetting it with a small sanctioned-roster swap is exactly what this test exists to do (see one-shot `yolo-true-blockers` carve-out). +PLATFORM: default to iOS. Provision the iOS sim, run the iOS maestro flow, and credit `iOS Sim` UNLESS the task EXPLICITLY calls out Android (task title/description says Android, the task is tagged Android, or the change is under `android/` only). For an Android-called-out task, run the ANDROID path instead of (or in addition to) iOS: `./gradlew :app:assembleDebug` from the gui worktree's `android/` is the build verification, and a successful APK is the terminal-success signal for a BUILD-ONLY fix (GitHub `pr-checks.yml` does NOT build Android, so these regressions are invisible to GitHub checks — the local assembleDebug is what catches them). Credit `Android Sim` and log the attempt via `log-attempt.sh --category test-drive --result success|failed:`. The Android build needs gitignored secrets the node_modules clone does not carry (`android/app/google-services.json`, `EdgeApiKey.java`, `android/app/src/main/assets/edge-core/plugin-bundle.js`, a generated `android/local.properties` with `sdk.dir`) — `setup-task-workspace.sh` copies them; and the Android SDK (`ANDROID_HOME`/`ANDROID_SDK_ROOT`) must be present in the env. Run gradle with `--no-daemon` (or a per-slot `GRADLE_USER_HOME`) for parallel-safety; assembleDebug is CPU/RAM-heavy, so do not run many concurrently. A genuine in-app Android drive (AVD + maestro) is a larger path; the build-only check closes the regression gap for build/native fixes. If a task touches BOTH platforms, exercise iOS and credit both. +DEFAULT to physically exercising the change in the running app on the sim. Almost ANY task can be tested in-app — a swap, a send, a settings toggle, an onboarding/account-creation flow, a specific wallet action, a bug repro. `tsc`/jest/build passing is NECESSARY BUT NOT SUFFICIENT: a change is not verified until you have driven the actual changed behavior in the app via maestro and seen the expected result — to its TERMINAL success, not a precursor (see `test-drives-the-real-action` for the exact bar: execute the real action, e.g. an actual swap, not just a quote). Do NOT skip the sim test because static analysis "looks right", because the diff is small, or because authoring a flow is effort (Rango shipped a swap-plugin change with NO in-app test — that is the failure this rule forbids). Specifically: before setting `blocked = Yes` with reason "can't verify / no defensible default" on a bug, repro, or investigation task, you MUST first attempt the most-specific RUNTIME REPRO you can construct — build the relevant flavor (e.g. `ENABLE_MAESTRO_BUILD=true` for test-server flows) and drive the precise maestro flow. "I can only trace it statically" is NOT a blocker. Block only if the repro is genuinely un-runnable here (missing creds/KYC/datastore the slot can't provide). For FUNDS specifically: the ONLY funds blocker is an OBSERVED TRUE LOSS — an attempted swap/send that failed AND lost principal. Fees/slippage NEVER count as loss (budgeted at $15 equivalent per run, per the playbook), and blocked-ness is established by ATTEMPTING, never predicted. +When the asset or feature under test genuinely CANNOT be driven in the sim — a default-disabled plugin that crashes the debug build on enable (e.g. `BOTANIX_INIT`), a date-gated change that only activates after a future date, or a non-GUI-dependency surface — the bar is NOT "skip the in-app test and declare `verified: not-run`". Do BOTH: (1) drive a PROXY that exercises the SAME mechanism your change routes through and capture proof of it reaching its terminal state (e.g. for a keys-only create-wallet exclusion, a hardcoded-enabled keys-only asset like `bitcoinsv` hits the identical exclusion path — see the sim-testing playbook); and (2) unit-test the gate/branch your change adds (the condition that routes the target asset) so the logic is covered even though the asset itself cannot run. Record both in the Testing section and set `verify_blockers: [precondition]` (asset un-runnable here) with the proxy drive + unit test as the evidence. That combination is a sanctioned PASS; `verified: not-run` with an empty `verify_blockers` is not. When the un-drivable surface is VISUAL, the combination is not sufficient on its own — add the hack-verified frame per `hack-verify-visual-changes`. +LOG every value-moving action and every test-drive/repro the moment it resolves, via `~/.config/agent-watcher/log-attempt.sh --gid --action "" --result success|failed:|loss:|blocked: --category swap|send|sweep|test-drive|repro`. This attempt-log (`$XDG_STATE_HOME/agent-watcher/attempts/.jsonl`) is the AUTHORITATIVE, agent-location-independent record of what the run actually attempted — the concession-validation gate reads it to tell a real wall (`loss:`/`failed:`/`blocked:` after an attempt) from a predicted one, on BOTH a formal `--blocked yes` AND a silent DOWNGRADE-finalize (completing or opening a PR without reaching the prescribed in-app success), and the eval reads it as ground truth for testing-depth instead of trusting transcript narration. RESULT semantics: `success` = reached terminal success; `failed:` = attempted, no success, principal safe (fees only); `loss:` = attempted, FAILED, principal unrecoverable (the ONLY funds condition that legitimizes a block); `blocked:` = attempted up to a precondition the slot genuinely cannot satisfy (real provider halt, geo-block confirmed by attempt). This matters most after the tester becomes its own agent: the drive will live in the tester's context, NOT this transcript, so the log is the only place a concession gate or the eval can see it — write it on EVERY attempt now so the contract is already in force when that split lands. +The test harness is YOURS to build — its absence is NEVER a blocker. When driving the real behavior needs scaffolding that does not exist yet, CREATE it locally and uncommitted: author a new maestro `.yaml` flow (per `maestro-flows-are-shortcuts` — expected, not exceptional), add a missing `testID` (`testids-over-coordinates`), trim unrelated plugins (`single-asset-plugin-trim`), disable a crashing module (the piratechain local-disable), or HARD-CODE the inputs the code path reads — fixtures, seed data, info-server/remote-config payloads, feature-flag state, a forced provider. "No maestro flow exists for this", "the data comes from a remote server I don't control", "there's no fixture", "the feature isn't enabled by default" are NOT blockers and NOT reasons to stop at static analysis — they are scaffolding to BUILD. KEY DISTINCTION: hard-code the INPUTS to REACH and exercise the real logic, never fake the OUTPUT to fabricate a pass. Injecting a disable-map into the store so the REAL `isSpendBrandDisabled` filter runs against controlled data is correct; hard-coding "this brand is hidden" to skip the filter is not — the changed code path must actually execute. LOCAL-ONLY discipline (same as `single-asset-plugin-trim`): this scaffolding is throwaway — it must NEVER land in a commit/PR. Revert it before any commit (verify `git status`/`git diff` is clean of it), or rely on it living only in the disposable test build. If you find yourself writing `blocked = Yes` or "could not test because doesn't exist", stop: build the scaffolding and drive the test. +Before the sim-test phase, READ `~/.cursor/skills/build-and-test/references/sim-testing-playbook.md` — it is short and holds the working knowledge (funding floors, account roster/switching, feature-enablement gotchas, investigation order) that otherwise gets re-learned every run. Then COMPOSE, don't re-derive: parameterized subflows live in `~/.cursor/skills/build-and-test/maestro/common/` (`login-if-needed`, `dismiss-startup-modals`, `select-swap-pair`, `confirm-slider` — the slider is SOLVED there; never re-derive the gesture). Copy the subflows you need next to your task flow and `runFlow` them; author NEW task-specific `.yaml` liberally for what the task actually changed (expected, not exceptional), keeping task flows LOCAL (`.syncignore`d from the agent repo; never committed to the gui repo — its `maestro/` is the heavyweight verification suite, reference-only for selectors). What DOES get committed to the gui: missing `testID`s, per `testids-over-coordinates` (add them the moment a selector is missing — never grind coordinates on an editable component). EXPLORATION vs PROOF: for exploring screens/selectors use the **maestro MCP tools** (persistent driver — no ~2-min `maestro test` startup per probe). CAUTION — the maestro MCP daemon binds to one device and DRIFTS. `maestro-mcp-wrapper.sh` start-pins it (`maestro --device "$AGENT_SIM_UDID" mcp`), but the start-pin does NOT survive slot-sim relaunch cycling: the daemon rebinds to a different booted pool sim mid-run, so it then drives/observes ANOTHER slot's sim (the confirmed cause of a red-springboard "proof" of the wrong device). Therefore: (1) the MCP is for EXPLORATION ONLY — re-verify its bound device (list_devices + an inspect matching your app's expected state) at the START OF EVERY observation block, not once, because it can rebind between blocks; (2) ALL PROOF evidence comes from the maestro CLI — canonical slot invocation: `maestro --device "$AGENT_SIM_UDID" --driver-host-port $((AGENT_METRO_PORT + 1000)) test ` (the per-slot driver port keeps parallel slots' iOS drivers off each other's ports; hook-enforced; the MCP daemon uses +2000 via its wrapper) — + `simctl io` against that SAME UDID — NEVER an MCP screenshot and NEVER a parallel `simctl io` against a different device than the CLI drove. If a verify shows the MCP on the wrong device, drop to CLI for the rest of the run. For the repeatable PROOF run compose ONE yaml flow and run it once via `capture-buy-quote.sh --flow ` (CLI-driven, auto-pins device + driver port from the slot env; that run produces the PR evidence screenshots). When a run teaches you something durable, PROPOSE it as a `[playbook]`-tagged bullet in your run report's Dev Notes & Gotchas section — do NOT edit the playbook directly. A NEW reusable drive sequence (not knowledge, a FLOW) is proposed the same way with a `[flow]` tag: name, params, one-line purpose, and the FULL yaml EMBEDDED in the report as a fenced block (worktrees are pruned on retention — a path reference dies with the worktree; the report attachment is the durable copy). A change to an EXISTING library flow (genericize, new param, split into subflows) uses `[flow-update]`: name the flow, the change, and the compatibility argument — new params MUST default to current behavior so existing callers are unaffected, and a rename/split must say so explicitly (callers get grepped at promotion). Promotion is NEVER done by task runs: proposals wait in reports for the eval's manually-triggered flow-consolidation pass. The playbook is operator-curated: entries in it may be trusted without re-verification precisely because every one was reviewed and promoted by the operator; agent-written proposals wait in reports until promoted. SCOPE a `[playbook]` proposal TIGHTLY — it earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE — if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), do NOT propose it as a recipe; propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything already covered by an existing rule. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately — miscategorizing them bloats the playbook and dilutes the load-bearing lines. +The test is COMPLETE only when the ACTUAL end-to-end user action the task is about has EXECUTED successfully in the app and you've captured proof of its terminal success state — NOT a precursor or a partial step. The repeated failure is stopping SHORT: Rango declared a swap-plugin change tested at the QUOTE; a quote is NOT an executed swap. Per-action bar: a SWAP is done at the executed-swap success scene (e.g. "Congratulations" / order submitted), NEVER at the quote; a SEND at the broadcast/confirmation screen, not an address entered; a feature at its real user-visible outcome, not "it builds" / "the plugin loaded". (Exception: when the task's deliverable IS the precursor — e.g. the buy-quote smoke test exists to render a quote — then that's the bar. Identify the actual user-facing outcome and drive to IT.) +ALL prerequisites to reach that terminal state are MANDATORY and not skippable — and the FIRST, most-skipped one is finishing the IMPLEMENTATION itself: complete EVERY code change across ALL required repos to make the feature integrated and actually RUNNABLE in the sim — the core/dep change AND the gui-side wiring it needs (plugin init / apiKey, provider/plugin registration, imports, config). A partial implementation that never makes the app fully runnable is the upstream failure here: agents do part of the work and stop before the feature even loads, let alone executes. Do NOT stop at "core change written" or "it compiles" — wire it all the way into the gui so the app RUNS with the feature. THEN: link the modified dep into core+gui (`gui-dependency-integration`), build, fund or switch accounts (`funded-test-accounts`), force the provider (`force-swap-provider-locally` — the local corePlugins hack, NOT in-app Exchange Settings), and execute. "The task isn't complete until you can execute a successful swap in the sim, so all parts are necessary including core/gui." Drive through every step EAGERLY; do not declare done, and do not block, until the real action has actually run — unless you hit a genuine precondition the slot truly cannot satisfy (real funds/KYC/finality), in which case capture what you have and `blocked = Yes` with the specific precondition. +CEILING (so you don't over-grind once you're actually there): the bar is the IN-APP success state, not EXTERNAL finality — once the success scene shows and you've captured proof, you are DONE; do NOT then wait for on-chain settlement / full balance sync / provider-side completion (minutes-to-never, out of scope). Run each maestro flow as a SINGLE bounded in-turn call (`timeout maestro test `), never backgrounded-and-polled (re-pays the ~2-min driver startup; it's the `blocking-in-turn-waits` footgun); bound every `extendedWaitUntil`. +A change with a USER-VISIBLE surface (a new row, badge, spinner, empty state, error copy, layout shift) is NEVER finalized on logic evidence alone. When the natural trigger cannot be reproduced on the sim — the window is too fast, the state needs a cold engine, the condition depends on a remote/timing precondition the slot cannot hold — you still owe the PIXELS: HACK-VERIFY it. Force the state with a TEMPORARY UNCOMMITTED edit in the worktree (hard-code the branch condition true, stub the slow call, pin the state field), drive to the screen, capture the frame, then REVERT the hack and prove the tree is clean (`git status --porcelain` empty for that file) before any commit. This is `build-the-test-harness` applied to rendering: you are forcing the INPUT state so the REAL component renders, never faking the output pixels — a mocked-up image or a design comp is not evidence. Unit tests that assert the element renders are complementary, not a substitute: they prove the condition wires up, not that the thing looks right on the device. LABEL THE RESULT so the provenance is never lost: save the frame as `/tmp/agent-proof---HACKED-.png` (the literal token `HACKED`), which makes `pr-attach-screenshots.sh` caption it 🪓 HACK-FORCED and banner the comment — the attach step requires `--hack-note` with a one-line description of the exact hack (pr-create `attach-test-evidence`), so write that line when you make the hack and reuse it; state in the report's Testing section which frames were hack-forced, what the hack was, and that it was reverted. A RE-TEST after any code change lands at a new head sha, which retires the previous frames: carry forward the ones the change left true and retire the ones it invalidated (pr-create `evidence-dispositions`), so a run never leaves pixels standing that its own fix made false. What a hack-verified frame does and does not buy: it PROVES the rendering (layout, copy, colors, no reflow); it does NOT prove the trigger fires in production, so the trigger still needs its own evidence (a unit test, a log line, a code-path argument) and stays named in `verify_blockers`. +The milestone screenshots you capture per `test-drives-the-real-action` are PR EVIDENCE, not throwaways — they get attached to the PR (by `/pr-create`'s `attach-test-evidence`). Capture them deliberately: MULTIPLE if needed to actually illustrate the change (typically: the feature enabled/visible, the action in progress, the terminal success state — e.g. provider-in-settings → quote → confirm → success scene). Save as `/tmp/agent-proof---.png` where `NN` orders them and `slug` is a short human-readable description (`01-provider-enabled`, `02-sonic-quote`, `03-swap-success`) — the slug becomes the image caption on the PR, so write it for a reviewer, not for yourself. One screenshot is enough only when the change is fully visible in a single frame. A frame captured under a forced state carries the `HACKED` token in its filename per `hack-verify-visual-changes` — never strip it to make the evidence look stronger. +VERIFY THE PIXELS before attaching: Read each captured PNG and confirm it actually renders the scene its slug claims. A springboard/home screen, a red error screen, a lock screen, or a blank frame is NOT proof of anything — recapture, never attach. A FAILED read is also a fail: if your own Read of the image errors or returns "could not be processed / removed", you have NOT pixel-verified it — never label such a screenshot "pixel-verified", never attach it, and never credit `Simulator` on it; recapture until a Read SUCCEEDS and shows the claimed terminal-success scene. "I drove the flow so the proof must be right" is precisely the trap (the maestro/MCP driver can observe a different device than simctl photographed); the successful read of a non-error scene is the only thing that substantiates the credit. CAPTURE FROM THE DEVICE YOU DROVE: when the maestro MCP drove the flow, capture via the MCP's own take_screenshot (it renders the daemon's bound device); when the maestro CLI drove it, `simctl io` against the SAME `--device` UDID. A parallel `simctl io` against a different device than the one maestro actually drove photographs the wrong simulator and produces sincere-but-false evidence. +For tasks needing a FUNDED asset (swaps/sends/sync-observation), the test-account ROSTER is the LOCAL-ONLY file `~/.config/edge-secrets/test-accounts.json`: roles `agent` (the agents' own account and **the default YOLO login**; 2FA ON, password + OTP key in its `credsFile`, so its 2FA is never a user-only-credential wall), `primary` (heavily funded but cluttered with leftover assets), `qa-a`, `qa-b` (region California/USA), each with username, PIN, and notes. Refer to accounts BY ROLE in anything synced, committed, or posted; usernames and PINs never leave that file. This roster is EXHAUSTIVE and the scope of any account search: the sim also contains many junk/leftover accounts — do NOT trawl beyond the roster. **Test on the `agent` account and acquire assets by SWAPPING on it**: (1) the agent account already holds the asset → use it; (2) otherwise swap into it on the agent account from a wallet it holds (creating the destination wallet per `create-missing-destination-wallet`), driven to the success scene like any tested swap; if no provider quotes the pair, enable every provider for the requote, worktree-locally only: undo any `force-swap-provider-locally` edit, and set any `*_INIT` that the worktree env.json has as `false` to `{}`. NEVER send funds from another roster account to the agent account: its balance grows only through swaps it executes. Size each acquiring swap per the minimum-viable-amounts rule (binding floor + 10-20% buffer); what it buys stays in the agent account for later runs. (3) Only when the agent account holds nothing that clears any provider floor for a route to the asset, run the test on the roster account that holds the asset and say so in the run report; "no test account holds X" is NOT a valid conclusion until each ROSTER account was actually checked. Move a test off the agent account otherwise ONLY when it needs another account's own state (qa-b's region, a per-account setting or history the task names). (`YOLO_*` auto-login starts every run on the `agent` account and re-asserts it on each relaunch.) **Switch accounts by EDITING the worktree's `env.json`** (`YOLO_USERNAME`/`YOLO_PIN` → target roster account, then `simctl terminate` + `launch`; auto-login does the rest) — NOT by driving the in-app account dropdown, which is slow and fumble-prone. Never brute-force a PIN prompt (exponential lockout) — look it up in the roster. +If you ALREADY hold a funded, provider-supported pair — both wallets present and the source funded above the floor (e.g. funded BTC + an existing FTM wallet → BTC→FTM on SideShift) — that swap is EXECUTABLE NOW and you MUST drive it to terminal success per `test-drives-the-real-action` (quote → confirm slider → broadcast/success scene). Do NOT abandon a held executable pair for a slower-settling or riskier alternative (a slower new-wallet path, a different target asset, etc.) — pick the executable pair you have and finish it. A slow off-chain settle is NOT a reason to abandon: the bar is the in-app success scene, not on-chain finality (`test-drives-the-real-action` CEILING). NEVER report `blocked = Yes` with reason "blocked on funding / no executable pair" while any roster account holds a funded, provider-supported wallet you have not driven to completion — that conclusion is invalid until you have actually attempted the held pair to the confirm slider. Sanctioned majors (BTC/ETH/USDC and similar, supported by nearly every provider) are spendable funding sources; a $2k+ funded roster account is never "blocked on funding". Falling back to direct-API / boot-init verification (sim-testing-playbook Fabric-SIGABRT entry) is permitted ONLY after a genuine funded attempt on a held executable pair is interrupted by the build crash — not as a first resort and not while an untried funded pair exists. SIZE every value-moving action (swap, send, sweep) per the playbook's minimum-viable-amounts rule: discover the binding floor first (provider pair minimum / dust limit), then move the smallest amount that clears it with a 10-20% buffer — never a round default like $20, and never repeat a successful value-moving action for extra evidence. +When a test SENDS to another wallet (a same-asset wallet on another roster account, a second wallet on the same account, a swap/transfer/sweep target) and that destination wallet does not exist, CREATE IT and continue: in whichever roster account the test needs it, via Add Wallet (or Add / Edit Tokens for a token). A receive-side wallet needs no funds, so the roster-search-first step of `funded-test-accounts` does not apply to it. "No destination wallet", "the receiving account has no X wallet", or "would need to create a wallet" is NEVER a blocker, a concession, or a reason to switch to a weaker verification; the completion judge and blocker validator deny it on sight. Wallet creation is a supported test path (playbook). +In a watcher slot (`$AGENT_SIM_UDID` set), resolve your simulator ONLY via `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"` — NEVER by `--runtime`/`--device`. Raw `simctl ... booted` is hook-blocked in slot sessions (multiple sims boot concurrently; `booted` is ambiguous) — pass `$AGENT_SIM_UDID` explicitly to every ad-hoc simctl call, and select that device in the maestro MCP before driving. By-name resolution targets the SHARED MASTER sim ("iPhone 16 Pro Max"); running builds/maestro on the master pollutes the golden image every clone is cut from. `select-ios-sim.sh` now refuses by-name in slot mode (override: `--allow-master`). Your slot clone DOES carry the Edge app + the logged-in test account (APFS copy-on-write from the master) once booted — note `get_app_container` returns NOTHING on a SHUT/never-booted clone, a FALSE negative; boot first (the scripts do) and trust the clone. Do NOT trigger a fresh rebuild on that false negative (it wastes minutes AND wipes the cloned login state → the onboarding screen). DISTINCT and LEGITIMATE: `ios-rn-build.sh` will itself force a rebuild when the cached app's NATIVE build drifted from the worktree (its `ios/Podfile.lock` stamp no longer matches — e.g. a slot image baked before develop's reanimated-4/new-arch migration, which otherwise throws at render on current-develop JS). That is the script self-healing a stale image, NOT the false-negative footgun — let it rebuild (a reinstall keeps the data container, so login survives); do not pass `--skip-install` to dodge it. A JS-only change leaves Podfile.lock identical and still hits the fast path. If `$AGENT_SIM_UDID` is set but `select-ios-sim --accept-udid` HARD-FAILS ("not found") — you are a RESUMED session whose slot sim was recycled after completion — you cannot fix it in-process (no self-respawn). Report it and set `blocked = Yes` noting the operator must re-provision via `~/.config/agent-watcher/resume-task.sh --task-gid ` (allocates a fresh slot+sim+port and relaunches you with working env). Do NOT fall back to the master or a by-name sim. + + + + +A real on-simulator UI test that logs into the pre-provisioned test account, navigates to the Buy tab, requests a $500 quote, and captures a proof screenshot. PASS requires the screenshot to actually render the resolved quote. + +**Parallel-session env contract:** when the agent-watcher spawns this session as one of several parallel slots, it exports `$AGENT_SIM_UDID` (the slot's cloned simulator) and `$AGENT_METRO_PORT` (the slot's Metro port) into the shell. The scripts below honor them automatically — `select-ios-sim.sh --accept-udid "$AGENT_SIM_UDID"` skips name/runtime resolution and trusts the clone, and `ios-rn-build.sh` falls back to `$AGENT_SIM_UDID` / `$AGENT_METRO_PORT` when `--udid` / `--port` are not passed (forwarding a non-8081 port to `react-native run-ios`). On a manual run with neither var set, behavior is unchanged: resolve the iOS 18 sim by name and use Metro 8081. + +### 0a. Prerequisites (check, install if missing) + +- `xcrun -version` → Xcode CLT +- `maestro --version` → install with `curl -Ls "https://get.maestro.mobile.dev" | bash`, then add `$HOME/.maestro/bin` to PATH. maestro needs JDK 11+; Temurin 17 works. + +### 0b. Resolve + boot the simulator + +There can be multiple "iPhone 16 Pro Max" devices across runtimes. **Only the iOS 18 device holds the test accounts** (the `funded-test-accounts` roster; default login: the `agent` account). The iOS 26.x device does NOT. + +```bash +UDID=$(~/.cursor/skills/build-and-test/scripts/select-ios-sim.sh \ + --runtime "iOS 18" --device "iPhone 16 Pro Max" --boot) +``` + +If the script exits 2 (ambiguous), narrow `--runtime` (e.g. `"iOS 18.6"`). + +### 0c. Build + install + launch the app + +```bash +~/.cursor/skills/build-and-test/scripts/ios-rn-build.sh \ + --udid "$UDID" --bundle-id co.edgesecure.app --detach +~/.cursor/skills/build-and-test/scripts/ios-rn-build-wait.sh --udid "$UDID" # re-run while it exits 7 +``` + +Skips the full RN build when the app is already installed (cached path: seconds; a fresh build is usually just a few minutes — the Hermes prebuilt is prefetched). Pass `--force-rebuild` to always rebuild. + +### 0d. Run the maestro capture + +```bash +~/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh +``` + +Drives `maestro/buy-quote-input.yaml` (login → Buy → $500), then captures via an external simctl screenshot burst — keeping the last frame taken while the app was alive. Retries up to 5 cycles. Writes `/tmp/agent-mvp-buy-quote-screenshot.png` on success. + +### 0e. PASS / FAIL contract + +On capture-buy-quote.sh exit 0, the screenshot must visibly show **USD 500**, a non-empty **Amount BTC**, and the **`1 BTC = USD`** line. Emit: + +``` +build-and-test: PASS (iOS maestro — Buy $500 quote) +screenshot: /tmp/agent-mvp-buy-quote-screenshot.png +``` + +On exit nonzero, emit FAIL with the last 30 lines of the script's output: + +``` +build-and-test: FAIL — Buy $500 quote not captured + +``` + +Return success exit only on PASS. + +### 0f. Critical gotchas baked into the flow (do not "fix" them) + +Edge's RN keypad drops digits tapped too fast → wrong PIN → exponential lockout (465s → 914s → …). Each PIN digit tap in `buy-quote-input.yaml` uses `waitToSettleTimeoutMs`. Never speed it up. If a run logs "Invalid PIN: Account locked for N seconds", wait — do NOT tap. +On this debug build, `hideKeyboard` reliably triggers an RN Fabric text-measure SIGABRT. The flow leaves the keyboard up. Do not add `hideKeyboard` steps. +`assertVisible`/`extendedWaitUntil` traverse the a11y hierarchy on a poll loop, provoking the same Fabric crash on the Buy scene. The flow stops polling once the amount is entered; the capture script uses external simctl screenshots (no hierarchy traversal). + + + + +Run, in order: + +```bash +[ -d node_modules ] || ~/.cursor/skills/pm.sh install +npx tsc --noEmit +``` + +Emit PASS: +``` +build-and-test: PASS (tsc --noEmit clean) +``` + +Or FAIL with the last 30 lines of failing output: +``` +build-and-test: FAIL — exit + +``` + + + +```bash +[ -d node_modules ] || ~/.cursor/skills/pm.sh install +~/.cursor/skills/pm.sh run test +``` + +Same PASS/FAIL contract as step 1. + + + +Emit exactly: +``` +build-and-test: placeholder mode — no commands executed (repo shape not auto-detected). +``` +Return success. + + + +Re-run `select-ios-sim.sh` with a more specific `--runtime` (e.g. `"iOS 18.6"`). If still ambiguous, surface the list to the caller and set `blocked = Yes` on the Asana task with the candidate UDIDs and ask which to use. +Run `xcrun simctl shutdown all && xcrun simctl erase ` is destructive — do NOT run it. Set `blocked = Yes` with the boot error. +Re-run step 0b. If it fails twice, set `blocked = Yes`. +Acceptable in --yolo. Watch loop should NOT timeout the iteration during a known cold-build window — that's handled by /one-shot's `iOS prep budget` policy. +Emit FAIL with the maestro tail. Do NOT set `blocked = Yes` unless the failure mode is clearly a true-blocker (e.g. simulator died entirely, app uninstalled). A normal capture exhaustion is a real test FAIL the caller (watch loop) should react to. +Set `blocked = Yes` with the install error and a note about JDK requirement. +Set `blocked = Yes` — the test relies on the roster accounts (default: the `agent` account) being present on the sim image. Re-provisioning is a human step. + From c427d8be0816b01cf5825f444cfbb2583930c5b0 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 15:21:30 -0700 Subject: [PATCH 02/13] Add XCUITest flow interpreter for iOS One generic XCUITest bundle (xcuitest/EdgeFlowRunner) interprets Maestro flow YAML at run time. The host converts the YAML to JSON (maestro-yaml-to-json.rb) and passes it through xcodebuild test-without-building via TEST_RUNNER_ env, so the bundle is built once per Xcode build, runtime and source hash (xcuitest-build.sh, cached under ~/Library/Caches/edge-flow-runner, mkdir lock). xcuitest-run.sh drives one flow on one UDID with per-run JSON, log and result bundle under TMPDIR keyed by UDID, and no host port. It stops the slot's maestro MCP daemon by explicit PID first. The runner preflights every command before step 1 and fails naming the flow and step for anything unsupported. ${} and evalScript run in JavaScriptCore. A swizzle caps XCUITest quiescence waits (default 1s, --quiescence-cap or env EDGE_QUIESCENCE_CAP per flow, 0 skips) and logs each cap hit with its step. --- .../scripts/maestro-yaml-to-json.rb | 98 +++ .../build-and-test/scripts/xcuitest-build.sh | 78 +++ .../build-and-test/scripts/xcuitest-run.sh | 106 +++ .../EdgeFlowRunner.xcodeproj/project.pbxproj | 185 ++++++ .../xcschemes/EdgeFlowRunner.xcscheme | 43 ++ .../EdgeFlowRunner-Bridging-Header.h | 1 + .../xcuitest/EdgeFlowRunner/EdgeQuiescence.h | 25 + .../xcuitest/EdgeFlowRunner/EdgeQuiescence.m | 124 ++++ .../EdgeFlowRunner/FlowInterpreter.swift | 625 ++++++++++++++++++ .../EdgeFlowRunner/FlowPreflight.swift | 147 ++++ .../EdgeFlowRunner/FlowRunnerTests.swift | 65 ++ .../xcuitest/EdgeFlowRunner/FlowScript.swift | 148 +++++ 12 files changed, 1645 insertions(+) create mode 100755 .cursor/skills/build-and-test/scripts/maestro-yaml-to-json.rb create mode 100755 .cursor/skills/build-and-test/scripts/xcuitest-build.sh create mode 100755 .cursor/skills/build-and-test/scripts/xcuitest-run.sh create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/project.pbxproj create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/xcshareddata/xcschemes/EdgeFlowRunner.xcscheme create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeFlowRunner-Bridging-Header.h create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.h create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.m create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowRunnerTests.swift create mode 100644 .cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift diff --git a/.cursor/skills/build-and-test/scripts/maestro-yaml-to-json.rb b/.cursor/skills/build-and-test/scripts/maestro-yaml-to-json.rb new file mode 100755 index 00000000..08849888 --- /dev/null +++ b/.cursor/skills/build-and-test/scripts/maestro-yaml-to-json.rb @@ -0,0 +1,98 @@ +#!/usr/bin/ruby +# maestro-yaml-to-json.rb +# +# Converts a Maestro flow to the JSON the XCUITest interpreter runs. The runner +# never parses YAML: this script does, on the host, with the system Ruby's +# Psych (no install needed). +# +# Output: {"file": , "config":
, "commands": [...]}, plus +# a top-level "env" object from the --env K=V arguments (Maestro's `-e`). +# Every `env:` map becomes a list of [key, value] pairs so the runner +# evaluates entries in file order (a later entry may read an earlier one). +# Every `runFlow` that references a file is inlined as `"_flow"` inside its +# args (the same shape, recursively), so the runner gets one self-contained +# document. A runFlow path that needs run-time interpolation (`${...}`) cannot +# be resolved here: the script exits 1 naming the flow and the path. +require 'yaml' +require 'json' + +def fail_with(message) + warn "maestro-yaml-to-json: #{message}" + exit 1 +end + +def load_flow(path, stack) + path = File.expand_path(path) + fail_with("flow not found: #{path}") unless File.file?(path) + fail_with("runFlow cycle: #{(stack + [path]).join(' -> ')}") if stack.include?(path) + + docs = begin + YAML.load_stream(File.read(path)) + rescue Psych::SyntaxError => e + fail_with("YAML error in #{path}: #{e.message}") + end + docs = docs.compact + config, commands = + case docs.length + when 1 then docs[0].is_a?(Array) ? [{}, docs[0]] : fail_with("#{path}: expected a command list") + when 2 then docs + else fail_with("#{path}: expected a config document and a command list, found #{docs.length} documents") + end + fail_with("#{path}: config is not a map") unless config.is_a?(Hash) + fail_with("#{path}: commands are not a list") unless commands.is_a?(Array) + config = config.merge('env' => env_pairs(config['env'], path)) if config.key?('env') + + dir = File.dirname(path) + { + 'file' => path, + 'config' => config, + 'commands' => commands.map { |c| inline_command(c, dir, path, stack + [path]) } + } +end + +# Walks one command, inlining runFlow file references and recursing into the +# nested command lists of runFlow, repeat and retry. +def inline_command(command, dir, path, stack) + return command unless command.is_a?(Hash) && command.length == 1 + + name, args = command.first + case name + when 'runFlow', 'retry' + args = { 'file' => args } if args.is_a?(String) + return command unless args.is_a?(Hash) + + args = args.dup + args['env'] = env_pairs(args['env'], path) if args.key?('env') + if args['file'] + file = args['file'].to_s + fail_with("#{path}: #{name} file '#{file}' needs run-time interpolation, which the host converter cannot resolve") if file.include?('${') + args['_flow'] = load_flow(File.expand_path(file, dir), stack) + end + args['commands'] = args['commands'].map { |c| inline_command(c, dir, path, stack) } if args['commands'].is_a?(Array) + { name => args } + when 'repeat' + return command unless args.is_a?(Hash) && args['commands'].is_a?(Array) + + { name => args.merge('commands' => args['commands'].map { |c| inline_command(c, dir, path, stack) }) } + else + command + end +end + +def env_pairs(env, path) + fail_with("#{path}: env is not a map") unless env.is_a?(Hash) + env.map { |key, value| [key.to_s, value.is_a?(String) ? value : value.to_s] } +end + +usage = 'usage: maestro-yaml-to-json.rb [--env KEY=VALUE ...] ' +cli_env = {} +args = ARGV.dup +while args.first == '--env' + args.shift + pair = args.shift.to_s + fail_with("--env expects KEY=VALUE, got '#{pair}'") unless pair.include?('=') + key, value = pair.split('=', 2) + cli_env[key] = value +end +fail_with(usage) unless args.length == 1 +puts JSON.generate(load_flow(args[0], []).merge('env' => cli_env)) diff --git a/.cursor/skills/build-and-test/scripts/xcuitest-build.sh b/.cursor/skills/build-and-test/scripts/xcuitest-build.sh new file mode 100755 index 00000000..fc5f1a42 --- /dev/null +++ b/.cursor/skills/build-and-test/scripts/xcuitest-build.sh @@ -0,0 +1,78 @@ +#!/usr/bin/env bash +# xcuitest-build.sh: build the EdgeFlowRunner XCUITest bundle once and cache it. +# +# The runner is generic (it interprets flow JSON at run time), so one build +# serves every flow and every slot sim on the same runtime. The cache key is +# Xcode build + target iOS runtime + a hash of the runner sources; a source +# change or Xcode upgrade builds a fresh copy. Concurrent callers (parallel +# slots) serialize on a mkdir lock and reuse the first build. +# +# Usage: xcuitest-build.sh [--udid ] [--force] +# --udid defaults to $AGENT_SIM_UDID; its runtime picks the cache entry. +# Output: last line is `XCTESTRUN=`. +# Exit: 0 built or cached; 1 usage or build failure. + +set -euo pipefail + +UDID="${AGENT_SIM_UDID:-}" +FORCE=0 +while [[ $# -gt 0 ]]; do + case "$1" in + --udid) UDID="$2"; shift 2 ;; + --force) FORCE=1; shift ;; + *) echo "Unknown arg: $1" >&2; exit 1 ;; + esac +done +[[ -n "$UDID" ]] || { echo "xcuitest-build: no --udid and \$AGENT_SIM_UDID unset" >&2; exit 1; } + +SRC_DIR="$(cd "$(dirname "$0")/../xcuitest" && pwd)" +RUNTIME="$(xcrun simctl list devices -j | /usr/bin/ruby -rjson -e ' + udid = ARGV[0] + JSON.parse($stdin.read)["devices"].each { |rt, devs| devs.each { |d| (puts rt.split(".").last; exit) if d["udid"] == udid } } + exit 1' "$UDID")" || { echo "xcuitest-build: sim $UDID not found" >&2; exit 1; } +XCODE_BUILD="$(xcodebuild -version | awk '/Build version/ {print $3}')" +SRC_HASH="$(cd "$SRC_DIR" && find . -type f \( -name '*.swift' -o -name '*.m' -o -name '*.h' -o -name '*.pbxproj' -o -name '*.xcscheme' \) -print0 | sort -z | xargs -0 shasum | shasum | cut -c1-12)" +CACHE_ROOT="$HOME/Library/Caches/edge-flow-runner" +CACHE="$CACHE_ROOT/$XCODE_BUILD-$RUNTIME-$SRC_HASH" +LOCK="$CACHE_ROOT/.lock-$XCODE_BUILD-$RUNTIME" +mkdir -p "$CACHE_ROOT" + +find_xctestrun() { find "$CACHE/Build/Products" -maxdepth 1 -name '*.xctestrun' 2>/dev/null | head -1; } + +if [[ "$FORCE" == 0 && -f "$CACHE/.done" ]]; then + echo "xcuitest-build: cached ($CACHE)" + echo "XCTESTRUN=$(find_xctestrun)" + exit 0 +fi + +# mkdir is atomic; a lock older than 15 minutes belongs to a dead build. +waited=0 +until mkdir "$LOCK" 2>/dev/null; do + if [[ -n "$(find "$LOCK" -maxdepth 0 -mmin +15 2>/dev/null)" ]]; then + echo "xcuitest-build: removing stale lock $LOCK" >&2 + rmdir "$LOCK" 2>/dev/null || true + continue + fi + (( waited % 30 == 0 )) && echo "xcuitest-build: waiting for another build ($LOCK)" + sleep 2; waited=$((waited + 2)) + (( waited < 900 )) || { echo "xcuitest-build: lock wait timed out" >&2; exit 1; } +done +trap 'rmdir "$LOCK" 2>/dev/null || true' EXIT + +if [[ "$FORCE" == 0 && -f "$CACHE/.done" ]]; then + echo "xcuitest-build: built by a concurrent caller ($CACHE)" + echo "XCTESTRUN=$(find_xctestrun)" + exit 0 +fi + +rm -rf "$CACHE" +echo "xcuitest-build: building runner for $RUNTIME (Xcode $XCODE_BUILD) into $CACHE" +if ! xcodebuild build-for-testing \ + -project "$SRC_DIR/EdgeFlowRunner.xcodeproj" -scheme EdgeFlowRunner \ + -destination "id=$UDID" -derivedDataPath "$CACHE" -quiet > "$CACHE_ROOT/build-$RUNTIME.log" 2>&1; then + grep -E "error:" "$CACHE_ROOT/build-$RUNTIME.log" | head -20 >&2 + echo "xcuitest-build: build failed, full log $CACHE_ROOT/build-$RUNTIME.log" >&2 + exit 1 +fi +touch "$CACHE/.done" +echo "XCTESTRUN=$(find_xctestrun)" diff --git a/.cursor/skills/build-and-test/scripts/xcuitest-run.sh b/.cursor/skills/build-and-test/scripts/xcuitest-run.sh new file mode 100755 index 00000000..5ed0664c --- /dev/null +++ b/.cursor/skills/build-and-test/scripts/xcuitest-run.sh @@ -0,0 +1,106 @@ +#!/usr/bin/env bash +# xcuitest-run.sh: run one Maestro flow YAML on a slot sim with the native +# XCUITest interpreter (EdgeFlowRunner) instead of the Maestro driver. +# +# Pipeline: maestro-yaml-to-json.rb inlines every runFlow/retry file into one +# JSON document; xcuitest-build.sh returns the cached runner; one +# `xcodebuild test-without-building` drives the sim named by UDID. Nothing +# binds a host port, so any number of slots can run side by side. The runner +# preflights the whole flow tree and fails before step 1 on any command it +# does not implement; there is no Maestro fallback. +# +# Before driving, the slot's maestro MCP daemon (the java process started +# with `--device ... mcp`) is stopped by explicit PID, and the +# Maestro XCTest driver app is terminated on this sim only: two automation +# sessions on one sim fight over the accessibility snapshot. +# +# Usage: xcuitest-run.sh --flow [--udid ] [--env K=V ...] +# [--quiescence-cap ] [--animations off|fast|on] [--keep-mcp] +# --udid defaults to $AGENT_SIM_UDID. +# --quiescence-cap: cap on XCUITest's app-idle wait per event (default 1, +# 0 disables the wait). A flow can override it with env EDGE_QUIESCENCE_CAP. +# --animations: test-mode animation switch passed to the app as +# `-EdgeTestAnimations ` on launchApp and persisted in the sim's +# defaults for relaunches (default off; `on` clears it). +# Relative takeScreenshot paths resolve against the current directory, as +# with `maestro test`. +# Output: `[edge-flow]` step lines, then RESULT=passed|failed, WALL=, +# RUN_DIR=. +# Exit: 0 flow passed; 1 usage/setup error; 2 flow failed. + +set -uo pipefail + +UDID="${AGENT_SIM_UDID:-}" +FLOW="" +CAP="1" +ANIMATIONS="off" +KEEP_MCP=0 +BUNDLE_ID="co.edgesecure.app" +ENV_ARGS=() +while [[ $# -gt 0 ]]; do + case "$1" in + --flow) FLOW="$2"; shift 2 ;; + --udid) UDID="$2"; shift 2 ;; + --env) ENV_ARGS+=(--env "$2"); shift 2 ;; + --quiescence-cap) CAP="$2"; shift 2 ;; + --animations) ANIMATIONS="$2"; shift 2 ;; + --keep-mcp) KEEP_MCP=1; shift ;; + *) echo "Unknown arg: $1" >&2; exit 1 ;; + esac +done +[[ -n "$FLOW" && -f "$FLOW" ]] || { echo "xcuitest-run: --flow required" >&2; exit 1; } +[[ -n "$UDID" ]] || { echo "xcuitest-run: no --udid and \$AGENT_SIM_UDID unset" >&2; exit 1; } +case "$ANIMATIONS" in off|fast|on) ;; *) echo "xcuitest-run: --animations must be off, fast or on" >&2; exit 1 ;; esac + +SCRIPTS="$(cd "$(dirname "$0")" && pwd)" +START=$(date +%s) +FLOW_NAME="$(basename "$FLOW" .yaml)" +RUN_DIR="${TMPDIR:-/tmp}/edge-flow-runs/$UDID/$(date +%Y%m%d-%H%M%S)-$FLOW_NAME-$$" +mkdir -p "$RUN_DIR" + +/usr/bin/ruby "$SCRIPTS/maestro-yaml-to-json.rb" ${ENV_ARGS[@]+"${ENV_ARGS[@]}"} "$FLOW" > "$RUN_DIR/flow.json" \ + || { echo "xcuitest-run: could not convert $FLOW" >&2; exit 1; } + +BUILD_OUT="$("$SCRIPTS/xcuitest-build.sh" --udid "$UDID")" || { echo "$BUILD_OUT" >&2; exit 1; } +XCTESTRUN="$(printf '%s\n' "$BUILD_OUT" | sed -n 's/^XCTESTRUN=//p' | tail -1)" +[[ -f "$XCTESTRUN" ]] || { echo "xcuitest-run: no xctestrun from xcuitest-build.sh" >&2; exit 1; } + +if [[ "$KEEP_MCP" == 0 ]]; then + for pid in $(pgrep -f "maestro.cli.AppKt --device $UDID .*mcp" || true); do + echo "xcuitest-run: stopping maestro MCP daemon pid $pid (bound to $UDID)" + kill "$pid" 2>/dev/null || true + done + xcrun simctl terminate "$UDID" dev.mobile.maestro-driver-iosUITests.xctrunner >/dev/null 2>&1 || true +fi + +if [[ "$ANIMATIONS" == on ]]; then + xcrun simctl spawn "$UDID" defaults delete "$BUNDLE_ID" EdgeTestAnimations >/dev/null 2>&1 || true +else + xcrun simctl spawn "$UDID" defaults write "$BUNDLE_ID" EdgeTestAnimations "$ANIMATIONS" >/dev/null 2>&1 || true +fi + +echo "xcuitest-run: $FLOW on $UDID (cap ${CAP}s, animations $ANIMATIONS), run dir $RUN_DIR" +TEST_RUNNER_EDGE_FLOW_FILE="$RUN_DIR/flow.json" \ +TEST_RUNNER_EDGE_FLOW_CWD="$(pwd)" \ +TEST_RUNNER_EDGE_QUIESCENCE_CAP="$CAP" \ +TEST_RUNNER_EDGE_TEST_ANIMATIONS="$ANIMATIONS" \ + xcodebuild test-without-building -xctestrun "$XCTESTRUN" -destination "id=$UDID" \ + -only-testing:EdgeFlowRunner/FlowRunnerTests/testFlow -parallel-testing-enabled NO \ + -resultBundlePath "$RUN_DIR/result.xcresult" > "$RUN_DIR/xcodebuild.log" 2>&1 & +XC_PID=$! +# Stream step lines as they land (the log is the full record). +tail -n +1 -f "$RUN_DIR/xcodebuild.log" 2>/dev/null > >(grep --line-buffered '\[edge-flow\]') & +TAIL_PID=$! +wait "$XC_PID" +STATUS=$? +sleep 0.5 +kill "$TAIL_PID" 2>/dev/null +wait "$TAIL_PID" 2>/dev/null + +if [[ "$STATUS" != 0 ]]; then + grep -E "error:|Failing tests|\*\* TEST" "$RUN_DIR/xcodebuild.log" | grep -v '\[edge-flow\]' | head -10 +fi +echo "RESULT=$([[ "$STATUS" == 0 ]] && echo passed || echo failed)" +echo "WALL=$(( $(date +%s) - START ))" +echo "RUN_DIR=$RUN_DIR" +[[ "$STATUS" == 0 ]] || exit 2 diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/project.pbxproj b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/project.pbxproj new file mode 100644 index 00000000..e22e8730 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/project.pbxproj @@ -0,0 +1,185 @@ +// !$*UTF8*$! +{ + archiveVersion = 1; + classes = { + }; + objectVersion = 56; + objects = { + +/* Begin PBXBuildFile section */ + E0F1000000000000000000B1 /* FlowRunnerTests.swift in Sources */ = {isa = PBXBuildFile; fileRef = E0F1000000000000000000A1 /* FlowRunnerTests.swift */; }; + E0F1000000000000000000B2 /* FlowInterpreter.swift in Sources */ = {isa = PBXBuildFile; fileRef = E0F1000000000000000000A2 /* FlowInterpreter.swift */; }; + E0F1000000000000000000B3 /* FlowPreflight.swift in Sources */ = {isa = PBXBuildFile; fileRef = E0F1000000000000000000A3 /* FlowPreflight.swift */; }; + E0F1000000000000000000B4 /* FlowScript.swift in Sources */ = {isa = PBXBuildFile; fileRef = E0F1000000000000000000A4 /* FlowScript.swift */; }; + E0F1000000000000000000B5 /* EdgeQuiescence.m in Sources */ = {isa = PBXBuildFile; fileRef = E0F1000000000000000000A5 /* EdgeQuiescence.m */; }; +/* End PBXBuildFile section */ + +/* Begin PBXFileReference section */ + E0F1000000000000000000A1 /* FlowRunnerTests.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = FlowRunnerTests.swift; sourceTree = ""; }; + E0F1000000000000000000A2 /* FlowInterpreter.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = FlowInterpreter.swift; sourceTree = ""; }; + E0F1000000000000000000A3 /* FlowPreflight.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = FlowPreflight.swift; sourceTree = ""; }; + E0F1000000000000000000A4 /* FlowScript.swift */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.swift; path = FlowScript.swift; sourceTree = ""; }; + E0F1000000000000000000A5 /* EdgeQuiescence.m */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.c.objc; path = EdgeQuiescence.m; sourceTree = ""; }; + E0F1000000000000000000A6 /* EdgeQuiescence.h */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.c.h; path = EdgeQuiescence.h; sourceTree = ""; }; + E0F1000000000000000000A7 /* EdgeFlowRunner-Bridging-Header.h */ = {isa = PBXFileReference; lastKnownFileType = sourcecode.c.h; path = "EdgeFlowRunner-Bridging-Header.h"; sourceTree = ""; }; + E0F1000000000000000000A8 /* EdgeFlowRunner.xctest */ = {isa = PBXFileReference; explicitFileType = wrapper.cfbundle; includeInIndex = 0; path = EdgeFlowRunner.xctest; sourceTree = BUILT_PRODUCTS_DIR; }; +/* End PBXFileReference section */ + +/* Begin PBXFrameworksBuildPhase section */ + E0F1000000000000000000C2 /* Frameworks */ = { + isa = PBXFrameworksBuildPhase; + buildActionMask = 2147483647; + files = ( + ); + runOnlyForDeploymentPostprocessing = 0; + }; +/* End PBXFrameworksBuildPhase section */ + +/* Begin PBXGroup section */ + E0F1000000000000000000D1 = { + isa = PBXGroup; + children = ( + E0F1000000000000000000D2 /* EdgeFlowRunner */, + E0F1000000000000000000D3 /* Products */, + ); + sourceTree = ""; + }; + E0F1000000000000000000D2 /* EdgeFlowRunner */ = { + isa = PBXGroup; + children = ( + E0F1000000000000000000A1 /* FlowRunnerTests.swift */, + E0F1000000000000000000A2 /* FlowInterpreter.swift */, + E0F1000000000000000000A3 /* FlowPreflight.swift */, + E0F1000000000000000000A4 /* FlowScript.swift */, + E0F1000000000000000000A6 /* EdgeQuiescence.h */, + E0F1000000000000000000A5 /* EdgeQuiescence.m */, + E0F1000000000000000000A7 /* EdgeFlowRunner-Bridging-Header.h */, + ); + path = EdgeFlowRunner; + sourceTree = ""; + }; + E0F1000000000000000000D3 /* Products */ = { + isa = PBXGroup; + children = ( + E0F1000000000000000000A8 /* EdgeFlowRunner.xctest */, + ); + name = Products; + sourceTree = ""; + }; +/* End PBXGroup section */ + +/* Begin PBXNativeTarget section */ + E0F1000000000000000000E1 /* EdgeFlowRunner */ = { + isa = PBXNativeTarget; + buildConfigurationList = E0F1000000000000000000F2 /* Build configuration list for PBXNativeTarget "EdgeFlowRunner" */; + buildPhases = ( + E0F1000000000000000000C1 /* Sources */, + E0F1000000000000000000C2 /* Frameworks */, + ); + buildRules = ( + ); + dependencies = ( + ); + name = EdgeFlowRunner; + productName = EdgeFlowRunner; + productReference = E0F1000000000000000000A8 /* EdgeFlowRunner.xctest */; + productType = "com.apple.product-type.bundle.ui-testing"; + }; +/* End PBXNativeTarget section */ + +/* Begin PBXProject section */ + E0F1000000000000000000E0 /* Project object */ = { + isa = PBXProject; + attributes = { + BuildIndependentTargetsInParallel = 1; + LastSwiftUpdateCheck = 2600; + LastUpgradeCheck = 2600; + }; + buildConfigurationList = E0F1000000000000000000F1 /* Build configuration list for PBXProject "EdgeFlowRunner" */; + compatibilityVersion = "Xcode 14.0"; + developmentRegion = en; + hasScannedForEncodings = 0; + knownRegions = ( + en, + Base, + ); + mainGroup = E0F1000000000000000000D1; + productRefGroup = E0F1000000000000000000D3 /* Products */; + projectDirPath = ""; + projectRoot = ""; + targets = ( + E0F1000000000000000000E1 /* EdgeFlowRunner */, + ); + }; +/* End PBXProject section */ + +/* Begin PBXSourcesBuildPhase section */ + E0F1000000000000000000C1 /* Sources */ = { + isa = PBXSourcesBuildPhase; + buildActionMask = 2147483647; + files = ( + E0F1000000000000000000B1 /* FlowRunnerTests.swift in Sources */, + E0F1000000000000000000B2 /* FlowInterpreter.swift in Sources */, + E0F1000000000000000000B3 /* FlowPreflight.swift in Sources */, + E0F1000000000000000000B4 /* FlowScript.swift in Sources */, + E0F1000000000000000000B5 /* EdgeQuiescence.m in Sources */, + ); + runOnlyForDeploymentPostprocessing = 0; + }; +/* End PBXSourcesBuildPhase section */ + +/* Begin XCBuildConfiguration section */ + E0F1000000000000000000F3 /* Debug */ = { + isa = XCBuildConfiguration; + buildSettings = { + CLANG_ENABLE_MODULES = YES; + CLANG_ENABLE_OBJC_ARC = YES; + CODE_SIGNING_ALLOWED = NO; + DEBUG_INFORMATION_FORMAT = dwarf; + ENABLE_TESTABILITY = YES; + GCC_OPTIMIZATION_LEVEL = 0; + IPHONEOS_DEPLOYMENT_TARGET = 16.0; + ONLY_ACTIVE_ARCH = YES; + SDKROOT = iphoneos; + SUPPORTED_PLATFORMS = "iphonesimulator iphoneos"; + SWIFT_OPTIMIZATION_LEVEL = "-Onone"; + SWIFT_VERSION = 5.0; + }; + name = Debug; + }; + E0F1000000000000000000F4 /* Debug */ = { + isa = XCBuildConfiguration; + buildSettings = { + CODE_SIGN_STYLE = Manual; + GENERATE_INFOPLIST_FILE = YES; + PRODUCT_BUNDLE_IDENTIFIER = app.edge.EdgeFlowRunner; + PRODUCT_NAME = "$(TARGET_NAME)"; + SWIFT_EMIT_LOC_STRINGS = NO; + SWIFT_OBJC_BRIDGING_HEADER = "EdgeFlowRunner/EdgeFlowRunner-Bridging-Header.h"; + TARGETED_DEVICE_FAMILY = "1,2"; + }; + name = Debug; + }; +/* End XCBuildConfiguration section */ + +/* Begin XCConfigurationList section */ + E0F1000000000000000000F1 /* Build configuration list for PBXProject "EdgeFlowRunner" */ = { + isa = XCConfigurationList; + buildConfigurations = ( + E0F1000000000000000000F3 /* Debug */, + ); + defaultConfigurationIsVisible = 0; + defaultConfigurationName = Debug; + }; + E0F1000000000000000000F2 /* Build configuration list for PBXNativeTarget "EdgeFlowRunner" */ = { + isa = XCConfigurationList; + buildConfigurations = ( + E0F1000000000000000000F4 /* Debug */, + ); + defaultConfigurationIsVisible = 0; + defaultConfigurationName = Debug; + }; +/* End XCConfigurationList section */ + }; + rootObject = E0F1000000000000000000E0 /* Project object */; +} diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/xcshareddata/xcschemes/EdgeFlowRunner.xcscheme b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/xcshareddata/xcschemes/EdgeFlowRunner.xcscheme new file mode 100644 index 00000000..fe34ab9a --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner.xcodeproj/xcshareddata/xcschemes/EdgeFlowRunner.xcscheme @@ -0,0 +1,43 @@ + + + + + + + + + + + + + + + + + + + diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeFlowRunner-Bridging-Header.h b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeFlowRunner-Bridging-Header.h new file mode 100644 index 00000000..b784a065 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeFlowRunner-Bridging-Header.h @@ -0,0 +1 @@ +#import "EdgeQuiescence.h" diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.h b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.h new file mode 100644 index 00000000..8586bdc3 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.h @@ -0,0 +1,25 @@ +#import + +NS_ASSUME_NONNULL_BEGIN + +// Hard cap on XCUITest's app-quiescence wait (the wait XCTest runs before and +// after every synthesized event). Modeled on WebDriverAgent's +// XCUIApplicationProcess+FBQuiescence: the private wait methods are swizzled +// to call the original under a bounded _XCTSetApplicationStateTimeout. +@interface EdgeQuiescence : NSObject + +// Seconds. Default 1. 0 skips the quiescence wait entirely. +@property (class, nonatomic) NSTimeInterval cap; + +// Label of the flow step currently running, printed with each cap hit. +@property (class, nonatomic, copy) NSString *currentStep; + +// Number of waits that ran into the cap during this process. +@property (class, nonatomic, readonly) NSInteger capHits; + +// Swizzles the wait methods. Safe to call more than once. ++ (void)install; + +@end + +NS_ASSUME_NONNULL_END diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.m b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.m new file mode 100644 index 00000000..ba225f7d --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/EdgeQuiescence.m @@ -0,0 +1,124 @@ +#import "EdgeQuiescence.h" + +#import +#import + +typedef void (*SetTimeoutFn)(double); +typedef double (*GetTimeoutFn)(void); + +static NSTimeInterval gCap = 1; +static NSString *gCurrentStep = @""; +static NSInteger gCapHits = 0; +static NSInteger gDepth = 0; +static NSRecursiveLock *gLock; +static SetTimeoutFn gSetTimeout; +static GetTimeoutFn gGetTimeout; + +static IMP gOrigIdle; +static IMP gOrigPreEvent; +static IMP gOrigActivity; + +// The application-state timeout is process-wide, so the set/wait/restore +// sequence runs under one lock. Nested waits (one wait method calling another) +// only apply the cap at the outermost level. +static void capped(NSString *name, void (^original)(void)) +{ + if (gCap <= 0) return; + [gLock lock]; + gDepth += 1; + if (gDepth > 1) { + original(); + gDepth -= 1; + [gLock unlock]; + return; + } + double previous = gGetTimeout != NULL ? gGetTimeout() : 0; + if (gSetTimeout != NULL) gSetTimeout(gCap); + NSDate *start = [NSDate date]; + @try { + original(); + } @finally { + if (gSetTimeout != NULL) gSetTimeout(previous); + NSTimeInterval elapsed = -[start timeIntervalSinceNow]; + if (elapsed >= gCap * 0.95) { + gCapHits += 1; + NSLog(@"[edge-flow] quiescence cap hit (%.2fs >= %.2fs) %@ step: %@", elapsed, gCap, name, gCurrentStep); + } + gDepth -= 1; + [gLock unlock]; + } +} + +static void swizzledIdle(id self, SEL _cmd, BOOL includingAnimations) +{ + capped(@"waitForQuiescenceIncludingAnimationsIdle:", ^{ + ((void (*)(id, SEL, BOOL))gOrigIdle)(self, _cmd, includingAnimations); + }); +} + +static void swizzledPreEvent(id self, SEL _cmd, BOOL includingAnimations, BOOL isPreEvent) +{ + capped(@"waitForQuiescenceIncludingAnimationsIdle:isPreEvent:", ^{ + ((void (*)(id, SEL, BOOL, BOOL))gOrigPreEvent)(self, _cmd, includingAnimations, isPreEvent); + }); +} + +// usingActivity is a BOOL, not an XCTActivity (Xcode 26 passes 0/1 there). +static void swizzledActivity(id self, SEL _cmd, BOOL includingAnimations, BOOL usingActivity, BOOL isPreEvent) +{ + capped(@"waitForQuiescenceIncludingAnimationsIdle:usingActivity:isPreEvent:", ^{ + ((void (*)(id, SEL, BOOL, BOOL, BOOL))gOrigActivity)(self, _cmd, includingAnimations, usingActivity, isPreEvent); + }); +} + +// Swaps only when every argument is a BOOL (`B` or `c`) and the method +// returns void, so a changed private signature skips the cap instead of +// crashing the runner. +static IMP swap(Class cls, NSString *selectorName, IMP replacement, unsigned int boolArgs) +{ + Method method = class_getInstanceMethod(cls, NSSelectorFromString(selectorName)); + if (method == NULL) return NULL; + unsigned int count = method_getNumberOfArguments(method); + char returnType[8]; + method_getReturnType(method, returnType, sizeof(returnType)); + BOOL matches = count == boolArgs + 2 && returnType[0] == 'v'; + for (unsigned int i = 2; matches && i < count; i++) { + char argType[8]; + method_getArgumentType(method, i, argType, sizeof(argType)); + matches = argType[0] == 'B' || argType[0] == 'c'; + } + if (!matches) { + NSLog(@"[edge-flow] quiescence cap skips %@ (unexpected signature %s)", selectorName, method_getTypeEncoding(method)); + return NULL; + } + return method_setImplementation(method, replacement); +} + +@implementation EdgeQuiescence + ++ (NSTimeInterval)cap { return gCap; } ++ (void)setCap:(NSTimeInterval)cap { gCap = cap; } ++ (NSString *)currentStep { return gCurrentStep; } ++ (void)setCurrentStep:(NSString *)step { gCurrentStep = [step copy]; } ++ (NSInteger)capHits { return gCapHits; } + ++ (void)install +{ + static dispatch_once_t once; + dispatch_once(&once, ^{ + gLock = [NSRecursiveLock new]; + gSetTimeout = (SetTimeoutFn)dlsym(RTLD_DEFAULT, "_XCTSetApplicationStateTimeout"); + gGetTimeout = (GetTimeoutFn)dlsym(RTLD_DEFAULT, "_XCTApplicationStateTimeout"); + Class cls = NSClassFromString(@"XCUIApplicationProcess"); + if (cls == Nil || gSetTimeout == NULL) { + NSLog(@"[edge-flow] quiescence cap unavailable (XCUIApplicationProcess or _XCTSetApplicationStateTimeout missing)"); + return; + } + gOrigIdle = swap(cls, @"waitForQuiescenceIncludingAnimationsIdle:", (IMP)swizzledIdle, 1); + gOrigPreEvent = swap(cls, @"waitForQuiescenceIncludingAnimationsIdle:isPreEvent:", (IMP)swizzledPreEvent, 2); + gOrigActivity = swap(cls, @"waitForQuiescenceIncludingAnimationsIdle:usingActivity:isPreEvent:", (IMP)swizzledActivity, 3); + NSLog(@"[edge-flow] quiescence cap installed (idle=%d preEvent=%d activity=%d)", gOrigIdle != NULL, gOrigPreEvent != NULL, gOrigActivity != NULL); + }); +} + +@end diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift new file mode 100644 index 00000000..0cba7255 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift @@ -0,0 +1,625 @@ +import Foundation +import XCTest + +/// Run-wide settings, read from the runner's environment (xcodebuild passes +/// `TEST_RUNNER_` through as ``). +struct RunOptions { + var quiescenceCap: TimeInterval = 1 + /// `off` or `fast` is passed to the app as `-EdgeTestAnimations ` on + /// launchApp; `on` passes nothing. + var animations = "off" + /// Maestro's defaults: 17s element lookup, 7s for optional lookups and + /// `when:` conditions. + var lookupTimeout: TimeInterval = 17 + var optionalLookupTimeout: TimeInterval = 7 + var workingDirectory = FileManager.default.currentDirectoryPath + + init(environment: [String: String]) { + if let cap = environment["EDGE_QUIESCENCE_CAP"].flatMap(Double.init) { quiescenceCap = cap } + if let mode = environment["EDGE_TEST_ANIMATIONS"], !mode.isEmpty { animations = mode } + if let ms = environment["EDGE_LOOKUP_MS"].flatMap(Double.init) { lookupTimeout = ms / 1000 } + if let ms = environment["EDGE_OPTIONAL_LOOKUP_MS"].flatMap(Double.init) { optionalLookupTimeout = ms / 1000 } + if let cwd = environment["EDGE_FLOW_CWD"], !cwd.isEmpty { workingDirectory = cwd } + } +} + +/// Element selector: Maestro `text` and `id` are case-insensitive regexes +/// (dot matches newline) that must match the whole attribute, or equal it +/// literally. `text` is checked against label, value and placeholderValue. +struct Selector: CustomStringConvertible { + var text: String? + var id: String? + var index: Int? + + var description: String { + var parts: [String] = [] + if let text = text { parts.append("text=\"\(text)\"") } + if let id = id { parts.append("id=\"\(id)\"") } + if let index = index { parts.append("index=\(index)") } + return parts.joined(separator: " ") + } +} + +final class FlowInterpreter { + private let options: RunOptions + private let script = FlowScript() + private var appId = "co.edgesecure.app" + private var app = XCUIApplication(bundleIdentifier: "co.edgesecure.app") + private var screenBounds: CGRect? + private var lastInteraction = Date() + private let started = Date() + + init(options: RunOptions) { + self.options = options + EdgeQuiescence.cap = options.quiescenceCap + } + + // MARK: Flows + + func run(_ document: [String: Any]) throws { + let cliEnv = (document["env"] as? [String: Any] ?? [:]).mapValues { "\($0)" } + for (key, value) in cliEnv { script.set(key, value) } + try runFlow(document, callerEnv: [], skipHeaderKeys: Set(cliEnv.keys)) + } + + /// Runs one flow document in its own env scope. Caller env applies first + /// and the flow's header `env:` is evaluated after it, which is Maestro + /// 2.x behavior (a subflow's header shadows the caller's value, hence the + /// `KEY: ${KEY || default}` idiom in the library flows). + private func runFlow(_ flow: [String: Any], callerEnv: [(String, String)], skipHeaderKeys: Set = []) throws { + let flowName = ((flow["file"] as? String).map { ($0 as NSString).lastPathComponent }) ?? "" + let config = flow["config"] as? [String: Any] ?? [:] + let scope = EnvScope(script) + let previousCap = EdgeQuiescence.cap + defer { + scope.restore() + EdgeQuiescence.cap = previousCap + } + for (key, value) in callerEnv { scope.set(key, value) } + for pair in config["env"] as? [[Any]] ?? [] { + guard pair.count == 2, let key = pair[0] as? String else { continue } + if skipHeaderKeys.contains(key) { continue } + if !script.isDefined(key) { scope.declare(key) } + scope.set(key, try script.interpolate("\(pair[1])")) + } + if let cap = script.value("EDGE_QUIESCENCE_CAP")?.toString().flatMap(Double.init) { + EdgeQuiescence.cap = cap + } + if let appId = config["appId"] as? String { + let bundleId = try script.interpolate(appId) + if bundleId != appId { selectApp(bundleId) } + } + try runCommands(flow["commands"] as? [Any] ?? [], flowName: flowName) + } + + private func runCommands(_ commands: [Any], flowName: String) throws { + for (offset, command) in commands.enumerated() { + try execute(command, at: "\(flowName) #\(offset + 1)") + } + } + + private func execute(_ command: Any, at: String) throws { + let name: String + let args: Any? + if let bare = command as? String { + name = bare + args = nil + } else if let map = command as? [String: Any], let entry = map.first { + name = entry.key + args = entry.value is NSNull ? nil : entry.value + } else { + throw FlowError("\(at): malformed command") + } + let map = args as? [String: Any] + let optional = map?["optional"] as? Bool ?? false + let step = "\(at) \(name) \(summary(args))" + EdgeQuiescence.currentStep = step + let begin = Date() + do { + let note = try dispatch(name, args, optional: optional) + log(step, begin, note ?? "ok") + } catch let error as FlowError where optional { + log(step, begin, "skipped (optional): \(error)") + } catch { + log(step, begin, "FAILED: \(error)") + throw FlowError("\(step): \(error)") + } + } + + // MARK: Commands + + private func dispatch(_ name: String, _ args: Any?, optional: Bool) throws -> String? { + let map = args as? [String: Any] ?? [:] + switch name { + case "launchApp": + if let appId = map["appId"] as? String { selectApp(try script.interpolate(appId)) } + if map["stopApp"] as? Bool == false, app.state == .runningForeground || app.state == .runningBackground { + app.activate() + } else { + app.launchArguments = options.animations == "on" ? [] : ["-EdgeTestAnimations", options.animations] + app.launch() + } + screenBounds = nil + interacted() + case "stopApp": + if let appId = map["appId"] as? String { selectApp(try script.interpolate(appId)) } + app.terminate() + interacted() + case "tapOn": + let selector = try parseSelector(args) + let timeout = optional ? options.optionalLookupTimeout : options.lookupTimeout + guard let frame = try waitForElement(selector, timeout: timeout) else { + throw FlowError("element not found: \(selector)") + } + let before = map["retryTapIfNoChange"] as? Bool == true ? screenPixels() : nil + tap(frame) + if let before = before, screenPixels() == before { tap(frame) } + if let ms = try optionalNumber(map["waitToSettleTimeoutMs"]) { waitForSettle(timeout: ms / 1000) } + case "assertVisible": + let selector = try parseSelector(args) + let timeout = adjusted(optional ? options.optionalLookupTimeout : options.lookupTimeout) + if try waitForElement(selector, timeout: timeout) == nil { + throw FlowError("assertion failed: \(selector) is not visible") + } + case "assertNotVisible": + let selector = try parseSelector(args) + let timeout = adjusted(optional ? options.optionalLookupTimeout : options.lookupTimeout) + if try !waitForAbsence(selector, timeout: timeout) { + throw FlowError("assertion failed: \(selector) is still visible") + } + case "extendedWaitUntil": + let timeout = (try optionalNumber(map["timeout"]) ?? 17000) / 1000 + if let visible = map["visible"] { + let selector = try parseSelector(visible) + if try waitForElement(selector, timeout: timeout) == nil { + throw FlowError("\(selector) not visible after \(timeout)s") + } + } + if let hidden = map["notVisible"] { + let selector = try parseSelector(hidden) + if try !waitForAbsence(selector, timeout: timeout) { + throw FlowError("\(selector) still visible after \(timeout)s") + } + } + case "runFlow": + if let when = map["when"], try !evaluateCondition(when) { return "skipped (condition false)" } + let env = try envPairs(map["env"]) + if let flow = map["_flow"] as? [String: Any] { + try runFlow(flow, callerEnv: env) + } else if let commands = map["commands"] as? [Any] { + try runFlow(["file": "inline", "commands": commands], callerEnv: env) + } + case "inputText": + let text = try script.interpolate(stringArg(args, key: "text")) + try typeKeys(text) + case "eraseText": + let count = Int(try optionalNumber(map.isEmpty ? args : map["charactersToErase"]) ?? 50) + try typeKeys(String(repeating: XCUIKeyboardKey.delete.rawValue, count: count)) + case "pressKey": + let key = try script.interpolate(stringArg(args, key: "key")).lowercased() + switch key { + case "enter": try typeKeys(XCUIKeyboardKey.return.rawValue) + case "backspace": try typeKeys(XCUIKeyboardKey.delete.rawValue) + case "home": + XCUIDevice.shared.press(.home) + interacted() + default: throw FlowError("unsupported pressKey '\(key)'") + } + case "scroll": + let bounds = screen() + drag(from: CGPoint(x: bounds.midX, y: bounds.height * 0.7), to: CGPoint(x: bounds.midX, y: bounds.height * 0.3), duration: 0.4) + case "scrollUntilVisible": + return try scrollUntilVisible(map) + case "swipe": + try swipe(map) + case "repeat": + let times = try optionalNumber(map["times"]).map { Int($0) } + var count = 0 + while times.map({ count < $0 }) ?? true { + if let condition = map["while"], try !evaluateCondition(condition) { break } + try runFlow(["file": "repeat", "commands": map["commands"] as? [Any] ?? []], callerEnv: []) + count += 1 + } + return "ran \(count)x" + case "retry": + let retries = Int(try optionalNumber(map["maxRetries"]) ?? 1) + var attempt = 0 + while true { + do { + if let flow = map["_flow"] as? [String: Any] { + try runFlow(flow, callerEnv: []) + } else { + try runFlow(["file": "retry", "commands": map["commands"] as? [Any] ?? []], callerEnv: []) + } + return attempt == 0 ? nil : "passed on retry \(attempt)" + } catch { + attempt += 1 + if attempt > retries { throw error } + print("[edge-flow] retry \(attempt)/\(retries) after: \(error)") + } + } + case "evalScript": + let source = try stringArg(args, key: "script").trimmingCharacters(in: .whitespaces) + if source.hasPrefix("${"), source.hasSuffix("}") { + _ = try script.evaluate(String(source.dropFirst(2).dropLast())) + } else { + _ = try script.evaluate(source) + } + case "waitForAnimationToEnd": + let timeout = (try optionalNumber(map["timeout"]) ?? 15000) / 1000 + if !waitForSettle(timeout: timeout) { return "still animating after \(timeout)s, continuing" } + case "takeScreenshot": + var path = try script.interpolate(stringArg(args, key: "path")) + if !path.hasPrefix("/") { path = (options.workingDirectory as NSString).appendingPathComponent(path) } + if !path.hasSuffix(".png") { path += ".png" } + try XCUIScreen.main.screenshot().pngRepresentation.write(to: URL(fileURLWithPath: path)) + return "saved \(path)" + default: + throw FlowError("unsupported command '\(name)'") + } + return nil + } + + private func scrollUntilVisible(_ map: [String: Any]) throws -> String? { + guard let element = map["element"] else { throw FlowError("scrollUntilVisible needs element") } + let selector = try parseSelector(element) + let direction = try script.interpolate("\(map["direction"] ?? "DOWN")").uppercased() + let timeout = (try optionalNumber(map["timeout"]) ?? 20000) / 1000 + let percent = (try optionalNumber(map["visibilityPercentage"]) ?? 100) / 100 + let settle = (try optionalNumber(map["waitToSettleTimeoutMs"])).map { $0 / 1000 } + let deadline = Date().addingTimeInterval(timeout) + var swipes = 0 + while true { + if let frame = try findElement(selector), visibleFraction(frame) >= percent { + return swipes == 0 ? nil : "visible after \(swipes) swipe(s)" + } + if Date() >= deadline { throw FlowError("\(selector) not visible after scrolling \(timeout)s") } + let bounds = screen() + let (from, to): (CGPoint, CGPoint) + switch direction { + case "UP": (from, to) = (CGPoint(x: bounds.midX, y: bounds.height * 0.3), CGPoint(x: bounds.midX, y: bounds.height * 0.7)) + case "LEFT": (from, to) = (CGPoint(x: bounds.width * 0.3, y: bounds.midY), CGPoint(x: bounds.width * 0.7, y: bounds.midY)) + case "RIGHT": (from, to) = (CGPoint(x: bounds.width * 0.7, y: bounds.midY), CGPoint(x: bounds.width * 0.3, y: bounds.midY)) + default: (from, to) = (CGPoint(x: bounds.midX, y: bounds.height * 0.7), CGPoint(x: bounds.midX, y: bounds.height * 0.3)) + } + drag(from: from, to: to, duration: 0.4) + swipes += 1 + if let settle = settle { waitForSettle(timeout: settle) } + } + } + + private func swipe(_ map: [String: Any]) throws { + let bounds = screen() + let duration = (try optionalNumber(map["duration"]) ?? 400) / 1000 + if let start = map["start"], let end = map["end"] { + drag(from: try point(start, in: bounds), to: try point(end, in: bounds), duration: duration) + return + } + let direction = try script.interpolate("\(map["direction"] ?? "")").uppercased() + var from = CGPoint(x: bounds.midX, y: bounds.midY) + if let element = map["from"] { + let selector = try parseSelector(element) + guard let frame = try waitForElement(selector, timeout: options.lookupTimeout) else { + throw FlowError("swipe origin not found: \(selector)") + } + from = CGPoint(x: frame.midX, y: frame.midY) + } + // Maestro swipes from the origin to the screen edge in the direction. + let to: CGPoint + switch direction { + case "LEFT": to = CGPoint(x: bounds.width * 0.02, y: from.y) + case "RIGHT": to = CGPoint(x: bounds.width * 0.98, y: from.y) + case "UP": to = CGPoint(x: from.x, y: bounds.height * 0.1) + case "DOWN": to = CGPoint(x: from.x, y: bounds.height * 0.9) + default: throw FlowError("swipe needs direction LEFT, RIGHT, UP or DOWN (or start and end)") + } + drag(from: from, to: to, duration: duration) + if let ms = try optionalNumber(map["waitToSettleTimeoutMs"]) { waitForSettle(timeout: ms / 1000) } + } + + // MARK: Conditions + + /// Maestro evaluates `visible` / `notVisible` with the optional lookup + /// timeout, reduced by the time since the last interaction. + private func evaluateCondition(_ raw: Any) throws -> Bool { + guard let condition = raw as? [String: Any] else { throw FlowError("condition must be a map") } + if let platform = condition["platform"] { + if try script.interpolate("\(platform)").lowercased() != "ios" { return false } + } + if let visible = condition["visible"] { + let selector = try parseSelector(visible) + if try waitForElement(selector, timeout: adjusted(options.optionalLookupTimeout)) == nil { return false } + } + if let hidden = condition["notVisible"] { + let selector = try parseSelector(hidden) + if try !waitForAbsence(selector, timeout: adjusted(options.optionalLookupTimeout)) { return false } + } + if let value = condition["true"] { + let result = try script.interpolate("\(value)") + let trimmed = result.trimmingCharacters(in: .whitespacesAndNewlines) + if trimmed.isEmpty || ["false", "undefined", "null", "0"].contains(trimmed.lowercased()) { return false } + } + return true + } + + private func adjusted(_ timeout: TimeInterval) -> TimeInterval { + max(0, timeout - Date().timeIntervalSince(lastInteraction)) + } + + // MARK: Elements + + private func parseSelector(_ raw: Any?) throws -> Selector { + if let map = raw as? [String: Any] { + var selector = Selector() + if let text = map["text"] { selector.text = try script.interpolate("\(text)") } + if let id = map["id"] { selector.id = try script.interpolate("\(id)") } + if let index = try optionalNumber(map["index"]) { selector.index = Int(index) } + if selector.text == nil, selector.id == nil { throw FlowError("selector needs text or id") } + return selector + } + guard let raw = raw else { throw FlowError("missing selector") } + return Selector(text: try script.interpolate("\(raw)")) + } + + private func waitForElement(_ selector: Selector, timeout: TimeInterval) throws -> CGRect? { + let deadline = Date().addingTimeInterval(timeout) + while true { + if let frame = try findElement(selector) { return frame } + if Date() >= deadline { return nil } + Thread.sleep(forTimeInterval: 0.15) + } + } + + private func waitForAbsence(_ selector: Selector, timeout: TimeInterval) throws -> Bool { + let deadline = Date().addingTimeInterval(timeout) + while true { + if try findElement(selector) == nil { return true } + if Date() >= deadline { return false } + Thread.sleep(forTimeInterval: 0.15) + } + } + + /// Fast path: one query with `.firstMatch`. `firstMatch` returns the + /// outermost match, and a container's label concatenates its children's + /// text, so the fast path only holds when nothing inside it matches too. + /// Otherwise (off screen, nested match, or an index) fall back to one + /// hierarchy snapshot walked with Maestro's rules (deepest match only, on + /// screen, ordered by position). + private func findElement(_ selector: Selector) throws -> CGRect? { + if selector.index == nil { + let first = query(selector).firstMatch + guard first.exists else { return nil } + let frame = first.frame + if visibleFraction(frame) > 0, !query(selector, in: first).firstMatch.exists { return frame } + } + let matches = try snapshotMatches(selector) + let index = selector.index ?? 0 + return index < matches.count ? matches[index] : nil + } + + private func query(_ selector: Selector, in root: XCUIElement? = nil) -> XCUIElementQuery { + var query = (root ?? app).descendants(matching: .any) + if let id = selector.id { query = query.matching(predicate(fields: ["identifier"], pattern: id)) } + if let text = selector.text { query = query.matching(predicate(fields: ["label", "value", "placeholderValue"], pattern: text)) } + return query + } + + private func predicate(fields: [String], pattern: String) -> NSPredicate { + var clauses: [String] = [] + var args: [Any] = [] + for field in fields { + clauses.append("\(field) == %@") + args.append(pattern) + } + if let regex = matchRegex(pattern) { + for field in fields { + clauses.append("\(field) MATCHES %@") + args.append(regex.pattern) + } + } + return NSPredicate(format: clauses.joined(separator: " OR "), argumentArray: args) + } + + /// The whole-string, case-insensitive, dot-matches-newline regex Maestro + /// builds from a selector (nil when the pattern is not a valid regex, in + /// which case only literal equality can match). + private func matchRegex(_ pattern: String) -> NSRegularExpression? { + try? NSRegularExpression(pattern: "(?ism)" + pattern) + } + + private func snapshotMatches(_ selector: Selector) throws -> [CGRect] { + let textRegex = selector.text.flatMap(matchRegex) + let idRegex = selector.id.flatMap(matchRegex) + func fullMatch(_ regex: NSRegularExpression?, _ pattern: String, _ value: String?) -> Bool { + guard let value = value, !value.isEmpty || pattern.isEmpty else { return false } + for candidate in [value, value.replacingOccurrences(of: "\n", with: " ")] { + if candidate == pattern { return true } + if let regex = regex, let match = regex.firstMatch(in: candidate, range: NSRange(candidate.startIndex..., in: candidate)), + match.range.location == 0, match.range.length == (candidate as NSString).length { + return true + } + } + return false + } + func matches(_ node: XCUIElementSnapshot) -> Bool { + if let id = selector.id, !fullMatch(idRegex, id, node.identifier) { return false } + if let text = selector.text { + let values = [node.label, node.value as? String, node.placeholderValue] + if !values.contains(where: { fullMatch(textRegex, text, $0) }) { return false } + } + return true + } + var found: [CGRect] = [] + // Returns whether the subtree holds a match; a node only counts when no + // descendant matches (Maestro's deepestMatchingElement). + func walk(_ node: XCUIElementSnapshot) -> Bool { + var childMatched = false + for child in node.children where walk(child) { childMatched = true } + let isMatch = matches(node) + if isMatch, !childMatched, visibleFraction(node.frame) > 0 { found.append(node.frame) } + return isMatch || childMatched + } + _ = walk(try app.snapshot()) + return found.sorted { $0.minY != $1.minY ? $0.minY < $1.minY : $0.minX < $1.minX } + } + + private func visibleFraction(_ frame: CGRect) -> CGFloat { + guard frame.width > 0, frame.height > 0 else { return 0 } + let overlap = frame.intersection(screen()) + if overlap.isNull { return 0 } + return (overlap.width * overlap.height) / (frame.width * frame.height) + } + + private func screen() -> CGRect { + if let bounds = screenBounds, bounds.width > 0 { return bounds } + let frame = app.frame + screenBounds = frame + return frame + } + + private func selectApp(_ bundleId: String) { + appId = bundleId + app = XCUIApplication(bundleIdentifier: bundleId) + screenBounds = nil + } + + // MARK: Gestures + + private func coordinate(_ point: CGPoint) -> XCUICoordinate { + app.coordinate(withNormalizedOffset: .zero).withOffset(CGVector(dx: point.x, dy: point.y)) + } + + private func tap(_ frame: CGRect) { + coordinate(CGPoint(x: frame.midX, y: frame.midY)).tap() + interacted() + } + + private func drag(from: CGPoint, to: CGPoint, duration: TimeInterval) { + let distance = hypot(to.x - from.x, to.y - from.y) + let velocity = XCUIGestureVelocity(rawValue: distance / CGFloat(max(duration, 0.05))) + coordinate(from).press(forDuration: 0.05, thenDragTo: coordinate(to), withVelocity: velocity, thenHoldForDuration: 0.05) + interacted() + } + + /// Accepts "x%,y%" (screen-relative) or "x,y" (points). + private func point(_ raw: Any, in bounds: CGRect) throws -> CGPoint { + let text = try script.interpolate("\(raw)") + let parts = text.split(separator: ",").map { $0.trimmingCharacters(in: .whitespaces) } + guard parts.count == 2 else { throw FlowError("bad point '\(text)'") } + func value(_ part: String, _ size: CGFloat) throws -> CGFloat { + if part.hasSuffix("%"), let percent = Double(part.dropLast()) { return size * CGFloat(percent) / 100 } + if let points = Double(part) { return CGFloat(points) } + throw FlowError("bad point '\(text)'") + } + return CGPoint(x: try value(parts[0], bounds.width), y: try value(parts[1], bounds.height)) + } + + private func interacted() { + lastInteraction = Date() + } + + // MARK: Keyboard + + /// Types into the focused field. React Native can hide the focused input + /// (Edge's PIN entry), so no element reports keyboard focus and `typeText` + /// fails with the keyboard up. Maestro synthesizes key events regardless of + /// focus; this falls back to tapping the on-screen keys instead. + private func typeKeys(_ text: String) throws { + let focused = app.descendants(matching: .any).matching(NSPredicate(format: "hasKeyboardFocus == true")).firstMatch + if focused.exists { + app.typeText(text) + interacted() + return + } + let keyboard = app.keyboards.firstMatch + guard keyboard.exists else { throw FlowError("no element has keyboard focus and no keyboard is showing") } + for character in text { + let labels: [String] + switch String(character) { + case XCUIKeyboardKey.delete.rawValue: labels = ["delete"] + case XCUIKeyboardKey.return.rawValue: labels = ["return", "done", "go", "search", "next", "send"] + case " ": labels = ["space"] + default: labels = [String(character)] + } + let match = labels.lazy.map { label in + keyboard.descendants(matching: .any).matching(NSPredicate(format: "label ==[c] %@", label)).firstMatch + }.first { $0.exists } + guard let key = match else { + throw FlowError("no element has keyboard focus, and the keyboard has no '\(labels[0])' key to tap") + } + key.tap() + } + interacted() + } + + // MARK: Screen settle + + private func screenPixels() -> Data? { + guard let image = XCUIScreen.main.screenshot().image.cgImage else { return nil } + return image.dataProvider?.data as Data? + } + + /// Returns once two consecutive screenshots are identical (Maestro's + /// waitForAnimationToEnd), or false at the timeout. + @discardableResult + private func waitForSettle(timeout: TimeInterval) -> Bool { + let deadline = Date().addingTimeInterval(timeout) + var previous = screenPixels() + while Date() < deadline { + let current = screenPixels() + if current != nil, current == previous { return true } + previous = current + } + return false + } + + // MARK: Arguments + + private func stringArg(_ args: Any?, key: String) throws -> String { + if let map = args as? [String: Any] { + guard let value = map[key] else { throw FlowError("missing '\(key)'") } + return "\(value)" + } + guard let args = args else { throw FlowError("missing argument") } + return "\(args)" + } + + private func optionalNumber(_ raw: Any?) throws -> Double? { + guard let raw = raw, !(raw is NSNull) else { return nil } + if let number = raw as? NSNumber { return number.doubleValue } + let text = try script.interpolate("\(raw)") + guard let value = Double(text.trimmingCharacters(in: .whitespaces)) else { + throw FlowError("expected a number, got '\(text)'") + } + return value + } + + private func envPairs(_ raw: Any?) throws -> [(String, String)] { + guard let pairs = raw as? [[Any]] else { return [] } + return try pairs.compactMap { pair in + guard pair.count == 2, let key = pair[0] as? String else { return nil } + return (key, try script.interpolate("\(pair[1])")) + } + } + + // MARK: Logging + + private func summary(_ args: Any?) -> String { + guard let args = args else { return "" } + var copy = args + if var map = args as? [String: Any] { + map["_flow"] = nil + map["commands"] = (map["commands"] as? [Any]).map { "[\($0.count) commands]" } + copy = map + } + let data = (try? JSONSerialization.data(withJSONObject: copy, options: [.sortedKeys, .fragmentsAllowed])) ?? Data() + let text = String(data: data, encoding: .utf8) ?? "\(copy)" + return text.count > 100 ? String(text.prefix(97)) + "..." : text + } + + private func log(_ step: String, _ begin: Date, _ note: String) { + let total = Date().timeIntervalSince(started) + let took = Date().timeIntervalSince(begin) + print(String(format: "[edge-flow] %7.2fs (+%5.2fs) %@ -> %@", total, took, step, note)) + } +} diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift new file mode 100644 index 00000000..5b61d328 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift @@ -0,0 +1,147 @@ +import Foundation + +/// Walks a converted flow (and every nested/inlined subflow) before any step +/// runs, and reports each command or argument the interpreter does not +/// implement. The run fails on any report: there is no Maestro fallback. +enum FlowPreflight { + static let common: Set = ["label", "optional"] + static let selectorKeys: Set = ["text", "id", "index"] + static let conditionKeys: Set = ["visible", "notVisible", "true", "platform"] + static let configKeys: Set = ["appId", "env", "name", "tags", "jsEngine"] + static let pressKeys: Set = ["enter", "backspace", "home"] + + /// Argument keys each command accepts in map form (plus `common`). + static let commandKeys: [String: Set] = [ + "launchApp": ["appId", "clearState", "stopApp"], + "stopApp": ["appId"], + "tapOn": selectorKeys.union(["waitToSettleTimeoutMs", "retryTapIfNoChange"]), + "assertVisible": selectorKeys, + "assertNotVisible": selectorKeys, + "extendedWaitUntil": ["visible", "notVisible", "timeout"], + "runFlow": ["file", "_flow", "when", "env", "commands"], + "inputText": ["text"], + "eraseText": ["charactersToErase"], + "pressKey": ["key"], + "scroll": [], + "scrollUntilVisible": ["element", "direction", "timeout", "visibilityPercentage", "speed", "waitToSettleTimeoutMs"], + "swipe": ["from", "direction", "duration", "start", "end", "waitToSettleTimeoutMs"], + "repeat": ["times", "while", "commands"], + "retry": ["maxRetries", "commands", "file", "_flow"], + "evalScript": ["script"], + "waitForAnimationToEnd": ["timeout"], + "takeScreenshot": ["path"] + ] + + /// Commands that may be written as a bare name or with a scalar argument. + static let scalarForms: Set = [ + "launchApp", "stopApp", "tapOn", "assertVisible", "assertNotVisible", "inputText", "eraseText", + "pressKey", "scroll", "evalScript", "waitForAnimationToEnd", "takeScreenshot" + ] + + static func problems(in flow: [String: Any]) -> [String] { + var found: [String] = [] + check(flow: flow, into: &found) + return found + } + + private static func name(of flow: [String: Any]) -> String { + ((flow["file"] as? String).map { ($0 as NSString).lastPathComponent }) ?? "" + } + + private static func check(flow: [String: Any], into found: inout [String]) { + let flowName = name(of: flow) + let config = flow["config"] as? [String: Any] ?? [:] + for key in config.keys.sorted() where !configKeys.contains(key) { + found.append("\(flowName): unsupported flow config key '\(key)'") + } + let commands = flow["commands"] as? [Any] ?? [] + check(commands: commands, flowName: flowName, into: &found) + } + + private static func check(commands: [Any], flowName: String, into found: inout [String]) { + for (offset, command) in commands.enumerated() { + let at = "\(flowName) #\(offset + 1)" + let name: String + let args: Any? + if let bare = command as? String { + name = bare + args = nil + } else if let map = command as? [String: Any], map.count == 1, let entry = map.first { + name = entry.key + args = entry.value is NSNull ? nil : entry.value + } else { + found.append("\(at): malformed command \(command)") + continue + } + guard let allowed = commandKeys[name] else { + found.append("\(at): unsupported command '\(name)'") + continue + } + guard let map = args as? [String: Any] else { + if !scalarForms.contains(name) { + found.append("\(at): '\(name)' needs a map argument") + } + if name == "pressKey", let key = args as? String { + checkKey(key, at: at, into: &found) + } + continue + } + for key in map.keys.sorted() where !allowed.contains(key) && !common.contains(key) { + found.append("\(at): unsupported argument '\(key)' for \(name)") + } + checkArguments(name: name, map: map, at: at, flowName: flowName, into: &found) + } + } + + private static func checkArguments(name: String, map: [String: Any], at: String, flowName: String, into found: inout [String]) { + switch name { + case "launchApp": + if let clear = map["clearState"] as? Bool, clear { + found.append("\(at): launchApp clearState: true is unsupported (it would wipe the sim's roster accounts)") + } + case "extendedWaitUntil": + for key in ["visible", "notVisible"] { + if let selector = map[key] { checkSelector(selector, at: "\(at) \(key)", into: &found) } + } + case "scrollUntilVisible": + if let selector = map["element"] { checkSelector(selector, at: "\(at) element", into: &found) } + case "swipe": + if let selector = map["from"] { checkSelector(selector, at: "\(at) from", into: &found) } + case "pressKey": + if let key = map["key"] as? String { checkKey(key, at: at, into: &found) } + case "runFlow", "retry", "repeat": + if let when = map["when"] { checkCondition(when, at: at, into: &found) } + if let condition = map["while"] { checkCondition(condition, at: at, into: &found) } + if let flow = map["_flow"] as? [String: Any] { check(flow: flow, into: &found) } + if let commands = map["commands"] as? [Any] { check(commands: commands, flowName: flowName, into: &found) } + default: + break + } + } + + private static func checkSelector(_ selector: Any, at: String, into found: inout [String]) { + guard let map = selector as? [String: Any] else { return } + for key in map.keys.sorted() where !selectorKeys.contains(key) { + found.append("\(at): unsupported selector key '\(key)'") + } + } + + private static func checkCondition(_ condition: Any, at: String, into found: inout [String]) { + guard let map = condition as? [String: Any] else { + found.append("\(at): condition must be a map") + return + } + for key in map.keys.sorted() where !conditionKeys.contains(key) { + found.append("\(at): unsupported condition key '\(key)'") + } + for key in ["visible", "notVisible"] { + if let selector = map[key] { checkSelector(selector, at: "\(at) \(key)", into: &found) } + } + } + + private static func checkKey(_ key: String, at: String, into found: inout [String]) { + if !key.contains("${"), !pressKeys.contains(key.lowercased()) { + found.append("\(at): unsupported pressKey '\(key)'") + } + } +} diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowRunnerTests.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowRunnerTests.swift new file mode 100644 index 00000000..6c345ea1 --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowRunnerTests.swift @@ -0,0 +1,65 @@ +import XCTest + +/// The single test the runner bundle holds: interpret the flow named by +/// `EDGE_FLOW_FILE` (JSON written by scripts/maestro-yaml-to-json.rb). +final class FlowRunnerTests: XCTestCase { + private var recordedIssue = false + + override func setUp() { + continueAfterFailure = false + } + + func testFlow() throws { + let environment = ProcessInfo.processInfo.environment + guard let path = environment["EDGE_FLOW_FILE"], !path.isEmpty else { + XCTFail("EDGE_FLOW_FILE is not set (pass TEST_RUNNER_EDGE_FLOW_FILE to xcodebuild)") + return + } + let data = try Data(contentsOf: URL(fileURLWithPath: path)) + guard let document = try JSONSerialization.jsonObject(with: data) as? [String: Any] else { + XCTFail("\(path) is not a flow JSON object") + return + } + let flowName = ((document["file"] as? String) ?? path as String) as NSString + + let problems = FlowPreflight.problems(in: document) + if !problems.isEmpty { + let list = problems.map { " - \($0)" }.joined(separator: "\n") + print("[edge-flow] PREFLIGHT FAILED for \(flowName.lastPathComponent):\n\(list)") + XCTFail("\(flowName.lastPathComponent) uses commands the interpreter does not support:\n\(list)") + return + } + + let options = RunOptions(environment: environment) + EdgeQuiescence.install() + print("[edge-flow] start \(flowName.lastPathComponent) cap=\(options.quiescenceCap)s animations=\(options.animations)") + let started = Date() + do { + try FlowInterpreter(options: options).run(document) + print(String(format: "[edge-flow] PASSED %@ in %.2fs", flowName.lastPathComponent, Date().timeIntervalSince(started))) + } catch { + print(String(format: "[edge-flow] FAILED %@ after %.2fs: %@", flowName.lastPathComponent, Date().timeIntervalSince(started), "\(error)")) + XCTFail("\(error)") + } + printCapHits() + } + + override func record(_ issue: XCTIssue) { + recordedIssue = true + super.record(issue) + } + + override func tearDown() { + // XCUI API failures (e.g. a tap on a vanished element) record an XCTest + // failure without throwing, so name the step that was running. + // testRun.hasSucceeded is not final yet inside tearDown, so track issues. + if recordedIssue, !EdgeQuiescence.currentStep.isEmpty { + print("[edge-flow] last step: \(EdgeQuiescence.currentStep)") + } + } + + /// Each hit was already logged with its step as it happened. + private func printCapHits() { + print("[edge-flow] quiescence cap hits: \(EdgeQuiescence.capHits)") + } +} diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift new file mode 100644 index 00000000..6010292c --- /dev/null +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift @@ -0,0 +1,148 @@ +import Foundation +import JavaScriptCore + +/// JavaScript engine for `${...}` interpolation, `evalScript` and `when: true`. +/// +/// Maestro evaluates these as JavaScript, and the library flows use real JS +/// (Math.random, JSON.parse, arrow functions, `||` defaults), so a hand-rolled +/// subset would have to guess. JavaScriptCore runs the exact expression; any +/// script error fails the step with the offending source. +final class FlowScript { + private let context: JSContext + private var lastException: String? + + init() { + context = JSContext() + context.exceptionHandler = { [weak self] _, exception in + self?.lastException = exception?.toString() ?? "unknown JavaScript error" + } + context.evaluateScript("var output = {};") + } + + /// Evaluates one JavaScript expression or statement. + func evaluate(_ source: String) throws -> JSValue { + lastException = nil + let value = context.evaluateScript(source) + if let error = lastException { + throw FlowError("JavaScript error in `\(source)`: \(error)") + } + guard let value = value else { throw FlowError("JavaScript returned nothing for `\(source)`") } + return value + } + + /// Replaces every `${expr}` in the string with the expression's result + /// (converted with JavaScript String()), the way Maestro does. + func interpolate(_ text: String) throws -> String { + guard text.contains("${") else { return text } + var result = "" + var index = text.startIndex + while index < text.endIndex { + if text[index] == "$", text.index(after: index) < text.endIndex, text[text.index(after: index)] == "{" { + let start = text.index(index, offsetBy: 2) + let end = try closingBrace(in: text, from: start) + let value = try evaluate(String(text[start.. String.Index { + var depth = 0 + var quote: Character? + var index = start + while index < text.endIndex { + let char = text[index] + if let open = quote { + if char == "\\" { + index = text.index(after: index) + } else if char == open { + quote = nil + } + } else if char == "\"" || char == "'" || char == "`" { + quote = char + } else if char == "{" { + depth += 1 + } else if char == "}" { + if depth == 0 { return index } + depth -= 1 + } + if index < text.endIndex { index = text.index(after: index) } + } + throw FlowError("unterminated ${ in `\(text)`") + } + + func isDefined(_ name: String) -> Bool { + context.globalObject.hasProperty(name) + } + + func value(_ name: String) -> JSValue? { + isDefined(name) ? context.globalObject.forProperty(name) : nil + } + + func set(_ name: String, _ value: Any?) { + context.globalObject.setValue(value ?? JSValue(undefinedIn: context) as Any, forProperty: name) + } + + func remove(_ name: String) { + context.globalObject.deleteProperty(name) + } + + /// Declares a name as `undefined` when nothing defines it yet, so a flow's + /// own `KEY: ${KEY || default}` header can read it without a ReferenceError. + func declare(_ name: String) { + if !isDefined(name) { set(name, nil) } + } +} + +/// Saves and restores the variables a flow scope overrides (runFlow `env:` +/// and the subflow's own header `env:`), so they do not leak to the caller. +final class EnvScope { + private var saved: [(String, JSValue?)] = [] + private let script: FlowScript + + init(_ script: FlowScript) { + self.script = script + } + + func set(_ name: String, _ value: String) { + save(name) + script.set(name, value) + } + + /// Declares `name` as undefined for this scope (removed again on restore). + func declare(_ name: String) { + save(name) + script.declare(name) + } + + private func save(_ name: String) { + if !saved.contains(where: { $0.0 == name }) { + saved.append((name, script.value(name))) + } + } + + func restore() { + for (name, value) in saved.reversed() { + if let value = value { + script.set(name, value) + } else { + script.remove(name) + } + } + saved = [] + } +} + +struct FlowError: Error, CustomStringConvertible { + let description: String + + init(_ description: String) { + self.description = description + } +} From c30736bf31d6f98252fe2484043a941fa3bcc4c2 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 15:21:35 -0700 Subject: [PATCH 03/13] Route iOS flows to the XCUITest interpreter build-and-test's new ios-flows-run-on-xcuitest rule makes xcuitest-run.sh the default iOS flow runner; the maestro CLI runs iOS flows only when a task asks or the interpreter's preflight rejects a command. Android stays on Maestro. capture-buy-quote.sh takes --driver xcuitest|maestro (default xcuitest). references/xcuitest-interpreter.md lists the supported commands, the Maestro semantics kept, the quiescence cap, the test-mode animation switch and the unsupported commands. --- .cursor/skills/build-and-test/SKILL.md | 4 +- .../skills/build-and-test/references/build.md | 2 +- .../skills/build-and-test/references/drive.md | 9 ++- .../references/sim-testing-playbook.md | 21 +++-- .../references/xcuitest-interpreter.md | 79 +++++++++++++++++++ .../scripts/capture-buy-quote.sh | 28 +++++-- .../hooks/require-skill-read-for-scripts.sh | 6 +- 7 files changed, 126 insertions(+), 23 deletions(-) create mode 100644 .cursor/skills/build-and-test/references/xcuitest-interpreter.md diff --git a/.cursor/skills/build-and-test/SKILL.md b/.cursor/skills/build-and-test/SKILL.md index 80554eb0..d2bc051e 100644 --- a/.cursor/skills/build-and-test/SKILL.md +++ b/.cursor/skills/build-and-test/SKILL.md @@ -41,8 +41,8 @@ CEILING: the bar is the IN-APP success state, not EXTERNAL finality. Once the su | Phase | Reference (under `~/.cursor/skills/build-and-test/`) | Rules it owns | Delivered by the gate at | |---|---|---|---| | Step 0a-0c: preflight, sim selection, RN build, linking a gui dependency into the app | `references/build.md` | `preflight-before-build-decisions`, `slot-sim-is-the-clone`, `gui-dependency-integration` | first `slot-preflight.sh`, `select-ios-sim.sh`, `ios-rn-build.sh` or `ios-rn-build-wait.sh` call | -| Step 0d-0f: driving the app | `references/drive.md` | `maestro-flows-are-shortcuts`, `testids-over-coordinates`, `single-asset-plugin-trim`, `force-swap-provider-locally`, `runtime-inspection-via-debugger`, `spaced-pin-taps`, `no-hideKeyboard`, `no-hierarchy-polling-on-buy` | first maestro drive or `capture-buy-quote.sh` call | -| Proof frames, forced visual states, un-runnable assets | `references/evidence.md` | `proof-screenshots-for-pr`, `hack-verify-visual-changes`, `unrunnable-asset-proxy-verification` | first maestro drive or `capture-buy-quote.sh` call | +| Step 0d-0f: driving the app | `references/drive.md` | `maestro-flows-are-shortcuts`, `ios-flows-run-on-xcuitest`, `testids-over-coordinates`, `single-asset-plugin-trim`, `force-swap-provider-locally`, `runtime-inspection-via-debugger`, `spaced-pin-taps`, `no-hideKeyboard`, `no-hierarchy-polling-on-buy` | first maestro drive, `xcuitest-run.sh` or `capture-buy-quote.sh` call | +| Proof frames, forced visual states, un-runnable assets | `references/evidence.md` | `proof-screenshots-for-pr`, `hack-verify-visual-changes`, `unrunnable-asset-proxy-verification` | first maestro drive, `xcuitest-run.sh` or `capture-buy-quote.sh` call | | Any test that needs a funded asset or moves value (swap, send, sweep) | `references/funding.md` | `funded-test-accounts`, `executable-pair-must-complete`, `create-missing-destination-wallet` | first `log-attempt.sh --category swap\|send\|sweep` (a backstop: read it yourself BEFORE choosing an account or a pair) | | Steps 1-3: Node/TypeScript, Node, placeholder | this file | n/a | n/a | diff --git a/.cursor/skills/build-and-test/references/build.md b/.cursor/skills/build-and-test/references/build.md index 72e363d5..cea9d26f 100644 --- a/.cursor/skills/build-and-test/references/build.md +++ b/.cursor/skills/build-and-test/references/build.md @@ -40,7 +40,7 @@ A real on-simulator UI test that logs into the pre-provisioned test account, nav ### 0a. Prerequisites (check, install if missing) - `xcrun -version` → Xcode CLT -- `maestro --version` → install with `curl -Ls "https://get.maestro.mobile.dev" | bash`, then add `$HOME/.maestro/bin` to PATH. maestro needs JDK 11+; Temurin 17 works. +- `maestro --version` (Android, or a flow the XCUITest interpreter rejects) → install with `curl -Ls "https://get.maestro.mobile.dev" | bash`, then add `$HOME/.maestro/bin` to PATH. maestro needs JDK 11+; Temurin 17 works. ### 0b. Resolve + boot the simulator diff --git a/.cursor/skills/build-and-test/references/drive.md b/.cursor/skills/build-and-test/references/drive.md index 3563bd1e..37b09a9c 100644 --- a/.cursor/skills/build-and-test/references/drive.md +++ b/.cursor/skills/build-and-test/references/drive.md @@ -4,19 +4,20 @@ Governs step 0d-0f of `/build-and-test` (driving the app on the sim: the flow li | Script | Purpose | Exit codes | |--------|---------|------------| -| `capture-buy-quote.sh [--flow ]` | Run one flow on the maestro CLI (device and driver port pinned from the slot env), then capture via an external simctl screenshot burst; up to 5 retry cycles | 0 = captured; nonzero = FAIL (step 0e) | +| `capture-buy-quote.sh [--flow ] [--driver xcuitest\|maestro]` | Run one flow on the XCUITest interpreter (default) or the maestro CLI (device and driver port pinned from the slot env), then capture via an external simctl screenshot burst; up to 5 retry cycles | 0 = captured; nonzero = FAIL (step 0e) | Before the sim-test phase, READ `~/.cursor/skills/build-and-test/references/sim-testing-playbook.md`: it holds the working knowledge (funding floors, account roster/switching, feature-enablement gotchas, investigation order) that otherwise gets re-learned every run. - **Compose, don't re-derive.** Parameterized subflows live in `~/.cursor/skills/build-and-test/maestro/common/` (`login-if-needed`, `dismiss-startup-modals`, `select-swap-pair`, `confirm-slider`; the slider gesture is SOLVED there, never re-derive it). Copy the subflows you need next to your task flow and `runFlow` them; author NEW task-specific `.yaml` liberally for what the task actually changed. Task flows stay LOCAL (`.syncignore`d from the agent repo; never committed to the gui repo, whose `maestro/` is the heavyweight verification suite, reference-only for selectors). What DOES get committed to the gui: missing `testID`s, per `testids-over-coordinates`. -- **Exploration on the MCP, proof on the CLI.** For exploring screens/selectors use the maestro MCP tools (persistent driver, no ~2-min `maestro test` startup per probe). The MCP daemon binds to one device and DRIFTS: `maestro-mcp-wrapper.sh` start-pins it, but the pin does not survive slot-sim relaunch cycling, and the daemon can rebind to another slot's sim mid-run. So: (1) re-verify its bound device (list_devices + an inspect matching your app's expected state) at the START OF EVERY observation block; (2) ALL PROOF evidence comes from the maestro CLI, canonical slot invocation `maestro --device "$AGENT_SIM_UDID" --driver-host-port $((AGENT_METRO_PORT + 1000)) test ` (the per-slot driver port keeps parallel slots' iOS drivers apart; hook-enforced; the MCP daemon uses +2000 via its wrapper), plus `simctl io` against that SAME UDID. NEVER an MCP screenshot, and NEVER a `simctl io` against a different device than the CLI drove. If a verify shows the MCP on the wrong device, drop to CLI for the rest of the run. +- **Exploration on the MCP, proof on the CLI.** For exploring screens/selectors use the maestro MCP tools (persistent driver, no ~2-min `maestro test` startup per probe). The MCP daemon binds to one device and DRIFTS: `maestro-mcp-wrapper.sh` start-pins it, but the pin does not survive slot-sim relaunch cycling, and the daemon can rebind to another slot's sim mid-run. So: (1) re-verify its bound device (list_devices + an inspect matching your app's expected state) at the START OF EVERY observation block; (2) ALL PROOF evidence comes from a CLI flow run: on iOS `xcuitest-run.sh` per `ios-flows-run-on-xcuitest`; on Android, or when the task asks for Maestro, the maestro CLI, canonical slot invocation `maestro --device "$AGENT_SIM_UDID" --driver-host-port $((AGENT_METRO_PORT + 1000)) test ` (the per-slot driver port keeps parallel slots' iOS drivers apart; hook-enforced; the MCP daemon uses +2000 via its wrapper), plus `simctl io` against that SAME UDID. NEVER an MCP screenshot, and NEVER a `simctl io` against a different device than the CLI drove. If a verify shows the MCP on the wrong device, drop to CLI for the rest of the run. - **One proof flow.** For the repeatable PROOF run compose ONE yaml flow and run it once via `capture-buy-quote.sh --flow `; that run produces the PR evidence screenshots. - **Proposals, never direct edits.** The playbook and the flow library are operator-curated: entries may be trusted without re-verification because the operator reviewed and promoted each one. Task runs never promote; proposals wait in run reports for the eval's manually-triggered flow-consolidation pass. In the report's Dev Notes & Gotchas section: - `[playbook]`: durable knowledge, one bullet. - `[flow]`: a NEW reusable drive sequence: name, params, one-line purpose, and the FULL yaml EMBEDDED as a fenced block (worktrees are pruned on retention; the report attachment is the durable copy). - `[flow-update]`: a change to an EXISTING library flow (genericize, new param, split into subflows): name the flow, the change, and the compatibility argument. New params MUST default to current behavior so existing callers are unaffected, and a rename/split must say so explicitly (callers get grepped at promotion). - **Scope a `[playbook]` proposal tightly.** It earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE: if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING instead. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything an existing rule already covers. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately. +On iOS, RUN flow YAML with the native XCUITest interpreter, not `maestro test`: `~/.cursor/skills/build-and-test/scripts/xcuitest-run.sh --flow [--env K=V ...]` (slot UDID from `$AGENT_SIM_UDID`; no host port, so parallel slots never collide). It runs the same flow files (library `common/` flows included) and writes `takeScreenshot` output to the same paths; `capture-buy-quote.sh` uses it by default. Use the maestro CLI only when the task asks for Maestro, on Android, or for a flow the interpreter rejects. The interpreter preflights the whole flow tree and exits 2 before step 1 when a command or argument is unsupported, naming the command and the flow: rewrite that step with supported commands, or run that one flow on the maestro CLI and name the rejected command in the run report. There is no automatic fallback. Pass `--animations on` when the task is about an animation (the default turns them off). The run kills this slot's maestro MCP daemon and the session's MCP tools do not come back: finish MCP exploration before the first interpreter run. Supported commands, semantics and the other flags: `references/xcuitest-interpreter.md`; usage, output and exit codes: the script header. Scoped exception to `no-mutation`, test-infrastructure only. TESTIDS FIRST, COORDINATES LAST. When a flow needs to drive an element that has no stable selector (text match fails and no `testID` exists), ADD the missing `testID` prop to that component in the gui worktree and drive via it. A `testID` is a JS-only prop: Metro reload picks it up in seconds (no native rebuild), so adding one is cheaper than a single round of coordinate trial-and-error, and it de-brittles the suite for every future run. Coordinate taps are permitted ONLY for surfaces you cannot edit (system dialogs, native pickers, third-party views that don't forward `testID`) or when a reload would destroy unrecoverable in-flight app state, and any coordinate tap that survives into the PROOF flow must be called out in the run report with why a testID was not possible. - **Commit discipline:** commit the testID additions as a SEPARATE commit, distinct from any feature commit; change ONLY `testID` props, never component logic; update the flow selector(s) to use them. - **Message names the surface:** subject `test: add testIDs to ` (e.g. `test: add testIDs to ExchangeScene swap pills`), with every added id listed in the body. Never a generic subject: identical subjects across runs make these commits indistinguishable when a human cherry-picks between branches. @@ -29,13 +30,13 @@ If no selector was missing, this rule is a no-op. -### 0d. Run the maestro capture +### 0d. Run the capture ```bash ~/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh ``` -Drives `maestro/buy-quote-input.yaml` (login → Buy → $500), then captures via an external simctl screenshot burst, keeping the last frame taken while the app was alive. Retries up to 5 cycles. Writes `/tmp/agent-mvp-buy-quote-screenshot.png` on success. +Drives `maestro/buy-quote-input.yaml` (login → Buy → $500) on the XCUITest interpreter (`--driver maestro` for the maestro CLI), then captures via an external simctl screenshot burst, keeping the last frame taken while the app was alive. Retries up to 5 cycles. Writes `/tmp/agent-mvp-buy-quote-screenshot.png` on success. ### 0e. PASS / FAIL contract diff --git a/.cursor/skills/build-and-test/references/sim-testing-playbook.md b/.cursor/skills/build-and-test/references/sim-testing-playbook.md index c1d6c0cc..6cd5854d 100644 --- a/.cursor/skills/build-and-test/references/sim-testing-playbook.md +++ b/.cursor/skills/build-and-test/references/sim-testing-playbook.md @@ -7,7 +7,9 @@ entry (the human audits and prunes it periodically — keep entries dense). ## Flow library — compose these, never re-derive them Parameterized subflows in `~/.cursor/skills/build-and-test/maestro/common/` -(compose via `runFlow` with `env:`; copy next to your task flow). Re-deriving +(compose via `runFlow` with `env:`; copy next to your task flow). Every flow +here runs on the iOS XCUITest interpreter (`scripts/xcuitest-run.sh`) and on the +maestro CLI. Re-deriving any of these inline is wasted derivation — two 2026-07 sessions independently rebuilt the entire swap-pair sequence tap-by-tap that `select-swap-pair` already encodes, params and gotchas included. @@ -486,18 +488,27 @@ debugging screenshots of the wrong device. you explored through the MCP, kill that slot's daemon and its `xcodebuild test-without-building` child (match both on YOUR udid, never another slot's), then run the CLI. (Promoted 2026-08-06, run 1216251688512498.) + `scripts/xcuitest-run.sh` does this itself: it kills the daemon PIDs matched + on `--udid` and terminates the maestro driver app before its own drive. The + MCP server does not come back after that kill: the session loses its maestro + MCP tools, so finish MCP exploration before the first interpreter run. - **Maestro `visible:` matches the WHOLE text node.** "Powered by Maya Protocol" renders inside a node whose text is `Powered by Maya ProtocolTap to Change Provider`, so the exact match never hits while `"Powered by .*"` does. Anchor quote-ready waits on `Slide to Confirm` instead — it appears only once a quote resolves. (Promoted 2026-07-29, run 1216518039073159.) -- **Maestro economics:** each `maestro test` invocation pays ~2 min driver +- **Driver economics:** each `maestro test` invocation pays ~2 min driver startup. For EXPLORATION (finding selectors, poking screens) use the **maestro - MCP tools** (persistent driver, per-command tap/swipe/hierarchy/screenshot — + MCP tools** (persistent driver, per-command tap/swipe/hierarchy/screenshot; select the device matching `$AGENT_SIM_UDID` first). For the REPEATABLE PROOF - run, compose ONE yaml flow and run it once — that run produces the evidence - screenshots for the PR. + run, compose ONE yaml flow and run it once: that run produces the evidence + screenshots for the PR. On iOS the proof run goes through + `scripts/xcuitest-run.sh --flow ` (the XCUITest interpreter, see + `references/xcuitest-interpreter.md`): same YAML, a cached runner, no Maestro + driver startup, and quiescence waits capped at 1s. The maestro CLI runs iOS + flows only when the task asks for Maestro or the interpreter's preflight + rejects a command the flow needs. Android stays on the maestro CLI. - Modal gauntlet, eraseText-before-inputText, spaced PIN taps: all encoded in the `common/` flows — use them instead of remembering. - **Fixed-port debug dev-servers are NOT slot-safe — use `updot` instead.** The diff --git a/.cursor/skills/build-and-test/references/xcuitest-interpreter.md b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md new file mode 100644 index 00000000..5942950d --- /dev/null +++ b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md @@ -0,0 +1,79 @@ +# XCUITest flow interpreter (iOS) + +`scripts/xcuitest-run.sh` runs Maestro flow YAML natively. One generic XCUITest +bundle (`xcuitest/EdgeFlowRunner`) is built once per Xcode build, iOS runtime +and source hash (`scripts/xcuitest-build.sh`, cached under +`~/Library/Caches/edge-flow-runner`). Each run is then one +`xcodebuild test-without-building` against the slot UDID. + +## Run sequence + +1. `maestro-yaml-to-json.rb` converts the flow and inlines every `runFlow` / + `retry` file (paths resolve relative to the including flow). `--env K=V` + values ride along as CLI env. +2. The runner reads the JSON (`TEST_RUNNER_EDGE_FLOW_FILE`) and preflights the + whole tree. Any unsupported command, argument, condition key, selector key + or `pressKey` value fails the run before step 1 with exit 2, naming the + flow and step. There is no Maestro fallback. +3. Commands run in order. Each prints + `[edge-flow] s (+s) # -> ok|skipped|FAILED`. + +## Supported commands + +| Command | Notes | +|---|---| +| `launchApp` | `appId`, `stopApp: false` (activate if running). `clearState: true` is rejected because it wipes the roster accounts. Passes `-EdgeTestAnimations ` unless `--animations on`. | +| `stopApp` | | +| `tapOn` | `text`, `id`, `index`, `waitToSettleTimeoutMs`, `retryTapIfNoChange` | +| `assertVisible` / `assertNotVisible` | Maestro timeouts: 17s (7s when `optional`), minus time since the last interaction | +| `extendedWaitUntil` | `visible` / `notVisible`, `timeout` | +| `runFlow` | `file`, inline `commands`, `env`, `when` (`visible`, `notVisible`, `true`, `platform`) | +| `inputText` / `eraseText` / `pressKey` | `pressKey`: Enter, Backspace, Home. `eraseText` defaults to 50 characters. When no element has keyboard focus (a hidden input, such as the PIN entry) the runner taps the on-screen keys instead, so only characters with their own key (digits, the current letter case, space) can be typed | +| `scroll` / `scrollUntilVisible` / `swipe` | `scrollUntilVisible`: `element`, `direction`, `timeout`, `visibilityPercentage`, `waitToSettleTimeoutMs`. `swipe`: `from` + `direction`, or `start` / `end` points (`"50%,80%"`), `duration` | +| `repeat` / `retry` | `repeat`: `times`, `while`. `retry`: `maxRetries`, `file` or `commands` | +| `evalScript` | Full JavaScript (JavaScriptCore). `output.*` persists for the whole run | +| `waitForAnimationToEnd` | Two consecutive identical screenshots, `timeout` default 15s | +| `takeScreenshot` | Same path rules as `maestro test`: relative to the current directory, `.png` appended | + +`optional: true` and `label` work on every command. + +## Maestro semantics the runner keeps + +- `text` and `id` selectors are case-insensitive regexes (dot matches newline) + that must match the WHOLE attribute, or equal it literally. `text` checks + label, value and placeholderValue. The deepest matching element wins; + `index` orders on-screen matches top to bottom, then left to right. +- `${...}` and `evalScript` are JavaScript. A script error fails the step and + prints the source. `when: { true: ... }` is false for blank, `false`, `0`, + `null` and `undefined`. +- A subflow's header `env:` is evaluated after the caller's `runFlow env:`, so + the header's `KEY: ${KEY || default}` idiom sees the caller's value. + +## Quiescence cap + +XCUITest waits for the app to go idle before and after every event. React +Native apps with timers or loading spinners can hold that wait for many +seconds. The runner swizzles `XCUIApplicationProcess`'s quiescence waits (the +WebDriverAgent approach) to give up after `--quiescence-cap` seconds (default +1; 0 skips the wait). Every wait that reaches the cap prints +`[edge-flow] quiescence cap hit ... step: `, and the last line of the run +counts them. A flow sets its own cap with env `EDGE_QUIESCENCE_CAP`. + +## Test-mode animations (edge-react-gui debug builds) + +`--animations off` (default) passes `-EdgeTestAnimations off` on `launchApp` +and writes it to the sim's defaults for manual relaunches. UIKit animations +then complete immediately, and Reanimated runs in reduced motion, so looping +shimmers and chart pulses stop after one cycle. `fast` runs Core Animation at +100x instead. Both modes also freeze native `ActivityIndicator` spinners in +place, so `waitForAnimationToEnd` settles on a loading screen (0.2s, versus +the full timeout with `on`). `on` clears the flag. Release builds ignore it. +Spinners do not hold the quiescence wait in any mode. + +## Not supported (preflight rejects) + +`inputRandomText`, `copyTextFrom`, `pasteText`, `longPressOn`, `hideKeyboard`, +`assertVisible enabled:`, `scrollUntilVisible centerElement:`, `tapOn point:`, +`launchApp clearState: true`, and any command missing from the table above. +To drive a flow that needs one, rewrite the step or run that flow on the +maestro CLI and name it in the run report. diff --git a/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh b/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh index fdc1861e..a37156e1 100755 --- a/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh +++ b/.cursor/skills/build-and-test/scripts/capture-buy-quote.sh @@ -13,8 +13,8 @@ # scene, so a fixed-delay single shot is either too early (still loading) # or too late (already crashed → springboard). # -# This wrapper drives the interaction with maestro (the input flow), which does -# no polling after entering the amount, then captures with an EXTERNAL simctl +# This wrapper drives the interaction flow (the XCUITest interpreter by +# default, or maestro), which does no polling after entering the amount, then captures with an EXTERNAL simctl # screenshot burst (pixel-only, no hierarchy traversal), keeping the LAST frame # taken while the app was still alive — i.e. the resolved quote, just before any # crash. Retries the whole cycle until it lands a frame from late enough to @@ -23,7 +23,7 @@ # Usage: # capture-buy-quote.sh [--out ] [--flow ] \ # [--bundle-id ] [--quote-secs N] [--window-secs N] [--cycles N] \ -# [--device ] [--driver-port N] +# [--device ] [--driver-port N] [--driver xcuitest|maestro] # # Defaults: # --out /tmp/agent-mvp-buy-quote-screenshot.png @@ -35,6 +35,7 @@ # --device $AGENT_SIM_UDID when set (slot session), else simctl "booted" # --driver-port $AGENT_METRO_PORT+1000 when set (per-slot maestro driver port, # keeps parallel slots' iOS drivers off each other), else unset +# --driver xcuitest (xcuitest-run.sh; needs a device) or maestro # # Device pinning: with multiple sims booted (parallel orch slots), an unpinned # maestro attaches to an arbitrary device and `simctl io booted` photographs an @@ -58,6 +59,7 @@ WINDOW_SECS=14 CYCLES=5 DEVICE="${AGENT_SIM_UDID:-}" DRIVER_PORT="" +DRIVER="xcuitest" [[ -n "${AGENT_METRO_PORT:-}" ]] && DRIVER_PORT=$((AGENT_METRO_PORT + 1000)) while [[ $# -gt 0 ]]; do @@ -70,6 +72,7 @@ while [[ $# -gt 0 ]]; do --cycles) CYCLES="$2"; shift 2 ;; --device) DEVICE="$2"; shift 2 ;; --driver-port) DRIVER_PORT="$2"; shift 2 ;; + --driver) DRIVER="$2"; shift 2 ;; *) echo "Unknown arg: $1" >&2; exit 1 ;; esac done @@ -81,7 +84,11 @@ MAESTRO_ARGS=() [[ -n "$DEVICE" ]] && MAESTRO_ARGS+=(--device "$DEVICE") [[ -n "$DRIVER_PORT" ]] && MAESTRO_ARGS+=(--driver-host-port "$DRIVER_PORT") -command -v maestro >/dev/null 2>&1 || { echo "maestro not found in PATH" >&2; exit 1; } +case "$DRIVER" in + xcuitest) [[ -n "$DEVICE" ]] || { echo "--driver xcuitest needs --device or \$AGENT_SIM_UDID" >&2; exit 1; } ;; + maestro) command -v maestro >/dev/null 2>&1 || { echo "maestro not found in PATH" >&2; exit 1; } ;; + *) echo "--driver must be xcuitest or maestro" >&2; exit 1 ;; +esac command -v xcrun >/dev/null 2>&1 || { echo "xcrun not found (need Xcode CLT)" >&2; exit 1; } [[ -f "$FLOW" ]] || { echo "Maestro flow not found: $FLOW" >&2; exit 1; } @@ -91,8 +98,13 @@ trap 'rm -rf "$TMP"' EXIT alive() { xcrun simctl spawn "$SIMCTL_DEVICE" launchctl list 2>/dev/null | grep -qi "${BUNDLE_ID#*.}"; } for ((cycle = 1; cycle <= CYCLES; cycle++)); do - echo "[capture] cycle $cycle/$CYCLES: maestro ${MAESTRO_ARGS[*]:-} $FLOW (simctl device: $SIMCTL_DEVICE) ..." - maestro ${MAESTRO_ARGS[@]+"${MAESTRO_ARGS[@]}"} test "$FLOW" >"$TMP/maestro.log" 2>&1 || true + if [[ "$DRIVER" == xcuitest ]]; then + echo "[capture] cycle $cycle/$CYCLES: xcuitest-run.sh $FLOW (device: $DEVICE) ..." + "$SCRIPT_DIR/xcuitest-run.sh" --flow "$FLOW" --udid "$DEVICE" >"$TMP/driver.log" 2>&1 || true + else + echo "[capture] cycle $cycle/$CYCLES: maestro ${MAESTRO_ARGS[*]:-} $FLOW (simctl device: $SIMCTL_DEVICE) ..." + maestro ${MAESTRO_ARGS[@]+"${MAESTRO_ARGS[@]}"} test "$FLOW" >"$TMP/driver.log" 2>&1 || true + fi best=""; best_t=0; SECONDS=0 while [[ "$SECONDS" -lt "$WINDOW_SECS" ]]; do alive || break @@ -109,6 +121,6 @@ for ((cycle = 1; cycle <= CYCLES; cycle++)); do echo "[capture] crashed before the quote resolved (last frame t=${best_t}s); retrying ..." done -echo "[capture] FAIL after $CYCLES cycles — last maestro output:" -tail -30 "$TMP/maestro.log" >&2 +echo "[capture] FAIL after $CYCLES cycles — last $DRIVER output:" +tail -30 "$TMP/driver.log" >&2 exit 1 diff --git a/agent-watcher/hooks/require-skill-read-for-scripts.sh b/agent-watcher/hooks/require-skill-read-for-scripts.sh index 5be5ad23..e0cd2950 100755 --- a/agent-watcher/hooks/require-skill-read-for-scripts.sh +++ b/agent-watcher/hooks/require-skill-read-for-scripts.sh @@ -145,9 +145,9 @@ need '(pr-land-automerge|pr-merge-watch|pr-land-merge|force-land-rationale)\.sh( need '(pr-land-publish|npm-publish-web|npm-auth-wait|upgrade-dep)\.sh([[:space:]]|$)' pr-land:publish need '(pr-bot-findings-sweep|pr-land-extract-asana-task)\.sh([[:space:]]|$)' pr-land:post-merge # /build-and-test phase slices, every session. All four build scripts and the -# capture script live in build-and-test/scripts, so the core is already required. +# two drive scripts live in build-and-test/scripts, so the core is already required. need '(slot-preflight|select-ios-sim|ios-rn-build|ios-rn-build-wait)\.sh([[:space:]]|$)' build-and-test:build -need 'capture-buy-quote\.sh([[:space:]]|$)' build-and-test:drive build-and-test:evidence +need '(capture-buy-quote|xcuitest-run)\.sh([[:space:]]|$)' build-and-test:drive build-and-test:evidence [ "$ORCH" = 1 ] || return 0 # Shared top-level scripts with one governing skill, and the /one-shot phase @@ -171,7 +171,7 @@ need 'set-tested\.sh([[:space:]]|$)' one-shot:testing # build-and-test's scripts that build or drive the app on the sim start the # /one-shot testing phase (select-ios-sim.sh / slot-preflight.sh only pick and # check a slot, so they pull build-and-test:build above and nothing here). -need '(capture-buy-quote|ios-rn-build|ios-rn-build-wait)\.sh([[:space:]]|$)' one-shot:testing +need '(capture-buy-quote|xcuitest-run|ios-rn-build|ios-rn-build-wait)\.sh([[:space:]]|$)' one-shot:testing need 'asana-review-field\.sh([[:space:]]|$)' one-shot:review need 'pr-create\.sh([[:space:]]|$)' one-shot:pr need 'watch-pr\.sh([[:space:]]|$)' one-shot:watch From b7e41bf91b059202b26dc6d005a203bc0f307c31 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 15:57:40 -0700 Subject: [PATCH 04/13] Retry Buy tab tap after login --- .../build-and-test/maestro/buy-quote-input.yaml | 14 ++++++++++---- .../skills/build-and-test/maestro/buy-quote.yaml | 14 ++++++++++---- 2 files changed, 20 insertions(+), 8 deletions(-) diff --git a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml index 52ea17e1..a7861e5b 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml @@ -43,9 +43,15 @@ env: visible: "Claim Your Web3 Handle" commands: - tapOn: "Not Now" -- tapOn: "Buy" -- extendedWaitUntil: - visible: "Amount USD" - timeout: 15000 +# The tab bar stays hidden for up to ~20s after login while "Buy" is already +# in the hierarchy, so a single tap can land on nothing. Retry until the Buy +# scene shows. +- retry: + maxRetries: 3 + commands: + - tapOn: "Buy" + - extendedWaitUntil: + visible: "Amount USD" + timeout: 10000 - tapOn: "Amount USD" - inputText: ${BUY_AMOUNT} diff --git a/.cursor/skills/build-and-test/maestro/buy-quote.yaml b/.cursor/skills/build-and-test/maestro/buy-quote.yaml index 47e00bac..8d3b8b82 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote.yaml @@ -59,10 +59,16 @@ env: commands: [{ tapOn: "Not Now" }] # --- navigate to the Buy (Buy Crypto) scene --- -- tapOn: "Buy" -- extendedWaitUntil: - visible: "Amount USD" - timeout: 15000 +# The tab bar stays hidden for up to ~20s after login while "Buy" is already +# in the hierarchy, so a single tap can land on nothing. Retry until the Buy +# scene shows. +- retry: + maxRetries: 3 + commands: + - tapOn: "Buy" + - extendedWaitUntil: + visible: "Amount USD" + timeout: 10000 # --- enter the $500 fiat amount --- # NOTE: do NOT call `hideKeyboard` here — on this debug build it forces a re-layout From 7de0ddccd2baad03dc029c69c9ae64edff8cb010 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 16:22:08 -0700 Subject: [PATCH 05/13] Fix select-swap-pair wallet picks The search field holds the typed text, so a plain name match hit the field instead of the first row; pick match 1 of a contains-regex. Header env uses ${KEY || default} so runFlow env wins, and the Exchange tap retries through the post-login hidden tab bar. --- .../maestro/common/select-swap-pair.yaml | 43 +++++++++++++------ 1 file changed, 30 insertions(+), 13 deletions(-) diff --git a/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml b/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml index 39f04cf1..5e45ed44 100644 --- a/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml +++ b/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml @@ -2,8 +2,8 @@ # enter fiat amount → Next → wait for a quote → optionally force a provider. # Assumes logged-in app (compose login-if-needed + dismiss-startup-modals first). # PARAMS: -# SRC_WALLET / DST_WALLET — text typed into "Search Wallets" AND the regex -# matched against the filtered result (e.g. ".*Bitcoin.*") +# SRC_WALLET / DST_WALLET — plain text typed into "Search Wallets" (e.g. +# "My Bitcoin"); the first filtered row is picked # FIAT_AMOUNT — fiat number typed into the amount field # PROVIDER — exact provider name to force (e.g. "Maya Protocol"). # Forcing works by SELECTING it in the provider sheet; @@ -23,15 +23,24 @@ appId: ${APP_ID} env: APP_ID: co.edgesecure.app - SRC_WALLET: ".*Bitcoin.*" - DST_WALLET: ".*Ethereum.*" - FIAT_AMOUNT: "16" - PROVIDER: "" + # `${KEY || default}` so a caller's runFlow env wins over the default. + SRC_WALLET: ${SRC_WALLET || "Bitcoin"} + DST_WALLET: ${DST_WALLET || "Ethereum"} + FIAT_AMOUNT: ${FIAT_AMOUNT || "16"} + PROVIDER: ${PROVIDER || ""} --- - extendedWaitUntil: visible: "Exchange" timeout: 30000 -- tapOn: "Exchange" +# The tab bar stays hidden for up to ~20s after login while "Exchange" is +# already in the hierarchy, so retry the tap until the Exchange scene shows. +- retry: + maxRetries: 3 + commands: + - tapOn: "Exchange" + - extendedWaitUntil: + visible: "Select Source Wallet|I have" + timeout: 10000 - runFlow: when: visible: "Select Source Wallet" @@ -44,12 +53,16 @@ env: visible: "Not Now" commands: - tapOn: "Not Now" + # The search field holds the typed text, so it is always match 0; the + # rows' labels contain the wallet name, so match 1 is the first row. - extendedWaitUntil: - visible: ${SRC_WALLET} + visible: + text: ".*${SRC_WALLET}.*" + index: 1 timeout: 15000 - tapOn: - text: ${SRC_WALLET} - index: 0 + text: ".*${SRC_WALLET}.*" + index: 1 - runFlow: when: visible: "Select Receiving Wallet" @@ -62,12 +75,16 @@ env: visible: "Not Now" commands: - tapOn: "Not Now" + # The search field holds the typed text, so it is always match 0; the + # rows' labels contain the wallet name, so match 1 is the first row. - extendedWaitUntil: - visible: ${DST_WALLET} + visible: + text: ".*${DST_WALLET}.*" + index: 1 timeout: 15000 - tapOn: - text: ${DST_WALLET} - index: 0 + text: ".*${DST_WALLET}.*" + index: 1 - extendedWaitUntil: visible: "Tap to edit" timeout: 15000 From 49cf984ba792f554eab450be44eccceb863cd87c Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 16:41:18 -0700 Subject: [PATCH 06/13] Gate PIN taps on the PIN screen YOLO auto-login can finish mid-entry and leave the PIN screen, so a later tap found no digit and failed the flow. Each tap now runs only while Exit PIN is visible. --- .../maestro/common/login-if-needed.yaml | 17 ++++++++++++++++- 1 file changed, 16 insertions(+), 1 deletion(-) diff --git a/.cursor/skills/build-and-test/maestro/common/login-if-needed.yaml b/.cursor/skills/build-and-test/maestro/common/login-if-needed.yaml index ba283dba..8a4e654e 100644 --- a/.cursor/skills/build-and-test/maestro/common/login-if-needed.yaml +++ b/.cursor/skills/build-and-test/maestro/common/login-if-needed.yaml @@ -6,10 +6,13 @@ # ~/.config/edge-secrets/test-accounts.json). # GOTCHA: taps MUST stay spaced (waitToSettleTimeoutMs) — fast taps drop digits → # wrong PIN → exponential lockout. Never tighten these. +# YOLO auto-login can finish mid-entry and leave the PIN screen, so each tap +# is gated on the PIN screen still showing. appId: ${APP_ID} env: APP_ID: co.edgesecure.app - PIN_DIGIT: "0" + # `${KEY || default}` so a caller's runFlow env wins over the default. + PIN_DIGIT: ${PIN_DIGIT || "0"} --- - runFlow: when: @@ -18,12 +21,24 @@ env: - tapOn: text: ${PIN_DIGIT} waitToSettleTimeoutMs: 900 +- runFlow: + when: + visible: "Exit PIN" + commands: - tapOn: text: ${PIN_DIGIT} waitToSettleTimeoutMs: 900 +- runFlow: + when: + visible: "Exit PIN" + commands: - tapOn: text: ${PIN_DIGIT} waitToSettleTimeoutMs: 900 +- runFlow: + when: + visible: "Exit PIN" + commands: - tapOn: text: ${PIN_DIGIT} waitToSettleTimeoutMs: 1500 From ce30bfcb41d551675350c64a952549babcff6a12 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 17:06:47 -0700 Subject: [PATCH 07/13] Clear wallet search before typing The Search Wallets text survives navigation, so inputText appended to a stale query and the row tap hit the search field. --- .cursor/skills/build-and-test/maestro/common/find-wallet.yaml | 2 ++ .../skills/build-and-test/maestro/common/send-to-address.yaml | 2 ++ 2 files changed, 4 insertions(+) diff --git a/.cursor/skills/build-and-test/maestro/common/find-wallet.yaml b/.cursor/skills/build-and-test/maestro/common/find-wallet.yaml index df5db1af..69633439 100644 --- a/.cursor/skills/build-and-test/maestro/common/find-wallet.yaml +++ b/.cursor/skills/build-and-test/maestro/common/find-wallet.yaml @@ -18,6 +18,8 @@ env: visible: "Search Wallets" timeout: 20000 - tapOn: "Search Wallets" +# The search text survives navigation; clear it or inputText appends. +- eraseText: 40 - inputText: ${SEARCH_TERM} - runFlow: when: diff --git a/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml b/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml index e3462039..daee59cb 100644 --- a/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml +++ b/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml @@ -10,6 +10,8 @@ appId: co.edgesecure.app - tapOn: "Assets" - extendedWaitUntil: {visible: "Search Wallets", timeout: 20000} - tapOn: "Search Wallets" +# The search text survives navigation; clear it or inputText appends. +- eraseText: 40 - inputText: ${WALLET_SEARCH} - extendedWaitUntil: {visible: ".*${WALLET_SEARCH}.*", timeout: 15000} - tapOn: {text: ".*${WALLET_SEARCH}.*", index: 0} From ea24bd773c5912d974378a760ab2f68cb1b3a9de Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 17:21:48 -0700 Subject: [PATCH 08/13] Gate buy-quote PIN taps, clear amount --- .../maestro/buy-quote-input.yaml | 22 +++++++----------- .../build-and-test/maestro/buy-quote.yaml | 23 +++++++------------ 2 files changed, 16 insertions(+), 29 deletions(-) diff --git a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml index a7861e5b..4c4bf277 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml @@ -11,20 +11,12 @@ env: - extendedWaitUntil: visible: "Exit PIN" timeout: 40000 -- assertVisible: "1" -- assertVisible: "0" -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 1500 +# YOLO auto-login can finish mid-entry, so the taps go through the gated +# login-if-needed flow instead of four unconditional taps. +- runFlow: + file: common/login-if-needed.yaml + env: + PIN_DIGIT: ${PIN_DIGIT} - extendedWaitUntil: visible: "Buy" timeout: 30000 @@ -54,4 +46,6 @@ env: visible: "Amount USD" timeout: 10000 - tapOn: "Amount USD" +# The Buy scene remembers the last amount; clear it or inputText appends. +- eraseText: 20 - inputText: ${BUY_AMOUNT} diff --git a/.cursor/skills/build-and-test/maestro/buy-quote.yaml b/.cursor/skills/build-and-test/maestro/buy-quote.yaml index 8d3b8b82..0bf2125b 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote.yaml @@ -28,21 +28,12 @@ env: - extendedWaitUntil: visible: "Exit PIN" timeout: 40000 -# make sure the whole keypad is rendered & interactive before tapping -- assertVisible: "1" -- assertVisible: "0" -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 900 -- tapOn: - text: ${PIN_DIGIT} - waitToSettleTimeoutMs: 1500 +# YOLO auto-login can finish mid-entry, so the taps go through the gated +# login-if-needed flow instead of four unconditional taps. +- runFlow: + file: common/login-if-needed.yaml + env: + PIN_DIGIT: ${PIN_DIGIT} # --- wait for the main tab bar, then dismiss optional post-login modals --- - extendedWaitUntil: @@ -75,6 +66,8 @@ env: # that reliably triggers the RN Fabric text-measure SIGABRT. The USD + converted-BTC # fields and the exchange-rate line are all above the keyboard, so leave it up. - tapOn: "Amount USD" +# The Buy scene remembers the last amount; clear it or inputText appends. +- eraseText: 20 - inputText: ${BUY_AMOUNT} # --- let the live quote resolve, then capture --- From ae0cf03bd78ee5b77a04ced09a58eb92c2563886 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Wed, 30 Sep 2026 17:21:48 -0700 Subject: [PATCH 09/13] Retry ramp region tap behind keyboard --- .../maestro/common/ramp-set-region-fiat.yaml | 14 +++++++++----- 1 file changed, 9 insertions(+), 5 deletions(-) diff --git a/.cursor/skills/build-and-test/maestro/common/ramp-set-region-fiat.yaml b/.cursor/skills/build-and-test/maestro/common/ramp-set-region-fiat.yaml index c84f74dc..ecb94422 100644 --- a/.cursor/skills/build-and-test/maestro/common/ramp-set-region-fiat.yaml +++ b/.cursor/skills/build-and-test/maestro/common/ramp-set-region-fiat.yaml @@ -11,11 +11,15 @@ env: FIAT_SEARCH: "USD" FIAT_ROW: "USD United States Dollar" --- -- tapOn: - text: ".*(Select your region|USA|Germany|United Kingdom|Canada|Australia).*" -- extendedWaitUntil: - visible: "Search region" - timeout: 10000 +# With the amount keyboard up, the first tap only dismisses the keyboard. +- retry: + maxRetries: 2 + commands: + - tapOn: + text: ".*(Select your region|USA|Germany|United Kingdom|Canada|Australia).*" + - extendedWaitUntil: + visible: "Search region" + timeout: 5000 - tapOn: "Search region" - inputText: ${COUNTRY_SEARCH} - waitForAnimationToEnd: From 24daa9187f8099e7f16aebc5e54095c1e7fa42ae Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Thu, 1 Oct 2026 02:04:04 -0700 Subject: [PATCH 10/13] Cover orch flow commands in XCUITest runner A preflight survey of the agent task flows under /tmp found flows the interpreter rejected, almost all on tapOn point: and hideKeyboard. Add those and the remaining parity commands so every surveyed flow runs natively: - tapOn point: as screen percentages or absolute points, and relative to the element when a selector is present - longPressOn (3s hold, as Maestro on iOS) - hideKeyboard with Maestro's iOS behavior: nothing when no keyboard is up, else a short swipe up from the screen center, then a short swipe left - copyTextFrom and pasteText, with maestro.copiedText in the script context - inputRandomText - enabled: on selectors - scrollUntilVisible centerElement: - back and pressKey: back as no-ops, as on Maestro iOS launchApp clearState: true stays rejected. --- .../references/xcuitest-interpreter.md | 16 +- .../EdgeFlowRunner/FlowInterpreter.swift | 187 +++++++++++++++--- .../EdgeFlowRunner/FlowPreflight.swift | 35 +++- .../xcuitest/EdgeFlowRunner/FlowScript.swift | 7 +- 4 files changed, 209 insertions(+), 36 deletions(-) diff --git a/.cursor/skills/build-and-test/references/xcuitest-interpreter.md b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md index 5942950d..83cb6437 100644 --- a/.cursor/skills/build-and-test/references/xcuitest-interpreter.md +++ b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md @@ -24,12 +24,16 @@ and source hash (`scripts/xcuitest-build.sh`, cached under |---|---| | `launchApp` | `appId`, `stopApp: false` (activate if running). `clearState: true` is rejected because it wipes the roster accounts. Passes `-EdgeTestAnimations ` unless `--animations on`. | | `stopApp` | | -| `tapOn` | `text`, `id`, `index`, `waitToSettleTimeoutMs`, `retryTapIfNoChange` | -| `assertVisible` / `assertNotVisible` | Maestro timeouts: 17s (7s when `optional`), minus time since the last interaction | +| `tapOn` / `longPressOn` | `text`, `id`, `index`, `enabled`, `point`, `waitToSettleTimeoutMs`, `retryTapIfNoChange`. `point` alone is a screen position: `"50%,80%"` (whole percentages of the screen) or `"120,640"` (points). `point` next to `text` or `id` is relative to the matched element. `longPressOn` holds for 3s, as Maestro does on iOS | +| `assertVisible` / `assertNotVisible` | `enabled: true/false` narrows the match to enabled or disabled elements. Maestro timeouts: 17s (7s when `optional`), minus time since the last interaction | | `extendedWaitUntil` | `visible` / `notVisible`, `timeout` | | `runFlow` | `file`, inline `commands`, `env`, `when` (`visible`, `notVisible`, `true`, `platform`) | -| `inputText` / `eraseText` / `pressKey` | `pressKey`: Enter, Backspace, Home. `eraseText` defaults to 50 characters. When no element has keyboard focus (a hidden input, such as the PIN entry) the runner taps the on-screen keys instead, so only characters with their own key (digits, the current letter case, space) can be typed | -| `scroll` / `scrollUntilVisible` / `swipe` | `scrollUntilVisible`: `element`, `direction`, `timeout`, `visibilityPercentage`, `waitToSettleTimeoutMs`. `swipe`: `from` + `direction`, or `start` / `end` points (`"50%,80%"`), `duration` | +| `inputText` / `eraseText` / `pressKey` | `pressKey`: Enter, Backspace, Home, and Back (does nothing, as on Maestro iOS). `eraseText` defaults to 50 characters. When no element has keyboard focus (a hidden input, such as the PIN entry) the runner taps the on-screen keys instead, so only characters with their own key (digits, the current letter case, space) can be typed | +| `inputRandomText` | `length` (default 8). Types random lowercase letters | +| `copyTextFrom` / `pasteText` | `copyTextFrom` takes a selector and stores the element's text (title, else value, else placeholder, else label), also as `maestro.copiedText` for `evalScript` and `${}`. `pasteText` types it, and types nothing when nothing was copied | +| `hideKeyboard` | Maestro's iOS behavior: nothing when no keyboard is up, else a short swipe up from the screen center, then a short swipe left if the keyboard is still there. Fails when the keyboard survives both, which Maestro also does; tap a non-interactive element in that case | +| `back` | Accepted and does nothing, the same as Maestro on iOS | +| `scroll` / `scrollUntilVisible` / `swipe` | `scrollUntilVisible`: `element`, `direction`, `timeout`, `visibilityPercentage`, `centerElement`, `waitToSettleTimeoutMs`. With `centerElement`, an element that is on screen but outside the center band is dragged by its own distance from the screen center (Maestro repeats the full swipe, which can carry the element past the band and off screen). `swipe`: `from` + `direction`, or `start` / `end` points (`"50%,80%"`), `duration` | | `repeat` / `retry` | `repeat`: `times`, `while`. `retry`: `maxRetries`, `file` or `commands` | | `evalScript` | Full JavaScript (JavaScriptCore). `output.*` persists for the whole run | | `waitForAnimationToEnd` | Two consecutive identical screenshots, `timeout` default 15s | @@ -72,8 +76,6 @@ Spinners do not hold the quiescence wait in any mode. ## Not supported (preflight rejects) -`inputRandomText`, `copyTextFrom`, `pasteText`, `longPressOn`, `hideKeyboard`, -`assertVisible enabled:`, `scrollUntilVisible centerElement:`, `tapOn point:`, -`launchApp clearState: true`, and any command missing from the table above. +`launchApp clearState: true` and any command missing from the table above. To drive a flow that needs one, rewrite the step or run that flow on the maestro CLI and name it in the run report. diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift index 0cba7255..e2573685 100644 --- a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift @@ -23,6 +23,14 @@ struct RunOptions { } } +/// One on-screen element a selector matched. `text` is what Maestro's iOS +/// driver calls the element's text: title, else value, else placeholder, else +/// label. +struct Match { + var frame: CGRect + var text: String? +} + /// Element selector: Maestro `text` and `id` are case-insensitive regexes /// (dot matches newline) that must match the whole attribute, or equal it /// literally. `text` is checked against label, value and placeholderValue. @@ -30,12 +38,14 @@ struct Selector: CustomStringConvertible { var text: String? var id: String? var index: Int? + var enabled: Bool? var description: String { var parts: [String] = [] if let text = text { parts.append("text=\"\(text)\"") } if let id = id { parts.append("id=\"\(id)\"") } if let index = index { parts.append("index=\(index)") } + if let enabled = enabled { parts.append("enabled=\(enabled)") } return parts.joined(separator: " ") } } @@ -46,6 +56,7 @@ final class FlowInterpreter { private var appId = "co.edgesecure.app" private var app = XCUIApplication(bundleIdentifier: "co.edgesecure.app") private var screenBounds: CGRect? + private var copiedText: String? private var lastInteraction = Date() private let started = Date() @@ -145,15 +156,24 @@ final class FlowInterpreter { if let appId = map["appId"] as? String { selectApp(try script.interpolate(appId)) } app.terminate() interacted() - case "tapOn": - let selector = try parseSelector(args) - let timeout = optional ? options.optionalLookupTimeout : options.lookupTimeout - guard let frame = try waitForElement(selector, timeout: timeout) else { - throw FlowError("element not found: \(selector)") + case "tapOn", "longPressOn": + let long = name == "longPressOn" + let target: XCUICoordinate + if let point = map["point"], map["text"] == nil, map["id"] == nil { + target = try screenCoordinate(point) + } else { + let selector = try parseSelector(args) + let timeout = optional ? options.optionalLookupTimeout : options.lookupTimeout + guard let frame = try waitForElement(selector, timeout: timeout) else { + throw FlowError("element not found: \(selector)") + } + // A point next to a selector is relative to the element's frame. + let offset = try map["point"].map { try point($0, in: frame) } ?? CGPoint(x: frame.width / 2, y: frame.height / 2) + target = coordinate(CGPoint(x: frame.minX + offset.x, y: frame.minY + offset.y)) } let before = map["retryTapIfNoChange"] as? Bool == true ? screenPixels() : nil - tap(frame) - if let before = before, screenPixels() == before { tap(frame) } + press(target, long: long) + if let before = before, screenPixels() == before { press(target, long: long) } if let ms = try optionalNumber(map["waitToSettleTimeoutMs"]) { waitForSettle(timeout: ms / 1000) } case "assertVisible": let selector = try parseSelector(args) @@ -192,6 +212,28 @@ final class FlowInterpreter { case "inputText": let text = try script.interpolate(stringArg(args, key: "text")) try typeKeys(text) + case "inputRandomText": + let length = Int(try optionalNumber(map.isEmpty ? args : map["length"]) ?? 8) + let letters = "abcdefghijklmnopqrstuvwxyz" + try typeKeys(String((0..<(length > 0 ? length : 8)).compactMap { _ in letters.randomElement() })) + case "copyTextFrom": + let selector = try parseSelector(args) + let timeout = optional ? options.optionalLookupTimeout : options.lookupTimeout + guard let match = try waitForMatch(selector, timeout: timeout) else { + throw FlowError("element not found: \(selector)") + } + guard let text = match.text else { throw FlowError("\(selector) has no text to copy") } + copiedText = text + script.setCopiedText(text) + return "copied \"\(text)\"" + case "pasteText": + guard let text = copiedText else { return "nothing copied, typed nothing" } + try typeKeys(text) + case "hideKeyboard": + return try hideKeyboard() + case "back": + // Maestro's iOS driver implements back as an empty function. + return "no-op on iOS" case "eraseText": let count = Int(try optionalNumber(map.isEmpty ? args : map["charactersToErase"]) ?? 50) try typeKeys(String(repeating: XCUIKeyboardKey.delete.rawValue, count: count)) @@ -203,6 +245,8 @@ final class FlowInterpreter { case "home": XCUIDevice.shared.press(.home) interacted() + // Maestro's iOS driver has no back key and ignores it. + case "back": return "no-op on iOS" default: throw FlowError("unsupported pressKey '\(key)'") } case "scroll": @@ -260,6 +304,29 @@ final class FlowInterpreter { return nil } + /// Maestro's iOS hideKeyboard: nothing when no keyboard is up, else a 3% + /// swipe up from the screen center, then a 3% swipe left if the keyboard + /// survived the first. Fails when the keyboard is still up afterwards. + private func hideKeyboard() throws -> String? { + if keyboardGone(within: 0) { return "no keyboard showing" } + let bounds = screen() + let center = CGPoint(x: bounds.width * 0.5, y: bounds.height * 0.5) + drag(from: center, to: CGPoint(x: center.x, y: bounds.height * 0.47), duration: 0.05) + if keyboardGone(within: 2) { return nil } + drag(from: center, to: CGPoint(x: bounds.width * 0.47, y: center.y), duration: 0.05) + if keyboardGone(within: 2) { return nil } + throw FlowError("keyboard still showing after hideKeyboard; tap a non-interactive element instead") + } + + private func keyboardGone(within timeout: TimeInterval) -> Bool { + let deadline = Date().addingTimeInterval(timeout) + while app.keyboards.firstMatch.exists { + if Date() >= deadline { return false } + Thread.sleep(forTimeInterval: 0.15) + } + return true + } + private func scrollUntilVisible(_ map: [String: Any]) throws -> String? { guard let element = map["element"] else { throw FlowError("scrollUntilVisible needs element") } let selector = try parseSelector(element) @@ -267,24 +334,64 @@ final class FlowInterpreter { let timeout = (try optionalNumber(map["timeout"]) ?? 20000) / 1000 let percent = (try optionalNumber(map["visibilityPercentage"]) ?? 100) / 100 let settle = (try optionalNumber(map["waitToSettleTimeoutMs"])).map { $0 / 1000 } + let center = map["centerElement"] as? Bool ?? false let deadline = Date().addingTimeInterval(timeout) var swipes = 0 + // Maestro gives up centering after 5 tries: the element may sit at the + // end of a list that cannot scroll further. + var centerTries = 0 while true { - if let frame = try findElement(selector), visibleFraction(frame) >= percent { - return swipes == 0 ? nil : "visible after \(swipes) swipe(s)" + // Set while the element is on screen but short of the center band. + var centering: CGRect? + if let frame = try findElement(selector) { + let visible = visibleFraction(frame) + let done: Bool + if center, visible > 0.1, centerTries <= 4 { + done = nearCenter(frame, scrolling: direction) + centerTries += 1 + if !done { centering = frame } + } else { + done = visible >= percent + } + if done { return swipes == 0 ? nil : "visible after \(swipes) swipe(s)" } } if Date() >= deadline { throw FlowError("\(selector) not visible after scrolling \(timeout)s") } let bounds = screen() + let vertical = direction != "LEFT" && direction != "RIGHT" let (from, to): (CGPoint, CGPoint) switch direction { + // Maestro repeats its full swipe here, which can carry an on-screen + // element clean past the band. Dragging by the element's own distance + // from the center lands it there instead. + case _ where centering != nil && vertical: + let offset = max(-bounds.height * 0.4, min(bounds.height * 0.4, centering!.midY - bounds.midY)) + (from, to) = (CGPoint(x: bounds.midX, y: bounds.midY + offset / 2), CGPoint(x: bounds.midX, y: bounds.midY - offset / 2)) + case _ where centering != nil: + let offset = max(-bounds.width * 0.4, min(bounds.width * 0.4, centering!.midX - bounds.midX)) + (from, to) = (CGPoint(x: bounds.midX + offset / 2, y: bounds.midY), CGPoint(x: bounds.midX - offset / 2, y: bounds.midY)) case "UP": (from, to) = (CGPoint(x: bounds.midX, y: bounds.height * 0.3), CGPoint(x: bounds.midX, y: bounds.height * 0.7)) case "LEFT": (from, to) = (CGPoint(x: bounds.width * 0.3, y: bounds.midY), CGPoint(x: bounds.width * 0.7, y: bounds.midY)) case "RIGHT": (from, to) = (CGPoint(x: bounds.width * 0.7, y: bounds.midY), CGPoint(x: bounds.width * 0.3, y: bounds.midY)) default: (from, to) = (CGPoint(x: bounds.midX, y: bounds.height * 0.7), CGPoint(x: bounds.midX, y: bounds.height * 0.3)) } - drag(from: from, to: to, duration: 0.4) + // A centering swipe holds at its end so the list stops under the + // finger: a fling would carry the element past the center band. + drag(from: from, to: to, duration: 0.4, hold: center ? 0.3 : 0.05) swipes += 1 - if let settle = settle { waitForSettle(timeout: settle) } + if let settle = settle ?? (center ? 2 : nil) { waitForSettle(timeout: settle) } + } + } + + /// Maestro's isElementNearScreenCenter: the element's center has reached + /// the middle of the screen, give or take a fifth of it, coming from the + /// side the scroll brings it in from. + private func nearCenter(_ frame: CGRect, scrolling direction: String) -> Bool { + let bounds = screen() + switch direction { + case "UP": return frame.midY > bounds.midY - bounds.height / 5 + case "LEFT": return frame.midX > bounds.midX - bounds.width / 5 + case "RIGHT": return frame.midX < bounds.midX + bounds.width / 5 + default: return frame.midY < bounds.midY + bounds.height / 5 } } @@ -354,6 +461,10 @@ final class FlowInterpreter { if let text = map["text"] { selector.text = try script.interpolate("\(text)") } if let id = map["id"] { selector.id = try script.interpolate("\(id)") } if let index = try optionalNumber(map["index"]) { selector.index = Int(index) } + if let enabled = map["enabled"] { + // JSON booleans arrive as NSNumber, which prints as 1 or 0. + selector.enabled = try (enabled as? Bool) ?? (script.interpolate("\(enabled)").lowercased() == "true") + } if selector.text == nil, selector.id == nil { throw FlowError("selector needs text or id") } return selector } @@ -370,6 +481,19 @@ final class FlowInterpreter { } } + /// Like waitForElement, but always walks a snapshot so the match carries + /// its text. + private func waitForMatch(_ selector: Selector, timeout: TimeInterval) throws -> Match? { + let deadline = Date().addingTimeInterval(timeout) + while true { + let matches = try snapshotMatches(selector) + let index = selector.index ?? 0 + if index < matches.count { return matches[index] } + if Date() >= deadline { return nil } + Thread.sleep(forTimeInterval: 0.15) + } + } + private func waitForAbsence(_ selector: Selector, timeout: TimeInterval) throws -> Bool { let deadline = Date().addingTimeInterval(timeout) while true { @@ -388,19 +512,21 @@ final class FlowInterpreter { private func findElement(_ selector: Selector) throws -> CGRect? { if selector.index == nil { let first = query(selector).firstMatch - guard first.exists else { return nil } - let frame = first.frame + // The frame comes from a throwing snapshot: reading `frame` of an + // element that left between the two calls records a test failure. + guard first.exists, let frame = try? first.snapshot().frame else { return nil } if visibleFraction(frame) > 0, !query(selector, in: first).firstMatch.exists { return frame } } let matches = try snapshotMatches(selector) let index = selector.index ?? 0 - return index < matches.count ? matches[index] : nil + return index < matches.count ? matches[index].frame : nil } private func query(_ selector: Selector, in root: XCUIElement? = nil) -> XCUIElementQuery { var query = (root ?? app).descendants(matching: .any) if let id = selector.id { query = query.matching(predicate(fields: ["identifier"], pattern: id)) } if let text = selector.text { query = query.matching(predicate(fields: ["label", "value", "placeholderValue"], pattern: text)) } + if let enabled = selector.enabled { query = query.matching(NSPredicate(format: "enabled == %@", NSNumber(value: enabled))) } return query } @@ -427,7 +553,7 @@ final class FlowInterpreter { try? NSRegularExpression(pattern: "(?ism)" + pattern) } - private func snapshotMatches(_ selector: Selector) throws -> [CGRect] { + private func snapshotMatches(_ selector: Selector) throws -> [Match] { let textRegex = selector.text.flatMap(matchRegex) let idRegex = selector.id.flatMap(matchRegex) func fullMatch(_ regex: NSRegularExpression?, _ pattern: String, _ value: String?) -> Bool { @@ -442,6 +568,7 @@ final class FlowInterpreter { return false } func matches(_ node: XCUIElementSnapshot) -> Bool { + if let enabled = selector.enabled, node.isEnabled != enabled { return false } if let id = selector.id, !fullMatch(idRegex, id, node.identifier) { return false } if let text = selector.text { let values = [node.label, node.value as? String, node.placeholderValue] @@ -449,18 +576,21 @@ final class FlowInterpreter { } return true } - var found: [CGRect] = [] + func text(_ node: XCUIElementSnapshot) -> String? { + [node.title, node.value as? String, node.placeholderValue, node.label].compactMap { $0 }.first { !$0.isEmpty } + } + var found: [Match] = [] // Returns whether the subtree holds a match; a node only counts when no // descendant matches (Maestro's deepestMatchingElement). func walk(_ node: XCUIElementSnapshot) -> Bool { var childMatched = false for child in node.children where walk(child) { childMatched = true } let isMatch = matches(node) - if isMatch, !childMatched, visibleFraction(node.frame) > 0 { found.append(node.frame) } + if isMatch, !childMatched, visibleFraction(node.frame) > 0 { found.append(Match(frame: node.frame, text: text(node))) } return isMatch || childMatched } _ = walk(try app.snapshot()) - return found.sorted { $0.minY != $1.minY ? $0.minY < $1.minY : $0.minX < $1.minX } + return found.sorted { $0.frame.minY != $1.frame.minY ? $0.frame.minY < $1.frame.minY : $0.frame.minX < $1.frame.minX } } private func visibleFraction(_ frame: CGRect) -> CGFloat { @@ -489,19 +619,30 @@ final class FlowInterpreter { app.coordinate(withNormalizedOffset: .zero).withOffset(CGVector(dx: point.x, dy: point.y)) } - private func tap(_ frame: CGRect) { - coordinate(CGPoint(x: frame.midX, y: frame.midY)).tap() + /// A screen point as Maestro writes it: "x%,y%" becomes a normalized + /// offset on the app, "x,y" an offset in points from its origin. + private func screenCoordinate(_ raw: Any) throws -> XCUICoordinate { + let text = try script.interpolate("\(raw)") + guard text.contains("%") else { return coordinate(try point(text, in: screen())) } + let unit = try point(text, in: CGRect(x: 0, y: 0, width: 1, height: 1)) + guard (0...1).contains(unit.x), (0...1).contains(unit.y) else { throw FlowError("bad point '\(text)'") } + return app.coordinate(withNormalizedOffset: CGVector(dx: unit.x, dy: unit.y)) + } + + /// Maestro holds a long press for 3 seconds on iOS. + private func press(_ target: XCUICoordinate, long: Bool) { + if long { target.press(forDuration: 3) } else { target.tap() } interacted() } - private func drag(from: CGPoint, to: CGPoint, duration: TimeInterval) { + private func drag(from: CGPoint, to: CGPoint, duration: TimeInterval, hold: TimeInterval = 0.05) { let distance = hypot(to.x - from.x, to.y - from.y) let velocity = XCUIGestureVelocity(rawValue: distance / CGFloat(max(duration, 0.05))) - coordinate(from).press(forDuration: 0.05, thenDragTo: coordinate(to), withVelocity: velocity, thenHoldForDuration: 0.05) + coordinate(from).press(forDuration: 0.05, thenDragTo: coordinate(to), withVelocity: velocity, thenHoldForDuration: hold) interacted() } - /// Accepts "x%,y%" (screen-relative) or "x,y" (points). + /// Accepts "x%,y%" (relative to `bounds`) or "x,y" (points). private func point(_ raw: Any, in bounds: CGRect) throws -> CGPoint { let text = try script.interpolate("\(raw)") let parts = text.split(separator: ",").map { $0.trimmingCharacters(in: .whitespaces) } diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift index 5b61d328..3b521910 100644 --- a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift @@ -5,16 +5,23 @@ import Foundation /// implement. The run fails on any report: there is no Maestro fallback. enum FlowPreflight { static let common: Set = ["label", "optional"] - static let selectorKeys: Set = ["text", "id", "index"] + static let selectorKeys: Set = ["text", "id", "index", "enabled"] static let conditionKeys: Set = ["visible", "notVisible", "true", "platform"] static let configKeys: Set = ["appId", "env", "name", "tags", "jsEngine"] - static let pressKeys: Set = ["enter", "backspace", "home"] + static let tapKeys: Set = ["point", "waitToSettleTimeoutMs", "retryTapIfNoChange"] + static let pressKeys: Set = ["enter", "backspace", "home", "back"] /// Argument keys each command accepts in map form (plus `common`). static let commandKeys: [String: Set] = [ "launchApp": ["appId", "clearState", "stopApp"], "stopApp": ["appId"], - "tapOn": selectorKeys.union(["waitToSettleTimeoutMs", "retryTapIfNoChange"]), + "tapOn": selectorKeys.union(tapKeys), + "longPressOn": selectorKeys.union(tapKeys), + "copyTextFrom": selectorKeys, + "pasteText": [], + "inputRandomText": ["length"], + "hideKeyboard": [], + "back": [], "assertVisible": selectorKeys, "assertNotVisible": selectorKeys, "extendedWaitUntil": ["visible", "notVisible", "timeout"], @@ -23,7 +30,7 @@ enum FlowPreflight { "eraseText": ["charactersToErase"], "pressKey": ["key"], "scroll": [], - "scrollUntilVisible": ["element", "direction", "timeout", "visibilityPercentage", "speed", "waitToSettleTimeoutMs"], + "scrollUntilVisible": ["element", "direction", "timeout", "visibilityPercentage", "centerElement", "speed", "waitToSettleTimeoutMs"], "swipe": ["from", "direction", "duration", "start", "end", "waitToSettleTimeoutMs"], "repeat": ["times", "while", "commands"], "retry": ["maxRetries", "commands", "file", "_flow"], @@ -35,7 +42,8 @@ enum FlowPreflight { /// Commands that may be written as a bare name or with a scalar argument. static let scalarForms: Set = [ "launchApp", "stopApp", "tapOn", "assertVisible", "assertNotVisible", "inputText", "eraseText", - "pressKey", "scroll", "evalScript", "waitForAnimationToEnd", "takeScreenshot" + "pressKey", "scroll", "evalScript", "waitForAnimationToEnd", "takeScreenshot", "longPressOn", + "copyTextFrom", "pasteText", "inputRandomText", "hideKeyboard", "back" ] static func problems(in flow: [String: Any]) -> [String] { @@ -99,6 +107,11 @@ enum FlowPreflight { if let clear = map["clearState"] as? Bool, clear { found.append("\(at): launchApp clearState: true is unsupported (it would wipe the sim's roster accounts)") } + case "tapOn", "longPressOn": + if let point = map["point"] { checkPoint("\(point)", at: at, into: &found) } + if map["point"] == nil, map["text"] == nil, map["id"] == nil { + found.append("\(at): \(name) needs text, id or point") + } case "extendedWaitUntil": for key in ["visible", "notVisible"] { if let selector = map[key] { checkSelector(selector, at: "\(at) \(key)", into: &found) } @@ -139,6 +152,18 @@ enum FlowPreflight { } } + /// "x%,y%" with both in 0...100, or "x,y" in points. + private static func checkPoint(_ point: String, at: String, into found: inout [String]) { + if point.contains("${") { return } + let parts = point.split(separator: ",", omittingEmptySubsequences: false).map { $0.trimmingCharacters(in: .whitespaces) } + let percent = point.contains("%") + let valid = parts.count == 2 && parts.allSatisfy { part in + guard part.hasSuffix("%") == percent, let value = Double(percent ? String(part.dropLast()) : part) else { return false } + return value >= 0 && (!percent || value <= 100) + } + if !valid { found.append("\(at): bad point '\(point)' (want \"50%,80%\" or \"120,640\")") } + } + private static func checkKey(_ key: String, at: String, into found: inout [String]) { if !key.contains("${"), !pressKeys.contains(key.lowercased()) { found.append("\(at): unsupported pressKey '\(key)'") diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift index 6010292c..1e0304b7 100644 --- a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowScript.swift @@ -16,7 +16,7 @@ final class FlowScript { context.exceptionHandler = { [weak self] _, exception in self?.lastException = exception?.toString() ?? "unknown JavaScript error" } - context.evaluateScript("var output = {};") + context.evaluateScript("var output = {}; var maestro = { platform: 'ios' };") } /// Evaluates one JavaScript expression or statement. @@ -77,6 +77,11 @@ final class FlowScript { throw FlowError("unterminated ${ in `\(text)`") } + /// `maestro.copiedText`, which Maestro sets on copyTextFrom. + func setCopiedText(_ text: String) { + context.globalObject.forProperty("maestro").setValue(text, forProperty: "copiedText") + } + func isDefined(_ name: String) -> Bool { context.globalObject.hasProperty(name) } From 8a391d033419689f78bc53bcd09be83c6a1b7f50 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Thu, 1 Oct 2026 14:37:53 -0700 Subject: [PATCH 11/13] Add openLink to the XCUITest interpreter --- .../references/xcuitest-interpreter.md | 1 + .../xcuitest/EdgeFlowRunner/FlowInterpreter.swift | 15 +++++++++++++++ .../xcuitest/EdgeFlowRunner/FlowPreflight.swift | 6 +++++- 3 files changed, 21 insertions(+), 1 deletion(-) diff --git a/.cursor/skills/build-and-test/references/xcuitest-interpreter.md b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md index 83cb6437..58416971 100644 --- a/.cursor/skills/build-and-test/references/xcuitest-interpreter.md +++ b/.cursor/skills/build-and-test/references/xcuitest-interpreter.md @@ -24,6 +24,7 @@ and source hash (`scripts/xcuitest-build.sh`, cached under |---|---| | `launchApp` | `appId`, `stopApp: false` (activate if running). `clearState: true` is rejected because it wipes the roster accounts. Passes `-EdgeTestAnimations ` unless `--animations on`. | | `stopApp` | | +| `openLink` | Scalar URL or `link:`. Hands the URL straight to the app under test (`XCUIApplication.open`), with no "Open in Edge?" system dialog, for custom-scheme (`edge://`) and `https://` links alike. Returns before the app navigates, so follow it with an `extendedWaitUntil` on the target scene. The Android-only keys `autoVerify` and `browser` are accepted and ignored. | | `tapOn` / `longPressOn` | `text`, `id`, `index`, `enabled`, `point`, `waitToSettleTimeoutMs`, `retryTapIfNoChange`. `point` alone is a screen position: `"50%,80%"` (whole percentages of the screen) or `"120,640"` (points). `point` next to `text` or `id` is relative to the matched element. `longPressOn` holds for 3s, as Maestro does on iOS | | `assertVisible` / `assertNotVisible` | `enabled: true/false` narrows the match to enabled or disabled elements. Maestro timeouts: 17s (7s when `optional`), minus time since the last interaction | | `extendedWaitUntil` | `visible` / `notVisible`, `timeout` | diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift index e2573685..6e3f63e9 100644 --- a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowInterpreter.swift @@ -156,6 +156,21 @@ final class FlowInterpreter { if let appId = map["appId"] as? String { selectApp(try script.interpolate(appId)) } app.terminate() interacted() + case "openLink": + // XCUIApplication.open hands the URL to the app under test itself, so + // custom schemes and https links both arrive without the system's + // "Open in ?" prompt and without an associated-domains lookup. + // autoVerify and browser are Android-only and ignored. + let link = try script.interpolate(stringArg(args, key: "link")) + guard let url = URL(string: link), url.scheme != nil else { + throw FlowError("not a URL: \(link)") + } + guard #available(iOS 16.4, *) else { + throw FlowError("openLink needs iOS 16.4 or later") + } + app.open(url) + screenBounds = nil + interacted() case "tapOn", "longPressOn": let long = name == "longPressOn" let target: XCUICoordinate diff --git a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift index 3b521910..a65aa093 100644 --- a/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift +++ b/.cursor/skills/build-and-test/xcuitest/EdgeFlowRunner/FlowPreflight.swift @@ -15,6 +15,7 @@ enum FlowPreflight { static let commandKeys: [String: Set] = [ "launchApp": ["appId", "clearState", "stopApp"], "stopApp": ["appId"], + "openLink": ["link", "autoVerify", "browser"], "tapOn": selectorKeys.union(tapKeys), "longPressOn": selectorKeys.union(tapKeys), "copyTextFrom": selectorKeys, @@ -41,7 +42,7 @@ enum FlowPreflight { /// Commands that may be written as a bare name or with a scalar argument. static let scalarForms: Set = [ - "launchApp", "stopApp", "tapOn", "assertVisible", "assertNotVisible", "inputText", "eraseText", + "launchApp", "stopApp", "openLink", "tapOn", "assertVisible", "assertNotVisible", "inputText", "eraseText", "pressKey", "scroll", "evalScript", "waitForAnimationToEnd", "takeScreenshot", "longPressOn", "copyTextFrom", "pasteText", "inputRandomText", "hideKeyboard", "back" ] @@ -89,6 +90,7 @@ enum FlowPreflight { if !scalarForms.contains(name) { found.append("\(at): '\(name)' needs a map argument") } + if name == "openLink", args == nil { found.append("\(at): openLink needs a link") } if name == "pressKey", let key = args as? String { checkKey(key, at: at, into: &found) } @@ -107,6 +109,8 @@ enum FlowPreflight { if let clear = map["clearState"] as? Bool, clear { found.append("\(at): launchApp clearState: true is unsupported (it would wipe the sim's roster accounts)") } + case "openLink": + if map["link"] == nil { found.append("\(at): openLink needs link") } case "tapOn", "longPressOn": if let point = map["point"] { checkPoint("\(point)", at: at, into: &found) } if map["point"] == nil, map["text"] == nil, map["id"] == nil { From 59d8da9edf7483062e4212f5312c9fa48083cb65 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Thu, 1 Oct 2026 14:37:53 -0700 Subject: [PATCH 12/13] Open committed flows by deep link --- .../maestro/buy-quote-input.yaml | 33 +++++-- .../maestro/common/select-swap-pair.yaml | 58 ++++++++++--- .../maestro/common/send-to-address.yaml | 85 +++++++++++++------ 3 files changed, 132 insertions(+), 44 deletions(-) diff --git a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml index 4c4bf277..f9092111 100644 --- a/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml +++ b/.cursor/skills/build-and-test/maestro/buy-quote-input.yaml @@ -1,4 +1,9 @@ -# Interaction-only companion to buy-quote.yaml: PIN login -> Buy tab -> enter $500. +# Interaction-only companion to buy-quote.yaml: PIN login -> Buy scene -> enter $500. +# The Buy scene opens with one deep link, `edge://exchange/buy?buyAsset=..`. +# PARAMS: PIN_DIGIT, BUY_AMOUNT, BUY_ASSET (deep-link asset: a plugin id or +# `_`, default bitcoin), MANUAL_PATH ("true" taps the Buy +# tab instead of the link and leaves the asset as the app has it: for a task +# whose change is on the tab bar or the Buy entry). # Captures nothing — capture-buy-quote.sh grabs the proof screenshot externally # (see that script for why an external burst beats an in-flow screenshot on this build). appId: ${APP_ID} @@ -6,6 +11,8 @@ env: APP_ID: co.edgesecure.app PIN_DIGIT: "1" BUY_AMOUNT: "500" + BUY_ASSET: ${BUY_ASSET || "bitcoin"} + MANUAL_PATH: ${MANUAL_PATH || ""} --- - launchApp - extendedWaitUntil: @@ -35,16 +42,28 @@ env: visible: "Claim Your Web3 Handle" commands: - tapOn: "Not Now" +- runFlow: + when: + true: ${MANUAL_PATH != 'true'} + commands: + - openLink: "edge://exchange/buy?buyAsset=${BUY_ASSET}" + - extendedWaitUntil: + visible: "Amount USD" + timeout: 30000 # The tab bar stays hidden for up to ~20s after login while "Buy" is already # in the hierarchy, so a single tap can land on nothing. Retry until the Buy # scene shows. -- retry: - maxRetries: 3 +- runFlow: + when: + true: ${MANUAL_PATH == 'true'} commands: - - tapOn: "Buy" - - extendedWaitUntil: - visible: "Amount USD" - timeout: 10000 + - retry: + maxRetries: 3 + commands: + - tapOn: "Buy" + - extendedWaitUntil: + visible: "Amount USD" + timeout: 10000 - tapOn: "Amount USD" # The Buy scene remembers the last amount; clear it or inputText appends. - eraseText: 20 diff --git a/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml b/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml index 5e45ed44..0fb19297 100644 --- a/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml +++ b/.cursor/skills/build-and-test/maestro/common/select-swap-pair.yaml @@ -1,9 +1,23 @@ -# common/select-swap-pair.yaml — Exchange tab → pick source/receiving wallets → +# common/select-swap-pair.yaml: open the Exchange scene with the pair selected → # enter fiat amount → Next → wait for a quote → optionally force a provider. # Assumes logged-in app (compose login-if-needed + dismiss-startup-modals first). +# The pair is set with one deep link, `edge://exchange/swap?sellAsset=..&buyAsset=..`, +# which works from any scene. The wallet pickers below it run only for a side +# the link did not settle. # PARAMS: +# SRC_ASSET / DST_ASSET: deep-link asset: a plugin id (`bitcoin`) or +# `_` (`ethereum_0xa0b8..`). +# Default bitcoin / ethereum. A side named only by +# SRC_WALLET / DST_WALLET is left out of the link +# and picked by name instead. +# MANUAL_PATH: "true" skips the link: Exchange tab, then both +# wallet pickers. For a task whose change is on the +# tab bar, the Exchange entry or the wallet picker. # SRC_WALLET / DST_WALLET — plain text typed into "Search Wallets" (e.g. -# "My Bitcoin"); the first filtered row is picked +# "My Bitcoin"); the first filtered row is picked. +# Used on the manual path, for a side with no asset +# param, and when the link raises the app's picker +# (the account holds several wallets for the asset) # FIAT_AMOUNT — fiat number typed into the amount field # PROVIDER — exact provider name to force (e.g. "Maya Protocol"). # Forcing works by SELECTING it in the provider sheet; @@ -24,23 +38,43 @@ appId: ${APP_ID} env: APP_ID: co.edgesecure.app # `${KEY || default}` so a caller's runFlow env wins over the default. + # The asset lines come first: they read whether the caller named a wallet + # (`typeof`, because an unset name is not declared yet at that point). + SRC_ASSET: '${SRC_ASSET || (typeof SRC_WALLET === "string" && SRC_WALLET ? "" : "bitcoin")}' + DST_ASSET: '${DST_ASSET || (typeof DST_WALLET === "string" && DST_WALLET ? "" : "ethereum")}' + MANUAL_PATH: ${MANUAL_PATH || ""} SRC_WALLET: ${SRC_WALLET || "Bitcoin"} DST_WALLET: ${DST_WALLET || "Ethereum"} FIAT_AMOUNT: ${FIAT_AMOUNT || "16"} PROVIDER: ${PROVIDER || ""} --- -- extendedWaitUntil: - visible: "Exchange" - timeout: 30000 -# The tab bar stays hidden for up to ~20s after login while "Exchange" is -# already in the hierarchy, so retry the tap until the Exchange scene shows. -- retry: - maxRetries: 3 +- runFlow: + when: + true: ${MANUAL_PATH != 'true'} + commands: + - openLink: "edge://exchange/swap?sellAsset=${SRC_ASSET}&buyAsset=${DST_ASSET}" + - extendedWaitUntil: + visible: "Select Source Wallet|Select Receiving Wallet|I have" + timeout: 30000 + # The link returns before the scene settles on the linked pair. + - waitForAnimationToEnd: + timeout: 5000 +- runFlow: + when: + true: ${MANUAL_PATH == 'true'} commands: - - tapOn: "Exchange" - extendedWaitUntil: - visible: "Select Source Wallet|I have" - timeout: 10000 + visible: "Exchange" + timeout: 30000 + # The tab bar stays hidden for up to ~20s after login while "Exchange" is + # already in the hierarchy, so retry the tap until the Exchange scene shows. + - retry: + maxRetries: 3 + commands: + - tapOn: "Exchange" + - extendedWaitUntil: + visible: "Select Source Wallet|I have" + timeout: 10000 - runFlow: when: visible: "Select Source Wallet" diff --git a/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml b/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml index daee59cb..c1cc1332 100644 --- a/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml +++ b/.cursor/skills/build-and-test/maestro/common/send-to-address.yaml @@ -1,30 +1,65 @@ # common/send-to-address.yaml -# Assets -> search wallet -> Send -> address -> amount -> confirm slider. -# Params: WALLET_SEARCH, ADDRESS, AMOUNT. Promoted 2026-08-06 (run 1217135300337949); -# proof-frame tail steps intentionally dropped for reuse. +# Open a Send scene with address and amount filled in -> confirm slider. +# With CURRENCY_CODE set, one payment-redirect deep link does it from any scene: +# `https://edge.app/redirect/payment/?baseCurrencyCode=..&depositWalletAddress=..&baseCurrencyAmount=..`. +# Params: +# CURRENCY_CODE currency code for the link (`btc`, `usdc`). Unset, the flow +# takes the manual path. +# ADDRESS, AMOUNT recipient and amount (on the link, in units of the currency). +# WALLET_SEARCH text typed into Search Wallets to find the sending wallet: +# always on the manual path, and on the link when the account +# holds several wallets for the code and the app asks which. +# MANUAL_PATH "true" forces Assets -> wallet -> Send -> address -> amount, +# for a task whose change is on the wallet scene, the address +# modal or the amount entry. +# Promoted 2026-08-06 (run 1217135300337949); proof-frame tail steps +# intentionally dropped for reuse. appId: co.edgesecure.app +env: + CURRENCY_CODE: ${CURRENCY_CODE || ""} + MANUAL_PATH: ${MANUAL_PATH || ""} + WALLET_SEARCH: ${WALLET_SEARCH || ""} --- -- extendedWaitUntil: {visible: "Assets", timeout: 90000} -- runFlow: {when: {visible: "Not Now"}, commands: [{tapOn: "Not Now"}]} -- runFlow: {when: {visible: {id: "modal-close-button"}}, commands: [{tapOn: {id: "modal-close-button"}}]} -- tapOn: "Assets" -- extendedWaitUntil: {visible: "Search Wallets", timeout: 20000} -- tapOn: "Search Wallets" -# The search text survives navigation; clear it or inputText appends. -- eraseText: 40 -- inputText: ${WALLET_SEARCH} -- extendedWaitUntil: {visible: ".*${WALLET_SEARCH}.*", timeout: 15000} -- tapOn: {text: ".*${WALLET_SEARCH}.*", index: 0} -- extendedWaitUntil: {visible: "Send", timeout: 15000} -- tapOn: "Send" -- extendedWaitUntil: {visible: ".*Enter.*", timeout: 15000} -- tapOn: {text: ".*Enter.*"} -- tapOn: "Address" -- inputText: ${ADDRESS} -- tapOn: "Next" -- extendedWaitUntil: {visible: "Enter Value", timeout: 15000} -- tapOn: "Enter Value" -- inputText: ${AMOUNT} -- tapOn: "Done" +- runFlow: + when: + true: ${MANUAL_PATH != 'true' && CURRENCY_CODE != ''} + commands: + - openLink: "https://edge.app/redirect/payment/?baseCurrencyCode=${CURRENCY_CODE}&depositWalletAddress=${ADDRESS}&baseCurrencyAmount=${AMOUNT}" + - extendedWaitUntil: {visible: "Slide to Confirm|Select Wallet", timeout: 30000} + # Several wallets hold the code (ETH on each L2): the app asks which. + - runFlow: + when: {visible: "Select Wallet"} + commands: + - tapOn: "Search Wallets" + - inputText: ${WALLET_SEARCH} + # The search field holds the typed text, so it is match 0. + - extendedWaitUntil: {visible: {text: ".*${WALLET_SEARCH}.*", index: 1}, timeout: 15000} + - tapOn: {text: ".*${WALLET_SEARCH}.*", index: 1} +- runFlow: + when: + true: ${MANUAL_PATH == 'true' || CURRENCY_CODE == ''} + commands: + - extendedWaitUntil: {visible: "Assets", timeout: 90000} + - runFlow: {when: {visible: "Not Now"}, commands: [{tapOn: "Not Now"}]} + - runFlow: {when: {visible: {id: "modal-close-button"}}, commands: [{tapOn: {id: "modal-close-button"}}]} + - tapOn: "Assets" + - extendedWaitUntil: {visible: "Search Wallets", timeout: 20000} + - tapOn: "Search Wallets" + # The search text survives navigation; clear it or inputText appends. + - eraseText: 40 + - inputText: ${WALLET_SEARCH} + - extendedWaitUntil: {visible: ".*${WALLET_SEARCH}.*", timeout: 15000} + - tapOn: {text: ".*${WALLET_SEARCH}.*", index: 0} + - extendedWaitUntil: {visible: "Send", timeout: 15000} + - tapOn: "Send" + - extendedWaitUntil: {visible: ".*Enter.*", timeout: 15000} + - tapOn: {text: ".*Enter.*"} + - tapOn: "Address" + - inputText: ${ADDRESS} + - tapOn: "Next" + - extendedWaitUntil: {visible: "Enter Value", timeout: 15000} + - tapOn: "Enter Value" + - inputText: ${AMOUNT} + - tapOn: "Done" - extendedWaitUntil: {visible: "Slide to Confirm", timeout: 15000} - swipe: {from: {id: "confirmSliderThumb"}, direction: LEFT, duration: 1500} From 1d19ad8dc7dfc275b5535a2ee07546c490246390 Mon Sep 17 00:00:00 2001 From: Jonathan Tzeng Date: Thu, 1 Oct 2026 14:37:54 -0700 Subject: [PATCH 13/13] Document deep-link drives per driver --- .../skills/build-and-test/references/drive.md | 5 ++-- .../references/sim-testing-playbook.md | 26 ++++++++++++------- 2 files changed, 19 insertions(+), 12 deletions(-) diff --git a/.cursor/skills/build-and-test/references/drive.md b/.cursor/skills/build-and-test/references/drive.md index 37b09a9c..d677ff07 100644 --- a/.cursor/skills/build-and-test/references/drive.md +++ b/.cursor/skills/build-and-test/references/drive.md @@ -16,7 +16,8 @@ Governs step 0d-0f of `/build-and-test` (driving the app on the sim: the flow li - `[playbook]`: durable knowledge, one bullet. - `[flow]`: a NEW reusable drive sequence: name, params, one-line purpose, and the FULL yaml EMBEDDED as a fenced block (worktrees are pruned on retention; the report attachment is the durable copy). - `[flow-update]`: a change to an EXISTING library flow (genericize, new param, split into subflows): name the flow, the change, and the compatibility argument. New params MUST default to current behavior so existing callers are unaffected, and a rename/split must say so explicitly (callers get grepped at promotion). -- **Scope a `[playbook]` proposal tightly.** It earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE: if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING instead. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything an existing rule already covers. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately. +- **Scope a `[playbook]` proposal tightly.** It earns a slot ONLY if it is (1) SIM-TESTING WORKING KNOWLEDGE: how to drive, fund, enable, or verify a change in the running app (a provider floor/geo-block, an executable test-pair recipe, a funding path, a feature-enablement gotcha, a crash mitigation, a flow/selector gotcha); AND (2) a STABLE EXTERNAL fact that would shorten or unblock a FUTURE sim test; AND (3) PARALLEL-SAFE: if the recipe relies on a shared host resource (a fixed localhost port, a single dev-server, the master sim, the maestro MCP daemon), propose the slot-safe variant (e.g. `updot` over a fixed-port debug dev-server) or a one-line WARNING instead. Do NOT propose as `[playbook]`: orchestration/watchdog/slot/revive/resource-release behavior, eval-tooling or rubric observations, one-off task specifics, or anything an existing rule already covers. Those belong in the report's Orchestration Issues or Skill Gaps sections, which the eval routes separately. +- **Do not deep link past the screen the task changed.** The committed flows reach their scene by deep link (`openLink`) by default; pass the flow's `MANUAL_PATH: "true"` so the drive walks through the changed screen. On iOS, RUN flow YAML with the native XCUITest interpreter, not `maestro test`: `~/.cursor/skills/build-and-test/scripts/xcuitest-run.sh --flow [--env K=V ...]` (slot UDID from `$AGENT_SIM_UDID`; no host port, so parallel slots never collide). It runs the same flow files (library `common/` flows included) and writes `takeScreenshot` output to the same paths; `capture-buy-quote.sh` uses it by default. Use the maestro CLI only when the task asks for Maestro, on Android, or for a flow the interpreter rejects. The interpreter preflights the whole flow tree and exits 2 before step 1 when a command or argument is unsupported, naming the command and the flow: rewrite that step with supported commands, or run that one flow on the maestro CLI and name the rejected command in the run report. There is no automatic fallback. Pass `--animations on` when the task is about an animation (the default turns them off). The run kills this slot's maestro MCP daemon and the session's MCP tools do not come back: finish MCP exploration before the first interpreter run. Supported commands, semantics and the other flags: `references/xcuitest-interpreter.md`; usage, output and exit codes: the script header. Scoped exception to `no-mutation`, test-infrastructure only. TESTIDS FIRST, COORDINATES LAST. When a flow needs to drive an element that has no stable selector (text match fails and no `testID` exists), ADD the missing `testID` prop to that component in the gui worktree and drive via it. A `testID` is a JS-only prop: Metro reload picks it up in seconds (no native rebuild), so adding one is cheaper than a single round of coordinate trial-and-error, and it de-brittles the suite for every future run. Coordinate taps are permitted ONLY for surfaces you cannot edit (system dialogs, native pickers, third-party views that don't forward `testID`) or when a reload would destroy unrecoverable in-flight app state, and any coordinate tap that survives into the PROOF flow must be called out in the run report with why a testID was not possible. - **Commit discipline:** commit the testID additions as a SEPARATE commit, distinct from any feature commit; change ONLY `testID` props, never component logic; update the flow selector(s) to use them. @@ -59,7 +60,7 @@ Return success exit only on PASS. ### 0f. Critical gotchas baked into the flow (do not "fix" them) Edge's RN keypad drops digits tapped too fast → wrong PIN → exponential lockout (465s → 914s → …). Each PIN digit tap in `buy-quote-input.yaml` uses `waitToSettleTimeoutMs`. Never speed it up. If a run logs "Invalid PIN: Account locked for N seconds", wait; do NOT tap. -On this debug build, `hideKeyboard` reliably triggers an RN Fabric text-measure SIGABRT. The flow leaves the keyboard up. Do not add `hideKeyboard` steps. +Driver-specific. On the Maestro driver, on this debug build, `hideKeyboard` reliably triggers an RN Fabric text-measure SIGABRT: do not add `hideKeyboard` to a flow that runs there (Android, `--driver maestro`), and leave the keyboard up. On the XCUITest interpreter `hideKeyboard` works and an iOS-only flow may use it. The committed flows run on both drivers, so they stay free of `hideKeyboard`. `assertVisible`/`extendedWaitUntil` traverse the a11y hierarchy on a poll loop, provoking the same Fabric crash on the Buy scene. The flow stops polling once the amount is entered; the capture script uses external simctl screenshots (no hierarchy traversal). diff --git a/.cursor/skills/build-and-test/references/sim-testing-playbook.md b/.cursor/skills/build-and-test/references/sim-testing-playbook.md index 6cd5854d..65fdba49 100644 --- a/.cursor/skills/build-and-test/references/sim-testing-playbook.md +++ b/.cursor/skills/build-and-test/references/sim-testing-playbook.md @@ -18,16 +18,16 @@ already encodes, params and gotchas included. |---|---|---| | `common/login-if-needed.yaml` | (YOLO env) | Land in a logged-in account, incl. PIN entry | | `common/dismiss-startup-modals.yaml` | - | Clear survey/notification/update modals | -| `common/select-swap-pair.yaml` | SRC_WALLET, DST_WALLET, FIAT_AMOUNT, PROVIDER | Exchange tab → pick wallets (via Search Wallets) → amount → quote (+ provider force, amount-field eraseText gotcha) | +| `common/select-swap-pair.yaml` | SRC_ASSET, DST_ASSET, MANUAL_PATH, SRC_WALLET, DST_WALLET, FIAT_AMOUNT, PROVIDER | Swap deep link sets the pair (MANUAL_PATH: Exchange tab → pick wallets via Search Wallets) → amount → quote (+ provider force, amount-field eraseText gotcha) | | `common/find-wallet.yaml` | SEARCH_TERM, MATCH_INDEX | Assets tab → search → open a wallet's detail scene (Receive/Send/Trade) | | `common/open-settings.yaml` | - | Side menu → Settings list (compose your own subpage nav after) | | `common/confirm-slider.yaml` | - | The confirm slider gesture (SOLVED — never re-derive) | | `common/ramp-set-region-fiat.yaml` | COUNTRY_ROW/SEARCH, STATE_ROW/SEARCH, FIAT_ROW/SEARCH | Set ramp region + fiat from Buy/Sell scene (row selectors are the COMBINED row string, e.g. "United States of America US") | -| `common/send-to-address.yaml` | WALLET_SEARCH, ADDRESS, AMOUNT | Assets → wallet → Send → address → amount → confirm slider | +| `common/send-to-address.yaml` | CURRENCY_CODE, ADDRESS, AMOUNT, WALLET_SEARCH, MANUAL_PATH | Payment-redirect link → pre-filled Send scene → confirm slider (no CURRENCY_CODE, or MANUAL_PATH: Assets → wallet → Send → address → amount) | | `common/create-throwaway-account.yaml` | NEW_USERNAME, NEW_PASSWORD, NEW_PIN, NEW_WALLETS, VERIFY_ACCOUNT_INFO | Login scene → new empty account, logged in (never `clearState`). For tests that would dirty a roster account's SYNCED state | | `common/delete-throwaway-account.yaml` | DELETE_USERNAME, DELETE_PASSWORD | Deletes the logged-in account. REQUIRED before the run ends for every throwaway the run created (throwaways are single-use) | -| `buy-quote-input.yaml` / `buy-quote.yaml` | (see file) | Canonical Buy $500 proof flow | -| `swap-quote-input.yaml` / `swap-confirm.yaml` | (see file) | Swap quote + confirm proof pair | +| `buy-quote-input.yaml` / `buy-quote.yaml` | BUY_ASSET, MANUAL_PATH (see file) | Canonical Buy $500 proof flow; opens Buy by deep link (MANUAL_PATH: Buy tab) | +| `swap-quote-input.yaml` / `swap-confirm.yaml` | (see file) | Maya swap quote + confirm pair. Local library only (`.syncignore`), not in the repo: compose `common/select-swap-pair.yaml` + `common/confirm-slider.yaml` instead | Wrote a NEW sequence a future task will plausibly need? Propose it for the library with a `[flow]`-tagged bullet in your run report's Dev Notes (name, @@ -460,12 +460,18 @@ debugging screenshots of the wrong device. worktree `env.json` (nulling only the username hits a light-account fallback that still auto-logs-in), then terminate + launch; restore after. (Promoted 2026-07-29, run 1215939017452141.) -- **Drive a deep link from cold start with `ENV.YOLO_DEEP_LINK` in the worktree - `env.json`** (read by `DeepLinkingManager` alongside `Linking.getInitialURL`), - then `simctl terminate` + `launch`. Deterministic, and avoids `simctl openurl`, - which pops an "Open in Edge?" system dialog that can background the app when - maestro taps it; maestro's own `openLink` also fails to deliver. (Promoted - 2026-08-06, run 1217224633446931.) +- **Drive a deep link with `openLink` on the XCUITest interpreter.** It hands + the URL to the running app directly (`edge://` and `https://edge.app/...` + alike) with no "Open in Edge?" dialog, from any scene, and the committed + flows use it by default. It returns before the app navigates: wait on the + target scene's text. Avoid `simctl openurl` and Maestro's own `openLink` on + iOS: the first raises the "Open in Edge?" system dialog, which can background + the app when tapped, and the second fails to deliver. `ENV.YOLO_DEEP_LINK` + in the worktree `env.json` (read by `DeepLinkingManager`, then + `simctl terminate` + `launch`) is only for a link that must arrive at cold + start. When the account holds several wallets for the linked asset the app + raises its own wallet picker; the flows handle it by wallet name. (Replaced + 2026-10-01, run 1219043297854606.) - **Nested `runFlow` with `env:` may NOT override a subflow's own `env:` defaults** (maestro 2.x, this host): `select-swap-pair` ran its built-in `.*Bitcoin.*` while the parent passed `.*Litecoin.*`, and `inputText` logged