Conversation
The Python bridge broadcasts approval.resolved for a gateway approval even when resolve_gateway_approval did not resolve anything, and the event carried no outcome. The chat-run forwarder then dropped the flag entirely, so a failed resolution reached the client indistinguishable from a successful one and closed the approval card. Carry the resolved outcome on the broadcast event and forward it when the runtime reports one.
Split the three-case Python broadcast test so each runtime condition fails independently, add the gateway-raises branch and an interrupt-path control, pin the non-boolean resolved values the typeof guard drops, and cover the live handleBridgeRun surface alongside the resume one.
VerificationAdversarial verification by a fresh run at head Holes found (3) — all now pinned
Added as declared controls (green on both arms, each with a distinct job): the four non-boolean Still open, stated in the body and not closed by any test: the client-store consumer ( 14 tests on the branch: 8 discriminating, 6 controls. Verification commit Base arm — both source files reverted to
|
sprayberry-redline
left a comment
There was a problem hiding this comment.
Automated review from the Sprayberry Labs fleet code reviewer.
Reviewed by the GPT gating lane (gating review).
Verdict: CHANGES_REQUESTED — not ready for the operator because this browser-visible approval outcome change lacks the upstream-required browser coverage and required full build validation.
Blocking finding — missing required browser-flow coverage and build validation
packages/server/src/modules/hermes/services/bridge/python/bridge_pool.py:2294
"resolved": resolved,
packages/server/src/modules/studio/services/chat-run/handle-bridge-run.ts:1535-1536
// Older runtimes omit resolved; forward only an explicit outcome.
...(typeof ev.resolved === 'boolean' ? { resolved: ev.resolved } : {}),
These changed lines deliberately alter the payload that makes the Web UI distinguish a failed approval resolution (resolved: false) from an older/absent outcome. The facts sheet confirms that the client-store consumer was only read, not executed, and that neither Playwright nor the full npm run build ran. That does not meet upstream DEVELOPMENT.md:62 ("Add Playwright coverage for browser-visible flows and routing/auth regressions") or DEVELOPMENT.md:64,123 (run the relevant tests plus npm run build; mark ready only after those pass). A server-only emitter assertion can pass while the actual card still dismisses, fails to show its stale/expiry state, or mishandles this field at the browser boundary.
Exercise an approval-resolved event with resolved: false through the client-visible path, assert that the pending card remains and exposes the stale/expired handling, and run the required build before re-submitting.
// Add a focused browser/client-flow regression that feeds an
// approval.resolved event with resolved: false and asserts the approval is
// retained and its stale/expiry handling is shown; then run `npm run build`.
What's good: the producer change publishes the already-computed boolean, and the typeof ev.resolved === 'boolean' guard correctly preserves a meaningful false while retaining legacy absent-key compatibility. The focused producer and both server forwarder surfaces are well covered. I also independently traced the base producer: it computes resolved but omits it from the event, so the reported defect is real. Fork CI reports no checks; that is absence rather than a passing signal.
I checked the candidate facts sheet, base behavior, changed producer/forwarder paths, boundary ledger, commit hygiene, policy excerpts, prior-art searches, and reported test/typecheck evidence. I did not run tests locally.
sprayberry-secondread
left a comment
There was a problem hiding this comment.
Automated review from the Sprayberry Labs fleet code reviewer.
Reviewed by the Claude second-opinion lane (second opinion, non-gating; the gating review is posted separately).
Verdict: no blocking issues found in the outcome-propagation fix at 9daa0b3bd3ce41f405e385a6794d033cbda601ed; ready within its deliberately narrow scope.
Independent correctness read
I can confirm the bug from the base code represented by the diff and surrounding source: a registered bridge approval whose gateway request has expired produces resolved = False, but the old event dictionary omits that value and the old TS forwarder independently drops it. With a matching pending card, clearPendingApproval therefore reaches deletion rather than its strict resolved === false branch.
The changed lines address both losses:
packages/server/src/modules/hermes/services/bridge/python/bridge_pool.py:2294:"resolved": resolved,packages/server/src/modules/studio/services/chat-run/handle-bridge-run.ts:1536:...(typeof ev.resolved === 'boolean' ? { resolved: ev.resolved } : {}),
The type guard is appropriate: forwarding false is essential, while defaulting an absent outcome would invent information. The other Python emit sites are semantically different: the terminal callback has finished waiting, and interruption terminates the pending interaction. Preserving their existing omission is defensible, not evidence that this narrow fix must become a lifecycle rewrite. The new interruption control exercises the interruption site, not the separate terminal-callback site.
Boundaries — rebuilt from the diff
There is one new production predicate, typeof ev.resolved === 'boolean'; the Python addition publishes an already-computed value without changing its computation.
| Changed expression / boundary | Fixed behavior | Test pin |
|---|---|---|
TS type predicate: false / true |
Preserve the exact boolean | Both variants in resume tests and live-run final-context tests |
Same predicate: absent / undefined |
Omit the key | Older-runtime resume control tests absent key; explicit undefined has the same type branch |
Same predicate: null, 0, "", "true" |
Omit the key, regardless of truthiness | Four non-boolean resume controls |
| Same predicate: empty array/object, negative number, 1, maximum finite number | Omit the key; no numeric threshold or indexing is involved | No individual fixtures; same non-boolean branch as the controls, not a distinct numeric limit |
| Python published value: nonempty request, one waiter resolved | Publish true | Successful gateway test |
| Python published value: zero waiters, empty request ID, gateway exception | Publish false | Three separate negative tests; empty-ID test also checks no resolver call |
| Other call surface | Same forwarding via live handleBridgeRun and reconnect resumeBridgeRun |
Both exports have true/false tests; tests drive stream and polling respectively, not every call site |
| Other producer path: interruption | Existing reason-bearing event remains outcome-free | Interruption control |
Test-only guards and indexing also have observable checks: the prelude's if queue_request drives the queued and unregistered cases; its session match selects the newly installed request. Python event-list filters and TS event-name filters isolate the emitted event, then assert exactly one match before indexing [0]. Each regression variant asserts the actual boolean on that event. The choice/return-value assertions are supporting invariants, not claimed discriminators. The declared controls legitimately pass on base while asserting that an event was emitted; they are not substitutes for the eight boolean-propagation regression cases.
Maintainer's-eye notes
- Idioms and scope: retaining older-runtime compatibility fits the touched module's recent #3018, which explicitly preserves old approval imports with a fallback. No new abstraction is needed for copying an already-computed field.
- Test shape: recent outside contributions #3033 and #3112 pair focused source changes with server regression coverage; EKKOLearnAI#3033 extends this same final-context test file. #3131 also records a failing-before regression plus retained compatibility controls. This candidate follows that pattern.
- Prior art: independently searched PRs for
approval resolvedandgateway approval, and issues forapproval resolved. #2681 addresses exact waiter identity, not publishing the outcome; EKKOLearnAI#2135 addresses persistence. Issues EKKOLearnAI#2558 and EKKOLearnAI#2894 describe adjacent failure symptoms/conditions. I did not find a duplicate outcome-propagation fix in these results; this is not an exhaustive historical search. - Operator validation notes:
DEVELOPMENT.mdasks for focused tests plusnpm run build, and Playwright coverage for browser-visible flows. The candidate reports server typechecking, not the full build. The client consumer was inspected but not executed: a bareresolved: falseretains the matching card; it does not itself generate an expiry notification withoutstaleor an expiry reason. A socket-mocked client regression would strengthen the user-facing claim, but the changed producer/forwarder contract already has direct tests. Do not describe this patch as fixing all expiry UX or runtime approval failures. - Submission shape: the proposed conventional
fix:title fits EKKOLearnAI#3018/EKKOLearnAI#3131. Keep the upstream description concise about the bug and validation, as those merged PRs do. I found no reason here to demand a release-version change; EKKOLearnAI#3131 includes topic documentation, while EKKOLearnAI#3033 has no changelog file.
What's good: this reuses the existing computed outcome, preserves false rather than using truthiness, and tests both public forwarder entry points without unrelated refactoring.
Scope of review: full five-file diff, Python producer context, client consumer, call-site enumeration, upstream history and the PRs cited above. No test suite was run. gh pr checks reports no checks; the submitted A/B transcript is evidence supplied by the author, not an independently executed result from this review.
SECOND READ: READY
Drive the bridge forwarder over a runtime approval.resolved event and feed what it emits to the chat store, so a failed gateway resolution is shown to keep its approval card and reach the expiry notice. Add browser coverage for what the card does with each outcome the forwarder can send.
The chat view dismisses an approval card as soon as the user answers, so a failed gateway resolution arrives with nothing pending for the session and clearPendingApproval takes its no-pending early return. That branch notifies the user only when the event reports resolved false, which is what the forwarder now sends; add the failing case and a successful-resolution control for the same branch.
VerificationAdversarial re-verification by a fresh run at head The body's claims reproduce exactlyBefore touching anything, I re-ran both arms of the A/B as the body describes them, at head $ npx vitest run tests/client/bridge-approval-outcome-contract.test.ts tests/server/run-chat-bridge-resume.test.ts tests/server/run-chat-bridge-final-context.test.ts tests/server/agent-bridge-python-concurrency.test.ts
Test Files 4 passed (4)
Tests 99 passed (99)
Duration 10.85s
$ git checkout 0cc8271e -- bridge_pool.py handle-bridge-run.ts # tests kept
$ npx vitest run <the same four files>
Test Files 4 failed (4)
Tests 11 failed | 88 passed (99)
Duration 11.73sSame 11 named discriminators, four of them aborting with Hole found (1) — Boundaries row 24, closed by argument rather than by a testThe previous head's ledger listed row 24 —
That argument does not hold, in the same shape as the It is also not an exotic path. It is what the main chat view does on every approval:
On base that branch is dead, so the most common path through the UI is also the one that fails most silently. Tests added (2), on the same branch
Both drive the real forwarder through Both arms at the new head
|
sprayberry-redline
left a comment
There was a problem hiding this comment.
Automated review from the Sprayberry Labs fleet code reviewer.
Reviewed by the GPT gating lane (gating review).
Verdict: APPROVE — no blocking issues found; ready for the operator to submit.
I reviewed the three-line production diff, its Python producer and TypeScript forwarder context, both exported call surfaces (handleBridgeRun and resumeBridgeRun), and the client consumer’s pending and submit-dismissed branches. bridge_pool.py:2294 now publishes the already-computed boolean outcome, and handle-bridge-run.ts:1535-1536 preserves both true and the bug-critical false while retaining the legacy omitted-key payload for older runtimes. The focused tests cover producer outcomes, both forwarder surfaces, the store branches, and browser controls; the PR body supplies current head/base A/B output (12 discriminators fail on base and pass on head) plus a historical successful full build. Fork CI currently reports no checks, which the facts sheet accurately records.
What's good: this remains a minimal, single-bug change. The typeof ev.resolved === 'boolean' guard is the correct compatibility boundary: a truthiness guard would drop false and recreate the reported UI failure. The rework also closes the previously missing client path where submitting locally dismisses the card before a failed resolution reaches clearPendingApproval.
Notes for the operator:
- Fork Actions are not enabled, so GitHub reports no PR checks. Enable Actions and run the relevant workflow(s) before upstream submission if available.
- The upstream PR body should follow any upstream template requirements; this fork facts sheet already records the required technical evidence.
What I did not run locally: the repository test suite, build, and Playwright suite. Per the reviewer environment policy, I used the PR’s recorded executed A/B evidence and CI status rather than rerunning local suites.
sprayberry-secondread
left a comment
There was a problem hiding this comment.
Automated review from the Sprayberry Labs fleet code reviewer.
Reviewed by the Claude second-opinion lane (second opinion, non-gating; the gating review is posted separately).
The production fix is sound on static inspection, but the test additions need submission-readiness cleanup.
Findings
Low — patch-history narration and review labels in shipped tests.
Exact examples from the added lines:
tests/client/bridge-approval-outcome-contract.test.ts:284-285:// The card is dismissed with or without the fix here; the forwarded true is/// what discriminates, so it is asserted first.- Same file,
:308-309:// its \!current` early return — a branch that only notifies when the event/// says resolved === false, which is precisely what the forwarder now sends.` This also introduces an em dash. - Same file,
:290and:342: test names end in(control);:333-334narrates the unfixed forwarder, and:340-341again sayswith or without the fix. tests/server/agent-bridge-python-concurrency.test.ts:1456andtests/server/run-chat-bridge-resume.test.ts:464: test names end in(control). The Python test's:1461-1463comment explains what isuntouched by the fixand holdson either arm.tests/e2e/chat-streaming.spec.ts:1607-1611explains the tests in terms of the forwarder fix and controls;:1612and:1690put(control)in the test names.
The new 368-line client test file also dwarfs the three-line production change. The cross-layer coverage has value, but comments arguing how the patch is verified belong in the evidence packet, not the permanent regression suite. These are concrete submission-readiness tells, not a claim that the boolean forwarding is incorrect. Keep the coverage while removing patch narration and investigating reuse of existing setup.
Suggested fix:
Rename controls after their contracts, e.g.:
'ignores non-boolean approval outcomes'
'dismisses approvals without an outcome'
'does not notify after successful approval'
Remove before/after and control-accounting commentary from test files.
Keep only comments explaining otherwise non-obvious runtime setup.
Move discrimination accounting to the fork PR evidence; reuse existing
fixtures where feasible rather than deleting boundary assertions.
Boundaries rebuilt from the change
| Changed element / input | Fixed behavior | Coverage read |
|---|---|---|
| Python event publishes the already computed boolean: queued waiter / zero waiters | Publishes true / false | Separate queued and never-registered Python cases |
| Empty request ID / gateway exception | Publishes false | Separate empty-ID and raising-gateway cases |
| Other approval path | No new field at the separate interrupt emitter | Interrupt control |
typeof ev.resolved === 'boolean': true / false |
Preserves both, including falsy false | Resume and live-run parameterized cases; client contract cases |
| Missing/undefined, null, zero, empty string, truthy string | Omits field | Legacy and four non-boolean resume cases |
| Negative / maximum number, empty array/object | Also omitted by the same typeof guard; no numeric threshold exists | Not individually tested; same non-boolean equivalence class |
| Client pending / already dismissed | False preserves a non-stale pending card; a stale attempted failure notifies in either branch | Pending failure, stale failure, submit-dismissed failure cases |
| Client true / absent field | Existing dismissal or silent no-pending behavior | Success, legacy, and submit-dismissed success cases |
Test-helper filters and index lookups select one emitted approval.resolved event, with length assertions before accessing its payload. The Python helper's queue_request guard is exercised both ways. Server true/false variants assert the forwarded field, so neither variant silently becomes a base-passing success test. Browser tests inject the outcome directly and are compatibility checks, not proof that the producer/forwarder fix works.
The ticket's two questions: control 22 is distinct from control 18: it checks that a successful outcome does not notify after a locally dismissed submission even with stale metadata; control 18 checks legacy omission with a pending card. Retain both behaviors, without the (control) names. Historical browser/build transcripts labelled at e70d31f3 are honest evidence, not current-head CI. A test-only successor does not itself demonstrate a production regression; I would not demand a browser rerun solely for that reason, but the operator must not present those transcripts as runs at 534cf3c9.
Maintainer context and what's good
I confirmed the bug from base code: respond_approval computes and returns resolved but omits it from the emitted event, while the consumer explicitly distinguishes resolved === false. Publishing the existing result and preserving only booleans is appropriately scoped. It does not repair the wider approval lifecycle. In particular, the client expiry tests add stale/error after forwarding; they do not prove that the bridge alone emits expiry metadata.
Recent module history includes #3132, which uses a focused fix: title and dedicated regression coverage. The merged outside changes #3133 and #3131 likewise preserve explicit settings boundaries without unrelated refactors. The proposed upstream title fits that idiom. These examples do not establish a mandatory changelog requirement. Independent approval searches surfaced #2681, #2648, and #2135; I did not establish from search results alone that they duplicate this narrow field-propagation fix.
Scope: read the two source hunks, the server/browser additions and new client contract tests, relevant base producer/consumer code, commit messages, and upstream history/merged-PR summaries. No test suite was run. GitHub reports no checks for this branch; reported local results remain author evidence. I did not review unrelated repository code or other review bodies.
SECOND READ: NOT READY — remove patch-history narration, (control) test labels, and the added em dash from the shipped tests.
The added cases carried patch-history narration: test names ending in (control), comments saying what holds with or without the change, and accounting of which assertion discriminates. Name each case after the behaviour it pins, fold the client and browser cases that differ only in outcome into tables, and keep only the comments that explain runtime setup. No assertion input is dropped.
VerificationAdversarial verification by a fresh run at head Two prior verification segments on this candidate were cut off at the iteration cap: the first pushed What I ran, myself, at this headFresh worktree ( $ npx vitest run tests/client/bridge-approval-outcome-contract.test.ts tests/server/run-chat-bridge-resume.test.ts tests/server/agent-bridge-python-concurrency.test.ts tests/server/run-chat-bridge-final-context.test.ts
Test Files 4 passed (4)
Tests 105 passed (105)Base arm ( Test Files 4 failed (4)
Tests 17 failed | 88 passed (105)The 17 base failures are exactly the 17 discriminating rows the body claims: the 13 from the previously-verified head plus the 4 new ones added at
Rebuilt boundary ledger and the new call surfaces
Mutant of the forwarder's guard (the single most important line in the diff): reverted Test Files 3 failed (3)
Tests 8 failed | 49 passed (57)killed by (among others) both new tests ( Two consumers outside the diff's blast radius, probed not testedThe body lists two further consumers reached by the same payload shape but not routed through by this diff (group-chat store, voice relay), measured on both arms rather than tested since their existing code already pins the $ npx vitest run tests/client/zz-probe-group.test.ts # copied in, removed after
PROBE group-chat store no resolved key: pending before=1 after=0
PROBE group-chat store resolved false: pending before=1 after=1
PROBE group-chat store resolved true: pending before=1 after=0
Tests 3 passed (3)
$ npx vitest run tests/server/zz-probe-relay.test.ts # copied in, removed after
PROBE voice relay no resolved key: [{"type":"tool.completed","interactionId":"voice-1","tool":"approval"}]
PROBE voice relay resolved false: [{"type":"tool.completed","interactionId":"voice-1","tool":"approval","error":"approval failed"}]
PROBE voice relay resolved true: [{"type":"tool.completed","interactionId":"voice-1","tool":"approval"}]
Tests 3 passed (3)Both byte-identical to the salvaged transcripts. Typecheck and fork CI$ npx tsc --noEmit -p packages/server/tsconfig.json
rc=0
Both green, no non-green job. The build workflow ran Prior art re-checked at the gate
BodyThe PR body was stale (still described head ResultEverything holds: 105/105 at head, 88/105 (17 failing) at base, the truthiness mutant killed including by both new tests, both extra-consumer probes reproduced, tsc clean, fork CI green on both workflows, no source change since the fix commit, no superseding upstream fix. Rules: reads-as-generated=covered(body/test-count reconciled to head, 26 tests against a 3-line/2-file diff across 5 files, each row traces to a specific test) | no-control-cases-in-the-suite=covered(all 9 declared controls re-verified passing on base per their own base-arm run, each with a distinct stated job) | mutate-the-rejected-alternatives=covered(truthiness-guard mutant, the alternative the Fix section's longest rejection paragraph names, killed by 8 tests including both new ones) | ledger-row-needs-its-fixture=covered(rows 16b/16c each own a dedicated fixture and test, not folded behind a nearby one) | prior-art-recheck-at-gate=covered(origin/main re-fetched to a485b57, -S search over the two changed files since base, both cited issues re-checked OPEN) | run-every-ci-step-not-just-the-red-one=covered(gh pr checks read in full: build and e2e both green, no step skipped) | comment-cites-its-own-review=covered(diff scanned for review/PR/issue numbers and gate words in added test lines, none found) | idempotence-test-asserts-only-agreement=unreachable(no round-trip or repeated-call test in this diff) | side-effect-change-needs-its-test=unreachable(no change to what a value is computed from, only to whether it is published) | static-row-vs-alias-stub=unreachable(no row declared static; group-chat/voice-relay rows are measured probes, not stubbed) |
sprayberry-redline
left a comment
There was a problem hiding this comment.
Automated review from the Sprayberry Labs fleet code reviewer.
Reviewed by the gating lane (gating review).
Verdict: APPROVE at head ca908be29b6ff7716fc9ccde837e0fb9f6340a35. No blocking findings; ready for the operator to submit.
What changed since my last verdict (534cf3c9, approved)
git diff 534cf3c9...ca908be2 touches five test files and no source file; git diff 027380b3 ca908be2 -- packages/ is empty, so the three-line production fix I already traced (bridge_pool.py:2294 "resolved": resolved, and handle-bridge-run.ts:1535-1536 ...(typeof ev.resolved === 'boolean' ? { resolved: ev.resolved } : {})) is byte-identical to what I approved. I re-read the full head diff rather than only the delta.
94a1c92f(tests only): removes the(control)suffixes, the "with or without the fix" / "on the unfixed forwarder" narration and the "asserted first" accounting from the client, e2e and Python tests, and folds same-shape cases intoit.eachtables. Every assertion input from the previous head is still present; I compared the removed blocks against the new table rows intests/client/bridge-approval-outcome-contract.test.ts:224-292.ca908be2(tests only): adds four cases for the two forwarder outputs no previous test read.tests/server/run-chat-bridge-resume.test.ts:790-864feeds theonEventobserver and assertsforegroundNotification('approval.resolved', observed[0][1], 10) !== nullequals the expected outcome;tests/server/run-chat-bridge-final-context.test.ts:646-701asserts thereplaceStatecall carriesresolved.
What I verified
- Live-head gate. Head
ca908be2differs from my standing verdict commit534cf3c9; this is a fresh review of the live head. - Facts sheet. All ten required sections present and reconciled to
ca908be2(the## Upstreamblock lists all six commits;## Verification methodquotes the fork CI run ids for both workflows at this head). - Bug real on upstream
maintoday.handle-bridge-run.ts:1557-1563onEKKOLearnAI/ekko-studio@mainstill builds theapproval.resolvedpayload fromrun_id, approval_id, choiceonly, so the fix has not landed upstream since the base sha. - New tests discriminate. Traced
foregroundNotificationat this head (foreground-notification.ts:5:if (event.endsWith('.resolved') && (value.resolved === false || value.stale === true)) return null). On base the observer payload has noresolvedkey, so the failed-resolution variant returns a notification object and the.toBe(false)assertion fails; on head it returnsnull. The successful variant fails on base throughtoMatchObject({ resolved: true })against a payload with no key. Same shape for thereplaceStatepair. The verification comment's base arm (17 failed | 88 passed, 13 prior + these 4) matches that trace. - Boundaries. The one new predicate,
typeof ev.resolved === 'boolean', is pinned fortrue,false, key absent,null,0,''and'true'; both exported entry points (handleBridgeRun,resumeBridgeRun) and all three consumers of the emitted payload (socket broadcast,onEvent,replaceState) now have a test. The Python side pins queued, never-registered, emptyrequest_id, raising gateway, and the interrupt path. - CI.
gh pr checks 1:build pass 5m18s,e2e pass 11m34sat this head. - Prior art, re-run.
gh pr list --repo EKKOLearnAI/ekko-studio --state open --search "approval"returns EKKOLearnAI#2135, EKKOLearnAI#2455, EKKOLearnAI#3087, EKKOLearnAI#2339, EKKOLearnAI#2233; none touches theapproval.resolvedpayload. No open PR for this bug. - Policy and hygiene. No CONTRIBUTING, CLA, DCO or AI policy in the upstream (404s confirmed in the body; AGENTS.md is an agent map). Commit messages and title carry no model name, no
Co-Authored-By, no em dash; commit style matches the upstream's lowercasefix:/test:convention. Branchfix/bridge-approval-resolved-flagfollowsDEVELOPMENT.md'scodex/fix-login-tokenshape. - Tell pass over the diff, test names, comments, commit messages and title: no em dashes, no
(control), no "with or without the fix", no "previously", no "ensure"/"gracefully"/"robust", no sleeps, no restating comments.
Notes for the operator (non-blocking)
- The upstream repository has been renamed.
gh api repos/EKKOLearnAI/hermes-studionow resolves toEKKOLearnAI/ekko-studio(GitHub redirects the old name). Open the upstream PR againstEKKOLearnAI/ekko-studioand update the## Upstreamline in your body; the fork's base sha and the bug are unaffected. - Upstream issue EKKOLearnAI#2894 ("审批点击确认后立即超时 ... bool(gateway_request_id) 条件导致 resolve_gateway_approval 未被调用", open) describes the EKKOLearnAI#2756 condition from the user's side and is not cited in the body. Worth adding alongside EKKOLearnAI#2558 and EKKOLearnAI#2756 in the upstream description.
- Test volume: 26 tests and roughly 920 added lines for a 3-line fix, across five files. Each case pins a distinct boundary or consumer and the narration is gone, but a maintainer skimming the file list may ask to trim; be ready to say which consumer each file covers (server forwarder x2 surfaces, Python broadcast, client store, browser card).
tests/client/bridge-approval-outcome-contract.test.ts:2-4opens with a three-line file comment explaining why the server forwarder is driven from a client test. It reads as setup rationale, not patch history, and I did not block on it; if the maintainer prefers bare files, it can go.
What's good
The production change is still the minimal honest fix: publish the outcome the Python side already computed, forward it only when the runtime reports a boolean so older runtimes keep their current shape. The test rework answered both prior review rounds without dropping an input, and the last commit closed the two outputs (onEvent, replayed state) that every earlier test had missed.
|
Submitted upstream for review. |
Summary
AgentPool.respond_approval()computes whether a gateway approval actually resolved, but theapproval.resolvedevent it broadcasts to the session omits that outcome, so a failed resolution is broadcast in exactly the same shape as a successful one.handle-bridge-run.tsforwards only{event, run_id, approval_id, choice}forapproval.resolved, dropping any outcome the runtime does report.resolved === false(packages/client/src/stores/hermes/chat.ts:3282), so the popup closes and the user sees a success for a command that was never approved.resolvedon the broadcast event and forwards it when the runtime reports a boolean; runtimes that do not report one keep today's payload shape.ca908be2): 17 discriminating, 9 controls. Both call surfaces of the forwarder (live run and resume), the forwarder's third output (theonEventobserver that feeds webhooks and push notices) and its fourth (the replayed run state a reconnecting client receives), all four Python outcome branches, both client-store consumer branches (card pending and card already dismissed), and the browser card are covered.Passes after (fix applied, head
ca908be2):(105 vitest tests in those four files; 24 of the 26 approval tests live here, the other 2 are the Playwright pair in
tests/e2e/chat-streaming.spec.ts.)Fails before (both changed source files reverted to base
0cc8271e, all tests kept), headca908be2:The four Python failures each abort with
KeyError: 'resolved': on base the broadcast event has no such key. The nine controls are among the 88 passing on base.The user-visible failure, on base, at the client store. This is the assertion that the approval card is gone when the resolution failed:
Upstream
EKKOLearnAI/hermes-studiomain0cc8271ed99bba1936b63a85b45fb35448892217ca908be29b6ff7716fc9ccde837e0fb9f6340a35(fix027380b3+ boundary tests9daa0b3b+ client/browser coveragee70d31f3+ submit-dismissed path534cf3c9+ test naming cleanup94a1c92f+ observer and replayed-state coverageca908be2; none of the five test commits touches a source file,git diff 027380b3 ca908be2 -- packages/is empty)packages/server/src/modules/hermes/services/bridge/python/bridge_pool.py,AgentPool.respond_approval()(gateway branch, around line 2289)packages/server/src/modules/studio/services/chat-run/handle-bridge-run.ts,applyBridgeChunkAsync(), theapproval.resolvedbranch (around line 1529)Bug
When the Web UI answers a gateway (dangerous-command) approval,
AgentPool.respond_approval()callsresolve_gateway_approval()and computesresolved = bool(gateway_request_id) and resolve_gateway_approval(...) > 0. That value is returned to the socket caller, but theapproval.resolvedevent appended to the run's event stream carries onlyevent,run_id,approval_idandchoice.applyBridgeChunkAsync()then rebuilds the outbound payload from exactly those four keys. The client'sclearPendingApproval()dismisses the pending approval unless the event saysresolved === false, so a resolution that failed (the runtime emitted norequest_id, hermes-agent older than v0.20.5, issue EKKOLearnAI#2756; the queue entry had already expired, issue EKKOLearnAI#2558; orresolve_gateway_approvalraised) closes the approval card exactly as a successful one does. The user sees their Allow/Deny accepted while the command is never executed and the agent thread keeps waiting for its own timeout. Blast radius: every Web UI user running the Hermes CLI bridge with gateway approvals whose runtime is version-skewed or whose approval has aged out. The socket path (sockets/chat-run.ts:1199-1211) already sends an honestresolvedflag, so only the bridge run-event broadcast is affected. Both of that forwarder's call surfaces are hit:handleBridgeRun(the live run, where the user clicks Allow mid-run) andresumeBridgeRun(reconnect, and any client that was not the responder).The main chat view makes the drop worse rather than milder.
MessageList.vue:368answers throughrespondApproval(), which dismisses the card locally the moment the choice is sent (chat.ts:3401). The failed resolution therefore arrives with nothing pending for the session, andclearPendingApprovaltakes its no-pending early return, a branch that notifies the user only when the event reportsresolved === false. On base that branch is dead, so the most common path through the UI is also the one that fails most silently.Repro
Feed
applyBridgeChunkAsyncanapproval.resolvedbridge event carryingresolved: false, and observe that the payload emitted to clients has noresolvedkey at all, and that the approval card in the store is consequently dismissed, or that the expiry notice never fires when the card was already dismissed on submit. This is pinned end to end at three layers (the Python producer, the TypeScript forwarder, and the client store consumer) by the 17 discriminating tests listed under## Test evidence. The verbatim base-arm transcripts in## Summaryare current for headca908be2.At the Python layer the base failure is a bare missing key:
At the client layer the base failure is the reported symptom itself: the card is already gone. That transcript is the third console block in
## Summary.Fix
Two lines of behaviour plus one comment.
resolvedis already computed on the line above the event; publishing it is the minimal change that makes the broadcast agree with the valuerespond_approvalreturns to its caller, and it matches the shape the socket path (sockets/chat-run.ts) and the ekko-agent path (handle-ekko-agent-run.ts:1272) already emit.Alternatives rejected:
resolvedtotruein the TS forwarder when absent. Would assert an outcome for runtimes that never reported one, turning the silent-success bug into a silent-success guarantee. The conditional spread keeps the payload byte-identical for those runtimes, so no client behaviour changes where there is no new information....(ev.resolved ? ...)) instead oftypeof. This is the single most important line in the diff. A truthiness guard dropsfalse, the falsy-but-valid case, and reinstates the exact bug being fixed. Pinned by the two'a failed gateway resolution'tests and by client tests 15 and 19, which fail on base and would fail again under a truthiness mutant. Test 7 is the truthy non-boolean"true", the one value such a guard would wrongly forward.approval.resolvedbroadcast entirely whenresolvedis false. The client's expired/stale handling (chat.ts:3276-3288) is driven by receiving the event withresolved === false; withholding it would leave the card up with no explanation and no path to the expiry notice. Client tests 18 and 19 execute both of that handling's branches.stale/errorfields as the socket path does. That is a larger contract change;resolvedalone is whatclearPendingApprovalgates on, and the rest belongs to the wider lifecycle work proposed in [Bug]: EKKOLearnAI/ekko-studio#3099.Test evidence
26 tests at head
ca908be2: 17 discriminating (fail on base) and 9 controls (green on both arms, each with a distinct job). Tests 1 and 2 and 8 shipped with the fix commit027380b3; tests 3 to 7 and 9 to 14 were added by adversarial verification at9daa0b3b; tests 15 to 17 and 21 to 22 were added ate70d31f3to answer the gating review's client-boundary finding; tests 18 to 20 were added at534cf3c9by a second adversarial verification that found the submit-dismissed consumer branch reachable and untested.94a1c92frenamed cases and folded same-shape cases into tables; it changed no assertion input. Tests 23 to 26 were added atca908be2by a third adversarial verification: the forwarder'semithas two consumers the earlier tests never read, theonEventobserver (webhooks, push, foreground notices) andreplaceState(the run state replayed to a reconnecting client), and each now has a row per outcome.forwards the bridge approval outcome for 'a failed gateway resolution'run-chat-bridge-resume.test.tsforwards the bridge approval outcome for 'a successful gateway resolution'run-chat-bridge-resume.test.tsomits resolved when an older bridge runtime does not report an approval outcomerun-chat-bridge-resume.test.tsdrops a non-boolean resolved value of 'null' rather than forwarding itrun-chat-bridge-resume.test.tstypeof null === 'object', distinct from the key-absent test 3drops a non-boolean resolved value of 'the number 0' rather than forwarding itrun-chat-bridge-resume.test.tsdrops a non-boolean resolved value of 'an empty string' rather than forwarding itrun-chat-bridge-resume.test.tsdrops a non-boolean resolved value of 'the string "true"' rather than forwarding itrun-chat-bridge-resume.test.tsreports a resolved gateway approval outcome on the broadcast approval.resolved eventagent-bridge-python-concurrency.test.tsKeyError: 'resolved')reports an unresolved gateway approval outcome when the runtime never registered the requestagent-bridge-python-concurrency.test.tsreports an unresolved gateway approval outcome for a runtime that sends no request idagent-bridge-python-concurrency.test.tsreports an unresolved gateway approval outcome when the approval gateway raisesagent-bridge-python-concurrency.test.tsexcept Exceptionbranchomits the outcome on the session-interrupt approval.resolved eventagent-bridge-python-concurrency.test.tsapproval.resolvedemit sites inbridge_pool.py(lines 1426 and 2156) were not disturbedforwards the bridge approval outcome on a live run for 'a failed gateway resolution'run-chat-bridge-final-context.test.tsforwards the bridge approval outcome on a live run for 'a successful gateway resolution'run-chat-bridge-final-context.test.tsforwards 'a failed gateway resolution' to the approval cardbridge-approval-outcome-contract.test.tsexpected undefined to be 'approval-bridge', i.e. the card is goneforwards 'a successful gateway resolution' to the approval cardbridge-approval-outcome-contract.test.tsexpected undefined to be true)forwards 'an older runtime that reports no outcome' to the approval cardbridge-approval-outcome-contract.test.tsreports the expiry when a failed gateway resolution is also stalebridge-approval-outcome-contract.test.tsexpected "spy" to be called 1 times, but got 0 timesreports 'a failed gateway resolution' to a card the view already dismissed on submitbridge-approval-outcome-contract.test.ts!currentearly-return branchexpected "spy" to be called 1 times, but got 0 timesreports 'a successful gateway resolution' to a card the view already dismissed on submitbridge-approval-outcome-contract.test.tsresolved: truetrue) fails on basekeeps the approval card when the runtime reports a failed resolutiontests/e2e/chat-streaming.spec.tsdismisses the approval card when the runtime reports a successful resolutiontests/e2e/chat-streaming.spec.tsreports the bridge approval outcome to the run event observer for 'a failed gateway resolution'run-chat-bridge-resume.test.tsonEventobserver (the hooksockets/chat-run.tsuses for webhooks, push and pending-interaction notices)foregroundNotification()returnsnullfor the failed outcome, the checkapp-events.ts:60makes before notifying a phonereports the bridge approval outcome to the run event observer for 'a successful gateway resolution'run-chat-bridge-resume.test.tsonEventobserverforegroundNotification()produces a notice for the successful outcomekeeps the bridge approval outcome in the replayed run state for 'a failed gateway resolution'run-chat-bridge-final-context.test.tsreplaceState(theapproval.resolvedentrybuildResumeEventsreplays on reconnect andbindAppEventSubscriptionreads atsockets/chat-run.ts:748)keeps the bridge approval outcome in the replayed run state for 'a successful gateway resolution'run-chat-bridge-final-context.test.tsreplaceStateTests 23 to 26 read the two outputs of the forwarder that tests 1 to 20 do not. The forwarder's
emitclosure (handle-bridge-run.ts:629and:1012) does three things with one payload object:replaceStateinto the session's event list,onEventto the socket layer's observer, and the socket broadcast. Tests 1, 2, 13 and 14 read only the broadcast. The observer is whatsockets/chat-run.ts:1622-1638wraps to feedobserveChatRunWebhookEventandemitPendingInteraction, andforegroundNotification()(foreground-notification.ts:5) returnsnullforresolved === false, so a phone gets a "resolved" notice on base for an approval that failed. The replayed state is what a reconnecting client receives and what the app-event subscription (sockets/chat-run.ts:748) uses to decide whether an approval is still pending: on base the entry has noresolvedkey and the pending card is dropped from the replay.Tests 15 to 20 answer the gating review's blocking finding. They do not use a hand-written payload fixture: each one drives the real forwarder (
resumeBridgeRun) over a runtimeapproval.resolvedevent and feeds whatever it emits into the real chat store via the global peer handler, so the server payload and the client card are checked against each other.Tests 19 and 20 cover the branch the main chat view actually takes. An earlier head listed that branch as reachable with no test and argued it was "unchanged by the diff". The argument does not hold: the diff publishes the very value the branch gates on. Test 19 drives
respondApproval()exactly asMessageList.vue:368does and then forwards a failed resolution.Per the multi-assert rule, the user-visible assertion is ordered first in each body, so it is the one proven on the base arm (a base run aborts at the first failing assert, which would otherwise leave later assertions unproven).
Tests 21 and 22 are declared controls, not regressions.
DEVELOPMENT.md:62asks for Playwright coverage of browser-visible flows, and these run in a real Chromium against the real Vue app. But the repo'smockChatSocketfixture injects event payloads at the browser boundary, downstream of the changed server code, so no e2e test in this harness can discriminate on the fix. Measured, not assumed: both pass on base and on head. They pin what the card must do with each outcome the forwarder can send; tests 15 to 20 pin that the forwarder actually sends it.Playwright: measured locally at head
e70d31f3, and by the fork's own CI atca908be2. The only change totests/e2e/chat-streaming.spec.tssincee70d31f3is the two test names and three comments (git diff e70d31f3 ca908be2 -- tests/e2etouches no assertion, no locator and no payload), and no commit since027380b3touches a source file. The fork's Playwright workflow atca908be2(run 35940638064, job 107447538023) ran the wholetests/e2esuite in Chromium:Running 212 tests using 1 worker...212 passed (10.6m), which includes both approval-card tests, in the dot reporter CI uses. The local transcript below is frome70d31f3, under the names the tests carried then, because the containers that ran the later rounds have no Chromium binary:Base arm at
e70d31f3(both source files reverted to0cc8271e), both still pass, which is why they are controls:pw.local.config.tsis an untracked local override that points Playwright at the container's system Chromium (the managed download is a glibc build and that container was musl). It is not part of the diff; upstream CI uses the repo's ownplaywright.config.tsunchanged.Typecheck at this head (
ca908be2):Full build, required by
DEVELOPMENT.md:64andDEVELOPMENT.md:123before marking a PR ready. Local transcript at94a1c92f(no source or build input changed since;ca908be2is two test files); the fork's Build workflow re-ran the samenpm run buildatca908be2(run 35940638067, job 107447538421) and reached[build-server] ESP32-C3 v2 firmware copied from release artifactafternpm run test:coveragereportedTest Files 693 passed | 7 skipped (700):All five
&&-chained stages completed:openapi:generate,vue-tsc -b(full client typecheck),vite build,tsc --noEmit -p packages/server/tsconfig.json, andscripts/build-server.mjs. The chain reachingbuild-serveris the proof that both typecheck legs exited 0.Verification method
executed, in the run container: Node v24.19.0, vitest 3.2.4, python3 3.14 (the Python tests run through the repo's existingexecFileSync('python3', ...)harness). Install:npm ci --ignore-scripts --no-audit --no-fund. Both arms of the A/B were run at headca908be2by checking the base copies of the two changed source files out over the worktree, running the four affected vitest files, then restoring withgit checkout HEAD -- <the two paths>.Six runs contributed: the fix run (
027380b3), an adversarial verification run (9daa0b3b, tests 3 to 7 and 9 to 14), a rework run answering the gating review (e70d31f3, tests 15 to 17 and 21 to 22 plus the first full build), a second adversarial verification (534cf3c9, tests 18 to 20), a rework answering a second-opinion review about test naming (94a1c92f, which re-ran both vitest arms and the full build), and a third adversarial verification (ca908be2, tests 23 to 26, both vitest arms re-run at that head, mutants re-run, the two consumer probes in## Boundariesrows 27 and 28, and this document reconciled to the new head).Historical and labelled as such above: the local Playwright transcript (tests 21 and 22) is from
e70d31f3and the localnpm run buildtranscript from94a1c92f; the fork CI atca908be2re-ran both.git diff 027380b3 ca908be2 -- packages/is empty, so no source file has changed since the fix commit.Not executed by any run, for the fork CI or the operator to confirm:
npm run test(whole vitest suite) and the wholenpm run test:e2esuite: only the affected test files were run here, per the repo's "run the smallest relevant tests" rule inDEVELOPMENT.md:64. The fork CI ran both whole suites atca908be2(see the next paragraph).Fork CI at
ca908be2:gh pr checks 1 --repo sprayberry-code/hermes-studioreports both workflows green,build pass 5m18s(run 35940638067:npm ci,harness:check,test:coveragewithTest Files 693 passed | 7 skipped (700),npm run buildto completion) ande2e pass 11m34s(run 35940638064:npx playwright install --with-deps chromium,npm run test:e2e,212 passed (10.6m)). No non-green job.Prior art
gh search prs --repo EKKOLearnAI/hermes-studio "approval resolved flag" --limit 10→[]gh search prs --repo EKKOLearnAI/hermes-studio "respond_approval" --limit 10→[]gh search prs --repo EKKOLearnAI/hermes-studio "bridge approval" --limit 10→ fix(agent-bridge): persist legacy execute_code approvals EKKOLearnAI/ekko-studio#2135 (open,fix(agent-bridge): persist legacy execute_code approvals, touchesbridge_pool.pybut only the execute_code allowlist memory patch), [codex] Sync bridge approval allowlist EKKOLearnAI/ekko-studio#1272 (merged), feat(group-chat): 建立 v2 actor/access 与运行时撤权基础 EKKOLearnAI/ekko-studio#2233 (open, group-chat v2 actor/access), [codex] fix tool approval flow EKKOLearnAI/ekko-studio#773 (merged)gh search prs --repo EKKOLearnAI/hermes-studio "bridge_pool.py" --state open --limit 10→ feat(bridge): deliver terminal notify_on_complete completions to Web UI sessions EKKOLearnAI/ekko-studio#2544, fix: persist Hermes Studio bridge interactions EKKOLearnAI/ekko-studio#2339, Web UI only recognizes hardcoded slash commands, custom bundles/skills don't appear or resolve EKKOLearnAI/ekko-studio#2152, fix(agent-bridge): persist legacy execute_code approvals EKKOLearnAI/ekko-studio#2135, feat: add session-scoped approval mode switcher to chat input EKKOLearnAI/ekko-studio#2455, none touchingrespond_approval's event payloadgh search prs --repo EKKOLearnAI/hermes-studio "handle-bridge-run.ts" --state open --limit 10→ fix: attribute shared-workspace run changes to the session that wrote them EKKOLearnAI/ekko-studio#2407, fix(chat): heal stale bridge session titles from Hermes state.db EKKOLearnAI/ekko-studio#2658, fix(client): show token usage on replies when Show Cost is enabled EKKOLearnAI/ekko-studio#2240, fix: goal kickoff runs losing conversation history due to excludeLastUser truncation EKKOLearnAI/ekko-studio#1946, none touching theapproval.resolvedbranchgh search issues --repo EKKOLearnAI/hermes-studio "approval.resolved resolved" --limit 10→[]gh pr list --search "<n> in:body"for each cited issue: Gateway approval misroute: respond_approval() discards approval_id, resolve_gateway_approval() pops FIFO — user's Allow/Deny lands on the wrong pending command EKKOLearnAI/ekko-studio#1992 and [Bug] 审批命令超时后,迟到点击"仅此次"仍被接口接收但命令不执行,弹窗却被关闭(approval lifecycle desync) EKKOLearnAI/ekko-studio#2558 both reference PR fix(approval): wait for authoritative runtime response EKKOLearnAI/ekko-studio#2648,fix(approval): wait for authoritative runtime response, CLOSED unmerged 2026-08-22, 2643 additions across 49 files, no review comments. It attempted the whole lifecycle at once and explicitly put Gateway approval misroute: respond_approval() discards approval_id, resolve_gateway_approval() pops FIFO — user's Allow/Deny lands on the wrong pending command EKKOLearnAI/ekko-studio#1992 out of scope. This change is the two-line subset of that surface that is provable in unit tests.mainby merged PR fix(write-gate): make approval results reliable and serialized EKKOLearnAI/ekko-studio#2295 and is merely unclosed.Policy
No
CONTRIBUTING.md,.github/CONTRIBUTING.md,CODE_OF_CONDUCT.md,.github/PULL_REQUEST_TEMPLATE*,AI_POLICY.md,.github/AI_POLICY.md,AI.mdorAGENT_POLICY.mdexists in this repository at0cc8271e(eachgh api repos/EKKOLearnAI/hermes-studio/contents/<path>returned 404; re-confirmed absent in the worktree at this head). There is no CLA and no DCO sign-off requirement.AGENTS.mdis titled "Agent Map" and opens: "This file is a short map for coding agents. Keep detailed guidance indocs/and keep this file small enough to fit into every task context." AI/agent contribution is anticipated by the repository's own tooling docs; there is no ban and no disclosure requirement stated anywhere in the repo.Rules followed, quoted verbatim with their file and line:
AGENTS.md, Hard Rules, line 49: "Do not mix unrelated refactors into a bug fix." The source diff is three lines in two files, one bug. All five test commits are tests only (git diff 027380b3 ca908be2 -- packages/is empty).AGENTS.md, line 36: "tests/client,tests/server,tests/shared- Vitest coverage." / line 37: "tests/e2e- Playwright browser coverage with mocked backend services." The new client tests are intests/client/, the browser tests extend the existingtests/e2e/chat-streaming.spec.ts.DEVELOPMENT.md, Coding Rules, line 35: "Do not mix unrelated refactors into feature or bugfix commits."DEVELOPMENT.md, Testing Rules, line 61: "Add focused Vitest coverage for server and store logic changes." Focused vitest added for both the server forwarder and the store consumer.DEVELOPMENT.md, Testing Rules, line 62: "Add Playwright coverage for browser-visible flows and routing/auth regressions." Tests 21 and 22, appended to the existing approval e2e file, executed in a real Chromium ate70d31f3. Their limits are stated above rather than glossed.DEVELOPMENT.md, Testing Rules, line 63: "For frontend browser tests, prefer API/socket mocks over real external services." They use the repo's ownmockHermesApi/mockChatSocketfixtures.DEVELOPMENT.md, Testing Rules, line 64: "Before opening a PR, run the smallest relevant tests plusnpm run build." The affected test files were run on both arms at this head; the fullnpm run buildcompleted locally at94a1c92fand in the fork CI at this head; transcripts in## Test evidence.DEVELOPMENT.md, Commit And PR Rules, line 123: "Mark a PR ready only after the relevant tests and build pass."DEVELOPMENT.md, Commit And PR Rules, lines 116-117: "Branch frommainfor new work." / "Use short, descriptive branch names such ascodex/fix-login-tokenorfeat/group-chat-copy." Branched frommainat0cc8271e; branchfix/bridge-approval-resolved-flag.DEVELOPMENT.md, Commit And PR Rules: "Use concise commit messages that describe the change" / "Commit only files that belong to the change." Six commits, source then tests, nothing unrelated.Tooling run at this head:
npm ci --ignore-scripts --no-audit --no-fund,npx vitest run <the four affected files>on both arms,npx tsc --noEmit -p packages/server/tsconfig.json;npm run buildlocally at94a1c92fand in the fork CI atca908be2.Disclosure facts for the operator
Plain facts about what AI did on this change, for you to write your own disclosure:
respond_approval()againstapplyBridgeChunkAsync()and the client'sclearPendingApproval(). It was not reported as such by any issue: [Bug] 审批命令超时后,迟到点击"仅此次"仍被接口接收但命令不执行,弹窗却被关闭(approval lifecycle desync) EKKOLearnAI/ekko-studio#2558 reports the symptom, Bug: WebUI approvals never resolve — respond_approval() requires request_id that hermes-agent <v0.20.5 never sends EKKOLearnAI/ekko-studio#2756 reports the condition.except Exceptionbranch was untested. It split that test and added tests 3 to 7 and 9 to 14.npm run build.94a1c92f. It found that every forwarder test read only the socket broadcast while the same payload also goes to theonEventobserver (webhooks, push, foreground notices) and into the replayed run state, and added tests 23 to 26 for those two outputs. It then re-ran both vitest arms and the five mutants atca908be2, probed the two other consumers of the outcome (group-chat store, voice relay) on both arms, and reconciled this document.e70d31f3) and are labelled as such; later containers had no Chromium available. The fork's CI ran the whole Playwright suite atca908be2(212 passed). No source file has changed since027380b3, and the e2e delta sincee70d31f3is two test names and three comments.ca908be2). Not run anywhere: a live test against a real hermes-agent runtime.Boundaries
Every predicate, comparison and truthiness check the diff adds or changes. Ledger rebuilt from the diff at head
ca908be2.Diff element 1,
"resolved": resolvedadded to the event dict (bridge_pool.py). No new predicate; it publishes the value bound on the preceding lines byresolved = bool(gateway_request_id) and resolve_gateway_approval(...) > 0, inside atrywhoseexcept Exceptionsetsresolved = False. Rows for every value that binding can take at this point:resolvedpublishedgateway_request_idnon-empty, runtime resolved 1 waiter (> 0true)Truegateway_request_idnon-empty, runtime resolved 0 waiters (> 0false, entry expired or never registered, EKKOLearnAI#2558)Falsegateway_request_idempty string (bool("")false, hermes-agent < v0.20.5, EKKOLearnAI#2756)False, short-circuit; test 10 also assertsresolve_gateway_approvalis not calledresolve_gateway_approvalraises (import error or runtime error)Falsevia the existingexcept ExceptionRuntimeErrorgateway_generation is None(approval id unknown)return {"resolved": False}above the changed lineresponse_queue is not None_approval_callbackappends its own event (line 1426) with noresolvedkey_cancel_pending_approvals_for_generation, line 2156)resolvedkey and carriesreason: "Session interrupted"Diff element 2,
...(typeof ev.resolved === 'boolean' ? { resolved: ev.resolved } : {})(handle-bridge-run.ts). Atypeof === 'boolean'guard, deliberately not a truthiness check, so thatfalseis forwarded rather than swallowed:ev.resolvedtrue{..., resolved: true}false(the falsy-but-valid case; a truthiness guard would drop it and reintroduce the bug){..., resolved: false}undefined/ key absent (runtime older than this change)resolvedkey, payload byte-identical to todaynull(typeof null === 'object')resolvedkey0, falsy non-booleanresolvedkey"", falsy non-booleanresolvedkey"true", truthy non-boolean, the case a truthiness guard would wrongly forwardresolvedkeyCall surfaces of the changed forwarder.
applyBridgeChunkAsyncis reached from four call sites in two exported entry points; the fix claims to cover both:handleBridgeRun, live run,bridge.streamOutputloop (lines 835, 881)resumeBridgeRun, reconnect, snapshot plusbridge.getOutputpoll loop (lines 1043, 1089)Outputs of the changed branch. The
payloadobject built in the diff is handed toreplaceStateand then toemit, andemit(lines 629 and 1012) forwards the same object to theonEventobserver and to the socket broadcast. Every output must carry the outcome, or a consumer keeps seeing base behaviour:nsp.to(session).emit)onEventobserversockets/chat-run.ts:1622wraps it intoobserveChatRunWebhookEventandemitPendingInteraction;app-events.ts:60callsforegroundNotification(), which returnsnullforresolved === false(foreground-notification.ts:5), so on base a failed approval produces a "resolved" push noticeforegroundNotification()resultreplaceStateintostate.eventsbuildResumeEventsreplays it to a reconnecting client;sockets/chat-run.ts:748keeps an approval pending in the app-event snapshot only whendata.resolved === falseSibling branches of the same
if/else ifchain, to show nothing else changed:approval.requestedresolvedfield involvedclarify.resolved/clarify.requestedapproval.resolvedemitted by the ekko-agent path (handle-ekko-agent-run.ts:1268)resolved: trueDownstream consumer, every branch now executed.
clearPendingApproval()(packages/client/src/stores/hermes/chat.ts:3268-3292) branches on(evt as any).resolved === falsestrictly, in two places: once in the no-pending early return and once with a card pending.resolvedkey (rows 10 to 14)resolved: false, nostale/error(row 9), card pendingclearPendingApprovaltakes theresolved === falsebranch and returns without deleting: the card stays upresolved: falseandstale: true(row 9 plus expiry), card pendingdismissPendingApprovalForruns and, because the user had attempted a response,notifyPendingInteractionExpired()firesresolved: true(row 8)resolvedisundefinedthereresolved: falseplusstale: truebut no pending card for the session (!current,chat.ts:3275-3280), the state the main chat view is always in, becauserespondApproval()dismisses the card on submit (chat.ts:3401, called fromMessageList.vue:368)notifyPendingInteractionExpired()fires because the user attempted and the event is stale and it reportsresolved === falseresolved: trueplusstale: true, no pending card (same early return, successful resolution)## Test evidenceOther consumers of
approval.resolvedreached by the same payload shape, probed on both arms rather than tested, because the diff does not route through them and their existing tests already pin theresolved === falsegate with hand-built payloads:group-chat.ts:1227(data.resolved === falsekeeps the pending approval)resolved: false: stays 1;resolved: true: 1 to 0. The group-chat server path (agent-clients.ts:1739-1745) rebuilds its own payload fromapproval_idandchoiceonly, so group rooms never see the outcome on either arm; out of scope for this change, noted for the follow-upoutbound-relay-client.ts:1233(event.resolved !== false, forwardserror: 'approval failed'to the device)tool.completedwith no error;resolved: false:error: 'approval failed';resolved: true: no error. Same shape as the chat store: on base the device is told the approval succeededSuggested upstream PR title
fix: report bridge approval outcome on approval.resolved