Skip to content

feat(collaboration): wake the requesting lead once a delegated result is accepted - #5304

Open
songoow wants to merge 22 commits into
loopx-project:mainfrom
songoow:codex/delegation-wake-lead
Open

songoow wants to merge 22 commits into
loopx-project:mainfrom
songoow:codex/delegation-wake-lead

Conversation

@songoow

@songoow songoow commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Why

Roadmap R3: automatic return should continue the lead without owner polling. When a delegated member's result became accepted, nothing scheduled the requester's next bounded Turn. The App's return service only appends a follow-up message and never starts a Turn, and delegation routes are peer routes that do not even receive that message. A Goal Chat LoopX lead therefore saw a result only if it happened to call read/wait during its own Turn.

Seam decision

The wake is emitted on the first transition to accepted, not on adoption. Adoption runs inside the lead's own Turn, where planChatMode refuses a new Turn because one is already active.

What changed

  • Intent. transitionDelegationObservation returns a wake_intent (requester identity, operation, request and accepted artifact digests hashed into intent_id) only on a non-accepted to accepted transition; accepted to accepted emits nothing.
  • Persistence. _observe stores the intent as wake: {..., state: "pending"} in the same write that records accepted, so result and intent are crash-atomic. An in-Turn read/wait/adopt that observes the accepted result marks it observed_in_turn.
  • Admission. LoopXMode.wake and a wake operation in planChatMode reuse the resume rules plus mode enabled and not paused, and accept only a host origin. On admission it submits /goal resume with client_turn_id = "wake-" + intent_id[:32], so turn_for_client makes a retry idempotent. The session lock is taken before the record lock, the same order the in-Turn tool uses.
  • Pump. A delegation wake pump runs beside the return service in the Chat server. It resolves the owner from the intent's requester, skips an intent stored under another requester's address, and isolates per-record failures.
  • Receipts. The result and the wake are separate facts: woken (session, Turn, created), pending (lead_turn_active, lead_paused, allowance_exhausted) or refused (goal_stopped, lead_unbound, binding_revoked, native_goal_complete, native_goal_absent, wake_identity_conflict, no_wake_owner). A wake never unpauses the lead, never starts a native Goal and never raises the allowance.
  • Placement. The pump lives in chat_loopx_mode.py beside the owner it drives. An earlier revision added a new top-level module, which broke the top-level module budget.
  • Census and docs. The registry I/O manifest is regenerated for the pump's Goal context read. goal-chat-continuation.md and local-delegation.md describe the receipts and limits in English and Chinese.

Behavior change disclosure

A running Chat service now starts one Turn for an enabled, unpaused Goal Chat LoopX lead after its delegated result is accepted. Nothing changes for leads outside Chat LoopX mode (no_wake_owner, the scheduler deadline recheck remains their continuation), for paused leads, or when the Chat service is not running.

Checks run on this head

Check Result
pytest tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py tests/test_chat_delegation_journey.py tests/test_collaboration_mcp.py tests/test_local_delegation.py tests/test_delegation_inventory.py 55 passed
New wake tests alone 11 passed; mutating both accepted-status guards fails 3, removing the existing-Turn recovery fails 2
node --test tests/control_plane_ts/delegation.test.ts tests/control_plane_ts/chat_mode.test.ts 21 passed, 0 failed
npm run typecheck:control-plane ok
tests/architecture/test_top_level_module_budget.py, import boundaries, registry census passed
examples/docs-governance-smoke.py, examples/semantic-vocabulary-drift-smoke.py ok
loopx canary premerge --from-git-diff passed: tier=standard, changed_files=13, surfaces=control_plane/docs_project_content/public_boundary/python; selected=13, failures=0

Not run: a live model Turn started by a wake, the packaged App and Lark. The wake starts a Turn through the existing submit_turn; the App shows it as an ordinary LoopX execution Turn.

Coordination with other open PRs

Rollback

Revert the PR, or stop the pump alone: pending intents stay pending and wake nobody. Receipts already written remain readable facts on the operation record.

Bounded future-facing refactor

Applied: the pump sits inside the owner it drives instead of a new top-level module. Deferred: converging the return service and the wake pump into one Chat-host tick, which would change the return service's cadence and belongs with R3's return-lifecycle unification.

This is a control-plane change; it is left for maintainer review and merge.

🤖 Generated with Claude Code

songoow and others added 12 commits September 29, 2026 12:31
The transition to accepted is the one durable moment a requester can be
continued without polling. transitionDelegationObservation now returns a
wake_intent beside the accepted status: schema, an intent id derived from
the requester, operation, request and accepted artifact digests, and the
requester identity. accepted -> accepted stays an idempotent readback and
emits nothing; rejection wakes nobody. The intent grants no Turn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…in-Turn

The accepted observation now stores row["wake"] = {intent, state: pending}
in the same atomic write that records accepted, so a result and its wake
receipt cannot diverge. record_wake settles a pending intent under the
same lock adopt_result uses and never touches a non-accepted row; the
Chat LoopX tool marks the intent observed_in_turn whenever the lead sees
an accepted result inside its own Turn, so no redundant wake follows.
read_delegation exposes the receipt as a fact distinct from the result.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…owner

planChatMode gains a host-only "wake" operation that reuses the owner
resume facts (Goal active, no running Turn, registered coordinator, valid
bindings, resumable native Goal, allowance above consumed tokens) plus
mode enabled and not paused, and returns a typed outcome instead of an
error: admitted, pending (lead_turn_active, lead_paused,
allowance_exhausted) or refused (goal_stopped, no_wake_owner,
lead_unbound, binding_revoked, native_goal_complete, native_goal_absent).

ChatLoopXMode.wake takes the session lock before the record lock, asks the
rule, and submits one idempotent "/goal resume" Turn keyed by the intent
id, so a crash between submit and receipt recovers through turn_for_client
with created=false. apply rejects the host-only operation. The wake Turn
tells the owner why it started and tells the lead which operation to
read; it never unpauses the lead or starts a native Goal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
DelegationWakeService scans conversations that configured a coordinator,
maps their observed operations (loopx_deliveries) to the requester-scoped
records, and asks ChatLoopXMode.wake to settle each pending intent under
the same record lock adopt_result uses. It prefers a conversation that is
still enabled over one that exited, changes a pending receipt only when
its reason changes, and skips records held by a worker until the next
tick. Requesters outside a Chat LoopX conversation are never visited.
Closing the Chat server stops the pump; intents left pending are the
rollback state.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…ting

The wake rule checked the native status before the running-Turn fence, so
an owner start that was still activating (Turn active, native "absent" or
the previous run's "complete") refused the wake terminally. The fence now
comes first; native complete/absent are refused only when no Turn runs.
The rule test covers every admitted, pending and refused outcome, and that
the host origin admits only wake.

The intent uses schema_version like every other collaboration contract and
requires the requester goal reference to be an object or null.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The pump visited only operations listed in a conversation's
loopx_deliveries, so requesters outside Chat LoopX mode stayed pending
forever and a lead with more than twenty operations lost wakes to the
bounded list. It now scans this runtime's operation records and resolves
the owner from the intent's requester identity, preferring a conversation
that can still continue and then the one that observed the operation; a
requester with no such conversation is refused with no_wake_owner. An
intent must name the requester that owns its storage address, and one
failing record no longer aborts the tick.

Receipts share one constructor that keeps only the typed intent and the
current state's facts. An existing wake-* Turn is recorded only when it
carries this intent; otherwise the wake is refused as
wake_identity_conflict rather than claiming another Turn. The Goal
context is read only after admission, so a removed Goal refuses instead
of raising. Adopting a result also retires the consumer's wake.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The typed rule accepted origin "web" for a wake, so the host-only
guarantee rested on the Python apply guard alone. Each origin now maps to
exactly one operation family: the host may only wake, and the owner's web
origin may never request one.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The pump was a new top-level module, which raised loopx/ from 147 to 148
top-level modules and failed the module budget. It only drives
LoopXMode.wake, so it now lives beside that owner in chat_loopx_mode. The
chat_runtime and collaboration_mcp dependencies are imported inside the pump
because chat_runtime imports this module. No behavior change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Expectations follow the wake contract: an accepted result with a pending
intent continues its requester's lead exactly once; a second tick, a crash
after submit, a client id owned by another Turn, an active or paused lead, a
stopped Goal, a requester without a lead conversation, an intent stored under
another requester and a non-accepted result each leave the matching receipt
without starting a Turn. Mutating both accepted-status guards or the
existing-Turn recovery fails these tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Goal Chat continuation now lists the wake receipt states and reasons and the
limits (no unpause, no new native Goal, no higher allowance, requires the
Chat service). The shell entrypoint note points to it. English and Chinese
sections stay aligned.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
The Chat server reads the registry to build the Goal context for a
host-initiated wake Turn, so that read is a new codec_api site and the
server's other sites moved. Regenerated with
scripts/generate_project_registry_io_manifest.py; no unclassified sites.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Regenerated with scripts/generate_project_registry_io_manifest.py after
rebasing onto main. Site ids and classifications are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: song <22676124+songoow@users.noreply.github.com>
@songoow
songoow force-pushed the codex/delegation-wake-lead branch from 4a72d07 to 3f99898 Compare September 29, 2026 16:32

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

English verdict: REQUEST_CHANGES — persisted queued Turns are mistaken for completed wake dispatch, and an unpinned conversation selector permits duplicate dispatch after receipt loss.

Reviewed exact head: 3f99898d65d2d0c880f56797d1e46f35efab87e2, against immutable merge base 9c29941559cff92f675c37cc80c4cf87a23f8448. This review follows the current LoopX PR-review capability execution contract.

[P1] queued Turn 存在不等于唤醒已派发;启动失败后会永久跳过恢复

位置:chat_loopx_mode.py:494–504。native admission 先持久化 queued Turn,再启动 adapter。若 adapter 在这两个阶段之间失败,intent 仍 pending,但下次 tick 只检查 Turn 的身份是否匹配,便直接写 woken,没有重新进入原生 dispatch/recovery。

独立复现使用真实 ChatRuntimeController.submit_turn、TS acceptance 和 durable Chat store,只在 adapter/worker 传输边界注入故障:第一次 adapter 初始化失败,Turn 为 queued、started_at=None、dispatch=0;第二次 tick 后 adapter 尝试次数仍为 1、dispatch 仍为 0,wake 却是 woken。后续 pump 不再处理它,结果无法触发承诺的后续 Turn。现有测试对 submit_turn 的替身不能证明这段恢复。

最小修复是复用既有原生 acceptance/dispatch 恢复 owner,而非把“Turn 存在”当作派发完成。chat_turn_acceptance.ts:453 的 dispatchFor 已区分 queued、running 等状态。请覆盖 acceptance 后 adapter 失败、准备 capsule 后失败、queued/running/terminal 重试和 receipt 丢失,并验证只恢复原 Turn、不另建 Turn,同时保留 pause/revocation 边界。

[P1] 每次重选会话会破坏原会话返回和跨重试的单次派发

位置:chat_loopx_mode.py:923–934。selector 按 enabled/usable 优先、其次 operation 观测、最后更新时间选择同一 Goal/Agent 的会话,但 intent 没有固定原始 conversation;而 turn_for_client 的去重范围是 session,不是整个 intent。

独立复现一:原会话拥有该 operation,随后退出模式;新建同一 Goal/Agent 的 enabled 会话。pump 向新会话提交,而不是保留原会话的关闭边界。复现二通过真实 native acceptance:首 tick 已在原会话接受并派发 Turn,仅 wake receipt 写入注入一次 OSError;随后原会话 disabled,创建新的 enabled 会话;重试产生第二个 session 的新 Turn 和第二次 dispatch,最终回执指向替代会话。两个调用使用完全相同 intent/client id。这里只替换 worker 传输以避免运行模型;观测到的是两次真实 admission 和两次派发请求,不声称已花费模型额度。

请复用既有原会话返回关系,在 intent/恢复记录中稳定绑定原 conversation 和已接受 Turn;转换 owner 必须是显式、受 fence 保护的 handoff,不能由 enabled 或 updated_at 隐式完成。补充多会话、退出模式、关闭源会话与 lost-receipt 后换会话的回归,证明同一 intent 不能在另一 session 再次派发。

动机

这是 roadmap R3 的实际交付缺口:member result 已 accepted,协调者却必须自己轮询才能继续。让 Chat host 在既有开启模式、授权绑定和总额度内继续同一个 lead,是有用的有界产品增量。但 R3 的“回原会话、只触发一次、丢回执可恢复”必须成立,accepted、Turn 被保存和实际派发不能互相替代。上述两处故障正好落在该验收路径上,而不是额外要求完成所有宿主或云端生命周期。

改动思路

TS observation owner 在第一次 accepted transition 生成稳定 intent;Python 将它与 accepted 同写,in-Turn 读取则标记 observed_in_turn。host pump 复用 ChatLoopXMode 与原生 submit_turn,不增加新的 Goal 或额度;typed chat-mode planner 保留 mode、pause、绑定、原生状态与 allowance 的拒绝规则。这些 seams 和同写边界合理。缺口在 Python 的 recovery shortcut 和动态会话选择,绕开了已有 Turn dispatch 与原会话关系,而非 intent hash 本身。

具体改动

共 12 个文件,新增 847 行、删除 23 行:Chat wake host/pump、collaboration intent/receipt、两个现有 TS owner、耐久回归、双语文档和派生 census。完整 diff、旧 caller 和 related return/delegation contract 已阅读,没有把另一 PR 的修复当作本 head 的证据。

关键代码讲解

  • transitionDelegationObservation(delegation.ts:334)仅在首次 accepted transition 生成 requester/result 绑定 intent;accepted 重读不会产生第二条 intent,终态非 accepted 不触发。
  • Delegations._observe(collaboration_mcp.py:689)把 accepted 与 pending wake 原子记录;record_wake(316)在 operation 锁下结算,工具读取与 host 使用一致锁序。
  • ChatLoopXMode._wake_decision(chat_loopx_mode.py:473)先恢复 client Turn,再走 typed host admission。现在前半段把 queued 存在直接当 woken,是第一个阻塞点。
  • _select_owner(923)与 pump_delegation_wakes(950)发现 pending operation、挑选目标并隔离单条故障;重试未固定 session,形成第二个阻塞点。

本人重跑六个 Python suite 共 55 项、TS delegation/chat-mode 共 21 项、architecture 28 项均通过,typecheck 通过。相同 fixture 在不可变 base/head 通过真实 mode apply → controller → TS acceptance 与持久化 readback,普通 configure-off、start、completed-resume refusal 的完整规范化结果一致;disabled 不派发,普通 start 单次派发,拒绝全文及零副作用保持。仅规范化随机 ID、时间、临时 fixture 根目录及由该目录参与的 digest,并同时保留完整 config/request,未删除语义字段。另加独立负例 3 项均失败,覆盖 queued 恢复、原会话隔离和 lost-receipt 跨 session 重复派发。

对主干的风险

这是 enabled Goal Chat 的默认行为变化,不是新 actor 生命周期或跨 Agent 权限。普通单会话 off 路径的 paired parity 已验证,但源会话 disabled 后另一会话仍可消费它的 wake,故不能认定 scoped default-off 已隔离。实际执行仍要经过注册绑定、原生 admission 与预算 gate;execution guidance 是模型提示,不是完成/收养 obligation 的替代。状态和错误文字保持 domain-neutral,没有新增 substring classifier。后台 service 自带启动停止边界,未证明已安装 App 或 Lark 的运行体验。

语义与 CI 对齐

远端 CI 不查询、不轮询。风险 canary 执行 11 个 selected 检查和 5 个 direct 检查,唯一失败为现有 semantic twin budget 44 independently maintained py/ts twins; budget is 43。同一原始检查在 immutable base 和本 head 的失败身份/细节相同,PR 不改其因果路径,census 全量生成字节一致,错误行 mutation 仍被拒绝;该既有红项不决定评审。此次阻塞来自独立复现的本 PR 唤醒语义:queued 不是 dispatched,动态选中的同身份新会话不是原 conversation,也不能使 session-local client id 成为跨会话单次派发证明。

我的整体评价

这是有用的 R3 slice,但恢复和返回绑定未达到标题承诺。bounded future-facing pass:将 pump 留在现有 Chat mode owner 而非新增顶层模块是合理的;当前必须复用原生 dispatch 状态机和既有原会话返回关系,避免另建 replay 决策源。更大的 return/wake tick 合并可以继续 deferred,不作为本次修复前提。未运行 live model、packaged App 或 Lark,不能把 transport 替身扩展成这些证据。修复两个 blocker 后重新审查完整 exact head;control-plane 变更留给 maintainer 合并。

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
A persisted queued Turn is not a completed wake dispatch.  The wake decision
now carries the Turn already accepted under the intent's client id to the
typed planner: only a Turn that started is recorded as woken, a still-queued
Turn is replayed through the same native acceptance and dispatch owner, and a
Turn that ended before it started is refused instead of being reported as a
wake.  Pending intents therefore recover from an adapter failure between
acceptance and dispatch.

Signed-off-by: song <liusongstep@gmail.com>
planDelegationWake now refuses an intent whose conversation is not this
conversation, or whose coordinator is not the configured one, and it reads the
wake Turn's own facts: a started Turn is dispatch evidence, a queued one is
replayed under the same admission (including its own active Turn id), and a
Turn that ended before it started refuses the wake.  The planner returns
woken/admitted/refused/pending with the exact dispatch the host must perform,
so the host can no longer infer dispatch from a Turn's existence.

Signed-off-by: song <liusongstep@gmail.com>
The trusted Chat host records the session and Turn that started a delegated
operation, once, on first creation; a replay under the same operation id never
rebinds it, and the model supplies no routing.  The accepted-result wake intent
carries that conversation, so a wake returns to the conversation that owes it
and to no other.

Signed-off-by: song <liusongstep@gmail.com>
The originating conversation is part of the intent identity, so the same
accepted result cannot be woken into a different conversation of the same
requester.

Signed-off-by: song <liusongstep@gmail.com>
Real ChatRuntimeController.submit_turn, TS acceptance and durable store with
faults injected only at the adapter and worker transport: adapter failure after
acceptance is re-dispatched as the same Turn, a prepared capsule is repaired,
a lost receipt replays only the original Turn, a Turn cancelled before it
started is not woken, a queued Turn keeps the pause boundary, and neither
exiting, closing nor replacing the origin conversation moves the wake to
another session.

Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

English verdict: both blockers addressed on the new exact head; please re-review.

Reviewed head: 3573ade2d7993e2c72b8b41741c3804ea52c6886 (upstream/main 3ec049e13 merged in, --signoff, no rebase). Both findings were reproduced by a failing test before the fix, on real ChatRuntimeController.submit_turn + TS acceptance + durable store, with faults injected only at the adapter/worker transport (no model, no provider).

[P1] queued Turn ≠ completed dispatch — fixed in 43fcdb834, 5e74c7beb

The decision no longer short-circuits on "a Turn exists". _wake_decision passes the Turn already accepted under the intent's client id to the typed planner (wake_turn: id, status, started_at, loopx_execution, operation, intent_id), and planDelegationWake decides:

  • started → woken (dispatch: "recorded"); dispatch is not repeated.
  • queued → admitted with dispatch: "replay", and the host re-enters submit_turn with the stored message and request, so chat_turn_acceptance.ts dispatchFor()/planSettledReplay/planPreparedRepair own the recovery. Same Turn, no second Turn. The wake's own queued Turn is excluded from the lead_turn_active check, so replay is not self-blocked.
  • ended before it started → refused: wake_turn_not_started (its client id cannot admit another Turn).

Evidence: test_adapter_failure_after_acceptance_is_redispatched_not_recorded (first tick: adapter init fails, Turn queued with started_at=None, 0 dispatches, intent stays pending; second tick: 2 adapter attempts, 1 dispatch, 1 wake Turn, receipt woken with created=False), test_failure_after_the_prepared_capsule_is_repaired_and_dispatched (prepared capsule → repaired → dispatched), test_lost_receipt_replays_only_the_original_turn[running|terminal], test_wake_turn_cancelled_before_it_started_is_not_woken, test_queued_wake_turn_keeps_the_pause_boundary (pause still refuses, zero dispatch).

[P1] conversation pinning — fixed in 03efe2150, be361feb(=be361c160), 5e74c7beb

The implicit selector is gone (_owners/_select_owner deleted). The trusted Chat host records conversation={session_id, turn_id} beside the operation on first creation only (Delegations.start(..., conversation=...), written by chat_loopx_mode.dispatch from its own session and Turn, never model-supplied), and delegationWakeIntent folds it into intent_id. The pump now wakes only that session: no conversation → refused: no_wake_owner; closed/exited/session missing → refused: no_wake_owner; another conversation or a reconfigured coordinator → refused: wake_identity_conflict. No implicit handoff, and the scheduler deadline recheck remains the fallback.

Evidence: test_wake_never_moves_to_another_conversation_after_exit, test_closed_origin_conversation_is_refused_without_a_substitute, test_intent_without_an_originating_conversation_wakes_nobody, test_pinned_conversation_rebound_to_another_requester_is_refused, test_lost_receipt_then_new_conversation_cannot_dispatch_twice (first tick accepts + dispatches in the origin session, receipt write lost; origin conversation exits; new same-identity conversation is enabled → retry lands in the origin session, created=False, exactly 1 dispatch, and the substitute session has no wake Turn), plus test_the_starting_conversation_is_pinned_beside_the_operation and test_the_in_turn_tool_pins_the_conversation_it_runs_in (a second start under the same operation id never rebinds the pin).

Verification on this head

  • Python: pytest tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py tests/test_chat_delegation_journey.py tests/test_collaboration_mcp.py tests/test_local_delegation.py tests/test_delegation_inventory.py tests/architecture/test_project_registry_io_census.py tests/architecture/test_top_level_module_budget.py → 81 passed. Wake suite alone: 22.
  • TS: node --test tests/control_plane_ts/{chat_mode,delegation,turn_acceptance}.test.ts → 22 passed, 0 failed; npm run typecheck:control-plane ok.
  • Mutation check (each fix reverted, its test fails, then restored): queued-Turn shortcut → 6 fail; not-started treated as dispatch evidence → 1 fail; conversation dropped from the planner identity → 1 fail; conversation dropped from intent_id → 1 fail; conversation pin not written by Delegations.start → 2 fail; pump ignoring the pin → 1 fail. All restored.
  • loopx canary premerge --from-git-diff --git-diff-base upstream/main: gate passed, tier standard, 5 direct + 11 selected checks, 0 failures, 0 manual holds.

Docs updated for both behavior changes (goal-chat-continuation.md, local-delegation.md, EN+zh).

Not run: a live model Turn started by a wake, the packaged App and Lark — unchanged from the previous head; the wake still goes through the existing submit_turn.

Control-plane change; leaving review and merge to the maintainer.

Review 修正已提交,新 head 3573ade2d:queued Turn 不再当成已派发(改为交回原生 acceptance/dispatch 恢复同一个 Turn,未启动就终止的 Turn 明确拒绝),唤醒按「启动该操作的会话」固定绑定,不再隐式选择其他同身份会话;两个阻塞点均有先失败后通过的回归与变异验证,请复审。

@huangruiteng huangruiteng left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

REQUEST_CHANGES — 来源 Session 的固定与已有 queued 回合重放已改善,但新建/重放 dispatch 后仍会提前写终态 woken。native worker 的首次启动事实落盘失败时,Turn 没有启动,存储恢复及 controller 重建后 pump 却都不再重试。这是当前恢复语义的 P1,不是外部 CI 红灯。

Exact head: 5304@3573ade2d7993e2c72b8b41741c3804ea52c6886;比较基线:3ec049e138917a8cce4f84197ba196d26445b2b0。按 LoopX PR-review capability policy 12 审阅完整差异和当前修复反馈,而不是沿用旧头结论。

动机

Roadmap R3 所需的是接收方完成独立验收之后,原 lead 能继续下一次有界 Turn,不要求 owner 反复轮询。此前 delegation 进入 accepted 不会调度发起它的 Goal Chat lead;追加消息也不等于开始回合。此 PR 面向已激活、本地管理的 Chat lead,是合理的 R3 增量,不是整个多 Agent 协作或 Goal 完成。正常路径减少“结果已接受但无人继续”的停顿,用户继续使用既有 start/resume、pause/exit。但投递终态必须对应可恢复的启动事实,否则服务故障仍会把自动继续变成 owner 手动救援。

改动思路

首个 accepted 转换在现有 delegation journal 上记录 intent;可信 host 在开始 delegation 时固定来源 Session/Turn,intent 身份包含这个不可变来源。typed TypeScript owner 判断是否合法继续,Python 负责已有 Chat store、锁、native 接受和 dispatch 传输,不另建决策源。pump 是投递器而非权限授予者:结果被接受不代表可以唤醒任意会话,更不代表 adopted、Todo 完成或父 Goal 完成。已 started 的同一 intent 回合可补记事实;仅 queued 时应通过原 native 路径修复未完成 dispatch。当前 typed 分支遵守这一区分,但 Python 在 submit 返回后无条件终结 intent,破坏了同一条规则。

具体改动

完整差异为 13 文件、+1164/-21,集中于既有 Chat mode/runtime/server、delegation journal、两处 TS owner、语义登记、回归测试和公开操作文档。新增状态是不可重算的 accepted-result identity/origin 与投递回执,不是第二套 canonical settlement。Session 与 journal 锁顺序一致,accepted replay 不重新发 intent,已有回合观察到结果时也不会再强行唤醒。

关键代码讲解

  • transitionDelegationObservation / delegationWakeIntent 只在首次进入 accepted、canonical completion 和 artifact 当前时产生 intent;requester、artifact identity 与可信来源共同确定身份,accepted→accepted 不重新调度。
  • planDelegationWake 验证原 Session、Goal、注册 lead、绑定、启用状态和剩余额度。其他活跃 Turn、pause 或预算耗尽保留 pending;关闭/解绑/原 Goal 完成拒绝。只有同 intent 的真正 started 回合记为 woken;queued 回合走 replay,不凭存在就宣称启动。
  • ChatLoopXMode.wake / _wake_decision 将 typed 决定送入原 submit_turn/native acceptance 和 worker 路径,稳定的 intent client id 把补偿锁在同一回合,而不是抢另一条 Session。不过 L569 无条件 _woken(...),没有验证异步 worker 的启动事实已持久化。
  • pump_delegation_wakes 从 canonical execution records 发现未终结 intent,校验 storage address,隔离坏记录;每 tick 限制的是实际改变的回执,不让前面 unchanged pending 记录遮住后面的可处理结果。它由现有 Chat server 托管并随之停止。

对主干的风险

[P1] 不要在启动事实尚未持久化时终结 wake intent。 位置 chat_loopx_mode.py:569。用真实 submit_turn、_start_accepted_turn_worker 和 _run_turn,只在隔离 Chat store 的第一次 queued→starting 写入注入一次 OSError。观察到 Turn 仍 queued、started_at=null,journal 却已 woken。恢复原 store writer 后再次 pump,随后用同一持久化 store/registry 重建真实 controller 再 pump,两次均返回空结果,原 Turn 仍未启动:record_wake 和 _pending_identity 都只接受 pending,终态回执使意图从恢复扫描永久消失。

这与只在“已有 Turn”分支区分 queued/started 不同:submit_turn 返回只证明接受及异步 thread dispatch,并不保证启动事实已落盘。最小修复应由现有 native dispatch/receipt owner 保留可恢复状态,直到同 intent 的 started 事实可读回;并让首次启动写入失败正确释放 worker single-flight 状态,以便恢复后继续同一个 Turn,而不是新建 Turn、换 Session 或提高额度。回归必须覆盖首次启动写入失败/进程中断、存储恢复、controller 重建、同 client id 重试、一次实际启动和之后终态去重。

来源改绑问题已验证修复:host 固定 conversation,模型参数不能写它,找不到原 owner 不会自动改绑。此前已 queued、adapter dispatch 尚未完成的重放和 worker single-flight 测试也通过;它们不能替代上述更晚的失败窗口。普通 Chat mode-off 的配置、start、完成 Goal 的 resume 拒绝与基线保持一致,包括完整返回、错误、配置、持久化和 dispatch 观察,而非只比较 decision code。

独立执行了 81 项 Python 测试、22 项 TS 测试、control-plane typecheck 和标准 premerge(5 direct + 11 selected),均通过;基线对应现有 Python 57 项、TS 19 项也通过。额外真实 Chat controller/store 对照的 3 个普通入口观测完全相同(标准化摘要 f02408b7130f9e5d71b7f80b7b09fe204571432373f7c7e8d5245924b);只去除随机 id、时刻及合成根路径派生摘要,保留字段、错误、持久化、消息和副作用。另两个独立检查覆盖真实 worker 的 queued replay 单次启动,以及 25 条 unchanged paused pending 前置后仍能处理后续记录,均通过。新增上述启动事实失败的独立检查则失败,准确暴露 premature terminal receipt。前两个检查的初版 fixture 有时序断言和 reason 名称错误,修正后重跑通过;故障检查初次有 import 路径设置错误,正确设置后才取得产品反例。未修改 PR 源码、放宽检查或把评审 setup 错误当成产品失败。

实际 mode/controller/store 和 journal 运行在隔离合成 Goal 中;付费模型/远端 delegated process 使用受控 transport substitute,未认证全部署 soak、远程 peer 生命周期或完整 R3 接收采纳旅程。pump 仍扫描 execution records;本次只证明发现完整性及变化上限,不宣称大规模长期扫描性能已达标。这是明确的残余验证边界,不应靠扩大超时或无关框架冒充解决。

语义与 CI 对齐

延伸现有 delegation observation、Chat mode/native acceptance 词汇,登记了真实生产 call site,没有新建跨 Agent authority。wake 是 host-only operation,不进入普通 web owner operation;只有已有显式启用的受管理 Goal Chat lead 可以被继续。公开文档明确“只有真正启动过的 Turn 才算已派发”,typed owner 也用 started 区分事实;当前 Python 终态写回与之不一致,不能只修改文案把 queued 说成 started。这里的 guidance 只是既有结果提示;入场、限额、身份是机器规则。修复后重跑 uv run --extra test python -m pytest -q tests/test_chat_delegation_wake.py tests/test_chat_loopx_mode.py,扩展真实 worker/重建恢复用例,再跑 TS owner 测试和对应 premerge。未查询、轮询或等待远程 CI。

我的整体评价

长期推进判断为“正常路径改善、故障恢复仍有缺口”,用户体验为“启用后的回传改善、关闭路径保留,但故障后仍可能要手动救援”;不把一条 woken 回执当作全系统持续运行证明。新增可靠性状态有真实 producer,不能仅靠当前 UI 状态重算;保留现有 TS 决策 owner 和 Python 传输边界是合适的,但终态必须受启动事实约束。未来向检查考虑了 session 选择、worker single-flight 和 journal 扫描:原 owner 已足够,不需要第二个唤醒框架;扫描优化若有实际规模证据再定位既有 owner,不为本 PR 造任务。完整范围与原问题相称、关闭路径兼容也已验证,但本 PR 自身的恢复义务尚未满足,因此要求修改;修复后重新核验新精确头。本次未合并。

English verdict: REQUEST_CHANGES — The current head fixes requester-session identity and earlier queued replay, but terminalizes wake after asynchronous submit without durable start evidence. A first worker-start write failure leaves the Turn queued and the wake woken; storage recovery and a fresh controller cannot rediscover it. Keep the same intent recoverable until matching start evidence, with native-worker cleanup and restart coverage.

…ct lands

`submit_turn` returning proves admission and an asynchronous dispatch, not
that the Turn's start fact is durable.  The wake path wrote a terminal
`woken` receipt right after it returned, so a first worker-start write
failure left the Turn queued with the intent already settled: storage
recovery and a freshly built controller both stopped rediscovering it.

Two changes, in their owning modules:

- `_wake_decision` now records `pending` / `wake_dispatch_pending` after
  admission and lets the next tick read the durable `started_at` back, so
  the receipt stays a statement about a start that actually persisted.
  The typed owner distinguishes `started_at` from a status guess, keeps a
  Turn that is still activating pending, and refuses only a Turn that
  ended without starting, reusing `isTerminalTurnStatus` rather than
  restating the terminal set.
- `_run_turn` released its single-flight guard only when `worker.start()`
  itself raised.  A failure writing `queued -> starting` escaped the
  worker body and left the key in `turn_done_events` forever, so every
  later dispatch of that Turn was a silent no-op.  The start write and
  the Turn body now run under one `finally` that hands the guard back.

Coverage pins the failure window itself: a one-shot `OSError` on the
first `queued -> starting` write, then the same Turn replaying through
native dispatch to one start, a rebuilt controller over the persisted
store rediscovering the intent, and terminal de-duplication after that.
Characterization of the pass path is unchanged.

Signed-off-by: song <liusongstep@gmail.com>
…-lead

Signed-off-by: song <liusongstep@gmail.com>
@songoow

songoow commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Re-review request — exact head b76fb6993efeddeb85b9b7dededc3a589fb7db0d

Thanks for the concrete failure window. I reproduced it on the reviewed head first, fixed the two causes, then re-ran the same fault path.

[P1] A terminal woken receipt before the start fact landed — fixed

chat_loopx_mode.py:569 wrote _woken(...) unconditionally after submit_turn returned. Your reproduction showed the consequence: the first queued -> starting write raises OSError, the Turn stays queued with started_at=null, and the intent is already terminal — so record_wake / _pending_identity only accept pending and both storage recovery and a freshly built controller stop rediscovering it.

Fix. _wake_decision now records settled("pending", "wake_dispatch_pending", turn_id=...) after admission. The receipt becomes a statement about a start that actually persisted: the next tick reads the durable started_at back through the same typed decision and records woken then. The typed owner distinguishes started_at from a status guess, keeps a Turn that is still activating pending, and refuses only a Turn that ended without starting — reusing isTerminalTurnStatus from chat_turn_acceptance.ts rather than restating the terminal set beside it.

The second cause: the worker single-flight guard was never released

Your review noted this should be handled by the existing native dispatch owner. While reproducing, the retry stayed refused even after the store was healthy, which the receipt path alone cannot explain:

  • _start_accepted_turn_worker registers turn_done_events[key] before the start fact is durable, and cleared it only when worker.start() itself raised.
  • _run_turn's queued -> starting write sits outside any try/finally in the worker body. An OSError there escaped and left the key in turn_done_events permanently, so every later dispatch of that Turn was a silent no-op.

Fix (chat_runtime.py): the start write and the Turn body now run under one try/finally that hands the guard back, so recovery continues the same Turn through the existing path rather than creating a new one, moving Session, or raising a limit.

Evidence

tests/test_chat_delegation_wake.py now covers the failure window itself, over the real submit_turn / _start_accepted_turn_worker / _run_turn with faults injected only at the worker transport and the store write — no model, no provider:

  • a one-shot OSError on the first queued -> starting write: the Turn stays queued, the receipt stays pending with no woken_at, and the single-flight key is released;
  • the next tick replays that same Turn through native dispatch to exactly one start, then a later tick reads the start fact back and records woken with created=false;
  • a rebuilt ChatRuntimeController over the same persisted store rediscovers the intent and starts the same Turn once;
  • terminal de-duplication: the following tick is a no-op.

Verification on the merged head: tests/test_chat_delegation_wake.py + tests/test_chat_loopx_mode.py 42 passed; chat_mode.test.ts 5 passed; control-plane typecheck clean. Latest origin/main merged with sign-off (--no-ff, no force-push).

The remaining CI red on this head is inherited from main, not from this diff: test_prompt_upgrade_hook.py::test_live_decision_adds_only_existing_required_read_channel[loopx_turn_run_once] and Execution chip is not a compact hairline row: 28px tall both reproduce on current main and are being fixed on their own PRs.

English verdict request: the intent now stays recoverable until matching start evidence, with the worker guard released, and both the same-process and rebuilt-controller paths are covered. Please re-review b76fb6993efeddeb85b9b7dededc3a589fb7db0d.

@cocolord cocolord left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

评审提交:b76fb6993efeddeb85b9b7dededc3a589fb7db0d。这是 policy-12 whole-PR、exact-head 复审。我重新核对了 GitHub 的完整 14 文件差异(+1424/-36)、提交与 checks,复用了上一轮在未变更边界上的证据并检查了失效条件,重点验证 3573ade2d 之后对 queued replay、首次 start 写失败和 worker single-flight 的修复。此前的 conversation pin 与 queued Turn 误判已实质修好;当前仍有一个更靠后的可靠性 blocker、一个公开协议漂移,以及一个由本 PR 引入的 required maintainability gate 失败。

动机

这个 PR 解决的用户问题明确且值得做:成员结果已经通过独立验收并进入 accepted 后,空闲的 Goal Chat LoopX 协调员此前不会自动继续,只能碰巧在自己的 Turn 中轮询,或等待 owner 再次操作。新行为限定在已显式开启、未暂停、绑定仍有效、原生 Goal 可恢复且额度充足的 managed Codex Goal Chat;它不触碰非 Chat/Lark/attached-host,不解除暂停,不新开 Goal,也不把 accepted 当作 adopted、Todo 完成或父 Goal 完成。正常路径能减少“结果已有、主线不动”的长期停顿,收益是真实且可观察的。

改动思路

首次进入 accepted 时,typed delegation owner 生成包含 requester、operation/request、artifact digests 和可信来源 conversation 的 intent;Python 在同一 requester-scoped operation record 中原子保存 wake: pending。Chat server 托管一个轻量 pump,只扫描 pending intent,固定回到启动 delegation 的原 Session,再由 planDelegationWake 复用现有 resume 的 Goal、身份、binding、pause、active Turn 和 allowance 规则。真正执行仍走 ChatRuntimeController.submit_turn 与 native acceptance;intent 派生的稳定 client id 让 prepared/queued 丢回执能够恢复同一个 Turn,而不是创建第二个 Turn。lead 在自己的 Turn 内 read/wait/adopt 时则写 observed_in_turn,避免多余唤醒。

这个 owner 切分总体正确:TS 决策、Python effect、现有 Chat store/native Turn、server lifecycle 都被复用,没有再造 scheduler 或跨 Agent authority。当前问题不在总体方向,而在 started_at 的含义跨层被放大:Python 把它写成 worker 开始处理的事实,TS 却把它解释成 provider 已实际接收 Turn 的终态证据。

具体改动

完整差异覆盖两份公开文档;chat_loopx_mode.py 的 guidance、host-only wake、pending 扫描与服务;collaboration_mcp.py 的 execution path、receipt、conversation pin 与 in-Turn suppression;chat_server.py 的启动/关闭;chat_runtime.py 的 wake display 和 worker guard;三处 typed contract;registry I/O manifest;以及四组 Python/TS 回归。测试占新增行的大部分,但生产边界仍增加了新的 durable intent/receipt 和周期性扫描语义。

关键代码讲解

  • transitionDelegationObservation / delegationWakeIntent 只在首次合法进入 accepted 时产生 intent;accepted→accepted 不重复调度,conversation 进入 identity,避免同 Goal/同 Agent 的另一会话接管。
  • Delegations._observe / record_wake / wake_observed_in_turn 把结果与 wake 分成不同事实,并在同一 record lock 下从 pending 走向 woken/refused/observed;因此任何 terminal receipt 都必须真实,因为后续不会再进入恢复扫描。
  • pump_delegation_wakes / ChatLoopXMode._wake_decision 校验 storage address 与 pinned Session,隔离单条坏记录,调用 typed owner,并把 queued/prepared 同 client id 交回既有 native acceptance 修复。当前 head 已把 submit_turn 返回后的回执从无条件 woken 改为 wake_dispatch_pending。
  • planDelegationWake 集中处理 identity、mode、Goal、binding、active Turn、native 状态和 allowance;也复用了 shared terminal Turn status。但 L94 把任意 started_at 直接当作 dispatch: recorded。
  • ChatRuntimeController._run_turn / _run_started_turn 新的外层 finally 正确修复了第一次 start 写失败时 single-flight key 泄漏;不过 started_at 在进入 _run_started_turn 之前就写入,后面仍要做 context、mode、driver 和 adapter/provider dispatch。真正的上游启动事件会在 _TurnEventBuffer 中把 Turn 改为 running 并记录 upstream_turn_id。

对主干的风险

[P1] started_at 不是 provider dispatch 证据,当前会产生不可恢复的假 woken。 chat_mode.ts:94 先看 started_at 就返回 woken,甚至早于 terminal status 分支;而 chat_runtime.py:1406 在 prepare_turn_context、loopx_mode.prepare、CodexGoalDriver.run / adapter.start_turn 之前就写入该字段。独立 exact-head 探针走 shipped store/controller/planner/receipt 两条路径:一条在写入后模拟进程退出并重建 controller;另一条让 prepare_turn_context 在 provider 调用前抛错并由真实错误路径把 Turn 标为 failed。两条都没有 turn.started 事件、没有 provider dispatch,但下一 tick 都把同一 intent 写成 terminal woken。之后 _pending_identity / record_wake 不再处理它,自动继续永久丢失。

最小修复应复用现有真实上游启动事实(例如 durable turn.started 对应的 running / upstream_turn_id),或新增一个明确但同 owner 的 dispatch receipt;对 pre-dispatch starting、failed、server_restarted 分别定义可恢复或可操作的终态。回归必须覆盖“start 写成功、provider 调用前失败/进程退出、重建 store/controller、不得 woken、同 intent 最终只实际启动一次、之后去重”,不能让 mock 直接把被验证的 postcondition 当作 dispatch。

[P2] 公开 reason vocabulary 与实现不一致。 当前 EN/ZH 文档仍列 wake_turn_not_started,而实现/测试已经改为 wake_turn_ended_unstarted;新增的 pending wake_dispatch_pending 也未列入文档。这里是 operation readback 的机器状态,不只是解释性文案。请确定稳定名称、同步两种语言,并用契约测试防止再次漂移。

Required gate 也由本 PR 引入失败。 精确 base 7e60e6999 的 maintainability ratchet 通过;精确 head 因 module_metric_budget:loopx/chat_runtime.py 未登记增长而失败。当前 loopx canary premerge --from-git-diff 共选 11 项,10 项通过,唯一失败就是该 catalog canary。请优先把新增职责放到最近的有界 owner;若维护者明确接受 hot-module 增长,再按仓库机制更新 reviewed ceiling 并说明收益。不能把 required red 仅归为 main 噪声。

其余验证:focused Python 42 passed;focused TS 22 passed;control-plane typecheck 和 changed-Python Ruff 通过;public boundary 与其他 9 个 selected canary/smoke 通过。远端当前为 19 success、11 failure、7 skipped;日志中的 28px UI、prompt-upgrade/generated-twin 等失败至少有主干既有成分,但 aggregate gate 仍红,我没有用这些噪声替代上述本 PR 反例。未运行付费 live-model、打包 App/Lark、跨平台 hard-kill 和大规模 journal 扫描性能,这些保留为残余验证边界。

语义与 CI 对齐

这个 PR 是对既有 delegation observation、Chat mode 与 native Turn 词汇的合理扩展;identity/permission 规则是 typed 且 domain-neutral,没有 substring denylist,也没有把强制 gate 写成 guidance。真正不对齐的是:公开文档声称“只有真正启动过的 Turn 才算已派发”,而 current typed rule 使用的是更早的 worker timestamp;同时文档枚举与实现 reason 不同。请修复事实边界和文档,再重跑 pre-dispatch fault/restart、focused Python/TS、typecheck、Ruff、reason contract、maintainability ratchet 与 premerge。

我的整体评价

结论是 REQUEST_CHANGES。功能动机明确,正常路径对长期推进和用户体验明显正向;conversation pin、queued replay、首次 start-write 失败和 single-flight cleanup 都已得到实质修复,现有 owner 选择也合理。但当前 exact head 仍会把“worker 准备开始”记录成“lead 已被唤醒”,这正好破坏本 PR 最核心的恢复承诺;公共 reason 也已漂移,且 required maintainability gate 是 head 新增失败。请在现有 planner/runtime owner 内做有界修复,不需要继续扩展框架;新 head 到来后我会重跑这两个 pre-provider 反例、reason 对齐和 exact-base/head gate。本次只发布 review,不修改作者代码,也不授予 merge authority。

English verdict: REQUEST_CHANGES on exact head b76fb6993efeddeb85b9b7dededc3a589fb7db0d. The feature has clear positive value and the prior conversation-pin, queued-replay, first start-write, and single-flight issues are materially fixed. However, started_at is persisted before context/provider dispatch, yet the planner treats it as terminal wake evidence; two real-store/controller fault probes produced woken with no turn.started event. Align success with durable provider-start evidence, synchronize the public reason vocabulary, and clear the PR-attributable chat_runtime.py maintainability ratchet before re-review.

if (turn.loopx_execution !== true || turn.operation !== "wake" || turn.intent_id !== intent.intent_id) {
return outcome("refused", "wake_identity_conflict");
}
if (typeof turn.started_at === "string" && turn.started_at) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] started_at cannot certify this terminal woken outcome. _run_turn persists it before _run_started_turn, so a process loss or prepare_turn_context failure can leave no turn.started event and no provider dispatch. On this exact head, both cases are nevertheless finalized as woken on the next pump, after which pending-intent recovery no longer sees them. Please base terminal success on the existing durable provider-start evidence (or an equivalent explicit dispatch receipt), define the pre-dispatch failed/restarted outcome, and cover controller rebuild plus exactly-one eventual start.

`pending` with `lead_turn_active`, `lead_paused` or `allowance_exhausted`, or
`refused` with `goal_stopped`, `lead_unbound`, `binding_revoked`,
`native_goal_complete`, `native_goal_absent`, `wake_identity_conflict`,
`wake_turn_not_started` or `no_wake_owner`. Its Turn is accepted once under a

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] The public reason list is stale on this head: the typed owner/tests now emit wake_turn_ended_unstarted, while this EN/ZH document still advertises wake_turn_not_started; the new pending reason wake_dispatch_pending is also omitted. Since these values are exposed by delegation readback, please synchronize both language sections and add a small contract check for the emitted reason set.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants