fix(workspace): wait for a running Turn before sending - #5261
huangruiteng merged 20 commits into
Conversation
The Chat service accepts one Turn per Session. After a reload the page only knew about its own in-flight send, so while a recovered Turn was still running the composer stayed open and a new message was rejected with HTTP 409, shown as the raw English service error. While the conversation shows a Turn in flight, the send button and quick prompts wait and the composer explains that the Turn can be adjusted or interrupted from the reply. LoopX mode keeps delivering into its running Turn. A 409 that still arrives is shown as the same localized message. Signed-off-by: song <liusongstep@gmail.com>
Sends a long Turn, reloads while it runs, and requires the composer to wait with its hint, send nothing, and reopen once the Turn completes. Signed-off-by: song <liusongstep@gmail.com>
huangruiteng
left a comment
There was a problem hiding this comment.
English verdict: REQUEST_CHANGES
Reviewed exact head: 4929e8fa9c240ca0e01e30f85ede9aeb02eee19c.
阻塞项:[P2] 新 composer guard 会把已取消的历史恢复占位当成仍运行的 Turn;切走再返回后,即使同一 Turn 已完成,仍无法正常发送下一轮。定位:conversationTurnRunning。
动机
目标是合理的:普通 Chat 恢复一个仍在执行的 Turn 时,不应因为页面的 sending 标记已重置,就允许再提交普通消息。用户仍应能调整或中断本轮,Turn 完成后继续下一轮;LoopX 模式的队列、inbox 和 steer 入口不能被这个普通对话限制一起封住。这个 head 修好了“reload 后运行中仍可发送”的局部问题,但没有保住恢复完成后的持续使用路径,所以尚不能认定该目标已完整交付。
改动思路
实现复用了现有 Dashboard Session/Turn 回调、PersonalWorkspace composer、双语状态文案和 chat-recovery 场景,没有另建服务或持久化状态。普通发送的 onTurn 给占位回复补充 sourceTurnId/sourceSessionId,composer 用当前上下文消息中的 pending 与 Turn id 判断是否应禁用发送。LoopX delivery 分支先处理自身 ingress,因此仍可接受队列与追加指令。边界方向是对的,但 pending 是界面投影,不是当前 Session 的权威存活状态;恢复 effect 在取消后返回时会留下旧占位,这个旧行为现在被新 guard 放大成发送锁。
具体改动
完整 PR 是五个文件,+48/-7:发送与快捷入口 guard、接受 Turn 后的身份投影、双语提示、提示样式,以及既有恢复 smoke 的增量;没有新增协议、provider 或全局调度状态。
关键代码讲解
- conversationTurnRunning / composerBlocked:输入是当前上下文的 managerMessages 和 loopxDeliveryOpen。新规则对任意 pending 且有 sourceTurnId 的消息返回运行中,并与 sending 合并。它统一驱动按钮禁用与状态提示,但没有核对该消息是否属于当前 Session 的当前活动 Turn,终态读回也不能覆盖旧 pending。
- sendMessage:保留 LoopX delivery 的 queue/steer/inbox 路由,再拒绝普通 running-Turn 提交;快捷按钮和键盘也使用同一 guard。不能通过移除 guard 来修复终态锁,那会重新打开本 PR 原来要修的重复提交路径。
- 恢复 effect 的 pending producer:恢复已接受 Turn 时追加 sourceTurnId 占位。切换 Goal 会取消旧恢复;其 success/catch 在 cancelled 时先返回,没有结算旧占位。返回同一 Goal 又创建新占位,新占位完成并不退休旧占位。这是 guard 的真实上游,不只是一个虚构的消息组合。
- 普通发送的 onTurn:把接受的 Session/Turn identity 写到正在显示的回复,便于现有活动控制与 guard 消费。这是投影元数据,不应单凭 id 存在推导仍在执行。
对主干的风险
[P2] 我通过实际 React 页面、原生 fetch 和隔离的 HTTP 合成后端复现了完整流程:发送一轮并保持执行 → reload 恢复 → 切到独立 Goal B → 在完成前返回 A → 后端完成同一个 Turn。head 显示完成回复,独立 Session 读回为 ready/active_turn_id=null,后端只接受了一轮,草稿仍在;但 Send 继续禁用,仍显示“本轮回答进行中”。相同输入和界面路径在不可变 base 完成后可以发送。B 的 composer 不被 A 阻塞,这个正向范围检查已通过;问题不是全局锁,而是当前上下文中的历史 pending 锁。reload 可以缓解,不能把它说成永久无法恢复。
最小修改是在现有恢复/Session owner 中绑定当前身份与生命周期,并在取消、重放、终态读回时退休旧占位;避免另外维护一个 busy flag。扩展既有 chat-recovery:A→B→A 时完成同一 Turn,验证草稿保留、下一轮可发送,同时保留真正 running 时禁用、调整/中断与 LoopX exemption。还应检查切换新 Session 不继承旧占位,不能靠只改提示文字隐藏失败。
语义与 CI 对齐
当前义务是“当前活动 Turn 排斥普通新发送,终态恢复下一轮”,不是“所有历史 pending 都是活动执行”。同一个状态事实的 base/head UI 反例证明了语义违约;请沿以上当前 owner 修复后重跑 LOOPX_PERSONAL_WORKSPACE_SCENARIO=chat-recovery node examples/personal-workspace-browser-smoke.mjs,并补上上述切换/终态场景。
已亲自运行:dashboard 的 npm run build(TS、Vite 及 packaged Chat assets)、source 与 packaged chat-recovery、loopx-mode,以及 uv run --extra test python examples/loopx-chat-runtime-smoke.py,均通过。运行时 smoke 使用隔离状态与仓库 fake executor,不是 live 模型证明。新独立终态 oracle 在 head 失败。
另有既有 source-contract 检查失败:line359 期望未改动代码中的 workspace_ref current 字面表达。不可变 base 与 head 同命令、同失败 identity、完整 stderr 仅规范化 checkout 根路径后相同(SHA256 bf9163b485948835873b89c6518def30af7112495647b0193ce7d8d99d016e18);相关代码/断言均不在 PR 改动内。这项保持 failed 记录,但不是此次 REQUEST_CHANGES 的理由。本轮没有查询、等待或推断远端 CI。
我的整体评价
这是一项规模合适、应该继续完善的默认 Chat 修复,不需要扩展成新的状态机、协议或 Python owner。我做了相邻 future-facing pass:最有价值的精简是让现有 Session/Turn 生命周期提供唯一派生状态,消除历史占位与当前执行的双重知识;评审-only 未代作者修改。long_horizon 与 user_experience 均存在已复现回归:本轮结束了,用户却不能自然进入下一轮。source/packaged 正向 smoke 通过不覆盖这个反例;修复取消占位、终态及新 Session 的边界并验证完整续聊后再重审。没有权限或 actor-lifecycle 扩张,也没有 default-off 声明;这是既有默认行为的变更。本轮不修复、合并或升级此 PR。
…ser-running-turn Signed-off-by: song <22676124+songoow@users.noreply.github.com>
Leaving a conversation cancels its Turn recovery, but the recovery's pending placeholder was never settled. Returning started a new recovery with its own placeholder, so after the Turn completed the stale one stayed pending and the composer guard kept Send disabled. A cancelled recovery now removes its placeholder; the next recovery streams the same Turn from the start into a fresh one. The composer guard also only counts pending replies from the current Session, so a new Session never inherits an older Session's placeholder. The chat-recovery scenario leaves the Goal during a recovered Turn, checks the other Goal's composer is free, returns before completion and requires Send to reopen when the Turn completes. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
已处理 exact-head 根因修复:在恢复 effect 里,如果恢复被取消(切走 Goal), guard 收紧: 验证(Node 24.21.0):扩展了 还有一个没有处理的相邻问题:如果离开期间 Turn 已经完成,回到原 Goal 时不会发起新恢复;因为 history 只在消息为空时加载,所以最终回答要 reload 后才出现。base 在这种情况下留下的是一个永远 pending 的气泡,现在是既不锁 composer 也不显示陈旧占位。补齐回答需要 history 合并,超出本 PR 范围。 |
cocolord
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES — [P2] exact head 4ee26f4 修复了上轮“切走再返回后旧 pending 永久锁住 composer”的问题,但 409 的权威 active_turn_id 仍只被翻译成文案,没有进入现有 Turn recovery/state owner。实测这条竞态下 Send 仍可用、没有调整/中断控件,第二次点击会再次 POST。
动机
目标本身明确且有实际收益:Chat service 每个 Session 只允许一个 active Turn。页面 reload 或恢复后,如果 composer 只看本页的 sending 标记,就会向正在运行的 Session 再发一个普通 Turn,得到 409 并丢失自然的连续对话体验。合理的结果应该是:当前 Turn 运行时普通发送等待,用户仍能调整或中断本轮;终态到达后继续下一轮;LoopX 的 queue/inbox/steer 通道不被普通发送规则误封。
当前 head 已让 snapshot 能看到的 running Turn 达到这个目标,也修复了上轮跨 Goal 取消恢复留下陈旧 pending 的 blocker。但 PR 还明确承诺“即使 409 先到,也显示同一状态”。这一分支目前只有同样的文字,没有同样的状态和能力,所以 scoped outcome 仍未完整交付。
改动思路
实现复用了既有 Session/Turn 身份与 pending reply 投影:PersonalWorkspacePage 从当前 conversationSessionId 对应的 pending message 推导 conversationTurnRunning,再统一驱动普通 Send、普通 quick prompts 和状态提示。LoopX active delivery 先走原 queue/inbox/steer 分支,因此不受普通 one-Turn guard 影响。
恢复侧由 dashboard-page 的现有 effect 读取 Session snapshot 的 active_turn_id,创建带 sourceSessionId/sourceTurnId 的 pending assistant message,并调用 resumeChatTurnStreaming。新 head 在 effect 被 Goal 切换取消时删除该 effect 自己创建的 placeholder;返回原 Goal 后,由新 recovery 重新承接同一个 Turn。这个 owner 方向正确,没有新增第二个 busy flag。
缺口在普通 POST 的 catch:服务端已经用 409 payload 告知 active_turn_id,但代码只把失败气泡文本替换为 composer.turnRunning,随后 finally 仍清除 activeTurnIds 并把 runtime binding 写成 ready。由于失败气泡是 pending=false 且没有 Session/Turn identity,composer guard、Adjust 和 Interrupt 都无法消费它。
具体改动
完整 PR 是五个文件、+69/-7:personal-workspace-page.tsx 增加当前 Session 的 running 推导、普通发送/quick prompt guard 与状态行;dashboard-page.tsx 增加 409 本地化分支和 cancelled recovery placeholder 清理;i18n.tsx 增加中英文提示;CSS 增加一行 muted status 样式;chat-recovery.mjs 增加 reload 与 Goal A→B→A 的恢复验证。没有新增后端协议、权限或持久化 schema,也不改变首屏。
关键代码讲解
- conversationTurnRunning / composerBlocked 以 pending、sourceTurnId 和 sourceSessionId === conversationSessionId 精确推导普通会话繁忙状态;它驱动按钮禁用、quick prompts、sendMessage 早退和 status 文案。它不使用 substring,但前提是上游完整投影 identity。
- running-Turn recovery effect 从 Session snapshot 读取 active_turn_id,登记 recoveringTurnKeys、activeTurnIds 与 runtime binding,再创建 pending placeholder。新 finally 中的 cancelled 分支按 streamingMessageId 删除旧占位,修复了上轮迟到占位污染。
- sendManagerQuestion 的 failureMessage 在 payloadError.active_turn_id 存在时只选择本地化文字;该对象仍是 pending=false,未写 sourceTurnId/sourceSessionId。随后的 finally 无条件 delete activeTurnIds 并记录 ready,恰好使新 guard 看不到服务端报告的 running Turn。
- chat-recovery 新场景验证 snapshot-known path:A 运行时禁用、切到 B 不继承、返回 A 重新恢复、终态后开放,并保证没有额外 Turn;这个场景有价值,但没有让 POST 本身返回 active_turn_id,因此覆盖不到上述竞态。
对主干的风险
[P2] 触发条件是正常并发边界:页面最后一次 snapshot 尚未显示 active Turn,但另一个页面或 actor 已让 Session busy,此时普通 POST 返回服务端现有的 409 + active_turn_id。代码在 dashboard-page.tsx 的 failureMessage 分支把它渲染成“本轮回答进行中。可在回答里调整或中断本轮”,却没有建立 pending/identity/recovery。
我用实际 React 页面、生产 fetch 路径和服务端相同 payload 做了隔离 headless Chrome 反例。第一次 409 后观察为 sendDisabled=false、composerStatusCount=0、adjustCount=0、interruptCount=0;再次点击后 POST 计数从 1 变成 2。原输入也已经在 sendMessage 调 callback 前从 composer 清空,而 dashboard callback 吞掉错误并正常返回,因而不会恢复草稿。结果既没有阻止本 PR 要消除的重复提交,也让“在回答里调整或中断”的新文案成为不可执行指导。
最小修复是把 active_turn_id 作为权威 receipt 交给现有 recovery/runtime owner:绑定当前 Session/Turn、创建或恢复带 identity 的 pending reply、保持普通 send blocked,并在 terminal readback 后结算;同时保留未被接受的用户草稿。不要只加另一个 busy boolean。请扩展 chat-recovery,让普通 POST 在 snapshot 更新前返回真实 409 payload,并断言只有一次 POST、草稿保留、当前 Turn 恢复/控件可见、终态前 Send 一直禁用。
其他路径证据是正向的:development 与 packaged chat-recovery 均通过,loopx-mode 通过,personal-workspace contract 通过,完整 dashboard build 通过且只有既有 chunk-size warning。risk-based premerge 的 3 项 direct、4 项 catalog 和 1 项 public-boundary check 全部通过,0 failure、0 manual hold。
远端 required checks 仍有四个 Python shard、aggregate pytest 和 merge-gate 失败;六个失败均在 quota projection/selection 与 scheduler acknowledgement,当前 head 本地逐项复现,且此前相同 immutable base 的同组对照已有一致失败签名。本 PR 只改 dashboard 与 browser smoke,因此这些红灯属于 pre-existing unrelated merge-readiness hold,不是本次 REQUEST_CHANGES 的原因。
语义与 CI 对齐
服务端 active_turn_id 是“当前 Session 已有运行中 Turn”的 typed authority;pending reply 是 UI projection。当前 snapshot producer 与 guard 对齐,但 409 producer 只映射文案,没有映射状态。新文案承诺用户可“调整或中断”,而 MessageActivity 只有拿到 sourceTurnId 才渲染这两个操作,实际反例为零个控件。这是具体的 state/authority semantic violation,不是一般性的测试建议。domain wording 保持中性,也没有把机器义务误称 guidance;问题是可见 guidance 没有可执行状态支持。
我的整体评价
当前改动规模合适、方向也基本正确:已覆盖的 reload 和 Goal 切换路径明显改善 long-horizon 与用户体验,取消占位修复能关闭上轮 blocker,LoopX delivery exemption 也保持了既有语义。没有必要扩张为新协议或新状态机,最有价值的 future-facing 修复就是让现有 Session/Turn owner 统一消费 snapshot 与 409 两种 active_turn_id 来源。
但 exact head 仍在其核心竞态上允许重复发送,并显示不存在的调整/中断入口;这是用户可以直接遇到、且现有正向 smoke 会漏掉的 P2。修复上述 receipt adoption 与 draft retention,并加入能杀死当前实现的浏览器回归后再复审。远端独立 CI 红灯仍需在 merge 前解决,但不改变这个 blocker 的归因。
English verdict: REQUEST_CHANGES — exact head 4ee26f4 fixes the prior cancelled-recovery placeholder lock and passes source/packaged recovery, LoopX-mode, build, contract, and premerge checks. However, when the ordinary Turn POST itself returns the server's 409 active_turn_id receipt, the UI only translates the error: Send stays enabled, no adjust/interrupt controls exist, the draft is not restored, and a second click issues a second POST. Route that receipt through the existing Session/Turn recovery owner, retain the input, and add this browser regression. The unrelated quota/control-plane CI failures remain a separate merge-readiness hold.
When another page started a Turn after this page's last snapshot, the ordinary send was the first to learn of it: the service answered 409 with active_turn_id. The page only translated that into text, then marked the Session ready, so Send stayed enabled, the reply had no adjust/interrupt controls, the draft was gone, and a second click posted again. A 409 carrying active_turn_id now withdraws the unaccepted message and its placeholder, asks the existing Turn recovery effect to re-read the Session and adopt the running Turn (pending reply with Session/Turn identity, runtime binding and controls), and rejects the send so the composer keeps the draft and images. The chat-recovery scenario makes the Session busy behind the page's back, answers the POST with the service's 409 payload, and requires one POST, a kept draft, visible adjust/interrupt controls, Send disabled until the Turn completes, and no stale user bubble. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
已处理 exact-head 修复:普通发送收到带
LoopX mode 的路径不走这个分支。原来只把 409 翻译成文字的分支已经删除,不会再出现。 验证:
用上一个 head 的源码跑这个场景,会在等待"本轮回答进行中"出现处超时。 远端 6 个 quota/control-plane 失败在最新 main 上同样复现,和本 PR 无关。 |
huangruiteng
left a comment
There was a problem hiding this comment.
English verdict: APPROVE — reviewed exact head 9741362697602d7ff4e555ca51f0f58ccad1325e. The previously reported cancelled-recovery and 409 state/draft blockers are fixed. Source and packaged browser recovery, LoopX-mode parity, dashboard build and focused contract pass. Remote CI was not queried; merge readiness is separate.
动机
普通 Chat 的服务端 Session 一次只能运行一个 Turn。之前页面 reload 后可能仍让普通 Send 可用;上轮修复了切换 Goal 留下旧 pending 占位,却在另一条真实竞态上只把 409 的 active_turn_id 译成提示文字,草稿丢失、没有调整/中断控件,第二次点击仍发 POST。当前 PR 的可观察目标是:当前 Turn 未结束时等它完成,同时保留草稿和控制权;终态能继续下一轮;LoopX 的 queue/inbox/steer 不被普通 Chat guard 误封。当前 head 的浏览器负例已覆盖这些路径。
改动思路
实现继续让服务端 Session/Turn identity 作权威,UI 的 pending 回复只作投影,没有增设第二个 busy flag。conversationTurnRunning 仅看当前 Session 的 pending Turn;普通发送和快捷入口共用 composerBlocked。409 分支撤回未被服务端接受的本地消息,触发既有 recovery effect 重新读取 Session 并承接同一个 active Turn;callback 抛错让 composer catch 复原草稿。LoopX active-delivery 分支仍在普通 guard 之前处理,权限和消息路由没有变。
具体改动
完整 diff 为五个文件、+123/-8:PersonalWorkspace 的 Send/快捷入口 guard 与状态行、双语文案和样式、Dashboard 409 回收与恢复触发,以及既有 browser recovery 场景的扩展。没有服务端协议、持久 schema 或新 provider。
关键代码讲解
sendMessage先保留 LoopX queue/inbox/steer,随后在普通 Chat 的当前 Turn 运行时早退。被 409 拒绝时,catch恢复原文字与附件;按钮禁用、键盘和 quick prompt 使用同一运行状态。- 409 处理 不再留下表示“失败”的终态气泡代替运行中 Turn;它删掉未接受的用户/助手占位、请求恢复,并把可读错误交给 composer。恢复 effect 重新读取当前 Session,不凭 409 文案猜状态。
- 取消恢复清理 只退休本次 recovery 创建的占位;A→B→A 回到同一 Turn 时,新 effect 可以承接,终态后不会被旧 pending 永久锁住。
对主干的风险
最强反例是 snapshot 尚未显示 busy,但普通 POST 收到服务端 409 + active_turn_id。我在真实 React/fetch 浏览器路径跑了 development 与 packaged chat-recovery:只发一次 POST,草稿保留,调整/中断控件跟随恢复的 Turn,终态前普通 Send 禁用,终态后继续可发送。A→B→A 取消/恢复场景也通过;单独 loopx-mode smoke 验证活动投递例外未退化。Dashboard npm run build(TS、Vite、packaged assets)及 PersonalWorkspace contract 1/1 通过,git diff --check 通过。build 有非阻断 chunk-size warning。测试使用隔离合成后端,不证明 live provider 的延迟;按 Goal 策略未查询或等待远端 CI,红 CI 若与本 PR 无关也不能单独构成 REQUEST_CHANGES。
语义与 CI 对齐
active_turn_id 是服务端 typed 状态;pending reply 是 UI 投影。当前 409 与正常 snapshot 两个 producer 都落到同一 recovery owner,文案所说“调整或中断”已有对应控件。默认普通 Chat 行为确实变为“等待当前 Turn”,PR 与 browser smoke 已披露;没有 substring 状态分类、default-off 隐式激活或把机器义务写成指导。LoopX 投递仍走单独的既有入口。
我的整体评价
本轮是对 9741362697602d7ff4e555ca51f0f58ccad1325e 的完整重审,不继承旧结论。它以小而可回滚的 UI/recovery 改动关闭了先前两个 blocker:连续对话不再因旧占位卡住,也不再因 409 重复提交并丢草稿。long_horizon 与用户体验均有实际改善。相邻 future-facing 检查的结论是保持 Session/Turn 单一身份 owner,不需再拆出状态机或额外 busy flag。基于当前完整 diff 和正反浏览器证据,我批准这条 PR;合并权限和 required checks 由独立门禁决定。
cocolord
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES — [P2] exact head 9741362697602d7ff4e555ca51f0f58ccad1325e 已修复上一轮 409 receipt 未进入 Turn recovery、草稿丢失和控件缺失的问题,但 recovery handoff 还不是原子的:慢 Session 重读时,旧 send 的 finally 会先解除 sending 并写回 ready,在 recovery effect 建立 pending Turn 前,Send 会短暂重新可用。我在真实 Chrome 页面中延迟这次重读并二次点击,观察到 POST 计数从 1 变成 2。
动机
目标是让 Personal Workspace 严格遵守 Chat service 的“一 个 Session 同时只有一个 active Turn”契约。无论 running Turn 是在 reload 时从 snapshot 发现,还是普通发送第一次通过 409 + active_turn_id 得知,用户都应该看到同一份可信状态:未被接受的草稿保留、普通发送持续阻塞、当前 Turn 可调整或中断,并在终态后恢复下一轮。这个目标对多窗口、重连和并发 actor 都有直接用户价值,也能避免反复发送必然被服务端拒绝的请求。
当前 head 比上一轮明显进步:409 分支会撤回未被接受的 user/placeholder 消息、向上抛出 ChatApiError 让 composer 恢复草稿,并触发现有 recovery effect 重新读取 Session;恢复完成后,running hint、Adjust/Interrupt、禁用 Send 与终态解锁都正确。剩余问题发生在 409 receipt 与 recovery effect 接管之间,因此核心结果仍未完整闭合。
改动思路
整体方向是复用既有 Session/Turn owner,而不是引入另一套字符串或持久状态。PersonalWorkspacePage 从当前 Session 对应的 pending message 精确推导 conversationTurnRunning,统一约束发送按钮、普通 quick prompts 和 sendMessage 入口;LoopX queue/inbox/steer 仍走原专用路径,不被普通 one-Turn guard 误封。dashboard-page 的 recovery effect 继续负责把权威 active_turn_id 投影成带 sourceSessionId/sourceTurnId 的 pending reply、runtime binding 和恢复流。
新 409 分支也选择了这个 owner:删除乐观消息、递增 turnRecoveryRequest、让 effect 读取 Session 并恢复 Turn。问题是该触发只是异步通知;catch 随后抛错,原 send 的 finally 仍无条件删除 active Turn、把 binding 写成 ready 并清除 sendingContextId。React 要到下一次 effect 及网络读取后才建立 pending identity,因此 machine-enforced guard 在这段过渡期没有权威状态可消费。
具体改动
完整 PR 为 5 个文件、+123/-8。i18n.tsx 增加中英文 running 提示;personal-workspace-page.tsx 增加当前 Session 的 running 推导、普通发送/quick-prompt guard、状态行和发送前早退;CSS 增加相应 muted status 样式;dashboard-page.tsx 修复取消 recovery 后遗留 placeholder,并在 409 时撤回未接受消息、触发恢复且重新抛错;chat-recovery.mjs 覆盖 reload、跨 Goal 取消/返回,以及 409 后最终 adoption。没有新增后端协议、权限或持久化 schema。
关键代码讲解
conversationTurnRunning/composerBlocked使用 pending、sourceTurnId与sourceSessionId === conversationSessionId的 typed identity 判断普通会话繁忙;它正确覆盖按钮、Enter/sendMessage早退和普通 quick prompts,同时保留 LoopX delivery exemption。- running-Turn recovery effect 从 Session snapshot 读取
active_turn_id,登记 recovery key、active Turn 和 runtime binding,创建带 identity 的 pending reply,再调用resumeChatTurnStreaming;取消时按 message id 清理自己的 placeholder,避免旧占位永久锁住 composer。 - 409 catch 现在识别
payloadError.active_turn_id,撤回未被服务端接受的两条乐观消息,递增turnRecoveryRequest,并重新抛出错误;这让上层恢复草稿,也最终复用了 recovery owner。 - 原 send 的
finally仍会立即activeTurnIds.delete、记录status: ready并清除 sending。因为turnRecoveryRequest只会在 render 后启动异步 effect,这里形成一个没有 pending identity 的可发送窗口。 - 新
composer-running-turn-409场景验证 recovery 完成后的状态,并在按钮已禁用后用 forced click 确认不会 POST;它没有延迟 recovery GET,因此未覆盖 handoff 前的 enabled 窗口。
对主干的风险
[P2] 触发条件是正常的慢网络/多窗口竞态:普通 POST 收到 409 + active_turn_id 后,Session list/snapshot 的恢复读取尚未返回。此时外层 composer 已因 reject 恢复草稿,但旧 send 的 finally 已把 sending 清掉,而 conversationTurnRunning 仍为 false;用户可以再次点击并产生第二个必然失败的 POST。服务端 one-Turn invariant 会阻止第二个 Turn 真正创建,所以不会造成双执行,但会制造重复请求、重复错误和不可预测的交互窗口。
我在 exact head 的现有 409 browser case 中只加入一个反证条件:第一次拒绝后阻塞 GET /api/chat/sessions,等待按钮重新 enabled 就点击一次,再释放恢复读取。真实 headless Chrome 观察到 rejectedPosts === 2;原测试在不延迟读取时通过,说明它只验证了最终状态,没有验证过渡期。最小修复是让 409 receipt 到 recovery owner 的交接保持 fail-closed:例如在 catch 中立即用现有 Session/Turn identity 建立可被 composer 消费的 recovering/pending 状态,或让 recovery owner 返回明确接管 receipt 后再解除 sending;同时避免旧 finally 写回 ready 覆盖该 ownership。不要只增加与 Session/Turn 分离的永久 busy boolean。
回归应在普通 POST 返回 409 后人为挂起 Session 重读,并在释放前断言 Send 和普通 quick prompts 始终禁用、二次点击不增加 POST;释放后再断言草稿仍在、当前 Turn 的 Adjust/Interrupt 可见、未接受消息不在 timeline、终态后恢复发送。这个测试应同时跑 development 与 packaged bundle。
其余验证是正向的:现有 chat-recovery 在 development 与 packaged Chrome 中通过;同一新增场景移植到旧 head 会在 running hint 处失败,证明本次修复确实关闭了上一轮缺陷;loopx-mode、personal-workspace contract、完整 Dashboard/Chat 生产构建通过。补齐站点与 Playwright 前置后,risk-based premerge 的 3 个 direct checks 和 6 个选择检查全部通过,0 failure、0 warning、0 manual hold,5 个公开候选文件边界扫描干净。
远端 required CI 仍红:6 个 quota/global-gate/scheduler 失败与固定 base 的既有签名一致;合并引用另有 2 个 registry-census 失败来自较新的 main 父提交,相关路径不在本 PR diff,且在 exact head 上独立通过。这些红灯仍单独阻塞 merge readiness,但不是本次 REQUEST_CHANGES 的原因。
语义与 CI 对齐
本 PR 没有 substring denylist;它复用 typed active_turn_id、Session id 和 Turn id,领域措辞保持中性,也没有把机器义务仅写成“建议”。但当前过渡期违反了这套 typed state contract:服务端已经明确报告 running Turn,UI 的机器 guard 却暂时回到 ready。可见文案、控件和按钮在最终状态对齐,在 recovery handoff 窗口仍不对齐。修复应继续收敛到现有 Session/Turn owner,不需要新增第二套 authority。
我的整体评价
REQUEST_CHANGES。这个 head 的收益已被严格验证:上一轮 blocker 确实修复,草稿保留、未接受消息撤回、最终 Turn adoption、控制按钮、跨 Goal 清理和 LoopX-mode 隔离均成立;改动范围也围绕同一 UI lifecycle,规模合理。future-facing pass 应集中在把 “snapshot 发现” 与 “409 receipt 发现” 的 ownership handoff 做成同一条原子、可验证的状态路径,而不是扩张抽象。
不过 exact head 仍允许 recovery 读取期间重复 POST,直接违背 PR 要保证的持续禁用语义。请关闭这一窗口并加入可控延迟回归后再复审。即使修复后获得 approval,也不等于 merge authority;required CI 和 branch freshness 仍需单独满足。
English verdict: REQUEST_CHANGES on exact head 9741362697602d7ff4e555ca51f0f58ccad1325e. The previous 409 adoption blocker is substantially fixed and the source/packaged Chrome scenarios, LoopX-mode isolation, production build, contract smoke, premerge suite, and public-boundary scan pass. However, delaying the post-409 Session read exposes a handoff gap: the original send's finally clears sending and records ready before recovery creates a pending Turn, so a second click issues a second POST. Keep the UI fail-closed from the authoritative 409 receipt until the existing Session/Turn recovery owner has taken over, and add this delayed-read regression. Unrelated required-CI failures remain a separate merge-readiness hold.
cocolord
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES — [P2] exact head 9741362697602d7ff4e555ca51f0f58ccad1325e 已修复上一轮 409 receipt 未进入 Turn recovery、草稿丢失和控件缺失的问题,但 recovery handoff 还不是原子的:慢 Session 重读时,旧 send 的 finally 会先解除 sending 并写回 ready,在 recovery effect 建立 pending Turn 前,Send 会短暂重新可用。我在真实 Chrome 页面中延迟这次重读并二次点击,观察到 POST 计数从 1 变成 2。
动机
目标是让 Personal Workspace 严格遵守 Chat service 的“一 个 Session 同时只有一个 active Turn”契约。无论 running Turn 是在 reload 时从 snapshot 发现,还是普通发送第一次通过 409 + active_turn_id 得知,用户都应该看到同一份可信状态:未被接受的草稿保留、普通发送持续阻塞、当前 Turn 可调整或中断,并在终态后恢复下一轮。这个目标对多窗口、重连和并发 actor 都有直接用户价值,也能避免反复发送必然被服务端拒绝的请求。
当前 head 比上一轮明显进步:409 分支会撤回未被接受的 user/placeholder 消息、向上抛出 ChatApiError 让 composer 恢复草稿,并触发现有 recovery effect 重新读取 Session;恢复完成后,running hint、Adjust/Interrupt、禁用 Send 与终态解锁都正确。剩余问题发生在 409 receipt 与 recovery effect 接管之间,因此核心结果仍未完整闭合。
改动思路
整体方向是复用既有 Session/Turn owner,而不是引入另一套字符串或持久状态。PersonalWorkspacePage 从当前 Session 对应的 pending message 精确推导 conversationTurnRunning,统一约束发送按钮、普通 quick prompts 和 sendMessage 入口;LoopX queue/inbox/steer 仍走原专用路径,不被普通 one-Turn guard 误封。dashboard-page 的 recovery effect 继续负责把权威 active_turn_id 投影成带 sourceSessionId/sourceTurnId 的 pending reply、runtime binding 和恢复流。
新 409 分支也选择了这个 owner:删除乐观消息、递增 turnRecoveryRequest、让 effect 读取 Session 并恢复 Turn。问题是该触发只是异步通知;catch 随后抛错,原 send 的 finally 仍无条件删除 active Turn、把 binding 写成 ready 并清除 sendingContextId。React 要到下一次 effect 及网络读取后才建立 pending identity,因此 machine-enforced guard 在这段过渡期没有权威状态可消费。
具体改动
完整 PR 为 5 个文件、+123/-8。i18n.tsx 增加中英文 running 提示;personal-workspace-page.tsx 增加当前 Session 的 running 推导、普通发送/quick-prompt guard、状态行和发送前早退;CSS 增加相应 muted status 样式;dashboard-page.tsx 修复取消 recovery 后遗留 placeholder,并在 409 时撤回未接受消息、触发恢复且重新抛错;chat-recovery.mjs 覆盖 reload、跨 Goal 取消/返回,以及 409 后最终 adoption。没有新增后端协议、权限或持久化 schema。
关键代码讲解
conversationTurnRunning/composerBlocked使用 pending、sourceTurnId与sourceSessionId === conversationSessionId的 typed identity 判断普通会话繁忙;它正确覆盖按钮、Enter/sendMessage早退和普通 quick prompts,同时保留 LoopX delivery exemption。- running-Turn recovery effect 从 Session snapshot 读取
active_turn_id,登记 recovery key、active Turn 和 runtime binding,创建带 identity 的 pending reply,再调用resumeChatTurnStreaming;取消时按 message id 清理自己的 placeholder,避免旧占位永久锁住 composer。 - 409 catch 现在识别
payloadError.active_turn_id,撤回未被服务端接受的两条乐观消息,递增turnRecoveryRequest,并重新抛出错误;这让上层恢复草稿,也最终复用了 recovery owner。 - 原 send 的
finally仍会立即activeTurnIds.delete、记录status: ready并清除 sending。因为turnRecoveryRequest只会在 render 后启动异步 effect,这里形成一个没有 pending identity 的可发送窗口。 - 新
composer-running-turn-409场景验证 recovery 完成后的状态,并在按钮已禁用后用 forced click 确认不会 POST;它没有延迟 recovery GET,因此未覆盖 handoff 前的 enabled 窗口。
对主干的风险
[P2] 触发条件是正常的慢网络/多窗口竞态:普通 POST 收到 409 + active_turn_id 后,Session list/snapshot 的恢复读取尚未返回。此时外层 composer 已因 reject 恢复草稿,但旧 send 的 finally 已把 sending 清掉,而 conversationTurnRunning 仍为 false;用户可以再次点击并产生第二个必然失败的 POST。服务端 one-Turn invariant 会阻止第二个 Turn 真正创建,所以不会造成双执行,但会制造重复请求、重复错误和不可预测的交互窗口。
我在 exact head 的现有 409 browser case 中只加入一个反证条件:第一次拒绝后阻塞 GET /api/chat/sessions,等待按钮重新 enabled 就点击一次,再释放恢复读取。真实 headless Chrome 观察到 rejectedPosts === 2;原测试在不延迟读取时通过,说明它只验证了最终状态,没有验证过渡期。最小修复是让 409 receipt 到 recovery owner 的交接保持 fail-closed:例如在 catch 中立即用现有 Session/Turn identity 建立可被 composer 消费的 recovering/pending 状态,或让 recovery owner 返回明确接管 receipt 后再解除 sending;同时避免旧 finally 写回 ready 覆盖该 ownership。不要只增加与 Session/Turn 分离的永久 busy boolean。
回归应在普通 POST 返回 409 后人为挂起 Session 重读,并在释放前断言 Send 和普通 quick prompts 始终禁用、二次点击不增加 POST;释放后再断言草稿仍在、当前 Turn 的 Adjust/Interrupt 可见、未接受消息不在 timeline、终态后恢复发送。这个测试应同时跑 development 与 packaged bundle。
其余验证是正向的:现有 chat-recovery 在 development 与 packaged Chrome 中通过;同一新增场景移植到旧 head 会在 running hint 处失败,证明本次修复确实关闭了上一轮缺陷;loopx-mode、personal-workspace contract、完整 Dashboard/Chat 生产构建通过。补齐站点与 Playwright 前置后,risk-based premerge 的 3 个 direct checks 和 6 个选择检查全部通过,0 failure、0 warning、0 manual hold,5 个公开候选文件边界扫描干净。
远端 required CI 仍红:6 个 quota/global-gate/scheduler 失败与固定 base 的既有签名一致;合并引用另有 2 个 registry-census 失败来自较新的 main 父提交,相关路径不在本 PR diff,且在 exact head 上独立通过。这些红灯仍单独阻塞 merge readiness,但不是本次 REQUEST_CHANGES 的原因。
语义与 CI 对齐
本 PR 没有 substring denylist;它复用 typed active_turn_id、Session id 和 Turn id,领域措辞保持中性,也没有把机器义务仅写成“建议”。但当前过渡期违反了这套 typed state contract:服务端已经明确报告 running Turn,UI 的机器 guard 却暂时回到 ready。可见文案、控件和按钮在最终状态对齐,在 recovery handoff 窗口仍不对齐。修复应继续收敛到现有 Session/Turn owner,不需要新增第二套 authority。
我的整体评价
REQUEST_CHANGES。这个 head 的收益已被严格验证:上一轮 blocker 确实修复,草稿保留、未接受消息撤回、最终 Turn adoption、控制按钮、跨 Goal 清理和 LoopX-mode 隔离均成立;改动范围也围绕同一 UI lifecycle,规模合理。future-facing pass 应集中在把 “snapshot 发现” 与 “409 receipt 发现” 的 ownership handoff 做成同一条原子、可验证的状态路径,而不是扩张抽象。
不过 exact head 仍允许 recovery 读取期间重复 POST,直接违背 PR 要保证的持续禁用语义。请关闭这一窗口并加入可控延迟回归后再复审。即使修复后获得 approval,也不等于 merge authority;required CI 和 branch freshness 仍需单独满足。
English verdict: REQUEST_CHANGES on exact head 9741362697602d7ff4e555ca51f0f58ccad1325e. The previous 409 adoption blocker is substantially fixed and the source/packaged Chrome scenarios, LoopX-mode isolation, production build, contract smoke, premerge suite, and public-boundary scan pass. However, delaying the post-409 Session read exposes a handoff gap: the original send's finally clears sending and records ready before recovery creates a pending Turn, so a second click issues a second POST. Keep the UI fail-closed from the authoritative 409 receipt until the existing Session/Turn recovery owner has taken over, and add this delayed-read regression. Unrelated required-CI failures remain a separate merge-readiness hold.
Duplicate submission from review readback; superseded by the identical exact-head review 5350048539.
After a 409 reported a running Turn, the send withdrew its reply and asked the recovery effect to re-read the Session. Until that read came back no pending reply carried the Turn, the send's finally had cleared sending, and Send was usable again, so a second click posted again. The 409 now turns the send's reply into the running Turn's pending reply at once (Session and Turn identity, adjust/interrupt controls), so the composer guard holds from the receipt on. The recovery effect adopts that reply instead of appending another, and a run that finds no such Turn retires it so the composer reopens. The chat-recovery scenario holds the post-409 Session re-read and requires Send and the sending quick prompts to stay disabled, a forced click to post nothing, and exactly one pending reply after adoption. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
已处理 exact-head 修复:交接改为从收到 409 起就保持关闭,没有新增与 Session/Turn 分离的 busy flag。
验证: 用上一个 head 的源码跑,失败在 "Send reopened before the recovery adopted the reported Turn"。development 和 packaged 的 |
huangruiteng
left a comment
There was a problem hiding this comment.
动机
这条 PR 想修复个人工作区在同一 Chat Session 的 Turn 尚在运行时仍可发送普通消息、得到服务端 409 的用户死路。前次评审在 9741362697602d7ff4e555ca51f0f58ccad1325e 发现:409 后等 Session 读取结束才关 composer,留下竞态。当前 6c56202029989c9328a3efc034cf22c81ec114c7 立即把服务端的 active_turn_id 变成待恢复的回复,确实缩短了这个窗口;但真实恢复路径遇到临时读取错误时仍会重新放开发送,因此本轮不能批准。
改动思路
PersonalGoalHome 在通常发送路径收到携带 active_turn_id 的 409 后,撤回未被接受的用户消息,将原回复标成该 Session/Turn 的 pending handoff,并触发 Session 恢复 effect。PersonalWorkspacePage 从当前 Session 中带 sourceTurnId 的 pending 回复派生 conversationTurnRunning,禁用发送和普通快捷提问,显示等待提示;LoopX mode 仍走既有 Turn 队列,不纳入普通发送禁用。这复用服务端的一 Session 一 Turn 权威事实,界面状态不应自行宣告该 Turn 结束。正常延迟 GET 能接管原回复,但读取失败与已结束不是同一状态,当前代码把它们折叠了。
具体改动
关键代码讲解
personal-workspace-page.tsx的conversationTurnRunning(约 line 1002)从当前 Session 的 pending assistant 消息派生,composerBlocked同时控制 Send 和多数会发送消息的快捷提问;安排 monitor、创建 Goal 等非发送动作不被误封。新增的 status 文案和一行 CSS 给出可见反馈。dashboard-page.tsx的turnHandoffs(约 line 1419)保存 409 所报告 Turn 与原回复的暂态关联;sendManagerMessage的 409 分支(约 line 2276)撤回未被接受的用户消息、保留草稿并立即关闭 composer,再要求恢复。- 同文件恢复 effect(约 line 1698)在成功读到相同
activeTurnId时复用 handoff 回复而非重复创建;chat-recovery.mjs新增延迟读取、接管后继续流式恢复的浏览器场景,旧的正常路径因此通过。 - 缺口在恢复 effect 的外层
finally(约 line 1823):只要没有 cancel,就删除所有未接管 handoff,注释将其解释为 Turn 已结束或属于其它 Session。外层 Session 读取异常同样进入此finally,但没有得到可作上述判断的权威 snapshot。
我在准确 head 上跑过 chat-recovery 与 loopx-mode 浏览器场景、tsc --noEmit、桌面 Vite build 和完整 diff 检查,均通过。随后仅临时修改测试 fixture,让 409 之后的 Session GET 返回一次 503 并断言 Send 仍禁用;该真实 React 页面回归在精确 head 失败,报出 Send 被重新开放。临时测试修改已撤销,生产源文件未改。
对主干的风险
这是 PR 自身的恢复状态回归,不是无关红 CI:触发条件为服务端 409 明确提供仍运行的 active_turn_id,紧随其后 /api/chat/sessions 读取暂时失败。外层 catch 没有可用 Session/Turn 结论,外层 finally 却移除 handoff 回复;PersonalWorkspacePage 于是得不到 pending Turn,Send 变成可点,再次重复 409 或让用户误以为可以开始新回合。正常 delayed GET 的新测试与 loopx-mode 均通过,恰好说明现有正例没有覆盖失败后重试。最小修复是仅在成功取得权威 Session snapshot 且确认 Turn 已终止或确属不同 Session 时退休 handoff;读取失败时保留阻塞及可操作的重试/恢复提示,再加一次 503→重试→终态的浏览器回归。不要以本地计时器或另一个散布的 busy boolean 猜测服务端状态。
当前改动只触及工作区 Chat UI/浏览器烟测,没有引入新公开权限、持久协议或通用 LoopX 控制面规则;typed Turn id 比文案 substring 可靠,普通消息与 LoopX 队列分流也是有意的。远端 CI 按本 Goal wait_for_ci=false 未查询或等待;即便远端有无关红灯,也不是本次 REQUEST_CHANGES 的理由。未来相关的小幅保形重构可考虑把“权威已终止”和“读取失败”用明确 union 结果表达,以免下一次恢复调整再次混淆。
我的整体评价
当前 head 对前次反馈做出了有价值的立即 handoff 修复,正常延迟读取也能避免双回复;但用户最需要的网络/重连恢复仍会在一次 GET 失败时失去保护。long_horizon 是未解决的重试循环风险,user_experience 是可复现的错误开放 Send;182 行新增代码的规模本身不是问题,问题是退休条件没有对应服务端证据。结论是 REQUEST_CHANGES,修复上述失败路径并以同一浏览器场景复测后再按新 head 复审;此结论不涉及合并授权。
English verdict: REQUEST_CHANGES - exact head 6c56202029989c9328a3efc034cf22c81ec114c7 drops the 409 Turn handoff after a transient Session GET failure, re-enabling Send without authoritative completion. Existing recovery and LoopX-mode browser scenarios, TypeScript and build pass; the added 503-after-409 counterexample fails as expected. Remote CI was not polled.
|
补充 本 head 已发布的 REQUEST_CHANGES review:我独立复核了完整 PR;正常延迟 GET 的 Send guard、development/packaged [P2] 409 交接期间,界面已显示的“调整本轮/中断本轮”仍不可用。 在 exact head 我在现有 409 浏览器场景挂起恢复的 English addendum: The already-blocking review covers the failed-read reopening. A separate exact-head browser probe shows that controls displayed during the 409 handoff reject the still-running Turn until recovery has rebound action authority. |
The recovery effect's outer finally retired an unadopted 409 handoff on any exit that was not a cancel, including a failed Session read. A read that failed says nothing about the reported Turn, so Send reopened while that Turn could still run. Retire the handoff only after an authoritative answer: a Session read that did not adopt the Turn, or the service refusing the resume. On any other failure keep the pending reply, with its Turn controls, and re-read the Session with a bounded backoff (1s doubling to 10s). The backoff only schedules re-reads; whether the Turn ended still comes from the snapshot. chat-recovery covers 409, a 503 Session read, a held retry with Send closed, adoption and completion; it fails on the previous head with "Send reopened after the Session read following the 409 failed". Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
Fixed the [P2] from the review on Cause. The recovery effect's outer Change (
Regression ( Validation
Future-facing pass: I considered replacing the flag with a typed read-outcome union, as the review suggests. I deferred it: with one caller, the |
cocolord
left a comment
There was a problem hiding this comment.
REQUEST_CHANGES — [P2] exact head 0dc1946c022f39d733927b7ff3620f84eaf297fa 已关闭上一轮“409 到 recovery 接管前 Send 会重开”的窗口,也在 Session read 失败时保持 pending handoff 并按 bounded backoff 重试;但这段 handoff 新显示的“调整本轮 / 中断本轮”在真正 adoption 前不可用。真实 headless Chrome 中,两项操作都被前端 typed-state guard 拒绝,interrupt/steer API 调用数均为 0。
动机
目标是让 Personal Workspace 严格遵守 Chat service 的“一个 Session 同时只有一个 active Turn”契约。无论 running Turn 是 reload 时从 snapshot 发现,还是普通发送首次通过 409 + active_turn_id 得知,用户都应得到连续且真实的状态:未被接受的草稿保留、普通发送保持阻塞、当前 Turn 可以调整或中断,并在终态后恢复下一轮。这对慢网络、多窗口和临时 Session read 故障都有直接价值。
当前 head 对前两轮问题有明确正向收益:409 receipt 到恢复期间不再产生第二个 POST;Session read 的 503 不再错误清除 handoff,后续读取会以 1s 到 10s 的上限回退重试;恢复成功后沿用同一个 placeholder,终态正确解锁 composer。剩余缺口是可见控制与实际可执行状态不一致,因此整体用户结果仍未闭合。
改动思路
整体方向仍然正确:复用既有 Session/Turn owner 和 typed identity,不引入平行的持久 authority。PersonalWorkspacePage 从当前 Session 对应的 pending message 推导 conversationTurnRunning,统一约束 Send、Enter 与普通 quick prompts;dashboard-page 的 recovery effect 读取权威 active_turn_id,建立 activeTurnIds、runtimeBindings 和恢复流;LoopX queue/inbox/steer 保持独立入口。
本轮补丁让 409 catch 立即把原 placeholder 转成带 sourceSessionId/sourceTurnId 的 pending reply,并写入 turnHandoffs,所以 composer 能在异步 Session read 前保持 fail-closed。读取失败时,effect 保留该 handoff、更新活动文案并调度 bounded retry;读取成功则复用同一 message id 接管 Turn。问题在于 pending reply 同时无条件渲染调整/中断控件,而控制 handler 只承认 recovery adoption 后才设置的 activeTurnIds 和 running runtimeBindings。因此展示状态先于执行 authority,形成一个新的假可用窗口。
具体改动
完整 PR 为 5 个文件、+253/-9。i18n.tsx 增加中英文 running 提示;personal-workspace-page.tsx 增加当前 Session 的 running 推导、普通发送/quick-prompt guard、状态行和发送前早退;CSS 增加对应 muted status 样式;dashboard-page.tsx 处理取消 recovery 的 placeholder、409 handoff、失败重试与最终 adoption;chat-recovery.mjs 覆盖 reload、跨 Goal 切换、409 期间连续禁用、一次 Session read 503、重试、adoption 和终态解锁。没有新增后端协议、权限或持久化 schema。
关键代码讲解
conversationTurnRunning/composerBlocked通过 pending、sourceTurnId和匹配的sourceSessionId对普通会话做精确阻塞;它正确关闭了上一轮 duplicate POST,并且没有误封 LoopX delivery。- 409 catch 现在保留 assistant placeholder,立即写入 Session/Turn identity,并用
turnHandoffs把该 message id 交给 recovery effect;这使现有 composer guard 在 Session read 尚未返回时也能生效。 - recovery effect 在读失败时增加
failedReads、保留 pending reply,并以min(1000 * 2^(n-1), 10000)触发下一次读取;在读到同一 active Turn 后才设置activeTurnIds和 runningruntimeBindings,随后复用 handoff message 恢复 stream。 MessageActivity仅根据 pending message 是否带sourceTurnId决定显示“调整本轮 / 中断本轮”;它不知道 recovery authority 是否已建立。onInterruptConversationTurn与onSteerConversationTurn则要求activeTurnIds[targetContextId] === turnId,且runtimeBindings的 Session/Turn 同时匹配。409 catch 未同步设置这些状态,旧 send 的finally还会写status: ready,所以 adoption 前点击必然在浏览器内被拒绝。
对主干的风险
[P2] 触发条件是普通 POST 收到 409 + active_turn_id 后,Session list/snapshot 读取较慢、被挂起或短暂失败。此时 composer 持续禁用是正确的,pending reply 也已经显示“调整本轮 / 中断本轮”;但点击“中断本轮”会得到“该回合已结束或已被新的回合取代”,提交调整会得到“本轮已结束或已被新的回合取代,追加指令未发送”。两项操作都不会调用后端。
我在 exact head 上运行真实 UI/request 边界:先完成一次普通会话,再让同一 Session 的 POST 返回带 foreign Turn id 的 409,并挂起后续 GET /api/chat/sessions。控件出现后分别执行 interrupt 和 steer,观察到 interruptApiCalls=0、steerApiCalls=0,同时收到上述两个前端错误。项目原生 development chat-recovery 仍通过,说明现有测试只断言这些控件 visible,没有在 adoption 前点击它们。
最小修复是让可见 affordance 与 typed execution authority 同步:要么 409 receipt 立即建立由同一 Session/Turn identity 约束的 recovery binding/active ownership,并确保旧 send 的 finally 不覆盖它;要么在 adoption 前不要把控件呈现为可用。鉴于文案明确承诺用户可调整或中断,更完整的修复是允许它们在 handoff/read-retry 窗口直接针对 receipt 给出的精确 Session/Turn 工作。回归需在挂起 Session read 时真实点击两项控制,并断言请求使用精确 id、回执处理正确、草稿与 pending 状态不丢失。
远端 CI 在本 exact head 仍红:typescript-core (3/3) 的失败是 closed_pipes descendant 生命周期用例,test shard 1/3 是既有 quota scheduler acknowledgement 断言;它们不用于证明本次 UI finding,但 required CI 仍是独立 merge-readiness hold。DCO、dashboard acceptance、chat bundle、kernel static、TypeScript 1/2、Stage 2C、desktop 与 build 等已通过;完整 workflow 尚有 shard 在运行。
语义与 CI 对齐
本 PR 没有 substring denylist,领域措辞保持中性,也没有把机器义务仅称为“建议”。它复用 typed active_turn_id、Session id 与 Turn id;但当前有两套不同步的“running”判定:展示层从 pending message 得出可控制,handler 从 activeTurnIds/runtimeBindings 得出不可控制。这是具体的 typed-state false positive。默认行为变化已在 PR 描述、双语文案和 smoke 中披露,LoopX delivery exemption 也有明确边界。
我的整体评价
REQUEST_CHANGES。这个 exact head 的收益是真实且重要的:上一轮重复 POST blocker 已修,失败读取也不再误判 Turn 结束;改动仍集中于同一 Session/Turn 生命周期,规模与兼容成本合理。future-facing pass 不需要扩张抽象,而应让 409 receipt、pending projection、runtime binding 和控制 handler 共享同一份 typed ownership。
目前 UI 在可能持续到多次 10 秒回退的窗口里向用户展示无法执行的核心控制,且错误地声称回合已结束或被替代,直接违背本 PR 的公开语义。请关闭这个窗口并加入真实点击回归后再复审。即使后续获得 approval,也不等于 merge authority;required CI 与 branch freshness 仍需独立满足。
English verdict: REQUEST_CHANGES on exact head 0dc1946c022f39d733927b7ff3620f84eaf297fa. The prior duplicate-POST gap is fixed, and Session-read failures now retain and retry the handoff. However, before recovery adopts the reported Turn, the UI exposes Adjust/Interrupt controls while both handlers reject locally because activeTurnIds and runtimeBindings still lack that exact Turn; a real headless-Chrome probe observed zero interrupt and steer API calls. Align visible controls with typed ownership and add a delayed-read interaction regression. Required CI is also still red as a separate merge-readiness hold.
From the 409 on, the pending reply shows Adjust and Interrupt for the reported Turn, but both handlers act only when activeTurnIds and the runtime binding name that Turn. Only recovery adoption set them, and the rejected send's finally reset them to ready, so until a Session read returned both controls failed locally with "already ended". The 409 receipt now records the reported Session and Turn as the context's active Turn and running binding, and the rejected send's finally leaves them to the handoff. The handoff carries its Session, so a Session read that finds the Turn ended retires that ownership with the reply. An interrupt during the handoff has no stream to settle the reply, so its receipt settles it and reopens Send. chat-recovery holds every Session read after the 409, then adjusts and interrupts from the handoff reply and checks both requests name the exact Session and Turn; a late read must not revive the Turn. The scenario fails on the previous head at the adjustment. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: song <22676124+songoow@users.noreply.github.com>
…ser-running-turn Signed-off-by: song <22676124+songoow@users.noreply.github.com>
|
Fixed the [P2] from @cocolord's review on Correction. My previous comment said the pending reply kept "its Turn controls" through a failed read. That was wrong: the controls were visible but rejected locally until a Session read had adopted the Turn. Cause. Change (
Regression ( Validation
Future-facing pass: the two "running" readings (the pending-reply projection and the handler ownership) are now written from the same 409 receipt. I did not merge them into one typed store. That would move ownership shared with the recovery effect and the drawer, which is larger than this fix. |
huangruiteng
left a comment
There was a problem hiding this comment.
动机
精确 head d6ad7754d9c4538ced39ebebd1a319b65ed0740e 想修复普通管理会话在一个 Turn 尚未结束时再次发送会遇到的冲突,并恢复 409 / 短暂 503 后的会话连续性。这个目标对 managed_runtime 会话合理:避免重复启动,仍能从同一条回复调整或中断。但同一 Goal Chat 也承载 attached_host 会话;它有明确的后续消息排队契约,不能把“当前 Turn 运行中”统一解释成“发送必被拒绝”。
改动思路
页面从当前 Session 的 pending 消息推导 conversationTurnRunning,禁用普通发送与快捷提问;LoopX 模式的专门投递入口保留开放。Dashboard 的 409 路径会采用服务端返回的 running Turn ID,重试短暂不可用的 Session 读取,待权威 readback 到来后撤掉临时 handoff。恢复场景同时验证调整和中断按钮可用。正向的 managed-runtime 恢复思路与现有状态 owner 一致;问题在于输入框门禁没有区分 Session 模式,它提前阻断了后端允许的 attached-host 入队路径。
具体改动
五个文件,+356/-11:i18n 和样式增加运行中提示;personal-workspace-page.tsx 在约 1004 行以 pending 消息决定通用 composerBlocked,约 1599 行又在 sendMessage 提前返回;dashboard-page.tsx 增加 409 handoff、503 后 readback 重试及控制按钮所需的临时 Turn 状态;chat-recovery.mjs 增加浏览器回归场景。构建、现有 chat-recovery、LoopX-mode、conversation-activity 浏览器场景、workspace contract smoke 和 diff 检查通过;附着会话后端的 25 个 broker 测试也通过。远端 CI 未轮询,不能把它写成绿灯。
关键代码讲解
conversationTurnRunning 只看当前 Session ID 与 pending Turn,除 LoopX 专门投递外没有 session_mode 分支。可是在 ChatRuntimeController.submit_turn 中,attached_host 首先进入 enqueue_attached_agent_turn;后者调用 create_queued_turn,在活动 Turn 存在时仍持久化有界后续消息,而非走 managed-runtime 的单 Turn 接受规则。dashboard-page.tsx 的 409 恢复是另一条路径,不能补回在按钮/handler 层已经被阻止、根本没有发出的附着会话请求。
对主干的风险
[P1,阻断] 附着宿主会话在当前 Turn 运行时无法排队下一条 Goal Chat 消息。 使用同一个合成 attached_host Goal Session、相同 active_turn_id 和待发送文本,对固定 base/head 做配对浏览器复现:base 的发送按钮可用,点击后 /turns 请求数从 0 到 1;此 head 的按钮禁用并显示“先调整或中断”的提示,请求数保持 0。底层 submit_turn -> enqueue_attached_agent_turn -> create_queued_turn 明确支持这类后续消息;现有 broker 测试也覆盖 Web/Lark 共用的附着会话队列。提示中的“中断”对活动 attached Turn 还会由 attached_session_interrupt_unavailable 拒绝,因此它不是可靠的替代入口。请把普通 managed-runtime 的运行中门禁限定在其适用模式,保留 attached-host 的合法队列路径,并增加该分支的 base/head 浏览器回归;若有附件限制,应单独按后端契约处理。
语义与 CI 对齐
本 PR 改的是默认发送行为,而不只是提示文案;“一个 Session 只能有一个 Turn”并非跨所有 Session 模式的真实规则。当前正向场景通过不覆盖上述反例;本地构建与现有 smoke 通过也不应抹去精确 head 的用户操作回退。这个阻断项完全来自 PR 改动,不是别的 PR 或红 CI 的归因。未查询远端 CI,合并就绪需在修复后另行判定。
我的整体评价
REQUEST_CHANGES。409/503 会话恢复和控制按钮修复有价值,相关现有验证通过;但通用输入门禁把受支持的 attached-host 后续队列关闭,影响正在运行 Goal 的继续协作。长期维护上应将“managed-runtime 单 Turn”与“attached-host 可排队”放在现有 Session 模式契约中,而非让 pending UI 消息充当所有模式的权限判定。相关的小型前瞻重构是复用已存在的 session_mode 作为门禁条件,不需要新增平行状态或框架。修复这个范围并补负例后即可重审,不要求扩大到无关恢复逻辑。
English verdict: REQUEST_CHANGES - exact head d6ad7754d9c4538ced39ebebd1a319b65ed0740e blocks the supported attached-host follow-up queue while an active Turn runs. On paired browser fixtures, base sends one queued request and head disables Send with zero requests. Focused build and existing smokes pass, but they do not cover this mode. Remote CI was not polled.
…ng Turn The running-Turn composer gate treated every Session like a managed runtime Session. An attached host Session queues a message sent while the host's Turn runs (submit_turn -> enqueue_attached_agent_turn -> create_queued_turn), so the gate closed a supported follow-up path and posted nothing. Read each bound Session's mode from the Session owner and apply the gate only to Sessions that admit one Turn at a time. The typed ChatSessionMode and chatSessionQueuesFollowUps helper name that service rule once. The composer-session-admission browser scenario runs the same Goal Session and running Turn in both modes: managed_runtime keeps Send closed with zero POSTs; attached_host keeps Send open and posts the follow-up exactly once. Signed-off-by: song <liusongstep@gmail.com>
|
This pull request has merge conflicts with Choose the remote for the base repository, not an out-of-date fork. git fetch upstream
git rebase upstream/main
# Resolve each conflict, git add the resolved files, then git rebase --continue.
git push --force-with-lease origin HEADFor a same-repository clone whose Keep the DCO |
Signed-off-by: song <liusongstep@gmail.com> # Conflicts: # apps/presentation/dashboard/src/views/dashboard-page.tsx # examples/personal-workspace-browser-smoke.mjs # examples/personal-workspace-browser/chat-recovery.mjs
Signed-off-by: song <liusongstep@gmail.com>
|
Addressed the open blocking findings on P1 — attached-host follow-up queue blocked (@huangruiteng)Cause. Change. The page now reads the bound Session's mode from the Session owner ( Evidence — new scenario
Mutation check (each reverted, then restored): gate without the mode branch → attached case fails; gate removed → managed case fails; recovery effect's admission read removed → attached case fails. P2 — Adjust/Interrupt unusable during the 409 handoff (@cocolord)Already fixed on the current head by Merge fallout worth flagging
Not covered by these tests: the |
huangruiteng
left a comment
There was a problem hiding this comment.
本轮对 d3f7de8adbc3b58b41ba39edd38cb82c106f109f 的完整 base-to-head 改动重新评审,结论:没有阻塞性发现,APPROVE。最近一轮指出的 attached-host 合法排队被阻断问题已经修复;此前取消恢复、409 接管及接管期间控制的修复也重新验证。
动机
managed-runtime Session 已有 Turn 时,普通发送会得到 409。刷新后只知道后端正在回答、却仍让用户发送,会丢掉有效反馈并诱发重复提交。目标是让 composer、回复控制和当前 Session 的真实接收规则一致,而不是把所有 Session 一概限制成“一次只能发一条”:attached-host 和已开启的 LoopX ingress 仍有各自合法的排队路径。
改动思路
复用既有 pending reply/Session/Turn identity 和后端提供的 session_mode。只在同一个 managed Session 的普通对话运行中关闭新发送,并提示使用已有调整/中断控件;不新增用户确认步骤,不改变后端接收权威。409 是接管正在运行的 Turn 的证据,不是本次消息已被接受;必须撤掉未接受的消息、保留草稿,并立刻绑定控件的目标。
具体改动
关键代码讲解
chatSessionQueuesFollowUps:输入 Session 读回中的ChatSessionMode,只有精确的 attached_host 才声明可排队。它只是消费后端接收规则,没有新的 permission/grant;缺失 mode 不猜测 attached 能力。recordSessionAdmission:在 create/resume/send/retry 的实际 Session 读回时维护按 session_id 的可排队集合;非 attached 读回会删除该条目。传给 composer 的值来自当前绑定 Session,而不是 Goal 名、所有历史消息或宿主是否存在。conversationTurnRunning:先排除合法 LoopX delivery 和 attached queue,再匹配 pending reply 的 sourceSessionId/sourceTurnId。普通 Send、Enter 和 quick prompts 共享这个 guard;另一个 Goal/Session 的回复不能封住当前输入。- 409 handoff / recovery:409 分支撤回未接受的 user message,把已有 reply 绑定到服务器报告的 exact Turn,立即写入 active/runtime binding,并保留草稿。后续 Session GET 失败会重读而不是把 Turn 当作结束;finally 不清掉 handed-off identity。取消的 recovery 会删除自己的 placeholder,终态/中断会释放 composer,因此离开再回来不会留下永久 pending。
独立 paired walkthrough 使用相同合成 Goals、same-status fixture、分别构建的 base/head assets,以及真实 Chat HTTP、ChatRuntimeController、durable ChatSessionStore 和 attached bind/claim owner;只有模型 provider 被脚本化,无付费模型调用。base 的 running managed Session 仍可点击发送,实际 POST 一次得到原 English 409;head 的同类 Session 禁止发送、没有 POST。head 在离开/返回并完成当前 Turn 后,原 Session 接收并持久完成第二个 Turn,两条 agent response 读回完整。attached case 在 base/head 都允许发送,实际 POST 一次写入跟进消息,原 host-owned active Turn 未被替换。检查了 populated packaged whole viewport,当前目标、等待原因和调整/中断路径一起可见。
对主干的风险
- 已运行:dashboard contract test、TS typecheck、打包构建、82 项真实 store/controller/broker 相关 Python 测试;packaged
composer-session-admission、chat-recovery、loopx-mode、conversation-activity、conversation-return-continuity五个场景通过。chat-recovery 包含 409、503 后重读、接管期间 exact-target 调整/中断及晚到读回不能复活已中断 Turn。完整 diff hygiene 和 premerge canary(8 项 selected checks,0 failures)通过。 - 单独的架构检查在 immutable base 与两个 heads 都因同一
generated == 1实际为 2 的 assertion 失败;本 PR 不触及其生成/测试因果路径,两个 Session 模式的 changed invariant 有独立真实后端证据。这是基线验证问题,不据此 request changes,也没有调整预算掩盖失败。未查询、轮询或等待远程 CI;代码批准不是合并准备证明。 - 真实后端 walkthrough 使用脚本化 provider,不证明部署级 Codex/Lark provider 或所有网络失效时序。独立测试代理在离开 SSE 页面后出现 stream-abort 清理错误,附加 attached completion 截图步骤未完成;已保留错误且不把它算作产品通过证据。此前完成的 actual attached queue/readback 与原生 broker completion 测试不受影响,脚本代理不是仓库改动。
- 此改动没有改变 CLI/Lark/Chat 后端接收、host binding、quota 或 authority;LoopX ingress 的既有分流以 packaged 场景验证。没有新安装步骤、协议写入或数据迁移。回滚可 revert 前端变更并重新构建。
我的整体评价
这是现有 Session/恢复 owner 内的修复,普通用户不再为必然被拒绝的请求重发;已知 Session mode 避免把安全等待变成过度阻断。整体 diff 的附加状态服务于同一个恢复交互,未引入第二个 Session authority。future-facing pass 采用小型 typed mode helper 和统一 admission 记录;保留不同 Session 接收语义,不做一套万能队列或跨上下文重写。
APPROVE;无阻塞性发现。最强剩余验证是部署级 provider/network recovery,而不是本轮已证明的接收/恢复不变量。本轮没有执行 merge,留给 maintainer。
English verdict: APPROVE
Keep the handed-off Turn's active identity across the send's finally while retaining main's manager capability-revision bump. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Signed-off-by: huangruiteng <huangrt01@163.com>
Both failures reproduce on a clean origin/main checkout, so the PR inherits them rather than causing them: - test_new_independent_twin_cannot_hide_behind_generated_pair pinned generated_verified == 1, but loopx-project#5322 made loopx/control_plane/content_digest.py a verified generated artifact, so the smoke now reports two generated twin pairs. The generated set is owned by the generators; the test now requires at least one generated pair and the raw split. - test_live_decision_adds_only_existing_required_read_channel compared every other payload key against the baseline, but the turn-start hook dispatch observation legitimately travels with the added required read. It is now part of the excluded set, so the assertion still proves the hint arrives on the existing required-read channel. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Signed-off-by: huangruiteng <huangrt01@163.com>
The packaged Personal Workspace smoke bound the execution chip at 26px. loopx-project#5341 turned the chip into the executor picker button with button.personal-execution-chip { min-height: 28px }, so build and chat-bundle-browser fail on main with "Execution chip is not a compact hairline row: 28px tall". Bind the check to the control's own compact height while still requiring the chip to stay inside the channel header's single row. Reproduced on a clean origin/main checkout and in dev mode before the change. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Signed-off-by: huangruiteng <huangrt01@163.com>
huangruiteng
left a comment
There was a problem hiding this comment.
本轮对 2d9fb282d6a0788b422ca54dbdf5e520173dec3b 的完整 base-to-head 改动重新评审,结论:没有阻塞性发现,APPROVE。
动机
d3f7de8ad 之前的评审结论保留:managed-runtime Session 运行中时,composer 与后端「一个 Session 一次一个 Turn」的接收规则一致;attached-host 与 LoopX ingress 的合法排队路径不被阻断。本轮额外处理两件与新 head 直接相关的事:与最新 main 的冲突,以及 main 上已经存在的红灯(它们让本 PR 的 CI 无法变绿)。
改动思路
不新增产品语义,只做三件必要的事:把 PR 分支与最新 main 合并并保留双方语义;把两处在干净 origin/main 上同样失败的陈陈旧测试期望改成跟随各自 owner;把 packaged 烟测对 execution chip 的高度绑定更新为交互控件的实际紧凑高度。
具体改动
- 冲突解决(dashboard-page.tsx):send 的
finally同时保留 PR 的if (!handedOff) activeTurnIds.current.delete(...)(接管中的 Turn 身份留给 recovery 或权威读回)和 main 新增的if (targetContextId === "manager") setCapabilityRevision(...)。两者语义独立,都未丢弃。 - main 红灯 1(test_turn_contract_generation.py):
#5322让loopx/control_plane/content_digest.py也成为生成物,generated_verified从 1 变 2。断言改为「至少一个生成对 + raw = maintained + generated」,生成集合的归属留给 generators,不再要求改这条测试。 - main 红灯 2(test_prompt_upgrade_hook.py):
turn_start_capability_hook_dispatch是产生该 read 的 turn-start hook 观测,随 read 一起下发;把它加入排除集合后,断言仍在证明「hint 只走既有 required-read 通道」。 - main 红灯 3(execution-chip.mjs):
#5341把 chip 变成 executor picker 按钮(button.personal-execution-chip { min-height: 28px }),packaged 烟测仍按 26px 判定,于是 main 的build/chat-bundle-browser报Execution chip is not a compact hairline row: 28px tall。阈值绑定到控件自身的紧凑高度,仍要求 chip 留在 channel header 同一行内。
对主干的风险
- 三处红灯均在干净
origin/main(7e60e6999)上复现:前两条我在 main worktree 里逐条跑出同样失败;第三条在 main 的 CIbuild任务日志里就是同一条 28px 报错,本地 dev 模式复现后修复。 - 本 head 本地验证:
tsc --noEmit通过;打包构建(vite build+build:chat)通过;packagedexecution-chip通过;dev 模式execution-chip、chat-recovery、composer-session-admission、loopx-mode、conversation-return-continuity通过;pytest test_attached_session_broker.py test_chat_session_active_turn.py test_chat_store_input_validation.py test_chat_codex_home.py82 项通过;被修的两份测试文件 38 项通过。 - 远程 CI(head
2d9fb282d):17 pass / 6 skipped / 0 fail。其中typescript-core (3/3)首次运行在host_process.test.ts:47 leader_exit上偶发失败('' !== '6',杀掉进程组时的在途写入竞态;本地连跑 5 次均 9/9 通过),重跑该 job 后通过,随后typescript-check/checks/merge-gate正常收敛。这条偶发不属于本 diff 的路径,但值得后续单独加固。 - 回滚:前端改动 revert + 重建即可;三处测试期望更新无运行时影响。
我的整体评价
PR 主体仍是既有 Session/恢复 owner 内的修复,未新增第二个 Session authority。本轮新增的四个 commit 全部是「合并最新 main」与「让继承自 main 的红灯停止阻塞」,没有把无关重构塞进这条 diff;对齐了各 owner(generators、prompt-upgrade hook、交互式 chip)已经做出的决定,而不是改产品行为去迁就旧测试。
APPROVE;无阻塞性发现。
English verdict: APPROVE
Keep this branch's composer-admission scenario and main's confirmed-operations scenario in the catalog, and take main's loopx-project#5327 hairline budget (28px for the button chip, 26px for the label) which supersedes this branch's interim bound. Signed-off-by: huangruiteng <14976749+huangruiteng@users.noreply.github.com> Signed-off-by: huangruiteng <huangrt01@163.com>
…ser-running-turn Signed-off-by: song <liusongstep@gmail.com> # Conflicts: # apps/presentation/dashboard/src/features/personal-workspace/personal-workspace-page.tsx # apps/presentation/dashboard/src/views/dashboard-page.tsx # examples/personal-workspace-browser-smoke.mjs
… it runs `managed_runtime` accepts one Turn at a time, so the composer closes Send while one runs. An `attached_host` Session is served by the host app, and `submit_turn -> enqueue_attached_agent_turn -> create_queued_turn` explicitly queues a follow-up instead of rejecting it. The composer read every running Turn as a rejected send, which closed a path the backend supports. The page already follows the Session's own `session_mode` through `recordSessionAdmission`. This adds the scenario that keeps that distinction honest, with the same Goal Session and the same running Turn and only the mode differing: - `managed_runtime`: no queued follow-up, one POST, the service's 409 with its `active_turn_id` is adopted by the recovery path. - `attached_host`: the follow-up is queued, the running Turn stays the Session's active Turn, and the composer stays open. Both Turns hold their stream for the whole check, so the assertion is about the admission rule rather than about timing. Signed-off-by: song <liusongstep@gmail.com>
Re-review request — exact head
|
A 409 handoff records the Session and Turn the service refused the send for. Once conversation history moved to the shared read hook, the recovery effect only saw that hook's earlier snapshot, so it could not find the reported Turn: it left the handoff reply pending beside a second pending reply and overwrote the bound running Turn with the stale one. Let the hook re-read the conversation on demand, and use that read as the evidence for adopting a handoff. Until it returns -- or while it keeps failing -- the handoff reply stays pending with its own Turn controls, which is the behaviour the 409 scenario already specifies. The scenario now holds and fails exactly that conversation read, so the check can no longer be satisfied by an unrelated session-listing poll. Signed-off-by: huangruiteng <huangrt01@163.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (maintainer review of an author-owned PR)
Reviewed exact head 64380925dac41055afe5e574f7434c3f07e802d5.
动机
上一版 head(77478f8c9)把 main 合并进来后,会话历史改由共享读取 hook
useConversationHistory 统一负责,409 接管路径因此失效:服务端已经在 409 里明确报告某个
Session 正在跑 Turn,但恢复流程只用 hook 早先缓存的历史快照,看不到那个 Turn,于是
updateConversationMessage 留下的待完成回执之外又追加了一条恢复回执("正在恢复进行中的
Agent 回合"),并把该上下文绑定的运行中 Turn 覆盖成快照里的旧 Turn。CI 的
Release Artifacts / build 打包烟测正是死在这里:chat-recovery.mjs:436 严格模式下
中断本轮 命中 2 个元素。
这不是风格问题,而是本 PR 自己引入的用户可见缺陷:Turn 已经结束、回执却仍是 pending,
composer 持续被锁住,只有重载页面或等下一次快照刷新才会恢复。
另外,main@21f89e4ad 上两条必检 lane 是红的
(tests/architecture/test_turn_contract_generation.py 把 generated twin 对写死为 1;
tests/control_plane/test_prompt_upgrade_hook.py 不允许 turn-start hook 的读取随其
required read 一起返回)。它们在同一个 checkout 上复现(3 failed),修好后本 PR 的
required checks 才可能通过,因此随本 PR 一并最小化修好。
改动思路
保持"会话历史只有一个 owner":不新增第二个历史读取者,而是让
useConversationHistory 提供一次真正的重读(refresh():不传 previous,重新读会话列表
并重新读取各 Session 快照),由 409 接管路径在采纳之前调用。
只有存在 409 交接时才重读;重载、切回会话等路径继续使用 hook 刚读到的快照,避免多余请求。
重读失败时保留待完成回执与其"调整本轮/中断本轮"控制,按退避重试;中断回执落定;晚到的读
不会复活已中断的 Turn(合成 store 的中断会清空 active_turn_id,重读到的是"没有运行中
Turn",直接返回)。
具体改动
src/data/use-conversation-history.ts:新增refresh()(重读列表与快照,写回共享
缓存与状态)与currentChannelSession()(把"频道中某执行器的最新 Session"这一判定从
hook 内联逻辑提出,恢复流程复用同一判定,不复制知识);hook 内部改为 ref 缓存,保留原有
"只补缺失快照、不重复轮询"的语义。src/views/dashboard-page.tsx:恢复 effect 在该上下文存在 409 交接时先
await conversationHistory.refresh(),再据此判定当前active_turn_id;latest用
currentChannelSession派生;409/交接/中断/重试的其余语义不变。examples/personal-workspace-browser/chat-recovery.mjs:把"409 后重读会话"的拦截收窄到
带channel_id的会话历史读(读失败步骤同样只作用于该读)。此前断言可被无关的会话列表
轮询"顺带满足",因此旧 head 才会出现"测试通过但接管语义已坏"的假绿。
对主干的风险
- 风险 1(低):409 交接会多一次频道历史重读(列表 + 快照)。触发面小(仅 409 或读失败
重试,退避上限 10s),且这是该页面在此前版本中的既有行为。 - 风险 2(低):未知
session_mode取值按 managed runtime 处理,即偏向"等待"而不是
"发进运行中的 Turn",失败方向保守。 - 风险 3(已知边界,非本 PR 引入):服务端 409 仍不带 typed
error_code,客户端靠
active_turn_id识别交接;PR 正文已把它记为可选后继。 - 已验证:开发态全量 29/29 场景;打包态全量 29/29 场景(即此前红灯的 lane);
tsc --noEmit与npm run build;test:conversation-returns与
smoke:conversation-history;两条 Python lane 38 passed(main@21f89e4ad上同两文件
3 failed)。 - 未验证:作者提供的 real entrypoint 证据停在
4929e8fa9,合并main后未重跑;浏览器
场景使用页内合成 chat store,不是真实 Python chat 服务;两种已声明取值以外的
session_mode。
我的整体评价
方向正确、改动小而准:它修的是自己带入的接管语义回归,并把"以服务端报告的 Session/Turn
为准"落到一次真实重读上,没有再造第二个历史来源,也没有为了过测试而放宽断言。没有阻塞性
发现,APPROVE。
非阻塞建议:features/personal-workspace/goal-activity.ts 仍直接比较
session.session_mode === "attached_host"(语义不同:宿主持有 Turn 的活动声明),若将来
要统一会话模式判定,可考虑复用 data/chat.ts 的类型化入口。
English verdict: APPROVE. The 409 handoff is now confirmed through a fresh, owner-owned
conversation read instead of a cached snapshot, so the reported Turn is adopted with a single
pending reply. Development and packaged suites pass 29/29 each at this exact head, along with
tsc/build, the conversation-history data smokes, and the two required Python lanes (38
passed; both files fail on main@21f89e4ad). No blocking findings. Residual risk: the handoff
re-read is exercised against a synthetic in-browser chat store rather than the real chat server,
and the author's real-entrypoint evidence predates the merge of main.
|
重新评审 根因: 修复:让历史 owner 提供一次真正的重读( 顺带加固: 验证: 精确 head |
Signed-off-by: huangruiteng <huangrt01@163.com>
huangruiteng
left a comment
There was a problem hiding this comment.
Approval conclusion (maintainer review of an author-owned PR)
Reviewed exact head b26878a0e937e3ff0dc3503d261263e255396a4b
(merge of main@80c75dd40a55a275ad01ad7f14851f63f3b1b2bc into the branch; the previous
64380925dac41055afe5e574f7434c3f07e802d5 review covered the same dashboard logic).
动机
上一版 head(77478f8c9)把 main 合并进来后,会话历史改由共享读取 hook
useConversationHistory 统一负责,409 接管路径因此失效:服务端已经在 409 里明确报告某个
Session 正在跑 Turn,但恢复流程只用 hook 早先缓存的历史快照,看不到那个 Turn,于是
updateConversationMessage 留下的待完成回执之外又追加了一条恢复回执("正在恢复进行中的
Agent 回合"),并把该上下文绑定的运行中 Turn 覆盖成快照里的旧 Turn。CI 的
Release Artifacts / build 打包烟测正是死在这里:chat-recovery.mjs:436 严格模式下
中断本轮 命中 2 个元素。
这不是风格问题,而是本 PR 自己引入的用户可见缺陷:Turn 已经结束、回执却仍是 pending,
composer 持续被锁住,只有重载页面或等下一次快照刷新才会恢复。
早期版本还顺带修过 main 当时的两条红灯 lane
(tests/architecture/test_turn_contract_generation.py 与
tests/control_plane/test_prompt_upgrade_hook.py)。当前 main 已自行修好这两处,
于是本次合并按 main 版本解决冲突,这两个文件已从本 PR 的改动中退出。
改动思路
保持"会话历史只有一个 owner":不新增第二个历史读取者,而是让
useConversationHistory 提供一次真正的重读(refresh():不传 previous,重新读会话列表
并重新读取各 Session 快照),由 409 接管路径在采纳之前调用。
只有存在 409 交接时才重读;重载、切回会话等路径继续使用 hook 刚读到的快照,避免多余请求。
重读失败时保留待完成回执与其"调整本轮/中断本轮"控制,按退避重试;中断回执落定;晚到的读
不会复活已中断的 Turn(合成 store 的中断会清空 active_turn_id,重读到的是"没有运行中
Turn",直接返回)。
具体改动
src/data/use-conversation-history.ts:新增refresh()(重读列表与快照,写回共享
缓存与状态)与currentChannelSession()(把"频道中某执行器的最新 Session"这一判定从
hook 内联逻辑提出,恢复流程复用同一判定,不复制知识);hook 内部改为 ref 缓存,保留原有
"只补缺失快照、不重复轮询"的语义。src/views/dashboard-page.tsx:恢复 effect 在该上下文存在 409 交接时先
await conversationHistory.refresh(),再据此判定当前active_turn_id;latest用
currentChannelSession派生;409/交接/中断/重试的其余语义不变。examples/personal-workspace-browser/chat-recovery.mjs:把"409 后重读会话"的拦截收窄到
带channel_id的会话历史读(读失败步骤同样只作用于该读)。此前断言可被无关的会话列表
轮询"顺带满足",因此旧 head 才会出现"测试通过但接管语义已坏"的假绿。
对主干的风险
- 风险 1(低):409 交接会多一次频道历史重读(列表 + 快照)。触发面小(仅 409 或读失败
重试,退避上限 10s),且这是该页面在此前版本中的既有行为。 - 风险 2(低):未知
session_mode取值按 managed runtime 处理,即偏向"等待"而不是
"发进运行中的 Turn",失败方向保守。 - 风险 3(已知边界,非本 PR 引入):服务端 409 仍不带 typed
error_code,客户端靠
active_turn_id识别交接;PR 正文已把它记为可选后继。 - 已验证:开发态全量 29/29 场景;打包态全量 29/29 场景(即此前红灯的 lane);
tsc --noEmit与npm run build;test:conversation-returns与
smoke:conversation-history;tests/architecture/test_turn_contract_generation.py与
tests/control_plane/test_prompt_upgrade_hook.py38 passed(现为main自己的版本)。 - 未验证:作者提供的 real entrypoint 证据停在
4929e8fa9,合并main后未重跑;浏览器
场景使用页内合成 chat store,不是真实 Python chat 服务;两种已声明取值以外的
session_mode。
我的整体评价
方向正确、改动小而准:它修的是自己带入的接管语义回归,并把"以服务端报告的 Session/Turn
为准"落到一次真实重读上,没有再造第二个历史来源,也没有为了过测试而放宽断言。没有阻塞性
发现,APPROVE。
非阻塞建议:features/personal-workspace/goal-activity.ts 仍直接比较
session.session_mode === "attached_host"(语义不同:宿主持有 Turn 的活动声明),若将来
要统一会话模式判定,可考虑复用 data/chat.ts 的类型化入口。
English verdict: APPROVE. The 409 handoff is now confirmed through a fresh, owner-owned
conversation read instead of a cached snapshot, so the reported Turn is adopted with a single
pending reply. Development and packaged suites pass 29/29 each at this exact head, along with
tsc/build, the conversation-history data smokes, and the two affected Python lanes (38
passed; main now carries their repair, so they are no longer part of this diff). No blocking findings. Residual risk: the handoff
re-read is exercised against a synthetic in-browser chat store rather than the real chat server,
and the author's real-entrypoint evidence predates the merge of main.
Goal And Delivered Outcome
another turn is already running for this session.本轮回答进行中。可在回答里调整或中断本轮,结束后再发送。until the Turn completes. A 409 that still arrives is adopted as the same pending reply, with its Adjust/Interrupt controls acting on the reported Session and Turn.main.Scope And Continuation
attached_hostqueues a bounded follow-up, a managed runtime admits one Turn at a time), so the composer follows that typed mode instead of reading every running Turn as a rejected send.error_code; the client keys onactive_turn_id. A typed code on the server would be a separate contract change.main(tests/architecture/test_turn_contract_generation.py,tests/control_plane/test_prompt_upgrade_hook.py). Currentmainfixed both upstream, so merging it resolved those conflicts inmain's favor and this PR no longer touches either file.Validation
staticpassednpx tsc --noEmitandnpm run build(tsc +vite build+build:chat) inapps/presentation/dashboard.integrationpassednode examples/personal-workspace-browser-smoke.mjs— 29/29 scenarios, includingchat-recovery(composer waits while a Turn runs; the 409 handoff adopts that Turn, keeps the draft, retries a failed Session read and settles on interrupt),composer-session-admission,attached-host-follow-up,conversation-history-recovery,loopx-mode,typed-actions,steward-journey.integrationpassednpm run build:chatthenLOOPX_PERSONAL_WORKSPACE_PACKAGED=1 node examples/personal-workspace-browser-smoke.mjs— 29/29 scenarios served from the built bundle. This is the lane that failed on the previous head (chat-recovery.mjs:436, two pending replies).unitpassednpm run test:conversation-returns(dedup, stream preservation, watch retirement) andnpm run smoke:conversation-history(real HTTP store, partial read, channel isolation, missing-only recovery, zero writes or Turns).integrationpassedPYTHONPATH=$PWD python -m pytest tests/architecture/test_turn_contract_generation.py tests/control_plane/test_prompt_upgrade_hook.py— 38 passed on the merged head (these lanes aremain's current version, not changed by this PR).regression_paritypassed中断本轮resolved to 2 elements, and the run ended atchat-recovery.mjs:436. The scenario now holds, and in the read-failure step fails, the conversation's own read, which only the requested re-read satisfies.real_entrypointpassedloopx serve-status+loopx chaton an isolated synthetic registry with a stub Codex app-server: a Turn left running, page reloaded;mainposted a 409 Turn and showed the English error, this branch posted nothing and showed the hint. Recorded at 4929e8f; not re-run after the merge ofmaininto this branch.managed_runtime/attached_hostis read as a managed runtime Session, which errs toward waiting rather than sending into a running Turn.Frontend / Visual Evidence
Type of Change
LoopX Area
Technical Direction
Shared-authority RFC fixture impact
N/A
Boundary Checklist
none.Signed-off-bytrailer (git commit -s).Future-facing refactor pass: considered deriving "Turn running" from the runtime binding in
dashboard-page.tsx; kept the timeline's pending message as the single source because it already drives the reply's Adjust/Interrupt controls, so the composer and those controls cannot disagree. Recorded the 409 confirmation as a re-read owned by the conversation-history hook, so the page keeps one history reader instead of two.