fix(ci): reconcile exact-target recovery and settled readback - #5567
huangruiteng merged 7 commits into
Conversation
Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
…failures-20261004 Signed-off-by: song <liusongstep@gmail.com>
Reuse the reviewed shared source-grant and settled-readback fixes from PR loopx-project#5533 by Duang777. Keep recipient policy in the TypeScript owner and expose only receipt-bound historical identity after settlement. Signed-off-by: song <liusongstep@gmail.com>
Exercise Agent identity drift rather than legal cross-Todo session reuse, include release identity in the independent fingerprint oracle, and execute local Goal success fixtures inside their registered workspace. Retain settled no-run/no-spend assertions. Signed-off-by: song <liusongstep@gmail.com>
Signed-off-by: song <liusongstep@gmail.com>
|
Repair validation on exact head The 14 failures from the previous shard-2 run now pass together: 9 exact-target authorization cases, 3 settled-monitor recovery cases, the runtime fingerprint case and failed-Turn session identity case. Their fixes are integrated into this candidate, reusing the relevant shared work from #5533 by @Duang777. Broader validation: 314 other related Python cases passed, then the complete 90-case Turn-driver module passed after preparing its success workspaces inside the registered Goal. The earlier 12 out-of-root fixture failures are preserved in the validation history. 124 typed policy/lifecycle/readback cases also passed. Revoked source no-send, restore-once, original instance readback, lost-response recovery, real File/SQLite monitor flows and no second spend are covered. Standard premerge: 19 selected checks plus direct checks passed. The inherited advisory The new head needs fresh CI and maintainer review. These local results neither certify the remaining #5533 supervisor changes nor constitute approval to merge this runtime change. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
精确版本:5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7;实际 fork 基础版本:1af7dbd。首次独立全 PR 评审,未查询、轮询或等待 CI。
动机
由管家向已注册的精确 Goal/Agent 交付工作、随后读取结果的使用者,以及重放已结算监控 Turn 的执行者。
例如合法的来源已获准向 builder 交付:旧版因再次走不支持严格注册表的通用目录而拒绝,新版走既有来源策略后成功;监控结算后,新版为仍可重建的工作项恢复历史投影,但完成后的工作项仍只在原回执中可读。
实际 base/head 对照确认合法精确目标恢复、越权目标继续拒绝;monitor 的原结算身份和无执行、无重复扣额保持。新增文档的无条件 selected_todo 读回承诺在 completed 反例上未兑现。
本轮不把历史 CI 统计、测试通过或单个投影等同于完整产品验收,不接管来源账户、执行权限、benchmark 或其他 agent,也不授权合并该运行时 PR。
改动思路
这个修复的大方向合理:精确目标先由 Goal lifetime 与注册成员校验,再复用既有 TypeScript 来源策略;Python 只观察来源与存储,不另建权限判断源。普通目录与已经核验的精确目标共享 _source_context_grant,分别提供完整目录或单个合法候选,不因为同一个 GitHub 身份而混淆请求者或权限。返回仍绑定原来源、原 instance、原会话,并按当前 sender/selected/blocked policy 重新准入。
监控读回复用既有 compact projector,保留 no-run/no-spend 字段。这里必须区分永久结算事实与当前工作 lane:前者由已提交回执拥有,后者可以在完成时消失。新增文档把两个工作投影的可用性写得比生产分支更强,因此当前整体结论不能批准。
具体改动
完整 14 文件 +190/−67:四个生产文件整合来源 grant adapter、精确 return 的 Agent 注册检查与 settled 历史投影;六个 Python 测试文件修正 sender 撤权、source release fingerprint、Agent/session scope 与合法 workspace fixture;workflow 同时固定 budget 的 base/main selector 为 event SHA;质量指南中英同步说明固定比较基线;monitor 协议新增历史投影说明;registry I/O manifest 仅调整四处行号。没有新 CLI 选项、grant 数据库或 quota/scheduler 状态。
先读取基础修订 1af7dbd43a1629b0356ecd5ee28a426ad45d1d7a 的 docs/reference/protocols/quota-monitor-observation-receipt-v0.md 与 docs/architecture/rfcs/goal-instance-identity-and-orphan-recovery-v0.md。原 monitor 契约的 Original-settlement-identity、Settled-no-execution、Original-historical-receipt 三条均通过:原回执仍可读,完成不会重开 Turn,无新操作或第二次扣额。instance RFC 的来源权威、精确 lifetime 与 existing grants 约束由原 owner 保留。新增文档的无条件 selected_todo 承诺单独评估,不能用它替自己制造通过标准。
关键代码讲解
source_grant_observation.py:36的_source_context_grant:把目录或单个目标交给同一个resolveSourceRecipients,保留 verified sender、source digest、selected/all_registered 与 blocked 规则。新增 exact adapter 不读取任意目标注册事实;调用方必须先持有合法 Goal scope。manager_context/__init__.py:99的deliver:在 active/registered/lifetime/lifecycle 校验之后才调用 exact target authority,随后按原 request lock 写 inbox/route 并读回。实际反例中 selected-other、blocked Goal、移除 sender、未注册 Agent 都拒绝且没有 entry/registry 修改;两个合法来源在 head 成功,base 错误拒绝。manager_context/roundtrip.py:388的_exact_return_context:先核验原 result/route/instance,再验证注册 Agent 与当前 grant,最后由现有 return admission/verification owner 处理重试。相关真实 store restart、原 HTTP snapshot、撤权后不发送、恢复后只返回一次均经过本轮回归;外部 provider sender 在测试中是模拟边界,不能声称真实 Lark 网络发送已验收。settlement_precedence.py:121的apply_settled_monitor_precedence:只在当前 work lane 已识别为 settled monitor 时恢复匹配的旧 action 与 compact selected 投影。Monitor 完成后该 lane 不可重建,127–129 行提前返回;最终 CLI 仍安全 settled,却没有新增文档保证的 selected 字段。
对主干的风险
[P2] 新增历史投影契约应与完成后的实际读回一致。 位置:docs/reference/protocols/quota-monitor-observation-receipt-v0.md:80。本轮用同一 synthetic fixture 和真实 CLI 在基础版本、精确 head 都执行 monitor admission → poll/settlement → 正常 completion → 同 Turn guard/duplicate poll。head 结果是:原 heartbeat_receipt.settlement_identity 正确,should_run=false、must_attempt_work=false、无 executable CLI actions、journal 不变、duplicate replay 不追加;但 selected_todo 与 agent_lane_next_action 都缺失,当前 lane 也不含原 monitor。open 与 28 项其他 Agent 工作的 crowded 情况则成功恢复投影。
旧版也没有 completed selected 字段,因此没有把它说成新引入的扣额或执行回归;阻塞的是本 PR 新发布、超出实际行为的共享契约。消费者按新文字实现历史读回会遇到正常生命周期下的缺字段。最小修复是把文档明确为:原回执始终可读,两个工作投影仅在对应 monitor lane 可重建时出现,并同步中文与回归断言;若必须无条件提供 selected_todo,则由已提交回执派生它,补 completed/superseded/archived 实际 CLI 覆盖,保持无执行、无第二次扣额。
独立验证:五组相关模块 183 passed;完整 Turn driver 90 passed;source grant/lifecycle/quota readback 19 TS passed;control-plane typecheck、Ruff、配置内 19 文件 mypy、diff hygiene 均通过。真实 legacy/File/SQLite monitor 回归及 source-session/ChatHTTPServer 路径包含在其中。配对探针 6 个来源 scope、3 个 monitor 生命周期场景:相同 fixture SHA-256 f7487f3bb7183441ffecaba26756c3776c2e5c8ef7bc876ed7d206634a646323,base/head 观察分别 50aaab28ef0c4cbdd942f45ed87538c78431cfa1b3be6c36f607ec3bd5affa7a / d1bfa6898e0cd0474220ceae8f4bf1b71c3aff1ffb3f82fe7ebfe03ba09dccd9。九个场景中 completed 新字段存在性 oracle 失败;三个 monitor 的无执行、无扣额 oracle 全部通过。没有删除或弱化失败断言。
验证设置历史保留:初版探针误读 deliver 返回字段、误用通用 update 完成;首次 TS 命令错误假定安装 tsx。改为真实返回字段、专用 completion 入口及仓库 native Node strip-types 命令后重跑;这些不是生产缺陷,也不抹去当前 completed 反例。全仓库套件、native Windows、实际外部 provider、安装 App/模型采纳与长期收益未测。
语义与 CI 对齐
复用现有 typed source scope、GoalRef、ReceiptBoundMonitorPhase 和 effective action,没有新增共享 vocabulary 或 substring 分类。先跑 changed-from advisory(0 个支持语法候选,动态构造不在其证明范围),再跑 semantic smoke;均通过。代码默认行为变化已由 PR、fixture 与协议披露,未宣称 opt-in/default-off;历史投影不授予权限。当前语义问题是新增文档与实际生命周期不一致,semantic inventory 通过不能证明此承诺。
workflow 的两个比较 ref 固定到 pull-request base、merge-group base 或当前事件提交,避免队列等待时 main 漂移;未获取历史运行日志,也未把作者历史 CI 统计当独立当前证据。实际固定 base/main 为上述不可变基础修订执行 CLI output differential budget,通过;registry I/O census 检查为 current(278 sites)。没有改任何 hard limit。公共差异仅包含可复用产品/开发契约;既有 untracked lockfile 不在 PR。
我的整体评价
REQUEST_CHANGES。 来源授权修复及 monitor 无重复副作用的价值已得到独立验证,结构与工作量也合适。long_horizon 在本次恢复/重放上保持;user_experience 因新文档指向完成后缺失的历史字段仍未证明一致。最小修复可只把新契约范围写准确,而不扩展生命周期或新增存储。未来重构检查认为共用来源 adapter 与既有 projection 已足够,拒绝为此加入平行权限或历史状态 owner。
修改后在新精确 head 重跑 completed readback 与相关 monitor/return 回归,再评审完整 PR。本结论不受 CI 是否成功影响,也不表示原回执丢失或发生第二次扣额;尚无授权由本 reviewer 合并这个通用运行时 PR。
English verdict: REQUEST_CHANGES - 5567@fcfb24c65472805d3a03d9d9fc200f1927edb3d7; align the new unconditional historical selected_todo contract with completed-monitor CLI readback. Original settlement identity and no-run/no-spend are preserved; 273 Python and 19 TS tests pass, but the independent completed-field oracle fails. CI not consulted.
The new contract text promised the receipt-bound identity unconditionally in `selected_todo` and `agent_lane_next_action`. Independent review reproduced a completed Monitor Turn where the committed receipt still reports `settlement_identity`, while both work-lane projections are absent because the lane is no longer reconstructible. State the durable and derived halves separately: the receipt identity always stays readable, and the two projections appear only while the bound Monitor still projects as current work. Synchronize the Chinese contract and pin the completed readback with a regression case that also proves one observation, no re-open and no second spend. Rejected the alternative of deriving `selected_todo` from the committed receipt: that would add a second historical selection authority beside the existing typed lane owner, rather than correcting the published contract. Signed-off-by: song <liusongstep@gmail.com>
|
Repair validation on exact head What the review foundThe P2 was a contract-vs-behavior mismatch: the added text promised the receipt-bound identity unconditionally in
The mechanism is the one named in the review: RepairApplied the minimal option the review proposed rather than the unconditional-
Validation
One correction to my own earlier work: I first asserted CI on this head is fresh and unconsumed. This is still a runtime-adjacent contract change, so it needs re-review on the new head; the local evidence here is not merge authorization. |
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh.
精确版本:5567@a952b9a;不可变共同基础:1af7dbd43a1629b0356ecd5ee28a426ad45d1d7a。当前完整 PR 独立复审,未查询、轮询或等待 CI。没有当前阻塞 finding。
动机
向已注册精确 Goal/Agent 交付工作并读取结果的调用者,以及重放已结算 monitor 的执行者。
合法 sender 向已注册 builder 交付时,基础版错误拒绝;新版成功而越权目标仍拒绝。monitor 完成后原回执仍可读,当前工作投影可以消失,文档现在准确说明这一条件。
当前配对探针的九个 head 场景全部通过;原实例、原结算身份、无新执行和无第二次扣额保持,上轮完成后字段承诺的阻塞已经修复。
这次只资格化来源恢复和结算读回,不代表实际外部 provider、安装 App、模型采纳或整个 Goal 已验收,也不授予本 PR 的合并权限。
改动思路
继续复用已经存在的 TypeScript 来源策略与 Goal lifetime/成员校验。精确目标先通过源注册表验证,再向同一策略提供单个合法候选;普通目录仍向它提供完整候选。sender、selected/all_registered、blocked-target 的规则共用,Python 只适配观察与 IO。合法来源不必绕回不支持严格源注册表的普通目录。结果仍绑定原实例、原请求和原受众,当前 grant 撤回会阻断发送,恢复后原结果只返回一次。
monitor 永久结算事实由原回执拥有,当前工作字段由可重建 lane 派生。新增文档现在明确二者不同,完成后字段消失不代表回执消失,也不重新授权执行。比添加历史数据库或第二套权限源更小,保留已有 typed owner。
具体改动
全 14 文件 +324/-70:四个生产文件共用来源 grant adapter、精确 return 的 Agent 注册检查与 settled compact projection;六个测试文件覆盖实际 sender 撤权、source instance fingerprint、session 身份及有效 workspace;workflow 固定 budget base/main 为事件 SHA;质量指南中英同步;monitor 协议修正条件;registry I/O manifest 只更新四处源行号。相对上轮 fcfb24c 仅 docs/test +142/-11,生产字节相同,但本轮仍重新执行了当前证据。
先读 docs/reference/protocols/quota-monitor-observation-receipt-v0.md,spec_revision 1af7dbd43a1629b0356ecd5ee28a426ad45d1d7a 及同修订 instance RFC。Original-settlement-identity、Settled-no-execution、Original-historical-receipt 均 implemented:完成不会重开已提交 Turn;原 receipt 始终保留身份;工作字段存在与否均不能发起新执行。新版协议 79–95 行同时准确说明 completed 后的字段条件,中英文一致。
关键代码讲解
source_grant_observation.py:36_source_context_grant共用 sender/digest、selected/all_registered 与 blocked 规则;只改变候选观察范围,没有 Python 平行权限判断。manager_context/__init__.py:99deliver在注册 Agent、active/lifetime 和 lifecycle 检查后调用 exact adapter;真实反例中 selected-other、blocked、sender removed、未注册 Agent 拒绝且零 entry/注册表写入。manager_context/roundtrip.py:388_exact_return_context先核验原实例、result、route,再验证当前来源 grant,之后进入已有 return admission 与恢复 owner。store restart/response loss/撤权后恢复由当前测试覆盖;外部 sender 是合成边界。settlement_precedence.py:121apply_settled_monitor_precedence只在 typed settled lane 内复用selected_todo_projection,同时should_run=false、must_attempt_work=false、无执行命令。completed lane 不可重建时回执身份仍在,工作投影不制造出来。
对主干的风险
上轮 P2 已修复:新文档将无条件 selected 字段承诺改为当前 lane 可重建条件,并增加真实 CLI completed 生命周期回归。本轮九个配对场景全部通过:六个来源 scope,加 open/completed/28 条其他 Agent 工作的 monitor 情况。相同 fixture f7487f3bb7183441ffecaba26756c3776c2e5c8ef7bc876ed7d206634a646323;base/head 观察摘要 8e447d276258a4c11b9b10552ceaa0a8dd2bf1410df63f908d36311f14f8c8f4 / 1a764f896c9a838709232d6abca73f00c2184417a2064eb45ea8f825956f9b1e。合法 builder 和 all_registered future 在基础版错误拒绝、当前 head 成功;相同 Goal 的越权目标和未注册 Agent 仍拒绝。completed 的 receipt 原身份、重复 poll replay、journal 不变和无第二次扣额全部成立。
独立验证:274 Python(163.26s)、124 原生 TS、tsc、76 semantic/IO 检查、Ruff、配置内 19 文件 mypy、固定不可变 base 的真实 CLI budget smoke、diff hygiene 均通过。最初 companion 的 3 项 census 失败和 tsc 缺失来自此 worktree 未安装 npm dev dependencies;安装后完整 7 census 项与 tsc 通过。早先误写两个 TS 文件名只运行了 8 项,已用实际存在的三个文件重跑 124 项;没有删除断言或提高预算。全仓库套件、native Windows、真实 provider 和安装 App/模型采纳未测。
语义与 CI 对齐
复用现有 source scope、GoalRef、ReceiptBoundMonitorPhase 与 effective action,没有新共享 vocabulary 或 substring 分类。先跑 diff advisory(0 个支持语法候选,动态构造不在证明范围),再跑 semantic 与 registry IO 检查。默认来源/历史投影变化已在协议、fixture 和 PR 披露,未宣称 opt-in;不同来源观测共用一个 typed 决策,永久历史不授予当前执行。工作流的 event/base SHA 对照通过现有 budget smoke,未读取 CI。公共候选扫描无私有状态、凭据或本机路径。
我的整体评价
APPROVE。这次来源恢复与诚实结算读回的有界任务已满足;long_horizon preserved,user_experience improved。当前整 PR 复用既有 owner,改动与真实失败成本相称,上轮阻塞解决;未来重构检查认为已有共享 adapter/projection 足够,无须新增权限或历史状态源。历史 compatibility 保留于真实 stored receipt/instance 边界,未扩展 actor 权限。实际外部 provider、安装 App/模型收益仍是未测边界,批准不表示 whole Goal 完成或取得运行时合并授权。发布后独立执行旧 review closeout,撤回成功以 DISMISSED 原生读回为准。
English verdict: APPROVE - 5567@a952b9aab1acc40bbdb8593f59a584aa580beca9; the previous completed-monitor contract finding is fixed, legal exact-source recovery and no-effect settlement remain independently validated. 274 Python, 124 TS and current native checks pass; CI not consulted.
Problem and outcome
The previous shard-2 run failed 14 cases: nine exact-Goal context return cases, three settled-monitor recovery cases, one source fingerprint oracle, and one failed-Turn session identity fixture. This PR now integrates their shared repairs into one candidate instead of requiring an unmerged dependency to make those cases pass.
Changes
selected_todoandagent_lane_next_actionafter settlement, while retaining skip/no-run/no-spend authority. Update the protocol and existing replay assertions to explain this readback.The exact-target, settled-monitor and fixture repairs reuse work from #5533 by @Duang777. This candidate does not incorporate that PR's process-supervisor changes. Its remaining independent fixes still need their own acceptance.
Validation
Source: latest main
1af7dbd43plus this PR, candidatefcfb24c65.module_metric_budget:loopx/extensions/lark/goal_topic_runtime.py. The same failure was reproduced in an untouched1af7dbd43checkout; neither that module nor its ceiling is changed here.1af7dbd43, semantic inventory/census, configured mypy (19 source files), CI Ruff scope, diff and public-boundary checks passed. No budget allowance was raised.Future-facing pass: share the source-grant adapter while leaving policy decisions in the existing typed owner; reuse the existing selected-Todo projector rather than another rule implementation. Runtime/API changes require maintainer review on the new head.
Historical scope
The audit covered all 100 merged PRs by the submitting author: 262 failed pull-request runs across 67 PRs, 494 failed job logs, and 418 pytest node IDs recurring across at least two PRs. This fixes the identified shared causes above; it does not claim every historical failure is a single bug or that all CI is green. Fifteen no-job runs expired awaiting maintainer approval, and 19 logged DCO failures lacked sign-offs; their checks remain enforced.