Repository navigation
✨ feat: guide proactive reviewer investigation - #48
Conversation
Make the main reviewer own its working understanding, choose evidence that can change a conclusion, and revisit affected conclusions when requirements change. Add an unhinted batch-export evaluation with removal, correction, and retention witnesses. Co-authored-by: Codex <noreply@openai.com>
Distinguish cross-batch regression inputs from ineffective same-batch cases, and treat cross-source checkpoint guards and CLI compatibility as contract questions unless requirements support them. Co-authored-by: Codex <noreply@openai.com>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Preserve current main guidance on human feedback and memory retrieval alongside proactive investigation and requirement revalidation. Co-authored-by: Codex <noreply@openai.com>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b38d8b8e41
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| @@ -0,0 +1,53 @@ | |||
| # Batch export evaluation | |||
There was a problem hiding this comment.
Add the required Codex co-author trailer
This commit was produced by Codex, but its message contains no Co-authored-by trailer. Add Co-authored-by: Codex <noreply@openai.com> to the final commit message before merging, as required for coding-agent-assisted commits.
AGENTS.md reference: AGENTS.md:L27-L27
Useful? React with 👍 / 👎.
将主 reviewer 的调查职责写清楚:建立对需求和调用链的理解,选择能改变结论的证据,通过追踪、委派或最小实验验证疑点,并在证据充分时结束调查。继续、恢复及发布前重新核对需求和相关讨论,避免仅因代码版本未变就复用已失效的结论。
补充一个不向 reviewer 提示缺陷的批量导出评测,分别验证应删除的内容去重账本、应修正的 checkpoint 提交顺序,以及必须保留的接收端重试语义。评分说明区分有效的跨批次复现、无效的同批次测试,以及缺乏需求依据的额外防御建议。
验证均在现有容器内完成:
glm-5.3对新旧指引各做两轮隔离运行。首轮均在 15 分钟超时;第二轮双方增加相同的报告时间预算提醒,新版 11 分 22 秒、31 次工具调用,旧版 12 分 30 秒、44 次工具调用。单组样本不能证明速度或质量提升。质量与成本仍待验证:双方都发现了核心问题,但旧版对保留边界的实验证据更强,新版最终建议中有一个无法复现缺陷的测试输入,双方也都有超出明确前提的防御建议。该离线实验未覆盖完整 GitHub 发布、托管子任务及独立设计生命周期,因此不能作为端到端质量提升或耗时可接受的证明。