coding-review-agent-loop is a local command-line orchestrator for GitHub code
review. One coding agent creates or updates a pull request, one or more other
agents review it, and the loop sends blocking feedback back to the coder until
the reviewers approve or the run reaches a clear stopping condition.
GitHub issue, task, or PR
|
v
coding agent <-------+
| |
v |
pull request |
| |
v |
reviewers ---- feedback
|
v
approved -> optional CI wait and merge
The loop runs on your machine and uses the local claude, codex, agy,
gemini, and gh programs you have already authenticated. It does not require
you to put model API keys into this project. You only need the agent CLIs used
for the roles you select; you do not need to install every supported backend.
The project is alpha software. It can let coding agents edit repositories, run commands, push branches, and write to GitHub. Start with a repository where you can inspect and revert the results.
- Replace the manual cycle of copying reviewer feedback between agent sessions.
- Assign coding and review to different models or providers.
- Start from a GitHub issue, an existing pull request, or a free-form task.
- Review an implementation plan before allowing code changes.
- Require multiple independent reviewers to approve the same PR head.
- Resume interrupted work from durable metadata recorded on GitHub.
- Optionally run local tests, wait for CI, and merge after approval.
The default workflow is deliberately conservative: Claude is the coder, Codex is the reviewer, the review limit is 10 rounds, and automatic merge is off.
- Python 3.11 or newer.
- Git and GitHub CLI with
gh auth statussucceeding. - Repository access sufficient for the requested issue, branch, PR, and comment operations.
- A local CLI for the coder and each reviewer you select:
Claude Code,
OpenAI Codex CLI,
Antigravity CLI (
agy), or the legacy Gemini CLI backend.
Each agent CLI has its own authentication, quota, terms, and model availability. Confirm those with the provider; this tool does not combine or replace provider subscriptions.
Agent subprocesses and repository test gates use a shared per-user containment
policy by default. On Linux with systemd 253+ and delegated cgroup v2,
agent-loop creates a foreground child scope inside agent-loop.slice and
applies aggregate plus role-specific MemoryHigh, MemoryMax,
MemorySwapMax, and TasksMax limits. The aggregate protects host headroom
across independent agent-loop processes; a child scope cannot exceed it. The
portable fallback uses process-group TERM/KILL for deterministic termination
but provides no memory ceiling. --containment-mode required fails closed
when cgroup preflight is unavailable, while auto reports the fallback.
Inspect the resolved policy and capabilities with
agent-loop containment-preflight --containment-mode auto. The systemd
launcher is the synchronous foreground form
systemd-run --user --scope --quiet; it deliberately does not use --wait,
--service, or --pipe. A target-start shim report is authoritative when
distinguishing launcher failures from target exit statuses. Optional cgroup
telemetry (peak memory, PSI, and swap counters) is capability-aware: missing
files are reported as not collected rather than treated as lost evidence.
OOM, hard memory/swap failures, and task-limit failures are
resource-exhausted; MemoryHigh/PSI pressure is diagnostic only.
The same-command test wrapper has a per-invocation lane lock. A duplicate is rejected before spawn until the earlier managed attempt exits or is explicitly terminated. Standalone commands not run through the wrapper cannot be deduplicated from free-form agent logs, but their complete agent tree remains resource-bounded when launched by agent-loop. The aggregate is per user manager, not cross-user host isolation. A skill host's in-session Claude turn and arbitrary descendants remain part of that host session; use the managed wrapper for test gates and external skill agents for the mechanical boundary.
In skill mode, the test-gate policy is selected with
AGENT_LOOP_CONTAINMENT_MODE=auto|required|off (default auto).
Clone the repository and install it into a virtual environment:
gh repo clone wwind123/coding-review-agent-loop
cd coding-review-agent-loop
python3 -m venv .venv
. .venv/bin/activate
python -m pip install -e .
agent-loop --helpCheck the CLIs for your chosen roles before the first run:
gh auth status
claude --version
codex --versionSubstitute agy --version or gemini --version when using those backends.
This is the smallest useful first run. The reviewers inspect the current PR; if they find blockers, the coder updates that same PR and review continues.
agent-loop pr 456 \
--repo OWNER/REPO \
--coder claude \
--reviewer codexIssue mode gives the issue title, body, and comments to the coder, validates the resulting PR, and then enters the same review loop.
agent-loop issue 123 \
--repo OWNER/REPO \
--coder claude \
--reviewer codexWithout --plan-first, issue mode asks the coder to implement immediately.
Use plan-first mode for work whose design should be challenged before files are
changed. --plan-first alone stops after plan approval. Add
--implement-after-approval to continue into implementation and PR review.
agent-loop issue 123 \
--repo OWNER/REPO \
--coder codex \
--reviewer claude \
--plan-first \
--implement-after-approvalagent-loop task "Add a health-check endpoint" \
--repo OWNER/REPO \
--coder codex \
--reviewer claude| Command | Use it when |
|---|---|
agent-loop issue |
A GitHub issue defines the work to implement or plan. |
agent-loop pr |
The implementation PR already exists. |
agent-loop task |
You have a direct task and do not need an issue first. |
agent-loop discuss |
You want agents to evaluate an open question without writing code. |
agent-loop managed-pr |
Code is pushed but no PR exists, and the repository uses managed exact-head CI. |
agent-loop managed-ci preflight |
You want a read-only readiness report for managed CI. |
Run agent-loop <command> --help for the complete options for one workflow.
The full CLI guide covers lifecycle and resume
behavior in detail.
- The coder implements the issue or updates the existing PR.
- Every configured reviewer reviews the same PR head.
- Blocking findings return to the coder as an explicit work ledger.
- The updated head is reviewed again.
- The run stops on unanimous approval, a terminal blocker, a clarification request, unavailable required input, or the round limit.
Repeat --reviewer to require multiple approvals. Use --review-parallel when
the reviewers have distinct workdirs and may run concurrently:
agent-loop pr 456 \
--repo OWNER/REPO \
--coder codex \
--reviewer claude \
--reviewer agy \
--review-parallelAgent-loop creates separate repo-scoped temporary checkouts for active agents
unless you provide workdirs. GitHub comments carry durable round and handoff
metadata, so a later run can reconstruct the active review state. When the PR
number is known, resume with agent-loop pr <number> instead of starting issue
implementation again.
Signed comments ending in -- Human Reviewer are treated as explicit human
requirements and remain approval-critical. See
Human requirements for the exact
contract.
Agent-loop publishes model reviews as structured comments in the pull
request's conversation. An Approved verdict means that the named model
approved the recorded PR head under agent-loop's protocol. It is not a native
GitHub pull-request review, does not populate GitHub's Reviewers or
Reviews panels, and does not satisfy a branch-protection rule that requires
approving GitHub reviews.
The agent CLIs normally share the GitHub identity authenticated through gh,
so their model signatures identify protocol participants rather than distinct
GitHub accounts. If a repository requires native GitHub approvals, obtain them
separately from an eligible human, bot, or GitHub App identity. GitHub's own
merge protections remain in force in addition to agent-loop's reviewer and CI
gates.
- Run only one active
agent-loopinvocation per repository per machine. The default workdirs and repo-scoped local state are shared, and the tool does not currently enforce a repository-wide process lock. Separate concurrent runs against the same repository can interfere with each other's checkouts and artifacts.--review-parallelis supported within one orchestrator run; it does not make multiple same-repository invocations safe. - Agent-loop is a local process, not a hosted service. The machine must remain
available for the run, and an interrupted in-flight agent turn may need to be
repeated. Once a PR exists, resume with
agent-loop pr <number>. - GitHub is the only supported forge, and every selected agent backend must be installed and authenticated locally.
- Agent CLIs can exhaust quota, time out, update themselves, or return malformed structured output. Retries, repair passes, and salvage reduce lost work but cannot guarantee unattended completion.
- Default agent checkouts live under the system temporary directory (
/tmpon Linux) and may disappear after a reboot or system cleanup. Configure explicit workdirs for long-lived installations.
Plan-first mode supports four post-approval choices:
| Mode | Result after plan approval |
|---|---|
plan-only |
Post the approved plan and stop. This is the default. |
implement-one-shot |
Implement the approved plan in one PR. |
decompose-only |
Create detailed child issues for the approved phases and stop. |
implement-by-phase |
Create the phase issues and implement only the first phase. |
Example:
agent-loop issue 123 --repo OWNER/REPO \
--plan-first \
--plan-execution-mode decompose-only--plan-execution-mode decompose-only and --materialize-split-issues are
different mechanisms. Do not combine them for the same decomposition: doing so
can create duplicate children. Use the former for detailed approved phases and
the latter for discuss-mode split proposals or eligible plan-only deferred
work. Read
Phased decomposition versus split materialization
before filing child issues.
When approved future follow-ups are summarized or filed, semantic reuse is
enabled by default after deterministic narrowing. Use
--no-semantic-followup-dedupe for deterministic-only operation. The provider
and its bounds are configurable with --semantic-followup-backend,
--semantic-followup-model, --semantic-followup-timeout-seconds,
--semantic-followup-max-calls, --semantic-followup-max-candidates, and
--semantic-followup-prompt-char-limit. Only high-confidence matches suppress
or merge work; uncertain matches are filed with a possible-duplicate note.
Discuss mode asks agents to evaluate an issue without modifying the repository. Use it for architecture choices, product decisions, feasibility questions, or whether work should be implemented or split.
agent-loop discuss 123 \
--repo OWNER/REPO \
--reviewer claude \
--reviewer codexThe default result contract is implementation triage: implement,
do-not-implement, needs-human, or split. For an open-ended recommendation
instead of an implementation vote, use --discuss-result-mode answer.
Useful optional controls include:
--discuss-analyzer AGENTfor a structured consensus/disagreement agenda.--discuss-research auto|required|nonefor current external facts.--discuss-parallelfor concurrent independent positions.--materialize-split-issuesto file agreed split proposals.
See Discuss mode for result semantics, research evidence, deadlocks, and resume behavior.
Answer-mode summaries now put a bounded executive state before the audit
transcript. With a configured --discuss-analyzer, completed non-final rounds
reuse the analyzer's enriched agenda when available and show cumulative current
consensus, active disagreements, changes, missing facts, and next-round focus.
The final comment similarly leads with outcome, agreed conclusions, residual
decisions, and the next action. Exact, semantic-equivalent, and debater-confirmed
results reuse their mechanically validated artifacts without a redundant final
synthesis call; only the configured discuss analyzer may perform the explicitly
bounded fallback or final synthesis call. Analyzer-less or invalid synthesis
falls back to the existing fail-closed result. Per-agent comments remain the
authoritative raw audit, and resume metadata carries only the latest validated
snapshot with bounded, spillable excerpts.
| Backend | CLI | Notes |
|---|---|---|
| Claude | claude |
Default coder. Select a model with --claude-model. |
| Codex | codex |
Default reviewer. Select a model with --codex-model. |
| Antigravity | agy |
Accepted as agy or antigravity; supports coder and reviewer roles. |
| Gemini | gemini |
Legacy, best-effort path for accounts that still have CLI access. |
The default Antigravity model chain is Gemini 3.7 Flash (High), then
Gemini 3.6 Flash (High), then Gemini 3.1 Pro (High) for eligible capacity
failures. Override it with --antigravity-model or
--antigravity-models. Antigravity turns are single-shot and its usage totals
are estimated because agy does not expose token counts.
Backend-specific authentication, model selection, fallback, timeout, and executable-replacement behavior are documented under Agent backends.
Agent-loop owns the effort setting for Codex and Claude. When no effort option
is supplied, every invocation explicitly receives medium; a local CLI config
file or inherited environment value cannot silently change that selection.
The precedence is the matching role override (reviewer or implementation), agent-wide option,
then the tool default. Use --codex-reasoning-effort xhigh (or the matching
implementation option) for an explicit higher-effort Codex/Luna run. Claude
accepts low, medium, high, xhigh, and max through --claude-effort.
Antigravity's selected model already carries its tier, such as
Gemini 3.7 Flash (High), and is not rewritten by this setting. Startup logs,
signatures, usage records, and new round metadata distinguish configured effort
from verified runtime observations; older records retain unknown values.
When Codex or Claude occupies both seats, select its reviewer independently:
agent-loop pr 123 --repo OWNER/REPO \
--coder codex --reviewer claude --reviewer codex \
--codex-model gpt-5.6-luna --codex-reasoning-effort xhigh \
--reviewer-codex-model gpt-5.6-sol --reviewer-codex-reasoning-effort mediumClaude has matching --reviewer-claude-model and --reviewer-claude-effort
options. These overrides apply to plan reviews, PR reviews, and discussion
participants, not the coder or discussion analyzer. Model and effort are
independent: omitting either retains that setting's existing fallback.
Reviewer overrides survive the issue-to-PR implementation handoff; repeat
them when starting a separate resume command. See the
issue-mode example and precedence.
Agents can run commands and change code. Keep their normal permission prompts unless you understand and accept the repository and machine-level risk.
For a trusted local environment, this flag supplies each backend's permission bypass option:
agent-loop pr 456 --repo OWNER/REPO \
--coder codex --reviewer claude \
--dangerous-agent-permissionsThe flag is intentionally explicit. It does not make agent output, fetched issue text, dependencies, shell commands, or generated code trustworthy.
Other important boundaries:
- Automatic merge is off unless
--auto-mergeis present. - Reviewer approval is not a substitute for project tests or human judgment.
--test-commandadds a local gate before review and again before auto-merge. The finite watchdog defaults to 1,800 seconds and is configurable with--coder-test-command-timeout-seconds SECONDS.- The tool validates assigned workdirs and reported test locations, but agent CLIs may still consume substantial CPU, memory, network, and provider quota.
- Raw subprocess logs and salvage artifacts can contain sensitive repository context. Protect access to the configured log directories and review their retention settings.
Read Workdirs, Agent permission flags, and Logs before unattended use.
Coder test commands may be run through the backend-neutral wrapper:
agent-loop run-tests --memory-dir /path/to/memory -- pytest tests/test_app.py -qThe --timeout-seconds value is the watchdog for that one whole command. When
omitted, the wrapper uses the inherited run ceiling, or 1,800 seconds when run
outside agent-loop. A positive finite override may be smaller than the ceiling;
values above it are rejected before the child starts. Agent backends inherit the
ceiling through AGENT_LOOP_CODER_TEST_TIMEOUT_CEILING_SECONDS.
When agent memory is enabled, the wrapper records measured outcomes, elapsed
time, the attempted cap, a privacy-preserving environment fingerprint, and
cheap lockfile/configuration hashes in test-runtime.json. Recent successful
runs produce advisory median/p95 recommendations with headroom; timeouts remain
lower-bound evidence and are never treated as successful durations. Data is
best-effort, retained to 20 samples per command/fingerprint cohort and 200
cohorts, and becomes stale after 30 days or when relevant inputs change.
Remembered commands are suggestions only: agents must inspect the checkout and
select focused tests. Framework per-test limits, the wrapper whole-command
watchdog, and the backend whole-turn timeout are separate. The backend turn
must leave headroom for analysis, edits, and reporting; split or shard healthy
browser/integration matrices when that improves diagnosis and retry cost.
Use --auto-merge only when the repository's CI and branch protections are
appropriate for unattended merging:
agent-loop pr 456 --repo OWNER/REPO --auto-mergeFor ordinary CI, auto-merge waits for a reliable, non-empty check board on the
current head. Without auto-merge, --watch-pending-ci can wait and report that
an approved PR is merge-ready without merging. Set the total watcher budget
with --ci-timeout-seconds (default 1200) and its polling interval with
--ci-poll-interval-seconds (default 30). GitHub runner stalls are bounded by
--ci-queued-grace-seconds; see
External CI infrastructure stalls.
Managed CI is an advanced, repository-integrated workflow that suppresses expensive intermediate CI and qualifies one reviewed SHA at the end. Do not enable it from a README example alone. First read Managed exact-head CI and run the read-only preflight:
agent-loop managed-ci preflight \
--repo OWNER/REPO \
--base main \
--trusted-actor LOGINFor code already pushed without an open PR, the pre-creation form begins with
agent-loop managed-pr --head BRANCH. Managed issue and PR recovery relies on
a canonical issue handoff, explicit --managed-ci intent, documented
draft/labeled and ready/unlabeled lifecycle states, immutable actor evidence,
and preserved base provenance. Historical records are audit evidence only and
never grant fresh authority.
Qualification and merge remain bound to the live head. The final merge uses
--match-head-commit and merges only that qualified SHA. --watch-pending-ci
and --no-watch-pending-ci do not alter managed exact-head qualification.
The repository also contains a Claude Code skill for running the orchestration inside an attended Claude Code session. In skill mode, Claude acts in the current interactive session while external agents still run through their local CLIs.
Use the standalone CLI for predictable or unattended runs. Use skill mode when you want conversational setup, active steering, and interactive recovery. Skill mode never auto-merges.
See SKILL.md for invocation instructions and
docs/skill_mode.md for its design and limitations.
agent-loop --help: full command and option reference.docs/local_agent_loop.md: architecture, lifecycle, protocol, recovery, CI, memory, and safety reference.docs/skill_mode.md: Claude Code skill architecture and operation.SKILL.md: executable instructions for Claude Code skill mode.
The detailed guide is intentionally the source for protocol schemas, durable markers, repair passes, fallback ladders, CI provenance, and compatibility behavior. Those internals are not required for a first successful run.
Install the development dependency and run the tests:
python -m pip install -e '.[dev]'
python -m pytestUse focused tests while changing one subsystem, for example:
python -m pytest tests/test_docs_guidance.py
python -m pytest tests/test_protocol.py
python -m pytest tests/test_orchestrator_pr.pyTests use fake subprocess runners and do not invoke real agent CLIs or GitHub.
Browse the focused test modules in tests/ and see the architecture
diagram in docs/local_agent_loop.md.
This project is a standalone local GitHub lifecycle orchestrator. Projects such as claude-review-loop, codex-review, and codex-plugin-cc integrate review or delegation into a particular agent host. Here, the orchestrator stays outside the agent hosts and can reverse coder/reviewer roles.