Skip to content

fix(zitadel): restart env-at-start OIDC consumers after a client rotation - #2149

Merged
Smana merged 1 commit into
mainfrom
fix/restart-oidc-consumers-after-rotation
Oct 1, 2026
Merged

Smana merged 1 commit into
mainfrom
fix/restart-oidc-consumers-after-rotation

Conversation

@Smana

@Smana Smana commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Why

After an OpenBao rebuild, zitadel-oidc-clients.sh sync rotates the ZITADEL OIDC client (#2078), but a Deployment that reads its client id from env keeps the old id until its pods restart. headlamp-oauth2-proxy (namespace tooling) showed "App not found" until someone ran kubectl rollout restart deployment headlamp-oauth2-proxy -n tooling, and every resource reported healthy the whole time.

What changes

restart_rotated_consumers runs at the end of cmd_sync, right after force_sync_mirrored. That covers every stage-3/4 call site on both clouds (aws/eks/init stage 4, gcp/gke/init stage 3, hosting and consuming), with no workflow change.

flowchart LR
    loop["consumer loop"] -->|"stored client id != ZITADEL's"| rot["rotated: key=new id"]
    rot --> ctx{"kubectl on --cluster?<br/>(*-CLUSTER-vars CM)"}
    ctx -- no --> warn["WARN, list keys"]
    ctx -- yes --> es["force-sync every ExternalSecret<br/>reading the key"]
    es --> wait{"Secret carries new id?<br/>(one 180s deadline)"}
    wait -- no --> warn2["WARN + manual command"]
    wait -- yes --> rs["rollout restart each Deployment<br/>with env/envFrom on that Secret"]
Loading
Property How
Scoped to real rotations A key counts as rotated only when it held a client id and this run replaced it with a different one, on the create path or the converge path. A first bootstrap (no previous id) restarts nothing, because those pods start once the Secret appears.
Idempotent A re-run with nothing rotated passes no keys and makes no kubectl calls.
Right cluster only Requires kubectl to see flux-system/*-<cluster>-vars. aws/eks/init's consuming sync for gcp-0 runs with kubectl on aws-0, so it only warns there.
New pods read the new id Force-syncs each ExternalSecret reading the key, then waits until the Secret holds the new id before restarting.
Readers found, not listed Deployments in the Secret's namespace that reference it via env[].valueFrom.secretKeyRef or envFrom[].secretRef. Volume mounts are excluded.
Non-fatal Warn-only, like force_sync_mirrored: the clients are already written.

ExternalSecret matching covers both store shapes: OpenBao at the key's bao-map.sh path (headlamp-envvars, grafana-envvars) and clustersecretstore under the bare key (the unmapped headlamp-oauth2-proxy).

Env-at-start readers this reaches today:

Cluster Deployment Secret
gcp-0 tooling/headlamp-oauth2-proxy headlamp-oauth2-proxy (secretKeyRef)
aws-0 tooling/headlamp headlamp-envvars
both Grafana, observability victoria-metrics-k8s-stack-grafana-envvars (envFromSecret)

Not covered: Harbor and the Flux UI take their client through HelmRelease valuesFrom, so a refresh reaches them through a Helm upgrade and not through a Deployment env reference. OpenBao is reconciled through its API by reconcile_openbao_oidc.

Known gap

With AWS as primary, gcp-0's suffixed clients are rotated by aws-0's stage 4, where kubectl points at aws-0. That run only warns. gcp-0's own stage 3 runs the sync later, sees the store already converged, and restarts nothing. Closing this needs the AWS stage to reach gcp-0's kubeconfig, which is out of scope here.

Evidence

  • New scripts/ci/tests/test-zitadel-oidc-clients-restart.sh: 14/14 ok. It runs the shipped function bodies against a kubectl stub. Two mutants are killed: dropping the context guard, and dropping the Secret-refreshed check.
  • test-zitadel-oidc-clients-convergence.sh gains a rotation case: headlamp-envvars=existing-client-id reaches the restart, while a same-id converge passes nothing.
  • bash scripts/ci/tests/run.sh: 40 passed, 1 skipped (vector not installed), 0 failed
  • shellcheck -x scripts/provision/zitadel-oidc-clients.sh: clean
  • validate-manifests.sh (MemoryMax=6G): exit 0, Valid: 2096, Invalid: 0, "All gates passed"
  • validate-doc-claims.sh: 30 claims match. validate-links.sh and verify-doc-paths.sh: clean
  • trivy config --exit-code=1 --ignorefile=./.trivyignore.yaml scripts/: exit 0
  • tofu validate: not applicable, because no .tf or .tm.hcl changed (the change is in the script the stages call). No terramate or tofu apply was run.

No ADR: this is a fix, not a technology choice.

…tion

After an OpenBao rebuild the sync rotates the ZITADEL OIDC client (#2078),
but a Deployment that read its client id from env at start keeps the dead
one: headlamp-oauth2-proxy answered "App not found" until restarted by hand.

cmd_sync now records each store key whose stored client id it replaced and,
with kubectl on --cluster, force-syncs the ExternalSecrets reading it, waits
until the Secret carries the new id, then rollout-restarts the Deployments
referencing that Secret through env or envFrom. A first bootstrap or a re-run
with nothing rotated restarts nothing.
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

🔍 Rendered manifest diff — this PR vs main (desired state)

No changes to the rendered desired state. ✅

@Smana
Smana merged commit 8385201 into main Oct 1, 2026
12 checks passed
@Smana
Smana deleted the fix/restart-oidc-consumers-after-rotation branch October 1, 2026 07:06
Smana added a commit that referenced this pull request Oct 7, 2026
…P2 phase 6) (#2226)

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(rooms): roomctl's ZITADEL client, oauth2-proxy bearer tokens, pin AP-6 (SP2 phase 6)

roomctl is a native ZITADEL app (no secret, device code and refresh token,
JWT access tokens) whose id reaches OpenBao's agents/roomctl. The broker reads
it from a file to tell roomctl's tokens from the web UI's (ruling P18), and
oauth2-proxy lets those bearers through against roomctl's own client. A native
app gets no secret back, so its create and its missing-key paths no longer
fail. The broker, its retention CronJob, crd-rooms.yaml and atlasSchema.ref
move to agent-platform#28 together (P33).

The rooms OIDC suite also loads stored_client_id and stubs
restart_rotated_consumers: its cmd_sync section had exited 127 since #2149
reached this stack.

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(rooms): force the proxy's reconcile after the roomctl sync

A reconcile of rooms-oauth2-proxy before zitadel-oidc-clients.sh has
created the roomctl client leaves the new pod waiting for
rooms-proxy-extra-issuers. After three failed upgrades Helm rolls back
to skip-jwt-bearer-tokens=false and the release stalls, and nothing
retries it once the Secret appears. The ordering note and Task 6.4 now
end with flux reconcile hr rooms-oauth2-proxy -n agent-system --force.

Review I2.

* test(rooms): load previous_client_id in the rooms sync harness

main's cmd_sync now reads the old client id through previous_client_id
(#2208's OpenBao fallback). The rooms harness extracts functions by name,
so after the merge every cmd_sync case exited 127. It now loads that
function and stubs mirror_read, as the openbao suite does.

* fix(rooms): pin the broker to AP-6's review fixes (agent-platform#28 at 8dcf11d)

The broker and the retention CronJob move to
v0.0.1-pr28.8dcf11d2@sha256:40a38465…: the fork's own rate and prefix
caps, roomctl's final errors and the equal-client-id guard. The CRD is
unchanged at 8dcf11d; its header names the new commit.

* fix(rooms): pin the broker to AP-6 with the owner's fork rulings (agent-platform#28 at 91a6fb3)

The broker and the retention CronJob move to
v0.0.1-pr28.91a6fb39@sha256:2889162f…: a fork by anyone but the source's
owners takes the stricter approvals (M2), and a sealed room forks only
for its owners and agents-admin (M3). The CRD is unchanged at 91a6fb3;
its header names the new commit.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>
Smana added a commit that referenced this pull request Oct 8, 2026
* feat(openbao): agents-secrets policy and JWT role scoped to platform/agents

* feat(agents): agent-sandbox controller from its pinned git chart

* feat(github): idempotent applier for the agents' branch ruleset

* fix(ci): render oidc_issuer_host with its /id path, as the live value has it

* feat(agents): gVisor pool, RuntimeClass and shared sandbox config

* feat(security): self-hosted octo-sts for the agents' GitHub App

* feat(security): Kyverno admission and GC for agent runs

* chore(scripts): throwaway identity probe for agent-router checks

* fix(github): bypass the write and maintain roles in the agents' branch ruleset

* fix(agents): drop the unpinnable digest from agent-sandbox's image tag

* feat(observability): Vector tolerates the agents gVisor pool

* feat(agents): agent-harness image on OpenHands agent-server 1.49.5

The AgentRun sandbox's harness: agent-server plus the agent-run driver,
git-credential-agent, gh, and a commit-msg hook adding Agent-Run: $RUN_ID.
The published agent-server image is upstream's PyInstaller binary, so the
SDK is installed into /agent-server/.venv from a hash-pinned universal lock.

Deviations from the plan's code, each found against the running server:
- agent-server binds 127.0.0.1 explicitly (plan P13, CC-1 composition),
  not 0.0.0.0: its API is unauthenticated.
- OH_ENABLE_VSCODE=false: otherwise agent-server starts openvscode-server
  on 0.0.0.0:8001.
- build_request dumps with expose_secrets: the SDK redacts the placeholder
  model key and drops it on reload, so litellm refused to call at all.
- A SIGTERM handler raises SystemExit, so pod deletion still runs the
  finally block that revokes the GitHub token and stops agent-server.

docker build --target test: 10 tests OK.

* feat(security): agents-secrets store for agent-system

* feat(security): reach octo-sts only through agent-router's sts listener

* feat(agents): agent-router Gateway with one JWT listener per data class

* fix(security): exempt octo-sts's key-path variable from Polaris's secret-name check

* fix(docs): drop the ADR-0041 reference to an unmerged spike doc path

* feat(observability): agent-platform VMRules and dashboard

Task 6.4 content. The umbrella child that applies it follows in its own commit,
once PR 2's agent-platform umbrella exists on this branch.

* feat(clusters): apply the agent-platform VMRules and dashboard behind the umbrella

Wires the agent-observability child Kustomization (observability/base/agent-platform)
into the aws-0-agent-platform umbrella, so it resumes with the rest of the platform.

* feat(agents): read-only flux-operator-mcp with a narrow ClusterRole

* feat(agents): VictoriaMetrics and VictoriaLogs MCP servers for agents

* feat(agents): per-class MCPRoutes with role-scoped tools

* fix(agents): fold in review findings for the agent-harness image

Cleanup order in agent_run.py's finally now stops agent-server before
revoking its GitHub token, and ignores a second SIGTERM so the revoke
can't be aborted. Disable autotitle to save a model call per run. Pin
agent-server's cwd to "/" so its state lands under /workspace whatever
the pod's workingDir is. Status polling tolerates up to 5 consecutive
network errors before failing, and logs the final execution_status.

git-credential-agent: revoke() no longer tracebacks on an already-dead
token (logs the status instead), a token being replaced on refresh is
revoked before the cache is overwritten, and concurrent get/token
calls serialize on an flock of the cache file.

Dockerfile: strip setgid alongside setuid, pin GIT_TERMINAL_PROMPT=0,
and install the locked requirements with --only-binary=:all: (with a
targeted exception for func-timeout, a pure-Python, wheel-less
transitive pin from openhands-tools -- verified every other locked
package ships a wheel for both amd64 and arm64).

* fix(ci): gate agent-router listener pinning and route-level SecurityPolicy merges

A5 in assert-ai-gateway.py: every route naming agent-system/agent-router
sets sectionName (without one it attaches to public, internal and sts),
and every SecurityPolicy on such a route sets mergeType (without one it
replaces the listener's JWT check). Unresolvable targets and label
selectors are assumed in scope; zero routes on the Gateway fails rather
than passing vacuously. Other Gateways' routes may still omit sectionName.

* fix(agents): let the identity proxy outlast agent-models' 600s request timeout

Both model listeners set stream_idle_timeout 300s and no route overrides
it, so a non-streamed GLM-5.2 completion taking 300-600s was reset by the
proxy before the router's 600s budget ran out. Raised to 900s. The route
comment no longer claims the proxy sets no timeout, and records that EG
v1.9.1 derives the router's per-route idle timeout as max(1h, request)
while no ClientTrafficPolicy sets streamIdleTimeout.

* fix(github): pin the ruleset fields that make it a control

The test asserted the bypass list, conditions and rule types but never
the name, target or enforcement, and never the update body. An edit to
`enforcement: evaluate` or `disabled`, `target: tag`, a renamed ruleset
or an empty PUT body all stayed green; each now fails, on both the
create and the update body.

The applier reads the ruleset name from the JSON instead of hard-coding
it, lists only the repository's own rulesets (includes_parents=false),
takes the first match, and warns when an update would drop an App from
the bypass list, so a re-run without FACTORY_APP_SLUG after SP3 is not
silent. Its header says to run it before the agents' App is installed.

* fix(security): keep octo-sts's Polaris exemption in .polaris.yaml

SPEC-007 keeps every exemption in .polaris.yaml, by controller and
rule, with its reason; the annotation on the Deployment was the only
one in the repo and invisible from that list. Same rule, same scope:
octo-sts, sensitiveContainerEnvVar. The section header now admits
false positives on a name the workload fixes, not only privilege.

* docs(adr): ADR-0043 says what is recorded and puts the ruleset first

octo-sts 0.10.0 puts the issuer, subject and token SHA-256 only in its
exchange event, which needs METRICS=true and a CloudEvents sink; neither
is set, so its logs carry the requested repository and trust policy and
nothing that ties a token to a run. The Pro now names what is recorded
(agent-router's sts access log, octo-sts's log, GitHub's own record of
what the App's bot did), and a Negative says what is not.

A repository opts in three times, ruleset first: main needs no
approval, so until the ruleset exists nothing stops the implementer
from merging its own green PR. The cluster README's Resume section
lists the owner prerequisites in that order, and what fails without
the App key.

Also recorded: contents:write reaches tags, releases and
repository_dispatch; PR CI still runs scripts an agent can edit; the
App key's exposure through openbao-platform (T14); a rotated key needs
a rollout restart; Dependabot is off on this repository.

* fix(security): close the TokenRequest and forged-owner gaps in the agent policies

- agent-audience-token-request: the TokenRequest API (serviceaccounts/token)
  minted the reserved agent audiences without a pod, so the pod rule alone
  left them open to `kubectl create token` and ESO's cluster-wide grant.
- agents-pod-creator: ownerReferences are client-supplied; the requester
  identity is not. Admission-only, since `request` is absent in background.
- agentrun-gc: delete terminal runs only a day after they ended, falling
  back to creation time when a run carries no finishedAt.

* fix(agents): narrow the controller scrape ingress and correct review nits

- CNP: admit only vmagent on :8080, not the whole observability namespace.
- Resume steps: wait for the crossplane-configuration pin that serves
  AgentRun, since agent-policies reports Ready either way.
- ADR-0041: the spike notes are a superpowers spec, landing with #2092.
- gen-catalog.sh: drop the stale source counts.

* fix(agents): make the agent-sandbox chart and image lockstep mechanical

Renovate's flux manager tracks the GitRepository tag while the helm-values
manager could not see the controller image (no `repository` in values), so
the patch/minor automerge rule would have moved the chart and left the
binary behind. Restate the image repository so the image is tracked, and
group both into one PR that is never automerged. The HelmRelease comment
no longer claims a digest cannot be pinned (a postRenderer could) and says
the chart still installs the extension CRDs with extensions off.

* fix(ci): render oidc_issuer_url with its /id path, like oidc_issuer_host

The live issuer URL carries /id/<ID>; the bare-host fixture rendered the
agent-router SecurityPolicies with an issuer shape the cluster never has,
and left configuration-aws's EnvironmentConfig with an inconsistent pair.

* fix(agents): revoke under the same lock as token refresh

revoke() now holds CACHE + ".lock" when reading, revoking, and deleting
the cached token. This closes the race with concurrent token() exchanges
that could leave a live, unrevoked token in the cache if revoke() ran
without the lock.

The fix ensures that revoke() atomically reads the cache, revokes the token
at GitHub, and deletes the file — serializing with any concurrent token()
call via the same flock that token() uses.

Also clarify the Dockerfile comment to mention that both setuid and setgid
bits are stripped, matching what the find command actually does (-perm /6000).

Test: new test_revoke_waits_for_the_cache_lock verifies revoke() blocks
on the lock held by token(), preventing the interleaving that would leave
an unrevoked token in the cache.

Ran 19 tests, all pass.

* fix(agents): run agent-router on two replicas behind a PDB

With one replica, every drain or Karpenter consolidation of its node cut
all in-flight agent completions at once, and the harness does not retry.
envoyPDB minAvailable 1 keeps one proxy serving through a voluntary
disruption.

* fix(agents): drop agent-router's host ingress and name its rate-limit egress

The fromEntities: host rule on 8080/8081 admitted nothing the data plane
needs: EG's kubelet probes target the readiness and shutdown-manager
ports, which it never listed, and Cilium's default allow-localhost
already admits them. It only implied host-network pods were an intended
caller, which ADR-0042 says is agents pods alone.

The header now names the envoy-ratelimit :8081 egress SP4 must add with
its budgets: a dropped rate-limit check fails open.

* fix(security): retry agent-secrets after 30s, not a whole interval

On first apply the SecretStore can go Ready=False before openbao-ca holds
its CA; the failed health check then waited the 5m interval, holding
agent-router back by as much. 30s matches security-openbao.

* fix(security): call agents-secrets agent-system's store by convention

"agent-system's only way to secrets" read as a control. Nothing enforces
it: no ClusterSecretStore sets namespace conditions, and openbao-ca itself
reads through clustersecretstore.

* fix(scripts): let the identity probe's Sandbox expire if its delete is forgotten

agent-sandbox v1.0.3 keeps a terminated pod rather than recreating it,
but creates a new one, with fresh tokens, if that pod is ever removed.
shutdownPolicy Delete plus the documented shutdownTime patch bound the
Sandbox's life; the delete tolerates the Sandbox having expired first.

* fix(agents): audience mismatches are 403, and sts's audiences are the per-repo opt-in

Envoy's jwt_authn answers JwtAudienceNotAllowed with 403 and every other
failure with 401 (filter.cc at v1.39.1), so a cross-class or octo-sts
token is a 403 on a listener, not the 401 the comments claimed.

securitypolicy-sts now says its four audiences are the per-repository
opt-in: an AgentRun on another repository fails closed at its first
exchange, and opting one in means adding its audiences (EG: 8 per
provider, 4 providers per policy).

* docs(agents): record why the verified sub still reaches Z.ai

Removing x-ar-agent with a headerMutation on the zai backendRef edits the
same request headers the access log reads, so it would blank the
Gateway's own x_ar_agent attribution. The forwarding stays, with the
reason next to the claim that sets the header.

* fix(scripts): validate agent-run.sh inputs and surface the applied claim

Phase 6 review fixes (I1, minors 3-5): create (never apply, so a runId
collision surfaces as AlreadyExists instead of a silent CEL-rejected
update), validate --role/--class/--size/--minutes and AGENT_PRINCIPAL
against the design's principal CEL (human:<id>|system:<name>), and fail
cleanly when git user.email is unset instead of a raw set -e abort.
Success now also echoes the applied principal/role/class/repo/branch/
minutes to stderr, so a stale AGENT_PRINCIPAL or a --role typo is
visible without a follow-up kubectl get; stdout still carries only the
run name.

Also note (I2) that the dashboard's ar_agent label is produced by SP4
PR 1's envoy-ai-gateway controller on another branch, not this one, so
"No data" on that panel is expected until it merges. No absent() alert:
no agent run existing is not an outage.

Extends test-agent-run.sh for every new validation path, the
create-not-apply switch, --branch landing in the claim, an unknown
flag, and the property that a failed kubectl create prints no run
name.

* fix(observability): count 403s as rejected agent-router tokens

Envoy's jwt_authn filter returns 403 for wrong audience tokens, not just
401. Update the AgentRouterUnauthorizedBurst alert to count both status
codes and clarify the description for T8 (replayed/foreign) and R2
(identity-proxy) failure modes.

* fix(security): generate the MCP session-encryption seed in-cluster

The Agent Router chart falls back to the published "default-insecure-seed"
for controller.mcp.sessionEncryption.seed unless a value is supplied, which
lets anyone decrypt or forge a client-facing MCP session ID (phase-5 review
I1). Generate it with the same ESO Password + ExternalSecret pattern already
used for the AI-gateway rate-limit Valkey password, and feed it to the
HelmRelease via valuesFrom/targetPath, since the chart only renders the seed
into a controller CLI argument that has no env or file alternative.

The seed still lands in the extproc sidecar args of every AI-gateway
data-plane pod, so it remains readable by anything that can read pod specs.
SP2's room broker must authorize every call on x-ar-agent regardless.

* fix(ci): close two more MCP token-passthrough paths, rename gate A5 to A6

check_mcp_token_passthrough only ever looked at backendRefs[].forwardHeaders.
Agent Router also lets a route hand Authorization to a backend through
spec.securityPolicy.oauth.claimToHeaders[].header and
spec.securityPolicy.apiKeyAuth.forwardClientIDHeader (review I2); the
docstring's "the one way" was wrong. Both are now checked, case-insensitively,
with a failing test per path added first.

Renamed A5 to A6 throughout (code, messages, docstring, tests,
scripts/AGENTS.md): a parallel branch adds its own A5 (route sectionName)
to the same module, and the two need distinct ids before they merge.

Also cross-referenced the MCPRoute-level oauth issuer/audiences with their
Gateway-listener SecurityPolicy counterparts (agent-router/securitypolicy-
{public,internal}.yaml): a route-level SecurityPolicy replaces rather than
merges with the listener one, so the two must be kept in sync by hand
(review M7).

* fix(security): narrow agent-mcp's RBAC, egress and identity pins

- flux-operator-mcp's ClusterRole enumerated resources per apiGroup instead
  of `resources: ["*"]` under 8 groups, so a future Kind (Flux's own, or a
  Configuration bump under cloud.ogenki.io) needs a diff before an agent can
  read it (review M1).
- mcp-victoriametrics/mcp-victorialogs egress to `observability` now selects
  the vmsingle/victoria-logs-single pods by label, not the whole namespace
  (review M6).
- Noted the residual on each server's `fromEntities: host` ingress rule,
  needed for kubelet probes since there's no separate health port or shell
  for an exec probe (review M2).
- Added the Agent Router's internal per-backend routing headers
  (x-ai-eg-mcp-backend, x-ai-eg-mcp-route) to the Gateway-wide identity
  header strip, as defense-in-depth against the per-backend relocation ever
  failing (review M4, optional half).
- Anchored the flux-operator-mcp OCIRepository's cosign matchOIDCIdentity
  regexes, which were an unintended substring match, and pinned its image
  to a digest like VM/VL already are (review M8).
- Warned at each MCP chart/image pin that a bump exposing a new resource,
  prompt or template ships unauthorized, since those bypass MCPRoute
  authorization entirely (review M5).

* fix(ops): harden the MCP probe script and exercise reviewer-only tools

agent-probe-mcp.sh used fixed /tmp/auth and /tmp/h paths, so two classes
probed in parallel would clobber each other's token and session headers; no
curl call had a timeout; and a 401 on `initialize` surfaced only as a
confusing empty session ID on the next call. Switched to mktemp with a trap
cleanup, added -m 20 to every curl call, print the `initialize` status, added
a best-effort session DELETE, and an optional 4th argument to override the
listener port (for a live cross-class 401 check).

Added a reviewer-audience token projection to agent-probe.yaml: the Allow
rules that only reviewer/tester/triager get (VictoriaLogs tools) were never
exercised by any probe run (review M9).

* docs(superpowers): correct two residual notes in the identity design spec

C5 ("whether identity reaches the MCP backends") was marked UNVERIFIED; the
phase-5 review pre-flight settled it from source: yes, as x-ar-agent, via
oauth.claimToHeaders. SP2 no longer needs a fallback for this path.

T12's residual named only logs and ConfigMaps; cluster-wide `get pods` also
exposes pod specs and, under FallbackToLogsOnError, a crash log tail,
reachable by every internal run rather than only reviewer/tester/triager
(review M10, M3).

* docs(superpowers): mirror the design docs from #2092

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(rooms): room_* tools on both MCPRoutes, behind a generated key

The room-broker backend (:8090, /mcp) joins agent-mcp-public and
agent-mcp-internal. agent-router injects x-room-mcp-key from the generated
Secret room-broker-mcp-key and the verified x-ar-agent; each role gets its
room_* tools per SP2 section 3 on both listeners. The broker reads the key as
ROOMS_MCP_KEY and admits :8090 from agent-router's data plane only.

Pins AP-3's broker pre-release (agent-platform#8) by digest in the App and the
retention CronJob, and re-vendors the Room CRD at 91b25ce (spec.dataClass is
immutable, ruling TD). test-agent-mcp-scope.sh pins the room grants.

* feat(rooms): factory App key, GitHub egress and alert for verdict comments

The broker reads the factory App's key (agents/factory-app: app_id,
private_key) through the agents-secrets store, mounted optionally at
ROOMS_GITHUB_APP_DIR with defaultMode 0440 so the pod's fsGroup reads it
(ruling P31): the broker runs before the owner writes the key, and kubelet
fills the volume once it lands. Its only GitHub egress is api.github.com:443.

RoomVerdictsNotReachingGitHub fires after 30 minutes of errors with no post.
Both clusters' READMEs name the key as an owner prerequisite.

* fix(rooms): verdict alert survives the backoff gap, MCP key never regenerates

- RoomVerdictsNotReachingGitHub counts errors over 20m: a stuck verdict
  retries every 15 min at most, so the 10m rate window emptied between
  retries and reset 'for' forever. The description no longer names the
  App's installation, which reaches the counter as not_posted, not error.
- room-broker-mcp-key: creationPolicy Orphan, Flux prune disabled and
  target.immutable. Deleting the ExternalSecret keeps the Secret, and a
  recreated one cannot overwrite it (CreatedOnce is tracked on the
  ExternalSecret's status), so the broker's start-time copy stays valid.
- Gate A6 fails a backend apiKey with no header or queryParam, or one
  named Authorization: it would go out as a bearer.
- The factory App volume's 0440 comment names the real dependency: the
  composition's fsGroup.

* docs(rooms): the MCP key's rotation steps under Orphan

* chore(rooms): pin both clouds to CC-S3 v0.7.2-pr34.dea51a5, schema from feat/room-tools

CC-S3 (crossplane-configuration#34 at 8cc8cc8) carries H-S3's harness
v0.2.0-pr2142.10c062c2, the one that appends the PR provenance footer. The
core package follows through dependsOn; the App Wizard stays on v0.7.1.

atlasSchema.ref moves to agent-platform feat/room-tools, AP-3's stack tip,
which holds the verdict index migration and matches the broker pin
(pr8.91b25cef) and crd-rooms.yaml (already vendored from 91b25ce).

* chore(rooms): pin both clouds to CC-S4 v0.7.2-pr35.465e19f, broker pr9, schema from feat/room-driver

* fix(rooms): a signable room-broker certificate on gcp-0

The OpenBao PKI role requires a common name and only signed names under the
private domain, so room-broker's Service-name certificate never issued. Add
the CN, and let gcp-0's role sign agent-system's cluster-local names.

* chore(rooms): pin the S4 broker and crossplane-configuration pre-releases (pr11.eb61ce7c, pr35.85a0fae)

agent-platform#11 (feat/room-approvals at eb61ce7) is the rooms stack's
reviewed tip: approvals, the SBB lease takeover, the F10/F11/F15 fixes.
crossplane-configuration#35 published v0.7.2-pr35.85a0fae with the matching
room-bridge pin. atlasSchema.ref moves to feat/room-approvals with the
broker pin; the vendored CRD is unchanged since feat/room-tools, only its
provenance comment moves.

* feat(rooms): alert on approvals pending too long, pin CC-S5 (SP2 phase 5)

RoomApprovalPendingTooLong fires when the oldest undecided approval has waited
15 minutes: the run is parked, holding its sandbox. Routed to Slack (ADR-0037).
crossplane-configuration v0.7.2-pr35.4852542 is CC-S5: the room bridge gets the
run's BRANCH (live finding A1) and a started run latches Running (finding B).

* feat(rooms): roomctl client, oauth2-proxy bearer tokens, AP-6 pins (SP2 phase 6) (#2226)

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(rooms): roomctl's ZITADEL client, oauth2-proxy bearer tokens, pin AP-6 (SP2 phase 6)

roomctl is a native ZITADEL app (no secret, device code and refresh token,
JWT access tokens) whose id reaches OpenBao's agents/roomctl. The broker reads
it from a file to tell roomctl's tokens from the web UI's (ruling P18), and
oauth2-proxy lets those bearers through against roomctl's own client. A native
app gets no secret back, so its create and its missing-key paths no longer
fail. The broker, its retention CronJob, crd-rooms.yaml and atlasSchema.ref
move to agent-platform#28 together (P33).

The rooms OIDC suite also loads stored_client_id and stubs
restart_rotated_consumers: its cmd_sync section had exited 127 since #2149
reached this stack.

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(rooms): force the proxy's reconcile after the roomctl sync

A reconcile of rooms-oauth2-proxy before zitadel-oidc-clients.sh has
created the roomctl client leaves the new pod waiting for
rooms-proxy-extra-issuers. After three failed upgrades Helm rolls back
to skip-jwt-bearer-tokens=false and the release stalls, and nothing
retries it once the Secret appears. The ordering note and Task 6.4 now
end with flux reconcile hr rooms-oauth2-proxy -n agent-system --force.

Review I2.

* test(rooms): load previous_client_id in the rooms sync harness

main's cmd_sync now reads the old client id through previous_client_id
(#2208's OpenBao fallback). The rooms harness extracts functions by name,
so after the merge every cmd_sync case exited 127. It now loads that
function and stubs mirror_read, as the openbao suite does.

* fix(rooms): pin the broker to AP-6's review fixes (agent-platform#28 at 8dcf11d)

The broker and the retention CronJob move to
v0.0.1-pr28.8dcf11d2@sha256:40a38465…: the fork's own rate and prefix
caps, roomctl's final errors and the equal-client-id guard. The CRD is
unchanged at 8dcf11d; its header names the new commit.

* fix(rooms): pin the broker to AP-6 with the owner's fork rulings (agent-platform#28 at 91a6fb3)

The broker and the retention CronJob move to
v0.0.1-pr28.91a6fb39@sha256:2889162f…: a fork by anyone but the source's
owners takes the stricter approvals (M2), and a sealed room forks only
for its owners and agents-admin (M3). The CRD is unchanged at 91a6fb3;
its header names the new commit.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* feat(agents): runs that survive a spot or preemptible reclaim (rooms stack) (#2215)

* feat(agent-harness): provenance footer on gh pr create

The gh wrapper hands `gh pr create` to pr_footer.py, which appends the
run's footer (Agent-Room, Agent-Run, Agent-Role, Agent-Task,
Agent-Task-URL, Agent-Model) after the pull request exists, so every way
of writing the body gets it (SP2 ruling P32).

The Agent-* keys are the harness's alone. Any model-written line that
starts with one of them, in any case, is prefixed "(agent-written) " in
both the PR body and the commit message, and the real values are added
last: an agent can no longer suppress Agent-Run with its own or a
lower-case copy (SP3 ruling SL), nor pre-empt the footer by containing
it. The commit-msg hook moves to `--if-exists replace` and also writes
Agent-Task from TASK_ID; the issue or PR URL lands on Agent-Task-URL
(SP3 ruling SW). Harness version v0.2.0.

* fix(agent-harness): harden the provenance footer and trailer hook

Review of H-S3 (e20d524a):
- commit-msg passes --no-divider, so a `---` line in the message no
  longer pushes Agent-Run out of the last paragraph.
- The whole agent-* key namespace is marked, not six exact keys, and
  the hook adds with --if-exists add: replace matched `Agent:` as a
  prefix and could delete it. A second pass (amend) keeps one trailer.
- Lone CR, VT, FF, NEL, U+2028 and U+2029 count as line breaks before
  marking.
- The gh wrapper also routes `pr new` and `pr -R/--repo <o/r> create`.
- A footer that cannot be written fails closed: exit 1 with the URL.
- Values drop C0/C1 controls and line separators, are capped at 256
  characters, and Agent-Task-URL must be https on github.com exactly.
- Docs and comments call trailers and footer provenance hints, never
  authorisation (SP3 ruling TB).

* docs(agents): design agent runs that survive a spot or preemptible reclaim

An external review found runs on reclaimable capacity lose their work
and are never resumed. The design budgets the pod's shutdown at 15 s
(GKE preemptible VMs cannot extend it): pause, checkpoint commit and
push, the bridge's final read, then stop and revoke. The composition
reads the pod to record Disrupted, PodLost or PodFailed, and the
factory resumes Disrupted and PodLost runs at most twice within the
task's token cap. An early warning on aws-0 waits on a fault-injection
test. The research records the facts at pinned commits.

* docs(agents): plan agent runs that survive a reclaim

Thirteen tasks across agent-platform, cloud-native-ref and
crossplane-configuration: the bridge's final read, agent-run's
15 s shutdown sequence, the composition reading the pod for
Disrupted/PodLost/PodFailed, GKE Spot's 120 s window, the factory's
automatic resume, docs, runbook 09 and live verification. The spec is
aligned with what the code showed: no XRD enum, a deleted pod reads
PodLost, Crossplane needs cluster-wide pod reads for required resources,
interrupt rather than pause, an in-process revoke, the trailer from the
commit-msg hook.

* docs(agents): record the owner's decisions on the disruption plan

The wider Crossplane pod read is accepted, and merging the factory resume
work may bring factory phases 2-9 to gcp-0. The two owner actions the
plan depends on are listed.

* docs(agents): order the agent factory's remaining work

One table from the open reviews to the merge wave, naming the plan that
holds each step: the disruption plan's runtime half rides the same re-pin
and rebuild as the F10-F18 fixes, the agentgateway migration (with prompt
caching) follows the live re-verification, then SP2 and SP3.

* docs(agents): link the kagent evaluation now that it is on main

* feat(agent-harness): a SIGTERM sequence that checkpoints, lets the room read, then revokes

* feat(agents): Crossplane reads run pods for the composition's required resource

* docs(agents): the run's end reasons, the SIGTERM checkpoint, R7 corrected

* feat(gcp): 120 s of graceful node shutdown on the gVisor Spot pool

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(agent-harness): survive a rejected reply to a server ping (F29)

A backend's ping reaches agent-server through agent-router with its numeric
id unrewritten, so the client's reply is answered 400 (envoyproxy/ai-gateway#2715).
mcp 1.28.1 sends that reply inline from StreamableHTTPTransport.post_writer;
_handle_post_request's raise_for_status() ends post_writer, which closes both
session streams while fastmcp's is_connected() stays True, so every later tool
call failed with an empty ClosedResourceError and the SDK's reconnect-once
never fired.

agent-run now starts agent-server with site/ first on PYTHONPATH; its
sitecustomize wraps _handle_post_request so a 4xx to a JSONRPCResponse or
JSONRPCError is logged once and dropped. Requests and notifications are
unchanged, the agent's own commands never load it, and a missing method is
said loudly at start-up. Belt and braces: agent-router is being fixed to answer
client replies 202 in parallel; the harness no longer depends on that.

* docs(agent-harness): the commit hook's checkpoint trailer in the README

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* feat(agents): pin the disruption pre-release (pr36.5f36168)

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* fix(agent-harness): keep a finished run's exit code, guard and box the shutdown

C1 (harness half): a SIGTERM that lands after the conversation ended (during
the final read, the close or the step log) exited 143. The composition reads
the harness's exit code, and 143 on a pod with DisruptionTarget reads
Disrupted, which the factory resumes: a finished task run twice. main() now
returns the run's own code from its SystemExit handler.

M1: the checkpoint refuses (logs and skips the commit) a staged tree holding a
GitHub token, the cached run token whatever its shape (redact()), more than
200 files or more than 5 MiB. The agent's own unpushed commits are still
pushed.

M2: stop() kills agent-server for the box's last second and never waits past
it; the root span's export after the revoke is boxed at 1 s on the signal path
(a silent collector held exit 5 s past the revoke).

M3: CHECKPOINT_S 8 -> 7: the five boxes sum to 14 s, so the revoke is done a
second before a 15 s window's SIGKILL.

* docs(agents): a revoked or deleted implementer pushes its checkpoint first (I1)

Recorded as known, not changed: the owner decides. The user guide's stop section and SC-07's procedure say that the SIGTERM sequence runs on a revoke or a delete too, so an implementer with uncommitted changes pushes a checkpoint commit before its token is revoked.

* chore(crossplane): point the controller's memory limit at Task 13's measurement (M4)

Crossplane now keeps a cluster-wide pod informer for the AgentRun composition's runPod; the disruption plan's Task 13 measures what it costs against the 512Mi limit.

* fix(agent-harness): a late SIGTERM still gets the final read; refuse on a failed git diff

Re-review minors of the disruption fixes:
- Exit 0 latches the run Succeeded at once, after which the broker refuses
  the bridge's own flush. A SIGTERM that interrupts the run's own final read
  (or lands before it) now asks for it again on the signal path, where nothing
  is paused or checkpointed. A read that had returned is not repeated.
  timed() reports whether its step is done.
- _unfit treats a failing git diff as unfit: the checkpoint is refused rather
  than checked against an empty diff.

* docs(agents): Succeeded latches at the harness's exit 0, the bridge's flush then refused

The disruption design's §3 table, after review C1: exit 0 is Succeeded whatever the pod went through, latched while the sidecars drain, so the harness's final read is the room's last; a crash sets the reason only; only the Sandbox's own pod is read.

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to …
Smana added a commit that referenced this pull request Oct 8, 2026
…actory-api

Merged with c4d62db as the base, the FR-4 commit this branch started from. The
broker image takes FR-4's release pin (room-broker v0.6.1). The rooms suite takes
main's: it already loads stored_client_id and previous_client_id and stubs
restart_rotated_consumers, which #2149's lines added a second time.
Smana added a commit that referenced this pull request Oct 8, 2026
…(SP3 phase 5) (#2185)

* feat(agents): agent-harness image on OpenHands agent-server 1.49.5

The AgentRun sandbox's harness: agent-server plus the agent-run driver,
git-credential-agent, gh, and a commit-msg hook adding Agent-Run: $RUN_ID.
The published agent-server image is upstream's PyInstaller binary, so the
SDK is installed into /agent-server/.venv from a hash-pinned universal lock.

Deviations from the plan's code, each found against the running server:
- agent-server binds 127.0.0.1 explicitly (plan P13, CC-1 composition),
  not 0.0.0.0: its API is unauthenticated.
- OH_ENABLE_VSCODE=false: otherwise agent-server starts openvscode-server
  on 0.0.0.0:8001.
- build_request dumps with expose_secrets: the SDK redacts the placeholder
  model key and drops it on reload, so litellm refused to call at all.
- A SIGTERM handler raises SystemExit, so pod deletion still runs the
  finally block that revokes the GitHub token and stops agent-server.

docker build --target test: 10 tests OK.

* feat(security): agents-secrets store for agent-system

* feat(security): reach octo-sts only through agent-router's sts listener

* feat(agents): agent-router Gateway with one JWT listener per data class

* fix(security): exempt octo-sts's key-path variable from Polaris's secret-name check

* fix(docs): drop the ADR-0041 reference to an unmerged spike doc path

* feat(observability): agent-platform VMRules and dashboard

Task 6.4 content. The umbrella child that applies it follows in its own commit,
once PR 2's agent-platform umbrella exists on this branch.

* feat(clusters): apply the agent-platform VMRules and dashboard behind the umbrella

Wires the agent-observability child Kustomization (observability/base/agent-platform)
into the aws-0-agent-platform umbrella, so it resumes with the rest of the platform.

* feat(agents): read-only flux-operator-mcp with a narrow ClusterRole

* feat(agents): VictoriaMetrics and VictoriaLogs MCP servers for agents

* feat(agents): per-class MCPRoutes with role-scoped tools

* fix(agents): fold in review findings for the agent-harness image

Cleanup order in agent_run.py's finally now stops agent-server before
revoking its GitHub token, and ignores a second SIGTERM so the revoke
can't be aborted. Disable autotitle to save a model call per run. Pin
agent-server's cwd to "/" so its state lands under /workspace whatever
the pod's workingDir is. Status polling tolerates up to 5 consecutive
network errors before failing, and logs the final execution_status.

git-credential-agent: revoke() no longer tracebacks on an already-dead
token (logs the status instead), a token being replaced on refresh is
revoked before the cache is overwritten, and concurrent get/token
calls serialize on an flock of the cache file.

Dockerfile: strip setgid alongside setuid, pin GIT_TERMINAL_PROMPT=0,
and install the locked requirements with --only-binary=:all: (with a
targeted exception for func-timeout, a pure-Python, wheel-less
transitive pin from openhands-tools -- verified every other locked
package ships a wheel for both amd64 and arm64).

* fix(ci): gate agent-router listener pinning and route-level SecurityPolicy merges

A5 in assert-ai-gateway.py: every route naming agent-system/agent-router
sets sectionName (without one it attaches to public, internal and sts),
and every SecurityPolicy on such a route sets mergeType (without one it
replaces the listener's JWT check). Unresolvable targets and label
selectors are assumed in scope; zero routes on the Gateway fails rather
than passing vacuously. Other Gateways' routes may still omit sectionName.

* fix(agents): let the identity proxy outlast agent-models' 600s request timeout

Both model listeners set stream_idle_timeout 300s and no route overrides
it, so a non-streamed GLM-5.2 completion taking 300-600s was reset by the
proxy before the router's 600s budget ran out. Raised to 900s. The route
comment no longer claims the proxy sets no timeout, and records that EG
v1.9.1 derives the router's per-route idle timeout as max(1h, request)
while no ClientTrafficPolicy sets streamIdleTimeout.

* fix(github): pin the ruleset fields that make it a control

The test asserted the bypass list, conditions and rule types but never
the name, target or enforcement, and never the update body. An edit to
`enforcement: evaluate` or `disabled`, `target: tag`, a renamed ruleset
or an empty PUT body all stayed green; each now fails, on both the
create and the update body.

The applier reads the ruleset name from the JSON instead of hard-coding
it, lists only the repository's own rulesets (includes_parents=false),
takes the first match, and warns when an update would drop an App from
the bypass list, so a re-run without FACTORY_APP_SLUG after SP3 is not
silent. Its header says to run it before the agents' App is installed.

* fix(security): keep octo-sts's Polaris exemption in .polaris.yaml

SPEC-007 keeps every exemption in .polaris.yaml, by controller and
rule, with its reason; the annotation on the Deployment was the only
one in the repo and invisible from that list. Same rule, same scope:
octo-sts, sensitiveContainerEnvVar. The section header now admits
false positives on a name the workload fixes, not only privilege.

* docs(adr): ADR-0043 says what is recorded and puts the ruleset first

octo-sts 0.10.0 puts the issuer, subject and token SHA-256 only in its
exchange event, which needs METRICS=true and a CloudEvents sink; neither
is set, so its logs carry the requested repository and trust policy and
nothing that ties a token to a run. The Pro now names what is recorded
(agent-router's sts access log, octo-sts's log, GitHub's own record of
what the App's bot did), and a Negative says what is not.

A repository opts in three times, ruleset first: main needs no
approval, so until the ruleset exists nothing stops the implementer
from merging its own green PR. The cluster README's Resume section
lists the owner prerequisites in that order, and what fails without
the App key.

Also recorded: contents:write reaches tags, releases and
repository_dispatch; PR CI still runs scripts an agent can edit; the
App key's exposure through openbao-platform (T14); a rotated key needs
a rollout restart; Dependabot is off on this repository.

* fix(security): close the TokenRequest and forged-owner gaps in the agent policies

- agent-audience-token-request: the TokenRequest API (serviceaccounts/token)
  minted the reserved agent audiences without a pod, so the pod rule alone
  left them open to `kubectl create token` and ESO's cluster-wide grant.
- agents-pod-creator: ownerReferences are client-supplied; the requester
  identity is not. Admission-only, since `request` is absent in background.
- agentrun-gc: delete terminal runs only a day after they ended, falling
  back to creation time when a run carries no finishedAt.

* fix(agents): narrow the controller scrape ingress and correct review nits

- CNP: admit only vmagent on :8080, not the whole observability namespace.
- Resume steps: wait for the crossplane-configuration pin that serves
  AgentRun, since agent-policies reports Ready either way.
- ADR-0041: the spike notes are a superpowers spec, landing with #2092.
- gen-catalog.sh: drop the stale source counts.

* fix(agents): make the agent-sandbox chart and image lockstep mechanical

Renovate's flux manager tracks the GitRepository tag while the helm-values
manager could not see the controller image (no `repository` in values), so
the patch/minor automerge rule would have moved the chart and left the
binary behind. Restate the image repository so the image is tracked, and
group both into one PR that is never automerged. The HelmRelease comment
no longer claims a digest cannot be pinned (a postRenderer could) and says
the chart still installs the extension CRDs with extensions off.

* fix(ci): render oidc_issuer_url with its /id path, like oidc_issuer_host

The live issuer URL carries /id/<ID>; the bare-host fixture rendered the
agent-router SecurityPolicies with an issuer shape the cluster never has,
and left configuration-aws's EnvironmentConfig with an inconsistent pair.

* fix(agents): revoke under the same lock as token refresh

revoke() now holds CACHE + ".lock" when reading, revoking, and deleting
the cached token. This closes the race with concurrent token() exchanges
that could leave a live, unrevoked token in the cache if revoke() ran
without the lock.

The fix ensures that revoke() atomically reads the cache, revokes the token
at GitHub, and deletes the file — serializing with any concurrent token()
call via the same flock that token() uses.

Also clarify the Dockerfile comment to mention that both setuid and setgid
bits are stripped, matching what the find command actually does (-perm /6000).

Test: new test_revoke_waits_for_the_cache_lock verifies revoke() blocks
on the lock held by token(), preventing the interleaving that would leave
an unrevoked token in the cache.

Ran 19 tests, all pass.

* fix(agents): run agent-router on two replicas behind a PDB

With one replica, every drain or Karpenter consolidation of its node cut
all in-flight agent completions at once, and the harness does not retry.
envoyPDB minAvailable 1 keeps one proxy serving through a voluntary
disruption.

* fix(agents): drop agent-router's host ingress and name its rate-limit egress

The fromEntities: host rule on 8080/8081 admitted nothing the data plane
needs: EG's kubelet probes target the readiness and shutdown-manager
ports, which it never listed, and Cilium's default allow-localhost
already admits them. It only implied host-network pods were an intended
caller, which ADR-0042 says is agents pods alone.

The header now names the envoy-ratelimit :8081 egress SP4 must add with
its budgets: a dropped rate-limit check fails open.

* fix(security): retry agent-secrets after 30s, not a whole interval

On first apply the SecretStore can go Ready=False before openbao-ca holds
its CA; the failed health check then waited the 5m interval, holding
agent-router back by as much. 30s matches security-openbao.

* fix(security): call agents-secrets agent-system's store by convention

"agent-system's only way to secrets" read as a control. Nothing enforces
it: no ClusterSecretStore sets namespace conditions, and openbao-ca itself
reads through clustersecretstore.

* fix(scripts): let the identity probe's Sandbox expire if its delete is forgotten

agent-sandbox v1.0.3 keeps a terminated pod rather than recreating it,
but creates a new one, with fresh tokens, if that pod is ever removed.
shutdownPolicy Delete plus the documented shutdownTime patch bound the
Sandbox's life; the delete tolerates the Sandbox having expired first.

* fix(agents): audience mismatches are 403, and sts's audiences are the per-repo opt-in

Envoy's jwt_authn answers JwtAudienceNotAllowed with 403 and every other
failure with 401 (filter.cc at v1.39.1), so a cross-class or octo-sts
token is a 403 on a listener, not the 401 the comments claimed.

securitypolicy-sts now says its four audiences are the per-repository
opt-in: an AgentRun on another repository fails closed at its first
exchange, and opting one in means adding its audiences (EG: 8 per
provider, 4 providers per policy).

* docs(agents): record why the verified sub still reaches Z.ai

Removing x-ar-agent with a headerMutation on the zai backendRef edits the
same request headers the access log reads, so it would blank the
Gateway's own x_ar_agent attribution. The forwarding stays, with the
reason next to the claim that sets the header.

* fix(scripts): validate agent-run.sh inputs and surface the applied claim

Phase 6 review fixes (I1, minors 3-5): create (never apply, so a runId
collision surfaces as AlreadyExists instead of a silent CEL-rejected
update), validate --role/--class/--size/--minutes and AGENT_PRINCIPAL
against the design's principal CEL (human:<id>|system:<name>), and fail
cleanly when git user.email is unset instead of a raw set -e abort.
Success now also echoes the applied principal/role/class/repo/branch/
minutes to stderr, so a stale AGENT_PRINCIPAL or a --role typo is
visible without a follow-up kubectl get; stdout still carries only the
run name.

Also note (I2) that the dashboard's ar_agent label is produced by SP4
PR 1's envoy-ai-gateway controller on another branch, not this one, so
"No data" on that panel is expected until it merges. No absent() alert:
no agent run existing is not an outage.

Extends test-agent-run.sh for every new validation path, the
create-not-apply switch, --branch landing in the claim, an unknown
flag, and the property that a failed kubectl create prints no run
name.

* fix(observability): count 403s as rejected agent-router tokens

Envoy's jwt_authn filter returns 403 for wrong audience tokens, not just
401. Update the AgentRouterUnauthorizedBurst alert to count both status
codes and clarify the description for T8 (replayed/foreign) and R2
(identity-proxy) failure modes.

* fix(security): generate the MCP session-encryption seed in-cluster

The Agent Router chart falls back to the published "default-insecure-seed"
for controller.mcp.sessionEncryption.seed unless a value is supplied, which
lets anyone decrypt or forge a client-facing MCP session ID (phase-5 review
I1). Generate it with the same ESO Password + ExternalSecret pattern already
used for the AI-gateway rate-limit Valkey password, and feed it to the
HelmRelease via valuesFrom/targetPath, since the chart only renders the seed
into a controller CLI argument that has no env or file alternative.

The seed still lands in the extproc sidecar args of every AI-gateway
data-plane pod, so it remains readable by anything that can read pod specs.
SP2's room broker must authorize every call on x-ar-agent regardless.

* fix(ci): close two more MCP token-passthrough paths, rename gate A5 to A6

check_mcp_token_passthrough only ever looked at backendRefs[].forwardHeaders.
Agent Router also lets a route hand Authorization to a backend through
spec.securityPolicy.oauth.claimToHeaders[].header and
spec.securityPolicy.apiKeyAuth.forwardClientIDHeader (review I2); the
docstring's "the one way" was wrong. Both are now checked, case-insensitively,
with a failing test per path added first.

Renamed A5 to A6 throughout (code, messages, docstring, tests,
scripts/AGENTS.md): a parallel branch adds its own A5 (route sectionName)
to the same module, and the two need distinct ids before they merge.

Also cross-referenced the MCPRoute-level oauth issuer/audiences with their
Gateway-listener SecurityPolicy counterparts (agent-router/securitypolicy-
{public,internal}.yaml): a route-level SecurityPolicy replaces rather than
merges with the listener one, so the two must be kept in sync by hand
(review M7).

* fix(security): narrow agent-mcp's RBAC, egress and identity pins

- flux-operator-mcp's ClusterRole enumerated resources per apiGroup instead
  of `resources: ["*"]` under 8 groups, so a future Kind (Flux's own, or a
  Configuration bump under cloud.ogenki.io) needs a diff before an agent can
  read it (review M1).
- mcp-victoriametrics/mcp-victorialogs egress to `observability` now selects
  the vmsingle/victoria-logs-single pods by label, not the whole namespace
  (review M6).
- Noted the residual on each server's `fromEntities: host` ingress rule,
  needed for kubelet probes since there's no separate health port or shell
  for an exec probe (review M2).
- Added the Agent Router's internal per-backend routing headers
  (x-ai-eg-mcp-backend, x-ai-eg-mcp-route) to the Gateway-wide identity
  header strip, as defense-in-depth against the per-backend relocation ever
  failing (review M4, optional half).
- Anchored the flux-operator-mcp OCIRepository's cosign matchOIDCIdentity
  regexes, which were an unintended substring match, and pinned its image
  to a digest like VM/VL already are (review M8).
- Warned at each MCP chart/image pin that a bump exposing a new resource,
  prompt or template ships unauthorized, since those bypass MCPRoute
  authorization entirely (review M5).

* fix(ops): harden the MCP probe script and exercise reviewer-only tools

agent-probe-mcp.sh used fixed /tmp/auth and /tmp/h paths, so two classes
probed in parallel would clobber each other's token and session headers; no
curl call had a timeout; and a 401 on `initialize` surfaced only as a
confusing empty session ID on the next call. Switched to mktemp with a trap
cleanup, added -m 20 to every curl call, print the `initialize` status, added
a best-effort session DELETE, and an optional 4th argument to override the
listener port (for a live cross-class 401 check).

Added a reviewer-audience token projection to agent-probe.yaml: the Allow
rules that only reviewer/tester/triager get (VictoriaLogs tools) were never
exercised by any probe run (review M9).

* docs(superpowers): correct two residual notes in the identity design spec

C5 ("whether identity reaches the MCP backends") was marked UNVERIFIED; the
phase-5 review pre-flight settled it from source: yes, as x-ar-agent, via
oauth.claimToHeaders. SP2 no longer needs a fallback for this path.

T12's residual named only logs and ConfigMaps; cluster-wide `get pods` also
exposes pod specs and, under FallbackToLogsOnError, a crash log tail,
reachable by every internal run rather than only reviewer/tester/triager
(review M10, M3).

* docs(superpowers): mirror the design docs from #2092

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(agent-factory): ADR-0048, the factory under the agent-platform umbrella

The factory's signed chart source, HelmRelease (Task CRD created on install,
replaced on upgrade), values with the strictly parsed config, the App key
through agents-secrets, a default-deny CNP and its scrape, as an umbrella
child on both aws-0 and gcp-0 with a GP-14 render root per cloud. The broker's
allowlist admits the factory's ServiceAccount, and the cloud-shape gate now
judges the gcp-0 factory overlay.

* feat(agent-factory): send task spans to the trace collector

* fix(agent-factory): pin the chart by digest, pair it with the image, state ADR-0048's reasons

- OCIRepository carries ref.digest, which Flux resolves before the tag
- render-bundle renders a digest-pinned chartRef as <url>@<digest>
- test-agent-factory-pins.py fails when chart and image pins disagree;
  Renovate groups the two and now tracks the image tag in the values
- VMRule runbook reads every replica, since only the leader polls
- ADR-0048 gives reasons in place of programme ledger labels

* chore(agent-factory): pin the factory to agent-platform#7 at 569dc1c, by digest

Chart and image are the signed pre-release of that head.

* chore(agent-factory): pin the phase 2 pre-releases

agent-platform#10 at 800fcd8: the factory image and chart, and the room-broker image, a superset of AP-4's, for the app and the retention CronJob. The vendored Room CRD moves with the broker pin, as its comment requires; it gains dataClass immutability.

* chore(agent-factory): pin the phase 3 pre-releases; pair by default

agent-platform#12 at f2e51f9: the factory image and chart, and the room-broker image for the app and the retention CronJob (FA-3 leaves the broker's code and the Room CRD unchanged, so crd-rooms.yaml stays). defaults.template becomes pair: every task gets a reviewer until phase 4's triage picks templates.

* docs(adr): 0048 separates the decided design from what phase 1 builds

Kueue is not deployed and the factory enforces only the run budget; no gateway per-run ceiling exists yet. Keep the decision, add a decided/built table, and name SP3 phase 4 for Kueue. The trace collector's CNP comment no longer calls the factory rule inert.

* chore(agent-factory): re-pin the phase 3 pre-releases to agent-platform#12 at ddb06e0

* feat(agents): Kueue under the agent-platform umbrella (SP3 FR-4)

Kueue admits every factory sandbox pod: pod-only integration scoped to the
agents namespace, two non-borrowing ClusterQueues (factory, interactive) and
their LocalQueues, default-deny CNP with webhook/metrics/health ingress and
DNS/API egress, TLS metrics scraped by vmagent through a dedicated
metrics-reader ClusterRole.

Per ruling SB the target cloud is gcp-0: the controller child points at an
infrastructure/gcp-0/kueue render root (GP-14) and substitutes
gke-gcp-0-vars; the queue child is cloud-neutral and stays on the base.
agent-factory now dependsOn kueue-queues so submitted runs find their
LocalQueue. The Kueue chart joins the factory's no-automerge Renovate rule.

* feat(agent-factory): triage config, classes, shadow budgets, classifier egress

* feat(agent-policies): the factory is the only AgentRun creator; its patches are annotations only

* fix(agent-policies): bound the stop-readonly rule by resourceNames; DELETE carries no object

* feat(agent-factory): POST /v1/runs on the tailnet, client ids, 429 lookup egress

* feat(agent-run): task agent:run asks the factory's API; the token names the principal

* test(zitadel-oidc-clients): load stored_client_id, mock restart_rotated_consumers (#2149)

* feat(rooms): the broker requests runs from the factory

* fix(kueue): drop the visibility APIServices the API server cannot reach (#2245)

Kueue's manager serves its on-demand visibility API on 8082, which neither the
EKS node security group nor agent-system's CiliumNetworkPolicy opens, so both
APIServices sit FailedDiscoveryCheck: every discovery client logs errors and
namespace deletion blocks. Nothing reads that API, and chart 0.19.6 renders
the APIServices with no values switch, so a post-renderer removes them.

(cherry picked from commit e744c44)

* chore(crossplane): pin both clouds to CC v0.9.2 (harness v0.3.0)

v0.9.2 carries core v0.9.1, which pins agent-harness v0.3.0 (the disruption
build S5 published), and raises the aws and gcp packages' core floor to
>=v0.9.1 so the new core reaches clusters. Package digests: core f0b29186,
aws bf1ef35a, gcp 589d1310. The App Wizard clone tag and the doc mirrors follow;
the inference KCL module is still 0.9.0.

* chore(agent-factory): pin agent-platform v0.7.0, release signatures only (R20)

The factory image and chart, and the room broker with its retention CronJob, CRD
and schema, move to the v0.7.0 release (tag 7140faef, run 37721493733):

- agent-factory v0.7.0@sha256:83a9adc1, chart 0.7.0@sha256:977928e1 (appVersion v0.7.0);
- room-broker v0.7.0@sha256:c7a33043 in app.yaml and the retention CronJob;
- crd-rooms.yaml from the v0.7.0 asset (sha256 980ee0c5; body unchanged since v0.6.1);
- atlasSchema.ref v0.7.0 (no migration changed since v0.6.1).

The OCIRepository's subject narrows to release.yaml on v* tags (R20): ci.yaml
pre-releases and pull request builds no longer verify. All three artifacts pass
cosign verify against it.
Smana added a commit that referenced this pull request Oct 8, 2026
…#2187)

* fix(github): pin the ruleset fields that make it a control

The test asserted the bypass list, conditions and rule types but never
the name, target or enforcement, and never the update body. An edit to
`enforcement: evaluate` or `disabled`, `target: tag`, a renamed ruleset
or an empty PUT body all stayed green; each now fails, on both the
create and the update body.

The applier reads the ruleset name from the JSON instead of hard-coding
it, lists only the repository's own rulesets (includes_parents=false),
takes the first match, and warns when an update would drop an App from
the bypass list, so a re-run without FACTORY_APP_SLUG after SP3 is not
silent. Its header says to run it before the agents' App is installed.

* fix(security): keep octo-sts's Polaris exemption in .polaris.yaml

SPEC-007 keeps every exemption in .polaris.yaml, by controller and
rule, with its reason; the annotation on the Deployment was the only
one in the repo and invisible from that list. Same rule, same scope:
octo-sts, sensitiveContainerEnvVar. The section header now admits
false positives on a name the workload fixes, not only privilege.

* docs(adr): ADR-0043 says what is recorded and puts the ruleset first

octo-sts 0.10.0 puts the issuer, subject and token SHA-256 only in its
exchange event, which needs METRICS=true and a CloudEvents sink; neither
is set, so its logs carry the requested repository and trust policy and
nothing that ties a token to a run. The Pro now names what is recorded
(agent-router's sts access log, octo-sts's log, GitHub's own record of
what the App's bot did), and a Negative says what is not.

A repository opts in three times, ruleset first: main needs no
approval, so until the ruleset exists nothing stops the implementer
from merging its own green PR. The cluster README's Resume section
lists the owner prerequisites in that order, and what fails without
the App key.

Also recorded: contents:write reaches tags, releases and
repository_dispatch; PR CI still runs scripts an agent can edit; the
App key's exposure through openbao-platform (T14); a rotated key needs
a rollout restart; Dependabot is off on this repository.

* fix(security): close the TokenRequest and forged-owner gaps in the agent policies

- agent-audience-token-request: the TokenRequest API (serviceaccounts/token)
  minted the reserved agent audiences without a pod, so the pod rule alone
  left them open to `kubectl create token` and ESO's cluster-wide grant.
- agents-pod-creator: ownerReferences are client-supplied; the requester
  identity is not. Admission-only, since `request` is absent in background.
- agentrun-gc: delete terminal runs only a day after they ended, falling
  back to creation time when a run carries no finishedAt.

* fix(agents): narrow the controller scrape ingress and correct review nits

- CNP: admit only vmagent on :8080, not the whole observability namespace.
- Resume steps: wait for the crossplane-configuration pin that serves
  AgentRun, since agent-policies reports Ready either way.
- ADR-0041: the spike notes are a superpowers spec, landing with #2092.
- gen-catalog.sh: drop the stale source counts.

* fix(agents): make the agent-sandbox chart and image lockstep mechanical

Renovate's flux manager tracks the GitRepository tag while the helm-values
manager could not see the controller image (no `repository` in values), so
the patch/minor automerge rule would have moved the chart and left the
binary behind. Restate the image repository so the image is tracked, and
group both into one PR that is never automerged. The HelmRelease comment
no longer claims a digest cannot be pinned (a postRenderer could) and says
the chart still installs the extension CRDs with extensions off.

* fix(ci): render oidc_issuer_url with its /id path, like oidc_issuer_host

The live issuer URL carries /id/<ID>; the bare-host fixture rendered the
agent-router SecurityPolicies with an issuer shape the cluster never has,
and left configuration-aws's EnvironmentConfig with an inconsistent pair.

* fix(agents): revoke under the same lock as token refresh

revoke() now holds CACHE + ".lock" when reading, revoking, and deleting
the cached token. This closes the race with concurrent token() exchanges
that could leave a live, unrevoked token in the cache if revoke() ran
without the lock.

The fix ensures that revoke() atomically reads the cache, revokes the token
at GitHub, and deletes the file — serializing with any concurrent token()
call via the same flock that token() uses.

Also clarify the Dockerfile comment to mention that both setuid and setgid
bits are stripped, matching what the find command actually does (-perm /6000).

Test: new test_revoke_waits_for_the_cache_lock verifies revoke() blocks
on the lock held by token(), preventing the interleaving that would leave
an unrevoked token in the cache.

Ran 19 tests, all pass.

* fix(agents): run agent-router on two replicas behind a PDB

With one replica, every drain or Karpenter consolidation of its node cut
all in-flight agent completions at once, and the harness does not retry.
envoyPDB minAvailable 1 keeps one proxy serving through a voluntary
disruption.

* fix(agents): drop agent-router's host ingress and name its rate-limit egress

The fromEntities: host rule on 8080/8081 admitted nothing the data plane
needs: EG's kubelet probes target the readiness and shutdown-manager
ports, which it never listed, and Cilium's default allow-localhost
already admits them. It only implied host-network pods were an intended
caller, which ADR-0042 says is agents pods alone.

The header now names the envoy-ratelimit :8081 egress SP4 must add with
its budgets: a dropped rate-limit check fails open.

* fix(security): retry agent-secrets after 30s, not a whole interval

On first apply the SecretStore can go Ready=False before openbao-ca holds
its CA; the failed health check then waited the 5m interval, holding
agent-router back by as much. 30s matches security-openbao.

* fix(security): call agents-secrets agent-system's store by convention

"agent-system's only way to secrets" read as a control. Nothing enforces
it: no ClusterSecretStore sets namespace conditions, and openbao-ca itself
reads through clustersecretstore.

* fix(scripts): let the identity probe's Sandbox expire if its delete is forgotten

agent-sandbox v1.0.3 keeps a terminated pod rather than recreating it,
but creates a new one, with fresh tokens, if that pod is ever removed.
shutdownPolicy Delete plus the documented shutdownTime patch bound the
Sandbox's life; the delete tolerates the Sandbox having expired first.

* fix(agents): audience mismatches are 403, and sts's audiences are the per-repo opt-in

Envoy's jwt_authn answers JwtAudienceNotAllowed with 403 and every other
failure with 401 (filter.cc at v1.39.1), so a cross-class or octo-sts
token is a 403 on a listener, not the 401 the comments claimed.

securitypolicy-sts now says its four audiences are the per-repository
opt-in: an AgentRun on another repository fails closed at its first
exchange, and opting one in means adding its audiences (EG: 8 per
provider, 4 providers per policy).

* docs(agents): record why the verified sub still reaches Z.ai

Removing x-ar-agent with a headerMutation on the zai backendRef edits the
same request headers the access log reads, so it would blank the
Gateway's own x_ar_agent attribution. The forwarding stays, with the
reason next to the claim that sets the header.

* fix(scripts): validate agent-run.sh inputs and surface the applied claim

Phase 6 review fixes (I1, minors 3-5): create (never apply, so a runId
collision surfaces as AlreadyExists instead of a silent CEL-rejected
update), validate --role/--class/--size/--minutes and AGENT_PRINCIPAL
against the design's principal CEL (human:<id>|system:<name>), and fail
cleanly when git user.email is unset instead of a raw set -e abort.
Success now also echoes the applied principal/role/class/repo/branch/
minutes to stderr, so a stale AGENT_PRINCIPAL or a --role typo is
visible without a follow-up kubectl get; stdout still carries only the
run name.

Also note (I2) that the dashboard's ar_agent label is produced by SP4
PR 1's envoy-ai-gateway controller on another branch, not this one, so
"No data" on that panel is expected until it merges. No absent() alert:
no agent run existing is not an outage.

Extends test-agent-run.sh for every new validation path, the
create-not-apply switch, --branch landing in the claim, an unknown
flag, and the property that a failed kubectl create prints no run
name.

* fix(observability): count 403s as rejected agent-router tokens

Envoy's jwt_authn filter returns 403 for wrong audience tokens, not just
401. Update the AgentRouterUnauthorizedBurst alert to count both status
codes and clarify the description for T8 (replayed/foreign) and R2
(identity-proxy) failure modes.

* fix(security): generate the MCP session-encryption seed in-cluster

The Agent Router chart falls back to the published "default-insecure-seed"
for controller.mcp.sessionEncryption.seed unless a value is supplied, which
lets anyone decrypt or forge a client-facing MCP session ID (phase-5 review
I1). Generate it with the same ESO Password + ExternalSecret pattern already
used for the AI-gateway rate-limit Valkey password, and feed it to the
HelmRelease via valuesFrom/targetPath, since the chart only renders the seed
into a controller CLI argument that has no env or file alternative.

The seed still lands in the extproc sidecar args of every AI-gateway
data-plane pod, so it remains readable by anything that can read pod specs.
SP2's room broker must authorize every call on x-ar-agent regardless.

* fix(ci): close two more MCP token-passthrough paths, rename gate A5 to A6

check_mcp_token_passthrough only ever looked at backendRefs[].forwardHeaders.
Agent Router also lets a route hand Authorization to a backend through
spec.securityPolicy.oauth.claimToHeaders[].header and
spec.securityPolicy.apiKeyAuth.forwardClientIDHeader (review I2); the
docstring's "the one way" was wrong. Both are now checked, case-insensitively,
with a failing test per path added first.

Renamed A5 to A6 throughout (code, messages, docstring, tests,
scripts/AGENTS.md): a parallel branch adds its own A5 (route sectionName)
to the same module, and the two need distinct ids before they merge.

Also cross-referenced the MCPRoute-level oauth issuer/audiences with their
Gateway-listener SecurityPolicy counterparts (agent-router/securitypolicy-
{public,internal}.yaml): a route-level SecurityPolicy replaces rather than
merges with the listener one, so the two must be kept in sync by hand
(review M7).

* fix(security): narrow agent-mcp's RBAC, egress and identity pins

- flux-operator-mcp's ClusterRole enumerated resources per apiGroup instead
  of `resources: ["*"]` under 8 groups, so a future Kind (Flux's own, or a
  Configuration bump under cloud.ogenki.io) needs a diff before an agent can
  read it (review M1).
- mcp-victoriametrics/mcp-victorialogs egress to `observability` now selects
  the vmsingle/victoria-logs-single pods by label, not the whole namespace
  (review M6).
- Noted the residual on each server's `fromEntities: host` ingress rule,
  needed for kubelet probes since there's no separate health port or shell
  for an exec probe (review M2).
- Added the Agent Router's internal per-backend routing headers
  (x-ai-eg-mcp-backend, x-ai-eg-mcp-route) to the Gateway-wide identity
  header strip, as defense-in-depth against the per-backend relocation ever
  failing (review M4, optional half).
- Anchored the flux-operator-mcp OCIRepository's cosign matchOIDCIdentity
  regexes, which were an unintended substring match, and pinned its image
  to a digest like VM/VL already are (review M8).
- Warned at each MCP chart/image pin that a bump exposing a new resource,
  prompt or template ships unauthorized, since those bypass MCPRoute
  authorization entirely (review M5).

* fix(ops): harden the MCP probe script and exercise reviewer-only tools

agent-probe-mcp.sh used fixed /tmp/auth and /tmp/h paths, so two classes
probed in parallel would clobber each other's token and session headers; no
curl call had a timeout; and a 401 on `initialize` surfaced only as a
confusing empty session ID on the next call. Switched to mktemp with a trap
cleanup, added -m 20 to every curl call, print the `initialize` status, added
a best-effort session DELETE, and an optional 4th argument to override the
listener port (for a live cross-class 401 check).

Added a reviewer-audience token projection to agent-probe.yaml: the Allow
rules that only reviewer/tester/triager get (VictoriaLogs tools) were never
exercised by any probe run (review M9).

* docs(superpowers): correct two residual notes in the identity design spec

C5 ("whether identity reaches the MCP backends") was marked UNVERIFIED; the
phase-5 review pre-flight settled it from source: yes, as x-ar-agent, via
oauth.claimToHeaders. SP2 no longer needs a fallback for this path.

T12's residual named only logs and ConfigMaps; cluster-wide `get pods` also
exposes pod specs and, under FallbackToLogsOnError, a crash log tail,
reachable by every internal run rather than only reviewer/tester/triager
(review M10, M3).

* docs(superpowers): mirror the design docs from #2092

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(agent-factory): ADR-0048, the factory under the agent-platform umbrella

The factory's signed chart source, HelmRelease (Task CRD created on install,
replaced on upgrade), values with the strictly parsed config, the App key
through agents-secrets, a default-deny CNP and its scrape, as an umbrella
child on both aws-0 and gcp-0 with a GP-14 render root per cloud. The broker's
allowlist admits the factory's ServiceAccount, and the cloud-shape gate now
judges the gcp-0 factory overlay.

* feat(agent-factory): send task spans to the trace collector

* fix(agent-factory): pin the chart by digest, pair it with the image, state ADR-0048's reasons

- OCIRepository carries ref.digest, which Flux resolves before the tag
- render-bundle renders a digest-pinned chartRef as <url>@<digest>
- test-agent-factory-pins.py fails when chart and image pins disagree;
  Renovate groups the two and now tracks the image tag in the values
- VMRule runbook reads every replica, since only the leader polls
- ADR-0048 gives reasons in place of programme ledger labels

* chore(agent-factory): pin the factory to agent-platform#7 at 569dc1c, by digest

Chart and image are the signed pre-release of that head.

* chore(agent-factory): pin the phase 2 pre-releases

agent-platform#10 at 800fcd8: the factory image and chart, and the room-broker image, a superset of AP-4's, for the app and the retention CronJob. The vendored Room CRD moves with the broker pin, as its comment requires; it gains dataClass immutability.

* chore(agent-factory): pin the phase 3 pre-releases; pair by default

agent-platform#12 at f2e51f9: the factory image and chart, and the room-broker image for the app and the retention CronJob (FA-3 leaves the broker's code and the Room CRD unchanged, so crd-rooms.yaml stays). defaults.template becomes pair: every task gets a reviewer until phase 4's triage picks templates.

* docs(adr): 0048 separates the decided design from what phase 1 builds

Kueue is not deployed and the factory enforces only the run budget; no gateway per-run ceiling exists yet. Keep the decision, add a decided/built table, and name SP3 phase 4 for Kueue. The trace collector's CNP comment no longer calls the factory rule inert.

* chore(agent-factory): re-pin the phase 3 pre-releases to agent-platform#12 at ddb06e0

* feat(agents): Kueue under the agent-platform umbrella (SP3 FR-4)

Kueue admits every factory sandbox pod: pod-only integration scoped to the
agents namespace, two non-borrowing ClusterQueues (factory, interactive) and
their LocalQueues, default-deny CNP with webhook/metrics/health ingress and
DNS/API egress, TLS metrics scraped by vmagent through a dedicated
metrics-reader ClusterRole.

Per ruling SB the target cloud is gcp-0: the controller child points at an
infrastructure/gcp-0/kueue render root (GP-14) and substitutes
gke-gcp-0-vars; the queue child is cloud-neutral and stays on the base.
agent-factory now dependsOn kueue-queues so submitted runs find their
LocalQueue. The Kueue chart joins the factory's no-automerge Renovate rule.

* feat(agent-factory): triage config, classes, shadow budgets, classifier egress

* feat(agent-policies): the factory is the only AgentRun creator; its patches are annotations only

* fix(agent-policies): bound the stop-readonly rule by resourceNames; DELETE carries no object

* feat(agent-factory): POST /v1/runs on the tailnet, client ids, 429 lookup egress

* feat(agent-run): task agent:run asks the factory's API; the token names the principal

* test(zitadel-oidc-clients): load stored_client_id, mock restart_rotated_consumers (#2149)

* feat(rooms): the broker requests runs from the factory

* docs(adr): ADR-0045 merge policy gate

* feat(ci): gate-path coverage (SC-12) and pull_request secrets lint (T8)

* fix(ci): the workflow-secrets lint catches permissions: write-all

* feat(merge-gate): namespace, OpenBao policy and role, namespaced store

* feat(policy-bot): the merge gate in merge-gate, its public hook and tailnet UI

* feat(merge-gate): .policy.yml: docs-links and revert live, gate paths unmergeable

* feat(merge-gate): agent-merge-gate ruleset source and its idempotent applier

* feat(merge-gate): policy-bot and its store under the agent-platform umbrella

* feat(merge-gate): approval-free agent rules require CI, secret scan included

* feat(merge-gate): the agent-merge ruleset and the agent-branches split, for after the wave

* feat(agent-factory): the merge gate in shadow, the merger key, link-rot

* fix(kueue): drop the visibility APIServices the API server cannot reach (#2245)

Kueue's manager serves its on-demand visibility API on 8082, which neither the
EKS node security group nor agent-system's CiliumNetworkPolicy opens, so both
APIServices sit FailedDiscoveryCheck: every discovery client logs errors and
namespace deletion blocks. Nothing reads that API, and chart 0.19.6 renders
the APIServices with no values switch, so a post-renderer removes them.

(cherry picked from commit e744c44)

* fix(openbao): the merge-gate mount and its secrets-admin grants (#2247)

* fix(openbao): create the merge-gate mount policy-bot's secret lives in

The merge-gate-secrets policy, merge-gate's SecretStore and its JWT role all
read a kv-v2 mount named merge-gate (SP3 R44) that no stack ever created: the
plan assumed SP2's Task 1.15a made it, but that task created agents/ only.
secrets-admin gains the same grant on it as on agents/, so a break-glass
administrator can write the App secret once per lineage.

(cherry picked from commit d594602)

* fix(openbao): give the shared secrets-admin policy the merge-gate grant too

d594602 added the merge-gate mount and its secrets-admin grant to the AWS
management stack only; validate-openbao-policies requires the shared copy to
match. GCP gains a grant on a mount it does not have yet, which is inert.

(cherry picked from commit 965e7f6)

* chore(crossplane): pin both clouds to CC v0.9.2 (harness v0.3.0)

v0.9.2 carries core v0.9.1, which pins agent-harness v0.3.0 (the disruption
build S5 published), and raises the aws and gcp packages' core floor to
>=v0.9.1 so the new core reaches clusters. Package digests: core f0b29186,
aws bf1ef35a, gcp 589d1310. The App Wizard clone tag and the doc mirrors follow;
the inference KCL module is still 0.9.0.

* chore(agent-factory): pin agent-platform v0.7.0, release signatures only (R20)

The factory image and chart, and the room broker with its retention CronJob, CRD
and schema, move to the v0.7.0 release (tag 7140faef, run 37721493733):

- agent-factory v0.7.0@sha256:83a9adc1, chart 0.7.0@sha256:977928e1 (appVersion v0.7.0);
- room-broker v0.7.0@sha256:c7a33043 in app.yaml and the retention CronJob;
- crd-rooms.yaml from the v0.7.0 asset (sha256 980ee0c5; body unchanged since v0.6.1);
- atlasSchema.ref v0.7.0 (no migration changed since v0.6.1).

The OCIRepository's subject narrows to release.yaml on v* tags (R20): ci.yaml
pre-releases and pull request builds no longer verify. All three artifacts pass
cosign verify against it.
Smana added a commit that referenced this pull request Oct 8, 2026
…8) (#2189)

* fix(ci): render oidc_issuer_url with its /id path, like oidc_issuer_host

The live issuer URL carries /id/<ID>; the bare-host fixture rendered the
agent-router SecurityPolicies with an issuer shape the cluster never has,
and left configuration-aws's EnvironmentConfig with an inconsistent pair.

* fix(agents): revoke under the same lock as token refresh

revoke() now holds CACHE + ".lock" when reading, revoking, and deleting
the cached token. This closes the race with concurrent token() exchanges
that could leave a live, unrevoked token in the cache if revoke() ran
without the lock.

The fix ensures that revoke() atomically reads the cache, revokes the token
at GitHub, and deletes the file — serializing with any concurrent token()
call via the same flock that token() uses.

Also clarify the Dockerfile comment to mention that both setuid and setgid
bits are stripped, matching what the find command actually does (-perm /6000).

Test: new test_revoke_waits_for_the_cache_lock verifies revoke() blocks
on the lock held by token(), preventing the interleaving that would leave
an unrevoked token in the cache.

Ran 19 tests, all pass.

* fix(agents): run agent-router on two replicas behind a PDB

With one replica, every drain or Karpenter consolidation of its node cut
all in-flight agent completions at once, and the harness does not retry.
envoyPDB minAvailable 1 keeps one proxy serving through a voluntary
disruption.

* fix(agents): drop agent-router's host ingress and name its rate-limit egress

The fromEntities: host rule on 8080/8081 admitted nothing the data plane
needs: EG's kubelet probes target the readiness and shutdown-manager
ports, which it never listed, and Cilium's default allow-localhost
already admits them. It only implied host-network pods were an intended
caller, which ADR-0042 says is agents pods alone.

The header now names the envoy-ratelimit :8081 egress SP4 must add with
its budgets: a dropped rate-limit check fails open.

* fix(security): retry agent-secrets after 30s, not a whole interval

On first apply the SecretStore can go Ready=False before openbao-ca holds
its CA; the failed health check then waited the 5m interval, holding
agent-router back by as much. 30s matches security-openbao.

* fix(security): call agents-secrets agent-system's store by convention

"agent-system's only way to secrets" read as a control. Nothing enforces
it: no ClusterSecretStore sets namespace conditions, and openbao-ca itself
reads through clustersecretstore.

* fix(scripts): let the identity probe's Sandbox expire if its delete is forgotten

agent-sandbox v1.0.3 keeps a terminated pod rather than recreating it,
but creates a new one, with fresh tokens, if that pod is ever removed.
shutdownPolicy Delete plus the documented shutdownTime patch bound the
Sandbox's life; the delete tolerates the Sandbox having expired first.

* fix(agents): audience mismatches are 403, and sts's audiences are the per-repo opt-in

Envoy's jwt_authn answers JwtAudienceNotAllowed with 403 and every other
failure with 401 (filter.cc at v1.39.1), so a cross-class or octo-sts
token is a 403 on a listener, not the 401 the comments claimed.

securitypolicy-sts now says its four audiences are the per-repository
opt-in: an AgentRun on another repository fails closed at its first
exchange, and opting one in means adding its audiences (EG: 8 per
provider, 4 providers per policy).

* docs(agents): record why the verified sub still reaches Z.ai

Removing x-ar-agent with a headerMutation on the zai backendRef edits the
same request headers the access log reads, so it would blank the
Gateway's own x_ar_agent attribution. The forwarding stays, with the
reason next to the claim that sets the header.

* fix(scripts): validate agent-run.sh inputs and surface the applied claim

Phase 6 review fixes (I1, minors 3-5): create (never apply, so a runId
collision surfaces as AlreadyExists instead of a silent CEL-rejected
update), validate --role/--class/--size/--minutes and AGENT_PRINCIPAL
against the design's principal CEL (human:<id>|system:<name>), and fail
cleanly when git user.email is unset instead of a raw set -e abort.
Success now also echoes the applied principal/role/class/repo/branch/
minutes to stderr, so a stale AGENT_PRINCIPAL or a --role typo is
visible without a follow-up kubectl get; stdout still carries only the
run name.

Also note (I2) that the dashboard's ar_agent label is produced by SP4
PR 1's envoy-ai-gateway controller on another branch, not this one, so
"No data" on that panel is expected until it merges. No absent() alert:
no agent run existing is not an outage.

Extends test-agent-run.sh for every new validation path, the
create-not-apply switch, --branch landing in the claim, an unknown
flag, and the property that a failed kubectl create prints no run
name.

* fix(observability): count 403s as rejected agent-router tokens

Envoy's jwt_authn filter returns 403 for wrong audience tokens, not just
401. Update the AgentRouterUnauthorizedBurst alert to count both status
codes and clarify the description for T8 (replayed/foreign) and R2
(identity-proxy) failure modes.

* fix(security): generate the MCP session-encryption seed in-cluster

The Agent Router chart falls back to the published "default-insecure-seed"
for controller.mcp.sessionEncryption.seed unless a value is supplied, which
lets anyone decrypt or forge a client-facing MCP session ID (phase-5 review
I1). Generate it with the same ESO Password + ExternalSecret pattern already
used for the AI-gateway rate-limit Valkey password, and feed it to the
HelmRelease via valuesFrom/targetPath, since the chart only renders the seed
into a controller CLI argument that has no env or file alternative.

The seed still lands in the extproc sidecar args of every AI-gateway
data-plane pod, so it remains readable by anything that can read pod specs.
SP2's room broker must authorize every call on x-ar-agent regardless.

* fix(ci): close two more MCP token-passthrough paths, rename gate A5 to A6

check_mcp_token_passthrough only ever looked at backendRefs[].forwardHeaders.
Agent Router also lets a route hand Authorization to a backend through
spec.securityPolicy.oauth.claimToHeaders[].header and
spec.securityPolicy.apiKeyAuth.forwardClientIDHeader (review I2); the
docstring's "the one way" was wrong. Both are now checked, case-insensitively,
with a failing test per path added first.

Renamed A5 to A6 throughout (code, messages, docstring, tests,
scripts/AGENTS.md): a parallel branch adds its own A5 (route sectionName)
to the same module, and the two need distinct ids before they merge.

Also cross-referenced the MCPRoute-level oauth issuer/audiences with their
Gateway-listener SecurityPolicy counterparts (agent-router/securitypolicy-
{public,internal}.yaml): a route-level SecurityPolicy replaces rather than
merges with the listener one, so the two must be kept in sync by hand
(review M7).

* fix(security): narrow agent-mcp's RBAC, egress and identity pins

- flux-operator-mcp's ClusterRole enumerated resources per apiGroup instead
  of `resources: ["*"]` under 8 groups, so a future Kind (Flux's own, or a
  Configuration bump under cloud.ogenki.io) needs a diff before an agent can
  read it (review M1).
- mcp-victoriametrics/mcp-victorialogs egress to `observability` now selects
  the vmsingle/victoria-logs-single pods by label, not the whole namespace
  (review M6).
- Noted the residual on each server's `fromEntities: host` ingress rule,
  needed for kubelet probes since there's no separate health port or shell
  for an exec probe (review M2).
- Added the Agent Router's internal per-backend routing headers
  (x-ai-eg-mcp-backend, x-ai-eg-mcp-route) to the Gateway-wide identity
  header strip, as defense-in-depth against the per-backend relocation ever
  failing (review M4, optional half).
- Anchored the flux-operator-mcp OCIRepository's cosign matchOIDCIdentity
  regexes, which were an unintended substring match, and pinned its image
  to a digest like VM/VL already are (review M8).
- Warned at each MCP chart/image pin that a bump exposing a new resource,
  prompt or template ships unauthorized, since those bypass MCPRoute
  authorization entirely (review M5).

* fix(ops): harden the MCP probe script and exercise reviewer-only tools

agent-probe-mcp.sh used fixed /tmp/auth and /tmp/h paths, so two classes
probed in parallel would clobber each other's token and session headers; no
curl call had a timeout; and a 401 on `initialize` surfaced only as a
confusing empty session ID on the next call. Switched to mktemp with a trap
cleanup, added -m 20 to every curl call, print the `initialize` status, added
a best-effort session DELETE, and an optional 4th argument to override the
listener port (for a live cross-class 401 check).

Added a reviewer-audience token projection to agent-probe.yaml: the Allow
rules that only reviewer/tester/triager get (VictoriaLogs tools) were never
exercised by any probe run (review M9).

* docs(superpowers): correct two residual notes in the identity design spec

C5 ("whether identity reaches the MCP backends") was marked UNVERIFIED; the
phase-5 review pre-flight settled it from source: yes, as x-ar-agent, via
oauth.claimToHeaders. SP2 no longer needs a fallback for this path.

T12's residual named only logs and ConfigMaps; cluster-wide `get pods` also
exposes pod specs and, under FallbackToLogsOnError, a crash log tail,
reachable by every internal run rather than only reviewer/tester/triager
(review M10, M3).

* docs(superpowers): mirror the design docs from #2092

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(agent-factory): ADR-0048, the factory under the agent-platform umbrella

The factory's signed chart source, HelmRelease (Task CRD created on install,
replaced on upgrade), values with the strictly parsed config, the App key
through agents-secrets, a default-deny CNP and its scrape, as an umbrella
child on both aws-0 and gcp-0 with a GP-14 render root per cloud. The broker's
allowlist admits the factory's ServiceAccount, and the cloud-shape gate now
judges the gcp-0 factory overlay.

* feat(agent-factory): send task spans to the trace collector

* fix(agent-factory): pin the chart by digest, pair it with the image, state ADR-0048's reasons

- OCIRepository carries ref.digest, which Flux resolves before the tag
- render-bundle renders a digest-pinned chartRef as <url>@<digest>
- test-agent-factory-pins.py fails when chart and image pins disagree;
  Renovate groups the two and now tracks the image tag in the values
- VMRule runbook reads every replica, since only the leader polls
- ADR-0048 gives reasons in place of programme ledger labels

* chore(agent-factory): pin the factory to agent-platform#7 at 569dc1c, by digest

Chart and image are the signed pre-release of that head.

* chore(agent-factory): pin the phase 2 pre-releases

agent-platform#10 at 800fcd8: the factory image and chart, and the room-broker image, a superset of AP-4's, for the app and the retention CronJob. The vendored Room CRD moves with the broker pin, as its comment requires; it gains dataClass immutability.

* chore(agent-factory): pin the phase 3 pre-releases; pair by default

agent-platform#12 at f2e51f9: the factory image and chart, and the room-broker image for the app and the retention CronJob (FA-3 leaves the broker's code and the Room CRD unchanged, so crd-rooms.yaml stays). defaults.template becomes pair: every task gets a reviewer until phase 4's triage picks templates.

* docs(adr): 0048 separates the decided design from what phase 1 builds

Kueue is not deployed and the factory enforces only the run budget; no gateway per-run ceiling exists yet. Keep the decision, add a decided/built table, and name SP3 phase 4 for Kueue. The trace collector's CNP comment no longer calls the factory rule inert.

* chore(agent-factory): re-pin the phase 3 pre-releases to agent-platform#12 at ddb06e0

* feat(agents): Kueue under the agent-platform umbrella (SP3 FR-4)

Kueue admits every factory sandbox pod: pod-only integration scoped to the
agents namespace, two non-borrowing ClusterQueues (factory, interactive) and
their LocalQueues, default-deny CNP with webhook/metrics/health ingress and
DNS/API egress, TLS metrics scraped by vmagent through a dedicated
metrics-reader ClusterRole.

Per ruling SB the target cloud is gcp-0: the controller child points at an
infrastructure/gcp-0/kueue render root (GP-14) and substitutes
gke-gcp-0-vars; the queue child is cloud-neutral and stays on the base.
agent-factory now dependsOn kueue-queues so submitted runs find their
LocalQueue. The Kueue chart joins the factory's no-automerge Renovate rule.

* feat(agent-factory): triage config, classes, shadow budgets, classifier egress

* feat(agent-policies): the factory is the only AgentRun creator; its patches are annotations only

* fix(agent-policies): bound the stop-readonly rule by resourceNames; DELETE carries no object

* feat(agent-factory): POST /v1/runs on the tailnet, client ids, 429 lookup egress

* feat(agent-run): task agent:run asks the factory's API; the token names the principal

* test(zitadel-oidc-clients): load stored_client_id, mock restart_rotated_consumers (#2149)

* feat(rooms): the broker requests runs from the factory

* docs(adr): ADR-0045 merge policy gate

* feat(ci): gate-path coverage (SC-12) and pull_request secrets lint (T8)

* fix(ci): the workflow-secrets lint catches permissions: write-all

* feat(merge-gate): namespace, OpenBao policy and role, namespaced store

* feat(policy-bot): the merge gate in merge-gate, its public hook and tailnet UI

* feat(merge-gate): .policy.yml: docs-links and revert live, gate paths unmergeable

* feat(merge-gate): agent-merge-gate ruleset source and its idempotent applier

* feat(merge-gate): policy-bot and its store under the agent-platform umbrella

* feat(merge-gate): approval-free agent rules require CI, secret scan included

* feat(merge-gate): the agent-merge ruleset and the agent-branches split, for after the wave

* feat(agent-factory): the merge gate in shadow, the merger key, link-rot

* feat(agent-factory): the factory's alerts and dashboard, inside the umbrella

* docs(runbooks): materialize agent-factory runbooks 01-08 and their README from integration/agent-factory

* docs(runbooks): 09 App key compromise, all four Apps

* feat(agent-factory): pin the factory to agent-platform#16; the control issue

* fix(kueue): drop the visibility APIServices the API server cannot reach (#2245)

Kueue's manager serves its on-demand visibility API on 8082, which neither the
EKS node security group nor agent-system's CiliumNetworkPolicy opens, so both
APIServices sit FailedDiscoveryCheck: every discovery client logs errors and
namespace deletion blocks. Nothing reads that API, and chart 0.19.6 renders
the APIServices with no values switch, so a post-renderer removes them.

(cherry picked from commit e744c44)

* fix(openbao): the merge-gate mount and its secrets-admin grants (#2247)

* fix(openbao): create the merge-gate mount policy-bot's secret lives in

The merge-gate-secrets policy, merge-gate's SecretStore and its JWT role all
read a kv-v2 mount named merge-gate (SP3 R44) that no stack ever created: the
plan assumed SP2's Task 1.15a made it, but that task created agents/ only.
secrets-admin gains the same grant on it as on agents/, so a break-glass
administrator can write the App secret once per lineage.

(cherry picked from commit d594602)

* fix(openbao): give the shared secrets-admin policy the merge-gate grant too

d594602 added the merge-gate mount and its secrets-admin grant to the AWS
management stack only; validate-openbao-policies requires the shared copy to
match. GCP gains a grant on a mount it does not have yet, which is inert.

(cherry picked from commit 965e7f6)

* chore(crossplane): pin both clouds to CC v0.9.2 (harness v0.3.0)

v0.9.2 carries core v0.9.1, which pins agent-harness v0.3.0 (the disruption
build S5 published), and raises the aws and gcp packages' core floor to
>=v0.9.1 so the new core reaches clusters. Package digests: core f0b29186,
aws bf1ef35a, gcp 589d1310. The App Wizard clone tag and the doc mirrors follow;
the inference KCL module is still 0.9.0.

* chore(agent-factory): pin agent-platform v0.7.0, release signatures only (R20)

The factory image and chart, and the room broker with its retention CronJob, CRD
and schema, move to the v0.7.0 release (tag 7140faef, run 37721493733):

- agent-factory v0.7.0@sha256:83a9adc1, chart 0.7.0@sha256:977928e1 (appVersion v0.7.0);
- room-broker v0.7.0@sha256:c7a33043 in app.yaml and the retention CronJob;
- crd-rooms.yaml from the v0.7.0 asset (sha256 980ee0c5; body unchanged since v0.6.1);
- atlasSchema.ref v0.7.0 (no migration changed since v0.6.1).

The OCIRepository's subject narrows to release.yaml on v* tags (R20): ci.yaml
pre-releases and pull request builds no longer verify. All three artifacts pass
cosign verify against it.
Smana added a commit that referenced this pull request Oct 8, 2026
)

* fix(security): generate the MCP session-encryption seed in-cluster

The Agent Router chart falls back to the published "default-insecure-seed"
for controller.mcp.sessionEncryption.seed unless a value is supplied, which
lets anyone decrypt or forge a client-facing MCP session ID (phase-5 review
I1). Generate it with the same ESO Password + ExternalSecret pattern already
used for the AI-gateway rate-limit Valkey password, and feed it to the
HelmRelease via valuesFrom/targetPath, since the chart only renders the seed
into a controller CLI argument that has no env or file alternative.

The seed still lands in the extproc sidecar args of every AI-gateway
data-plane pod, so it remains readable by anything that can read pod specs.
SP2's room broker must authorize every call on x-ar-agent regardless.

* fix(ci): close two more MCP token-passthrough paths, rename gate A5 to A6

check_mcp_token_passthrough only ever looked at backendRefs[].forwardHeaders.
Agent Router also lets a route hand Authorization to a backend through
spec.securityPolicy.oauth.claimToHeaders[].header and
spec.securityPolicy.apiKeyAuth.forwardClientIDHeader (review I2); the
docstring's "the one way" was wrong. Both are now checked, case-insensitively,
with a failing test per path added first.

Renamed A5 to A6 throughout (code, messages, docstring, tests,
scripts/AGENTS.md): a parallel branch adds its own A5 (route sectionName)
to the same module, and the two need distinct ids before they merge.

Also cross-referenced the MCPRoute-level oauth issuer/audiences with their
Gateway-listener SecurityPolicy counterparts (agent-router/securitypolicy-
{public,internal}.yaml): a route-level SecurityPolicy replaces rather than
merges with the listener one, so the two must be kept in sync by hand
(review M7).

* fix(security): narrow agent-mcp's RBAC, egress and identity pins

- flux-operator-mcp's ClusterRole enumerated resources per apiGroup instead
  of `resources: ["*"]` under 8 groups, so a future Kind (Flux's own, or a
  Configuration bump under cloud.ogenki.io) needs a diff before an agent can
  read it (review M1).
- mcp-victoriametrics/mcp-victorialogs egress to `observability` now selects
  the vmsingle/victoria-logs-single pods by label, not the whole namespace
  (review M6).
- Noted the residual on each server's `fromEntities: host` ingress rule,
  needed for kubelet probes since there's no separate health port or shell
  for an exec probe (review M2).
- Added the Agent Router's internal per-backend routing headers
  (x-ai-eg-mcp-backend, x-ai-eg-mcp-route) to the Gateway-wide identity
  header strip, as defense-in-depth against the per-backend relocation ever
  failing (review M4, optional half).
- Anchored the flux-operator-mcp OCIRepository's cosign matchOIDCIdentity
  regexes, which were an unintended substring match, and pinned its image
  to a digest like VM/VL already are (review M8).
- Warned at each MCP chart/image pin that a bump exposing a new resource,
  prompt or template ships unauthorized, since those bypass MCPRoute
  authorization entirely (review M5).

* fix(ops): harden the MCP probe script and exercise reviewer-only tools

agent-probe-mcp.sh used fixed /tmp/auth and /tmp/h paths, so two classes
probed in parallel would clobber each other's token and session headers; no
curl call had a timeout; and a 401 on `initialize` surfaced only as a
confusing empty session ID on the next call. Switched to mktemp with a trap
cleanup, added -m 20 to every curl call, print the `initialize` status, added
a best-effort session DELETE, and an optional 4th argument to override the
listener port (for a live cross-class 401 check).

Added a reviewer-audience token projection to agent-probe.yaml: the Allow
rules that only reviewer/tester/triager get (VictoriaLogs tools) were never
exercised by any probe run (review M9).

* docs(superpowers): correct two residual notes in the identity design spec

C5 ("whether identity reaches the MCP backends") was marked UNVERIFIED; the
phase-5 review pre-flight settled it from source: yes, as x-ar-agent, via
oauth.claimToHeaders. SP2 no longer needs a fallback for this path.

T12's residual named only logs and ConfigMaps; cluster-wide `get pods` also
exposes pod specs and, under FallbackToLogsOnError, a crash log tail,
reachable by every internal run rather than only reviewer/tester/triager
(review M10, M3).

* docs(superpowers): mirror the design docs from #2092

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(agent-factory): ADR-0048, the factory under the agent-platform umbrella

The factory's signed chart source, HelmRelease (Task CRD created on install,
replaced on upgrade), values with the strictly parsed config, the App key
through agents-secrets, a default-deny CNP and its scrape, as an umbrella
child on both aws-0 and gcp-0 with a GP-14 render root per cloud. The broker's
allowlist admits the factory's ServiceAccount, and the cloud-shape gate now
judges the gcp-0 factory overlay.

* feat(agent-factory): send task spans to the trace collector

* fix(agent-factory): pin the chart by digest, pair it with the image, state ADR-0048's reasons

- OCIRepository carries ref.digest, which Flux resolves before the tag
- render-bundle renders a digest-pinned chartRef as <url>@<digest>
- test-agent-factory-pins.py fails when chart and image pins disagree;
  Renovate groups the two and now tracks the image tag in the values
- VMRule runbook reads every replica, since only the leader polls
- ADR-0048 gives reasons in place of programme ledger labels

* chore(agent-factory): pin the factory to agent-platform#7 at 569dc1c, by digest

Chart and image are the signed pre-release of that head.

* chore(agent-factory): pin the phase 2 pre-releases

agent-platform#10 at 800fcd8: the factory image and chart, and the room-broker image, a superset of AP-4's, for the app and the retention CronJob. The vendored Room CRD moves with the broker pin, as its comment requires; it gains dataClass immutability.

* chore(agent-factory): pin the phase 3 pre-releases; pair by default

agent-platform#12 at f2e51f9: the factory image and chart, and the room-broker image for the app and the retention CronJob (FA-3 leaves the broker's code and the Room CRD unchanged, so crd-rooms.yaml stays). defaults.template becomes pair: every task gets a reviewer until phase 4's triage picks templates.

* docs(adr): 0048 separates the decided design from what phase 1 builds

Kueue is not deployed and the factory enforces only the run budget; no gateway per-run ceiling exists yet. Keep the decision, add a decided/built table, and name SP3 phase 4 for Kueue. The trace collector's CNP comment no longer calls the factory rule inert.

* chore(agent-factory): re-pin the phase 3 pre-releases to agent-platform#12 at ddb06e0

* feat(agents): Kueue under the agent-platform umbrella (SP3 FR-4)

Kueue admits every factory sandbox pod: pod-only integration scoped to the
agents namespace, two non-borrowing ClusterQueues (factory, interactive) and
their LocalQueues, default-deny CNP with webhook/metrics/health ingress and
DNS/API egress, TLS metrics scraped by vmagent through a dedicated
metrics-reader ClusterRole.

Per ruling SB the target cloud is gcp-0: the controller child points at an
infrastructure/gcp-0/kueue render root (GP-14) and substitutes
gke-gcp-0-vars; the queue child is cloud-neutral and stays on the base.
agent-factory now dependsOn kueue-queues so submitted runs find their
LocalQueue. The Kueue chart joins the factory's no-automerge Renovate rule.

* feat(agent-factory): triage config, classes, shadow budgets, classifier egress

* feat(agent-policies): the factory is the only AgentRun creator; its patches are annotations only

* fix(agent-policies): bound the stop-readonly rule by resourceNames; DELETE carries no object

* feat(agent-factory): POST /v1/runs on the tailnet, client ids, 429 lookup egress

* feat(agent-run): task agent:run asks the factory's API; the token names the principal

* test(zitadel-oidc-clients): load stored_client_id, mock restart_rotated_consumers (#2149)

* feat(rooms): the broker requests runs from the factory

* docs(adr): ADR-0045 merge policy gate

* feat(ci): gate-path coverage (SC-12) and pull_request secrets lint (T8)

* fix(ci): the workflow-secrets lint catches permissions: write-all

* feat(merge-gate): namespace, OpenBao policy and role, namespaced store

* feat(policy-bot): the merge gate in merge-gate, its public hook and tailnet UI

* feat(merge-gate): .policy.yml: docs-links and revert live, gate paths unmergeable

* feat(merge-gate): agent-merge-gate ruleset source and its idempotent applier

* feat(merge-gate): policy-bot and its store under the agent-platform umbrella

* feat(merge-gate): approval-free agent rules require CI, secret scan included

* feat(merge-gate): the agent-merge ruleset and the agent-branches split, for after the wave

* feat(agent-factory): the merge gate in shadow, the merger key, link-rot

* feat(agent-factory): the factory's alerts and dashboard, inside the umbrella

* docs(runbooks): materialize agent-factory runbooks 01-08 and their README from integration/agent-factory

* docs(runbooks): 09 App key compromise, all four Apps

* feat(agent-factory): pin the factory to agent-platform#16; the control issue

* feat(runlore): findings reach the agent factory's intake, inside the umbrella

* fix(runlore): the intake env survives Flux's values merge order

* feat(agents): the aws-0 umbrella carries Kueue too (aws-0 is the live target again)

* fix(rooms): the retention CronJob joins the broker's pr14 pin (one image, one version)

* feat(agent-factory): resume runs lost to a spot reclaim or an eviction (#2222)

* docs(agents): design agent runs that survive a spot or preemptible reclaim

An external review found runs on reclaimable capacity lose their work
and are never resumed. The design budgets the pod's shutdown at 15 s
(GKE preemptible VMs cannot extend it): pause, checkpoint commit and
push, the bridge's final read, then stop and revoke. The composition
reads the pod to record Disrupted, PodLost or PodFailed, and the
factory resumes Disrupted and PodLost runs at most twice within the
task's token cap. An early warning on aws-0 waits on a fault-injection
test. The research records the facts at pinned commits.

* docs(agents): plan agent runs that survive a reclaim

Thirteen tasks across agent-platform, cloud-native-ref and
crossplane-configuration: the bridge's final read, agent-run's
15 s shutdown sequence, the composition reading the pod for
Disrupted/PodLost/PodFailed, GKE Spot's 120 s window, the factory's
automatic resume, docs, runbook 09 and live verification. The spec is
aligned with what the code showed: no XRD enum, a deleted pod reads
PodLost, Crossplane needs cluster-wide pod reads for required resources,
interrupt rather than pause, an in-process revoke, the trailer from the
commit-msg hook.

* docs(agents): record the owner's decisions on the disruption plan

The wider Crossplane pod read is accepted, and merging the factory resume
work may bring factory phases 2-9 to gcp-0. The two owner actions the
plan depends on are listed.

* docs(agents): order the agent factory's remaining work

One table from the open reviews to the merge wave, naming the plan that
holds each step: the disruption plan's runtime half rides the same re-pin
and rebuild as the F10-F18 fixes, the agentgateway migration (with prompt
caching) follows the live re-verification, then SP2 and SP3.

* docs(agents): link the kagent evaluation now that it is on main

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(agent-factory): pin automatic resume (agent-platform#25), resume.maxPerTask, resumes on the dashboards

* docs(runbooks): runbook 10, disruption live checks; the factory resumes a lost run

* docs(agents): the disruption runbook is 10; 09 is the App key-compromise runbook

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(runbooks): runbook 10 waits on resumes and new runs, reads resumes_exhausted; resumes are per task

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(agents): pin the factory to agent-platform#25 at 642a5b12; enforce the task cap

642a5b12 carries the resume review fixes and #30's whole-run task cap.
enforceTask: true matches integration (owner, 2026-10-07).

* test(zitadel): the rooms-proxy suite loads previous_client_id after #2208

cmd_sync reads the previous client through previous_client_id since #2208;
the suite did not load it (exit 127). Same lines as integration/agent-factory.

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(agents): pin the factory to agent-platform#25 at ef31783e

ef31783e quotes a review queued while a lost run was live in the resumed run's brief,
and escalates a lost reviewer past the enforced task cap.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* feat(eks): manage the agent-run FIS role and experiment template as code (#2233)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(eks): manage the agent-run FIS role and experiment template as code

Runbook 10 Step 7 (the aws-0 Spot-interruption test, disruption plan Task 13)
told the owner to create an IAM role and an FIS experiment template with aws CLI
commands. Both now live in opentofu/aws/eks/init, the stack that owns aws-0's
node and Karpenter IAM, so an aws-0 rebuild creates them.

The role trusts fis.amazonaws.com only for this account's experiments in the
region (aws:SourceAccount, aws:SourceArn) and may interrupt only instances
tagged agents.ogenki.io/fis-target=true. The runbook now tags the run's node,
starts the template, and untags the node in cleanup.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ci): restore the programme's merge-gate and pre-release XRD CRD steps (#2235)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release external-secrets to v2.12.0 (#2232)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(ci): restore the programme's merge-gate and pre-release XRD CRD steps

#2233 resolved its main merge to main's ci.yaml, which never had them: the
Merge-gate invariants (SP3 SC-12) and the pre-release XRD CRD fetch (SP2 P40)
left feat/factory-runlore. Restored byte-identical to c7fb11c8 and integration;
main's newer checkov and trufflehog pins stay.

* test(ci): a main merge cannot drop the programme's CI steps unnoticed

Nothing read ci.yaml for them, so their loss in #2233 passed every local gate.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* fix(agent-factory): drift detection corrects a kubectl scale of the factory (#2237)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release external-secrets to v2.12.0 (#2232)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(agent-factory): drift detection corrects a kubectl scale of the factory

A kill-switch drill scaled the factory to 0; on resume Helm's 3-way merge re-applied
only what the chart changed, so it stayed at 0 for 8.5 h. Drift detection restores the
chart's 2 replicas on the next reconcile. No ignore rules: nothing else writes the
chart's objects. The plan's drill now suspends the HelmRelease, which drift detection
would otherwise undo mid-drill, and checks 2/2 after the restore: rollout status alone
passes at 0 of 0.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(runbooks): runbook 10's aws-0 results; the Karpenter reason label; capture logs live (#2243)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(c…
Smana added a commit that referenced this pull request Oct 8, 2026
…enderer (SP3 10.1) (#2192)

* fix(ai-gateway): retry envoy-ai-gateway upgrades cancelled by the seed Secret

On aws-0 the first upgrade carrying the MCP session-seed valuesFrom was
cancelled when ESO wrote the new Secret a second time (watch label), and with
no upgrade remediation the HelmRelease stalled with RetriesExceeded.

* fix(agents): strip the build suffix from agent-sandbox's chart label

reconcileStrategy: Revision versions the chart 0.1.0+<git sha>, and the chart
copies that verbatim into every object's helm.sh/chart label. `+` is illegal
in a label value, so the API server rejected the release on aws-0
(ServiceMonitor first; 5 objects affected). A postRenderer replaces the label.
Proven with helm template + kustomize on the v1.0.3 chart: 5 illegal values
before, 0 after, 7 objects kept.

* fix(agent-router): spread envoy replicas across zones and nodes

Two replicas had no topologySpreadConstraints, so both could land on
the same node with nothing to catch the loss.

* fix(agent-github): fail loudly on a missing ruleset name

jq -r on a missing .name silently reads as the string "null" and the
script proceeds; -e makes jq exit nonzero instead.

* fix(agent-harness): escape literal dots in cosign identity regexes

The flux-operator-mcp OCIRepository verify block anchors its issuer and
subject regexes but leaves the domain dots unescaped, so any character
would match in place of a literal ".".

* fix(agent-e2e): pin the owning gateway's namespace in the rejected-token alert

The AgentRouterUnauthorizedBurst LogsQL selector matched on
owning-gateway-name alone, so a same-named Gateway in another
namespace would feed the same counter.

* fix(agent-e2e): reject an overflowing --minutes value before arithmetic

A digit string too large for bash's integer comparison makes both
`[ -lt ]` and `[ -gt ]` fail their own check instead of the range test,
so the value slips through under set -e. Reject anything longer than
3 digits first, using the script's existing usage-error path. Adds a
regression case to test-agent-run.sh.

* fix(agent-harness): retry agent-mcp fast on rollout (M1)

Same reasoning as agent-router and octo-sts: the Deployments/HelmRelease can
still be rolling out on first apply, so retry in 30s rather than waiting a
full 5m interval.

* fix(agent-harness): block automerge on the MCP servers, restate the mcp image repo (I2)

- renovate.json: automerge:false packageRule for flux-operator-mcp (chart and
  image), mcp-victoriametrics, mcp-victorialogs and octo-sts/app, placed after
  the blanket automerge rule so it wins. Each server's release notes must be
  read before bumping: resources/prompts bypass MCPRoute authorization.
  Validated with renovate-config-validator and python3 -m json.tool.
- flux-operator-mcp-helmrelease.yaml: restate image.repository (chart default,
  confirmed via `helm pull`) next to image.tag so the helm-values manager
  tracks the image alongside the chart instead of only the digest moving.

* fix(agent-harness): run each image's test stage in CI (M4)

Generic detection (grep for 'AS test' in the Dockerfile), not agent-harness
specific: build --target test before build-push and fail the job on failure.
Without it, a Renovate-automerged FROM-digest bump rebuilds and republishes
agent-harness untested. actionlint's remaining findings are pre-existing
shellcheck info/style notes on other steps, unchanged by this diff.

* fix(agent-router): retry fast on rollout (M1)

Envoy Gateway reports Programmed=False while the proxy Deployment is still
rolling out; retryInterval 30s mirrors agent-secrets rather than waiting a
full 5m interval.

* docs(agent-router): document existing-cluster owner prereqs (M3)

Resume needs opentofu/aws/openbao/management then opentofu/aws/eks/configure
applied first on an existing cluster, or SecretStore agents-secrets never
goes Ready. Feature-branch clusters also need
TF_VAR_flux_git_ref=refs/heads/<branch> for eks/configure.

* docs(agent-router): confirm identity reaches MCP backends (M6)

ADR-0042 still called this UNVERIFIED (C5); the design spec now records it as
confirmed from source: x-ar-agent, via the MCPRoute's
securityPolicy.oauth.claimToHeaders. SP2 no longer needs the fallback for
this path.

* fix(agent-github): retry octo-sts fast on rollout (M1)

Same reasoning as agent-router: retryInterval 30s so a first-apply rollout
does not wait a full 5m interval.

* docs(agent-github): correct octo-sts ingress claim (M6)

octo-sts's CNP also admits node-local host on :8080 (security/base/octo-sts/
network-policy.yaml), not just agent-router's data plane; that reach can
already read the mounted App key, so it's harmless but the safety comment
should say so.

* fix(agent-e2e): match Karpenter 1.14.1's plural NodePool metric names (I1)

karpenter_nodepool_usage/limit match nothing at the pinned 1.14.1
(NodePoolSubsystem = "nodepools"); rename to karpenter_nodepools_usage/limit
in the AgentGvisorPoolNearLimit rule and dashboard panel 3. Labels nodepool
and resource_type are unchanged.

* fix(observability): match Karpenter 1.14.1's plural NodePool metric names (I1)

Same bug as agent-platform's vmrule.yaml, in the copy source: at the pinned
1.14.1, the subsystem is nodepools, so KarpenterNodepoolAlmostFull matches no
series and never fires.

* fix(agent-e2e): add a token-spend watchdog for agent-router (I3)

No per-run cap is enforced until SP4 PR 2 (design spec line 162): a run is
bounded only by the gateway's 5M ceiling today. AgentRunTokenSpendHigh (warning,
per ar_agent, 5M/8h) and AgentFleetTokenSpendHigh (critical, fleet-wide,
10M/1h) reuse the dashboard's ar_agent selector and the gen_ai_token_type
input|output filter from llm-gateway's FrontierSpendGuardTripped. Descriptions
carry the runbook's manual-revoke command
(agents.ogenki.io/revoked=budget-run).

* docs(agent-e2e): drop stale PR-merge note, pin gateway namespace on panel 4 (M6)

SP4 PR 1 is part of this package, so 'No data until that PR merges' no longer
applies. Panel 4's log selector now also pins
owning-gateway-namespace:"agent-system", matching vmrule-logs.yaml, so a
same-named Gateway in another namespace cannot match.

* docs(agent-e2e): note that internal runs have no model route yet (M2)

--class internal 404s on every model call until SP4 PR 2 lands
agent-models-internal; still accepted (not refused) because the runbooks use
it to test the internal listener. shellcheck clean.

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* docs(superpowers): mirror the design docs from #2092

* fix(agents): give the VictoriaMetrics and VictoriaLogs MCP servers startup memory

mcp-victoriametrics 1.20.2 was OOMKilled (exit 137) on every start under a 128Mi
limit on aws-0; mcp-victorialogs idled at 89Mi of the same limit.

* fix(agents): let the MCP servers finish loading before liveness applies

mcp-victoriametrics opens :8081 only after loading its docs index. At a 200m CPU
limit that outlasted the liveness window, so the kubelet restarted it in a loop
(7 restarts, never Ready, on aws-0). A startupProbe (up to 5 min) holds liveness
off, and a 1-CPU limit shortens the load.

* fix(agents): size the VictoriaMetrics MCP server for its resident docs index

Its documentation tool indexes the whole embedded docs site in memory at startup
and keeps it: ~700MiB resident, measured locally on v1.20.2 (listening after 14s
at 1 CPU). The 512Mi limit was OOMKilled 18s after start on aws-0.

* docs(adr): ADR-0043 records the hosted octo-sts App as a risk

The trust policies accept any eu-west-3 EKS issuer, safe only because agent-router's
sts listener verifies this cluster's issuer first. Chainguard's hosted octo-sts reads
the same files without that check, so it must never be installed alongside them.

* docs(adr): ADR-0043 records that every human merge to main is a ruleset bypass

agent-branches covers main, so a plain merge is refused and the owner must bypass
explicitly (seen merging #2113); branch protection still applies.

* fix(agents): stop agent-router answering /v1/models before authentication

* feat(agents): serve agent-default with GLM-5.3

Same switch as SP4 PR 1 (#2105) for the llm-gateway files, byte-identical, plus the
agents' own route. Verified on the agents' key: HTTP 200, served model glm-5.3.

* fix(agents): OpenHands 1.49.6 on its tested dependency set, litellm below 1.95.1

litellm 1.102.1 (what `>=1.93.0` resolved to) deletes an unset cache_creation_tokens, and the
SDK's usage accounting raises AttributeError on the first GLM response, failing every run
(OpenHands/software-agent-sdk#5213 and five duplicates, open). Dependencies now resolve against
OpenHands' own uv.lock for v1.49.6, the latest release, and litellm is capped below 1.95.1.

* feat(agents): a per-pod OH_SECRET_KEY and per-token prices in the harness

agent-server warned on every start that OH_SECRET_KEY was unset (its stored secrets and MCP
OAuth state went unencrypted), and the SDK warned on every call that litellm could not price the
agent-default alias. The key is now generated per pod; the LLM carries input/output prices
(GLM-5.3 list price by default, LLM_*_USD_PER_MTOK override).

* feat(agents): a step log -- the harness prints each agent step to stdout

Nothing showed what an agent was doing: its steps lived only in agent-server's conversation API
inside the pod, and died with it. The harness now prints one line per action (tool, summary,
command or path; never outputs), each agent message, errors, and at the end a summary with the
final message in full -- a read-only role's report. kubectl logs shows it live, and VictoriaLogs
keeps it after the pod is gone. Best effort: it never fails the run.

* fix(agents): send reasoning_effort on every model call

An agent took about a minute per step, and #2112's run hit its 30-minute deadline before editing
anything: 98% of its time was inside model calls (median 25 s, p90 249 s, two 300 s timeouts).
litellm drops reasoning_effort for the unknown agent-default alias, so GLM-5.3 fell back to maximum
thinking: 150 s for a reply that takes 10.8 s at "high". The harness now sends it in the body
(LLM_REASONING_EFFORT, default high), with a test that it reaches the wire.

* fix(agents): let Crossplane read run pods for the CNP Usage

The AgentRun composition now keys the Usage that holds a run's CNP on the
run's Pod instead of its Sandbox (crossplane-configuration#29). The
Sandbox leaves the API at once, while the pod still revokes its GitHub
token. Crossplane GETs that pod, so it needs `get pods` in `agents`, and
nothing wider.

* ci: validate manifests against a pinned pre-release's XRD CRDs

* chore(docs): pin the agent-platform umbrella's suspend in doc-claims

* fix(ops): the agent probe resolves DNS over TCP too

* fix(agent-sandbox): enumerate Crossplane's verbs on sandboxes

* fix(observability): runbook and dashboard links on every agent-platform alert

* fix(agent-harness): redact GitHub tokens from the step log

* fix(ci): gate A3 needs every listener of a Gateway stripped

* fix(agent-mcp): trim introspection tools and cluster-wide reads from internal runs

* test(agent-mcp): allowlist the MCP scope test instead of denylisting three tool names

* fix(ci): move XRD CRDs fetch right before manifest validation

No suite reads XRD_CRDS_FILE, so the fetch only needs to happen before
validate-manifests consumes it. Running it after the test suites, gated on
!cancelled(), keeps it from being skipped when an earlier suite fails.

* fix(agent-mcp): drop VictoriaLogs flags tool, fix stale comments

- Point agent-platform runbook_url links at the integration branch, where
  the runbooks actually live until Phase 7 re-points them (Ruling G).
- Remove the VictoriaLogs `flags` tool from the backend toolSelector and
  from the reviewer/tester/triager grants: it is operator introspection,
  the same class already trimmed from VictoriaMetrics (M2). Update the
  scope test's expected sets accordingly.
- Reflow the truncated comment in flux-operator-mcp-rbac.yaml.
- Reword the scope test's header comment: it only parses
  flux-operator-mcp-rbac.yaml, not every RoleBinding in the repo.

* chore(crossplane): pin CC-H1's pre-release so the evidence gate runs on its XRD CRDs

* fix(openbao): the agents' secrets on their own mount, on both clouds (SP2 P38)

* fix(openbao): the agents mount is documented and its boundary enforced by test

* fix(agents): the token issuer and its JWKS are per-cloud variables

* docs(secrets): point the prose at the External Secrets row, not the last row

* feat(gcp): a GKE Sandbox pool for agent runs, and per-packet LB for gVisor

* feat(gcp): AgentRun's pre-release package and Kyverno on gcp-0

* feat(ci): one render root per cloud for every substituted agent base; gcp-0's CA key and gateway keys

* fix(gcp): gVisor runs tolerate the Cilium taint, so the sandbox pool can scale from zero

A gcp-0 Kyverno policy adds the node.cilium.io/agent-not-ready toleration to every gVisor pod (ADR-0006). The pool moves to pd-standard; the smoke probe tolerates the taint, asserts gVisor and has a deadline. socketLB.hostNamespaceOnly stays as an explicit guard: the chart already forces it with Gateway API since Cilium 1.20. Also fixes three 6.4 review comments.

* fix(gcp): workloads reach the metadata server by CIDR, not the host entity

On GKE 169.254.169.254 is never node-local: iptables DNATs it to gke-metadata-server after Cilium has classified it as world, so toEntities host never matches. The barman plugin and the OpenBao snapshot job on gcp-0 now use toCIDR 169.254.169.254/32 on TCP 80, the rule runlore and image-gallery run live. security/AGENTS.md rule 3 is split per cloud, and a test fails on any gcp-0 CNP reaching host:80.

* feat(gcp): gcp-0's ai-gateway umbrella, with the rate limit its budgets need

* feat(gcp): gcp-0's agent-platform umbrella

* docs(gcp): fix 6.6 loose ends after review

- rate-limit CNP comment now says the service exists on aws-0 and gcp-0,
  not aws-0 only, so the 18001 rule isn't trimmed as aws-0-specific
- point gcp-0/ai-gateway.yaml at aws-0-ai-gateway/README.md, since gcp-0
  has no README of its own for this umbrella
- rewrap two lines left overlong after 6.6's rewrap

* feat(ci): gate gcp-0's agent renders on GKE-shaped values

* feat(ci): assert gcp-0's gateway keys, CA key, rate limit and chart metadata egress

The GP-24 and GP-26 patches must apply, not just leave no AWS string: ai-gateway-api-keys is Password-generated with refreshPolicy CreatedOnce and no store, and openbao-ca reads openbao-priv-gcp-ca-chain. gcp-0's envoy-gateway must render the rate-limit KVStore while llm-gateway is an ai-gateway child. No gcp-0 chart CNP may reach the metadata server through toEntities host, the part test-gcp-metadata-server-cidr.py cannot build.

* fix(gcp): gcp-0's sandbox waits for the toleration policy; the README carries gcp-0's prerequisites

- agent-sandbox now dependsOn security-sandbox-policies too: its Kyverno
  gVisor toleration mutate is CREATE-only with background: false, so a run
  pod admitted before the policy exists never gets the Cilium toleration
  and stays Pending forever (aws-0 has no equivalent hazard: Karpenter
  ignores startupTaints when simulating scheduling)
- gcp-0-agent-platform/README.md Resume section now ports aws-0's
  prerequisites with gcp-0 paths: openbao/management + gke/configure,
  and the owner's branch ruleset + github-app/zai/factory-app keys on
  gcp-0's agents mount
- restored the sibling-of-clusters/gcp-0/ guard comment in
  clusters/gcp-0/agent-platform.yaml, matching ai-gateway.yaml and aws-0
- clusters/gcp-0/ai-gateway.yaml's README pointer now sends operators to
  the Teardown section only, substituting gcp-0 paths -- the aws-0 README's
  "what it reads" section is an AWS Secrets Manager bootstrap that does
  not apply on gcp-0
- reworded two comments that named runtimeclass-gvisor literally, which
  test-gcp-agents-pool.sh (added in 6.8) now greps clusters/gcp-0* for and
  fails on

* fix(ci): the cloud-shape gate fails when what it checks disappears

check_bundle names the eight overlays it expects and requires a GKE issuer in each run-token overlay; check_umbrellas fails on a missing or childless umbrella; check_ratelimit keys on the child's path, not its name. A host rule on port 0 or a range spanning 80 now counts as reaching the metadata server, here and in test-gcp-metadata-server-cidr.py, which shares the definition. The website's gate lists name five gates.

* fix(gcp): the gVisor mutate only sees gVisor pods, and the agent pool has room to scale

A webhook matchCondition keeps every non-gVisor Pod create away from Kyverno, so a Kyverno outage no longer blocks the cluster. The cluster autoscaler ceiling rises to 48 vCPU / 192 GiB so the fixed pools and the L4 allowance fit beside NAP, and a test sums them. The Karpenter pool alert is documented as aws-0-only, and the cloud-shape umbrella check no longer passes silently when the vars ConfigMap is renamed.

* feat(observability): the agent trace collector, metadata only

An OpenTelemetry Collector (otelcol-k8s 0.160.0, chart 0.173.1) that takes the run id from the sending pod's label only, drops spans no run sent, and keeps an allowlist of metadata keys. transform/cap also bounds event names and scope strings to 256 characters. A ReferenceGrant lets agent-router's EnvoyProxy in agent-system reference the collector Service.

* feat(observability): admit the factory's task spans on the platform port

* fix(observability): links and tracestate never leave the collector, and the router pipeline is capped

transform/cap drops span links and tracestate, which no attribute processor sees, and cuts names on a UTF-8 boundary. traces/router now runs transform/cap too. The suite compares the collector CNP, RBAC, pipelines and caps whole, so an extra peer, entity, exporter or ClusterRole fails it.

* fix(observability): span links are actually cleared

On contrib 0.160 'set' drops a nil value unless the alpha gate ottl.set.allowNil is on, so set(span.links, nil) was a silent no-op, and an empty list literal fails the links setter's type check. The collector now runs with --feature-gates=ottl.set.allowNil. The suite asserts the gate, and ties releaseName and the VMServiceScrape selector to the CNP's instance label.

* test(observability): run the trace collector's filter against a content fixture

Replays the HelmRelease's own agents pipeline and extraArgs in the digest-pinned otelcol-k8s, with k8s_attributes stubbed and debug plus file exporters. Content, spoofed run ids, unattributed spans, links and tracestate must not come out, and the names and allowlisted values are capped on a UTF-8 boundary. A control run without the allowNil gate and the tracestate statement must leak both, so those checks can fail.

* docs(adr): 0044 room session protocol

* feat(rooms): vendor the Room CRD and add it to the schema catalog

* feat(agent-router): trace every request to the agent trace collector

* feat(observability): kube-state-metrics series for AgentRun state

* feat(rooms): the log's SQLInstance with generated credentials and its CNPG policy

* feat(observability): the run's tier on agentrun_info

* test(observability): the router's trace egress and sampling are pinned

* feat(rooms): room log alerts

* feat(observability): the Agent run dashboard

* feat(observability): the run's tier, and a step line's trace link, on the run page

* test(observability): the AgentRun series read the right fields

* chore(crossplane): pin CC-S2's pre-release (SQLInstance credentials, room bridge)

* feat(observability): the Agent fleet dashboard

* docs(adr): 0044 fans out with Postgres LISTEN/NOTIFY, not Valkey (Ruling AJ)

* feat(observability): tier chosen vs tokens and steps, on the fleet page

* fix(rooms): migrate the room log from agent-platform main

* fix(rooms): room alerts follow the running-run gauge and the split stub reasons

* docs(agents): ADR-0051 and the umbrella README rows for the per-run view

* feat(ops): task agent:run prints the run's dashboard link on stderr

* feat(agent-harness): root the run's trace and carry its id in the step log

* fix(agent-harness): SIGTERM revokes within the grace, and main() is tested with tracing on

* fix(agent-harness): the factory always samples, and the traced test cleans up

* docs(rooms): ADR-0044 states its reasons, not ledger labels

* feat(agent-run): --room joins a run to a room on the room's branch

* docs(agents): ADR-0051 and the comments state their reasons, not ledger labels

Ledger-only labels meant nothing outside the plan's progress file; each is replaced by its reason. ADR-0051 now records why the harness always samples (lmnr's span context carries no sampled flag) and that SP3's factory spans in traces/router must stay metadata-only. The AI Gateway extproc comment no longer advises exporting prompts straight to VictoriaTraces, and the EnvoyProxy comment admits Envoy's default tags come from the request.

* fix(observability): the step log's trace link reads the run's own id, and scopes are proven redacted

extract_regexp takes the first match, and the harness appends the trace id after agent-written text, so a run could point its step-log links at another trace. The regex is now anchored on the line's last field. The replay fixture gains a scope attribute: redaction walks scopes, but nothing proved it; a scratch run allowlisting that key goes red.

* chore(crossplane): pin CC-O1's pre-release carrying harness v0.1.2 on aws and gcp

* chore(crossplane): pin the pr33 pre-release carrying the harness v0.1.2 and room-bridge pins

* test(agent-run): a room id is exactly 8 characters of [a-z2-7]

* feat(rooms): room-broker App, RBAC, policies, retention and umbrella children

The broker (room-broker v0.0.1-pr5.f3ac98ce, by digest) as an App claim with
TLS on :8443 from the openbao ClusterIssuer, its config, least-privilege RBAC,
default-deny CNPs for the broker and the retention job, the daily retention
CronJob and a VMServiceScrape. The run bridges' CA is copied into agents as
room-broker-ca, the one ExternalSecret agents-no-secret-import now admits.

Umbrella children on both clouds, each on a per-cloud render root; gcp-0
patches room-broker-ca to Secret Manager's entry, and assert-cloud-shape.py
and the metadata-server test now cover the room-broker overlay.

* fix(rooms): room-broker-ca exception admits allowlisted fields only

agents-no-secret-import let a room-broker-ca ExternalSecret override its store
per item (data[].sourceRef.storeRef) or write a non-Secret (target.manifest),
and accepted an AWS key with no property: enough to pull a whole OpenBao KV
entry into agents. The exception now allowlists the fields of spec, target,
the one data item and its remoteRef, requires property ca on the AWS key and
none on the GCP one, and refuses metadataPolicy Fetch.

* chore(rooms): record the room CRD's source as AP-1's pinned commit

The CRD at agent-platform f3ac98c (agent-platform#5, the room-broker pin) is
byte-identical to the one vendored from 363626a; only the source line moves.

* docs(adr): 0049 room client and human auth

* feat(zitadel): agent groups, rooms-proxy with JWT tokens mirrored to OpenBao, --grant

* feat(rooms): oauth2-proxy in front of the room UI

The rooms-proxy payload now carries the project id: oauth2-proxy's audience scope and the broker's aud check need it on both clouds, and aws-0's vars have no zitadel_project_id. gcp-0's overlay opens ZITADEL egress with toEntities all (the Gateway hairpin, ruling AU).

* fix(zitadel): a refused grant fails, grants are paged, token-type drift is repaired

Review of 2.8: grant_role runs under || in cmd_sync, so its writes now fail explicitly; both v1 searches page past 200 and fail loudly on a listing that never ends; an existing app with the wrong token type is repaired like a stale redirect. ADR-0049 states why the secret goes through the mirror and that only a hosting --mirror-openbao sync produces it.

* feat(rooms): rooms.<private domain> on the tailnet gateway

* feat(rooms): two broker replicas and the human listener

No Valkey (ruling AT): the broker fans out with Postgres LISTEN/NOTIFY, so the brief's KVStore, its password and the broker's Valkey env and egress are left out. Pins the AP-2 broker pre-release, whose config requires the human block; the client and project ids are mounted from room-broker-oidc and read at use (ruling AS-a). Adds RoomRejectedActionsSpike and an absent() alert for a human issuer never fetched.

* fix(rooms): oauth2-proxy refreshes sessions hourly, no basic auth, explicit broker grace

cookie-refresh 1h is oauth2-proxy's only session-expiry check; without it humans get 401s from the broker after ZITADEL's token lifetime until the 168 h cookie expires. rooms-proxy is created with the refresh_token grant, so offline_access joins the scope and the refresh is silent (a test pins the grant). pass-basic-auth false, as headlamp's proxy. terminationGracePeriodSeconds 30 pinned on the broker.

* chore(crossplane): pin both clouds to CC-S2 v0.7.2-pr33.00e6520

CC-S2 a96da7ea carries room-bridge v0.0.1-pr6.e335dd32@sha256:814e637c, the AP-2 bridge that matches the broker pin. The core package follows through dependsOn.

* feat(agent-factory): ADR-0048, the factory under the agent-platform umbrella

The factory's signed chart source, HelmRelease (Task CRD created on install,
replaced on upgrade), values with the strictly parsed config, the App key
through agents-secrets, a default-deny CNP and its scrape, as an umbrella
child on both aws-0 and gcp-0 with a GP-14 render root per cloud. The broker's
allowlist admits the factory's ServiceAccount, and the cloud-shape gate now
judges the gcp-0 factory overlay.

* feat(agent-factory): send task spans to the trace collector

* fix(agent-factory): pin the chart by digest, pair it with the image, state ADR-0048's reasons

- OCIRepository carries ref.digest, which Flux resolves before the tag
- render-bundle renders a digest-pinned chartRef as <url>@<digest>
- test-agent-factory-pins.py fails when chart and image pins disagree;
  Renovate groups the two and now tracks the image tag in the values
- VMRule runbook reads every replica, since only the leader polls
- ADR-0048 gives reasons in place of programme ledger labels

* chore(agent-factory): pin the factory to agent-platform#7 at 569dc1c, by digest

Chart and image are the signed pre-release of that head.

* chore(agent-factory): pin the phase 2 pre-releases

agent-platform#10 at 800fcd8: the factory image and chart, and the room-broker image, a superset of AP-4's, for the app and the retention CronJob. The vendored Room CRD moves with the broker pin, as its comment requires; it gains dataClass immutability.

* chore(agent-factory): pin the phase 3 pre-releases; pair by default

agent-platform#12 at f2e51f9: the factory image and chart, and the room-broker image for the app and the retention CronJob (FA-3 leaves the broker's code and the Room CRD unchanged, so crd-rooms.yaml stays). defaults.template becomes pair: every task gets a reviewer until phase 4's triage picks templates.

* docs(adr): 0048 separates the decided design from what phase 1 builds

Kueue is not deployed and the factory enforces only the run budget; no gateway per-run ceiling exists yet. Keep the decision, add a decided/built table, and name SP3 phase 4 for Kueue. The trace collector's CNP comment no longer calls the factory rule inert.

* chore(agent-factory): re-pin the phase 3 pre-releases to agent-platform#12 at ddb06e0

* feat(agents): Kueue under the agent-platform umbrella (SP3 FR-4)

Kueue admits every factory sandbox pod: pod-only integration scoped to the
agents namespace, two non-borrowing ClusterQueues (factory, interactive) and
their LocalQueues, default-deny CNP with webhook/metrics/health ingress and
DNS/API egress, TLS metrics scraped by vmagent through a dedicated
metrics-reader ClusterRole.

Per ruling SB the target cloud is gcp-0: the controller child points at an
infrastructure/gcp-0/kueue render root (GP-14) and substitutes
gke-gcp-0-vars; the queue child is cloud-neutral and stays on the base.
agent-factory now dependsOn kueue-queues so submitted runs find their
LocalQueue. The Kueue chart joins the factory's no-automerge Renovate rule.

* feat(agent-factory): triage config, classes, shadow budgets, classifier egress

* feat(agent-policies): the factory is the only AgentRun creator; its patches are annotations only

* fix(agent-policies): bound the stop-readonly rule by resourceNames; DELETE carries no object

* feat(agent-factory): POST /v1/runs on the tailnet, client ids, 429 lookup egress

* feat(agent-run): task agent:run asks the factory's API; the token names the principal

* test(zitadel-oidc-clients): load stored_client_id, mock restart_rotated_consumers (#2149)

* feat(rooms): the broker requests runs from the factory

* docs(adr): ADR-0045 merge policy gate

* feat(ci): gate-path coverage (SC-12) and pull_request secrets lint (T8)

* fix(ci): the workflow-secrets lint catches permissions: write-all

* feat(merge-gate): namespace, OpenBao policy and role, namespaced store

* feat(policy-bot): the merge gate in merge-gate, its public hook and tailnet UI

* feat(merge-gate): .policy.yml: docs-links and revert live, gate paths unmergeable

* feat(merge-gate): agent-merge-gate ruleset source and its idempotent applier

* feat(merge-gate): policy-bot and its store under the agent-platform umbrella

* feat(merge-gate): approval-free agent rules require CI, secret scan included

* feat(merge-gate): the agent-merge ruleset and the agent-branches split, for after the wave

* feat(agent-factory): the merge gate in shadow, the merger key, link-rot

* feat(agent-factory): the factory's alerts and dashboard, inside the umbrella

* docs(runbooks): materialize agent-factory runbooks 01-08 and their README from integration/agent-factory

* docs(runbooks): 09 App key compromise, all four Apps

* feat(agent-factory): pin the factory to agent-platform#16; the control issue

* feat(runlore): findings reach the agent factory's intake, inside the umbrella

* fix(runlore): the intake env survives Flux's values merge order

* feat(agent-factory): scripted developer walkthrough and its journey renderer

* feat(agents): the aws-0 umbrella carries Kueue too (aws-0 is the live target again)

* fix(rooms): the retention CronJob joins the broker's pr14 pin (one image, one version)

* feat(agent-factory): resume runs lost to a spot reclaim or an eviction (#2222)

* docs(agents): design agent runs that survive a spot or preemptible reclaim

An external review found runs on reclaimable capacity lose their work
and are never resumed. The design budgets the pod's shutdown at 15 s
(GKE preemptible VMs cannot extend it): pause, checkpoint commit and
push, the bridge's final read, then stop and revoke. The composition
reads the pod to record Disrupted, PodLost or PodFailed, and the
factory resumes Disrupted and PodLost runs at most twice within the
task's token cap. An early warning on aws-0 waits on a fault-injection
test. The research records the facts at pinned commits.

* docs(agents): plan agent runs that survive a reclaim

Thirteen tasks across agent-platform, cloud-native-ref and
crossplane-configuration: the bridge's final read, agent-run's
15 s shutdown sequence, the composition reading the pod for
Disrupted/PodLost/PodFailed, GKE Spot's 120 s window, the factory's
automatic resume, docs, runbook 09 and live verification. The spec is
aligned with what the code showed: no XRD enum, a deleted pod reads
PodLost, Crossplane needs cluster-wide pod reads for required resources,
interrupt rather than pause, an in-process revoke, the trailer from the
commit-msg hook.

* docs(agents): record the owner's decisions on the disruption plan

The wider Crossplane pod read is accepted, and merging the factory resume
work may bring factory phases 2-9 to gcp-0. The two owner actions the
plan depends on are listed.

* docs(agents): order the agent factory's remaining work

One table from the open reviews to the merge wave, naming the plan that
holds each step: the disruption plan's runtime half rides the same re-pin
and rebuild as the F10-F18 fixes, the agentgateway migration (with prompt
caching) follows the live re-verification, then SP2 and SP3.

* docs(agents): link the kagent evaluation now that it is on main

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(agent-factory): pin automatic resume (agent-platform#25), resume.maxPerTask, resumes on the dashboards

* docs(runbooks): runbook 10, disruption live checks; the factory resumes a lost run

* docs(agents): the disruption runbook is 10; 09 is the App key-compromise runbook

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(runbooks): runbook 10 waits on resumes and new runs, reads resumes_exhausted; resumes are per task

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(agents): pin the factory to agent-platform#25 at 642a5b12; enforce the task cap

642a5b12 carries the resume review fixes and #30's whole-run task cap.
enforceTask: true matches integration (owner, 2026-10-07).

* test(zitadel): the rooms-proxy suite loads previous_client_id after #2208

cmd_sync reads the previous client through previous_client_id since #2208;
the suite did not load it (exit 127). Same lines as integration/agent-factory.

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(agents): pin the factory to agent-platform#25 at ef31783e

ef31783e quotes a review queued while a lost run was live in the resumed run's brief,
and escalates a lost reviewer past the enforced task cap.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* feat(eks): manage the agent-run FIS role and experiment template as code (#2233)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* feat(eks): manage the agent-run FIS role and experiment template as code

Runbook 10 Step 7 (the aws-0 Spot-interruption test, disruption plan Task 13)
told the owner to create an IAM role and an FIS experiment template with aws CLI
commands. Both now live in opentofu/aws/eks/init, the stack that owns aws-0's
node and Karpenter IAM, so an aws-0 rebuild creates them.

The role trusts fis.amazonaws.com only for this account's experiments in the
region (aws:SourceAccount, aws:SourceArn) and may interrupt only instances
tagged agents.ogenki.io/fis-target=true. The runbook now tags the run's node,
starts the template, and untags the node in cleanup.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* fix(ci): restore the programme's merge-gate and pre-release XRD CRD steps (#2235)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release external-secrets to v2.12.0 (#2232)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(ci): restore the programme's merge-gate and pre-release XRD CRD steps

#2233 resolved its main merge to main's ci.yaml, which never had them: the
Merge-gate invariants (SP3 SC-12) and the pre-release XRD CRD fetch (SP2 P40)
left feat/factory-runlore. Restored byte-identical to c7fb11c8 and integration;
main's newer checkov and trufflehog pins stay.

* test(ci): a main merge cannot drop the programme's CI steps unnoticed

Nothing read ci.yaml for them, so their loss in #2233 passed every local gate.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* fix(agent-factory): drift detection corrects a kubectl scale of the factory (#2237)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release external-secrets to v2.12.0 (#2232)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(agent-factory): drift detection corrects a kubectl scale of the factory

A kill-switch drill scaled the factory to 0; on resume Helm's 3-way merge re-applied
only what the chart changed, so it stayed at 0 for 8.5 h. Drift detection restores the
chart's 2 replicas on the next reconcile. No ignore rules: nothing else writes the
chart's objects. The plan's drill now suspends the HelmRelease, which drift detection
would otherwise undo mid-drill, and checks 2/2 after the restore: rollout status alone
passes at 0 of 0.

---------

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <334403746+ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(runbooks): runbook 10's aws-0 results; the Karpenter reason label; capture logs live (#2243)

* chore(deps): update bridgecrewio/checkov-action action to v12.3129.0 (#2202)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.0 (#2204)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release aws-load-balancer-controller to v3.6.0 (#2216)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* fix(zitadel): count a restored OpenBao mirror's client as the previous one (#2208)

On a rebuild whose OpenBao was restored while the managed store was not, a
mirrored consumer's previous client id lives only in OpenBao: that is what its
ExternalSecret reads. The create path read the store alone, found nothing, and
treated the new client as a first bootstrap, so the restart step reported
"nothing to restart" while the env reader kept the dead client and ZITADEL
answered App.NotFound (aws-0, 2026-10-06: rooms-oauth2-proxy).

previous_client_id reads the store first and, on a mirrored key, falls back to
OpenBao's copy, so that case counts as a rotation and its readers restart.

* fix(cilium): keep agent run pods out of the unmanaged-pod restarts (#2225)

cilium-operator's unmanaged-pod watcher ran with an empty --pod-restart-selector,
so it deleted every pod without a Cilium endpoint. A finished agent run pod
loses its endpoint while its native sidecars drain: the watcher deleted it,
agent-sandbox created a replacement, and the AgentRun ended Failed/PodLost
although the run had succeeded (aws-0, 2026-10-06). Every unmanaged pod except
agent runs is still restarted, as the bootstrap needs (EBS CSI, CoreDNS).

* docs(coding-clients): add Authorization header to /v1/models smoke test (#2206)

The gateway enforces apiKeyAuth (SecurityPolicy
infrastructure/base/envoy-ai-gateway/security-policy.yaml), so the
smoke-test curl returned 401 without a Bearer token. The other curls on
the page already send it; this one was the odd one out.

Agent-Run: uzdtmemb
Agent-Task: brhe2ia4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: update Polaris version in validation.md Requirements to 10.2.5 (#2218)

mise.toml pins github:FairwindsOps/polaris at 10.2.5, so the
Requirements section should name that version, not 8.5.0.

Fixes #2213

Agent-Run: ulvcv5hh
Agent-Task: 5ato23w4

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* docs: sre-agent.md runlore GitRepository pin is v0.16.2 (#2219)

The Source row in the deployment table said the GitRepository is pinned
to tag v0.16.1, but flux/sources/gitrepo-runlore.yaml pins v0.16.2.


Agent-Run: vbrwt3fo
Agent-Task: yaq5gmdm

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: correct crossplane-configuration tag in app-wizard.md to v0.7.1 (#2220)

Agent-Run: qwruq4bg
Agent-Task: lnu2pdwy

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: update crossplane-configuration-aws pin to v0.7.1 in developer-platform page (#2221)

The developer-platform _index.md still showed v0.4.6 while
infrastructure/base/crossplane/configuration-aws/configuration-packages.yaml
pins v0.7.1.

Fixes #2211


Agent-Run: py5thsfa
Agent-Task: w74kdnyf

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: metrics.md victoria-metrics-k8s-stack chart pin is 0.95.0 (#2227)

The chart version in website/content/docs/platform/observability/metrics.md
said 0.91.2 while both HelmReleases in observability/base/victoria-metrics-k8s-stack/
pin 0.95.0.

Fixes #2209


Agent-Run: sg433pli
Agent-Task: dnx5nu5v

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs: commands.md describe verify-doc-paths.sh as backticked path existence check (#2228)

Fixes #2210


Agent-Run: zjv5r62l
Agent-Task: b7kelb63

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* docs(scripts): name the three teardown helpers in scripts/README.md (#2141)

Closes #2140

Agent-Run: 4iv2rpdq

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>

* chore(deps): update bridgecrewio/checkov-action action to v12.3130.0 (#2229)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update dependency external-secrets to v2.12.0 (#2230)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(security): correct kyverno controller chart version to 3.9.1 (#2224)

* docs(security): correct kyverno controller chart version to 3.9.1

The policies.md page listed the kyverno controller chart as 3.9.0, but
security/base/kyverno/helmrelease-controller.yaml pins 3.9.1.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: imsmupzi
Agent-Task: msgctzzw

* docs(security): correct kyverno-policies chart version to 3.9.1

Per review feedback on #2224: security/base/kyverno/helmrelease-policies.yaml
pins version 3.9.1, so the doc on line 18 should say 3.9.1, not 3.9.0.

Fixes #2223

Co-authored-by: openhands <openhands@all-hands.dev>
Agent-Run: ptvuqyg3
Agent-Task: msgctzzw

---------

Co-authored-by: ogenki-agents[bot] <ogenki-agents[bot]@users.noreply.github.com>
Co-authored-by: openhands <openhands@all-hands.dev>

* chore(deps): update trufflesecurity/trufflehog action to v3.98.1 (#2231)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* chore(deps): update helm release external-secrets to v2.12.0 (#2232)

Co-authored-by: renovate[bot] <29139614+renovate[bot]@users.noreply.github.com>

* docs(runbooks): runbook 10's aws-0 results; the Karpenter reason label; capture logs live…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant