Skip to content

Add an agentcore_gateway model backend: AgentCore Gateway in front of LiteLLM, per-session credentials - #40

Merged
odinwang merged 4 commits into
aws-samples:mainfrom
zzkamzn:feat/agentcore-gateway-model-backend
Oct 1, 2026
Merged

odinwang merged 4 commits into
aws-samples:mainfrom
zzkamzn:feat/agentcore-gateway-model-backend

Conversation

@zzkamzn

@zzkamzn zzkamzn commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Update (maintainer, 2026-10-01): gateway now fronts LiteLLM

Two commits pushed on top: a merge of current main, and 292d41b.

  • Target is an inference provider in front of LiteLLM, not the bedrock-mantle connector. The LiteLLM virtual key lives in an AgentCore Identity API-key credential provider and the gateway injects it, so LiteLLM keeps its multi-provider routing and cost accounting and llm-edge is not needed. Operations declare /v1/messages, /v1/chat/completions and /v1/responses. The gateway does no protocol translation (it matches path + model, swaps the key, relays body and SSE verbatim), so which model works on which path is LiteLLM's call. All four test models (claude-opus-5-5, claude-sonnet-5-5, gpt-6-astra, kimi-k3) answered on all three paths.
  • No interceptor. drop_params on the LiteLLM model entries strips Claude Code's context_management. An interceptor adds a synchronous Lambda call per model request and caps bodies at about 4.5 MB (Lambda's 6 MB invoke payload after base64): 3.1 MB passed, 4.7 MB was refused, 26 MB went straight through without it.
  • Private path to LiteLLM. privateEndpoint.managedVpcResource (top-level, next to targetConfiguration) puts VPC Lattice resource-gateway ENIs in the LiteLLM VPC and connects to an internal ALB via routingDomain, keeping the public cert name as SNI. ALB access logs show every request arriving from the two Lattice ENIs.
  • Limit found: Lattice's resource-mode TCP idle timeout is 350 s (VPC Lattice quotas page). A non-streaming call taking 312 s returned; one taking 368 s never did, and the gateway gave the caller no error. Streaming is unaffected (634 s stream completed). Clients on this path must stream; Claude Code and the SDK kernel already do.
  • Smaller items: credentialPrefix is "Bearer" with no trailing space (the gateway adds one); the execution role also needs GetWorkloadAccessToken; the gateway rejects any path outside those three, including count_tokens, at target-creation time.
  • Re-verified end to end in ap-northeast-1 through the private endpoint: chain, multi-turn tool task, multi-model, failfast, scoping, revoke (8.6 s).

Runtime → gateway traffic can stay inside the VPC too, via the gateway interface endpoint in #45.


What

A third model backend, agentcore_gateway, alongside bedrock and litellm. It reaches the same goal as llm-edge — no upstream model credential ever lands in a container whose user is root — with one fewer service to run: the provider credential lives in the AgentCore Gateway's token vault, so there is nothing on the platform side to hold either.

Two commits: code, then docs.

How a session is authorized

The credential a kernel receives is an STS session tagged with its runtime session id:

  • a session policy narrows it to InvokeGateway on the one gateway the backend points at, so a credential read out of kernel memory cannot reach any other gateway in the account;
  • the session tag is what revoke() conditions a Deny on — that is how one session is stopped without disturbing another. The Deny set is rebuilt from an LLMREVOKE partition and pruned once entries outlive the credential they revoke, keeping the inline policy inside its 10 KB limit;
  • STS cannot extend a session, so rotate() re-mints from the routing already recorded on the token item, preserving the "config edits apply at next warmup" rule the edge path has.

Claude Code cannot sign SigV4, which is exactly why the signing belongs in the kernel shim rather than the CLI subprocess. Both shims now sign per request — botocore in the SDK kernel, ~40 lines of node:crypto in the contract server so it keeps its single dependency. Only a minimal header set is signed and the caller's headers are attached afterwards, so Claude Code's anthropic-* and x-stainless-* headers cannot invalidate a signature; its own x-amz-* headers are dropped rather than folded in.

What this backend deliberately does not claim

IAM conditions cannot see a request body, so the per-session model allowlist is not expressible as an IAM condition. The shim's check is a fail-fast optimisation and says so in the code — the session's user is root in that microVM and can bypass the shim. Real enforcement needs a gateway request interceptor or one gateway per model. The allowlist is recorded on the token item either way so the decision stays auditable.

Verification

scripts/e2e_agentcore_gateway.py provisions a live gateway and caller role, then drives the real Claude Code CLI through the real kernel shim and tears everything down. Nothing in the chain is stubbed; the only simulation is that the shim runs in the test process rather than inside AgentCore Runtime, which is the same place it runs in production.

It asserts, and currently passes:

single turn Claude Code completes it through the chain, and the shim observed the request (so it cannot be satisfied by some ambient credential)
multi-turn task with tools write a file → run it → report a line: 3 model calls, tool_result blocks accumulating 0→1→2, request body growing 112 KB → 114 KB, all streaming
scoping the same credential is refused (403) on a second gateway
revocation session A stops working ~9 s after revoke(); session B keeps working
fail-fast the shim refuses a model the session was not routed to (403)

Claude Code sent exactly one path throughout: POST /v1/messages?beta=true. It never calls /v1/messages/count_tokens, which AgentCore Gateway does not accept inbound — so that gap is not a blocker.

Things operators need to know (all documented)

  • ⚠️ The kernel roles already hold InvokeGateway on gateway/* so they can reach MCP tool gateways, and that wildcard covers an inference gateway too. A root user in the microVM could call it on the role's own identity, with no session tag to revoke. Enabling this backend requires an explicit Deny for the inference gateway's ARN — without it the per-session credential is decorative. Written up as a prerequisite in deployment.md and permissions.md.
  • ⚠️ A gateway interceptor's escaping exception is relayed to the caller with its full stack trace, regardless of exceptionLevel — and the caller is the tenant's microVM. Interceptors must catch everything themselves. The reassuring half: interceptor failure is fail-closed (verified).
  • ⚠️ allowedRequestHeaders is not optional. Left unset the gateway relays the caller's own x-amz-security-token; once set, a connector target stops relaying content-type unless it is listed.

The docs commit also corrects two claims security-explainer.zh.md §9.4/§9.5 had been overstating, independent of this backend: the platform's quota counts invocations rather than tokens (so there is no per-user token ceiling on either gateway mode), and credential reuse across sessions is not prevented by either mode, because nothing proves a caller is the session it claims to be.

Not included

  • Terraform for the caller role. The role, its MaxSessionDuration, the gateway and the kernel-role Deny are described in the docs but not yet a module. Deliberate: I could not exercise a Terraform path end to end here.
  • Nothing changes for existing deployments. agentcore_gateway ships disabled, and with the caller role ARN unset the backend refuses to route rather than falling back to a shared credential.

Note on stacking

This branch is rebased directly onto main and is independent of #39, which touches llm_credentials_service.py too. Whichever merges second will want a look at the ownership ConditionExpression in mint() / mint_agentcore() — the two express the same rule in their own paths.

🤖 Generated with Claude Code

zzkamzn and others added 4 commits September 18, 2026 03:02
…tials

Adds an agentcore_gateway model backend alongside bedrock and litellm. It
reaches the same goal as llm-edge — no upstream model credential ever lands in
a container whose user is root — with one fewer service to run: the provider
credential lives in the gateway's token vault, so there is nothing on the
platform side to hold either.

A session's credential is an STS session tagged with its runtime session id,
minted by the backend and delivered in the existing llm_credentials block:

- a session policy narrows it to InvokeGateway on the one gateway the backend
  points at, so a credential read out of kernel memory cannot reach any other
  gateway in the account;
- the session tag is what revoke() conditions a Deny on, which is how one
  session is stopped without disturbing another. The Deny set is rebuilt from a
  LLMREVOKE partition and pruned once entries outlive the credential they
  revoke, keeping the inline policy inside its 10 KB limit;
- STS cannot extend a session, so rotate() re-mints from the routing already
  recorded on the token item, preserving the "config edits apply at next
  warmup" rule the edge path already has.

Claude Code cannot sign SigV4, which is exactly why the signing belongs in the
kernel shim rather than the CLI subprocess: both shims now sign per request
(botocore in the SDK kernel, ~40 lines of node:crypto in the contract server,
which keeps its single dependency). Only a minimal header set is signed and the
caller's headers are attached afterwards, so Claude Code's anthropic-* and
x-stainless-* headers cannot invalidate a signature; its own x-amz-* headers
are dropped rather than folded in. The SigV4 path buffers the request body
because it needs the payload hash, and is capped; the edge path still streams
it through unread.

What this backend deliberately does not claim: IAM conditions cannot see a
request body, so the per-session model allowlist is not expressible as an IAM
condition. The shim's check is a fail-fast optimisation and says so — the
session's user is root in that microVM and can bypass it. Real enforcement
needs a gateway request interceptor or one gateway per model; the allowlist is
recorded on the token item either way so the decision stays auditable.

scripts/e2e_agentcore_gateway.py provisions a live gateway and drives the real
Claude Code CLI through the real shim, asserting the chain end to end: a
multi-turn tool-using task completes over three model calls, the credential is
refused on a second gateway, and revoking one session leaves another working.
Terraform for the caller role is not included yet.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Folds the third model backend into the existing model-access sections rather
than adding a new page: architecture.md's backend table, security-explainer
§9 (which gains the second call chain), deployment.md §2 as Option A2,
permissions.md's principal table, the user guide's Governance card, and the
EXTENDING.md configuration index.

Three things the evaluation turned up that operators need in writing:

- The kernel roles already hold InvokeGateway on gateway/* so they can reach
  MCP tool gateways, and that wildcard covers an inference gateway too — a root
  user in the microVM could call it on the role's own identity, with no session
  tag to revoke. Enabling this backend requires an explicit Deny for the
  inference gateway's ARN; without it the per-session credential is decorative.
  Documented as a prerequisite in deployment.md and permissions.md.
- A gateway interceptor's escaping exception is relayed to the caller with its
  full stack trace, regardless of exceptionLevel — and the caller is the
  tenant's microVM. Interceptors must catch everything themselves. The
  reassuring half: interceptor failure is fail-closed.
- allowedRequestHeaders is not optional. Left unset the gateway relays the
  caller's own x-amz-security-token; once set, a connector target stops
  relaying content-type unless it is listed.

Also corrects two claims §9.4 and §9.5 had been overstating, independent of
this backend: the platform's quota counts invocations rather than tokens, so
there is no per-user token ceiling on either gateway mode; and credential reuse
across sessions is not prevented by either, because nothing proves a caller is
the session it claims to be.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Repoint the gateway target from the bedrock-mantle connector to an
inference *provider* target in front of LiteLLM, which is what this
backend is for: the LiteLLM virtual key lives in an AgentCore Identity
API-key credential provider and the gateway injects it, so neither the
platform nor the container holds it and llm-edge is not needed.

- Target: endpoint = LiteLLM, operations /v1/messages,
  /v1/chat/completions and /v1/responses with the models the virtual key
  permits. The gateway does no protocol translation (it matches path and
  model, swaps the key, relays body and SSE verbatim), so which model
  works on which path is LiteLLM's call; all four test models answered on
  all three. credentialPrefix "Bearer", no trailing space (the gateway
  adds it). allowedRequestHeaders keeps anthropic-beta, which LiteLLM
  accepts.
- Gateway execution role: GetWorkloadAccessToken as well as
  GetResourceApiKey (the documented policy alone fails with "Failed to get
  workload identity token"), plus the provider's secret.
- Private path: a top-level privateEndpoint.managedVpcResource places VPC
  Lattice resource-gateway ENIs in the LiteLLM VPC and connects to an
  internal ALB (routingDomain) while keeping the public cert name as SNI.
  Driven by LITELLM_ROUTING_DOMAIN / LITELLM_VPC_ID / LITELLM_SUBNET_IDS /
  LITELLM_LATTICE_SG_IDS.
- No REQUEST interceptor: drop_params on the LiteLLM model entries strips
  Claude Code's first-party-only fields (context_management). An
  interceptor adds a Lambda hop per call and caps bodies at ~4.5 MB.
- e2e: [multimodel] smoke over every model on the key; SSO reserved roles
  trusted via the account root; teardown waits for private-endpoint
  targets; LITELLM_ENDPOINT / LITELLM_CREDENTIAL_PROVIDER_ARN required.
- deployment.md Option A2 rewritten for LiteLLM, including the 350-second
  Lattice idle timeout that makes long non-streaming calls hang.

Verified end to end in ap-northeast-1 through the private endpoint: chain,
multi-turn tool task, multi-model, failfast, scoping, revoke (8.6 s).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@odinwang odinwang changed the title Add an agentcore_gateway model backend: per-session credentials, no key on the platform side Add an agentcore_gateway model backend: AgentCore Gateway in front of LiteLLM, per-session credentials Oct 1, 2026
@odinwang
odinwang merged commit 7689b93 into aws-samples:main Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants