Every route, field, status code and header below is taken from the code under
src/main/scala, with defaults from src/main/resources/application.conf.
- Base URL (docker compose):
http://localhost:8080. The port isSERVER_PORT(default 8080), and compose publishes it on 8080. - Bodies: request bodies are parsed as JSON whatever the
Content-Typesays, but sendContent-Type: application/json. JSON responses areapplication/json. - Fields: outside the demo dashboard, every response field shown below is
always present. An optional value that does not apply is
null(responses are printed with nulls kept). X-Request-Id: every response echoes the request'sX-Request-Id, or a generated UUID when the request had none. That includes the 401, 403 and 429 answers from authentication, the 400/422 body-decoding answers, and a 500.- Errors: every answer that is not a 2xx has a JSON body with
error(a stable code) andmessage(text for a person). See Errors. - Tenancy: every key a caller names (rate-limit key, idempotency key, quota user/agent/org, reservation ID) is scoped to the authenticated client (ADR-005). Two clients using the same key string never share state and cannot see each other's.
| Method | Path | Permission (name in the keys secret) | Purpose |
|---|---|---|---|
POST |
/v1/ratelimit/check |
RateLimitCheck (ratelimit_check) |
Consume tokens from a bucket |
GET |
/v1/ratelimit/status/{key} |
RateLimitStatus (ratelimit_status) |
Read a bucket without consuming |
POST |
/v1/idempotency/check |
IdempotencyCheck (idempotency_check) |
Claim a key, or learn its state |
POST |
/v1/idempotency/{key}/complete |
IdempotencyComplete (idempotency_complete) |
Store the response to replay |
POST |
/v1/idempotency/{key}/fail |
IdempotencyComplete (idempotency_complete) |
Release a pending key for retry |
POST |
/v1/quota/check |
QuotaCheck (quota_check) |
Reserve estimated LLM tokens |
POST |
/v1/quota/reconcile |
QuotaReconcile (quota_reconcile) |
Replace a reservation's estimate with actual usage |
GET |
/health |
none | Liveness |
GET |
/ready |
none | Readiness |
GET |
/metrics |
AdminMetrics (admin_metrics) |
Prometheus scrape |
| various | /dashboard, /dashboard/api/*, /v1/ratelimit/dashboard/stats |
none | Demo dashboard, only with DASHBOARD_ENABLED=true (below) |
The quota routes exist only with TOKEN_QUOTA_ENABLED=true (compose sets it;
the default is false).
The middleware reads the key from, in order:
Authorization: Bearer <key>orAuthorization: ApiKey <key>(scheme matched case-insensitively);X-Api-Key: <key>, used when there is noAuthorizationheader with one of those schemes.
/health, /ready and the dashboard routes are matched before
authentication and need no key. Every other request goes through it, including
requests for paths that do not exist.
| Case | Status | Body |
|---|---|---|
| No key | 401 |
{"error": "unauthorized", "message": "Missing API key in Authorization header"} |
| An unknown key | 401 |
{"error": "unauthorized", "message": "Invalid API key"} |
| Valid key, sending faster than the auth throttle allows | 429 + Retry-After |
{"error": "rate_limited", "message": "Rate limited. Retry after 42 seconds", "retryAfter": 42} |
| Any key, from a source over its unknown-key limit | 429 + Retry-After |
The same body |
| Valid key without the route's permission | 403 |
{"error": "forbidden", "message": "Insufficient permissions: QuotaCheck required"} |
| Valid key, path that no route matches (or wrong method) | 404 |
{"error": "not_found", "message": "No route for GET /v1/nope"} |
The permission check runs before the route touches any state. The 403 message
names the permission by its code name (QuotaCheck), not its secret name
(quota_check).
Auth throttle. Each key may authenticate security.authentication.rate-limit-per-minute
times per minute (AUTH_RATE_LIMIT_PER_MINUTE, default 1000; compose sets
10,000,000 for load tests). The count is per key ID, in memory on each
instance, over a one-minute window that starts at the key's first request.
Past it, the answer is 429, not 401: back off for Retry-After seconds, do not
fix credentials. It counts every authenticated request, including ones that
end in 403 or 404.
Unknown-key throttle. Each source may present
security.authentication.failed-attempts-per-minute (20) unknown keys per
minute. After that, every request from it that carries a key, valid or not,
answers 429 with Retry-After until its minute ends, and no key is looked up.
A missing key is not counted. The source is the connecting address, or, with
AUTH_TRUST_FORWARDED_FOR=true (Terraform sets it, since the ALB fronts every
task), the last X-Forwarded-For entry, which the ALB appends. The count is in
memory on each instance, so it slows a guesser at one address, not one spread
across many; that is the ALB's or WAF's job.
Each authenticated route needs one permission; the route table
names it and its name in the keys secret. Secret names are matched
case-insensitively, the forms without the underscore (ratelimitcheck) are
accepted too, and an unrecognized name is ignored.
The service needs one key source, or it refuses to start.
Secrets Manager (SECRETS_MANAGER_ENABLED=true; Terraform always sets it).
The secret is named <secret-prefix>/<environment>/<api-keys-secret-name>
(SECRETS_PREFIX, SECRETS_ENVIRONMENT, API_KEYS_SECRET_NAME; default
rate-limiter/dev/api-keys) and holds a JSON list. Every field is required:
[
{
"apiKey": "<random secret>",
"apiKeyId": "key_acme_001",
"clientName": "Acme",
"tier": "premium",
"permissions": ["ratelimit_check", "ratelimit_status", "idempotency_check",
"idempotency_complete", "quota_check", "quota_reconcile"],
"active": true
}
]tierisfree,basic,premiumorenterprise. An entry with"active": falseis skipped. An unknown tier or permission name, in any entry, is an error that names the entry'sapiKeyIdand the value.- The tenant is the entry's
apiKeyId. Keep it when you rotateapiKeyto keep the client's state. - Startup fails if the secret is missing, is not a list in this shape, has an unknown tier or permission, or has no active key.
- The keys are re-read about every 5 minutes (
cache-ttl), so adding or revoking a key takes effect within minutes, not at once. If a re-read fails, the current keys stay. If it finds an unknown tier or permission, the valid entries take effect (so revocations do), the invalid ones cannot authenticate, and the error is logged.
Built-in development keys (ALLOW_BUILT_IN_KEYS=true, set only by docker
compose and make run). They are public, so never use them outside local
development. When Secrets Manager is enabled it wins, even with this flag set.
| Key | Tier | Permissions |
|---|---|---|
test-api-key |
premium | the six standard permissions (all but AdminMetrics) |
free-api-key |
free | the six standard permissions |
admin-api-key |
enterprise | the six standard permissions and AdminMetrics |
Permission: RateLimitCheck. Consumes cost tokens from the caller's
bucket for key, or refuses without consuming.
{ "key": "user:12345", "cost": 1 }| Field | Type | Required | Description |
|---|---|---|---|
key |
string | yes | Bucket identifier, scoped to your client |
cost |
integer | no, default 1 |
Tokens to consume; must be at least 1 |
profile |
string | no | A named profile to apply instead of your tier's (see Profiles) |
endpoint |
string | no | A label copied onto the decision event; it does not affect the decision |
Allowed (200):
{ "allowed": true, "tokensRemaining": 999, "retryAfter": null, "limit": 1000,
"resetAt": "2026-09-27T12:00:00.010Z", "message": null, "error": null }Headers: X-RateLimit-Limit (the profile's capacity), X-RateLimit-Remaining,
X-RateLimit-Reset (resetAt as epoch seconds).
Refused (429): nothing was consumed. Header: Retry-After (seconds, same
as retryAfter). No X-RateLimit-* headers.
{ "allowed": false, "tokensRemaining": null, "retryAfter": 1, "limit": 1000,
"resetAt": "2026-09-27T12:00:10.000Z", "message": "Rate limit exceeded",
"error": "rate_limit_exceeded" }| Field | Type | Description |
|---|---|---|
allowed |
boolean | Whether the tokens were consumed |
tokensRemaining |
integer or null | Tokens left after this check; null when refused |
retryAfter |
integer or null | Seconds until cost tokens should be available; null when allowed |
limit |
integer | Capacity of the profile applied |
resetAt |
string (ISO-8601) | When the bucket would be full again, per the configured algorithm |
message |
string or null | "Rate limit exceeded" when refused |
error |
string or null | rate_limit_exceeded when refused; degraded when the store could not answer and the degradation mode refused |
Other answers:
| Status | Body | Cause |
|---|---|---|
400 |
{"error": "validation_error", "message": "cost must be positive"} |
cost below 1 |
400 |
{"error": "validation_error", "message": "unknown profile 'gold'"} |
profile is not configured |
400 |
{"error": "validation_error", "message": "cost 6 exceeds the profile's capacity 5, so it can never be admitted"} |
cost above the profile's capacity; nothing consumed |
403 |
{"error": "profile_not_permitted", "message": "profile 'enterprise' exceeds the free tier's limits"} |
profile is above your tier; nothing consumed |
400 / 422 |
text | Body not JSON / key missing (Errors) |
This route does not answer 503. When the store cannot answer, the result follows the degradation mode (below).
curl -s -X POST http://localhost:8080/v1/ratelimit/check \
-H "Content-Type: application/json" \
-H "Authorization: Bearer test-api-key" \
-d '{"key": "user:12345", "cost": 1}'Permission: RateLimitStatus. Reads the caller's bucket for key without
consuming. Percent-encode reserved characters in key; the path segment is
decoded before use.
200:
{ "key": "user:12345", "tokensRemaining": 997, "limit": 1000, "resetAt": "2026-09-27T12:00:01Z" }| Field | Type | Description |
|---|---|---|
key |
string | The key as you sent it |
tokensRemaining |
integer | Tokens in the bucket now |
limit |
integer | Your tier's capacity |
resetAt |
string (ISO-8601) | When the bucket would be full again, per the configured algorithm: the same value a check on it would report |
Status always uses your tier's profile, even if checks on this key named a
narrower one. A key never seen reads as full, with resetAt the current time:
there is nothing to wait for. No X-RateLimit-* headers.
503: the store could not be read. The service does not guess "full":
{ "error": "storage_unavailable", "message": "The rate-limit store could not be read; the status is unknown, not full" }curl -s http://localhost:8080/v1/ratelimit/status/user:12345 \
-H "Authorization: Bearer test-api-key"Your key's tier picks your profile: the entry under rate-limit.profiles named
after the tier, or rate-limit.default-capacity / default-refill-rate-per-second
(100, 10/s) if there is none. The shipped profiles:
| Profile | capacity |
refillRatePerSecond |
|---|---|---|
free |
20 | 2 |
basic |
100 | 10 |
premium |
1,000 | 100 |
enterprise |
10,000 | 1,000 |
A request's profile may only narrow your limits: it is accepted when its
capacity and its refill rate are both at most your tier's, and its
ttlSeconds at least your tier's (under sliding-window that is the window,
and a shorter window is a wider limit). A premium key may
name free or basic; a free key only free. The bucket itself is identified
by client and key alone, so naming a profile applies different limits to the
same bucket, not a new one.
The algorithm is set per deployment by RATE_LIMIT_ALGORITHM: token-bucket
(default), leaky-bucket (the refill rate is the leak rate) or
sliding-window (capacity per window of the profile's ttlSeconds, 3600 in
the shipped profiles). The request and response shapes are the same for all
three; resetAt and retryAfter come from the algorithm in use.
Each attempt at the rate-limit store has a timeout (TIMEOUT_RATE_LIMIT_CHECK,
default 2 s; compose sets 10 s). DynamoDB service errors are retried up to 3
times; a timeout is not. The store sits behind a bulkhead and a circuit
breaker, which opens after 20 consecutive failed calls
(CIRCUIT_BREAKER_MAX_FAILURES) and stays open 30 s. OCC conflicts do not
count toward it: a check that loses every OCC retry is refused with a 429, not
degraded.
When the breaker is open, the bulkhead is full, or the call fails, a check is
answered by DEGRADATION_MODE:
| Mode | Answer |
|---|---|
reject-all (default) |
429, retryAfter: 60, Retry-After: 60, resetAt 60 s from now, error: "degraded", message: "The rate-limit store could not answer; refused by the degradation mode". |
allow-all |
200, tokensRemaining and limit both the profile's capacity, resetAt 60 s from now. Nothing is counted, so spend is unbounded while degraded. |
Both carry the header X-Gate-Degraded: true, which no other answer has: it
is the way to tell a degraded 429 from an empty bucket, and a degraded 200
from a counted one. Any other mode stops startup. Degraded answers are counted
in the gate_degraded_total metric.
The flow:
checkthe key.newmeans you now own it: keep theclaimIdit returns, and run the operation.- When it succeeds,
completethe key with the response to replay. When it fails without effect,failthe key so a retry can claim it. Send theclaimIdwith either (why). - Later checks answer
in_progresswhile the key is pending,duplicatewith the stored response once it is completed, andconflictwhen therequestBodydiffers from the one that claimed it.
A pending key stays in_progress until it is completed, failed, or its TTL
passes. Only the client that claimed a key can complete or fail it; for any
other client the key does not exist.
Idempotency store calls (check, complete, fail) have a timeout
(TIMEOUT_IDEMPOTENCY_CHECK, default 2 s; compose sets 10 s) and a bulkhead;
a failed call is not retried. A timeout, a full bulkhead or a store error
answers 503 storage_unavailable.
Permission: IdempotencyCheck.
{ "idempotencyKey": "payment:abc-123", "ttl": 86400, "requestBody": "{\"amount\":100}" }| Field | Type | Required | Description |
|---|---|---|---|
idempotencyKey |
string | yes | The operation's key, scoped to your client |
ttl |
integer | no | Seconds the record should live. Default IDEMPOTENCY_DEFAULT_TTL (86400); capped at IDEMPOTENCY_MAX_TTL_SECONDS (86400). Zero or negative is a 400. |
requestBody |
string | no | Any string that identifies the request, usually its body. Only its SHA-256 is stored. The hash is over the exact UTF-8 bytes, so reordered or reformatted JSON hashes differently. |
The record's TTL is a DynamoDB TTL attribute. DynamoDB removes expired items
some time after they expire, so the service checks expiry itself: once the TTL
passes, the key answers new to the next check, whether or not the item has
been removed.
Answers:
| Status | status |
Meaning |
|---|---|---|
200 |
new |
You claimed the key (or reclaimed one that was failed). Run the operation. |
202 |
in_progress |
The key is pending. Do not run the operation. |
200 |
duplicate |
The operation completed; originalResponse holds what was stored. |
409 |
conflict |
This check's requestBody hash differs from the stored one. |
400 |
validation_error |
ttl is zero or negative. Nothing is claimed. |
A conflict needs both hashes: a check without requestBody, or a key claimed
without one, never conflicts.
{
"status": "duplicate",
"idempotencyKey": "payment:abc-123",
"originalResponse": {
"statusCode": 201,
"body": "{\"paymentId\":\"pay_xyz789\"}",
"headers": { "Content-Type": "application/json" }
},
"firstSeenAt": "2026-09-27T11:58:02.114Z",
"message": null,
"error": null,
"claimId": null
}| Field | Type | Description |
|---|---|---|
status |
string | new, in_progress, duplicate or conflict |
idempotencyKey |
string | The key as you sent it |
originalResponse |
object or null | For duplicate: statusCode (integer), body (string), headers (object, {} if none were stored). null otherwise. |
firstSeenAt |
string (ISO-8601) or null | When the key was claimed; set for in_progress and duplicate |
message |
string or null | in_progress: "Operation is currently being processed". conflict: "Request body does not match the original request for this idempotency key". |
error |
string or null | idempotency_conflict on the 409; null otherwise |
claimId |
string or null | Set on new only: a UUID naming this claim. Send it back on complete or fail. |
A new answer is {"status": "new", "idempotencyKey": "...", "originalResponse": null, "firstSeenAt": null, "message": null, "error": null, "claimId": "0b9d6c1e-6a1f-4c55-9b0e-2f5a8f4f7d21"}.
A key can be claimed more than once: after it is failed, and after its TTL
passes. Both claims are yours, so without the ID a slow first run's late
complete would store its response on the second run's claim, and its late
fail would release that claim while the second run is still going. With the
claimId, complete and fail apply only to the claim that was issued it;
for any other they answer 409 not_pending
(ADR-006).
The ID is optional. A call without one is accepted on any pending claim of yours for the key, as before.
503: storage_unavailable (the store failed or timed out, or the claim
lost its retries to concurrent changes of the record) or storage_corruption
(the stored record cannot be decoded; whether the operation ran is unknown, so
do not proceed):
{ "error": "storage_corruption", "message": "Idempotency record corrupted: unknown status value: 'Done'" }curl -s -X POST http://localhost:8080/v1/idempotency/check \
-H "Content-Type: application/json" \
-H "Authorization: Bearer test-api-key" \
-d '{"idempotencyKey": "payment:abc-123", "ttl": 3600}'Permission: IdempotencyComplete. Stores the response to replay, and moves
the key from pending to completed. The service stores and replays these values
as given; it does not interpret them.
{ "statusCode": 201, "body": "{\"paymentId\":\"pay_xyz789\"}", "headers": { "Content-Type": "application/json" } }| Field | Type | Required | Description |
|---|---|---|---|
statusCode |
integer | yes | Status to replay |
body |
string | yes | Body to replay |
headers |
object of strings | no | Headers to replay; {} if omitted |
claimId |
string | no | The claimId from the new answer. When sent, only that claim is completed. |
| Status | Body | Cause |
|---|---|---|
200 |
{"idempotencyKey": "payment:abc-123", "status": "completed", "message": null, "error": null} |
Stored |
409 |
{"idempotencyKey": "...", "status": "conflict", "message": "Could not store response - key may not exist or is not pending", "error": "not_pending"} |
No pending key of yours by that name: never claimed, already completed, failed, expired, or another client's. With a claimId: also when the key was claimed again since that ID was issued (the message says so). |
413 |
{"error": "response_too_large", "message": "The response to store is 412345 bytes encoded; ..."} |
The stored response would exceed 358,400 bytes (350 KiB). The key stays pending: store a smaller response, such as a reference to the result, or fail the key. |
503 |
{"error": "storage_unavailable", ...} |
Store failure or timeout |
The size limit counts the response as the store encodes it: JSON with
statusCode, body, headers and a completion time, so escaping counts (a
" in body costs two bytes). The stored response is replayed inline in the
check answer, not streamed. The size is checked before the store is touched,
so an oversized response answers 413 whatever the key's state.
curl -s -X POST http://localhost:8080/v1/idempotency/payment:abc-123/complete \
-H "Content-Type: application/json" \
-H "Authorization: Bearer test-api-key" \
-d '{"statusCode": 201, "body": "{\"paymentId\":\"pay_xyz789\"}"}'Permission: IdempotencyComplete. Releases a pending key after its
operation failed without effect, so the next check answers new and the
caller can retry. A completed key is never reopened: its operation ran.
The body is optional: none, or {"claimId": "..."} with the claimId from
the new answer. When sent, only that claim is released.
| Status | Body | Cause |
|---|---|---|
200 |
{"idempotencyKey": "payment:abc-123", "status": "failed", "message": null, "error": null} |
Released |
409 |
{"idempotencyKey": "...", "status": "conflict", "message": "Could not mark failed - key may not exist or is not pending", "error": "not_pending"} |
No pending key of yours by that name; or, with a claimId, the key was claimed again since that ID was issued |
503 |
{"error": "storage_unavailable", ...} |
Store failure or timeout |
curl -s -X POST http://localhost:8080/v1/idempotency/payment:abc-123/fail \
-H "Authorization: Bearer test-api-key"Quotas cap LLM tokens (input plus output) per user, per agent and per org. A check charges its estimate to every level it names at once: if any level would go over its limit, nothing is charged anywhere. Levels are checked in the order user, agent, org, and the first that would overflow is reported.
Each counter has a fixed window that starts at its first charge and lasts the
level's window length; after it lapses, the next charge starts a new window.
Counters are scoped to your client, so two clients naming the same userId
meter separate counters.
| Setting (env) | Default |
|---|---|
TOKEN_QUOTA_ENABLED |
false (compose: true) |
TOKEN_QUOTA_USER_LIMIT / TOKEN_QUOTA_USER_WINDOW |
1,000,000 tokens / 3600 s |
TOKEN_QUOTA_AGENT_LIMIT / TOKEN_QUOTA_AGENT_WINDOW |
500,000 tokens / 3600 s. With quotas on, startup fails if the limit is above 80% of the user limit. |
TOKEN_QUOTA_ORG_LIMIT / TOKEN_QUOTA_ORG_WINDOW |
10,000,000 tokens / 86400 s |
TOKEN_QUOTA_RESERVATION_TTL_SECONDS |
3600 s |
TIMEOUT_QUOTA_CHECK |
5 s per check or reconcile, including its retries (compose: 10 s) |
With quotas disabled, both routes answer 404 with an empty body. The
permission check still runs first, so a key without the permission gets 403.
Permission: QuotaCheck. Reserves the estimate before the LLM call.
{ "userId": "user:alice", "agentId": "agent:planner", "orgId": "org:acme",
"estimatedInputTokens": 1200, "estimatedOutputTokens": 400 }| Field | Type | Required | Description |
|---|---|---|---|
userId |
string | yes | User charged; always a level |
agentId |
string | no | Adds the agent level |
orgId |
string | no | Adds the org level |
estimatedInputTokens |
integer | yes | Expected prompt tokens, at least 0 |
estimatedOutputTokens |
integer | no, default 0 |
Expected completion tokens, at least 0 |
Allowed (200):
{ "allowed": true, "remainingTokens": { "user": 998400, "agent": 498400, "org": 9998400 },
"exceededLevel": null, "retryAfter": null, "message": null,
"reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "error": null }Exceeded (429): nothing was charged. Header: Retry-After (same as
retryAfter).
{ "allowed": false, "remainingTokens": {}, "exceededLevel": "org", "retryAfter": 1740,
"message": "org quota exceeded: 9999000/10000000 tokens used", "reservationId": null,
"error": "quota_exceeded" }| Field | Type | Description |
|---|---|---|
allowed |
boolean | Whether the estimate was charged |
remainingTokens |
object | Per requested level (user, agent, org): the limit minus usage after this charge. {} unless allowed. |
exceededLevel |
string or null | user, agent or org on a 429 |
retryAfter |
integer or null | 429: seconds until the exceeded level's window ends (at least 1) |
message |
string or null | Why it was refused |
reservationId |
string or null | Set when allowed; pass it to /v1/quota/reconcile |
error |
string or null | quota_exceeded on the 429; null when allowed |
400, never fits: the estimate alone is larger than a requested level's
limit, so no window could admit it. Nothing is reserved, and there is no
Retry-After: {"error": "validation_error", "message": "estimate of 1001 tokens exceeds the user limit of 1000, so it can never be admitted"}.
503, contended: the store lost every conditional write in its 25 attempts,
or the reservation record could not be written (the charge is then released).
Nothing was reserved; retry after Retry-After: 1.
{ "error": "contended", "message": "quota state contended after 25 attempts; nothing was reserved, retry shortly" }503, store failure: a timeout, a full bulkhead or an SDK error, with
Retry-After: 1. A check that timed out may still have charged; that charge
cannot be reconciled and stays counted until its window ends.
{ "error": "storage_unavailable", "message": "The quota store did not answer; the check did not complete, retry it" }400: {"error": "validation_error", "message": "token estimates must be non-negative"}.
curl -s -X POST http://localhost:8080/v1/quota/check \
-H "Content-Type: application/json" \
-H "Authorization: Bearer test-api-key" \
-d '{"userId": "user:alice", "orgId": "org:acme", "estimatedInputTokens": 1200, "estimatedOutputTokens": 400}'Permission: QuotaReconcile. After the LLM call, replaces the reservation's
estimate with actual usage. The estimate comes from the reservation the check
stored, never from the request, so a reconcile can only give back its own
reservation's charge.
{ "reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "actualInputTokens": 1000, "actualOutputTokens": 550 }| Field | Type | Required | Description |
|---|---|---|---|
reservationId |
string | yes | From the check's 200 |
actualInputTokens |
integer | yes | Prompt tokens used, at least 0 |
actualOutputTokens |
integer | yes | Completion tokens used, at least 0 |
Other fields are ignored; a body without reservationId answers 422.
For each counter the check charged, actual - estimate is applied (it may be
negative) together with marking the reservation reconciled, in one
transaction. A counter whose window has rolled over since the check never held
the estimate, so it gets only usage above the estimate, never a refund.
Actual usage is recorded even past the limit; later checks for that identity
are then refused until the window ends.
| Status | Body | Meaning |
|---|---|---|
200 |
{"status": "reconciled", "inputDelta": -200, "outputDelta": 150} |
Applied. Deltas are actual minus the stored estimate. Sending the same usage again returns the same 200 and changes nothing, so a retry is safe. |
404 |
{"error": "reservation_not_found", "message": "No live reservation with this ID for this key: unknown, expired, or another client's"} |
The three cases are deliberately indistinguishable. Nothing was written; the estimate stays counted. |
409 |
{"error": "already_reconciled", "message": "Reservation already reconciled with 1000 input and 550 output tokens"} |
Reconciled before with different usage |
503 + Retry-After: 1 |
{"error": "contended", "message": "reconciliation contended after 25 attempts; nothing was recorded, retry shortly"} |
Not recorded; send the same request again |
503 + Retry-After: 1 |
{"error": "storage_unavailable", "message": "The quota store did not answer; retry with the same usage"} |
Store failure or timeout; whether it was recorded is unknown, and a repeat with the same usage is safe |
400 |
{"error": "validation_error", "message": "token counts must be non-negative"} |
A negative count |
A reservation can be reconciled for TOKEN_QUOTA_RESERVATION_TTL_SECONDS
after the check; after that the answer is 404 and the estimate stays counted.
curl -s -X POST http://localhost:8080/v1/quota/reconcile \
-H "Content-Type: application/json" \
-H "Authorization: Bearer test-api-key" \
-d '{"reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "actualInputTokens": 1000, "actualOutputTokens": 550}'No authentication. Liveness only: 200 whenever the process serves HTTP, with
no dependency checks. version is the build's version (sbt-buildinfo), and
commit is the git commit the build was made from, or "unknown" when the
build was not told (an image built without the GIT_COMMIT build argument).
{ "status": "healthy", "version": "0.1.0", "commit": "23fb85b" }No authentication. Checks each component and answers 503 only when a
required one fails.
| Component | Required | Check |
|---|---|---|
dynamodb_ratelimit |
yes | The rate-limit table (DescribeTable, bounded only by the SDK's own timeouts) |
dynamodb_idempotency |
yes | The idempotency table |
dynamodb_quota |
yes | The quota table; listed only when quotas are enabled |
kinesis |
no | The event stream. Events are fire-and-forget and no request waits on them. Reads ok when Kinesis is disabled. |
status |
HTTP | Meaning |
|---|---|---|
ok |
200 | Every component is ok |
degraded |
200 | Only optional components fail |
unavailable |
503 | A required component fails |
{
"status": "degraded",
"components": [
{ "name": "dynamodb_ratelimit", "status": "ok", "required": true, "details": null },
{ "name": "dynamodb_idempotency", "status": "ok", "required": true, "details": null },
{ "name": "dynamodb_quota", "status": "ok", "required": true, "details": null },
{ "name": "kinesis", "status": "error", "required": false, "details": "<error message>" }
]
}A component's status is ok or error; details is the error message, or
null.
Permission: AdminMetrics (the built-in admin-api-key). Prometheus text
exposition, text/plain; charset=UTF-8. Gate's own series are prefixed
gate_. With PROMETHEUS_ENABLED=false an authorized request gets 404 with
an empty body.
curl -s -H "Authorization: Bearer admin-api-key" http://localhost:8080/metricsMounted only with DASHBOARD_ENABLED=true (default false; docker compose
sets it). Every dashboard route is unauthenticated: the config POST
rewrites the demo bucket's profile for everyone on that instance, and the
decision stream carries every client's key ID. With the flag off these paths
fall through to authentication (401 without a key, 404 with one).
The demo bucket is the fixed key dashboard-demo in the real rate-limit
store, starting from rate-limit.default-* (100 tokens, 10/s).
| Method | Path | Answer |
|---|---|---|
GET |
/dashboard |
The HTML page |
GET |
/dashboard/api/config |
{"capacity", "refillRatePerSecond", "ttlSeconds"} of the demo bucket |
POST |
/dashboard/api/config |
Body with the same three fields, each above 0. Answers them plus "message": "Configuration updated successfully", or 400 {"error": "validation_error", "message": "<reason>"}. |
POST |
/dashboard/api/check |
Consumes 1 token. Always 200: allowed, tokensRemaining, limit, resetAt, plus retryAfter when refused |
GET |
/dashboard/api/status |
tokensRemaining, limit, and resetAt (the store's value; the current time for a bucket not yet created) |
GET |
/dashboard/api/stats |
Server-sent events every 500 ms: {"tokensRemaining", "limit", "timestamp"} (epoch ms) |
GET |
/v1/ratelimit/dashboard/stats |
Server-sent events: each event the service publishes (rate-limit decisions, idempotency and quota events, audit events), as JSON with an event_type field. They pass through a 512-event queue: events are dropped while it is full, and concurrent viewers split the stream between them. |
Every answer that is not a 2xx has a JSON body with two fields: error, a
stable code to branch on, and message, text for a person. That holds for
every route and every status, including the answers from authentication, an
unknown path, a body that cannot be decoded, and an unhandled failure.
| Shape | Where |
|---|---|
{"error": "<code>", "message": "<text>"} |
Refusals and failures (codes below) |
The same, plus "retryAfter": N |
Auth throttle 429 |
The route's own response shape, with error and message set |
Rate-limit 429; idempotency conflict 409 and the complete/fail 409s; quota check 429. The outcome is still in allowed or status. |
/ready is not an error body: its 503 is the same health document as its 200.
| Status | error |
Route | Cause |
|---|---|---|---|
| 400 | invalid_request |
Any route with a body | The body is not JSON, or is empty |
| 400 | validation_error |
POST /v1/ratelimit/check |
cost below 1 or above the profile's capacity, or an unknown profile |
| 400 | validation_error |
POST /v1/idempotency/check |
ttl zero or negative |
| 400 | validation_error |
POST /v1/quota/check, /reconcile |
A negative token count, or an estimate above a level's limit |
| 400 | validation_error |
POST /dashboard/api/config |
A missing field, or one not above 0 |
| 401 | unauthorized |
Any authenticated route | No key, or an unknown key |
| 403 | forbidden |
Any authenticated route | The key lacks the route's permission |
| 403 | profile_not_permitted |
POST /v1/ratelimit/check |
profile above the key's tier |
| 404 | not_found |
Any path | No route matches; or a quota route with quotas disabled, or /metrics with Prometheus disabled |
| 404 | reservation_not_found |
POST /v1/quota/reconcile |
Unknown, expired, or another client's reservation |
| 409 | idempotency_conflict |
POST /v1/idempotency/check |
requestBody differs from the one that claimed the key |
| 409 | not_pending |
POST /v1/idempotency/{key}/complete, /fail |
No pending key of yours by that name, or the claimId sent is not the current claim's |
| 409 | already_reconciled |
POST /v1/quota/reconcile |
Reconciled before with different usage |
| 413 | response_too_large |
POST /v1/idempotency/{key}/complete |
Stored response over 358,400 bytes encoded |
| 422 | invalid_request |
Any route with a body | JSON without a required field, or a field of the wrong type. message names the field's path (.cost), never its value. |
| 429 | rate_limited |
Any authenticated route | The auth throttle, or a source over its unknown-key limit |
| 429 | rate_limit_exceeded |
POST /v1/ratelimit/check |
The bucket cannot cover cost |
| 429 | degraded |
POST /v1/ratelimit/check |
The store could not answer and reject-all refused; also X-Gate-Degraded: true |
| 429 | quota_exceeded |
POST /v1/quota/check |
A level's quota is spent |
| 500 | internal_error |
Any route | An unhandled failure; whether the request took effect is unknown |
| 503 | storage_unavailable |
Rate-limit status, all idempotency routes, both quota routes (quota adds Retry-After: 1) |
Store error, timeout, or full bulkhead |
| 503 | storage_corruption |
POST /v1/idempotency/check |
The stored record cannot be decoded |
| 503 | contended |
POST /v1/quota/check, /reconcile |
Every conditional write lost; nothing reserved or recorded |
Each point follows from the behavior above.
- Rate limit: read
allowed, not just the status code, and waitRetry-Afterseconds after a 429.X-Gate-Degraded: truemarks an answer the degradation mode gave because the store could not: a 429 that is not an empty bucket, or a 200 that was not counted. The auth throttle's 429 has"error": "rate_limited"and noallowed; it also means back off. - Idempotency: run the operation only on
new, and finish every claimed key withcompleteorfail, or it answersin_progressuntil its TTL passes. A 503 means the state is unknown: do not run the operation. Choose a TTL longer than the operation can take: once it passes, the next check answersnewand the operation runs again. Send theclaimIdoncompleteandfail, so a run that outlived its TTL gets a 409 instead of ending the claim of the run that replaced it. - Quota: keep the check's
reservationIdand reconcile with it. After a 503 on reconcile, send the same request again afterRetry-After; a repeat with the same usage is safe. A 503 on check means you were not admitted.