Skip to content

Latest commit

 

History

History
763 lines (606 loc) · 37.9 KB

File metadata and controls

763 lines (606 loc) · 37.9 KB

API Reference

Every route, field, status code and header below is taken from the code under src/main/scala, with defaults from src/main/resources/application.conf.

Overview

  • Base URL (docker compose): http://localhost:8080. The port is SERVER_PORT (default 8080), and compose publishes it on 8080.
  • Bodies: request bodies are parsed as JSON whatever the Content-Type says, but send Content-Type: application/json. JSON responses are application/json.
  • Fields: outside the demo dashboard, every response field shown below is always present. An optional value that does not apply is null (responses are printed with nulls kept).
  • X-Request-Id: every response echoes the request's X-Request-Id, or a generated UUID when the request had none. That includes the 401, 403 and 429 answers from authentication, the 400/422 body-decoding answers, and a 500.
  • Errors: every answer that is not a 2xx has a JSON body with error (a stable code) and message (text for a person). See Errors.
  • Tenancy: every key a caller names (rate-limit key, idempotency key, quota user/agent/org, reservation ID) is scoped to the authenticated client (ADR-005). Two clients using the same key string never share state and cannot see each other's.
Method Path Permission (name in the keys secret) Purpose
POST /v1/ratelimit/check RateLimitCheck (ratelimit_check) Consume tokens from a bucket
GET /v1/ratelimit/status/{key} RateLimitStatus (ratelimit_status) Read a bucket without consuming
POST /v1/idempotency/check IdempotencyCheck (idempotency_check) Claim a key, or learn its state
POST /v1/idempotency/{key}/complete IdempotencyComplete (idempotency_complete) Store the response to replay
POST /v1/idempotency/{key}/fail IdempotencyComplete (idempotency_complete) Release a pending key for retry
POST /v1/quota/check QuotaCheck (quota_check) Reserve estimated LLM tokens
POST /v1/quota/reconcile QuotaReconcile (quota_reconcile) Replace a reservation's estimate with actual usage
GET /health none Liveness
GET /ready none Readiness
GET /metrics AdminMetrics (admin_metrics) Prometheus scrape
various /dashboard, /dashboard/api/*, /v1/ratelimit/dashboard/stats none Demo dashboard, only with DASHBOARD_ENABLED=true (below)

The quota routes exist only with TOKEN_QUOTA_ENABLED=true (compose sets it; the default is false).


Authentication

Sending a key

The middleware reads the key from, in order:

  1. Authorization: Bearer <key> or Authorization: ApiKey <key> (scheme matched case-insensitively);
  2. X-Api-Key: <key>, used when there is no Authorization header with one of those schemes.

/health, /ready and the dashboard routes are matched before authentication and need no key. Every other request goes through it, including requests for paths that do not exist.

Outcomes

Case Status Body
No key 401 {"error": "unauthorized", "message": "Missing API key in Authorization header"}
An unknown key 401 {"error": "unauthorized", "message": "Invalid API key"}
Valid key, sending faster than the auth throttle allows 429 + Retry-After {"error": "rate_limited", "message": "Rate limited. Retry after 42 seconds", "retryAfter": 42}
Any key, from a source over its unknown-key limit 429 + Retry-After The same body
Valid key without the route's permission 403 {"error": "forbidden", "message": "Insufficient permissions: QuotaCheck required"}
Valid key, path that no route matches (or wrong method) 404 {"error": "not_found", "message": "No route for GET /v1/nope"}

The permission check runs before the route touches any state. The 403 message names the permission by its code name (QuotaCheck), not its secret name (quota_check).

Auth throttle. Each key may authenticate security.authentication.rate-limit-per-minute times per minute (AUTH_RATE_LIMIT_PER_MINUTE, default 1000; compose sets 10,000,000 for load tests). The count is per key ID, in memory on each instance, over a one-minute window that starts at the key's first request. Past it, the answer is 429, not 401: back off for Retry-After seconds, do not fix credentials. It counts every authenticated request, including ones that end in 403 or 404.

Unknown-key throttle. Each source may present security.authentication.failed-attempts-per-minute (20) unknown keys per minute. After that, every request from it that carries a key, valid or not, answers 429 with Retry-After until its minute ends, and no key is looked up. A missing key is not counted. The source is the connecting address, or, with AUTH_TRUST_FORWARDED_FOR=true (Terraform sets it, since the ALB fronts every task), the last X-Forwarded-For entry, which the ALB appends. The count is in memory on each instance, so it slows a guesser at one address, not one spread across many; that is the ALB's or WAF's job.

Permissions

Each authenticated route needs one permission; the route table names it and its name in the keys secret. Secret names are matched case-insensitively, the forms without the underscore (ratelimitcheck) are accepted too, and an unrecognized name is ignored.

Where keys come from

The service needs one key source, or it refuses to start.

Secrets Manager (SECRETS_MANAGER_ENABLED=true; Terraform always sets it). The secret is named <secret-prefix>/<environment>/<api-keys-secret-name> (SECRETS_PREFIX, SECRETS_ENVIRONMENT, API_KEYS_SECRET_NAME; default rate-limiter/dev/api-keys) and holds a JSON list. Every field is required:

[
  {
    "apiKey": "<random secret>",
    "apiKeyId": "key_acme_001",
    "clientName": "Acme",
    "tier": "premium",
    "permissions": ["ratelimit_check", "ratelimit_status", "idempotency_check",
                    "idempotency_complete", "quota_check", "quota_reconcile"],
    "active": true
  }
]
  • tier is free, basic, premium or enterprise. An entry with "active": false is skipped. An unknown tier or permission name, in any entry, is an error that names the entry's apiKeyId and the value.
  • The tenant is the entry's apiKeyId. Keep it when you rotate apiKey to keep the client's state.
  • Startup fails if the secret is missing, is not a list in this shape, has an unknown tier or permission, or has no active key.
  • The keys are re-read about every 5 minutes (cache-ttl), so adding or revoking a key takes effect within minutes, not at once. If a re-read fails, the current keys stay. If it finds an unknown tier or permission, the valid entries take effect (so revocations do), the invalid ones cannot authenticate, and the error is logged.

Built-in development keys (ALLOW_BUILT_IN_KEYS=true, set only by docker compose and make run). They are public, so never use them outside local development. When Secrets Manager is enabled it wins, even with this flag set.

Key Tier Permissions
test-api-key premium the six standard permissions (all but AdminMetrics)
free-api-key free the six standard permissions
admin-api-key enterprise the six standard permissions and AdminMetrics

Rate limiting

POST /v1/ratelimit/check

Permission: RateLimitCheck. Consumes cost tokens from the caller's bucket for key, or refuses without consuming.

{ "key": "user:12345", "cost": 1 }
Field Type Required Description
key string yes Bucket identifier, scoped to your client
cost integer no, default 1 Tokens to consume; must be at least 1
profile string no A named profile to apply instead of your tier's (see Profiles)
endpoint string no A label copied onto the decision event; it does not affect the decision

Allowed (200):

{ "allowed": true, "tokensRemaining": 999, "retryAfter": null, "limit": 1000,
  "resetAt": "2026-09-27T12:00:00.010Z", "message": null, "error": null }

Headers: X-RateLimit-Limit (the profile's capacity), X-RateLimit-Remaining, X-RateLimit-Reset (resetAt as epoch seconds).

Refused (429): nothing was consumed. Header: Retry-After (seconds, same as retryAfter). No X-RateLimit-* headers.

{ "allowed": false, "tokensRemaining": null, "retryAfter": 1, "limit": 1000,
  "resetAt": "2026-09-27T12:00:10.000Z", "message": "Rate limit exceeded",
  "error": "rate_limit_exceeded" }
Field Type Description
allowed boolean Whether the tokens were consumed
tokensRemaining integer or null Tokens left after this check; null when refused
retryAfter integer or null Seconds until cost tokens should be available; null when allowed
limit integer Capacity of the profile applied
resetAt string (ISO-8601) When the bucket would be full again, per the configured algorithm
message string or null "Rate limit exceeded" when refused
error string or null rate_limit_exceeded when refused; degraded when the store could not answer and the degradation mode refused

Other answers:

Status Body Cause
400 {"error": "validation_error", "message": "cost must be positive"} cost below 1
400 {"error": "validation_error", "message": "unknown profile 'gold'"} profile is not configured
400 {"error": "validation_error", "message": "cost 6 exceeds the profile's capacity 5, so it can never be admitted"} cost above the profile's capacity; nothing consumed
403 {"error": "profile_not_permitted", "message": "profile 'enterprise' exceeds the free tier's limits"} profile is above your tier; nothing consumed
400 / 422 text Body not JSON / key missing (Errors)

This route does not answer 503. When the store cannot answer, the result follows the degradation mode (below).

curl -s -X POST http://localhost:8080/v1/ratelimit/check \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer test-api-key" \
  -d '{"key": "user:12345", "cost": 1}'

GET /v1/ratelimit/status/{key}

Permission: RateLimitStatus. Reads the caller's bucket for key without consuming. Percent-encode reserved characters in key; the path segment is decoded before use.

200:

{ "key": "user:12345", "tokensRemaining": 997, "limit": 1000, "resetAt": "2026-09-27T12:00:01Z" }
Field Type Description
key string The key as you sent it
tokensRemaining integer Tokens in the bucket now
limit integer Your tier's capacity
resetAt string (ISO-8601) When the bucket would be full again, per the configured algorithm: the same value a check on it would report

Status always uses your tier's profile, even if checks on this key named a narrower one. A key never seen reads as full, with resetAt the current time: there is nothing to wait for. No X-RateLimit-* headers.

503: the store could not be read. The service does not guess "full":

{ "error": "storage_unavailable", "message": "The rate-limit store could not be read; the status is unknown, not full" }
curl -s http://localhost:8080/v1/ratelimit/status/user:12345 \
  -H "Authorization: Bearer test-api-key"

Profiles and tiers

Your key's tier picks your profile: the entry under rate-limit.profiles named after the tier, or rate-limit.default-capacity / default-refill-rate-per-second (100, 10/s) if there is none. The shipped profiles:

Profile capacity refillRatePerSecond
free 20 2
basic 100 10
premium 1,000 100
enterprise 10,000 1,000

A request's profile may only narrow your limits: it is accepted when its capacity and its refill rate are both at most your tier's, and its ttlSeconds at least your tier's (under sliding-window that is the window, and a shorter window is a wider limit). A premium key may name free or basic; a free key only free. The bucket itself is identified by client and key alone, so naming a profile applies different limits to the same bucket, not a new one.

The algorithm is set per deployment by RATE_LIMIT_ALGORITHM: token-bucket (default), leaky-bucket (the refill rate is the leak rate) or sliding-window (capacity per window of the profile's ttlSeconds, 3600 in the shipped profiles). The request and response shapes are the same for all three; resetAt and retryAfter come from the algorithm in use.

When the store cannot answer

Each attempt at the rate-limit store has a timeout (TIMEOUT_RATE_LIMIT_CHECK, default 2 s; compose sets 10 s). DynamoDB service errors are retried up to 3 times; a timeout is not. The store sits behind a bulkhead and a circuit breaker, which opens after 20 consecutive failed calls (CIRCUIT_BREAKER_MAX_FAILURES) and stays open 30 s. OCC conflicts do not count toward it: a check that loses every OCC retry is refused with a 429, not degraded.

When the breaker is open, the bulkhead is full, or the call fails, a check is answered by DEGRADATION_MODE:

Mode Answer
reject-all (default) 429, retryAfter: 60, Retry-After: 60, resetAt 60 s from now, error: "degraded", message: "The rate-limit store could not answer; refused by the degradation mode".
allow-all 200, tokensRemaining and limit both the profile's capacity, resetAt 60 s from now. Nothing is counted, so spend is unbounded while degraded.

Both carry the header X-Gate-Degraded: true, which no other answer has: it is the way to tell a degraded 429 from an empty bucket, and a degraded 200 from a counted one. Any other mode stops startup. Degraded answers are counted in the gate_degraded_total metric.


Idempotency

The flow:

  1. check the key. new means you now own it: keep the claimId it returns, and run the operation.
  2. When it succeeds, complete the key with the response to replay. When it fails without effect, fail the key so a retry can claim it. Send the claimId with either (why).
  3. Later checks answer in_progress while the key is pending, duplicate with the stored response once it is completed, and conflict when the requestBody differs from the one that claimed it.

A pending key stays in_progress until it is completed, failed, or its TTL passes. Only the client that claimed a key can complete or fail it; for any other client the key does not exist.

Idempotency store calls (check, complete, fail) have a timeout (TIMEOUT_IDEMPOTENCY_CHECK, default 2 s; compose sets 10 s) and a bulkhead; a failed call is not retried. A timeout, a full bulkhead or a store error answers 503 storage_unavailable.

POST /v1/idempotency/check

Permission: IdempotencyCheck.

{ "idempotencyKey": "payment:abc-123", "ttl": 86400, "requestBody": "{\"amount\":100}" }
Field Type Required Description
idempotencyKey string yes The operation's key, scoped to your client
ttl integer no Seconds the record should live. Default IDEMPOTENCY_DEFAULT_TTL (86400); capped at IDEMPOTENCY_MAX_TTL_SECONDS (86400). Zero or negative is a 400.
requestBody string no Any string that identifies the request, usually its body. Only its SHA-256 is stored. The hash is over the exact UTF-8 bytes, so reordered or reformatted JSON hashes differently.

The record's TTL is a DynamoDB TTL attribute. DynamoDB removes expired items some time after they expire, so the service checks expiry itself: once the TTL passes, the key answers new to the next check, whether or not the item has been removed.

Answers:

Status status Meaning
200 new You claimed the key (or reclaimed one that was failed). Run the operation.
202 in_progress The key is pending. Do not run the operation.
200 duplicate The operation completed; originalResponse holds what was stored.
409 conflict This check's requestBody hash differs from the stored one.
400 validation_error ttl is zero or negative. Nothing is claimed.

A conflict needs both hashes: a check without requestBody, or a key claimed without one, never conflicts.

{
  "status": "duplicate",
  "idempotencyKey": "payment:abc-123",
  "originalResponse": {
    "statusCode": 201,
    "body": "{\"paymentId\":\"pay_xyz789\"}",
    "headers": { "Content-Type": "application/json" }
  },
  "firstSeenAt": "2026-09-27T11:58:02.114Z",
  "message": null,
  "error": null,
  "claimId": null
}
Field Type Description
status string new, in_progress, duplicate or conflict
idempotencyKey string The key as you sent it
originalResponse object or null For duplicate: statusCode (integer), body (string), headers (object, {} if none were stored). null otherwise.
firstSeenAt string (ISO-8601) or null When the key was claimed; set for in_progress and duplicate
message string or null in_progress: "Operation is currently being processed". conflict: "Request body does not match the original request for this idempotency key".
error string or null idempotency_conflict on the 409; null otherwise
claimId string or null Set on new only: a UUID naming this claim. Send it back on complete or fail.

A new answer is {"status": "new", "idempotencyKey": "...", "originalResponse": null, "firstSeenAt": null, "message": null, "error": null, "claimId": "0b9d6c1e-6a1f-4c55-9b0e-2f5a8f4f7d21"}.

The claim ID

A key can be claimed more than once: after it is failed, and after its TTL passes. Both claims are yours, so without the ID a slow first run's late complete would store its response on the second run's claim, and its late fail would release that claim while the second run is still going. With the claimId, complete and fail apply only to the claim that was issued it; for any other they answer 409 not_pending (ADR-006).

The ID is optional. A call without one is accepted on any pending claim of yours for the key, as before.

503: storage_unavailable (the store failed or timed out, or the claim lost its retries to concurrent changes of the record) or storage_corruption (the stored record cannot be decoded; whether the operation ran is unknown, so do not proceed):

{ "error": "storage_corruption", "message": "Idempotency record corrupted: unknown status value: 'Done'" }
curl -s -X POST http://localhost:8080/v1/idempotency/check \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer test-api-key" \
  -d '{"idempotencyKey": "payment:abc-123", "ttl": 3600}'

POST /v1/idempotency/{key}/complete

Permission: IdempotencyComplete. Stores the response to replay, and moves the key from pending to completed. The service stores and replays these values as given; it does not interpret them.

{ "statusCode": 201, "body": "{\"paymentId\":\"pay_xyz789\"}", "headers": { "Content-Type": "application/json" } }
Field Type Required Description
statusCode integer yes Status to replay
body string yes Body to replay
headers object of strings no Headers to replay; {} if omitted
claimId string no The claimId from the new answer. When sent, only that claim is completed.
Status Body Cause
200 {"idempotencyKey": "payment:abc-123", "status": "completed", "message": null, "error": null} Stored
409 {"idempotencyKey": "...", "status": "conflict", "message": "Could not store response - key may not exist or is not pending", "error": "not_pending"} No pending key of yours by that name: never claimed, already completed, failed, expired, or another client's. With a claimId: also when the key was claimed again since that ID was issued (the message says so).
413 {"error": "response_too_large", "message": "The response to store is 412345 bytes encoded; ..."} The stored response would exceed 358,400 bytes (350 KiB). The key stays pending: store a smaller response, such as a reference to the result, or fail the key.
503 {"error": "storage_unavailable", ...} Store failure or timeout

The size limit counts the response as the store encodes it: JSON with statusCode, body, headers and a completion time, so escaping counts (a " in body costs two bytes). The stored response is replayed inline in the check answer, not streamed. The size is checked before the store is touched, so an oversized response answers 413 whatever the key's state.

curl -s -X POST http://localhost:8080/v1/idempotency/payment:abc-123/complete \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer test-api-key" \
  -d '{"statusCode": 201, "body": "{\"paymentId\":\"pay_xyz789\"}"}'

POST /v1/idempotency/{key}/fail

Permission: IdempotencyComplete. Releases a pending key after its operation failed without effect, so the next check answers new and the caller can retry. A completed key is never reopened: its operation ran.

The body is optional: none, or {"claimId": "..."} with the claimId from the new answer. When sent, only that claim is released.

Status Body Cause
200 {"idempotencyKey": "payment:abc-123", "status": "failed", "message": null, "error": null} Released
409 {"idempotencyKey": "...", "status": "conflict", "message": "Could not mark failed - key may not exist or is not pending", "error": "not_pending"} No pending key of yours by that name; or, with a claimId, the key was claimed again since that ID was issued
503 {"error": "storage_unavailable", ...} Store failure or timeout
curl -s -X POST http://localhost:8080/v1/idempotency/payment:abc-123/fail \
  -H "Authorization: Bearer test-api-key"

Token quotas

Quotas cap LLM tokens (input plus output) per user, per agent and per org. A check charges its estimate to every level it names at once: if any level would go over its limit, nothing is charged anywhere. Levels are checked in the order user, agent, org, and the first that would overflow is reported.

Each counter has a fixed window that starts at its first charge and lasts the level's window length; after it lapses, the next charge starts a new window. Counters are scoped to your client, so two clients naming the same userId meter separate counters.

Setting (env) Default
TOKEN_QUOTA_ENABLED false (compose: true)
TOKEN_QUOTA_USER_LIMIT / TOKEN_QUOTA_USER_WINDOW 1,000,000 tokens / 3600 s
TOKEN_QUOTA_AGENT_LIMIT / TOKEN_QUOTA_AGENT_WINDOW 500,000 tokens / 3600 s. With quotas on, startup fails if the limit is above 80% of the user limit.
TOKEN_QUOTA_ORG_LIMIT / TOKEN_QUOTA_ORG_WINDOW 10,000,000 tokens / 86400 s
TOKEN_QUOTA_RESERVATION_TTL_SECONDS 3600 s
TIMEOUT_QUOTA_CHECK 5 s per check or reconcile, including its retries (compose: 10 s)

With quotas disabled, both routes answer 404 with an empty body. The permission check still runs first, so a key without the permission gets 403.

POST /v1/quota/check

Permission: QuotaCheck. Reserves the estimate before the LLM call.

{ "userId": "user:alice", "agentId": "agent:planner", "orgId": "org:acme",
  "estimatedInputTokens": 1200, "estimatedOutputTokens": 400 }
Field Type Required Description
userId string yes User charged; always a level
agentId string no Adds the agent level
orgId string no Adds the org level
estimatedInputTokens integer yes Expected prompt tokens, at least 0
estimatedOutputTokens integer no, default 0 Expected completion tokens, at least 0

Allowed (200):

{ "allowed": true, "remainingTokens": { "user": 998400, "agent": 498400, "org": 9998400 },
  "exceededLevel": null, "retryAfter": null, "message": null,
  "reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "error": null }

Exceeded (429): nothing was charged. Header: Retry-After (same as retryAfter).

{ "allowed": false, "remainingTokens": {}, "exceededLevel": "org", "retryAfter": 1740,
  "message": "org quota exceeded: 9999000/10000000 tokens used", "reservationId": null,
  "error": "quota_exceeded" }
Field Type Description
allowed boolean Whether the estimate was charged
remainingTokens object Per requested level (user, agent, org): the limit minus usage after this charge. {} unless allowed.
exceededLevel string or null user, agent or org on a 429
retryAfter integer or null 429: seconds until the exceeded level's window ends (at least 1)
message string or null Why it was refused
reservationId string or null Set when allowed; pass it to /v1/quota/reconcile
error string or null quota_exceeded on the 429; null when allowed

400, never fits: the estimate alone is larger than a requested level's limit, so no window could admit it. Nothing is reserved, and there is no Retry-After: {"error": "validation_error", "message": "estimate of 1001 tokens exceeds the user limit of 1000, so it can never be admitted"}.

503, contended: the store lost every conditional write in its 25 attempts, or the reservation record could not be written (the charge is then released). Nothing was reserved; retry after Retry-After: 1.

{ "error": "contended", "message": "quota state contended after 25 attempts; nothing was reserved, retry shortly" }

503, store failure: a timeout, a full bulkhead or an SDK error, with Retry-After: 1. A check that timed out may still have charged; that charge cannot be reconciled and stays counted until its window ends.

{ "error": "storage_unavailable", "message": "The quota store did not answer; the check did not complete, retry it" }

400: {"error": "validation_error", "message": "token estimates must be non-negative"}.

curl -s -X POST http://localhost:8080/v1/quota/check \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer test-api-key" \
  -d '{"userId": "user:alice", "orgId": "org:acme", "estimatedInputTokens": 1200, "estimatedOutputTokens": 400}'

POST /v1/quota/reconcile

Permission: QuotaReconcile. After the LLM call, replaces the reservation's estimate with actual usage. The estimate comes from the reservation the check stored, never from the request, so a reconcile can only give back its own reservation's charge.

{ "reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "actualInputTokens": 1000, "actualOutputTokens": 550 }
Field Type Required Description
reservationId string yes From the check's 200
actualInputTokens integer yes Prompt tokens used, at least 0
actualOutputTokens integer yes Completion tokens used, at least 0

Other fields are ignored; a body without reservationId answers 422.

For each counter the check charged, actual - estimate is applied (it may be negative) together with marking the reservation reconciled, in one transaction. A counter whose window has rolled over since the check never held the estimate, so it gets only usage above the estimate, never a refund. Actual usage is recorded even past the limit; later checks for that identity are then refused until the window ends.

Status Body Meaning
200 {"status": "reconciled", "inputDelta": -200, "outputDelta": 150} Applied. Deltas are actual minus the stored estimate. Sending the same usage again returns the same 200 and changes nothing, so a retry is safe.
404 {"error": "reservation_not_found", "message": "No live reservation with this ID for this key: unknown, expired, or another client's"} The three cases are deliberately indistinguishable. Nothing was written; the estimate stays counted.
409 {"error": "already_reconciled", "message": "Reservation already reconciled with 1000 input and 550 output tokens"} Reconciled before with different usage
503 + Retry-After: 1 {"error": "contended", "message": "reconciliation contended after 25 attempts; nothing was recorded, retry shortly"} Not recorded; send the same request again
503 + Retry-After: 1 {"error": "storage_unavailable", "message": "The quota store did not answer; retry with the same usage"} Store failure or timeout; whether it was recorded is unknown, and a repeat with the same usage is safe
400 {"error": "validation_error", "message": "token counts must be non-negative"} A negative count

A reservation can be reconciled for TOKEN_QUOTA_RESERVATION_TTL_SECONDS after the check; after that the answer is 404 and the estimate stays counted.

curl -s -X POST http://localhost:8080/v1/quota/reconcile \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer test-api-key" \
  -d '{"reservationId": "5f0c2a6e-8a53-4c43-9f53-8f0d7e1f3b0a", "actualInputTokens": 1000, "actualOutputTokens": 550}'

Health and monitoring

GET /health

No authentication. Liveness only: 200 whenever the process serves HTTP, with no dependency checks. version is the build's version (sbt-buildinfo), and commit is the git commit the build was made from, or "unknown" when the build was not told (an image built without the GIT_COMMIT build argument).

{ "status": "healthy", "version": "0.1.0", "commit": "23fb85b" }

GET /ready

No authentication. Checks each component and answers 503 only when a required one fails.

Component Required Check
dynamodb_ratelimit yes The rate-limit table (DescribeTable, bounded only by the SDK's own timeouts)
dynamodb_idempotency yes The idempotency table
dynamodb_quota yes The quota table; listed only when quotas are enabled
kinesis no The event stream. Events are fire-and-forget and no request waits on them. Reads ok when Kinesis is disabled.
status HTTP Meaning
ok 200 Every component is ok
degraded 200 Only optional components fail
unavailable 503 A required component fails
{
  "status": "degraded",
  "components": [
    { "name": "dynamodb_ratelimit", "status": "ok", "required": true, "details": null },
    { "name": "dynamodb_idempotency", "status": "ok", "required": true, "details": null },
    { "name": "dynamodb_quota", "status": "ok", "required": true, "details": null },
    { "name": "kinesis", "status": "error", "required": false, "details": "<error message>" }
  ]
}

A component's status is ok or error; details is the error message, or null.

GET /metrics

Permission: AdminMetrics (the built-in admin-api-key). Prometheus text exposition, text/plain; charset=UTF-8. Gate's own series are prefixed gate_. With PROMETHEUS_ENABLED=false an authorized request gets 404 with an empty body.

curl -s -H "Authorization: Bearer admin-api-key" http://localhost:8080/metrics

Demo dashboard

Mounted only with DASHBOARD_ENABLED=true (default false; docker compose sets it). Every dashboard route is unauthenticated: the config POST rewrites the demo bucket's profile for everyone on that instance, and the decision stream carries every client's key ID. With the flag off these paths fall through to authentication (401 without a key, 404 with one).

The demo bucket is the fixed key dashboard-demo in the real rate-limit store, starting from rate-limit.default-* (100 tokens, 10/s).

Method Path Answer
GET /dashboard The HTML page
GET /dashboard/api/config {"capacity", "refillRatePerSecond", "ttlSeconds"} of the demo bucket
POST /dashboard/api/config Body with the same three fields, each above 0. Answers them plus "message": "Configuration updated successfully", or 400 {"error": "validation_error", "message": "<reason>"}.
POST /dashboard/api/check Consumes 1 token. Always 200: allowed, tokensRemaining, limit, resetAt, plus retryAfter when refused
GET /dashboard/api/status tokensRemaining, limit, and resetAt (the store's value; the current time for a bucket not yet created)
GET /dashboard/api/stats Server-sent events every 500 ms: {"tokensRemaining", "limit", "timestamp"} (epoch ms)
GET /v1/ratelimit/dashboard/stats Server-sent events: each event the service publishes (rate-limit decisions, idempotency and quota events, audit events), as JSON with an event_type field. They pass through a 512-event queue: events are dropped while it is full, and concurrent viewers split the stream between them.

Errors

Error shapes

Every answer that is not a 2xx has a JSON body with two fields: error, a stable code to branch on, and message, text for a person. That holds for every route and every status, including the answers from authentication, an unknown path, a body that cannot be decoded, and an unhandled failure.

Shape Where
{"error": "<code>", "message": "<text>"} Refusals and failures (codes below)
The same, plus "retryAfter": N Auth throttle 429
The route's own response shape, with error and message set Rate-limit 429; idempotency conflict 409 and the complete/fail 409s; quota check 429. The outcome is still in allowed or status.

/ready is not an error body: its 503 is the same health document as its 200.

Error codes

Status error Route Cause
400 invalid_request Any route with a body The body is not JSON, or is empty
400 validation_error POST /v1/ratelimit/check cost below 1 or above the profile's capacity, or an unknown profile
400 validation_error POST /v1/idempotency/check ttl zero or negative
400 validation_error POST /v1/quota/check, /reconcile A negative token count, or an estimate above a level's limit
400 validation_error POST /dashboard/api/config A missing field, or one not above 0
401 unauthorized Any authenticated route No key, or an unknown key
403 forbidden Any authenticated route The key lacks the route's permission
403 profile_not_permitted POST /v1/ratelimit/check profile above the key's tier
404 not_found Any path No route matches; or a quota route with quotas disabled, or /metrics with Prometheus disabled
404 reservation_not_found POST /v1/quota/reconcile Unknown, expired, or another client's reservation
409 idempotency_conflict POST /v1/idempotency/check requestBody differs from the one that claimed the key
409 not_pending POST /v1/idempotency/{key}/complete, /fail No pending key of yours by that name, or the claimId sent is not the current claim's
409 already_reconciled POST /v1/quota/reconcile Reconciled before with different usage
413 response_too_large POST /v1/idempotency/{key}/complete Stored response over 358,400 bytes encoded
422 invalid_request Any route with a body JSON without a required field, or a field of the wrong type. message names the field's path (.cost), never its value.
429 rate_limited Any authenticated route The auth throttle, or a source over its unknown-key limit
429 rate_limit_exceeded POST /v1/ratelimit/check The bucket cannot cover cost
429 degraded POST /v1/ratelimit/check The store could not answer and reject-all refused; also X-Gate-Degraded: true
429 quota_exceeded POST /v1/quota/check A level's quota is spent
500 internal_error Any route An unhandled failure; whether the request took effect is unknown
503 storage_unavailable Rate-limit status, all idempotency routes, both quota routes (quota adds Retry-After: 1) Store error, timeout, or full bulkhead
503 storage_corruption POST /v1/idempotency/check The stored record cannot be decoded
503 contended POST /v1/quota/check, /reconcile Every conditional write lost; nothing reserved or recorded

Client guidance

Each point follows from the behavior above.

  • Rate limit: read allowed, not just the status code, and wait Retry-After seconds after a 429. X-Gate-Degraded: true marks an answer the degradation mode gave because the store could not: a 429 that is not an empty bucket, or a 200 that was not counted. The auth throttle's 429 has "error": "rate_limited" and no allowed; it also means back off.
  • Idempotency: run the operation only on new, and finish every claimed key with complete or fail, or it answers in_progress until its TTL passes. A 503 means the state is unknown: do not run the operation. Choose a TTL longer than the operation can take: once it passes, the next check answers new and the operation runs again. Send the claimId on complete and fail, so a run that outlived its TTL gets a 409 instead of ending the claim of the run that replaced it.
  • Quota: keep the check's reservationId and reconcile with it. After a 503 on reconcile, send the same request again after Retry-After; a repeat with the same usage is safe. A 503 on check means you were not admitted.