Shadow Warden AI — Gateway v7.9 — Explore the API Reference
Home / Doc / Rate limits

Shadow Warden AI Rate Limits

Two limits, published on every response, so you never have to discover one by being refused

The two limits

requests-per-minute

A sliding window, keyed on your API key for authenticated routes — so a shared egress address does not collapse tenants into one bucket. 60/min by default; enterprise keys carry their own rate. The one unauthenticated route, /demo/filter, is keyed on client IP instead, at a fixed 10/min.

requests-per-month

The plan allowance, counted on POST /filter and /filter/batch, reset at the start of each UTC calendar month. Enterprise is unlimited and publishes no monthly policy.

What every response carries

Structured fields from draft-ietf-httpapi-ratelimit-headers-09, alongside the earlier draft spelling and the pre-draft X- spelling. Three generations of client parse three different things, and the cost of emitting all of them is a handful of bytes.

RateLimit-Policy
"requests-per-minute";q=60;w=60, "requests-per-month";q=5000;w=2592000
The static contract. One member per policy: q is the quota, w the window in seconds. Present on every response.
RateLimit
"requests-per-minute";r=59;t=41
The live reading. r is what is left, t is seconds until that policy resets. Present only when the request actually consumed the window.
RateLimit-Limit
60
The earlier draft spelling of the same quota, for clients that parse it.
RateLimit-Remaining
59
The earlier draft spelling of r.
RateLimit-Reset
41
The earlier draft spelling of t — delta-seconds, not a timestamp.
X-RateLimit-*
X-RateLimit-Reset: 1788426730
The pre-draft spelling, kept for compatibility. Note that X-RateLimit-Reset is a Unix epoch, not a delta.
Retry-After
41
Delta-seconds. Sent on 429 only. When both are present it names the same instant as the reset, and it takes precedence.

See it

curl -D-
$ curl -sS -D- -o/dev/null https://api.shadow-warden-ai.com/health

RateLimit-Policy: "requests-per-minute";q=60;w=60
RateLimit-Limit: 60
X-RateLimit-Limit: 60

A route that consumes no quota — such as /health — publishes the policy and no counter. There is no bucket reading that belongs to that request, and inventing one would be a number with nothing behind it.

How to pace against it

  • Read RateLimit on each response and slow down as r approaches zero, rather than waiting for a 429.
  • On 429, honour Retry-After — it takes precedence over the reset parameter. Do not retry sooner.
  • A monthly 429 carries a long Retry-After and an upgrade URL in the body. Retrying will not clear it.
  • Both limits are per API key on authenticated routes. Running two workers on one key halves each worker's share.

The same conventions are described in the OpenAPI document so a generated client sees them too.