Rate limits & quotas
Rate limits & quotas
Per-key throughput tiers, daily/monthly quotas, and per-IP ceilings — every number generated from config, with self-describing headers and a fail-open posture.
Drazill enforces limits at two levels: a per-key tier (throughput + quota) and a per-IP ceiling (a coarse shared-address guard that runs before authentication). Every table here is generated from the enforcing code/config and drift-gated — the numbers cannot rot.
Per-key throughput (tier)
Each API key has a rate-limit tier with a requests-per-minute ceiling. Tier changes are
admin-reviewed (self-serve requests via POST /api/v1/api-keys/{id}/tier-request).
Daily & monthly quotas
On top of the per-minute ceiling, keys carry daily and monthly request quotas (part of the
Data & API usage layer). Usage is metered when DATA_API_USAGE_METERING_ENABLED is on.
Metering fails open. If the usage store (Redis) has a blip, metering and quota
enforcement fail open — a cache outage never denies a paying customer a request. The
per-minute tier ceiling still applies. This is a deliberate customer-favourable posture
(app/core/config.py).
Per-IP ceilings
Before a request is authenticated, a per-IP limiter applies a coarse ceiling by path category. These are per-IP, not per-user — many legitimate users share an address behind carrier-grade NAT, corporate proxies, and VPNs, so the ceilings are generous. Fine-grained throttling is the per-key tier above.
The per-IP limiter also fails open if its store is unavailable — a blip must never block signup or trading.
Per-endpoint limits
Some endpoints carry their own per-route limit on top of the two levels above. These are the
values enforced by the route decorators in app/api/routes/orders.py — hand-maintained here
because the generator reads config and enums, not decorators, so treat that file as the
source if the two ever disagree (a test asserts this table’s numbers against the mounted
routes, so they cannot drift silently).
A batch costs what its legs cost. POST /orders/batch is one HTTP request but N order
placements, so each leg is charged individually against the shared per-IP orders bucket
(120 / 60s, in the table above) in addition to the 10 / min on the route itself. A 20-leg
batch therefore consumes 20 of that bucket, exactly as 20 single placements would —
batching is a round-trip saving, not a rate-limit loophole. The legs are charged before
any of them runs, so an over-quota batch places nothing rather than a prefix, and returns
429.
Response headers
Limits are self-describing — you don’t have to guess:
These are exposed via CORS so a first-party browser client can read them
(docs/api/cors-and-auth.md).
Handling 429
On a 429, back off for the Retry-After duration and retry. The SDKs do this for you —
configure the retry budget and they retry 5xx and 429 responses, honouring Retry-After:
If you handle HTTP directly, read Retry-After and sleep before retrying; use an
Idempotency-Key on any mutation you retry
so a retry can never double-apply.
The full, per-plan rate/quota sheet is maintained in the repo at
docs/legal/api-rate-plans.md — which carries a PENDING LEGAL REVIEW banner: the
numbers there (and here) are the technical enforcement values from app/core/config.py
and app/models/api_key.py, not a price sheet or a service-level guarantee. Pricing
and contractual commitments are owner-defined and pending counsel review.

