Limits on our servers
Each throttle is a bucket of tokens. A request spends tokens from its bucket. The bucket holds up to its capacity, which is your burst, and refills at its refill rate, in tokens per second.ThrottleBurstCapacityRefill
place Hzcancel Hzread Hzaccount Hzstream Hzhistory Hzhistory bucket meters the record reads: fills, transactions, settled orders, and the history routes. Each such read costs a base plus one token per page of rows, so a larger page spends more. GET /v3/limits reports every bucket’s numbers.
Websocket subscriptions
Thestream bucket meters how fast you subscribe. A second limit caps how much you hold: a connection may watch up to markets at once. An event counts as the markets it contains. A subscribe that would take you over that answers SUBSCRIPTION_LIMIT_EXCEEDED, and it subscribes to nothing. GET /v3/limits reports the cap.
The key routes and the subaccount routes spend the account bucket, one token a request. A transfer spends from it too.
Limits at the edge
The edge is a filter in front of our servers. It checks every request on every route before the request reaches them.Pace your requests
A throttle refills over time, up to its capacity: Here is the capacity, is the refill rate, and is the time that has passed.- Model each throttle in your client.
- When you send a request, subtract its cost from your model.
- While your model holds too few tokens, queue the request instead of sending it.
Error response
When you go over a limit, we answer429 with a Retry-After header:
- Wait
Retry-Afterseconds before you retry. Retrying sooner hits the same empty throttle. - Split your work across keys. The throttle counts per key, not per connection.
- Handle a
429even if you model the throttles. Our servers share token counts with each other gradually, so your model is never exact.

