Rate Limits
What the limits are, how a key learns it hit one, and which operations cost extra
Limits exist per key and per target organisation. A tech provider has a further ceiling across all its clients.
The numbers below are the defaults. They are configurable, so treat them as the shape of the system rather than constants to hard-code. Build against the signals - retryAfterMs and Retry-After - not against the figures.
The Buckets
Per key, per target org:
| Bucket | Limit |
|---|---|
| Reads | 10/s, burst 30 |
| Writes | 1/s, burst 60 (so 60 per minute) |
| Expensive operations | 2/s, burst 10 - on top of the read or write bucket |
Per organisation, across every key: 300 writes/min.
A tech provider also has a provider-wide ceiling that grows with the number of managed clients, up to 100 writes/s and 500 reads/s. One client cannot consume the whole provider allowance, and adding clients raises the ceiling.
Request Shape
| Limit | Value |
|---|---|
| Root fields or aliases per GraphQL operation | 30 |
| Operations per GraphQL batch | 10 |
| MCP request body | 256 KB |
| Page size on any list operation, for keys | clamped to 100 |
The page clamp is silent - ask for 500 conversations and you get 100, not an error. Page with the cursor rather than assuming your requested size was honoured.
Expensive Operations
These draw on the extra bucket as well as the normal one, because each does real work:
| Area | Operations |
|---|---|
| Search | search_messages / searchMessages, searchChats, searchConversations |
| Filtered lists | list_conversations with rules or a collection; list_contacts / contactsApi with search, channel or fields |
| Counts | inbox_counts / inboxCounts, chatCollectionCounts |
| Deletes | delete_chat, delete_conversation |
| Channels | resume_channel |
| WhatsApp templates | sync_, create_, delete_whatsapp_template |
| Telegram | sync_telegram_bot_settings |
A plain list_conversations is a normal read; adding rules makes it expensive. If you are polling, poll the cheap form and use the expensive one only when the filter is the point.
When You Hit a Limit
| Surface | Response |
|---|---|
| GraphQL | An error with extensions.code = "RATE_LIMITED" and extensions.retryAfterMs |
HTTP and /mcp | 429 with a Retry-After header, in seconds |
| MCP tool text | Rate limited - retry after Ns |
Honour the value you are given rather than backing off on a schedule of your own - it reflects when the bucket actually refills.
The same RATE_LIMITED shape is reused for two non-traffic limits: the WhatsApp template daily cap and a full async-send backlog.
Failed Authentication
Repeatedly presenting a bad key is throttled separately: 20 failures per minute per IP for malformed keys, and the same per known keyId. A deploy shipping a stale key will lock itself out rather than hammer the door.
WhatsApp Template Daily Cap
Creating templates with a key is capped per org: 20 per rolling 24 hours by default.
| It is a token bucket | It refills gradually. There is no midnight reset to wait for. |
| A token is spent late | Only after ownership, media, duplicate-name and structure checks pass. A malformed request costs nothing. |
| A Meta rejection still spends one | The attempt reached Meta, so it counts. |
| Over the cap | RATE_LIMITED with retryAfterMs. |
| Limiter unreachable | temporarily unavailable; retry shortly - retry, do not treat it as a rejection. |
Template updates are not capped, and members are not capped at all.
Errors Other Than Limits
What a key sees when something goes wrong:
| Situation | GraphQL | MCP |
|---|---|---|
| Bad input | 400 | the tool's error text |
| Missing scope, or a foreign org | 403 | the tool's error text |
| Unknown id | See below | a not-found text |
| A contact mid-merge | 409 | a contact of this conversation is being merged; retry shortly |
| Over a limit | RATE_LIMITED | Rate limited - retry after Ns |
| Plan does not allow it | Plan error | Not available on the current plan: ... |
| Anything else | - | Internal error |
MCP passes 400, 403, 404, 409 and 429 through as readable text an agent can act on. Everything else becomes the opaque Internal error - the detail is logged server-side, not returned. So an agent seeing Internal error should surface it rather than retry in a loop.
There Is No Existence Oracle
A foreign id behaves exactly like a missing one:
- Operations behind the chat-access guard answer 403 uniformly for foreign, missing and malformed chat ids.
- Operations reading org-scoped data return
nullor an empty page. - MCP returns a not-found text for any id outside the target org.
You cannot use error responses to discover whether something exists in another organisation.
Webhook Receivers
Delivery to your endpoint has its own rules:
| Rule | Detail |
|---|---|
| Scheme and host | Public https:// only. An org-own webhook URL is validated when you save it. |
| Redirects | Not followed. A receiver answering 301 or 308 fails permanently. |
| Timeout | 5 seconds. |
| Retries | 429 and 5xx are retried, honouring your Retry-After. |
| Expiry | Events older than 24 hours are dropped. |
| Subscriptions | At most 10 per provider. |
The redirect rule catches people out. If your endpoint is behind something that normalises /webhook to /webhook/ with a 301, every delivery fails - and it fails permanently, not as a retryable error. Point the subscription at the final URL.
See Envelope & delivery for the retry and deduplication contract.
Designing Within the Limits
- Bulk sending: use
async: trueso each send returns immediately, and watch the write bucket rather than wall-clock time. - Mirroring an inbox: take the initial state from paged reads, then keep up with webhooks instead of polling. Polling expensive reads is what exhausts the extra bucket.
- Backfills: page with cursors at 100 per call and pace yourself against the read bucket.
- Many clients: remember the per-org write limit applies to each client separately, so work spread across clients scales better than the same volume aimed at one.