Rate Limits

What the limits are, how a key learns it hit one, and which operations cost extra

Limits exist per key and per target organisation. A tech provider has a further ceiling across all its clients.

The numbers below are the defaults. They are configurable, so treat them as the shape of the system rather than constants to hard-code. Build against the signals - retryAfterMs and Retry-After - not against the figures.

The Buckets

Per key, per target org:

BucketLimit
Reads10/s, burst 30
Writes1/s, burst 60 (so 60 per minute)
Expensive operations2/s, burst 10 - on top of the read or write bucket

Per organisation, across every key: 300 writes/min.

A tech provider also has a provider-wide ceiling that grows with the number of managed clients, up to 100 writes/s and 500 reads/s. One client cannot consume the whole provider allowance, and adding clients raises the ceiling.

Request Shape

LimitValue
Root fields or aliases per GraphQL operation30
Operations per GraphQL batch10
MCP request body256 KB
Page size on any list operation, for keysclamped to 100

The page clamp is silent - ask for 500 conversations and you get 100, not an error. Page with the cursor rather than assuming your requested size was honoured.

Expensive Operations

These draw on the extra bucket as well as the normal one, because each does real work:

AreaOperations
Searchsearch_messages / searchMessages, searchChats, searchConversations
Filtered listslist_conversations with rules or a collection; list_contacts / contactsApi with search, channel or fields
Countsinbox_counts / inboxCounts, chatCollectionCounts
Deletesdelete_chat, delete_conversation
Channelsresume_channel
WhatsApp templatessync_, create_, delete_whatsapp_template
Telegramsync_telegram_bot_settings

A plain list_conversations is a normal read; adding rules makes it expensive. If you are polling, poll the cheap form and use the expensive one only when the filter is the point.

When You Hit a Limit

SurfaceResponse
GraphQLAn error with extensions.code = "RATE_LIMITED" and extensions.retryAfterMs
HTTP and /mcp429 with a Retry-After header, in seconds
MCP tool textRate limited - retry after Ns

Honour the value you are given rather than backing off on a schedule of your own - it reflects when the bucket actually refills.

The same RATE_LIMITED shape is reused for two non-traffic limits: the WhatsApp template daily cap and a full async-send backlog.

Failed Authentication

Repeatedly presenting a bad key is throttled separately: 20 failures per minute per IP for malformed keys, and the same per known keyId. A deploy shipping a stale key will lock itself out rather than hammer the door.

WhatsApp Template Daily Cap

Creating templates with a key is capped per org: 20 per rolling 24 hours by default.

It is a token bucketIt refills gradually. There is no midnight reset to wait for.
A token is spent lateOnly after ownership, media, duplicate-name and structure checks pass. A malformed request costs nothing.
A Meta rejection still spends oneThe attempt reached Meta, so it counts.
Over the capRATE_LIMITED with retryAfterMs.
Limiter unreachabletemporarily unavailable; retry shortly - retry, do not treat it as a rejection.

Template updates are not capped, and members are not capped at all.

Errors Other Than Limits

What a key sees when something goes wrong:

SituationGraphQLMCP
Bad input400the tool's error text
Missing scope, or a foreign org403the tool's error text
Unknown idSee belowa not-found text
A contact mid-merge409a contact of this conversation is being merged; retry shortly
Over a limitRATE_LIMITEDRate limited - retry after Ns
Plan does not allow itPlan errorNot available on the current plan: ...
Anything else-Internal error

MCP passes 400, 403, 404, 409 and 429 through as readable text an agent can act on. Everything else becomes the opaque Internal error - the detail is logged server-side, not returned. So an agent seeing Internal error should surface it rather than retry in a loop.

There Is No Existence Oracle

A foreign id behaves exactly like a missing one:

  • Operations behind the chat-access guard answer 403 uniformly for foreign, missing and malformed chat ids.
  • Operations reading org-scoped data return null or an empty page.
  • MCP returns a not-found text for any id outside the target org.

You cannot use error responses to discover whether something exists in another organisation.

Webhook Receivers

Delivery to your endpoint has its own rules:

RuleDetail
Scheme and hostPublic https:// only. An org-own webhook URL is validated when you save it.
RedirectsNot followed. A receiver answering 301 or 308 fails permanently.
Timeout5 seconds.
Retries429 and 5xx are retried, honouring your Retry-After.
ExpiryEvents older than 24 hours are dropped.
SubscriptionsAt most 10 per provider.

The redirect rule catches people out. If your endpoint is behind something that normalises /webhook to /webhook/ with a 301, every delivery fails - and it fails permanently, not as a retryable error. Point the subscription at the final URL.

See Envelope & delivery for the retry and deduplication contract.

Designing Within the Limits

  • Bulk sending: use async: true so each send returns immediately, and watch the write bucket rather than wall-clock time.
  • Mirroring an inbox: take the initial state from paged reads, then keep up with webhooks instead of polling. Polling expensive reads is what exhausts the extra bucket.
  • Backfills: page with cursors at 100 per call and pace yourself against the read bucket.
  • Many clients: remember the per-org write limit applies to each client separately, so work spread across clients scales better than the same volume aimed at one.

On this page