Skip to contentKumoDocs
Sections
On this page
The gateway

Rate limits

The ceilings a call is measured against, what a call that reaches one gets back, and how a higher ceiling is asked for.

View as Markdown

What a new key carries

A key is issued with three ceilings, and a call has to be under every one of them. Nothing is switched on to get them, nothing in a request asks for them, and a key that has just been minted already has them. Your organization carries the two rate ceilings too, and its are higher; the third — how many streams the key holds open at once — belongs to the key alone:

RequestsTokensConcurrent streams
Key20 per second500,000 per minute16
Organization200 per second5,000,000 per minutenone of its own

A call is measured against the organization's ceilings first, then against the key's. An organization's default is ten keys' worth, which is where its rate ceilings come from; issuing an eleventh key does not raise them.

The organization's budget is shared, not a sum of what the keys are allowed: a call can be well under both of its own key's ceilings and still be refused for either of the organization's, on traffic other keys generated — a quiet key can be turned away for the organization's request ceiling as readily as for its token one.

The ceilings are counted separately and on different clocks — one over a second, the other over a minute — so a workload can sit comfortably inside the request ceiling and still be over the token one. A handful of very large calls every second is exactly the shape that does it, and it is the shape that surprises people, because the request count looks fine.

How many streams are open at once

The rate ceilings count how often calls arrive. The stream ceiling counts how many are open right now: a stream takes a slot from the moment it is admitted until it ends, is abandoned by the client, or expires. A key's seventeenth simultaneous stream is refused with a 429, and the sixteen already running carry on.

Only streams take a slot. An ordinary non-streaming call takes none and frees none, however long it runs, and counting tokens takes nothing at all.

Sixteen is chosen for a client that keeps several calls in flight by itself: an agent running subagents or parallel tool calls reaches about ten requests on one key, and the ceiling sits above that rather than against it. It sits below the request ceiling so that the whole of it can be taken at once: sixteen streams are opened by sixteen requests, and a client with no other traffic fills them inside a single second. Nothing in the arithmetic forbids more — concurrency accumulates across seconds — but a stream ceiling above the request ceiling would be filled in instalments rather than in a burst, and the rate ceiling would refuse on the way.

The stream ceiling a key is issued with belongs to the key and not to the organization: two keys of one organization each hold sixteen streams and do not take them from each other. The organization carries no stream default, unlike both rate ceilings, which it does.

A configured organization-wide stream ceiling does exist as a possibility: if one was set for you it is shared by every key and is checked before the key's own. It never appears by itself and nothing above assumes one — but where there is one, a second key does not get around it, and it is raised separately.

Which is also why one key shared between several agents is not the same as a key per agent: the sixteen slots are divided among everything using that key. A client that runs out of them gets a second key more cheaply than a raise — unless its organization carries a stream ceiling of its own, which the keys share.

When a ceiling is reached

A call that would cross any of the ceilings is refused rather than queued: the status is 429, and the envelope's type is rate_limit_error. The refusal is decided before the call reaches a provider, so nothing upstream sees it. It names the ceiling the call crossed, that ceiling's value, the room left under it, and the instant before which retrying is pointless.

The request ceiling can be counted as calls arrive. The token ceiling cannot, because what a call will cost is not known until it has been answered.

So the platform estimates the call's tokens before it goes upstream and counts the estimate against the minute, then reconciles that estimate against the usage the answer actually reports.

When that reconciliation lands, the minute carries the real figure rather than the guess, and a run of calls that came in far under their estimates leaves room behind it.

Two cases end differently, and both are deliberate:

  • The minute has already rolled over — The correction is aimed at the exact minute the call was charged to, so once a later call has carried a counter into a new minute there is nothing left to correct. The counter is no longer on that minute. Time passing does not do this; a further call does. And the two counters move on independently, so a correction can still land on one of them and find nothing on the other.
  • Images — The estimate really does stand: an operation that reports no tokens at all keeps what it was charged. Image generation is priced in images, not tokens, so correcting it to zero would refund the whole estimate and let that surface cost nothing against the ceiling.

Streaming does not change the arithmetic: a stream that has begun is a call that has been counted. Counting tokens does not enter it at all — that operation calls no provider and consumes no quota, which is what makes it safe to run in front of a large request.

Back off before retrying a 429, lengthen the wait with every attempt, and cap the number of attempts.

What a refusal looks like →

Sizing a call in advance

The token ceiling runs into the fact that what a call costs is not known until it has been answered. One half of that can be learned before you send: counting tokens on the Messages dialect takes a subset of an ordinary call's members — model, messages, system, tools, tool_choice, thinking, metadata and context_management — and answers with a deterministic local estimate of the input tokens. The members that describe the generation rather than the input are refused on this operation: max_tokens, stream, output_config, stop_sequences, temperature and top_p.

It calls no provider, creates no request and no reservation, touches no balance and spends no rate quota. So you can put it in front of a large request without paying for it in money or in room under the ceiling.

It does not count the output half and cannot: that half is bounded by you, through the output ceiling the request names.

How to call it →

Asking for more

A raise is a conversation, not a setting. None of the three ceilings is editable in the console, and no member of a request asks for a higher one. Write to support with the traffic you need admitted — the shape of it, not only the peak — and the ceilings on your key are changed for you.