Streaming
SSE across the gateway's three dialects — how to turn a stream on, which events arrive, what marks a whole answer, and why a call is never moved after the first byte.
Quick
Add "stream": true to the same call and the answer arrives as text/event-stream. On Chat Completions, ask for the usage too: stream_options.include_usage.
curl -N https://api.kumorouter.com/v1/chat/completions \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"max_tokens": 128,
"stream": true,
"stream_options": { "include_usage": true },
"messages": [{ "role": "user", "content": "Write a haiku about latency." }]
}'A stream without its terminal frame is a failed call
A refusal decided before the first event arrives as an ordinary JSON error with a status — the same envelope a unary call would get, and nothing streamed at all. A failure after the first event cannot be that: the status line has already gone out, and it cannot be rewritten.
So each dialect names its terminal frame, and a whole answer is marked by that frame and nothing else.
| Dialect | Terminal frame | If it never arrives |
|---|---|---|
| Chat Completions | [DONE] | The stream simply ended. The answer is truncated. |
| Anthropic Messages | message_stop | This protocol's error event arrived, and the stream ended without message_stop. |
| Responses | response.completed | response.failed arrived, carrying the error envelope in place of the finished body. |
Read for that absence. A client that treats "the connection closed" as "the answer finished" will hand truncated text to whatever comes next, and nothing in the bytes it received says otherwise. A stream that ends before its terminal frame is a failed call, not a short one.
A practical rule: keep a "terminal frame seen" flag and check it first when the loop ends, before you pass the text on.
What arrives in the stream
Chat Completions
chat.completion.chunk events in arrival order. Each carries the same id, model and object a unary reply would, and its content sits in choices[0].delta: role on the first chunk, then content, refusal or tool_calls in fragments. finish_reason is stated once, on the chunk that closes the answer, and is null until then.
When stream_options.include_usage asks for it, a usage chunk arrives before [DONE]: an event with an empty choices array carrying only usage. It does not change what the platform settles from — usage is always collected — it only decides whether your own stream shows it.
data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Latency"},"finish_reason":null}]}
data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":18,"total_tokens":30}}
data: [DONE]Anthropic Messages
Six named events in order: message_start, the opening envelope with the answer's identity and the input half of the account; content_block_start, content_block_delta and content_block_stop for each block; message_delta, the closing facts — stop_reason and the settled usage; and message_stop. Alongside them arrive ping and error.
A block delta comes in two kinds: text_delta carries a fragment of text, and input_json_delta carries a fragment of a tool call's arguments as partial JSON in partial_json.
Responses
This protocol's own named events: response.created, response.output_item.added, response.output_text.delta, response.refusal.delta, response.function_call_arguments.delta, response.output_item.done, and exactly one terminal event — response.completed or response.failed. The event name is repeated in the SSE frame's event: field, so a client may dispatch on either.
Usage arrives on response.completed. response.failed carries this protocol's error envelope in its error member.
A stream on the Messages dialect
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.kumorouter.com",
api_key=os.environ["KUMO_API_KEY"],
)
with client.messages.stream(
model="<model>",
max_tokens=256,
messages=[{"role": "user", "content": "Write a haiku about latency."}],
) as stream:
for text in stream.text_stream:
print(text, end="", flush=True)The first byte decides everything
A model is usually reachable through more than one channel, and a channel failing before any of the answer has reached you moves the call to the next channel for the same model.
After the first byte, there is no move and there cannot be one: you cannot be handed a second answer halfway through one you are already reading. All of the stream's behavior follows from that — a cut instead of an error, a missing frame instead of an envelope.
Cancellation
Closing the connection cancels the call: the platform sees the disconnect and records it as your withdrawal. A canceled call is not moved to another channel — you are no longer waiting for it.
Usage the supplier had already confirmed is still settled: your disconnect racing the result does not make the work free. Where there is nothing confirmed to settle, the call releases what it held.
Where there is no streaming
- Embeddings do not stream: that surface answers with whole vectors.
- Image generation has no
streammember at all. - Counting tokens calls no provider, so it has no stream to give.
Streaming is a proven capability of an exact model and channel, not a property of the protocol. A request with stream: true will not be routed to a pair that has not proven streaming: it is refused before anything upstream is contacted. What a given model has proven is on its card in the catalog.