---
title: Streaming
description: SSE across the gateway's three dialects — how to turn a stream on, which events arrive, what marks a whole answer, and why a call is never moved after the first byte.
keywords: streaming, sse, server-sent events, done, message_stop, response.completed, cancellation
group: gateway
---

## Quick {#quick keywords="stream, include_usage, curl, python, node"}

Add `"stream": true` to the same call and the answer arrives as `text/event-stream`. On Chat Completions, ask for the usage too: `stream_options.include_usage`.

:::code-group
```bash title=curl
curl -N https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "max_tokens": 128,
    "stream": true,
    "stream_options": { "include_usage": true },
    "messages": [{ "role": "user", "content": "Write a haiku about latency." }]
  }'
```
```python title=Python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.kumorouter.com/v1",
    api_key=os.environ["KUMO_API_KEY"],
)

stream = client.chat.completions.create(
    model="<model>",
    max_tokens=128,
    stream=True,
    stream_options={"include_usage": True},
    messages=[{"role": "user", "content": "Write a haiku about latency."}],
)
for chunk in stream:
    for choice in chunk.choices:
        print(choice.delta.content or "", end="", flush=True)
```
```javascript title=Node
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.kumorouter.com/v1",
  apiKey: process.env.KUMO_API_KEY,
});

const stream = await client.chat.completions.create({
  model: "<model>",
  max_tokens: 128,
  stream: true,
  stream_options: { include_usage: true },
  messages: [{ role: "user", content: "Write a haiku about latency." }],
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
```
:::

## A stream without its terminal frame is a failed call {#terminal keywords="done, message_stop, response.failed, truncated"}

A refusal decided **before** the first event arrives as an ordinary JSON error with a status — the same envelope a unary call would get, and nothing streamed at all. A failure **after** the first event cannot be that: the status line has already gone out, and it cannot be rewritten.

So each dialect names its terminal frame, and a whole answer is marked by that frame and nothing else.

:::matrix
| Dialect | Terminal frame | If it never arrives |
| --- | --- | --- |
| Chat Completions | `[DONE]` | The stream simply ended. The answer is truncated. |
| Anthropic Messages | `message_stop` | This protocol's `error` event arrived, and the stream ended without `message_stop`. |
| Responses | `response.completed` | `response.failed` arrived, carrying the error envelope in place of the finished body. |
:::

Read for that absence. A client that treats "the connection closed" as "the answer finished" will hand truncated text to whatever comes next, and nothing in the bytes it received says otherwise. A stream that ends before its terminal frame is a failed call, not a short one.

A practical rule: keep a "terminal frame seen" flag and check it first when the loop ends, before you pass the text on.

## What arrives in the stream {#events keywords="events, chunks, deltas, usage"}

### Chat Completions

`chat.completion.chunk` events in arrival order. Each carries the same `id`, `model` and `object` a unary reply would, and its content sits in `choices[0].delta`: `role` on the first chunk, then `content`, `refusal` or `tool_calls` in fragments. `finish_reason` is stated once, on the chunk that closes the answer, and is `null` until then.

When `stream_options.include_usage` asks for it, a **usage chunk** arrives before `[DONE]`: an event with an empty `choices` array carrying only `usage`. It does not change what the platform settles from — usage is always collected — it only decides whether your own stream shows it.

```text title=Frames
data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}

data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{"content":"Latency"},"finish_reason":null}]}

data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-8f2b7e10c9","object":"chat.completion.chunk","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":18,"total_tokens":30}}

data: [DONE]
```

### Anthropic Messages

Six named events in order: `message_start`, the opening envelope with the answer's identity and the input half of the account; `content_block_start`, `content_block_delta` and `content_block_stop` for each block; `message_delta`, the closing facts — `stop_reason` and the settled usage; and `message_stop`. Alongside them arrive `ping` and `error`.

A block delta comes in two kinds: `text_delta` carries a fragment of text, and `input_json_delta` carries a fragment of a tool call's arguments as partial JSON in `partial_json`.

### Responses

This protocol's own named events: `response.created`, `response.output_item.added`, `response.output_text.delta`, `response.refusal.delta`, `response.function_call_arguments.delta`, `response.output_item.done`, and exactly one terminal event — `response.completed` or `response.failed`. The event name is repeated in the SSE frame's `event:` field, so a client may dispatch on either.

Usage arrives on `response.completed`. `response.failed` carries this protocol's error envelope in its `error` member.

## A stream on the Messages dialect {#messages-example keywords="anthropic, sdk, stream, example"}

```python title=Python
import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.kumorouter.com",
    api_key=os.environ["KUMO_API_KEY"],
)

with client.messages.stream(
    model="<model>",
    max_tokens=256,
    messages=[{"role": "user", "content": "Write a haiku about latency."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
```

## The first byte decides everything {#first-byte keywords="failover, channel, first byte, retry"}

A model is usually reachable through more than one channel, and a channel failing **before** any of the answer has reached you moves the call to the next channel for the same model.

After the first byte, there is no move and there cannot be one: you cannot be handed a second answer halfway through one you are already reading. All of the stream's behavior follows from that — a cut instead of an error, a missing frame instead of an envelope.

## Cancellation {#cancellation keywords="cancel, disconnect, reservation"}

Closing the connection cancels the call: the platform sees the disconnect and records it as your withdrawal. A canceled call is not moved to another channel — you are no longer waiting for it.

Usage the supplier had already confirmed is still settled: your disconnect racing the result does not make the work free. Where there is nothing confirmed to settle, the call releases what it held.

## Where there is no streaming {#no-stream keywords="embeddings, images, count_tokens"}

- Embeddings do not stream: that surface answers with whole vectors.
- Image generation has no `stream` member at all.
- Counting tokens calls no provider, so it has no stream to give.

:::note
Streaming is a proven capability of an exact model and channel, not a property of the protocol. A request with `stream: true` will not be routed to a pair that has not proven streaming: it is refused before anything upstream is contacted. What a given model has proven is on its card in the catalog.
:::

## Next {#next keywords="errors, chat completions, messages, responses"}

:::cards
- [Errors](/en/errors) — the refusal envelope, the status codes and failover between channels.
- [Chat Completions](/en/chat-completions) — the unary call of the same protocol.
- [Messages](/en/messages) — the Anthropic dialect in full.
- [Responses](/en/responses) — typed output items.
- [Model catalog](/en/models) — what each model has proven.
:::
