---
title: Responses
description: The POST /v1/responses call — how it differs from Chat Completions, the request members, the typed output items and the stream's terminal event.
keywords: responses, input, output, instructions, response.completed, response.failed
group: gateway
---

## Quick {#quick keywords="curl, python, node, first call"}

The same key and the same base URL as every other gateway operation. The conversation travels in `input`, and the system prompt in `instructions`.

:::code-group
```bash title=curl
curl https://api.kumorouter.com/v1/responses \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "instructions": "Answer in one sentence.",
    "input": [
      {
        "type": "message",
        "role": "user",
        "content": [{ "type": "input_text", "text": "Explain tokens." }]
      }
    ]
  }'
```
```python title=Python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.kumorouter.com/v1",
    api_key=os.environ["KUMO_API_KEY"],
)

answer = client.responses.create(
    model="<model>",
    instructions="Answer in one sentence.",
    input=[
        {
            "type": "message",
            "role": "user",
            "content": [{"type": "input_text", "text": "Explain tokens."}],
        }
    ],
)
print(answer.output[0].content[0].text)
```
```javascript title=Node
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.kumorouter.com/v1",
  apiKey: process.env.KUMO_API_KEY,
});

const answer = await client.responses.create({
  model: "<model>",
  instructions: "Answer in one sentence.",
  input: [
    {
      type: "message",
      role: "user",
      content: [{ type: "input_text", text: "Explain tokens." }],
    },
  ],
});

const reply = answer.output[0].content[0].text;
process.stdout.write((reply ?? "") + "\n");
```
:::

## How it differs {#differences keywords="differences, input, output, status, instructions"}

This is its own protocol, not another spelling of Chat Completions. Five differences show up immediately.

- The conversation is `input`: one ordered list of typed items. A turn, a tool call and a tool result are three kinds of item rather than members of one message.
- The system prompt is `instructions`, a member of the request. It becomes the conversation's leading system turn, ahead of every item of `input`.
- The answer is `output`: a list of items, not `choices`.
- The answer states a `status` — a lifecycle stage, `completed` or `incomplete` — rather than a finish reason. The two vocabularies are not translations of each other.
- The output ceiling is optional: a request stating no `max_output_tokens` is answered under this surface's published default of 32768 tokens. An explicit zero is refused: it is a request for no output at all.

Some members of the protocol are **accepted and not forwarded**, and each says so itself: `parallel_tool_calls`, `reasoning`, `prompt_cache_key`, `client_metadata` and `text.verbosity`. They are neither refused nor silently honored — they do nothing. `include` is half of that: the values this surface publishes an answer for are forwarded, and the rest do nothing.

:::note
`store` is accepted only as `false`, and its absence means `false`. This platform retains no prompt, response or chunk, so there is nothing here to keep and hand back, and `true` is refused rather than accepted and ignored. For the same reason `previous_response_id` is validated as belonging to your organization but is not honored as server-side context: there is nothing to continue from, and the caller sends its own context.
:::

## Request members {#request keywords="members, input, instructions, max_output_tokens, text, tools"}

| Member | What it is |
| --- | --- |
| `model` | Required. The model to answer with: a canonical name or an alias from the published catalog. |
| `input` | Required. The conversation in order, as typed items; from one up to 512. |
| `instructions` | The system prompt. It becomes the conversation's leading system turn. |
| `max_output_tokens` | The output ceiling. Without it the surface's default of 32768 tokens applies; an explicit zero is refused. |
| `stream` | Whether to answer as a stream of this protocol's own named events. |
| `text` | How the answer's text is shaped. A structured output requirement lives here too, under `text.format`. |
| `tools` | This turn's tools. They are declared flat: `type`, `name`, `parameters`. |
| `tool_choice` | What this request requires of its tool list: the string `auto`, `none` or `required`, or an object naming one declared tool. |
| `store` | Accepted only as `false`. |
| `previous_response_id` | A response of your own organization that this request follows. It is validated for ownership and is not server-side context. |
| `parallel_tool_calls` | Accepted and not forwarded. |
| `reasoning` | Accepted and not forwarded. |
| `include` | Forwarded for the values this surface publishes an answer for — today exactly `reasoning.encrypted_content`, which is what makes the supplier mint the encrypted record on the `reasoning` items in your answer. Send those items back in `input` on your next turn and the model continues from an account it signed itself; without the record it does not. Every other value is accepted and does nothing: `include` asks the supplier to add items to its answer, and an item this surface does not publish is one it refuses as a break of the supplier's contract — so forwarding such a value would turn a request that works into one that fails. The member itself is never refused, because the clients that send it send it on every request. |
| `prompt_cache_key` | Accepted and not forwarded. |
| `client_metadata` | Accepted and not forwarded: this platform does not describe your sessions to a supplier. |

### Kinds of input item

| `type` | What it carries |
| --- | --- |
| `message` | A turn: `role` (`system`, `user`, `developer`, `assistant`) and `content` in `input_text` or `output_text` parts. A `developer` turn is carried as a system turn: one role under two names. |
| `function_call` | A call the model made earlier: `call_id`, `name` and `arguments` as JSON text, plus the `namespace` it was declared in when the request declared its tools in groups. |
| `function_call_output` | The result of running a tool: `call_id` and `output` as text. |
| `custom_tool_call` | A call the model made to a tool declared `custom`: `call_id`, `name`, the optional `namespace`, and `input` — the model's answer in the tool's own language rather than JSON. An empty `input` is a legal answer and is sent as an empty string. |
| `custom_tool_call_output` | The result of running that one: `call_id` and `output`. A result answers a call of its own kind: a `custom_tool_call_output` answers a `custom_tool_call` and a `function_call_output` answers a `function_call`. |
| `reasoning` | The model's own account of an earlier turn, replayed: `id`, the required `summary` of `summary_text` parts, the optional `content` of `reasoning_text` parts, `encrypted_content` and `status`. Nothing here is read by this platform; it is carried to the supplier that produced it. |
| `web_search_call` | A search the supplier already performed, replayed: `id`, `status` and the optional `action`. Nobody answers it. |
| `additional_tools` | Tool declarations grouped into namespaces, which is where the newest clients of this protocol put them instead of in `tools`. Admitted as the first item of the input and only once. |

### The second turn

This platform stores no response, so your client carries its own history — and the history it has is what the answer gave it. **Every item this surface answers with is admitted back in `input`**, in the members it was published under, so replaying an answer verbatim is a request this surface serves. The one exception is a `refusal` content part: a refusal ends the exchange it belongs to, and this surface does not take one back as a turn.

## The response {#response keywords="output, status, usage, message, function_call"}

| Reply member | What it is |
| --- | --- |
| `id` | This request's Kumo identity — the same one a `previous_response_id` names. |
| `object` | Always `response`. |
| `created_at` | The instant the response was answered, as whole seconds since the Unix epoch. |
| `model` | The name the request sent, echoed back exactly. |
| `status` | `completed` or `incomplete`. |
| `output` | What the model produced, in order: `message`, `function_call`, `custom_tool_call`, `web_search_call` and `reasoning` items. |
| `usage` | `input_tokens`, `output_tokens`, `total_tokens`. Absent when the supplier reported no usage at all. |

A `message` item's parts come in two kinds: `output_text`, the answer's text, and `refusal`, the model's own refusal to answer.

A `reasoning` item is the model's own account of how it reached the answer, and a reasoning model states one first, ahead of the message it explains. It carries `id`, the required `summary` in `summary_text` parts — present and empty when the model summarised nothing — the optional `content` in `reasoning_text` parts, `encrypted_content` when you asked for it under `include`, and `status`. Nothing in it is read by this platform and nothing in it is stored. **Send it back in `input` on your next turn**, in the members it was published under: that is what this protocol asks a client managing its own context to do, and it is what lets a reasoning model continue where it left off.

```json title=Response
{
  "id": "resp-4d19a0b7c2",
  "object": "response",
  "created_at": 1756900000,
  "model": "<model>",
  "status": "completed",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Tokens are the pieces of text a model measures input and output in." }]
    }
  ],
  "usage": { "input_tokens": 14, "output_tokens": 21, "total_tokens": 35 }
}
```

## Streaming {#stream keywords="stream, response.created, response.completed, response.failed"}

`"stream": true` gives you `text/event-stream` in this protocol's own named events: `response.created` first, then the output items and their deltas, then **exactly one** terminal event — `response.completed` carrying the finished body and usage, or `response.failed` carrying this protocol's error envelope.

> [Every event and the terminal-frame rule →](/en/streaming)

## Next {#next keywords="tools, structured output, errors, reference"}

:::cards
- [Streaming](/en/streaming) — the events of all three dialects and what a whole answer ends with.
- [Tools](/en/tools) — the flat declaration, `function_call` and `function_call_output`.
- [Structured output](/en/structured-output) — `text.format` with a named strict schema.
- [Errors](/en/errors) — the refusal envelope and the status codes.
- [API reference](/en/api-reference) — the operation member by member, straight from the API description.
:::
