Skip to contentKumoDocs
Sections
On this page
The gateway

Responses

The POST /v1/responses call — how it differs from Chat Completions, the request members, the typed output items and the stream's terminal event.

View as Markdown

Quick

The same key and the same base URL as every other gateway operation. The conversation travels in input, and the system prompt in instructions.

curl https://api.kumorouter.com/v1/responses \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "instructions": "Answer in one sentence.",
    "input": [
      {
        "type": "message",
        "role": "user",
        "content": [{ "type": "input_text", "text": "Explain tokens." }]
      }
    ]
  }'

How it differs

This is its own protocol, not another spelling of Chat Completions. Five differences show up immediately.

  • The conversation is input: one ordered list of typed items. A turn, a tool call and a tool result are three kinds of item rather than members of one message.
  • The system prompt is instructions, a member of the request. It becomes the conversation's leading system turn, ahead of every item of input.
  • The answer is output: a list of items, not choices.
  • The answer states a status — a lifecycle stage, completed or incomplete — rather than a finish reason. The two vocabularies are not translations of each other.
  • The output ceiling is optional: a request stating no max_output_tokens is answered under this surface's published default of 32768 tokens. An explicit zero is refused: it is a request for no output at all.

Some members of the protocol are accepted and not forwarded, and each says so itself: parallel_tool_calls, reasoning, prompt_cache_key, client_metadata and text.verbosity. They are neither refused nor silently honored — they do nothing. include is half of that: the values this surface publishes an answer for are forwarded, and the rest do nothing.

Note

store is accepted only as false, and its absence means false. This platform retains no prompt, response or chunk, so there is nothing here to keep and hand back, and true is refused rather than accepted and ignored. For the same reason previous_response_id is validated as belonging to your organization but is not honored as server-side context: there is nothing to continue from, and the caller sends its own context.

Request members

modelRequired. The model to answer with: a canonical name or an alias from the published catalog.
inputRequired. The conversation in order, as typed items; from one up to 512.
instructionsThe system prompt. It becomes the conversation's leading system turn.
max_output_tokensThe output ceiling. Without it the surface's default of 32768 tokens applies; an explicit zero is refused.
streamWhether to answer as a stream of this protocol's own named events.
textHow the answer's text is shaped. A structured output requirement lives here too, under `text.format`.
toolsThis turn's tools. They are declared flat: `type`, `name`, `parameters`.
tool_choiceWhat this request requires of its tool list: the string `auto`, `none` or `required`, or an object naming one declared tool.
storeAccepted only as `false`.
previous_response_idA response of your own organization that this request follows. It is validated for ownership and is not server-side context.
parallel_tool_callsAccepted and not forwarded.
reasoningAccepted and not forwarded.
includeForwarded for the values this surface publishes an answer for — today exactly `reasoning.encrypted_content`, which is what makes the supplier mint the encrypted record on the `reasoning` items in your answer. Send those items back in `input` on your next turn and the model continues from an account it signed itself; without the record it does not. Every other value is accepted and does nothing: `include` asks the supplier to add items to its answer, and an item this surface does not publish is one it refuses as a break of the supplier's contract — so forwarding such a value would turn a request that works into one that fails. The member itself is never refused, because the clients that send it send it on every request.
prompt_cache_keyAccepted and not forwarded.
client_metadataAccepted and not forwarded: this platform does not describe your sessions to a supplier.

Kinds of input item

messageA turn: `role` (`system`, `user`, `developer`, `assistant`) and `content` in `input_text` or `output_text` parts. A `developer` turn is carried as a system turn: one role under two names.
function_callA call the model made earlier: `call_id`, `name` and `arguments` as JSON text, plus the `namespace` it was declared in when the request declared its tools in groups.
function_call_outputThe result of running a tool: `call_id` and `output` as text.
custom_tool_callA call the model made to a tool declared `custom`: `call_id`, `name`, the optional `namespace`, and `input` — the model's answer in the tool's own language rather than JSON. An empty `input` is a legal answer and is sent as an empty string.
custom_tool_call_outputThe result of running that one: `call_id` and `output`. A result answers a call of its own kind: a `custom_tool_call_output` answers a `custom_tool_call` and a `function_call_output` answers a `function_call`.
reasoningThe model's own account of an earlier turn, replayed: `id`, the required `summary` of `summary_text` parts, the optional `content` of `reasoning_text` parts, `encrypted_content` and `status`. Nothing here is read by this platform; it is carried to the supplier that produced it.
web_search_callA search the supplier already performed, replayed: `id`, `status` and the optional `action`. Nobody answers it.
additional_toolsTool declarations grouped into namespaces, which is where the newest clients of this protocol put them instead of in `tools`. Admitted as the first item of the input and only once.

The second turn

This platform stores no response, so your client carries its own history — and the history it has is what the answer gave it. Every item this surface answers with is admitted back in `input`, in the members it was published under, so replaying an answer verbatim is a request this surface serves. The one exception is a refusal content part: a refusal ends the exchange it belongs to, and this surface does not take one back as a turn.

The response

idThis request's Kumo identity — the same one a `previous_response_id` names.
objectAlways `response`.
created_atThe instant the response was answered, as whole seconds since the Unix epoch.
modelThe name the request sent, echoed back exactly.
status`completed` or `incomplete`.
outputWhat the model produced, in order: `message`, `function_call`, `custom_tool_call`, `web_search_call` and `reasoning` items.
usage`input_tokens`, `output_tokens`, `total_tokens`. Absent when the supplier reported no usage at all.

A message item's parts come in two kinds: output_text, the answer's text, and refusal, the model's own refusal to answer.

A reasoning item is the model's own account of how it reached the answer, and a reasoning model states one first, ahead of the message it explains. It carries id, the required summary in summary_text parts — present and empty when the model summarised nothing — the optional content in reasoning_text parts, encrypted_content when you asked for it under include, and status. Nothing in it is read by this platform and nothing in it is stored. Send it back in `input` on your next turn, in the members it was published under: that is what this protocol asks a client managing its own context to do, and it is what lets a reasoning model continue where it left off.

{
  "id": "resp-4d19a0b7c2",
  "object": "response",
  "created_at": 1756900000,
  "model": "<model>",
  "status": "completed",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "Tokens are the pieces of text a model measures input and output in." }]
    }
  ],
  "usage": { "input_tokens": 14, "output_tokens": 21, "total_tokens": 35 }
}

Streaming

"stream": true gives you text/event-stream in this protocol's own named events: response.created first, then the output items and their deltas, then exactly one terminal event — response.completed carrying the finished body and usage, or response.failed carrying this protocol's error envelope.

Every event and the terminal-frame rule →

Next