---
title: API reference
description: Every operation of the Kumo gateway, generated from the published API description: the address, how it is authenticated, what the request carries and what comes back.
keywords: api, reference, endpoints, operations, openapi
group: gateway
generated: from the published API description this build was compiled against
---

This page is generated from the API description this build was compiled against, so it says what the gateway answers rather than what it was once documented to answer. Every member below is a member of the wire.

## POST /v1/chat/completions {#chat-completions-create keywords="Create a chat completion."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/chat/completions |
| Auth | Authorization: Bearer <key> |
| Operation | public.chat_completions.create |
:::

The OpenAI-compatible chat surface: send a conversation, get one completion back or a stream of chunks.

Create a chat completion.

Answers one Chat Completions request against a model of the published catalog, authenticated by a Kumo API key. It is this protocol served natively and not a translation of another: the request members, the response shape, the usage vocabulary and the error envelope are this protocol's own.

A stream=true request is answered as text/event-stream — chat.completion.chunk events, a usage chunk when stream_options.include_usage asks for one, then the terminal [DONE] frame; a refusal decided before the first event is answered as this protocol's ordinary JSON error, and a failure after it ends the stream without its terminal frame.

An output ceiling is OPTIONAL — max_tokens or max_completion_tokens, interchangeably — and a request stating neither is answered under this surface's published default of 32768 output tokens, which is the finite bound the reservation is taken against before the upstream call. Stating both with different values is refused rather than resolved by a precedence rule.

:::code-group
```json title=Request
{
  "max_tokens": 128,
  "messages": [
    {
      "role": "user"
    }
  ],
  "model": "<model>"
}
```
```json title=Response
{
  "choices": [
    {
      "finish_reason": "stop",
      "index": 1,
      "message": {
        "content": "<content>",
        "role": "assistant"
      }
    }
  ],
  "created": 1,
  "id": "<id>",
  "model": "<model>",
  "object": "chat.completion"
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `max_completion_tokens` | integer, optional — The output ceiling, in the current spelling. |
| `max_tokens` | integer, optional — The output ceiling, in the spelling long-established clients send. |
| `messages` | array of [ChatMessage](#schema-chatmessage), required — The conversation, in order. |
| `model` | string, required — The model to answer with: a canonical name or an alias the published catalog carries. |
| `n` | integer, optional — How many completions to answer with. |
| `reasoning_effort` | string, optional — Accepted and NOT CARRIED: this states how much reasoning to spend on the answer, as OpenAI-compatible clients fill it in for reasoning models, but no supplier request on this platform has a member for a reasoning budget, so the answer comes back at whatever budget the model itself defaults to. |
| `response_format` | [ChatResponseFormat](#schema-chatresponseformat), optional — A structured output requirement. |
| `stop` | array of string, optional — Sequences whose appearance ends the answer, at most four, as a LIST — the bare-string spelling this protocol also defines is not accepted on this surface and is refused rather than ignored. |
| `stream` | boolean, optional — Whether to stream the answer. |
| `stream_options` | [ChatStreamOptions](#schema-chatstreamoptions), optional — Options that apply only when stream is true. |
| `temperature` | number, optional — How much randomness to use, from 0 to 2 — this protocol's own interval, and not the [0, 1] of the Anthropic Messages surface. |
| `tool_choice` | "auto" or "none" or "required" or object, optional — What this request requires of its tool list, in either spelling this protocol defines: the bare mode word "auto", "none" or "required", or an object naming one declared tool. |
| `tools` | array of [ChatTool](#schema-chattool), optional — The tools this turn may call. |
| `top_p` | number, optional — Nucleus sampling, from 0 to 1. |
| `user` | string, optional — An opaque label for the end user this request is made on behalf of, as OpenAI-compatible clients send it. |
:::

### ChatMessage {#schema-chatmessage}

:::matrix
| member | what it is |
| --- | --- |
| `cache_control` | ChatCacheControl, optional — Accepted and NOT CARRIED: this marks the turn as a prompt-cache anchor, the way OpenAI-compatible clients write it when they address an Anthropic model, but the supplier wire this protocol is served by has no member for an anchor, so the cached-token counts in usage come back as though it had not been sent. |
| `content` | string or array of ChatContentPart, optional — The turn's text, in either spelling this protocol defines: a bare string, or a list of typed parts joined in order. |
| `name` | string, optional — The tool this result came from, as OpenAI-compatible clients write it on a tool turn. |
| `role` | "system" or "developer" or "user" or "assistant" or "tool", required — Who is speaking. |
| `tool_call_id` | string, optional — The call this result answers. |
| `tool_calls` | array of ChatToolCall, optional — The calls this assistant turn made, replayed back into the conversation so that a tool result has something in the history to answer. |
:::

### ChatResponseFormat {#schema-chatresponseformat}

:::matrix
| member | what it is |
| --- | --- |
| `json_schema` | StructuredSchema, optional — The named schema an answer must satisfy. |
| `type` | string, required — The kind of format. |
:::

### ChatStreamOptions {#schema-chatstreamoptions}

:::matrix
| member | what it is |
| --- | --- |
| `include_usage` | boolean, optional — Whether the stream ends with a usage chunk — an event with an empty choices array carrying only usage — before the terminal [DONE] frame. |
:::

### ChatTool {#schema-chattool}

:::matrix
| member | what it is |
| --- | --- |
| `function` | ChatToolFunction, required |
| `type` | "function", required — The kind of tool. |
:::

The answer:

| what | shape |
| --- | --- |
| `application/json` | ChatCompletionsReply |
| `text/event-stream` | ChatCompletionsChunk |
| `on failure` | ChatCompletionsError |

### ChatCompletionsReply {#schema-chatcompletionsreply}

:::matrix
| member | what it is |
| --- | --- |
| `choices` | array of [ReplyChoice](#schema-replychoice), required — The answer. |
| `created` | integer, required — The instant this completion was answered, as whole seconds since the Unix epoch. |
| `id` | string, required — This request's Kumo identity. |
| `model` | string, required — The model this request named, echoed back exactly as it was sent. |
| `object` | "chat.completion", required — The kind of object this is. |
| `usage` | [ReplyUsage](#schema-replyusage), optional — What the supplier reported this request consumed. |
:::

### ReplyChoice {#schema-replychoice}

:::matrix
| member | what it is |
| --- | --- |
| `finish_reason` | "stop" or "length" or "tool_calls" or "content_filter", required — Why the model stopped. |
| `index` | integer, required — The position of this choice. |
| `message` | ReplyMessage, required |
:::

### ReplyUsage {#schema-replyusage}

:::matrix
| member | what it is |
| --- | --- |
| `completion_tokens` | integer, optional — Output tokens the supplier counted. |
| `prompt_tokens` | integer, optional — Input tokens the supplier counted, including any it served from its own cache. |
| `prompt_tokens_details` | PromptTokensDetails, optional — How the input divides, when the supplier said. |
| `total_tokens` | integer, optional — Input and output together. |
:::

### ChatCompletionsChunk {#schema-chatcompletionschunk}

:::matrix
| member | what it is |
| --- | --- |
| `choices` | array of [ChunkChoice](#schema-chunkchoice), required — This chunk's delta. |
| `created` | integer, required — The instant this completion was answered, as whole seconds since the Unix epoch. |
| `id` | string, required — This request's Kumo identity, identical on every chunk of one stream. |
| `model` | string, required — The model this request named, echoed back exactly as it was sent, on every chunk. |
| `object` | "chat.completion.chunk", required — The kind of object this is. |
| `usage` | [ReplyUsage](#schema-replyusage), optional — Present only on the usage chunk — the last event before [DONE] when stream_options.include_usage asked for one — and absent from every content chunk. |
:::

### ChunkChoice {#schema-chunkchoice}

:::matrix
| member | what it is |
| --- | --- |
| `delta` | ChunkDelta, required |
| `finish_reason` | "stop" or "length" or "tool_calls" or "content_filter", required — Why the model stopped, stated once on the chunk that closes the answer and null until then. |
| `index` | integer, required — The position of this choice. |
:::

### ReplyUsage

`ReplyUsage` — see above.

### ChatCompletionsError {#schema-chatcompletionserror}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [ChatCompletionsErrorBody](#schema-chatcompletionserrorbody), required |
:::

### ChatCompletionsErrorBody {#schema-chatcompletionserrorbody}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, optional — The machine-readable reason. |
| `message` | string, required — What went wrong, in a fixed safe sentence. |
| `param` | string, optional — The request member at fault, when one member is at fault. |
| `type` | "invalid_request_error" or "not_found_error" or "authentication_error" or "permission_error" or "rate_limit_error" or "api_error", required — The class of failure, in this protocol's own closed vocabulary. |
:::

## POST /v1/embeddings {#embeddings-create keywords="Create embedding vectors for a batch of inputs."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/embeddings |
| Auth | Authorization: Bearer <key> |
| Operation | public.embeddings.create |
:::

Turn a batch of inputs into vectors, one vector per input.

Create embedding vectors for a batch of inputs.

Returns one vector per input, in the order the inputs were given. Billing counts input tokens only: this surface produces no output tokens and no image units, and the ledger records zero for both.

The batch is bounded — at most 128 inputs and 131072 bytes in total — and a request outside those bounds is refused before any provider is called. Streaming is not part of this surface. `encoding_format` must be stated and must be "base64": vectors are returned as base64-encoded little-endian float32, and this surface does not serve this protocol's "float" default.

A request asking for float, or asking for nothing and therefore for float, is refused by name rather than answered in another format.

:::code-group
```json title=Request
{
  "encoding_format": "base64",
  "input": [
    "<input>"
  ],
  "model": "<model>"
}
```
```json title=Response
{
  "data": [
    {
      "embedding": "<embedding>",
      "index": 1,
      "object": "embedding"
    }
  ],
  "model": "<model>",
  "object": "list",
  "usage": {
    "prompt_tokens": 1,
    "total_tokens": 1
  }
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `encoding_format` | "base64", required — Must be "base64": this surface does not serve the protocol's "float" default. |
| `input` | array of string, required — The batch of inputs to embed, at most 128 members and 131072 bytes in total. |
| `model` | string, required — The catalog model to embed with, as the customer names it. |
:::

The answer:

| what | shape |
| --- | --- |
| `application/json` | EmbeddingsResponseBody |
| `on failure` | EmbeddingsError |

### EmbeddingsResponseBody {#schema-embeddingsresponsebody}

:::matrix
| member | what it is |
| --- | --- |
| `data` | array of [EmbeddingsVector](#schema-embeddingsvector), required — One vector per input, in the order the inputs were given. |
| `model` | string, required — The model the vectors were produced with. |
| `object` | "list", required — Always "list". |
| `usage` | [EmbeddingsUsage](#schema-embeddingsusage), required — What the request cost. |
:::

### EmbeddingsVector {#schema-embeddingsvector}

:::matrix
| member | what it is |
| --- | --- |
| `embedding` | string, required — The vector components as base64-encoded little-endian float32 values, which is this protocol's own "base64" encoding format. |
| `index` | integer, required — The position of the input this vector is for. |
| `object` | "embedding", required — Always "embedding". |
:::

### EmbeddingsUsage {#schema-embeddingsusage}

:::matrix
| member | what it is |
| --- | --- |
| `prompt_tokens` | integer, required — The input tokens charged for this request. |
| `total_tokens` | integer, required — The total tokens charged, which on this surface equals prompt_tokens. |
:::

### EmbeddingsError {#schema-embeddingserror}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [EmbeddingsErrorBody](#schema-embeddingserrorbody), required — The failure. |
:::

### EmbeddingsErrorBody {#schema-embeddingserrorbody}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, required — The machine-readable reason, shared with every Kumo surface. |
| `message` | string, required — A safe description of the failure. |
| `request_id` | string, optional — The Kumo request identifier, for support. |
| `type` | string, required — The class of failure. |
:::

## POST /v1/images/generations {#images-generate keywords="Generate images from a prompt."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/images/generations |
| Auth | Authorization: Bearer <key> |
| Operation | public.images.generate |
:::

Generate one or more images from a text prompt.

Generate images from a prompt.

Generates one or more images from a text prompt with an image-capable model. The request and response envelopes are OpenAI-compatible, and so is the error envelope: refusals carry `error.message`, `error.type` and `error.code` rather than Kumo's REST envelope. Billing is per image unit, debited through the same admission, reservation and settlement kernel as every other model surface.

Edits and variations are not served.

:::code-group
```json title=Request
{
  "model": "<model>",
  "prompt": "<prompt>"
}
```
```json title=Response
{
  "created": 1,
  "data": [
    {}
  ]
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `model` | string, required — The public model name to generate with. |
| `n` | integer, optional — How many images to generate. |
| `prompt` | string, required — The prompt to generate an image from. |
| `size` | string, optional — The image size as WIDTHxHEIGHT, for example 1024x1024. |
:::

The answer:

| what | shape |
| --- | --- |
| `application/json` | ImagesResponseBody |
| `on failure` | ImagesErrorBody |

### ImagesResponseBody {#schema-imagesresponsebody}

:::matrix
| member | what it is |
| --- | --- |
| `created` | integer, required — When the images were generated, as a Unix timestamp in seconds. |
| `data` | array of [ImagesDataEntry](#schema-imagesdataentry), required — The generated images, in the order the supplier answered them. |
:::

### ImagesDataEntry {#schema-imagesdataentry}

:::matrix
| member | what it is |
| --- | --- |
| `b64_json` | string, optional — The image bytes, base64-encoded. |
| `url` | string, optional — A URL the image can be fetched from. |
:::

### ImagesErrorBody {#schema-imageserrorbody}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [ImagesErrorPayload](#schema-imageserrorpayload), required — The refusal, in the OpenAI-compatible error shape. |
:::

### ImagesErrorPayload {#schema-imageserrorpayload}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, required — The machine-readable reason. |
| `message` | string, required — A safe human-readable description of the refusal. |
| `param` | string, required — Always null: this surface never names a field of the request. |
| `type` | string, required — The coarse OpenAI error class. |
:::

## POST /v1/messages {#anthropic-messages-create keywords="Create a message."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/messages |
| Auth | Authorization: Bearer <key> or x-api-key: <key> |
| Operation | public.anthropic_messages.create |
:::

The Anthropic Messages surface: send a conversation in that protocol's own shape, get one message back or a stream of events.

Create a message.

Answers one Anthropic Messages request against a model of the published catalog, authenticated by a Kumo API key.

It is this protocol served NATIVELY and not a translation of another: max_tokens is required and has no default, the system prompt is a member of the request rather than a turn of the conversation, tools carry input_schema with no function wrapper, the answer states stop_reason and content blocks, and refusals arrive in this protocol's own error envelope.

A tool exchange is carried natively — an assistant's calls as tool_use blocks with input as the JSON object the tool schema describes, and their results as tool_result blocks of the following user turn.

A stream=true request is answered as text/event-stream carrying this protocol's own named events — message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop — with a refusal decided before the first event answered as this protocol's ordinary JSON error, and a failure after it expressed in-stream as this protocol's error event, the stream then ending without message_stop.

The credential may be presented as `Authorization: Bearer <key>` or, as the native API does it, bare in the `x-api-key` header; either admits the caller. When both are present the Authorization header is the one read and the bare header is ignored. Sampling is carried to the supplier unchanged: temperature, top_p and stop_sequences.

Four members are ACCEPTED AND HAVE NO EFFECT, and each says so in its own description rather than being refused or silently honoured — metadata, which this platform does not forward because it would describe a customer's own user to a supplier, and thinking, output_config and context_management, which are beta configuration this platform proves on no tuple.

:::code-group
```json title=Request
{
  "max_tokens": 128,
  "messages": [
    {
      "content": "<content>",
      "role": "user"
    }
  ],
  "model": "<model>"
}
```
```json title=Response
{
  "content": [
    {
      "type": "text"
    }
  ],
  "id": "<id>",
  "model": "<model>",
  "role": "<role>",
  "stop_reason": "end_turn",
  "type": "<type>"
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `context_management` | object, optional — This protocol's context-management configuration, as the beta spells it. |
| `max_tokens` | integer, required — The maximum number of tokens to generate. |
| `messages` | array of [MessagesInputMessage](#schema-messagesinputmessage), required — The conversation, oldest turn first. |
| `metadata` | [MessagesMetadata](#schema-messagesmetadata), optional — This protocol's request metadata. |
| `model` | string, required — The catalog model to answer with, as the customer names it. |
| `output_config` | object, optional — This protocol's output-effort configuration, as the beta spells it. |
| `stop_sequences` | array of string, optional — Sequences that end the answer when the model produces one. |
| `stream` | boolean, optional — Stream the answer as this protocol's own named events over text/event-stream: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop. |
| `system` | string or array of [MessagesSystemBlock](#schema-messagessystemblock), optional — The system prompt, beside the conversation rather than as a turn of it. |
| `temperature` | number, optional — How much randomness to use, from 0 to 1. |
| `thinking` | object, optional — This protocol's extended-thinking configuration, as the beta spells it. |
| `tool_choice` | [MessagesToolChoice](#schema-messagestoolchoice), optional — Forces one of the declared tools. |
| `tools` | array of [MessagesTool](#schema-messagestool), optional — The tools this turn may call. |
| `top_k` | integer, optional — Keep only the K most likely tokens when sampling. |
| `top_p` | number, optional — Nucleus sampling, from 0 to 1. |
:::

### MessagesInputMessage {#schema-messagesinputmessage}

:::matrix
| member | what it is |
| --- | --- |
| `content` | string or array of MessagesContentBlock, required — What the turn says. |
| `role` | "user" or "assistant" or "system", required — Who is speaking. |
:::

### MessagesMetadata {#schema-messagesmetadata}

:::matrix
| member | what it is |
| --- | --- |
| `user_id` | string, optional — An opaque identifier the caller keeps for its own end user. |
:::

### MessagesSystemBlock {#schema-messagessystemblock}

:::matrix
| member | what it is |
| --- | --- |
| `cache_control` | MessagesCacheControl, optional — Marks this block as a prompt-cache anchor. |
| `text` | string, required — The block's text. |
| `type` | "text", required — Always "text". |
:::

### MessagesToolChoice {#schema-messagestoolchoice}

:::matrix
| member | what it is |
| --- | --- |
| `disable_parallel_tool_use` | boolean, optional — Whether at most one tool may be called in one answer; the default is that several may be. |
| `name` | string, optional — The tool to force. |
| `type` | "auto" or "any" or "tool" or "none", required — What the request requires: "auto" leaves the choice to the model, "any" requires a call to some declared tool, "tool" requires a call to the one "name" states, and "none" forbids a call while the declarations stay visible. |
:::

### MessagesTool {#schema-messagestool}

:::matrix
| member | what it is |
| --- | --- |
| `allowed_callers` | array of string, optional — The beta's list of tools permitted to call this one. |
| `cache_control` | MessagesCacheControl, optional — Marks this declaration as a prompt-cache anchor; the tool list is part of the cached prefix. |
| `defer_loading` | boolean, optional — The beta's request to load this declaration only when it is first needed. |
| `description` | string, optional — What the tool does. |
| `eager_input_streaming` | boolean, optional — The beta's request to begin streaming this tool's arguments before they are complete. |
| `input_examples` | array of object, optional — The beta's example arguments for this tool, each the JSON object input_schema describes. |
| `input_schema` | object, optional — The tool's parameters, as the JSON Schema the caller wrote for its own tool. |
| `max_uses` | integer, optional — The beta's ceiling on how many times a server-side tool may run. |
| `name` | string, required — The tool's name. |
| `strict` | boolean, optional — The beta's request that the arguments conform exactly to input_schema. |
| `type` | string, optional — The kind of declaration. |
:::

The answer:

| what | shape |
| --- | --- |
| `application/json` | MessagesReply |
| `text/event-stream` | MessagesStreamEvent |
| `on failure` | MessagesError |

### MessagesReply {#schema-messagesreply}

:::matrix
| member | what it is |
| --- | --- |
| `content` | array of [MessagesReplyBlock](#schema-messagesreplyblock), required — The answer's blocks, in order. |
| `id` | string, required — The Kumo request identifier for this answer. |
| `model` | string, required — The model the customer named. |
| `role` | string, required — Always "assistant". |
| `stop_reason` | "end_turn" or "max_tokens" or "stop_sequence" or "tool_use", required — Why the answer ended. |
| `type` | string, required — Always "message". |
| `usage` | [MessagesReplyUsage](#schema-messagesreplyusage), optional — What the request cost, exactly as this platform settles it. |
:::

### MessagesReplyBlock {#schema-messagesreplyblock}

:::matrix
| member | what it is |
| --- | --- |
| `id` | string, optional — On a tool use, its identifier. |
| `input` | object, optional — On a tool use, the arguments as the JSON object the supplier produced. |
| `name` | string, optional — On a tool use, the tool called. |
| `text` | string, optional — On a text block, the text. |
| `type` | "text" or "tool_use", required — Which kind of block this is. |
:::

### MessagesReplyUsage {#schema-messagesreplyusage}

:::matrix
| member | what it is |
| --- | --- |
| `cache_creation_input_tokens` | integer, required — Input tokens written to the prompt cache. |
| `cache_read_input_tokens` | integer, required — Input tokens served from the prompt cache. |
| `input_tokens` | integer, required — Input tokens charged, excluding cache reads and writes. |
| `output_tokens` | integer, required — Output tokens charged. |
:::

### MessagesStreamEvent {#schema-messagesstreamevent}

:::matrix
| member | what it is |
| --- | --- |
| `content_block` | [MessagesStreamBlock](#schema-messagesstreamblock), optional — On content_block_start, the opening block. |
| `delta` | [MessagesStreamDelta](#schema-messagesstreamdelta), optional — On content_block_delta, the fragment; on message_delta, the closing facts. |
| `error` | [MessagesErrorBody](#schema-messageserrorbody), optional — On error, the failure, in this protocol's own envelope. |
| `index` | integer, optional — On the block events, which block. |
| `message` | [MessagesStreamMessage](#schema-messagesstreammessage), optional — On message_start, the opening envelope: the answer's identity and the account's input half. |
| `type` | "message_start" or "content_block_start" or "content_block_delta" or "content_block_stop" or "message_delta" or "message_stop" or "ping" or "error", required — Which event this is; it is also the SSE frame's event name. |
| `usage` | [MessagesReplyUsage](#schema-messagesreplyusage), optional — On message_delta, what the request cost — exactly as this platform settles it, and exactly what the unary reply would state. |
:::

### MessagesStreamBlock {#schema-messagesstreamblock}

:::matrix
| member | what it is |
| --- | --- |
| `id` | string, optional — On a tool use, its identifier. |
| `input` | object, optional — On a tool use, the opening input — the empty object; the arguments arrive as input_json_delta fragments. |
| `name` | string, optional — On a tool use, the tool called. |
| `text` | string, optional — On a text block, the opening text — empty; the text arrives as deltas. |
| `type` | "text" or "tool_use", required — Which kind of block opened. |
:::

### MessagesStreamDelta {#schema-messagesstreamdelta}

:::matrix
| member | what it is |
| --- | --- |
| `partial_json` | string, optional — On input_json_delta, the fragment of the call's input object, as partial JSON text. |
| `stop_reason` | "end_turn" or "max_tokens" or "stop_sequence" or "tool_use", optional — On message_delta, why the answer ended. |
| `text` | string, optional — On text_delta, the fragment of the answer's text. |
| `type` | "text_delta" or "input_json_delta", optional — On content_block_delta, which fragment this is; absent on message_delta. |
:::

### MessagesErrorBody {#schema-messageserrorbody}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, optional — The machine-readable reason, shared with every Kumo surface. |
| `message` | string, required — A safe description of the failure. |
| `param` | string, optional — The member the failure is about, when it is about one. |
| `type` | "invalid_request_error" or "not_found_error" or "authentication_error" or "permission_error" or "rate_limit_error" or "api_error" or "overloaded_error", required — The class of failure. |
:::

### MessagesStreamMessage {#schema-messagesstreammessage}

:::matrix
| member | what it is |
| --- | --- |
| `content` | array of [MessagesReplyBlock](#schema-messagesreplyblock), required — Always empty here: the blocks arrive as events. |
| `id` | string, required — The Kumo request identifier for this answer, identical on every event of one stream. |
| `model` | string, required — The model the customer named. |
| `role` | string, required — Always "assistant". |
| `type` | string, required — Always "message". |
| `usage` | [MessagesReplyUsage](#schema-messagesreplyusage), required — The account's input half, as the supplier stated it on opening; the closing message_delta states the settled whole. |
:::

### MessagesReplyUsage

`MessagesReplyUsage` — see above.

### MessagesError {#schema-messageserror}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [MessagesErrorBody](#schema-messageserrorbody), required — The failure. |
| `type` | "error", required — Always "error". |
:::

### MessagesErrorBody

`MessagesErrorBody` — see above.

## POST /v1/messages/count_tokens {#anthropic-messages-count-tokens keywords="Count message input tokens."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/messages/count_tokens |
| Auth | Authorization: Bearer <key> or x-api-key: <key> |
| Operation | public.anthropic_messages.count_tokens |
:::

Count how many tokens a Messages request would spend, without spending them.

Count message input tokens.

Authenticates a Kumo API key and returns a deterministic local estimate for a subset of the native Messages input accepted by /v1/messages: model, messages, system, tools, tool_choice, thinking, metadata and context_management. It performs no provider call, creates no customer request or reservation, touches no balance, and consumes no customer RPS/RPM quota.

The generation-only members are refused on this operation: max_tokens, stream, output_config, stop_sequences, temperature, top_p and top_k. The credential may be presented as `Authorization: Bearer <key>` or, as the native API does it, bare in the `x-api-key` header; either admits the caller. When both are present the Authorization header is the one read and the bare header is ignored.

:::code-group
```json title=Request
{
  "messages": [
    {
      "content": "<content>",
      "role": "user"
    }
  ],
  "model": "<model>"
}
```
```json title=Response
{
  "estimated": true,
  "input_tokens": 1
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `context_management` | object, optional — This protocol's context-management configuration, as the beta spells it. |
| `messages` | array of [MessagesInputMessage](#schema-messagesinputmessage), required — The conversation, oldest turn first. |
| `metadata` | [MessagesMetadata](#schema-messagesmetadata), optional — This protocol's request metadata. |
| `model` | string, required — The catalog model to answer with, as the customer names it. |
| `system` | string or array of [MessagesSystemBlock](#schema-messagessystemblock), optional — The system prompt, beside the conversation rather than as a turn of it. |
| `thinking` | object, optional — This protocol's extended-thinking configuration, as the beta spells it. |
| `tool_choice` | [MessagesToolChoice](#schema-messagestoolchoice), optional — Forces one of the declared tools. |
| `tools` | array of [MessagesTool](#schema-messagestool), optional — The tools this turn may call. |
:::

### MessagesInputMessage

`MessagesInputMessage` — see above.

### MessagesMetadata

`MessagesMetadata` — see above.

### MessagesSystemBlock

`MessagesSystemBlock` — see above.

### MessagesToolChoice

`MessagesToolChoice` — see above.

### MessagesTool

`MessagesTool` — see above.

The answer:

| what | shape |
| --- | --- |
| `application/json` | AnthropicTokenCountReply |
| `on failure` | MessagesError |

### AnthropicTokenCountReply {#schema-anthropictokencountreply}

:::matrix
| member | what it is |
| --- | --- |
| `estimated` | boolean, required — Always true until an exact local tokenizer is available. |
| `input_tokens` | integer, required — The deterministic local estimate of input tokens. |
:::

### MessagesError

`MessagesError` — see above.

### MessagesErrorBody

`MessagesErrorBody` — see above.

## GET /v1/models {#models-list keywords="List enabled, evidence-backed models."}

:::deflist
| field | value |
| --- | --- |
| Method | **GET** |
| Path | /v1/models |
| Auth | None — this operation is open. |
| Operation | public.models.list |
:::

List every model this gateway currently answers for.

List enabled, evidence-backed models.

Returns only active catalog models that have at least one enabled capability in the currently published provider configuration. Protocol and modality support are the exact intersection of catalog declarations and routable provider evidence; disabled or unproven tuples are absent.

```json title=Response
{
  "catalog_revision": 1,
  "data": [
    {
      "banner_seed": 1,
      "canonical_name": "<canonical_name>",
      "capabilities": [
        {
          "modality_code": "text_generation",
          "protocol_code": "responses",
          "status": "enabled",
          "streaming_status": "unsupported",
          "strict_semantics_status": "unsupported",
          "structured_output_mode": "none",
          "structured_output_status": "unsupported",
          "structured_output_streaming_status": "unsupported",
          "tools_status": "unsupported"
        }
      ],
      "description_en": "<description_en>",
      "description_ru": "<description_ru>",
      "display_name_en": "<display_name_en>",
      "display_name_ru": "<display_name_ru>",
      "id": "<id>",
      "modality_codes": [
        "text_generation"
      ],
      "model_id": "<model_id>",
      "object": "model",
      "protocol_codes": [
        "responses"
      ],
      "supplier_count": 1,
      "vendor": {
        "code": "<code>",
        "display_name_en": "<display_name_en>",
        "display_name_ru": "<display_name_ru>",
        "id": "<id>"
      }
    }
  ],
  "loaded_at": "<loaded_at>",
  "object": "list",
  "stale": true
}
```

The request carries no body.

The answer:

| what | shape |
| --- | --- |
| `application/json` | PublicModelCatalog |
| `on failure` | ErrorEnvelope |

### PublicModelCatalog {#schema-publicmodelcatalog}

:::matrix
| member | what it is |
| --- | --- |
| `catalog_revision` | integer, required — The append-only catalog revision. |
| `data` | array of [PublicCatalogModel](#schema-publiccatalogmodel), required |
| `loaded_at` | string, required |
| `object` | "list", required |
| `stale` | boolean, required — True only when a database refresh failed; the response then retains revision diagnostics but advertises no model capabilities. |
:::

### PublicCatalogModel {#schema-publiccatalogmodel}

:::matrix
| member | what it is |
| --- | --- |
| `aliases` | array of string, optional |
| `banner_seed` | integer, required — The seed of the model's halftone banner. |
| `canonical_name` | string, required |
| `capabilities` | array of ModelCapability, required |
| `context_window_tokens` | integer, optional — How many tokens this model accepts in one request. |
| `description_en` | string, required |
| `description_ru` | string, required |
| `display_name_en` | string, required |
| `display_name_ru` | string, required |
| `header_badge` | "top" or "value" or "fast", optional — The card this model fills in the header models panel. |
| `id` | string, required — The name an API request names this model by — the canonical name, so an OpenAI-compatible client that lists models and sends back data[].id as model is answered. |
| `max_output_tokens` | integer, optional — How many tokens this model may answer with in one call. |
| `modality_codes` | array of "text_generation" or "embeddings" or "image_generation", required |
| `model_id` | string, required — The catalog's UUID for this model, the key that pricing, key scopes and usage rows join on. |
| `object` | "model", required |
| `protocol_codes` | array of "responses" or "chat_completions" or "anthropic_messages" or "embeddings" or "images", required |
| `released_on` | string, optional — The day this model's maker released it, as yyyy-mm-dd. |
| `showcase_position` | integer, optional — The model's place on the landing model carousel, ascending. |
| `supplier_count` | integer, required — How many distinct suppliers currently carry this model. |
| `typical_request` | CatalogTypicalRequest, optional — What one average request to this model actually spends, measured over the platform's own finished traffic. |
| `vendor` | CatalogVendor, required |
:::

### ErrorEnvelope {#schema-errorenvelope}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [ErrorBody](#schema-errorbody), required — The envelope payload. |
:::

### ErrorBody {#schema-errorbody}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, required — Machine-readable error code. |
| `field_details` | array of FieldDetail, optional — Per-field rejections, when the error is a validation error. |
| `limit` | LimitDetail, optional — Present when the error is a limit or funding-source rejection. |
| `message` | string, required — Safe human-readable fallback. |
| `promotion` | PromotionDetail, optional — Present when the error is a named promotion refusal. |
| `request_id` | string, required — Matches the X-Request-Id response header. |
| `version` | string, required — Envelope version. |
:::

## POST /v1/responses {#responses-create keywords="Create a response."}

:::deflist
| field | value |
| --- | --- |
| Method | **POST** |
| Path | /v1/responses |
| Auth | Authorization: Bearer <key> |
| Operation | public.responses.create |
:::

The newer OpenAI Responses surface: send input items, get one response object back or a stream of events.

Create a response.

Answers one Responses request against a model of the published catalog, authenticated by a Kumo API key. It is this protocol served natively and not a translation of another: the input items, the output items, the usage vocabulary and the error envelope are this protocol's own.

A stream=true request is answered as text/event-stream in this protocol's own named events — response.created, the output items and their deltas, then exactly one terminal event: response.completed carrying the finished body and usage, or response.failed carrying this protocol's error envelope. A failure before the first event is answered as this protocol's ordinary JSON error; cancellation releases what the request held.

The output ceiling — max_output_tokens — is optional, and a request stating none is answered under this surface's published default of 32768 tokens; this platform requires a FINITE bound before the upstream call, and a published default is one the caller can read in advance. An explicit zero is refused: it is a request for no output at all.

instructions is implemented and becomes the conversation's leading system turn; a developer turn is carried as a system turn, the two being one role under two names. store is accepted only as false, because this platform retains no response and accepting true would promise a retrieval that cannot happen.

Five members are ACCEPTED AND HAVE NO EFFECT, each saying so in its own description rather than being refused or silently honoured — parallel_tool_calls, reasoning, prompt_cache_key, client_metadata and text.verbosity.

include is half of that: a value this surface publishes an answer for is forwarded to the supplier — today reasoning.encrypted_content, which is what lets a client replay a reasoning item the supplier itself signed — and any other value has no effect, because asking a supplier for an item this surface does not publish would turn a served answer into a refused one.

A previous_response_id is validated for OWNERSHIP and is not honoured as server-side context: this platform persists no prompt, response or chunk, so there is nothing to continue from and the caller sends its own context.

:::code-group
```json title=Request
{
  "input": "<input>",
  "model": "<model>"
}
```
```json title=Response
{
  "created_at": 1,
  "id": "<id>",
  "model": "<model>",
  "object": "response",
  "output": [
    {
      "type": "message"
    }
  ],
  "status": "completed"
}
```
:::

The request carries:

:::matrix
| member | what it is |
| --- | --- |
| `client_metadata` | object, optional — Metadata the client keeps about its own session. |
| `include` | array of string, optional — Extra members the caller asks the answer to carry. |
| `input` | string or array of [ResponsesInputItem](#schema-responsesinputitem), required — The conversation, in either spelling this protocol defines: a bare string, or a list of typed items in order. |
| `instructions` | string, optional — The system prompt, as this protocol carries it: a member of the request rather than a turn of the conversation. |
| `max_output_tokens` | integer, optional — The output ceiling. |
| `model` | string, required — The model to answer with: a canonical name or an alias the published catalog carries. |
| `parallel_tool_calls` | boolean, optional — Whether the model may make several tool calls in one turn. |
| `previous_response_id` | string, optional — A response of this organization that this request follows. |
| `prompt_cache_key` | string, optional — An opaque key the caller uses to group requests for prompt caching. |
| `reasoning` | object, optional — This protocol's reasoning configuration. |
| `store` | boolean, optional — Whether the supplier should retain this response for later retrieval. |
| `stream` | boolean, optional — Whether to stream the answer. |
| `text` | [ResponsesTextConfig](#schema-responsestextconfig), optional — How the answer's text is shaped. |
| `tool_choice` | "auto" or "none" or "required" or object, optional — What this request requires of its tool list, as a bare mode word or as an object naming one declared tool. |
| `tools` | array of [ResponsesTool](#schema-responsestool), optional — The tools this turn may call. |
:::

### ResponsesInputItem {#schema-responsesinputitem}

:::matrix
| member | what it is |
| --- | --- |
| `action` | object, optional — What a supplier-side search DID, as the supplier described it and as this surface published it. |
| `arguments` | string, optional — The arguments the call was made with, as the JSON text the model produced. |
| `call_id` | string, optional — The call this item is or answers. |
| `content` | string or array of ResponsesInputContentPart, optional — The text of a message item, and the reasoning prose of a reasoning item, in either spelling this protocol defines: a bare string, or a list of typed parts. |
| `encrypted_content` | string, optional — The supplier's own encrypted record of a reasoning item, as it was given to the client. |
| `id` | string, optional — The identity this item carried when the caller last saw it. |
| `input` | string, optional — The model's FREEFORM answer to a custom tool, as the text it produced: a patch, a query, whatever the tool's own grammar admits. |
| `name` | string, optional — The tool that was called. |
| `namespace` | string, optional — The group the called tool was declared in, on a function_call or a custom_tool_call, when the request that produced the call declared its tools in namespaces. |
| `output` | string or array of ResponsesInputContentPart, optional — The result of running the tool, in either spelling this protocol defines: a bare string, or a list of parts of kind input_text, which is what a tool result's parts are. |
| `role` | "system" or "user" or "developer" or "assistant", optional — Who is speaking. |
| `status` | string, optional — How far the item got, in this protocol's own vocabulary and carried as the caller stated it. |
| `summary` | array of ResponsesInputContentPart, optional — A reasoning item's summary, in parts of kind summary_text. |
| `tools` | array of ResponsesToolNamespace, optional — The namespaced tool declarations of an additional_tools item. |
| `type` | "message" or "function_call" or "function_call_output" or "additional_tools" or "custom_tool_call" or "custom_tool_call_output" or "reasoning" or "web_search_call", optional — Which kind of item this is. |
:::

### ResponsesTextConfig {#schema-responsestextconfig}

:::matrix
| member | what it is |
| --- | --- |
| `format` | ResponsesTextFormat, optional — The named schema an answer must satisfy. |
| `verbosity` | "low" or "medium" or "high", optional — How much the answer should say. |
:::

### ResponsesTool {#schema-responsestool}

:::matrix
| member | what it is |
| --- | --- |
| `description` | string, optional — What the tool does. |
| `execution` | string, optional — Which side runs a TOOL_SEARCH — the reference spells "client" and "server". |
| `external_web_access` | boolean, optional — Whether a WEB_SEARCH may reach the open web. |
| `filters` | ResponsesWebSearchFilters, optional — What a WEB_SEARCH is limited to. |
| `format` | ResponsesToolFormat, optional — How a CUSTOM tool's freeform answer is shaped. |
| `name` | string, optional — The tool's name, as it will come back on the call. |
| `parameters` | object, optional — The tool's parameters, as the JSON Schema the caller wrote for its own tool. |
| `search_content_types` | array of string, optional — Which kinds of content a WEB_SEARCH may return. |
| `search_context_size` | string, optional — How much search context a WEB_SEARCH should gather. |
| `strict` | boolean, optional — Whether the supplier must enforce the parameter schema rather than be encouraged toward it. |
| `type` | "function" or "custom" or "tool_search" or "web_search", required — The kind of tool. |
| `user_location` | ResponsesWebSearchLocation, optional — Where a WEB_SEARCH should answer as though it were. |
:::

The answer:

| what | shape |
| --- | --- |
| `application/json` | ResponsesReply |
| `text/event-stream` | ResponsesStreamEvent |
| `on failure` | ResponsesError |

### ResponsesReply {#schema-responsesreply}

:::matrix
| member | what it is |
| --- | --- |
| `created_at` | integer, required — The instant this response was answered, as whole seconds since the Unix epoch. |
| `id` | string, required — This request's Kumo identity. |
| `model` | string, required — The model this request named, echoed back exactly as sent. |
| `object` | "response", required — The kind of object this is. |
| `output` | array of [ResponsesOutputItem](#schema-responsesoutputitem), required — What the model produced, in order, as typed items. |
| `status` | "completed" or "incomplete", required — How the answer ended. |
| `usage` | [ResponsesUsage](#schema-responsesusage), optional — What the supplier reported this request consumed. |
:::

### ResponsesOutputItem {#schema-responsesoutputitem}

:::matrix
| member | what it is |
| --- | --- |
| `action` | object, optional — What a supplier-side search DID, as the supplier described it. |
| `arguments` | string, optional — The arguments, as the JSON text the supplier produced. |
| `call_id` | string, optional — The call's identity, to quote when answering it on the next turn. |
| `content` | array of [ResponsesOutputContentPart](#schema-responsesoutputcontentpart), optional — The message's content, in parts, and a reasoning item's own reasoning text, in parts of kind reasoning_text. |
| `encrypted_content` | string, optional — The supplier's own encrypted record of a reasoning item. |
| `id` | string, optional — The item's own identity, as the supplier stated it. |
| `input` | string, optional — The model's freeform answer to a CUSTOM tool, as the text it produced. |
| `name` | string, optional — The tool that was called. |
| `namespace` | string, optional — The group the called tool was declared in, on a function_call or a custom_tool_call, when this request declared its tools in namespaces. |
| `role` | "assistant", optional — Who is speaking. |
| `status` | string, optional — How far a supplier-side search has got, and how far a reasoning item got. |
| `summary` | array of [ResponsesOutputContentPart](#schema-responsesoutputcontentpart), optional — A reasoning item's summary, in parts of kind summary_text. |
| `type` | "message" or "function_call" or "custom_tool_call" or "web_search_call" or "reasoning", required — Which kind of item this is. |
:::

### ResponsesUsage {#schema-responsesusage}

:::matrix
| member | what it is |
| --- | --- |
| `input_tokens` | integer, optional — Input tokens the supplier counted, including any it served from its own cache. |
| `input_tokens_details` | ResponsesInputTokensDetails, optional — How the input divides, when the supplier said. |
| `output_tokens` | integer, optional — Output tokens the supplier counted. |
| `total_tokens` | integer, optional — Input and output together. |
:::

### ResponsesStreamEvent {#schema-responsesstreamevent}

:::matrix
| member | what it is |
| --- | --- |
| `content_index` | integer, optional — The index of the content part this event is about, within its own output item. |
| `delta` | string, optional — The fragment a delta event carries: text on response.output_text.delta, the model's own refusal on response.refusal.delta, argument text on response.function_call_arguments.delta, and a custom tool's freeform input on response.custom_tool_call_input.delta. |
| `item` | [ResponsesOutputItem](#schema-responsesoutputitem), optional — The item, on response.output_item.added — opened, so a call carries its identity and no arguments yet — on response.output_item.done, finished, and on the three web_search_call stage events, carrying the status the search has reached. |
| `output_index` | integer, optional — The index of the output item this event is about. |
| `part` | [ResponsesOutputContentPart](#schema-responsesoutputcontentpart), optional — The content part, on response.content_part.added — opened, so it carries its kind and no text yet — and on response.content_part.done, assembled, carrying exactly the text the fragments between the two events delivered. |
| `response` | [ResponsesStreamSnapshot](#schema-responsesstreamsnapshot), optional — The response, on the response-scoped events: opening on response.created, complete with output and usage on response.completed, and carrying this protocol's error envelope on response.failed. |
| `sequence_number` | integer, required — This event's position in the stream that carried it, counting from zero and rising by one on every event the customer is sent. |
| `type` | "response.created" or "response.output_item.added" or "response.content_part.added" or "response.output_text.delta" or "response.refusal.delta" or "response.function_call_arguments.delta" or "response.custom_tool_call_input.delta" or "response.web_search_call.in_progress" or "response.web_search_call.searching" or "response.web_search_call.completed" or "response.content_part.done" or "response.output_item.done" or "response.completed" or "response.failed", required — Which event this is. |
:::

### ResponsesOutputItem

`ResponsesOutputItem` — see above.

### ResponsesOutputContentPart {#schema-responsesoutputcontentpart}

:::matrix
| member | what it is |
| --- | --- |
| `refusal` | string, optional — The model's own refusal, on a refusal part. |
| `text` | string, optional — The answer's text, on an output_text part, and the reasoning prose on a summary_text or reasoning_text part. |
| `type` | "output_text" or "refusal" or "summary_text" or "reasoning_text", required — The kind of part. |
:::

### ResponsesStreamSnapshot {#schema-responsesstreamsnapshot}

:::matrix
| member | what it is |
| --- | --- |
| `created_at` | integer, required — The instant this response was answered, as whole seconds since the Unix epoch. |
| `error` | [ResponsesErrorBody](#schema-responseserrorbody), optional — This protocol's own error envelope, on response.failed. |
| `id` | string, required — This request's Kumo identity, identical on every event of one stream. |
| `model` | string, required — The model this request named, echoed back exactly as sent, on every snapshot. |
| `object` | "response", required — The kind of object this is. |
| `output` | array of [ResponsesOutputItem](#schema-responsesoutputitem), optional — What the model produced, in order, on response.completed. |
| `status` | "in_progress" or "completed" or "incomplete" or "failed", required — Where the answer stands. |
| `usage` | [ResponsesUsage](#schema-responsesusage), optional — What the supplier reported this request consumed, on response.completed. |
:::

### ResponsesError {#schema-responseserror}

:::matrix
| member | what it is |
| --- | --- |
| `error` | [ResponsesErrorBody](#schema-responseserrorbody), required |
:::

### ResponsesErrorBody {#schema-responseserrorbody}

:::matrix
| member | what it is |
| --- | --- |
| `code` | string, optional — The machine-readable reason. |
| `message` | string, required — What went wrong, in a fixed safe sentence. |
| `param` | string, optional — The request member at fault, when one member is at fault. |
| `type` | "invalid_request_error" or "not_found_error" or "authentication_error" or "permission_error" or "rate_limit_error" or "api_error", required — The class of failure, in this protocol's own closed vocabulary. |
:::

> Every operation above is emitted from the API description this build was compiled against. The whole of this documentation as one Markdown document is described on the [machine-readable page](/en/machine-readable).
