Skip to contentKumoDocs
Sections
On this page
The gateway

Chat Completions

The POST /v1/chat/completions call — request members, conversation roles, the response shape, and pointers to streaming, tools and structured output.

View as Markdown

Quick

Point an OpenAI-compatible client at https://api.kumorouter.com/v1, give it your key and name a model. An output ceiling is required: max_tokens or max_completion_tokens.

curl https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "max_tokens": 128,
    "messages": [{ "role": "user", "content": "Explain tokens in one line." }]
  }'

Request members

The body is one JSON object. A member that is not in the table is not accepted by this endpoint: an extra member is refused rather than ignored.

modelRequired. The model to answer with: a canonical name or an alias the published catalog carries. The reply echoes the same name back.
messagesRequired. The conversation in order, from one turn up to 512.
max_tokensThe output ceiling, in the spelling long-established clients send.
max_completion_tokensThe same ceiling in the current spelling. Send one of the two, or both carrying the same value; a request naming neither is refused.
streamWhether to stream the answer. `true` gives you `text/event-stream`.
stream_optionsOnly alongside `stream: true`. Its one member is `include_usage`.
toolsThe tools this turn may call; up to 128 declarations.
tool_choiceA requirement that one named declared tool be called. Omit it and the model chooses.
response_formatA structured output requirement: a named JSON Schema.

Turns of the conversation

A turn carries a role and a content. There are four roles: system, user, assistant, tool.

content is the turn's text, or null. null is legal on an assistant turn whose whole answer was a tool call — the form every OpenAI-compatible client replays. On a tool turn, content is the result of running the tool.

This surface carries text. A turn is not assembled from parts and takes no images: content is a string or null, and there is no other shape for it.

Two members belong to the tool exchange.

tool_callsThe calls this assistant turn made, replayed back into the conversation so that a tool result has something to answer. Only an assistant makes a call.
tool_call_idThe call this result answers. Required on a `tool` turn and refused on every other: a supplier matches results to calls by it rather than by position.

What a whole tool round looks like →

The response

The reply carries exactly one choice — this surface offers no way to ask for more — and its index is always zero.

idThis request's Kumo identity. It is the one to quote in a support question.
objectAlways `chat.completion`.
createdThe instant the completion was answered, as whole seconds since the Unix epoch.
modelThe name the request sent, echoed back exactly: a request that named an alias reads that alias back.
choicesThe answer: a `message` with the `assistant` role, and a `finish_reason`.
usageWhat the supplier reported the request consumed. Absent when the supplier reported no usage at all.

finish_reason takes four values: stop, length, tool_calls, content_filter.

Besides content, the message may carry refusal — the model's own refusal to answer. That is an answer and not an error: the request was served and is settled like any other.

usage counts prompt_tokens, completion_tokens and total_tokens; prompt_tokens_details.cached_tokens is part of the input tokens, not additional to them.

{
  "id": "chatcmpl-8f2b7e10c9",
  "object": "chat.completion",
  "created": 1756900000,
  "model": "<model>",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Tokens are the small pieces of text a model reads and writes."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 18,
    "total_tokens": 30
  }
}

Next