Chat Completions
The POST /v1/chat/completions call — request members, conversation roles, the response shape, and pointers to streaming, tools and structured output.
Quick
Point an OpenAI-compatible client at https://api.kumorouter.com/v1, give it your key and name a model. An output ceiling is required: max_tokens or max_completion_tokens.
curl https://api.kumorouter.com/v1/chat/completions \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"max_tokens": 128,
"messages": [{ "role": "user", "content": "Explain tokens in one line." }]
}'Request members
The body is one JSON object. A member that is not in the table is not accepted by this endpoint: an extra member is refused rather than ignored.
modelRequired. The model to answer with: a canonical name or an alias the published catalog carries. The reply echoes the same name back.messagesRequired. The conversation in order, from one turn up to 512.max_tokensThe output ceiling, in the spelling long-established clients send.max_completion_tokensThe same ceiling in the current spelling. Send one of the two, or both carrying the same value; a request naming neither is refused.streamWhether to stream the answer. `true` gives you `text/event-stream`.stream_optionsOnly alongside `stream: true`. Its one member is `include_usage`.toolsThe tools this turn may call; up to 128 declarations.tool_choiceA requirement that one named declared tool be called. Omit it and the model chooses.response_formatA structured output requirement: a named JSON Schema.Turns of the conversation
A turn carries a role and a content. There are four roles: system, user, assistant, tool.
content is the turn's text, or null. null is legal on an assistant turn whose whole answer was a tool call — the form every OpenAI-compatible client replays. On a tool turn, content is the result of running the tool.
This surface carries text. A turn is not assembled from parts and takes no images: content is a string or null, and there is no other shape for it.
Two members belong to the tool exchange.
tool_callsThe calls this assistant turn made, replayed back into the conversation so that a tool result has something to answer. Only an assistant makes a call.tool_call_idThe call this result answers. Required on a `tool` turn and refused on every other: a supplier matches results to calls by it rather than by position.What a whole tool round looks like →
The response
The reply carries exactly one choice — this surface offers no way to ask for more — and its index is always zero.
idThis request's Kumo identity. It is the one to quote in a support question.objectAlways `chat.completion`.createdThe instant the completion was answered, as whole seconds since the Unix epoch.modelThe name the request sent, echoed back exactly: a request that named an alias reads that alias back.choicesThe answer: a `message` with the `assistant` role, and a `finish_reason`.usageWhat the supplier reported the request consumed. Absent when the supplier reported no usage at all.finish_reason takes four values: stop, length, tool_calls, content_filter.
Besides content, the message may carry refusal — the model's own refusal to answer. That is an answer and not an error: the request was served and is settled like any other.
usage counts prompt_tokens, completion_tokens and total_tokens; prompt_tokens_details.cached_tokens is part of the input tokens, not additional to them.
{
"id": "chatcmpl-8f2b7e10c9",
"object": "chat.completion",
"created": 1756900000,
"model": "<model>",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Tokens are the small pieces of text a model reads and writes."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 12,
"completion_tokens": 18,
"total_tokens": 30
}
}