---
title: Chat Completions
description: The POST /v1/chat/completions call — request members, conversation roles, the response shape, and pointers to streaming, tools and structured output.
keywords: chat completions, openai, messages, choices, finish_reason, usage
group: gateway
---

## Quick {#quick keywords="curl, python, node, first call"}

Point an OpenAI-compatible client at `https://api.kumorouter.com/v1`, give it your key and name a model. An output ceiling is required: `max_tokens` or `max_completion_tokens`.

:::code-group
```bash title=curl
curl https://api.kumorouter.com/v1/chat/completions \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "max_tokens": 128,
    "messages": [{ "role": "user", "content": "Explain tokens in one line." }]
  }'
```
```python title=Python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.kumorouter.com/v1",
    api_key=os.environ["KUMO_API_KEY"],
)

answer = client.chat.completions.create(
    model="<model>",
    max_tokens=128,
    messages=[{"role": "user", "content": "Explain tokens in one line."}],
)
print(answer.choices[0].message.content)
```
```javascript title=Node
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.kumorouter.com/v1",
  apiKey: process.env.KUMO_API_KEY,
});

const answer = await client.chat.completions.create({
  model: "<model>",
  max_tokens: 128,
  messages: [{ role: "user", content: "Explain tokens in one line." }],
});

const reply = answer.choices[0].message.content;
process.stdout.write((reply ?? "") + "\n");
```
:::

## Request members {#request keywords="members, fields, body, max_tokens, stream"}

The body is one JSON object. A member that is not in the table is not accepted by this endpoint: an extra member is refused rather than ignored.

| Member | What it is |
| --- | --- |
| `model` | Required. The model to answer with: a canonical name or an alias the published catalog carries. The reply echoes the same name back. |
| `messages` | Required. The conversation in order, from one turn up to 512. |
| `max_tokens` | The output ceiling, in the spelling long-established clients send. |
| `max_completion_tokens` | The same ceiling in the current spelling. Send one of the two, or both carrying the same value; a request naming neither is refused. |
| `stream` | Whether to stream the answer. `true` gives you `text/event-stream`. |
| `stream_options` | Only alongside `stream: true`. Its one member is `include_usage`. |
| `tools` | The tools this turn may call; up to 128 declarations. |
| `tool_choice` | A requirement that one named declared tool be called. Omit it and the model chooses. |
| `response_format` | A structured output requirement: a named JSON Schema. |

## Turns of the conversation {#messages keywords="roles, system, user, assistant, tool, content"}

A turn carries a `role` and a `content`. There are four roles: `system`, `user`, `assistant`, `tool`.

`content` is the turn's text, or `null`. `null` is legal on an assistant turn whose whole answer was a tool call — the form every OpenAI-compatible client replays. On a `tool` turn, `content` is the result of running the tool.

This surface carries text. A turn is not assembled from parts and takes no images: `content` is a string or `null`, and there is no other shape for it.

Two members belong to the tool exchange.

| Turn member | What it is |
| --- | --- |
| `tool_calls` | The calls this assistant turn made, replayed back into the conversation so that a tool result has something to answer. Only an assistant makes a call. |
| `tool_call_id` | The call this result answers. Required on a `tool` turn and refused on every other: a supplier matches results to calls by it rather than by position. |

> [What a whole tool round looks like →](/en/tools)

## The response {#response keywords="choices, finish_reason, usage, id, reply"}

The reply carries exactly one choice — this surface offers no way to ask for more — and its `index` is always zero.

| Reply member | What it is |
| --- | --- |
| `id` | This request's Kumo identity. It is the one to quote in a support question. |
| `object` | Always `chat.completion`. |
| `created` | The instant the completion was answered, as whole seconds since the Unix epoch. |
| `model` | The name the request sent, echoed back exactly: a request that named an alias reads that alias back. |
| `choices` | The answer: a `message` with the `assistant` role, and a `finish_reason`. |
| `usage` | What the supplier reported the request consumed. Absent when the supplier reported no usage at all. |

`finish_reason` takes four values: `stop`, `length`, `tool_calls`, `content_filter`.

Besides `content`, the message may carry `refusal` — the model's own refusal to answer. That is an answer and not an error: the request was served and is settled like any other.

`usage` counts `prompt_tokens`, `completion_tokens` and `total_tokens`; `prompt_tokens_details.cached_tokens` is part of the input tokens, not additional to them.

```json title=Response
{
  "id": "chatcmpl-8f2b7e10c9",
  "object": "chat.completion",
  "created": 1756900000,
  "model": "<model>",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Tokens are the small pieces of text a model reads and writes."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 18,
    "total_tokens": 30
  }
}
```

## Next {#next keywords="streaming, tools, structured output, errors"}

:::cards
- [Streaming](/en/streaming) — `stream: true`, the event names, and the terminal frame that marks a whole answer.
- [Tools](/en/tools) — `tools`, `tool_choice` and the turn that carries a result.
- [Structured output](/en/structured-output) — `response_format` with a named strict schema.
- [Errors](/en/errors) — the refusal envelope, the status codes, and one 401 for anything wrong with a key.
- [Rate limits](/en/limits) — the ceilings a call has to stay under.
- [API reference](/en/api-reference) — the same operation member by member, straight from the API description.
:::
