---
title: Messages
description: The Anthropic Messages dialect — the bare base URL, two carriers for the key, request members, reply blocks, stream events, and a token count that spends no quota.
keywords: anthropic, messages, x-api-key, count_tokens, stop_reason, message_stop
group: gateway
---

## Quick {#quick keywords="curl, python, anthropic sdk, base url"}

The Anthropic Messages dialect is served **natively** and not translated from another: `max_tokens` is required and has no default, the system prompt is a member of the request, tools carry `input_schema` with no function wrapper, and the answer speaks in blocks and a `stop_reason`.

:::warning
The base URL for Anthropic clients is the **bare origin** `https://api.kumorouter.com`, without `/v1`. Such a client composes the path itself: it takes the base URL and appends `/v1/messages`. A base URL carrying a prefix of its own sends it to a path nothing serves, and you get a 404 instead of an answer.
:::

:::code-group
```bash title=curl
curl https://api.kumorouter.com/v1/messages \
  -H "x-api-key: $KUMO_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "<model>",
    "max_tokens": 256,
    "system": "Answer in one sentence.",
    "messages": [{ "role": "user", "content": "Explain tokens." }]
  }'
```
```python title=Python
import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.kumorouter.com",
    api_key=os.environ["KUMO_API_KEY"],
)

message = client.messages.create(
    model="<model>",
    max_tokens=256,
    system="Answer in one sentence.",
    messages=[{"role": "user", "content": "Explain tokens."}],
)
print(message.content[0].text)
```
:::

## Two carriers for the key {#credential keywords="x-api-key, authorization, bearer, 401"}

The key is presented either as `Authorization: Bearer <key>` or, as the native API does it, bare in the `x-api-key` header. Both name the same Kumo key, and either admits the call. An Anthropic SDK sends the second and no `Authorization` header at all, which is why it is declared on these operations.

Presenting both is legal. `Authorization` wins: when it is present and well formed, it is the one that authenticates, and the bare `x-api-key` is not read at all in that case. An ordinary client sends exactly one of the two — the second header shows up only in a hand-built wrapper that layers headers on top of a ready-made SDK.

> [How the header is written →](/en/authentication) [Check your key →](/en/key-check)

## Request members {#request keywords="members, max_tokens, system, tools, temperature"}

| Member | What it is |
| --- | --- |
| `model` | Required. The catalog model, under the name the client uses for it. |
| `max_tokens` | Required. The maximum number of tokens to generate; this surface has no default. |
| `messages` | Required. The conversation, oldest turn first; from one up to 512. |
| `system` | The system prompt, beside the conversation rather than as a turn of it. Either spelling the protocol defines is accepted: a bare string, or a list of text blocks. A string is carried as the equivalent one-block list. |
| `stream` | Whether to answer as a stream of this protocol's own named events. |
| `tools` | This turn's tools; up to 128 declarations. |
| `tool_choice` | An object whose `type` is `auto`, `any`, `none` or `tool`. `auto` lets the model choose, `any` requires some declared tool, and `none` forbids a tool call. Only `tool` requires `name`, which must name a declared tool; the other modes reject `name`. |
| `temperature` | Randomness, from 0 to 1. Carried to the supplier unchanged. |
| `top_p` | Nucleus sampling, from 0 to 1. Carried to the supplier unchanged. |
| `top_k` | Integer from 1 to 1048576. Carried unchanged to a compatible supplier; omission stays absent, not zero. A route that cannot represent it refuses it before an upstream call. |
| `stop_sequences` | Sequences whose appearance ends the answer; up to 16. One that fires comes back as `stop_reason: "stop_sequence"`; which of them fired is not published. |
| `metadata` | Accepted and not forwarded: this platform does not describe your user to a supplier. |
| `thinking` | Accepted and not forwarded. |
| `output_config` | Accepted and not forwarded. |
| `context_management` | Accepted and not forwarded: there is no server-side context here, and the caller sends its own. |

A turn carries a `role` (`user`, `assistant`, `system`) and `content` — a string, or a list of typed blocks: `text`, `tool_use`, `tool_result`. A `system` turn is not a turn of the conversation: its text is lifted into the system prompt, after the blocks the `system` member itself states.

`tool_result.content` accepts a string or up to 32 blocks of type `text` or `image`. Image bytes use a `base64` source with `media_type` and `data`; URL sources are rejected. Native Messages carries these blocks unchanged. A Chat Completions route cannot represent an image inside a tool result and refuses that result rather than discarding the image.

## The response {#response keywords="content, stop_reason, usage, blocks"}

| Reply member | What it is |
| --- | --- |
| `id` | This request's Kumo identity. |
| `type` | Always `message`. |
| `role` | Always `assistant`. |
| `model` | The model the client named. |
| `content` | The answer's blocks in order: `text` and `tool_use`. |
| `stop_reason` | Why the answer ended: `end_turn`, `max_tokens`, `stop_sequence` or `tool_use`. |
| `usage` | `input_tokens`, `output_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`. |

```json title=Response
{
  "id": "msg-71c3d0aa4e",
  "type": "message",
  "role": "assistant",
  "model": "<model>",
  "content": [{ "type": "text", "text": "Tokens are the pieces of text a model measures input and output in." }],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 16,
    "output_tokens": 22,
    "cache_read_input_tokens": 0,
    "cache_creation_input_tokens": 0
  }
}
```

## Streaming {#stream keywords="message_start, content_block_delta, message_stop, error"}

`"stream": true` answers as `text/event-stream` in the protocol's own named events: `message_start`, `content_block_start`, `content_block_delta`, `content_block_stop`, `message_delta`, `message_stop`. Alongside them arrive `ping` and `error`.

A refusal decided before the first event is this protocol's ordinary JSON error. A failure after the first event can only be expressed in the stream — as the `error` event — and the stream then ends **without** `message_stop`. So here a whole answer is marked by `message_stop`, and by nothing else.

> [Every event of all three dialects →](/en/streaming)

## Counting tokens {#count-tokens keywords="count_tokens, estimate, quota, ahead of time"}

`POST /v1/messages/count_tokens` takes a subset of the `/v1/messages` members — `model`, `messages`, `system`, `tools`, `tool_choice`, `thinking`, `metadata` and `context_management` — and answers with a deterministic local estimate of the input tokens. It calls no provider, creates no request or reservation, touches no balance and consumes no rate quota — which is what makes it easy to put in front of a large call.

Everything else describes the generation rather than the input, and is refused on this operation: `max_tokens`, `stream`, `output_config`, `stop_sequences`, `temperature` and `top_p`.

```bash title=curl
curl https://api.kumorouter.com/v1/messages/count_tokens \
  -H "x-api-key: $KUMO_API_KEY" \
  -H "content-type: application/json" \
  -d '{
    "model": "<model>",
    "system": "Answer in one sentence.",
    "messages": [{ "role": "user", "content": "Explain tokens." }]
  }'
```

The reply carries two members: `input_tokens`, the estimate itself, and `estimated`, which stays `true` until an exact local tokenizer is available.

```json title=Response
{ "input_tokens": 16, "estimated": true }
```

## Next {#next keywords="tools, errors, rate limits, reference"}

:::cards
- [Tools](/en/tools) — `input_schema`, and the `tool_use` and `tool_result` blocks.
- [Streaming](/en/streaming) — the events of all three dialects and the terminal-frame rule.
- [Errors](/en/errors) — the refusal envelope and the status codes.
- [Rate limits](/en/limits) — the ceilings, and how to size a call in advance.
- [Claude Code](/en/claude-code) — a ready recipe for a client that speaks this dialect.
- [API reference](/en/api-reference) — the operation member by member, straight from the API description.
:::
