Messages
The Anthropic Messages dialect — the bare base URL, two carriers for the key, request members, reply blocks, stream events, and a token count that spends no quota.
Quick
The Anthropic Messages dialect is served natively and not translated from another: max_tokens is required and has no default, the system prompt is a member of the request, tools carry input_schema with no function wrapper, and the answer speaks in blocks and a stop_reason.
The base URL for Anthropic clients is the bare origin https://api.kumorouter.com, without /v1. Such a client composes the path itself: it takes the base URL and appends /v1/messages. A base URL carrying a prefix of its own sends it to a path nothing serves, and you get a 404 instead of an answer.
curl https://api.kumorouter.com/v1/messages \
-H "x-api-key: $KUMO_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "<model>",
"max_tokens": 256,
"system": "Answer in one sentence.",
"messages": [{ "role": "user", "content": "Explain tokens." }]
}'Two carriers for the key
The key is presented either as Authorization: Bearer <key> or, as the native API does it, bare in the x-api-key header. Both name the same Kumo key, and either admits the call. An Anthropic SDK sends the second and no Authorization header at all, which is why it is declared on these operations.
Presenting both is legal. Authorization wins: when it is present and well formed, it is the one that authenticates, and the bare x-api-key is not read at all in that case. An ordinary client sends exactly one of the two — the second header shows up only in a hand-built wrapper that layers headers on top of a ready-made SDK.
How the header is written → Check your key →
Request members
modelRequired. The catalog model, under the name the client uses for it.max_tokensRequired. The maximum number of tokens to generate; this surface has no default.messagesRequired. The conversation, oldest turn first; from one up to 512.systemThe system prompt, beside the conversation rather than as a turn of it. Either spelling the protocol defines is accepted: a bare string, or a list of text blocks. A string is carried as the equivalent one-block list.streamWhether to answer as a stream of this protocol's own named events.toolsThis turn's tools; up to 128 declarations.tool_choiceAn object whose `type` is `auto`, `any`, `none` or `tool`. `auto` lets the model choose, `any` requires some declared tool, and `none` forbids a tool call. Only `tool` requires `name`, which must name a declared tool; the other modes reject `name`.temperatureRandomness, from 0 to 1. Carried to the supplier unchanged.top_pNucleus sampling, from 0 to 1. Carried to the supplier unchanged.top_kInteger from 1 to 1048576. Carried unchanged to a compatible supplier; omission stays absent, not zero. A route that cannot represent it refuses it before an upstream call.stop_sequencesSequences whose appearance ends the answer; up to 16. One that fires comes back as `stop_reason: "stop_sequence"`; which of them fired is not published.metadataAccepted and not forwarded: this platform does not describe your user to a supplier.thinkingAccepted and not forwarded.output_configAccepted and not forwarded.context_managementAccepted and not forwarded: there is no server-side context here, and the caller sends its own.A turn carries a role (user, assistant, system) and content — a string, or a list of typed blocks: text, tool_use, tool_result. A system turn is not a turn of the conversation: its text is lifted into the system prompt, after the blocks the system member itself states.
tool_result.content accepts a string or up to 32 blocks of type text or image. Image bytes use a base64 source with media_type and data; URL sources are rejected. Native Messages carries these blocks unchanged. A Chat Completions route cannot represent an image inside a tool result and refuses that result rather than discarding the image.
The response
idThis request's Kumo identity.typeAlways `message`.roleAlways `assistant`.modelThe model the client named.contentThe answer's blocks in order: `text` and `tool_use`.stop_reasonWhy the answer ended: `end_turn`, `max_tokens`, `stop_sequence` or `tool_use`.usage`input_tokens`, `output_tokens`, `cache_read_input_tokens`, `cache_creation_input_tokens`.{
"id": "msg-71c3d0aa4e",
"type": "message",
"role": "assistant",
"model": "<model>",
"content": [{ "type": "text", "text": "Tokens are the pieces of text a model measures input and output in." }],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 16,
"output_tokens": 22,
"cache_read_input_tokens": 0,
"cache_creation_input_tokens": 0
}
}Streaming
"stream": true answers as text/event-stream in the protocol's own named events: message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop. Alongside them arrive ping and error.
A refusal decided before the first event is this protocol's ordinary JSON error. A failure after the first event can only be expressed in the stream — as the error event — and the stream then ends without message_stop. So here a whole answer is marked by message_stop, and by nothing else.
Every event of all three dialects →
Counting tokens
POST /v1/messages/count_tokens takes a subset of the /v1/messages members — model, messages, system, tools, tool_choice, thinking, metadata and context_management — and answers with a deterministic local estimate of the input tokens. It calls no provider, creates no request or reservation, touches no balance and consumes no rate quota — which is what makes it easy to put in front of a large call.
Everything else describes the generation rather than the input, and is refused on this operation: max_tokens, stream, output_config, stop_sequences, temperature and top_p.
curl https://api.kumorouter.com/v1/messages/count_tokens \
-H "x-api-key: $KUMO_API_KEY" \
-H "content-type: application/json" \
-d '{
"model": "<model>",
"system": "Answer in one sentence.",
"messages": [{ "role": "user", "content": "Explain tokens." }]
}'The reply carries two members: input_tokens, the estimate itself, and estimated, which stays true until an exact local tokenizer is available.
{ "input_tokens": 16, "estimated": true }