Claude Haiku 4.5
Anthropic Claude Haiku 4.5 text model. Available over the Anthropic Messages and Chat Completions protocols.
Overview
200Ktokens
64Ktokens
0.6 USDper 1M tokens
3 USDper 1M tokens
15 October 2025
Start with Claude Haiku 4.5
The model name is already filled in. Mint a key in the console, put it in an environment variable, and the call below runs as it stands — provided the organization’s wallet holds funds: a call with nothing to pay with answers 402.
curl https://api.kumorouter.com/v1/chat/completions \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-haiku-4-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Explain tokens in one line."
}
]
}'The Anthropic-shaped call
curl https://api.kumorouter.com/v1/messages \
-H "x-api-key: $KUMO_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-haiku-4-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Explain tokens in one line."
}
]
}'OpenAI — API parameters
Authorization: Bearer $KUMO_API_KEY
| Parameter | Type | Default / range | Description |
|---|---|---|---|
| model | string | Required | Published model identifier. |
| messages | array | 1–512 | Conversation messages in chronological order. |
| max_tokens | integer | ≥ 1 | Optional. A request stating neither max_tokens nor max_completion_tokens is answered under the surface default of 32768 output tokens. Model limits also apply. |
| temperature | number | 0–2 | Sampling randomness. Omission leaves the choice to the model. |
| top_p | number | 0–1 | Nucleus sampling probability mass. Omission leaves the choice to the model. |
| stop | array | ≤ 4 | Array of strings that stop generation. |
| max_completion_tokens | integer | ≥ 1 | Optional alternative to max_tokens. Stating both with DIFFERENT values is refused: a precedence rule would silently discard one of two numbers the caller deliberately wrote. |
| stream | boolean | false | Stream response events using Server-Sent Events. |
| stream_options | object | {include_usage: boolean} | Only with stream: true. include_usage adds a final usage event before [DONE]; its default is false. |
| tools | array | ≤ 128 | Function declarations. parameters carries the original JSON Schema object, up to 65536 bytes per tool, including nested objects and arrays. |
| tool_choice | string | object | auto | none | required | {type: "function", function: {name}} | auto leaves the choice to the model; none forbids a call while the declarations stay visible; required requires a call to some declared tool; the object form requires the named one. Omission leaves the choice to the model. |
One completion per request. The gateway rejects frequency_penalty, presence_penalty, seed, logprobs, top_logprobs and logit_bias: each of them would change the generation or the shape of the reply, and accepting one silently would misdescribe the answer you got. n is accepted only at 1 — the protocol default and exactly what the gateway does — and an n above one is rejected. user is accepted but is not forwarded to the model, not stored and not echoed back. The same applies to strict on a function declaration and name on a tool-role message: both are accepted and not forwarded. reasoning_effort is accepted and not forwarded either: no supplier request on this platform carries a reasoning budget, so the answer comes back at the default of the model itself whatever the field says. cache_control on a message, on a content part or on a replayed tool call is accepted and not forwarded either: the supplier wire this protocol is served by has no member for a prompt-cache anchor, so the cached-token counts come back as though the anchor had not been sent. messages[].content is accepted both as a string and as a list of parts: text parts are joined in order into one text, so a request written in parts and the same request written as a string are one request. A part of kind image_url, input_audio, file or refusal is rejected naming its own type: no multimodal capability is verified on this platform for any pair, and a replayed refusal would reach the supplier as ordinary assistant speech. The developer role is accepted and forwarded as system — it is the newer name this protocol gives the instruction role. stop is array-only; tool_choice is the string mode auto, none or required, or an object naming a declared tool. json_object is refused by name: promising valid JSON says nothing about its shape.
Anthropic — API parameters
x-api-key: $KUMO_API_KEY
| Parameter | Type | Default / range | Description |
|---|---|---|---|
| model | string | Required | Published model identifier. |
| messages | array | 1–512 | Conversation messages in chronological order. |
| max_tokens | integer | 1–1048576 | Required output token ceiling, subject to model limits. |
| temperature | number | 0–1 | Sampling randomness. Omission leaves the choice to the model. |
| top_p | number | 0–1 | Nucleus sampling probability mass. Omission leaves the choice to the model. |
| stop_sequences | array | ≤ 16 | Array of strings that stop generation. |
| system | string | array | — | System instruction as a string or text blocks. |
| top_k | integer | 1–1048576 | Keep the K most likely tokens when sampling. Forwarded unchanged; omission is not replaced by zero. |
| stream | boolean | false | Stream response events using Server-Sent Events. |
| tools | array | ≤ 128 | Tool declarations. input_schema carries the original JSON Schema object, up to 65536 bytes per tool. |
| tool_choice | object | {type: auto | any | none | tool} | auto leaves the choice to the model; any requires a call; none forbids calls; tool selects a named tool. name is required only for tool and forbidden in other modes. |
Uses the native Messages format. OpenAI fields frequency_penalty, presence_penalty, n, seed and response_format do not belong to it. metadata, thinking, output_config and context_management are accepted but not forwarded to the model and are outside verified support. The same applies to tool declaration fields eager_input_streaming, strict, defer_loading, allowed_callers, input_examples and max_uses; server-side tools are unsupported. Tool input_schema is carried as the original JSON Schema object. cache_control: {type: "ephemeral"} is supported on text blocks and tool declarations; cache availability depends on the model.
Identifiers
Capabilities
| Protocol and modality | Streaming | Tools | Structured output | Strict semantics | Structured streaming |
|---|---|---|---|---|---|
| Anthropic Messages text generation | Yes | Yes | Unverified | Unverified | Unverified |
| Chat Completions text generation | Yes | Yes | Unverified | Unverified | Unverified |
| Chat Completions text generation | Yes | Unverified | Unverified | Unverified | Unverified |
| Chat Completions text generation | Unverified | Yes | Unverified | Unverified | Unverified |
A row describes one protocol-and-modality pair rather than the model as a whole: a capability is proved on a particular surface, and the answer beside it may differ.
How it compares
| Model | Context | Max output | Input $/M | Output $/M | Modalities |
|---|---|---|---|---|---|
| Claude Haiku 4.5 | 200K | 64K | 0.6 USD | 3 USD | text generation |
| Claude Sonnet 5 | 1M | 128K | 1.2 USD | 6 USD | text generation |
| Claude Opus 4.8 | 1M | 128K | 3 USD | 15 USD | text generation |
| Claude Fable 5 | 1M | 128K | 6 USD | 30 USD | text generation |
| Claude Fable 5.1 | 1M | 128K | 6 USD | 30 USD | text generation |
| Claude Opus 4.5 | 200K | 64K | 3 USD | 15 USD | text generation |
| Claude Opus 4.6 | 1M | 128K | 3 USD | 15 USD | text generation |
| Claude Opus 4.7 | 1M | 128K | 3 USD | 15 USD | text generation |
| Claude Opus 5.5 | 1M | 128K | 2.2 USD | 11 USD | text generation |
| Claude Opus 5 | 1M | 128K | 3 USD | 15 USD | text generation |
| Claude Sonnet 4.5 | 1M | 64K | 1.8 USD | 9 USD | text generation |
| Claude Sonnet 4.6 | 1M | 128K | 1.8 USD | 9 USD | text generation |
| Gemini 3 Flash Preview | 1M | 65.5K | 0.3 USD | 1.8 USD | text generation |
| Gemini 3.1 Flash Lite | 1M | 65.5K | 0.2 USD | 0.9 USD | text generation |
| Gemini 3.6 Flash High | 1M | 65.5K | 0.5 USD | 2.3 USD | text generation |
| Gemini 3.7 Flash High | 1M | 65.5K | 0.5 USD | 2.3 USD | text generation |
| Gemini 3.8 Flash High | 1M | 65.5K | 0.5 USD | 2.3 USD | text generation |
| Codex Auto Review | 1.1M | 128K | 3 USD | 18 USD | text generation |
| GPT-5.3 Codex Spark | 128K | 32K | 1.1 USD | 8.4 USD | text generation |
| GPT-5.5 | 1.1M | 128K | 3 USD | 18 USD | text generation |
| GPT-5.6 Luna | 1.1M | 128K | 0.1 USD | 0.7 USD | text generation |
| GPT-5.6 Sol | 1.1M | 128K | 2.4 USD | 12 USD | text generation |
| GPT-6 Astra | 1.1M | 128K | 6 USD | 30 USD | text generation |
| GPT-5.6 Terra | 1.1M | 128K | 1.2 USD | 7.2 USD | text generation |
| GPT-6 Sol | 1.1M | 128K | 1.1 USD | 5.5 USD | text generation |
| Grok 4.5 | 500K | 500K | 1.2 USD | 3.6 USD | text generation |
| Grok 4.6 | 500K | 500K | 1.2 USD | 3.6 USD | text generation |
| Grok 4.7 | 500K | 500K | 1.1 USD | 3.3 USD | text generation |
| GLM-5.2 | 1M | 131.1K | 0.8 USD | 2.6 USD | text generation |
| GLM-5.3 | 1M | 131.1K | 0.8 USD | 2.6 USD | text generation |
Prices
| Dimension | Rate |
|---|---|
| Input tokens | 0.6 USD per 1M tokens |
| Cache write | 0.8 USD per 1M tokens |
| Cache read | 0.06 USD per 1M tokens |
| Output tokens | 3 USD per 1M tokens |
Integrations
Every link opens its own guide with this model already chosen.