Documentation
One key for every model
Change the base URL and the key — the rest of your code stays as it is. The platform itself publishes which models are available, what they can do, and what they cost.
curl https://api.kumorouter.com/v1/chat/completions \
-H "Authorization: Bearer $KUMO_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "<model>",
"max_tokens": 128,
"messages": [{ "role": "user", "content": "Explain tokens in one line." }]
}'Get started
QuickstartThree minutes from the browser to a model’s answer — a key in the console, a base URL in your client, a first call. The same thing in detail below.AuthenticationOne key, the header that carries it, and the three things the platform records about a key: what it is called, what funds it and what it is allowed to reach.Check your keyThe one live operation on this site: paste a key, and the platform reads back what it is and what it may do.ModelsThe platform publishes its own model catalog: what is available right now, how to name it in a request, what every capability means, and how to choose a model for your task.
Integrations
IntegrationsPoint the tool you already use at the gateway: a base URL and a key. The platform's live recipes, step-by-step guides, and a prompt that lets your agent do the setup.Claude CodeThree environment variables and Claude Code runs through the gateway. Plus the settings.json route, the key precedence to watch for, and a one-command check.Codex CLIA provider in config.toml, the key from an environment variable — Codex CLI runs through the gateway without changing a single habit of yours.CursorYour own key and an overridden base URL in Cursor's model settings, and the editor's chat runs through the gateway. Plus an honest list of what does not.VS Code extensionsCline, Roo Code and Continue all take the same OpenAI-compatible provider: base URL, key, model ID. Here is where those fields live in each of them.Other toolsAider, Zed, OpenClaw — and the way to connect a tool that is not on this list: hand the connection details to a coding agent and let it do the setup.SDKsThe official OpenAI and Anthropic SDKs, Python and Node: install, a client with the gateway's base URL and the key from the environment, one call. Your code stays your code.Kumo Usage CLI BarA spend-and-remainder panel under the Claude Code and Codex prompt — one command to install, it finds your key itself, and it shows what is left: a package, the balance, or your own ceiling.
The gateway
Chat CompletionsThe POST /v1/chat/completions call — request members, conversation roles, the response shape, and pointers to streaming, tools and structured output.ResponsesThe POST /v1/responses call — how it differs from Chat Completions, the request members, the typed output items and the stream's terminal event.MessagesThe Anthropic Messages dialect — the bare base URL, two carriers for the key, request members, reply blocks, stream events, and a token count that spends no quota.EmbeddingsThe POST /v1/embeddings call — a batch of inputs, one base64 vector per input, and what the call costs.ImagesThe POST /v1/images/generations call — the prompt, how many images, the size, the response shape, and why generation is counted in images rather than tokens.StreamingSSE across the gateway's three dialects — how to turn a stream on, which events arrive, what marks a whole answer, and why a call is never moved after the first byte.ToolsFunction calling across the gateway's three dialects — declaring a tool, the call-and-result round, and why tool support is proven per model.Structured outputAn answer to a named JSON Schema — what format the API takes, what strict means, how it is spelled on the three dialects, and why support is proven separately for every model.ErrorsThe envelope a refusal arrives in, the status codes the gateway answers with, and the two behaviors every client must handle.Rate limitsThe ceilings a call is measured against, what a call that reaches one gets back, and how a higher ceiling is asked for.API referenceEvery operation of the Kumo gateway, generated from the published API description: the address, how it is authenticated, what the request carries and what comes back.
Account and billing
AccountSigning up, signing in, the profile and the organization — what the platform remembers about an account, what of it you change in the console, and what you cannot change yourself.KeysThe whole life of a key: what is chosen at creation and never again, the one showing of the secret, the switch, the trash, and erasing for good.BillingThe organization wallet, a top-up through a payment provider, the pinned rate, what a charge is built from, and what a call answers on an empty wallet.PackagesOne package per account: a limit in percent of the plan, top-ups and the minimum purchase, which keys spend the package, freezing, notices, and how a package differs from the wallet.Usage and logsConsumption over a window, breakdowns by model, key and project, the request log, and how to find one particular call in it.ProjectsThe label by which keys are grouped among themselves: creating a project, moving a key into one, what its card shows, and how archiving differs from deleting.
Resources
Questions and answersShort answers to what people trip over most, each with the page where it is worked through in full.For AI and agentsThis documentation as Markdown — an index, the whole corpus in one document, every page addressable as its own source, and a control that puts a ready-made prompt on the clipboard.SupportWhere to write, what to attach to a message, where to see the state of the services, and where the legal documents live.
Running an agent? The whole of this documentation is available as Markdown at/llms-full.txt.