Skip to contentKumoDocs
Sections
On this page
Integrations

SDKs

The official OpenAI and Anthropic SDKs, Python and Node: install, a client with the gateway's base URL and the key from the environment, one call. Your code stays your code.

View as Markdown

Quick

You do not change SDKs. Install the same official package you always would; two constructor arguments change — the base URL and the key.

# The OpenAI-compatible dialect
pip install openai
npm install openai

# The native Anthropic dialect
pip install anthropic
npm install @anthropic-ai/sdk

Keep the key in an environment variable. None of the examples below writes a secret into source — they all read KUMO_API_KEY, so export it first in the shell you run the code from.

export KUMO_API_KEY="kumo_sk_..."

A new organization's wallet is empty: a first call without a top-up answers 402 (billing).

The OpenAI SDK

The base URL is https://api.kumorouter.com/v1, including the /v1. The SDK appends the operation path itself.

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.kumorouter.com/v1",
    api_key=os.environ["KUMO_API_KEY"],
)

answer = client.chat.completions.create(
    model="<model>",
    max_tokens=128,
    messages=[{"role": "user", "content": "Explain tokens in one line."}],
)
print(answer.choices[0].message.content)

The output ceiling is required: a request without max_tokens (equivalently max_completion_tokens) is refused rather than left unbounded.

The Anthropic SDK

The base URL is https://api.kumorouter.com, the bare origin, with no `/v1`: the SDK appends /v1/messages itself. The key travels in the x-api-key header, and the SDK puts it there without your help.

import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.kumorouter.com",
    api_key=os.environ["KUMO_API_KEY"],
)

answer = client.messages.create(
    model="<model>",
    max_tokens=128,
    messages=[{"role": "user", "content": "Explain tokens in one line."}],
)
print(answer.content[0].text)
Note

Not every model answers on both dialects. Which protocols a given model supports is stated on its card in the catalog, beside its capabilities. Read the card before you pick an SDK.

How the key header is written → What a model can do →

Streaming

Both SDKs stream, and the gateway answers each in the shape of its own protocol: events in arrival order, then a terminating frame. There is one difference from a unary call — a failure after the stream has begun cuts it off with no terminating frame, and such an answer counts as a failed call, not a short one.

Streaming in full →

Libraries built on the OpenAI client

Frameworks such as LangChain or LlamaIndex usually reach models through that same OpenAI client. If a library lets you configure the OpenAI client's base URL, it works with the gateway exactly as a direct call does: set https://api.kumorouter.com/v1 and the key, and leave everything else alone.

The exact parameter name differs per library, and it is deliberately not named here: names are renamed between releases, and a wrong name looks like a setting that was silently ignored. Check your library's own documentation — what you are looking for is the place where the OpenAI client's base URL is set.

The easiest proof that the setting took effect is the console log: the call is there, so the library went through the gateway.