SDKs
The official OpenAI and Anthropic SDKs, Python and Node: install, a client with the gateway's base URL and the key from the environment, one call. Your code stays your code.
Quick
You do not change SDKs. Install the same official package you always would; two constructor arguments change — the base URL and the key.
# The OpenAI-compatible dialect
pip install openai
npm install openai
# The native Anthropic dialect
pip install anthropic
npm install @anthropic-ai/sdkKeep the key in an environment variable. None of the examples below writes a secret into source — they all read KUMO_API_KEY, so export it first in the shell you run the code from.
export KUMO_API_KEY="kumo_sk_..."A new organization's wallet is empty: a first call without a top-up answers 402 (billing).
The OpenAI SDK
The base URL is https://api.kumorouter.com/v1, including the /v1. The SDK appends the operation path itself.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.kumorouter.com/v1",
api_key=os.environ["KUMO_API_KEY"],
)
answer = client.chat.completions.create(
model="<model>",
max_tokens=128,
messages=[{"role": "user", "content": "Explain tokens in one line."}],
)
print(answer.choices[0].message.content)The output ceiling is required: a request without max_tokens (equivalently max_completion_tokens) is refused rather than left unbounded.
The Anthropic SDK
The base URL is https://api.kumorouter.com, the bare origin, with no `/v1`: the SDK appends /v1/messages itself. The key travels in the x-api-key header, and the SDK puts it there without your help.
import os
from anthropic import Anthropic
client = Anthropic(
base_url="https://api.kumorouter.com",
api_key=os.environ["KUMO_API_KEY"],
)
answer = client.messages.create(
model="<model>",
max_tokens=128,
messages=[{"role": "user", "content": "Explain tokens in one line."}],
)
print(answer.content[0].text)Not every model answers on both dialects. Which protocols a given model supports is stated on its card in the catalog, beside its capabilities. Read the card before you pick an SDK.
How the key header is written → What a model can do →
Streaming
Both SDKs stream, and the gateway answers each in the shape of its own protocol: events in arrival order, then a terminating frame. There is one difference from a unary call — a failure after the stream has begun cuts it off with no terminating frame, and such an answer counts as a failed call, not a short one.
Libraries built on the OpenAI client
Frameworks such as LangChain or LlamaIndex usually reach models through that same OpenAI client. If a library lets you configure the OpenAI client's base URL, it works with the gateway exactly as a direct call does: set https://api.kumorouter.com/v1 and the key, and leave everything else alone.
The exact parameter name differs per library, and it is deliberately not named here: names are renamed between releases, and a wrong name looks like a setting that was silently ignored. Check your library's own documentation — what you are looking for is the place where the OpenAI client's base URL is set.
The easiest proof that the setting took effect is the console log: the call is there, so the library went through the gateway.