Skip to contentKumoDocs
Sections
On this page
Get started

Models

The platform publishes its own model catalog: what is available right now, how to name it in a request, what every capability means, and how to choose a model for your task.

View as Markdown

What the catalog is

The catalog is the list of models the platform can route right now. It is not written into this documentation and not maintained by hand: the page below reads it from the gateway, exactly as your code would.

curl https://api.kumorouter.com/v1/models

The answer is an OpenAI-shaped list: object: "list" and a data array whose every element is an object: "model". Reading it needs no key.

Not everything the platform holds gets in — only what can actually be called: the model is active and has at least one enabled capability in the published supplier configuration. A model's protocols and modalities are the exact intersection of what the catalog declares and what real supplier access has proven. So the list moves on its own when a supplier channel is withdrawn or added.

canonical_nameThe model's canonical name — what a request names it by.
aliasesThe aliases that lead to this same model.
vendorWho made the model — code, display names, brand color.
context_window_tokensHow many tokens the model accepts in one request. Absent when the catalog publishes no figure.
max_output_tokensHow many tokens it may answer with. Declared separately and never derived from the context window.
released_onThe day the model's maker released it, as `yyyy-mm-dd`. Not the day it was added here.
supplier_countHow many suppliers currently carry the model. Derived, so it moves when a channel is withdrawn.
protocol_codesThe protocols the model is reachable by.
modality_codesThe kind of work: text generation, embeddings, image generation.
capabilitiesWhat the model can do on each pair of protocol and modality — see below.
catalog_revisionA member of the list itself rather than of a model — the catalog revision, which grows on every change.
{
  "object": "model",
  "canonical_name": "<model>",
  "aliases": ["<alias>"],
  "vendor": { "code": "<vendor>" },
  "protocol_codes": ["chat_completions"],
  "modality_codes": ["text_generation"],
  "capabilities": [
    {
      "protocol_code": "chat_completions",
      "modality_code": "text_generation",
      "streaming_status": "proven",
      "tools_status": "unproven",
      "structured_output_mode": "json_schema"
    }
  ]
}

The catalog

Every model on the platform in one table, one row per model. The whole row opens that model's page; the single exception is the chip beside the identifier, which copies the name to the clipboard and leads nowhere.

ModelThe name, and under it the canonical name — exactly the string you will put in the `model` member. The maker stands beside it.
TypeWhat the model does: text, embeddings, images.
ProtocolsWhich protocols it can be called on.
ContextHow many tokens fit into one request together with its answer.
Max outputThe ceiling on the answer, in tokens.
InputThe price of a million input tokens.
OutputThe price of a million output tokens.
ReleasedThe day the platform dated the model's release.

The filters sit above the table. The search box looks at the name, the canonical name and the maker. Four rails follow, each narrowing on the thing it names: the maker, the type, the protocol, and a context window of "at least this much". The rails add up — choose a type and a protocol and you are asking for a model that does both. Pressing the chosen chip again clears it, and "Reset the filters" puts every rail back at once. The counter under them says how many rows are left, out of how many.

A column head sorts the table: the first press ascending, the second descending. A model the platform has published nothing for in that column sinks to the bottom in both directions — a dash means "not published", not "zero".

None of the figures in the table are written into this article. The platform serves the models and everything about them except money as a live catalog, and the rates as a published price list; the page reads both documents every time it is opened. So a model an operator adds or withdraws appears and disappears here on its own. If the price list cannot be read at this moment, the model list stays where it is and the rates stand as dashes.

Model catalog

A live list: the models come from the platform catalog and the rates from the published price list. A row opens the model’s own page; a column head sorts the table.

Loading the catalog

How to name a model

A request names the model as a string in the model member. Either the canonical name or any alias that leads to the same model will do.

  • A canonical name belongs to one model and to no other.
  • An alias is a second name for the same model. At any moment an alias leads to exactly one model; over time it can be pointed elsewhere, and that is the whole point of it — code that names an alias moves with it.
  • Canonical names and aliases live in one namespace: the same string cannot be both, and cannot belong to two models.

A name no live model answers to is refused straight away — the platform says plainly that it has no such model instead of guessing at one.

Note

A key may carry a scope. If that scope lists models, the key may name only those: a scope is a whitelist, not a preference. A call that names something outside the list is refused rather than routed.

What the platform records about a key →

How to read the capabilities

Capabilities are not declared for a model in general. They are declared for each pair of a protocol and a modality, so one model can do more on one protocol than on another.

Protocols

chat_completionsThe OpenAI-shaped conversational operation: messages in, `choices` out.
responsesThe second operation of that ecosystem, with a request and answer shape of its own.
anthropic_messagesThe Anthropic-native shape, understood by clients and SDKs written against that API.
embeddingsVectors for text.
imagesImage generation.

Modalities

A modality is the kind of work rather than the shape of the request: text_generation, embeddings, image_generation. The protocol says how you ask; the modality says what you get. A key's scope restricts modalities, which is why the pair of them decides whether a call is allowed at all.

The three states

Every capability — streaming, tools, structured output, strict semantics, streaming of structured output — stands in one of three states.

StateWhat it meansWhat the gateway does
ProvenThe platform checked it on a real supplier channel.The call goes upstream.
Not provenThe platform has not checked it. Not "it fails", but "we have not proven it".A request that needs the capability is refused before the supplier is contacted.
UnsupportedThe capability is not there.Such a request is refused in the same way.

The difference between "not proven" and "unsupported" is in where it comes from, not in what you see: both are refused alike and refused early. The refusal happens without reaching the supplier and does not change if you repeat the same request — if you need a capability, pick a model that has it proven.

Structured output carries a mode as well: json_object (the answer must be a JSON object) or json_schema (the answer must match your schema); none means there is no mode at all.

How a response schema is stated → Tools →

The model page

Every model in the catalog has a page of its own at /models/<canonical name>. It collects what you would otherwise assemble from several sources:

  • the figures: context window, output ceiling and release date;
  • a quickstart with this model's identifier already substituted — copy the code as it stands;
  • a capability matrix per protocol: what is proven and what is not;
  • the wallet's prices per dimension;
  • links onward to the integration pages with that model chosen.

The catalogue table above leads to these pages, and the documentation search indexes them.

How to choose a model

The catalog answers "what is available"; the choosing is yours. An order that saves time:

  1. Filter by protocolKeep the models that speak the protocol your client speaks. Changing the client costs more than changing the model.
  2. Check the capabilitiesIf you need tools, streaming or a strict response schema, the capability has to be proven rather than merely mentioned.
  3. Compare the volumesThe context window decides how much you can send, the output ceiling how much you can get back. They are two independent figures.
  4. Work out the priceRates are published per million tokens, separately for input and for output, so no single number compares two models.
  5. Check measured performanceThe model card shows request latency, generation speed and availability over the last 24 hours. Missing measurements are not a zero delay or perfect availability.
  6. Note the release dateIt is the day the maker released the model, so it tells you the model's generation, not when it was added here.

Do not choose based on a single request. Take a dozen of your real tasks, run two or three candidates over them, and compare the answers and the spend — the catalog gives you a shortlist, your own run decides.

The first call → Limits →