Skip to contentKumoDocs
Sections
On this page
The gateway

Embeddings

The POST /v1/embeddings call — a batch of inputs, one base64 vector per input, and what the call costs.

View as Markdown

Quick

One operation takes a batch of inputs and answers with one vector per input, in the order the inputs were given. encoding_format is required and is "base64".

curl https://api.kumorouter.com/v1/embeddings \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "encoding_format": "base64",
    "input": ["the first text", "the second text"]
  }'

Request members

modelRequired. The catalog model to embed with, under the name the client uses for it.
inputRequired. The batch of inputs: at most 128 members and at most 131,072 bytes in total. An empty input is refused.
encoding_formatRequired, and it is `"base64"`.
Note

"float" is not served by this surface — neither named outright nor arrived at through the protocol's default. A request asking for float, or asking for nothing and therefore for float, is refused by name rather than answered in another format. That is what keeps the contract positional: a vector cannot be mistaken for the wrong input, and it cannot be quietly dropped.

Streaming is not part of this surface.

The response

objectAlways `list`.
dataOne vector per input, in the order the inputs were given.
modelThe model the vectors were produced with.
usage`prompt_tokens` and `total_tokens`; on this surface they are equal.

Every member of data carries object set to embedding, an index — the position of its input — and embedding: the vector components as base64-encoded little-endian float32 values.

{
  "object": "list",
  "model": "<model>",
  "data": [
    { "object": "embedding", "index": 0, "embedding": "AACAPwAAAEAAAEBA" },
    { "object": "embedding", "index": 1, "embedding": "AABAwAAAgL8AAAAA" }
  ],
  "usage": { "prompt_tokens": 9, "total_tokens": 9 }
}

The batch and its bounds

A batch is not a convenience but the way to count: 128 inputs in one call cost less in overhead than 128 calls, and the order of the answer is guaranteed. The bounds are checked before any provider is called, so a batch that exceeds them costs you a refusal rather than money.

Billing counts input tokens only: this surface produces no output tokens and no images, and the ledger records zero for both.

Which models answer

Only a model whose published catalog entry carries the embeddings protocol answers here. The catalog is the only source of that fact; it also carries the model's names and aliases and what it can do.

The model catalog → The price list →

Next