---
title: Embeddings
description: The POST /v1/embeddings call — a batch of inputs, one base64 vector per input, and what the call costs.
keywords: embeddings, vectors, base64, batch, encoding_format, usage
group: gateway
---

## Quick {#quick keywords="curl, python, node, vector"}

One operation takes a batch of inputs and answers with one vector per input, in the order the inputs were given. `encoding_format` is required and is `"base64"`.

:::code-group
```bash title=curl
curl https://api.kumorouter.com/v1/embeddings \
  -H "Authorization: Bearer $KUMO_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "<model>",
    "encoding_format": "base64",
    "input": ["the first text", "the second text"]
  }'
```
```python title=Python
import base64
import os
import struct

import requests

response = requests.post(
    "https://api.kumorouter.com/v1/embeddings",
    headers={"Authorization": f"Bearer {os.environ['KUMO_API_KEY']}"},
    json={
        "model": "<model>",
        "encoding_format": "base64",
        "input": ["the first text", "the second text"],
    },
    timeout=60,
)
body = response.json()

vectors = []
for entry in sorted(body["data"], key=lambda item: item["index"]):
    raw = base64.b64decode(entry["embedding"])
    vectors.append(struct.unpack(f"<{len(raw) // 4}f", raw))
```
```javascript title=Node
const response = await fetch("https://api.kumorouter.com/v1/embeddings", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.KUMO_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "<model>",
    encoding_format: "base64",
    input: ["the first text", "the second text"],
  }),
});

const body = await response.json();
const vectors = body.data
  .sort((a, b) => a.index - b.index)
  .map((entry) => {
    const bytes = Buffer.from(entry.embedding, "base64");
    return new Float32Array(bytes.buffer, bytes.byteOffset, bytes.length / 4);
  });
```
:::

## Request members {#request keywords="input, encoding_format, model, batch"}

| Member | What it is |
| --- | --- |
| `model` | Required. The catalog model to embed with, under the name the client uses for it. |
| `input` | Required. The batch of inputs: at most 128 members and at most 131,072 bytes in total. An empty input is refused. |
| `encoding_format` | Required, and it is `"base64"`. |

:::note
`"float"` is not served by this surface — neither named outright nor arrived at through the protocol's default. A request asking for float, or asking for nothing and therefore for float, is refused by name rather than answered in another format. That is what keeps the contract positional: a vector cannot be mistaken for the wrong input, and it cannot be quietly dropped.
:::

Streaming is not part of this surface.

## The response {#response keywords="data, embedding, index, usage"}

| Reply member | What it is |
| --- | --- |
| `object` | Always `list`. |
| `data` | One vector per input, in the order the inputs were given. |
| `model` | The model the vectors were produced with. |
| `usage` | `prompt_tokens` and `total_tokens`; on this surface they are equal. |

Every member of `data` carries `object` set to `embedding`, an `index` — the position of its input — and `embedding`: the vector components as base64-encoded little-endian float32 values.

```json title=Response
{
  "object": "list",
  "model": "<model>",
  "data": [
    { "object": "embedding", "index": 0, "embedding": "AACAPwAAAEAAAEBA" },
    { "object": "embedding", "index": 1, "embedding": "AABAwAAAgL8AAAAA" }
  ],
  "usage": { "prompt_tokens": 9, "total_tokens": 9 }
}
```

## The batch and its bounds {#batching keywords="batch, 128, order, index"}

A batch is not a convenience but the way to count: 128 inputs in one call cost less in overhead than 128 calls, and the order of the answer is guaranteed. The bounds are checked before any provider is called, so a batch that exceeds them costs you a refusal rather than money.

Billing counts input tokens only: this surface produces no output tokens and no images, and the ledger records zero for both.

## Which models answer {#models keywords="catalog, protocol, capabilities"}

Only a model whose published catalog entry carries the embeddings protocol answers here. The catalog is the only source of that fact; it also carries the model's names and aliases and what it can do.

> [The model catalog →](/en/models) [The price list →](https://kumorouter.com/en/pricing)

## Next {#next keywords="errors, rate limits, reference"}

:::cards
- [Errors](/en/errors) — the refusal envelope and this operation's status codes.
- [Rate limits](/en/limits) — the request and token ceilings.
- [API reference](/en/api-reference) — the operation member by member, straight from the API description.
:::
