Skip to contentKumoDocs
Sections
On this page
Account and billing

Usage and logs

Consumption over a window, breakdowns by model, key and project, the request log, and how to find one particular call in it.

View as Markdown

Find one call

  1. Take the identifier from the answerEvery answer of both runtimes carries an `X-Request-Id` header.
  2. Open the request logFind the call by time, model and key.
  3. Open the rowThe call's card carries its identifier, its tokens and its latency.

The identifier is always the platform's own: it never echoes back what you sent, so the value in the header is a real lookup key for an entry in the platform log. That is the thing to attach to a message to support.

What a refusal looks like → What to send support →

Activity

The activity screen shows settled consumption over a window: exact token quantities by day, the totals over the window, and — as a separate figure — how many calls are still in flight holding a reservation. Reservations are never added into the totals: until a call settles, its cost is an estimate rather than a fact.

There are four windows: the last 24 hours, 7 days, 30 days and 90 days. They are trailing windows, not calendar months.

The window narrows on three axes — project, key, model. The first two are a real narrowing of the read: the chosen projects resolve into their keys, the window is read again against them, and everything on the screen — the daily chart included — becomes about those alone. A project holding no key can narrow nothing, and the screen says so rather than showing you an empty result.

What is counted

TokensThe five dimensions separately: input at the uncached rate, input from cache, input into cache, output, and image units.
Spent walletDollars debited from the wallet. Zero when nothing was wallet-funded.
RequestsSettled calls in the window.

Cached input tokens are a figure of their own and are never folded into the uncached one: they are charged at a different rate, and adding them would hide exactly the saving the cache was turned on for.

Spend is printed in two units and stays in two: package calls in package tokens, wallet calls in dollars, and the daily chart draws one of them at a time. That is not an interface inconvenience but the absence of a rate between them.

The breakdowns are by model, by key and by project, busiest first.

The request log

The log is the organization's immutable history of calls, newest first, with each call's confirmed token quantities and what it actually cost.

ModelThe model name as the caller wrote it, with an outcome mark beside it.
Tokens costPrompt and completion, plus the cost — in the unit the call was charged in.
KeyThe key the call was admitted with. An erased key has no name left, and the cell empties.
TimeThe instant of the call and its latency.

A refused call prints its reason in place of tokens and cost.

The outcome filter offers all, answered and refused. Answered calls are requested from the platform; refusals are picked out of what has already been loaded, because a refusal has six different endings and the platform takes one at a time.

An opened call shows nine fields: time, model, key, prompt tokens, cached prompt tokens, completion tokens, latency, cost, and the call ID.

What the log does not hold

A call is recorded by its shape and never by its content: no prompt, no answer, no headers, no client address. They are not hidden from the screen — they are not stored at all, so there is nothing there that could be shown.

That has a practical consequence worth planning for: neither you nor support can reconstruct what you asked from the log. If the content of your requests matters to your own investigations, record it on your side.

Money movements → How keys are grouped →