Features

Run Cost

How Appstrate measures token usage and dollar cost per run, live and after the fact, and what a zero really means.

Every run reports what it consumed. The figures come from one place, a usage ledger written by the platform, so they stay comparable whether the run executed in a platform sandbox or on a remote runner.

What a run exposes

FieldMeaning
costTotal cost in US dollars. It is stored only when greater than zero, so a run with no priced usage has null.
cost_pricing_statusHow much of cost is backed by real rates: priced, partial, unpriced, or null.
token_usageinput_tokens, output_tokens, cache_creation_input_tokens, cache_read_input_tokens.
model_label, model_sourceThe model the run used, and whether it came from the platform (system) or from the organization's own credentials (org).

GET /api/runs/{id} and the run list return them. While the run is in flight, the same figures stream over realtime as run_metric events (costSoFar, tokenUsage, costPricingStatus), throttled per run.

How the cost is computed

Each model call adds a row to a usage ledger. A run's cost is the sum of its rows, cached on the run when it finishes. There is one writer and one read path, so a UI total and an invoice cannot disagree.

A row is priced from four disjoint token buckets, at the rates of the model:

cost = input x input_rate + output x output_rate
     + cache_read x cache_read_rate + cache_write x cache_write_rate

Rates are in US dollars per million tokens. They come from, in order of precedence:

  1. a cost object set on the organization model (an operator override),
  2. the model registry bundled with the runtime, or the price list of OpenRouter for OpenRouter models.

A rate card can define price tiers: when the input of a single request exceeds a threshold, the whole request is repriced at the tier rates.

Where the rows come from depends on how inference reaches the model:

  • API-key models are served through the platform's LLM proxy, which writes one row per call.
  • Subscription models (OAuth credentials of the optional @appstrate/module-claude-code and @appstrate/module-codex modules) are metered from the runner's cumulative token counters, priced server-side. They are priced at the public API rates even though the user pays a flat subscription, so the figure is an estimate of the equivalent API cost.
  • Remote runs report their usage through signed events. When their inference goes through the platform's LLM proxy, the proxy rows are the source of truth.
  • Chat turns are metered per turn and attributed to the chat session, not to a run (see Chat).

The platform prices usage itself and does not trust a figure computed inside the sandbox. Cost data is stored as floating point numbers, so a re-summation can differ from the stored total by sub-cent amounts.

A zero is not always free

A cost of zero, or no cost at all, does not always mean the run was free. cost_pricing_status says which case you are looking at:

StatusMeaning
pricedEvery token bucket that carried usage had a rate. The figure is complete.
partialSome usage, typically cached input, had no rate and counted as zero. The figure is a floor.
unpricedThe model has no rates at all, for example a custom gateway model without a cost override. A 0 here means "not priced", not "free".
nullNo claim. Runs finalized before the field existed, runs that produced no usage rows. Never read null as priced.

The web app withholds the amount of an unpriced run instead of showing $0.00, and marks a partial run as a lower bound. API consumers should do the same. To price a gateway model, set its cost when you create the model (see LLM Models).

Every successful call that reached the provider ends up in the ledger. When the usage of a successful call cannot be parsed, the platform records a row with zero tokens so the call is still countable.

Limits and quotas

cost is reporting. Spending limits are not part of the open-source core. A deployment can load a billing module (the source-available @appstrate/module-ee) that admits or refuses usage before a run or chat turn starts. A refusal answers 402 with quota_exceeded or subscription_blocked.

Reference

The full design, including settlement rules for billing consumers, is in RUN_COST.md.

On this page