> ## Documentation Index
> Fetch the complete documentation index at: https://docs.aixy-gateway.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Pricing and cost attribution

> Understand price precedence, configure provider cost contracts and the price book, and verify estimates before enforcing budgets.

Aixy records the USD cost of model usage and the evidence behind that amount. By default, it uses
**the amount reported in the provider response, then a configured price-book rule, then an automatic
estimate from the provider's price catalog or models.dev**.

Both price-book calculations and catalog calculations are **estimates**, including rates taken from
a negotiated contract or an official provider price list. A provider-reported amount is stronger
per-request evidence, but it is still distinct from reconciliation against authoritative billing.
This page describes inference cost attribution, not the price of an Aixy subscription.

## How Aixy chooses the cost

The default **Hybrid** strategy follows this order for each observed upstream attempt:

```mermaid theme={null}
flowchart TD
  A{"Trusted response amount available?"} -->|Yes| B[Use provider-reported amount]
  A -->|No| C{"Matching price-book rule?"}
  C -->|Yes| D[Estimate using that rule]
  C -->|No| E{"Provider reference price available?"}
  E -->|Yes| F[Estimate using provider catalog]
  E -->|No| G{"models.dev reference price available?"}
  G -->|Yes| H[Estimate using models.dev]
  G -->|No| I[Leave cost unattributed]
```

1. **Provider response:** accept a valid per-request USD amount only from the response field trusted
   by the effective provider credential's cost contract. Token counts alone are not a reported cost.
2. **Price book:** select an effective organization rule first, then a platform rule. Use its rates
   with the observed usage. The price book is optional; a matching rule takes priority over catalogs.
3. **Provider catalog:** use a versioned reference price from an implemented provider pricing source.
   Aixy currently has dedicated sources for regional Amazon Bedrock public offers and the native
   Charm Hyper catalog at reviewed Hyper endpoints.
4. **models.dev:** use the provider/model entry in the shared metadata catalog when no provider
   reference price is available. Coverage and rate dimensions depend on that entry.

The diagram describes source selection, not a guarantee of complete accounting. An incomplete
response or missing billable rate can still produce `partial`. Catalog sources use persisted
versions and caches; Aixy does not make a live price request for every inference.

<Warning>
  Aixy selects one source. It does not add a price-book charge on top of a provider-reported amount,
  or fill a selected rule's missing rates from another source. An incomplete organization rule can
  therefore override a complete platform or catalog price and leave the result partial.
</Warning>

The same rule applies between reference sources: a provider price with missing cache rates is not
completed using models.dev. Once a provider-catalog version exists for a target, a temporary source
outage retains it instead of replacing it with a models.dev estimate.

## Read the amount and its evidence together

| Status | Meaning |
| - | - |
| `provider_reported` | A trusted response supplied the amount and the observation completed. Source: `provider_response`. |
| `estimated` | Available usage was fully calculated under the selected versioned rates. Price-book sources are `contract`, `manual`, or platform `official`; reference sources are `provider_catalog` or `models.dev`. |
| `partial` | An amount was recorded, but the observation is incomplete, a used rate is missing, or observed billable work is unsupported. Its source still identifies where the amount came from. |
| No status / unattributed | No usable amount could be assigned. Inspect the recorded reason; do not interpret this as free usage. |

**Source and status answer different questions.** `official` means the rate came from a platform
price-book entry labelled official; its calculated cost is still `estimated`. A provider-catalog
rate is also estimated, despite coming directly from the provider. A response amount observed on an
incomplete request retains source `provider_response` with status `partial`.

An explicit zero rate or response amount is valid. A blank rate, absent token count, missing amount,
or unavailable analytics source is not evidence of zero cost. Even a partial amount of zero must be
read with its status. Positive dashboard totals below one cent display as **less than \$0.01**;
decimal amounts remain available for aggregation and request investigation.

## Configure the provider cost contract

Open **Providers**, choose the organization or project connection, and use its credential row's
pencil to open **Edit provider**, then review **Cost tracking**. You can save cost changes while
keeping the saved credential, or use **Replace credential** to update both together. These settings
also appear when adding a provider. Configure the credential marked **Effective** for the traffic
you want to measure: a project credential overrides the organization's credential and uses its own
cost contract.

Owners and Org admins can manage organization connections and project connections. Project admins
can change connections only in their assigned projects. Saving increments the cost-contract version
without requiring a replacement provider secret.

### Select the strategy

| Dashboard strategy | Behavior |
| - | - |
| **Hybrid** (`hybrid`) | Provider response → price book → provider reference price → models.dev. Recommended default. |
| **Provider-reported only** (`provider_reported`) | Accept only a supported response amount. Missing or invalid amounts stay unattributed, even if rates exist. |
| **Price book only** (`pricing_rules`) | Ignore response amounts; use explicit rules, then automatic reference prices. Despite the dashboard label, this mode still permits provider-catalog and models.dev estimates. |
| **Do not attribute cost** (`disabled`) | Leave cost unattributed for this credential. This does not prevent the provider from charging for requests. |

### Let Aixy manage response formats

New credentials use **Automatic** for cost extraction and streaming usage. Choose your attribution
strategy; Aixy handles the protocol details. Provider-reported only can be saved before any requests
have run. Requests without a valid reported amount remain unattributed in that mode.

Aixy validates cost evidence in each response. Reviewed provider contracts recognize formats such
as OpenRouter's `usage.cost` and the reviewed Charm Hyper endpoints' `usage.cost.usd`. On unknown
endpoints, automatic extraction accepts the supported `usage.cost.usd` shape, which explicitly
denominates the request amount in USD. A scalar `cost` field alone is ambiguous and is ignored.
Negative, non-finite, malformed, or missing amounts are ignored; an explicit valid zero is retained.
Aixy does not convert currencies or execute arbitrary JSONPath expressions.

Automatic streaming settings request final usage on providers known to support it. Other protocols
retain their native usage handling. Unknown endpoints are observed without adding an unverified
streaming parameter, making probe calls, or retrying billable requests. A missing usage field in one
response does not establish that a provider never supports it.

The form describes the configured behavior. **Activity** shows actual cost, source, token usage, and
coverage after requests run. No response is validated merely by saving the credential. Missing
terminal evidence, malformed usage, interruptions, and client cancellation can leave accounting
partial even when the client receives useful output. See [Stream responses](/guides/streaming).

### Advanced protocol settings

Existing credentials keep their saved settings. Select **Use automatic settings** in the credential's
cost configuration to adopt automatic handling without replacing the secret.

For a custom integration, expand **Advanced protocol settings** to override cost extraction or
streaming usage. Select `usage.cost` or `usage.cost.usd` only after confirming that it is a per-request
USD amount. **Ignore provider cost amounts** disables response amounts; with Provider-reported only,
that leaves requests unattributed. **Request OpenAI final usage** adds `stream_options.include_usage`
for Chat Completions streams. **Do not add a usage request** leaves client-supplied parameters and
native protocol usage handling unchanged. Overrides do not create evidence a provider never returns.

## Configure the price book

Use an explicit price when you have negotiated rates, a private or custom connection, an uncovered
model, or public rates that do not describe your billing terms. For automatic estimates alone, no
manual price or dashboard catalog visit is required.

In **Cost control → Price book**, choose **Add custom price**. Owners and Org admins can create
organization rules and deactivate them; Project admins can inspect the price book but cannot edit
it. Platform rules are maintained by the Aixy operator and cannot be edited from an organization's
price book.

1. **Provider:** enter the canonical Aixy provider identifier, such as `openai` or `bedrock-runtime`.
2. **Model or prefix:** enter the exact upstream model ID, preferably a versioned ID. For a gateway
   model `bedrock-runtime/openai.gpt-oss-120b-1:0`, use provider `bedrock-runtime` and model
   `openai.gpt-oss-120b-1:0`. Use the actual upstream target rather than a virtual routing name. A single trailing `*` matches a
   prefix; other wildcard forms are unsupported.
3. **Service tier:** leave blank for a rule that can match any tier, or enter the exact named tier
   whose rates you reviewed. Nondefault tiers need an exact tier rule for hard-budget admission.
4. **Token rates:** enter USD **per one million tokens**, separately for uncached input, cache reads,
   cache writes, and output. Do not paste a per-token or per-thousand-token rate without converting
   its units. Leave an unknown rate blank; enter zero only when that category is known to be free.
   For cache reads/writes, use **Charged at the input rate** only when those tokens are billed at
   the input price; use **Not applicable to this model** only when that token class cannot occur.
   These declarations need an exact model ID and supporting evidence URL.
5. **Per-request charge:** optionally add a flat USD charge per matched upstream attempt. It is
   additional to this rule's token calculation, not a percentage or markup on reported cost.
6. **Price source and Evidence URL:** select **Provider contract** for negotiated terms or **Manual
   operating estimate** for other operating rates. Retain an HTTPS link to the supporting evidence;
   do not put credentials in the URL. Then select **Create price version**.

At least one rate is required to save. Saving a rule does not prove it covers every billable
dimension. A per-request-only rule can still be partial when observed token usage has no rates.
To apply your own rates even when a provider reports an amount, select **Price book only** on the
effective credential.

<Warning>
  Organization rules apply across that organization's projects and connections matching the same
  provider, model, and tier. They are not scoped to a project, API key, endpoint, or AWS region.
  A project credential override does not create a separate price book. Before overriding prices
  for traffic spanning regions or different contracts, verify that one rule is appropriate for all
  matching traffic; automatic reference histories do distinguish connection endpoints.
</Warning>

### Understand matching and replacements

Only rules effective at the upstream attempt's start time are eligible. Among matching rules, the
resolver prioritizes, in order:

1. Organization scope over platform scope.
2. An exact service-tier match over a rule with no tier restriction.
3. An exact model over a prefix, then a longer matching prefix over a shorter one.
4. The later effective start if the preceding criteria tie.

Consequently, an organization prefix can override an exact platform rule, and a tier-specific
prefix can override an unrestricted exact model within the same scope. A blank tier is a wildcard,
not an alias for the provider's `standard` tier. An explicitly named tier matches only that value.

Create another price with the same provider, model pattern, and tier to replace an organization
rule. Aixy appends a version and closes the earlier effective window. The dashboard uses the save
time; the generated API reference documents explicit effective dates and history retrieval. A new
version must start after the latest version for that same scope. Deactivating a rule stops future
matches and may expose a lower-priority rule or reference fallback. Neither operation recalculates
costs already emitted.

### Example calculation

These are illustrative rates, not a provider quotation. Assume a complete response reports 750
uncached input tokens, 250 cache-read tokens, 100 cache-write tokens, and 100 output tokens:

| Component | Example rate | Calculated USD |
| - | - | - |
| Uncached input | \$2 / 1M tokens | \$0.00150 |
| Cache reads | \$0.20 / 1M tokens | \$0.00005 |
| Cache writes | \$5 / 1M tokens | \$0.00050 |
| Output | \$8 / 1M tokens | \$0.00080 |
| Per-request charge | \$0.001 / attempt | \$0.00100 |
| **Total** | | **\$0.00385 estimated** |

The calculation is `sum(tokens × rate / 1,000,000) + per-request charge`, using decimal arithmetic.
Aixy normalizes supported token formats so cached input is not charged again as uncached input.
Reasoning counts are retained when reported; the price book has no separate reasoning-rate field.
Confirm whether reasoning is already included in the provider's output billing.

If the cache-write rate were missing, the calculable amount would be **$0.00335 partial**; Aixy would
not fill that rate from a catalog. If Hybrid instead received a trusted USD amount of `0.004` on a
complete response, it would record **$0.004 provider\_reported**, without adding the rule's charges.

## Automatic reference prices and history

Reference estimates are stored separately from the editable price book. Their history is keyed by
the exact provider, model, and a fingerprint of the connection endpoint. Regional or custom
connections therefore do not share price versions accidentally. This isolation preserves which
source was used; it does not establish that a generic public price matches a private contract.

Bedrock's source selects supported public on-demand token rates for the configured region. Charm
Hyper's native source supplies input, output, cache-read and cache-write rates for reviewed Hyper
endpoints. If a dedicated source is absent or has no price, Aixy consults
[models.dev](https://models.dev), including its `amazon-bedrock` alias for Aixy's `bedrock-runtime` provider.
These public lookups do not require tenant billing credentials or send inference keys to catalogs.

Active targets become eligible for refresh after **six hours**, driven by exported usage,
hard-budget preparation, or Playground preparation. Previously watched targets are also refreshed
in the background after expiry. Repeated manual refreshes are bounded to protect source availability.
Unchanged rates and provenance retain the same version; changes append a version. Failed refreshes
retry after a short delay, normally one minute, while retaining the last known estimate. A stale
source cache does not count as a fresh confirmation.

The first observation establishes an estimation baseline that can cover the request which caused
it to be fetched. It is not proof of what the provider charged before observation. Later versions
start at their source observation time because these catalogs do not provide a reliable historical
change date. Attribution selects the version effective when the upstream attempt started, including
streams that finish after a price change. Events before the initial baseline can remain
unattributed, and already emitted events are not repriced with today's catalog.

Automatic references do not support nonstandard service tiers, flattened context-dependent tiers,
or separately priced reasoning. These cases report `reference_pricing_unsupported`; use an explicit
rule only if its supported dimensions faithfully represent the relevant tariff. An explicit rule
does not add support for arbitrary meters. Observed audio charges, server-side tools, and one-hour
cache creation can make rate-based accounting partial. Public estimates do not establish private
discounts, credits, taxes, or charges for work absent from the observed usage.

Pricing persistence, catalog requests, and ordinary usage enrichment run outside model forwarding.
Reports update asynchronously. Catalog or management database availability is not a synchronous
dependency of ordinary inference; hard-budget admission has additional requirements below.

## Prices and hard budgets

A response amount arrives too late to authorize spend. Before calling a provider, a **hard** budget
needs a conservative reservation based on an explicit rule or a prepared reference price, even
when the credential normally uses provider-reported cost. A **monitor** budget reports usage and
warnings without blocking requests.

* Provide input/output rates and all applicable cache rates, plus a positive output-token bound
  such as `max_tokens`, `max_completion_tokens`, or `max_output_tokens`. When several bounds are
  supplied, all must be valid and the largest is used.
* Cache billing is explicit pricing data. Enter a separate numeric rate, or declare that a cache
  token class is **Charged at the input rate** or **Not applicable to this model**. Blank rates
  remain unknown and cannot authorize spend. Declarations require an exact model ID and an
  evidence URL. Do not infer free or unsupported caching from a missing catalog field.
  New models with complete applicable prices do not need a gateway update. Unexpected tokens in
  a class declared not applicable leave accounting partial.
* Reference prices not confirmed within **24 hours** cannot authorize new spend. They can still
  support historical estimates. A cold or stale target queues preparation and can return
  `hard_budget_pricing_warming` before any upstream call; retry after preparation succeeds. An
  explicit complete rule avoids catalog preparation.
* Admission supports locally bounded text requests. Stored conversation references, files,
  multimodal work, multiple candidates, server-side tools, and one-hour cache writes cannot be
  bounded by this contract and are rejected under hard budgets.

Reservations use a conservative request-size and output-limit bound, rounded up to micro-USD; they
can exceed the eventual token-based cost. Each applicable hard limit must have enough capacity.
Complete accounting can replace the reservation with the attributed amount. Incomplete or missing
accounting retains conservative liability; a timeout does not make uncertain spend free. Disabling
cost attribution does not disable hard budgets or prevent reservations from remaining unresolved.

Routing fallbacks can run several upstream attempts, each with its own provider, model, price and
usage evidence. Reporting counts observed attempts, not unique prompts. The budget reserves for
possible attempts and retains a conservative total when multiple attempts run. Do not compare one
successful completion's cost with the whole routed request's reservation as if they were equivalent.
See [Routing](/guides/routing).

## Resolve a Playground pricing warning

The Playground checks price coverage when you select a model, key or protocol. It checks all
potential targets of a routing model. This does not invoke the provider or reserve budget. The
final request still goes through normal admission, including balance and supported-input checks.

If prices are incomplete, the notice identifies the model, missing billing information and source.
Use **Recheck prices** after a source update or configuration change. Your message remains in place.
Pricing administrators can choose **Complete pricing** to open Price Book in another tab. Known
rates are loaded when available; review them, fill the missing rates or documented cache billing
declarations, and save the new version. Return to the Playground and recheck. Members without
pricing permission can select another model or ask an administrator to complete the entry.

An absent cache-write price does not prove there is no cache-write charge. If neither the catalog
nor provider documentation establishes the billing behavior, the model remains blocked under a
hard budget. Do not fill invented zero prices or disable a hard budget to bypass the warning.
Unsupported context tiers and separately priced reasoning still require a faithfully representable
pricing contract; the form does not add support for arbitrary charges.

Actual hard-budget admission rejections are recorded in Activity with their request ID, zero
upstream attempts and zero provider cost. A price-coverage check itself is not a model request.

## Verify configuration and investigate gaps

Before relying on totals or enabling a hard budget:

1. Confirm the effective credential, strategy, trusted response field, upstream model ID, connection
   region, and service tier.
2. Send bounded ordinary and streaming test requests. Exercise cache reads/writes and routing
   fallbacks if your application uses them.
3. In **Activity**, open each attempt's **Cost attribution** details. Review **Amount**, **Evidence**,
   **Source**, **Pricing rule** and **Contract**. Reference estimates also use the pricing-rule ID
   and version fields; their absence from the editable price book is expected.
4. Check the arithmetic against observed token categories and the correct rates, then review
   attribution coverage in **Cost control**. Include partial and unattributed counts when assessing
   spend, rather than relying only on the displayed total.
5. For hard budgets, verify admission and settlement separately. Prepare pricing before sending
   production traffic and confirm that completed test requests release excess reservation capacity.

| Symptom or reason | What to check |
| - | - |
| `provider_cost_missing` | Provider-reported only is selected, but no valid configured USD field was observed. Check the provider contract and streaming response. |
| `provider_cost_and_pricing_missing` / `pricing_rule_missing` | Check rule scope, upstream model, tier, effective date, and reference coverage or availability. |
| `usage_missing` | A price exists but no usable usage or priced per-request charge could produce an amount. Rates cannot reconstruct missing token usage. |
| `attribution_disabled` | Enable an appropriate strategy on the effective credential if cost reporting is intended. |
| `reference_pricing_unsupported` | Inspect the service tier, context-dependent pricing, or separate reasoning rate. Supply a representative explicit rule only when its dimensions suffice. |
| `partial` | Inspect stream completion, missing input/output usage, used cache categories with missing rates, and unsupported meters. A different fallback is not automatically selected. |
| `hard_budget_pricing_warming` | Wait for preparation, then retry. Persistent failures require checking catalog coverage/freshness or configuring a complete explicit price. |
| `hard_budget_pricing_incomplete` | Supply all required rates and an exact nondefault tier rule; inspect unsupported reference dimensions. |
| `hard_budget_output_limit_required` | Set a valid positive output-token bound for your protocol. |

Customer [telemetry exports](/guides/telemetry-exports) carry the amount, status, source, rule version,
cost-contract version, attribution mode and unattributed reason when applicable. Retain this
provenance with exported usage. Organization, project, team, user and API-key totals are different
views of the same events; adding those views together double-counts spend. Dashboard totals cover
retained analytics, so retain a separate export for accounting periods longer than your retention.

## Compare with provider billing

Use the provider's billing records for invoice totals. When comparing them with Aixy, align the
provider account, project scope, time period, currency, and billed dimensions. Include provider
adjustments such as credits and discounts in the comparison.

Price-book calculations, reference estimates, and settled reservations describe recorded requests;
they are not invoice totals. Preserve their source and quality evidence alongside exported usage.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.