Skip to main content
Aixy records the USD cost of model usage and the evidence behind that amount. By default, it uses the amount reported in the provider response, then a configured price-book rule, then an automatic estimate from the provider’s price catalog or models.dev. Both price-book calculations and catalog calculations are estimates, including rates taken from a negotiated contract or an official provider price list. A provider-reported amount is stronger per-request evidence, but it is still distinct from reconciliation against authoritative billing. This page describes inference cost attribution, not the price of an Aixy subscription.

How Aixy chooses the cost

The default Hybrid strategy follows this order for each observed upstream attempt:
  1. Provider response: accept a valid per-request USD amount only from the response field trusted by the effective provider credential’s cost contract. Token counts alone are not a reported cost.
  2. Price book: select an effective organization rule first, then a platform rule. Use its rates with the observed usage. The price book is optional; a matching rule takes priority over catalogs.
  3. Provider catalog: use a versioned reference price from an implemented provider pricing source. Aixy currently has dedicated sources for regional Amazon Bedrock public offers and the native Charm Hyper catalog at reviewed Hyper endpoints.
  4. models.dev: use the provider/model entry in the shared metadata catalog when no provider reference price is available. Coverage and rate dimensions depend on that entry.
The diagram describes source selection, not a guarantee of complete accounting. An incomplete response or missing billable rate can still produce partial. Catalog sources use persisted versions and caches; Aixy does not make a live price request for every inference.
Aixy selects one source. It does not add a price-book charge on top of a provider-reported amount, or fill a selected rule’s missing rates from another source. An incomplete organization rule can therefore override a complete platform or catalog price and leave the result partial.
The same rule applies between reference sources: a provider price with missing cache rates is not completed using models.dev. Once a provider-catalog version exists for a target, a temporary source outage retains it instead of replacing it with a models.dev estimate.

Read the amount and its evidence together

Source and status answer different questions. official means the rate came from a platform price-book entry labelled official; its calculated cost is still estimated. A provider-catalog rate is also estimated, despite coming directly from the provider. A response amount observed on an incomplete request retains source provider_response with status partial. An explicit zero rate or response amount is valid. A blank rate, absent token count, missing amount, or unavailable analytics source is not evidence of zero cost. Even a partial amount of zero must be read with its status. Positive dashboard totals below one cent display as less than $0.01; decimal amounts remain available for aggregation and request investigation.

Configure the provider cost contract

Open Providers, choose the organization or project connection, and use its credential row’s pencil to open Edit provider, then review Cost tracking. You can save cost changes while keeping the saved credential, or use Replace credential to update both together. These settings also appear when adding a provider. Configure the credential marked Effective for the traffic you want to measure: a project credential overrides the organization’s credential and uses its own cost contract. Owners and Org admins can manage organization connections and project connections. Project admins can change connections only in their assigned projects. Saving increments the cost-contract version without requiring a replacement provider secret.

Select the strategy

Let Aixy manage response formats

New credentials use Automatic for cost extraction and streaming usage. Choose your attribution strategy; Aixy handles the protocol details. Provider-reported only can be saved before any requests have run. Requests without a valid reported amount remain unattributed in that mode. Aixy validates cost evidence in each response. Reviewed provider contracts recognize formats such as OpenRouter’s usage.cost and the reviewed Charm Hyper endpoints’ usage.cost.usd. On unknown endpoints, automatic extraction accepts the supported usage.cost.usd shape, which explicitly denominates the request amount in USD. A scalar cost field alone is ambiguous and is ignored. Negative, non-finite, malformed, or missing amounts are ignored; an explicit valid zero is retained. Aixy does not convert currencies or execute arbitrary JSONPath expressions. Automatic streaming settings request final usage on providers known to support it. Other protocols retain their native usage handling. Unknown endpoints are observed without adding an unverified streaming parameter, making probe calls, or retrying billable requests. A missing usage field in one response does not establish that a provider never supports it. The form describes the configured behavior. Activity shows actual cost, source, token usage, and coverage after requests run. No response is validated merely by saving the credential. Missing terminal evidence, malformed usage, interruptions, and client cancellation can leave accounting partial even when the client receives useful output. See Stream responses.

Advanced protocol settings

Existing credentials keep their saved settings. Select Use automatic settings in the credential’s cost configuration to adopt automatic handling without replacing the secret. For a custom integration, expand Advanced protocol settings to override cost extraction or streaming usage. Select usage.cost or usage.cost.usd only after confirming that it is a per-request USD amount. Ignore provider cost amounts disables response amounts; with Provider-reported only, that leaves requests unattributed. Request OpenAI final usage adds stream_options.include_usage for Chat Completions streams. Do not add a usage request leaves client-supplied parameters and native protocol usage handling unchanged. Overrides do not create evidence a provider never returns.

Configure the price book

Use an explicit price when you have negotiated rates, a private or custom connection, an uncovered model, or public rates that do not describe your billing terms. For automatic estimates alone, no manual price or dashboard catalog visit is required. In Cost control → Price book, choose Add custom price. Owners and Org admins can create organization rules and deactivate them; Project admins can inspect the price book but cannot edit it. Platform rules are maintained by the Aixy operator and cannot be edited from an organization’s price book.
  1. Provider: enter the canonical Aixy provider identifier, such as openai or bedrock-runtime.
  2. Model or prefix: enter the exact upstream model ID, preferably a versioned ID. For a gateway model bedrock-runtime/openai.gpt-oss-120b-1:0, use provider bedrock-runtime and model openai.gpt-oss-120b-1:0. Use the actual upstream target rather than a virtual routing name. A single trailing * matches a prefix; other wildcard forms are unsupported.
  3. Service tier: leave blank for a rule that can match any tier, or enter the exact named tier whose rates you reviewed. Nondefault tiers need an exact tier rule for hard-budget admission.
  4. Token rates: enter USD per one million tokens, separately for uncached input, cache reads, cache writes, and output. Do not paste a per-token or per-thousand-token rate without converting its units. Leave an unknown rate blank; enter zero only when that category is known to be free. For cache reads/writes, use Charged at the input rate only when those tokens are billed at the input price; use Not applicable to this model only when that token class cannot occur. These declarations need an exact model ID and supporting evidence URL.
  5. Per-request charge: optionally add a flat USD charge per matched upstream attempt. It is additional to this rule’s token calculation, not a percentage or markup on reported cost.
  6. Price source and Evidence URL: select Provider contract for negotiated terms or Manual operating estimate for other operating rates. Retain an HTTPS link to the supporting evidence; do not put credentials in the URL. Then select Create price version.
At least one rate is required to save. Saving a rule does not prove it covers every billable dimension. A per-request-only rule can still be partial when observed token usage has no rates. To apply your own rates even when a provider reports an amount, select Price book only on the effective credential.
Organization rules apply across that organization’s projects and connections matching the same provider, model, and tier. They are not scoped to a project, API key, endpoint, or AWS region. A project credential override does not create a separate price book. Before overriding prices for traffic spanning regions or different contracts, verify that one rule is appropriate for all matching traffic; automatic reference histories do distinguish connection endpoints.

Understand matching and replacements

Only rules effective at the upstream attempt’s start time are eligible. Among matching rules, the resolver prioritizes, in order:
  1. Organization scope over platform scope.
  2. An exact service-tier match over a rule with no tier restriction.
  3. An exact model over a prefix, then a longer matching prefix over a shorter one.
  4. The later effective start if the preceding criteria tie.
Consequently, an organization prefix can override an exact platform rule, and a tier-specific prefix can override an unrestricted exact model within the same scope. A blank tier is a wildcard, not an alias for the provider’s standard tier. An explicitly named tier matches only that value. Create another price with the same provider, model pattern, and tier to replace an organization rule. Aixy appends a version and closes the earlier effective window. The dashboard uses the save time; the generated API reference documents explicit effective dates and history retrieval. A new version must start after the latest version for that same scope. Deactivating a rule stops future matches and may expose a lower-priority rule or reference fallback. Neither operation recalculates costs already emitted.

Example calculation

These are illustrative rates, not a provider quotation. Assume a complete response reports 750 uncached input tokens, 250 cache-read tokens, 100 cache-write tokens, and 100 output tokens: The calculation is sum(tokens × rate / 1,000,000) + per-request charge, using decimal arithmetic. Aixy normalizes supported token formats so cached input is not charged again as uncached input. Reasoning counts are retained when reported; the price book has no separate reasoning-rate field. Confirm whether reasoning is already included in the provider’s output billing. If the cache-write rate were missing, the calculable amount would be 0.00335partial∗∗;Aixywouldnotfillthatratefromacatalog.IfHybridinsteadreceivedatrustedUSDamountof‘0.004‘onacompleteresponse,itwouldrecord∗∗0.00335 partial**; Aixy would not fill that rate from a catalog. If Hybrid instead received a trusted USD amount of `0.004` on a complete response, it would record **0.004 provider_reported, without adding the rule’s charges.

Automatic reference prices and history

Reference estimates are stored separately from the editable price book. Their history is keyed by the exact provider, model, and a fingerprint of the connection endpoint. Regional or custom connections therefore do not share price versions accidentally. This isolation preserves which source was used; it does not establish that a generic public price matches a private contract. Bedrock’s source selects supported public on-demand token rates for the configured region. Charm Hyper’s native source supplies input, output, cache-read and cache-write rates for reviewed Hyper endpoints. If a dedicated source is absent or has no price, Aixy consults models.dev, including its amazon-bedrock alias for Aixy’s bedrock-runtime provider. These public lookups do not require tenant billing credentials or send inference keys to catalogs. Active targets become eligible for refresh after six hours, driven by exported usage, hard-budget preparation, or Playground preparation. Previously watched targets are also refreshed in the background after expiry. Repeated manual refreshes are bounded to protect source availability. Unchanged rates and provenance retain the same version; changes append a version. Failed refreshes retry after a short delay, normally one minute, while retaining the last known estimate. A stale source cache does not count as a fresh confirmation. The first observation establishes an estimation baseline that can cover the request which caused it to be fetched. It is not proof of what the provider charged before observation. Later versions start at their source observation time because these catalogs do not provide a reliable historical change date. Attribution selects the version effective when the upstream attempt started, including streams that finish after a price change. Events before the initial baseline can remain unattributed, and already emitted events are not repriced with today’s catalog. Automatic references do not support nonstandard service tiers, flattened context-dependent tiers, or separately priced reasoning. These cases report reference_pricing_unsupported; use an explicit rule only if its supported dimensions faithfully represent the relevant tariff. An explicit rule does not add support for arbitrary meters. Observed audio charges, server-side tools, and one-hour cache creation can make rate-based accounting partial. Public estimates do not establish private discounts, credits, taxes, or charges for work absent from the observed usage. Pricing persistence, catalog requests, and ordinary usage enrichment run outside model forwarding. Reports update asynchronously. Catalog or management database availability is not a synchronous dependency of ordinary inference; hard-budget admission has additional requirements below.

Prices and hard budgets

A response amount arrives too late to authorize spend. Before calling a provider, a hard budget needs a conservative reservation based on an explicit rule or a prepared reference price, even when the credential normally uses provider-reported cost. A monitor budget reports usage and warnings without blocking requests.
  • Provide input/output rates and all applicable cache rates, plus a positive output-token bound such as max_tokens, max_completion_tokens, or max_output_tokens. When several bounds are supplied, all must be valid and the largest is used.
  • Cache billing is explicit pricing data. Enter a separate numeric rate, or declare that a cache token class is Charged at the input rate or Not applicable to this model. Blank rates remain unknown and cannot authorize spend. Declarations require an exact model ID and an evidence URL. Do not infer free or unsupported caching from a missing catalog field. New models with complete applicable prices do not need a gateway update. Unexpected tokens in a class declared not applicable leave accounting partial.
  • Reference prices not confirmed within 24 hours cannot authorize new spend. They can still support historical estimates. A cold or stale target queues preparation and can return hard_budget_pricing_warming before any upstream call; retry after preparation succeeds. An explicit complete rule avoids catalog preparation.
  • Admission supports locally bounded text requests. Stored conversation references, files, multimodal work, multiple candidates, server-side tools, and one-hour cache writes cannot be bounded by this contract and are rejected under hard budgets.
Reservations use a conservative request-size and output-limit bound, rounded up to micro-USD; they can exceed the eventual token-based cost. Each applicable hard limit must have enough capacity. Complete accounting can replace the reservation with the attributed amount. Incomplete or missing accounting retains conservative liability; a timeout does not make uncertain spend free. Disabling cost attribution does not disable hard budgets or prevent reservations from remaining unresolved. Routing fallbacks can run several upstream attempts, each with its own provider, model, price and usage evidence. Reporting counts observed attempts, not unique prompts. The budget reserves for possible attempts and retains a conservative total when multiple attempts run. Do not compare one successful completion’s cost with the whole routed request’s reservation as if they were equivalent. See Routing.

Resolve a Playground pricing warning

The Playground checks price coverage when you select a model, key or protocol. It checks all potential targets of a routing model. This does not invoke the provider or reserve budget. The final request still goes through normal admission, including balance and supported-input checks. If prices are incomplete, the notice identifies the model, missing billing information and source. Use Recheck prices after a source update or configuration change. Your message remains in place. Pricing administrators can choose Complete pricing to open Price Book in another tab. Known rates are loaded when available; review them, fill the missing rates or documented cache billing declarations, and save the new version. Return to the Playground and recheck. Members without pricing permission can select another model or ask an administrator to complete the entry. An absent cache-write price does not prove there is no cache-write charge. If neither the catalog nor provider documentation establishes the billing behavior, the model remains blocked under a hard budget. Do not fill invented zero prices or disable a hard budget to bypass the warning. Unsupported context tiers and separately priced reasoning still require a faithfully representable pricing contract; the form does not add support for arbitrary charges. Actual hard-budget admission rejections are recorded in Activity with their request ID, zero upstream attempts and zero provider cost. A price-coverage check itself is not a model request.

Verify configuration and investigate gaps

Before relying on totals or enabling a hard budget:
  1. Confirm the effective credential, strategy, trusted response field, upstream model ID, connection region, and service tier.
  2. Send bounded ordinary and streaming test requests. Exercise cache reads/writes and routing fallbacks if your application uses them.
  3. In Activity, open each attempt’s Cost attribution details. Review Amount, Evidence, Source, Pricing rule and Contract. Reference estimates also use the pricing-rule ID and version fields; their absence from the editable price book is expected.
  4. Check the arithmetic against observed token categories and the correct rates, then review attribution coverage in Cost control. Include partial and unattributed counts when assessing spend, rather than relying only on the displayed total.
  5. For hard budgets, verify admission and settlement separately. Prepare pricing before sending production traffic and confirm that completed test requests release excess reservation capacity.
Customer telemetry exports carry the amount, status, source, rule version, cost-contract version, attribution mode and unattributed reason when applicable. Retain this provenance with exported usage. Organization, project, team, user and API-key totals are different views of the same events; adding those views together double-counts spend. Dashboard totals cover retained analytics, so retain a separate export for accounting periods longer than your retention.

Compare with provider billing

Use the provider’s billing records for invoice totals. When comparing them with Aixy, align the provider account, project scope, time period, currency, and billed dimensions. Include provider adjustments such as credits and discounts in the comparison. Price-book calculations, reference estimates, and settled reservations describe recorded requests; they are not invoice totals. Preserve their source and quality evidence alongside exported usage.