How Aixy chooses the cost
The default Hybrid strategy follows this order for each observed upstream attempt:- Provider response: accept a valid per-request USD amount only from the response field trusted by the effective provider credential’s cost contract. Token counts alone are not a reported cost.
- Price book: select an effective organization rule first, then a platform rule. Use its rates with the observed usage. The price book is optional; a matching rule takes priority over catalogs.
- Provider catalog: use a versioned reference price from an implemented provider pricing source. Aixy currently has dedicated sources for regional Amazon Bedrock public offers and the native Charm Hyper catalog at reviewed Hyper endpoints.
- models.dev: use the provider/model entry in the shared metadata catalog when no provider reference price is available. Coverage and rate dimensions depend on that entry.
partial. Catalog sources use persisted
versions and caches; Aixy does not make a live price request for every inference.
The same rule applies between reference sources: a provider price with missing cache rates is not
completed using models.dev. Once a provider-catalog version exists for a target, a temporary source
outage retains it instead of replacing it with a models.dev estimate.
Read the amount and its evidence together
Source and status answer different questions.
official means the rate came from a platform
price-book entry labelled official; its calculated cost is still estimated. A provider-catalog
rate is also estimated, despite coming directly from the provider. A response amount observed on an
incomplete request retains source provider_response with status partial.
An explicit zero rate or response amount is valid. A blank rate, absent token count, missing amount,
or unavailable analytics source is not evidence of zero cost. Even a partial amount of zero must be
read with its status. Positive dashboard totals below one cent display as less than $0.01;
decimal amounts remain available for aggregation and request investigation.
Configure the provider cost contract
Open Providers, choose the organization or project connection, and use its credential row’s pencil to open Edit provider, then review Cost tracking. You can save cost changes while keeping the saved credential, or use Replace credential to update both together. These settings also appear when adding a provider. Configure the credential marked Effective for the traffic you want to measure: a project credential overrides the organization’s credential and uses its own cost contract. Owners and Org admins can manage organization connections and project connections. Project admins can change connections only in their assigned projects. Saving increments the cost-contract version without requiring a replacement provider secret.Select the strategy
Let Aixy manage response formats
New credentials use Automatic for cost extraction and streaming usage. Choose your attribution strategy; Aixy handles the protocol details. Provider-reported only can be saved before any requests have run. Requests without a valid reported amount remain unattributed in that mode. Aixy validates cost evidence in each response. Reviewed provider contracts recognize formats such as OpenRouter’susage.cost and the reviewed Charm Hyper endpoints’ usage.cost.usd. On unknown
endpoints, automatic extraction accepts the supported usage.cost.usd shape, which explicitly
denominates the request amount in USD. A scalar cost field alone is ambiguous and is ignored.
Negative, non-finite, malformed, or missing amounts are ignored; an explicit valid zero is retained.
Aixy does not convert currencies or execute arbitrary JSONPath expressions.
Automatic streaming settings request final usage on providers known to support it. Other protocols
retain their native usage handling. Unknown endpoints are observed without adding an unverified
streaming parameter, making probe calls, or retrying billable requests. A missing usage field in one
response does not establish that a provider never supports it.
The form describes the configured behavior. Activity shows actual cost, source, token usage, and
coverage after requests run. No response is validated merely by saving the credential. Missing
terminal evidence, malformed usage, interruptions, and client cancellation can leave accounting
partial even when the client receives useful output. See Stream responses.
Advanced protocol settings
Existing credentials keep their saved settings. Select Use automatic settings in the credential’s cost configuration to adopt automatic handling without replacing the secret. For a custom integration, expand Advanced protocol settings to override cost extraction or streaming usage. Selectusage.cost or usage.cost.usd only after confirming that it is a per-request
USD amount. Ignore provider cost amounts disables response amounts; with Provider-reported only,
that leaves requests unattributed. Request OpenAI final usage adds stream_options.include_usage
for Chat Completions streams. Do not add a usage request leaves client-supplied parameters and
native protocol usage handling unchanged. Overrides do not create evidence a provider never returns.
Configure the price book
Use an explicit price when you have negotiated rates, a private or custom connection, an uncovered model, or public rates that do not describe your billing terms. For automatic estimates alone, no manual price or dashboard catalog visit is required. In Cost control → Price book, choose Add custom price. Owners and Org admins can create organization rules and deactivate them; Project admins can inspect the price book but cannot edit it. Platform rules are maintained by the Aixy operator and cannot be edited from an organization’s price book.- Provider: enter the canonical Aixy provider identifier, such as
openaiorbedrock-runtime. - Model or prefix: enter the exact upstream model ID, preferably a versioned ID. For a gateway
model
bedrock-runtime/openai.gpt-oss-120b-1:0, use providerbedrock-runtimeand modelopenai.gpt-oss-120b-1:0. Use the actual upstream target rather than a virtual routing name. A single trailing*matches a prefix; other wildcard forms are unsupported. - Service tier: leave blank for a rule that can match any tier, or enter the exact named tier whose rates you reviewed. Nondefault tiers need an exact tier rule for hard-budget admission.
- Token rates: enter USD per one million tokens, separately for uncached input, cache reads, cache writes, and output. Do not paste a per-token or per-thousand-token rate without converting its units. Leave an unknown rate blank; enter zero only when that category is known to be free. For cache reads/writes, use Charged at the input rate only when those tokens are billed at the input price; use Not applicable to this model only when that token class cannot occur. These declarations need an exact model ID and supporting evidence URL.
- Per-request charge: optionally add a flat USD charge per matched upstream attempt. It is additional to this rule’s token calculation, not a percentage or markup on reported cost.
- Price source and Evidence URL: select Provider contract for negotiated terms or Manual operating estimate for other operating rates. Retain an HTTPS link to the supporting evidence; do not put credentials in the URL. Then select Create price version.
Understand matching and replacements
Only rules effective at the upstream attempt’s start time are eligible. Among matching rules, the resolver prioritizes, in order:- Organization scope over platform scope.
- An exact service-tier match over a rule with no tier restriction.
- An exact model over a prefix, then a longer matching prefix over a shorter one.
- The later effective start if the preceding criteria tie.
standard tier. An explicitly named tier matches only that value.
Create another price with the same provider, model pattern, and tier to replace an organization
rule. Aixy appends a version and closes the earlier effective window. The dashboard uses the save
time; the generated API reference documents explicit effective dates and history retrieval. A new
version must start after the latest version for that same scope. Deactivating a rule stops future
matches and may expose a lower-priority rule or reference fallback. Neither operation recalculates
costs already emitted.
Example calculation
These are illustrative rates, not a provider quotation. Assume a complete response reports 750 uncached input tokens, 250 cache-read tokens, 100 cache-write tokens, and 100 output tokens:
The calculation is
sum(tokens × rate / 1,000,000) + per-request charge, using decimal arithmetic.
Aixy normalizes supported token formats so cached input is not charged again as uncached input.
Reasoning counts are retained when reported; the price book has no separate reasoning-rate field.
Confirm whether reasoning is already included in the provider’s output billing.
If the cache-write rate were missing, the calculable amount would be 0.004 provider_reported, without adding the rule’s charges.
Automatic reference prices and history
Reference estimates are stored separately from the editable price book. Their history is keyed by the exact provider, model, and a fingerprint of the connection endpoint. Regional or custom connections therefore do not share price versions accidentally. This isolation preserves which source was used; it does not establish that a generic public price matches a private contract. Bedrock’s source selects supported public on-demand token rates for the configured region. Charm Hyper’s native source supplies input, output, cache-read and cache-write rates for reviewed Hyper endpoints. If a dedicated source is absent or has no price, Aixy consults models.dev, including itsamazon-bedrock alias for Aixy’s bedrock-runtime provider.
These public lookups do not require tenant billing credentials or send inference keys to catalogs.
Active targets become eligible for refresh after six hours, driven by exported usage,
hard-budget preparation, or Playground preparation. Previously watched targets are also refreshed
in the background after expiry. Repeated manual refreshes are bounded to protect source availability.
Unchanged rates and provenance retain the same version; changes append a version. Failed refreshes
retry after a short delay, normally one minute, while retaining the last known estimate. A stale
source cache does not count as a fresh confirmation.
The first observation establishes an estimation baseline that can cover the request which caused
it to be fetched. It is not proof of what the provider charged before observation. Later versions
start at their source observation time because these catalogs do not provide a reliable historical
change date. Attribution selects the version effective when the upstream attempt started, including
streams that finish after a price change. Events before the initial baseline can remain
unattributed, and already emitted events are not repriced with today’s catalog.
Automatic references do not support nonstandard service tiers, flattened context-dependent tiers,
or separately priced reasoning. These cases report reference_pricing_unsupported; use an explicit
rule only if its supported dimensions faithfully represent the relevant tariff. An explicit rule
does not add support for arbitrary meters. Observed audio charges, server-side tools, and one-hour
cache creation can make rate-based accounting partial. Public estimates do not establish private
discounts, credits, taxes, or charges for work absent from the observed usage.
Pricing persistence, catalog requests, and ordinary usage enrichment run outside model forwarding.
Reports update asynchronously. Catalog or management database availability is not a synchronous
dependency of ordinary inference; hard-budget admission has additional requirements below.
Prices and hard budgets
A response amount arrives too late to authorize spend. Before calling a provider, a hard budget needs a conservative reservation based on an explicit rule or a prepared reference price, even when the credential normally uses provider-reported cost. A monitor budget reports usage and warnings without blocking requests.- Provide input/output rates and all applicable cache rates, plus a positive output-token bound
such as
max_tokens,max_completion_tokens, ormax_output_tokens. When several bounds are supplied, all must be valid and the largest is used. - Cache billing is explicit pricing data. Enter a separate numeric rate, or declare that a cache token class is Charged at the input rate or Not applicable to this model. Blank rates remain unknown and cannot authorize spend. Declarations require an exact model ID and an evidence URL. Do not infer free or unsupported caching from a missing catalog field. New models with complete applicable prices do not need a gateway update. Unexpected tokens in a class declared not applicable leave accounting partial.
- Reference prices not confirmed within 24 hours cannot authorize new spend. They can still
support historical estimates. A cold or stale target queues preparation and can return
hard_budget_pricing_warmingbefore any upstream call; retry after preparation succeeds. An explicit complete rule avoids catalog preparation. - Admission supports locally bounded text requests. Stored conversation references, files, multimodal work, multiple candidates, server-side tools, and one-hour cache writes cannot be bounded by this contract and are rejected under hard budgets.
Resolve a Playground pricing warning
The Playground checks price coverage when you select a model, key or protocol. It checks all potential targets of a routing model. This does not invoke the provider or reserve budget. The final request still goes through normal admission, including balance and supported-input checks. If prices are incomplete, the notice identifies the model, missing billing information and source. Use Recheck prices after a source update or configuration change. Your message remains in place. Pricing administrators can choose Complete pricing to open Price Book in another tab. Known rates are loaded when available; review them, fill the missing rates or documented cache billing declarations, and save the new version. Return to the Playground and recheck. Members without pricing permission can select another model or ask an administrator to complete the entry. An absent cache-write price does not prove there is no cache-write charge. If neither the catalog nor provider documentation establishes the billing behavior, the model remains blocked under a hard budget. Do not fill invented zero prices or disable a hard budget to bypass the warning. Unsupported context tiers and separately priced reasoning still require a faithfully representable pricing contract; the form does not add support for arbitrary charges. Actual hard-budget admission rejections are recorded in Activity with their request ID, zero upstream attempts and zero provider cost. A price-coverage check itself is not a model request.Verify configuration and investigate gaps
Before relying on totals or enabling a hard budget:- Confirm the effective credential, strategy, trusted response field, upstream model ID, connection region, and service tier.
- Send bounded ordinary and streaming test requests. Exercise cache reads/writes and routing fallbacks if your application uses them.
- In Activity, open each attempt’s Cost attribution details. Review Amount, Evidence, Source, Pricing rule and Contract. Reference estimates also use the pricing-rule ID and version fields; their absence from the editable price book is expected.
- Check the arithmetic against observed token categories and the correct rates, then review attribution coverage in Cost control. Include partial and unattributed counts when assessing spend, rather than relying only on the displayed total.
- For hard budgets, verify admission and settlement separately. Prepare pricing before sending production traffic and confirm that completed test requests release excess reservation capacity.
Customer telemetry exports carry the amount, status, source, rule version,
cost-contract version, attribution mode and unattributed reason when applicable. Retain this
provenance with exported usage. Organization, project, team, user and API-key totals are different
views of the same events; adding those views together double-counts spend. Dashboard totals cover
retained analytics, so retain a separate export for accounting periods longer than your retention.