Skip to main content
Routing models let you expose an organization or project name such as aixy/customer-support while managing the concrete provider models behind it. Change targets in Aixy without reconfiguring every client. Use a direct provider/model ID when you want to choose a particular provider explicitly.

Create a routing model

  1. Open Routing. Select a project for a project route, or choose Organization for a route inherited by every project. Choose Create routing model.
  2. In Route, set the display name and API model ID. customer-support becomes aixy/customer-support. Choose whether to prefer one model, share requests equally, share them by weight, or choose by task. Client model names accepts exact names sent by clients such as Claude Code without a provider prefix.
  3. In Models, use Add destination and order the models. Weighted routes expose traffic weights. Task-aware routes ask when each model should be chosen. Expand a model’s connection and timeout settings only when the defaults need adjusting.
  4. Choosing Choose by task adds a Decision model step. Use the default Managed by Aixy option or select your own provider and model ID. Other routes skip this step.
  5. In Reliability, choose whether failures should try another model. Add session preferences, health-based pauses or capability checks when your application needs them. Retry limits and capability declarations expand on demand.
  6. In Review, check the selection, destinations and failure behavior. Use Edit to revisit a step without losing the draft. Confirm the client model names, add a description, choose whether to enable the route, then save. Applications can use its aixy/… ID after configuration refresh.
The editor validates each step before continuing and checks the complete configuration before saving. Moving between steps does not save or invoke a provider. Preview with an example request on the review step checks local planning without invoking the classifier or generation models. Only models allowed by the project and backed by an effective provider connection are eligible at request time. Existing destinations remain visible when catalog discovery temporarily fails; review their current account access before relying on them.

Publish client model aliases

In the first Route step, enter Client model names, one exact model ID per line. For example, publish claude-sonnet-4-5 and claude-sonnet-4-5-20250929 for a route whose destinations serve that model. Clients can send either alias as their model without the aixy/ prefix. The canonical aixy/… ID continues to address the same route. Review lists all accepted names alongside the destinations; use Edit in its Route section to change them. Aliases are case-sensitive, unique within their organization or project scope, and limited to 20 per route. Each ID contains 1–128 letters, numbers, dots, underscores, colons or hyphens and starts with a letter or number. The same name can resolve differently in another project; the API key selects the project. Explicit provider/model IDs and X-Provider retain direct provider selection. Aliases use the route’s existing destinations, restrictions, budgets, guardrails and fallback policy. Available aliases appear in the project’s authenticated model catalog. Removing an alias, disabling an organization route or deleting it stops its inherited resolution after configuration refresh. Deleting a project route restores matching organization names. Disabled project names block inherited names rather than falling back to them. Disabled routes retain their alias reservations until those aliases are removed or the route is deleted. Aliases do not translate model families, choose versions automatically or add provider protocol capabilities. Confirm the identity and supported features of every destination before publishing a client’s model ID. See Claude Code for client configuration.

Organization routes and project overrides

Organization routes provide shared names without copying configuration into every project. For example, define aixy/haiku at organization scope, add the Claude Haiku names your client sends under Client model names, and choose the intended Bedrock destination and Organization connection. Every project’s API key can use those names when that destination is allowed for the project. Each requested name resolves independently: a matching project API model ID or client name takes precedence over its organization counterpart. Other organization names remain inherited. A disabled or unavailable project route never falls back to the organization route with the same requested name. Scope is fixed after creation; create a separate route to change scope. Only users with organization policy authority can create, edit or delete organization routes. Project administrators can read inherited configuration and manage their assigned projects’ own routes. Inheritance does not grant provider or model access: credentials resolve through the target’s selected connection scope, while the calling project’s model restrictions, guardrails, budgets and accounting remain in force. New organization destinations default to the organization connection; Effective can use the calling project’s connection override. Select a project to preview or view performance for an organization route. Both use that project’s authorization context, and performance includes only its retained request evidence.

Choose a strategy

The strategy chooses the first target. Eligible failures then advance through the remaining ordered targets. Set compatible models behind one alias: check tool support, context limits, output behavior, regions, and price for every possible target. With reliability options enabled, capability, connection and health checks determine the eligible destinations before the strategy and Maximum attempts apply. Excluded destinations do not distort the relative weights or rotation of the remaining targets.

Select models by task

Choose Choose by task in the Route step. Describe each destination in Models, then choose one of the two Decision model source options in Decision model:
  • Managed by Aixy is selected by default. Aixy uses Jev through Cloudflare with a managed connection and fixed limits. No provider connection, model ID, credentials or limits need to be configured by your team. Aixy does not store the decision input or output. Cloudflare lists Jev as a zero data retention service.
  • Use your own provider selects one of your existing server-side connections and its model ID. The model must support structured Chat Completions. Decision limits and fallback selection contains connection scope, timeout, character limit and fallback strategy. Follow your provider’s Setup guide; credentials belong in the provider connection.
For each destination, fill When should this model be chosen? with the tasks it handles well and its limits. For example, one destination can handle short extraction and translation while another handles complex code and reasoning. These are administrator-authored criteria, not measured quality scores or guarantees. The classifier sees destinations that passed project access, native protocol, enabled capability checks and health checks. It can select any of those before Maximum attempts limits the generation chain. Use strict capability validation for tools, images or structured output. Managed decisions have a five-second default deadline, controlled by Aixy’s operator, and an 8,000-character total limit across the newest three user text turns. With your own provider, use Decision timeout (ms) and Task character limit to change these bounds. Aixy sends this text after guardrails along with destination criteria. System prompts, assistant turns, tool output and images are excluded; complete system-reminder blocks are removed. Older user text is discarded first. Changing the classifier can change both the provider receiving that text and the decision’s latency and cost. Successful classification moves the chosen destination to the front. The configured strategy orders the other eligible destinations. Timeout, malformed or out-of-list choices, missing connections, blocked decision models and provider errors use that strategy without another classification attempt. No classification runs for token counting, requests without human task text, or a single eligible destination. Send full conversation history; upstream-owned conversation identifiers are unsupported. Task-aware routing disables Keep a stable session destination. Change Decision model source, or your own provider’s model ID, to replace the classifier. Choose another selection mode in Route to pause task-aware routing; completed settings and destination criteria are retained. Returning to Choose by task restores them. An unfinished decision configuration with missing model details or invalid limits is discarded when leaving task-aware mode. Disabling the entire routing model still disables its aliases. Decision content is processed transiently, including when request capture is enabled. Operational metadata such as usage, timing, status and the selected destination remains available for routing and accounting. This privacy rule applies to the decision call; your generation request follows your normal activity-capture policy and destination provider’s terms. Decision calls add provider cost and latency. Their separate Activity usage evidence, labelled routing/classify, joins the same route execution and contributes to observed route cost. They do not count as generation attempts, an extra gateway billing unit or monthly request admission. Missing pricing or usage remains unknown. Hard budgets require complete classifier pricing and reserve for classification plus eligible generation destinations before classification. Budget or quota denial initiates no classifier call. Timeout or cancellation cannot prove that a provider did no billable work. Review pricing and cost attribution before enabling it.

Validate destination capabilities

In Models, expand the destination’s connection and timeout settings to choose the effective connection (project override, then organization), the organization connection, or the project connection only. A project-only target is unavailable when that connection is missing. Project model restrictions apply to every choice. In Reliability, expand Tools, images and structured output, enable Validate destination capabilities, and confirm each model’s support for Tools, Structured output and Images against its selected connection. Unverified is the default. Catalog metadata is informational and does not confirm a capability. Changing the selected model or connection clears its declarations; review them again when changing the provider endpoint or account configuration. Enable Validate destination capabilities to discard targets with an unverified or unsupported feature required by the request. Aixy inspects standard Chat Completions, Messages and Responses input, including tool history and JSON output formats. Native endpoint and streaming restrictions still apply. Audio, file and video input have no confirmed contract in this release and are rejected in strict mode. This validation does not measure model quality, token context limits or data residency. If no destination qualifies, the request returns routing_targets_unavailable before provider execution. Discarded destinations consume no provider attempts. Existing routes keep their behavior until reliability options are enabled.

Keep conversations on a stable destination

Enable Keep a stable session destination and send an opaque X-Aixy-Session value with every turn. Use 1–128 ASCII letters, numbers, dots, underscores, colons or hyphens. Invalid values return 400 when affinity is enabled. Aixy never forwards or records this header. Round robin and weighted routes derive a stable preference from the session, authenticated caller, alias and eligible destinations. Weighted routes distribute sessions according to configured weights. Priority keeps its first eligible destination. Omitting the header retains normal selection. Send the complete conversation history on each request. Affinity does not retain conversation content or provider state. Reliability-enabled routes reject upstream-owned previous_response_id and conversation identifiers. Replicas choose the same preference when their eligible destinations agree; replica-local health can temporarily change that set. Fallback is not permanently pinned: the original preferred destination can be selected again after recovery.

Avoid unhealthy destinations

Enable Avoid unhealthy destinations to pause a connection/model after three consecutive connection failures, timeouts, HTTP 429 or HTTP 5xx responses within 60 seconds. The pause lasts 30 seconds. One subsequent request probes recovery; concurrent requests use other eligible targets. A qualifying probe failure starts another pause. Cancellation releases the probe permit. Health observations end when response headers arrive. Functional HTTP 4xx responses do not count as circuit failures. Later stream interruptions are recorded as request failures, without restarting the stream or changing destinations. Circuits use bounded local memory (up to 4,096 entries), expire idle state after five minutes, and reset on replica restart or connection revision. Eviction can forget older health. Circuit state is local to each replica and is not a global provider status.

Configure failure behavior

In Reliability, enable Try another model when a request fails with at least two destinations. Expand Retry limits and failure conditions to choose whether to advance after connection failures/timeouts, provider rate limits (HTTP 429), or provider server errors (HTTP 5xx). Other upstream 4xx responses are not retried. Maximum attempts includes the first target, is limited by the number of destinations, and cannot exceed five. A target timeout bounds that attempt, while the gateway also bounds the complete request. Fallback selection happens before delivering the chosen response; an interrupted response stream is not transparently restarted on another model. Your client should detect incomplete streams. Repeated attempts may increase latency and cost, and a timeout does not prove that the upstream did no work. Keep the chain short and configure retry behavior appropriate for your application. External tool execution has its own approval and recovery flow in Agent tools.

Call and verify the route

Use the routing model as the model field in the normal client request:
Use the project’s API key and test the protocols and tools your client needs. Open Activity to compare the requested alias, effective provider/model, attempt outcomes, timing, and destination decisions. Activity explains capability, connection, model-access, health and attempt-limit exclusions separately from provider execution. Historical events may have no candidate decisions. A successful response alone does not tell you whether a fallback occurred. Each observed upstream attempt has its own usage and cost evidence. Hard budgets reserve for the possible attempt chain; review pricing and cost attribution before directing budgeted traffic to the route. Disabling a route stops its alias from resolving after configuration refresh. Update callers or enable a replacement before retiring an alias in use.

Preview a routing decision

Use Decision preview in the routing editor to check the current draft before saving, or open a route’s name in Routing to preview its saved configuration. Choose the protocol, enter an example JSON request and optionally a session ID, then select Preview decision. The result shows required capabilities, guardrail outcomes, candidate decisions and possible attempt order, including connection scope and timeout. Drafts use the same validation and permissions as saving a route. Disabled routes show the prospective plan if enabled. Examples and session IDs are not saved, captured or sent to providers; preview does not reserve budgets, consume quota, record API-key usage, advance rotation or acquire a circuit recovery probe. For task-aware routes, preview shows the fallback order and explicitly states that the decision model was not called. It does not predict the model a live classifier would choose. The preview checks static eligibility against the current configuration snapshot. It does not read the inference service’s live circuit health or rotation position; request compatibility is validated again when inference executes. The dashboard uses your user identity for affinity, so an application’s API key can choose a different destination. The API reference also supports selecting an active key you own in the same project for that scope. Rotation, concurrent traffic and health can change before execution. Preview does not predict budget admission, latency, response quality or live provider availability. Change the example or draft and preview again to update the result.

Compare observed route performance

Open a route’s name in Routing to see Observed performance. Select the retained time range, My usage or Project usage, and a recorded configuration version. Analytics permissions and your plan’s retention apply. Project usage requires access to other users’ metrics. The page shows request outcomes, fallback recovery, attributed cost, request duration and first upstream byte at p50/p95, recorded destinations, and observed candidate exclusions. Latency samples come from completed requests. First upstream byte measures a transport observation, not the first generated token. It does not evaluate model quality. A generated execution identifier joins the recorded attempts of each request, even when a client reuses its correlation ID. Costs sum all recorded attempts. Missing attempt events and partial attribution remain incomplete; unknown costs are not zero. Cost per recorded request includes partial attribution, so use the coverage counts when comparing configurations. Attributed cost is not a verified invoice, and budget reservations are not charges. An interrupted HTTP 200 stream is a failure; an intermediate fallback event without the final event has an unverified request outcome. Historical events without an execution identifier appear separately with their attributed cost and are excluded from request metrics. Asynchronous telemetry can arrive late or be dropped; these are observed results rather than an exhaustive billing ledger. Tables show up to 100 destinations and exclusion groups and the most recent 20 versions; totals cover the complete selected window/version.