> ## Documentation Index
> Fetch the complete documentation index at: https://docs.aixy-gateway.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Provider prompt caching

> Configure provider-managed prompt caching and understand estimated savings.

Open **Cost optimization** in the dashboard to manage prompt caching and review savings.
Owners and organization admins can manage organization connections. Project admins can manage
their project connections and see savings within their assigned projects. Regular project
users cannot access this section.

Use the project selector to choose an organization or project view, then select a reporting
period. Reporting availability follows your plan's analytics retention.

## Choose a caching policy

For supported Claude models on Anthropic and Amazon Bedrock, each connection offers:

* **Follow client**: keep the behavior your application requests. This is the default.
* **Enabled**: let Aixy add 5-minute provider cache checkpoints to reusable system context and
  the latest user turn. Native requests with existing client checkpoints keep those settings.
* **Disabled**: stop Aixy adding checkpoints and remove cache controls from supported native
  requests. This does not delete entries already stored by the provider or disable implicit
  caching, which Bedrock can still apply to eligible Claude models.

Changes save immediately. A project connection takes precedence over the organization's
connection. Changing its key preserves the caching preference. Without a project connection,
requests inherit the organization connection and its preference.

OpenAI, Azure OpenAI, Gemini, Vertex AI, DeepSeek, and xAI display **Automatic** for their
provider-managed caching behavior on eligible models. Their caching is managed by the provider.
Configure explicit checkpoints on the supported Anthropic and Bedrock Claude connections shown
in Cost optimization; other models retain their existing caching behavior.

A cache hit depends on the model, region, minimum prompt length, repeated prefix, and provider
expiry rules. Keep stable instructions before changing content. A setting alone does not
ensure a hit or a saving. Your provider holds and bills the cache; Aixy does not store a cache
of prompts or responses for this strategy.

## Understand savings

**Net savings · estimated** compares the same measured input at standard input rates with its
estimated cost including uncached input, cache reads, and cache writes. The percentage refers
to measured input cost. Writes can make the net saving negative when content is not reused.

The panel includes all observed provider-cache savings, including caching requested by your
application and automatic provider caching. It does not claim that activating Aixy caused all
of those savings. Provider minimums and different write lifetimes can affect the outcome.

Request counts show measurement coverage. Missing prices, unsupported billing dimensions,
incomplete usage, and older records without comparable evidence remain excluded. A dash
means savings cannot yet be measured; it does not mean there was no saving. Historical
observations keep their original pricing evidence. These estimates are not invoice totals.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.