What Is Prompt Caching? How It Cuts AI API Cost and Latency

What is prompt caching?
Prompt caching is an AI API feature that stores the repeated beginning of a prompt (a system instruction, a document, tool definitions) on the provider side and reuses it in later requests without processing it again. As a result, the model does not reread the same text, so cost and latency drop while answer quality stays the same.
First, a simple analogy. Imagine an employee who rereads the same long work manual for every new customer. If that employee read it once and kept a summary on the desk, they would only read the new question each time. In short, prompt caching is exactly that shortcut.
However, one detail matters here. The provider does not store the model's answer; it stores the processed state of the first part of your prompt. In other words, the model still writes a fresh answer on every request, and it only skips the repeated reading work.
In this guide we answer what is prompt caching at the concept level. Also, we deliberately avoid numbers that age quickly, such as model names, prices and time limits. For current values, check the provider's official pages.
How does prompt caching work?
A language model first splits text into pieces called tokens. Then it computes intermediate states (key-value states) for each piece, which the attention mechanism uses later. In a long prompt, this preprocessing is most of the total effort, so skipping it pays off.
Then the provider keeps these intermediate states for a short time. When a new request starts with the same content, the provider does not recompute that part. Instead, it loads the saved state and processes only the new part. The match must be exact, because a single changed character ends the match at that point.
- A request arrives, and the provider compares its beginning with content it has seen before.
- If it finds the longest matching prefix, it reuses the saved intermediate states.
- It processes the remaining, unmatched part in the normal way and produces the answer.
- When needed, it also saves the new prefix, so the next request benefits from it.
However, nothing changes for the end user in this flow. Instead, the same request returns the same kind of answer, with or without a cache. The difference shows up only in the compute load behind the scenes and on your invoice.
What are a stable prefix and a variable suffix?
First, you can split every prompt into two parts. The first part stays the same in every request, and we call it the stable prefix. System instructions, role definitions, rules, examples, tool definitions and reference documents belong here. The second part changes with every request, and we call it the variable suffix: the user's question, the current date and session-specific data.
Also, the cache works only on the shared part that starts at the very beginning. For this reason, order is decisive. Put stable content first and variable content last. Otherwise, a small change at the top breaks the match for everything after it.
- Stable prefix: brand voice, safety rules, output format and frequently used documents.
- Semi-stable part: conversation history that stays the same during a session.
- Variable suffix: the new user message and information tied to that moment.
For example, if you write the current time on the first line of the system instruction, every request starts differently. When you move the same detail to the suffix, the prefix stays fixed and the cache works.
What is prompt caching and why does it lower cost?
Because AI APIs usually charge for the input tokens they process and the output tokens they generate, resend costs add up. When you resend a long system instruction with every request, you pay again and again for the same tokens. So a cache lowers the price of that repetition.
In practice, the logic at most providers is similar. Writing a prefix to the cache for the first time can be priced differently from normal input, while later reads are clearly cheaper than normal input. Therefore, your gain depends on how many times you reuse the same prefix. Do not expect savings from a one-off request.
We deliberately do not give discount rates or write prices here. They differ by provider and model, and providers update them over time. Check the current rate on the provider's official pricing page.
To test the savings with your own price data, you can use our percentage calculator. First enter your monthly input token volume, then the share of the stable prefix. That way you get a rough view of the potential gain from caching.
How does prompt caching affect response time?
Also, latency drops along with cost. When the model does not have to process a long prefix from scratch, it starts producing the first token sooner. Users feel this as an assistant that starts answering faster.
In practice, the effect is strongest in jobs with very long input and short output. For example, consider a flow that reads a long contract and returns a one-sentence summary. Most of the waiting time comes from processing the input, so caching makes a clear difference there.
On the other hand, the effect stays limited when the output is very long. Generating an answer is a separate job from reading the input, and caching speeds up only the reading side. So set your expectations by the shape of your workload.
So do not guess the latency gain; measure it. Sending the same request twice and comparing the two response times is a good first step.
When does prompt caching help?
In short, caching creates value wherever the same large content is sent again and again. The common pattern is simple: the prefix is long, requests arrive one after another, and the changing part is small.
- Long system instructions: assistants with brand voice, policy text and output rules.
- Fixed documents: flows where many questions target the same manual, contract or product catalog.
- Multi-turn chat: bots that resend the unchanged start of the history on each turn.
- Prompts with many examples: classification and formatting jobs that keep the example pairs fixed.
- Tool-using agents: loops where long tool definitions stay the same at every step.
However, the gain is especially large in agent scenarios. An agent calls the model many times to finish one task, and it sends the same instructions and tool list on every call. So caching cuts the cost of that repetition sharply. This is why keeping tool definitions stable matters in flows built on function calling.
When does prompt caching not help?
A cache catches nothing when the prompt differs from the very first token in every request. For example, one-off jobs that generate a completely separate text for each user share no common prefix.
A short prompt is also a problem. Providers set a minimum length threshold for caching, and prompts below it never enter the cache. Because the threshold changes by model, check the current value in the official documentation.
The third case is a gap between requests that is longer than the cache lifetime. Once the entry expires, the provider writes the prefix again, and the first request comes back at the full price or the write price. In an application with sparse traffic, the gain may stay low.
- A short prefix brings no gain.
- A prefix that changes in every request never matches.
- Sparse requests let the entry expire.
- User-specific data at the very top leaves no shared part.
How do the major providers offer prompt caching?
Providers differ in the details, but two common paths exist: automatic caching and explicit control. First, in automatic mode, the provider catches matching prefixes by itself. In explicit mode, you mark where the cacheable section ends.
OpenAI documentation recommends placing stable instructions and shared reference material first and changing content last. Anthropic documentation describes a cache order that follows tools, system and messages, with marked breakpoints. Google documentation says implicit caching is on by default and suggests putting large shared content at the start of the prompt.
These three pages are a good place to start: the OpenAI prompt caching guide, the Anthropic prompt caching documentation and the Google Gemini caching documentation. Field names and settings can change over time, so read the current versions before you build.
If you want a broader API starting point, see our guide to the OpenAI API.
Why do cache lifetime and minimum length matter?
First, a cache does not live forever. The provider keeps the entry for a limited time, and if a new request uses the same prefix within that time, the lifetime often extends. Also, the lifetime itself depends on the provider, the model and the option you choose.
In practice, this means the following. When traffic is dense, the entry stays fresh because it is used often. When traffic is sparse, the entry drops out between requests. Some providers offer a longer retention option; check the official pricing page to see whether it carries an extra cost.
However, minimum length matters just as much. If your prefix is below the threshold, no entry is created even when you mark it. In that case you may think the cache is broken, but the prompt is simply not long enough.
Do not hard-code lifetime and threshold values in your product documentation. Instead, keep them in a configuration note and review them whenever the provider changes something.
How do you measure cache hits?
Do not guess whether the cache works; read it from the usage fields of the response. Also, providers report how many tokens came from the cache with each reply. At OpenAI this appears in the `cached_tokens` field, at Anthropic in `cache_read_input_tokens` and `cache_creation_input_tokens`, and at Google in a usage field that shows cached token counts.
However, field names can change between versions. Therefore, check the official response schema before you write the integration.
You can track the hit rate with simple logic: divide the cached tokens by the total input tokens. If the rate is low, suspect that something in the prefix keeps changing.
- Log the usage information of every response.
- Sum the cached tokens and the newly written tokens separately.
- Watch the daily or weekly hit rate on a dashboard.
- When the rate drops, check the latest change (instruction, tool list, ordering) step by step.
What is the difference between prompt caching and an app cache like Redis?
First, the two ideas share a word but work at different layers. An application cache lives on your server and usually stores an answer or a database result. Prompt caching lives on the provider side and stores the effort of processing the model input.
Tools such as Redis or Memcached are ideal for returning the exact same answer to the exact same request. We covered that topic in our article on how caching works with Redis and Memcached. Prompt caching, in contrast, does not freeze the answer; it only makes the repeated part of the input cheaper.
| Feature | Prompt caching | App cache (like Redis) | Semantic cache |
|---|---|---|---|
| Where it runs | On the provider's servers | In your infrastructure | In your infrastructure, usually with vector search |
| What it stores | The processed state of the prompt prefix | A ready answer or a computed result | Answers given to similar questions |
| Match type | Exactly the same prefix from the start | Exactly the same key | A question close in meaning |
| Does the answer change | The model generates it fresh each time | The same answer returns | An old answer may return |
| Main gain | Input cost and latency | Skipping the model call | Skipping the model call |
| Main risk | Lost hits when the prefix changes | Stale answers | Wrong matches |
The two do not replace each other, and they can work together. Use the app cache for frequent, identical questions, and use prompt caching for the remaining requests that carry a long stable instruction.
How does prompt caching relate to tokens, the context window and RAG?
It is easy to mix up neighboring terms, so let us draw the relationship briefly. First, a token is the unit of measure for both the bill and the limit. The context window is the total amount of text the model can see at once. Prompt caching does not enlarge that window; it processes the repeated part inside the window more cheaply.
RAG finds the needed document pieces at query time and adds them to the prompt. Therefore, RAG and caching are not rivals. The shared instruction and template go into the cache, while the retrieved pieces specific to each question go into the suffix.| Term | What it does | Relation to prompt caching |
|---|---|---|
| Token | Unit of measure for text | The cache lowers the processing cost of tokens |
| Context window | Limit on text visible at once | Does not change the limit, makes the filled part cheaper |
| RAG | Finds and adds relevant document pieces | Stable part in the cache, retrieved pieces in the suffix |
| Structured output | Fits the answer into a given schema | The schema definition can sit in the stable prefix |
| Fine-tuning | Changes model behavior through training | Caching is a runtime optimization, not training |
If you want the answer in a fixed format, read our article on structured outputs. Because the schema stays constant, it often sits in the prefix and benefits from the cache.
Example scenario: what does prompt caching give a customer support assistant?
This is an example scenario, not a real client result. For example, think of a store's support assistant. The system instruction holds the brand voice, the return policy, the shipping rules and the answer format. In addition, you attach a summary of the product catalog to every request.
Without a cache, the model also rereads this long text on every customer message. As a result, every short question creates a long input bill. With caching, the instruction and the catalog sit in the stable prefix, and the customer's question sits in the suffix.
- Put the system instruction, the policy text and the catalog summary at the very start of the request.
- Move variables such as the customer name, the order number and the current date to the end.
- The first request writes the prefix, and later requests read it.
- Track the hit rate from the usage fields.
We do not give a percentage or a savings figure here, because it depends on your prefix length, your traffic and the provider's current price. Still, the structure is clear: the longer the prefix and the denser the requests, the larger the gain.
If you want to run this kind of assistant on your own company documents, our RAG development solution describes this architecture.
Is prompt caching safe, and what should you watch for in privacy?
Provider documentation states that caches are not shared across organizations. For example, OpenAI documentation says caches are not shared across organizations and cannot be reused across regional processing boundaries. Anthropic documentation also says caches are not shared between workspaces.
Even so, good practice still matters. Do not put user-specific personal data in the stable prefix. The prefix is the common point of all requests inside the same organization or workspace, so you need to be clear about the scope of your data.
- Keep passwords, keys and confidential documents out of the prefix.
- Move personal data to the suffix when possible, and do not send it without need.
- Do not merge the data of different customers into the same prefix.
- Read the retention and processing terms in the provider's official data documents.
This section is not legal advice. For an application under GDPR or similar rules, review the data processing terms with your legal advisor.
Which mistakes do teams make most often with prompt caching?
Most failed attempts, in fact, trace back to the same few causes. Most of these mistakes are not coding errors; they come from changing the prefix without noticing.
- Writing a timestamp or a random ID at the top of the system instruction.
- Changing the order or the content of the tool list on every request.
- Mixing user-specific data into the stable prefix.
- Appending content to an existing message and breaking the match.
- Switching request settings, such as thinking or image options, from request to request.
- Leaving the prefix below the threshold and still expecting a cache.
Anthropic documentation states clearly that changes to tool definitions, system content and some request settings can invalidate the cache partly or fully. Therefore, think about the cache effect before you change a setting.
Another common mistake is assuming the gain without measuring it. Moreover, some providers charge extra to write a prefix, so marking a prefix that you never reuse can even cost you money.
What are the common misunderstandings about prompt caching?
Whoever asks what is prompt caching often carries one of four myths. The first misunderstanding is that the cache stores the answer and returns it again. It does not; the model generates a new answer every time, and only the input processing gets shorter. Anyone who expects identical answers will be disappointed.
The second misunderstanding is that caching gives the model memory. It does not. The model does not remember the earlier conversation; you resend the history with every request, and the provider shortens the work on the repeated part.
The third misunderstanding is that caching always saves money. With short prompts, sparse traffic and changing prefixes, there is no gain. The fourth is that output quality changes; however, the documentation presents caching only as an efficiency feature.
Also, caching does not remove inference. The model still runs; it only reuses the computation it already did for the first part of the input.
What should a practical prompt caching checklist include?
If you still wonder what is prompt caching worth for your own product, review the list below once before you roll it out. Each item is a decision point, and you can check most of them in half an hour.
- Separate the stable and variable parts of the prompt, and write them down.
- Place stable content first and variable content last.
- Freeze the order and the text of tool definitions.
- Make sure the prefix is above the provider's minimum length threshold.
- Check whether your request frequency fits the cache lifetime.
- Log cache hits from the usage fields and watch them on a dashboard.
- Remove personal data and secrets from the prefix.
- Read the current discount and write price on the official pricing page.
This list works for developers and managers alike. A manager looks at the first and last items, which cover the cost calculation and the price check. A developer focuses on the technical items in the middle.
After that, watch the hit rate for a week. If it is lower than you expected, inspect the ordering first, then the setting changes, and request frequency last.
Which terms should you know when you talk about prompt caching?
In the documentation, you will see the same idea under different names. A small glossary reduces the confusion. In addition, it helps your team build a shared language.
- Prefix: the part of the prompt that starts at the beginning and stays the same across requests.
- Breakpoint: a marker that shows where the cacheable section ends.
- Cache hit: the prefix is found in the cache and reused.
- Cache miss: there is no match, so the prefix is processed from scratch.
- Write and read: the first save and the later use are two operations that can be priced separately.
- Lifetime (TTL): how long an entry survives if nobody reuses it.
All of these terms connect to one idea: using the same beginning without paying for it twice. Provider setting names differ, but the logic stays the same.
How does prompt caching grow in multi-turn chats?
In chatbots, the full history goes out again with every new message. As the conversation gets longer, the text you send grows, but the growth is always at the end. Because the start stays the same, the cache fits this structure very well.
For example, the text you send in the third turn is the second turn plus one new message. The provider finds the entry of the second turn and processes only the new message. So even as the chat grows, the extra cost of each turn stays relatively flat.
The only trap here is rewriting the history. Summarizing or shortening old messages, or inserting content in between, changes the prefix and wastes the cache. If you must shorten the history, do it rarely and on purpose.
What does warming up the cache do?
The first request to a cache is always slower, because the prefix is not saved yet. Some providers let you send a warm-up request that writes the prefix before real users arrive. Anthropic documentation describes this approach as pre-warming.
Pre-warming makes sense mainly on user-facing screens that are sensitive to latency. For example, you send one warm-up request before the working day starts, so the first real user meets a ready cache. However, it means one extra request and one extra cost.
For this reason, do not warm up every application. Measure first: does the slow first request really bother your users? If it does not, warming up is unnecessary complexity.
How do you build a prompt template when you ask what is prompt caching?
The technical answer to what is prompt caching in practice is a template that splits the prompt into layers. The layers run from the least changing to the most changing. This order matches the matching logic of the cache exactly.
- First, tool definitions and the system instruction, which rarely change.
- Second, reference documents that stay the same for days.
- Then the conversation history that grows during a session.
- Finally, the user message and the live data for that request.
Define this structure in one place in your code. That way, when someone on the team updates the instruction, they know it resets the cache and make the change on purpose. Putting a frequently changing part into an upper layer breaks the match of all the layers below it.
Adding a version number or a change note to the template is tempting, but do not write it into the prefix. If you need version information, keep it in your logs, not inside the request.
How does our team approach prompt caching for businesses?
At Talha Aslan and team, we first understand the workflow in AI projects, and then we break the cost into line items. Prompt caching is a small but effective item in that analysis, because the input cost has a large share in assistants with long instructions.
In practice, we follow this order. First we simplify the prompt, then we separate the stable and variable parts, and finally we switch on the cache and measure the hits. The order matters, because caching an unnecessarily long prompt locks in an expensive habit instead of solving the problem.
If you are planning an AI integration that fits your business, take a look at our AI automation services. There we cover architecture, cost and security together, based on your use case.
Remember that what we describe here is a general framework. For each project, we evaluate the provider choice, the data rules and the budget separately.
What should you do next after learning what is prompt caching?
Let us summarize briefly. Put simply, prompt caching stores the repeated beginning of a prompt on the provider side and lowers cost and latency. For it to work, the prefix must be long, stable and exactly the same across requests.
As a first step, open your current prompt and mark the stable and variable parts. In the second step, move the stable part to the top. In the third step, watch the hit rate for a week and confirm it with the official usage fields.
Do not treat any price, duration or threshold in this article as fixed. Provider documentation changes, so read the official pages before you decide.
For neighboring topics, our articles on tokens and API cost, the context window and large language models are a good next read.



