Is Prompt Caching Cutting Your Automation Bill?
Why is your AI automation bill higher than the token count suggests?
As of October 2026, often because you are paying full price for the same prompt prefix over and over. Every major provider now discounts repeated input heavily, and the discount is large. Anthropic's documentation puts cache reads at 0.1 times the base input price on most Claude models. If you are not hitting that path, you are leaving most of a bill on the table.
This matters more for automations than for chat. A chat session is one person typing. An automation runs the same instruction block a thousand times a day with a different record pasted at the end, which is exactly the shape caching rewards.
So this is a practical look at what prompt caching does, what the three big providers actually promise, and the design choices that decide whether you get the discount or quietly miss it.
What is prompt caching?
A way for a provider to reuse the work it already did on the start of your prompt. The first call processes the whole thing and stores the processed prefix. Later calls that begin with the identical text skip that work and pay a reduced rate for those tokens instead of the full input rate.
The key word is prefix. Caching works from the beginning of the prompt forward, so the stable part has to come first and the variable part last. A prompt that puts the record at the top and the instructions at the bottom gets no benefit at all, even though it contains exactly the same words.
That single ordering decision is usually the difference between a cheap automation and an expensive one. It costs nothing to get right and it is almost never the way people naturally write a prompt.
How much does prompt caching actually save?
Enough to change what is worth building. Anthropic's prompt caching documentation lists cache read tokens at 0.1 times the base input price on most models, dropping to 0.05 times on Claude Opus 5.5 and 0.025 times on Claude Fable 5.1 and Claude Mythos 5.1.
OpenAI's prompt caching guide describes a similar structure, saying you pay the model's reduced cached-input rate for reused tokens, discounted up to 95 percent. For GPT-5.6 and later it puts reads at 0.1 times the input rate on most of those models and 0.05 times on GPT-6.1 Sol.
Google's Gemini documentation is less specific about the number and more specific about the mechanism, saying it automatically passes on cost savings if your request hits caches, and that implicit caching is enabled by default for all Gemini 2.5 and newer models. The common thread across all three is that reuse is cheap and the first write is not.
How long does a cache entry survive?
Minutes by default, with longer options. Anthropic's documentation gives a default cache lifetime of 5 minutes, with an extended 1 hour option, and prices cache writes at 1.25 times the base input price for the 5 minute window and 2 times for the 1 hour window.
OpenAI's guide describes a sliding window rather than a fixed one. For GPT-5.6 and later it says a cached prefix remains eligible for reuse for 30 minutes after its most recent write or reuse, and notes that earlier models typically retain entries for around 5 to 10 minutes of inactivity, up to one hour.
The practical consequence is about cadence, not duration. An automation that runs every few seconds keeps its cache warm for free. One that runs twice an hour pays the write premium almost every time and gets no read discount at all, which can make caching a net loss. Our notes on batch versus real-time automation cover how that scheduling choice ripples outward.
How do the three major providers differ?
They agree on the idea and differ on the dials. All three now cache by default to some degree, all three charge less for reused input, and all three require a minimum prefix length before anything is cached at all. Below is what each one's own documentation states as of October 2026.
| Behaviour | Anthropic Claude | OpenAI | Google Gemini |
|---|---|---|---|
| Minimum cacheable prefix | 512 tokens on Claude Opus 5.5, Opus 5, Sonnet 5.5 and Fable 5.1; up to 4,096 on some older models | 1,024 tokens on GPT-5.6 and later | Stated per model, from 2,048 tokens on Gemini 2.5 models to 4,096 on newer versions |
| Cache read price | 0.1 times base input, 0.05 on Opus 5.5, 0.025 on Fable 5.1 and Mythos 5.1 | 0.1 times input on most GPT-5.6 and later, 0.05 on GPT-6.1 Sol, described as discounted up to 95 percent | Savings passed on automatically, figure not stated on the caching page |
| Cache write cost | 1.25 times base input for the 5 minute window, 2 times for 1 hour | Not stated as a separate multiplier on the guide | Not stated on the caching page |
| Lifetime | 5 minutes by default, 1 hour option | 30 minutes from the most recent write or reuse on GPT-5.6 and later | Not stated on the caching page |
| Control | Up to 4 cache breakpoints per request | Enabled by default, with implicit and explicit modes on GPT-5.6 and later | Implicit caching on by default for Gemini 2.5 and newer; explicit caching needs the generateContent API |
Read that table as a snapshot rather than a law. These numbers change, and the per-model lists change faster than the mechanisms do. Check the vendor page before you build a cost model on it.
What breaks a cache hit?
Any change to the prefix, however small. A timestamp injected into the system prompt, a record ID at the top, a randomly ordered list of tools, a trailing space that varies: each one produces a different prefix, and a different prefix is a cache miss.
This is the most common self-inflicted cost we see. A team adds the current date to the instructions for good reasons, and every single call now writes a fresh cache entry at the write premium and reads none of them. The bill goes up and nothing in the code looks wrong.
The fix is to put everything volatile after everything stable. Instructions, schemas, examples, and tool definitions first. The specific record, the timestamp, and anything per-run last. If you need the date in the prompt, put it in the user message, not the system block.
Which automations benefit most from caching?
The repetitive, instruction-heavy ones. Classification, extraction, content QA, and triage all share one shape: a long fixed rubric plus a short variable input. That is the ideal case, because the expensive part of the prompt is the part that never changes.
Agentic workflows benefit for a second reason. A tool-using agent resends the whole conversation on every turn, so the prefix grows and repeats constantly. Anthropic's documentation allows up to 4 cache breakpoints per request, which exists precisely so you can cache a growing conversation in layers rather than all or nothing.
The automations that benefit least are the one-off and the tiny. If your prompt sits under the minimum cacheable length, nothing is cached regardless of how often you call it, and the minimums are not small: 512 tokens at the lowest end and 4,096 at the highest. Our piece on choosing a small model over a frontier one is often the better lever for those cases.
When does caching cost you more?
When you write far more often than you read. Anthropic prices a cache write above the normal input rate, at 1.25 times for the short window and 2 times for the long one. If a workflow runs once every two hours with a 5 minute cache, every call is a write at a premium and no call is ever a read.
The 1 hour option changes that arithmetic but not for free. You are choosing to pay double on writes to make reads possible for an hour. That pays off for a workflow running several times an hour and loses for one running twice a day.
So the honest version is that caching is not a free win, it is a bet on cadence. We would rather see a team measure its actual call pattern for a week than assume the discount applies. Our notes on AI cost control for web teams cover the rest of that measurement.
How do you measure whether caching is working?
By reading the token counts in the response, not the invoice. Every provider reports cached and uncached input tokens separately per call. If your cached-read count is near zero on a workflow you expected to cache, the prefix is changing and you can find out why in minutes.
Log those two numbers per run alongside the prompt version. That gives you a cache hit rate you can watch, and it tells you immediately when a prompt edit quietly destroys reuse. A prompt change that improves quality and triples cost is a trade worth making knowingly and not by accident.
We treat this as part of prompt change management rather than as a cost exercise. If you are already versioning prompts, adding a cache hit rate to the same record is almost free. Our notes on versioning and testing prompts describe the surrounding habit.
How do we think about this in our own pipelines?
Prefix discipline first, model choice second, cadence third. We write prompts with the stable block at the top from the start, because retrofitting that order into a working automation is more annoying than it sounds and the saving is the same either way.
We are also wary of treating a cheap read rate as permission to stop thinking about prompt size. A 0.1 times rate on a very long prompt is still a real cost, and a shorter prompt is usually also a clearer one. Caching rewards repetition, not bloat.
Where we stay cautious: these prices and minimums are vendor settings that move. We would not design an automation whose business case only works at one provider's current cache read multiple. The same caution we apply to falling model prices applies here, in both directions.
Where is prompt caching heading?
Towards being invisible and assumed. Two of the three providers already cache by default, the discounts are converging around a tenth of input price, and the interesting differences are moving to control rather than existence. Within a year, not getting the discount will look like a configuration mistake rather than a missed optimisation.
What we expect to stay hard is prompt architecture. Deciding what belongs in the stable prefix, what belongs per call, and where the breakpoints go is a design question about your automation, and no amount of default caching answers it for you.
If your AI automation costs more than the token maths suggests and you want someone to look at the prompt structure, we are happy to dig in. We are at phoenix.studio.
Want a site that performs like this?
Tell us about your project. We will come back with a clear next step, no pressure.
This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
Have a project like this?
Tell us where you want to go. We'll tell you how we'd get you there.