Look at an AI API price table today and you rarely see just Input and Output. Columns like Cached input and Cache write sit next to them.

It looks complicated at first, but for repeated work the difference is quite large.

Why the repeated prefix feels expensive

Say your agent sends the following on every request:

System instructions       10K
Repository summary        30K
Policies / docs           40K
Current user message       1K
-----------------------------
Total                     81K

Only 1K of user message changes, yet more than 80K of shared context goes along every time.

Prompt caching reduces the cost of reprocessing this repeated prefix.

Read cache reads and cache writes separately

The OpenAI GPT-6 price table lists regular input, cached input, and cache write as separate lines.

The core structure looks like this:

First request
→ may include the cost of creating the cache

Later requests
→ reuse the same cached prefix at a lower unit price

So if you only make one request, caching helps little or not at all. The more repetitions, the more it matters.

CodeBridge Mini Lab: calculate the break-even point yourself

Assume this:

Shared context: 100,000 tokens
Changing input: 2,000 tokens
Repetitions: N

Compare the two approaches.

A. Uncached input every time:

cost_A = N × 102K × normal_input_rate

B. First-request cache write plus later cache hits:

cost_B = first_write + (N-1) × cached_rate + changing_input

Plug in real per-model prices and try N = 1, 2, 5, 10, 50 to see where the savings begin.

Even simple Python is enough:

for n in [1, 2, 5, 10, 50]:
    print(n)

The point is not memorizing exact numbers. It is seeing how repetitive your own request pattern is.

Some structures barely benefit from the cache

If the front of your prompt changes a lot on every request, the reuse rate stays low.

A bad structure:

[Current time]
[Dynamically changing user data]
[Long fixed documents]
[System instructions]

When volatile content sits in front of the fixed parts, building a cache-friendly prefix gets hard.

If you can, keep long fixed instructions and documents in a stable structure and put frequently changing information elsewhere.

Caching does not compete with RAG

You can shrink context with retrieval and still cache the remaining shared instructions:

Retrieval
→ select only the documents you need
→ cache repeated system/tool instructions
→ call the model

They are different layers of cost optimization.

Cost per Token vs Cost per Workflow

Do not look at token prices alone when you evaluate prompt caching.

If an over-complicated prompt structure built for the cache raises your agent failure rate, the total cost can actually grow.

So this is the better final metric:

Total workflow cost
──────────────────
Number of successful tasks

Conclusion: if your inputs are long and repeated, put the cache in your cost model

In a short chatbot the difference can be small. But when you reuse the same context dozens of times — a repository agent, document analysis, long system prompts — the cached input price can matter as much as the model choice.

When you read a price table, look at all three lines together:

Input
Cached input
Cache write

And the most accurate approach is calculating with your real repetition counts.

Further reading

References

Go deeper with a course

If you want to build the habit of picking the right model and settings for each situation, a hands-on course on using AI tools by scenario helps.