Cut costs with prompt caching
Prompt caching lets you reuse the unchanged prefix of a conversation across API calls, reducing token costs and latency. This article explains how caching works, why it matters for long multi-turn chats, and how to structure prompts to benefit on Claude and other models.
What is prompt caching?
Prompt caching is a technique where the API provider stores parts of your prompt that are repeated across multiple requests. Instead of processing the same text from scratch each time, the model reuses the cached representation, saving computation.
When you send a request, the provider checks if an identical prefix has been seen recently. If so, it can skip re-encoding that portion and only process the new text. This reduces both the number of tokens billed and the response time.
Why caching matters for multi-turn conversations
In a typical chat, you resend the entire conversation history with each new message. That history grows with every turn, so later turns become expensive and slow.
With caching:
- The system prompt and earlier turns can be cached.
- Only the latest user message and any new assistant output are processed fresh.
- Cost and latency stay roughly constant per turn instead of growing linearly.
This is especially valuable for long-running chats, document-based Q&A, or any workflow where a large, static instruction set or context is reused.
How caching works conceptually
Most implementations use a prefix-matching approach:
- You mark a section of your prompt as cacheable (e.g., using a special parameter or breakpoint).
- The provider stores the internal state for that prefix after the first request.
- On subsequent requests, if the prefix is identical, the provider reuses the stored state for the cached portion.
- If the prefix changes, the cache for that part is invalidated and rebuilt.
The cache is typically scoped to your account and has a short time-to-live, so it works best for bursts of related requests.
Structuring prompts for effective caching
To get the most out of caching, organize your prompts so that the static parts come first and the dynamic parts come last.
- Put stable content early: system instructions, examples, and long documents should be at the beginning.
- Keep the prefix identical: even a small change in the cached region can break the cache. Avoid timestamps or random IDs in that part.
- Use explicit cache breakpoints: some APIs let you mark where the cacheable prefix ends.
- Batch similar requests: if you have many queries against the same context, send them close together so the cache stays warm.
Caching with Claude
Claude supports prompt caching via explicit cache control markers. You can designate up to a few breakpoints in your prompt to indicate which parts should be cached. This is ideal for:
- Long system prompts that define a persona or rules.
- Large reference documents that are reused across many questions.
- Multi-turn conversations where the history is resent.
The cache has a limited lifetime, so it's most effective when requests are made in quick succession.
Cost implications
Caching changes the billing structure: cached tokens are usually cheaper than regular input tokens, while writing to the cache may incur a small extra cost. The net effect is that repeated use of the same prefix becomes significantly cheaper than processing it fresh each time.
On our platform, you pay the official price multiplied by 1.3 for all tokens, including cached ones. However, because caching reduces the number of tokens that need to be processed at full price, your overall spend can still drop. If you're a key contributor, you earn credits at official price × 1.1 (or × 1.2 for premium models) in USDC, so caching also helps you maximize your earnings by reducing the effective cost per request.
Other models and providers
While Claude offers explicit caching controls, other models like GPT, DeepSeek, Qwen, GLM, and Kimi may have their own caching mechanisms or automatic prefix caching. The principles remain the same: put static content first, keep prefixes stable, and group related requests.
Best practices
- Design prompts with a clear separation between static and dynamic parts.
- Monitor cache hit rates if your provider exposes them.
- Avoid changing the beginning of your prompt unnecessarily.
- Use caching for long conversations and document-based tasks.
- Combine caching with other cost-saving measures like shorter prompts and streaming.
Summary
Prompt caching is a straightforward way to cut costs and latency in multi-turn conversations. By reusing the unchanged prefix of your prompt, you avoid paying full price for the same tokens over and over. Structure your prompts with static content first, use cache breakpoints where available, and keep requests grouped to keep the cache warm. This works across many models, with Claude offering explicit controls for fine-grained optimization.