Prompt Caching Order Sensitivity: Why Putting the Wrong Block First Kills Your Hit Rate
Prompt caching only works when the cacheable prefix is identical across requests. If you put a changing block early, every request gets a new prefix and the cache never hits. This article explains the ordering rules and shows before/after layouts to keep hit rates high.
How Prompt Caching Sees a Prompt
Prompt caching matches the longest common prefix between the current request and a previous one. The provider hashes the prompt from the start and reuses the cached state until it reaches a difference.
So the rule is simple: everything before your first changing token can be cached; everything after it cannot. If you place a volatile block early, you break the cache for every token that follows, including the parts that never change.
Many providers also cache in chunks. A tiny difference near the top can invalidate a large block, even if the remaining text is identical. Ordering is therefore not cosmetic—it directly controls your hit rate.
The Prime Directive: Stable First, Volatile Last
Structure your prompt as a pipeline from most stable to least stable:
- System instructions and role definition
- Static reference material (documentation, schemas, examples)
- Tool or function definitions (if they rarely change)
- Long-term context that is fixed per session or per user
- The actual user query
- Per-request metadata, timestamps, IDs, or random values
If a block changes on every call, it belongs at the bottom—or outside the cached prefix entirely.
Before/After: A Layout You Can Copy
Before (bad ordering)
[System]: You are a support agent.
[Session ID]: 7f3a9c2e <-- changes every request
[Knowledge]: <long static product documentation>
[User]: How do I reset my password?
The session ID appears before the documentation. Because the prefix diverges at the session ID, the entire documentation block is never cached. You pay full price for it every time.
After (good ordering)
[System]: You are a support agent.
[Knowledge]: <long static product documentation>
[User]: How do I reset my password?
[Session ID]: 7f3a9c2e <-- moved to the end
Now the system instructions and knowledge block form a stable prefix. Only the user query and session ID change, so the expensive part is cached across requests. If the session ID is not needed by the model, omit it entirely.
Common Ordering Mistakes
- Timestamps or request IDs at the top. These change constantly and destroy the prefix. Move them to the end or drop them.
- Rotating few-shot examples near the front. If examples are shuffled per request, the cache breaks. Keep a fixed set, or put variable examples after the stable prefix.
- Injecting user profile data before static instructions. Profile data changes per user. Place it after the shared system prompt and static knowledge, or accept per-user cache entries.
- Reordering tool definitions. Tool schemas often serialize in an order that depends on a dictionary or set. Sort them deterministically and keep them stable.
- Template whitespace or formatting drift. A trailing space, a different newline, or a changed separator can make two prompts that look identical hash differently. Normalize your template.
Keep the Prefix Byte-Identical
Cache matching is usually exact at the token level. Two prompts that differ only in whitespace, capitalization, or JSON key order may not share a cache entry.
- Build the stable prefix once and reuse the exact same string.
- Avoid string interpolation inside the prefix; interpolate only in the volatile tail.
- If you use a templating engine, verify that it does not inject timestamps, random IDs, or unstable ordering.
- Log the exact prefix you send so you can diff it when hit rates drop.
Design for Cache Segments
If different parts of your prompt change at different rates, order them by frequency of change:
- Almost never changes: system prompt, safety rules, static docs.
- Changes per session: user profile, conversation history summary.
- Changes per request: the current question, retrieved snippets, tool outputs.
Place the almost-never-changing content first. This maximizes the shared prefix across the most requests. Per-session content can still be cached within a session if it stays identical across turns.
Measure and Iterate
Caching is empirical. Track cache hit rate and cached token usage from your provider's response metadata, then test layout changes.
- Compare hit rate before and after moving a block.
- Watch for drops after template or dependency updates.
- If a block is large but changes frequently, consider whether it belongs in the cached prefix at all.
- Remember that cache entries expire. A layout that works for a burst of traffic may hit less during idle periods.
Cost Impact
Cached input tokens are typically billed at a lower rate than uncached input tokens, but the discount applies only to the matched prefix. A bad layout means you pay full price for tokens you could have cached.
On this platform, you pay the official price multiplied by 1.3, and key contributors are credited official price multiplied by 1.1 (premium models ×1.2) in USDC. Because caching reduces the official token cost, it also reduces the base that the multiplier applies to—so good ordering lowers what you pay without changing the multiplier itself.
Quick Checklist
- Stable instructions and reference material first.
- Volatile metadata, timestamps, and IDs last.
- Deterministic ordering for tools and examples.
- No interpolation inside the cached prefix.
- Monitor hit rate and diff your prefix when it falls.