Prompt Caching Across a Long Claude Code Session: What Actually Gets Cached
Learn how prompt caching works in Claude Code's long sessions, which parts of the prompt prefix are cached (system prompt, tool schemas, file context) and which are rewritten each turn, and how to structure sessions for optimal cache hits.
How Prompt Caching Works in Claude Code
Claude Code manages a growing conversation with the model, sending the entire conversation history each turn. Prompt caching lets the API reuse previously processed parts of that history, reducing latency and cost. Understanding what gets cached helps you structure sessions efficiently.
The Anatomy of a Claude Code Prompt
Each request Claude Code sends consists of:
- System prompt: The core instructions that define Claude's behavior (e.g., "You are a helpful coding assistant..."). This is static across turns.
- Tool schemas: JSON definitions of tools like file readers, editors, and search. These rarely change during a session.
- Conversation history: User messages, assistant responses, and tool call outputs.
- File context: Code snippets, file contents, and search results injected as tool outputs.
What Gets Cached
The cache operates on the prefix of the prompt. If a prefix is identical to a previous request, the API can reuse its internal representation.
- System prompt: Always cached. It's the first thing in the prompt and never changes within a session.
- Tool schemas: Cached as long as the set of tools and their definitions remain unchanged. In Claude Code, tools are fixed per session, so this part is stable.
- Early conversation history: The initial user request and early assistant responses are cached. As the session grows, these become part of the cached prefix.
- File context from earlier turns: If a file was read early and its content is still in the history, that portion can be cached. However, file contents injected later may not benefit until subsequent turns.
What Gets Rewritten Each Turn
- New user input: Each new message from the user is appended after the cached prefix and is not cached on the first turn it appears.
- Latest assistant response: The most recent assistant output is new and not yet cached.
- New tool outputs: When Claude Code reads a file or runs a command, the output is appended to the history. This new content is not cached initially but becomes part of the prefix for future turns.
- Dynamic file changes: If a file is modified and re-read, the new content replaces or adds to the history. The old content may still be cached if it remains in the prefix, but the new content is fresh.
Cache Invalidation
Caching is prefix-based. If any part of the prefix changes, the cache for that portion and everything after it is invalidated.
- Changing the system prompt or tool schemas mid-session would invalidate the entire cache, but Claude Code does not do this.
- Editing a file that was previously read and is still in the history does not invalidate the cache for the old content; the new content is simply appended. However, if Claude Code decides to truncate or rewrite history (e.g., to fit context limits), the cache may be partially invalidated.
- Starting a new session resets the cache entirely.
Implications for Long Sessions
- Early turns are most expensive: The first request processes the full system prompt and tool schemas without cache benefit. Subsequent turns reuse that prefix.
- Cache hits grow with session length: As more of the conversation becomes static prefix, the proportion of cached tokens increases.
- File context is most valuable when stable: Reading a file once and referencing it across many turns maximizes cache reuse. Re-reading the same file unnecessarily can bloat the prefix without adding cached value.
- Avoid frequent context resets: Tools that clear or summarize history may reduce cache efficiency. Let the session grow naturally when possible.
Best Practices
- Keep the system prompt and tool schemas consistent within a session.
- Load large files early if you anticipate multiple references.
- Avoid editing files that are already in the context unless necessary; each edit adds new tokens that must be processed anew.
- For long-running tasks, prefer continuing the same session over starting fresh to leverage the accumulated cache.
- Be aware that context window limits may force truncation, which can invalidate parts of the cache.
By understanding these mechanics, you can structure your Claude Code sessions to maximize cache hits, reducing both latency and cost.