pricing-cache-vs-multiplier · EN · 2026-10-07

Does the 1.3x Multiplier Apply to Cached Tokens Too? Break-Even Math for Prompt Caching

This article explains how the platform's 1.3x multiplier applies to cached input tokens and provides a break-even analysis for prompt caching. It clarifies that the multiplier is applied after the provider's cache discount, so cached tokens are billed at 1.3 times the official cached rate. The break-even point for enabling caching depends on the cache write cost and the number of reuses; a simple formula is provided to calculate the minimum reuse count.

How the 1.3x Multiplier Works with Prompt Caching

On this platform, you pay the official provider price multiplied by 1.3 for all token usage. That includes cached input tokens. The multiplier is applied after any provider-specific cache discount. For example, if a provider charges 10% of the normal input price for a cached token, you pay 1.3 × (0.1 × normal price) = 0.13 × normal price per cached token.

This means the relative savings from caching remain the same as with the provider: you still pay 1.3 times whatever the provider charges, but the cached rate itself is lower. The multiplier does not erase the benefit of caching.

Break-Even Math for Prompt Caching

Prompt caching typically involves a one-time cost to write the cache (the cache write or storage cost) and a reduced cost for subsequent uses (cache reads). The break-even point is the number of reuses at which the total cost with caching equals the total cost without caching.

Let:

  • P = official input price per token (non-cached)
  • C_write = official cache write price per token (often higher than P)
  • C_read = official cache read price per token (often much lower than P)
  • N = number of times the cached prefix is reused (including the first use)

Without caching, the cost for N uses of the same prefix is:
Costnocache = N × P

With caching, the first use incurs the cache write cost, and subsequent uses incur the cache read cost:
Costcache = Cwrite + (N - 1) × C_read

Since the platform applies a 1.3 multiplier uniformly, the multiplier cancels out in the comparison. The break-even condition Costcache < Costno_cache simplifies to:
Cwrite + (N - 1) × Cread < N × P

Solving for N:
N > (Cwrite - Cread) / (P - C_read)

So the minimum number of reuses needed for caching to pay off is the smallest integer N that satisfies this inequality.

Example

Suppose a provider charges:

  • P = $3 per million input tokens
  • C_write = $3.75 per million tokens (25% premium for cache write)
  • C_read = $0.30 per million tokens (90% discount for cache read)

Then the break-even N is:
N > (3.75 - 0.30) / (3 - 0.30) = 3.45 / 2.70 ≈ 1.28

Since N must be an integer, N = 2 reuses (i.e., using the cached prefix twice) already makes caching cheaper. The platform's 1.3 multiplier does not change this threshold.

Important Considerations

  • Provider-specific pricing: The exact break-even point depends on the provider's cache pricing. Always check the provider's official documentation for Cwrite and Cread.
  • Cache lifetime: Caches may expire after a certain time. If the reuse happens after expiration, the cache write cost is incurred again, affecting the break-even calculation.
  • Multiplier uniformity: Because the 1.3 multiplier applies to all token types, it does not alter the relative economics of caching. It only scales the absolute costs.
  • Other factors: Some providers may have minimum token requirements for caching or charge for cache storage separately. These should be included in C_write or as additional terms.

Conclusion

Yes, the 1.3x multiplier applies to cached tokens, but it does not change the break-even point for prompt caching. The decision to use caching should be based on the provider's pricing and your expected reuse count. Use the formula above to determine if caching is beneficial for your use case.