Prompt Caching Savings Calculator

See the real monthly savings from caching a repeated system prompt or context block, at your provider's cache discount and how often the cache actually hits.

$0 saved / month

Cache-write requests (the first call that fills the cache) often cost slightly more than a normal call on some providers; this estimate treats them as full price, which is close enough for planning.

Why the cached prefix, not total tokens, is what decides the savings

Prompt caching only discounts the part of a request that's identical to a recent previous call — usually a system prompt, a tool schema, or a big shared context block pasted at the top of every request. The unique part (the actual user message, the part that changes every time) never benefits, no matter how the cache is configured. That's why the two numbers that matter most here are the split between cached and unique tokens, and the cache hit rate: a large shared prefix reused across thousands of near-identical calls (a fixed system prompt, a long document being queried repeatedly) can cut input cost by most of what caching discounts, while a workload with little repeated context barely moves.

Hit rate matters because caches expire — most providers keep a cache warm for a few minutes, so a burst of requests using the same prefix hits it repeatedly, while sparse, spread-out traffic keeps paying full price to refill it. Providers also differ in mechanics: some cache automatically and discount by default, others require you to mark the cached portion explicitly and may charge a small premium the first time they write it. Check your provider's current caching docs for the exact numbers; the preset here gets you in the right ballpark for planning.

FAQ

What should I put as the 'cached prefix'?

Anything sent identically on most requests: system prompts, tool/function definitions, a pasted document or knowledge base chunk, few-shot examples. If it changes every call, it belongs in 'unique tokens' instead.

What's a realistic cache hit rate?

Depends on your traffic pattern and how long the provider keeps a cache warm (often a few minutes). Bursty, high-volume traffic against the same prefix can see 80-95% hit rates; sparse or spread-out traffic sees much less, since the cache goes cold between calls.

Does caching affect output token cost?

No. Caching discounts apply to input (prompt) tokens only — the model still generates and charges for output tokens normally, which is why the output rate field doesn't change between the with/without scenarios here.

Doing this manually every week?

The n8n Launch Pack automates the operations behind these numbers, and The Prompt Operating System covers the decisions. Both are a one-time payment away (PayPal or USDT).

Send exactly USDT

After sending, paste your transaction ID (TxID / hash) from your wallet or exchange withdrawal. Payment is verified on-chain and your download opens instantly.

or
Pay with PayPal

Paying with PayPal? Your download link is emailed to your PayPal address within 24h — or email your receipt for faster delivery.

Paid more than 48h ago or trouble verifying? Email your TxID with order code and we deliver within 24h.