Cost · AI FinOps
Semantic Cache Hit-Rate Savings Calculator
A semantic cache answers a repeated or very similar question from memory. You pay a small lookup instead of a full model call. This page turns that idea into dollars.
Enter how busy the API is, how long each answer is, and what percent of questions the cache already knows. An in-browser AI agent shows the dollars avoided. No API key.
Three words that drive the bill
- Miss. The question is new. The model writes a full answer. You pay input tokens and output tokens.
- Hit. The question is close to one already stored. You pay a lookup. You do not pay the model again.
- Hit rate. Out of 100 questions, how many are hits. You type this. The page does not read your logs.
A worked example, before you touch a slider
2 questions per second. 800 tokens in, 400 tokens out. $2.50 per million input tokens and $10 per million output tokens. 40% hit rate. Lookup costs $0.50 per 1,000 questions. Setup costs $15,000 once.
- 2 questions per second is about 5.2 million questions in a 30-day month.
- One full answer costs about $0.006. Without a cache the month is about $31,100.
- A 40% hit rate avoids about $12,400 of that model spend.
- Lookups on every question cost $2,592. Net savings are $9,850 a month, $118,195 a year.
- $15,000 of setup pays back in about 1.5 months at these defaults. Move the sliders and the AI agent recalculates your own case.
How the model works
- Monthly questions = questions per second × 2,592,000 (30 days).
- One full answer = (input tokens ÷ 1,000,000 × input price) + (output tokens ÷ 1,000,000 × output price).
- Bill without a cache = monthly questions × one full answer.
- Model spend avoided = bill without a cache × hit rate.
- Lookup bill = monthly questions × (lookup $ per 1,000 ÷ 1,000). Charged on hits and misses.
- Monthly net = model spend avoided − lookup bill. Annual net is that times 12.
- Break-even hit rate = lookup cost of one question ÷ cost of one full answer.
The AI agent fills this line after it runs.
Unlock a printable summary
The numbers above stay visible with no email. A work email unlocks Print / Save as PDF. Same subscribe flow as the newsletter.
FAQ
What is a semantic cache?
A store of earlier model answers. When a new question is close enough to one already stored, the system returns the stored answer instead of calling the model again. “Close enough” is a similarity match, not only an exact copy of the words.
How does this calculator work?
You set questions per second, input and output tokens, price per million tokens, hit rate, lookup cost, and a one-time setup cost. The page turns questions per second into a monthly volume, prices a full model call, then subtracts lookup cost from the calls the cache avoids. Payback is setup cost divided by monthly net savings.
Do I need an API key?
No. No OpenAI, Anthropic, Google, or other API key. The formulas run in your browser. Results stay on screen before any email ask. Email is only for the printable summary.
Is the AI agent calling a cloud model?
No. The AI agent on this page is the in-browser analysis experience: a badge, a short thinking step, and a result stamp. It uses rule-based heuristics. No prompt or number is sent to a third-party model.
Are the savings guaranteed?
No. Results are directional heuristics. The hit rate is a number you type, not a measurement from your logs. A cheap wrong answer is still wrong if the cached text is stale. Validate with a pilot and your real token invoice.
When does a cache fail to save money?
When the hit rate is below the break-even line shown on the page. That happens if lookups are expensive or almost every question is new. The agent shows that line next to your hit rate so you can see the gap.
Vatsal Shah
AI Leader · Solution Architect · TPM
https://shahvatsal.com
Semantic Cache Hit-Rate Savings
Prepared via shahvatsal.com/tools/semantic-cache-hit-rate-savings-calculator
Directional heuristic. The AI agent on this page is rule-based and in-browser. Not an invoice. © Vatsal Shah — shahvatsal.com