Dashboard

What Is a Semantic Cache? Cheaper AI Answers Explained

Semantic caching skips the model when a new question means the same as an old one. The savings are real and so are the false hits.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
1 September 20261 min read

A semantic cache stores past questions and answers, and returns a stored answer when a new question means roughly the same thing as an old one. Not the same characters. The same meaning. "How do I cancel my plan" and "where's the cancel button for my subscription" are different strings, and a semantic cache treats them as the same lookup. That is the whole idea, and it is both the reason it saves money and the reason it occasionally returns something wrong.

How the match actually happens

The mechanism is short. Every incoming question gets converted into an embedding, a list of numbers that positions the text by meaning. That vector is compared against the vectors of questions you have answered before. If the closest stored question scores above a similarity threshold, you skip the model entirely and return the stored answer.

Three moving parts, and each one is a decision:

  • The embedding model. Cheap and fast, and it does not need to be the model that writes your answers.

  • The similarity threshold. Usually cosine similarity between about 0.80 and 0.95, depending on how much you trust a near-match.

  • The eviction rule. How long a cached answer stays valid before it is thrown away.

The threshold is where the engineering lives. Everything else is plumbing.

This is not prompt caching, and it is not a KV cache

Three things share the word cache in this space and they operate at completely different layers. Mixing them up leads to expecting savings that do not arrive.

Cache type

What it stores

Who decides a hit

Saves

KV cache

Attention state inside one generation

The runtime, automatically

Compute during a single response

Prompt caching

An exact reusable prefix of your prompt

The provider, on an exact prefix match

Input token cost on repeated context

Semantic cache

Whole question and answer pairs

Your code, on an embedding similarity score

The entire model call

Prompt caching requires the beginning of your prompt to be byte-identical to a previous call, and it still runs the model. Anthropic's prompt caching documentation is explicit about this: cache hits require 100% identical prompt segments up to the cached block, and cache reads are billed at 0.1x the base input price rather than free. A semantic cache does not run the model at all when it hits. The KV cache is a lower-level detail inside the generation itself and you generally do not control it. They compose fine, and they are not substitutes.

The arithmetic, which is more favourable than people expect

Take a support assistant handling 20,000 questions a month. Say a model call costs you 1.4 cents in tokens, and an embedding call costs 0.002 cents. Assume nothing else changes.

At a 35% hit rate, 7,000 questions never reach the model. You still pay for 20,000 embeddings, which is about 4 cents in total and effectively free. You save roughly 98 dollars a month against a 280 dollar bill.

The reason support workloads hit that high is Zipf's law showing up in customer questions: a small number of questions account for a very large share of the volume. Password resets, refund policy, delivery times, and where the invoice is. Once those are cached, a third of your traffic never touches a model again.

The reason a general purpose assistant does not hit 35% is the same law running the other way. When every question is genuinely different, there is nothing to hit. If cost is what brought you here, the broader options are in how to reduce AI API costs.

Latency is often the bigger win anyway. A cache hit answers in tens of milliseconds against a second or more for a generated reply.

Where semantic caching goes wrong

Negation is nearly invisible to embeddings. "Can I cancel after the trial" and "can I cancel before the trial" sit very close together in vector space and mean opposite things. So do "is X included" and "is X not included". This is the classic false hit, and it is not fixed by nudging the threshold, because the two sentences are genuinely similar as text.

Personal context leaks. "What is my current balance" is a question whose correct answer differs per user. If your cache key is the question alone, you will eventually serve one customer another customer's answer. Cache keys must include the identity and scope that the answer depends on, or those questions must be excluded from caching entirely.

Answers go stale silently. A cached response to "how much does the Pro plan cost" survives your price change perfectly happily. Cached answers need a time to live and an explicit invalidation hook on the events that change them.

The threshold trap. Set it low and you get false hits, which are worse than no cache because they are confidently wrong. Set it high and you get almost no hits, which is a lot of infrastructure for nothing. There is no universal correct number, which is why you measure rather than pick.

How to set the threshold without guessing

Log real questions for a week with caching turned off. Then, offline, compute similarity for every pair and look at the pairs that fall between 0.80 and 0.95. Read fifty of them by hand and mark each one as "same answer would be fine" or "different answer needed".

The threshold you want is the point above which almost none of your hand-marked pairs need different answers. For most support corpora that lands somewhere near 0.90, but the number matters far less than the fact that you derived it from your own traffic instead of a blog post.

Then keep a small holdout: send a random 5% of hits to the model anyway and compare the fresh answer with what the cache would have returned. That is your ongoing false hit rate, and it is the only honest measure of whether the cache is still safe.

When it is worth building

A semantic cache pays for itself when your traffic is repetitive, the answers are the same for everyone, and staleness is bounded. Support assistants, documentation search, and internal FAQ bots all qualify.

It is a poor fit when answers are personalised, when they depend on live data, or when volume is low enough that the whole model bill is smaller than the effort of maintaining a cache. Under a few thousand calls a month, do something else. Understanding how AI models work at the level of what each call actually costs you tends to make that judgement easier.

FAQ

What is the difference between a semantic cache and a normal cache?

A normal cache matches on an exact key. A semantic cache matches on meaning, using embedding similarity, so differently worded questions with the same intent can hit the same entry.

Does a semantic cache reduce hallucinations?

Indirectly, on the cached paths, because a reviewed cached answer is returned verbatim instead of being regenerated. It also introduces its own error mode, since a false hit returns a confident answer to a question that was never asked.

What similarity threshold should I use?

Derive it from your own logged traffic rather than adopting a default. Many support workloads land near 0.90 cosine similarity, but the correct value depends entirely on how similar your near-miss questions are.

Can a semantic cache serve one user another user's data?

Yes, if the cache key is only the question text. Any answer that depends on who is asking must include that identity in the key, or be excluded from caching.

How do I stop cached answers going stale?

Give every entry a time to live, and invalidate explicitly on the events that change the underlying facts, such as a price change or a policy update. Time to live alone is not enough for anything that can change suddenly.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.

What Is a Semantic Cache? Cheaper AI Answers Explained | swarmz.net