What a KV cache is
A key-value (KV) cache is the saved result of the work a large language model (LLM) has already done on the words it has read.
A model reads text in small pieces called tokens. For every token it computes a pair of vectors, a key and a value, that later tokens use to look back at it. Writing a reply means looking back at everything that came before, again and again. Without a cache, the model would redo that work for the whole conversation each time it added one more word. With a cache, it does the work once and keeps the result.
Serving a request therefore has two phases:
- Prefill. The model reads the whole prompt in one pass and builds the cache. A long prompt costs real time here on a graphics processing unit (GPU), and prefill decides how long the user waits for the first word of the reply.
- Decode. The model writes the reply one token at a time and adds a little to the cache at each step.
How prefix caching reuses it
Prefix caching reuses a saved KV cache when a new prompt starts with exactly the same text as an earlier prompt.
It has to be the opening because the key and value for a token depend on every token before it. A saved result is valid only for the same beginning. Change one early word, and everything after it must be computed again.
Serving engines turn this into a lookup. vLLM, an open-source serving engine described in its design notes, cuts a prompt into fixed-size blocks of tokens. It names each block with a hash, a short fingerprint made from the tokens in the block and the fingerprint of the block before it. Two prompts with the same opening produce the same chain of fingerprints, so the engine can find the saved blocks and skip the work. Only full blocks are cached.
- reused from the cache
- computed fresh
Two requests that open the same way. The second finds the saved blocks for the shared opening and computes only what is new.
When a lookup finds a saved block, that is a hit. When it does not, that is a miss, and the block is computed. The share of lookups that hit is the hit rate. The same notes state that prefix caching does not change model outputs.
The chain of fingerprints has a practical consequence for anyone who writes prompts. Text that changes on every request, such as a timestamp or a user name, breaks the chain if it sits near the top. Put the stable text first and the changing text last, and more of each prompt can be reused.
The fingerprint can also carry more than tokens. The notes list extra values, such as the identity of a fine-tuned adapter, and a cache salt that isolates caches in multi-tenant environments. That last item is the subject of the next post on leaks.
Where a KV cache is stored
A KV cache starts in the memory of the GPU, and larger systems spill it to main memory, to disk or to other servers.
- GPU memory is fast but small, and the cache takes a share of it that grows with the length of every prompt in flight.
- Main memory holds more but is slower to reach.
- Solid-state drives hold more again.
- Other servers add capacity across a network.
A fleet of servers has all of these, spread over many machines. That spread opens the way to a bet. If fetching a saved block from a slower place is cheaper than computing it again, a provider can keep far more saved work than one GPU could hold. The Mooncake paper builds a production serving system on this idea. It pools spare main memory, solid-state drives and network capacity across a GPU cluster into one cache, and its title states the trade: more storage for less computation.
The cost of the bet is movement. A block fetched from another machine crosses a network, and the time that takes counts against the time saved. The post on placement looks at where a fleet should keep its blocks.
Why some providers share it between customers
Those that share it do so because the cheapest computation is the computation they skip, and the more requests that share an opening, the more they skip.
Many requests open with the same text. Four common sources are:
- a long system prompt that sets the assistant’s role;
- a list of tool descriptions that follows it;
- a document that many users read;
- the earlier turns of a chat, which repeat at every step.
Each repeat is work that was already done.
A cache kept on one server helps only the requests that land on that server. A provider with many servers can pool the cache instead, so work done on one machine helps all of them. A pool that also spans customers saves more again, because openings common to many customers, such as a shared template, are computed once for all of them.
Whether a pool spans customers is a policy choice, and providers differ. One major provider’s prompt caching guide states that its caches are not shared across organizations. Researchers who audited public APIs by timing found evidence that several providers shared cache entries across users. Sharing saves money, and it has a price in what one customer can learn about another. The next post takes that price apart.
What limits the savings
The savings are limited by how often prompts repeat and by how much cache memory the provider can afford to keep.
A cache can hit only on an opening it has seen before, so the first appearance of any opening is a miss. A cache with finite memory must also forget something. Many engines forget the block used longest ago first, a rule called least recently used (LRU). The vLLM notes describe its eviction that way. A recency rule works well when repeats come soon after the last use, and badly when they come after a long gap.
Two questions follow for anyone who sizes a cache. How high can the hit rate go on this traffic at all? And how close does a given design come? The hit-rate ceiling post answers the first on a public trace. The placement post covers part of the second for a fleet.
What this means for you
The mechanism matters most to anyone who pays for, runs or checks shared inference.
- Inference providers. Decide what the cache key includes and who the cache is shared with. Sharing saves computation. Separation protects customers. Both are choices, and customers will ask which one you made.
- GPU clouds. Tenants share your machines, so the engine’s cache setting belongs in your isolation story. Know whether your tenants share one cache.
- Enterprise AI platforms. Repeated system prompts and shared documents are where caching helps most. Put stable text first in your prompts, and ask your provider what is shared and with whom.
- Auditors. Ask for the sharing policy in writing, and for a test that could show it is wrong.
What this does not show
This post explains a mechanism. It does not measure any provider.
- It reports what documentation and published audits say, linked in the sources below. It does not test any provider’s behavior.
- Savings depend on traffic. A workload with few repeated openings gains little from any cache.
- It does not say that sharing a cache is unsafe. Whether a shared cache leaks depends on how the key is built and on who can send requests.
- It explains the idea in general terms. The lab’s own results are on the research pages and are not needed to follow it.