How a shared cache can leak
A shared cache leaks when what one customer sent earlier changes what another customer observes later.
The cache does not hand over the other customer’s text. It hands over something smaller, a signal. A request whose opening is already cached is answered faster than one whose opening is not. An attacker who can send requests and time the replies can therefore ask a yes-or-no question: has anyone sent this exact opening recently?
Two conditions make this work:
- The attacker must be able to send requests to the same cache the victim uses.
- The attacker must be able to tell a fast reply from a slow one.
One guess tells little. Many guesses tell a lot. If the target prompt follows a known template, an attacker can try candidate next words, keep the ones that come back fast, and build the prompt up piece by piece. Researchers call this a token-by-token search.
Network noise and load blur the signal, so the studies below repeat their probes and apply statistical tests. Some serving engines also report how many prompt tokens came from the cache, and vLLM is one of them. Where a customer can read that counter, the signal is exact instead of noisy.
What public research has shown
Several research groups have shown that the timing signal is real and can be measured on real systems.
- Song and colleagues described timing side channels that arise from shared caches and from GPU memory allocation. They tested popular online services as black boxes and proposed a token-by-token search that recovers shared prompt prefixes.
- Wu and colleagues showed that the way engines such as vLLM and SGLang share a KV cache between users may allow unauthorized reconstruction of user prompts.
- Gu and colleagues built statistical audits for public APIs. They found timing evidence of cache sharing across users at several providers, and they argue that providers should be open about their caching policies.
- Zheng and colleagues described InputSnatch, a timing attack that steals user inputs from cache-sharing inference services.
Providers have started to say how their caches are scoped. One major provider’s guide states that its caches are not shared across organizations. That is a statement of policy. It does not tell a customer whether the policy is tested.
How serving engines close the gap
The standard fix is to make the cache key depend on who is asking, so that one customer’s opening can never match another’s.
vLLM added an optional cache salt for each request. The change, titled “Prevent side-channel attacks via cache salting,” mixes the salt into the fingerprint of the first block. Only requests that carry the same salt can then reuse one another’s blocks. The field already has this mitigation. The idea is not new, and it is not ours.
A salt is a separator, not encryption. Nothing is hidden. The cache is split into separate spaces. Two limits follow:
- The salt is optional, so a request that arrives without one lands in the shared space.
- Someone has to choose the salt, usually the gateway in front of the engine. A gateway that forgets, or that gives many customers one salt, brings the leak back.
The choice of salt also sets how much sharing survives. A salt per organization keeps reuse inside each organization, and a salt per user keeps less. A narrower salt gives a stronger separation and a lower hit rate. That is the trade, and the right setting depends on who in your system can see whose prompts.
The price of any salt is lost sharing. Salting by customer gives up reuse across customers, which was the saving that made sharing attractive. A pooled cache shared across a fleet raises the same question, because a pool is shared state.
How to test a salt
A test of cache isolation has to be able to fail, so it starts by showing the leak and only then the fix.
- Show the leak first. Send a prompt as one tenant, then send a request with the same opening as a second tenant, with no salt. If the second tenant gets no cached work, the test cannot tell a working salt from a cache that was never shared.
- Then switch the salt on. Repeat with a salt tied to each tenant. The second tenant should now get no cached work from the first.
- Check that caching still works. The first tenant, repeating its own prompt, should still be served from the cache. Otherwise a zero only means the cache was off.
- Read the engine’s counter, then time the replies. The counter shows exact matches. Timing shows what an outside attacker could see.
A clean result on a few prompts bounds the leftover rate; it does not prove it is zero.
What this means for you
The fix is cheap to adopt and easy to get wrong, so the people who run, buy or check shared inference each have a part in it.
- Inference providers. Put the tenant into the cache key at the gateway. Treat a request with no salt as an error. Test that a second tenant gets no cached tokens from the first.
- GPU clouds. If tenants share one engine, its cache is shared state. Isolation there is a setting you can test, not a promise you can assume.
- Enterprise AI platforms. Ask how your provider separates caches between customers and how it tests that. Ask what happens to a request that arrives without a salt.
- Auditors. Ask for a test that reproduces the leak first and the fix second. A test that cannot fail proves nothing.
What this does not show
This post describes published research and a test method. It does not report a measurement of its own.
- It does not show a timing attack, and it says nothing about any provider’s systems.
- A salt separates exact matches only. Caches that match on meaning instead of exact text, which the research above also covers, need their own checks.
- A salt works only if the gateway sets it for every request. It is not encryption.
- The lab’s results on hit-rate ceilings and on pooled placement are separate from this topic and do not depend on it.