What the lab recomputed
The lab recomputed a hit rate that SemiAnalysis published for its own public dataset, using nothing but the block identifiers inside that dataset.
The dataset holds 739 sessions of an AI agent, each recorded as a trace of requests. Every request lists its prompt as a row of blocks, 64 tokens each, and gives every block an identifier. Two requests that share an identifier share a block that a prefix cache could reuse. The identifiers are local, so they can be compared inside one trace but not across traces.
The dataset card states one headline figure. SemiAnalysis published a hit rate of 96.57% across all requests and all traces. The card defines a hit as the opening of a request that matches the request before it in the same session. It counts the first request of every trace as a miss, because nothing can be cached before it.
The earlier post on hit-rate ceilings explains what a hit rate is, and why the first lookup of any block must miss.
How the check was set up
The check was fixed before it ran: the definitions, the pass mark and the controls were written down first, and a changed pass mark would have stopped the run.
The lab wrote a pre-registration before the run. It had three parts:
- Two definitions. The first is the publisher’s own: the common opening of a request with the request before it. The second is weaker: any block seen before in the same trace, as if a cache never forgot anything.
- A pass mark. The recomputed rate under the first definition must fall within 0.005 of the published rate, as a fraction, and the block totals must match. If it does not, the lab publishes which definition reproduces the figure, or neither.
- Two controls. They run first and are built to break the method if it is wrong. In one, every block identifier is replaced by a fresh one, so both definitions must read zero. In the other, the identifiers inside each request are shuffled, so the first definition must collapse and the second must not.
The runner checks a fingerprint of the pass mark before it spends anything. A changed pass mark stops the run instead of quietly redefining success.
What the recomputation found
Under the publisher’s own definition, the recomputed rate matches the published one, well inside the pass mark.
- Published: 0.9657, as a fraction.
- Recomputed, publisher’s definition: 0.9657, the same at the precision the publisher printed.
- Recomputed, weaker definition: 0.9678, which is 0.002132 from the published figure.
The block totals match the publisher’s count exactly. The gap under the publisher’s definition is far smaller than the gap under the weaker one, and even the weaker one lands inside the pass mark. The lab scored 59,204 requests. A further 70 carry no block identifiers, so they could not be scored.
A number that someone else published was checked here by arithmetic on their own files, and the check could have failed. Had the figure not recomputed, the rule fixed in advance would have had the lab report which definition reproduces it, or neither. The agreement shows that the published figure is arithmetically sound under the definition on the card.
What the same tools say about four Alibaba traces
The same tooling, with controls run first, shows how much of the reuse in four public Alibaba Bailian traces a cache on the GPU alone captures.
Alibaba published four anonymized traces from its Bailian service: a chat-style interactive service, API-driven task automation, reasoning-focused chat and code generation. Each carries identifiers for 16-token blocks, hashed with a salt that the provider chose.
The question here differs from the first check. It asks how much of each trace’s reuse a GPU-only cache of 65,536 tokens captures, compared with the ceiling set by repeats. Two controls ran first:
- A control trace, the one the receipt names as mooncake_timed_600, had to reproduce the gap the receipt records as expected for it, 13.0039 percentage points, and the audit produced 13.004.
- A trace with every identifier made fresh had to read zero reuse and be refused, and it was.
The results, one trace at a time, are the hit rate of a least-recently-used (LRU) cache on the GPU alone and the ceiling for that trace:
- Chat-style service. LRU reaches 0.152303 against a ceiling of 0.597975.
- API task automation. LRU reaches 0.305622 against 0.615273.
- Reasoning chat. LRU reaches 0.25042 against 0.461842.
- Code generation. LRU reaches 0.147633 against 0.663723.
The pass mark, set in advance in its own pre-registration, was a gap of at least 12.0 percentage points between the ceiling and the GPU-only cache. All four traces cleared it. That gap is the most that memory beyond the GPU could add on these traces, for a demand cache counted from an empty start and measured against the LRU stand-in. It is a limit on the benefit, not a forecast. The ceiling is the same idea as in the lab’s published ceiling result.
The audit also asked what spreading requests over a fleet costs. Round-robin across the largest fleet it tried, eight replicas, cut the hit rate in every trace. Keeping each session on one replica won a little back in three traces and lost more in the code trace, where the hit rate fell from 0.147633 on one replica to 0.025134 with round-robin and 0.006353 with each session kept together. The post on placement explains why a fleet’s hit rate depends on where requests land.
What this means for you
A hit rate that others can recompute is worth more than one they must take on trust, and each kind of reader can ask for that.
- Inference providers. When you quote a hit rate, publish the block identifiers and the definition with it. Say whether the first turn counts, what the denominator is, and whether the scope is one session or the whole fleet.
- GPU clouds. Compare any quoted hit rate with the ceiling for the customer’s own traffic. The gap between a GPU-only cache and that ceiling is the room that extra memory tiers could fill.
- Enterprise AI platforms. When a vendor quotes a hit rate, ask whether it counts blocks, tokens or dollars, and whether the figure can be recomputed from files you can see.
- Auditors. Ask for the pass mark in advance and for controls that could have failed. A figure that reproduces shows the arithmetic is right. It does not show the definition suits your purpose.
What this does not show
This is arithmetic on published files, and it says nothing about anyone’s engine.
- The sharing question. The block identifiers in the SemiAnalysis dataset are local to each trace, so sharing across sessions is not measurable on this data.
- Blocks, not dollars. The rate counts blocks, not tokens or money. The first turn of every trace counts as a miss, exactly as the publisher counted it.
- Not the engine. The recomputation says nothing about any serving engine. That is why it can grade the publisher’s number: it uses only the publisher’s identifiers.
- Hits, not GPU hours. The Alibaba audit counts block hits. Turning them into GPU time needs the share of GPU time that is bandwidth-bound, which was not measured.
- A stand-in policy. LRU stands in for an engine’s real prefix tree, and that stand-in was not checked on these traces.
- No customer labels. The Alibaba traces carry no customer identifier, so sessions are rebuilt from chat links, and keeping a customer’s requests together could not be tested.
- Short windows. Each trace covers a short sampling window, and the provider anonymized the identifiers and chose the salt.