Skip to content

Evidence 7 min read

A published cache hit rate, recomputed from the publisher's own files

A published prefix-cache hit rate recomputed from the publisher's own dataset matched, under controls that had to fail first and a pass mark fixed in advance.

A dotted rust underline marks a number read straight from a published file when this page was built.

In this post
  1. What the lab recomputed
  2. How the check was set up
  3. What the recomputation found
  4. What the same tools say about four Alibaba traces
  5. What this means for you
  6. What this does not show

What the lab recomputed

The lab recomputed a hit rate that SemiAnalysis published for its own public dataset, using nothing but the block identifiers inside that dataset.

The dataset holds 739 sessions of an AI agent, each recorded as a trace of requests. Every request lists its prompt as a row of blocks, 64 tokens each, and gives every block an identifier. Two requests that share an identifier share a block that a prefix cache could reuse. The identifiers are local, so they can be compared inside one trace but not across traces.

The dataset card states one headline figure. SemiAnalysis published a hit rate of 96.57% across all requests and all traces. The card defines a hit as the opening of a request that matches the request before it in the same session. It counts the first request of every trace as a miss, because nothing can be cached before it.

The earlier post on hit-rate ceilings explains what a hit rate is, and why the first lookup of any block must miss.

How the check was set up

The check was fixed before it ran: the definitions, the pass mark and the controls were written down first, and a changed pass mark would have stopped the run.

The lab wrote a pre-registration before the run. It had three parts:

  • Two definitions. The first is the publisher’s own: the common opening of a request with the request before it. The second is weaker: any block seen before in the same trace, as if a cache never forgot anything.
  • A pass mark. The recomputed rate under the first definition must fall within 0.005 of the published rate, as a fraction, and the block totals must match. If it does not, the lab publishes which definition reproduces the figure, or neither.
  • Two controls. They run first and are built to break the method if it is wrong. In one, every block identifier is replaced by a fresh one, so both definitions must read zero. In the other, the identifiers inside each request are shuffled, so the first definition must collapse and the second must not.

The runner checks a fingerprint of the pass mark before it spends anything. A changed pass mark stops the run instead of quietly redefining success.

What the recomputation found

Under the publisher’s own definition, the recomputed rate matches the published one, well inside the pass mark.

  • Published: 0.9657, as a fraction.
  • Recomputed, publisher’s definition: 0.9657, the same at the precision the publisher printed.
  • Recomputed, weaker definition: 0.9678, which is 0.002132 from the published figure.

The block totals match the publisher’s count exactly. The gap under the publisher’s definition is far smaller than the gap under the weaker one, and even the weaker one lands inside the pass mark. The lab scored 59,204 requests. A further 70 carry no block identifiers, so they could not be scored.

A number that someone else published was checked here by arithmetic on their own files, and the check could have failed. Had the figure not recomputed, the rule fixed in advance would have had the lab report which definition reproduces it, or neither. The agreement shows that the published figure is arithmetically sound under the definition on the card.

What the same tools say about four Alibaba traces

The same tooling, with controls run first, shows how much of the reuse in four public Alibaba Bailian traces a cache on the GPU alone captures.

Alibaba published four anonymized traces from its Bailian service: a chat-style interactive service, API-driven task automation, reasoning-focused chat and code generation. Each carries identifiers for 16-token blocks, hashed with a salt that the provider chose.

The question here differs from the first check. It asks how much of each trace’s reuse a GPU-only cache of 65,536 tokens captures, compared with the ceiling set by repeats. Two controls ran first:

  • A control trace, the one the receipt names as mooncake_timed_600, had to reproduce the gap the receipt records as expected for it, 13.0039 percentage points, and the audit produced 13.004.
  • A trace with every identifier made fresh had to read zero reuse and be refused, and it was.

The results, one trace at a time, are the hit rate of a least-recently-used (LRU) cache on the GPU alone and the ceiling for that trace:

  • Chat-style service. LRU reaches 0.152303 against a ceiling of 0.597975.
  • API task automation. LRU reaches 0.305622 against 0.615273.
  • Reasoning chat. LRU reaches 0.25042 against 0.461842.
  • Code generation. LRU reaches 0.147633 against 0.663723.

The pass mark, set in advance in its own pre-registration, was a gap of at least 12.0 percentage points between the ceiling and the GPU-only cache. All four traces cleared it. That gap is the most that memory beyond the GPU could add on these traces, for a demand cache counted from an empty start and measured against the LRU stand-in. It is a limit on the benefit, not a forecast. The ceiling is the same idea as in the lab’s published ceiling result.

The audit also asked what spreading requests over a fleet costs. Round-robin across the largest fleet it tried, eight replicas, cut the hit rate in every trace. Keeping each session on one replica won a little back in three traces and lost more in the code trace, where the hit rate fell from 0.147633 on one replica to 0.025134 with round-robin and 0.006353 with each session kept together. The post on placement explains why a fleet’s hit rate depends on where requests land.

What this means for you

A hit rate that others can recompute is worth more than one they must take on trust, and each kind of reader can ask for that.

  • Inference providers. When you quote a hit rate, publish the block identifiers and the definition with it. Say whether the first turn counts, what the denominator is, and whether the scope is one session or the whole fleet.
  • GPU clouds. Compare any quoted hit rate with the ceiling for the customer’s own traffic. The gap between a GPU-only cache and that ceiling is the room that extra memory tiers could fill.
  • Enterprise AI platforms. When a vendor quotes a hit rate, ask whether it counts blocks, tokens or dollars, and whether the figure can be recomputed from files you can see.
  • Auditors. Ask for the pass mark in advance and for controls that could have failed. A figure that reproduces shows the arithmetic is right. It does not show the definition suits your purpose.

What this does not show

This is arithmetic on published files, and it says nothing about anyone’s engine.

  • The sharing question. The block identifiers in the SemiAnalysis dataset are local to each trace, so sharing across sessions is not measurable on this data.
  • Blocks, not dollars. The rate counts blocks, not tokens or money. The first turn of every trace counts as a miss, exactly as the publisher counted it.
  • Not the engine. The recomputation says nothing about any serving engine. That is why it can grade the publisher’s number: it uses only the publisher’s identifiers.
  • Hits, not GPU hours. The Alibaba audit counts block hits. Turning them into GPU time needs the share of GPU time that is bandwidth-bound, which was not measured.
  • A stand-in policy. LRU stands in for an engine’s real prefix tree, and that stand-in was not checked on these traces.
  • No customer labels. The Alibaba traces carry no customer identifier, so sessions are rebuilt from chat links, and keeping a customer’s requests together could not be tested.
  • Short windows. Each trace covers a short sampling window, and the provider anonymized the identifiers and chose the salt.

Results behind this post

All results from AxiomLimit

Keep reading

All posts

Sources and further reading

Where the numbers come from

Each number with a dotted rust underline was read from one of these published files, field by field, when the page was built.

Every file this site publishes

Ask about a result, or check one yourself

Every result on this site comes with the files it was measured from. Acquisition, licensing and partnership enquiries go to one address, and a person reads it.

Write to us Read the research results

How the numbers on this page are checked

Every number on this page links to the file it comes from. All published files.

  • Loadingshown only once its file has loaded in your browser
  • Not checkeda question we have not checked would say so, with no number
  • Checked, nothing founda search that found nothing would say so, with no number
  • No valuea file that holds no value for the question would say so
  • File missinga number whose file is missing or altered would be hidden
  • Unclear subjecttwo files that disagree about what a number describes would both be shown
  • Small samplea number from a small sample would carry its sample size
  • Conflicting filestwo files giving different values would both be shown
  • Not publishablea file we may not publish would be named by its fingerprint only
  • Out of datea measurement older than a week would carry its age
  • Does not applya question that does not apply to this page would say so
  • Run faileda measurement whose program failed would say so, with no number