Research
Each result below has its own page. Every page states the result in plain words and says what it does not claim. It names the published work it builds on and links the files it was checked against.
Receiver reorder memory for multipath transport
The Mooncake trace: hit-rate ceiling versus LRU
Placing KV cache across a GPU fleet
More results are on the way. Each gets its own page when its evidence is ready to publish.
Questions, answered plainly
- What is a KV cache?
- When a language model reads a prompt, it stores intermediate results from every layer so it does not have to recompute them for each new word it writes. That stored working state is the KV cache (short for key-value cache). Serving systems reuse it between requests that begin the same way, which saves GPU time.
- What does "machine-checked" mean?
- Our proofs are written in a proof assistant called Lean. A program checks every step, so a result does not depend on anyone’s reading of the argument. A proof still covers only what it states, and each result page says what it does not claim.
- What is a cache hit-rate ceiling?
- It is the largest share of lookups any cache of its kind could answer from memory on a given stream of traffic. A cache that starts empty and fetches only on request cannot serve more than the traffic repeats. The bound is proved in Lean; its value on a test slice of one public production trace comes from our counts of that slice, and it is a property of that trace, not of any implementation. A cache that starts warm, or one that prefetches, can go higher.
- Are these simulations or measurements on hardware?
- It depends on the result, and each page says which. The reorder-memory result is a network simulation, not silicon. The placement result is an optimum on one fixed test instance, confirmed by an outside solver on our model of that instance. The hit-rate ceiling belongs to one public trace.
- Can I check the work myself?
- Every figure on this site opens to the file it was measured from, and the checker on the home page tests a file in your own browser, without sending anything over the network. The checker confirms that a file is unaltered. Re-running the underlying experiments needs the lab’s code, which is not public yet.
- Who is AxiomLimit for?
- Teams that run or buy AI inference at scale, makers of network and accelerator hardware, and the auditors who examine them. To license or acquire a result, write to the address at the bottom of any page.
Every file this site publishes
How the numbers on this page are checked
This page prints no measured number. All published files.
- Loadingshown only once its file has loaded in your browser
- Not checkeda question we have not checked would say so, with no number
- Checked, nothing founda search that found nothing would say so, with no number
- No valuea file that holds no value for the question would say so
- File missinga number whose file is missing or altered would be hidden
- Unclear subjecttwo files that disagree about what a number describes would both be shown
- Small samplea number from a small sample would carry its sample size
- Conflicting filestwo files giving different values would both be shown
- Not publishablea file we may not publish would be named by its fingerprint only
- Out of datea measurement older than a week would carry its age
- Does not applya question that does not apply to this page would say so
- Run faileda measurement whose program failed would say so, with no number