Each post explains one idea for a technical investor or an infrastructure engineer, and links to the published result behind it. Every number in the posts is read from a published file when the pages are built. Subscribe by feed.
A published cache hit rate, recomputed from the publisher's own files
A published prefix-cache hit rate recomputed from the publisher's own dataset matched, under controls that had to fail first and a pass mark fixed in advance.
What a KV cache is, and why some AI providers share it between customers
A key-value cache saves a language model's work on a prompt, and some providers share it because many prompts start with text that others have already paid for.
When shared caches leak: how one customer's prompt can show up for another
A shared cache can tell one customer whether another sent a given prompt opening, and a per-customer salt, built into vLLM, stops exact-match reuse across customers when every request carries one.
No demand cache that starts empty can beat your traffic: the hit-rate ceiling on a public trace
On a test slice of our excerpt of a public trace of real requests, no demand cache that starts empty can beat a proven hit-rate ceiling, and a simulated least-recently-used cache lands far below it.
Where should the KV cache live? Pooled placement across a GPU fleet
On one fixed test instance, pooling cached work across a fleet reached the fewest misses the lab's model allows, and an outside solver confirmed it.
Out-of-order data and the memory it costs: what a receiver must remember
When a network spreads a transfer over many paths, data arrives out of order, and a receiver that tracks the whole window needs memory that grows with it.