<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AxiomLimit blog</title>
    <link>https://axiomlimit.com/blog/</link>
    <atom:link href="https://axiomlimit.com/blog/rss.xml" rel="self" type="application/rss+xml" />
    <description>Plain-English posts on KV-cache sharing, hit-rate ceilings, pooled placement and network reordering, each tied to a published result.</description>
    <language>en</language>
    <lastBuildDate>Thu, 01 Oct 2026 00:00:00 GMT</lastBuildDate>
    <item>
      <title>A published cache hit rate, recomputed from the publisher's own files</title>
      <link>https://axiomlimit.com/blog/recomputing-a-published-hit-rate/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/recomputing-a-published-hit-rate/</guid>
      <pubDate>Thu, 01 Oct 2026 00:00:00 GMT</pubDate>
      <category>Evidence</category>
      <description>A published prefix-cache hit rate recomputed from the publisher's own dataset matched, under controls that had to fail first and a pass mark fixed in advance.</description>
    </item>
    <item>
      <title>What a KV cache is, and why some AI providers share it between customers</title>
      <link>https://axiomlimit.com/blog/what-is-a-kv-cache/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/what-is-a-kv-cache/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <category>Explainer</category>
      <description>A key-value cache saves a language model's work on a prompt, and some providers share it because many prompts start with text that others have already paid for.</description>
    </item>
    <item>
      <title>When shared caches leak: how one customer's prompt can show up for another</title>
      <link>https://axiomlimit.com/blog/when-shared-caches-leak/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/when-shared-caches-leak/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <category>Security</category>
      <description>A shared cache can tell one customer whether another sent a given prompt opening, and a per-customer salt, built into vLLM, stops exact-match reuse across customers when every request carries one.</description>
    </item>
    <item>
      <title>No demand cache that starts empty can beat your traffic: the hit-rate ceiling on a public trace</title>
      <link>https://axiomlimit.com/blog/cache-hit-rate-ceiling/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/cache-hit-rate-ceiling/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <category>Evidence</category>
      <description>On a test slice of our excerpt of a public trace of real requests, no demand cache that starts empty can beat a proven hit-rate ceiling, and a simulated least-recently-used cache lands far below it.</description>
    </item>
    <item>
      <title>Where should the KV cache live? Pooled placement across a GPU fleet</title>
      <link>https://axiomlimit.com/blog/where-should-the-kv-cache-live/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/where-should-the-kv-cache-live/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <category>Systems</category>
      <description>On one fixed test instance, pooling cached work across a fleet reached the fewest misses the lab's model allows, and an outside solver confirmed it.</description>
    </item>
    <item>
      <title>Out-of-order data and the memory it costs: what a receiver must remember</title>
      <link>https://axiomlimit.com/blog/out-of-order-data-receiver-memory/</link>
      <guid isPermaLink="true">https://axiomlimit.com/blog/out-of-order-data-receiver-memory/</guid>
      <pubDate>Wed, 30 Sep 2026 00:00:00 GMT</pubDate>
      <category>Networking</category>
      <description>When a network spreads a transfer over many paths, data arrives out of order, and a receiver that tracks the whole window needs memory that grows with it.</description>
    </item>
  </channel>
</rss>
