Resources
The paper argues that AI agents should stop recomputing the same prompt over and over, and instead buy and reuse precomputed KV caches to make LLM serving much cheaper.
Key results
greedy decoding equivalence between loaded KV and from-scratch prefill
largest absolute logits deviation in the correctness check
reported upper bound of N⋆ for KV reuse
fp16 KV size for the 3774-token context
hosted priced cache-read cost for one hot 3774-token document
What the paper found
In “Can I Buy Your KV Cache,” Luoyuan Zhang of Harbin Institute of Technology, Shenzhen argues that a large language model’s key–value cache should be treated as a reusable commodity rather than recomputed for every reader. The paper shows that, for the safe shared-prefix case, loading a precomputed cache and continuing generation is token-exact with from-scratch prefill under greedy decoding, matching 24/24 tokens and the logits up to a maximum absolute difference of 0.02 on Qwen3-4B. On Qwen3-4B in fp16, the compute cost of reusing a resident cache is 9–50× lower than re-prefilling across context lengths from 255 to 3774 tokens, with the speedup widening as length grows because attention cost scales super-linearly with sequence length. The measured break-even is essentially immediate, with N⋆ in the range 1.02–1.13, so the second read already pays off. The paper also quantifies the economics of distribution: a 3774-token document produces a 557 MB KV artifact, and serving it to 80M agents would cost about $1.5M if re-prefilled, about $0.03M in reuse compute, or about $150,960 under a 0.1pin cache-read tariff. By contrast, shipping the raw fp16 artifact is unattractive because lossless egress is nearly incompressible and costs about $3.9M at commodity transfer prices. A simple int8 test halves a 102.1 MB artifact to 51.1 MB but breaks exactness, confirming that lossless or selective compression is needed if reproducibility matters. The result is a concrete proposal for a provider-hosted “prefill CDN” that caches computation, not bytes, and lets agents buy access to a model-bound KV artifact instead of paying to recompute it.
Original abstract
Right now, across the world, AI agents are repeating the same absurd act: to read one document, they each recompute it from scratch. Every agent re-runs prefill, the most compute-intensive step a large model takes, over identical text, only to rebuild a key-value (KV) cache identical to the one the agent before it just built. The same answer, computed a million times. We make a proposal that is almost offensively simple: compute it once. Let a publisher precompute a document's KV cache, and let every other agent buy the right to load it and skip prefill. It works, and it is token-exact: loading a precomputed KV and continuing matches prefilling from scratch (24/24 greedy tokens, and at the logits level), with no accuracy cost. On Qwen3-4B, reuse is 9-50x cheaper in compute than prefill, and the gap widens with length (prefill's attention scales with L^2), so a single reuse already pays it back. Then the part that matters: where the KV lives. Shipping it fails, because KV is nearly incompressible, so per-load egress costs more than the prefill it saves. Hosting it provider-side, exactly as production prompt-caching works, removes egress entirely. The size of the prize is set by our measured compute saving: serving one hot 3774-token document to 80M agents costs ~$1.5M to re-prefill but only ~$0.03M of reuse compute (49.7x less). The 0.1x cache-read tariff APIs charge passes a 10x discount to users while sitting inside this measured envelope, so the 10x is a floor that the measured ~50x compute saving clears, and the gap to the physical ~50x is provider margin: millions of dollars per popular document. We frame the resulting agent-native prefill CDN and leave lossless KV compression and a cross-party payment layer as the open problems.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.