Token Reduction Is Not Cost Reduction
AuthorsSarel Weinberger, Amir Hozez
Resources
For coding agents, cutting tokens is not the same as cutting costs—and aggressive compression can even break successful software patches.
Key results
Provider-billed runs used for the main cost and success analysis
Cache creation and reads dominated the four-component cost reconstruction
Dollar-weighted portion not explained by retained telemetry
Estimated reduction measured with the local tiktoken o200k_base tokenizer
Billed cost increased relative to baseline despite token reduction
Cost per successful execution relative to the Claude Code baseline
What the paper found
PointFive researchers Sarel Weinberger and Amir Hozez tested whether context compression actually lowers the cost of API-based coding agents. In a paired campaign of 2908 provider-billed Anthropic Claude Code runs, with 2848 analyzed across 103 tasks, 7 repositories, and Claude Haiku 4.5, Sonnet 5, and Opus 4.8, prompt-cache creation and reads dominated spending, accounting for approximately 87% of reconstructed cost, while 8.7% of billed dollars remained unattributed. The central result is a decoupling between token and dollar efficiency: RTK-ML, a hook-based system combining query-aware retrieval, lexical and BGE embedding ranking, ast-grep structural hints, and preservation gates, removed 38.4% of estimated raw tool-output tokens but increased paired billed cost by 6.8%; per-task token reduction predicted cost change only weakly, with Pearson r = 0.154. The mechanism was trajectory inflation: compressed agents issued additional diagnosis, testing, and retrieval turns that retransmitted cached context. By contrast, deterministic RTK was nearly cost-neutral, while the API-boundary Headroom v0.27.0 proxy raised cost per successful execution to 1.464 times the Claude Code baseline. A separate SWE-bench-derived Go study showed that compression could destroy byte-exact SEARCH/REPLACE anchors, reducing patch application from 27/40 to 15/40. The authors therefore recommend evaluating compression through production-path activation, task success, provider-billed cost, trajectory length, and ultimately success-adjusted cost, while preserving dense evidence such as tracebacks, test logs, structured shell streams, and edit anchors.
Original abstract
Context-reduction layers for API-based coding agents, including command-output compressors, retrieval rankers, and API-boundary proxies, are commonly evaluated by how much context or tool output they remove. We ask a different question: which interventions actually reduce end-to-end billed cost while preserving task success? Our primary evidence is a pre-specified, hash-frozen, paired campaign of 2,908 provider-billed Claude Code runs, of which 2,848 were analyzed, covering 103 tasks, seven repositories, and three models. The campaign compared a baseline with two generations of hook-based compression and an API-boundary proxy within a broader measured program of roughly 5,500 billed executions. Three findings emerge. First, prompt-cache traffic dominated cost composition, accounting for about 87% of reconstructed four-component cost (about 80% of the actual bill), with an 8.7% dollar-weighted residual not attributable from retained telemetry. Second, local payload reduction was not a reliable predictor of end-to-end billed cost. An arm that removed 38% of estimated raw tool-output tokens incurred 6.8% higher paired cost (95% CI: +2.8% to +11.3%), while per-task reduction showed only a weak association with cost change (Pearson r = 0.15). Third, aggressive compression can remove action-critical evidence: on SWE-bench-derived Go tasks, compression reduced successful patch application from 27/40 to 15/40 by corrupting verbatim edit anchors. We propose evaluating context-reduction systems by success-adjusted billed cost rather than token reduction alone.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.