NTH

Token Reduction Is Not Cost Reduction

AuthorsSarel Weinberger, Amir Hozez

August 2, 2026 2 min read
Watch on YouTube
The one-line take

For coding agents, cutting tokens is not the same as cutting costs—and aggressive compression can even break successful software patches.

Key results

2848
Analyzed Claude Code runs

Provider-billed runs used for the main cost and success analysis

87%
Cache share of reconstructed cost

Cache creation and reads dominated the four-component cost reconstruction

8.7%
Unattributed billed-cost residual

Dollar-weighted portion not explained by retained telemetry

38.4%
RTK-ML raw tool-token reduction

Estimated reduction measured with the local tiktoken o200k_base tokenizer

6.8%
RTK-ML paired cost change

Billed cost increased relative to baseline despite token reduction

1.464
Headroom cost-per-success ratio

Cost per successful execution relative to the Claude Code baseline

What the paper found

PointFive researchers Sarel Weinberger and Amir Hozez tested whether context compression actually lowers the cost of API-based coding agents. In a paired campaign of 2908 provider-billed Anthropic Claude Code runs, with 2848 analyzed across 103 tasks, 7 repositories, and Claude Haiku 4.5, Sonnet 5, and Opus 4.8, prompt-cache creation and reads dominated spending, accounting for approximately 87% of reconstructed cost, while 8.7% of billed dollars remained unattributed. The central result is a decoupling between token and dollar efficiency: RTK-ML, a hook-based system combining query-aware retrieval, lexical and BGE embedding ranking, ast-grep structural hints, and preservation gates, removed 38.4% of estimated raw tool-output tokens but increased paired billed cost by 6.8%; per-task token reduction predicted cost change only weakly, with Pearson r = 0.154. The mechanism was trajectory inflation: compressed agents issued additional diagnosis, testing, and retrieval turns that retransmitted cached context. By contrast, deterministic RTK was nearly cost-neutral, while the API-boundary Headroom v0.27.0 proxy raised cost per successful execution to 1.464 times the Claude Code baseline. A separate SWE-bench-derived Go study showed that compression could destroy byte-exact SEARCH/REPLACE anchors, reducing patch application from 27/40 to 15/40. The authors therefore recommend evaluating compression through production-path activation, task success, provider-billed cost, trajectory length, and ultimately success-adjusted cost, while preserving dense evidence such as tracebacks, test logs, structured shell streams, and edit anchors.

Original abstract

Context-reduction layers for API-based coding agents, including command-output compressors, retrieval rankers, and API-boundary proxies, are commonly evaluated by how much context or tool output they remove. We ask a different question: which interventions actually reduce end-to-end billed cost while preserving task success? Our primary evidence is a pre-specified, hash-frozen, paired campaign of 2,908 provider-billed Claude Code runs, of which 2,848 were analyzed, covering 103 tasks, seven repositories, and three models. The campaign compared a baseline with two generations of hook-based compression and an API-boundary proxy within a broader measured program of roughly 5,500 billed executions. Three findings emerge. First, prompt-cache traffic dominated cost composition, accounting for about 87% of reconstructed four-component cost (about 80% of the actual bill), with an 8.7% dollar-weighted residual not attributable from retained telemetry. Second, local payload reduction was not a reliable predictor of end-to-end billed cost. An arm that removed 38% of estimated raw tool-output tokens incurred 6.8% higher paired cost (95% CI: +2.8% to +11.3%), while per-task reduction showed only a weak association with cost change (Pearson r = 0.15). Third, aggressive compression can remove action-critical evidence: on SWE-bench-derived Go tasks, compression reduced successful patch application from 27/40 to 15/40 by corrupting verbatim edit anchors. We propose evaluating context-reduction systems by success-adjusted billed cost rather than token reduction alone.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis