TokTier: Exact Stateful Tokenization for Agentic LLM Serving
AuthorsZhenyu Zhang, Zhichao Cao
Resources
TokTier makes agentic LLM serving faster by reusing and repairing tokenization state exactly instead of repeatedly tokenizing long conversation transcripts from scratch.
Key results
Interactive Claude Code and Codex serving calls used to characterize session continuations.
Fleet-level token-weighted cache hit rate, despite full-text rescanning by conventional tokenizers.
Observed divergence across tested tokenizer families, real-text sweeps, and replay campaigns.
Milliseconds to tokenize a 1 M-character request with array output.
Requests per second from a four-core repair pool plus one GPU under a 50 ms P99 objective.
What the paper found
TokTier, from Arizona State University researchers Zhenyu Zhang and Zhichao Cao, targets a bottleneck in agentic LLM serving: Claude Code and Codex repeatedly resend long transcripts after small tool-result appends, while vLLM’s prefix caching only helps after tokenization. Across 153951 calls, the fleet prompt-cache hit rate was 94.1%, yet tokenization rose to 64% of time to first token at very high cache-hit levels. TokTier maintains per-session token IDs and byte spans, retokenizes the append plus a 512-character suffix, and splices results only after a tokenizer-specific stable-boundary certificate proves reference equivalence; failed checks widen the window or fall back to full tokenization. For state misses, its NVIDIA GPU implementation reformulates GPT-family regex pre-tokenization into parallel character-class runs and applies exact GPU BPE, covering configurations including Meta’s Llama 3.1, Qwen3, gpt-oss, and DeepSeek families. Differential campaigns over 17 tokenizer families and 12.4 TB of real text observed 0 divergence. Incremental repair remains 0.5–1.1 ms from 100 K to 3 M characters, while GPU full tokenization handles a 1 M-character request in 0.87 ms, 23.4× faster than the strongest prior CPU configuration. Integrated with vLLM, TokTier reduces median time to first token by 16–34% in loaded regimes and enables 1821 requests/s under a 50 ms P99 target, versus 40 requests/s for a 16-core stateless CPU front end.
Original abstract
LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end of the previous sequence. Across 153,951 calls from two agent ecosystems, the median call appends about 1.4K characters, and only 1.0-3.6% of calls start or rebuild a session with contexts of millions of characters. At a 94.1% fleet prompt-cache hit rate, tokenization reaches up to 64% of time to first token. TokTier is a stateful tokenization service with one contract: emitted token IDs are always identical to full reference tokenization of the request text. For a session continuation, it re-tokenizes a small window around the append and splices only after a per-request stable-boundary check, widening the window or falling back to full tokenization on failure. For a call without a reusable prefix, it decomposes GPT-family regex pre-tokenization into run-local rules and runs exact pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU full tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method. With vLLM, median time to first token drops 16-34% and P99 drops 23% under recorded bursts. Under a 50 ms P99 objective, four repair cores plus one GPU sustain 1,821 requests/s where a 16-core stateless front end saturates at 40.
Read the original paperMore in Efficient AI
Browse all 55 papers →Decoding Looped Transformers Better for (Almost) Free
Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.
Scaling Laws for Looped Mixture of Experts
Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi
This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
Yiyu Liu, Minlan Yu, Juncheng Yang
For LLM prefix caches, simple recency may beat fancy eviction rules, especially when workloads follow predictable session patterns.