NTH

TokTier: Exact Stateful Tokenization for Agentic LLM Serving

AuthorsZhenyu Zhang, Zhichao Cao

August 8, 2026 3 min read
Watch on YouTube
The one-line take

TokTier makes agentic LLM serving faster by reusing and repairing tokenization state exactly instead of repeatedly tokenizing long conversation transcripts from scratch.

Key results

153951
Agent calls analyzed

Interactive Claude Code and Codex serving calls used to characterize session continuations.

94.1%
Prompt-cache hit rate

Fleet-level token-weighted cache hit rate, despite full-text rescanning by conventional tokenizers.

0
Differential divergence

Observed divergence across tested tokenizer families, real-text sweeps, and replay campaigns.

0.87
GPU full-tokenization latency

Milliseconds to tokenize a 1 M-character request with array output.

1821
Hybrid serving capacity

Requests per second from a four-core repair pool plus one GPU under a 50 ms P99 objective.

What the paper found

TokTier, from Arizona State University researchers Zhenyu Zhang and Zhichao Cao, targets a bottleneck in agentic LLM serving: Claude Code and Codex repeatedly resend long transcripts after small tool-result appends, while vLLM’s prefix caching only helps after tokenization. Across 153951 calls, the fleet prompt-cache hit rate was 94.1%, yet tokenization rose to 64% of time to first token at very high cache-hit levels. TokTier maintains per-session token IDs and byte spans, retokenizes the append plus a 512-character suffix, and splices results only after a tokenizer-specific stable-boundary certificate proves reference equivalence; failed checks widen the window or fall back to full tokenization. For state misses, its NVIDIA GPU implementation reformulates GPT-family regex pre-tokenization into parallel character-class runs and applies exact GPU BPE, covering configurations including Meta’s Llama 3.1, Qwen3, gpt-oss, and DeepSeek families. Differential campaigns over 17 tokenizer families and 12.4 TB of real text observed 0 divergence. Incremental repair remains 0.5–1.1 ms from 100 K to 3 M characters, while GPU full tokenization handles a 1 M-character request in 0.87 ms, 23.4× faster than the strongest prior CPU configuration. Integrated with vLLM, TokTier reduces median time to first token by 16–34% in loaded regimes and enables 1821 requests/s under a 50 ms P99 target, versus 40 requests/s for a 16-core stateless CPU front end.

Original abstract

LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small tool result, and reuse is hard because even a short append can change token boundaries near the end of the previous sequence. Across 153,951 calls from two agent ecosystems, the median call appends about 1.4K characters, and only 1.0-3.6% of calls start or rebuild a session with contexts of millions of characters. At a 94.1% fleet prompt-cache hit rate, tokenization reaches up to 64% of time to first token. TokTier is a stateful tokenization service with one contract: emitted token IDs are always identical to full reference tokenization of the request text. For a session continuation, it re-tokenizes a small window around the append and splices only after a per-request stable-boundary check, widening the window or falling back to full tokenization on failure. For a call without a reusable prefix, it decomposes GPT-family regex pre-tokenization into run-local rules and runs exact pre-tokenization and BPE on a GPU. A sampled shadow verifier re-checks live traffic. Across 17 tokenizer families, differential campaigns cover 1.5x10^10 split checks, a 12.4 TB real-text corpus, and 93,000+ replayed agent steps, with zero divergence. Incremental repair takes 0.5-1.1 ms from 100K to 3M characters, up to 437x faster than HF tokenization and 2.1x faster at 1M than the strongest cache-based baseline (Gigatoken) fully prewarmed. GPU full tokenization encodes a 1M-character request in 0.87 ms, up to 491x below HF and 23.4x below the fastest published CPU method. With vLLM, median time to first token drops 16-34% and P99 drops 23% under recorded bursts. Under a 50 ms P99 objective, four repair cores plus one GPU sustain 1,821 requests/s where a 16-core stateless front end saturates at 40.

Read the original paper

More in Efficient AI

Browse all 55 papers →
01Efficiency

Decoding Looped Transformers Better for (Almost) Free

Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang

LoopCD turns the partially computed states of looped Transformers into free guidance, improving accuracy while often cutting inference compute nearly in half.

Read analysis
02Efficiency

Scaling Laws for Looped Mixture of Experts

Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthi

This work develops scaling laws that explain how looping and sparse experts can be combined to build more capable models with less training and inference compute.

Read analysis