Trace as State: Reasoning Traces as Conditional States for Long-Context Transformers
AuthorsXu Zou, Jie Tang
Resources
The paper shows that putting a model’s reasoning trace before a long document can dramatically improve its ability to solve long-context reasoning tasks.
Key results
TRACE AS STATE outperformed TRACE APPEND in 26 of 27 reported model-task-metric combinations.
The main comparison covered 27 combinations across three models, tasks, and metrics.
TRACE AS STATE exact match for DeepSeek V4 Pro Preview on GraphWalks Parents.
TRACE AS STATE exact match for GLM-5.2 on GraphWalks Parents.
What the paper found
TRACE AS STATE treats a model’s reasoning traces as an imperfect textual representation of task state. After an initial pass over a long prompt, it serializes one or more traces and prepends them to the same context on a fresh causal pass, so the model can use information discovered late while rereading earlier tokens. Its control, TRACE APPEND, places the identical traces after the context; they can influence later generation but not representations already formed. The paper formalizes this distinction with conditional state-update tasks, showing that revealing the condition before a sequence can require exponentially less working memory than revealing it afterward. Across three models—DeepSeek V4 Pro Preview, GLM-5.2, and Qwen 3.7 Max—and three long-context benchmarks—OpenAI’s GraphWalks 256K, MRCRv2 8-needle, and NUB-1M—TRACE AS STATE wins in 26 of 27 model-task-metric combinations. On GraphWalks Parents exact match, DeepSeek V4 Pro Preview improves from 29.2% initially and 43.0% with TRACE APPEND to 81.8% with TRACE AS STATE; GLM-5.2 rises from 66.4% and 83.2% to 100.0%. Ablations show that problem-specific traces outperform answer feedback, random traces, rereading alone, and even selecting the best of five first-pass answers. The method preserves causal attention within each pass but adds inference cost, latency, and dependence on APIs that expose reasoning traces or another reusable state interface.
Original abstract
Transformers process information causally, but long-context reasoning may depend on task state discovered only later. We formalize this mismatch through conditional state update tasks. For causal state update processors, providing the condition first can require exponentially less memory in the worst case than providing it last. Motivated by this principle, we introduce Trace as State. We use collected reasoning traces as a textual proxy for task state and place it before the long-context block on a fresh pass, allowing information derived previously to guide rereading. We conduct extensive experiments on Trace as State and Trace Append, a matched control that uses the same task state proxy but put it after the context. Across three models and three long-context datasets, Trace as State outperforms Trace Append in 26 of 27 reported combinations of model, task, and metric. On GraphWalks Parents, exact match lifts DeepSeek V4 Pro Preview from 29.2% on the initial pass and 43.0% with Trace Appendto 81.8% with Trace as State, and from 66.4% and 83.2% to 100.0% for GLM-5.2. These results show that placing traces before the context can improve long-context reasoning while retaining the causal transformer structure.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.