Can Computation from Earlier Problems Help LLMs Solve New Ones?
AuthorsJipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song
AffiliationsGaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · University of Modena and Reggio Emilia, Italy*Corresponding author: rsong@ruc.edu.cn
STAIR helps language models reuse useful computation from earlier questions, improving performance on later problems with only a tiny trainable module.
Key results
Percentage-point improvement over Native across the evaluated Qwen models and benchmarks.
Only the learned query-reflection directions are trained; the backbone stays frozen.
Percentage-point Avg@4 improvement over Native for Qwen3.5-4B at T2.
Median additional seconds in the fixed-workload runtime test.
Additional GiB versus Native in the fixed-workload runtime test.
What the paper found
A conversation’s earlier problems can change how an LLM solves a new one: retained history sometimes helps and sometimes hurts. The paper introduces STAIR, which saves attention keys and values from earlier reasoning in a read-only bank, then uses learned Householder reflections to redirect current queries across that bank. It subtracts the original read from the redirected read, adding only that difference during prompt processing; the Qwen backbone and stored states remain frozen. Trained on DAPO-Math-17k, STAIR improves mean later-turn Avg@4 by up to 11.67 percentage points over Native across three Qwen models and four benchmarks, using just 12,288 trainable parameters. In a matched-history AIME 2025 test with Qwen3.5-4B, STAIR scored 48.33% Avg@4 versus 21.67% for Native, a 26.67-point gain. The method’s full-bank read has a measured median overhead of 0.81 seconds and 6.60 GiB of peak memory in a fixed workload, so the accuracy gains come with additional prefill cost. The results suggest earlier computation can be reused productively, but they cover four-turn sessions and show that benefits vary by model and task.
Original abstract
Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
Read the original paperMore in Attention Mechanisms
Browse all 21 papers →Mechanics of Long-Context Hybrid Models Part 1.1: From Hybrid Attention to Hybrid Position
Xiaoran Liu, Ziwei He, Xipeng Qiu
This paper explains why different hybrid attention designs succeed or fail at long-context modeling and introduces a method that extends context length efficiently without additional training.
Universal interpolation for deep residual self-attention networks
Sibylle Marcotte, Joan Bruna
This work shows that remarkably small, frozen attention systems can still transform arbitrary token sequences into one another simply by choosing how—and how long—to apply them.
CoWindow Attention: Full Causal Coverage Is a Collective Property
Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo
CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.