NTH

Can Computation from Earlier Problems Help LLMs Solve New Ones?

AuthorsJipei He, Wenhui Tan, Xiaoyi Yu, Enver Sangineto, Fiorenzo Parascandolo, Rita Cucchiara, Ruihua Song

AffiliationsGaoling School of Artificial Intelligence, Renmin University of China, Beijing, China · University of Modena and Reggio Emilia, Italy*Corresponding author: rsong@ruc.edu.cn

October 11, 2026 2 min read
Watch on YouTube
The one-line take

STAIR helps language models reuse useful computation from earlier questions, improving performance on later problems with only a tiny trainable module.

Key results

11.67
Maximum later-turn Avg@4 gain

Percentage-point improvement over Native across the evaluated Qwen models and benchmarks.

12,288
STAIR trainable parameters

Only the learned query-reflection directions are trained; the backbone stays frozen.

26.67
Matched-history AIME 2025 gain

Percentage-point Avg@4 improvement over Native for Qwen3.5-4B at T2.

0.81
Full-bank runtime overhead

Median additional seconds in the fixed-workload runtime test.

6.60
Full-bank peak memory overhead

Additional GiB versus Native in the fixed-workload runtime test.

What the paper found

A conversation’s earlier problems can change how an LLM solves a new one: retained history sometimes helps and sometimes hurts. The paper introduces STAIR, which saves attention keys and values from earlier reasoning in a read-only bank, then uses learned Householder reflections to redirect current queries across that bank. It subtracts the original read from the redirected read, adding only that difference during prompt processing; the Qwen backbone and stored states remain frozen. Trained on DAPO-Math-17k, STAIR improves mean later-turn Avg@4 by up to 11.67 percentage points over Native across three Qwen models and four benchmarks, using just 12,288 trainable parameters. In a matched-history AIME 2025 test with Qwen3.5-4B, STAIR scored 48.33% Avg@4 versus 21.67% for Native, a 26.67-point gain. The method’s full-bank read has a measured median overhead of 0.81 seconds and 6.60 GiB of peak memory in a fixed workload, so the accuracy gains come with additional prefill cost. The results suggest earlier computation can be reused productively, but they cover four-turn sessions and show that benefits vary by model and task.

Original abstract

Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.

Read the original paper

More in Attention Mechanisms

Browse all 21 papers →
03Attention

CoWindow Attention: Full Causal Coverage Is a Collective Property

Jingze Shi, Zhangyang Peng, Xianduo Li, Yanlin Qi, Xiaotian Lin, Haoxian Chen, Liangdong Wang, Guang Liu, Yuyu Luo

CoWindow Attention makes long-context transformers faster by letting attention heads collectively cover the past instead of redundantly reading all of it.

Read analysis