Represented Is Not Computed: A Causal Test of Candidate Algorithmic Intermediates in a Transformer
AuthorsIshita Darade, Sushrut Thorat
Resources
This paper shows that a Transformer can represent the right intermediate arithmetic steps without actually using them in the way probes suggest, highlighting the gap between what models encode and what they compute.
Key results
Across three independently trained 10-layer decoder-only Transformers, held-out exact-answer accuracy on number–base intersections was 99.83%.
The intermediate B^D was linearly decoded best from the Dones stream at layer 0.
Both N/B^D and floor(N/B^D) reached mean linear probe performance of R^2 = 0.96 around layer 2.
The final answer quantity floor(N/B^D) mod B was strongest late in the output stream, at O[1] layer 9.
In cumulative attention ablation, masking Dones-to-output attention at layers 0–1 reduced exact-answer accuracy from 99.83% to 73.08%.
Masking all Dones-to-output attention across layers reduced exact-answer accuracy to 32.06%.
What the paper found
This paper tests whether a Transformer that performs base-digit extraction actually computes the closed-form intermediates suggested by the exact solution, y = floor(N/B^D) mod B, or merely represents them. The authors train 10-layer decoder-only Transformers from scratch on inputs of the form N, B, and D, with held-out number–base intersections, and obtain 99.83% mean exact-answer accuracy across three seeds. Linear probes then show that the algorithmic quantities are strongly decodable from residual streams: B^D is recovered with R^2 = 1.00 from the Dones stream at layer 0, N/B^D and floor(N/B^D) reach R^2 = 0.96 around layer 2, and the final answer term floor(N/B^D) mod B reaches R^2 = 0.94 in the output stream at layer 9. But causal tests overturn the staged-computation story. Targeted attention ablations on the Dones → O route show that performance depends mainly on the earliest D-selective communication: masking layers 0–1 drops exact accuracy from 99.83% to 73.08%, and full ablation reduces it to 32.06%. Key/value patching at layer 1 transfers behavior almost entirely by changing D, not N or B, with source-exact accuracy falling to 0.00% and donor-exact rising to 99.84% when only D differs. A sparse circuit search further finds mostly factorized N, B, and D pathways that converge late at the output, preserving 93.40% exact accuracy while retaining only 28.33% of candidate source-target relations. The central result is a decoupling: the model linearly represents candidate algorithmic intermediates, but the localized causal route to the answer does not use them as the main computation.
Original abstract
Structured prompts require integrating components according to task-relevant relations. How a network implements this integration is often hard to judge in language or vision, where those relations are rarely specified precisely enough to define a candidate internal algorithm. Arithmetic offers a cleaner setting. We study a Transformer trained on base-digit extraction: given $N$, $B$, and $D$, it must report the coefficient of $B^D$ in the base-$B$ expansion of $N$. The closed-form solution, $\lfloor N/B^D \rfloor \bmod B$, provides explicit candidate algorithmic intermediates. Across three seeds, the model reaches 99.83% exact-answer accuracy on held-out number-base intersections, establishing reliable task competence. Linear probes decode the intermediates, making staged arithmetic computation plausible. Causal tests then separate representation from use: within the localized route from the stream with $D$ as input to the output positions, behavior depends on early $D$-selective communication, independent of $N$ and $B$. Relatedly, a sparse circuit search finds mostly separate $N$, $B$, and $D$ routes that combine late rather than the staged route suggested by the probes. Thus, the model represents the intermediates that make the closed-form solution plausible, but the identified localized causal route does not transmit them to the output stream. This case shows that probe-based conclusions can diverge sharply from causal observations, even when explicit algorithmic hypotheses are available.
Read the original paperMore in AI Reasoning
Browse all 39 papers →Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
FlyBy teaches small reasoning models to recognize when more internal thinking will not help and instead ask a stronger model for missing knowledge.
On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller
RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.
Principled Thoughts for Latent Recursive LLM Systems
Fahd Seddik, Fatemeh Fard
REST teaches latent LLM agents to form more causal, minimal, separable, and stable internal thoughts, improving reasoning accuracy and interpretability.