NTH

Represented Is Not Computed: A Causal Test of Candidate Algorithmic Intermediates in a Transformer

AuthorsIshita Darade, Sushrut Thorat

May 22, 2026 3 min read
Watch on YouTube
The one-line take

This paper shows that a Transformer can represent the right intermediate arithmetic steps without actually using them in the way probes suggest, highlighting the gap between what models encode and what they compute.

Key results

99.83%
Model exact-answer accuracy

Across three independently trained 10-layer decoder-only Transformers, held-out exact-answer accuracy on number–base intersections was 99.83%.

R^2 = 1.00
B^D decodability

The intermediate B^D was linearly decoded best from the Dones stream at layer 0.

R^2 = 0.96
Quotient-like decodability

Both N/B^D and floor(N/B^D) reached mean linear probe performance of R^2 = 0.96 around layer 2.

R^2 = 0.94
Final answer decodability

The final answer quantity floor(N/B^D) mod B was strongest late in the output stream, at O[1] layer 9.

73.08%
Masked layers 0–1 accuracy

In cumulative attention ablation, masking Dones-to-output attention at layers 0–1 reduced exact-answer accuracy from 99.83% to 73.08%.

32.06%
Full Dones-to-output ablation accuracy

Masking all Dones-to-output attention across layers reduced exact-answer accuracy to 32.06%.

What the paper found

This paper tests whether a Transformer that performs base-digit extraction actually computes the closed-form intermediates suggested by the exact solution, y = floor(N/B^D) mod B, or merely represents them. The authors train 10-layer decoder-only Transformers from scratch on inputs of the form N, B, and D, with held-out number–base intersections, and obtain 99.83% mean exact-answer accuracy across three seeds. Linear probes then show that the algorithmic quantities are strongly decodable from residual streams: B^D is recovered with R^2 = 1.00 from the Dones stream at layer 0, N/B^D and floor(N/B^D) reach R^2 = 0.96 around layer 2, and the final answer term floor(N/B^D) mod B reaches R^2 = 0.94 in the output stream at layer 9. But causal tests overturn the staged-computation story. Targeted attention ablations on the Dones → O route show that performance depends mainly on the earliest D-selective communication: masking layers 0–1 drops exact accuracy from 99.83% to 73.08%, and full ablation reduces it to 32.06%. Key/value patching at layer 1 transfers behavior almost entirely by changing D, not N or B, with source-exact accuracy falling to 0.00% and donor-exact rising to 99.84% when only D differs. A sparse circuit search further finds mostly factorized N, B, and D pathways that converge late at the output, preserving 93.40% exact accuracy while retaining only 28.33% of candidate source-target relations. The central result is a decoupling: the model linearly represents candidate algorithmic intermediates, but the localized causal route to the answer does not use them as the main computation.

Original abstract

Structured prompts require integrating components according to task-relevant relations. How a network implements this integration is often hard to judge in language or vision, where those relations are rarely specified precisely enough to define a candidate internal algorithm. Arithmetic offers a cleaner setting. We study a Transformer trained on base-digit extraction: given $N$, $B$, and $D$, it must report the coefficient of $B^D$ in the base-$B$ expansion of $N$. The closed-form solution, $\lfloor N/B^D \rfloor \bmod B$, provides explicit candidate algorithmic intermediates. Across three seeds, the model reaches 99.83% exact-answer accuracy on held-out number-base intersections, establishing reliable task competence. Linear probes decode the intermediates, making staged arithmetic computation plausible. Causal tests then separate representation from use: within the localized route from the stream with $D$ as input to the output positions, behavior depends on early $D$-selective communication, independent of $N$ and $B$. Relatedly, a sparse circuit search finds mostly separate $N$, $B$, and $D$ routes that combine late rather than the staged route suggested by the probes. Thus, the model represents the intermediates that make the closed-form solution plausible, but the identified localized causal route does not transmit them to the output stream. This case shows that probe-based conclusions can diverge sharply from causal observations, even when explicit algorithmic hypotheses are available.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis