NTH

Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation

AuthorsNura Aljaafari, Andre Freitas

September 1, 2026 2 min read
Watch on YouTube
The one-line take

This paper offers a four-part language for explaining how transformers align, combine, suppress, and route information across different models.

Key results

15
Models evaluated

Number of models tested across five architecture families.

0.59
Routing-copy correlation

Strongest held-out Spearman correlation between a weight-space signature and its activation-level role measure.

10
Alignment-ablation effect

Models in which ablating alignment heads reduced downstream suppression beyond random-head controls.

14
Contradiction-sensitive models

Models showing a significant raw coherence-proxy shift under explicit contradictions.

0.98
Induction preservation

Lowest base-to-instruction-tuned induction-head correlation; reported correlations ranged from 0.98 to 1.00.

0.68
Suppression cross-task stability

Mean cross-task correlation for suppression, compared with 0.45 for unification.

What the paper found

This paper introduces Coherentist Probabilistic Compositionalism, or CPC, a shared vocabulary for transformer computation based on four roles: alignment proposes relations through query-key compatibility, unification integrates compatible information through additive OV or feedforward writes, suppression subtracts incompatible alternatives, and routing transports selected content to the readout. The framework was evaluated across 15 models from five architecture families, including GPT-2, Pythia, Qwen 2.5, Gemma 2, and LLaMA 3.2, using indirect object identification, greater-than comparison, and CounterFact factual recall. Weight-space signatures for suppression, unification, and routing correlated positively with held-out activation measures, while the copy-based routing score was the strongest predictor at Spearman r=0.59. Ablating alignment heads reduced downstream suppressive activity beyond random-head controls in 10 models, but comparable effects on no-conflict prompts suggest a general upstream dependency rather than contradiction-specific circuitry. Explicit contradictions significantly altered the raw coherence proxy in 14 models; after whitening residual covariance, the predicted direction appeared in every model, indicating that unprocessed similarity mainly captures global state changes. Base and instruction-tuned variants preserved induction-head structure with correlations from 0.98 to 1.00, although operator shifts were architecture-dependent rather than consistently concentrated in later layers. Suppression was also more stable across tasks than unification, with mean correlations of 0.68 versus 0.45. CPC therefore offers a useful cross-model interpretive language, but its role assignments remain graded hypotheses rather than mechanistic proofs.

Original abstract

Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis