Align, Unify, Suppress, Route: A Coherentist View of Transformer Computation
AuthorsNura Aljaafari, Andre Freitas
Resources
This paper offers a four-part language for explaining how transformers align, combine, suppress, and route information across different models.
Key results
Number of models tested across five architecture families.
Strongest held-out Spearman correlation between a weight-space signature and its activation-level role measure.
Models in which ablating alignment heads reduced downstream suppression beyond random-head controls.
Models showing a significant raw coherence-proxy shift under explicit contradictions.
Lowest base-to-instruction-tuned induction-head correlation; reported correlations ranged from 0.98 to 1.00.
Mean cross-task correlation for suppression, compared with 0.45 for unification.
What the paper found
This paper introduces Coherentist Probabilistic Compositionalism, or CPC, a shared vocabulary for transformer computation based on four roles: alignment proposes relations through query-key compatibility, unification integrates compatible information through additive OV or feedforward writes, suppression subtracts incompatible alternatives, and routing transports selected content to the readout. The framework was evaluated across 15 models from five architecture families, including GPT-2, Pythia, Qwen 2.5, Gemma 2, and LLaMA 3.2, using indirect object identification, greater-than comparison, and CounterFact factual recall. Weight-space signatures for suppression, unification, and routing correlated positively with held-out activation measures, while the copy-based routing score was the strongest predictor at Spearman r=0.59. Ablating alignment heads reduced downstream suppressive activity beyond random-head controls in 10 models, but comparable effects on no-conflict prompts suggest a general upstream dependency rather than contradiction-specific circuitry. Explicit contradictions significantly altered the raw coherence proxy in 14 models; after whitening residual covariance, the predicted direction appeared in every model, indicating that unprocessed similarity mainly captures global state changes. Base and instruction-tuned variants preserved induction-head structure with correlations from 0.98 to 1.00, although operator shifts were architecture-dependent rather than consistently concentrated in later layers. Suppression was also more stable across tasks than unification, with mean correlations of 0.68 versus 0.45. CPC therefore offers a useful cross-model interpretive language, but its role assignments remain graded hypotheses rather than mechanistic proofs.
Original abstract
Mechanistic interpretability has identified transformer circuits, but lacks a shared vocabulary for describing how their functions compose across tasks and architectures. We introduce Coherentist Probabilistic Compositionalism (CPC), an interpretive framework that grounds transformer computation in coherentist theories of interpretation and describes it through four operator roles. Alignment identifies candidate relations, unification integrates supporting information, suppression reduces incompatible alternatives, and routing carries selected information to the output. Across 15 models from five architecture families, the suppression, unification, and routing weight-space signatures correlate with held-out activation-level role measures above random baselines. Suppression is more stable across tasks than unification. Ablating alignment heads reduces downstream suppressive activity beyond a random-head control in 10 models, but similar effects on no-conflict prompts indicate a general upstream dependency, not contradiction-specific coupling. Explicit contradictions significantly shift a layerwise coherence proxy in 14 models; after removing shared residual covariance, the gap has the predicted direction in every model. Base and instruction-tuned variants preserve induction-head score structure ($r{\geq}0.98$) without a consistent shift of operator signatures towards later layers. These results support CPC as a shared vocabulary for comparing transformer mechanisms while showing that their depth and geometric expression remain architecture-specific.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.