Explaining Attention with Program Synthesis
AuthorsAmiri Hayes, Belinda Li, Jacob Andreas
Resources
This paper turns transformer attention heads into small executable Python programs, offering a new way to explain and partially replace opaque neural behavior with human-readable code.
Key results
one synthesized program per head across four models
fewer than this many candidates were synthesized
total Claude Sonnet 4 API cost in dollars
best-program alignment on GPT-2
best-program alignment on TinyLlama-1.1B
best-program alignment on Llama-3B
What the paper found
Explaining Attention with Program Synthesis, from NJIT and MIT EECS, reframes transformer interpretability as executable code search rather than natural-language labeling: the authors extract attention maps from BERT-Base, GPT-2-Small, TinyLlama-1.1B, and Llama-3B, then use Claude Sonnet 4 to synthesize Python programs that approximate each head’s behavior on TinyStories and rank them with Jensen-Shannon distance and Intersection over Union. The resulting library contains 1,664 programs built from fewer than 4,000 candidates at about $150 API cost, and the best-fit symbolic proxies can reach up to 99% mean IoU on some heads. Across models, mean best-program IoU rises with scale, reaching 69% for GPT-2, 74% for TinyLlama-1.1B, and 79% for Llama-3B, while BERT-base is hardest to capture because bidirectional attention is less structurally constrained. The causal test is stronger than alignment alone: replacing as many as 25% of attention heads with synthesized programs increases perplexity by only 16%, and substituting 30–40% of heads does not significantly degrade performance on HellaSwag, PIQA, SciQ, ARC-Easy, Social IQA, or COPA. The key result is that many attention heads are not just describable but functionally substitutable by symbolic Python programs, giving a direct path from mechanistic interpretability to model editing.
Original abstract
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected training examples. Next, we prompt a pre-trained language model with a summary of these matrices, and instruct it to generate a set of Python programs that can reproduce the associated attention patterns given only text from the input sentence. Finally, we re-rank programs according to how well our final set of programs predict behavior on held-out inputs. We demonstrate that a set of fewer than 1,000 such generated programs can reproduce the attention patterns of heads in GPT-2, TinyLlama-1.1B, and Llama-3B, achieving an average Intersection-over-Union similarity above 75% on TinyStories. Moreover, the best-fit programs can replace neural attention heads without substantially affecting model behavior: replacing 25% of attention heads with programmatic surrogates across the three models incurs only a 16% average perplexity increase, while maintaining performance on a variety of downstream question answering benchmarks. This work contributes a scalable pipeline for reverse-engineering attention heads in transformer models using human-readable, executable code, advancing a path toward symbolic transparency in neural models.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.