NTH

Explaining Attention with Program Synthesis

AuthorsAmiri Hayes, Belinda Li, Jacob Andreas

July 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper turns transformer attention heads into small executable Python programs, offering a new way to explain and partially replace opaque neural behavior with human-readable code.

Key results

1664
program library size

one synthesized program per head across four models

4000
candidate programs

fewer than this many candidates were synthesized

150
API cost

total Claude Sonnet 4 API cost in dollars

69%
GPT-2 mean best-program IoU

best-program alignment on GPT-2

74%
TinyLlama mean best-program IoU

best-program alignment on TinyLlama-1.1B

79%
Llama-3B mean best-program IoU

best-program alignment on Llama-3B

What the paper found

Explaining Attention with Program Synthesis, from NJIT and MIT EECS, reframes transformer interpretability as executable code search rather than natural-language labeling: the authors extract attention maps from BERT-Base, GPT-2-Small, TinyLlama-1.1B, and Llama-3B, then use Claude Sonnet 4 to synthesize Python programs that approximate each head’s behavior on TinyStories and rank them with Jensen-Shannon distance and Intersection over Union. The resulting library contains 1,664 programs built from fewer than 4,000 candidates at about $150 API cost, and the best-fit symbolic proxies can reach up to 99% mean IoU on some heads. Across models, mean best-program IoU rises with scale, reaching 69% for GPT-2, 74% for TinyLlama-1.1B, and 79% for Llama-3B, while BERT-base is hardest to capture because bidirectional attention is less structurally constrained. The causal test is stronger than alignment alone: replacing as many as 25% of attention heads with synthesized programs increases perplexity by only 16%, and substituting 30–40% of heads does not significantly degrade performance on HellaSwag, PIQA, SciQ, ARC-Easy, Social IQA, or COPA. The key result is that many attention heads are not just describable but functionally substitutable by symbolic Python programs, giving a direct path from mechanistic interpretability to model editing.

Original abstract

A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected training examples. Next, we prompt a pre-trained language model with a summary of these matrices, and instruct it to generate a set of Python programs that can reproduce the associated attention patterns given only text from the input sentence. Finally, we re-rank programs according to how well our final set of programs predict behavior on held-out inputs. We demonstrate that a set of fewer than 1,000 such generated programs can reproduce the attention patterns of heads in GPT-2, TinyLlama-1.1B, and Llama-3B, achieving an average Intersection-over-Union similarity above 75% on TinyStories. Moreover, the best-fit programs can replace neural attention heads without substantially affecting model behavior: replacing 25% of attention heads with programmatic surrogates across the three models incurs only a 16% average perplexity increase, while maintaining performance on a variety of downstream question answering benchmarks. This work contributes a scalable pipeline for reverse-engineering attention heads in transformer models using human-readable, executable code, advancing a path toward symbolic transparency in neural models.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis