NTH

The Emergent Symbolic Structure of Artificial Neural Networks

AuthorsR. Thomas McCoy, Paul Soulos, Tal Linzen, Paul Smolensky

September 6, 2026 2 min read
Watch on YouTube
The one-line take

The paper argues that neural networks, including LLMs, may secretly organize their continuous representations according to discoverable symbolic structures.

Key results

7
LLMs analyzed

Open-weight models in which DISCOVER identified emergent role-filler structure.

0.973
Lowest sequence-model approximation

Lowest average bidirectional DISCOVER accuracy among 12 architecture-task combinations.

99.98%
Reversing GRU approximation

Lowest approximation accuracy across 10 reruns of the reversing GRU.

0.76
GPT-OSS weakest task accuracy

Accuracy on syllogistic reasoning.

0.903
GPT-OSS intervention accuracy

Average accuracy across 31 causal intervention types.

What the paper found

This paper argues that neural networks can develop implicit symbolic structure rather than relying only on unstructured continuous vectors. Its DISCOVER method fits internal representations with linearly transformed Tensor Product Representations, encoding fillers such as words, numbers, or code tokens together with roles such as subject, object, or sequence position. On synthetic copying, reversing, and interleaving tasks, the method successfully reconstructed representations from MLPs, GRUs, Transformers, and bottleneck Transformers; the weakest bidirectional result still reached 0.973 approximation accuracy, while a reversing GRU reached 99.98%. The analysis extended to 7 open-weight language models—Gemma-3-27b, GPT-2-XL, OpenAI's GPT-OSS-20b, Pythia-12b, Qwen3-14b, OLMo-2-13B, and Llama-3.1-8b—showing role-filler structure in list and sentence representations. In GPT-OSS-20b, DISCOVER approximations closely preserved performance across arithmetic, syllogisms, Python code execution, passivization, tense reinflection, and question formation; the model's weakest original task was syllogistic reasoning at 0.76 accuracy. Causal “constituent surgery” interventions changed fillers or structural roles and produced the intended outputs with 0.903 average accuracy across 31 intervention types, including edits to arithmetic expressions, code variables, and sentence syntax. The authors conclude that modern neural systems are approximately symbolic: learned vector spaces can encode systematic compositional structure, while retaining the flexibility and noise tolerance of differentiable computation.

Original abstract

Modern systems in artificial intelligence (AI) somehow excel in domains for which they seem poorly suited. Intelligence has traditionally been modeled as operating over structured combinations of symbols, such as logical formulas. However, the strongest modern AI systems are based on neural networks, which instead represent information in continuous vectors. Vectors seem inadequate for capturing the structure of language, logic, and other cognitive domains, yet neural networks achieve impressive performance in these areas. How do they do it? In this work, we propose a potential answer: Despite appearances, perhaps the internal representations of neural networks implicitly realize symbolic structure. In support of this hypothesis, we show that the vector representations of a variety of neural networks can be closely approximated with symbolic structures: we can replace the network's entire representation-generating process with a closed-form equation instantiating a symbolic structure, and the network's behavior remains largely unchanged. This finding holds for both small-scale neural networks trained to manipulate lists as well as large language models (LLMs) operating in four domains that are central in symbolic traditions: arithmetic, logic, computer code, and language. Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations, showing that the LLM's behavior is reliant on the symbolic structures we have identified. This work provides a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI.

Read the original paper

More in AI Reasoning

Browse all 39 papers →
02Reasoning

On Language Drift during RLVR Post-Training

Michael Sullivan, Alexander Koller

RLVR can make reasoning models increasingly use strange internal languages, and preventing that drift may require sacrificing some performance.

Read analysis