NTH

On the Architectural Complexity of Neural Networks

AuthorsNicholas J. Cooper, François G. Meyer, Michael L. Roberts, Carlos Zapata-Carratalá, Lijun Chen, Danna Gurari

July 6, 2026 2 min read
Watch on YouTube
The one-line take

This paper gives a new theory for measuring how architecturally complex a neural network is and uses it to generate thousands of previously unexplored designs.

Key results

3028
architectures_dataset_size

novel higher-complexity architectures publicly released

3
self_attention_Calpha

single-head self-attention arity complexity

65.52%
red_star_CIFAR100_accuracy

highlighted sampled architecture accuracy on CIFAR-100

64.29%
MobileNetV2_CIFAR100_accuracy

MobileNetV2 comparison on CIFAR-100

What the paper found

Nicholas J. Cooper, François G. Meyer, Michael L. Roberts, Carlos Zapata-Carratalá, Lijun Chen, and Danna Gurari present a combinatorial theory for neural networks that models architectures as rank-5 hierarchical complexes built from elements, generalized tensors, mode maps, tensor operations, and full network diagrams. The key technical move is to represent tensor operations explicitly rather than hiding them behind high-level layer abstractions, which lets the authors define architectural complexity through five signatures: operation complexity, tensor complexity, arity complexity, order complexity, and coupling-arity complexity. Using this framework, they analyze eight canonical architectures over 40 years and show that major jumps, such as from ResNet to Transformer, align with the first appearance of higher arity operations; in their encoding, a single-head self-attention block has Cop 3, CT 7, Cα 3, CO 4, and CA 2, while a residual block has Cop 3, CT 6, Cα 2, CO 6, and CA 2. They then systematically generate 3,028 novel higher-complexity architectures, evaluate them on CIFAR-10, CIFAR-100, and Tiny Imagenet, and find that some small sampled blocks are highly parameter efficient: one highlighted CIFAR-100 model reaches 65.52% accuracy with fewer than 200,000 parameters, surpassing MobileNetV2’s 64.29% with an order of magnitude fewer parameters, and the same model reaches 66.32% after 120 epochs. The paper also introduces a neural-network engine and a tensor-operation calculus, including tensor operation matrices and tensor equation matrices, to make these constructions executable and searchable.

Original abstract

We introduce a unified theoretical framework for the rigorous analysis and systematic construction of deep neural networks (DNNs). This framework addresses a gap in existing theory by explicitly modeling the structure of tensor operations -- lower level information that is often abstracted. Our framework enables two novel objectives: (1) analysis of the evolution of architectural complexity over deep learning history, and (2) automatic construction of novel architectures based on new types of tensor operations. Our study of DNNs introduced over the past 40 years reveals a connection between groundbreaking architectures and increases in different types of architectural complexity. Moreover, we identify several large classes of higher complexity architectures that have not yet been explored. We then collect a dataset of 3,000+ higher complexity architectures, which we publicly release at: https://github.com/combinatoriallabs/ArchitecturalComplexity.

Read the original paper

More in Neural Networks

Browse all 22 papers →
02Neural Network

Retrieving Individual Stems from Music Mixtures with Slot Embeddings

David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein

Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.

Read analysis
03Neural Network

The Linear Representation Hypothesis Needs a Group Action

Louie Hong Yao, Yuhao Li, Shengchao Liu

This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.

Read analysis