On the Architectural Complexity of Neural Networks
AuthorsNicholas J. Cooper, François G. Meyer, Michael L. Roberts, Carlos Zapata-Carratalá, Lijun Chen, Danna Gurari
This paper gives a new theory for measuring how architecturally complex a neural network is and uses it to generate thousands of previously unexplored designs.
Key results
novel higher-complexity architectures publicly released
single-head self-attention arity complexity
highlighted sampled architecture accuracy on CIFAR-100
MobileNetV2 comparison on CIFAR-100
What the paper found
Nicholas J. Cooper, François G. Meyer, Michael L. Roberts, Carlos Zapata-Carratalá, Lijun Chen, and Danna Gurari present a combinatorial theory for neural networks that models architectures as rank-5 hierarchical complexes built from elements, generalized tensors, mode maps, tensor operations, and full network diagrams. The key technical move is to represent tensor operations explicitly rather than hiding them behind high-level layer abstractions, which lets the authors define architectural complexity through five signatures: operation complexity, tensor complexity, arity complexity, order complexity, and coupling-arity complexity. Using this framework, they analyze eight canonical architectures over 40 years and show that major jumps, such as from ResNet to Transformer, align with the first appearance of higher arity operations; in their encoding, a single-head self-attention block has Cop 3, CT 7, Cα 3, CO 4, and CA 2, while a residual block has Cop 3, CT 6, Cα 2, CO 6, and CA 2. They then systematically generate 3,028 novel higher-complexity architectures, evaluate them on CIFAR-10, CIFAR-100, and Tiny Imagenet, and find that some small sampled blocks are highly parameter efficient: one highlighted CIFAR-100 model reaches 65.52% accuracy with fewer than 200,000 parameters, surpassing MobileNetV2’s 64.29% with an order of magnitude fewer parameters, and the same model reaches 66.32% after 120 epochs. The paper also introduces a neural-network engine and a tensor-operation calculus, including tensor operation matrices and tensor equation matrices, to make these constructions executable and searchable.
Original abstract
We introduce a unified theoretical framework for the rigorous analysis and systematic construction of deep neural networks (DNNs). This framework addresses a gap in existing theory by explicitly modeling the structure of tensor operations -- lower level information that is often abstracted. Our framework enables two novel objectives: (1) analysis of the evolution of architectural complexity over deep learning history, and (2) automatic construction of novel architectures based on new types of tensor operations. Our study of DNNs introduced over the past 40 years reveals a connection between groundbreaking architectures and increases in different types of architectural complexity. Moreover, we identify several large classes of higher complexity architectures that have not yet been explored. We then collect a dataset of 3,000+ higher complexity architectures, which we publicly release at: https://github.com/combinatoriallabs/ArchitecturalComplexity.
Read the original paperMore in Neural Networks
Browse all 22 papers →End-to-End Hard-Label Cryptanalytic Model Extraction Using Efficient Sign Recovery
Akira Ito, Takayuki Miura, Yosuke Todo
A new query-efficient technique makes it possible to steal the parameters of small black-box neural networks using only their predicted labels.
Retrieving Individual Stems from Music Mixtures with Slot Embeddings
David Braun, Junyi Fan, Pranay Manocha, Donald S. Williamson, Adam Finkelstein
Stembed lets music producers search for individual instrument sounds hidden inside a full song by representing the mixture as multiple searchable stem-like embeddings.
The Linear Representation Hypothesis Needs a Group Action
Louie Hong Yao, Yuhao Li, Shengchao Liu
This paper argues that claims about linear representations only become meaningful once we specify which transformations leave a representation essentially unchanged.