The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
AuthorsZhihua Liang
Resources
This paper recasts Transformer behavior as continuous geometry and thermodynamics, offering an intriguing but currently insufficiently substantiated mathematical lens.
Key results
Measured ϵ scaling exponent across Qwen3, Gemma-3, and LLaMA layers.
Coefficient of determination for all 15 scaling regressions.
Independent Jacobian measurements in the conical-singularity experiment.
Median cosine similarity between observed splitting drift and the Lie bracket.
Residual-norm amplification factor after forced FFN symmetrization.
Sequence-length interval where random-token perplexity underwent a sharp collapse.
What the paper found
Zhihua Liang proposes a continuous geometric interpretation of the Transformer, representing token sequences as a one-dimensional measure lattice and hidden states as sections of a semantic fiber bundle. The framework translates RMSNorm into a smooth radial embedding, RoPE into a flat gauge connection, Softmax Attention into a nonlocal Urysohn–Volterra operator, the FFN into a local Hodge reaction field, and SGD into an Itô diffusion that generally violates detailed balance. Across Qwen3, LLaMA-3.1, Gemma-3, GPT-2 from OpenAI, and Mistral models spanning 124M to 8B parameters, six experiments test predictions about stability, representation drift, context limits, and optimization. The RMSNorm study measured 4500 Jacobians across Qwen3, Gemma, and LLaMA, recovering an exact ϵ^-1/2 law with exponent -0.5000 and R2 = 1.000000. Lie–Trotter interference aligned observed drift with the Attention–FFN commutator, producing median cosine similarities of 0.793 for Qwen3 and 0.842 for GPT-2. Symmetrizing mature FFNs caused Qwen3-0.6B residual norms to grow 142955-fold, while anti-symmetric perturbations restored stability, supporting the proposed role of geometric vorticity. Under random-token inputs, Qwen3, LLaMA-3.1, and Gemma-3 showed a sharp perplexity transition between sequence lengths 32 and 48, interpreted as Attention Sink evaporation and thermodynamic amnesia. Finally, persistent parameter-space vortices under both AdamW and pure SGD suggest that non-equilibrium optimization is a general property of singular neural parameter manifolds, not an artifact of momentum.
Original abstract
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token sequence forms a discrete $1$-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning $124$M to $8$B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the $ε^{-1/2}$ Lipschitz scaling calibration at machine precision ($R^2 = 1.000$), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the $\calO(1/\sqrt{k})$ thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.
Read the original paperMore in Transformers
Browse all 42 papers →Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva
LIFT teaches Transformers to pass rich hidden-state information across steps, potentially making language models more efficient and capable than standard feed-forward designs.
The Geometry of Inference in Transformer Residual Streams
Timur Mudarisov, Mikhail Burtsev, Radu State
This paper shows how Transformer hidden states gradually geometrically converge toward the correct prediction while eliminating competing possible outcomes.
Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It
Zehao Jin, Ruixuan Deng, Junran Wang
A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.