NTH

The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture

AuthorsZhihua Liang

July 25, 2026 3 min read
Watch on YouTube
The one-line take

This paper recasts Transformer behavior as continuous geometry and thermodynamics, offering an intriguing but currently insufficiently substantiated mathematical lens.

Key results

-0.5000
RMSNorm scaling exponent

Measured ϵ scaling exponent across Qwen3, Gemma-3, and LLaMA layers.

1.000000
RMSNorm regression fit

Coefficient of determination for all 15 scaling regressions.

4500
RMSNorm measurements

Independent Jacobian measurements in the conical-singularity experiment.

0.793
Qwen3 torsion median

Median cosine similarity between observed splitting drift and the Lie bracket.

142955
Qwen3 symmetric-ablation growth

Residual-norm amplification factor after forced FFN symmetrization.

32-48
Context phase-transition interval

Sequence-length interval where random-token perplexity underwent a sharp collapse.

What the paper found

Zhihua Liang proposes a continuous geometric interpretation of the Transformer, representing token sequences as a one-dimensional measure lattice and hidden states as sections of a semantic fiber bundle. The framework translates RMSNorm into a smooth radial embedding, RoPE into a flat gauge connection, Softmax Attention into a nonlocal Urysohn–Volterra operator, the FFN into a local Hodge reaction field, and SGD into an Itô diffusion that generally violates detailed balance. Across Qwen3, LLaMA-3.1, Gemma-3, GPT-2 from OpenAI, and Mistral models spanning 124M to 8B parameters, six experiments test predictions about stability, representation drift, context limits, and optimization. The RMSNorm study measured 4500 Jacobians across Qwen3, Gemma, and LLaMA, recovering an exact ϵ^-1/2 law with exponent -0.5000 and R2 = 1.000000. Lie–Trotter interference aligned observed drift with the Attention–FFN commutator, producing median cosine similarities of 0.793 for Qwen3 and 0.842 for GPT-2. Symmetrizing mature FFNs caused Qwen3-0.6B residual norms to grow 142955-fold, while anti-symmetric perturbations restored stability, supporting the proposed role of geometric vorticity. Under random-token inputs, Qwen3, LLaMA-3.1, and Gemma-3 showed a sharp perplexity transition between sequence lengths 32 and 48, interpreted as Attention Sink evaporation and thermodynamic amnesia. Finally, persistent parameter-space vortices under both AdamW and pure SGD suggest that non-equilibrium optimization is a general property of singular neural parameter manifolds, not an artifact of momentum.

Original abstract

We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle $\calE = \calM \times \R^d$. Beginning from a single geometric axiom -- that the token sequence forms a discrete $1$-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning $124$M to $8$B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the $ε^{-1/2}$ Lipschitz scaling calibration at machine precision ($R^2 = 1.000$), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the $\calO(1/\sqrt{k})$ thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.

Read the original paper

More in Transformers

Browse all 42 papers →
03Transformer

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang

A small LoRA update appears to make transformers carry information through many more layers, dramatically extending their ability to follow long chains without retraining the full model.

Read analysis