The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
Abstract
We present a continuous geometric framework that models the discrete algebraic operations of the Transformer architecture as an integro-differential equation (IDE) on a semantic fiber bundle calE = calM times R^d. Beginning from a single geometric axiom -- that the token sequence forms a discrete 1-manifold equipped with a canonical measure lattice -- we translate every core component of the modern Transformer (RMSNorm, RoPE, Softmax Attention, FFN, Residual Stream, SGD, Weight Decay) into a cohesive vocabulary of differential geometry, measure theory, and stochastic calculus. The resulting framework yields quantitative predictions spanning entropic optimal transport (Attention as a Schrödinger bridge) and non-equilibrium thermodynamics (SGD as Itô diffusion violating detailed balance). We conduct a six-part experimental campaign across five architectures (Qwen3, LLaMA\nobreakdash-3.1, Gemma\nobreakdash-3, GPT-2, Mistral) spanning 124M to 8B parameters. The empirical observables are quantitatively consistent with the geometric predictions: the ε^{-1/2} Lipschitz scaling calibration at machine precision (R^2 = 1.000), the Lie--Trotter operator-splitting torsion, the symmetric ablation instability confirming the Dual-Law of Topological Stability, the calO(1/k) thermodynamic suppression of Poincaré recurrence on the RoPE torus, the thermodynamic context-limit phase transition, and the Non-Equilibrium Steady State parameter vortex -- verified across two optimizers (AdamW and Pure SGD) to exclude momentum artifacts. The results demonstrate that analyzing Transformers through the lens of continuous stochastic differential geometry provides a predictive descriptive vocabulary for the stability limits, context bounds, and optimization dynamics of Large Language Models.
Community
Author here. This paper writes the entire Transformer as one continuous geometric system — RMSNorm, RoPE, softmax attention, the gated FFN, residual stream, and SGD each mapped to a geometric object on a semantic fiber bundle — and then puts the framework through six falsifiable prediction→measurement pairs, spanning 124M to 8B parameters:
- RMSNorm acts as a topological mollifier of a conical singularity; the predicted ε^(−1/2) activation scaling is measured at exactly −0.5.
- Layer stacking is a non-commuting flow; the torsion is measured directly with a Lie–Trotter interferometer.
- Symmetrizing the FFN of a mature LLM (W_down := W_upᵀ) detonates the forward pass (~12,000× norm amplification; Jacobian eigenspectrum collapses from 88.5% complex to 0%) — yet the same constraint trained from initialization is perfectly stable. The instability law classifies weight configurations, not architectures (onset near 1.3–2B tokens; 743× at trillion-token scale).
- RoPE is a rigid gauge connection on a torus: exact Poincaré recurrences exist, are observed at small feature dimension, and are thermodynamically suppressed as O(1/√k).
- Attention sinks behave as a Dirichlet boundary defect, with the context horizon as a thermodynamic phase boundary.
- In an FP64 sandbox, SGD shows a nonzero commutator — a dissipative non-equilibrium steady-state vortex over the gauge orbit.
The claim is deliberately modest: the geometry is the map, not the territory — a descriptive, falsifiable vocabulary rather than a new architecture. Feedback and attempts to break the predictions are very welcome.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- The Infinitesimal Structure of Quantum Information (2026)
- On Phase-Space Orthogonality for Higher-Order Distributions: Obstructions and Algebraic Resolutions (2026)
- Geometric Gradient Flows from Elliptic Level Sets: Normal Decomposition and Reflection Dynamics (2026)
- Learning as a Geometric Phase Transition: Renormalization Group Flow and Anisotropic Symmetry Breaking in Deep Networks (2026)
- Chern Character for Discrete Spectrum Partition Function (2026)
- The Geometry of Memorization: Finite-Time Spectral Sensitivity as a Diagnostic for Flow Matching Models (2026)
- The Token Is a Group Element: On Lie-Algebra Attention over Matrix Lie Groups (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.17146 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper