REVIEW 28 cited by
A mathematical perspective on Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A mathematical perspective on Transformers
read the original abstract
Transformers play a central role in the inner workings of large language models. We develop a mathematical framework for analyzing Transformers based on their interpretation as interacting particle systems, which reveals that clusters emerge in long time. Our study explores the underlying theory and offers new perspectives for mathematicians as well as computer scientists.
Forward citations
Cited by 28 Pith papers
-
Attention as Frustrated Synchronization
FSN achieves lower validation loss (1.5953) than a RoPE-SwiGLU transformer (1.611) on character-level tasks at 1M parameters by implementing next-token prediction as synchronization frustrated by data transitions.
-
A Unified Framework for Critical Scaling of Inverse Temperature in Self-Attention
The upper-tail accumulation scale derived from the gap-counting function N_n sets the critical inverse temperature for softmax attention concentration, unifying prior conflicting laws as special cases of different N_n.
-
Transient Reserves, Sink Dampers, and the Failure of Eigenvalue Reasoning in the Attention Propagator
Resolvent analysis of trained causal attention shows sinks act as transient dampers, routing heads carry excess Kreiss reserve, and eigenvalue depth predictions fail by 7–11 orders of magnitude.
-
Structure Before Collapse: Transient semantic geometry in next-token prediction
Semantic geometry emerges transiently early in next-token prediction training before collapsing to Neural Collapse symmetry in synthetic settings with latent semantic factors.
-
Kuramoto Attention: Synchronizing Self-Attention on the Torus
Kuramoto attention reformulates self-attention as phase synchronization on the torus and matches a matched RoPE+SwiGLU transformer within 0.02 BPC on enwiki8 at 1M-5M parameters.
-
Polynomial Context-Truncation Sensitivity in Autoregressive Language Models: Sequential Wyner-Ziv Bounds for KV Cache Compression
Under a polynomial context-truncation sensitivity assumption, suffix-only KV cache policies require per-token memory scaling as Θ(ε^{-1/α}) to achieve distortion ε.
-
The physics of AI weather models
AI weather models may simulate the atmosphere via particle positions in latent space whose updates follow gradient flow on a learned free energy functional rather than conventional physical equations.
-
Transformer-like Inference from Optimal Control
Derives transformer-like dual-filter inference layers from first-principles optimal control on nonlinear discrete and linear Gaussian sequence models.
-
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Transformer hidden states encode facts as attractor basins; hallucinations occur from basin absence and conflicts from basin competition, detected cleanly by geometric margin rather than entropy.
-
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Attractor basins in transformer hidden states unify conflict and hallucination as basin competition or absence, with geometric margin outperforming entropy for detection and a scaling law governing confident hallucina...
-
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Conflict and hallucination in transformers are basin competition versus basin absence in hidden-state space; geometric margin detects them with zero false refusals while entropy cannot, and confident hallucinations sc...
-
Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention
Multi-head self-attention is modeled as a gradient flow with a non-decreasing energy functional under conditions on score matrices, yielding closed-form clustering thresholds in simplified regimes and monotonic entrop...
-
Robust Filter Attention: Self-Attention as Precision-Weighted State Estimation
Robust Filter Attention models self-attention as consistency-based state estimation under a linear SDE for token trajectories, matching standard attention complexity while showing lower perplexity and better zero-shot...
-
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
Hardmax transformers converge to leader-determined clusters, enabling an interpretable model for sentiment analysis.
-
Self-Attention Dynamics with Rotary Position Embeddings: Twisted States and Explicit Consensus Rates on the Sphere
Normalized query/key-only RoPE attention on the sphere has reversible consensus kernels with exact Bessel-aliasing spectra, explicit regional contraction rates from a sharp softmax floor, and RoPE-selected twisted equ...
-
On Transformer Dynamics
A universal, finitely parametrized family of geometric interaction laws realizes any prescribed attention digraph, with cost governed by the biclique cover number and a new hub-chromatic index.
-
Patnaik-Pearson intrinsic dimension for internal representations of neural networks
Introduces the Patnaik-Pearson intrinsic dimension estimator, proves some of its properties, relates it to HTSR/SETOL for Pareto spectra, and applies it to track embedding dimension evolution in BERT-base and DeepSeek...
-
Patnaik-Pearson intrinsic dimension for internal representations of neural networks
Introduces the Patnaik-Pearson intrinsic dimension estimator, relates it to HTSR/SETOL for Pareto spectral densities, and applies it to measure embedding dimension evolution in BERT-base and DeepSeek-R1-Distill-Qwen-1.
-
Kuramoto Attention: Synchronizing Self-Attention on the Torus
Kuramoto Attention replaces the standard value aggregation in self-attention with a Kuramoto synchronization step on per-token phase states living on a torus, producing small average improvements over matched RoPE/Swi...
-
Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time
RCA is a training-free module that boosts input context signal strength in the residual stream of LLMs by orthogonal decoupling of attention routing from value magnitude.
-
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
Training installs a depth-dependent spectral gradient and low-rank bottleneck in LLM residual streams whose amplification or suppression of graph communities is predicted by local operator type.
-
Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention
Multi-head self-attention dynamics admit a non-decreasing energy functional under suitable score-matrix conditions, with closed-form clustering thresholds and monotonic entropy production in simplified regimes.
-
Two failure modes of deep transformers and how to avoid them: a unified theory of signal propagation at initialisation
Analytical theory of signal propagation in deep transformers at initialization yields quantitative prescriptions for weights and residuals to avoid rank and entropy collapse via Random Energy Model analogy.
-
Chaos in reason: How chain-of-thought LLMs can look for an answer
Greedy LLM inference shows bounded, jump-like sensitivity to sub-token perturbations that the authors interpret as chaotic, with attention expanding and normalization suppressing perturbations.
-
Path-Measure Dynamics of Attention-Driven World Models: A Nonlocal Onsager--Machlup Approach
Derives that attention-induced non-Markovian dynamics yield a nonlocal Onsager-Machlup action whose short-memory expansion recovers the local action of a companion paper.
-
A Path-Space Formulation of Prediction in World Models: From a Single Action to Prediction, Planning, and Irreversibility
Path-space formulation of world-model prediction via Onsager-Machlup action, with attention-based models acquiring asymmetry proportional to data irreversibility.
-
Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
Proposes a semantic information theory for LLMs that substitutes the token for the bit as the atomic carrier of meaning, recasts the Transformer as an energy-based model, and derives directed rate-distortion and rate-...
-
A Mathematical Explanation of Transformers
The Transformer is interpreted as discretization of a structured integro-differential equation in continuous domains for tokens and features, unifying attention, feedforward, and normalization via operator and variati...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.