Fixed-depth Transformers with softmax attention approximate Hölder and Sobolev sequence functions at parameter rates epsilon^{-dxn/gamma} and epsilon^{-dxn}, and yield regression rates under beta-mixing data.
Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Transformers are a type of neural network that have demonstrated remarkable performance across various domains, particularly in natural language processing tasks. Motivated by this success, research on the theoretical understanding of transformers has garnered significant attention. A notable example is the mathematical analysis of their approximation power, which validates the empirical expressive capability of transformers. In this study, we investigate the ability of transformers to approximate column-symmetric polynomials, an extension of symmetric polynomials that take matrices as input. Consequently, we establish an explicit relationship between the size of the transformer network and its approximation capability, leveraging the parameter efficiency of transformers and their compatibility with symmetry by focusing on the algebraic properties of symmetric polynomials.
fields
stat.ML 1years
2025 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Approximation Bounds for Transformer Networks with Application to Regression
Fixed-depth Transformers with softmax attention approximate Hölder and Sobolev sequence functions at parameter rates epsilon^{-dxn/gamma} and epsilon^{-dxn}, and yield regression rates under beta-mixing data.