Fixed-depth Transformers with softmax attention approximate Hölder and Sobolev sequence functions at parameter rates epsilon^{-dxn/gamma} and epsilon^{-dxn}, and yield regression rates under beta-mixing data.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
stat.ML 1years
2025 1verdicts
ACCEPT 1representative citing papers
citing papers explorer
-
Approximation Bounds for Transformer Networks with Application to Regression
Fixed-depth Transformers with softmax attention approximate Hölder and Sobolev sequence functions at parameter rates epsilon^{-dxn/gamma} and epsilon^{-dxn}, and yield regression rates under beta-mixing data.