Pith. sign in

Paper Citation Record · LEDGER

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

As of 17 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2602.20555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20555 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:28:35.859983Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T05:08:00.711576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:38:40.636427Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e501e7c-54b3-422f-bb71-a1e2a5cf5896 · outbound

This paper cites Pearson, 2 edition, 1974.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Pearson, 2 edition, 1974

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.605327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.605327Z digest=sha256:5bd2eba4803ad854397b89cccfb84f1a3acdcecd01843169e023372d4720b807

Observation 2cb93dd9-59ed-4fef-814b-6c73a60efb61 · outbound

This paper cites Low-rank bottleneck in multi-head attention models.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Low-rank bottleneck in multi-head attention models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.662144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.662144Z digest=sha256:496bc66e25f3921c2889b458af4cdbf1446093f849370c15e239e204ef27432c

Observation 9dcec2f4-ee2f-438d-8a71-8b2b4c50cf3f · outbound

This paper cites an unresolved cited work.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.738468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.738468Z digest=sha256:fc7d4bd227be10f429a466f2cad131a496fa5826b7aff92abfe6a7dc80b34e63

Observation 3a154508-52de-4dab-b3d7-8c63331d65bd · outbound

This paper cites A unified framework for establishing the universal approximation of transformer-type architectures.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets A unified framework for establishing the universal approximation of transformer-type architectures

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:27.890674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:27.890674Z digest=sha256:6d3e5eb33af860f0e57f33e57165feeaebe0ea1ca30b781e220eb0b39874e4ab

Observation 0631de7d-335e-43e3-b836-7243964006c8 · outbound

This paper cites Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.002565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.002565Z digest=sha256:c34643b258ffa49e54d2581cc1d4a559a2d68c6559fdcbfbc4bb78a63a741660

Observation 3015893b-4405-4162-b14b-292db3e5d535 · outbound

This paper cites Bert: Pre- training of deep bidirectional transformers for language understanding.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Bert: Pre- training of deep bidirectional transformers for language understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.198176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.198176Z digest=sha256:795f692b1e420bb2a5207c9de19b7d33bceea962eb5b75028de97f0d25846e27

Observation 74b568fe-6acc-4a49-a4bd-22c86ff064c3 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.302711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.302711Z digest=sha256:abb98fa423a4b564d607414aaa332f174cf44141bf78bf72b31e2cb810bc2424

Observation 8460b952-662f-4668-992d-b767a01fc7ef · outbound

This paper cites Inductive biases and variable creation in self-attention mechanisms.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Inductive biases and variable creation in self-attention mechanisms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.421735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.421735Z digest=sha256:8d4bef2b5de0304d812137593b1e8bc98b828c9f69161f47e9480d6e07ad9858

Observation c42b2c03-6729-453e-b86e-560979d2eaeb · outbound

This paper cites How do noise tails impact on deep relu networks?The Annals of Statistics, 52(4):1845–1871, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets How do noise tails impact on deep relu networks?The Annals of Statistics, 52(4):1845–1871, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.506109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.506109Z digest=sha256:08a26655e5e116d7e88924f3b5ce63615e6f760c3b519cae0b41977b3bac2cea

Observation 21fca2ad-42cd-474f-8f6e-442a5c6a24a2 · outbound

This paper cites Deep neural networks for esti- mation and inference.Econometrica, 89(1):181–213, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep neural networks for esti- mation and inference.Econometrica, 89(1):181–213, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.660201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.660201Z digest=sha256:64bae643c3c5a11dcc0c8f28819fa15bdbe76851cebd3832ed6d4506b3d18efd

Observation 52cf793c-4063-4ea4-a189-cc84f8b183f8 · outbound

This paper cites Cambridge university press, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Cambridge university press, 2021

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.761626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.761626Z digest=sha256:de86e46a2e0a5ca5ab2a52051d87c9b1b0df807b3310c2d2083ac8e278e7b118

Observation 50bfc956-8029-4e2f-9fa0-23f6530e8275 · outbound

This paper cites Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation rates for neural networks with encodable weights in smoothness spaces.Neural Networks, 134:107–130, 2021

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.870976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.870976Z digest=sha256:0fd27d26f40e5eda4b2638474823eac5dbce5ce7da7ada2f5e7d9057904fb3b0

Observation bcfbd33e-b72c-4476-9626-f494cb8692e3 · outbound

This paper cites On the rate of convergence of a classifier based on a transformer encoder.IEEE Transactions on Information Theory, 68(12):8139–8155, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the rate of convergence of a classifier based on a transformer encoder.IEEE Transactions on Information Theory, 68(12):8139–8155, 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:28.932299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:28.932299Z digest=sha256:9e1a5e49d861f2310aad6edfda4a841e441312810b0b5cdd6c533dd09fafde04

Observation 29a4e6c4-1d11-4e20-8767-9cf148ade33f · outbound

This paper cites Understanding scaling laws with statisti- cal and approximation theory for transformer neural networks on intrinsically low- dimensional data.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Understanding scaling laws with statisti- cal and approximation theory for transformer neural networks on intrinsically low- dimensional data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.050031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.050031Z digest=sha256:fda545a7f689b3de7cb1378822979e948c95ed672b9bd9e1f649ab02d88517a0

Observation 84ba33a6-a5b1-478a-a558-35c34085ea7d · outbound

This paper cites Minimal width for universal property of deep rnn.Journal of Machine Learning Research, 24(121):1–41, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Minimal width for universal property of deep rnn.Journal of Machine Learning Research, 24(121):1–41, 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.159551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.159551Z digest=sha256:ba9d93382f42754e9e91208f3a40aefd29deb6ee5d9f86f01532b8282fa8a4cb

Observation 59978333-edd3-4ecd-a1e5-d7b4ee9e32ca · outbound

This paper cites Universal approximation with softmax attention.arXiv preprint arXiv:2504.15956, 2025.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Universal approximation with softmax attention.arXiv preprint arXiv:2504.15956, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.291247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.291247Z digest=sha256:1175efda60b62c5275b221fcebeb071cba5b912ca4b781142da65689c23b8a32

Observation a1ac003f-e4f8-457f-a66e-5c90c914d2ce · outbound

This paper cites Approximation rate of the transformer architecture for sequence modeling.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation rate of the transformer architecture for sequence modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.408007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.408007Z digest=sha256:59b7a239276dda9f9e0a29a2f430d6ee5b8634ffa2851468c2bc3ea892a756cb

Observation 4ae2c7f0-c677-4dca-b5c1-d7f38b197cb3 · outbound

This paper cites Approximation Bounds for Transformer Networks with Application to Regression.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation Bounds for Transformer Networks with Application to Regression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.502605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.502605Z digest=sha256:fb1b73e82eaaf82904152b277f675749060ee9b5ceb8c7da6d267d4d86ff00bf

Observation 53c04e93-8632-4176-9f9f-cd0403ff7dca · outbound

This paper cites Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation Perspective

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.665918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.665918Z digest=sha256:00525fb6fcdfb9849c83b9ac01ce69bfc10f0d13ccb261069376c25e0c9f05ef

Observation edb1e354-523b-40eb-93c1-43e29e926da5 · outbound

This paper cites Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep nonparametric regression on approximate manifolds: Nonasymptotic error bounds with polynomial prefactors.The Annals of Statistics, 51(2):691–716, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.779421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.779421Z digest=sha256:2ccfaf5c30656947c07dd86409b11d536213b96e391d7916bedc3fa58de12ebf

Observation 4a49078c-b132-492e-a8b4-c2e8951e1e92 · outbound

This paper cites Approximation bounds for recurrent neural networks with application to regression.arXiv preprint arXiv:2409.05577, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation bounds for recurrent neural networks with application to regression.arXiv preprint arXiv:2409.05577, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:29.879046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:29.879046Z digest=sha256:d5ea53be2200f1b218dc99a093461c681f43f57533367022d6b6e36564489280

Observation 03d88569-eff9-45f2-9fd9-c2b44b11131b · outbound

This paper cites Are transformers with one layer self-attention using low-rank weight matrices universal approximators? InThe Twelfth International Conference on Learning Representations, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Are transformers with one layer self-attention using low-rank weight matrices universal approximators? InThe Twelfth International Conference on Learning Representations, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.030422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.030422Z digest=sha256:6d8abdd9b838f4690485210495d45eb8c610bcb1c99d4af9bd2871985a120424

Observation 5781e412-70e1-47c4-9cbf-75993ee51e79 · outbound

This paper cites On the optimal memorization capacity of trans- formers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal memorization capacity of trans- formers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.176620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.176620Z digest=sha256:bcb78aa6407ac7b144b3e0fb72c887c2a18c42fd7ac949947fb6efd6c5ddcbd8

Observation b82921ca-47a1-4901-b929-5422a6d5df4a · outbound

This paper cites The lipschitz constant of self-attention.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets The lipschitz constant of self-attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.303270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.303270Z digest=sha256:f5273ca25dc1dc4078716734ac7c008b0058facf2b9be7d78a4a6af6a15acf59

Observation 650bce6b-9287-4c58-9aca-20b450104f33 · outbound

This paper cites Provable memorization capac- ity of transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Provable memorization capac- ity of transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.429646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.429646Z digest=sha256:7f0e3cd8bf89323a98bb851783d20636d2b1b913a777f5ed136e1d408e51358e

Observation 3e335400-d249-410d-a1f5-27555456dd6b · outbound

This paper cites Transformers are minimax optimal non- parametric in-context learners.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Transformers are minimax optimal non- parametric in-context learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.607066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.607066Z digest=sha256:96a05d714299e679cdb364d8b08ca26fb99f72bf3b8ef81ce7943ba5b155d225

Observation 4f45b3b4-d6e9-4c0e-8bb9-0d60c0de16be · outbound

This paper cites On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the rate of convergence of fully connected deep neural network regression estimates.The Annals of Statistics, 49(4):2231–2249, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.744229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.744229Z digest=sha256:0490fea8426f7351bd722e6465bd7cd3bc3d9df1ca1f50d4315c943f0d2e42c9

Observation b668f606-30f5-4b2e-ae07-52c39ce70c3f · outbound

This paper cites Univer- sal approximation under constraints is possible with transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Univer- sal approximation under constraints is possible with transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:30.872899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:30.872899Z digest=sha256:aa22494fb3bb0cfe3cc3d6c5feffac3e29d617056a662d59df33078ad86417de

Observation 260ffc38-d2a9-45d3-ae22-ae304c2de87f · outbound

This paper cites Approximation and optimization theory for linear continuous-time recurrent neural networks.Journal of Machine Learning Research, 23(42):1–85, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation and optimization theory for linear continuous-time recurrent neural networks.Journal of Machine Learning Research, 23(42):1–85, 2022

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.044785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.044785Z digest=sha256:ff2274d2e907f426eea8c7f0d7f0b9a9c4e72e06aa4418b50240917c01986e3a

Observation bda47ccb-5703-41b4-b655-844792366c2d · outbound

This paper cites Generalization analysis of transformers in distribu- tion regression.Neural Computation, 37(2):260–293, 2025.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Generalization analysis of transformers in distribu- tion regression.Neural Computation, 37(2):260–293, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.160241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.160241Z digest=sha256:ba100b53903a68ca54a08f05954144b9c696a87b9c3847e1cc51696fa2d9b162

Observation 108ea5fe-2b55-4c77-9472-2eaa21ea1ebd · outbound

This paper cites Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep network approxi- mation for smooth functions.SIAM Journal on Mathematical Analysis, 53(5):5465– 5506, 2021

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.326706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.326706Z digest=sha256:f8e8d7fa21c561e3e99e94190ee0a3df06b5a84440ed54a3764f155641225496

Observation 5839742c-0e87-4728-bc4d-4e84d1d4e510 · outbound

This paper cites Upper and lower memory capacity bounds of transformers for next- token prediction.arXiv preprint arXiv:2405.13718, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Upper and lower memory capacity bounds of transformers for next- token prediction.arXiv preprint arXiv:2405.13718, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.429497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.429497Z digest=sha256:3cb4cdfb15ba5ed410db7fc2debddd151b673977fcb1edfba37f6f52a05086df

Observation 251460d0-5f51-4561-9b09-e150acfe6811 · outbound

This paper cites Memorization capacity of multi-head attention in transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Memorization capacity of multi-head attention in transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.589919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.589919Z digest=sha256:fcfac4578aa97e66bd18247b4033527cf91472756ec822bde08d52ff5fa9a102

Observation 5c1cf47c-8ded-4184-bbb6-c2965444c604 · outbound

This paper cites Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Rates of approximation by relu shallow neural networks.Journal of Complexity, 79:101784, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.739676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.739676Z digest=sha256:80f04e1bcb1f1c31f0de06d4e28386d6a438fcf8418d8d373af9cd78291ef338

Observation cf525ea2-cadb-48a4-95e8-8298ca2ab187 · outbound

This paper cites Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Adaptive approximation and generalization of deep neural network with intrinsic dimensionality.Journal of Machine Learning Research, 21(174):1–38, 2020

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:31.921512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:31.921512Z digest=sha256:a60995bc6bcad222a64f63d706a9bd378642ba8f390d7a1866f38a239f98da02

Observation 43b61371-d162-4bca-97b7-97550f2f9e86 · outbound

This paper cites Provable memorization via deep neural networks using sub-linear parameters.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Provable memorization via deep neural networks using sub-linear parameters

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.126644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.126644Z digest=sha256:f4cf6cd6f69851d0788cb37b105e475c04dd9a51d84986a3b712f615662fc58b

Observation 742af754-03ef-4462-9503-40a2cfa02e69 · outbound

This paper cites Equivalence of approximation by convolu- tional neural networks and fully-connected networks.Proceedings of the American Mathematical Society, 148(4):1567–1581, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Equivalence of approximation by convolu- tional neural networks and fully-connected networks.Proceedings of the American Mathematical Society, 148(4):1567–1581, 2020

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.361578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.361578Z digest=sha256:65f0090d6fa6875fa97a6f39a83edae02a2e79d51159e902b654b085a5953ac4

Observation d868eeef-7558-442d-b0dc-6b00fee595bd · outbound

This paper cites Representational strengths and limitations of transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Representational strengths and limitations of transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.551226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.551226Z digest=sha256:6e122311ff9d5db80e73eb67dc28f03e30611ab13d1a01946a852a792b9b01ca

Observation 7eb5b0ad-7a9a-465e-a2b1-a5decf1adfa0 · outbound

This paper cites Nonparametric regression using deep neural net- works with relu activation function.Annals of statistics, 48(4):1875–1897, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Nonparametric regression using deep neural net- works with relu activation function.Annals of statistics, 48(4):1875–1897, 2020

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.650600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.650600Z digest=sha256:527f02ccf02e4213c168b2ca459b7b7ed01e9f146c49896f7eeb1590b0fd5d77

Observation dacd7337-14ac-4d67-9a69-ff3c8d659772 · outbound

This paper cites The kolmogorov–arnold representation theorem revisited.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets The kolmogorov–arnold representation theorem revisited

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.770414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.770414Z digest=sha256:10dc9c0a29a579cd87631f0ebece103b222085df3f0d4e295a7e09aae8a22cc1

Observation de14478b-5b9a-48a7-8353-971145ea79f6 · outbound

This paper cites Understanding in- context learning on structured manifolds: Bridging attention to kernel methods.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Understanding in- context learning on structured manifolds: Bridging attention to kernel methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:32.935975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:32.935975Z digest=sha256:e5cc2925b7bc6653e7c228f7efd23e1e470d277c144b1443ed86e4a2c8f7b868

Observation 1bf983f1-d5f5-433a-a9ec-1c0b5da649fa · outbound

This paper cites Deep network approximation char- acterized by number of neurons.Communications in Computational Physics, 28(5), 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Deep network approximation char- acterized by number of neurons.Communications in Computational Physics, 28(5), 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.054527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.054527Z digest=sha256:cfcb01227d3fa048b695dbbfe4588f1aa2a1386373ea1c27b08f1b8ba4e051d3

Observation d17c58a7-8caf-4b31-9793-6dd6a74c6613 · outbound

This paper cites Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation rate of relu networks in terms of width and depth.Journal de Math´ ematiques Pures et Appliqu´ ees, 157:101–135, 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.200221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.200221Z digest=sha256:b0327ade5eefb68d29fefdf635b776e364518d96373eae9dff4ca3c26366a732

Observation e4e39434-fbd4-46d0-a726-e7376d1f7765 · outbound

This paper cites Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation rates for deep relu neural networks on sobolev and besov spaces.Journal of Machine Learning Research, 24(357):1–52, 2023

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.378677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.378677Z digest=sha256:c5c5b0cc53a39bd3a793d1925d0f51b49a4e2747b602efa051a3599ccac61631

Observation 73080394-9a7f-4dec-bd37-0cb148359b01 · outbound

This paper cites Optimal rates of convergence for nonparametric estimators.The annals of Statistics, pages 1348–1360, 1980.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal rates of convergence for nonparametric estimators.The annals of Statistics, pages 1348–1360, 1980

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.503352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.503352Z digest=sha256:f505ea344148c7f5b40cfe208738950a0cf0a8406e102a5faa5d50163ef7fcba

Observation 0dcf055d-49fd-40e5-87b4-b8361a039fd6 · outbound

This paper cites Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Adaptivity of deep reLU network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.641883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.641883Z digest=sha256:4e1e3aaa49a562d24dd204269756d53267ed440a1e4a9aa0bd26b11574c46e7c

Observation d7ad0ad1-320c-4f5a-96d9-0591b72e1610 · outbound

This paper cites Approximation and estimation ability of trans- formers for sequence-to-sequence functions with infinite dimensional input.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation and estimation ability of trans- formers for sequence-to-sequence functions with infinite dimensional input

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.794727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.794727Z digest=sha256:648b57c48385f1a097e137280dfdb818dec0533a4885236a52d33a4c5b6ad189

Observation ec702a1a-f943-401f-8d3b-7efa7b87a55f · outbound

This paper cites Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Approximation of Permutation Invariant Polynomials by Transformers: Efficient Construction in Column-Size

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:33.955158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:33.955158Z digest=sha256:f230a34e028066a83a155114eac5095ec98c31c2d5ba1aaf940ffb5e5944eee9

Observation 83aa814b-aab6-44a8-9754-a959086b241a · outbound

This paper cites Weak convergence.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Weak convergence

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.094329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.094329Z digest=sha256:4d33fe0764c771029f4f0b5e300d1609c2a84f39a28ca695591b611bde751ca0

Observation 99685651-1abc-433d-86a7-2760e342b9b3 · outbound

This paper cites On the optimal memorization power of reLU neural networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal memorization power of reLU neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.244520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.244520Z digest=sha256:6529584c5116ed18cdefceeda4ca9d54783ddd3e8725e9654db997eb309d2e35

Observation 89b4f471-10f4-4ce3-9e09-92b1fd3deeec · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.372655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.372655Z digest=sha256:f203a91e98853c61cbe90ae5189b58237fccc0d018e8d3050bc19db7be76ee29

Observation 56d032e3-a846-4c70-a37c-2f027e3cdcb7 · outbound

This paper cites Prompt tuning transformers for data memorization.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Prompt tuning transformers for data memorization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.546484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.546484Z digest=sha256:f8de720af3cfb96267d74013d27ab5ac801c0311d81a4d73f248b60df0da6cc3

Observation c22ce1a9-de37-476e-9d5d-3a8213da738f · outbound

This paper cites an unresolved cited work.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.685239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.685239Z digest=sha256:c15cb932fb19a0311dd1683a0c6bb1b88ec816307b91dabb30b6aea163b82492

Observation f2100755-2050-4fab-b5ba-e3c77e50f39b · outbound

This paper cites Statistically meaningful approximation: a case study on approximating turing machines with transformers.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Statistically meaningful approximation: a case study on approximating turing machines with transformers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.856079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.856079Z digest=sha256:3c182dc647c203fdfbe95d564811bd0de321acb04339a664dc29184462f97ee4

Observation 986e253b-7c69-4463-8e21-357e00d1a999 · outbound

This paper cites On the optimal approximation of Sobolev and Besov functions using deep ReLU neural networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets On the optimal approximation of Sobolev and Besov functions using deep ReLU neural networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:34.968470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:34.968470Z digest=sha256:40ad28ee66d15909095ececaf74ebeabbfe92149aa78689aed8222aea3255985

Observation 14739f9b-4530-47ab-b1cb-451be7a90070 · outbound

This paper cites Nonparametric regression using over- parameterized shallow relu neural networks.Journal of Machine Learning Research, 25(165):1–35, 2024.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Nonparametric regression using over- parameterized shallow relu neural networks.Journal of Machine Learning Research, 25(165):1–35, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.048382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.048382Z digest=sha256:3618c4c75c7193b60f8dc1f090d81b3cfcbe18f0a23830132a06540ca02e4512

Observation f319b704-d617-4bd8-85bb-28ec3cf8b13d · outbound

This paper cites Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Error bounds for approximations with deep relu networks.Neural networks, 94:103–114, 2017

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.137363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.137363Z digest=sha256:ba39f5c134b4b5176ee8ccea0f0826f27ad92eedbda979be1d01bd8c054e6adc

Observation f64b5c00-4260-436c-9af7-3131b2b60151 · outbound

This paper cites Optimal approximation of continuous functions by very deep relu networks.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Optimal approximation of continuous functions by very deep relu networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.306182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.306182Z digest=sha256:2ce9f18ac3b1991fbcf8c178104547582bc85e55bffe9ad8b68e51993a5aebf1

Observation e94b9401-45af-4c6b-8d36-f462d85f6de9 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence func- tions? InInternational Conference on Learning Representations, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Are transformers universal approximators of sequence-to-sequence func- tions? InInternational Conference on Learning Representations, 2020

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.487880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.487880Z digest=sha256:5d499e7df68dc9e7f917c09e43b0b4fcb354c4d3f3ca5909be566a07599b3cbb

Observation ce2b942a-2dfd-40b5-8197-b5f2af50374d · outbound

This paper cites O (n) connections are expressive enough: Universal approximability of sparse transformers.Advances in Neural Information Processing Systems, 33:13783–13794, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets O (n) connections are expressive enough: Universal approximability of sparse transformers.Advances in Neural Information Processing Systems, 33:13783–13794, 2020

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.601184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.601184Z digest=sha256:230d32ad20c0901f2bf0b38714b1c9c30b0274906f796606c48c0a48e857e5ac

Observation 4ac02717-490a-42a6-bc5b-d1643ab170dc · outbound

This paper cites Theory of deep convolutional neural networks: Downsampling.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Theory of deep convolutional neural networks: Downsampling

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.721401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.721401Z digest=sha256:b640acbeda2fa9b6a739218c76e0ba68e9ffd46bd1a9aac44d3c9306a017f710

Observation 89d356a7-abf0-415b-bbb9-5f6fefffdbca · outbound

This paper cites Universality of deep convolutional neural networks.Applied and computational harmonic analysis, 48(2):787–794, 2020.

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets Universality of deep convolutional neural networks.Applied and computational harmonic analysis, 48(2):787–794, 2020

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T21:28:35.859983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:28:35.859983Z digest=sha256:94270a899eb056b405fe31ed34f802ad4990b0702e8ddeca85b2dbba1e7b8bd1

Pith citing papers

Observation 68c41094-657e-4173-97e4-2c0684c89151 · inbound

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model cites this paper.

Generalization Bounds for Transformer-Based Next-Token Prediction in a Language Model Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,\lambda}$ Targets

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-29T02:24:15.834827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T05:08:00.711576Z digest=sha256:6a7bc79e1946fb4fc7a94b16e00161a13889a4686d53107e9ca97cbf8383aa7c