Pith. sign in

Paper Citation Record · LEDGER

Are Transformers universal approximators of sequence-to-sequence functions?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:1912.10077.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1912.10077 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:13:25.754027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.309933Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 70a371b3-fb9d-458f-818c-9b2bd6025f81 · inbound

Transformers versus the EM Algorithm in Multi-class Clustering cites this paper.

Transformers versus the EM Algorithm in Multi-class Clustering Are Transformers universal approximators of sequence-to-sequence functions?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:25.754027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:13:25.754027Z digest=sha256:0023c186af6638c10b2d0cae77bac2a30b838f070b707d7201fc863d59b44ceb

Observation db5b7be2-258f-440b-a1a8-592ca5d26ecb · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.396306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.396306Z digest=sha256:0a7ae7492d6c3ee092b08b3b882c01ebf63df6a03e6bab8fe5322bb9429b6b2d

Observation deadf153-e3be-404f-ad12-0701ed62cc0b · inbound

Rethinking Causal Mask Attention for Vision-Language Inference cites this paper.

Rethinking Causal Mask Attention for Vision-Language Inference Are Transformers universal approximators of sequence-to-sequence functions?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:33.425489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:33.425489Z digest=sha256:b6d7f68373c0a76e60344d91ff5677290b42569f1943d87115b288f53cc6f2e7

Observation 908a9664-6277-4c40-80e4-df1aa557d70f · inbound

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor cites this paper.

GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor Are Transformers universal approximators of sequence-to-sequence functions?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:07.766917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:07.766917Z digest=sha256:34d341665dddc7ce8f0ed37bc51db30ad97b455726f7c4a168b25e37837b3bd4

Observation f8a408ed-f361-4636-a630-e614ce3f8825 · inbound

Transformers Are Universally Consistent cites this paper.

Transformers Are Universally Consistent Are Transformers universal approximators of sequence-to-sequence functions?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:06.521514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:06.521514Z digest=sha256:b357745cc68878022e4c285203804413b4ce3461171e108dc96a1b8f90833490

Observation a2084ee9-b7bd-45a0-8515-4d9860c7f7c7 · inbound

Existing Large Language Model Unlearning Evaluations Are Inconclusive cites this paper.

Existing Large Language Model Unlearning Evaluations Are Inconclusive Are Transformers universal approximators of sequence-to-sequence functions?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.351628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.351628Z digest=sha256:544b9b9ecab649802324b1d4725c7daacfcdcfec1d23add454335de84fc5558b

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:1db04c340fd22b71180011b9fae0b866105701d917c0cc0170bc43506de6dac9

Observation 84a07a37-1edc-468f-b252-8735743043ca · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:31.995796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:31.995796Z digest=sha256:13cc70bcc38db7527800d28c6340350418cbc81879cae30f00ef6dbaaa3db425

Observation 036772fa-2944-4984-b8d5-206c295f6300 · inbound

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques cites this paper.

Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:38:35.295805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:38:35.295805Z digest=sha256:8795a4f2529f74497321ef35748ab5375134ded919236f222e3c785c3cc51f3e

Observation e24742c2-24f2-4c94-9fb9-1d5775b6f79a · inbound

Time Resolution Independent Operator Learning cites this paper.

Time Resolution Independent Operator Learning Are Transformers universal approximators of sequence-to-sequence functions?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:34:09.205665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:34:09.205665Z digest=sha256:077e0483e9440d85e8cea5664519c5cb9f26db59ed78ea03366bf0b6e6e4c5d6

Observation 36941140-0c54-4d2b-a716-20ac10fdcd4c · inbound

On the Mathematical Impossibility of Safe Universal Approximators cites this paper.

On the Mathematical Impossibility of Safe Universal Approximators Are Transformers universal approximators of sequence-to-sequence functions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:41:55.587713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:41:55.587713Z digest=sha256:3a8f0e57816d1acc43b04d568161777d3599cf35b475346c8e488f70e9e26fb6

Observation 53fa0876-b32a-4d23-8dda-130b4925efb4 · inbound

Decoding Consumer Preferences Using Attention-Based Language Models cites this paper.

Decoding Consumer Preferences Using Attention-Based Language Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:51:53.078612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:51:53.078612Z digest=sha256:5b2d3cb3ef7b9d07e546c63dbf582677db75abc4d63d1329bbe8dfcd035a0ed5

Observation 50e79fa4-7a8d-4670-8b8a-57c90e4b4820 · inbound

Context-aware Rotary Position Embedding cites this paper.

Context-aware Rotary Position Embedding Are Transformers universal approximators of sequence-to-sequence functions?

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:09:08.432049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:09:08.432049Z digest=sha256:c7040c4f750281e8e049040ae2c17501f786e660dabed6fd0a1a8f209d4a9d7c

Observation 5f78fd59-44ff-486a-a621-39ee8527af08 · inbound

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling cites this paper.

Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling Are Transformers universal approximators of sequence-to-sequence functions?

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:51:50.868831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:49:26.966293Z digest=sha256:86d29f06f9f5d99de74f20828f97edda39aec924301e191148e91e069283e99d

Observation fcd3909e-358c-4cc6-a20f-584c92128715 · inbound

Unraveling Syntax: Language Modeling and the Substructure of Grammars cites this paper.

Unraveling Syntax: Language Modeling and the Substructure of Grammars Are Transformers universal approximators of sequence-to-sequence functions?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:47:21.283300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:47:21.283300Z digest=sha256:f924a90cf51b75adfc21fca0f743d9d7fc1b52aa99dcf138a12862298d9a2c41

Observation 0f10c16c-47ba-4cd8-a5e9-c11b220203b2 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Are Transformers universal approximators of sequence-to-sequence functions?

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:40:50.496470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:367144e4f2c6f8f4cb1260ba54a3475477b49ce054b1253e24e7c2df2b71dedf

Observation d6286ac3-109e-4068-89e0-8b0de8876a56 · inbound

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees cites this paper.

Universal Approximation Theorems for Dynamical Systems with Infinite-Time Horizon Guarantees Are Transformers universal approximators of sequence-to-sequence functions?

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-03T03:21:49.824957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:21:49.824957Z digest=sha256:78ef83ffb1af6e5f6f16af8f52b6efbb10d84e83767d6ec9e9fc5bc9506ed4e9

Observation f01a6dbb-b2cd-48d4-9662-d8b2c6245f2a · inbound

Gating Enables Curvature: A Geometric Expressivity Gap in Attention cites this paper.

Gating Enables Curvature: A Geometric Expressivity Gap in Attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:20:11.376181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:10:47.526833Z digest=sha256:32666124bf374a1690e77990fb02c3ba7bad982459057d88a5e8682cc7c45572

Observation 2ecaae8f-ef3e-4e72-b0e9-bd0d77e79bfc · inbound

Continuous transformations of probability measures and their transport representations cites this paper.

Continuous transformations of probability measures and their transport representations Are Transformers universal approximators of sequence-to-sequence functions?

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.232493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T06:56:24.472027Z digest=sha256:28ba8508806ee9567c2f9a494ded88db3cd2299cf773542ac71cfea0542d2ef7

Observation 79ee8e96-593f-4e46-a052-e5a04f30d8ae · inbound

Progressive Approximation in Deep Residual Networks: Theory and Validation cites this paper.

Progressive Approximation in Deep Residual Networks: Theory and Validation Are Transformers universal approximators of sequence-to-sequence functions?

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.148117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:35:14.636981Z digest=sha256:beb194862529d70913fcf66a27c724705778ff89bae384e615669a0a666b07db

Observation ec91d9a6-9208-4372-b079-1b6ddd27492c · inbound

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models cites this paper.

How Does Attention Help? Insights from Random Matrices on Signal Recovery from Sequence Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:05:58.700918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:52:08.399401Z digest=sha256:99cfb2e1e7f71596b02f62a83367861c21ae4baffefd109ce65397deaf863c9e

Observation c2348a66-465a-4e22-8a00-5423acb01eb5 · inbound

A generative pre-trained transformer with Kerr-soliton attention cites this paper.

A generative pre-trained transformer with Kerr-soliton attention Are Transformers universal approximators of sequence-to-sequence functions?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:04:46.486774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:57:08.127650Z digest=sha256:ecf9e0cfdb8c8b8410acdf3dfaab21f3beac48e44ddfbc2b492634284e38d0fd

Observation 31e7d9a7-ca2e-478e-a3d4-09830604fd58 · inbound

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail cites this paper.

The Variance Brain Foundation Models Forgot: Third-Order Statistics Predict Cognition Where Billion-Parameter Models Fail Are Transformers universal approximators of sequence-to-sequence functions?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.913911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T19:07:34.240009Z digest=sha256:fd89da5c681e30007b42a455f7f746f5100e0daacded14cefa5cf312b655ee45

Observation d4e3d0f4-2e32-4f57-a427-31644db06327 · inbound

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair cites this paper.

Imbuing Large Language Models with Bidirectional Logic for Robust Chain Repair Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.999772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T05:46:17.481698Z digest=sha256:b5ba3820903081bee492905cea7ac9b3c3adc214f2c8cba28569ef570a183714

Observation f16d0fac-d2c0-41da-9f12-fc68533b825a · inbound

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime cites this paper.

Convergence of Gradient Descent for General Neural Network Architectures Beyond the NTK Regime Are Transformers universal approximators of sequence-to-sequence functions?

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.311684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T08:53:46.285233Z digest=sha256:af9302fc182636756d1f59e81035133837af4d0fbea183dddd18928b2e94fe95

Observation 21bb62ce-7cc0-4e08-8803-e89fb4644a7a · inbound

Pre-Strings Lectures on Artificial Intelligence cites this paper.

Pre-Strings Lectures on Artificial Intelligence Are Transformers universal approximators of sequence-to-sequence functions?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:03.658427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:03.658427Z digest=sha256:559faf67638b514b69d091f681b622bbe1b64fdd7e44d326945ea9529cbea255

Observation f95da0d6-cff4-4aa6-8be6-00989bc4ecdb · inbound

On Transformer Dynamics cites this paper.

On Transformer Dynamics Are Transformers universal approximators of sequence-to-sequence functions?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:49:28.691589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:49:28.691589Z digest=sha256:1dd147af462d0bf1a0f5ab6c291afea8ed125b05bfe807092586e4510a3d7e9f

Observation 847e8761-c387-4425-be05-f1e3f3272ba1 · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Are Transformers universal approximators of sequence-to-sequence functions?

Reference 213

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:08.233130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:08.233130Z digest=sha256:4f76ec36538c6e8aa3fb2c7727c30a6d1d55ae0ec5ee04f603ca4ef0edf74eeb

Observation 4257aa79-2604-4872-89a4-6e4b6ef5cb9a · inbound

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems cites this paper.

Identifiability-Aware Source Apportionment in City-Scale Advection-Diffusion Systems Are Transformers universal approximators of sequence-to-sequence functions?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-04T01:37:40.900908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:37:40.900908Z digest=sha256:d7d6e562b5a2faccbed752037742dede3aa772dc190c8fd876b955315da97ca5