Pith. sign in

Paper Citation Record · LEDGER

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.06179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06179 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact8
  • verified fuzzy7
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b186e2f3-97c6-49a0-869d-367bed3f96c2 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.045295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.045295Z digest=sha256:ac9117f7ead668c6397848fd9591225b1bad0d996529273ec4cb68eff381b4a4

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:325311e9db0b90b752f9dbd48b7c7ea568b736c3abd11489676b3389d3d55002

Observation b86348a3-4561-4195-b6ea-de5fbd1c5f02 · outbound

This paper cites Block coordinate descent for neural networks provably finds global minima.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Block coordinate descent for neural networks provably finds global minima

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.632599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.308487Z digest=sha256:37d44db5920fe0088bc39ded1599960e311f2fb4638007eacdc51199bf624a72

Observation ed9c1163-000a-4186-bcff-5fc84f90fc4f · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.430307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.430307Z digest=sha256:4ef29b7a10fee79e797cfbb08832be5cd913a8fb288f8dbbc6777f807caac4c4

Observation e07cc56c-df49-44ff-8f9e-95ae7a63a9e7 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.520290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.520290Z digest=sha256:e7c7fbea8078640d47ed95711abd32e4708b37a256349a892c9186516bb71586

Observation 10cf59d9-00e8-4344-8bf5-a1783691c633 · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.604932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.604932Z digest=sha256:55200ae615d46640f4de1cb645322e01d679fd497e875132bf5ae2b9ea526886

Observation 3da02f76-adfb-40e5-a345-516b08524857 · outbound

This paper cites On the Computational Power of Transformers and its Implications in Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Computational Power of Transformers and its Implications in Sequence Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.691808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.691808Z digest=sha256:e106549d9535bbf3ebde9bdcbed05f3ec14516c741c7d7435fffe0df48b97e88

Observation 7e95b74c-6263-4093-bc2a-4b82c380afc1 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:10:25.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.756489Z digest=sha256:6685c69814910363311868ac272097d64aae86fc2c2fd91ad059d11cfdcb7f10

Observation d1e61cd4-1248-4fe4-9c20-fc4aaaf28195 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.825748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.825748Z digest=sha256:ce3fe5256035f40f0e1460482c82b8959f58f442edb39095bfcc89b189ab9ae1

Observation 536d2bdd-a47a-4ba5-935b-198c6aa00d06 · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.886542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.886542Z digest=sha256:d0918fd47f3d159b8696a38640bd32f319dc1516930aedd5f0cd4924b7005f56

Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · outbound

This paper cites Provably learning a multi-head attention layer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.973138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.973138Z digest=sha256:8396dbfcb610b811680f9a71015db38f94243c86d9aef1283e28a88b5af46f10

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:25c0dde1207e9ddb64776638edeec53f394e7a8dfd8325a79cfcdbb7370d791d

Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.138258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.138258Z digest=sha256:da13bde78a49310f88c106a1e95386b8dd16033fa728d9c6135acb7e1a839e3e

Observation 6a550f40-a8a8-4ee9-b9cd-7b0e0e2fd991 · outbound

This paper cites Rethinking Attention with Performers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.231632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.231632Z digest=sha256:cf30b7767f880bd1eb3cd15bbd2592b0729ed59f6c02bd2cf39cfa6c466f54f9

Observation cd45c953-249f-42de-9f57-f256a71e4bd0 · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Optimization and Generalization of Multi-head Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.345237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.345237Z digest=sha256:dd48d0b40da1e10ea6a266b878887dbe0fa2711ef92e4bbef5354fab425ac641

Observation d9661c3a-1b2a-4d0c-a859-bb38637c1eae · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.408915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.408915Z digest=sha256:9c35d093b484ab88e24632c34987b3269a9de7387d3c666f2ceff4aeed22c944

Observation f99701b9-2adf-497f-9573-589492dc5b38 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.511563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.511563Z digest=sha256:26898c94fbcfc2a41c9c08b5b3b49023e40a914d2383b922e5481de060ef6d06

Observation 4d4f7717-7ac1-4bc6-b735-738760779656 · outbound

This paper cites Inductive Biases and Variable Creation in Self-Attention Mechanisms.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Inductive Biases and Variable Creation in Self-Attention Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.582759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.582759Z digest=sha256:70cea04a07d67435cdf84a842e6baba8637a61f384a9fc680e6e2c2a31117aff

Observation 88ce5238-6b0a-4348-94a3-7f4802a095c6 · outbound

This paper cites A mathematical framework for transformer circuits.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A mathematical framework for transformer circuits

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.672307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.672307Z digest=sha256:d1d3802e68d93ac8c27498b62a0cdccf6ccddc445c6aed124b4c2c9fc001cb7c

Observation 31ae237e-f46c-40aa-868c-c80d2b355632 · outbound

This paper cites Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T06:10:22.175983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.767563Z digest=sha256:0ccaf18ac23031cb38b1b2fc2a5f4ba205e919fa4a26b0d50368fea9dc9d5f6c

Observation 22bbf9eb-1c4c-4bc6-92ea-a01c9e52edb3 · outbound

This paper cites M., and Fan, J.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization M., and Fan, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.117893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.829721Z digest=sha256:4b5710835dd1bea41bdb742b92e0b7152b9007bd26eb2f58f6ab668ae096e6de

Observation 6362e547-8525-4924-bdcc-060998abbc7a · outbound

This paper cites On Limitation of Transformer for Learning HMMs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On Limitation of Transformer for Learning HMMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.900610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.900610Z digest=sha256:2fc2098afbccfbfe95e335f290aebaa9a01214cb890b5ed1e649a1d3c320c9ca

Observation 7a8e3379-4fcd-4c62-a39f-9cf93381c4c3 · outbound

This paper cites In-context convergence of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization In-context convergence of transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.856363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.012937Z digest=sha256:1db8892e5df0076d340adf131baad41bca95d246f18854b34ac6b62389d24468

Observation 0f2c682a-00e9-44c6-927e-990e382e8067 · outbound

This paper cites How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.576290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.140528Z digest=sha256:1c2bc2509c694cbb4edc99fa55719192b0d4c71f3d4ee1ae94cd21c7ef694500

Observation a492e44c-f911-481a-84ba-313a963d527c · outbound

This paper cites Vision Transformers provably learn spatial structure.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Vision Transformers provably learn spatial structure

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.476399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.267665Z digest=sha256:b55494334a826aa9f9213bfd8cafd71bc19c9e906e9d8087fddf7f53d826e203

Observation 3bc12942-9cb9-4d32-8cc7-2ca58c640eff · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-07T06:10:21.958846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.493685Z digest=sha256:862c87028184a6081eeb2d1bbbc23fafe0d003e70bee7081bb4ed8781199ed46

Observation daa10853-5dcd-416e-a272-2b2af1fbdbb7 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.579936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.579936Z digest=sha256:ec8b4d0bb8944bac85dd39bd45a2e77e563af2488d28a64823967d6b2543657a

Observation f43d11c9-bdd9-46cc-ab9f-22c1271550ef · outbound

This paper cites PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.702105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.702105Z digest=sha256:dec2ad80e63fc67154d3f6da680d385f244b78bbc1c15619980344a962fa2af9

Observation 1f9a66d0-e2b3-4bc2-8315-fd07ad268ec9 · outbound

This paper cites and Sato, I.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization and Sato, I

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.310452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.824938Z digest=sha256:d685d130eaa16b071be3e1a030ee742b6ac5809f93bf3cccdaa6626601547481

Observation 690e44ed-757d-411e-868e-ae61c42ee29c · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.920713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.920713Z digest=sha256:512c4ee2e77069867a9dfe50fe609673c905811a2be7c3379a0a91e0bba6ed54

Observation 6a9d517c-5ac8-4e1f-a088-3fbba6efc302 · outbound

This paper cites SimA: Simple Softmax-free Attention for Vision Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization SimA: Simple Softmax-free Attention for Vision Transformers

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.317828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.033907Z digest=sha256:ed4f45f3de38d9c0358984afac471be8d055fc823adab2d73249db0c2a986aa6

Observation 4e9af55a-da66-4b5e-b0a8-190a9c3b0978 · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.105823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.105823Z digest=sha256:77835dfdc139a5174a0ac57d80813675f8c7000751cb0dddef48eb63676eeb4f

Observation 10d5a435-75ec-4bac-ab9e-e3aad03ec028 · outbound

This paper cites The Closeness of In-Context Learning and Weight Shifting for Softmax Regression.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization The Closeness of In-Context Learning and Weight Shifting for Softmax Regression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.205176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.205176Z digest=sha256:f8cd212b77865216c31ce4014c2becd207fc2d218c2276abac80878411651eb5

Observation a19f856e-914a-4962-bedf-4c7694a05efa · outbound

This paper cites On the Expressive Power of Self-Attention Matrices.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Expressive Power of Self-Attention Matrices

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.335874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.335874Z digest=sha256:7c5b7d91d619bf40aa2e56517514a0e1b5e318de7b87f9d9891a141001a1ad80

Observation 4de1ab7e-9de2-452b-9150-05f154133b75 · outbound

This paper cites Transformers Learn Shortcuts to Automata.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Learn Shortcuts to Automata

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.422720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.422720Z digest=sha256:a9ff5dd0c1084cd8d2754cdaf2d3fa4bca43eb6fd13550c69004767ee5fae800

Observation 4a00c977-f4d6-4b39-9d34-8f98f198fe77 · outbound

This paper cites Rethinking Transformers in Solving POMDPs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Transformers in Solving POMDPs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.099785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.505098Z digest=sha256:a94dfbd58542097375ff2dd336a13ce3934aa91ea0fb19a4c9286b028963f231

Observation 609770b0-6828-4366-8570-e6ee8d580761 · outbound

This paper cites Your transformer may not be as powerful as you expect.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Your transformer may not be as powerful as you expect

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.114111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.623244Z digest=sha256:a87e1fee32abbdbabfb12375a4c9b5246e04d39c3f68cf4a4bd0feafa5552a86

Observation eccb2b48-c5b2-4ff5-abaa-3473c20fd9ba · outbound

This paper cites Transformers are Expressive, But Are They Expressive Enough for Regression?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Expressive, But Are They Expressive Enough for Regression?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.708151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.708151Z digest=sha256:bba59bac2bae8dad3478152743edb6c9702c34777aabe42c3728d6a6778152c2

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · outbound

This paper cites Theory, Analysis, and Best Practices for Sigmoid Self-Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:a9914fd38a4a43179ded0716840a42129ca7cd2c16112ee213c0bf53a7bca420

Observation 9e29a8a8-5a31-4110-b43b-ed1dbccdd87e · outbound

This paper cites Representational Strengths and Limitations of Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Representational Strengths and Limitations of Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.882192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.882192Z digest=sha256:54ec3514f9d1d1bbed838934df5c3e091bc9e30e82d13326ce4d6c5931d08e5e

Observation a1ee5037-49bf-4eaa-844f-c146f8fa6315 · outbound

This paper cites Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:21.767656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.014640Z digest=sha256:7db3a865db277c26a47c8ca2db18189d42cd71bdfa0cc39ec400b9526779503b

Observation db8e1eb0-9a49-4aa4-8bee-0609d5641170 · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unraveling the gradient descent dynamics of transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:23.887917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.096084Z digest=sha256:a1babe9eff7c43262bf066eb4a965b064039e17a089cbe3ee165908045d8ccdb

Observation 67c703c1-8a6b-4f82-be50-7ba0e106a85f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.168393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.168393Z digest=sha256:7ae3f6edd60c8dffd228893972a323eb1babf28114f108526601709d5b68f419

Observation 2b50d96b-e00c-4ba8-9ca7-4a4b3b64e704 · outbound

This paper cites Transformers as Support Vector Machines.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers as Support Vector Machines

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.263719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.263719Z digest=sha256:3ffaacdbc1b0fc0403650966c40b4fa18a2bacbace9eb0bb492189534dda9f78

Observation ec4bcdc5-53d4-47c7-a369-99fd34080d22 · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.436184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.436184Z digest=sha256:a2906cbb23fb6ec593c344f1d896df0dd17dbc50954e361031fba18ea4be14cc

Observation c2bc851c-fdb9-4894-8e9b-b6dbb3bdb473 · outbound

This paper cites An Introduction to Matrix Concentration Inequalities.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Introduction to Matrix Concentration Inequalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.545869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.545869Z digest=sha256:1f768ded9dd944c8a803e837c5b807cb4a8962027644bd793c60e5812cce510c

Observation 7d115a87-ac16-4df7-9565-0ad748ae8882 · outbound

This paper cites Attention Is All You Need.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Attention Is All You Need

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.677748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.677748Z digest=sha256:87e5595f819350f8787efd70f44cffde1659c8967580eca23af7b3c5ce45516a

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · outbound

This paper cites Transformers learn in-context by gradient descent.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:3b85679a72a728ea2e4d4ca0e403b051dbe511614da174829828328589ddb65e

Observation ce0005a2-a080-41ca-ba0c-86ce5636c97f · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linformer: Self-Attention with Linear Complexity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.871924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.871924Z digest=sha256:b15bb7e192d0f4c07c8c5346c287f6359603f24c8a48f918ba0f641c1834ed6c

Observation 30eb5080-6990-49a8-882f-2efd1e737368 · outbound

This paper cites Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.984325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.984325Z digest=sha256:21f4a1a1ea514ad1e7b5bf0df87cf3245a78151084fa8ddd08592ca1516dfdb9

Observation 4559da0e-1c67-4936-926e-00f186f34955 · outbound

This paper cites Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.092976Z digest=sha256:07883ff9148c62aff40f8fca9ea29f092352af3e65a98c8a675810e189b25d71

Observation 3c471eee-41ae-4717-a298-d15b0b9f958f · outbound

This paper cites Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.715533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.173804Z digest=sha256:da2c930353f710feb11854499c5ab7ed25b47b04c3266675bb308d2dea5535c2

Observation 65e87cbc-8b73-4884-8d06-560981f90fff · outbound

This paper cites Self-Attention Networks Can Process Bounded Hierarchical Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Self-Attention Networks Can Process Bounded Hierarchical Languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.255804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.255804Z digest=sha256:e96e2bda1416aec03668f84e359e1a7829cabbc62fbdbdcc7b6ae6faa44f2cbb

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:ed44eda78bbea3de002aad41cf498044e21e105a73fe10b8084daa57f3e482e7

Observation 811b1712-1d5b-43dd-8bef-2b8dfe8355f3 · outbound

This paper cites Global Convergence of Block Coordinate Descent in Deep Learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Global Convergence of Block Coordinate Descent in Deep Learning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T06:10:22.512246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.459603Z digest=sha256:acf61d94649504d28f255fdc9633d72b1a00fc86e19a4723580a3cffa26486e4

Observation 1374c329-c8d2-4784-9261-e9b11af7be25 · outbound

This paper cites Transformers are Efficient Compilers, Provably.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Efficient Compilers, Provably

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.360182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.549139Z digest=sha256:5fe99dc64ab63169013999569b213a6410dfe26a1bd1dcc68dcab992f2e7ed17

Pith citing papers

No inbound Pith citation observations are available.