Pith. sign in

Paper Citation Record · LEDGER

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization

As of 19 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2506.06179.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06179 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:21.549139Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact8
  • verified fuzzy7
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b186e2f3-97c6-49a0-869d-367bed3f96c2 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.045295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.045295Z digest=sha256:04831d1a58bd2c46febbb578794e1940ceaaed77f8f0d0877146585c54e0025b

Observation 2829d5ae-d210-43c7-a57f-8b0f2fb32faf · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.118380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.118380Z digest=sha256:be2b150bade29d8ec8076a11e9cea2feb014b473a542290f031b6eedbc2c6b34

Observation b86348a3-4561-4195-b6ea-de5fbd1c5f02 · outbound

This paper cites Block coordinate descent for neural networks provably finds global minima.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Block coordinate descent for neural networks provably finds global minima

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.632599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.308487Z digest=sha256:4a5d149d7d8de8d6b130d42b047ea7dbf1a97e57550c0dc46cf598ac9dd6ed1d

Observation ed9c1163-000a-4186-bcff-5fc84f90fc4f · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.430307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.430307Z digest=sha256:3b6287b9310ac902b2695f25ef52d1148b32f871446b12b0ea086fa8e8ba039c

Observation e07cc56c-df49-44ff-8f9e-95ae7a63a9e7 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Neural Machine Translation by Jointly Learning to Align and Translate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.520290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.520290Z digest=sha256:c1c4d6632179a2120ef41059ea7c93529fd9a211e5e5d2b57afc005381d7b597

Observation 10cf59d9-00e8-4344-8bf5-a1783691c633 · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.604932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.604932Z digest=sha256:0801d5b52376b8f88f1a8799ebc5a2a6615ba93147456e058e962643186c3143

Observation 3da02f76-adfb-40e5-a345-516b08524857 · outbound

This paper cites On the Computational Power of Transformers and its Implications in Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Computational Power of Transformers and its Implications in Sequence Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.691808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.691808Z digest=sha256:bf9ce7320386a69b04df20d0f368f54699e5e8a0c44eea4b5fc7373dd683d0bc

Observation 7e95b74c-6263-4093-bc2a-4b82c380afc1 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:10:25.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:16.756489Z digest=sha256:e5f22a457c3192bdab980ec8f331ce2defdff186ce7c43c56d6ed3536437775f

Observation d1e61cd4-1248-4fe4-9c20-fc4aaaf28195 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.825748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.825748Z digest=sha256:7284fb50006cc4527224383a3bef6398cc1c1afea4dfd2d75fbbeecf7a93f684

Observation 536d2bdd-a47a-4ba5-935b-198c6aa00d06 · outbound

This paper cites Decision Transformer: Reinforcement Learning via Sequence Modeling.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Decision Transformer: Reinforcement Learning via Sequence Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.886542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.886542Z digest=sha256:950bbd08b189adc978187a67764e11fe1f4231b1d4ee4624c910673c324e6b6c

Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · outbound

This paper cites Provably learning a multi-head attention layer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.973138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.973138Z digest=sha256:4de84dd70015a9288c3d1d013edd5ddc7ab80c8cd7f4e112685d8d5fe7e59161

Observation 3c02b589-bf51-4c11-a0ad-75d7b76b6c0f · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.068890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.068890Z digest=sha256:417c868aba5dc5c56937552ff34f33234f1ae8e12720573ec44939ec9dff8ddc

Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.138258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.138258Z digest=sha256:6355ad95f8299fb8b902179c4c413e05abe3e2376e704b9458794bede91fe05b

Observation 6a550f40-a8a8-4ee9-b9cd-7b0e0e2fd991 · outbound

This paper cites Rethinking Attention with Performers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Attention with Performers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.231632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.231632Z digest=sha256:11f56e05fa0ffb4a83d0eaa154b6251e5c37fe820bf0bff2fb4213892ebdf37c

Observation cd45c953-249f-42de-9f57-f256a71e4bd0 · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Optimization and Generalization of Multi-head Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.345237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.345237Z digest=sha256:3c1f4c18109a8f2bed4a59bb850f4d10b915ecd882205f6ad4b3bfed8ba2f572

Observation d9661c3a-1b2a-4d0c-a859-bb38637c1eae · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.408915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.408915Z digest=sha256:4c2b448ad62f80c1d2f8d5a1c006793a7aa6001cbdf22c0ad5fa132d15e96a78

Observation f99701b9-2adf-497f-9573-589492dc5b38 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.511563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.511563Z digest=sha256:3eccadba1eaf1bccaf8bf0e15586112358fcf129abd70d180fa0921fe4ffae0e

Observation 4d4f7717-7ac1-4bc6-b735-738760779656 · outbound

This paper cites Inductive Biases and Variable Creation in Self-Attention Mechanisms.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Inductive Biases and Variable Creation in Self-Attention Mechanisms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.582759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.582759Z digest=sha256:6322680026544ede89caacc56a1aa5a5b92663819f2ea5464e7745289fe8181e

Observation 88ce5238-6b0a-4348-94a3-7f4802a095c6 · outbound

This paper cites A mathematical framework for transformer circuits.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A mathematical framework for transformer circuits

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.672307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.672307Z digest=sha256:68e05dd08f16c46db49dec9cc28dfe66ac8aad78cf35f24d0a47ecd99a58b3d0

Observation 31ae237e-f46c-40aa-868c-c80d2b355632 · outbound

This paper cites Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Phenotypes and Genotypes: The Search for Influential Genes, volume 18 of Computational Biology

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T06:10:22.175983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.767563Z digest=sha256:fe36de41840c0a0b199e841d54edcf04c82d42c5342847f97817b7e824abed92

Observation 22bbf9eb-1c4c-4bc6-92ea-a01c9e52edb3 · outbound

This paper cites M., and Fan, J.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization M., and Fan, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:25.117893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:17.829721Z digest=sha256:ba3209b565937069ad85e95f6c76e91801011159e7b3273d7f00abcefe15e66c

Observation 6362e547-8525-4924-bdcc-060998abbc7a · outbound

This paper cites On Limitation of Transformer for Learning HMMs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On Limitation of Transformer for Learning HMMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:17.900610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:17.900610Z digest=sha256:a05183dd31fdd7b86831a25993819a3d3ec7c76f11d26ca3211b4bce05d38f88

Observation 7a8e3379-4fcd-4c62-a39f-9cf93381c4c3 · outbound

This paper cites In-context convergence of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization In-context convergence of transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.856363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.012937Z digest=sha256:5c053835ab15991dc2e2499984ab55dce87f62084e6afea76cc147e397d7e0d9

Observation 0f2c682a-00e9-44c6-927e-990e382e8067 · outbound

This paper cites How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization How Transformers Learn Diverse Attention Correlations in Masked Vision Pretraining

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.576290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.140528Z digest=sha256:4d1c7a7a20d24f597d33752d83591e6f16d346f8e5a61f803ed183a8c37225b7

Observation a492e44c-f911-481a-84ba-313a963d527c · outbound

This paper cites Vision Transformers provably learn spatial structure.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Vision Transformers provably learn spatial structure

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.476399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.267665Z digest=sha256:dee479b1d104fb7d7ea67ff68e66ddb5a0a40ba64b6b01083d4541c17cc97177

Observation 3bc12942-9cb9-4d32-8cc7-2ca58c640eff · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 26

Resolution
verified exact
doi, observed 2026-08-07T06:10:21.958846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.493685Z digest=sha256:acfa8c1d53388dec70d90f130295af0c8f1de4b8abd053c91bb6fbbc01142457

Observation daa10853-5dcd-416e-a272-2b2af1fbdbb7 · outbound

This paper cites an unresolved cited work.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.579936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.579936Z digest=sha256:6350c15e7bc17e4c3a2ee073879b32003c93f20b4b8a6f94703c710e5c416c66

Observation f43d11c9-bdd9-46cc-ab9f-22c1271550ef · outbound

This paper cites PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.702105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.702105Z digest=sha256:11241342b31ce59e3dd0a5ff9f0b24f65519c5303d833647202c48e2dfb3f0ed

Observation 1f9a66d0-e2b3-4bc2-8315-fd07ad268ec9 · outbound

This paper cites and Sato, I.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization and Sato, I

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.310452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:18.824938Z digest=sha256:a31a4297b5845afc2b62d5b037d5da22b239786760c377adcdf23898b50ca4df

Observation 690e44ed-757d-411e-868e-ae61c42ee29c · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:18.920713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:18.920713Z digest=sha256:35243c2a7f59541afeb3bd361782fd73225ab33c1401eac6a0b7301e9688937a

Observation 6a9d517c-5ac8-4e1f-a088-3fbba6efc302 · outbound

This paper cites SimA: Simple Softmax-free Attention for Vision Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization SimA: Simple Softmax-free Attention for Vision Transformers

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.317828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.033907Z digest=sha256:48c8e97bbcc82dcd19d520d39b8a92c157d36adbd68909c648ea353294c90788

Observation 4e9af55a-da66-4b5e-b0a8-190a9c3b0978 · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.105823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.105823Z digest=sha256:c5f5459f05bc7ed03b34ee2381cf63a8a5b136bbb229f99d53a1a98acf06ff3c

Observation 10d5a435-75ec-4bac-ab9e-e3aad03ec028 · outbound

This paper cites The Closeness of In-Context Learning and Weight Shifting for Softmax Regression.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization The Closeness of In-Context Learning and Weight Shifting for Softmax Regression

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.205176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.205176Z digest=sha256:c7d6795aaf471e1d082106a6cd565d05d93804ac429d7c67aa5f683b5aaa4abb

Observation a19f856e-914a-4962-bedf-4c7694a05efa · outbound

This paper cites On the Expressive Power of Self-Attention Matrices.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization On the Expressive Power of Self-Attention Matrices

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.335874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.335874Z digest=sha256:b976e830851a473b134b8e55a268283e7b7794b32389d409a7ea2b5eea256025

Observation 4de1ab7e-9de2-452b-9150-05f154133b75 · outbound

This paper cites Transformers Learn Shortcuts to Automata.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Learn Shortcuts to Automata

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.422720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.422720Z digest=sha256:49c207f06bba5fd964c341632abc1ea688a4a6ffad52dd57142db22cacc8dbf5

Observation 4a00c977-f4d6-4b39-9d34-8f98f198fe77 · outbound

This paper cites Rethinking Transformers in Solving POMDPs.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Rethinking Transformers in Solving POMDPs

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:23.099785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.505098Z digest=sha256:d1abf99458556315b50d97f0e5489dc34680b00241667811c28251199c407cea

Observation 609770b0-6828-4366-8570-e6ee8d580761 · outbound

This paper cites Your transformer may not be as powerful as you expect.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Your transformer may not be as powerful as you expect

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:24.114111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:19.623244Z digest=sha256:d7af1f43009cd7681ee4b988cfdd8dbb693ae21ef3d63a7186f843c499eda7d4

Observation eccb2b48-c5b2-4ff5-abaa-3473c20fd9ba · outbound

This paper cites Transformers are Expressive, But Are They Expressive Enough for Regression?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Expressive, But Are They Expressive Enough for Regression?

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.708151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.708151Z digest=sha256:2867b72261b8581fe5b73d01359988740ef80e8623fc158444d175ed5f56df85

Observation 2737a320-d6ae-4494-b67f-3d9d3bf4cc1d · outbound

This paper cites Theory, Analysis, and Best Practices for Sigmoid Self-Attention.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.778678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.778678Z digest=sha256:79580888c010f07242410c0aad8da843968995e8832cce0210c3adf476855b72

Observation 9e29a8a8-5a31-4110-b43b-ed1dbccdd87e · outbound

This paper cites Representational Strengths and Limitations of Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Representational Strengths and Limitations of Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:19.882192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:19.882192Z digest=sha256:8dc5ad8d638b7d282b8a68e3dbcb076847d6358e3052b73bb388b91083f2f428

Observation a1ee5037-49bf-4eaa-844f-c146f8fa6315 · outbound

This paper cites Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Effects of Depth, Width, and Initialization: A Convergence Analysis of Layer-wise Training for Deep Linear Neural Networks

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:21.767656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.014640Z digest=sha256:ebaff8e86659a8344850025a4828b5af0c725a256f11d510609c2efe89d6996e

Observation db8e1eb0-9a49-4aa4-8bee-0609d5641170 · outbound

This paper cites Unraveling the gradient descent dynamics of transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Unraveling the gradient descent dynamics of transformers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:10:23.887917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:20.096084Z digest=sha256:dd29967de688a8423f15e57f767b26ec52124a3af584ae134c33d37e3e4eec71

Observation 67c703c1-8a6b-4f82-be50-7ba0e106a85f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.168393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.168393Z digest=sha256:4fe87972e3f2fc853cae89250994f94f793788c435ed8ee648d30e11b2941ae8

Observation 2b50d96b-e00c-4ba8-9ca7-4a4b3b64e704 · outbound

This paper cites Transformers as Support Vector Machines.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers as Support Vector Machines

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.263719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.263719Z digest=sha256:d4964979b595214f6d2785a674bbd670e3ff28ef74c23d949747bcaa7a1ddda5

Observation ec4bcdc5-53d4-47c7-a369-99fd34080d22 · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.436184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.436184Z digest=sha256:c0dcb55fc94226a9c352967d6b2edff48459ffc549a4090fe5ab519d73e1f8f2

Observation c2bc851c-fdb9-4894-8e9b-b6dbb3bdb473 · outbound

This paper cites An Introduction to Matrix Concentration Inequalities.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization An Introduction to Matrix Concentration Inequalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.545869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.545869Z digest=sha256:dc34ba0f4f95329f00d6d7e7ee6e56b6ac76a8e18b774a13e5e6f806d2a122cd

Observation 7d115a87-ac16-4df7-9565-0ad748ae8882 · outbound

This paper cites Attention Is All You Need.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Attention Is All You Need

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.677748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.677748Z digest=sha256:5bbf193db57cc07f5e8720def8316e3f0a2e3d0f9c1b7e3b631eb002d5d4c38a

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · outbound

This paper cites Transformers learn in-context by gradient descent.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:26d239cb2a6f860087a9aafeececcf12f61b118dc7728fdefc5c626f87746ed8

Observation ce0005a2-a080-41ca-ba0c-86ce5636c97f · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Linformer: Self-Attention with Linear Complexity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.871924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.871924Z digest=sha256:dd06393c3acabfec4342a83274657da916017e8d2a73f01f630baa2f81686d81

Observation 30eb5080-6990-49a8-882f-2efd1e737368 · outbound

This paper cites Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.984325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.984325Z digest=sha256:953c0c3976bc15d5ab09b3d445e13e97d74629b9ce2798db8e65abf39010bb23

Observation 4559da0e-1c67-4936-926e-00f186f34955 · outbound

This paper cites Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.092976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.092976Z digest=sha256:dbefa6e17a275d27f70b5abbdde1ec9505a06ca95f2ebce268b8839580e29c87

Observation 3c471eee-41ae-4717-a298-d15b0b9f958f · outbound

This paper cites Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Training Dynamics of Transformers to Recognize Word Co-occurrence via Gradient Flow Analysis

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.715533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.173804Z digest=sha256:ea6c2853e6afbad2fd2335e46d255e859f6ab38e5eb46cdc8ea6fec7d8874a95

Observation 65e87cbc-8b73-4884-8d06-560981f90fff · outbound

This paper cites Self-Attention Networks Can Process Bounded Hierarchical Languages.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Self-Attention Networks Can Process Bounded Hierarchical Languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.255804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.255804Z digest=sha256:ea4ffae59f3cbc05b36f2dc680e78ddfbfc1f4690621fab5d49ef57de0e1143b

Observation e11c36a9-ba77-42ee-894a-2d3458285132 · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Are Transformers universal approximators of sequence-to-sequence functions?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:21.358940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:21.358940Z digest=sha256:2217962f683aa22a292e691cb8765700954f4ebf51b7b301acea056363a2495b

Observation 811b1712-1d5b-43dd-8bef-2b8dfe8355f3 · outbound

This paper cites Global Convergence of Block Coordinate Descent in Deep Learning.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Global Convergence of Block Coordinate Descent in Deep Learning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T06:10:22.512246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.459603Z digest=sha256:910fea55787d8a7caa2c177362f6fecc1e039c6e4a64c407d8c8e02289ea5a59

Observation 1374c329-c8d2-4784-9261-e9b11af7be25 · outbound

This paper cites Transformers are Efficient Compilers, Provably.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers are Efficient Compilers, Provably

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:10:22.360182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T06:10:21.549139Z digest=sha256:41f5953ea2dac5ef711c2a0017aed5c7200dc48b18492ea704be824847b03e7f

Pith citing papers

No inbound Pith citation observations are available.