Pith. sign in

Paper Citation Record · LEDGER

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

As of 21 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 2 inbound Pith citation observations for arXiv:2505.00926.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00926 v3

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:55.113761Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:57:10.776717Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:42:49.961290Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cdcfbd51-fa8a-40ac-97a3-b0915cb0119c · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.990906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.990906Z digest=sha256:ecdc3b0c154beac2443a70d7de18a95a121d6f61e464e637c91997ba0d798d3b

Observation 18158418-9dd5-489e-8a7a-012a1f23fc06 · outbound

This paper cites Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.995663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.995663Z digest=sha256:bf0d694d8f9277333eff9534f59bb9a701cc106039f227ea1f95882bdfc0f6e6

Observation 799d9f0c-6b1f-413f-b6b3-c91d44b9317d · outbound

This paper cites Overcoming a Theoretical Limitation of Self-Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Overcoming a Theoretical Limitation of Self-Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.000654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.000654Z digest=sha256:a70e1969b1bca235974aa0e82e6040d23b1830f8d00e76d66aca6f87c6da17a5

Observation a7bf4afd-cb6f-4825-8f2a-7f3396c30c9d · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Optimization and Generalization of Multi-head Attention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.010451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.010451Z digest=sha256:aaa657ac0de1ddc2006bef739441db15cfec656eaa717c73eb0a6da6db96350b

Observation 8f1bf8ba-7431-4872-8a8b-7c2865077a0f · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.014925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.014925Z digest=sha256:e3939e7a5a81c966307cbee9937eb5c56ea395f152dca570d5bbb1476369081d

Observation 3232a6ae-4d9f-43c1-aee1-9f9a9fff8c40 · outbound

This paper cites On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Properties of the Softmax Function with Application in Game Theory and Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.024520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.024520Z digest=sha256:1564811482657ff528702165c4a44ad6527150e3ed3a3a6fff61e9c3b3c1211e

Observation b1de9036-333e-45a3-826a-9c832158e0a3 · outbound

This paper cites In-Context Convergence of Transformers.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias In-Context Convergence of Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.033194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.033194Z digest=sha256:2d8759e4a127b9ee83b67220b809a236e630f10c1109e5d290b70cbcdccbc19d

Observation 11777658-631e-4e64-a941-589384b125ef · outbound

This paper cites A Theoretical Analysis of Self-Supervised Learning for Vision Transformers.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Analysis of Self-Supervised Learning for Vision Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.038985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.038985Z digest=sha256:6fc4f0a47e0794c520c94222e814cc7762e0e32a33a42fd674732132780a11bf

Observation 456a1a31-bddc-4d14-ade5-db6d5590bd62 · outbound

This paper cites Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.048091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.048091Z digest=sha256:8f5024a688079ea8e2054cd644ffde82b7b3c28c9bab67f07a5e96de2b3c459c

Observation 06b3ef20-65e2-4861-94a3-5e285381d92a · outbound

This paper cites A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias A Theoretical Understanding of Shallow Vision Transformers: Learning, Generalization, and Sample Complexity

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.052535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.052535Z digest=sha256:03b734830657a296c2df2661dcb6b5fa882c402308d93660adeb27a3b109d974

Observation a87cda34-ab35-4ad1-91e6-5dfdb3acde01 · outbound

This paper cites Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.057589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.057589Z digest=sha256:f4093b869917bf125bf5145ff72343d7fbc538b01a37c7d9c7504ffcd6f6030c

Observation b139b00c-2f00-4bc6-aeb4-4bcc0712611d · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.061423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.061423Z digest=sha256:114d51fd9e236477c584ee3bf274a7085f1916d4dea4ac3467a3b0594826e9c3

Observation 3849a41a-2628-416c-a463-68a4e5478a53 · outbound

This paper cites The Expressive Power of Transformers with Chain of Thought.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias The Expressive Power of Transformers with Chain of Thought

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.065880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.065880Z digest=sha256:7cb404aee3ea7fff931b3e51238511cdc361721b45b1ee359424392e380e5568

Observation 579ba2f9-5294-4944-a8ef-4175aa5d9664 · outbound

This paper cites an unresolved cited work.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:39:55.504372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:39:55.070888Z digest=sha256:4b271004a875b2b93c92d34791efc36e9a5909ba17fb0051e4f9eaa45bdbea3d

Observation 27895490-f6d5-404a-ac19-e0b2a9814f1a · outbound

This paper cites Benign Overfitting in Token Selection of Attention Mechanism.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Benign Overfitting in Token Selection of Attention Mechanism

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.078557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.078557Z digest=sha256:d233fabdeabeb6b69b73bfca61589bf359ba6da7413d95d0045d2040b4cba778

Observation 439fd820-6ffa-4b94-a19f-4ca19600d8d0 · outbound

This paper cites Implicit Regularization of Gradient Flow on One-Layer Softmax Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Regularization of Gradient Flow on One-Layer Softmax Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.082521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.082521Z digest=sha256:aa95cced11ab069a8083796f243512933325527d632c37763bc1b9c18e5ea628

Observation 7c15a587-fc9f-4354-9f9f-9ef84f2a3456 · outbound

This paper cites Transformers as Support Vector Machines.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Transformers as Support Vector Machines

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.086987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.086987Z digest=sha256:9e93c67a73b0b137b8b89a3fe003f1584deacf4bb56bf24c889803c33f9a465c

Observation 379d2c9c-b268-487a-aa25-0cdb2c8605b7 · outbound

This paper cites JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.092241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.092241Z digest=sha256:a3c39dcef1931f190df38f4ca466e8eafbd52e150b37a0814140fcaa87c89b50

Observation 6a95aef2-5781-4583-a80b-64ac21e2cd4e · outbound

This paper cites Implicit Bias and Fast Convergence Rates for Self-attention.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias and Fast Convergence Rates for Self-attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.096841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.096841Z digest=sha256:7a0477855449c4d8283883ca13598256267adfdd8a2d17b6d850dec5d5d381dd

Observation 471940bf-60e2-48b6-bc36-428bcb39696a · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.101972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.101972Z digest=sha256:f4e9c41f503561efcd0190b1d286f266681a745feaa66f2f4d1946fd21c79501

Observation 2aebcc4d-0092-45f9-8b07-40dd5df9ed25 · outbound

This paper cites Trained Transformers Learn Linear Models In-Context.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Trained Transformers Learn Linear Models In-Context

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.105878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.105878Z digest=sha256:3c62cdac0ba37813d464124681617887bf2e76be5ab959fd80780f26bf5796c3

Observation b0c9f10c-2b93-4a6e-91aa-7cbb25e75fc3 · outbound

This paper cites Auxiliary Lemmas and Equations Lemma A.1 (Gao & Pavel (2017)).

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Auxiliary Lemmas and Equations Lemma A.1 (Gao & Pavel (2017))

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:39:55.490887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:39:55.109869Z digest=sha256:d1fe907f50190103b34c825cee35929f95834dce255029ee4f3e6ffa2a904d81

Observation 704aca31-5cd0-4508-b99a-4b474790b8d9 · outbound

This paper cites This can be done by noting that⟨u2,Ew 1 −E2 ℓ⟩≥ Ω(η) forℓ̸=ℓ0.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias This can be done by noting that⟨u2,Ew 1 −E2 ℓ⟩≥ Ω(η) forℓ̸=ℓ0

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:39:55.478138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:39:55.113761Z digest=sha256:ce0f86c636b1bfca341af10ef3ff63f9d9a2e020aea72ffe26fad68be0c502a3

Observation 0beb3973-5f06-4848-9285-64e2c0825a93 · outbound

This paper cites Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.019942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.019942Z digest=sha256:13470d99887f474252d6d655c25e705f994fcc9dce51e64dd865d60f7a10c7ed

Observation 67f1ecc1-9c23-48f5-b26a-ca3dbb3303aa · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias How Transformers Learn Causal Structure with Gradient Descent

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.074860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.074860Z digest=sha256:d78a5054de436c7e843ca5483fe70fa3ae50b9d3574b2ac4d2d589429ec359ef

Observation dd26e19c-2269-43b7-b139-cc5fea7b213b · outbound

This paper cites Why are Sensitive Functions Hard for Transformers?.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Why are Sensitive Functions Hard for Transformers?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.028824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.028824Z digest=sha256:3061d44e0c3bfb61e9e472f64be8a974f8b5a17244d8f0654e162ef072bebda9

Observation 267c3f93-77d8-43d0-ba85-389b4d88ce17 · outbound

This paper cites Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:39:55.311900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:39:55.043170Z digest=sha256:f36d62064e435217c03db34bd1e8ea39f9673b63cf1287871966db6e658d8358

Observation fb8d970e-59c1-4cb6-a089-55a7073a6925 · outbound

This paper cites Superiority of Multi-Head Attention in In-Context Linear Regression.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Superiority of Multi-Head Attention in In-Context Linear Regression

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:39:55.407581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:39:55.005255Z digest=sha256:eddeb2e7d1e07162f4bf53461e422152cbdc2aa185ae44817aec2e236100844e

Observation 6d55574b-e675-4f26-ac77-97799f6e2fa2 · outbound

This paper cites Provably learning a multi-head attention layer.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Provably learning a multi-head attention layer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.985634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.985634Z digest=sha256:f55a79ffc7f46cf83a1e3bfd4b1133df614b86098e5b0837ebb3eaf13d261537

Observation 8fbe3cd4-788c-44c5-8cde-21f2c77b66ea · outbound

This paper cites On the Ability and Limitations of Transformers to Recognize Formal Languages.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias On the Ability and Limitations of Transformers to Recognize Formal Languages

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.980591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.980591Z digest=sha256:46b6213da97832a7498fb061f9c0572b9ea6f7b8b370cb11467e708eb98ee45c

Pith citing papers

Observation b54f4b11-435e-45a8-9ac1-9f884743d67c · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:10.776717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:10.776717Z digest=sha256:5e6b26eeb67941710763d26363376d920a53f83f7da8d871ec023238c940d8ae

Observation 16c8f18c-5384-421d-95be-6c837185879a · inbound

Agentic Transformers Provably Learn to Search via Reinforcement Learning cites this paper.

Agentic Transformers Provably Learn to Search via Reinforcement Learning How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:42:49.962665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T23:26:28.158991Z digest=sha256:8cf3277e2858881d72d6d214a61a2a3c1052bd095ce6c275be95ce8901aa3aa4