Pith. sign in

Paper Citation Record · LEDGER

How Transformers Learn Causal Structure with Gradient Descent

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2402.14735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.14735 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:39:55.074860Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2a8c7430-2f72-4c3c-876f-d16d315c907c · inbound

Transformers with Sparse Attention for Granger Causality cites this paper.

Transformers with Sparse Attention for Granger Causality How Transformers Learn Causal Structure with Gradient Descent

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:44:59.240801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:44:59.240801Z digest=sha256:cfb5b47e8e7d49bb1545d0e6e272d8341b52e786a13368c08d6e2a617ddaf08c

Observation 8c183ed2-280b-45ad-925c-ea1580ad48cc · inbound

Rethinking Associative Memory Mechanism in Induction Head cites this paper.

Rethinking Associative Memory Mechanism in Induction Head How Transformers Learn Causal Structure with Gradient Descent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T15:06:30.765076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:06:30.765076Z digest=sha256:565f5813c6b2d76a5e000e32d79d43a764eecee9e45e5179a138b6b7d0a89982

Observation b4b8a0cc-5aff-4b01-91e0-42c87d69b3f1 · inbound

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation cites this paper.

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation How Transformers Learn Causal Structure with Gradient Descent

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:40.730834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:37:40.730834Z digest=sha256:e5e757997a4c5cb0fcc86021694e96c68068b59b62f08dc59878dfe08b9306ae

Observation 64b2ac71-9cd4-42de-b4fe-26e85ddf615e · inbound

Transformers and Their Roles as Time Series Foundation Models cites this paper.

Transformers and Their Roles as Time Series Foundation Models How Transformers Learn Causal Structure with Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:32.372660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:01:32.372660Z digest=sha256:ce4078bc35b50b910ed2514b064cf2f98c6683f5d218ccc1f8cc2d511775b6f6

Observation 096930cc-88e8-41d1-8ee2-b3f090ea8a2c · inbound

Spectral Journey: How Transformers Predict the Shortest Path cites this paper.

Spectral Journey: How Transformers Predict the Shortest Path How Transformers Learn Causal Structure with Gradient Descent

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T23:44:10.834946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:44:10.834946Z digest=sha256:3fb9c30073531790bb8c62182c113e89a148a81e8a315082954fbf78a1dab704

Observation 67f1ecc1-9c23-48f5-b26a-ca3dbb3303aa · inbound

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias cites this paper.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias How Transformers Learn Causal Structure with Gradient Descent

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:55.074860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:55.074860Z digest=sha256:d78a5054de436c7e843ca5483fe70fa3ae50b9d3574b2ac4d2d589429ec359ef

Observation ff75ee10-68d8-4aa0-8153-f544be1c01b3 · inbound

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models cites this paper.

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models How Transformers Learn Causal Structure with Gradient Descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:57.895555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:40:57.895555Z digest=sha256:05286c1b0cfe3289a1036e63fda0f6a20170ceb601aed184c7dbe855b0502ad3

Observation e6ad2f75-2630-4327-84e0-ed74ac02d5f6 · inbound

The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations cites this paper.

The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations How Transformers Learn Causal Structure with Gradient Descent

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:31:21.883499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:31:21.883499Z digest=sha256:aa99d887b1c9164aa857b3fe1ecfcc989fcb31bd4f0e91b3c9b93d7a47ad3d13

Observation 0f83e290-35b7-4ccd-80a0-69c93bc572ee · inbound

Provable Knowledge Acquisition and Extraction in One-Layer Transformers cites this paper.

Provable Knowledge Acquisition and Extraction in One-Layer Transformers How Transformers Learn Causal Structure with Gradient Descent

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:10:44.830980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T23:05:56.687644Z digest=sha256:40cf84c5882bfeb975f915fbe015e0f1a5d5392af7b99ae20315270cf0bb1f92

Observation 7ff6388f-9ea2-47c8-bb47-61d49b1e97cd · inbound

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points cites this paper.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points How Transformers Learn Causal Structure with Gradient Descent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.890264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.890264Z digest=sha256:871b46cfc7c231710abb6a91c3b05eeacd4bd842285e706299a6194627ad1bf7

Observation 3d1807cc-20b7-44dc-b906-a36d47a21dcd · inbound

Fast Wasserstein rates for estimating probability distributions of probabilistic graphical models cites this paper.

Fast Wasserstein rates for estimating probability distributions of probabilistic graphical models How Transformers Learn Causal Structure with Gradient Descent

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:51:33.487252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T12:46:40.176991Z digest=sha256:f579ff821e71b1945408d4239580316966aeeab79bc6c98758ab3fca5e240ec5

Observation 6936fe17-362f-450a-b57f-52294ab5cbd8 · inbound

A Global Characterization of $f$-Divergences Yielding PSD Mutual-Information Matrices cites this paper.

A Global Characterization of $f$-Divergences Yielding PSD Mutual-Information Matrices How Transformers Learn Causal Structure with Gradient Descent

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T14:27:59.464866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T14:23:44.320981Z digest=sha256:d2f9af9f2355d2827ecefaa942fa49fab01936a4de1d58425c01137e6b7a26a2

Observation 17079616-3aae-4d65-863d-992f1b4ff796 · inbound

Visual prompting reimagined: The power of the Activation Prompts cites this paper.

Visual prompting reimagined: The power of the Activation Prompts How Transformers Learn Causal Structure with Gradient Descent

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:45:54.010348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T18:52:10.770345Z digest=sha256:19321cba4574d4aed7031fb1ec6c81c71fb4c2c0b6b919c7094729dd6fb64453

Observation 3b4e5ebb-8cb4-4630-a8bb-253aef7587e7 · inbound

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification cites this paper.

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification How Transformers Learn Causal Structure with Gradient Descent

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T12:05:21.985681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T12:04:13.987667Z digest=sha256:0f21d54696e762ba5d8fe2310dbaa55b46fc1eac5eea421032b8cb8e822221a9

Observation a30481fc-c454-4a69-a8ad-fd407bac2e32 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective How Transformers Learn Causal Structure with Gradient Descent

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.441942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:71168f08f462f233a04f46d903ca1082257ce02edccdd68c34695564deaec79c

Observation d0e6d8b0-b8aa-481b-aeca-d8473989c797 · inbound

Rigorous uncertainty quantification of probabilistic AI weather forecasts with conformal prediction cites this paper.

Rigorous uncertainty quantification of probabilistic AI weather forecasts with conformal prediction How Transformers Learn Causal Structure with Gradient Descent

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T18:29:41.671020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T18:21:41.472210Z digest=sha256:934ef840f34793d50d45fce641fc294edb486df97e228ebcda1f4c77c2966f1a