Pith. sign in

Paper Citation Record · LEDGER

Provably learning a multi-head attention layer

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2402.04084.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.04084 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:49:39.588423Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:29.167982Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 51982e89-d015-491b-8279-7a1ff550281b · inbound

Training Dynamics of In-Context Learning in Linear Attention cites this paper.

Training Dynamics of In-Context Learning in Linear Attention Provably learning a multi-head attention layer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T13:42:08.863656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:42:08.863656Z digest=sha256:a67c3a897c7932ae4ba57884039107e27d2edd86545b9de50259988191da92bb

Observation 26894e0e-f0a5-48e1-9329-2652a5012270 · inbound

Attention Mechanism, Max-Affine Partition, and Universal Approximation cites this paper.

Attention Mechanism, Max-Affine Partition, and Universal Approximation Provably learning a multi-head attention layer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T05:49:39.588423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:49:39.588423Z digest=sha256:e392f40d320da1457cc97108a5dd3c73307910c655b5d9fbcf51aae1b934ae28

Observation 6d55574b-e675-4f26-ac77-97799f6e2fa2 · inbound

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias cites this paper.

How Transformers Learn Regular Language Recognition: A Theoretical Study on Training Dynamics and Implicit Bias Provably learning a multi-head attention layer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:39:54.985634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:39:54.985634Z digest=sha256:9fed1d09449afc84a6daed79d717f8a7d8266b273c8ed673dfcb55a57e4e9e32

Observation f5fee61a-8b01-4b22-9bcd-6d619cbfed03 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Provably learning a multi-head attention layer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:35.509713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:35.509713Z digest=sha256:b150e5f895b3687ad1752d34bace92663afa92afd25c5635fb82a8334a6af9ac

Observation 56b1e1c8-31ae-4eb2-a06f-9c493c7e724e · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Provably learning a multi-head attention layer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:16.973138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:16.973138Z digest=sha256:4de84dd70015a9288c3d1d013edd5ddc7ab80c8cd7f4e112685d8d5fe7e59161

Observation e0a40042-207e-41fd-babe-981385d2535f · inbound

Tight Sample Complexity of Transformers cites this paper.

Tight Sample Complexity of Transformers Provably learning a multi-head attention layer

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:57:29.169339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T17:33:48.373948Z digest=sha256:50bb929d8037e4caacf40c402fe14bcce1ffcbdc711a3abbccaf3bbbe7dea8ff