Pith. sign in

Paper Citation Record · LEDGER

Toward Understanding Why Adam Converges Faster Than SGD for Transformers

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2306.00204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.00204 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:18.115103Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:30.098265Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d3d1554-d265-4b76-8d1e-8be221b1f8a0 · inbound

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization cites this paper.

Revisiting the Initial Steps in Adaptive Gradient Descent Optimization Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:55:18.115103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:55:18.115103Z digest=sha256:dd020abe090969c4c9754c060ba2024599bd2a6437e9f3f4a0ae0e71eb3b2e78

Observation 0b515489-1a8d-484b-9237-153d3ac5cfce · inbound

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond cites this paper.

Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T20:09:35.207510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:09:35.207510Z digest=sha256:46159e16e2c24212d438ad593326e65aadc7dccedad99f705f42267361f7cd40

Observation de7bce79-ad55-4a69-8474-051b04a9546c · inbound

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism cites this paper.

Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T23:08:28.222610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:08:28.222610Z digest=sha256:8565af6bb0f126bc1bf4fc20d4f63f4734febc2b2c59babb42b3ac0cae444e1e

Observation 443d15c7-ca51-4408-bf6b-4c2bf2a238e3 · inbound

Artificial Neural Networks for Magnetoencephalography: A review of an emerging field cites this paper.

Artificial Neural Networks for Magnetoencephalography: A review of an emerging field Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 174

Resolution
unresolved
no resolver link, observed 2026-08-10T18:10:55.058541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:10:55.058541Z digest=sha256:11f5f92d977fd279bd79ab20f12f2068280e883c4a3ef829625c7d49ace80a39

Observation 35b501fc-e92c-452e-9964-c7f63f063938 · inbound

Avoiding spurious sharpness minimization broadens applicability of SAM cites this paper.

Avoiding spurious sharpness minimization broadens applicability of SAM Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T12:21:20.501096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:21:20.501096Z digest=sha256:5b100a9e92e17878c10cd1b1e4709d7cdb0837ae57b3a9c40a22dd9dbe64c4d3

Observation 86065612-bebb-4c33-b7ee-6d21d2f4474e · inbound

Mechanistic Insights into Grokking from the Embedding Layer cites this paper.

Mechanistic Insights into Grokking from the Embedding Layer Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:32.846977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:32.846977Z digest=sha256:c2f7aeedfa79457762d2e9b4769858629cea872174b166527e62524991d3e027

Observation c90daf0d-ba2c-4c9c-8d1c-cfba25a472cb · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.615997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.615997Z digest=sha256:03539155ecfe910968752cb45b0d31cd3d736172eca47017f39f90f530efb1ed

Observation 94da035a-eaff-4fde-bff5-c5e39899b738 · inbound

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails cites this paper.

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.446775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T16:34:31.388219Z digest=sha256:e5b21c0df76bc5eb6d8c7832567e8ec4c9e726770ff02804575ca162c11b4ea0

Observation a94a8ef7-bd88-443b-ab9a-14d63ed5e166 · inbound

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails cites this paper.

Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:50:03.763820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T11:47:27.345223Z digest=sha256:6bec6053f04c63013c56b57a156b3e3480cc544bcce1e2852ba8d98923b5da6d

Observation 23181093-ef2c-42ff-804d-ff41c271be07 · inbound

Spatial navigation in preclinical Alzheimer's disease: A review cites this paper.

Spatial navigation in preclinical Alzheimer's disease: A review Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-13T19:53:53.219607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:53:53.219607Z digest=sha256:280a73c531dc82927baea72d26841cc9683d5016e47156dafa026b0825928b9a

Observation f02d5820-7681-4c60-b805-390ab01bc326 · inbound

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent cites this paper.

Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:09.672347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T12:22:27.265973Z digest=sha256:100d0bec2ee68336e9ecc7fc82650dc19a2fd6ab2fb1cdec52387ec2dd5dc3bc

Observation fbdbb427-e811-45ef-b2ed-c306509e69e1 · inbound

Convergence of difference inclusions via a diameter criterion cites this paper.

Convergence of difference inclusions via a diameter criterion Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:23:31.750140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T02:21:30.228735Z digest=sha256:782eeaa36eabc8d97cdeacfbd1406fd4bcb74a8542e531a0f9692f18024dceb1

Observation 0ca5b382-defb-4f32-aac0-d706b0f72e82 · inbound

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates cites this paper.

Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:38:19.452751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T13:34:09.129379Z digest=sha256:eafaf7feb50a7b28fe2fe704262390b2b7dbdc91923da31707e3c70096156a95

Observation 29350e22-a680-4dad-b17d-d3242da6f806 · inbound

Looped Transformers with Layer Normalization Provably Learn the Power Method cites this paper.

Looped Transformers with Layer Normalization Provably Learn the Power Method Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:35.276216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T19:03:51.212121Z digest=sha256:b9fc09c23d99c7cc9cf25a5912748363a1acc0dedf7f7559d8e51c7b2947486d

Observation 8f1ec825-57ee-4075-b27d-5c225b58b0a2 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:16:44.460115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:f4a9ca48fba24c343672b092c1a26781312913e74e7f63a421b945c135a538eb

Observation c0eb6d77-ba4d-4e16-be4b-ff2dc4a299bb · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Toward Understanding Why Adam Converges Faster Than SGD for Transformers

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.099527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:f4a6cbbcfb4014471f60d0dcc38c01c487930c0e635bceeb86c7bb9375e90d0b