Pith. sign in

Paper Citation Record · LEDGER

Transformers without Tears: Improving the Normalization of Self-Attention

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1910.05895.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.05895 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:32.086306Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T12:34:28.397146Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b62bf687-faa3-4741-ba03-c1d63af740f4 · inbound

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model cites this paper.

Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Transformers without Tears: Improving the Normalization of Self-Attention

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:14:26.535552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T12:10:49.690618Z digest=sha256:8d1403afc66439f97c5ea435f072d2939aacbe755acb8f53ae0fc07ecbe6c289

Observation 22ff5b8e-f7cb-4e49-9e28-ca5b745135f7 · inbound

GPT-NeoX-20B: An Open-Source Autoregressive Language Model cites this paper.

GPT-NeoX-20B: An Open-Source Autoregressive Language Model Transformers without Tears: Improving the Normalization of Self-Attention

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-24T12:34:28.399984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T12:33:37.701655Z digest=sha256:ffe124c44aff9d8c71a3f9ab91188cd8e7c2bcaa2cd1b448a1c616bd7c736edd

Observation e0bc94ef-b68a-4d93-8816-7bd065a5c783 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Transformers without Tears: Improving the Normalization of Self-Attention

Reference 298

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.297137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:48290eccb2cc17ed87c037dc1c6fa8545561bfdfbce1996ad5a8b217cba2e955

Observation c2e63405-5085-4b8d-b3c8-131abe7a4d76 · inbound

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization cites this paper.

Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Transformers without Tears: Improving the Normalization of Self-Attention

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:32.086306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:32.086306Z digest=sha256:ceb180ca5c671b6dd0f35c5fc1575c774c4284d1cf568864869e23fee72ace3d

Observation f176b608-2526-4295-95c2-1538cb7575ee · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention

Reference 1986

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.609473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.609473Z digest=sha256:961deaac2e7da1969981429719fd5039d7cc13b0e8e3e0fd8cc475f9ff91e7e6

Observation c1b805f2-67bb-42a1-9bab-a59829c5c5f5 · inbound

Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation cites this paper.

Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation Transformers without Tears: Improving the Normalization of Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:07:00.041601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:07:00.041601Z digest=sha256:c5b2696a0a231cc72259845c7b3039f7990e26a2dba7119ed040cc0a09bb9b16

Observation 0b698c83-5376-4f15-832e-f29ac974cc19 · inbound

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning cites this paper.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Transformers without Tears: Improving the Normalization of Self-Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.074129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.074129Z digest=sha256:4f04e4a448444b3a69e6ec50c191c89260996997e9ac026733e7a98d35246e59

Observation d4414fba-f60a-4526-8073-2b743fece639 · inbound

Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks cites this paper.

Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks Transformers without Tears: Improving the Normalization of Self-Attention

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T19:20:50.271527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T19:20:50.271527Z digest=sha256:c982e2796aabb34e9048af674c21351e41d31cf38a7ea512279e8f6a3a7244ba

Observation 793975f9-0e15-4db5-b78f-b8d48533e0e6 · inbound

Long-Term Embeddings for Balanced Personalization cites this paper.

Long-Term Embeddings for Balanced Personalization Transformers without Tears: Improving the Normalization of Self-Attention

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.380626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:39:41.045848Z digest=sha256:256510ce5b9022c46d0920af885045ad6cd9fc28d9e6ea5ca56200a7782e44fe