Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1910.05895.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:32.086306Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-24T12:34:28.397146Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation b62bf687-faa3-4741-ba03-c1d63af740f4 · inbound
Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model Transformers without Tears: Improving the Normalization of Self-Attention
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 22ff5b8e-f7cb-4e49-9e28-ca5b745135f7 · inbound
GPT-NeoX-20B: An Open-Source Autoregressive Language Model Transformers without Tears: Improving the Normalization of Self-Attention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0bc94ef-b68a-4d93-8816-7bd065a5c783 · inbound
A Comprehensive Overview of Large Language Models Transformers without Tears: Improving the Normalization of Self-Attention
Reference 298
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2e63405-5085-4b8d-b3c8-131abe7a4d76 · inbound
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization Transformers without Tears: Improving the Normalization of Self-Attention
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f176b608-2526-4295-95c2-1538cb7575ee · inbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Transformers without Tears: Improving the Normalization of Self-Attention
Reference 1986
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b805f2-67bb-42a1-9bab-a59829c5c5f5 · inbound
Efficient and Effective Query Context-Aware Learning-to-Rank Model for Sequential Recommendation Transformers without Tears: Improving the Normalization of Self-Attention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b698c83-5376-4f15-832e-f29ac974cc19 · inbound
UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Transformers without Tears: Improving the Normalization of Self-Attention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4414fba-f60a-4526-8073-2b743fece639 · inbound
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks Transformers without Tears: Improving the Normalization of Self-Attention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793975f9-0e15-4db5-b78f-b8d48533e0e6 · inbound
Long-Term Embeddings for Balanced Personalization Transformers without Tears: Improving the Normalization of Self-Attention
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.