Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.819526Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2502.06742.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.819526Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.221296Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-13T05:52:21.848754Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b08ad4a5-16fe-4a52-8ad6-c06e6506952c · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Layer Normalization
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f122e6d1-83f1-4c54-b2a7-b5a8d3177886 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative bregman projections for regularized transportation problems
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b0d4c41-37f0-483c-943c-122dde9a1987 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Old Optimizer, New Norm: An Anthology
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a520790-e2ea-4bb8-8431-e9247760aaaa · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training signsgd: Compressed optimisation for non-convex problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c1c0c26-fcb7-4c9e-abc0-a04d482caf08 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8c2f0518-867a-4c4a-9849-f4602c19c6fa · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Distributed optimization and statistical learning via the alternating direction method of multipliers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6cf160a8-e1af-4778-b2c9-83d0377eb545 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Stochastic spectral descent for restricted boltzmann machines
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bbf7d13b-7496-4394-bab7-87a183480732 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training and Pock, T
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4299bd35-014a-4e46-b469-d40de0536ef7 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daea1544-84d4-4f64-b7fd-68adce40a54d · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Symbolic discovery of optimization algorithms
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3125a71-49a0-4a47-af5e-3d31edbf12e4 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training and Mehta, H
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bed42d3e-476b-4ee1-bf47-18a02c229b59 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training On hilbert’s metric for simplices
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7bf6bb02-4816-494e-a22a-7d422156ceb6 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1036d1-bd39-485a-bbb4-42a88518df53 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 807a43a5-376f-44f6-a5ff-ff3e66350ff1 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training and Lorenz, J
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a3d308f-309a-4bda-8ce9-3dee363a3b86 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Eigenvalue-corrected natural gradient based on a new approximation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3a89ebb-4335-405f-ac92-2ae195999eb1 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb83fe5-3150-4ccf-81ac-2e61517b0bdc · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d529124-8416-4357-8a69-f4d6d4058af4 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Beyond convexity: Stochastic quasi-convex optimization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 820b1f11-b874-4e60-a1fb-a7959e808908 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training LoRA: Low-Rank Adaptation of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6927afa-6bcc-4b13-ba94-53b1afe52fe4 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative normalization: Beyond standardization towards efficient whitening
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3245607-4341-4ef3-8480-8707711b816a · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Muon: An optimizer for hidden layers in neural networks, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation add29701-088d-434a-b17a-76bad058594f · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6080fe88-34be-44b1-8865-08a9657870ae · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8d51ea9-6283-4132-a32b-95ef53b7d28e · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a40cd0-c212-4fd7-ab57-f29d81426256 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56561ed3-f3da-4622-bb5c-cde77f1bfe3b · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02956aa1-6bba-4a64-b1e0-378a505245c6 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91f59945-4410-454e-9dd0-6fce281b283b · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Relora: High-rank training through low-rank updates
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226a59d9-2e4a-4ced-a9f0-f4b43badc942 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Decomposition through formalization in a product space
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6aa30c5-ec4d-443b-9791-a2c525b24b47 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 39ffadbf-194d-4c97-8e4e-3b406d8b9716 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Zero: Memory optimizations toward training trillion parameter models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 945ca777-367e-43f9-8627-88b80a777228 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5dd8a654-35d7-4b34-9180-f3bfd7b25fe4 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training A relationship between arbitrary positive matrices and doubly stochastic matrices
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bbad764-6fb7-4ec2-8bdd-239f32ec1b26 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training and Knopp, P
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e29277-8a95-484d-b646-8346a7ffa4bf · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Fast differentiable matrix square root and inverse square root
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1907b5da-4d76-4ef1-87d2-df117dc1f8e6 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5217120b-dfbe-4c6c-bb3c-8a46fd1c51cd · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training LLaMA: Open and Efficient Foundation Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555089af-2e8b-4842-91c8-0301ec813625 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Functional operators: Measures and integrals, volume 1
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0943772a-8700-42f1-93ce-bd84ff49c3bc · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4028448-fec2-4f51-9294-9d575a3c015a · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3dcb6c0-fd3a-4990-ae1b-f914be422b5b · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Training of Convolutional Networks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8841d1e-9bc4-4414-bfe0-5f8a69154ff8 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9cfafa-3b90-4834-9d6b-fdada3dad04f · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training and Sennrich, R
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a10613-9aab-45f7-a17a-6476b27fdf96 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d77efdc-e81c-4111-b116-cb3594c20dbf · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Adam-mini: Use Fewer Learning Rates To Gain More
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14e30325-302c-4d31-b315-0b89b822b86b · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 552f658d-4508-401a-9dea-25d0d0515d53 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6b8083a-4cd7-47bb-ad53-6e19cd7f3df8 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training APOLLO: SGD-like Memory, AdamW-level Performance
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ac2474-aae8-426b-a383-655c9cc14123 · outbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training write newline
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · inbound
Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 903bcba7-1617-42cb-b2b1-21f27dad111d · inbound
Demystifying Manifold Constraints in LLM Pre-training Gradient Multi-Normalization for Stateless and Scalable LLM Training
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e25f7357-df6e-4089-bfd0-46d14d6b6413 · inbound
Optimistic Dual Averaging Unifies Modern Optimizers Gradient Multi-Normalization for Stateless and Scalable LLM Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.