Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:28:29.503764Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 8 inbound Pith citation observations for arXiv:2412.13148.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:28:29.503764Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.643471Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T14:29:54.502956Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fad8d5b4-5ab3-480d-b0b1-ae043dcf7a3b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375f9e7d-20f4-4454-a14d-bc26b8c76ce9 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Scalable Second Order Optimization for Deep Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6b617b-414f-4f52-8947-255aaf655679 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Layer Normalization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130dbf43-533e-41eb-bfe6-67052f66da7b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a29298a-891d-4d77-baca-217d986a2364 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Modular Duality in Deep Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d453c04-24ab-4e5c-b210-1cd9640bcbff · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Old Optimizer, New Norm: An Anthology
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 037d451a-1734-4c56-8dcf-c6aecf768210 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training signsgd: Compressed optimisation for non-convex problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 880f976d-a632-4f37-a7c7-e7a7111df26f · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 483f8f95-5dd9-4282-80ec-d0140a041dc4 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad29e9b1-c6f9-46ff-bb4e-a0757af56314 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Preconditioned spectral descent for deep learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 43bac816-83fd-4506-9aa9-ab833454669b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b3d67e7-eba6-4e5a-8237-6fdfd2a0b217 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Robustness to unbounded smoothness of generalized signsgd
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51bb9488-1910-4ced-a052-97dc5f16e618 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Momentum improves normalized sgd
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2d7a6c7f-9f4c-47e9-93b7-5bae4abb1bbc · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A general system of differential equations to model first-order adaptive algorithms
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 297e43fb-e6f9-4af6-a931-464318b052ae · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f62afd6-8290-4170-b690-05f277f6044b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Duchi, Elad Hazan, and Yoram Singer
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3fc39d8d-795a-4fb1-b5e7-0e16756eee74 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kronecker-factored approximate curvature for modern neural network architectures
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6616db8b-1b10-40b1-bb1d-985e6fc574c7 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A trace-restricted kronecker-factored approximation to natural gradient
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1d1dc410-000d-4a2a-a607-84d73850b15d · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Eigenvalue-corrected natural gradient based on a new approximation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 62e675b6-ce37-4d73-ad63-51ec6435c85b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fast approximate natural gradient descent in a kronecker factored eigenbasis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d44018d-acef-4d0c-9fe3-adc9af8dc4df · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7660468d-b3b9-4def-a6d9-b55e2c2d2c3e · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2135bc15-fa3c-4380-beb5-2e1983674a79 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training LoRA: Low-Rank Adaptation of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b346cecc-1dcb-4502-b4df-4f3f53f9ef8a · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Decorrelated batch normalization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 687ce154-7c3c-4382-833c-f198c4017be5 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Iterative normalization: Beyond standardization towards efficient whitening
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9ddb77ce-a62d-475a-a7f8-239610115fa8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4cc3cf-6e69-4e06-9f35-4e65fa3a0560 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training An Isometric Stochastic Optimizer
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4b192fc3-cabc-4b70-913a-0c737c7d0a76 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Three Factors Influencing Minima in SGD
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b2dc94f-b4ed-4b0a-b1bd-e2fadb99c281 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb5ec09-42f5-4035-9007-cfea6e913dcc · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Muon: An optimizer for hidden layers in neural networks, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 48896bd7-6170-4848-b749-1ff378b5e378 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No train no gain: Revisiting efficient training algorithms for transformer-based language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f55ac609-a53b-47bb-b9c0-f03fa3f1cccb · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Exploring Low Rank Training of Deep Neural Networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6525654-fd46-4476-995f-46c04712cb89 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kingma and Jimmy Ba
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ed262b40-e69e-451b-a28b-995efa0b4727 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Efficient Approximations of the Fisher Matrix in Neural Networks using Kronecker Product Singular Value Decomposition
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7cc8bfb-00d7-4254-b6a6-ab19965059c3 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Reducing activation recomputation in large transformer models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8342c03-3d74-48e8-83a4-86663878fd95 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27594f91-5cf2-4a8b-b430-4fbd50d73a4e · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8534924f-03bf-4527-b9ad-9114558c45b8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5a127efb-8843-4917-87c3-3eab6bac0872 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Preconditioned stochastic gradient descent
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dbf2de18-9eaf-434f-aa26-493a00180d7f · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Relora: High-rank training through low-rank updates
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b1c10969-318b-4b0e-9838-ba034372db6e · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 816c38e3-096c-4bc3-8bd0-0af2dfd509f5 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74acc1a3-0091-456c-b28e-b78b8d6d97d8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Decoupled weight decay regularization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90870da8-9c28-444b-a0d4-b680b9f28c53 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Optimizing neural networks with kronecker-factored approximate curvature
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ea789dc-f0a1-47b1-804e-691e8b483c9b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kronecker-factored curvature approximations for recurrent neural networks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 31a91584-c553-4fa0-9209-2bf3f648c475 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Kradagrad: kronecker approximation-domination gradient preconditioned stochastic optimization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d6f21902-5e7a-4295-96b8-7fee0b54d973 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Theory on Adam Instability in Large-Scale Machine Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 373bb1a4-8514-41a3-bc60-c5f167e0e9b8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Introductory lectures on convex optimization: A basic course, volume 87
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce538f7e-81c3-4510-a8d7-9a0d02364d7d · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The AdEMAMix Optimizer: Better, Faster, Older
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f93ca78-6eaf-4d4c-b50e-b54a6db651a8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fishy: Layerwise fisher approximation for higher-order neural network optimization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5721afeb-b1a6-4472-bfef-5dd3671c81f8 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Curvature-Informed SGD via General Purpose Lie-Group Preconditioners
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10beb0ba-f15c-48e7-a456-43d095ef6c0e · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69378cf6-8b94-4ca7-9808-c9d3df570ddb · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Adafactor: Adaptive learning rates with sublinear memory cost
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 50362495-f249-4a7d-bd41-0e1ed9c1767d · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7722d98-497e-4693-8d11-9e5252a4494a · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Fast differentiable matrix square root and inverse square root
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e99ac229-0fe9-4cdf-a4fd-197fe7b6e2be · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training JoMA: Demystifying Multilayer Transformers via JOint Dynamics of MLP and Attention
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f6baea-eee4-487b-9d2a-3207e91d89f3 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5f96d0d0-39bf-4a1f-8fcb-22feb894ac6c · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training LLaMA: Open and Efficient Foundation Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d981dbb4-8874-4227-884a-969946d17287 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33285297-0730-4f02-98af-685372b2a3d2 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Orthogonalising gradients to speed up neural network optimisation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 360f6a6e-38c1-4698-8e1b-8a7ce24975dd · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training SOAP: Improving and Stabilizing Shampoo using Adam
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d810eae4-64ff-4333-8d25-24231d45a8c2 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training 4-bit Shampoo for Memory-Efficient Network Training
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42800392-4b12-4e39-aad4-2377e9ba907e · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1cc1d3-5c7d-4581-a3b4-e827aa1919e5 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Principal whitened gradient for information geometry
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3ce154d-4c22-4367-9712-890425faea5a · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deaf2b4f-0d60-45c4-be8e-0e094577709d · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Root mean square layer normalization
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a1919a-af36-4aa3-a511-2361c4d54eee · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe381d5a-8155-4bd1-8787-fb38cd144f1a · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training OPT: Open Pre-trained Transformer Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b9c2ff-b849-46d1-b73a-0bb669575807 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Why Transformers Need Adam: A Hessian Perspective
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8ef6745-d8a0-4b4d-a5a6-c269d390b9ee · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Adam-mini: Use Fewer Learning Rates To Gain More
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de82ffc-8709-4fda-99e0-56ad1c91c1b2 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e65589d3-b5f4-4b86-9757-59787ff81fe7 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Deconstructing What Makes a Good Optimizer for Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a38ff1a-5543-4da6-9d7f-a057b390944b · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training APOLLO: SGD-like Memory, AdamW-level Performance
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d48760b1-3cb5-477d-a3c9-2a160f539ca6 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ba84744-5121-4f25-9c76-0ff64573fdfb · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training @esa (Ref
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0553f1-7cad-4189-8e31-66f0d1f92661 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40712ffe-5059-4790-bab5-9e0e781f6725 · outbound
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · inbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 860c6b52-a841-4eaf-9efe-f4947bd814be · inbound
Towards Efficient Optimizer Design for LLM via Structured Fisher Approximation with a Low-Rank Extension SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff737671-9801-4638-be91-b3d68b07fc22 · inbound
Memory-Efficient LLM Pretraining via Minimalist Optimizer Design SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b33f2a32-f26b-4cab-bc92-e0b243e5deb4 · inbound
Low-rank Momentum Factorization for Memory Efficient Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7736c809-eda6-4a2d-846a-f5e47244863e · inbound
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c1ad9e-fe70-49b7-82be-fedba31a45a1 · inbound
Demystifying Manifold Constraints in LLM Pre-training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5ef100f-8376-4513-948d-4dc002dbe515 · inbound
Hierarchical Muon: Tiled Newton-Schulz Updates for Efficient Muon Optimization SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 76fda286-b3b7-4c75-8c6d-04316cf8a1cb · inbound
Muse: Representation Geometry of Muon Beyond Normalized Momentum SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.