Pith. sign in

Paper Citation Record · LEDGER

Gradient Multi-Normalization for Stateless and Scalable LLM Training

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2502.06742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06742 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.819526Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.221296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:52:21.848754Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b08ad4a5-16fe-4a52-8ad6-c06e6506952c · outbound

This paper cites Layer Normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.413471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.413471Z digest=sha256:37e0df12d5e452df72d4747ee1b62eeb4ddc5b29bfc192a0f8c46c35f9adc844

Observation f122e6d1-83f1-4c54-b2a7-b5a8d3177886 · outbound

This paper cites Iterative bregman projections for regularized transportation problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative bregman projections for regularized transportation problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.906346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.451252Z digest=sha256:7f246cb31e488cb31987cb385fa330148b91954c326c2413a7e74fd69af57541

Observation 4b0d4c41-37f0-483c-943c-122dde9a1987 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Old Optimizer, New Norm: An Anthology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.455661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.455661Z digest=sha256:3b5034a0146a7e5a68de2d9a94f47f70953fdcccd2d2afad5694eb0e350e67a1

Observation 4a520790-e2ea-4bb8-8431-e9247760aaaa · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training signsgd: Compressed optimisation for non-convex problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.459883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.459883Z digest=sha256:0e15ce41745ab189cea02528985f33e93b976874a2fcd457c8a478f71fe4a927

Observation 2c1c0c26-fcb7-4c9e-abc0-a04d482caf08 · outbound

This paper cites Proximal alternating linearized minimization for nonconvex and nonsmooth problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Proximal alternating linearized minimization for nonconvex and nonsmooth problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.888347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.463975Z digest=sha256:beac66db620f1014b34f262feda963857da8cfada73eca20a50f0787b541eaa6

Observation 8c2f0518-867a-4c4a-9849-f4602c19c6fa · outbound

This paper cites Distributed optimization and statistical learning via the alternating direction method of multipliers.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Distributed optimization and statistical learning via the alternating direction method of multipliers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.876277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.468048Z digest=sha256:65fdc45a7fc8a2fc4aa782194ce7a512acd4eafc1bf91c4253f234d918b60ab9

Observation 6cf160a8-e1af-4778-b2c9-83d0377eb545 · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Stochastic spectral descent for restricted boltzmann machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.865195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.472410Z digest=sha256:b0b7e2073a7d7eab09e9fad89d773d7dca838548fafbcdbd5ded7129938ae3e3

Observation bbf7d13b-7496-4394-bab7-87a183480732 · outbound

This paper cites and Pock, T.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Pock, T

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.854101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.475776Z digest=sha256:eef89f8569539718d865e79ab1a1dc9ec927c12bd001899be681c33cf6e3b469

Observation 4299bd35-014a-4e46-b469-d40de0536ef7 · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.482602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.482602Z digest=sha256:f639e33389d321fe571f9a1c7b9c9ef0f616768b1e19a83bcc6a6694adf16477

Observation daea1544-84d4-4f64-b7fd-68adce40a54d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Symbolic discovery of optimization algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.843957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.486131Z digest=sha256:ac4797135861d71766f2d9b6295dc87a1fd47982b087015f578616fc44e37b03

Observation f3125a71-49a0-4a47-af5e-3d31edbf12e4 · outbound

This paper cites and Mehta, H.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Mehta, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.832640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.489575Z digest=sha256:aecfafdb4bf4f14c22b5a6a312783b0d96f9236c987b0c26a8423b178042cb84

Observation bed42d3e-476b-4ee1-bf47-18a02c229b59 · outbound

This paper cites On hilbert’s metric for simplices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training On hilbert’s metric for simplices

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.821392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.492994Z digest=sha256:12af35824eb03f15c52eee510ecff13e92327e1becd30323fb73eb0a7fcd01f5

Observation 7bf6bb02-4816-494e-a22a-7d422156ceb6 · outbound

This paper cites The Llama 3 Herd of Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.496672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.496672Z digest=sha256:153206c3ea3fa1e652d21fbe7dd2f61ef7a05dcc697176e37c699296abb84730

Observation 5a1036d1-bd39-485a-bbb4-42a88518df53 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.810400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.500426Z digest=sha256:85a0be74716f63e2b01e0a584f472429a4180b8fda95eba413062d2e627e4706

Observation 807a43a5-376f-44f6-a5ff-ff3e66350ff1 · outbound

This paper cites and Lorenz, J.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Lorenz, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.799506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.504111Z digest=sha256:f61cde49b729911ea6b7e18092d0f8570ffa140fda3c522c9707679f41ab9e87

Observation 0a3d308f-309a-4bda-8ce9-3dee363a3b86 · outbound

This paper cites Eigenvalue-corrected natural gradient based on a new approximation.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Eigenvalue-corrected natural gradient based on a new approximation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.787737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.507644Z digest=sha256:b30268834f0dbd40cc31b1bd69272263d59af378a3c9c4d5d3b08caef0d27516

Observation f3a89ebb-4335-405f-ac92-2ae195999eb1 · outbound

This paper cites Shampoo: Preconditioned Stochastic Tensor Optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.510817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.510817Z digest=sha256:a6b1790be5ef3ac4a4311fffca0579d88b9dd92b66a0efd719463ba4092f15e8

Observation bbb83fe5-3150-4ccf-81ac-2e61517b0bdc · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.514931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.514931Z digest=sha256:6a2a6080df8bbf0aaa8848e860ee47752e59ad32beff000276f94ba112d0f63d

Observation 6d529124-8416-4357-8a69-f4d6d4058af4 · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Beyond convexity: Stochastic quasi-convex optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.518833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.518833Z digest=sha256:123a4bcbf0073367b4e5bca9ecdcc5aaa9eac4b9dee7fc5db68c17df1f5c57bd

Observation 820b1f11-b874-4e60-a1fb-a7959e808908 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.522219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.522219Z digest=sha256:992ef2670f1122e065c1de6cbafd89a3524986f72e5326560184b5e35d6113b7

Observation f6927afa-6bcc-4b13-ba94-53b1afe52fe4 · outbound

This paper cites Iterative normalization: Beyond standardization towards efficient whitening.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative normalization: Beyond standardization towards efficient whitening

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.735222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.525891Z digest=sha256:bcfa84aaedc748e676f7528e1d0a20a9f9493d1c186f0a5dfffd41d6dd06b039

Observation b3245607-4341-4ef3-8480-8707711b816a · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Muon: An optimizer for hidden layers in neural networks, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.681640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.529508Z digest=sha256:b943233a3d4eb8732b7d563e0c89d5d44a7d46bb3bdea8ec09b353987f172858

Observation add29701-088d-434a-b17a-76bad058594f · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.623494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.532971Z digest=sha256:a1494aa498b6cbc0c62f44b92af109c4c9be53944b3358cfdef2726a5c82f58f

Observation 6080fe88-34be-44b1-8865-08a9657870ae · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.536830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.536830Z digest=sha256:86c769d6d24fd502f6de51dbc4bc277290b0335de2525c57e3caf6d6073c5b6a

Observation f8d51ea9-6283-4132-a32b-95ef53b7d28e · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.540265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.540265Z digest=sha256:dda126c48b34a31ab601f51e2b3ab5066fd9ba0cf46451c176b61621d1bc0c72

Observation 34a40cd0-c212-4fd7-ab57-f29d81426256 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.544235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.544235Z digest=sha256:3e564eb5a6e926252f12149c91ff2927a142564e63d19fc8adba7f9f626651a9

Observation 56561ed3-f3da-4622-bb5c-cde77f1bfe3b · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.558644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.561000Z digest=sha256:a2285ab2ad091824b878f78dde50279d0c5eb26f9b86d93d7605546da8d3ac1d

Observation 02956aa1-6bba-4a64-b1e0-378a505245c6 · outbound

This paper cites Towards faster training of global covariance pooling networks by iterative matrix square root normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.540119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.586474Z digest=sha256:a51e6b4e9c1ceb480022775cb03524cf4b6867e6812748c1a85db956915b57b0

Observation 91f59945-4410-454e-9dd0-6fce281b283b · outbound

This paper cites Relora: High-rank training through low-rank updates.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Relora: High-rank training through low-rank updates

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.529355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.613591Z digest=sha256:fb59f16cc4042198278ac88e32beb8b66b65724c5ab013e886007bf43db0e4a8

Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.643471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.643471Z digest=sha256:384b4631de493ef4601ee9382cc296f189a5c994a68aeda605c732036b8df371

Observation 226a59d9-2e4a-4ced-a9f0-f4b43badc942 · outbound

This paper cites Decomposition through formalization in a product space.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Decomposition through formalization in a product space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.518803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.669813Z digest=sha256:99729a992b2b1a0bb15bdecde8bded1ab0195592e232607dd81ff090251b94c9

Observation c6aa30c5-ec4d-443b-9791-a2c525b24b47 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.508248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.685068Z digest=sha256:605ba3e41bd799060cc5d4babc450d934c919d12b9d2daab0b1e591fcc521841

Observation 39ffadbf-194d-4c97-8e4e-3b406d8b9716 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Zero: Memory optimizations toward training trillion parameter models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.713139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.713139Z digest=sha256:ebfc0cf1c29f8cc6ae9111e6dfe3f307c9314adf87ad7c6b6a358fc31c5bf218

Observation 945ca777-367e-43f9-8627-88b80a777228 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.492141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.744517Z digest=sha256:bb7624d132bce808bbef9dacd9d8a5c4b6fa95406ba6f22046dd9f57910e7bb1

Observation 5dd8a654-35d7-4b34-9180-f3bfd7b25fe4 · outbound

This paper cites A relationship between arbitrary positive matrices and doubly stochastic matrices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A relationship between arbitrary positive matrices and doubly stochastic matrices

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.757416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.757416Z digest=sha256:7096ddf25d19d383bbd90b943d84ef00a84c96be6d5d331cf731559efcabae55

Observation 8bbad764-6fb7-4ec2-8bdd-239f32ec1b26 · outbound

This paper cites and Knopp, P.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Knopp, P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.761340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.761340Z digest=sha256:0d83332eebff70c621f2abdccf81862e05ae72a05fad8b716881a1bfb41b8945

Observation b6e29277-8a95-484d-b646-8346a7ffa4bf · outbound

This paper cites Fast differentiable matrix square root and inverse square root.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fast differentiable matrix square root and inverse square root

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.467626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.765130Z digest=sha256:b9c8c3e444a1ac54945780a2b7bebf74b7ebcf32c8a88a4acf32dd07127be522

Observation 1907b5da-4d76-4ef1-87d2-df117dc1f8e6 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.456973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.768652Z digest=sha256:f115194420ba03f9567fc27b06075287da106e1ddc255be739448819e4a8f52b

Observation 5217120b-dfbe-4c6c-bb3c-8a46fd1c51cd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.772296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.772296Z digest=sha256:7bd36e213b3c1c10568bbfe0b33d450d6e32b095f138c6bcde673a977dc9b10e

Observation 555089af-2e8b-4842-91c8-0301ec813625 · outbound

This paper cites Functional operators: Measures and integrals, volume 1.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Functional operators: Measures and integrals, volume 1

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.446304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.776322Z digest=sha256:940eff1f25db9d3e77899bf0c4718222ede1047f2a2856352e9aa2832b262f2a

Observation 0943772a-8700-42f1-93ce-bd84ff49c3bc · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.434492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.779538Z digest=sha256:5ce1726ac240c2dcf522da1c1f19d0ef591a439292c9386ca47a2cf78cc4cfad

Observation f4028448-fec2-4f51-9294-9d575a3c015a · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.786570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.786570Z digest=sha256:c36bcf8f4e4a33bf71f65c93faf988f189908bab16517e7171d12b505079dc0d

Observation c3dcb6c0-fd3a-4990-ae1b-f914be422b5b · outbound

This paper cites Large Batch Training of Convolutional Networks.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Training of Convolutional Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.789854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.789854Z digest=sha256:b8c7ea465d6f906c50bac032cae53cb3b862c0c7132b35a1246ce7cba1aebd4a

Observation e8841d1e-9bc4-4414-bfe0-5f8a69154ff8 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.793583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.793583Z digest=sha256:5f9b6167051d9070d17226c04741bb4f910cc31b60b1c1c1e66e901f196919c7

Observation 8a9cfafa-3b90-4834-9d6b-fdada3dad04f · outbound

This paper cites and Sennrich, R.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Sennrich, R

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.797421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.797421Z digest=sha256:abfb2dd59eb79122d544a92974c6985cb05b5967dd37cf8de49e8bcabc0bfdb6

Observation 47a10613-9aab-45f7-a17a-6476b27fdf96 · outbound

This paper cites P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S.

Gradient Multi-Normalization for Stateless and Scalable LLM Training P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.417347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.801164Z digest=sha256:6264480bfbb1ce8be6d2a7c74207268cf5944e83d4eaa8aec311531b40b5a0e4

Observation 6d77efdc-e81c-4111-b116-cb3594c20dbf · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Adam-mini: Use Fewer Learning Rates To Gain More

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.804426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.804426Z digest=sha256:8d2cf9942e00e3c036f02babc3f5a78d3e5abe2988a6a39d9009dade81429227

Observation 14e30325-302c-4d31-b315-0b89b822b86b · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Gradient Multi-Normalization for Stateless and Scalable LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.808168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.808168Z digest=sha256:a906444172998102392d2fcbf4c366c6939cbfee810eb2c4e4986438464a27d0

Observation 552f658d-4508-401a-9dea-25d0d0515d53 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.812206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.812206Z digest=sha256:114c7af87ac4070e88cfe7ab06247c6e07dc96b622ef07d2e1d5bc93105b7644

Observation a6b8083a-4cd7-47bb-ad53-6e19cd7f3df8 · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Gradient Multi-Normalization for Stateless and Scalable LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.815685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.815685Z digest=sha256:c220f5fca84e2540c44c665a0b459ea207cc1e263cc01884d5a1eabdcc73a23a

Observation 69ac2474-aae8-426b-a383-655c9cc14123 · outbound

This paper cites write newline.

Gradient Multi-Normalization for Stateless and Scalable LLM Training write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.819526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.819526Z digest=sha256:cc4188226674b70656f72f2361da1987b1e5339fb385ee67ada10aeb0179555d

Pith citing papers

Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.221296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.221296Z digest=sha256:e52a57426d3450514a06c4b42f409faf91a70b7e8d56eeb6d687098967846099

Observation 903bcba7-1617-42cb-b2b1-21f27dad111d · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.860137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:a1851c54c8d48f020a993ec4c921bf6e04c8480cfde544885fb3215072ded6d8

Observation e25f7357-df6e-4089-bfd0-46d14d6b6413 · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.850767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:a2046627143135389c54307361534da1284344dd7c8984783c6bbc594d22e658