Pith. sign in

Paper Citation Record · LEDGER

Gradient Multi-Normalization for Stateless and Scalable LLM Training

As of 9 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 3 inbound Pith citation observations for arXiv:2502.06742.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06742 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.819526Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.221296Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T05:52:21.848754Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b08ad4a5-16fe-4a52-8ad6-c06e6506952c · outbound

This paper cites Layer Normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.413471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.413471Z digest=sha256:291e6808053cece0fdab508326fb161a4a769b9ed470b18734e44ab7a576f4c8

Observation f122e6d1-83f1-4c54-b2a7-b5a8d3177886 · outbound

This paper cites Iterative bregman projections for regularized transportation problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative bregman projections for regularized transportation problems

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.906346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.451252Z digest=sha256:5edabbdd3e13666d69af76e878e79bcfc9c9b38e1a040d4f73ee600f32a3c915

Observation 4b0d4c41-37f0-483c-943c-122dde9a1987 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Old Optimizer, New Norm: An Anthology

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.455661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.455661Z digest=sha256:5322deb28a8323e8cf9abf5c650d0fd3c754d997e7f345a7b456fc3c2b1fce57

Observation 4a520790-e2ea-4bb8-8431-e9247760aaaa · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training signsgd: Compressed optimisation for non-convex problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.459883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.459883Z digest=sha256:8ac83868dffdef7280d7aec5397301ba5aa8ebfeaef8eabf70f647a4d8916687

Observation 2c1c0c26-fcb7-4c9e-abc0-a04d482caf08 · outbound

This paper cites Proximal alternating linearized minimization for nonconvex and nonsmooth problems.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Proximal alternating linearized minimization for nonconvex and nonsmooth problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.888347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.463975Z digest=sha256:238226490f8d3041c7455912d273bd1ff72dc7d1516fd1a622f7804b4695cc16

Observation 8c2f0518-867a-4c4a-9849-f4602c19c6fa · outbound

This paper cites Distributed optimization and statistical learning via the alternating direction method of multipliers.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Distributed optimization and statistical learning via the alternating direction method of multipliers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.876277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.468048Z digest=sha256:01cac7d0d1580518da55035fd703a6ccff8347b07ae2e7014e35d9d85e39f3ec

Observation 6cf160a8-e1af-4778-b2c9-83d0377eb545 · outbound

This paper cites Stochastic spectral descent for restricted boltzmann machines.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Stochastic spectral descent for restricted boltzmann machines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.865195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.472410Z digest=sha256:d3d07af6227fa413f6caa72e0af3818a15495aee29871683d6ecc2c4126aac11

Observation bbf7d13b-7496-4394-bab7-87a183480732 · outbound

This paper cites and Pock, T.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Pock, T

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.854101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.475776Z digest=sha256:93831b1a052b0beb3f435728d26143a7a28473ba0a64496ccaa79aa9cae9d1b2

Observation 4299bd35-014a-4e46-b469-d40de0536ef7 · outbound

This paper cites Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fira: Can we achieve full-rank training of llms under low-rank constraint?, 2024 b

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.482602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.482602Z digest=sha256:f044d3ac8dd8da2ba27b77395a4ed4c11b9582444109df1a8e3d8be1c8fe8950

Observation daea1544-84d4-4f64-b7fd-68adce40a54d · outbound

This paper cites Symbolic discovery of optimization algorithms.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Symbolic discovery of optimization algorithms

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.843957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.486131Z digest=sha256:e533966665e4ef0e838ff1bdf65fe643430be80b4b641c556eec36f8565a9272

Observation f3125a71-49a0-4a47-af5e-3d31edbf12e4 · outbound

This paper cites and Mehta, H.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Mehta, H

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.832640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.489575Z digest=sha256:39f5572c4cbac0570d31b8079b333d5d80f2ff8039d3152fa24804cfbbc0b837

Observation bed42d3e-476b-4ee1-bf47-18a02c229b59 · outbound

This paper cites On hilbert’s metric for simplices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training On hilbert’s metric for simplices

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.821392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.492994Z digest=sha256:09ad3ce9ef2514c154f39db7c96a81b7ae1ecd7c9342efeb8d9c61ed6d9edbb9

Observation 7bf6bb02-4816-494e-a22a-7d422156ceb6 · outbound

This paper cites The Llama 3 Herd of Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.496672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.496672Z digest=sha256:b93d5254ba3aec1054b13a98d7358c109b035080b108272ce153478e5eac34e0

Observation 5a1036d1-bd39-485a-bbb4-42a88518df53 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.810400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.500426Z digest=sha256:54563a835ff793f839fac8aa474cac03721781c9c731551dc1f3734b83920821

Observation 807a43a5-376f-44f6-a5ff-ff3e66350ff1 · outbound

This paper cites and Lorenz, J.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Lorenz, J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.799506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.504111Z digest=sha256:37cc6154920103bd8cd6f2304997edf610228f9848f11cb95c2ed8d57e707221

Observation 0a3d308f-309a-4bda-8ce9-3dee363a3b86 · outbound

This paper cites Eigenvalue-corrected natural gradient based on a new approximation.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Eigenvalue-corrected natural gradient based on a new approximation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.787737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.507644Z digest=sha256:ede80c677311478e1e61f5f1d350b75ea306e04ddc94a4c1f5a82ca173448388

Observation f3a89ebb-4335-405f-ac92-2ae195999eb1 · outbound

This paper cites Shampoo: Preconditioned Stochastic Tensor Optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Shampoo: Preconditioned Stochastic Tensor Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.510817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.510817Z digest=sha256:eddc281b9efa63e45179096d3efe72a3b5f0c24e56a855ff82d76c01d1b6a96f

Observation bbb83fe5-3150-4ccf-81ac-2e61517b0bdc · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.514931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.514931Z digest=sha256:76784fe1af01985318989b6e9f78e7d8b371c0ad94a8d1a73f0f66ce383ed578

Observation 6d529124-8416-4357-8a69-f4d6d4058af4 · outbound

This paper cites Beyond convexity: Stochastic quasi-convex optimization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Beyond convexity: Stochastic quasi-convex optimization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.518833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.518833Z digest=sha256:fb9fbf325d843631e20395af5fc15234948038e3485f14c79931f8a3c7d5f146

Observation 820b1f11-b874-4e60-a1fb-a7959e808908 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.522219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.522219Z digest=sha256:682e186a49d66d7d39a7dbaf6b3380bc7a8d748c35f2121e456842c6eecf403f

Observation f6927afa-6bcc-4b13-ba94-53b1afe52fe4 · outbound

This paper cites Iterative normalization: Beyond standardization towards efficient whitening.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Iterative normalization: Beyond standardization towards efficient whitening

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.735222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.525891Z digest=sha256:21e1f4da134c5a228b1468f5be3119e00feab8200260239a5822363f8f3ebbde

Observation b3245607-4341-4ef3-8480-8707711b816a · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Muon: An optimizer for hidden layers in neural networks, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.681640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.529508Z digest=sha256:9424dfbc377c2ec45a3d6e4f5c06a00539ec2161e69f25199e0f2c7b10ba8169

Observation add29701-088d-434a-b17a-76bad058594f · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.623494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.532971Z digest=sha256:b01e02c4c8b81f31bf86e4272d6ad4ee72a774f7c47f66b4d3d0884a10aafd40

Observation 6080fe88-34be-44b1-8865-08a9657870ae · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.536830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.536830Z digest=sha256:5647c246e3d440f8d5e82086899a573345c931d8d13ec7fc092a3817e1c49cb4

Observation f8d51ea9-6283-4132-a32b-95ef53b7d28e · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.540265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.540265Z digest=sha256:8957a0670a595381f37b2545ac11da1e2dcbad42712e6625f5f6a2b31b11d7d1

Observation 34a40cd0-c212-4fd7-ab57-f29d81426256 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.544235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.544235Z digest=sha256:7b74f7c54271fa20abb9c3c5e9792ed137f5f0add9865fb6ae570a41a75340fb

Observation 56561ed3-f3da-4622-bb5c-cde77f1bfe3b · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.558644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.561000Z digest=sha256:462bdc3351d3bf97cc6233a2af7a98b29b26472388b30d1ca45e14b02869cbe8

Observation 02956aa1-6bba-4a64-b1e0-378a505245c6 · outbound

This paper cites Towards faster training of global covariance pooling networks by iterative matrix square root normalization.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Towards faster training of global covariance pooling networks by iterative matrix square root normalization

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.540119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.586474Z digest=sha256:cecf547d58c87e34486b79f661597dabb9b4aad338289c812b9e322cd12ff6c0

Observation 91f59945-4410-454e-9dd0-6fce281b283b · outbound

This paper cites Relora: High-rank training through low-rank updates.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Relora: High-rank training through low-rank updates

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.529355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.613591Z digest=sha256:eabc9ae3c01423d4fde09d8af6718b089ff63663aec2fa98b8c9b7f31bb0f7a8

Observation 5897e39f-a8d8-44e2-863d-cbf31763c4dc · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Gradient Multi-Normalization for Stateless and Scalable LLM Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.643471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.643471Z digest=sha256:096e1af97a9c8b12351c86319f5dcf1c8e51df4b954b5f1f573528da2eedf50a

Observation 226a59d9-2e4a-4ced-a9f0-f4b43badc942 · outbound

This paper cites Decomposition through formalization in a product space.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Decomposition through formalization in a product space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.518803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.669813Z digest=sha256:7f147090620c326e23371e65fccb61a2164b9c5a523d6929aa1ea50cecd44dca

Observation c6aa30c5-ec4d-443b-9791-a2c525b24b47 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.508248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.685068Z digest=sha256:5609f3778f67b67b67456b48330c166a28c5ef1fcb0d139d058a78c706d16dd8

Observation 39ffadbf-194d-4c97-8e4e-3b406d8b9716 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Zero: Memory optimizations toward training trillion parameter models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.713139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.713139Z digest=sha256:4cd07f5e55fa4c4ebc900acf3428d95df1de891000e63ffed76ecbfaf37d1a51

Observation 945ca777-367e-43f9-8627-88b80a777228 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.492141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.744517Z digest=sha256:2cde27de8dc96340e7db61c0ce3ae392670009e0caeca0e0a1460079b29ba5eb

Observation 5dd8a654-35d7-4b34-9180-f3bfd7b25fe4 · outbound

This paper cites A relationship between arbitrary positive matrices and doubly stochastic matrices.

Gradient Multi-Normalization for Stateless and Scalable LLM Training A relationship between arbitrary positive matrices and doubly stochastic matrices

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.757416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.757416Z digest=sha256:b80d3cebc1e6e3008cdee8013b6c2106866c96b2d8859e6b5c19852eee6364b8

Observation 8bbad764-6fb7-4ec2-8bdd-239f32ec1b26 · outbound

This paper cites and Knopp, P.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Knopp, P

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.761340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.761340Z digest=sha256:fee2db5d44b02460d5f4924655daaa47c88e41ab71483b1770c8239883705094

Observation b6e29277-8a95-484d-b646-8346a7ffa4bf · outbound

This paper cites Fast differentiable matrix square root and inverse square root.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Fast differentiable matrix square root and inverse square root

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.467626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.765130Z digest=sha256:eb4ac25bfb2ffa023f62e75ae818e0e17dad045f5970dde9f7ca3adbdfd12946

Observation 1907b5da-4d76-4ef1-87d2-df117dc1f8e6 · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.456973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.768652Z digest=sha256:b01cd7169dae9eae001cc5c6d2c6e8eda320a5d22abaf20b9443bce5fa73e7ed

Observation 5217120b-dfbe-4c6c-bb3c-8a46fd1c51cd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.772296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.772296Z digest=sha256:38602718fa4bb334e54e50d6d51ee73abaf91433d5eabb568bbc504d81d81b90

Observation 555089af-2e8b-4842-91c8-0301ec813625 · outbound

This paper cites Functional operators: Measures and integrals, volume 1.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Functional operators: Measures and integrals, volume 1

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.446304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.776322Z digest=sha256:c09a76c20888f60de8396c265218faf0f006c9ce0270ea703f6bd12f3bb84375

Observation 0943772a-8700-42f1-93ce-bd84ff49c3bc · outbound

This paper cites an unresolved cited work.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:37:24.434492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.779538Z digest=sha256:760900a6aaa183dfc1c4cd2fb6aa010571b325b38d24e2bb39a24da379b0b4e2

Observation f4028448-fec2-4f51-9294-9d575a3c015a · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

Gradient Multi-Normalization for Stateless and Scalable LLM Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.786570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.786570Z digest=sha256:cfff7ef55b851992802ac32e4b731490f007b4a029811278599ce746cab50415

Observation c3dcb6c0-fd3a-4990-ae1b-f914be422b5b · outbound

This paper cites Large Batch Training of Convolutional Networks.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Training of Convolutional Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.789854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.789854Z digest=sha256:934485f8702305609c340f787de13767d22e7cd7aa47daf74ac42fbff17badc0

Observation e8841d1e-9bc4-4414-bfe0-5f8a69154ff8 · outbound

This paper cites Large Batch Optimization for Deep Learning: Training BERT in 76 minutes.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.793583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.793583Z digest=sha256:864aefc786c54a9ab8a1e1fbe73115ed4ad6bfd946c293e5e66e7c85fdeadda9

Observation 8a9cfafa-3b90-4834-9d6b-fdada3dad04f · outbound

This paper cites and Sennrich, R.

Gradient Multi-Normalization for Stateless and Scalable LLM Training and Sennrich, R

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.797421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.797421Z digest=sha256:911b5bda7e3cfda2337750d3bfd9a65ae2ba6d903cade4d422b2d3c5cceaf8d5

Observation 47a10613-9aab-45f7-a17a-6476b27fdf96 · outbound

This paper cites P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S.

Gradient Multi-Normalization for Stateless and Scalable LLM Training P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:37:24.417347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T14:37:23.801164Z digest=sha256:3b2cd9aac2ac41141970c6e98cc6f55698b2d80bfab31e84ba865b54c02ced2c

Observation 6d77efdc-e81c-4111-b116-cb3594c20dbf · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Adam-mini: Use Fewer Learning Rates To Gain More

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.804426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.804426Z digest=sha256:364d1a2207b273c825c9799562e7decb012d969559afb79fc3704ef7bf84c23c

Observation 14e30325-302c-4d31-b315-0b89b822b86b · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Gradient Multi-Normalization for Stateless and Scalable LLM Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.808168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.808168Z digest=sha256:ef4e3c40549e5d7437ff23e7c003a461624e9ee9dbf2747c71b2e880af62137f

Observation 552f658d-4508-401a-9dea-25d0d0515d53 · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.812206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.812206Z digest=sha256:86d589eafcceb3bb50867f6ce1b45d4f9eb5c36d5b7f5bcfadd3ddf139baaea6

Observation a6b8083a-4cd7-47bb-ad53-6e19cd7f3df8 · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Gradient Multi-Normalization for Stateless and Scalable LLM Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.815685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.815685Z digest=sha256:a10a80b1ea06de2dd1497c6977225faf4d03247632af179b74a430c681eb4cb5

Observation 69ac2474-aae8-426b-a383-655c9cc14123 · outbound

This paper cites write newline.

Gradient Multi-Normalization for Stateless and Scalable LLM Training write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.819526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.819526Z digest=sha256:80c9fcafd55c6efbd8124e7f8487a464609f5bef9241fb6465ee41683aa569bc

Pith citing papers

Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · inbound

Low-rank Momentum Factorization for Memory Efficient Training cites this paper.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.221296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.221296Z digest=sha256:3424ca9efc9faf9ff8a48f58a8796249e2a1e2db0a1526eddc97055e19099ab7

Observation 903bcba7-1617-42cb-b2b1-21f27dad111d · inbound

Demystifying Manifold Constraints in LLM Pre-training cites this paper.

Demystifying Manifold Constraints in LLM Pre-training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:08.860137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:44:44.438637Z digest=sha256:109378e24b0462451515951a1e1b0e2ae9e4e7270e853d99ade522fd68af809d

Observation e25f7357-df6e-4089-bfd0-46d14d6b6413 · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.850767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:68920ec080fd77d9411f1328866a9f09ee8b9112cacf4d91f856dc8a99e87851