Pith. sign in

Paper Citation Record · LEDGER

Low-rank Momentum Factorization for Memory Efficient Training

As of 15 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2507.08091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08091 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:35:40.252196Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact3
  • verified fuzzy20
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3d091fdb-94ec-4d6e-8c7c-3f6b9ada94e2 · outbound

This paper cites Memory efficient adaptive optimization.

Low-rank Momentum Factorization for Memory Efficient Training Memory efficient adaptive optimization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.844330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.092595Z digest=sha256:6cf3c9fc24ea88a35c1ac9bd9581af6427026255551d2b6056dd3c1339f0a3d8

Observation b532b623-71d0-40f4-b010-b66bc957a580 · outbound

This paper cites Lower bounds for non-convex stochastic optimization.

Low-rank Momentum Factorization for Memory Efficient Training Lower bounds for non-convex stochastic optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.095813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.095813Z digest=sha256:c0c6af35b044ef2bbed9cb0533cd3583509dd41c9c3c4ef29a9e607128ab2eb3

Observation 3fe12a6e-83c1-43dd-9681-700cffe0ce6e · outbound

This paper cites Modular Duality in Deep Learning.

Low-rank Momentum Factorization for Memory Efficient Training Modular Duality in Deep Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.098289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.098289Z digest=sha256:b660080c3ff03f4b7160ac084ffddb3455e89150a1308f87ba113ad0592af34d

Observation ce0bfadf-c388-47e9-b543-633023dd2d42 · outbound

This paper cites Old Optimizer, New Norm: An Anthology.

Low-rank Momentum Factorization for Memory Efficient Training Old Optimizer, New Norm: An Anthology

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.101337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.101337Z digest=sha256:02aa918cb185ca9b0ed5f8430c033a0f88caead93948c8eeb5c9d3f811ac83fb

Observation ca466eb8-9b8a-426e-aba8-48cdb118eb72 · outbound

This paper cites signsgd: Compressed optimisation for non-convex problems.

Low-rank Momentum Factorization for Memory Efficient Training signsgd: Compressed optimisation for non-convex problems

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.832449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.104099Z digest=sha256:11dbe3fdace7f250cf07e1fa87768a42565d83ca12ce7e5fa5e36f40ec330ba4

Observation bbb69756-4f58-421b-9dca-bf1f25d5c1a4 · outbound

This paper cites Automatic Gradient Descent: Deep Learning without Hyperparameters.

Low-rank Momentum Factorization for Memory Efficient Training Automatic Gradient Descent: Deep Learning without Hyperparameters

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.106492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.106492Z digest=sha256:385f1fc5dd4fcbe079e15a0e54220e6819e15a74dbdbc0fc8eef59f29f343df8

Observation 01992ade-7bfc-4cf9-bad7-33c5713dd188 · outbound

This paper cites Symbolic discovery of optimization algorithms.

Low-rank Momentum Factorization for Memory Efficient Training Symbolic discovery of optimization algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.825256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.109327Z digest=sha256:f2888fdee985a3f068c853fffc046b04e5866722ef0d522f1326250335ce5d23

Observation 37add5ea-1f3f-4a0b-8e0b-d16bbe160b7b · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

Low-rank Momentum Factorization for Memory Efficient Training 8-bit Optimizers via Block-wise Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.111493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.111493Z digest=sha256:472c013e797b86c83ab4b7764c5a4da44ddadb0232aba3c6e7b18495ac0753be

Observation 1e8c6ba0-3238-45b3-965b-ea5cae56aedf · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

Low-rank Momentum Factorization for Memory Efficient Training Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.114360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.114360Z digest=sha256:f203c82b65545ed9a38aa0226d14054d536661a19690ef0f599b741143a59c40

Observation 92ea956e-c64b-40a3-8cc9-9336f072b503 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

Low-rank Momentum Factorization for Memory Efficient Training Adaptive subgradient methods for online learning and stochastic optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.117442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.117442Z digest=sha256:026a9b3f166725a192568df380984eaea0fbef87987c9d6d72579fe8a4cd37e4

Observation fba15dfa-01fe-458d-bf68-7558f6d01345 · outbound

This paper cites Combining axes preconditioners through kronecker approximation for deep learning.

Low-rank Momentum Factorization for Memory Efficient Training Combining axes preconditioners through kronecker approximation for deep learning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.808185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.119531Z digest=sha256:a40e6e934a08beb1cb62010bcd7675ba196353c8f7e34cd4e83fe74a2f834b10

Observation d8b1410a-a235-4978-bcc1-8cd76cb2083a · outbound

This paper cites Sketchy: Memory-efficient adaptive regularization with frequent directions.

Low-rank Momentum Factorization for Memory Efficient Training Sketchy: Memory-efficient adaptive regularization with frequent directions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.800933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.121654Z digest=sha256:53d1dbddedaf8a8ea9c0f0fa14cf31b183a2a3cb822839a9d45e571ca1055ff4

Observation f3523b01-4493-4be9-9ec8-8377958da5fd · outbound

This paper cites Fast approximate natural gradient descent in a kronecker factored eigenbasis.

Low-rank Momentum Factorization for Memory Efficient Training Fast approximate natural gradient descent in a kronecker factored eigenbasis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.792974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.123726Z digest=sha256:02923911500385fc91bea69f3b3cb9903334b68e06d5c499b9c585f0b39aabf7

Observation 43c30931-6709-46ec-ba7b-63fdf9549231 · outbound

This paper cites Improving neural network training in low dimensional random bases.

Low-rank Momentum Factorization for Memory Efficient Training Improving neural network training in low dimensional random bases

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.785399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.125769Z digest=sha256:92c7e90b774b18925338e3fed1a8d47138c2e64efae032f4665d1f5fef45022b

Observation 46d2730b-4bb7-4c8c-83c8-1e7073e1420b · outbound

This paper cites A kronecker-factored approximate fisher matrix for convolution layers.

Low-rank Momentum Factorization for Memory Efficient Training A kronecker-factored approximate fisher matrix for convolution layers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.777756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.127897Z digest=sha256:5cec22bc7d9fb2940aa441b3e60f698696774ae96aa9723e1209f4198e1d3b46

Observation d2334a67-9cef-41e9-bad2-19cb0492beaa · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Low-rank Momentum Factorization for Memory Efficient Training OLMES: A Standard for Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.130029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.130029Z digest=sha256:588a86ce957f4841522db6cefdfac093c40f917304dea4600f3c21e6966d2238

Observation cc7931f7-ca09-446e-b00e-37a06f26928d · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Low-rank Momentum Factorization for Memory Efficient Training Shampoo: Preconditioned stochastic tensor optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.770861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.132495Z digest=sha256:78f01e3e7441af76747e7ea54cd80f27facf2a3cc30f906d3dd0c6691c99b6b5

Observation 71536acc-4a68-4f4f-bfcb-ea96cc205e28 · outbound

This paper cites Gradient Descent Happens in a Tiny Subspace.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Descent Happens in a Tiny Subspace

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.134826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.134826Z digest=sha256:1bd764196ba8c0476fdc668523e20fa33e3a80f4c082b75aefccdedba81e39ae

Observation 7d13c518-a176-44f3-a921-2569a2abe3a7 · outbound

This paper cites Flora: Low-Rank Adapters Are Secretly Gradient Compressors.

Low-rank Momentum Factorization for Memory Efficient Training Flora: Low-Rank Adapters Are Secretly Gradient Compressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.137244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.137244Z digest=sha256:d10b28621c41f432f02373a1e964cceb58d7e6bf311e8f4fdf6eed186d8b3fce

Observation bd10f289-2c68-460e-b447-ed4f2b5479b8 · outbound

This paper cites Topics in matrix analysis.

Low-rank Momentum Factorization for Memory Efficient Training Topics in matrix analysis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.763802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.139626Z digest=sha256:16c5e420a65a48b3e8d5498d295ef7d6e0f2c4930cd15eb98593a6a66d54dca9

Observation a823097d-4846-4f24-9830-cd0abd3df947 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Low-rank Momentum Factorization for Memory Efficient Training Parameter-efficient transfer learning for nlp

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.141663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.141663Z digest=sha256:d7891f21dcf2d6654134ce1fe9b66c0190c714eed348736c9f5eb0a32593e6f0

Observation 81fac90f-2158-41c7-b5b7-fdaa8185c668 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Low-rank Momentum Factorization for Memory Efficient Training LoRA: Low-Rank Adaptation of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.143830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.143830Z digest=sha256:483bfbe7042a8736050ef61e0db4292dcb3fd0be804b97d9ca88658827ea80ef

Observation 305cc2a1-a26b-4291-9bf4-8fd0fb450d06 · outbound

This paper cites modded-nanogpt: Speedrunning the nanogpt baseline, 2024 a.

Low-rank Momentum Factorization for Memory Efficient Training modded-nanogpt: Speedrunning the nanogpt baseline, 2024 a

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.751685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.146130Z digest=sha256:545ac9bba141b0f134cbcc5c027babca2647cf374e701376b4f69cba43c28069

Observation 86df6ef4-909a-4fda-b1dd-0e25d75367f4 · outbound

This paper cites Muon: An optimizer for hidden layers in neural networks, 2024 b.

Low-rank Momentum Factorization for Memory Efficient Training Muon: An optimizer for hidden layers in neural networks, 2024 b

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.743734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.148307Z digest=sha256:f3f047ee848fad3f86bab33b63924f880d9676b63d50ce03bdb4d2e3b28f6348

Observation 705f936e-4b9c-4cdc-84c2-2819e351c7c1 · outbound

This paper cites A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA.

Low-rank Momentum Factorization for Memory Efficient Training A Rank Stabilization Scaling Factor for Fine-Tuning with LoRA

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.150396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.150396Z digest=sha256:f154bededef566a247286a134462643ee93dbecf76daca0a397b10b8839b6f44

Observation b05009ba-c43f-40bb-8ca5-f9d0c7d3d80f · outbound

This paper cites Scaling Laws for Neural Language Models.

Low-rank Momentum Factorization for Memory Efficient Training Scaling Laws for Neural Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.153241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.153241Z digest=sha256:2c579d23d6c1ac9bc6842242d22cc3f336ca0e91f9a6b7cce4df1c10273593a9

Observation 70e53d3a-442f-4abb-aa35-4c9f769be0fb · outbound

This paper cites Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition.

Low-rank Momentum Factorization for Memory Efficient Training Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.577672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.155578Z digest=sha256:7755a6b1a0153760ec1d00581f0fabbb70ee3a25127b8dcc091760c5a6669886

Observation 58235faa-8bb5-44c0-b417-5c000c6c6e7f · outbound

This paper cites Kingma and Jimmy Ba.

Low-rank Momentum Factorization for Memory Efficient Training Kingma and Jimmy Ba

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.735655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.157919Z digest=sha256:70a838592f3a86253e53d4516f793a43491c1a03d4d2c662b30bf58c53addc7b

Observation 5c145465-698e-4d14-940f-d61f01f4069e · outbound

This paper cites VeRA: Vector-based Random Matrix Adaptation.

Low-rank Momentum Factorization for Memory Efficient Training VeRA: Vector-based Random Matrix Adaptation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.160096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.160096Z digest=sha256:b5be329ba04879dda790da93e0fc5033b294a238fa48509c1b46be9c98c991dd

Observation b72f3667-9544-45d7-a43d-53d61327f79e · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Low-rank Momentum Factorization for Memory Efficient Training Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.162460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.162460Z digest=sha256:d6766ee142c3efa314a7a8ab28f6043f4fffb902f2689ee9a1f6803093f88d2f

Observation 33ae95ce-59c6-4ccc-a4b7-14e26ed11296 · outbound

This paper cites Scalable optimization in the modular norm.

Low-rank Momentum Factorization for Memory Efficient Training Scalable optimization in the modular norm

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.728448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.165018Z digest=sha256:a3b21730a8512db37a74479968654b1f53103f28fd17b5c61394ad973d849eea

Observation 48215550-78b8-43a5-ba07-9e7dadf5e798 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Low-rank Momentum Factorization for Memory Efficient Training The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.167207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.167207Z digest=sha256:466132a17fa1a3c72a49453b396f70fdd36c133b1c377555ff8bc9a2fc3b8bfd

Observation 7b5e13d4-b336-4951-b0b3-e6aac2eafcf5 · outbound

This paper cites Memory efficient optimizers with 4-bit states.

Low-rank Momentum Factorization for Memory Efficient Training Memory efficient optimizers with 4-bit states

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.721169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.169646Z digest=sha256:41c8dee13f0a9f48daca8f71ed36bf8a2189801ab96773d113f89d38a366865b

Observation cd7943ca-2435-4c8e-9273-49dbd1343839 · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Low-rank Momentum Factorization for Memory Efficient Training Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.171944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.171944Z digest=sha256:e8f294b59b481b74bb83a62664a93f827e685c3b9a998db52f1b864172338da7

Observation 805a5003-36a7-43bb-bef0-dc9b00e5268b · outbound

This paper cites Relora: High-rank training through low-rank updates.

Low-rank Momentum Factorization for Memory Efficient Training Relora: High-rank training through low-rank updates

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.713960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.174393Z digest=sha256:d60b2d18355d41bcdd0fcd52c1e4bcd3278d325f0b4988119380ce5075f79fe5

Observation 2c3af3d2-a44b-4111-bd5e-3e79e5bc9429 · outbound

This paper cites On the limited memory bfgs method for large scale optimization.

Low-rank Momentum Factorization for Memory Efficient Training On the limited memory bfgs method for large scale optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.176629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.176629Z digest=sha256:5e33427ecc275ea30a07160183d8efd57a1328b8b47ec070c1e8a0312e9eb848

Observation 5bd23c7b-b0fa-4202-ba6c-cbb38d1c8ce9 · outbound

This paper cites Muon is Scalable for LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training Muon is Scalable for LLM Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.178866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.178866Z digest=sha256:eb0afea7e79c6c30696efc0f59d513380ccf1cd1740383c725f1592a57ae7009

Observation 8e34f1bb-6d34-4aaa-b435-50c22f551eb7 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Low-rank Momentum Factorization for Memory Efficient Training DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.181226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.181226Z digest=sha256:6a41bdf8f0abecade8202b141e8b513db3294d18bff8feb42cb6ebc4a26d6daa

Observation bdd016c8-8623-473e-b4f0-d65feca6e5fd · outbound

This paper cites Decoupled Weight Decay Regularization.

Low-rank Momentum Factorization for Memory Efficient Training Decoupled Weight Decay Regularization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.183763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.183763Z digest=sha256:8b5b9771672a49eff41784a3dd8f1c5b2d78fed06b1ab9df41c0a786def889c9

Observation 41f8934f-c0ca-4531-9d2b-174c4e6cfcac · outbound

This paper cites BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models.

Low-rank Momentum Factorization for Memory Efficient Training BAdam: A Memory Efficient Full Parameter Optimization Method for Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.186047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.186047Z digest=sha256:af51013f0d858185a07e4622a466b4b98c76ddaa063880cb9db776e0123069d5

Observation bbe63c38-2c25-47b1-a48b-e8cd765c3073 · outbound

This paper cites CAME: Confidence-guided Adaptive Memory Efficient Optimization.

Low-rank Momentum Factorization for Memory Efficient Training CAME: Confidence-guided Adaptive Memory Efficient Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.188746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.188746Z digest=sha256:48d4d2bee04f0f031e989ecf53631982a838f93886f9c894900198a289553f1e

Observation 5758d07c-1abf-450f-9b2b-0ada3efa269a · outbound

This paper cites AdaLomo: Low-memory Optimization with Adaptive Learning Rate.

Low-rank Momentum Factorization for Memory Efficient Training AdaLomo: Low-memory Optimization with Adaptive Learning Rate

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.191264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.191264Z digest=sha256:66e2485fe12b7a09e7dec334f620138f3aa7736097bc3f601eebd7e965ea43a6

Observation aa8b78af-79f6-4412-8c45-a7c3df309054 · outbound

This paper cites Full Parameter Fine-tuning for Large Language Models with Limited Resources.

Low-rank Momentum Factorization for Memory Efficient Training Full Parameter Fine-tuning for Large Language Models with Limited Resources

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.193728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.193728Z digest=sha256:bd620c352ad5892ef9ba8fc033633c237f2704c81d18792dc85797f09a87af5a

Observation b33f2a32-f26b-4cab-bc92-e0b243e5deb4 · outbound

This paper cites SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.196266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.196266Z digest=sha256:ba1aa8b3d1e28bba8db83a94bfea8e72de87de47b71c54d5baed8235fe4920cc

Observation 1ecbedd0-ba15-4420-b2ea-843301b4e023 · outbound

This paper cites New insights and perspectives on the natural gradient method.

Low-rank Momentum Factorization for Memory Efficient Training New insights and perspectives on the natural gradient method

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.199023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.199023Z digest=sha256:630c36b7b11bf40902c6f57e518613519397b9d20e6722b0bd53867df3f96a8e

Observation 535a1cff-93fd-4595-b5dc-57ac99e1d34f · outbound

This paper cites Optimizing neural networks with kronecker-factored approximate curvature.

Low-rank Momentum Factorization for Memory Efficient Training Optimizing neural networks with kronecker-factored approximate curvature

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.201548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.201548Z digest=sha256:e95d7e9b662af6f3a3501d6110d96c5d9d739aeab7e667230e7e03dbe28f3815

Observation d96bc255-68ef-4318-8d27-7152607d0cfa · outbound

This paper cites MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence.

Low-rank Momentum Factorization for Memory Efficient Training MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.279717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.203626Z digest=sha256:4ab9cf9510495ca8274794cfbbf39e6c9f927ce13a22c7c419ede8d9eacd5818

Observation 3f9068b4-d42c-48fe-8539-1770cd4965d8 · outbound

This paper cites A New Perspective on Shampoo's Preconditioner.

Low-rank Momentum Factorization for Memory Efficient Training A New Perspective on Shampoo's Preconditioner

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.206232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.206232Z digest=sha256:cba43413641abdf7578945142e11ca3fea99061f5285e461afcc055bd399ca74

Observation 1efd01c2-0826-46a9-bcfa-d986b1cb4467 · outbound

This paper cites Training language models to follow instructions with human feedback.

Low-rank Momentum Factorization for Memory Efficient Training Training language models to follow instructions with human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.208989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.208989Z digest=sha256:09939b22f5f6e8332e16948f539591f857f69348e9235d1b7b758ecb8b979e3b

Observation 6f316196-de61-474c-baa8-f2798e19a7d8 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

Low-rank Momentum Factorization for Memory Efficient Training The fineweb datasets: Decanting the web for the finest text data at scale

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.688519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.211977Z digest=sha256:35f99751639c01198a086c00adc1e6a07930a43a058d593ca502ae60294cd532

Observation af92c5c3-9423-4e42-906b-a077d1a73022 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Low-rank Momentum Factorization for Memory Efficient Training Zero: Memory optimizations toward training trillion parameter models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.214420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.214420Z digest=sha256:e377476e4413e64a7d1ebb32480d5055573e10e72248708f2f13d06ef663fc24

Observation 9adae0d8-c7be-4909-8e96-53e62d9dc309 · outbound

This paper cites Adarankgrad: Adaptive gradient-rank and moments for memory-efficient llms training and fine-tuning.

Low-rank Momentum Factorization for Memory Efficient Training Adarankgrad: Adaptive gradient-rank and moments for memory-efficient llms training and fine-tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.216796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.216796Z digest=sha256:55c26011b3ffe3bc8ec7fbfbb55860361902c7ad56ebfcfedab9f1778ce09649

Observation 7e249672-1596-48c3-943e-0e72ef748e99 · outbound

This paper cites Ldadam: Adaptive optimization from low-dimensional gradient statistics.

Low-rank Momentum Factorization for Memory Efficient Training Ldadam: Adaptive optimization from low-dimensional gradient statistics

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.676755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.219217Z digest=sha256:3513766a3872327d0022ac503b7651d7d51f956ca01721fdd5332bcc55153701

Observation b23d3601-7f59-4904-8d2e-0039c1fa3197 · outbound

This paper cites Gradient Multi-Normalization for Stateless and Scalable LLM Training.

Low-rank Momentum Factorization for Memory Efficient Training Gradient Multi-Normalization for Stateless and Scalable LLM Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.221296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.221296Z digest=sha256:0c07ed50a2788f121e9800db1738e096f0d446c4d610c2c71829000a1c94f435

Observation 34e4caf4-1165-4a4f-9045-929d96a595e8 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Low-rank Momentum Factorization for Memory Efficient Training Adafactor: Adaptive learning rates with sublinear memory cost

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.223758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.223758Z digest=sha256:5fa412aec6af4d6cc92931596f48d921f180e8da2a20463f5ce23dc03314ec4a

Observation 841a71e7-81fb-4fbb-85f4-c201c60929ff · outbound

This paper cites Tieleman and G.

Low-rank Momentum Factorization for Memory Efficient Training Tieleman and G

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.664744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.225792Z digest=sha256:68df0900ee86e52d69f84420541a0b8ac0b46a20c8ce2ecea94b6395dda03921

Observation 110f4a2b-637f-488c-ae65-3bfc5ac28b6f · outbound

This paper cites Practical low-rank communication compression in decentralized deep learning.

Low-rank Momentum Factorization for Memory Efficient Training Practical low-rank communication compression in decentralized deep learning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:35:40.657442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.227937Z digest=sha256:3824a2defbbd6a58b267d0b728b6a5da034087a5d979287435958def83907a75

Observation c37c9c3d-7ce6-4a10-95b5-e88c51857c64 · outbound

This paper cites SOAP: Improving and Stabilizing Shampoo using Adam.

Low-rank Momentum Factorization for Memory Efficient Training SOAP: Improving and Stabilizing Shampoo using Adam

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.230139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.230139Z digest=sha256:fbe4c512b66d64d392b204e160fa2e48cc88f31441d63507dfec17fe04d72cfc

Observation 47ef6f09-ff35-46fa-80a5-2adbad33c61f · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Low-rank Momentum Factorization for Memory Efficient Training GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.232517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.232517Z digest=sha256:70a7ddba508205cb1c443d8d007b2b96a7041e98039edee30a44d153f79c99f9

Observation 70c1579f-dfb6-489b-ae7b-dac080c96305 · outbound

This paper cites How far can camels go? exploring the state of instruction tuning on open resources.

Low-rank Momentum Factorization for Memory Efficient Training How far can camels go? exploring the state of instruction tuning on open resources

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.234934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.234934Z digest=sha256:dbc918570b827d48a8d7ce6f0931d6d5e6432a175cd878d7e7c6e6f99a03ff6d

Observation 0d750cd9-6d15-4d64-8b9f-3f5302b61e32 · outbound

This paper cites A Spectral Condition for Feature Learning.

Low-rank Momentum Factorization for Memory Efficient Training A Spectral Condition for Feature Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.237394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.237394Z digest=sha256:900437ee28b0f111931cfc6598f86eba2cf082900db909d704c452940c3fd02e

Observation 44fa1c87-6636-4c72-b2e2-b83704a4c2f4 · outbound

This paper cites Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models.

Low-rank Momentum Factorization for Memory Efficient Training Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.239930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.239930Z digest=sha256:6c1cba8e5dd56e2bf88c42a768457fbef6f0c1c3dbd10e8b3233963d57fa12da

Observation f90f451e-8e67-4718-ab89-7507bd97bf7f · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Low-rank Momentum Factorization for Memory Efficient Training AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.242074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.242074Z digest=sha256:cabd197095922bde34b11b58fadf82314f6a865d065b29da5bfca715227aca13

Observation ce95a6af-a70d-440f-91d7-1eabe494034a · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection.

Low-rank Momentum Factorization for Memory Efficient Training GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.244756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.244756Z digest=sha256:151bda626d3de154b78891ff535834585cd2edd80b85ecfb0066cb2d9542e767

Observation 0b28fa6a-0895-48f6-8cb4-0bdc270963a8 · outbound

This paper cites Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices.

Low-rank Momentum Factorization for Memory Efficient Training Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:35:40.307682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T18:35:40.247371Z digest=sha256:bb5aab96a0625c9e9a37c93848752abbb93ff5e37e2b9266143f31722e5a2cfc

Observation 008ee817-5a97-436a-97d1-a07d2e8019f7 · outbound

This paper cites APOLLO: SGD-like Memory, AdamW-level Performance.

Low-rank Momentum Factorization for Memory Efficient Training APOLLO: SGD-like Memory, AdamW-level Performance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.249822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.249822Z digest=sha256:12bf6787c4300522ca3ad88c1e6f0bb198e6b4f8295801fe2dd301bc2d3344f2

Observation 9b521a52-776b-4ec0-9b32-ea53ebece1f2 · outbound

This paper cites write newline.

Low-rank Momentum Factorization for Memory Efficient Training write newline

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:35:40.252196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:35:40.252196Z digest=sha256:b6a6adf16647610a89c758cf90db48ce13349c0a4f405d3be4988a1746b8c0c1

Pith citing papers

No inbound Pith citation observations are available.