Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:43:02.115828Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 20 inbound Pith citation observations for arXiv:2411.10438.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:43:02.115828Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T17:20:51.021889Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:18:56.218978Z
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 32587236-7371-4bc6-8cae-c4c9317b1f20 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Yuan, Y
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c23494a3-6110-4f42-b1fb-ef0b3e02ddac · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Scalable Second Order Optimization for Deep Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 489e2f04-193e-4726-9e4f-70c67a967204 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models C., Foster, D
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3cade9e-05c8-4b36-a45c-aa930683b285 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Glynn, P
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75163b56-d8a4-4c74-b8e2-2ff5c75ab8cd · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Old Optimizer, New Norm: An Anthology
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34e7797b-eb8b-4859-a7d8-4b4e61f5733d · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models L., Gao, J., and Choi, Y
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0af577b4-de4b-447b-a9c0-132775e4e360 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Language Models are Few-Shot Learners
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18cd5d9d-fd4e-4d79-9728-7799a3146168 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic spectral descent for restricted boltzmann machines
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f271a367-fd5f-4c90-be6b-be2c2a5d4c4f · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic spectral descent for discrete graphical models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a327634-7f86-4bdd-abcb-41771812372d · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a69e144-f8f0-4d1e-a106-01c139073854 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Symbolic discovery of optimization algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0da958-9185-4240-ab0b-e12a0a20c986 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models W., Sutton, C., Gehrmann, S., et al
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dfdbe6d-757e-4d09-aded-a15b2248b3ce · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Boolq: Exploring the surprising difficulty of natural yes/no questions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6986463-4542-4b2e-9862-06881e5d84fe · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Orabona, F
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11856636-9353-4a77-a17b-83886cbbd8a4 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Bottou, L
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383cc93d-6e19-4562-b8f7-bf25784f1e2f · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473f2f71-6b58-4f5c-baaf-687600880774 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models The Road Less Scheduled
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311715d3-4e4b-4421-a965-f5f1ccda8f22 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic variance-reduced newton: Accelerating finite-sum minimization with large batches
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ce96ff-9c55-4f14-b1e6-4d48be2963c5 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c084cd7b-690e-4278-8d2c-dd4f5fcfbf75 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models The Llama 3 Herd of Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f877d8-2988-4dd1-9e2e-4e596f507325 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adaptive subgradient methods for online learning and stochastic optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7450097-b671-4ddc-96e0-bb465215a037 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models J., Lin, Z., and Zhang, T
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2867a9b-db15-4931-a3b8-dbaa05ed9173 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Promise: Preconditioned stochastic optimization methods by incorporating scalable curvature estimates
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 28f611b6-5a1a-40d5-8130-e98ebebb5e23 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A framework for few-shot language model evaluation, 07 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 332c3eeb-a786-460f-91d7-5586729a2c81 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Lan, G
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a9c5b38-f97a-4ac3-ad29-1b5a50e3f11e · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Openwebtext corpus
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0090491-0933-4f80-86d8-14c98d7ad1e8 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Graves, A
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99be9de-a6d6-4e29-ab8a-075002b83bb4 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Shampoo: Preconditioned stochastic tensor optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 837873a5-4868-487b-9676-6051d1199ade · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Beyond convexity: Stochastic quasi-convex optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eb16ae7-421a-4793-818f-b9012f330f17 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Deep residual learning for image recognition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365be778-5806-4500-81d9-2650b84c20ee · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Measuring massive multitask language understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7c0c1a4f-42c9-4135-80fe-d10eab8c40bc · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a44fc610-5fee-424c-903f-eeff8f848111 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f860c1b5-000b-43e3-90d6-ad060a880bd7 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Super-adam: faster and universal framework of adaptive gradients
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 095ffd0c-484f-441a-958f-0ccf4ac6f5a4 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Zhang, T
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6ac791e-5db5-4467-b101-9f19dc00ae43 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Muon: An optimizer for hidden layers in neural networks, 2024
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22458ebd-08da-48ea-9dec-10abb76e7a98 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d99fb412-a9cf-421e-93ec-20e084cb9f06 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7db134d0-4a1a-44a0-8ce9-914866c1da56 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models T., and Cevher, V
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e0978da1-f323-430f-9617-c1a31cc3e799 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579e6a5e-ef91-48af-ae00-7b31d2e79891 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Learning multiple layers of features from tiny images
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce1d6db-0167-4dcb-b47f-59bb14ca9253 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models On the computation of the matrix k-th root
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6b6aa78d-cce2-4b06-9429-5f6a6637dc43 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models S., Moeller, T
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 564e5370-2bc4-40a4-bf8c-75b95e12b00b · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Gradient-based learning applied to document recognition
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29b4ed67-9353-458f-8251-93feac329703 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Storm+: Fully adaptive sgd with recursive momentum for nonconvex optimization
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 49389a72-448f-4e61-b719-5bd1f2b428dc · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Smoothness and Adaptivity in Nonlinear Optimization for Machine Learning Applications
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 28f3b14e-136d-4d1c-a3ae-94871d9d4cae · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Convergence of adam under relaxed assumptions
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 476a60c7-8422-4d96-af61-7dabfc07b7ce · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6704f74b-430b-4ac6-aabf-bd407257fa08 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 929c6321-4e29-4615-9ed1-0bff2eb4f95e · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adam$^+$: A Stochastic Method with Adaptive Variance Reduction
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98bfc228-49f1-4c1d-90ea-7fcf22fbdc3f · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Hutter, F
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7feaa9b5-925b-4996-a482-6bb4a927d258 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Fineweb-edu: the finest collection of educational content, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d21d3e8-2f8b-4f1a-b690-c0ab786f0e98 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Grosse, R
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7297947-5172-401e-93a6-a3807bcf4d5b · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Tropp, J
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c3cd4b70-9543-4b57-8d1f-e0488227881d · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models An Empirical Model of Large-Batch Training
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c25452-1b46-4637-9aff-884ce3bff0e5 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adaptive Bound Optimization for Online Convex Optimization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68ea52d1-0b27-4be9-be39-b1dd82794446 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Can a suit of armor conduct electricity? A new dataset for open book question answering
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af18d8a0-4378-440b-9257-b35ebf83d8c2 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A New Perspective on Shampoo's Preconditioner
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85cfd310-4490-42d3-a135-aa6de65968f7 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A method for solving the convex programming problem with convergence rate o(1/k^2)
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2845d882-48fc-4566-8ce3-ed4e2923103f · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Introductory lectures on convex optimization: A basic course, volume 87
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4ff75ab9-714f-4a67-b7f3-503bb0df5637 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models M., Liu, J., Scheinberg, K., and Tak \'a c , M
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 3f2a9ec6-1ccb-468c-bb19-2d2e95021808 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic Recursive Gradient Algorithm for Nonconvex Optimization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16aaf30a-2629-4f06-990b-25a4de6a3f11 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Language models are unsupervised multitask learners
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28c67352-ff2d-457b-b4a0-f5a51d15074e · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models J., Hefny, A., Sra, S., Poczos, B., and Smola, A
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 64e703b0-6003-4ae7-ac02-cbd442a56063 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models On the Convergence of Adam and Beyond
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43854652-eacd-442b-9a6d-4ecb8e5021ab · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Braun, H
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eea572f5-4d30-4393-b5c5-794629c95ca1 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A stochastic gradient method with an exponential convergence \_rate for finite training sets
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7709864d-3762-4722-a6f1-245e4196a9ce · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models L., Bhagavatula, C., and Choi, Y
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ed306893-a7be-47ee-b441-72d8d69ab381 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models English Conversational Telephone Speech Recognition by Humans and Machines
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4caae8a-a9b1-4331-a710-0b162f46e2bc · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Smola, A
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85de6cc9-7d30-49f6-baaa-a23a13a6ab46 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Iterative berechnung der reziproken matrix
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8bbaeb57-1417-41df-9d90-d1be52929e5c · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Zhang, T
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c262343c-682c-4473-8583-183ee134a835 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models and Stern, M
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 840b886c-4ca4-4976-9d25-026b41b288c8 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A Distributed Data-Parallel PyTorch Implementation of the Distributed Shampoo Optimizer for Training Neural Networks At-Scale
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d974669-e5d9-4404-b931-c8068638f8c8 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Dropout: a simple way to prevent neural networks from overfitting
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa80dc0-fdc0-4942-8de4-99571633fcc4 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0024312a-7550-4bd0-ada2-61fad6a32166 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Attention is all you need
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a752e3d-2adc-4526-a48a-4cfed46f2a60 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models SOAP: Improving and Stabilizing Shampoo using Adam
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 012af07a-9e4d-4b3d-836a-77b90ec463da · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Spiderboost and momentum: Faster variance reduction algorithms
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b3dc8bbf-c6f3-4315-8dad-6e039b3dea6a · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be36d960-c952-4b93-986c-5cdb9aa09499 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3a7227-59be-4ac9-a24e-864cf24dc490 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Quick and (not so) dirty: Unsupervised selection of justification sentences for multi-hop question answering
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c21778-335f-47a5-9962-26cecaa098da · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models A Coefficient Makes SVRG Effective
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a3c723bf-65ea-48d1-8cee-0424688e03d4 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6dde87-279b-4b4a-b5b7-79bfaf0dd268 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models ADADELTA: An Adaptive Learning Rate Method
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b4d049-9e91-4b71-bc3f-4cd504d7f7fc · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd7335a-5e17-4e65-8923-dc475745f55a · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models OPT: Open Pre-trained Transformer Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107f6011-4904-4dc7-b21d-236b708babfd · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Adam can converge without any modification on update rules
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c62ad04f-5834-4868-a3e1-459c08c6a12a · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Deconstructing What Makes a Good Optimizer for Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e1daf55-ba58-4897-a71d-419a408daae3 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Stochastic nested variance reduction for nonconvex optimization
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 54a450a1-df44-4fc5-b952-58b06cdd5cb6 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Towards understanding convergence and generalization of adamw
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 5465fbce-eb1c-4d54-a499-e15f24becaf1 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models C., Dvornek, N., Papademetris, X., and Duncan, J
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad9cbca0-2e2f-4589-ae04-316cdc5154a9 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models @esa (Ref
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209059a6-ecd1-410f-926f-bbbd5bc28d06 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bafc3be5-4fca-4f01-9a61-c7cbb5b25d94 · outbound
MARS: Unleashing the Power of Variance Reduction for Training Large Models Unresolved cited work
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82111d87-d0d8-4a25-a1e8-c4c8250ce563 · inbound
Physics of Skill Learning MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c46eee-9c29-44dd-865b-18abd9153a49 · inbound
Training Deep Learning Models with Norm-Constrained LMOs MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 218
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 073a7acf-1bb7-471c-953e-9a2431184c4a · inbound
Continuous Cardiac Arrest Prediction in ICU using PPG Foundation Model MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b38aa2-0c1f-4d77-81d3-f81653fca415 · inbound
Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a11c4c39-2f8d-4a92-a779-1caf95de60bb · inbound
DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96305ccd-7657-4bf1-b514-55b65d9d72a7 · inbound
On the Convergence of Muon and Beyond MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 837ee271-41b3-4b54-aa56-2611a1f51e82 · inbound
Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9d406b4a-3007-4676-bb2c-6be5b7626b11 · inbound
Optimal Sample Complexity for Single Time-Scale Actor-Critic with Momentum MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a767a7f1-1c97-42be-904b-2253a4f48c5c · inbound
HTMuon: Improving Muon via Heavy-Tailed Spectral Correction MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8f4fb87c-5a45-480f-a4f1-3e547fba4e02 · inbound
CLion: Efficient Cautious Lion Optimizer with Enhanced Generalization MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 020babd5-7115-47a9-b824-8019276a58f3 · inbound
Personalized Federated Learning for Gradient Alignment MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 538e25c1-2441-4600-8506-41ad000ad01b · inbound
Optimizer-Model Consistency: Full Finetuning with the Same Optimizer as Pretraining Forgets Less MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1c50ea8b-05e4-4edd-a635-6d520343531d · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e4e9ed29-c096-40da-a023-b20670e9989e · inbound
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8c1a22ae-c71f-4321-af06-e8f23ac2b9cf · inbound
Why SGD is not Brownian Motion: A New Perspective on Stochastic Dynamics MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 158
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b4b54b21-6276-4641-8006-bcd108a796c8 · inbound
MGUP: A Momentum-Gradient Alignment Update Policy for Stochastic Optimization MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 72e002c3-65e5-4a39-875c-59bb0765fa72 · inbound
One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 92f28a8e-159e-4574-89fd-27f8a1ea7be4 · inbound
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 128
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0f53ac-6b1c-47ec-96bf-825f239333de · inbound
Muse: Representation Geometry of Muon Beyond Normalized Momentum MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f32997d-8304-4d9a-9fa0-e4c2397f32b3 · inbound
OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining MARS: Unleashing the Power of Variance Reduction for Training Large Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.