Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.06464.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e944ecc4-1240-4dbb-8a4b-37159c5297a8 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam: A method for stochastic optimization,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6e02f0c-d61c-4d6f-80e1-ea04028cd716 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Attention is all you need,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 130b3ed3-5f82-4fe6-8a9f-e2350eb1faee · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language mod- els are few-shot learners,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5ec753-fa43-4e58-b12d-202711d646bf · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Palm: Scal- ing language modeling with pathways,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7091e28a-58fc-431a-b9f6-edc9e96e4b74 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The Llama 3 Herd of Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f18002-e37c-421a-998c-93439602f0dd · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam DeepSeek-V3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c344c854-4c34-4e60-8ee5-4d7fa5be9619 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Learning transferable visual models from natural language supervision,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99a15b1d-518d-42bd-9632-e3469bc87f4b · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Segment Anything
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5b1c58-700a-4826-b81c-372942e56358 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A convnet for the 2020s,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea6ca0d-f1a1-4b7c-acfd-602228e86852 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convnext v2: Co-designing and scaling convnets with masked autoen- coders,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1825759-aff8-4476-ad47-2cf31d09d429 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Imagenet classification with deep convolutional neural networks,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c49b5738-5629-45bf-9d39-f9530cbe0292 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Deep residual learning for image recognition,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95fdae10-305a-4e29-92b0-50fbe1e9c141 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Symbolic Discovery of Optimization Algorithms
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 387290f7-6e61-4fbb-9549-3f7e6b456b20 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c549bf-2a0d-45a2-921e-0411c3fa38b6 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adaptive subgradient methods for online learning and stochastic optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9034aef6-356b-4b25-bf4d-b3a7e75faeb4 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ba9cd53-5f3c-4235-a11e-ac7c73ec22b0 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam ADADELTA: An Adaptive Learning Rate Method
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fad584c-c762-4759-acd9-db97b754c220 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Incorporating nesterov momentum into adam,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b424d17a-bdd7-4d6d-84f5-938ee379f839 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the convergence of adam and beyond,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf20935-3052-4582-a149-eba073699fa4 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Decoupled Weight Decay Regularization
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 882ac5b3-d99a-4380-afc0-42040410dad4 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6cc7300-be44-4848-8327-9827f0c8109d · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adafactor: Adaptive learning rates with sub- linear memory cost,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d39444e8-0734-4f72-98ad-07c53c256862 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebea6fa2-4ce9-47dd-be8e-5ba659eea44f · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signsgd: Compressed optimisation for non-convex problems,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e926c48a-d74f-4e96-8149-97437b6e9db3 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Momentum ensures convergence of SIGNSGD under weaker assumptions,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b4b94f-b472-4f8f-addb-70900f8590bd · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A direct adaptive method for faster backpropagation learning: The rprop algorithm,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8824a63e-45a6-459f-97fd-1ec59887e172 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afb2ed8b-b8f9-47ed-bb3d-4140897acc52 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6036972d-a558-4688-963c-b919c3331931 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Dissecting Adam: The sign, magnitude and variance of stochastic gradients,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03e8a653-b280-4c7a-9d79-158e223e8596 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7b1394-3226-48cb-a8e0-b8b9b09a1f9a · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d6fc65c-1ed3-4412-add0-8e36cd225a20 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Springer Science & Business Media, 2013, vol
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f9894cf-5e9c-4c60-9bb3-92b9b905aecf · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17c3485f-f796-4d31-97c8-1e15f18f969d · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e6bf4ef-121e-415e-a5f8-4efda35ca0fd · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam GLM-130B: An Open Bilingual Pre-trained Model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fef1eb67-c197-4ee5-84c5-067f88733fc4 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4389efc2-d313-4583-8363-821f45cddab5 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Baichuan 2: Open Large-scale Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c436ba2-b562-42f2-af4e-208e2df607d3 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Theory on Adam Instability in Large-Scale Machine Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9a43cc-6d72-4ab3-a56a-d57d4b7aaa05 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Mean Field Theory of Batch Normalization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81412f7d-d3c9-443f-b776-78d06309f9d7 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Understanding the Difficulty of Training Transformers
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae162317-639b-4c10-8aa0-87babd01ece8 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On layer normalization in the trans- former architecture,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2652fda8-ca7c-4f19-9ce7-dd05b222d562 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The lipschitz constant of self- attention,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8dfbbdde-7c7e-47bd-8fd7-2af0221cade8 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LipsFormer: Introducing Lipschitz Continuity to Vision Transformers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c0381e7-1410-476e-9265-b2dbab2a4363 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1702c3e-1070-4189-83a0-ceba8119b40e · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2e7d3e-cdc5-4a94-90b3-ffc1a0050369 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Loss Spike in Training Neural Networks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f8d309-aa58-4cf9-ab76-89e4d31a7e70 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0665f40e-517a-40fa-9b6b-57ba587a69f1 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa8ccb70-02cc-4385-9382-1b1296c7d6a0 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Simple Convergence Proof of Adam and Adagrad
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da32d6b-006e-4c34-9b9f-d5b321245e07 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c793896d-c0f0-4540-b1f4-ba65969027a9 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam can converge without any modification on update rules,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edb696f8-338f-4603-bc9b-65aa87994a2a · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence of adam under relaxed assumptions,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aef5e537-b67a-4191-baf8-44b8ee68e314 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d412bbce-b75e-4338-85a1-8d95ec0c4527 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lower bounds for non-convex stochastic optimization,
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8aa29741-fb40-45e8-9b6f-6eb49c4d653a · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A stochastic approximation method,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37318d86-763a-4520-a46e-43cf4fa59af7 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef273616-c7fa-4b26-957c-cf3618c69d17 · outbound
SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language models are unsupervised multitask learners,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.