Pith. sign in

Paper Citation Record · LEDGER

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam

As of 13 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.06464.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06464 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:10:16.631122Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e944ecc4-1240-4dbb-8a4b-37159c5297a8 · outbound

This paper cites Adam: A method for stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam: A method for stochastic optimization,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.336936Z digest=sha256:a975f5499ca8711fd29a93c0a9a040edeb42379feb0e9d0ae2e82354dfe2c80a

Observation e6e02f0c-d61c-4d6f-80e1-ea04028cd716 · outbound

This paper cites Attention is all you need,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Attention is all you need,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.191786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.455074Z digest=sha256:6c6265fc7ba1f1928e19c3a3d62429deae3170aea3b109d80725b7b80b8b51b5

Observation 130b3ed3-5f82-4fe6-8a9f-e2350eb1faee · outbound

This paper cites Language mod- els are few-shot learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language mod- els are few-shot learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.458605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.458605Z digest=sha256:4ddef490e61695f211fa51d9abf52717c0b3522a55fd5331dbd19d104d69d9e1

Observation 1a5ec753-fa43-4e58-b12d-202711d646bf · outbound

This paper cites Palm: Scal- ing language modeling with pathways,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Palm: Scal- ing language modeling with pathways,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.462151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.462151Z digest=sha256:c0d0f85cc4877da2ee6922a6ebe7a0dd7f1197431296df45d07b27e54de40f57

Observation 7091e28a-58fc-431a-b9f6-edc9e96e4b74 · outbound

This paper cites The Llama 3 Herd of Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.465335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.465335Z digest=sha256:2ba79671609dd9a2720324c1cd88b51cf03c6232f211c44122b49781f19b8476

Observation a7f18002-e37c-421a-998c-93439602f0dd · outbound

This paper cites DeepSeek-V3 Technical Report.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.468721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.468721Z digest=sha256:afb4f70ed5112fb82ce29ed52d8359f897c6d08e82e196679d6635a3c4d7aa82

Observation c344c854-4c34-4e60-8ee5-4d7fa5be9619 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Learning transferable visual models from natural language supervision,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.471944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.471944Z digest=sha256:f124659d307708e15d418ecf474d774997189eb02422d7f04cc04bc9c37eefd6

Observation 99a15b1d-518d-42bd-9632-e3469bc87f4b · outbound

This paper cites Segment Anything.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.474845Z digest=sha256:e4240a8fd0fe4f20028faa10be69c010964b6e1831cb62e92e8cf77f70ad494a

Observation 8c5b1c58-700a-4826-b81c-372942e56358 · outbound

This paper cites A convnet for the 2020s,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A convnet for the 2020s,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.478341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.478341Z digest=sha256:b90ffe0fe8e71a7b254610889eaeb592f93de0b092d3feea2a13c4d557e48527

Observation 7ea6ca0d-f1a1-4b7c-acfd-602228e86852 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoen- coders,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convnext v2: Co-designing and scaling convnets with masked autoen- coders,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.481830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.481830Z digest=sha256:3de9d023b6d3fc8bceccb0daf496593d8e0de849674321470d8cc3d69ac8102a

Observation c1825759-aff8-4476-ad47-2cf31d09d429 · outbound

This paper cites Imagenet classification with deep convolutional neural networks,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Imagenet classification with deep convolutional neural networks,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.485548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.485548Z digest=sha256:591d08439813cdb01a0a8c8bd3f9c034a36d51c1a2a6c3165f29a0f67b2eaebd

Observation c49b5738-5629-45bf-9d39-f9530cbe0292 · outbound

This paper cites Deep residual learning for image recognition,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Deep residual learning for image recognition,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.488397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.488397Z digest=sha256:8df4bd53d9cfd24a897417882f97cd523bafc41d02c11f2ab07ec51f5635d113

Observation 95fdae10-305a-4e29-92b0-50fbe1e9c141 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Symbolic Discovery of Optimization Algorithms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.490853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.490853Z digest=sha256:0c50d74ef02673e43ab98939dcd9af57f78782d9c548f4dd5aaf2da7f08d0e0b

Observation 387290f7-6e61-4fbb-9549-3f7e6b456b20 · outbound

This paper cites Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.493852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.493852Z digest=sha256:70c01c2fde85482285577adcf848d0595272bc5b895bd87e800ed0b7cdd6a232

Observation 64c549bf-2a0d-45a2-921e-0411c3fa38b6 · outbound

This paper cites Adaptive subgradient methods for online learning and stochastic optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adaptive subgradient methods for online learning and stochastic optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.497361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.497361Z digest=sha256:99afc5feb408cb43373250aaa31641570a3032ef4d14f2fd45c8acac5efc635c

Observation 9034aef6-356b-4b25-bf4d-b3a7e75faeb4 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Neural networks for machine learning lecture 6a overview of mini-batch gradient descent,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.137025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.499783Z digest=sha256:2089dd0bc549b4af32b333a7e8a8a1bdb8737add9058eab75134bdb4046b8e1c

Observation 8ba9cd53-5f3c-4235-a11e-ac7c73ec22b0 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam ADADELTA: An Adaptive Learning Rate Method

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.502866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.502866Z digest=sha256:71c8e3936c43da2c829c79b7cb727d5176c9daeeaade3755c72388c0e3bf595d

Observation 8fad584c-c762-4759-acd9-db97b754c220 · outbound

This paper cites Incorporating nesterov momentum into adam,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Incorporating nesterov momentum into adam,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.124513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.506673Z digest=sha256:fc832da3f6e36537a9dbff616a2b82f776ae3211ae7dbcc27261907f89e3c0ef

Observation b424d17a-bdd7-4d6d-84f5-938ee379f839 · outbound

This paper cites On the convergence of adam and beyond,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the convergence of adam and beyond,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.509896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.509896Z digest=sha256:47884cb83ebb2d6cd0f5e77f19790931e46bfd7724753d50dadcb400d7b602fb

Observation 6bf20935-3052-4582-a149-eba073699fa4 · outbound

This paper cites Decoupled Weight Decay Regularization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Decoupled Weight Decay Regularization

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.513195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.513195Z digest=sha256:b4a641ebab6a98796fb8d51caa34e6137139d80d8b600d00f94fd506e991419e

Observation 882ac5b3-d99a-4380-afc0-42040410dad4 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adabelief optimizer: Adapting stepsizes by the belief in observed gradients,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.109013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.516431Z digest=sha256:b36f499a0b2d54c1c01a555c60408a8f4c0c621d9cbc153072d98e804c945856

Observation f6cc7300-be44-4848-8327-9827f0c8109d · outbound

This paper cites Adafactor: Adaptive learning rates with sub- linear memory cost,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adafactor: Adaptive learning rates with sub- linear memory cost,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.097992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.518984Z digest=sha256:6f423aeafd2e34be23fa099ee6554f7245051a8b6d5a5ffcbbfbdeead76ab392

Observation d39444e8-0734-4f72-98ad-07c53c256862 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.088592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.521637Z digest=sha256:1a10852d4b820ab35f929984512292fdd41ccdee85767721436960fbc5dc748b

Observation ebea6fa2-4ce9-47dd-be8e-5ba659eea44f · outbound

This paper cites Signsgd: Compressed optimisation for non-convex problems,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signsgd: Compressed optimisation for non-convex problems,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.078021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.524327Z digest=sha256:c4c009ba1957acb2e420757783b97c8de88926649064782d83aeb09c112fff22

Observation e926c48a-d74f-4e96-8149-97437b6e9db3 · outbound

This paper cites Momentum ensures convergence of SIGNSGD under weaker assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Momentum ensures convergence of SIGNSGD under weaker assumptions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.067344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.526948Z digest=sha256:335a27d53d1e28d62f059a25551d27ed058a086ba8322117c8474b39cf91a476

Observation 8a4e6230-10b8-41d4-9d9c-120d5a4c958b · outbound

This paper cites Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.530074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.530074Z digest=sha256:f676b3602e5852ba5597bd5e19d034f109b449c0ca83482c8459fdfdd72e3429

Observation 82b4b94f-b472-4f8f-addb-70900f8590bd · outbound

This paper cites A direct adaptive method for faster backpropagation learning: The rprop algorithm,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A direct adaptive method for faster backpropagation learning: The rprop algorithm,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.055958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.533409Z digest=sha256:046eee46f84e6a7d7c226d7f47b46bfffa44ddc58c0527f8a16fabeaf6aa8654

Observation 8824a63e-45a6-459f-97fd-1ec59887e172 · outbound

This paper cites 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.045610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.535835Z digest=sha256:624801439d3c0a820600085dc1db29d5f846760c63de0549bc5fee343a591810

Observation afb2ed8b-b8f9-47ed-bb3d-4140897acc52 · outbound

This paper cites Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lion Secretly Solves Constrained Optimization: As Lyapunov Predicts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.538230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.538230Z digest=sha256:2bf517871f9f72a1869736e0363cf615c20cc45f19cb098ed99f5a2f6194ac65

Observation 6036972d-a558-4688-963c-b919c3331931 · outbound

This paper cites Dissecting Adam: The sign, magnitude and variance of stochastic gradients,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Dissecting Adam: The sign, magnitude and variance of stochastic gradients,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.035995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.541329Z digest=sha256:87c31bb2e0376ead33fa22bf642ef28b4e7f06a098fc3ed43c1b204c29162e84

Observation 03e8a653-b280-4c7a-9d79-158e223e8596 · outbound

This paper cites Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.543726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.543726Z digest=sha256:66ca6c1f63a8b37ab6b24aca6b2ba44c674d69047c2d4eb1b6ed0cbd478cae56

Observation 0b7b1394-3226-48cb-a8e0-b8b9b09a1f9a · outbound

This paper cites A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A method of solving a convex programming problem with convergence rate o\bigl(kˆ2\bigr),

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.025018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.546651Z digest=sha256:1b90a2951052d1c476265615104c0b00ce96ed6e58029e76e5914fe3f2046185

Observation 9d6fc65c-1ed3-4412-add0-8e36cd225a20 · outbound

This paper cites Springer Science & Business Media, 2013, vol.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Springer Science & Business Media, 2013, vol

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.015378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.549317Z digest=sha256:37b516cebfc3ef2b35454d4ab13d5751261315b67eff18c5eb768a6c794ff453

Observation 1f9894cf-5e9c-4c60-9bb3-92b9b905aecf · outbound

This paper cites Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adan: Adaptive nesterov momentum algorithm for faster optimizing deep models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:17.003645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.551946Z digest=sha256:69ddbd1c5a2308957f277c29499587ad5f7ae5281263010a25bc77af8cf93682

Observation 17c3485f-f796-4d31-97c8-1e15f18f969d · outbound

This paper cites Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Win: Weight-decay-integrated nesterov acceleration for adaptive gradient algorithms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.991917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.555136Z digest=sha256:13b8620dcf51f4af54854fde255f6171a227520ad8fde0dc6b2acd4e20c22b61

Observation 7e6bf4ef-121e-415e-a5f8-4efda35ca0fd · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam GLM-130B: An Open Bilingual Pre-trained Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.558707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.558707Z digest=sha256:55057ac9f33d72489621a5122e4a99d0010c36542c6b57b99e60df0637744dfe

Observation fef1eb67-c197-4ee5-84c5-067f88733fc4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.561798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.561798Z digest=sha256:f7934e96914caeee4c259c54217dca543ec554a6b1b00ff906f686d35993dd74

Observation 4389efc2-d313-4583-8363-821f45cddab5 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Baichuan 2: Open Large-scale Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.564494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.564494Z digest=sha256:e135cbc3104496c8195f3aab30b6e4a843b38349a550b58eb3834223376a63fc

Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · outbound

This paper cites What Language Model to Train if You Have One Million GPU Hours?.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.567629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.567629Z digest=sha256:e7152bf19851c1aae188bb1c28e77e00bbff14ca217615234dcef04c26d8986b

Observation 6c436ba2-b562-42f2-af4e-208e2df607d3 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Theory on Adam Instability in Large-Scale Machine Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.570768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.570768Z digest=sha256:e4319df0ac98e68bdbc4b587251ac0461315c9064c8ee4001cbfa353e44c9bb1

Observation 4a9a43cc-6d72-4ab3-a56a-d57d4b7aaa05 · outbound

This paper cites A Mean Field Theory of Batch Normalization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Mean Field Theory of Batch Normalization

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:10:16.753519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.573719Z digest=sha256:9279a87c730fb358099df9ae8df6c58ec6124be9cb6a68b6795330515d0a3f5c

Observation 81412f7d-d3c9-443f-b776-78d06309f9d7 · outbound

This paper cites Understanding the Difficulty of Training Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Understanding the Difficulty of Training Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.577022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.577022Z digest=sha256:46d51b53053ca97f6c3168766ab609cc8f0874eb50b55c7e81f5b35b9780a5f7

Observation ae162317-639b-4c10-8aa0-87babd01ece8 · outbound

This paper cites On layer normalization in the trans- former architecture,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On layer normalization in the trans- former architecture,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.981559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.580127Z digest=sha256:ba18192b260251ad2d3acd4bf9c6920858c2983164068f4764ebfc5d8a27c9f0

Observation 2652fda8-ca7c-4f19-9ce7-dd05b222d562 · outbound

This paper cites The lipschitz constant of self- attention,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam The lipschitz constant of self- attention,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.969465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.583966Z digest=sha256:b7198625e75f1f2ee3decc77cf089f4f9d504153754cbe7c65dfaad848482bb6

Observation 8dfbbdde-7c7e-47bd-8fd7-2af0221cade8 · outbound

This paper cites LipsFormer: Introducing Lipschitz Continuity to Vision Transformers.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam LipsFormer: Introducing Lipschitz Continuity to Vision Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.586759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.586759Z digest=sha256:b91fd89e1085f92638a5452027f807e61b2a1072648fdef4306d79bd6e54bcd8

Observation 8c0381e7-1410-476e-9265-b2dbab2a4363 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Signal propagation in transformers: Theoretical perspectives and the role of rank collapse,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.955532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.589251Z digest=sha256:1d69f2848356e1bf5cfe4b0c920761a509e44b6c43ebda70a36ea721bbe152b2

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · outbound

This paper cites On the Variance of the Adaptive Learning Rate and Beyond.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:7d8b632fd917c5b00d0957f36d71f40df3d8cb2634e896a575f778fb8e3ee235

Observation b1702c3e-1070-4189-83a0-ceba8119b40e · outbound

This paper cites Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.595651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.595651Z digest=sha256:7828be14adf308a472d726295ed138b407e714cb1a40830a90bd04b19b1aafb8

Observation ee2e7d3e-cdc5-4a94-90b3-ffc1a0050369 · outbound

This paper cites Loss Spike in Training Neural Networks.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Loss Spike in Training Neural Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.598782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.598782Z digest=sha256:460966986acaa5fa72c8206b5d85ebb86a076f444d46a01c01f592593f74cbf5

Observation 05f8d309-aa58-4cf9-ab76-89e4d31a7e70 · outbound

This paper cites Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.945621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.601925Z digest=sha256:2d36272160af30d3eb2ae516e208e1aa37fb5f41cfbb74719f111b3c3b12915b

Observation 0665f40e-517a-40fa-9b6b-57ba587a69f1 · outbound

This paper cites On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.605575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.605575Z digest=sha256:db51abd3a16623c715b6bdfe4dbe659ca661b8a7c8242f021fbdd2e97ce5102f

Observation aa8ccb70-02cc-4385-9382-1b1296c7d6a0 · outbound

This paper cites A Simple Convergence Proof of Adam and Adagrad.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A Simple Convergence Proof of Adam and Adagrad

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.608689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.608689Z digest=sha256:3b870f19e69a7acd5a1418ca744507218841d1341f7d31981188d27ce4a1f1df

Observation 2da32d6b-006e-4c34-9b9f-d5b321245e07 · outbound

This paper cites Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence guarantees for RMSProp and ADAM in non-convex optimization and an empirical comparison to Nesterov acceleration

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.611740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.611740Z digest=sha256:8bcfa94ca490b136ef62a353919f2ab6e4f198eaa857a761a84b7fe847f6afe9

Observation c793896d-c0f0-4540-b1f4-ba65969027a9 · outbound

This paper cites Adam can converge without any modification on update rules,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Adam can converge without any modification on update rules,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.935092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.614982Z digest=sha256:6a9194903d22947102e83a9b45da4fdd28214efc5e6bcbe385aed42dc02e762a

Observation edb696f8-338f-4603-bc9b-65aa87994a2a · outbound

This paper cites Convergence of adam under relaxed assumptions,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Convergence of adam under relaxed assumptions,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.925114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.617374Z digest=sha256:58695b9180e9045936b7cf4a1eaea3728288e2bc2be8512d952f18334b44bc26

Observation aef5e537-b67a-4191-baf8-44b8ee68e314 · outbound

This paper cites On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On Convergence of Adam for Stochastic Optimization under Relaxed Assumptions

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.619880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.619880Z digest=sha256:09fc3520f0bb85074d2d4041dcbe3ddee894c560e57284ddf2aecb701e12e3fd

Observation d412bbce-b75e-4338-85a1-8d95ec0c4527 · outbound

This paper cites Lower bounds for non-convex stochastic optimization,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Lower bounds for non-convex stochastic optimization,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.915141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.622696Z digest=sha256:e9f4a39d46b5f8ab34cb2dd5737d0fed1f43c46409cf7990dfb2d41e04cdbe6d

Observation 8aa29741-fb40-45e8-9b6f-6eb49c4d653a · outbound

This paper cites A stochastic approximation method,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam A stochastic approximation method,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.625262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.625262Z digest=sha256:c8d4f39a6e562d2893f5fbf908476d1a3ed371eb4b5a316c39a4bfa647f6e043

Observation 37318d86-763a-4520-a46e-43cf4fa59af7 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.628081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.628081Z digest=sha256:edec0e289430fcb396204461524f2ba6297f7367493e2ef2a8a2f62042a37e8e

Observation ef273616-c7fa-4b26-957c-cf3618c69d17 · outbound

This paper cites Language models are unsupervised multitask learners,.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam Language models are unsupervised multitask learners,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:10:16.898796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:10:16.631122Z digest=sha256:bab87abe68476445a4710b623513492d1635e24df2805317c0209b08ae554bf0

Pith citing papers

No inbound Pith citation observations are available.