Pith. sign in

Paper Citation Record · LEDGER

Insights from Gradient Dynamics: Gradient Autoscaled Normalization

As of 18 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2509.03677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03677 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:51:39.015136Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved13
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e8a9f65-878d-44ad-9fdb-c2c37f4b8f47 · outbound

This paper cites Layer Normalization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.901558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.901558Z digest=sha256:f4f94eea73e37d0a9f46287021e9e596f0612ad1e0e22be506eeb947b8f7eaff

Observation 1407376e-2cff-4088-9ec5-946660a7db70 · outbound

This paper cites Large-scale machine learning with stochastic gradient descent.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Large-scale machine learning with stochastic gradient descent

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.906194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.906194Z digest=sha256:1b2bffefd0e34dbff172d5a7441d6e1ab6cd7cce73a95df3e82954fb44423796

Observation 1cca347b-b28b-4c88-b834-82b04607df9a · outbound

This paper cites The tradeoffs of large scale learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization The tradeoffs of large scale learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.384027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.909953Z digest=sha256:fb3146161b5f7650f66b642b10f9c87af583c6ed5c223528266bca06d33ad8ac

Observation 5eb220e4-b18d-4c94-87c2-831deb57219a · outbound

This paper cites Curtis, and Jorge Nocedal.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Curtis, and Jorge Nocedal

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.913948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.913948Z digest=sha256:f7656854f9a01cba6647d85cb6ea2f815793dfaedfa191d24dc7b819fd0bfa0b

Observation 508674f8-9780-4928-8f92-23f6ddc6c4fd · outbound

This paper cites Entropy-sgd: biasing gradient descent into wide valleys.Journal of Statistical Mechanics: Theory and Experiment, 2019 (12):124018, 2019.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Entropy-sgd: biasing gradient descent into wide valleys.Journal of Statistical Mechanics: Theory and Experiment, 2019 (12):124018, 2019

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.917822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.917822Z digest=sha256:cdce6b4573ee8da87da77c4be63e984cd29687423fb7559a25f5ad25a52efa85

Observation 50dd9eb9-c3a2-4dea-bfc4-8127695d03bd · outbound

This paper cites Gradnorm: Gra- dient normalization for adaptive loss balancing in deep multitask networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradnorm: Gra- dient normalization for adaptive loss balancing in deep multitask networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.372116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.921632Z digest=sha256:aad2eaf9778f084b5937bab66a61516139d2c575b7eeabf5f7b69e092747d3be

Observation daf8b0b3-9c90-4ece-9d20-2d14d4ef2401 · outbound

This paper cites A Study of Gradient Variance in Deep Learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization A Study of Gradient Variance in Deep Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.925694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.925694Z digest=sha256:64e2894392a81b66dbc7961e80fd80fcbb6fdf0a675d818a5e55e0c69670c9d5

Observation 39c137a0-1bb3-4d63-a904-e524da2847fc · outbound

This paper cites Understanding the difficulty of training deep feedfor- ward neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Understanding the difficulty of training deep feedfor- ward neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.358803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.929387Z digest=sha256:f8163b3c78b88919721ac64c53ebb5e242b51f6fbe6370d26f5a9ce2368f0eba

Observation 2b9b9e64-56f9-45ad-82a7-5bf40561b76f · outbound

This paper cites Take a shortcut back: Mitigating the gradient vanishing for training spiking neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Take a shortcut back: Mitigating the gradient vanishing for training spiking neural networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.346789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.933179Z digest=sha256:732b5008f8342e5319ed96b830768b43959005a6235968afa63464aca78382a7

Observation 0ded4a05-1d21-438d-876b-2a6eb660e9db · outbound

This paper cites Stable architectures for deep neural networks.Inverse Prob- lems, 34(1):014004, 2018.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Stable architectures for deep neural networks.Inverse Prob- lems, 34(1):014004, 2018

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.936717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.936717Z digest=sha256:3a7b3951925285f06f7a47a016375e31894872272df57d0c2f30a0c7084ab3e6

Observation 5b8fee62-d60a-4c24-982a-08ffbde2ad7a · outbound

This paper cites Deep residual learning for im- age recognition.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Deep residual learning for im- age recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.940301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.940301Z digest=sha256:2e31ce19c6a17d53e20d7dcf307276701e84d090d0903b8f48fd2f068ea49ada

Observation 277ef9dd-4430-4ac2-84ce-f950cd0255d2 · outbound

This paper cites Densely con- nected convolutional networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Densely con- nected convolutional networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.326434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.943653Z digest=sha256:caefcaee00dfd771261a522994a5bc24192dcbfb13411da341784326c9eb2544

Observation 8680cf61-593d-422f-a7ab-24d101621042 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal covariate shift.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Batch normalization: accelerating deep network training by reducing internal covariate shift

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.314758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.947919Z digest=sha256:5716a53938bbcd3bef39b9b443fd3c4857120e2a2b0eb8aaa6a1ceb6dfe06510

Observation 14022fec-0fee-4ade-b581-4aeadc1fc45d · outbound

This paper cites Adam: A method for stochastic optimization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Adam: A method for stochastic optimization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.302393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.951799Z digest=sha256:fa3d14e3a07dc77a4db252cf3207b17a587ccf3b9ea0b8f1c253d2c7fd02361f

Observation 136da08e-f23b-47f1-9cbe-f0acc8558cf5 · outbound

This paper cites Learning multiple layers of features from tiny images.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Learning multiple layers of features from tiny images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.955468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.955468Z digest=sha256:a442588dc4a91e4677b75bd5ea42d5a79368a1a920003b8aac0ddcb62af04557

Observation 0b18b826-dd66-4c09-9ecc-2c100c3d7a92 · outbound

This paper cites Decoupled weight decay regularization.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Decoupled weight decay regularization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.958937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.958937Z digest=sha256:30d93b6bcdbc136729c25cd75c3c1616af3732af2e8170876dd36103ec0bb946

Observation 30333d8f-2e0f-45f9-a5e4-418c21ce1513 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:51:39.273980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.962218Z digest=sha256:2478fbd7b794c31c4ca7ea05c7d26a9473e6faf07408d9aa1b97ef8ff7e183f7

Observation d105d30b-bd0c-4d82-b4f0-7a71b96498ea · outbound

This paper cites ViT-CIFAR: PyTorch implementation for Vision Transformer on CIFAR datasets.https://github.com/omihub777/ViT-CIFAR, 2021.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization ViT-CIFAR: PyTorch implementation for Vision Transformer on CIFAR datasets.https://github.com/omihub777/ViT-CIFAR, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.965544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.965544Z digest=sha256:b908b8bbf1bd51af796d10e5488e1ed9a1236fe3e8cc471158ad5534b6d1f1b4

Observation cab82ace-f151-41f4-8b72-45abeeb5da04 · outbound

This paper cites On the difficulty of training recur- rent neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization On the difficulty of training recur- rent neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.253197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.969144Z digest=sha256:f4a44c72544c1cfc236750cfbeb238be8f57db646996f91a00d873b68160d051

Observation 44fd442c-17b7-4fc6-a684-1050325c3feb · outbound

This paper cites How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Information Processing Systems, page 2488–2498, 2018.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization How does batch normalization help optimization? InProceedings of the 32nd International Conference on Neural Information Processing Systems, page 2488–2498, 2018

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.240820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.972958Z digest=sha256:b8b3a0f24fc476eacd512d516203b4f3e1c0d842c7e7d2afad75220f8b602c29

Observation 5e225c31-4624-4cfe-a521-1c1bde9e2861 · outbound

This paper cites Very deep convolutional networks for large-scale image recognition.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Very deep convolutional networks for large-scale image recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:51:38.976502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.976502Z digest=sha256:9dd7e1c8291f1728afa24939a33b71c8a53a562372a0525553d0803e484228f9

Observation eeccafcb-e5fc-4df6-aa39-dfcdf2de9fec · outbound

This paper cites Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Lecture 6.5 - RMSProp: Divide the gradient by a running average of its recent magnitude

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.220382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.980056Z digest=sha256:85da8e103844f27ddff9389cdff6dd3f95cab6cce61f624b39f2c27151a4557d

Observation 41d7e801-aebb-46f8-bee6-af6aa73c5519 · outbound

This paper cites Aggregated residual transformations for deep neural networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Aggregated residual transformations for deep neural networks

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.199732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.987147Z digest=sha256:3c9fb5a1ce15e95ce79d3a41bf31d6a0e2240343fae7ac9e3df3da10dad96820

Observation 043e8490-a3aa-401d-8c54-cea17189a075 · outbound

This paper cites Gradient centraliza- tion: A new optimization technique for deep neural networks.Proceedings of the Eu- ropean Conference on Computer Vision (ECCV), pages 635–651, 2020.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Gradient centraliza- tion: A new optimization technique for deep neural networks.Proceedings of the Eu- ropean Conference on Computer Vision (ECCV), pages 635–651, 2020

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T10:51:39.187647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.990904Z digest=sha256:857e510dfb2d53b64b4a85d4ce43992b7698e27b8f029a7a5b9935724c3be46c

Observation 6e4fcd63-838c-4864-b798-6e7bacc0f7c0 · outbound

This paper cites Znorm: Z-score gradient normalization accelerating skip-connected network training without architectural modification.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Znorm: Z-score gradient normalization accelerating skip-connected network training without architectural modification

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.175325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.994138Z digest=sha256:3bbed84d47ba4477d4e4f47ad1d4faf3bc7ba812f09ab0b28ad97e8872b0583d

Observation 1f6f9682-e8f3-43ab-b0f5-5f048a893f88 · outbound

This paper cites Cutmix: Regularization strategy to train strong classifiers with localizable features.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Cutmix: Regularization strategy to train strong classifiers with localizable features

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.162901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:38.997441Z digest=sha256:726750c7e393c34586f93fd5cc36426d62e0fa8cbd11473be10d8f3c255f8b70

Observation 03ba7a4c-fef0-45a3-bab9-bfdc4d3a0626 · outbound

This paper cites Wide residual networks.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Wide residual networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.150168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:39.000807Z digest=sha256:117843dcd60cad4442550a07dd0d7169b934813d4d7cb3f0a30dd42c1426af06

Observation b3896c69-dee6-45ef-87f8-c291c0611120 · outbound

This paper cites When will gradient regularization be harmful? In Forty-first International Conference on Machine Learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization When will gradient regularization be harmful? In Forty-first International Conference on Machine Learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.138025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:39.004136Z digest=sha256:e639c2a8ae64a2338a1614c5b4ef43d1f7c3f9852200ee0881a9e0fc37b11e05

Observation a4417bfb-ccbe-4eef-b9b1-300ea2cf1166 · outbound

This paper cites Penalizing gradient norm for efficiently improving generalization in deep learning.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Penalizing gradient norm for efficiently improving generalization in deep learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.125921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:39.007740Z digest=sha256:75af640d7e1074cf3ef1d03475af5aa9728dec58dbff82ebc6a87bbba48238cb

Observation 4a462906-6865-4436-bb8f-37482170f4b7 · outbound

This paper cites Recurrent neural networks: vanishing and exploding gradients are not the end of the story.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Recurrent neural networks: vanishing and exploding gradients are not the end of the story

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:51:39.113707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:39.011129Z digest=sha256:21f540fdc43cb50bf70ee65dea787397246952a0af5c6860d87e769e67ef4e4b

Observation cba70997-9462-43ef-88ca-6d9c72dc12f8 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-05T10:51:39.100030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T10:51:39.015136Z digest=sha256:8e6cdd2a22474ca3e71215e073bd03f2d8e55f1d42ab4238b2e510ff65022f06

Observation dd4adf9b-d6e3-4efe-8629-5e515d7ffd98 · outbound

This paper cites an unresolved cited work.

Insights from Gradient Dynamics: Gradient Autoscaled Normalization Unresolved cited work

Reference 2012

Resolution
parse uncertain
no resolver link, observed 2026-08-05T10:51:38.983623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:51:38.983623Z digest=sha256:3ab3d51cdf2b5ed608ce0dbb4eda0c7e6dc5b589f388125b7663ed2035ab8cc4

Pith citing papers

No inbound Pith citation observations are available.