Pith. sign in

Paper Citation Record · LEDGER

On the Variance of the Adaptive Learning Rate and Beyond

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 63 inbound Pith citation observations for arXiv:1908.03265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03265 v4

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:28:07.436964Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:39:43.033568Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

608
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 269d722e-495f-4336-9488-0203991af3f0 · outbound

This paper cites Layer Normalization.

On the Variance of the Adaptive Learning Rate and Beyond Layer Normalization

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.396921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.396921Z digest=sha256:fec9af014192a43627d1188a822c1a5312867c3cbd5f407ef26c3c30e4bf93d0

Observation 15f1a38e-20cf-4bea-abba-01cb525356b5 · outbound

This paper cites an unresolved cited work.

On the Variance of the Adaptive Learning Rate and Beyond Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:28:07.545148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.426485Z digest=sha256:aa95e04e17e3d5794d47a0c811cd3bdc78996ad1a6dea067d142379dca29fcf2

Observation 87c0e687-3b0d-4ea7-9322-1399296c4a71 · outbound

This paper cites Taylor series methods.

On the Variance of the Adaptive Learning Rate and Beyond Taylor series methods

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:28:07.565493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.416170Z digest=sha256:ef9698ecb34c4a90320f79a9fbd6ccfcb54262fa5f82004ad12d59aff4c2a6d7

Observation 602da088-cebb-4951-b49f-8e457b676686 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

On the Variance of the Adaptive Learning Rate and Beyond ADADELTA: An Adaptive Learning Rate Method

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.419627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.419627Z digest=sha256:c788b39e0ae3a453045e2b448a52f65af6c3e5b4b2049b916270ac485dc9406a

Observation 71763c4f-5d5f-4b41-8b3a-93478d91d605 · outbound

This paper cites an unresolved cited work.

On the Variance of the Adaptive Learning Rate and Beyond Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:28:07.555175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.423191Z digest=sha256:18aeec3fb9605eb4c5c989d5346b9d753b96e58d2f1fe7f886b93f41ee7af7cf

Observation 0c791f5f-41e4-4148-9413-35784ff1c924 · outbound

This paper cites Specifically, we use two-layer LSTMs with 2048 hidden states with adaptive softmax to conduct experiments on the one billion words dataset.

On the Variance of the Adaptive Learning Rate and Beyond Specifically, we use two-layer LSTMs with 2048 hidden states with adaptive softmax to conduct experiments on the one billion words dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:28:07.534816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.429935Z digest=sha256:b824dc0bedf64e64a8db31a79ec3ddbaed7a4dc6c7584acaa530b8a0dededd72

Observation e3bd9a61-e2bd-45ca-85a3-b08e6377c96f · outbound

This paper cites an unresolved cited work.

On the Variance of the Adaptive Learning Rate and Beyond Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:28:07.523189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.433100Z digest=sha256:4192edf549bf565b6767eb4b78345d387898cb15c33b02a61633da07f5d45279

Observation 982da36c-afa4-4f91-bb9b-8a1690968aa5 · outbound

This paper cites Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks.

On the Variance of the Adaptive Learning Rate and Beyond Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.404944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.404944Z digest=sha256:9d9a90d74991f2c72a5e53239a5579245372a938f9d2ed315894c4f8b290cfd5

Observation 2d26b5b2-40e0-4306-8e35-50f07e2ac403 · outbound

This paper cites Advances in optimizing recurrent networks.

On the Variance of the Adaptive Learning Rate and Beyond Advances in optimizing recurrent networks

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:28:07.575739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.401429Z digest=sha256:9665cba4d94c5cc6eed6d9714fde70ca6559d5d31cf9932011a40332331c103e

Observation 171e4bee-62fa-47a7-ba21-135b5701fd6f · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

On the Variance of the Adaptive Learning Rate and Beyond RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.412371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.412371Z digest=sha256:be0376061f80a4e320f3d0bb34c57235e70bf925bc2c0ea27a54aa16d23d4574

Observation 5c73596d-2d4a-4e58-8042-726d0ef29b30 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

On the Variance of the Adaptive Learning Rate and Beyond Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T14:28:07.408678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:28:07.408678Z digest=sha256:87775183661a4f4c2bc5a2da1476ed7cee8a7bc8ea2b6dbd2e8dd1b7a69326b1

Observation 06f93314-7402-4f5b-83e1-f3d9cd89158f · outbound

This paper cites an unresolved cited work.

On the Variance of the Adaptive Learning Rate and Beyond Unresolved cited work

Reference 8196

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:28:07.512790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-14T14:28:07.436964Z digest=sha256:fe128133fb9c811d0a4752f771dc0e7575ece34a85c7775d4ff5139f08e7b49e

Pith citing papers

Observation 9bb73dcc-e9fd-4371-ab76-5469c992a53c · inbound

Raw-to-End Name Entity Recognition in Social Media cites this paper.

Raw-to-End Name Entity Recognition in Social Media On the Variance of the Adaptive Learning Rate and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T13:20:19.628026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:20:19.628026Z digest=sha256:432017ae612069dcd10ae83953ef3aa8e4b8d04eefec719698c97ea902e78000

Observation 1d4ed142-b2db-4c7d-8e9a-e1023b444557 · inbound

Kinematic Single Vehicle Trajectory Prediction Baselines and Applications with the NGSIM Dataset cites this paper.

Kinematic Single Vehicle Trajectory Prediction Baselines and Applications with the NGSIM Dataset On the Variance of the Adaptive Learning Rate and Beyond

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T10:18:08.719361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:18:08.719361Z digest=sha256:622a2a882f7d7931d5d051d9462f86393d37684ae13696c273b6ed5e80dc8a93

Observation 43ec225f-5dfe-4f7b-9cc2-091a0ccb5c79 · inbound

Port-Hamiltonian Approach to Neural Network Training cites this paper.

Port-Hamiltonian Approach to Neural Network Training On the Variance of the Adaptive Learning Rate and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:48:03.852071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:48:03.852071Z digest=sha256:bc4865cf5e40ecb4d0d558c3906f25fa038dd17502578fca32e4a9df05cafdcc

Observation 43527666-8bab-44bb-b011-16202f68355f · inbound

Consistency Models cites this paper.

Consistency Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:47:28.659081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T15:47:28.598205Z digest=sha256:e38f1700167aad2969046a92bbef1dc2c657040947ec1cb6e64408483203cf7a

Observation 760fe633-d1de-41b2-9955-934cca7e5fef · inbound

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models cites this paper.

H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-17T18:00:50.197627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T18:00:50.053377Z digest=sha256:05b8249f6e87593a037e3cfc673b6bee422faed70dca021f6e1b3f06f36a7759

Observation 216fd231-8fdc-4802-a6af-9bb99ed6622e · inbound

Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference cites this paper.

Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference On the Variance of the Adaptive Learning Rate and Beyond

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:15:54.867502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T04:15:54.681913Z digest=sha256:c4573ed548bbef447acfce4d0760e15fe86073a854c73679e9a69d2ec1502744

Observation 18eb65d8-34fc-45ea-a560-afdec281900c · inbound

Improved Techniques for Training Consistency Models cites this paper.

Improved Techniques for Training Consistency Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:04:10.662000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T05:04:10.600062Z digest=sha256:976bdaa3e7c40a4e4181f005863c2b680f12c83e017a136056c46bf246b766b8

Observation ab8f389a-b751-480d-901d-8f36cad52591 · inbound

Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks cites this paper.

Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:11:16.550216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:11:16.550216Z digest=sha256:32e7e04d3b30eece598a003397df65b19cc02ae16a207458a3b3cebb87b9e752

Observation e000a8d5-f4b0-4d2b-88e4-efff3ef24cd6 · inbound

Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization cites this paper.

Beyond adaptive gradient: Fast-Controlled Minibatch Algorithm for large-scale optimization On the Variance of the Adaptive Learning Rate and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:39.847015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:39.847015Z digest=sha256:41ceed133635b295cd1c461443a3dec1d0a0e68e7cef8cc194c03ecf0cf84062

Observation 4a93ff98-854e-4555-9890-62befd0919e2 · inbound

Quantity versus Diversity: Influence of Data on Detecting EEG Pathology with Advanced ML Models cites this paper.

Quantity versus Diversity: Influence of Data on Detecting EEG Pathology with Advanced ML Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:29:22.668886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:29:22.668886Z digest=sha256:20ddbd88ce7463dadb900ece7c1256ee1272dd7961198939cff72cd2d879ea26

Observation 7e4ce5f8-ae17-43a7-a779-efda4c0986a0 · inbound

CAdam: Confidence-Based Optimization for Online Learning cites this paper.

CAdam: Confidence-Based Optimization for Online Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T06:03:51.315324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T06:03:51.315324Z digest=sha256:1f1b2308bf4f0c4e10134902b56a022b14aedf42a7f8e942ab05e3bbb542efb7

Observation f72ffc1a-e400-4bde-a3a1-d8ddbcfa11fe · inbound

Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows cites this paper.

Automatic Differentiation-based Full Waveform Inversion with Flexible Workflows On the Variance of the Adaptive Learning Rate and Beyond

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T05:26:44.820508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:26:44.820508Z digest=sha256:df205b7e30aaadb6743c0644b3efe0644eb4d57f72f58f11242da38163a1ce31

Observation 355fd820-b7cc-4138-a443-29807ff5ffbf · inbound

Some Best Practices in Operator Learning cites this paper.

Some Best Practices in Operator Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:03.709142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:03.709142Z digest=sha256:30bf800b3e087bce05f2e75297860c1a3f4898b5ed9679e49c9b9f61103fcff3

Observation ddff11e7-46f3-4f80-83da-3fa88889e118 · inbound

No More Adam: Learning Rate Scaling at Initialization is All You Need cites this paper.

No More Adam: Learning Rate Scaling at Initialization is All You Need On the Variance of the Adaptive Learning Rate and Beyond

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:41:00.366975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:41:00.366975Z digest=sha256:2e6f7131b47c7c111aa418688218d50597e1876f6f7fcc2663646fd67b1d5615

Observation 7aa1e192-1da6-423c-b252-6cab0852ad95 · inbound

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data cites this paper.

Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:32:54.039495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:32:54.039495Z digest=sha256:b0d7954a4d159424f1cb0a1da40b765850123228c217a5821b5ed1eea974b39d

Observation 128c70fd-0c9b-4d1b-8255-ca8d832252b5 · inbound

From Histopathology Images to Cell Clouds: Learning Slide Representations with Hierarchical Cell Transformer cites this paper.

From Histopathology Images to Cell Clouds: Learning Slide Representations with Hierarchical Cell Transformer On the Variance of the Adaptive Learning Rate and Beyond

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T10:24:04.385407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:24:04.385407Z digest=sha256:321e6dc876c6288b694af8a8947a98056554137931b314725010c660f1e78d86

Observation 9d0f8ad0-3632-4a27-9641-b8111129220b · inbound

ReFlow6D: Refraction-Guided Transparent Object 6D Pose Estimation via Intermediate Representation Learning cites this paper.

ReFlow6D: Refraction-Guided Transparent Object 6D Pose Estimation via Intermediate Representation Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T23:14:17.622839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:14:17.622839Z digest=sha256:71b4de7434f303904598f5e52559cb682f81797589c8402d5334f3ab9ef0ea30

Observation 101104e3-08fe-4b5d-b2a8-05102d73849c · inbound

Celo: Training Versatile Learned Optimizers on a Compute Diet cites this paper.

Celo: Training Versatile Learned Optimizers on a Compute Diet On the Variance of the Adaptive Learning Rate and Beyond

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:01:13.908163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:01:13.908163Z digest=sha256:5f2b54603f714d59e8e7264afd56ee4adbc997368dc8afb208438a07bd92df62

Observation 41337159-851c-4c38-8ebb-f4896e234754 · inbound

Certified Guidance for Planning with Deep Generative Models cites this paper.

Certified Guidance for Planning with Deep Generative Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:51:20.102451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:51:20.102451Z digest=sha256:1ee7cd9056cb1b25d090ad222e6fa2f3bcdb2ba94f015086aaa8fa6f3131e98c

Observation 6c36cb10-3e04-4742-978e-6ea2c9d7138b · inbound

Unlearning Clients, Features and Samples in Vertical Federated Learning cites this paper.

Unlearning Clients, Features and Samples in Vertical Federated Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:46:20.139219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:46:20.139219Z digest=sha256:0e44d127ec655f8d9fe5627286134dd5417a99e2043bf0cb2aff1f3243ac7cde

Observation f125ef0f-c08d-4f99-a0e3-9afff26d105e · inbound

CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation cites this paper.

CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation On the Variance of the Adaptive Learning Rate and Beyond

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T20:27:53.458185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:27:53.458185Z digest=sha256:6b4c3d9a48ce2b86d5a14bd6174854a6ce9c11eb784571fde91e5c83988a3d8d

Observation e9a33574-bd16-48cc-873d-846c532b0f7a · inbound

Financial Data Analysis with Robust Federated Logistic Regression cites this paper.

Financial Data Analysis with Robust Federated Logistic Regression On the Variance of the Adaptive Learning Rate and Beyond

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:39:43.033568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:39:43.033568Z digest=sha256:d317fa11f105890bef006c32ce5816b7c546e7f19f7ac46152126dd4387c4461

Observation b219a8ae-9827-4f9a-adb3-fbfa6549f7a3 · inbound

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries cites this paper.

TopicVD: A Topic-Based Dataset of Video-Guided Multimodal Machine Translation for Documentaries On the Variance of the Adaptive Learning Rate and Beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:02:43.805122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:02:43.805122Z digest=sha256:7c2ae03d4b724e4bebcc9f0ca4612aca2255b0b3a510699363ca7e3b6002f3ac

Observation 2544e12b-bf62-438b-bdca-1dcf81fcfbba · inbound

Unified Continuous Generative Models cites this paper.

Unified Continuous Generative Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:25:31.076238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:25:31.076238Z digest=sha256:299552cc3df7118d09b6953c297a5ae5a3c14c8b2ecac56a349961e101a0221f

Observation 79aa1f88-a06f-44a5-9600-d996fb18e693 · inbound

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades cites this paper.

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades On the Variance of the Adaptive Learning Rate and Beyond

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:40.655389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:53:40.655389Z digest=sha256:7eed476e2bf72e4481073c36c913b8d29addc3fb6a800b6f2765346e6d7fa02a

Observation 9b8548ed-6392-44bd-aba0-c0307459498a · inbound

Rapid training of Hamiltonian graph networks using random features cites this paper.

Rapid training of Hamiltonian graph networks using random features On the Variance of the Adaptive Learning Rate and Beyond

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:17:15.702752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T10:17:00.343410Z digest=sha256:54dd850361ddfd929440ef798f401eabadb04cf979b6337eace679092871b9d1

Observation 05fd9f67-538a-4a5b-975b-c91c9b65fd08 · inbound

Deep Potential: Recovering the gravitational potential and local pattern speed in the solar neighborhood with GDR3 using normalizing flows cites this paper.

Deep Potential: Recovering the gravitational potential and local pattern speed in the solar neighborhood with GDR3 using normalizing flows On the Variance of the Adaptive Learning Rate and Beyond

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:07:26.301836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:07:26.301836Z digest=sha256:7a95b235c804ec867c85fc040fae147105f990301c740b6ba27974b0f09ceb61

Observation 218ef676-cdc5-4fff-9afb-a88f3cad0e79 · inbound

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam cites this paper.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.591929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.591929Z digest=sha256:82cfe976d05fd411ad54396bf2de4c0bb0aef5a791c5fccca47e781f451a89fb

Observation 9e6a5159-a688-48c3-9649-7e32726407b1 · inbound

Feature-Enhanced TResNet for Fine-Grained Food Image Classification cites this paper.

Feature-Enhanced TResNet for Fine-Grained Food Image Classification On the Variance of the Adaptive Learning Rate and Beyond

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T16:43:51.780376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:43:51.780376Z digest=sha256:1e76fac552d986d60feb601d35e745ee9a7d041fe658b57f6516a68442aa2744

Observation 80f58d11-afa8-4b7e-b467-8b8a79d71466 · inbound

Criticality analysis of nuclear binding energy neural networks cites this paper.

Criticality analysis of nuclear binding energy neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T06:01:11.597771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:01:11.597771Z digest=sha256:9d87a5ec1b9bdd21aa84ef051df2b766e103892653b3afe150c25b0a2220d4e5

Observation d42c4cf1-b7c5-41ff-983f-823e41ef7f6f · inbound

From Next Token Prediction to (STRIPS) World Models cites this paper.

From Next Token Prediction to (STRIPS) World Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.462305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T16:21:22.805473Z digest=sha256:4e639b04811f77e7a4667e25d716c905cc5f59c04d5ff523eb06d6bcd4917e34

Observation 37157420-cf3a-4228-9c3f-f79600f54f5b · inbound

From Next Token Prediction to (STRIPS) World Models cites this paper.

From Next Token Prediction to (STRIPS) World Models On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T16:41:18.775357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:41:18.775357Z digest=sha256:3c1f8b9e18df9f24cfeb2f525c0cec1baa26d1e46397d6ab768616c597d1a027

Observation 53416c7b-10c1-493a-a408-472c2dd13e13 · inbound

On the Convergence of Muon and Beyond cites this paper.

On the Convergence of Muon and Beyond On the Variance of the Adaptive Learning Rate and Beyond

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:56:34.102980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T15:56:30.602824Z digest=sha256:1b8354cb32796e6930ae280f974404e017216d2b8ab0e93874cd0b7ccab94921

Observation 4a969c6c-cba2-4130-8c93-7308ff18d135 · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective On the Variance of the Adaptive Learning Rate and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.819676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:58.819676Z digest=sha256:3f3a77079b5c13e9630539aa926b55fbca32f4c63796f78b92032a17d178f8d6

Observation 8606aaef-40c8-497f-ba60-d25e8cb2dc9c · inbound

Building Deep Graph Predictors with Graph Imitation Learning cites this paper.

Building Deep Graph Predictors with Graph Imitation Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T15:45:18.726240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T15:44:58.327091Z digest=sha256:5f79ef0aab45a90a7ce60d468b8a73453ae361b8056579a35d3f767ab0f8b808

Observation 1cceda00-a532-4936-a3b8-d125000f1e9a · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:45:29.768850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T07:41:48.634351Z digest=sha256:832231d6c24844603b4821a2107fec08111585030cde13be20d1d91131b35b55

Observation 6999988d-3e60-41a8-b3a6-590f5e3cf82e · inbound

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning cites this paper.

TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T05:37:38.449157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:37:38.449157Z digest=sha256:0e9b020ec3f9e0025ccc2de44b6868a8510cef5e640a918807e8da66aa5a5d2d

Observation becc2260-cb50-4253-b2ff-542c3fcc3f4b · inbound

Characterizing the Instrumental Profile of LAMOST cites this paper.

Characterizing the Instrumental Profile of LAMOST On the Variance of the Adaptive Learning Rate and Beyond

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:10:03.181485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T14:09:12.016309Z digest=sha256:9fc50b83b0fe89232e84e0f31d55d61c6f138c3a37030ebfbcfbb0b3b3fccdef

Observation 6b7d67b1-d767-4ff2-851f-29266263f10f · inbound

Video-guided Machine Translation with Global Video Context cites this paper.

Video-guided Machine Translation with Global Video Context On the Variance of the Adaptive Learning Rate and Beyond

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:01.141997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:02:31.374359Z digest=sha256:a537ed3ddb026db1e4cb2765b3e7b0e6d76175eacfcb8076a8b0d0bfa9d62f75

Observation e23ddd4e-e4dc-4372-856d-6bf8144e9d3e · inbound

Delve into the Applicability of Advanced Optimizers for Multi-Task Learning cites this paper.

Delve into the Applicability of Advanced Optimizers for Multi-Task Learning On the Variance of the Adaptive Learning Rate and Beyond

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:20:59.066359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:42:48.902337Z digest=sha256:aa051248678b5d668a5628dfe13977df73ea868cc9fe286412bedc1bc35573a6

Observation c01c0deb-37e6-4e97-bab7-0c246942c6dc · inbound

Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons cites this paper.

Particle transformers for identifying Lorentz-boosted Higgs bosons decaying to a pair of W bosons On the Variance of the Adaptive Learning Rate and Beyond

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:01:00.835334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:18:12.735038Z digest=sha256:3f43f7adae05f337bacb0f3bd25a141c393ba02301d1563d75b0cf3db1e1fead

Observation 8ae7bd9e-455d-4181-9297-823f6433ac9f · inbound

Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning cites this paper.

Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-Tuning On the Variance of the Adaptive Learning Rate and Beyond

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:22.069640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T03:26:09.751493Z digest=sha256:c2dfc8b17005215748e3404a8acd1fc060b4981c54444156c499fe6c4828d563

Observation 59fe7434-96b1-48a5-8a05-df485960b583 · inbound

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments cites this paper.

AdaMeZO: Adam-style Zeroth-Order Optimizer for LLM Fine-tuning Without Maintaining the Moments On the Variance of the Adaptive Learning Rate and Beyond

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-09T19:50:50.653184Z digest=sha256:1b2cb517b40f39e7cc3aba99395eee910ea83a1cbe84a987cf0d7bb2f1b592a9

Observation e984c67a-2fbf-4c5a-a644-babd64eadc37 · inbound

Deep neural networks with Fisher vector encoding for medical image classification cites this paper.

Deep neural networks with Fisher vector encoding for medical image classification On the Variance of the Adaptive Learning Rate and Beyond

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:00:58.899036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:22:00.015928Z digest=sha256:2f96e06142cbddc3c19a558e2bca24aed763ab5959ddc2733373c597b07387b6

Observation 4e6ec95d-267c-4f21-8d9a-e2e9a6cd9a9c · inbound

Anon: Extrapolating Adaptivity Beyond SGD and Adam cites this paper.

Anon: Extrapolating Adaptivity Beyond SGD and Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:09.614347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T16:25:36.073417Z digest=sha256:165b6abff278d376be8434dd0e75b5617410e7f0138e6d8765a5f550a8e4d809

Observation 00736419-3d20-4b7a-a440-5d677ac61205 · inbound

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior cites this paper.

Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior On the Variance of the Adaptive Learning Rate and Beyond

Reference 297

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:16:06.693844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T17:47:09.591001Z digest=sha256:16fd870fb099b55c85277f76881a629e5f157f2d2dadcf6d111664ab195ef9c3

Observation 2918b059-9733-48f5-a3cd-47bd6101ba2e · inbound

Neural Network-Based Virtual Wheel-Speed Sensor for Enhanced Low-Velocity State Estimation cites this paper.

Neural Network-Based Virtual Wheel-Speed Sensor for Enhanced Low-Velocity State Estimation On the Variance of the Adaptive Learning Rate and Beyond

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:37:15.837211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T04:32:39.053186Z digest=sha256:f4957f5506aa601b7fb97a7616d5c910971d36811559c6f52153c85e5884e340

Observation 59e925a4-f420-47fe-a40e-8169504af7f9 · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling On the Variance of the Adaptive Learning Rate and Beyond

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:24:27.799007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T00:21:15.933174Z digest=sha256:b3c83c14baba199d28aac688d030fff8422997e9f305964d3e1cb6eab6d2c368

Observation 200df501-70ac-4759-a340-d81bee41b7fe · inbound

Scalable Reinforcement Learning via Adaptive Batch Scaling cites this paper.

Scalable Reinforcement Learning via Adaptive Batch Scaling On the Variance of the Adaptive Learning Rate and Beyond

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:24:57.664358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T17:17:59.698127Z digest=sha256:f0e963eb1001734cccbf8dfa9751579d51c3b7f348a6a29f171de595d38ebcd3

Observation 431879f9-244d-4bb3-bfd4-38146499c5ea · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective On the Variance of the Adaptive Learning Rate and Beyond

Reference 169

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.922851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:cbdfd35be953847c8e1a7c674b42b16e5e12912af4fff7cb5a6b8df0ac492722

Observation a73ac89d-0f45-4267-a7d4-3024d2c63041 · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam On the Variance of the Adaptive Learning Rate and Beyond

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:30.094776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:2a59ea94c40dba305b036c5c4190b6df56a76cfec37e7e862c73b52cdfa87ecf

Observation 5b67d29d-83ba-4c55-af17-ac27193e4821 · inbound

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning cites this paper.

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning On the Variance of the Adaptive Learning Rate and Beyond

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:41.024926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T05:03:06.472027Z digest=sha256:6637e5f68d1107af06536112577300f053837dc24ec80048bbdd84c9f209e1d3

Observation 8e1790d6-d0bb-4204-a2d9-f13af6cbcd0c · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors On the Variance of the Adaptive Learning Rate and Beyond

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.698743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:8124fd28722b4d89e52958f527a51d983413bb55e356538409be767ca89e5cbe

Observation bf816f9e-848a-4f3d-a30e-92ee06128d26 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors On the Variance of the Adaptive Learning Rate and Beyond

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:08.871888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:08.871888Z digest=sha256:bfacaee0958c10b185f67b98713e9947566e9b803dc4827fed946632f30d2399

Observation 8aa828d9-8e2a-41c4-81a9-a577794beac8 · inbound

Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference cites this paper.

Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference On the Variance of the Adaptive Learning Rate and Beyond

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:14:25.179388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T08:09:50.192952Z digest=sha256:b287bfd3909369484be70ac255bd4653ef2b3ec7b220410078854cf8e9496bdc

Observation 3909f6fb-bb73-4b1a-8d41-6a0245ba64c8 · inbound

Point spread function wavefront recovery from in-focus stellar observations cites this paper.

Point spread function wavefront recovery from in-focus stellar observations On the Variance of the Adaptive Learning Rate and Beyond

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:36:40.238509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-02T05:34:06.325999Z digest=sha256:a9955247ca9079b9d2173b0b8110789080ed13d3dc707a68ab55cc708a616442

Observation 9acfc265-1011-43dd-a2a7-b148e9610e63 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers On the Variance of the Adaptive Learning Rate and Beyond

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:cbd45ca57b98322ddcc9947fcc93e7342304f776b941c0af2d3658d560a772d5

Observation e6b32ccb-1afa-420f-801f-d94cfd168b29 · inbound

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks cites this paper.

Unified convergence analysis for gradient descent optimization methods in the training of deep neural networks On the Variance of the Adaptive Learning Rate and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T20:46:05.467029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:46:05.467029Z digest=sha256:d0d2d1640c09816f04b2982247b51153f92a6ac4c1ac0fb7f0855174bc1a8055

Observation f9e1f0fb-c62e-4366-b4a9-d290f097a92e · inbound

A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction cites this paper.

A machine-learned probability distribution in the phase space of turbulent channel flow for synthetic turbulence and flow reconstruction On the Variance of the Adaptive Learning Rate and Beyond

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-01T16:19:27.144952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:19:27.144952Z digest=sha256:17bb5ba327e2729403fc3d092a7f7dcc3a6538e19325026340ca84d7f3021307

Observation a8553a71-a379-47ab-a86f-bd17140d0fb5 · inbound

GCR Spectra Reconstructed with Neutron Monitor Yield Function and Artificial Neural Networks: Comparison of Two Methods cites this paper.

GCR Spectra Reconstructed with Neutron Monitor Yield Function and Artificial Neural Networks: Comparison of Two Methods On the Variance of the Adaptive Learning Rate and Beyond

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T08:51:11.204445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:51:11.204445Z digest=sha256:3156f1808d3f721f01a38b7431f865da6b0df766f9756d53c8ebca996ddc6ffd

Observation 4aa982dc-4ad3-4a8a-aace-348d101318c1 · inbound

Payne4GAIN: NLTE Corrections for Red Giants in Milky Way Mapper using H-Band Neural Network Emulators cites this paper.

Payne4GAIN: NLTE Corrections for Red Giants in Milky Way Mapper using H-Band Neural Network Emulators On the Variance of the Adaptive Learning Rate and Beyond

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T04:40:18.200294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:40:18.200294Z digest=sha256:1ec277bc1f41ea6985f77314f66b40e518dbb605dcbfe1b9f05bb90558fbae68

Observation 320d20e7-082b-4fe6-a440-833a7c068a46 · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation On the Variance of the Adaptive Learning Rate and Beyond

Reference 104

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.223966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.223966Z digest=sha256:02b52f3dbf40f7d2e4c91158d29a097ec47ddd2179cd8915183ee4f2c71e6998

Observation 30ab29ce-e1d7-4d81-a48a-c360c4b99e42 · inbound

HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data cites this paper.

HealthCAT: An Interpretable Encoder-only Transformer Framework for Health Indicator Prediction and Temporal Interpretation of Wearable Sensor Data On the Variance of the Adaptive Learning Rate and Beyond

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T04:27:10.763773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:27:10.763773Z digest=sha256:1a7999138a0522544dac6d7ddbbcd4db1f0f3fca3901e7ee7d21d10cd74f6dde