Pith. sign in

Paper Citation Record · LEDGER

Gradient Methods with Online Scaling Part I. Theoretical Foundations

As of 9 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2505.23081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23081 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:18.346914Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:35:24.103094Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:46:45.040439Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 582a86a0-fd6c-4ea8-b92a-5d0679058bde · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.894473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:06.894473Z digest=sha256:3072e44adadd21947f51380b73c3ff8739bfad685ecd9dc8dcfe10413df6f400

Observation 211f7c72-ff18-4c09-98d4-96bdf3d248f3 · outbound

This paper cites Parameter adaptation in stochastic optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Parameter adaptation in stochastic optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.100819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:06.990413Z digest=sha256:462ca893abf789a69b3f20dd77348dd525c389e71703687681d0d7d2f1e6c6e3

Observation 599fe075-96bd-486d-81b3-a6cdd78f65b6 · outbound

This paper cites Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.891809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.143282Z digest=sha256:48c0ab17ef29b5a973ed9b16043643567931c5e4fe1b294de6c19373b66a1e26

Observation 97a905fd-c4f7-41c4-8a17-67139bb6e03d · outbound

This paper cites Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.749701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.264191Z digest=sha256:58be2d454ab0a96ad2fd7cc2f24e4db86b0917a791a918b493d7ef9cf6f35024

Observation c5b313fa-e5df-4908-9b14-6441dfa666ba · outbound

This paper cites Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.470056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.400967Z digest=sha256:18c5fc233c4e16566b2770bf4e8d849984ce562a52a33045a1c6bfbcd0571f41

Observation 685e8275-53dd-47c0-af85-624b89fa7cc4 · outbound

This paper cites Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.237528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.573874Z digest=sha256:ffdfacf3d45c932c87585568cfb46d461171153010af1145032e83779750a0d4

Observation e00dd785-aae5-44e8-8e37-adb003f682b6 · outbound

This paper cites Online learning rate adaptation with hypergradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning rate adaptation with hypergradient descent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.987009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.744515Z digest=sha256:debb49f75ccd5fe565277cf6ab57fd69993fa66857f745591ad5774f836e6970

Observation e0b9c6f3-bd1d-4b8a-9d02-fa5d750a7e55 · outbound

This paper cites Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.773172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:07.900330Z digest=sha256:c167a2edb5c78bea47c740c183443b4dd0ad34a1ec35652a5b3f732cee85249f

Observation 08be7905-0102-40c6-a315-5309d64fbfbe · outbound

This paper cites Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.974495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:07.974495Z digest=sha256:7fab61a0e665bae30835d2e31a901f791eb2b14a2880d7a27921b79b39fc0abc

Observation cc825b92-a85e-441e-ac2f-0b5ca3016d0c · outbound

This paper cites Non-monotonebehavioroftheheavyballmethod.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-monotonebehavioroftheheavyballmethod

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.467420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:08.166282Z digest=sha256:49a0374c462bb0a42799f887899f609116ecb7b91905a8414e26d7ccab854176

Observation 698dab72-3f8c-4753-9886-28607abc2b74 · outbound

This paper cites An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.295597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:08.339118Z digest=sha256:3f90fb3848a6804e8e6ad580225dcf5ceb99cf65f93653b76206016c071282a0

Observation ad823aab-24c0-474f-b3ef-d0b9e2921766 · outbound

This paper cites Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.488432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:08.488432Z digest=sha256:7c0fa94b8d84ed7ee9635c84cb707bc099d3998ce4dff9572a29d52e4705f2a4

Observation 700875c0-508d-452e-81d6-ddfa3eec3544 · outbound

This paper cites Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.076594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:08.656273Z digest=sha256:d9260856ab3bc2c13b7ad660889abc27bcc8652561e1d39c1f195f73e8941317

Observation 36c42a4e-7125-48b8-9783-19a765fb6ca4 · outbound

This paper cites John Wiley & Sons, 2000.

Gradient Methods with Online Scaling Part I. Theoretical Foundations John Wiley & Sons, 2000

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.834841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:08.809382Z digest=sha256:b4e9ea13d968c0b4c3520eab92f6b4432a94191f417e8f7855012ba907fe7a6e

Observation 3ea42a4d-5fdf-4220-b05e-a03902fc660f · outbound

This paper cites Gradient Methods with Online Scaling.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient Methods with Online Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.008540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.008540Z digest=sha256:b7c462e0ebb4a5c12ce422e5174fd5af7d10130c667e51f50e60766ba2e3d7b9

Observation e3cc866c-4d62-4b8a-b00e-b6e99c6b6d3f · outbound

This paper cites Scalable Approximate Optimal Diagonal Preconditioning.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Scalable Approximate Optimal Diagonal Preconditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.165212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.165212Z digest=sha256:99ecd1cb10ad4925844fe4f9676cc31ce922dee8affa85dd69fe92b0576771ed

Observation d7242c41-90e9-4302-98c0-6a9148e4e5ac · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.347627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.347627Z digest=sha256:4ae711c596e10c34dbbcea95c8125f8af608dc6e8aaeec8949eaffe08656344d

Observation 34e87645-c7c9-4b16-a7d6-cb95dde4dbfd · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Shampoo: Preconditioned stochastic tensor optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.552063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:09.577230Z digest=sha256:b490febd8866abb7b72be1d544674596280dd9d311e0acf384dd7c27594fd2ad

Observation e82d314b-437f-4135-a23d-8f672afdbcf5 · outbound

This paper cites Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.256654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:09.773950Z digest=sha256:1799690dd4c5198fcbdc578490e222fe00e9f8715f86ca703d797a359bd1072d

Observation 830e8eff-9523-47ee-986c-3c8d79cc32c5 · outbound

This paper cites Revisiting the Polyak step size.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Revisiting the Polyak step size

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.962195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.962195Z digest=sha256:d8380fbc4b0ea54a61b4b311ad082821f1645e4d1078b7f9ad7899dd2cf88e3b

Observation da8727a2-61f7-4d95-a314-16abdb8e8882 · outbound

This paper cites Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.949825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.093964Z digest=sha256:c9e0391e37f53c228aa5d90c32bc2fff880a0c95429d8599ae0ebcd594ece106

Observation c3541b60-4e1e-4b6e-bb89-7d2e4984a221 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.627509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.226107Z digest=sha256:a50e696664dd9a600681558bb06cc6e4f0f859469068b81efc344a34a284deff

Observation e6011afc-fe50-4103-97f7-d1c2d57e712e · outbound

This paper cites Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:10.367209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:10.367209Z digest=sha256:0da12fcbaba71ab5cc532507560db2a766448d899f189cd949b4571207c7bef1

Observation 242d1643-d6d3-4f2e-a90a-ef7b5f1ef31c · outbound

This paper cites Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.321487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.441402Z digest=sha256:88768f959141763960ca680164aec379285a3e27439dfada475943a7f40e89fe

Observation d1204d1f-fbf2-41b4-ab6e-d79bf1ea9980 · outbound

This paper cites Unconstrained online learning with unbounded losses.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unconstrained online learning with unbounded losses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.152637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.614661Z digest=sha256:72e5f012d7e27b123f735cdd8b0162d818599d42ec58e30e1faacbd7900835ae

Observation 21918e13-066d-48bb-9e7c-8b7f63fc7b0b · outbound

This paper cites Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.917531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.776182Z digest=sha256:0fc6a9518cad1e478b5cd79c51bcaf5f941d53cc5818059c8eac3aa9594c7d43

Observation e4d11c3a-afaf-4ef0-bdf8-fd298e008b85 · outbound

This paper cites Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:19.028351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:10.986354Z digest=sha256:5692565e12f369214b1f01045a4702f2ceae9716f1f7333184e447af4e3805f9

Observation 1a392729-b1cc-4389-b592-503af395f9af · outbound

This paper cites Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.781008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:11.149792Z digest=sha256:a5b345adf0a54d82af96b883a5698a6cfcd05a7000929ea86e1d67a7ee3aa109

Observation b509697c-f221-4fff-983a-ed592bab0dd8 · outbound

This paper cites Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.348183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.348183Z digest=sha256:082a1d8dca5216cd0be7f74cc606ea3bf06227d8aa794e57a102d0a977ac822e

Observation e0306c6a-5ece-4e3e-996c-773dbf121f8e · outbound

This paper cites Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.500660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.500660Z digest=sha256:0942b91e46e0af113873444c6cd32b6e89fd572e8aa1fb7e5d7051d12e84f3dc

Observation 2b6fc337-097e-43af-b5ff-a957bed4e696 · outbound

This paper cites Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.586015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:11.664806Z digest=sha256:d47f5a6736df308c81ca5f7cbcc4c6c9dcc9f2a8ed72dc68e3e5dbb2d938776a

Observation fd73d695-ef08-4ce8-833b-5df61ebae305 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam: A Method for Stochastic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.796601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.796601Z digest=sha256:f11bbf0d4e8ea93a5ab082fdb8425667cd76944c67bac2980a2f7f0369d753b7

Observation cda76596-1567-42f0-935a-41cdfc408c9f · outbound

This paper cites Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.347020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:11.949488Z digest=sha256:7a30c8d88586adcc8d724c513fb750b1f3b53a0ac061d9afb5f433a7551d448d

Observation 2f51d06b-a0fb-4662-b0ab-143ebea5d09d · outbound

This paper cites Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.138443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.138443Z digest=sha256:d02375a2dacbe104b02532d64c11ba04d9776682f14a874d2342e7e61732156c

Observation 5f05fe23-b761-4118-9756-402baffd1343 · outbound

This paper cites A simple uniformly optimal method without line search for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A simple uniformly optimal method without line search for convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.329475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.329475Z digest=sha256:3e1753fc700f4a7a778027dbd6e77a6f591f9c4c52dd5516ba5dab6e0bedda88

Observation 22afb333-41fb-4049-a4ec-f93bc7ff20bc · outbound

This paper cites A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.024072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:12.491677Z digest=sha256:c9184a6ac77377af71a58b96e8b37b12f1cebf697942a80590a99a7d773326ab

Observation 58b31885-133b-402d-846b-f502f5e517c4 · outbound

This paper cites An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.490108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:12.806062Z digest=sha256:4d1abe3d77ab694ac54e09b92cba5fd65d0734ed39a2a1a54e81f00b5397517d

Observation 10ffd99f-c795-4c93-9a25-2bc8a81907fd · outbound

This paper cites Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.017954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.017954Z digest=sha256:887107888c8a55dce271c74e4f46255e37081cb0084acb946f2ba5da96b386df

Observation 13043e44-b86d-454d-b3e1-a3d1cf0a2a7e · outbound

This paper cites cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.221669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.221669Z digest=sha256:5b8455057976133e7ec4c5792fe885992fb9e758c20675ebe3e9842b32165525

Observation 86ae7d01-ff31-4cda-8697-9c946a456fef · outbound

This paper cites cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.373008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.373008Z digest=sha256:cd01668fa3a7209d40ff524da01206d91a5205214404643f04873ff3deafb718

Observation 8cb11ec2-df2b-4407-8d46-3c97fdc6e0e8 · outbound

This paper cites Tuning-freestep-size adaptation.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Tuning-freestep-size adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.135518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:13.492046Z digest=sha256:e51e13d6fb7e61a4d6940139dab710f909cd92d16ac2d79e92e234de4c7b277b

Observation d537402d-9002-461e-9a01-fd926b69f4ad · outbound

This paper cites Adaptive gradient descent without descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive gradient descent without descent

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.619015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:13.621772Z digest=sha256:cbe7cc13c0c55453ecb6a0c3b59f55ced4c825a5b54208c77d84c4cf8864e278

Observation 690cc81a-7386-4bb0-b398-2d4fd274d894 · outbound

This paper cites Adaptive proximal gradient method for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive proximal gradient method for convex optimization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.349459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:13.774680Z digest=sha256:6b811c5a754856d08d95b80dd2b09e1a6765462a8a6bbfa0643c2efd76b8c818

Observation 405bcadf-473c-4c53-892d-48321954815b · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive Bound Optimization for Online Convex Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.932005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.932005Z digest=sha256:f323c47f51c4d80f74b64ef238fc0ff1050c735f8d53db36c35d1790dca29bf0

Observation 3eea46ed-ab74-43c5-9a9f-6f1a7e1fb9f5 · outbound

This paper cites an unresolved cited work.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:26.037314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.074627Z digest=sha256:badc8025652b2725379dec68c6653c1ee1f07131f2181373cf43dce27e4ee862

Observation cda48efa-20d9-41cb-84ca-4dcc1c763e04 · outbound

This paper cites Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.725934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.243941Z digest=sha256:46cdb966b829de0572ec992f5f9eaa8e142614b5994fd0e7159639dfccd6aaba

Observation a97f9952-f43a-4b26-ab3b-4bad7f9a6d42 · outbound

This paper cites A method for solving the convex programming problem with convergence rate o (1/k2).

Gradient Methods with Online Scaling Part I. Theoretical Foundations A method for solving the convex programming problem with convergence rate o (1/k2)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.422538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.377838Z digest=sha256:fb0a0be11afad133fcb3a9129cd8fe4a83fd6cd64e3c2566c757e24c3ffdbe47

Observation af122e1a-fd8f-45d9-8581-b8e70602774c · outbound

This paper cites Springer Science & Business Media, 2013.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer Science & Business Media, 2013

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.146817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.540097Z digest=sha256:3f6fd757fa7019fbac517ea3d2b8381abd9ecd9246daa993fa774a692e6ce22d

Observation b6fd424a-635e-4a24-b603-d5275a7a25a8 · outbound

This paper cites Springer, 1999.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer, 1999

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.847880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.688071Z digest=sha256:f046920c39498e7ebda91beeacdc2d01cd9f0b12f26bd672f45999a08c349e1f

Observation 073b82e1-a03a-4f53-9384-b6868aad5493 · outbound

This paper cites Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.596628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:14.887592Z digest=sha256:1a7d2c2e7faf9ec46a8e0f700d008c14770e44d357a10d847ef26cdedbf8d3a7

Observation 8fad7c9b-8518-43ef-9b02-cd5c18cbbc9c · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning: A Modern Introduction Using Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.037964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:15.037964Z digest=sha256:612c2e4670295f4a5b567221eb5c70405b245da002420ee7a45a959493e3acb6

Observation 11c62399-0f76-4282-9aeb-3c8a61671e7f · outbound

This paper cites Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.353932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:15.203695Z digest=sha256:3db7a8a92bcc9329905619b937deb4c1270e01480f1e735b1aac715e2cd48924

Observation a2d47ced-c099-4894-b9ba-99dab4fdbcd4 · outbound

This paper cites MADA: Meta-adaptive optimizers through hyper-gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations MADA: Meta-adaptive optimizers through hyper-gradient descent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.997951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:15.404845Z digest=sha256:c5229e20145a008d9f3dc9a32b57d216b9df0cecc4bcc2ea9e85319f15d89a9d

Observation 7d629dba-c740-40a2-b598-0f19e07ecefa · outbound

This paper cites Introduction to optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.506836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:15.614007Z digest=sha256:ead7e3376797662e4b11b3002bc6c17b4643c83fa23132e51b965d9c794f4e96

Observation 0473bbe0-f075-45ac-a88c-d90e57b30df4 · outbound

This paper cites Optimal diagonal precondi- tioning.Operations Research, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal diagonal precondi- tioning.Operations Research, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.187929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:15.729222Z digest=sha256:eb9dd90991b9d9c32d83748725405d25e0f25f50577458b12b2d3dfc93c56689

Observation e227455b-df14-4f60-8c1d-8ac694052ba0 · outbound

This paper cites Lecture notes on online learning draft, 2009.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Lecture notes on online learning draft, 2009

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.888468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:15.885167Z digest=sha256:9867167a5ea0a78492607df9acc69b488a5abbd17ce322d8a316f16a66c32aaa

Observation deef3e93-5088-46c0-b9e4-f524ea99ca99 · outbound

This paper cites On the Convergence of Adam and Beyond.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the Convergence of Adam and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:16.063965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:16.063965Z digest=sha256:26b240652732cd9c77f57f530830e6dabf5a66d629d3d25ec1c59fefa82e0eb0

Observation 07821767-069b-4dad-ada2-b1eb6a819236 · outbound

This paper cites Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.657859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:16.251549Z digest=sha256:aa305f8020c922fd7e1d7deeba74b0b837f0e020f756edab7df5ce1fe9669969

Observation fcd11103-fce6-4d6e-9177-588f55060dc1 · outbound

This paper cites New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.492178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:16.373687Z digest=sha256:60569ba50314039580bb337cc360d341234a0720b114ba99604c4f83c8e97eee

Observation 50bcfa29-c047-477c-aa49-627546b67153 · outbound

This paper cites Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.203383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:16.529064Z digest=sha256:a43a501ba8d210ccbfc14a6cc7200288d636543429dab49373d65d04395f3b31

Observation 28a70dfa-4977-4735-8bce-83f2423e0fd4 · outbound

This paper cites Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.869002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:16.744234Z digest=sha256:f8b6c9336d0eb11bbc57a1569420c10e625c040ac5fc9b085421c244e5827a46

Observation 86f09eaa-efb1-48da-a833-c69adcb76bb9 · outbound

This paper cites Local gain adaptation in stochastic gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Local gain adaptation in stochastic gradient descent

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.591671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:16.855472Z digest=sha256:d96e84aa9acd7382dc0879048ee4c33cd02057c9e913c3b7ec91d2d6c7e9b0b2

Observation 6a17c6f7-b9c3-4b92-853e-850ce34d27e5 · outbound

This paper cites Adapting bias by gradient descent: An incremental version of delta-bar-delta.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adapting bias by gradient descent: An incremental version of delta-bar-delta

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.248493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:17.038377Z digest=sha256:0a34d98f50cc28208184a79ebd1048e499c605d7080e8a0d76d556ed3b3c4082

Observation c4be4a74-e36d-4bec-be13-b0d9120ff377 · outbound

This paper cites No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.916902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:17.226881Z digest=sha256:735daaa5046b73f4e0c36d58fd50580fd2380c03da803adc24fe0b1e4abb3f21

Observation 2b5126e1-b40e-4efb-b46c-ee6866f94cb9 · outbound

This paper cites On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.605338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:17.431702Z digest=sha256:fc7de54aeb579e778bbbbbad809b10f6db9194afbb2676762521d7e7006f6913

Observation 81a47c91-72de-40b5-8262-7b20494b7c7e · outbound

This paper cites The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.588000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.588000Z digest=sha256:21a4fbc143324cae90c5dd8a41591e224fc3c9fa3487b8a60805d1ba338433b1

Observation 09a18a02-c2b9-413e-a58c-08ea1777735d · outbound

This paper cites Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.317776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:17.719711Z digest=sha256:dece1a681d1d44ef3b2f3bbcd8a0c020193bea21e5e5193074574b28848f0712

Observation ccd1611f-fda8-426f-b2c8-293a42ac63b6 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam-mini: Use Fewer Learning Rates To Gain More

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.855272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.855272Z digest=sha256:2eb568295041f8677bf0263cd6ae43e5971b94e5f96f3c9069b8bca0242ba888

Observation acf4447c-13d9-48d8-bdf1-9b7663236653 · outbound

This paper cites Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.001214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:18.009345Z digest=sha256:7f2a5450c1e26661f70156017b21a98e46d55213ee6899641039f4efef7b0a2b

Observation 1caf9fe7-598d-494a-a3d9-884a50a16de7 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:18.193092Z digest=sha256:803b8d160b3318350b8013cf19d0bd85dd318c62a834d57de949b22a371d9582

Observation f8f4b0a7-8547-4732-9036-7ad4c23feadd · outbound

This paper cites Surrogate losses for online learning of stepsizes in stochastic non-convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Surrogate losses for online learning of stepsizes in stochastic non-convex optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.409231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:18.346914Z digest=sha256:3a1f89354a51b3f263e8cb43e23360ccb6a4513e7e84af2e45c223ed1a49ec0d

Observation 9400cacb-1509-4471-be06-d6ca5c05c077 · outbound

This paper cites (cited on 17).

Gradient Methods with Online Scaling Part I. Theoretical Foundations (cited on 17)

Reference 6564

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.789881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:10:12.619868Z digest=sha256:74c7e9b98da84e77c536f7b7d907e9f9aacc9610953f465f50e34dc15b438653

Pith citing papers

Observation ff55bcce-22c1-47d3-8546-124f443adfee · inbound

Enhanced PDHG for Linear Programming with Online Preconditioning cites this paper.

Enhanced PDHG for Linear Programming with Online Preconditioning Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:24.103094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:24.103094Z digest=sha256:44f392408c5beb9661c6d4245ecaa63bbf926ffdfcdcb6675cdbaf7ddf1e9160

Observation d0b4f173-62f7-41a0-89ae-6cc9a6d201e3 · inbound

Learning to accelerate distributed ADMM using graph neural networks cites this paper.

Learning to accelerate distributed ADMM using graph neural networks Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:46:45.043079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:46:36.084583Z digest=sha256:b4a3e2ad1b38843073a969214893d607d8aaa01d760b38528eacb3014f42f8ef