Pith. sign in

Paper Citation Record · LEDGER

Gradient Methods with Online Scaling Part I. Theoretical Foundations

As of 20 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2505.23081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23081 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:10:18.346914Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:38:41.132738Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:46:45.040439Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy49
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 582a86a0-fd6c-4ea8-b92a-5d0679058bde · outbound

This paper cites Disentangling Adaptive Gradient Methods from Learning Rates.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Disentangling Adaptive Gradient Methods from Learning Rates

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:06.894473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:06.894473Z digest=sha256:538abb6214b17e5ef27cfaccc513c47b6e4566a10cd9b8b4dc988d6e0b6c93a8

Observation 211f7c72-ff18-4c09-98d4-96bdf3d248f3 · outbound

This paper cites Parameter adaptation in stochastic optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Parameter adaptation in stochastic optimization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:33.100819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:06.990413Z digest=sha256:001b06873cc8cbe2b08921133ac001f8f6e080bcbd98f607cbbc76635fb79525

Observation 599fe075-96bd-486d-81b3-a6cdd78f65b6 · outbound

This paper cites Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Silver stepsize schedule for smooth convex optimization.Mathematical Programming, pages 1–14, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.891809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.143282Z digest=sha256:3dfccc5698ec11a73d33d5f816f0e72720f049768ca956d33c7dca96e36143f7

Observation 97a905fd-c4f7-41c4-8a17-67139bb6e03d · outbound

This paper cites Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Acceleration by stepsize hedging: Multi-step descent and the silver stepsize schedule.Journal of the ACM, 72(2):1–38, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.749701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.264191Z digest=sha256:0f5999ca564d36365abdf1725be6ff23e717d6683e6ca4d0d86ce6ea4983a257

Observation c5b313fa-e5df-4908-9b14-6441dfa666ba · outbound

This paper cites Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Practical large-scale linear programming using primal-dual hybrid gradient.Advances in Neural Information Processing Systems, 34:20243–20257, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.470056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.400967Z digest=sha256:f843d23a7b00cce9be42a0ed8c2d5c19bed0cc18fe068f122c80ab756380261f

Observation 685e8275-53dd-47c0-af85-624b89fa7cc4 · outbound

This paper cites Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Two-point step size gradient methods.IMA journal of numerical analysis, 8(1):141–148, 1988

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:32.237528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.573874Z digest=sha256:0e0185ed7039250227cfc161ecb62466405199434d6619e94adcb58a4c580114

Observation e00dd785-aae5-44e8-8e37-adb003f682b6 · outbound

This paper cites Online learning rate adaptation with hypergradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning rate adaptation with hypergradient descent

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.987009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.744515Z digest=sha256:9cebbdf1f62ee594175b7d88c79b8c97d97ddd5e3b55848302b3c90c1ec2602c

Observation e0b9c6f3-bd1d-4b8a-9d02-fa5d750a7e55 · outbound

This paper cites Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient descent: The ultimate optimizer.Advances in Neural Information Processing Systems, 35:8214–8225, 2022

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.773172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:07.900330Z digest=sha256:60345a123f292732391d343fda5576732110af9393bc5061cf90327f34d8ee64

Observation 08be7905-0102-40c6-a315-5309d64fbfbe · outbound

This paper cites Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Provable and Practical Online Learning Rate Adaptation with Hypergradient Descent

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:07.974495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:07.974495Z digest=sha256:99181f226ec6d2eb9cb53f514f1c1d83245353e60880128139b707e8cab53e22

Observation cc825b92-a85e-441e-ac2f-0b5ca3016d0c · outbound

This paper cites Non-monotonebehavioroftheheavyballmethod.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-monotonebehavioroftheheavyballmethod

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.467420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:08.166282Z digest=sha256:152d1a3d6ddf43cf08ef52cc315227aa849786741fd77eab833471b793795a49

Observation 698dab72-3f8c-4753-9886-28607abc2b74 · outbound

This paper cites An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An enhanced alternating direction method of multipliers-based interior point method for linear and conic optimization.INFORMS Journal on Computing, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.295597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:08.339118Z digest=sha256:66e9f7871f16b0968424657fd4de1994e52f744999e998b6fb9a63677d2837b7

Observation ad823aab-24c0-474f-b3ef-d0b9e2921766 · outbound

This paper cites Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Uniformly Optimal and Parameter-free First-order Methods for Convex and Function-constrained Optimization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:08.488432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:08.488432Z digest=sha256:785d6f764d75e83ed889a6e6a9e516fee76e981f665d7bed2dfc68c2c01ac0d8

Observation 700875c0-508d-452e-81d6-ddfa3eec3544 · outbound

This paper cites Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive subgradient methods for online learning and stochas- tic optimization.Journal of machine learning research, 12(7), 2011

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:31.076594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:08.656273Z digest=sha256:1d7ef556ef67f5682ab081610b2e89ae2173a592639947187d9cc6b9fc66261a

Observation 36c42a4e-7125-48b8-9783-19a765fb6ca4 · outbound

This paper cites John Wiley & Sons, 2000.

Gradient Methods with Online Scaling Part I. Theoretical Foundations John Wiley & Sons, 2000

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.834841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:08.809382Z digest=sha256:e193aef999004210f36bdd1676ffbd6603b4887909fc75e85b06c771e987500a

Observation 3ea42a4d-5fdf-4220-b05e-a03902fc660f · outbound

This paper cites Gradient Methods with Online Scaling.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Gradient Methods with Online Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.008540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.008540Z digest=sha256:42463dd94ccc29b4d857a0009fa98c68bf3e1f0af2cb1d045cd160f589358c4c

Observation e3cc866c-4d62-4b8a-b00e-b6e99c6b6d3f · outbound

This paper cites Scalable Approximate Optimal Diagonal Preconditioning.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Scalable Approximate Optimal Diagonal Preconditioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.165212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.165212Z digest=sha256:4f54b25e8af8bd80fb4fd63dc1017b586043108dde9d6bf61672124e09a09007

Observation d7242c41-90e9-4302-98c0-6a9148e4e5ac · outbound

This paper cites Clarabel: An interior-point solver for conic programs with quadratic objectives.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Clarabel: An interior-point solver for conic programs with quadratic objectives

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.347627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.347627Z digest=sha256:daeaa990cae35949295cb26e277966a0d06e2053e5048e77b34753844507fe7e

Observation 34e87645-c7c9-4b16-a7d6-cb95dde4dbfd · outbound

This paper cites Shampoo: Preconditioned stochastic tensor optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Shampoo: Preconditioned stochastic tensor optimization

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.552063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:09.577230Z digest=sha256:ae6f55910b94460a675585d41733535f78adcfeeea4d91b6ff1c6c5721ac79d7

Observation e82d314b-437f-4135-a23d-8f672afdbcf5 · outbound

This paper cites Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to online convex optimization.Foundations and Trends®in Optimization, 2(3-4):157–325, 2016

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:30.256654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:09.773950Z digest=sha256:e9dbd3011c616a6fd1295adb4be00caf56ea28ccea5e72cc1f7ceb6a93ddc15f

Observation 830e8eff-9523-47ee-986c-3c8d79cc32c5 · outbound

This paper cites Revisiting the Polyak step size.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Revisiting the Polyak step size

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:09.962195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:09.962195Z digest=sha256:bc06ade02bd5bc47e93deb455f140da47f1a335ad526faa2a6469c947c333442

Observation da8727a2-61f7-4d95-a314-16abdb8e8882 · outbound

This paper cites Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive online gradient descent.Advances in neural information processing systems, 20, 2007

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.949825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.093964Z digest=sha256:a2925a8fd03076b6d360ace7318b2e72b664f72321179773edae37fd0df4b202

Observation c3541b60-4e1e-4b6e-bb89-7d2e4984a221 · outbound

This paper cites Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Neural networks for machine learning lecture 6a overview of mini-batch gradient descent.Cited on, 14(8):2, 2012

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.627509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.226107Z digest=sha256:feed5f087dc00ae6feebcf2069ea466cd79c7460c62253d00134ab4aba15e180

Observation e6011afc-fe50-4103-97f7-d1c2d57e712e · outbound

This paper cites Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Restarted Primal-Dual Hybrid Conjugate Gradient Method for Large-Scale Quadratic Programming

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:10.367209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:10.367209Z digest=sha256:bf5a2d5ee2ece81770e213a5443d097d402ec728744f1e8522fe2ac87c653827

Observation 242d1643-d6d3-4f2e-a90a-ef7b5f1ef31c · outbound

This paper cites Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Increased rates of convergence through learning rate adaptation.Neural networks, 1(4):295–307, 1988

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.321487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.441402Z digest=sha256:3d930c697e0f87e013aea15e74c908048e6ada1a18868f75e99642fcb8cd2db2

Observation d1204d1f-fbf2-41b4-ab6e-d79bf1ea9980 · outbound

This paper cites Unconstrained online learning with unbounded losses.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unconstrained online learning with unbounded losses

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:29.152637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.614661Z digest=sha256:4ec20208785ab07c7813044ae75db65f019b62a3269885c76428b06673274c0e

Observation 21918e13-066d-48bb-9e7c-8b7f63fc7b0b · outbound

This paper cites Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online learning guided curvature approximation: A quasi-newton method with global non-asymptotic superlinear convergence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.917531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.776182Z digest=sha256:c723cbdf2e6b1824d50672fa64e23a7e08e5e7b96c6c8b76b038d9025222d74e

Observation e4d11c3a-afaf-4ef0-bdf8-fd298e008b85 · outbound

This paper cites Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning Guided Quasi-Newton Methods with Global Non-Asymptotic Convergence

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:10:19.028351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:10.986354Z digest=sha256:c87e564d601e3975b95aff20962e443cb1a9f988da0972393442203dd1ceba79

Observation 1a392729-b1cc-4389-b592-503af395f9af · outbound

This paper cites Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive hierarchical hyper-gradient descent.International Journal of Machine Learning and Cybernetics, 13(12):3785–3805, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.781008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:11.149792Z digest=sha256:ea2ddcefa8a3cfaa2cc8edc73c2974006a420bb29fa6e95530b6bdef8d7568c3

Observation b509697c-f221-4fff-983a-ed592bab0dd8 · outbound

This paper cites Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Analysis of BFGS with the Armijo-Wolfe Line Search

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.348183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.348183Z digest=sha256:7a39853f8c4828827797a260696c0fb2734eeb994d8fda2fec741e56d7f20363

Observation e0306c6a-5ece-4e3e-996c-773dbf121f8e · outbound

This paper cites Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic Global Convergence Rates of BFGS with Exact Line Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.500660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.500660Z digest=sha256:ba535741c635653be42f46150ff47f0da02ad9033afe71a06e3350fc1ba9d9bb

Observation 2b6fc337-097e-43af-b5ff-a957bed4e696 · outbound

This paper cites Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Non-asymptotic superlinear convergence of standard quasi-newton methods.Mathematical Programming, 200(1):425–473, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.586015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:11.664806Z digest=sha256:66e8485c78a9e4a78cd37b88b55fa419a80ed66582232f8cc0e28fe57c8e19a9

Observation fd73d695-ef08-4ce8-833b-5df61ebae305 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam: A Method for Stochastic Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:11.796601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:11.796601Z digest=sha256:35d0dc9ffcc780d558f7287dd286cb8e2a91e3ff52681833b219419d258a6107

Observation cda76596-1567-42f0-935a-41cdfc408c9f · outbound

This paper cites Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Searching for optimal per-coordinate step-sizes with multidimensional backtracking.Advances in Neural Information Processing Systems, 36, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.347020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:11.949488Z digest=sha256:cefa532903731423030889d674aa81644bba29c2d1ede902cde1c830b3ae0c19

Observation 2f51d06b-a0fb-4662-b0ab-143ebea5d09d · outbound

This paper cites Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.138443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.138443Z digest=sha256:a9de2e93f6cb6d6107e132ac721590f7473da31ecd27e0ce74183dea6c80db56

Observation 5f05fe23-b761-4118-9756-402baffd1343 · outbound

This paper cites A simple uniformly optimal method without line search for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A simple uniformly optimal method without line search for convex optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:12.329475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:12.329475Z digest=sha256:a1d6c87305dc1f4ec87eb034103c9a3edc11634e6fafbb0ba2095df57fa89bee

Observation 22afb333-41fb-4049-a4ec-f93bc7ff20bc · outbound

This paper cites A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance.

Gradient Methods with Online Scaling Part I. Theoretical Foundations A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:28.024072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:12.491677Z digest=sha256:73c71647a7d0e147b11e13eb3f89155353b5521085960d030b5450e203e4392c

Observation 58b31885-133b-402d-846b-f502f5e517c4 · outbound

This paper cites An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations An admm-based interior-point method for large-scale linear programming.Optimization Methods and Software, 36(2-3):389–424, 2021

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.490108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:12.806062Z digest=sha256:3a6451e2c85a98d5ce161d3f48e3c1b1d8a065a9cb574eee4d04678443632413

Observation 10ffd99f-c795-4c93-9a25-2bc8a81907fd · outbound

This paper cites Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Pdcs: A primal-dual large-scale conic programming solver with gpu enhancements.arXiv preprint arXiv:2505.00311, 2025

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.017954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.017954Z digest=sha256:412267ea2cf0471e780789c790978c44bee6553067ef271619ba09f315afac0b

Observation 13043e44-b86d-454d-b3e1-a3d1cf0a2a7e · outbound

This paper cites cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP.jl: A GPU Implementation of Restarted Primal-Dual Hybrid Gradient for Linear Programming in Julia

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.221669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.221669Z digest=sha256:dc77bdef389f5351b2d95308c7a2c3aba50adbce774fb709b866611a9fd8d138

Observation 86ae7d01-ff31-4cda-8697-9c946a456fef · outbound

This paper cites cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language.

Gradient Methods with Online Scaling Part I. Theoretical Foundations cuPDLP-C: A Strengthened Implementation of cuPDLP for Linear Programming by C language

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.373008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.373008Z digest=sha256:6332f806b5093bb0b7156d231b9bbd6f51dae6de4c48946dc492de596ca477b6

Observation 8cb11ec2-df2b-4407-8d46-3c97fdc6e0e8 · outbound

This paper cites Tuning-freestep-size adaptation.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Tuning-freestep-size adaptation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.135518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:13.492046Z digest=sha256:b0593fbedeffe1694959f593a60e29cf5b406963afeda473fc0ca54ddae8d3e6

Observation d537402d-9002-461e-9a01-fd926b69f4ad · outbound

This paper cites Adaptive gradient descent without descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive gradient descent without descent

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.619015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:13.621772Z digest=sha256:10f33bc762b502f11e1111b3f920e28d4911250f6685a1cf3d0ec357f7b80a2a

Observation 690cc81a-7386-4bb0-b398-2d4fd274d894 · outbound

This paper cites Adaptive proximal gradient method for convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive proximal gradient method for convex optimization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:26.349459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:13.774680Z digest=sha256:179630babe580071a3393235cbcfad7278b13e01fdf68d0dea1c3cd25841102e

Observation 405bcadf-473c-4c53-892d-48321954815b · outbound

This paper cites Adaptive Bound Optimization for Online Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive Bound Optimization for Online Convex Optimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:13.932005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:13.932005Z digest=sha256:f7a06be21395393b84005083c8d80fe1bc98ff992e16c9e94be8aab4ddfd209a

Observation 3eea46ed-ab74-43c5-9a9f-6f1a7e1fb9f5 · outbound

This paper cites an unresolved cited work.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:10:26.037314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.074627Z digest=sha256:c43cfa6c15f07e0d4cac297a50b0cda256e36b8b9470ebdd10126071cec116f8

Observation cda48efa-20d9-41cb-84ca-4dcc1c763e04 · outbound

This paper cites Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Linear convergence of first order methods for non-strongly convex optimization.Mathematical Programming, 175:69–107, 2019

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.725934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.243941Z digest=sha256:0686154ddc8faa188356236279212645825d8bb8777a3f6b635ceddb69e13f07

Observation a97f9952-f43a-4b26-ab3b-4bad7f9a6d42 · outbound

This paper cites A method for solving the convex programming problem with convergence rate o (1/k2).

Gradient Methods with Online Scaling Part I. Theoretical Foundations A method for solving the convex programming problem with convergence rate o (1/k2)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.422538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.377838Z digest=sha256:701726937d5de4cc58996ad8e46ad654f76cdf5bf68475b682ceea18abbc3df5

Observation af122e1a-fd8f-45d9-8581-b8e70602774c · outbound

This paper cites Springer Science & Business Media, 2013.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer Science & Business Media, 2013

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:25.146817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.540097Z digest=sha256:058982026b8269ffec502dff38729f73f853260f4c31f744b8aad0c6003bef91

Observation b6fd424a-635e-4a24-b603-d5275a7a25a8 · outbound

This paper cites Springer, 1999.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Springer, 1999

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.847880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.688071Z digest=sha256:ecd1dfa39a826775016e9c52da3783be9a2043e7767c63d168b5add356dcb21d

Observation 073b82e1-a03a-4f53-9384-b6868aad5493 · outbound

This paper cites Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Conic optimization via operator splitting and homogeneous self-dual embedding.Journal of Optimization Theory and Applications, 169:1042–1068,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.596628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:14.887592Z digest=sha256:66dcff1c7fdae4898f98a9fb915aa22906cbdc49392e45ff16d6558b05989a0f

Observation 8fad7c9b-8518-43ef-9b02-cd5c18cbbc9c · outbound

This paper cites Online Learning: A Modern Introduction Using Convex Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Online Learning: A Modern Introduction Using Convex Optimization

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:15.037964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:15.037964Z digest=sha256:6193a56c11217e08227a0dd1bcd220e04543fd5a7429389a6ea01d6b47519c5e

Observation 11c62399-0f76-4282-9aeb-3c8a61671e7f · outbound

This paper cites Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Coin betting and parameter-free online learning.Advances in Neural Information Processing Systems, 29, 2016

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:24.353932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:15.203695Z digest=sha256:604f6a05ddf535fda52797cb25f32351a6dac62c6c23f3eb2139499805127d99

Observation a2d47ced-c099-4894-b9ba-99dab4fdbcd4 · outbound

This paper cites MADA: Meta-adaptive optimizers through hyper-gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations MADA: Meta-adaptive optimizers through hyper-gradient descent

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.997951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:15.404845Z digest=sha256:78a2a7943ff254d32a6d4f3ba69ceeb7e84d6a2ce9548a24cd6f6808f13211d5

Observation 7d629dba-c740-40a2-b598-0f19e07ecefa · outbound

This paper cites Introduction to optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Introduction to optimization

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.506836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:15.614007Z digest=sha256:1958f8bcc8a2b7a8badd082e85e03fe115f4a3ff8a83225dc95386d4e958878f

Observation 0473bbe0-f075-45ac-a88c-d90e57b30df4 · outbound

This paper cites Optimal diagonal precondi- tioning.Operations Research, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Optimal diagonal precondi- tioning.Operations Research, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:23.187929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:15.729222Z digest=sha256:938f78e50ab41124b309ccc48d7fd98a577000d2db1e7db28ae9a676cbdc90f6

Observation e227455b-df14-4f60-8c1d-8ac694052ba0 · outbound

This paper cites Lecture notes on online learning draft, 2009.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Lecture notes on online learning draft, 2009

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.888468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:15.885167Z digest=sha256:0d14fa940b7b4a9e7784fc9d92a544b431e98b69ccaf37515dae855eeed275d1

Observation deef3e93-5088-46c0-b9e4-f524ea99ca99 · outbound

This paper cites On the Convergence of Adam and Beyond.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the Convergence of Adam and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:16.063965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:16.063965Z digest=sha256:4b4ca139a97cdb2ae6f72d4511741cc1c85c47819b0d01569e7fd24c82852c0b

Observation 07821767-069b-4dad-ada2-b1eb6a819236 · outbound

This paper cites Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Greedy quasi-newton methods with explicit superlinear conver- gence.SIAM Journal on Optimization, 31(1):785–811, 2021

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.657859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:16.251549Z digest=sha256:2b3dab569e55124329755bee8e9eba307dfb215f01de70f6399090fa305cde35

Observation fcd11103-fce6-4d6e-9177-588f55060dc1 · outbound

This paper cites New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021.

Gradient Methods with Online Scaling Part I. Theoretical Foundations New results on superlinear convergence of classical quasi-newton methods.Journal of optimization theory and applications, 188:744–769, 2021

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.492178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:16.373687Z digest=sha256:01cd8aafcdcde0fdd4f34d6e555cf850768b84abffa4faaf22103fa73caaead8

Observation 50bcfa29-c047-477c-aa49-627546b67153 · outbound

This paper cites Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Ratesofsuperlinearconvergenceforclassicalquasi-newtonmethods

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:22.203383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:16.529064Z digest=sha256:79f83e5e7a46451c7447c74be9262932674f20b7fa8cb6768ffeda5ddc3e6961

Observation 28a70dfa-4977-4735-8bce-83f2423e0fd4 · outbound

This paper cites Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Convergence analysis of an adaptive method of gradient descent.University of Oxford, Oxford, M

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.869002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:16.744234Z digest=sha256:5d6933420b08038f77528fde6c899446126a6941b32db082c094f0345438ef0a

Observation 86f09eaa-efb1-48da-a833-c69adcb76bb9 · outbound

This paper cites Local gain adaptation in stochastic gradient descent.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Local gain adaptation in stochastic gradient descent

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.591671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:16.855472Z digest=sha256:c812792db7d039b749838a01405ba52dafd0435693ee3a0c13f2bd59aff32b39

Observation 6a17c6f7-b9c3-4b92-853e-850ce34d27e5 · outbound

This paper cites Adapting bias by gradient descent: An incremental version of delta-bar-delta.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adapting bias by gradient descent: An incremental version of delta-bar-delta

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:21.248493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:17.038377Z digest=sha256:988c5279cd63e617184020e1e9fed05fd4b4535e43255fc0a84e454c669a65dc

Observation c4be4a74-e36d-4bec-be13-b0d9120ff377 · outbound

This paper cites No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024.

Gradient Methods with Online Scaling Part I. Theoretical Foundations No-regret dynamics in the fenchel game: A unified framework for algorithmic convex optimization.Mathematical Programming, 205(1):203–268, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.916902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:17.226881Z digest=sha256:334268f89ed0e4127ccf8dfd765fd50ff152cdafa27a089fc25389dc72233cdc

Observation 2b5126e1-b40e-4efb-b46c-ee6866f94cb9 · outbound

This paper cites On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations On the convergence of stochastic gradient descent with bandwidth-based step size.Journal of Machine Learning Research, 24(48):1–49, 2023

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.605338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:17.431702Z digest=sha256:862390818cf82d1983fd932ceaee0446bdf8ca053f090a2bd5a13c502df1dcc8

Observation 81a47c91-72de-40b5-8262-7b20494b7c7e · outbound

This paper cites The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations The Role of Level-Set Geometry on the Performance of PDHG for Conic Linear Optimization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.588000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.588000Z digest=sha256:fa3a3abea282b93da34ac0171efd1ab7fe48a5e81508a093e58cea96ec04e0b3

Observation 09a18a02-c2b9-413e-a58c-08ea1777735d · outbound

This paper cites Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adaptive powerball stochastic conjugate gradient for large-scale learning.IEEE Transactions on Big Data, 9(6):1598–1606, 2023

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.317776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:17.719711Z digest=sha256:4b36a38df744e4adba7a0f93793b082122eaae03fcafd0a4dc96667270dcce09

Observation ccd1611f-fda8-426f-b2c8-293a42ac63b6 · outbound

This paper cites Adam-mini: Use Fewer Learning Rates To Gain More.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adam-mini: Use Fewer Learning Rates To Gain More

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:17.855272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:17.855272Z digest=sha256:8ae252f43bd19fdadb9133251f9cfbb27f3942bdb1fd01f694e0ca0a7a62a837

Observation acf4447c-13d9-48d8-bdf1-9b7663236653 · outbound

This paper cites Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Algorithm 778: L-bfgs-b: Fortran sub- routines for large-scale bound-constrained optimization.ACM Transactions on mathematical software (TOMS), 23(4):550–560, 1997

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:20.001214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:18.009345Z digest=sha256:579b6ed0417f0fc2e1fb681175b4224ac3af6b77ccfc22f28f016da35389efa5

Observation 1caf9fe7-598d-494a-a3d9-884a50a16de7 · outbound

This paper cites Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Adabelief optimizer: Adapting stepsizes by the belief in observed gradients.Advances in neural information processing systems, 33:18795–18806, 2020

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:18.193092Z digest=sha256:7613d580b14630abfbf6afbb33c5a510ada48bf8cddbac78f30813e76986ef8e

Observation f8f4b0a7-8547-4732-9036-7ad4c23feadd · outbound

This paper cites Surrogate losses for online learning of stepsizes in stochastic non-convex optimization.

Gradient Methods with Online Scaling Part I. Theoretical Foundations Surrogate losses for online learning of stepsizes in stochastic non-convex optimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:19.409231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:18.346914Z digest=sha256:e6dfd67a69ea535e3ed25d6274f32d3cb8a765cc809305b570ea30132f306e6f

Observation 9400cacb-1509-4471-be06-d6ca5c05c077 · outbound

This paper cites (cited on 17).

Gradient Methods with Online Scaling Part I. Theoretical Foundations (cited on 17)

Reference 6564

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:10:27.789881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T13:10:12.619868Z digest=sha256:6ce0fc8c71fd852d7dc5becea1c908389029568af1b9cd26647c1b5a9dc42f02

Pith citing papers

Observation ff55bcce-22c1-47d3-8546-124f443adfee · inbound

Enhanced PDHG for Linear Programming with Online Preconditioning cites this paper.

Enhanced PDHG for Linear Programming with Online Preconditioning Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:24.103094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:24.103094Z digest=sha256:c255c29186c6a3660cb1e4d8d155fc637e07ad3509018279e61ecc86b114c0e0

Observation d0b4f173-62f7-41a0-89ae-6cc9a6d201e3 · inbound

Learning to accelerate distributed ADMM using graph neural networks cites this paper.

Learning to accelerate distributed ADMM using graph neural networks Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:46:45.043079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T18:46:36.084583Z digest=sha256:e486eddcb2b709091d5f8e5fd1120dc8d4f0291fcb3cf33b8719fcc74d115c1f

Observation a17a2056-3d84-44f7-a89b-1f954def582f · inbound

Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp cites this paper.

Tight Nonasymptotic Local Convergence of Sinkhorn-Knopp Gradient Methods with Online Scaling Part I. Theoretical Foundations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:38:41.132738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:38:41.132738Z digest=sha256:aa6cdd6ba852f8ac5e8f68c83e2e15f5f68e154e9744801df94d6fb7fcd61811