Pith. sign in

Paper Citation Record · LEDGER

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2501.18965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18965 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:48:53.764373Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:40:24.434095Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 524ba4f5-5eea-4805-bd48-9394a5ddabe6 · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.764373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.764373Z digest=sha256:a8b5d148c994d4a1c9468c91d63c494e033f35927773dbff9f2756a94fa019d3

Observation ce80e345-2c2a-46c4-b820-4fbd67f88085 · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.978009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.978009Z digest=sha256:ad51e87059f2d4d1a8ea17da3570ace53dffb3ebafab616c686cf013085142b0

Observation cf95fe02-b66a-4973-9e6e-e17359c23ff9 · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.853745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:88ee8e9fa21eb5c29f3753f31f0fe99ca8703eb8b91bca92d13e19e76ea574c4

Observation 1b620e58-78fa-470b-81b9-c9e4586f43a6 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.438563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:adef5e20cc10f3ac35cd1df4d0fac17deee9137d50e1c8220a227d16d0ff64e6

Observation 2b84bfe1-2116-411a-b4fe-13fde4b8f7f5 · inbound

WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training cites this paper.

WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T08:04:06.432613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:04:06.432613Z digest=sha256:4c190911cffdd892707ed809d0d7af32bb1a18fbd0d8696afb2ca7d64c2350df

Observation 4c08e596-2b3d-4b4c-bc89-faa5e206c143 · inbound

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model cites this paper.

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:16:36.364978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:16:36.364978Z digest=sha256:5042a9e5fc2e870db9ed017aa90501fc21905085d698212498c8eae2a194632e

Observation c397482a-d1dc-4a09-a687-86898400f35b · inbound

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model cites this paper.

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-01T12:16:45.602773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:16:45.602773Z digest=sha256:0b84b3c4c89c7745527586e12e3a3899423b73cf969e66a43c5940c03ad33287

Observation 61cc0a2b-020d-4971-aca3-9a3efae113db · inbound

A Defense of the Quadratic Model cites this paper.

A Defense of the Quadratic Model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:58:11.864874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:58:11.864874Z digest=sha256:3b45e1b281eb3a448d4da73dcb6a590df77bda981ce2facf5ef77f34ac3dd929