Pith. sign in

Paper Citation Record · LEDGER

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 8 inbound Pith citation observations for arXiv:2501.18965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18965 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:56:50.876409Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:48:53.764373Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:40:24.434095Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6dd21981-5183-4264-a3d5-78c4a7bbfeeb · outbound

This paper cites Why you don't overfit, and don't need Bayes if you only train for one epoch.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Why you don't overfit, and don't need Bayes if you only train for one epoch

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-09T21:56:51.068688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.743708Z digest=sha256:70e3a719163925f51b3e2b43169f90444a5e2099e407736ce8e0e97a459eedbb

Observation 679b2f08-7bae-4f38-a898-433e752cb21b · outbound

This paper cites For validation set metrics, we display a running average over five epoch in thick to smoothen the plot, and the original data in thin.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training For validation set metrics, we display a running average over five epoch in thick to smoothen the plot, and the original data in thin

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T21:56:51.164759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.856860Z digest=sha256:15222386564f37a7b788fb23ce5df20899e2d16d027953f1b29dccee12e869d4

Observation a3b291a1-492c-4fff-8e31-19e824404fff · outbound

This paper cites Loss Landscape Characterization of Neural Networks without Over-Parametrization.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Loss Landscape Characterization of Neural Networks without Over-Parametrization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.773093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.773093Z digest=sha256:ba3bdf1432339938608a00389e6fcb297a804e98be30d36957e34b9321f28677

Observation 63c08d1b-7f65-464c-9cc5-dde6561344ff · outbound

This paper cites Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.797601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.797601Z digest=sha256:71fb962169fb24c96ad89d08851931ad1f3e328103947d9904fe250123b687a1

Observation 16e15bbc-102e-42d2-ac99-0f888910f114 · outbound

This paper cites No More Adam: Learning Rate Scaling at Initialization is All You Need.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training No More Adam: Learning Rate Scaling at Initialization is All You Need

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.802232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.802232Z digest=sha256:d7134136a05597d9c7044c6bd9dd4b607586538ff07528a8a9dd32b91b423fd2

Observation 3e8e6c92-2518-48cc-b0f9-d75f5f17155d · outbound

This paper cites Exact convergence rate of the last iterate in subgradient methods.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Exact convergence rate of the last iterate in subgradient methods

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.806915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.806915Z digest=sha256:fbe102c020a05662dd18299b07eba0e664da474ea0cf6e10a3e22be8ca1d2570

Observation 8b6da73c-238a-49bb-b9df-bb02ea17d8e7 · outbound

This paper cites • Appendix B: supplementary information on our experiments.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training • Appendix B: supplementary information on our experiments

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.305651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.811658Z digest=sha256:061123d0156e58ba0bf425873f93e9f2ab16dd469a77976f8fea39b78c0540ab

Observation 89f7914a-b082-4f20-847d-b31863ba5220 · outbound

This paper cites 2, but with Ωt from (11) The bound on the best-so-far bound has a very different shape of the last-iterate bound.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training 2, but with Ωt from (11) The bound on the best-so-far bound has a very different shape of the last-iterate bound

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.290719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.816379Z digest=sha256:982864288abb524298b27062bade960616665bd7d951f590edc3d33f0f1ae41d

Observation e643bf3e-f10b-4357-b32d-ae064670ba1c · outbound

This paper cites Convergence is plotted with the optimal base learning-rate γ⋆ (chosen individually for each schedule).

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Convergence is plotted with the optimal base learning-rate γ⋆ (chosen individually for each schedule)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.261522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.825923Z digest=sha256:2b4f212f868be1e978c2ff532a65f913f51c07fc975e78ac1a00d1054480add4

Observation 0ed11f91-b193-4379-9e99-d2f2333c5d1a · outbound

This paper cites Compare to Figure A1 in Hoffmann et al.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Compare to Figure A1 in Hoffmann et al

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.246198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.830267Z digest=sha256:6049ec74e0265b09648447ae148c907f9dcf2b34342df6a244302277a85ab5fb

Observation 1f0b1871-10d9-4c91-a674-e23fe81577af · outbound

This paper cites Details on Experiments in Fig.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Details on Experiments in Fig

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.232415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.835059Z digest=sha256:63cd6ea1a8483b36326e1baaf3806f6dcf2cabe53ca9199969873da6654b02c4

Observation aa67bb27-c4a6-4553-a890-8a8ebdc4f619 · outbound

This paper cites Dark grey marks the bound of the constant schedule.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Dark grey marks the bound of the constant schedule

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.192616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.848873Z digest=sha256:0318f71242e24f5a249dccbd43445b79752f5e3c5ed7d34337e3904b5567962c

Observation 70898871-83c4-4d1c-a77e-69c6cef1d03e · outbound

This paper cites We observe the same characteristic drop of the loss for wsd, as well as matching performance of wsd and cosine.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training We observe the same characteristic drop of the loss for wsd, as well as matching performance of wsd and cosine

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.148945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.860706Z digest=sha256:61f5e5a5841b43f8309776209fde1760b2bd8e6a18b8a5cc6f0f5b4470ddad18

Observation b1d64691-5d65-47b3-a0a8-76e3c4edb4e6 · outbound

This paper cites an unresolved cited work.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:56:51.132093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.864678Z digest=sha256:e62f1e2526df768e2c5432c99d6e6a62cbdfa9644f19a338d4fd8f5cb523cc90

Observation 96b1f77e-41a0-4eca-b531-c2ed1d4c5045 · outbound

This paper cites an unresolved cited work.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:56:51.116057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.868890Z digest=sha256:4d6ed0c3d28ee25b4095d5713962696e62a0b2b0710fe957218ae5ecfa8352d7

Observation 5ecfb48b-b270-45d3-b404-f23bf64a44fb · outbound

This paper cites Theorem 3.1 follows from applying Theorem E.2 with ˆηt := γηt.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Theorem 3.1 follows from applying Theorem E.2 with ˆηt := γηt

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.099589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.872663Z digest=sha256:29d62182f5e002ce63e1d71f63ce698ce0aca576b2a59fa14a0335e03df95345

Observation 5dcbfc1e-52df-47dc-a6ef-1c0aa8709ad4 · outbound

This paper cites From the (generalized) Cauchy-Schwarz inequality combined with Young’s inequality, we have s3 ≤ µ 2 ∥xt+1 − xt∥2 + η2 t 2µ ∥gt∥2 ∗.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training From the (generalized) Cauchy-Schwarz inequality combined with Young’s inequality, we have s3 ≤ µ 2 ∥xt+1 − xt∥2 + η2 t 2µ ∥gt∥2 ∗

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.084392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.876409Z digest=sha256:635a5081f167354c54024646259cd304725a9733132a17f8af7b0f234c620438

Observation bc4820f7-d152-4063-b115-95492681ba8d · outbound

This paper cites an unresolved cited work.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work

Reference 400

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:56:51.276044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.820993Z digest=sha256:8d9ff7acc597119ce77d68ecbfec220ee4fb32d3f6a29a5d1a8d6ce24706bd0c

Observation 40e4be6d-56a9-4089-8150-8e11153d32d5 · outbound

This paper cites Training trajectories, mini-batch losses and the curious role of the learning rate.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Training trajectories, mini-batch losses and the curious role of the learning rate

Reference 1970

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.782654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.782654Z digest=sha256:c8813e7f4140bdc6a301c7d3f65bd56713dbff6074663fc9a22a98475fe8cc3c

Observation 3915129f-23cd-42b6-9143-eec1894a883f · outbound

This paper cites Optimal Linear Decay Learning Rate Schedules and Further Refinements.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Optimal Linear Decay Learning Rate Schedules and Further Refinements

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.758379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.758379Z digest=sha256:10e96fe69f2211a3e5574ff00a4c04cd19954f0f08738187998e2054d6a42ec8

Observation 8d100234-d5ee-4cf5-ae7b-91a8a8b4af6a · outbound

This paper cites Chinchilla Scaling: A replication attempt.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Chinchilla Scaling: A replication attempt

Reference 2003

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.749126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.749126Z digest=sha256:e8f4e769f1858526a007dba41d813ab63f19c6001020bc364a0c29591d7cf358

Observation 8840c951-3c13-4241-86cc-eb1f689b32d8 · outbound

This paper cites Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.787977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.787977Z digest=sha256:6c376e9a4070f646d4f2d66dd28410cbd5a47e053c15b42fdf59cd35f9304880

Observation 5604a1e5-f2ac-4fce-8373-30cd4b0d2f00 · outbound

This paper cites We train all models with SGD with heavy-ball momentum.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training We train all models with SGD with heavy-ball momentum

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.179031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.852913Z digest=sha256:7b94cc12ea7ccfa1f933a29475ca75d89a06f7f3b5f6dafed25cab46a0e53094

Observation 163e6581-3c76-4b8c-9004-a3c6b2493bb0 · outbound

This paper cites Deep residual learning for image recognition.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Deep residual learning for image recognition

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.768440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.768440Z digest=sha256:855803e7ccbdadec7268086d3fc24e96336f35259ad1883bcfbbf02f986d55bf

Observation 79b3def1-f6fd-4824-bf96-ced9b94521f0 · outbound

This paper cites For all further details we refer to Hägele et al.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training For all further details we refer to Hägele et al

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.218222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.839609Z digest=sha256:be6de7701cebadecb2e604723572776c8e6ba09f04c2d944d2750b35cc0b3cc0

Observation 4dbf9e95-152d-4ef3-8b95-50e5dd62eaca · outbound

This paper cites Rockafellar, R.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Rockafellar, R

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T21:56:51.319527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.777860Z digest=sha256:3427483aa5ce18eaf8f141f6ada8f10008935d345c8e2c8d3035c9c6df4fa95a

Observation 27202c2b-dd3a-46fe-81e5-8d96359ea414 · outbound

This paper cites an unresolved cited work.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-09T21:56:51.344700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.753704Z digest=sha256:c7411895c3bfe5436cd72c77559b1dc31dc7a40c80e29b29dfeaa59ab5107a18

Observation c47b089a-3349-4086-9f92-679beef3cf3d · outbound

This paper cites More concretely, assume we have trained a model of sizeN1 for D1 tokens.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training More concretely, assume we have trained a model of sizeN1 for D1 tokens

Reference 2022

Resolution
malformed identifier
raw_fallback, observed 2026-08-09T21:56:51.205598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T21:56:50.844094Z digest=sha256:2e918d473089390f473a9678f91d42728440fced9c58b471c822c3d12d7a2e7e

Observation b818474e-760d-41e5-b3cd-ab77679033ef · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training LLaMA: Open and Efficient Foundation Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.792621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.792621Z digest=sha256:db483aafc52563748ed2de2ca349b523596aafd003cc1aa6526e51b72db8b6d9

Observation 2e7eb3c5-c3a8-4d26-a2c4-dac886695c21 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T21:56:50.763500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:56:50.763500Z digest=sha256:6682ab1c24e3940db95d6b388ff6e2333f112976c598f71a45dc78090021f15e

Pith citing papers

Observation 524ba4f5-5eea-4805-bd48-9394a5ddabe6 · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:53.764373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:53.764373Z digest=sha256:8feeb643f66561665f648f7069b587e880039146d29efe4739d3fe336ed6b914

Observation ce80e345-2c2a-46c4-b820-4fbd67f88085 · inbound

Why Do We Need Warm-up? A Theoretical Perspective cites this paper.

Why Do We Need Warm-up? A Theoretical Perspective The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.978009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:38:59.978009Z digest=sha256:e0d16e3b697969ee3088b1f8cda47297169f0e05dc7502d8367a0b7a3a32d3eb

Observation cf95fe02-b66a-4973-9e6e-e17359c23ff9 · inbound

Optimistic Dual Averaging Unifies Modern Optimizers cites this paper.

Optimistic Dual Averaging Unifies Modern Optimizers The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:21.853745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:52:16.805180Z digest=sha256:41dad701866ab5d41b1e0bba3db7f7f06ab48df8b01e0bd717ff4798bd50b06c

Observation 1b620e58-78fa-470b-81b9-c9e4586f43a6 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.438563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:fb83dae1181b89411cb57a3409888b75dc421986d974e844bc8453cfae2d9394

Observation 2b84bfe1-2116-411a-b4fe-13fde4b8f7f5 · inbound

WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training cites this paper.

WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T08:04:06.432613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T08:04:06.432613Z digest=sha256:2d7821015328003843b8a324a814bbc851e8a9a054f327b9ba7a937a17e7126f

Observation 4c08e596-2b3d-4b4c-bc89-faa5e206c143 · inbound

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model cites this paper.

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T12:16:36.364978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:16:36.364978Z digest=sha256:0afb9296fd4a13a668180383185c82a332256e89be281affeac5d127f88b21c3

Observation c397482a-d1dc-4a09-a687-86898400f35b · inbound

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model cites this paper.

AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 154

Resolution
unresolved
no resolver link, observed 2026-08-01T12:16:45.602773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:16:45.602773Z digest=sha256:1f3002526c090785b032c9bb74e799f52114e33ae21c2068e628867c8dd9b367

Observation 61cc0a2b-020d-4971-aca3-9a3efae113db · inbound

A Defense of the Quadratic Model cites this paper.

A Defense of the Quadratic Model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T06:58:11.864874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:58:11.864874Z digest=sha256:e9d16d322c70d9b309a9f8b7523cefb293183f6738b19bf7224c57d616007ec4