Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T21:56:50.876409Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 8 inbound Pith citation observations for arXiv:2501.18965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-09T21:56:50.876409Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:48:53.764373Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-25T05:40:24.434095Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6dd21981-5183-4264-a3d5-78c4a7bbfeeb · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Why you don't overfit, and don't need Bayes if you only train for one epoch
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 679b2f08-7bae-4f38-a898-433e752cb21b · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training For validation set metrics, we display a running average over five epoch in thick to smoothen the plot, and the original data in thin
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3b291a1-492c-4fff-8e31-19e824404fff · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Loss Landscape Characterization of Neural Networks without Over-Parametrization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63c08d1b-7f65-464c-9cc5-dde6561344ff · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Rethinking Conventional Wisdom in Machine Learning: From Generalization to Scaling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e15bbc-102e-42d2-ac99-0f888910f114 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training No More Adam: Learning Rate Scaling at Initialization is All You Need
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e8e6c92-2518-48cc-b0f9-d75f5f17155d · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Exact convergence rate of the last iterate in subgradient methods
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6da73c-238a-49bb-b9df-bb02ea17d8e7 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training • Appendix B: supplementary information on our experiments
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89f7914a-b082-4f20-847d-b31863ba5220 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training 2, but with Ωt from (11) The bound on the best-so-far bound has a very different shape of the last-iterate bound
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e643bf3e-f10b-4357-b32d-ae064670ba1c · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Convergence is plotted with the optimal base learning-rate γ⋆ (chosen individually for each schedule)
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ed11f91-b193-4379-9e99-d2f2333c5d1a · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Compare to Figure A1 in Hoffmann et al
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f0b1871-10d9-4c91-a674-e23fe81577af · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Details on Experiments in Fig
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aa67bb27-c4a6-4553-a890-8a8ebdc4f619 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Dark grey marks the bound of the constant schedule
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 70898871-83c4-4d1c-a77e-69c6cef1d03e · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training We observe the same characteristic drop of the loss for wsd, as well as matching performance of wsd and cosine
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b1d64691-5d65-47b3-a0a8-76e3c4edb4e6 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 96b1f77e-41a0-4eca-b531-c2ed1d4c5045 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5ecfb48b-b270-45d3-b404-f23bf64a44fb · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Theorem 3.1 follows from applying Theorem E.2 with ˆηt := γηt
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5dcbfc1e-52df-47dc-a6ef-1c0aa8709ad4 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training From the (generalized) Cauchy-Schwarz inequality combined with Young’s inequality, we have s3 ≤ µ 2 ∥xt+1 − xt∥2 + η2 t 2µ ∥gt∥2 ∗
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc4820f7-d152-4063-b115-95492681ba8d · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work
Reference 400
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40e4be6d-56a9-4089-8150-8e11153d32d5 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Training trajectories, mini-batch losses and the curious role of the learning rate
Reference 1970
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3915129f-23cd-42b6-9143-eec1894a883f · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Optimal Linear Decay Learning Rate Schedules and Further Refinements
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d100234-d5ee-4cf5-ae7b-91a8a8b4af6a · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Chinchilla Scaling: A replication attempt
Reference 2003
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8840c951-3c13-4241-86cc-eb1f689b32d8 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5604a1e5-f2ac-4fce-8373-30cd4b0d2f00 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training We train all models with SGD with heavy-ball momentum
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 163e6581-3c76-4b8c-9004-a3c6b2493bb0 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Deep residual learning for image recognition
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b3def1-f6fd-4824-bf96-ced9b94521f0 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training For all further details we refer to Hägele et al
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4dbf9e95-152d-4ef3-8b95-50e5dd62eaca · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Rockafellar, R
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27202c2b-dd3a-46fe-81e5-8d96359ea414 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c47b089a-3349-4086-9f92-679beef3cf3d · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training More concretely, assume we have trained a model of sizeN1 for D1 tokens
Reference 2022
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b818474e-760d-41e5-b3cd-ab77679033ef · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training LLaMA: Open and Efficient Foundation Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7eb3c5-c3a8-4d26-a2c4-dac886695c21 · outbound
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524ba4f5-5eea-4805-bd48-9394a5ddabe6 · inbound
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce80e345-2c2a-46c4-b820-4fbd67f88085 · inbound
Why Do We Need Warm-up? A Theoretical Perspective The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf95fe02-b66a-4973-9e6e-e17359c23ff9 · inbound
Optimistic Dual Averaging Unifies Modern Optimizers The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1b620e58-78fa-470b-81b9-c9e4586f43a6 · inbound
Anytime Training with Schedule-Free Spectral Optimization The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2b84bfe1-2116-411a-b4fe-13fde4b8f7f5 · inbound
WSqD: A Horizon-Free Learning Rate Schedule for Large Model Training The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c08e596-2b3d-4b4c-bc89-faa5e206c143 · inbound
AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c397482a-d1dc-4a09-a687-86898400f35b · inbound
AGNFormer I: Reconstruction of AGN spectra using a probabilistic transformer model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 154
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61cc0a2b-020d-4971-aca3-9a3efae113db · inbound
A Defense of the Quadratic Model The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.