Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2407.07972.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.812206Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:27:30.162768Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d37440de-1ffc-48ba-8fb9-af58f132c7e8 · inbound
Old Optimizer, New Norm: An Anthology Deconstructing What Makes a Good Optimizer for Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 552f658d-4508-401a-9dea-25d0d0515d53 · inbound
Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c06a0683-cccb-4e75-8f15-627151518656 · inbound
Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness Deconstructing What Makes a Good Optimizer for Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b7b1c12-83bc-45c6-834c-061fb5332e91 · inbound
Better Embeddings with Coupled Adam Deconstructing What Makes a Good Optimizer for Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · inbound
Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · inbound
Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12a81c7-2229-430d-9787-4b99a6c15a39 · inbound
On Design Principles for Private Adaptive Optimizers Deconstructing What Makes a Good Optimizer for Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea7beba-0650-4b10-aaf8-17de7a19d88a · inbound
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Deconstructing What Makes a Good Optimizer for Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2572600-d266-46aa-9b73-c18d7d38e2fe · inbound
Prototype Transformer: Towards Language Model Architectures Interpretable by Design Deconstructing What Makes a Good Optimizer for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38d8f07-0ca0-46f3-863d-c89996a5d9e8 · inbound
Why Muon Outperforms Adam: A Curvature Perspective Deconstructing What Makes a Good Optimizer for Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d6158db1-28c9-472b-918e-6470b164ddee · inbound
Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Deconstructing What Makes a Good Optimizer for Language Models
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac0f7cd7-60a2-45d1-a928-e2242878a79c · inbound
Muon Learns More Robust and Transferable Features than Adam Deconstructing What Makes a Good Optimizer for Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c531ab18-8baf-4ef5-a237-ab573d950f87 · inbound
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Deconstructing What Makes a Good Optimizer for Language Models
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e96dacf-3e2f-44a5-9d04-d6acdee852fa · inbound
Muon Meets Mamba: Spectral Optimization for State Space Models Deconstructing What Makes a Good Optimizer for Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.