Pith. sign in

Paper Citation Record · LEDGER

Deconstructing What Makes a Good Optimizer for Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2407.07972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07972 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:37:23.812206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:27:30.162768Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d37440de-1ffc-48ba-8fb9-af58f132c7e8 · inbound

Old Optimizer, New Norm: An Anthology cites this paper.

Old Optimizer, New Norm: An Anthology Deconstructing What Makes a Good Optimizer for Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:27:52.955453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T07:27:52.883335Z digest=sha256:a948b39e81c0d2361fdbc4cf2bc539a244bbaa4c5d34dc6411cbce006ea1a389

Observation 552f658d-4508-401a-9dea-25d0d0515d53 · inbound

Gradient Multi-Normalization for Stateless and Scalable LLM Training cites this paper.

Gradient Multi-Normalization for Stateless and Scalable LLM Training Deconstructing What Makes a Good Optimizer for Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:37:23.812206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:37:23.812206Z digest=sha256:86d589eafcceb3bb50867f6ce1b45d4f9eb5c36d5b7f5bcfadd3ddf139baaea6

Observation c06a0683-cccb-4e75-8f15-627151518656 · inbound

Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness cites this paper.

Sign Operator for Coping with Heavy-Tailed Noise in Non-Convex Optimization: High Probability Bounds Under $(L_0, L_1)$-Smoothness Deconstructing What Makes a Good Optimizer for Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-08T11:35:29.795327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:35:29.795327Z digest=sha256:48b77a282f67ddb5219b6226f39030061df11d40d63fbb2ac49f6b6fb9065146

Observation 0b7b1c12-83bc-45c6-834c-061fb5332e91 · inbound

Better Embeddings with Coupled Adam cites this paper.

Better Embeddings with Coupled Adam Deconstructing What Makes a Good Optimizer for Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:12:33.621237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:12:33.621237Z digest=sha256:fd423cd6b95cfa9bae6b094554ebce21a47ef1eebab5d7014da4f122b5865525

Observation b19eaa3e-d997-4475-af1e-df2e6a124e0c · inbound

Taming LLMs by Scaling Learning Rates with Gradient Grouping cites this paper.

Taming LLMs by Scaling Learning Rates with Gradient Grouping Deconstructing What Makes a Good Optimizer for Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:29.263146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:29.263146Z digest=sha256:bd465fdead2da0f54df27bae0bae395cb8fcb8c8a11a2de6b33dbe7a7de7131e

Observation d6d36155-b1fa-4519-8f90-785c1d93fa03 · inbound

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling cites this paper.

Is your batch size the problem? Revisiting the Adam-SGD gap in language modeling Deconstructing What Makes a Good Optimizer for Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:30.636520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:30.636520Z digest=sha256:2d350675019d0b3cccecbc17e7ec15407e08f106bb447e3e99d26925ee6b5eca

Observation d12a81c7-2229-430d-9787-4b99a6c15a39 · inbound

On Design Principles for Private Adaptive Optimizers cites this paper.

On Design Principles for Private Adaptive Optimizers Deconstructing What Makes a Good Optimizer for Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:09:33.200101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:09:33.200101Z digest=sha256:2d747647ec392f6d296da8017138ec75711c4eada5b035e862ac4df67a2671b0

Observation 7ea7beba-0650-4b10-aaf8-17de7a19d88a · inbound

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling cites this paper.

Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling Deconstructing What Makes a Good Optimizer for Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T09:38:51.648559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:38:51.648559Z digest=sha256:3db707ae290379aca6b8e118a662cae35bf7b3c98a1c91cd4cb4df2648745a8c

Observation e2572600-d266-46aa-9b73-c18d7d38e2fe · inbound

Prototype Transformer: Towards Language Model Architectures Interpretable by Design cites this paper.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.444321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.444321Z digest=sha256:64da1978bad59d041ed9301c98c75aa2c6d706ccb2ae18825b468198b5eee5ad

Observation d38d8f07-0ca0-46f3-863d-c89996a5d9e8 · inbound

Why Muon Outperforms Adam: A Curvature Perspective cites this paper.

Why Muon Outperforms Adam: A Curvature Perspective Deconstructing What Makes a Good Optimizer for Language Models

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:06:44.884155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T07:04:21.012269Z digest=sha256:e2faa25543a25a7369e27ddb27de26622a8ae1a34ab698477e407c9fe7ca4134

Observation d6158db1-28c9-472b-918e-6470b164ddee · inbound

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss cites this paper.

Double Preconditioning (DoPr): Optimization for Test-Time Performance, not Validation Loss Deconstructing What Makes a Good Optimizer for Language Models

Reference 283

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:56:56.098113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T02:35:39.845487Z digest=sha256:0efe300746f0221921e48859a8ad6ca3e1f7c5bf2a94dbad1c3bb0133e5ca524

Observation ac0f7cd7-60a2-45d1-a928-e2242878a79c · inbound

Muon Learns More Robust and Transferable Features than Adam cites this paper.

Muon Learns More Robust and Transferable Features than Adam Deconstructing What Makes a Good Optimizer for Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.164103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:08:30.717799Z digest=sha256:2a7b964c294aea1da1db6143c738dbd293da08eea333ae2982f4e7a665d82d0d

Observation c531ab18-8baf-4ef5-a237-ab573d950f87 · inbound

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers cites this paper.

OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers Deconstructing What Makes a Good Optimizer for Language Models

Reference 139

Resolution
unresolved
no resolver link, observed 2026-07-11T22:10:49.683444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:10:49.683444Z digest=sha256:e56d107c189d32091956b158f581dd827ca8417bac5bd9c86eb3c3f3bbbd93a2

Observation 4e96dacf-3e2f-44a5-9d04-d6acdee852fa · inbound

Muon Meets Mamba: Spectral Optimization for State Space Models cites this paper.

Muon Meets Mamba: Spectral Optimization for State Space Models Deconstructing What Makes a Good Optimizer for Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:27:16.990675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:27:16.990675Z digest=sha256:5b4547d0cdb45f169b7cce74a42ae2215996bf21325aa9481655af38bc6b58e1