Pith. sign in

Paper Citation Record · LEDGER

Scaling Exponents Across Parameterizations and Optimizers

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2407.05872.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.05872 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T01:03:53.041974Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28c08736-595d-4735-9602-dae2f7e9964a · inbound

GWT: Scalable Optimizer State Compression for Large Language Model Training cites this paper.

GWT: Scalable Optimizer State Compression for Large Language Model Training Scaling Exponents Across Parameterizations and Optimizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:57:36.655891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T05:57:09.276224Z digest=sha256:979e8fa7604a07307d2c379f467ae6b2667c933ace4f173fe0a5e20a479c759e

Observation b70c9695-3fb7-44af-abc1-b35ac67cf3bb · inbound

Avoiding spurious sharpness minimization broadens applicability of SAM cites this paper.

Avoiding spurious sharpness minimization broadens applicability of SAM Scaling Exponents Across Parameterizations and Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:21:20.377243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:21:20.377243Z digest=sha256:fbf22e53da27a715e938fde6600301e4d1629ce26b70aa894e4b9e950acc5a69

Observation a7e52749-50eb-49f8-8199-6388b3bf013b · inbound

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer cites this paper.

Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer Scaling Exponents Across Parameterizations and Optimizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:54:36.874008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:54:36.874008Z digest=sha256:57d047d4908bf83583a157da7d01e4855c783f2ea10c00ed7c717df56c0af46c

Observation 54ecae26-47be-45b3-96f0-58a71145a502 · inbound

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models cites this paper.

Two-Point Deterministic Equivalence for Stochastic Gradient Dynamics in Linear Models Scaling Exponents Across Parameterizations and Optimizers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:17:30.972203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T04:16:04.110552Z digest=sha256:dec5dc26ee988ad45bd1bdc6ae10c7f98571eeae8215216160eeaee9d62d56a3

Observation 2d965bc0-1928-4efa-a0e7-9ba8e7c6afd4 · inbound

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach cites this paper.

Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach Scaling Exponents Across Parameterizations and Optimizers

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:39:41.136195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T15:39:40.845703Z digest=sha256:d34feea65f348f77494411edb113b72ba38a07ac7506161cf6f7c89678116400

Observation 1936ef49-33e9-404f-ae58-372b5fffb8ac · inbound

Practical Efficiency of Muon for Pretraining cites this paper.

Practical Efficiency of Muon for Pretraining Scaling Exponents Across Parameterizations and Optimizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T01:03:53.041974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T01:03:53.041974Z digest=sha256:13a1fd9f0452371bb1b2a6513ae7936d359a0832f97b82f1e7cd531ffc1d820e

Observation f225e47c-d49a-40d2-8c6e-9d2916046295 · inbound

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning cites this paper.

Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning Scaling Exponents Across Parameterizations and Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:20.364789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:20.364789Z digest=sha256:196b32507961f906f731cffa6c2c6448566fbb0a8ca23d3b42c92458f444bb43

Observation 3c8dd48f-a720-4c59-839c-43c18097175c · inbound

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size cites this paper.

Optimal Embedding Learning Rate in LLMs: The Effect of Vocabulary Size Scaling Exponents Across Parameterizations and Optimizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:42.469149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:42.469149Z digest=sha256:8ba73b0d704b0df5c84cdc83950fe5b40a267d9e16a1c2df63233f16e86766b1

Observation 0231a66f-5898-4deb-a78f-3976de64ab0d · inbound

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks cites this paper.

Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks Scaling Exponents Across Parameterizations and Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:48:50.834417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:48:50.834417Z digest=sha256:f95a499b3c958aff04dd20b22ef6bee380869bc0b58dd40df1adeddb00971f31

Observation e0385b0d-7ebe-4c00-8d84-3431dec3bbdc · inbound

Decoupled Relative Learning Rate Schedules cites this paper.

Decoupled Relative Learning Rate Schedules Scaling Exponents Across Parameterizations and Optimizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:11.824040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:14:11.824040Z digest=sha256:fdddba1fbdc6f1341a2ec7a9e6ee13d3f3e3a9fc6fffcb23927a0837554db406

Observation 1af5ec08-b7dd-43b3-acae-d35d9369cca7 · inbound

Feature learning is decoupled from generalization in high capacity neural networks cites this paper.

Feature learning is decoupled from generalization in high capacity neural networks Scaling Exponents Across Parameterizations and Optimizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:17:37.081920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:17:37.081920Z digest=sha256:8f6e60d2b5620b0cc158e40159bbbd7b559125e91744ed2b6d28d3560e4a480f

Observation f93461aa-05fe-4da9-adf6-69f66ffbcde1 · inbound

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model cites this paper.

Theory of Optimal Learning Rate Schedules and Scaling Laws for a Random Feature Model Scaling Exponents Across Parameterizations and Optimizers

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:00:43.299777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T06:58:38.927268Z digest=sha256:4e63b3b4219c463b59a846ba10a70c19bb5754d11f8f783dab28138b5710d0bc

Observation 9102af60-12fc-4a52-b523-61953e1188a0 · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models Scaling Exponents Across Parameterizations and Optimizers

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:21:01.021466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:2fac6b2f062bf792ae0c2cef3717b75e14b4bc8096dc5cb62816a485c8aea0bc

Observation a223976d-c983-44ac-85da-6472e97be7d7 · inbound

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions cites this paper.

C-voting: Confidence-Based Test-Time Voting without Explicit Energy Functions Scaling Exponents Across Parameterizations and Optimizers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:05:24.996003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:02:52.920735Z digest=sha256:3b6011007afc73a3c12d5205ae0e5794307f66213872e53be95de645a79d32de

Observation 805efe36-d7e0-4353-96ba-7b0aba39edc1 · inbound

Learning Rate Transfer in Normalized Transformers cites this paper.

Learning Rate Transfer in Normalized Transformers Scaling Exponents Across Parameterizations and Optimizers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:51:28.083967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T09:01:45.365126Z digest=sha256:d4617fc0da286b5262698161480b89d1ae15330664b79c54a96034744e008239

Observation 2e55164d-4c29-441a-a4e6-ec14cbc6e6b6 · inbound

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data cites this paper.

Unlearning with Asymmetric Sources: Improved Unlearning-Utility Trade-off with Public Data Scaling Exponents Across Parameterizations and Optimizers

Reference 141

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:57:21.512477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T05:56:38.042978Z digest=sha256:0ab8dd447cdfbbf5019854856c24f4f1b3f7e1fc28e6a20396b910789f0193cb

Observation d85931f2-99b1-4294-9198-7a845be45fa1 · inbound

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization cites this paper.

How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization Scaling Exponents Across Parameterizations and Optimizers

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:49:44.837389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T04:45:20.091598Z digest=sha256:7305b89399a1082509e806e3f3516d6c518a8249fa748e85e75abf01769ac8fb

Observation fa8b0839-ae38-430c-8d99-827db6b9e8fc · inbound

GQA-{\mu}P: The maximal parameterization update for grouped query attention cites this paper.

GQA-{\mu}P: The maximal parameterization update for grouped query attention Scaling Exponents Across Parameterizations and Optimizers

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T16:37:39.842088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:35:40.231293Z digest=sha256:ca6dfa7b74b358b0666f6b9902c4057873e703d3f0d2b8be30c24d7723d1475c

Observation 08b3d673-2af3-4b48-a3b2-09dd84cd135e · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T15:39:33.195350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T15:35:51.654392Z digest=sha256:f6aebcb51a0a60da1d9bb649793ae43bd59d4b91a711de1edadaf1c82edc358f

Observation d0848018-1c61-4497-a063-900a48ac9109 · inbound

Statistical Properties of Training & Generalization cites this paper.

Statistical Properties of Training & Generalization Scaling Exponents Across Parameterizations and Optimizers

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:57:25.317838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T21:51:13.457071Z digest=sha256:ab61244b5b5e69e1be3371ff23f293aed865c85937e2525bbbbd7fbdbeb9ef1a

Observation 38c7f7b7-d4c8-4b42-8683-bb38aed6a0d7 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.676652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:f4381ffe06f8b02d83232ebf37c49743382bb8d1728964c52b2c3392f83335f1

Observation 44235205-ee6f-4780-be28-54aab9ef7635 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Exponents Across Parameterizations and Optimizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:06.198299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:06.198299Z digest=sha256:422f34f09ee5a1fb81e5320cad6c4d4433062e381c68cc381a0dfbe0b8341b86

Observation 0ca474ff-08db-4668-8d38-4b65e51d4f23 · inbound

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales cites this paper.

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales Scaling Exponents Across Parameterizations and Optimizers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T06:45:17.238514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:45:17.238514Z digest=sha256:caf6212f8cb392132cf1227b6e8b4a2b957a83215bc5b456fe974dfdaec18408