Pith. sign in

Paper Citation Record · LEDGER

Resolving Discrepancies in Compute-Optimal Scaling of Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.19146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.19146 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T20:06:23.567012Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T08:24:26.777027Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 34722a42-095f-427d-812c-4673fd61b4f3 · inbound

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient cites this paper.

Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T20:06:23.567012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:06:23.567012Z digest=sha256:15bcee819b1ddd2b761497349e239b46d91ba9a12a46261f46da865b8e83af5b

Observation 774c7d91-55e4-49be-8a31-36fbb86bdd82 · inbound

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection cites this paper.

Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:01:36.166405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:01:36.166405Z digest=sha256:145dfe59948bfd64cacbe3029bcd150cac093a564f544cec32976952d124020a

Observation 405b5901-2f94-4530-892e-a0a92d1feb18 · inbound

Distillation Scaling Laws cites this paper.

Distillation Scaling Laws Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:33:08.872000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:33:08.872000Z digest=sha256:30a5bf1ed899e3903f2233a8d1137efb41a7a2de72604088431d91d1c86068bc

Observation 67dd0421-289c-416f-89ea-365cd9570f7d · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:52:27.048190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:5fe5c331b955f4e9fd0c85e67ff5e11133f78dbc9d1f2a3ca911c7e81fb3cb42

Observation 3655e6a8-347f-4e19-8a15-b59acbea6277 · inbound

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs cites this paper.

Sub-Scaling Laws: On the Role of Data Density and Training Strategies in LLMs Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:56:44.065789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:56:44.065789Z digest=sha256:1076267437b1e8214750b43484ddb20ee3c2a54c1d629ce9a32e355a580d89ae

Observation 33c67b5a-9d66-4a3c-abb8-35214a292016 · inbound

The Art of Scaling Reinforcement Learning Compute for LLMs cites this paper.

The Art of Scaling Reinforcement Learning Compute for LLMs Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T16:29:14.020476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:29:13.954029Z digest=sha256:753e3334eebc4dc2e9989dea3e805bca20bad34690a2b9519456d97be38e0930

Observation 476b2924-8eda-4b34-8fd6-bcd00ef32e7f · inbound

Deriving Neural Scaling Laws from the statistics of natural language cites this paper.

Deriving Neural Scaling Laws from the statistics of natural language Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T03:40:04.086576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:40:04.086576Z digest=sha256:89147f1524420302e772b8e4396bc2770d2c87b0c17edf9cf33b1d114c70378b

Observation d1f9fc82-2e64-42e7-a7ae-822593694116 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.928165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:2a94598d63ab6d493f0bf1e56e1ebf54f477ce166d727bb790d37e6797e13cc1

Observation 0d3ff1a0-82a2-4c11-8b99-2105eb2b147b · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.501686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:629626ed1d590d1f827c8acc24fc55b3cde2a4d533e64899374d637e6f087ee2

Observation 787c5515-67ed-498e-a9cb-6ed590107e5b · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:14.783193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:33:05.952601Z digest=sha256:88cec73de6ab510b095a600fcea31c19a4e02bc1e7b790b567921066da424764

Observation 9b75355a-6218-459a-b27a-92394c8be7e1 · inbound

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions cites this paper.

How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:43.631762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:43.631762Z digest=sha256:8795a263760161771e9c83bd7973a626f5f6b71ad14d19e1a5bebab8911dbdb6

Observation 70f0f246-979c-4d3a-9506-8d9675bd2dab · inbound

On the Nonlinearity of Learning Rate Scaling for LLM Training cites this paper.

On the Nonlinearity of Learning Rate Scaling for LLM Training Resolving Discrepancies in Compute-Optimal Scaling of Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T08:24:26.779384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T08:15:20.191222Z digest=sha256:2f0208d41bc8f4dfda847d79a5167a3e56aefe76e938da92655a6fd7fe2a850a