Pith. sign in

Paper Citation Record · LEDGER

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency

As of 23 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2601.06649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.06649 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T11:25:09.571487Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4393baa4-451a-4bb1-9e56-ee8a2fef9a44 · outbound

This paper cites Explaining Neural Scaling Laws.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Explaining Neural Scaling Laws

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.522988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.522988Z digest=sha256:903dcdef4ceebffadb85fdd4064109e05a0269f7a9dcda75857be3e1836d5b96

Observation df46a85e-5b9d-4365-ae8a-08bf006d23d2 · outbound

This paper cites Unified Neural Network Scaling Laws and Scale-time Equivalence.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Unified Neural Network Scaling Laws and Scale-time Equivalence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.527105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.527105Z digest=sha256:155ab888fd8bb13abecf2590faa217086b5f445ebe16c3262183380fb73345b3

Observation 560b46eb-1a4c-49a5-87b3-8c6f3308af66 · outbound

This paper cites Language Models are Few-Shot Learners.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.530457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.530457Z digest=sha256:30e0bb0ba950c3ae3a3b7c0e0adbf8b4bec61a86a3c50cb2abddb95777fd6f12

Observation d46e6b39-cccf-4723-a0d9-44aca733766f · outbound

This paper cites K., & Koppula, R.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency K., & Koppula, R

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.533847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.533847Z digest=sha256:2a5c404a13715017354270434ee0e3d0a35d91ff04c6f65ae3e0c5a04e3138d6

Observation 9be43537-dca3-4942-8b62-7bf225639c74 · outbound

This paper cites The rising costs of training frontier AI models.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency The rising costs of training frontier AI models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.536774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.536774Z digest=sha256:3ad72698d98c241d37d049582635ca76c8547c68af25355587375034117db9d0

Observation 6cbd7b7e-00cf-4880-9092-b02329af101c · outbound

This paper cites (2026).Follow up power study.Zenodo.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency (2026).Follow up power study.Zenodo

Reference 6

Resolution
verified exact
doi, observed 2026-08-03T11:28:37.697542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-03T11:25:09.540025Z digest=sha256:21c9d15dbd3c3236a855cc34a74cfa869a6659c22f0bff85821fd1b7b33c0dd8

Observation 5d2829c3-f343-4e93-815e-3e837bda0749 · outbound

This paper cites an unresolved cited work.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.543309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.543309Z digest=sha256:a9a86fb06a9d7e78854d7b840b707662db5fb1d21a54c4d7d677f20cb813fcdd

Observation 34af1852-b6db-4409-ab7a-bb45f63d0733 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Training Compute-Optimal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.546220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.546220Z digest=sha256:b318dfc05bd60f3b791b1c4c00e9ee9e6c220b9f0224836773891249cc6eaedf

Observation ec6cd263-e8f8-4662-8120-67b5e62c6794 · outbound

This paper cites an unresolved cited work.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.549452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.549452Z digest=sha256:98b70811a85218dcb3aa031195abf00c7ab54368c8581c0ca9723098ca42a723

Observation df1854ed-86e2-44ce-bf9f-b221abdf9a65 · outbound

This paper cites Scaling Laws for Neural Language Models.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Scaling Laws for Neural Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.552337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.552337Z digest=sha256:75293038aba54847941ff2b3eeb1dd3bbb17ef8bf03c0ed1674045bd5937bd8d

Observation 8c35ae69-f996-438c-b8e0-023fa7214395 · outbound

This paper cites (2022).Overhead-communication exploration in large-scale machine learning frameworks(Tech.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency (2022).Overhead-communication exploration in large-scale machine learning frameworks(Tech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.555551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.555551Z digest=sha256:c39b8466cf02b416ad0e64b0917f34f43c815613de9ed522f75b91a4c7bf5867

Observation 586d93f0-6620-4dd7-8d48-38a7159a6a97 · outbound

This paper cites A Solvable Model of Neural Scaling Laws.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency A Solvable Model of Neural Scaling Laws

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.558686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.558686Z digest=sha256:a2de6832bc1945f7a07ed6d3b92fd48ed4f5cfdbc962b71b324387128807491b

Observation 19c9e4ea-19e8-4fd7-82ab-82d278e94d87 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.562026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.562026Z digest=sha256:692fcdb0f8ccb6b6c6ec600c15b2d1e5bd9d854451e810c684e8332053bb0957

Observation 92cb1cd5-ba43-4c9d-b888-2ead73b2ced4 · outbound

This paper cites Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.565155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.565155Z digest=sha256:6322390b124b53ee42654a204bf775ad1edc89a0644dbe8d4cbd1bc0dea2f8da

Observation fbc304b8-c642-4630-8fe1-8c32cdb27d01 · outbound

This paper cites Beyond neural scaling laws: beating power law scaling via data pruning.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Beyond neural scaling laws: beating power law scaling via data pruning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.568218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.568218Z digest=sha256:f52c7a421054502be679293a8d884ffc182a1bd894ed15f3d51359291ac8a14e

Observation 97c729b3-34fd-4d99-9081-b058935585ac · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T11:25:09.571487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:25:09.571487Z digest=sha256:50e002f1b1d6c0f3b73836d4b6924ec48d4aed497a6e275be89f2eb285b48e37

Pith citing papers

No inbound Pith citation observations are available.