Pith. sign in

Paper Citation Record · LEDGER

Training Trajectories of Language Models Across Scales

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2212.09803.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.09803 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:41.259204Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation eca6dce9-0940-423c-bb4c-a39f29078657 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Trajectories of Language Models Across Scales

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:45:17.903448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:d8506d5be5f5380714b0030637ff32bcddb74fe651c767aaa9d3a92fd1ccb426

Observation 50b8710f-c728-4e18-9c14-b2f87676ecc5 · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling Training Trajectories of Language Models Across Scales

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:45:17.674417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:eb3bbc6f5740dca1689fe06a85ff72fa52501fb87276910a42b9f33bc689824d

Observation 7e2a865d-2e55-4b95-a984-98098a30b8f6 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Training Trajectories of Language Models Across Scales

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.479135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:64370cbdfe8b444259f2dbc7da3dcc47de629dda882d13f7f5e78a9ac9c740f7

Observation 9bd4b689-2dfa-4773-81db-8605e8572abc · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Training Trajectories of Language Models Across Scales

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.501537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:b03459965c5846981bda9846e038fcc910e5555e94cc912577f7ffe20cff35b4

Observation 50842ce7-c236-41b2-a326-fb7d0b61af28 · inbound

How much do language models memorize? cites this paper.

How much do language models memorize? Training Trajectories of Language Models Across Scales

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:41.259204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:41.259204Z digest=sha256:ff214f0b0a25da5a0f77da42a5ce4db10b8a2fb9ed4fdf2f434b8320eda22fe7

Observation 472b5137-f1ae-45ce-90c0-e4df3b90f4dd · inbound

Fairness Dynamics During Training cites this paper.

Fairness Dynamics During Training Training Trajectories of Language Models Across Scales

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.924492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.924492Z digest=sha256:b9a6808b63b9cea9c0698701010c30cbfef92e7ad055b4a325065456ae0c3162

Observation 87b1c8c5-13f9-4c96-aac4-392dacadbba9 · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Training Trajectories of Language Models Across Scales

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:02.700113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:02.700113Z digest=sha256:303725ddb37c8697e9b5e6e480e34bf3cd1bb4d06864068efe41681f0dc6f147

Observation e71ca646-0868-4ca5-873d-b501270a77b0 · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Training Trajectories of Language Models Across Scales

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:32:55.873440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T01:29:14.555216Z digest=sha256:bc7717bcc1fcfddc40d47184e2df48aa509a7325c11463b03b784912011947e4

Observation 0a19b94e-d449-4f28-9d41-7f6768ef0bdb · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Training Trajectories of Language Models Across Scales

Reference 226

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:40:24.922360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T06:39:16.246591Z digest=sha256:c1faf68d983b6f8a676879548a97b8edb869fb16c1487d0bb9d69573cc3759b2