Pith. sign in

Paper Citation Record · LEDGER

What Language Model to Train if You Have One Million GPU Hours?

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2210.15424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.15424 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:50.770942Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T01:35:21.345442Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 61e030e2-6750-4d34-8d0e-a272c29193ed · inbound

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling cites this paper.

Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling What Language Model to Train if You Have One Million GPU Hours?

Reference 232

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:45:17.998688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-15T17:45:17.540282Z digest=sha256:76a1d405aefb929d7190deb44f8405d944931e6d5b16640c8df07efb2794bb43

Observation 1d76c07f-1ff9-4630-a381-4bb3ff401c88 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models What Language Model to Train if You Have One Million GPU Hours?

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.348948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:deef0a86f5d68d1732701d9f6ee3eb32316ccb00480df3658c62a4bcb796e35a

Observation 988db356-63c8-4340-897a-92518cfc7e01 · inbound

The Falcon Series of Open Language Models cites this paper.

The Falcon Series of Open Language Models What Language Model to Train if You Have One Million GPU Hours?

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:46:10.119349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T09:46:09.701440Z digest=sha256:ec967094186819b003d2215fd65a5b5c2b6e797f6aeac4ef0427ceb7d43eb0c5

Observation 8e9d6a9b-80c6-4f3d-a745-da2f5c871800 · inbound

StarCoder 2 and The Stack v2: The Next Generation cites this paper.

StarCoder 2 and The Stack v2: The Next Generation What Language Model to Train if You Have One Million GPU Hours?

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:28:22.683849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T17:28:22.353355Z digest=sha256:da15b435631e8c0f8e881a15b67a23543567b64b8d60b5d2248b5883581ec8f4

Observation cfd28a2c-8728-4634-9af2-5ed19faba107 · inbound

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN cites this paper.

Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN What Language Model to Train if You Have One Million GPU Hours?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:53:33.472454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:53:33.472454Z digest=sha256:d5995f7f55c7f8d6ed99ff7eb57cd229f345486175f4df0f2d477a50f22176ac

Observation dd4fec3b-5f25-4c4d-860d-e6cf23ba19dc · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model What Language Model to Train if You Have One Million GPU Hours?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.834003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.834003Z digest=sha256:2e4ea01ec42d7766606b3d29bb3e3ead4a440ac3b141b3a3931239827093d9bd

Observation 69231a50-059b-4c37-96f3-12b761480a4a · inbound

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training cites this paper.

SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training What Language Model to Train if You Have One Million GPU Hours?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T20:56:25.892133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:56:25.892133Z digest=sha256:27957d95a0f75370b74389ca8b473708632d88af64319a6a79a25bd7904e97ff

Observation 053e27b0-90d8-4de3-bdbf-9800e74e486f · inbound

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch cites this paper.

LLM360 K2: Building a 65B 360-Open-Source Large Language Model from Scratch What Language Model to Train if You Have One Million GPU Hours?

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:02.378178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:02.378178Z digest=sha256:6a2b0a14fde33c561a1b88b4bbbbed1fdb0c5e70cd26e4ed7277e0c631e6f693

Observation 4df4f656-6100-4cdc-a03e-1fbc6acb3228 · inbound

GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning cites this paper.

GreenMind: A Next-Generation Vietnamese Large Language Model for Structured and Logical Reasoning What Language Model to Train if You Have One Million GPU Hours?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:50.770942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:59:50.770942Z digest=sha256:c26624a6133c8aa8aa7244a3508a2befd0eb5dbad7f5eec79953440cae20478e

Observation a65fb9e5-510a-45de-9c7b-e09761f73eb6 · inbound

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam cites this paper.

SoftSignSGD(S3): An Enhanced Optimizer for Practical DNN Training and Loss Spikes Minimization Beyond Adam What Language Model to Train if You Have One Million GPU Hours?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:10:16.567629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:10:16.567629Z digest=sha256:648d69de9b4ee6573ba418b0a89ce5e55fb80418808e33bcdd14bf190229c6c1