Pith. sign in

Paper Citation Record · LEDGER

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2507.22250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22250 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:59:52.253496Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b89c444b-297d-4afe-a89e-d9536ae460d4 · outbound

This paper cites Nemotron-4 340B Technical Report.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Nemotron-4 340B Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.143341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.143341Z digest=sha256:01b618a5b82ed1a05b2120a506471459c2ce026d992bfc1eecc16da79db17f7b

Observation 02dbe5b1-a4fd-4c6d-b02c-5a6e981c77c0 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.155662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.155662Z digest=sha256:f9238f8b13a8f928400dd090ebabdf8e29e3da6b0979c34cf063402e1c0d89af

Observation 7f6c5223-f9ce-42a6-b4ba-59cf61fbcb53 · outbound

This paper cites Does your data spark joy? Performance gains from domain upsampling at the end of training.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Does your data spark joy? Performance gains from domain upsampling at the end of training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.159243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.159243Z digest=sha256:53eb46bfa38a0f8e05ac17f6a7e28714644109fa50532df2288bb0539ccb8fed

Observation 3a768c44-459e-44ac-b5f3-ef59d77b3083 · outbound

This paper cites Adapting Large Language Models to Domains via Reading Comprehension.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Adapting Large Language Models to Domains via Reading Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.170050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.170050Z digest=sha256:6ec6d925b488e6fc11c3f475c23695fdd23e202060f92df6a6bec9d89bc52a96

Observation 5a60c380-0052-4dfe-843a-a1bd50bc5f5b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.176891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.176891Z digest=sha256:3b29297d663a44084badb892947f6da8128e85d65752263e1070b9da9b28a97c

Observation abc81ed8-9caa-4b79-8a7e-819ee77b1f01 · outbound

This paper cites Data Filtering Networks.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data Filtering Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.179703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.179703Z digest=sha256:5c83b5d034ce873a66678d458dbfcebbed49cb8581ad8b312f4567881f127f18

Observation 008bf5d2-e8fb-4668-9e55-502b98c6b04e · outbound

This paper cites Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.182429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.182429Z digest=sha256:060d3c7737c3aade7a07083380cbaef44337489f1d62ef64b265808da012e2fb

Observation 5827bea8-e2f5-4fc0-ac50-f3017293fa59 · outbound

This paper cites Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.185886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.185886Z digest=sha256:b48cfce34d3f34fbea4638cfb593acdb5a36198be137a2a27600257f35ca988a

Observation 5fe2ea4d-b368-438d-9c34-eff553e03047 · outbound

This paper cites The Llama 3 Herd of Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.188900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.188900Z digest=sha256:cb0591e96f7ff13ae871766aa65e4f9d1ee4303263765d9eaa2a49ba08f334d1

Observation e357fd1e-b583-4fc4-b41d-8863f89cdc53 · outbound

This paper cites Mistral 7B.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.191496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.191496Z digest=sha256:4f83078a4cc5f5d9cad222e6d81ca5e36f1bbd320140d6d63297e4425abcb643

Observation 8e257326-75b0-4b0e-8204-617beb0d6ee4 · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.195074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.195074Z digest=sha256:b2ba42a83ca9f0913c3c87a8fea784c352ec3acb7412d988785ac0ede5cdd00f

Observation 89d5f78d-2df8-44a9-9882-61348e07e846 · outbound

This paper cites Scaling Laws for Neural Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.198222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.198222Z digest=sha256:aa122dbaa9f2bc60413c773845cfe2c99207cf5824b1810335a908a0112c1570

Observation 5ddd77f2-d528-43cb-b4dc-f5824a7ed1da · outbound

This paper cites Downstream Datasets Make Surprisingly Good Pretraining Corpora.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Downstream Datasets Make Surprisingly Good Pretraining Corpora

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.201021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.201021Z digest=sha256:775f1ba68e975d4d15d8980e4c333a3af94a03d0845910780db2ff4da660442a

Observation 9a12fec1-1f79-4c10-8444-b4d7206b91b3 · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.204776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.204776Z digest=sha256:408e0d5c0da40c6566c96eee4b7fa307d18ad668552c45115410f48dae950e9a

Observation 23b9081d-2131-494b-a5fd-80b5a08b5111 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.211197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.211197Z digest=sha256:01f7d7efe1991d3de9cd4c924b73d559db44503a19e32b48af68ef2823c16054

Observation 6119b7e9-03cc-420c-8198-ee7929aee534 · outbound

This paper cites Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.215099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.215099Z digest=sha256:a86f7257fa5b9c69403b6cc7956a2de3dfc563bfb2af1daa23f166ad642121b4

Observation 44c5a5b8-c283-45be-956f-065425973391 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training OLMoE: Open Mixture-of-Experts Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.217911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.217911Z digest=sha256:207aebf59d84e7d6fc60dcfc1100d550afc9054a1f4c5fb9da96dee1c7bf6093

Observation 93cb03ac-6bab-406d-b554-468247dd7f54 · outbound

This paper cites 2 OLMo 2 Furious.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training 2 OLMo 2 Furious

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.220699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.220699Z digest=sha256:78615af784342928979afedac2e97c8a8dccbd0fdb383ea9e5a5077cd276b7f7

Observation 0e3c9fc1-d6c2-4106-9c37-4bac7b08cd4a · outbound

This paper cites Data, Data Everywhere: A Guide for Pretraining Dataset Construction.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data, Data Everywhere: A Guide for Pretraining Dataset Construction

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:59:52.334238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:59:52.223959Z digest=sha256:700ca4327a7a51652bf5eef19802e4bfcc88e9b564ffec43cf92618375308550

Observation 32da7489-51d8-4fe6-8351-1654f7de2af6 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.227860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.227860Z digest=sha256:6d9edad9f85380688c33e5e4873269b58253fa1842047bc6dba0774879471fc9

Observation 6739f4ac-2bbc-4dc4-ac02-d695dcfdcde5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.230921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.230921Z digest=sha256:8298a9393c785aee060e5073f12d6a494fb3b7fa725aa6f4618214aea32f7d6e

Observation 2a8aa4d1-3820-4d9f-9b86-27663c35ac52 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.235178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.235178Z digest=sha256:e93badca373d5d125e247c59f67403c5edaa35ab51488b77cd27eb9fed129115

Observation 183b2ff1-2632-41bc-b5af-4c1a8308a847 · outbound

This paper cites Qwen2.5 Technical Report.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.238055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.238055Z digest=sha256:536123a17d6be5f0a55e9a7ad2f1ff907f3e031e0b995a3fbd9774d77f116e1b

Observation ef31d0b3-fa9a-46c8-b56c-d30b7e1a2402 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.240980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.240980Z digest=sha256:80f48c51b8151789dd4bf64320e43cf73f11afbd3ce734d2daf88c27978deac1

Observation 5b186594-ea56-4032-b466-0f044ac360ab · outbound

This paper cites It is trained with AdamW (Loshchilov & Hutter, 2017), using a sequence length of 8192 tokens and 256 sequences par minibatch, for a total of 2.1M tokens.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training It is trained with AdamW (Loshchilov & Hutter, 2017), using a sequence length of 8192 tokens and 256 sequences par minibatch, for a total of 2.1M tokens

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.754817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:59:52.244309Z digest=sha256:d3a1d42845cfc9c1ecf241a0389176df9aef9b4fe26a44b86ef9233ff59f1ae4

Observation dfef40b5-34fa-4761-8662-c3fc6d9859b2 · outbound

This paper cites an unresolved cited work.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T11:59:52.743284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:59:52.247321Z digest=sha256:af7f78b149ecffc1dd3b9e7685b48820b6285c27c36c02d6b48f6b3fdacd56d1

Observation 9451691a-d894-4242-85e4-8faa3b7bf930 · outbound

This paper cites We also conducted ablations on classifier training, comparing binary classification with regression and exploring up-sampling vs.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training We also conducted ablations on classifier training, comparing binary classification with regression and exploring up-sampling vs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.732655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:59:52.250088Z digest=sha256:638554c13b90ea0f8e697f8920d9ced8762e1e8e1e2b84687623dbfa41796439

Observation b34aa0aa-8ecf-47b2-a85d-be46e000181d · outbound

This paper cites Question:.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Question:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.721221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:59:52.253496Z digest=sha256:866b738285729519cc0ca88d57b47e3d61864fa0219e216af45311af9e524df7

Observation 9dcf3fd6-4d6b-4227-a543-c3c799165923 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 1950

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.163085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.163085Z digest=sha256:e5bc9b66d4ea14c3195e49b4de66b546d5ece192144464cdf2772c5973860c63

Observation 93e6082f-3a78-42ea-b405-969c83a9a2f7 · outbound

This paper cites ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.208204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.208204Z digest=sha256:6f7c97e55a06d5cabd240816a78c2e65ece7dd43c3a1ee79e03cb31e19018456

Observation f5b23fb9-4805-4d89-8d05-f9945896ddbb · outbound

This paper cites Scaling Parameter-Constrained Language Models with Quality Data.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Scaling Parameter-Constrained Language Models with Quality Data

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.166832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.166832Z digest=sha256:3ca7ac0155db9962e867398a40731cfd3bdf7dd0b48e41bb602c96fdf1c69624

Observation 1750ef90-b37e-47b8-aae6-84a7d3c63dfc · outbound

This paper cites Instruction Pre-Training: Language Models are Supervised Multitask Learners.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Instruction Pre-Training: Language Models are Supervised Multitask Learners

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.173406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.173406Z digest=sha256:2bf4720b72f329a4cb1d5b8bee65f395682680aed275bcc697b129dc05b0bb41

Observation 13e3b9ac-5f4b-409c-9d2a-c78bebc28867 · outbound

This paper cites MIND: Math Informed syNthetic Dialogues for Pretraining LLMs.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training MIND: Math Informed syNthetic Dialogues for Pretraining LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.151511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.151511Z digest=sha256:579471e97b9592ec3de9cb9224dea611b00c97f1b6eb62aaa38a83b3949d30e3

Observation 7cccc167-9e29-4ab5-bfb1-bc9ddb9eec1b · outbound

This paper cites DELIFT: Data Efficient Language model Instruction Fine Tuning.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training DELIFT: Data Efficient Language model Instruction Fine Tuning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.147629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.147629Z digest=sha256:3605b722376457d17d75f843f5310bab524e2f36f25a05d4f50530dd60026b71

Pith citing papers

No inbound Pith citation observations are available.