Pith. sign in

Paper Citation Record · LEDGER

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2507.22250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22250 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:59:52.253496Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b89c444b-297d-4afe-a89e-d9536ae460d4 · outbound

This paper cites Nemotron-4 340B Technical Report.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Nemotron-4 340B Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.143341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.143341Z digest=sha256:d8fb7c12b7c6de541da66dcd809d48a1f6465068636f625bb4029ce57702a3be

Observation 02dbe5b1-a4fd-4c6d-b02c-5a6e981c77c0 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.155662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.155662Z digest=sha256:07f5c388668c00c0eb5d13aa73d0ba07371776fae7b458127f65665021e6c1ea

Observation 7f6c5223-f9ce-42a6-b4ba-59cf61fbcb53 · outbound

This paper cites Does your data spark joy? Performance gains from domain upsampling at the end of training.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Does your data spark joy? Performance gains from domain upsampling at the end of training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.159243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.159243Z digest=sha256:98929e596c122e522b8c6618493df6810c9dd2fdb6576430cd15aeccff11b483

Observation 3a768c44-459e-44ac-b5f3-ef59d77b3083 · outbound

This paper cites Adapting Large Language Models to Domains via Reading Comprehension.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Adapting Large Language Models to Domains via Reading Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.170050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.170050Z digest=sha256:daff29ae6936e2912b87dc1f02d372759724184f4e61ad5b18aa197a378879b3

Observation 5a60c380-0052-4dfe-843a-a1bd50bc5f5b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.176891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.176891Z digest=sha256:04c309ad0dd1a63295cdb29b395a2e3f8751b7fb91da1559423889046b48c190

Observation abc81ed8-9caa-4b79-8a7e-819ee77b1f01 · outbound

This paper cites Data Filtering Networks.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data Filtering Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.179703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.179703Z digest=sha256:b2f8a7f44a9e7d9859d8317576838af39e3f12f23c5c3ccfc8dd74820f494f77

Observation 008bf5d2-e8fb-4668-9e55-502b98c6b04e · outbound

This paper cites Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Paolo Glorioso, Quentin Anthony, Yury Tokpanov, James Whittington, Jonathan Pilault, Adam Ibrahim, and Beren Millidge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.182429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.182429Z digest=sha256:060d3c7737c3aade7a07083380cbaef44337489f1d62ef64b265808da012e2fb

Observation 5827bea8-e2f5-4fc0-ac50-f3017293fa59 · outbound

This paper cites Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Why do small language models underperform? Studying Language Model Saturation via the Softmax Bottleneck

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.185886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.185886Z digest=sha256:96d54815f2170e52d84529d3517dda3f06a9bbc6a5f976e62290fae4e238aea9

Observation 5fe2ea4d-b368-438d-9c34-eff553e03047 · outbound

This paper cites The Llama 3 Herd of Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.188900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.188900Z digest=sha256:c733ef203a8688c80e6c78770c37f77280915f7e9747532b0a77f91fd53a4747

Observation e357fd1e-b583-4fc4-b41d-8863f89cdc53 · outbound

This paper cites Mistral 7B.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Mistral 7B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.191496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.191496Z digest=sha256:69fa542766a435aa2b4549cb61b03e57731ddc5aaeaad62aba42afe92102749b

Observation 8e257326-75b0-4b0e-8204-617beb0d6ee4 · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.195074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.195074Z digest=sha256:8f45eb660ec5e2e66dd508e3df09b75cb54ad8d4ea36e2aaa4486963a15bc589

Observation 89d5f78d-2df8-44a9-9882-61348e07e846 · outbound

This paper cites Scaling Laws for Neural Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Scaling Laws for Neural Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.198222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.198222Z digest=sha256:f21acdb1e888ccdf602cc58cf52eb020512c8647f80477323c5a4930e8801a97

Observation 5ddd77f2-d528-43cb-b4dc-f5824a7ed1da · outbound

This paper cites Downstream Datasets Make Surprisingly Good Pretraining Corpora.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Downstream Datasets Make Surprisingly Good Pretraining Corpora

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.201021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.201021Z digest=sha256:3f679d38286b5669cbb541f674c0ef58ef3e2b729118d2160931289d1931d1d0

Observation 9a12fec1-1f79-4c10-8444-b4d7206b91b3 · outbound

This paper cites A Diversity-Promoting Objective Function for Neural Conversation Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training A Diversity-Promoting Objective Function for Neural Conversation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.204776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.204776Z digest=sha256:e696a7128dd7d726b3edcd2c81dbdff6a07dde6a5c33484785e0e64f0eddf310

Observation 23b9081d-2131-494b-a5fd-80b5a08b5111 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.211197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.211197Z digest=sha256:17416e7a4c9012d0163c88adc9a28869db078ddfd7ed2529ecb30757c88259cd

Observation 6119b7e9-03cc-420c-8198-ee7929aee534 · outbound

This paper cites Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.215099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.215099Z digest=sha256:8b21b747513d6a20e64547bea105ec0b8c436a31563b4f2eb3bd6bcc516fbc40

Observation 44c5a5b8-c283-45be-956f-065425973391 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training OLMoE: Open Mixture-of-Experts Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.217911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.217911Z digest=sha256:0a0c0b2d0bd3718e1e88eee67c21ab0affa471479c276cb013162c9a4c5af6df

Observation 93cb03ac-6bab-406d-b554-468247dd7f54 · outbound

This paper cites 2 OLMo 2 Furious.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training 2 OLMo 2 Furious

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.220699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.220699Z digest=sha256:c3f4e5a8967898882861861c3ccae9e291cd130bcce3fc093580f9495520fed2

Observation 0e3c9fc1-d6c2-4106-9c37-4bac7b08cd4a · outbound

This paper cites Data, Data Everywhere: A Guide for Pretraining Dataset Construction.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data, Data Everywhere: A Guide for Pretraining Dataset Construction

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:59:52.334238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:59:52.223959Z digest=sha256:de575652fd1b4b34802bef5775b9b1852ff866a28612fb99f975f3a13c64c558

Observation 32da7489-51d8-4fe6-8351-1654f7de2af6 · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.227860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.227860Z digest=sha256:6d9edad9f85380688c33e5e4873269b58253fa1842047bc6dba0774879471fc9

Observation 6739f4ac-2bbc-4dc4-ac02-d695dcfdcde5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.230921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.230921Z digest=sha256:8298a9393c785aee060e5073f12d6a494fb3b7fa725aa6f4618214aea32f7d6e

Observation 2a8aa4d1-3820-4d9f-9b86-27663c35ac52 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.235178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.235178Z digest=sha256:67bed550fc7103db68a5ce9cf9ac20bd914bf05066599d539b63729c4142e094

Observation 183b2ff1-2632-41bc-b5af-4c1a8308a847 · outbound

This paper cites Qwen2.5 Technical Report.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.238055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.238055Z digest=sha256:536123a17d6be5f0a55e9a7ad2f1ff907f3e031e0b995a3fbd9774d77f116e1b

Observation ef31d0b3-fa9a-46c8-b56c-d30b7e1a2402 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.240980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.240980Z digest=sha256:6cac00b4ad731392d10075b252b31586aef519141951725bb2dbacab0085adb5

Observation 5b186594-ea56-4032-b466-0f044ac360ab · outbound

This paper cites It is trained with AdamW (Loshchilov & Hutter, 2017), using a sequence length of 8192 tokens and 256 sequences par minibatch, for a total of 2.1M tokens.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training It is trained with AdamW (Loshchilov & Hutter, 2017), using a sequence length of 8192 tokens and 256 sequences par minibatch, for a total of 2.1M tokens

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.754817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:59:52.244309Z digest=sha256:8646c655de6df16bb0c7d026bdc4972fd8ae859d8482a8a6ab9b8588864f6599

Observation dfef40b5-34fa-4761-8662-c3fc6d9859b2 · outbound

This paper cites an unresolved cited work.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T11:59:52.743284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:59:52.247321Z digest=sha256:6582e9bfe3b2469b1d5b077a701896945b7a4e833eb9744865419dad2a3fb677

Observation 9451691a-d894-4242-85e4-8faa3b7bf930 · outbound

This paper cites We also conducted ablations on classifier training, comparing binary classification with regression and exploring up-sampling vs.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training We also conducted ablations on classifier training, comparing binary classification with regression and exploring up-sampling vs

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.732655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:59:52.250088Z digest=sha256:7ceb2b0d51c9d98c4de2026aaa7fc203be7110fd4b5939cfdb8909fcfae20e08

Observation b34aa0aa-8ecf-47b2-a85d-be46e000181d · outbound

This paper cites Question:.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Question:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:59:52.721221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:59:52.253496Z digest=sha256:f0f5855d093cc2a9930b82592553f57a6f59181c75d7e6fa8ba99c09fa8bc1fb

Observation 9dcf3fd6-4d6b-4227-a543-c3c799165923 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 1950

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.163085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.163085Z digest=sha256:e5bc9b66d4ea14c3195e49b4de66b546d5ece192144464cdf2772c5973860c63

Observation 93e6082f-3a78-42ea-b405-969c83a9a2f7 · outbound

This paper cites ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.208204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.208204Z digest=sha256:e9c8ac5ed66a5b1806ce4b7a671b2080109130ed1b630550bf8988407ba34c10

Observation f5b23fb9-4805-4d89-8d05-f9945896ddbb · outbound

This paper cites Scaling Parameter-Constrained Language Models with Quality Data.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Scaling Parameter-Constrained Language Models with Quality Data

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.166832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.166832Z digest=sha256:1e539a553ac098e1b7a44cee9074d48bc3fcb06326141a421b8ff628959d2014

Observation 1750ef90-b37e-47b8-aae6-84a7d3c63dfc · outbound

This paper cites Instruction Pre-Training: Language Models are Supervised Multitask Learners.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Instruction Pre-Training: Language Models are Supervised Multitask Learners

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.173406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.173406Z digest=sha256:bffcce1f4026e944cd3a14bc0439f348b89e8a207699f549ab2b4d26786a3aa0

Observation 13e3b9ac-5f4b-409c-9d2a-c78bebc28867 · outbound

This paper cites MIND: Math Informed syNthetic Dialogues for Pretraining LLMs.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training MIND: Math Informed syNthetic Dialogues for Pretraining LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.151511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.151511Z digest=sha256:19840fc01e79fc8c0289d99661e9283ed220b35cff5f9743cf9ef20b42dff6a0

Observation 7cccc167-9e29-4ab5-bfb1-bc9ddb9eec1b · outbound

This paper cites DELIFT: Data Efficient Language model Instruction Fine Tuning.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training DELIFT: Data Efficient Language model Instruction Fine Tuning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.147629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.147629Z digest=sha256:45ffeb26d3ef2ee05226b08835ab541c99d53dd1b16ad0ea5609df720fb013c1

Pith citing papers

No inbound Pith citation observations are available.