Pith. sign in

Paper Citation Record · LEDGER

Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2403.16952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.16952 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:26:07.021278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:19:57.811085Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 55cf0bf7-c786-460d-8a03-aada044e2867 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T18:00:53.498646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:2063c55917eabb178f1992156d3fda6165839d46e7c5c882497701750a48e32a

Observation 740895cd-dda9-4473-8fc9-dbb30f9541cd · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:07.021278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:07.021278Z digest=sha256:bfa603046990e5fb7912b8e812dd0605e78ea6c2fb6cd1a429f67f8bd0c32ea1

Observation 6573b0bb-16db-45b0-b561-b91f1d432115 · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.440209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.440209Z digest=sha256:a64eb5d430af59305fdf1ff2666882cf98722caf5e2b763468b43280a0b4890c

Observation bccb8b40-7571-460c-a0f2-cfd7d46edc43 · inbound

Organize the Web: Constructing Domains Enhances Pre-Training Data Curation cites this paper.

Organize the Web: Constructing Domains Enhances Pre-Training Data Curation Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 657

Resolution
unresolved
no resolver link, observed 2026-08-07T18:33:11.510347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:33:11.510347Z digest=sha256:2952f12eb707b54fc6ede170dbc0af5ecf4f5ace1ca31fa743c3e2891b9d4fc9

Observation f3992b5e-320c-44e4-97a6-1b4e7a591ed3 · inbound

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training cites this paper.

MegaScale-Data: Scaling Dataloader for Multisource Large Foundation Model Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:15:09.608695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T21:12:22.201810Z digest=sha256:53a9a3ab51c66a8ae07d1262589bac123b6c44164010b05a1e94cc2a56d11d44

Observation dccf1a4d-75eb-4da0-b3d6-2ca9c8f997b2 · inbound

Merge to Mix: Mixing Datasets via Model Merging cites this paper.

Merge to Mix: Mixing Datasets via Model Merging Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:39.052347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:39.052347Z digest=sha256:3ac85c914254163dfd4e6c74ce06716e14775879fb386598192e16864ed4a2ab

Observation eb873359-d146-41a6-a21c-3727d824a5d8 · inbound

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives cites this paper.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.384371Z digest=sha256:5ef0d55c019962c4056f8fded07358102685ce2c563d5b807758fa60860fcd8f

Observation 20fa538f-7c97-471f-a9e3-a64eb288af94 · inbound

Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning cites this paper.

Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:26.870092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:26.870092Z digest=sha256:508b2216be0f88ab81ab89b4e5cad3c4dd4af380a224933b90689b5e5d0c47a5

Observation a6e03e86-85cd-4257-912e-4fb29755758a · inbound

Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning cites this paper.

Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:44.851956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:44.851956Z digest=sha256:9158db64014b0be9ad75657930cc2a8659227d127e1b3867b551646a8fa46aca

Observation 191b6167-856d-46d8-9199-6fe42cb3ccd3 · inbound

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning cites this paper.

MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:33.456884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:33.456884Z digest=sha256:495c2444f48a44013b5192f54ab897169387247cba379efbb41b02e190e06b0f

Observation 024a254e-8ef7-4f2c-b59d-0c32fcacb218 · inbound

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training cites this paper.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.793759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.793759Z digest=sha256:20cc67504f8b657ffd01bdf4065947e39772973242de5e6289613d7e229e44e5

Observation a2699ca4-c700-4482-b3f7-f4859ccf8536 · inbound

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training cites this paper.

Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:37:00.829310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T03:36:50.366757Z digest=sha256:9a42f8f34476c16515af5d50f90ac0875c4dad449041cb47428ab031ef429007

Observation ef31d0b3-fa9a-46c8-b56c-d30b7e1a2402 · inbound

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training cites this paper.

Using Scaling Laws for Data Source Utility Estimation in Domain-Specific Pre-Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:59:52.240980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:59:52.240980Z digest=sha256:80f48c51b8151789dd4bf64320e43cf73f11afbd3ce734d2daf88c27978deac1

Observation 486b6b78-bd24-4905-841b-b6487ca3b1de · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.614578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.614578Z digest=sha256:0839a9aff9cf284119a2a77a3e15e9f1c9b7b754df634879a16c768bcedd77f6

Observation 89741ca0-f45b-4bd2-b4e4-1bafc3b8a197 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:06.999694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:877be0782d51405025dfd2aa6436aac79be8fa0d0ef4dcf761272d3aeb31553b

Observation fef432e1-3217-462a-b59d-b0ca16be6637 · inbound

Knowledge Transfer Scaling Laws for 3D Medical Imaging cites this paper.

Knowledge Transfer Scaling Laws for 3D Medical Imaging Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:45:51.622263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T01:30:50.324360Z digest=sha256:21cc4fb6adfe580aac5693879ae592f0c10abe2c9a0c1c3197137d2bc4cfa9c7

Observation 0fb22db8-e891-4c72-bfaf-c83f20e2eb62 · inbound

On the Invariance and Generality of Neural Scaling Laws cites this paper.

On the Invariance and Generality of Neural Scaling Laws Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:56.035239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:34:14.087140Z digest=sha256:84a529143b6fda9426e735db8ce25c0f7f2c37d4f3c50eb895960710a9c49f02

Observation 854f6c11-7a77-439d-b3c0-053bd22b24d3 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.954431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:441f3202d343cb15c209aa284cb53e4f4fce2707ae7d5ff3c157ab2e5bf9202d

Observation a69322f0-547a-4290-83bf-0595bd573e66 · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.486738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:b030918af996adc4fcfaba359b8b926e3c69b4b98de49997939d4786a995f51b

Observation 1e6235db-4944-4a0b-bcb3-0c2bd92e20cb · inbound

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training cites this paper.

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:42:46.231130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:41:12.453128Z digest=sha256:f44ab8a61e24d13d9ccf61dee1bee8df46155b8d37ea31c8f210bbbdc827f168

Observation 8f41a52d-1312-4e67-9349-18a28f33a787 · inbound

Validity Threats for Foundation Model Research cites this paper.

Validity Threats for Foundation Model Research Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:36:44.897985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:52:41.653304Z digest=sha256:afa446598fc6535b902f9a5146fd5c7e49aa05e7427e2007645c50642491f7dc

Observation 60a76257-86aa-4335-9722-75b17b21b171 · inbound

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning cites this paper.

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:19:57.812566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:46:13.587037Z digest=sha256:1bba772bfb7f1cb3b6d46d0f5098ba380da02c668583f43c0f3839c8e566376d

Observation 7cefba8d-d3b0-4c9e-b540-1fdc4c684421 · inbound

Data and Evaluation Closed-Loop for Model Capability Enhancement cites this paper.

Data and Evaluation Closed-Loop for Model Capability Enhancement Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:25:48.051684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T01:33:08.134048Z digest=sha256:4b85006c98a71c177a600bba9846db7fd1bc06220cc6962b569637f1d9362b35

Observation a752bcf8-4bee-40a3-8504-d02272321cae · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:48:39.450809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:2e3d90c0d3ca7308e30536541bdad53d6b549cccc533408fb326744d0f06a16a

Observation 808aa92c-8e34-43b1-9180-58d2f58b0605 · inbound

Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation cites this paper.

Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:54:44.256835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:54:44.256835Z digest=sha256:ca8958d6b86b38a1dea2152e49a37a42f69065c74dac84c608a5a9e3c983a7af

Observation 72a3f635-59e7-4db1-9179-1a7d6df13ef4 · inbound

Domain-Aware Scaling Laws Uncover Data Synergy cites this paper.

Domain-Aware Scaling Laws Uncover Data Synergy Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T07:24:27.255815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:24:27.255815Z digest=sha256:987a192b39181b9e902da47edbc96f244fdc0fbceec75139155625c022d09ea4

Observation 22de8f3c-08d4-4808-a226-b81939f715aa · inbound

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation cites this paper.

Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T14:24:46.118299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:24:46.118299Z digest=sha256:e6439cfc8811e160ea1da36a703c3c5471d53896f19407e3239a1b2fe197e417

Observation ddca782f-f0ad-4a5e-b0fb-b577d2f1ceab · inbound

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models cites this paper.

The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T10:34:41.807413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:34:41.807413Z digest=sha256:05b8a74fa17de5e59c64e27f66aca10f97a4a04a6088de21624e5617958d0219