Pith. sign in

Paper Citation Record · LEDGER

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining

As of 18 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2504.16511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16511 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:07:15.189614Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T12:45:04.789254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.391752Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18113be9-a352-4253-9da4-c406b66c7660 · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.934745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.934745Z digest=sha256:9596b0fb69b3a34919e34b273701194cb6467dd22f8c5d785d9b14831c84d674

Observation 53644dcf-6db3-41ef-899c-d7d36e6b8c89 · outbound

This paper cites Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.940766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.940766Z digest=sha256:c8e8ba2d9a084a879a36408fa104149d41e657e320fb96bff6b0be615ab2faf2

Observation dd1a1174-8a4a-4701-b098-4f978042910f · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.944819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.944819Z digest=sha256:f7752726f5de383841cc7d88f88b4e249f62869740af26caee0228d89fc9d986

Observation c103f5ec-1966-4a95-8bed-7a6ff0ab299b · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.967105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:14.948776Z digest=sha256:d1fd409bd6d91e95b6c4f0a0b3258579aab57b0f71144d14ce6ee4bed71c6293

Observation 1e1db54a-4db0-466e-b923-a2b370876a0f · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.952639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.952639Z digest=sha256:e5e361c7f41aa838bd6b2b2cd21d78f0938d6bf6635e7a38bc74d5c6afd20802

Observation aa13d2ec-4e9f-47a3-ac44-1abcd9e1b42c · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.953433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:14.956467Z digest=sha256:a375d853c3e391f922ef1f6934bbc900efcf9e5338f08e887f7f6b0eee9d95c1

Observation 515475fb-f3ad-4856-8fe4-984fa26bf65a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.960502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.960502Z digest=sha256:ed48e5f04a259782f06753b6999beccec39e6f6a3536d102178114a930ffcf3e

Observation 06875fb2-614c-4730-8895-81964ac991dc · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.941015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:14.964421Z digest=sha256:ba294d5133a7850077417c4ba8b217b6096d024c6fa021b9a938de79e72d9b7f

Observation 7a6747e1-5539-4034-a383-98d4183db7cb · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.928837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:14.968271Z digest=sha256:58c37b1713cf4346e5aed6d95d364d5481a65c23b0b41e5bc15b381cbf4fd1fd

Observation bae800b3-9f38-430d-ab30-627b501f6c1c · outbound

This paper cites DoGE: Domain Reweighting with Generalization Estimation.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining DoGE: Domain Reweighting with Generalization Estimation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.972335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.972335Z digest=sha256:0d9ef7480066063b815a524d9b88ebf5ec5eaa2f4255418a40387eb036c11347

Observation c583bee2-491e-476f-9152-7a08cf2b7f65 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.977220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.977220Z digest=sha256:304ab501f29797d9ee56e6fe972626c799e3a0a4aa46fc16a2dc3dc5080bf1be

Observation 1d5d5b1a-3581-45e5-8fd2-dffb1a53d0ef · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.981495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.981495Z digest=sha256:d9a009c4a2a96af3fc31cd31a35f12e32e0c6f4340307da6c2a4854095592dcd

Observation 42191d73-36c5-46a2-8906-9bdf5902a81a · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.985726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.985726Z digest=sha256:aa6cdf21e9ab8916b8f7dd8a7a594232ac7ba6878ffa4f4d18b7fe6083a2ba59

Observation 8585257c-2f0b-41ba-950e-62b43148ab2f · outbound

This paper cites Data Selection via Optimal Control for Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Data Selection via Optimal Control for Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.990281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.990281Z digest=sha256:3f2318733af1bbcc27293def30c8c6f2f91d74593892c2e6b04d69350bf2cd50

Observation 0a251fb5-b23d-4742-af69-7a37871175d0 · outbound

This paper cites Textbooks Are All You Need.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Textbooks Are All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.994688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.994688Z digest=sha256:3c095966d2d848e5279366008fc03fdc22c01c9b14f2d9725fdd58d748e1a97f

Observation 518af1e8-c7d3-49df-ae7b-fd6a8a23fd38 · outbound

This paper cites DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:14.998885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:14.998885Z digest=sha256:eb5484f4b7258b996075f06b7affd7e298dce93f671409acd9413a98d12dc91f

Observation 227da26c-ee52-41ce-8369-8c4e13b754e3 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.003738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.003738Z digest=sha256:5c59500e834f47201bad91c1bba66a383de58a4b25e50caf78ea60fd0e3cd7d5

Observation 0d3be270-f7a4-4350-b558-1049e5ea7fa5 · outbound

This paper cites Scaling Laws and Interpretability of Learning from Repeated Data.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Scaling Laws and Interpretability of Learning from Repeated Data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.010383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.010383Z digest=sha256:5a0f0f0822b63fa03974613d9c92b2ccd5780296fb5d8421055a4e94cd4713e3

Observation a38c3afa-f65d-414d-8415-9a5ff70e89f1 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.016235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.016235Z digest=sha256:6432f240fca3f7b7c6c9aa6e3ad822aa97ad95a46b937d3fca21182d21d3106c

Observation 3231b085-952b-4924-aa80-c1960fa82af7 · outbound

This paper cites https://github.com/NVIDIA/NeMo-Curator NeMo-Curator: a toolkit for data curation.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining https://github.com/NVIDIA/NeMo-Curator NeMo-Curator: a toolkit for data curation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:07:15.901942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.022367Z digest=sha256:5c1fd64a2a828ad46cef163b02078342d404b73f0cbcbbc5aef33cb3e9ebbdc6

Observation ac0648ef-c5bb-4f6a-80da-a611b2dd5e38 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.029060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.029060Z digest=sha256:f98a2c64d1c2a976008b894336fd21808d12a5273b4a5015848388c4f70bc557

Observation 06748844-227f-4e56-a1d0-fd44982a70e9 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.033350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.033350Z digest=sha256:09ad5ee8a33565762baaeb70433e716c5871946e4c456ebc560d0117a436ae52

Observation 592da973-b2d3-4b95-bf78-42c75d0fc2ed · outbound

This paper cites Scaling Laws for Neural Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Scaling Laws for Neural Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.038483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.038483Z digest=sha256:2e935da801cd03772ce762a9cba1862e8054200082ef415d32b5eb15eaa96347

Observation b04539dc-8965-430e-a645-5af2e8867a3c · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.887730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.042655Z digest=sha256:b63214827b1a570aacf96e0f5e6a6e25f15f352fbfe208697a937309db29f4e0

Observation cb01819a-2dc8-4315-a467-5448f2036b0a · outbound

This paper cites Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.046668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.046668Z digest=sha256:8a52b5bed370a1e339cadc01296b39edf51ec4aa2b1ec9051748b5a62b017a04

Observation 3d59d27a-d9f5-42a6-b07e-a98c20968b30 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.875861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.050955Z digest=sha256:eeefc08c49892898d71d62e8d136415eff1654ce64269343a5b35c386aa013b2

Observation 959392f4-954d-4c79-a6ee-f94b664e779e · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.055003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.055003Z digest=sha256:c1e3351e219f5675fe3fa26f570d09dc2f1ad6223c11090eaa041896a3369ea1

Observation efba8946-9f2d-4e4d-bd65-2eb3c2c7a607 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining DataComp-LM: In search of the next generation of training sets for language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.058757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.058757Z digest=sha256:372ebd476fbd2f94750837b607049d2aa523d5615bcaa991fa16b8d685c0950d

Observation f9b0a59f-8084-49cc-b4c9-555bdd345cc8 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Rho-1: Not All Tokens Are What You Need

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.062662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.062662Z digest=sha256:204a121013174fccb12c70badb93c02e3044e43a889abdbc47d30dfa653c015d

Observation e9ea5e23-6343-4879-8bd5-5bdcc4785d38 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.066525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.066525Z digest=sha256:8d9b5df5a10b8a674d98857b707cb168599d5265855af23f4468bc703fdb4567

Observation 9d106167-c249-4804-8958-b8d32deb7558 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.070570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.070570Z digest=sha256:b3a2755bdb5127d8d1de1c18308b397d86237e7baa46aa2f714e2bc5900fe201

Observation 21c594f9-ba2a-48e2-9a20-609bda0ae654 · outbound

This paper cites When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.074219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.074219Z digest=sha256:f72e9c7c074228c920f62d19cabda5308b7673f34e48f87b1415cd573d571ac4

Observation a44a1ece-457f-469a-adf6-0c6392be2950 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.078021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.078021Z digest=sha256:aec2d1d5160f44dea165620cff88fd340b43fbf90cabd93e9a17fc658a319e60

Observation 2c218b1e-e5ef-42d6-80de-ffed2865bc0f · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.863865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.081989Z digest=sha256:aaab605dc0798d38fcd3341740f51c7dd8d7206e0f01d5d0c870f7cbc0e81cec

Observation 8144be20-d2cb-42a5-8b67-3851fdd051bb · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.085887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.085887Z digest=sha256:997a9b25c9f8583a88bbb6a7693bb7e697423f22906b4caa674621ff7f96b517

Observation 5ff4f8c8-a7f0-419f-957f-8e80a027c3be · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.094718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.094718Z digest=sha256:14ec63af483eb0b05c5fcdd78ad0ea19312d41cc45b2e52b782d32f429ac75ec

Observation dcc1d83b-a290-4f36-8938-9c5b8a9a59b6 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.850696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.098832Z digest=sha256:d51a6252163e81be276b8eb150833f84a8bb8d502b058bcc187b4382753c9f6a

Observation 96195f1e-59e2-4380-b87a-8b138b1f40fd · outbound

This paper cites D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.102679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.102679Z digest=sha256:d68c0999fd97f5a7034cc6eb3680d82139eb9f0283263c92fd930cf0b721f8d5

Observation e4196d29-ba01-4108-b793-afebda3cd471 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.106536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.106536Z digest=sha256:911c9c17c4b47346978175858bf4b5e24b93a40413367c7c5ef0e7c8ad48e2fd

Observation 53d97daa-b63f-4413-9dc6-e5900e812450 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.110636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.110636Z digest=sha256:d4f8b046b58023dd781c4ddf1d43c1bfcd8ce22fc17ba84fd2f63e65f2a40787

Observation f9324099-b849-4562-bf63-7ae68ab4c3ed · outbound

This paper cites How to Train Data-Efficient LLMs.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining How to Train Data-Efficient LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.114897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.114897Z digest=sha256:a4134c3f44d452193b2e40ab6ead44543ba747efab627da300e14c18f3b60d2e

Observation 2ec67987-421c-4e8f-821f-bde8968c86f5 · outbound

This paper cites Balanced Data Sampling for Language Model Training with Clustering.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Balanced Data Sampling for Language Model Training with Clustering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.119207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.119207Z digest=sha256:ba481210b9075afad677c0edc73664d7f806f31796e6f0a40f7803e33d6ac49e

Observation e802ce7e-4647-4071-9ea1-e07ed36321d9 · outbound

This paper cites GLU Variants Improve Transformer.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining GLU Variants Improve Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.123162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.123162Z digest=sha256:c749804ef6a3f660b508d0e85fb32dc000439197d01bc532b96ea24aed2dc028

Observation cce0c089-d6d5-4dbe-a44b-1cda77513f0c · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.127037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.127037Z digest=sha256:f79b4a79b3968bd8ab41782955a49b11a907af84d956d8b45a4b569f74ece4b6

Observation 66fc7ed9-934f-4f09-8446-f2e955029ff5 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.130819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.130819Z digest=sha256:3af49298865b58a138a03528464b6f5b248eb1ceb6a0c5f94f4f11216f936188

Observation fec343f3-b789-4c59-b17d-43a7f8739bbf · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.828228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.135006Z digest=sha256:a034262dcf904f8db9b8570e9753aaa7b7e1801b8f218923cb3398bf7f2b482a

Observation c2d4c18f-fac1-47d3-8585-c867dde88af8 · outbound

This paper cites Improving Pretraining Data Using Perplexity Correlations.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Improving Pretraining Data Using Perplexity Correlations

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.139084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.139084Z digest=sha256:9e0e93b8b9eafa84550fc5a2ba102586a33ecfe74b8f8e5e56eff57347fe70d7

Observation 1f473210-55aa-4c96-9121-b1f5252556e4 · outbound

This paper cites D4: Improving LLM Pretraining via Document De-Duplication and Diversification.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining D4: Improving LLM Pretraining via Document De-Duplication and Diversification

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.143211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.143211Z digest=sha256:991932b7702fc5dad23b323aab61fa90e1a1e963aafa0a302b24808052ecd770

Observation 2fb81605-8b96-49a9-9125-50ad5e314ce9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.147185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.147185Z digest=sha256:f95f50717aeacbbb862581761bbdee926de32fe2f6b5c15713e3ff97452f5949

Observation b47b022f-1d8d-420f-a45c-94ebe958bb07 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.151237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.151237Z digest=sha256:9d139ad44e0e13dafe70304d35582c4df38c3a831179198ab52e37a3a33e36b5

Observation 61f33809-279a-4c28-a0b1-a1e9d95fc373 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining RedPajama: an Open Dataset for Training Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.154851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.154851Z digest=sha256:6907d78f65da232554bba0bed9c24ee5cb3b059c4aace494688d3960ef1501bf

Observation 98c1c595-e1ab-4d6a-ac8c-2e9a1bfa100f · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.808351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.158677Z digest=sha256:1d3c9dadd40981c2bcef8166264bcc415a9d0bf69a25dffb4a00f01f726f3df3

Observation 5caec618-ad28-438a-8ef9-d13f7bc61416 · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining QuRating: Selecting High-Quality Data for Training Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.162119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.162119Z digest=sha256:4393fbc33bd5fa2c88a0dc17f1bc92cca2934ee0d3e043d39c4de4924cd69694

Observation e9c15c65-2606-44e7-8e55-e8262d91ff72 · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.166101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.166101Z digest=sha256:d6daf58ac2d962d20a60cc8a66f7c39148c0b3457eeacb8fd8000cb71e94cecd

Observation 5e40ea66-1aed-4df5-8e41-ad525714be2d · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:07:15.788126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T11:07:15.170134Z digest=sha256:7430da28025cc453510d20622fd70a3c02d8ac4725f2c44f0217b000e684a3a6

Observation 42297567-3221-4515-aabc-97cc9480bf14 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.173517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.173517Z digest=sha256:548b219d4af6aeac824bdfb4b9ae972d6ed3eb5a840ad2adc6c582e7e3bd5332

Observation 44a041b2-6238-4687-8ab0-71a56e4a2e29 · outbound

This paper cites MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.177393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.177393Z digest=sha256:b0b0d81f060adbfb946a2bab3874b02b10165bf700b3d8ef941fb743b186b35f

Observation 29557480-3042-46cf-914d-701c4d86b18b · outbound

This paper cites an unresolved cited work.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.181415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.181415Z digest=sha256:f9047e102afe72c04ac73f6c18ea133f28201ce3a10aee65d3d9506c34f664c3

Observation 2d04cb42-9ce1-45c5-9dcb-e3404df875ab · outbound

This paper cites online" 'onlinestring :=.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining online" 'onlinestring :=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.185297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.185297Z digest=sha256:088dea2de7052d2bad452b62fa81e2f55e2a07597b2827fa2ca8d75c507dce01

Observation 4e909089-ddfb-4dc3-971d-35b2c56b3b61 · outbound

This paper cites write newline.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.189614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.189614Z digest=sha256:74133eaaf3a4992d9641b39f9725814fd0aed5dc11c76b590342a814d303b71a

Pith citing papers

Observation 3b319f6d-5ca3-4aa1-897a-89221e9f7c32 · inbound

Data Mixing for Large Language Models Pretraining: A Survey and Outlook cites this paper.

Data Mixing for Large Language Models Pretraining: A Survey and Outlook QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.666412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T00:56:04.958757Z digest=sha256:54d74b3f429c6a176fc6286cdf2c1f97ec3fd4608e61b4024642919cd5f350f2

Observation 810dd919-b834-4c7b-aa5e-a33660d5f189 · inbound

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition cites this paper.

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:15:37.219533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T18:45:52.380042Z digest=sha256:8ed463be253d1576a0619fbcfd56ffdb17cdd3cb21ac8908d4d9de9e184f2665

Observation d5e2a428-68bd-4552-ac55-1b37d70192f8 · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.393155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:4bc0850c9de82698f4ecc1450f1bd36effb985d9ad1682d92ef2aecced170a38

Observation 85ef347f-29e8-45c7-84c1-9e86482d7926 · inbound

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes cites this paper.

DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T12:45:04.789254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T12:45:04.789254Z digest=sha256:563801093493af582cb5b787c802091da3ba774e066c1d280c1e30eaab42582c