Pith. sign in

Paper Citation Record · LEDGER

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives

As of 8 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.21598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21598 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.004931Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36bc18a6-cb30-40c1-a97f-4897f475f9cf · outbound

This paper cites A Survey on Data Selection for Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.370150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.370150Z digest=sha256:b1dfd2749e626e138a6c30ef45674610cd7a1fe41725226aa4203a01c4ebaee8

Observation 62f76901-e1cd-4dd5-b557-a4ee9abbe3a4 · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Efficient Online Data Mixing For Language Model Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.487199Z digest=sha256:30c6f098c8c9793a8ffcdfc8a1192b073f2ebcdea3fa2361d4493587e5de6857

Observation 882719b5-0d86-4d79-951b-57fdb5e0f51b · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.307035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.603636Z digest=sha256:11cee872cb0311d60f5d69e7269922b7211c317decca874682804d55c9f66e28

Observation b4addb51-6306-41eb-b8c1-294a5259ce6b · outbound

This paper cites Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:57.368547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.730422Z digest=sha256:49a8d083c9f338b294319fc9229ca665722d5b2e34b4027aef6897546ed9d8af

Observation ab01c965-32fd-4ebb-b0ae-849adfda1d2f · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.893722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.893722Z digest=sha256:28ec51b80630990ff4c9cb2f7c53c9fecf7f2ab4363b780f73788d3b30c75023

Observation 74c29320-b6e8-4c06-b69c-a39e6b5ba3d9 · outbound

This paper cites Aioli: A Unified Optimization Framework for Language Model Data Mixing.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Aioli: A Unified Optimization Framework for Language Model Data Mixing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.082128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.082128Z digest=sha256:aae644cbb983459534df581ca395b698e5b941314776501c0e8fdbe5400be858

Observation bcc8f024-6663-461c-8527-49b1cecfbe75 · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.258242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.258242Z digest=sha256:f46cfd664572e454a26dc81ae225056984fff57caa3285ce834461cc75ee27b1

Observation e3ce5e90-f245-477f-a761-efa5cf0f381d · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Instruction-Finetuned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.385365Z digest=sha256:cfb55fff64e71fd3273e2d6c3ea818b46d645dabd0ce65f459e187ef80fd0f68

Observation 23546973-3ab1-4667-a1c8-176bd45dbb1e · outbound

This paper cites Training GANs with Optimism.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training GANs with Optimism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.511514Z digest=sha256:ced0e64906b5bcbd9750946b0d6442df5c956b41795222f1d47b5168e794f03c

Observation f28c872d-09c2-4910-96ae-4e6d2f031454 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:33:57.103547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:49.614639Z digest=sha256:38e322b68a5ebffd2e65b2079422567fac4c182247411da85bea985f07340407

Observation 963250a7-090d-4ab7-9b92-d715f9f2a4b9 · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.759942Z digest=sha256:b550be879208e69d2a0403761eeae01c294b24b54c1741fe61aa94c1981274fa

Observation fd576d84-729d-4da9-a0a7-2ed46f8e60e4 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.915397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.915397Z digest=sha256:b68edfd816d8b8d30e0373d45d3670cdb357a8dad2c33b43bf60c3d39cdb2ddd

Observation b5ee6c42-924a-406c-ba0f-feb001fb6ba3 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.127646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.039377Z digest=sha256:36cf548e2a66ff5cbcc71c711cf2921b1107f29b752b59fc5a1f85396b7efa00

Observation ccbd187d-b65b-45d9-b2fd-9eee1d0eddbc · outbound

This paper cites Forward and Reverse Gradient-Based Hyperparameter Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Forward and Reverse Gradient-Based Hyperparameter Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.797037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.178732Z digest=sha256:36519d6f014514cfb268745d5b4ba07e542a3b33c03a562f88870f09c577b202

Observation 1e4ebb40-ae5c-4daf-960c-187a2e144079 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.325132Z digest=sha256:d813f1225071d622372d9a4e77832e845b3f8a50ed5b363713c5e22ad3e723b9

Observation df30824e-1b19-4511-92de-89ab5d0166ef · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.472683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.472683Z digest=sha256:1bf2806f2fefe5e478f6de239f8f291f00580c7198d6adc1e625ca43527fb818

Observation 66437165-7de4-4550-accc-8ec38fb8c18a · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.604941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.604941Z digest=sha256:464aae78fc1441b1d87e6a0c44976550c1a882099f3841784032578ace0d921c

Observation dea4f49c-d8e5-4686-86dd-5cdaa69a665a · outbound

This paper cites Scaling Expert Language Models with Unsupervised Domain Discovery.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Expert Language Models with Unsupervised Domain Discovery

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.715368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.715368Z digest=sha256:0ea37f009b8c46a203ad19a318ef1d0cb02f02adfe24d26ffcdcad6ab28788d8

Observation 40a15f17-2fd7-4935-88c7-255670c03bb2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.850028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.850028Z digest=sha256:061c6dea0fd894e3b46cba91661aa645eccef5a85cb4c8b479aee32cd0cd084d

Observation b311690c-0dce-4614-b305-f00efb81c32c · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.960916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.960916Z digest=sha256:9820f10c9e7ccc54f29f336a4e1180d2e26aac21b00cef292c118df38e5c0437

Observation 108619fd-ffb3-4a8e-a9bb-e1fe273c38d7 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.071911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.071911Z digest=sha256:0e109b525b8c7e0ebf9ca7121c0c120f71d3355905b3c06ff5dc243e6cefcc26

Observation c14ea77e-2f43-4978-9bd0-6ac2aab41cba · outbound

This paper cites Scaling Laws for Neural Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.178477Z digest=sha256:57f4c209c63f81ee18439174b853743deb88c33b2e215698cb8d136f505582dc

Observation 7f11cfa6-090e-4ecf-a0a8-b8fda047eca8 · outbound

This paper cites A Fully First-Order Method for Stochastic Bilevel Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Fully First-Order Method for Stochastic Bilevel Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.298763Z digest=sha256:9a1bb7ae30ac38c1840d66e5ecc0b3475843891bd7080df3ae70c1a778baf5a2

Observation e5192bdd-f44b-4a8a-9ca8-ee3e999773ef · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.971782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.414889Z digest=sha256:a8a62afab3c203a012ed7118eb8c395eb038921186a6029e21f9efa4f5f1137d

Observation 72f6d54c-d65e-473f-b53d-f57aa729b933 · outbound

This paper cites MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.548263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.548263Z digest=sha256:3f2a6c23e4fe1440cf693dc10aac32fa710a1ab2bf78cbf9fc4dce0736f20e7a

Observation 1004f965-6e3d-4c85-8453-fc39bb0c4af4 · outbound

This paper cites FAMO: Fast Adaptive Multitask Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives FAMO: Fast Adaptive Multitask Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.661519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.661519Z digest=sha256:7fb1820c15fb76c2876f2ffa53c0e2d86b136867d9dd7f0815c6044b5fdb600f

Observation ae44435e-3711-4976-ac39-7797fd167ca0 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.813615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.772616Z digest=sha256:74282cce0f95b091975464563487cca14081ca44887d69f81c80c9b27f59b170

Observation ea6153a1-ef08-4c32-bce6-d6fbdf085acc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.884698Z digest=sha256:861513a99177df401242cae25dbaedef5ba5511d453346edd8fc48002a5208a4

Observation 5f0f9683-56eb-49ac-9295-c930cdaa501e · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.644662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.984488Z digest=sha256:48ef4bacee2b91f1b6959d974d09194da0178e3369437cbc91f49ac993d9e6a0

Observation 02ee085f-bef2-4e81-9cac-ed70b0ea99f1 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.124575Z digest=sha256:b9a33e83d3924dcece1a2847bd2b010e3ec43d3888bf6c6c30df25c20208fc4d

Observation 5b8c3ab1-ee23-4537-86aa-73cb9e83704b · outbound

This paper cites Optimizing Millions of Hyperparameters by Implicit Differentiation.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Millions of Hyperparameters by Implicit Differentiation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.282275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.282275Z digest=sha256:a3d472f17a48bd1e52d5420d87835bf344b5741e4a275ca19a7828eb1f222645

Observation a8254859-93c3-45ee-ba9e-011c5d60762b · outbound

This paper cites Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.387698Z digest=sha256:c4fe3c023a0055455f2035f63b1dd580273892321d8a1d2a3f6fb06ca40991ba

Observation 583b578c-c295-45fa-840c-1ee032eaad82 · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.514060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.514060Z digest=sha256:2f3fb80107012edd8ef583b162d4a45360aa2e66ecb93228e5f17bd136a0f811

Observation 9d15914f-4f31-498a-8937-b25454cee021 · outbound

This paper cites Hashimoto, and Percy Liang.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Hashimoto, and Percy Liang

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.622870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.622870Z digest=sha256:24e427693ada4231eba2fd6aea3b13bbf9beaa02df66eed5a27fa4b561070f8a

Observation 47ac372b-45d2-4da2-843c-777b6e795982 · outbound

This paper cites LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.754333Z digest=sha256:d0f7de2bfce89366dc90bed226b766f00c995f2293dcb955faee3116d49f905d

Observation 29e0cfed-d8b5-454a-b07d-51453c2f516d · outbound

This paper cites ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.895760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.895760Z digest=sha256:e6eb644f806111b705a748ddb82c3331c3adf5bd7e8f8b14e272ecace2664090

Observation 16c2a519-496c-4d04-b4fa-96a085159c1a · outbound

This paper cites D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.994596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.994596Z digest=sha256:80f6905cda091fe4f5e71a082b95a362f0be28240def16e7866eb426debedc73

Observation 6914823b-34a4-443b-93df-1ad0ec8df797 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.139442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.139442Z digest=sha256:7efed6b15681870ded70897684d6d69aaad0c77bd50eb8ca176c663ba2bd0a1c

Observation af4c5844-478f-4d49-979d-0b99705c36cc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 39

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:33:55.904518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:53.242787Z digest=sha256:3525bc1f7a0e6ae319dc965522d0ea93322bc6e195de280aa2fddc7a94709e51

Observation 70d827fa-3469-459e-8020-9b16f8b81515 · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.325891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.325891Z digest=sha256:5ef7fa4d8094cb01222827e2a0c647c97a4813e6b800e2489c740fb7d1d200d2

Observation 7fe70179-7f88-46f0-8088-731f97cbb0cf · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.419429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.419429Z digest=sha256:760a7d0c7e59f57b4866a3f2e8467394a7d94d40a25c7ee10ae68127ace6a16c

Observation b6821ead-e6e2-44de-843c-4a7f822e84a6 · outbound

This paper cites The Vizier Gaussian Process Bandit Algorithm.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Vizier Gaussian Process Bandit Algorithm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.554390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.554390Z digest=sha256:a2342f5fa19f9f86b6d3a22c5750c7b638419e373232aefbbb029e5bed1e5f2d

Observation 66e1e1d6-b288-4248-a255-e95759f080ca · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.715579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.715579Z digest=sha256:2ed9443d6fc22d0f44d3901db9b23b896d208d8f355d4bf3014261f4b14286aa

Observation 1c3b19b2-514e-4dc4-8c72-bee9fdeb6b64 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Finetuned Language Models Are Zero-Shot Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.834351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.834351Z digest=sha256:6ed8d38f24da113761185fc6814af531f35f975c6bc9a29c24c66d24f273e74b

Observation 3597c35a-4d91-4682-82bc-7cb147cc2dab · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.990057Z digest=sha256:eea3c6f988096dde0ced0e5d5780d314d6331e6bf20b5c1195dc951df0fcafee

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:421a0d21c3031c6fa845a0c638a2a7e33501dab466669109f734f14a976e75e9

Observation 70accda3-7105-445a-8a53-0fa773379f6d · outbound

This paper cites A Unified Perspective on Multi-Domain and Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Unified Perspective on Multi-Domain and Multi-Task Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.253956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.253956Z digest=sha256:c6db7eccca0acac8a73ba71f0e323abe812d7c679763b74a3a9dd7d1b0fc55dc

Observation eb873359-d146-41a6-a21c-3727d824a5d8 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.384371Z digest=sha256:95603539e45d6a17a16f363042fa8b48318a8dcbad081f1e72a24b6d1b836e46

Observation a18e95ba-7dc5-42a9-b983-6530db2b0d41 · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Gradient Surgery for Multi-Task Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.509096Z digest=sha256:6b8c45ea74c3db3e2857ca94115638dffd9cf81ed30d00aaf18f67028969d464

Observation 34000aaf-f730-4f6e-9659-43a4e2ce50f5 · outbound

This paper cites A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:55.333064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:33:54.674767Z digest=sha256:fc59280bf07073a7af308a12f9e48daf200c62cfb4bdbf12bfe36fb2e85b5bc3

Observation 141ad7d3-e01b-4495-b403-50e15e80cb69 · outbound

This paper cites An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.796497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.796497Z digest=sha256:9135aaf0650f754a388a1e5701eb29db44258f555fd3d17bea6730f8623f3124

Observation 5bb3652e-81b6-4dc2-9b03-43b5ae33a944 · outbound

This paper cites online" 'onlinestring :=.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.889501Z digest=sha256:c52dd3b4704613f30358f0d98b6e649822c3ec489f7e169c0b2b211da4998f35

Observation b1e40183-72ef-4dd2-9328-0352283a0f8c · outbound

This paper cites write newline.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:55.004931Z digest=sha256:8a19171be19565459530f9eeabf694aff519a7c9629c2e9cd757cf3f6b1155c2

Pith citing papers

No inbound Pith citation observations are available.