Pith. sign in

Paper Citation Record · LEDGER

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives

As of 18 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2505.21598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21598 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:33:55.004931Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved47
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36bc18a6-cb30-40c1-a97f-4897f475f9cf · outbound

This paper cites A Survey on Data Selection for Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Survey on Data Selection for Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.370150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.370150Z digest=sha256:3ead70bd25b436a4b87eff2a66e3b70cee71e81846d45dd62f3d9d1a30485934

Observation 62f76901-e1cd-4dd5-b557-a4ee9abbe3a4 · outbound

This paper cites Efficient Online Data Mixing For Language Model Pre-Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Efficient Online Data Mixing For Language Model Pre-Training

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.487199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.487199Z digest=sha256:bfc965b2a7372e546d63f56b87915b71a10d7d3500bf8df88f500c9399f672c7

Observation 882719b5-0d86-4d79-951b-57fdb5e0f51b · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.307035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.603636Z digest=sha256:a63e469ab5c9fc736f81667d54fd2ae103633d18496897a97b4ca8cebf4bc791

Observation b4addb51-6306-41eb-b8c1-294a5259ce6b · outbound

This paper cites Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:57.368547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:48.730422Z digest=sha256:b3f040dd7b9fc5c6cc9fe5baf51db157c75961ff1b4376dc5f2f27b89064578b

Observation ab01c965-32fd-4ebb-b0ae-849adfda1d2f · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:48.893722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:48.893722Z digest=sha256:009c32dd9502e62d402b0c75e7a54c6b8c817e88ce7aca82db314d8c1f6f5f41

Observation 74c29320-b6e8-4c06-b69c-a39e6b5ba3d9 · outbound

This paper cites Aioli: A Unified Optimization Framework for Language Model Data Mixing.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Aioli: A Unified Optimization Framework for Language Model Data Mixing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.082128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.082128Z digest=sha256:cb3d22cf6c876741d231adfa453a096b7ee07b0f858f1cada5571d3fb9cbb310

Observation bcc8f024-6663-461c-8527-49b1cecfbe75 · outbound

This paper cites Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Skill-it! A Data-Driven Skills Framework for Understanding and Training Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.258242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.258242Z digest=sha256:2aea9973d393a17362a130bcdf9d1f9e2ac27d5c2d662d7455eb4ff1e52e7640

Observation e3ce5e90-f245-477f-a761-efa5cf0f381d · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Instruction-Finetuned Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.385365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.385365Z digest=sha256:88d8beceb4743c3b1f7e7fbad807ad9078019f8a4bbfdbed5dd8c2ba281182e8

Observation 23546973-3ab1-4667-a1c8-176bd45dbb1e · outbound

This paper cites Training GANs with Optimism.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training GANs with Optimism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.511514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.511514Z digest=sha256:11d0594c6098c208d37794e308c57a7a40dba917942dc303e0ccb1b1e5b0b285

Observation f28c872d-09c2-4910-96ae-4e6d2f031454 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:33:57.103547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:49.614639Z digest=sha256:9a17a47cff3451e1b75453c0777137f54f9765eebd5afd1de3b43a95011efc79

Observation 963250a7-090d-4ab7-9b92-d715f9f2a4b9 · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.759942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.759942Z digest=sha256:813662df75d45f711269a43545c830f426ea0102f06e7ba69a030ce1650e6d05

Observation fd576d84-729d-4da9-a0a7-2ed46f8e60e4 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:49.915397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:49.915397Z digest=sha256:b5178dba96e31fb3aea2f2745049c064234c9fa99002d2d7d11ff43e95d5427a

Observation b5ee6c42-924a-406c-ba0f-feb001fb6ba3 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:58.127646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.039377Z digest=sha256:6aa802d423b73acc791c5e51bf3ce9fafdd37fc0b623fddcc0efd791a9df1dac

Observation ccbd187d-b65b-45d9-b2fd-9eee1d0eddbc · outbound

This paper cites Forward and Reverse Gradient-Based Hyperparameter Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Forward and Reverse Gradient-Based Hyperparameter Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.797037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:50.178732Z digest=sha256:6a766f6a283d7f389dba711d4602ef800f14a7be0581bd512b07cdca628f8212

Observation 1e4ebb40-ae5c-4daf-960c-187a2e144079 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.325132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.325132Z digest=sha256:0fa14c938212ff70e80f93614368242454b86509202c86801930b4035667ef4c

Observation df30824e-1b19-4511-92de-89ab5d0166ef · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.472683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.472683Z digest=sha256:32d29ea51b17cca411a05e0173262f95a8141053cf7a78c3744d46112fdc9762

Observation 66437165-7de4-4550-accc-8ec38fb8c18a · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.604941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.604941Z digest=sha256:c53c48009a52475ea8de2e8ddcab808cd51df9f4d3b19dab10d91742afd8d9da

Observation dea4f49c-d8e5-4686-86dd-5cdaa69a665a · outbound

This paper cites Scaling Expert Language Models with Unsupervised Domain Discovery.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Expert Language Models with Unsupervised Domain Discovery

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.715368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.715368Z digest=sha256:c65eb382a75e35cb624a971d8a7e772612af3f0fa80d040e8282c297a14ae755

Observation 40a15f17-2fd7-4935-88c7-255670c03bb2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.850028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.850028Z digest=sha256:d2b0fd205f36d38b775163379a01e00d7840d7c55a510463cc4fe3733361f285

Observation b311690c-0dce-4614-b305-f00efb81c32c · outbound

This paper cites Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:50.960916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:50.960916Z digest=sha256:cba1d3da9a2e8df556ec770ad8f7f5d23b7bf9c5dd56b847c9ddac05941d46da

Observation 108619fd-ffb3-4a8e-a9bb-e1fe273c38d7 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.071911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.071911Z digest=sha256:d27f10a0e190df14985bc67a6bf2958a9836058e3cf70483f25aaba4b7a97579

Observation c14ea77e-2f43-4978-9bd0-6ac2aab41cba · outbound

This paper cites Scaling Laws for Neural Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.178477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.178477Z digest=sha256:acb13df3128579aefb2d7f02e8fa659ebdd0302543f3e08ecfc954bbb8aa14e8

Observation 7f11cfa6-090e-4ecf-a0a8-b8fda047eca8 · outbound

This paper cites A Fully First-Order Method for Stochastic Bilevel Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Fully First-Order Method for Stochastic Bilevel Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:56.324802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.298763Z digest=sha256:63d231bc46a55c9270ff3b993812a46de4b07b2887c81b677f9d58b876b33fc1

Observation e5192bdd-f44b-4a8a-9ca8-ee3e999773ef · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.971782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.414889Z digest=sha256:bea6954deddc4ef079078cf5f96b2c679772778b2697535cc362b25312d5be65

Observation 72f6d54c-d65e-473f-b53d-f57aa729b933 · outbound

This paper cites MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.548263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.548263Z digest=sha256:8a55f37acae7dc91d20de5b5d6e385e8ab43be4ebbc8c30cbe1bf876f151c8e4

Observation 1004f965-6e3d-4c85-8453-fc39bb0c4af4 · outbound

This paper cites FAMO: Fast Adaptive Multitask Optimization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives FAMO: Fast Adaptive Multitask Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.661519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.661519Z digest=sha256:23e6b797804d3302a8d96d50b0e2e5aa8922c75653670f52f2d97f3c73894f96

Observation ae44435e-3711-4976-ac39-7797fd167ca0 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.813615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.772616Z digest=sha256:f6061592e3240005748354f5699e45b03451f6cb1851103a968f2ec3b3785ff6

Observation ea6153a1-ef08-4c32-bce6-d6fbdf085acc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:51.884698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:51.884698Z digest=sha256:263e5bb3db10f4b23f5188263c00e273179a8fc5db5b795da17f0ed95f117fb8

Observation 5f0f9683-56eb-49ac-9295-c930cdaa501e · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:33:57.644662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:51.984488Z digest=sha256:15b99b6dccb50c90faa7a6cc960e9b48fd18f72e5caf682684c8aa491d725f4b

Observation 02ee085f-bef2-4e81-9cac-ed70b0ea99f1 · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.124575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.124575Z digest=sha256:62cef6eeae909c888758dd8a48d6e92fed7ea9cc0d24a8876e1ab039acba69d1

Observation 5b8c3ab1-ee23-4537-86aa-73cb9e83704b · outbound

This paper cites Optimizing Millions of Hyperparameters by Implicit Differentiation.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Optimizing Millions of Hyperparameters by Implicit Differentiation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.282275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.282275Z digest=sha256:c900d978508581e5503a1ff7db9c564afee7559786bab19d804d7f597318a8e2

Observation a8254859-93c3-45ee-ba9e-011c5d60762b · outbound

This paper cites Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.387698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.387698Z digest=sha256:bf45153af5cd8ba63ae4aa042cb4df8ce03c7fc07f3400d505f7be6fa901dc2f

Observation 583b578c-c295-45fa-840c-1ee032eaad82 · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.514060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.514060Z digest=sha256:ccbd3a8541c5b4e2e3cf770b9ae8e762f992fa4f3641ffba3c471bc9f37bc001

Observation 9d15914f-4f31-498a-8937-b25454cee021 · outbound

This paper cites Hashimoto, and Percy Liang.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Hashimoto, and Percy Liang

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.622870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.622870Z digest=sha256:2064fd41166f521e5329ed1bc277ae73efeef1d957f77c7e92b68ff30efb9d8c

Observation 47ac372b-45d2-4da2-843c-777b6e795982 · outbound

This paper cites LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.754333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.754333Z digest=sha256:27fbf9db8198560f47943a7c290b87edfada2b116fcc85afb08fedc177105c9a

Observation 29e0cfed-d8b5-454a-b07d-51453c2f516d · outbound

This paper cites ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.895760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.895760Z digest=sha256:8504b916498970adb699eecf5b594af64d7ac276f16058047b177efa5a7aadf1

Observation 16c2a519-496c-4d04-b4fa-96a085159c1a · outbound

This paper cites D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:52.994596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:52.994596Z digest=sha256:5001a1f998a31370fffd0a090ac6519b87218ab6dc80b7ac930898729efae46b

Observation 6914823b-34a4-443b-93df-1ad0ec8df797 · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.139442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.139442Z digest=sha256:cee5e5862be810e6d64bc750c4806e0d0b7e06756246bc506b9a3b540a368c58

Observation af4c5844-478f-4d49-979d-0b99705c36cc · outbound

This paper cites an unresolved cited work.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Unresolved cited work

Reference 39

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:33:55.904518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:53.242787Z digest=sha256:4c78d2f4a0bad3d953b2b25057c547f68214a5f4bd794d71263812b56a14b8c6

Observation 70d827fa-3469-459e-8020-9b16f8b81515 · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.325891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.325891Z digest=sha256:45201eb42b2c674a824594a592e503763d3ca2c248e8e70136e3ff8d3515e41c

Observation 7fe70179-7f88-46f0-8088-731f97cbb0cf · outbound

This paper cites SlimPajama-DC: Understanding Data Combinations for LLM Training.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives SlimPajama-DC: Understanding Data Combinations for LLM Training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.419429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.419429Z digest=sha256:3a70dd72308c6304f0acd3baf11f7bcc70d09e673db12d4a56343e48c8df35a7

Observation b6821ead-e6e2-44de-843c-4a7f822e84a6 · outbound

This paper cites The Vizier Gaussian Process Bandit Algorithm.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives The Vizier Gaussian Process Bandit Algorithm

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.554390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.554390Z digest=sha256:9d2721a505e4dea3930d26c2a35241d7583e0f2f8bf765fb684aa4a6d7677bc8

Observation 66e1e1d6-b288-4248-a255-e95759f080ca · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.715579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.715579Z digest=sha256:29b9dffc82d9b2f704243f6784a23278d012a5ec80f6a507a60f61f5173db33d

Observation 1c3b19b2-514e-4dc4-8c72-bee9fdeb6b64 · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Finetuned Language Models Are Zero-Shot Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.834351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.834351Z digest=sha256:881a3afa12979d64a7c01ada724707c49c900b8d42df8eaf80afc0793df26d53

Observation 3597c35a-4d91-4682-82bc-7cb147cc2dab · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:53.990057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:53.990057Z digest=sha256:8f35d11200ceb8e5173917d237b282a47d4dd248dd8cf3c3f00513526a3c4da5

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · outbound

This paper cites DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:d5242dbbcdebe514e00f40a00459e54f2348f3550d630462a1e894c42b433569

Observation 70accda3-7105-445a-8a53-0fa773379f6d · outbound

This paper cites A Unified Perspective on Multi-Domain and Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Unified Perspective on Multi-Domain and Multi-Task Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.253956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.253956Z digest=sha256:0bb28fb870735f787bdaf68bcd9b25fa92f9facf9774b3d3ecded29464a1e95a

Observation eb873359-d146-41a6-a21c-3727d824a5d8 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.384371Z digest=sha256:ae41fdb333a9cf7904a6049e236f56beb5a682273f888d6ee1731cad2d6a6a15

Observation a18e95ba-7dc5-42a9-b983-6530db2b0d41 · outbound

This paper cites Gradient Surgery for Multi-Task Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives Gradient Surgery for Multi-Task Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.509096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.509096Z digest=sha256:273779aaed6d3c70185629858dfdf77cb9eec542e59f7b771bc2ad3a36a8b9df

Observation 34000aaf-f730-4f6e-9659-43a4e2ce50f5 · outbound

This paper cites A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives A Single-Loop Smoothed Gradient Descent-Ascent Algorithm for Nonconvex-Concave Min-Max Problems

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:33:55.333064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-07T13:33:54.674767Z digest=sha256:ed940f2ed39bf4d003dc8a647ab0b86807b74ffb55f5b49faeb510b36d84f855

Observation 141ad7d3-e01b-4495-b403-50e15e80cb69 · outbound

This paper cites An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives An Introduction to Bi-level Optimization: Foundations and Applications in Signal Processing and Machine Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.796497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.796497Z digest=sha256:233185aafd8dc9a40f7656f4a0b717ee76203e75403407d81d416bf2777e5c6c

Observation 5bb3652e-81b6-4dc2-9b03-43b5ae33a944 · outbound

This paper cites online" 'onlinestring :=.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives online" 'onlinestring :=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.889501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.889501Z digest=sha256:3973b58820e5bbe4e00db3897b6f035c29d286a7eed3578401b1059b369931d3

Observation b1e40183-72ef-4dd2-9328-0352283a0f8c · outbound

This paper cites write newline.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:55.004931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:55.004931Z digest=sha256:20e2ca09724fc931fb2b3ec8b6b8772e8fd673d7620edcab718b96427285c21e

Pith citing papers

No inbound Pith citation observations are available.