Pith. sign in

Paper Citation Record · LEDGER

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2305.10429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10429 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:36:36.566504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.134017Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ddfd013d-5f9e-44ca-b6eb-4fb10d7b5e35 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.480274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:c66ad353b861d8444569c13076fbc05b2c976ece1ab4f7484d9fbc69307f1bc5

Observation f908cae6-8cff-4bac-8b2c-effc1ca44a04 · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:46.419903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:a11631829eca824644ee698640e96234c34879d49ee6eb54537b11ffc4bb67ee

Observation f0addaf2-1bbc-44bb-a1ac-08568136d1bd · inbound

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining cites this paper.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.566504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.566504Z digest=sha256:9c4ab1ce7d06fab5fae89d06568e3ca20de7f732160e2dd7a36a46efce8b7c6b

Observation 3daf7f01-a17b-4f35-84cb-77d48c8e660b · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:34.937484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:34.937484Z digest=sha256:99355639a1cf81e6cd270d6872bbd793146310cc055f4312bdc2129ad8586e18

Observation 06eaf234-dbfa-4b6e-8877-ee3e34034ffb · inbound

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects cites this paper.

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:03.006603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:03.006603Z digest=sha256:58b7f963cd57b7a235fc0c8727f0a5150228c901ac09d5440ea8ad164e96f44b

Observation b4303d09-6850-428d-9ed4-32fa963c15ef · inbound

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining cites this paper.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.550042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.550042Z digest=sha256:e8b849d2d267978c5567ce466e9412456074e3ddfe61dec2a862f206106050c9

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · inbound

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives cites this paper.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:421a0d21c3031c6fa845a0c638a2a7e33501dab466669109f734f14a976e75e9

Observation dde6cae9-73d6-4065-836e-a0cc34f9af89 · inbound

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning cites this paper.

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:28.811231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:28.811231Z digest=sha256:c787375681d2eb0d50ce8007f4909eebea169f9cad657d34381fc19790383b21

Observation 41496677-dd96-420c-a33b-7e258923b5e6 · inbound

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models cites this paper.

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:41:53.033134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T22:41:26.047957Z digest=sha256:e7143a188e3fdf7d05eb09a9c93874bd1caefabc9ffd7ed9d7f41c3360bcfe7e

Observation d1cb1622-3927-4ef0-8546-d8c72aa83fd5 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.127336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.127336Z digest=sha256:6972014a556ce50a6c87b9fdeb9423c4900e454639718bb4abff2bfa2f9eaf55

Observation 3fd118da-1829-4887-8de4-76839afb8443 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.534644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.534644Z digest=sha256:084b3356658d07c3ba01e766119a1970dcc9c397b39bbdd04e76c2d2f04ad588

Observation dbceb4dc-a267-431b-bb0f-2776d01a6a1e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:27.466265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:33cbb628247a053edab02af26fb9a5fc1afcb855d065ea1800e6a8196229d732

Observation 38133877-a792-4574-a582-bc635b36b31a · inbound

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics cites this paper.

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:57.184917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:52:38.167538Z digest=sha256:8420dd23889fb454048662e5c04fa576599675d0da48f542434fe93e30811be9

Observation a4e4bb3b-96a9-41fa-b64a-769f3af999c7 · inbound

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining cites this paper.

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.135670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T19:04:11.976747Z digest=sha256:53ec536538dc28c26fd206fd834f36bda1f72d12ba5dfb31b93f18540aee9377

Observation c4ea9473-5c1b-4fbe-83dd-040053d71e34 · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.445070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:df24157523352da6fa7ad12c1ca87ee4ab2c4d881227d8cdd00b3b45509b4a73

Observation 1fc5d40f-f130-4a86-9e41-b585abd7572d · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:20.625282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:20.625282Z digest=sha256:2314b2d741ccf88897ce87ea0ee07e77fcb6afd6411b7c23f971e5ff26d99134

Observation 66f4baed-efb2-4bff-8050-86e767bad1c8 · inbound

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement cites this paper.

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:15:07.227138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:15:07.227138Z digest=sha256:82574e64d94e330a4886497cef979b7e241f6b42b5a3e8a6927103556880ed3c