Pith. sign in

Paper Citation Record · LEDGER

DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2305.10429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.10429 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.557385Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.134017Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ddfd013d-5f9e-44ca-b6eb-4fb10d7b5e35 · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:46:40.480274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:194fefd872586bf92ac77dac249623d3a7d8664db45ad47b4d32ab63bb873b4b

Observation f908cae6-8cff-4bac-8b2c-effc1ca44a04 · inbound

Llemma: An Open Language Model For Mathematics cites this paper.

Llemma: An Open Language Model For Mathematics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 197

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:17:46.419903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T08:17:46.055279Z digest=sha256:08136682d7e518ef43264c34e3ea026698b7ce49663622b331337a29b754d9d5

Observation aa434cf1-ebd4-4aef-9ead-0802752e0920 · inbound

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging cites this paper.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.557385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.557385Z digest=sha256:0ec08bfa9654773d3430615acc57675a778c6c49e1eb40048b940240c57af673

Observation f0addaf2-1bbc-44bb-a1ac-08568136d1bd · inbound

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining cites this paper.

Dynamic Loss-Based Sample Reweighting for Improved Large Language Model Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:36:36.566504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:36:36.566504Z digest=sha256:966ad59f8965f4c0ec0f1b1598910219b67501653afcde358624ff5aaf1b879a

Observation 3daf7f01-a17b-4f35-84cb-77d48c8e660b · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:34.937484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:34.937484Z digest=sha256:eb21e4b30754e751c5e9d90b58b8516ea3dcee3ba59f6f0f75384611700e750f

Observation 06eaf234-dbfa-4b6e-8877-ee3e34034ffb · inbound

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects cites this paper.

ExeSQL: Self-Taught Text-to-SQL Models with Execution-Driven Bootstrapping for SQL Dialects DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:03.006603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:56:03.006603Z digest=sha256:d43da9dfdf8820f4daf24b8ec994850f91e1423d33fb9ee7e5ff85d8b3a9ecb6

Observation b4303d09-6850-428d-9ed4-32fa963c15ef · inbound

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining cites this paper.

GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:22.550042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:04:22.550042Z digest=sha256:05256b577821ba5c88b3567d7139f1161cec077994afdd36fc320e59664e90c3

Observation 8d4d2737-a203-4e5d-a27c-6133fb09a45f · inbound

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives cites this paper.

Rethinking Data Mixture for Large Language Models: A Comprehensive Survey and New Perspectives DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:54.125833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:54.125833Z digest=sha256:93ffb6e85767be8eb65fdd781896067941b3071201b45333d68362f0ddf1c843

Observation dde6cae9-73d6-4065-836e-a0cc34f9af89 · inbound

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning cites this paper.

SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:10:28.811231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:10:28.811231Z digest=sha256:13fb369ad08d13df108503ba7f906abbfdf396d9af0be063fa2ee11799655a7b

Observation 41496677-dd96-420c-a33b-7e258923b5e6 · inbound

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models cites this paper.

Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T22:41:53.033134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T22:41:26.047957Z digest=sha256:123b550f281acf53b7e5671037021bea896a154b57c32d4956a10d9d25c9558f

Observation d1cb1622-3927-4ef0-8546-d8c72aa83fd5 · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.127336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.127336Z digest=sha256:7ab4645b2c2b13a91c501c3176126a4101ddfaf1340580b8daddc4c979d583f1

Observation 3fd118da-1829-4887-8de4-76839afb8443 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.534644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.534644Z digest=sha256:4c3651b1d2c96e4e47ecd2ae494b166cf4ac1465ea6a8edd219e85b3bb147d73

Observation dbceb4dc-a267-431b-bb0f-2776d01a6a1e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:19:27.466265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:45815bcbf93b6ccb91e4b7c33a1c1a8141463d5acd80e163b7edac7a950515df

Observation 38133877-a792-4574-a582-bc635b36b31a · inbound

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics cites this paper.

Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:57.184917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:52:38.167538Z digest=sha256:61059898a769bcce4552eaf0f769372c8b604561e4dcfd6baa3624d8dd4b968d

Observation a4e4bb3b-96a9-41fa-b64a-769f3af999c7 · inbound

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining cites this paper.

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.135670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T19:04:11.976747Z digest=sha256:f845e8e33741fb0ee7da6767d38145b2ee206c3e0e169cbdcd845f82708f6227

Observation c4ea9473-5c1b-4fbe-83dd-040053d71e34 · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.445070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:ef644b10f5747a05b51f600d64545f539d03036726083b4349b8b0fe95372774

Observation 1fc5d40f-f130-4a86-9e41-b585abd7572d · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:20.625282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:20.625282Z digest=sha256:0f03ff97c2bea007c26a4f68d1cd5da83895db430c73cb09ca37a837bcce26ef

Observation 66f4baed-efb2-4bff-8050-86e767bad1c8 · inbound

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement cites this paper.

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:15:07.227138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:15:07.227138Z digest=sha256:0c97a25f78b69b59cd7a0fc9fb8db89091b1b690c03bcfed86c01923617b86e1