Pith. sign in

Paper Citation Record · LEDGER

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

As of 9 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.10952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10952 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:11.124746Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T00:56:04.958757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T00:58:25.710039Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5150f381-6e00-4937-a600-9f7df201e13c · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:04.749969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:04.749969Z digest=sha256:fde220b67faa32160a8aa61a1cd9af764a1c716a7af0d9f117ae1ff88cf3f73f

Observation 20a927a0-ef93-482f-9dce-ace099b5b482 · outbound

This paper cites and Vassilvitskii, S.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Vassilvitskii, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.885756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.830069Z digest=sha256:a1496b201fdbaa7dfe406796ea9c56855b7b32e5dffd499685f52dcc2670ba6f

Observation e622fe09-9584-446d-8696-f8235fbc7bf7 · outbound

This paper cites L., Gao, J., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Gao, J., and Choi, Y

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.575514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.986925Z digest=sha256:ba386edf6ec136b0811a0a020d0af108505629d90414a639579c07bf46d66e98

Observation ca118c46-9f2f-4d5f-8f25-5724cf2b0f37 · outbound

This paper cites Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:12.832198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.168093Z digest=sha256:e894c6e36e56102a5776cf59a5f830414afc0eb31236a0314bbad57e59f1f277

Observation 6e43ce2c-2b94-47ff-88f7-1355d77d3822 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.292132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.292132Z digest=sha256:518d722f5cac4872f14be92113dcd6449fe19357c30184c3e5d27cba7f3bc548

Observation b67c0f7c-b241-45bd-9131-ad2d9b7e24a1 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.428897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.428897Z digest=sha256:5587d0871adc56bce80f3279f298ac46f11facd1bdf49ef1a3b080f0c87a3457

Observation 4fc679f5-287e-45e9-84f6-f825520e8cd5 · outbound

This paper cites DOGE : Domain reweighting with generalization estimation.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training DOGE : Domain reweighting with generalization estimation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.238643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.546635Z digest=sha256:5c2eb20a0351c4256e0433349ebe28320dfd42360bb1eb1522848a35f286e5b5

Observation d7b03bc9-e2ae-4f46-a8bf-54061c6bc317 · outbound

This paper cites Unearthing Large Scale Domain-Specific Knowledge from Public Corpora.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.673354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.673354Z digest=sha256:d8ba1ff6c262dbe863a51203d1c538aed1afcef967b6277d0af78b4bb9f9615b

Observation 795ba27a-4d48-4890-8f46-906a2005cd65 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.821830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.821830Z digest=sha256:22ea8a865eb7d980545f4d562d8164ab7cc9bced8da477d9b75f03db97aec778

Observation eb5ad669-b4d0-44a5-9ca1-045173519542 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A framework for few-shot language model evaluation, 07 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.974554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.974554Z digest=sha256:9b6d091c0010151dd553ce17a5142a0343e0f451322398ee3b7f1d8505d9d112

Observation ee14808c-4cd9-46ba-ba7e-91720ffc2156 · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.110067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.110067Z digest=sha256:4a9fe8c3f1c9490d976ff1dc500e7560bb7722d3a435050cbd29793ed66b08e1

Observation 6b3fe130-c88d-48fb-acfe-7a3ef6219159 · outbound

This paper cites S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.773659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.244720Z digest=sha256:3b4c05c02d413862cf4a000a0971a1347af632be5433ad7f36f0fe795aa2d04b

Observation 1a311c98-2493-46fa-a0c2-74b275e9c9fd · outbound

This paper cites The Llama 3 Herd of Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.370050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.370050Z digest=sha256:a263143c21c30288784aff8a85137f983ba589cf9a6a481e02032dadc82da287

Observation 4fa1f33a-c468-46fb-8b96-451521052e48 · outbound

This paper cites CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T04:19:11.583751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.514933Z digest=sha256:87f67afac3af309100ab8e130655635433c8f9b7063dce8a006463e8c082e050

Observation c9b8c866-aae0-4ab3-bb09-4e4a55895a47 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Selection via Optimal Control for Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.688345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.688345Z digest=sha256:320cf9b08e77a4202d7faed06706ef33be16bcd0457ab015dd0b1bfb047c30a2

Observation 3037b8d0-57e6-4420-bf01-73a4c64c7fc2 · outbound

This paper cites V., and Smith, K.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training V., and Smith, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.402051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.781962Z digest=sha256:07ae76e31d32e2c55dec266b0854e04d07bc3bf427d8462e4de7353922ae74df

Observation bc0ecc10-2344-4fb4-bd57-c67cda0083c9 · outbound

This paper cites H., and Friedman, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training H., and Friedman, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.947572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.947572Z digest=sha256:a0a0338a403100b4f2fee7d8271f686fd990eedf9d9696c33a38b788338b66b6

Observation 2b4c0efb-c935-4e9f-9c93-bdc89412ff13 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.085550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.085550Z digest=sha256:2ed7bf29ef373ee0ecb43037c397220a73710897cc169a6ded4634fc3d5b12ad

Observation a61d6b29-1bab-42bc-8290-daa94f7ee6d0 · outbound

This paper cites A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.067075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.146299Z digest=sha256:f1d1b35462b7e22151677151de002e4f5400ed988cfa179276563b2faa14547f

Observation b0aca12f-c72f-4a13-a262-ec07c2960b10 · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.273463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.273463Z digest=sha256:8a18d1aab37a67af830ba0983b05c6c84be3985a5b6ef13afa18c91c3a4937a5

Observation 5db010b9-0585-4256-a111-bd59bea0bf09 · outbound

This paper cites S., Schmidt-Thieme, L., and Grabocka, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S., Schmidt-Thieme, L., and Grabocka, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.661402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.373042Z digest=sha256:cadbb3d6fca5f8fea4b6d3dab0f350dc80ff848302490355674b71ff7100fe15

Observation 3d05bc31-1160-47d7-85c3-f0b2bef852bc · outbound

This paper cites Scaling Laws for Neural Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.509003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.509003Z digest=sha256:639ffd4f6ee6d8c3ac57d8470d39de1db4b5ea0771424847c89bdffb0148b0de

Observation 28b1bdb6-0513-417c-a542-9ff30d953169 · outbound

This paper cites Lightgbm: A highly efficient gradient boosting decision tree.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Lightgbm: A highly efficient gradient boosting decision tree

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.348865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.670501Z digest=sha256:b7368eed632dd6f4931c0ffdd6c3f4073987231880280c24e9e1dda7e517788a

Observation fa5eec87-5d99-4c32-adfa-3dd4131dab22 · outbound

This paper cites Looking beyond the surface: A challenge set for reading comprehension over multiple sentences.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Looking beyond the surface: A challenge set for reading comprehension over multiple sentences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.849579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.849579Z digest=sha256:cb1dd5d2558e6066befeda2359be890002d179cbce8a68c2168fdb4176c2fad5

Observation d515f4e5-b32b-4dbf-80d4-dad73e26adaf · outbound

This paper cites RACE : Large-scale R e A ding comprehension dataset from examinations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RACE : Large-scale R e A ding comprehension dataset from examinations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.967559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.967559Z digest=sha256:500487ef2cc27c52f2830f448374453fe40a1f080aab9e0585a227674bf70dfc

Observation 2a9418f5-9f61-4c41-aca4-fbc309ea8ab9 · outbound

This paper cites Not all tokens are what you need for pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Not all tokens are what you need for pretraining

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.966571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.109415Z digest=sha256:3a255610daf52f49996bc2c4f79364b2c62415f863b4466f87ab94c081373a20

Observation 3641c731-e492-4180-a9ec-9cab909a1615 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.255118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.255118Z digest=sha256:164e42487de36e3d89fe1d47cc673ab404e68074aa1d5b547c8c97efec5969e5

Observation a96b27b8-a264-4bd2-9990-f3f620b06f1f · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.394746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.394746Z digest=sha256:b62256229e61f16460299ba6eb7c5d6868416d489db07f888180cc313fe7d12b

Observation 15ebcb6b-aab6-4b5b-b14e-3ffaaee57b9a · outbound

This paper cites Decoupled Weight Decay Regularization.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.491601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.491601Z digest=sha256:c6863c37026087bf8dffd9bdbaef8047fe88b8940fb36e986ec5abddc33fb8e2

Observation e81c65ec-f9c6-461e-842d-d440cb8af013 · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Some methods for classification and analysis of multivariate observations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.582732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.599517Z digest=sha256:e72621d28982120acfcc657784a1458e9d5f7334d5ebc8c705beeeac055529ec

Observation cda02639-cf84-467b-b204-4ac489126ff3 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.801169Z digest=sha256:2c8e043eb1610e4e4fddefdbce21f2291a0fa1e7ccb4fa8beb79d37deb58dc45

Observation 3986d171-07d7-4d1c-bec3-6d88314a4849 · outbound

This paper cites GPT-4 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.901499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.901499Z digest=sha256:2c0b9c3acc7176ea51187f3af625cb1190cd757a3cfbca5af65571d587900a37

Observation e8bb77d4-9e4d-41cc-918e-7953c2e48fe8 · outbound

This paper cites Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.056334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.056334Z digest=sha256:48c20ae59db8db7bbdc686d6e9f552c8cae2fc1c7e56ecb6fa70a5d88f51f3fa

Observation 909bae5a-a467-4000-9335-6ef00d7e9f3e · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.079631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.206680Z digest=sha256:5600d726aa45af2a23be29960e778c909513c45e08ab72fd1364c32f77acbda5

Observation 51f31caa-db4d-4702-af24-3e3dcc179a9e · outbound

This paper cites D-cpt law: Domain-specific continual pre-training scaling law for large language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training D-cpt law: Domain-specific continual pre-training scaling law for large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.801841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.316683Z digest=sha256:ffd04681246fe6e15644d234da1226be9c2d683cf81ba2febe3d9787c1589e0b

Observation df65be68-87cf-4d85-b930-c297b992b76f · outbound

This paper cites Qwen2 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.397800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.397800Z digest=sha256:be75f8192dda9a104f44be49556ae1b2565ae8dd8530f525e12459db2e6db0a8

Observation 128ba91b-2acf-4fdf-93c8-924dedff6d98 · outbound

This paper cites Improving language understanding by generative pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Improving language understanding by generative pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.508809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.496631Z digest=sha256:ee8ba691a4c540688066525374f8ea55164aaa4922f1e44c7e6d738ec78f5de4

Observation 7ae6989d-d90e-4cca-80c3-d7965b74c2c7 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.584692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.584692Z digest=sha256:6de28b3559b45fa8af1f71c7a8c07303d8c9282319c83d61d07af4117ed2019a

Observation 9bb47e5a-a264-4656-84c5-30e91f78cbdd · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.654342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.654342Z digest=sha256:511026c14503e17f46171edfc61458030ad486785e1c988f8d48298266994c57

Observation c7c0f0f9-6a3a-437a-9426-cb6192962d62 · outbound

This paper cites W., Hashimoto, T.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training W., Hashimoto, T

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.271649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.742647Z digest=sha256:7b1a5e3c7f1b43d518e0da948b766c502140641a3e5058840500b3975e5a2bd8

Observation 02d10b5b-4148-4fd0-9184-c40bedc80ebd · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Bhagavatula, C., and Choi, Y

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.846389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.846389Z digest=sha256:3879bd5a5093a3014f10a11db922eb48785cf7ea5e93d838f54a998d8ebc43b2

Observation 68496c35-c5de-45b2-9d60-9de0dc544f55 · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Social IQ a: Commonsense reasoning about social interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.950458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.950458Z digest=sha256:aa43df9a141b6c05bdea4c2a993d2f14fa8fe12a46843ebfee3e227c38f1540a

Observation b7c78fb7-fa74-402b-a247-29b0773f37ec · outbound

This paper cites Self-influence guided data reweighting for language model pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Self-influence guided data reweighting for language model pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.942881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.037663Z digest=sha256:64901f1e913f8c71b983f4eab501bf00a0990548f740359a0315385c86306885

Observation 453e0082-ae34-4f93-85eb-2bc47ef8a6ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.099522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.099522Z digest=sha256:4e67a2b0ee88c2bc6ff55ba78bd9e9e98386284013d0197a1caf70c8c5ac57ff

Observation 567f32ab-526b-40cb-b391-f820d1d0e369 · outbound

This paper cites and Hinton, G.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Hinton, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.178365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.178365Z digest=sha256:7f93c0ffea21a227a6419c6df7648a12d45b5ae4e811a3e1022747f30b4ba49c

Observation f881e0e2-8d5b-4e42-93fc-65d87c471d6c · outbound

This paper cites Learning Dynamics in Continual Pre-Training for Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Learning Dynamics in Continual Pre-Training for Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:11.911632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.282840Z digest=sha256:eec4f0f9828d91919fed6d2e730ef3c43cc17176fa586c8734e4b64bb984fbe5

Observation 2a633f51-b59d-4224-8915-6a1fad4c7ad1 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RedPajama: an Open Dataset for Training Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.350194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.350194Z digest=sha256:c63b745b8c45c6178bd775b3293e8982b373d55638e6e5eb346dd1984d17b4d5

Observation 94fc3b50-0d5a-4d2e-846a-11952fb1e08a · outbound

This paper cites F., and Gardner, M.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training F., and Gardner, M

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.444975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.444975Z digest=sha256:0e2d753593c9356301f1263d55225f47eb98a55fec776e8ffcbb0b0a6df3be7a

Observation 5c98e9f7-60a3-4e21-8647-30badaad4d2f · outbound

This paper cites C-pack: Packaged resources to advance general chinese embedding, 2023.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training C-pack: Packaged resources to advance general chinese embedding, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.606366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.532719Z digest=sha256:392a64ccfa0c165a7cdb50780d72dfda36b97f1d476e15e3900d61e8eb83da08

Observation 2f36a394-6086-419d-92fd-cff21e86821e · outbound

This paper cites M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.317126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.626223Z digest=sha256:5a5ad04da409b4eb548cd9333002cad87cbb569da3c3b93ac9bb386960f56437

Observation 968da2f7-e792-43ea-8804-346f988730db · outbound

This paper cites M., Santurkar, S., Ma, T., and Liang, P.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Santurkar, S., Ma, T., and Liang, P

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.932011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.703150Z digest=sha256:ee4af8184d3cf6ee1fd7785e50bba751ec244a57cd36870db9d3f05d7b261446

Observation 024a254e-8ef7-4f2c-b59d-0c32fcacb218 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.793759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.793759Z digest=sha256:20cc67504f8b657ffd01bdf4065947e39772973242de5e6289613d7e229e44e5

Observation ed166475-b0e1-41f5-aae6-c00ec04ecbac · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.596553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.914414Z digest=sha256:032619e75cb55833f8ce9ff798e74ce4208c2393bf60fd34b8a9872cea06d9d1

Observation dbc2900e-8802-4a8e-b9b1-867931298325 · outbound

This paper cites LIMA : Less is more for alignment.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training LIMA : Less is more for alignment

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.143966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T04:19:11.007446Z digest=sha256:6ac48d6dad11d231dc22abc4b30daf4b680147d45d8e930538303075d87726cf

Observation 9e46c7fd-fb5b-4c1a-809a-a2f94e6b0792 · outbound

This paper cites write newline.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:11.124746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:11.124746Z digest=sha256:0fce61db8682cfb149570e154977de146c0d576e622aad6c72b54381ae524742

Pith citing papers

Observation 75014b2e-a9cb-4c56-b770-895e9292b063 · inbound

Data Mixing for Large Language Models Pretraining: A Survey and Outlook cites this paper.

Data Mixing for Large Language Models Pretraining: A Survey and Outlook Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.711591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T00:56:04.958757Z digest=sha256:213f596737753f5aeccf65c8142de4204b330bc8c92b15732715838996afa380