Pith. sign in

Paper Citation Record · LEDGER

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2506.10952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10952 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:19:11.124746Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T00:56:04.958757Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T00:58:25.710039Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5150f381-6e00-4937-a600-9f7df201e13c · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:04.749969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:04.749969Z digest=sha256:83c6bd3f65b8e36554edc6bc3dbb1bb3e99686a4b68f49da731abfb00d425944

Observation 20a927a0-ef93-482f-9dce-ace099b5b482 · outbound

This paper cites and Vassilvitskii, S.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Vassilvitskii, S

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.885756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.830069Z digest=sha256:1d75256715cd988b4a842075a8f1145683377344177d7a8372d1b1a68343d92d

Observation e622fe09-9584-446d-8696-f8235fbc7bf7 · outbound

This paper cites L., Gao, J., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Gao, J., and Choi, Y

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.575514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:04.986925Z digest=sha256:3cdce22f64d3b68644bd2edd601d6840f6a7dc8bbe71f1b03bf6361f0a17b6ac

Observation ca118c46-9f2f-4d5f-8f25-5724cf2b0f37 · outbound

This paper cites Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Cross-Table Pretraining towards a Universal Function Space for Heterogeneous Tabular Data

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:12.832198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.168093Z digest=sha256:d01704b1139da562dd828627ae54ef86bdcbe475fdb9daf87196285cbf9670fd

Observation 6e43ce2c-2b94-47ff-88f7-1355d77d3822 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.292132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.292132Z digest=sha256:e260c6bbf4708af7b6943944fe8d586dd4c5ab127d98f2619e70cb57ef930e27

Observation b67c0f7c-b241-45bd-9131-ad2d9b7e24a1 · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.428897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.428897Z digest=sha256:3b1fa049f72523d2674c1f40d27042f102bd49144097cb09c4fac370a52614e0

Observation 4fc679f5-287e-45e9-84f6-f825520e8cd5 · outbound

This paper cites DOGE : Domain reweighting with generalization estimation.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training DOGE : Domain reweighting with generalization estimation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:19.238643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:05.546635Z digest=sha256:9c7be5929e21c87f295ab01f64a92ac979f7664e0e33b7bbf7a73e052e91e55d

Observation d7b03bc9-e2ae-4f46-a8bf-54061c6bc317 · outbound

This paper cites Unearthing Large Scale Domain-Specific Knowledge from Public Corpora.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unearthing Large Scale Domain-Specific Knowledge from Public Corpora

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.673354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.673354Z digest=sha256:75065f590dfb63e810c53f605cbe975f6a843307b8c18aa75620606591523f60

Observation 795ba27a-4d48-4890-8f46-906a2005cd65 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.821830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.821830Z digest=sha256:e4682b50e223a19d4e267282fab58bd394f505a2798edeaf3d3f7bcf26777083

Observation eb5ad669-b4d0-44a5-9ca1-045173519542 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A framework for few-shot language model evaluation, 07 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:05.974554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:05.974554Z digest=sha256:57405b806b43eb222c2e27a4d55b84fb6fc1880ca6dec562ce31624cfd779fc6

Observation ee14808c-4cd9-46ba-ba7e-91720ffc2156 · outbound

This paper cites BiMix: A Bivariate Data Mixing Law for Language Model Pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training BiMix: A Bivariate Data Mixing Law for Language Model Pretraining

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.110067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.110067Z digest=sha256:4915d949688caf9bbfb3b185d96b3e90e7208df62b3d94242f744beb316fefca

Observation 6b3fe130-c88d-48fb-acfe-7a3ef6219159 · outbound

This paper cites S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S em E val-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.773659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.244720Z digest=sha256:6d2454c9a0a954fe0cb74e17a7e5b1ebed56e200c56aedd1449b54fbbf3bc903

Observation 1a311c98-2493-46fa-a0c2-74b275e9c9fd · outbound

This paper cites The Llama 3 Herd of Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.370050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.370050Z digest=sha256:f1fde813692a36ed6df9f9d843738b672a56d3a67eb5945d5f520d768922a917

Observation 4fa1f33a-c468-46fb-8b96-451521052e48 · outbound

This paper cites CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training CMR scaling law: Predicting critical mixture ratios for continual pre-training of language models

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T04:19:11.583751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.514933Z digest=sha256:48906bb9e4068fcaaf25185b5f81427f0366a587b1a57666bec4690620b49ae4

Observation c9b8c866-aae0-4ab3-bb09-4e4a55895a47 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Selection via Optimal Control for Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.688345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.688345Z digest=sha256:d3946fe4586b38cf787bd717b250e2d9c947a13b452dc73fce161c9deb1f04d9

Observation 3037b8d0-57e6-4420-bf01-73a4c64c7fc2 · outbound

This paper cites V., and Smith, K.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training V., and Smith, K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.402051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:06.781962Z digest=sha256:97592de4dbbd3f19b4c81a6fa28a8d0ddf914b3c915a2800515d97e5b75318db

Observation bc0ecc10-2344-4fb4-bd57-c67cda0083c9 · outbound

This paper cites H., and Friedman, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training H., and Friedman, J

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:06.947572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:06.947572Z digest=sha256:1df3f54f1a08134ebbf5268a5ecba8a70154eaac6e313de3fca4b966a5be9cd0

Observation 2b4c0efb-c935-4e9f-9c93-bdc89412ff13 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.085550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.085550Z digest=sha256:27b085716e0a3d5c77b4214be80a7c957ac3a60dcdd9599970fc213c755872f2

Observation a61d6b29-1bab-42bc-8290-daa94f7ee6d0 · outbound

This paper cites A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:18.067075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.146299Z digest=sha256:2f7655b94bf07909d877198ff555ae1b7228a803eda83674ff7fe3e2202a6c5d

Observation b0aca12f-c72f-4a13-a262-ec07c2960b10 · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.273463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.273463Z digest=sha256:2bf2d33e0157573de421ec172e441dc2216a77cc1194268eabfcc35021dfdf30

Observation 5db010b9-0585-4256-a111-bd59bea0bf09 · outbound

This paper cites S., Schmidt-Thieme, L., and Grabocka, J.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training S., Schmidt-Thieme, L., and Grabocka, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.661402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.373042Z digest=sha256:a783bed5a4afa5bbaf5b9b8c2a0a296312ada6c203b116dbb613713aa6ceb466

Observation 3d05bc31-1160-47d7-85c3-f0b2bef852bc · outbound

This paper cites Scaling Laws for Neural Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Laws for Neural Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.509003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.509003Z digest=sha256:e770a9f41b5f8e22c3412349f2cd3a2042c51b2a89105d6e0e580e3f64e95f01

Observation 28b1bdb6-0513-417c-a542-9ff30d953169 · outbound

This paper cites Lightgbm: A highly efficient gradient boosting decision tree.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Lightgbm: A highly efficient gradient boosting decision tree

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:17.348865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:07.670501Z digest=sha256:0cdd3f321dfb358d6593cece49c9d5473bd0ab2e2d6a54ff8484dd777454647c

Observation fa5eec87-5d99-4c32-adfa-3dd4131dab22 · outbound

This paper cites Looking beyond the surface: A challenge set for reading comprehension over multiple sentences.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Looking beyond the surface: A challenge set for reading comprehension over multiple sentences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.849579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.849579Z digest=sha256:05c6e1829bf0e0458265fa99cb4eac4224d6377d1ef6bafc59e38b38e65e7041

Observation d515f4e5-b32b-4dbf-80d4-dad73e26adaf · outbound

This paper cites RACE : Large-scale R e A ding comprehension dataset from examinations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RACE : Large-scale R e A ding comprehension dataset from examinations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:07.967559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:07.967559Z digest=sha256:036c09738a438674f93f6a9d25d5bfa20ffb2ddb46c504cd6de01327443b32cb

Observation 2a9418f5-9f61-4c41-aca4-fbc309ea8ab9 · outbound

This paper cites Not all tokens are what you need for pretraining.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Not all tokens are what you need for pretraining

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.966571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.109415Z digest=sha256:d5ed2a43310f7b75f6cd057892ac1c61a9293539d45be8759f82daca0ba91031

Observation 3641c731-e492-4180-a9ec-9cab909a1615 · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Logiqa: a challenge dataset for machine reading comprehension with logical reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.255118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.255118Z digest=sha256:b58a34b31195ebd298c3ea7feaae5dc33f3ba52f717da50c654a10be2588876d

Observation a96b27b8-a264-4bd2-9990-f3f620b06f1f · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.394746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.394746Z digest=sha256:e172e4a7113ba2102c5cae95093cc5352e7acca5c47ebe6782efe08754b212bf

Observation 15ebcb6b-aab6-4b5b-b14e-3ffaaee57b9a · outbound

This paper cites Decoupled Weight Decay Regularization.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Decoupled Weight Decay Regularization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.491601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.491601Z digest=sha256:723e24c3ba91b6c094a526b573129299f3f415af38428b859e94b0aeb49a5e86

Observation e81c65ec-f9c6-461e-842d-d440cb8af013 · outbound

This paper cites Some methods for classification and analysis of multivariate observations.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Some methods for classification and analysis of multivariate observations

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.582732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.599517Z digest=sha256:df7d23b9c24cebba8b153954b6d30ef55311353bc025ad668efb81840526dafe

Observation cda02639-cf84-467b-b204-4ac489126ff3 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.337236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:08.801169Z digest=sha256:8c927120adcea5dbd0516f35e8bff3c5189912be9769821a3c369b504adc9b60

Observation 3986d171-07d7-4d1c-bec3-6d88314a4849 · outbound

This paper cites GPT-4 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training GPT-4 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:08.901499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:08.901499Z digest=sha256:aed7c634bcbf3b017d81c34d900171b6a1212915526d6e8927e370f5a37c4fb9

Observation e8bb77d4-9e4d-41cc-918e-7953c2e48fe8 · outbound

This paper cites Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fern \'a ndez, R

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.056334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.056334Z digest=sha256:ff59bb10aef97ecb96fe0ba100dbf392a141198ee6cb820b547ff6cab637ea1d

Observation 909bae5a-a467-4000-9335-6ef00d7e9f3e · outbound

This paper cites B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:16.079631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.206680Z digest=sha256:a1d49935a2d918b65e2ec1f307cdb7297127f3ce7a34090eec22439a5bf4e908

Observation 51f31caa-db4d-4702-af24-3e3dcc179a9e · outbound

This paper cites D-cpt law: Domain-specific continual pre-training scaling law for large language models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training D-cpt law: Domain-specific continual pre-training scaling law for large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.801841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.316683Z digest=sha256:df3194b126bb2256a6b2cef3ad7f5ac93eac62d1eecfabd17dd4f14506b32b5b

Observation df65be68-87cf-4d85-b930-c297b992b76f · outbound

This paper cites Qwen2 Technical Report.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.397800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.397800Z digest=sha256:071429c1a5028d10ade50527a3509ac0e9502f12ed3d0b4341da373ad053f727

Observation 128ba91b-2acf-4fdf-93c8-924dedff6d98 · outbound

This paper cites Improving language understanding by generative pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Improving language understanding by generative pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.508809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.496631Z digest=sha256:38dc036076dc8c0025ba08a7308d9dda89241acdd984aa146b3796b885f8974a

Observation 7ae6989d-d90e-4cca-80c3-d7965b74c2c7 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.584692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.584692Z digest=sha256:fcf93743c988a0fdabec2cf2cb3737df6fd05bd51644be2d143d7c8608faa756

Observation 9bb47e5a-a264-4656-84c5-30e91f78cbdd · outbound

This paper cites an unresolved cited work.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.654342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.654342Z digest=sha256:0eacf4351a9bfed52a3876632732aa28f791b3dd42336274ba9108e266af9c70

Observation c7c0f0f9-6a3a-437a-9426-cb6192962d62 · outbound

This paper cites W., Hashimoto, T.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training W., Hashimoto, T

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:15.271649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:09.742647Z digest=sha256:e9564ab271b275a52ff6cc805d24d37641d278c862eccbbdae33e79e9076db09

Observation 02d10b5b-4148-4fd0-9184-c40bedc80ebd · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training L., Bhagavatula, C., and Choi, Y

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.846389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.846389Z digest=sha256:cb45b53dfb24366eaa9179fa75d8fc83a53c87d2e764b0904be74cb34cbdc785

Observation 68496c35-c5de-45b2-9d60-9de0dc544f55 · outbound

This paper cites Social IQ a: Commonsense reasoning about social interactions.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Social IQ a: Commonsense reasoning about social interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:09.950458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:09.950458Z digest=sha256:0e09bfc01943c377cb1ca65c4eee591e6d1060441402725e515df49aab0a740b

Observation b7c78fb7-fa74-402b-a247-29b0773f37ec · outbound

This paper cites Self-influence guided data reweighting for language model pre-training.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Self-influence guided data reweighting for language model pre-training

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.942881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.037663Z digest=sha256:7a6abed3f87ae9290aaee1a6717d7356afe165470d27388bd6260203da0f7daa

Observation 453e0082-ae34-4f93-85eb-2bc47ef8a6ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.099522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.099522Z digest=sha256:fc667579e3daec4cd6498540f5b6bdd042efe4463b04fb5d477386428bb22436

Observation 567f32ab-526b-40cb-b391-f820d1d0e369 · outbound

This paper cites and Hinton, G.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training and Hinton, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.178365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.178365Z digest=sha256:d074a398c30ea60efd74604ffc4c72ceff2b338c2253a5c335a80f5df6e04bb1

Observation f881e0e2-8d5b-4e42-93fc-65d87c471d6c · outbound

This paper cites Learning Dynamics in Continual Pre-Training for Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Learning Dynamics in Continual Pre-Training for Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:19:11.911632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.282840Z digest=sha256:0cad27c49dfc5e496704a5f137cbb98276f450e384dc6d48b024bf90bd4ce67a

Observation 2a633f51-b59d-4224-8915-6a1fad4c7ad1 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training RedPajama: an Open Dataset for Training Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.350194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.350194Z digest=sha256:590b5150760573602e47e4075fbe7df6a96648f72c06fb3210dc2551d907c4a7

Observation 94fc3b50-0d5a-4d2e-846a-11952fb1e08a · outbound

This paper cites F., and Gardner, M.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training F., and Gardner, M

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.444975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.444975Z digest=sha256:c6c6576b487d2c5a92c0fedb617668d54071c49ee171ac70854b173857d69221

Observation 5c98e9f7-60a3-4e21-8647-30badaad4d2f · outbound

This paper cites C-pack: Packaged resources to advance general chinese embedding, 2023.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training C-pack: Packaged resources to advance general chinese embedding, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.606366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.532719Z digest=sha256:ffc51b02ddb148e26fe032a92506bc502eafa80d3a08003155c224b6d3dfa6e7

Observation 2f36a394-6086-419d-92fd-cff21e86821e · outbound

This paper cites M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:14.317126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.626223Z digest=sha256:37282afa36d8fd1d23c37280d64ebf704c2c5952fe81a7606ddebcbdc2cc21ef

Observation 968da2f7-e792-43ea-8804-346f988730db · outbound

This paper cites M., Santurkar, S., Ma, T., and Liang, P.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training M., Santurkar, S., Ma, T., and Liang, P

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.932011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.703150Z digest=sha256:789d6dd3435aebb081e112d519fcd6cdd57f0f5a71b0c830ba5e337625769d6d

Observation 024a254e-8ef7-4f2c-b59d-0c32fcacb218 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:10.793759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:10.793759Z digest=sha256:6e20426e0cd0273d32199673410230a6cea147989edd04dc63f90adbcabfa643

Observation ed166475-b0e1-41f5-aae6-c00ec04ecbac · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training Hellaswag: Can a machine really finish your sentence? In Annual Meeting of the Association for Computational Linguistics, 2019

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.596553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:10.914414Z digest=sha256:ba0e7701a65eff98d0ba0abbc63f62d53450f883ec30e35d20d1be5f18bfde6d

Observation dbc2900e-8802-4a8e-b9b1-867931298325 · outbound

This paper cites LIMA : Less is more for alignment.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training LIMA : Less is more for alignment

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:19:13.143966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T04:19:11.007446Z digest=sha256:5ac54c4a0b41acceb79b3d0c5d5e1778a64f9714f970daeabc5debcb01167946

Observation 9e46c7fd-fb5b-4c1a-809a-a2f94e6b0792 · outbound

This paper cites write newline.

Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:19:11.124746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:19:11.124746Z digest=sha256:e5cf90fada1bd7a7e6e292e0a1fbc2bc865f852c9597e09399ed1d482c2d8715

Pith citing papers

Observation 75014b2e-a9cb-4c56-b770-895e9292b063 · inbound

Data Mixing for Large Language Models Pretraining: A Survey and Outlook cites this paper.

Data Mixing for Large Language Models Pretraining: A Survey and Outlook Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:58:25.711591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T00:56:04.958757Z digest=sha256:c0676d1d19d22817240d1ed5ec85a19d2968e71a2203383920856b85b3512ec1