Pith. sign in

Paper Citation Record · LEDGER

Position: The Most Expensive Part of an LLM should be its Training Data

As of 17 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 5 inbound Pith citation observations for arXiv:2504.12427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12427 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:34:53.589828Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:29:44.273599Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:39:46.610216Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ffee95a-85f2-44e2-8647-575ce81577e8 · outbound

This paper cites https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025.

Position: The Most Expensive Part of an LLM should be its Training Data https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.222109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.416080Z digest=sha256:79f99a030fab5a52d56d741172acc00b42944af7c64e850b08b850ae307ee19c

Observation c0836bd0-436d-4232-a7b5-5f1de43df576 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.420387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.420387Z digest=sha256:5885b20f0ed73889963570fbffa6969e3489d4b9032a189c0c6fc9eb1f209f04

Observation 29a1ac42-8e9a-48f9-bc48-2966b1f50e70 · outbound

This paper cites Phi-4 Technical Report.

Position: The Most Expensive Part of an LLM should be its Training Data Phi-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.424961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.424961Z digest=sha256:1280967d77d782a95b92a25aa8eb735bfa8881e30005f24254eb726b87a50b90

Observation 94a2aa78-8802-47da-b446-a0bfd1abc6e2 · outbound

This paper cites Alden newspapers v.

Position: The Most Expensive Part of an LLM should be its Training Data Alden newspapers v

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-16T12:34:53.976485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.429188Z digest=sha256:5652019cdb0f4d3d4e6bcabd9288644426e289021f15a40f027b44b956d9d8de

Observation 60d5f8ad-9b5f-4da0-bfcd-50cdfb622b42 · outbound

This paper cites M., and Weber, G.

Position: The Most Expensive Part of an LLM should be its Training Data M., and Weber, G

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.433029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.433029Z digest=sha256:900a9acd878486d808addd59e2195871c70544d4978241671ace3ef6ef5b0882

Observation 7adfd76b-b4c3-45f7-b7cc-788c976aecab · outbound

This paper cites Authors guild v.

Position: The Most Expensive Part of an LLM should be its Training Data Authors guild v

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.206262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.437311Z digest=sha256:969951bb795b3e23415228060af1606beae6e64333027a9ba4f813f3ee49562a

Observation 72bb5d94-82ea-4e6d-abb9-10b563348510 · outbound

This paper cites What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.441434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.441434Z digest=sha256:b072f0e9f461e8340d0a8a4ad2c7d26753e46fb193737d8f2705e992053fb79a

Observation 558c1ceb-306d-4caa-86a9-1ef6e5f74c7d · outbound

This paper cites Common crawl dataset.

Position: The Most Expensive Part of an LLM should be its Training Data Common crawl dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.196771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.445464Z digest=sha256:417cc43ea2f97bfc65bf6c8480734d3462af237422f3cda7456ac3efc42ef1f5

Observation 3518d52d-387a-4a8d-9541-f23ebcd8f2f9 · outbound

This paper cites Concord music group v.

Position: The Most Expensive Part of an LLM should be its Training Data Concord music group v

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.187530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.449194Z digest=sha256:92d9685a3079f2cfbd93bb88437738b5a12715fb2322b8d5c1b4b5fbc7d38ff1

Observation 252eac39-7388-4717-be9e-6a70a8dc6d8e · outbound

This paper cites The rising costs of training frontier AI models.

Position: The Most Expensive Part of an LLM should be its Training Data The rising costs of training frontier AI models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.452722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.452722Z digest=sha256:bbe69cdca34e8247b709d05e2a97fe8e181edd3a96596e73861e11c7cf523b8f

Observation cbefb914-e2ca-4d11-93c0-b8d709d75485 · outbound

This paper cites Ai is a lot of work.

Position: The Most Expensive Part of an LLM should be its Training Data Ai is a lot of work

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.177432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.457066Z digest=sha256:f76c17022770657c43043c3e12667cdae8dc21277340980dbae1c6f1e1b2183b

Observation 0b9777a9-53d8-4c17-988e-4a47343008d5 · outbound

This paper cites Encyclopedia britannica 15th edition.

Position: The Most Expensive Part of an LLM should be its Training Data Encyclopedia britannica 15th edition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.167706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.460931Z digest=sha256:2d46d03f9d0284902a786c3c56ca9b0a601833494ebcc74a7a288c710423271f

Observation 9ac7fb88-d47c-487b-9be8-f0936f3ab916 · outbound

This paper cites Data on notable ai models, 2024.

Position: The Most Expensive Part of an LLM should be its Training Data Data on notable ai models, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.157863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.464492Z digest=sha256:b75cc6a84d3d996926cf8ca3d0e205adf2a622bbf8145809befac05ef851df95

Observation c957d152-6d94-4247-b56b-21a7ecc89e37 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp: In search of the next generation of multimodal datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.468171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.468171Z digest=sha256:e6ee44551355bb7113880b84ccc28cc9c2e440fbc3a65619413e7c13e6564216

Observation 01f09962-6c0c-46ab-a7ef-e7905f4ce083 · outbound

This paper cites Language models scale reliably with over-training and on downstream tasks.

Position: The Most Expensive Part of an LLM should be its Training Data Language models scale reliably with over-training and on downstream tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.471844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.471844Z digest=sha256:a95cc5518abd5a766912414108ad908744dd085ddbe94763674f8ad423fab7bb

Observation cba7a639-6920-48ad-a543-dbde1f5b53a7 · outbound

This paper cites Data Shapley: Equitable Valuation of Data for Machine Learning.

Position: The Most Expensive Part of an LLM should be its Training Data Data Shapley: Equitable Valuation of Data for Machine Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.475484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.475484Z digest=sha256:82d3a55aafa1b9415f4c61127a883b6e959042a52396debf497c9cf6c70e18bb

Observation 03d72fee-497e-469b-a1cd-5e46266b1faa · outbound

This paper cites Evaluation of Similarity-based Explanations.

Position: The Most Expensive Part of an LLM should be its Training Data Evaluation of Similarity-based Explanations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.479173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.479173Z digest=sha256:5e977d7b53181eaab64dd56dd43a7783672a0784bca2ab48a6f7ac091075208c

Observation 95e6842e-d01d-40fb-95c2-875558897655 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Training Compute-Optimal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.482888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.482888Z digest=sha256:cbd05e2e0ca914bf2c209cfd1720a689e1a39bc7d7bc708d99bd7e4ec822bc83

Observation e0e7818b-ac76-4f46-a3a7-5c125e939bf3 · outbound

This paper cites Statistics on wages.

Position: The Most Expensive Part of an LLM should be its Training Data Statistics on wages

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.147863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.487229Z digest=sha256:7c0fbcc87a658022aee70d67a7734c8d7444f9b1d91eca509bcc2153b3ed9683

Observation 6bad627c-5881-4537-af10-2065619543c6 · outbound

This paper cites Scaling Laws for Neural Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Laws for Neural Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.490580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.490580Z digest=sha256:205a0d3b639e40be97bd98e1c3426aa45b4d5fa8f35b2885129c69a4fe2e0d18

Observation b779aefe-5e3e-4bf2-9355-d243a13effb7 · outbound

This paper cites Understanding Black-box Predictions via Influence Functions.

Position: The Most Expensive Part of an LLM should be its Training Data Understanding Black-box Predictions via Influence Functions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.494514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.494514Z digest=sha256:5a58d1a492876a2d95d9e9c7273404a5391d563bd1130b9f32874760cab86630

Observation dc20d4d5-255e-42cc-80ea-f82e155d7c85 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Position: The Most Expensive Part of an LLM should be its Training Data OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.498197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.498197Z digest=sha256:9f0ef6ceced7937fb2def3d7f08b335ab2ca7f05c350c0c5cac69dbed145d1fe

Observation 2d526e30-7940-40ee-9d3b-5bf491476136 · outbound

This paper cites Releasing Common Corpus: the largest public domain dataset for training LLMs.

Position: The Most Expensive Part of an LLM should be its Training Data Releasing Common Corpus: the largest public domain dataset for training LLMs

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.138207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.501877Z digest=sha256:93682b1848d9563f134802c57901580d3b0a970e81b5d121cd63e2104f46ee53

Observation b173a557-7554-43ba-9c09-4b38c6b0d91f · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Position: The Most Expensive Part of an LLM should be its Training Data DataComp-LM: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.505422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.505422Z digest=sha256:80d529077c08dfb0771183234075187c05f39bd1a9291f887f8e5fa3935fec98

Observation 589d0d14-d539-4794-a32c-62452bcf0ab2 · outbound

This paper cites Consent in Crisis: The Rapid Decline of the AI Data Commons.

Position: The Most Expensive Part of an LLM should be its Training Data Consent in Crisis: The Rapid Decline of the AI Data Commons

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.509363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.509363Z digest=sha256:57645ed96cf441eaa37a436baf6df575dba951092aca01819e68891ad49371eb

Observation b9175506-313c-4a0e-88e3-645e166e0898 · outbound

This paper cites SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore.

Position: The Most Expensive Part of an LLM should be its Training Data SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.513176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.513176Z digest=sha256:b3676d22dc83dbfc8869f9c4beb01e4878fd2af0b618d57a4eb9bc3cba9d67ef

Observation 50021800-f7b3-4840-8bfa-76e96383ddf3 · outbound

This paper cites AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO.

Position: The Most Expensive Part of an LLM should be its Training Data AI models that cost \ 1 billion to train are underway, \ 100 billion models coming — largest current models take “only” \ 100 million to train: Anthropic CEO

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.128186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.516963Z digest=sha256:b3ef6033ab9adc67d58acb06628abd28fa19cac5abafcbb19f0b991d2955500c

Observation 34d09960-edb5-43cd-87e2-dfe3162eaac7 · outbound

This paper cites Scaling Data-Constrained Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data Scaling Data-Constrained Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.520465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.520465Z digest=sha256:b599bfc314781376968e3adc2b40aa434ecddcf115c0d26d070503d153dce003

Observation 51e8f85b-56e4-42ca-806b-f3d77227f000 · outbound

This paper cites New york times v.

Position: The Most Expensive Part of an LLM should be its Training Data New york times v

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.117493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.524159Z digest=sha256:1dd602b462eaf58a3331b225a4a92fb759b212cdfdaafcd0455ff13dd9689b0f

Observation 42b77bc3-054d-4116-bc9e-b719feca545f · outbound

This paper cites TRAK: Attributing Model Behavior at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data TRAK: Attributing Model Behavior at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.527627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.527627Z digest=sha256:b1cec388942c2f081690798bd99c04f9f5a06a3fff297b90b844cab092170535

Observation 1c16cddf-21d2-4a9b-abad-7c7caead9e09 · outbound

This paper cites an unresolved cited work.

Position: The Most Expensive Part of an LLM should be its Training Data Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:34:54.107235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.531585Z digest=sha256:8320d0ac37203b253743c9fddb2db592c414680020f86c715de1e034b6ac2d85

Observation 4e90d540-21fa-4d3f-a022-c765ce57d782 · outbound

This paper cites Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says.

Position: The Most Expensive Part of an LLM should be its Training Data Multiple ai companies bypassing web standard to scrape publisher sites, licensing firm says

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.096999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.535354Z digest=sha256:d1fccd0dbf2c75deb19974593e79019c62d3ac11a1025944cbbd3f40b4c9f30d

Observation 6bc0e423-6cf3-4234-9b54-9de913a1b9e9 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Position: The Most Expensive Part of an LLM should be its Training Data The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.538809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.538809Z digest=sha256:94d61810b8ffbd4d82be928be55a1d32c683262da6338031678fff44c86db760

Observation 117ed8c8-2b11-4ca1-9282-57fa9ad774a7 · outbound

This paper cites Estimating Training Data Influence by Tracing Gradient Descent.

Position: The Most Expensive Part of an LLM should be its Training Data Estimating Training Data Influence by Tracing Gradient Descent

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.542296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.542296Z digest=sha256:574022a713d096a26879eb58e5e0c23fefb505b77a1cd143961d55df8819e313

Observation 8f50bfd0-16cd-479b-85d2-b1164801cce4 · outbound

This paper cites Reddit and openai build partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Reddit and openai build partnership

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.086656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.546327Z digest=sha256:062295eddb0d9bba8daab094df7c012e8dedd008d682017c62476772ffb6b4a1

Observation a27c7275-92b1-4929-a0ce-de699e929c1d · outbound

This paper cites Thomson reuters' adjusted eps beats expectations, ai boosts results.

Position: The Most Expensive Part of an LLM should be its Training Data Thomson reuters' adjusted eps beats expectations, ai boosts results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.075904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.549780Z digest=sha256:af55abc064033044bab8382418cb013d0a078bfebc67eb371f6d71b695230046

Observation 22b3cdde-bdef-4cd2-b62d-524a157ec2fd · outbound

This paper cites Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data.

Position: The Most Expensive Part of an LLM should be its Training Data Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.064402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.553071Z digest=sha256:c51375be51f1a8dc561c03347498e5f103d889251f4f2e466d10a9d057d0d513

Observation 56e88207-67cb-4cda-b630-7c7928ad4f1e · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Position: The Most Expensive Part of an LLM should be its Training Data Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.556356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.556356Z digest=sha256:ce86df829f2a2b6d1b68ecd2135c6ba5e15731648b6bab99d2988389273d2201

Observation 5c5e27f6-786f-482f-8476-120d83c243b4 · outbound

This paper cites The atlantic announces product and content partnership with openai.

Position: The Most Expensive Part of an LLM should be its Training Data The atlantic announces product and content partnership with openai

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.052727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.560435Z digest=sha256:faceff4e2d3242f9f7ca3d581c4e545b591dbb6aeb07c68f42c5034f77a2c112

Observation 02da8847-5bea-46ec-90e4-ab9716fc5df4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Position: The Most Expensive Part of an LLM should be its Training Data LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.563650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.563650Z digest=sha256:43df59ad028668affdd950cfc44afec7e90b8a9153e22902a104c09f2b028222

Observation dde8342c-841c-44e7-9816-8d01bed89acc · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Position: The Most Expensive Part of an LLM should be its Training Data Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.567083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.567083Z digest=sha256:a6f1c1f0e6165a7d3ae732d0b62355ee246d287f59179dcec180cb153a2aabff

Observation cbf0bcf5-7040-43f4-ae36-73e83a968d80 · outbound

This paper cites Will we run out of data? Limits of LLM scaling based on human-generated data.

Position: The Most Expensive Part of an LLM should be its Training Data Will we run out of data? Limits of LLM scaling based on human-generated data

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.570304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.570304Z digest=sha256:b688f7c9199a6babaf81879aa85fdb4633fcd54aac649a0c377972e08987d490

Observation d74ac334-eff0-41e9-a858-6d7d0eb6f15b · outbound

This paper cites Vox media and openai form strategic content and product partnership.

Position: The Most Expensive Part of an LLM should be its Training Data Vox media and openai form strategic content and product partnership

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.042052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.573730Z digest=sha256:4de61c33d5d6138ddba948fdba65b5984dec02d440537793771aa94fd923f48f

Observation 3b58b3df-6705-40f0-bd7c-321f26fd55e5 · outbound

This paper cites Html standard.

Position: The Most Expensive Part of an LLM should be its Training Data Html standard

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.030794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.576828Z digest=sha256:8d28b7bd43cf945b75d9b73a2538e07b0e181f93546e7c7fe082d2b3c546a7a1

Observation 78226ba2-210d-4061-9217-5374472983b3 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Position: The Most Expensive Part of an LLM should be its Training Data Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.580062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.580062Z digest=sha256:4e8257e19fad66fb4d8f25a8e3a4dfddeb329426f720e6cce1f144e170fd99cf

Observation d9761ef1-9d90-41fb-b3ac-18d1a352da19 · outbound

This paper cites Breaking down emerging segments in the ai content licensing landscape.

Position: The Most Expensive Part of an LLM should be its Training Data Breaking down emerging segments in the ai content licensing landscape

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.019864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.583589Z digest=sha256:80bda6c0ba01f4d11f6ee39ec20ac039cd1f0194daf5e51b04635a5d4865a940

Observation e321c3e6-f2cd-44a7-b3d7-749b33eb06b1 · outbound

This paper cites Articulating value from data.

Position: The Most Expensive Part of an LLM should be its Training Data Articulating value from data

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:34:54.007464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T12:34:53.586790Z digest=sha256:871199af915a0c11e20787c4bedf0a6067301dd4639b255cf8651e58a13ca196

Observation d301cb7b-75c5-43ff-bab2-da3c74223efc · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Position: The Most Expensive Part of an LLM should be its Training Data WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:34:53.589828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:34:53.589828Z digest=sha256:458903437671e09af904f619f9500fd58445c9e06ccc4c4c692e66374f4a4dd5

Pith citing papers

Observation 179fac02-40ff-4f0c-a983-85f16f6422af · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Position: The Most Expensive Part of an LLM should be its Training Data

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.273599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.273599Z digest=sha256:3385e2e136d63cf1c9057b1f28b4c7c72a6a7fce4829563d9d944d4b4bd96351

Observation 7895a035-1447-469b-a995-8b83e5474abe · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T14:25:55.503750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T14:21:53.325788Z digest=sha256:3ed933cf7f6ffa78937f558d61d2cd34652552cc5ac52697a7f7d6507e82db2e

Observation de189bdf-7b49-496d-8753-7ab89b8483e5 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-15T12:17:47.123519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:17:47.123519Z digest=sha256:0fdbefbb74503744a7cbc69f33d786c1d5e30ee9d6d15e98648c21bfbd85399f

Observation f2438a06-593e-4df0-9cc9-bd8ff63152eb · inbound

On the Fragility of Data Attribution When Learning Is Distributed cites this paper.

On the Fragility of Data Attribution When Learning Is Distributed Position: The Most Expensive Part of an LLM should be its Training Data

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:17:39.452630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T15:16:32.125215Z digest=sha256:b7fb73e612b3a3a3eab9731016d623fdd4cf80867f1e60a612bb40902640bbb6

Observation fc6c92ef-bb74-4ea7-8367-1d7781beffa3 · inbound

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation cites this paper.

FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation Position: The Most Expensive Part of an LLM should be its Training Data

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:46.611909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T07:52:24.053501Z digest=sha256:312aedfef5cd174fbc097e74e618013a1e41f3e54ce5c2b173c55a49477da066