{"id":"e3dc11ab-dd55-4636-b42d-5a88331eb056","arxiv_id":"2504.12427","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"Even at conservative wages, recreating LLM training data from scratch would cost 10 to 1000 times more than the compute and energy used to train the models.","lead":"This position paper estimates what it would cost to pay people to write the text inside 64 large language models' training datasets, and finds the bill would be 10 to 1000 times larger than the cost of training the models. The authors argue this unpaid human labor should be the biggest expense in LLM production, and outline research directions for fair compensation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 10-1000x dataset-to-training cost ratio is asserted without a sensitivity analysis; plausible lower-bound wage/speed and token-to-word conversion choices could push some models below 10x, so the headline claim rests on unspecified parameters.","rationale":"The paper is a transparent position piece: the assumptions behind the cost estimate are explicitly stated, the authors deliberately bias toward underestimation, and they acknowledge limitations. The normative conclusion, that data compensation should be the largest cost, is an ethical stance that does not require the exact multiplier to be 10-1000x. However, the empirical support for the position is the claimed '10-1000x' range, and that range depends on a few unexamined quantitative choices: the definition of 'word units' (token-to-word conversion), the point values for wage and typing speed, and the use of point rather than upper-bound training costs. The reader correctly identifies the valuation method as the weakest assumption, but a more specific and testable soft spot is the lack of sensitivity analysis around these parameters. The paper's own intended conservatism makes the magnitude plausible, and the direction of the gap would likely survive, but the headline multiplier is not yet demonstrated to be robust across reasonable parameter choices. Since the missing per-model table and sensitivity analysis are addressable without changing the core argument, the conditional verdict remains appropriate.","tokens_in":12390,"tokens_out":7631,"duration_ms":79054,"concrete_test":"Recompute all 64 dataset-to-training cost ratios using the per-model table that should be released as supplementary material, over a grid: token-to-word conversion in {0.7, 1.0}; writing speed in {30, 60} words per minute; wage in {$1.00, $3.85, $7.25} per hour; and training cost using Cottier et al.'s upper-bound estimates (including engineering labor) in addition to the point estimates. Report the minimum ratio across all models for each combination. If the minimum ratio remains above 10 for every combination, the headline claim is robust. If any combination yields a ratio below 10, the claim must be qualified or the parameter choices defended.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 2.2, the paper sets dataset cost per word as wage divided by typing speed, using $3.85/hour and 30 words per minute, explicitly to underestimate data cost. Section 3.1 then claims that dataset costs are 10-1000x training costs for all 64 models. Two unspecified choices control this ratio. First, training dataset sizes are typically reported in tokens, but Section 2.2 refers to 'word units' without defining a token-to-word conversion; using tokens as words inflates costs by roughly 1.3-1.5x for English. Second, the wage and speed are point estimates, not a conservative lower bound; a $1/hour wage and 60 words per minute speed would cut estimated data costs by a factor of about 7.7 relative to the paper's assumptions. The underlying per-model data table is not included, so the minimum ratio across the 64 models cannot be checked against these alternatives. If the minimum ratio falls below 10 under equally defensible parameters, the headline '10-1000x' claim, the 'even under highly conservative estimates' framing, and the affordability analysis in Section 3.2 lose their stated force. The direction of the gap is likely to persist, but the specific quantitative claim is load-bearing for the position and is not yet robustly established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the monetary value of the human labor required to produce LLM training datasets dwarfs the cost of training the models themselves. It estimates dataset production cost as the number of word units in each training set multiplied by an assumed hourly wage divided by an assumed writing speed, using $3.85/hour and 30 words per minute as deliberately conservative point estimates. Applying this valuation to 64 LLMs released between 2016 and 2024, the authors report that dataset costs are 10–1000x larger than training costs, that several recent datasets would be valued above $10 billion, and that compensating data creators at even these low rates would be unaffordable for most of the ten companies examined. The paper concludes with research directions on data collection, data efficiency, and compensation structures.","tokens_in":12618,"tokens_out":2816,"duration_ms":30390,"significance":"If the quantitative claim is robust, the paper makes an important contribution to the policy and economics of AI training data, connecting labor valuation, copyright debates, and affordability in a way that few prior works do. Its strengths include a transparent closed-form calculation that is not fitted to a target outcome, an explicit intent to bias assumptions downward, and a fair presentation of opposing views in Section 5. The paper is also careful to describe its valuation as a deliberate lower bound rather than a market-price estimate. The main risk is not circularity but parameter sensitivity: the headline 10–1000x range depends on two point estimates and an undefined token-to-word conversion, and the paper does not currently provide enough data or sensitivity analysis for the quantitative claim to be independently verified. The directional claim that dataset labor value greatly exceeds training cost is likely to persist under many alternative assumptions, but the specific magnitudes and the phrase “even under highly conservative estimates” are not yet fully established.","major_comments":[{"comment":"The headline „10–1000x“ claim rests on two point estimates, $3.85/hour and 30 words per minute, with no sensitivity analysis. The calculation is simply dataset words × wage ÷ speed, so a wage of $1/hour and a speed of 60 wpm would reduce estimated dataset costs by a factor of about 7.7 relative to the paper’s numbers. Because the paper asserts that the result holds even under conservative assumptions, the authors should report the ratio range across a plausible grid of wage and speed values, and should state the minimum ratio across the 64 models under each parameter choice.","section":"Section 2.2 and Section 3.1"},{"comment":"The conversion from tokens to word units is never defined. Training dataset sizes are typically reported in tokens, while Section 2.2 refers to “word units,” and treating tokens as words can inflate English text cost by roughly 1.3–1.5x. The paper should state explicitly whether dataset sizes are token counts, and if so provide the token-to-word conversion used and justify it, since this choice directly scales every dataset cost estimate.","section":"Section 2.2 and Section 3.1"},{"comment":"The per-model data underlying Figure 1 and the 10–1000x range are not included. Without a table of each model’s training cost, dataset size, and computed ratio, the claim that “training data costs are at least an order of magnitude larger than training costs” for all 64 models cannot be checked, and the minimum ratio cannot be evaluated under alternative parameter choices. An appendix with these values would make the central quantitative claim reproducible.","section":"Section 3.1"},{"comment":"The paper equates the value of a training dataset with its replacement cost, and then uses that value as an estimate of compensation owed to training data producers. This conflates the cost of re-creating text from scratch with the amount owed to the original authors of text that already exists for other purposes. The authors should either reframe the analysis as a hypothetical replacement-cost valuation, clearly separated from a claim about legal or moral compensation owed, or add a discussion of why replacement cost is the right basis for compensation. The direction of the gap would likely persist under this reframing, but the direct policy meaning changes.","section":"Section 2.2"}],"minor_comments":[{"comment":"There are several typographical errors, including “heteregeneous” (Section 2.2), “the the optimal balance” (Section 4.2), inconsistent spelling of “Britannica” as “Brittanica” (Table 1 and Section 3.1), and “V o” in the Vox Media reference.","section":"Throughout"},{"comment":"The figure lacks clear axis labels and a visible scale; the annotated ratios 300x and 6000x are not evident from the plotted points. The y-axis should be labeled with units and the annotations should be tied to specific points.","section":"Figure 1"},{"comment":"The ten organizations in Figure 3 are described in the text, but the figure does not state the revenue year or the source for each company’s revenue; adding these details would make the affordability comparison easier to audit.","section":"Section 3.2"},{"comment":"The wage and speed assumptions are described qualitatively as “conservative,” but no evidence is given that 30 wpm is a lower bound for coherent text production rather than a typical value; a reference for adult typing speeds and for the $1–$50 annotation wage range would help.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position paper, and its core argument is timely and likely correct in direction. My main concern is that the quantitative claim is presented as robust when it depends on unspecified parameter choices and missing per-model data. These issues are entirely fixable within the manuscript’s scope by adding a sensitivity analysis, defining the token-to-word conversion, and including the underlying table. I do not see a basis for rejection, and I do not think the paper suffers from circular reasoning or fitted assumptions. One small editorial note: since the paper is under review and preprint-only, the authors should be aware that the version of record may need updated citations for the ongoing legal cases and revenue figures."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I'd say this is a well-meaning and genuinely useful paper, not a rigorous empirical study. What's new is the systematic application of a simple replacement-cost calculation to 64 models released from 2016 to 2024, and the explicit comparison against training costs estimated via Cottier et al. That's a real service: it puts a concrete number on something people gesture at. The authors are also admirably transparent about their assumptions, and they deliberately err on the side of underestimating data costs. The direction of the gap is almost certainly correct, and the policy relevance is obvious.\n\nThe soft spots are real but not fatal. The biggest issue is the absence of any sensitivity analysis. The cost estimate depends directly on wage ($3.85/hr) and typing speed (30 wpm), and the paper gives no indication of how the 10-1000x range responds to plausible variations. The stress-test note is right: using a $1/hr wage and 60 wpm would cut estimated data costs by a factor of about 7.7. Some models near the low end of the range could easily drop below 10x under those settings. Without the per-model data table, the reader can't check. The authors also say \"word units\" without clarifying how they converted token counts to words; if they used tokens as words, they've inflated costs by 30-50% for English. These choices don't threaten the qualitative conclusion, but they do mean the specific \"even under highly conservative estimates\" framing is overstated.\n\nThe deeper conceptual issue is that replacement cost is not the same as compensation owed. Most training text was written before anyone knew it would be used for AI, so the cost to recreate it from scratch is a fair upper bound on what creators might claim, not a lower bound. The authors acknowledge this implicitly by calling it a \"cost-based approach,\" but they don't grapple with the gap between replacement cost and the counterfactual world where authors write for AI training. That distinction matters for the legal and policy discussion.\n\nThe paper deserves a serious referee. It's a position piece, so it should be judged as one: the quantitative claim needs to be robust, and right now it isn't fully. A revision with a sensitivity analysis and the underlying data table would make it much stronger.\n\nFor me: I'd bring it to a reading group if the topic came up, and I'd cite it as a reference point for the economic framing, but I'd be careful to note the caveats. Send it to review, but expect the reviewers to ask for the missing robustness checks.","headline":"A transparent, honestly-scoped position paper whose headline 10-1000x claim needs a sensitivity analysis before it can carry the weight it is asked to carry.","tokens_in":13141,"tokens_out":1326,"would_cite":true,"duration_ms":16265,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that the human labor behind LLM training data is worth 10 to 1,000 times more than the compute, energy, and engineering used to train the models, so compensating data creators should be the largest cost of building an LLM.","keywords":["large language models","training data cost","data labor","replacement cost valuation","data compensation","copyright and fair use","data licensing","AI economics"],"falsifier":"For one named frontier model, tally every actual or legally required payment to training-data creators, including licensing fees, opt-in program payouts, and court-ordered settlements, and compare that total with the publicly reported training-run cost; the paper's claim that data compensation should be the dominant cost fails if the realized total is smaller than the training cost.","tokens_in":12161,"feed_emoji":"💰","tokens_out":8630,"duration_ms":86643,"temperature":0.7,"pith_summary":"Large language models are usually priced by the cost of compute, energy, and engineering, but this paper argues that the largest hidden cost is the human effort that produced the words they learn from. To measure that effort, the authors estimate what it would cost to pay writers to recreate the training corpora of 64 LLMs released between 2016 and 2024, using a deliberately low typing speed of 30 words per minute and a $3.85 per hour median global minimum wage. Across every model, the replacement labor cost of the training data is 10 to 1,000 times larger than the estimated training-run cost, with ratios such as about 300 times for GPT-4 and 6,000 times for DeepSeek-V3. The paper concludes that fair compensation for data producers should be the dominant expense in LLM production, and that current web-scraping practices leave LLM providers carrying an enormous financial and ethical liability.","feed_headline":"Data labor costs 10-1,000x the training run","feed_subtitle":"Across 64 models, even minimum-wage pay for writing training text would swamp compute, energy, and engineering costs.","key_machinery":"The argument runs on a deliberately conservative word-to-dollar conversion: a training dataset is valued by dividing its word count by 30 words per minute and multiplying by $3.85 per hour, the median statutory minimum wage across 167 countries. This replacement-cost formula converts an unmeasured stock of writing labor into a dollar amount that can be set directly against hardware, energy, and engineering costs from existing training-cost estimates. The machinery intentionally ignores text quality, expertise, research, and editing, which means every ratio it produces is meant to be a lower bound on the true gap.","core_discovery":"The central discovery is a quantitative mismatch between what it costs to run an LLM training job and what it would cost to produce the text that makes the model useful. Taking a replacement-cost view of data, the authors compute dataset value as the total number of words times the time needed to type them at 30 words per minute at the median global minimum wage, and compare that figure with prior estimates of training cost for the same models. The comparison shows data labor exceeding training cost for all 64 models, by two to three orders of magnitude in the most extreme recent cases. From this the authors draw the positional claim: because the human labor in training data is an order-of-magnitude larger than every other expense combined, compensation to data creators should be the most expensive part of producing an LLM, even though it is almost never paid.","pith_inferences":["The same word-to-dollar formula could be standardized as a labor-intensity metric on model cards, letting researchers, regulators, and consumers compare the hidden human cost of different models alongside FLOPs and parameter counts.","If training data were priced in, the economically rational scaling strategy might favor far smaller, higher-value datasets or synthetic-data-heavy pipelines, meaning current web-scale scraping may be inefficient as well as uncompensated.","The paper's numbers provide a lower-bound anchor for legal settlements and licensing negotiations: actual fair-use rulings or copyright settlements could convert this implicit labor debt into a real balance-sheet liability for LLM providers.","Using professional writer wages or realistic research-and-editing speeds would push the reported ratios even higher, so the 10-to-1,000-fold gap is best read as a floor on the true disparity."],"forward_implications":["Fair compensation at the paper's deliberately low rates would make frontier LLM training unaffordable for most current providers, with data costs exceeding the total annual revenue of 3 of the 10 major companies examined.","The implicit labor debt backed by training data is growing, roughly doubling every eight months, so the gap between data value and data compensation will widen unless data collection or algorithmic efficiency changes.","LLM providers would need to restructure data collection around permissively licensed and public-domain text, opt-in contribution, and verified authorship metadata to make any compensation scheme practical.","Research should shift from compute-optimal to price-optimal language models, allocating a fixed financial budget between acquiring data and training rather than treating scraped data as free.","Sustainable compensation is more likely to come from royalty or revenue-sharing structures than from upfront payment, because upfront payment at these rates is infeasible for all but the wealthiest organizations."],"supporting_citations":[{"why":"Supplies the training-cost estimation method and the hardware, energy, and engineering figures that the dataset costs are compared against.","marker":"(Cottier et al., 2024)"},{"why":"Provides the cross-country statutory minimum-wage statistics used to set the $3.85 per hour base wage.","marker":"(International Labor Organization, 2025)"},{"why":"Provides the cost-based data valuation approach the authors adopt to equate a dataset's value with its replacement cost.","marker":"(World Economic Forum, 2021)"},{"why":"Supplies the $1 to $50 per hour range for data annotation work cited to justify the conservative wage choice.","marker":"(Dzieza, 2023)"},{"why":"Gives the actual cost of creating the 15th edition, used to show the paper's per-word estimate is about 400 times lower than reality.","marker":"(Encyclopedia Britannica)"},{"why":"Supplies the database of notable models and training-dataset sizes underlying the 64-model analysis and the eight-month dataset doubling figure.","marker":"(Epoch AI, 2024)"},{"why":"Supports the projection that training data is only a small fraction of all available text, so dataset growth can continue.","marker":"(Villalobos et al., 2024)"},{"why":"Frames the scaling laws that motivate growing both model and dataset sizes, underlying the projection that dataset labor costs will keep rising.","marker":"(Kaplan et al., 2020)"},{"why":"Establishes compute-optimal scaling between model and dataset size, cited as the economic driver of growing training datasets.","marker":"(Hoffmann et al., 2022)"}],"fun_headline_variants":["Data labor costs 10-1000x AI training","Paying data writers would swamp AI compute costs","The untold cost of AI: human data labor","LLM data costs 1000x training compute"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a training corpus is worth what it would cost to type every word of it from scratch at a very low wage, even though most of that text was already produced for other purposes and a replacement cost is not the same thing as the compensation actually owed to its creators.","fun_headline_variants_meta":{"raw":{"variants":["Data labor costs 10-1000x AI training","Paying data writers would swamp AI compute costs","The untold cost of AI: human data labor","LLM data costs 1000x training compute"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1399,"prompt_tokens":937,"completion_tokens":462,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":553,"tokens_out":462,"duration_ms":5034,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:30:39.789471+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For one named frontier model, tally every actual or legally required payment to training-data creators, including licensing fees, opt-in program payouts, and court-ordered settlements, and compare that total with the publicly reported training-run cost; the paper's claim that data compensation should be the dominant cost fails if the realized total is smaller than the training cost.","supporting_citations":[{"cited_title":"Statistics on wages","cited_arxiv_id":null,"evidence_quote":"Provides the cross-country statutory minimum-wage statistics used to set the $3.85 per hour base wage."},{"cited_title":"Articulating value from data","cited_arxiv_id":null,"evidence_quote":"Provides the cost-based data valuation approach the authors adopt to equate a dataset's value with its replacement cost."},{"cited_title":"Ai is a lot of work","cited_arxiv_id":null,"evidence_quote":"Supplies the $1 to $50 per hour range for data annotation work cited to justify the conservative wage choice."},{"cited_title":"Data on notable ai models, 2024","cited_arxiv_id":null,"evidence_quote":"Supplies the database of notable models and training-dataset sizes underlying the 64-model analysis and the eight-month dataset doubling figure."}],"review_version":1}