Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Auxiliary generative models are now embedded in nearly all stages of the AI development pipeline, with synthetic data increasingly replacing human-generated data and labels — and validation is mostly informal spot-checking.

desk verdict A careful, honest qualitative study of synthetic data practices; the sample self-selects toward adoption, so read the 'ubiquitous' claims as hypothesis-generating. read the letter →

arxiv 2501.18493 v2 pith:4DOD7SRO submitted 2025-01-30 cs.HC

classification cs.HC
keywords syntheticdataauxiliarymodelsgenerativeAIdevelopmentpipelineLLM-as-a-judgevalidationresponsiblequalitativeinterviews
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports on a qualitative interview study with 19 AI practitioners and 10 responsible-AI experts. Its central claim is that one generative model (an 'auxiliary model') is now routinely used to produce data that trains or evaluates another model, so synthetic data has become embedded across nearly all stages of the AI development pipeline. The authors argue this matters because evaluation and scoring of models now often depend on outputs from other models, and because validation of that synthetic data is largely informal — practitioners describe 'spot-checking' and 'eyeballing' rather than systematic methods. The paper also catalogs limitations: difficulty controlling auxiliary model outputs, poor depiction of underrepresented groups, risks from chaining the same model across stages, and ethical concerns about agency, stereotyping, and organizational pressure to prioritize scale over rigor.

What carries the argument

The central object is the 'auxiliary model': a large, off-the-shelf generative model whose outputs — training examples, test cases, simulated user interactions, or numeric scores as in LLM-as-a-judge — are treated as synthetic data for a separate 'primary model.' The argument is carried methodologically by a two-phase interview study: 29 semi-structured interviews with 19 practitioners and 10 responsible-AI experts, analyzed with reflexive thematic analysis, using three vignettes in the second phase to surface ethical trade-offs. The conceptual move that organizes the findings is to count model-generated scores and labels as synthetic data, which reveals that evaluation itself has become a site where models judge models.

What would settle it

A systematic audit of production AI pipelines across many organizations — for example, reading public technical reports, model cards, or data documentation — that found most pipelines do not use model-generated data for training or evaluation would contradict the paper's central claim of near-universal embeddedness.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that one large generative model — the 'auxiliary model' — now routinely produces data, labels, and scores used to train or evaluate another 'primary model,' and this is happening across nearly every stage of the AI pipeline. Practitioners describe synthetic data as a way to overcome data scarcity, cost, and compliance friction, and as the only feasible way to evaluate generative systems at scale. But the promised control is often not realized: outputs are brittle to prompt changes, underrepresented groups are rendered as 'caricatures,' and most validation consists of manual spot-checking because teams cannot define what 'good' synthetic data means. Responsible-AI experts add structural concerns: chaining the same or similar models across stages creates feedback loops that can amplify bias and risk model collapse, and synthetic data removes pathways for data subjects to exercise agency.

Load-bearing premise

The interviews were with 29 people, all based in the United States and mostly working on text-based systems, and the paper's broad claims about synthetic data being ubiquitous across the industry depend on that sample being representative enough to generalize.

Editorial extensions

If this is right

  • If auxiliary models are as widespread as claimed, then benchmark comparisons between systems increasingly compare one model's outputs through another model's judgment, so scores embed the scorer's biases.
  • Evaluation at scale of generative systems would be impractical without auxiliary models, so policy should focus on transparent and validated use rather than simply removing synthetic data.
  • Validation practices need to move from spot-checking to systematic methods, including rubrics, calibration experiments, and sampling that prioritizes edge cases.
  • Documentation of auxiliary model choice, prompts, generation parameters, and rejected outputs becomes necessary for accountability.
  • Chaining the same or similar models across training and evaluation risks feedback loops and model collapse, so model selection should be deliberate rather than defaulting to the newest model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the ubiquity claim holds, then an independent 'audit' market could develop around auxiliary models as data generators, benchmarking how different generators perform at producing diverse, faithful data for a fixed downstream task.
  • The reliance on visual spot-checking suggests a concrete, testable gap: one could measure how many defects in synthetic datasets survive manual inspection at typical sample sizes, which would quantify the risk the interviews describe.
  • The finding that identity-related prompting produces 'caricatures' implies that generating synthetic data about marginalized groups is not a substitute for data from those groups, supporting a design rule that synthetic data should be validated against community input before use.
  • A longitudinal extension of this study could track whether practitioners' 'full control' expectations erode as chaining becomes more common, for example by re-interviewing participants as their pipelines accumulate more synthetic data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a two-phase interview study (19 AI practitioners and 10 responsible-AI experts, all US-based and predominantly working with text data) on how synthetic data is used across the AI development pipeline. The authors define synthetic data to include labels and scores produced by auxiliary generative models, and they document use cases spanning training, fine-tuning, evaluation, red-teaming, and model scoring. The findings describe practitioners' motivations (cost, scale, controllability, competitive edge), generation practices and desiderata (diversity, naturalness), validation practices (mostly manual spot-checking), and broader ethical limitations raised by RAI experts (chaining, stereotyping, loss of agency, organizational pressure). The paper concludes with considerations for responsible use, documentation, model selection, validation, and community engagement.

Significance. If the empirical claims are accepted after appropriate recalibration, this is a valuable contribution for FAccT and the broader HCI/CSCW community. The study complements concurrent work such as Qian et al. [78] by spanning multiple organizations and explicitly covering evaluation-stage uses such as LLM-as-a-judge. The paper is unusually transparent methodologically: the appendix contains full interview protocols and vignettes, the analysis follows a named reflexive thematic analysis approach, and the authors disclose positionality and potential adverse impacts. The participant quotes are instructive, and the discussion of the AI supply chain (Section 5.1) is a useful framing for future work. The main weakness is that the breadth of the central prevalence claims—'ubiquitous,' 'embedded in nearly all stages,' 'increasingly substituted'—goes beyond what the sampling and elicitation design can support, as detailed in the major comments.

major comments (3)
  1. [Section 3 (Recruitment) and Section 1] The inference from the interview sample to the abstract's claim that auxiliary models are 'now widely used across the AI development pipeline' and 'embedded in nearly all stages' is underdetermined by the sampling design. Recruitment was conducted through advertisements on X and LinkedIn, direct professional emails, and snowball sampling for a study whose title and abstract foreground synthetic data. Practitioners who do not use synthetic data, or use it rarely, have little reason to volunteer, and snowball sampling compounds the bias by drawing from the networks of existing synthetic-data users. The Limitations subsection explicitly acknowledges the US-only and text-focused skew but does not discuss self-selection on the outcome. As written, the prevalence claims in Section 4.1 and the abstract report the experiences of a purposive sample of adopters, not an industry-wide trend. I ask the authors to rescope these claims to the interviewed sample, or to provide explicit evidence about non-adopters and discuss the likely direction and magnitude of self-selection bias.
  2. [Section 4.1.1 and Appendix A.1] The interview protocol primes the outcome before asking about pipeline practices. Appendix A.1 asks, 'Have you ever or do you currently use synthetic, simulated, model-generated, or augmented data in your work?' before asking about the participant's actual pipeline, and later asks, 'Do you ever use models to evaluate other models or systems?' In addition, Section 4.1.1 reports that several participants initially did not consider their use of auxiliary models for scoring model outputs to be synthetic data and came to identify these labels as 'synthetic evaluation data' only after the interviewer's framing. Because the paper's definition in Section 1 counts labels and scores produced by auxiliary models as synthetic data, the empirical claim that scoring is a synthetic-data use is partly definitional rather than a participant-generated finding. This is load-bearing for the claim that synthetic data is 'increasingly substituted for traditional human-generated data and labels.' Please report separately which uses were spontaneously described by participants and which were accepted after the interviewer's definition, and qualify the substitution claim accordingly.
  3. [Section 4.3.1 and Section 4.2.2] Several descriptive claims use frequency language that the qualitative analysis does not appear to support with counts or systematic comparison. For example, Section 4.3.1 says manual inspection 'emerged as the most common validation method,' and Section 4.2.2 says the most common approach to generating scores 'was to specify annotation guidelines.' In a sample of 19 practitioners with varied roles and projects, 'most common' implies a quantitative distribution that is not reported anywhere in the paper. If no counts were computed, the authors should either report the relevant frequencies or soften these statements to 'the most frequently described method in our interviews' or similar. This matters because the paper's contribution includes characterizing current practices, and the reader should be able to see how widely each practice was distributed among participants.
minor comments (5)
  1. [Abstract and Section 5.2] The abstract promises 'concrete steps towards the development of best practices,' but Section 5.2 explicitly says the authors 'focus on foregrounding key considerations' rather than offering prescriptive recommendations. Aligning these two statements would avoid overstating the practical guidance.
  2. [Section 4.1.1, P11 quote] The quote from P11 contains the hedge 'which I think could count as synthetic data insofar as it's making a judgment on the output of the [primary model].' The surrounding text should preserve this hedge and attribute the categorization to the interviewer's definitional frame, since P11's own framing is tentative.
  3. [Section 3 (Analysis)] The statement that the team 'surfaced 762 first-level codes' would be more useful if accompanied by a description of how saturation was assessed or a table of the final themes and their definitions, since the nine initial domain categories and four top-level categories are described only briefly.
  4. [Section 4.4.1] The discussion of 'chaining' auxiliary models is compelling but relies heavily on hypothetical examples (P10's GPT-4-as-scorer example) and expert concerns. Clarifying which observations are documented practices versus expert warnings would strengthen the analytical separation between findings and interpretation.
  5. [Table 1 and Section 3] The paper would benefit from reporting the distribution of participants by organization type (e.g., large technology companies, startups, academia) and by years of experience, since the paper's claims about organizational pressure and 'studying up' depend on the institutional positions of the participants.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: central claims derive from new interviews; self-citations are minor and non-load-bearing.

full rationale

The paper's claimed derivation chain is empirical: 29 semi-structured interviews are coded via reflexive thematic analysis into findings and then into recommendations. I checked each circularity pattern against this chain. (1) No self-definitional step: the paper transparently scopes 'synthetic data' to include model-generated scores and labels (Section 2.1: 'We view this as a form of synthetic data as the scores produced by the auxiliary model take the place of labels...'), and then reports practitioners' practices under that scoping; the definition is an input framing, not something the findings are used to prove. (2) No fitted-input-called-prediction: the paper makes no quantitative predictions and fits no parameters; the 'ubiquitous integration' claim is a thematic summary of interview data. (3) Self-citations in related work (e.g., Heger et al. [37], Holstein et al. [40], Madaio et al. [61,62], Sambasivan et al. [83]) are contextual and not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work. (4) No renaming of a known result as new organization: 'LLM-as-a-judge' is cited to external work and explicitly reframed as synthetic data for the paper's scope, which is transparent. The Limitations paragraph in Section 3 acknowledges the US-only, text-heavy sample; that is an external-validity limitation, not circularity. One participant's reported use of an auxiliary model's confidence to validate its own output (Section 4.3.1) is an object-level finding, not the paper's own reasoning. Thus the central claim rests on new interview evidence and is not circular; only minor non-load-bearing self-citations appear.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on the assumption that self-reported interviews reflect actual practices and that a US-based, text-focused sample supports general claims. No free parameters or invented entities are involved.

assumptions (3)
  • domain assumption Participant self-reports during interviews correspond to actual synthetic data practices.
    The findings are based entirely on what participants say; no observational or artifact-based verification is included. This is standard for interview research but is a premise that could fail if participants misremember or present idealized accounts.
  • domain assumption The sample of 29 participants from 14 US organizations is sufficiently representative to support claims of widespread integration of synthetic data.
    The paper's language of 'ubiquitous' and 'common practice' generalizes from this sample. The Limitations section acknowledges the US-only, mostly text-focused skew.
  • ad hoc to paper Labels and scores produced by auxiliary models count as synthetic data.
    Section 2.1 explicitly treats LLM-as-a-judge outputs as a form of synthetic data. Participants initially did not always classify them this way, so this framing shapes what the study counts as evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline." pith.science (2026). https://pith.science/paper/4DOD7SRO

@misc{pith2026250118493,
  author       = {Pith},
  title        = {Pith review of: Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4DOD7SRO}},
  note         = {Machine review of arXiv:2501.18493}
}
read the original abstract

Alongside the growth of generative AI, we are witnessing a surge in the use of synthetic data across all stages of the AI development pipeline. It is now common practice for researchers and practitioners to use one large generative model (which we refer to as an auxiliary model) to generate synthetic data that is used to train or evaluate another, reconfiguring AI workflows and reshaping the very nature of data. While scholars have raised concerns over the risks of synthetic data, policy guidance and best practices for its responsible use have not kept up with these rapidly evolving industry trends, in part because we lack a clear picture of current practices and challenges. Our work aims to address this gap. Through 29 interviews with AI practitioners and responsible AI experts, we examine the expanding role of synthetic data in AI development. Our findings reveal how auxiliary models are now widely used across the AI development pipeline. Practitioners describe synthetic data as crucial for addressing data scarcity and providing a competitive edge, noting that evaluation of generative AI systems at scale would be infeasible without auxiliary models. However, they face challenges controlling the outputs of auxiliary models, generating data that accurately depict underrepresented groups, and scaling data validation practices that are based primarily on manual inspection. We detail general limitations of and ethical considerations for synthetic data and conclude with a proposal of concrete steps towards the development of best practices for its responsible use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic CVs To Build and Test Fairness-Aware Hiring Tools

    cs.CY 2025-08 conditional novelty 6.0 of 10

    A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.

Reference graph

Works this paper leans on

116 extracted references · 52 canonical work pages · cited by 1 Pith paper

  1. [78]

    Crystal Qian, Michael Xieyang Liu, Emily Reif, Grady Si mon, Nada Hussein, Nathan Clement, James Wexler, Carrie J Cai, Michael Terry, and Minsuk Kahng

  2. [1]

    Nazmiye Ceren Abay, Yan Zhou, Murat Kantarcioglu, Bhava ni Thuraisingham, and Latanya Sweeney. 2019. Privacy preserving synthetic da ta release using deep learning. In Machine Learning and Knowledge Discovery in Databases: Eu- ropean Conference, ECML PKDD 2018, Dublin, Ireland, Septemb er 10–14, 2018, Proceedings, Part I 18 . Springer, 510–526

  3. [2]

    Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bube ck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905 (2024)

  4. [3]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad , Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sa m Altman, Shya- mal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)

  5. [4]

    William Agnew, A Stevie Bergman, Jennifer Chien, Mark Dí az, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R McKee. 2024. The illusion of artificial inclusion. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–12

  6. [5]

    Samuel A Assefa, Danial Dervovic, Mahmoud Mahfouz, Robe rt E Tillman, Prashant Reddy, and Manuela Veloso. 2020. Generating synth etic data in fi- nance: opportunities, challenges and pitfalls. In Proceedings of the First ACM International Conference on AI in Finance . 1–8

  7. [6]

    Pierre Azoulay, Joshua L Krieger, and Abhishek Nagaraj. 2024. Old moats for new models: Openness, control, and competition in generati ve ai. Technical Re- port. National Bureau of Economic Research

  8. [7]

    Gwangbin Bae, Martin de La Gorce, Tadas Baltrušaitis, Ch arlie Hewitt, Dong Chen, Julien Valentin, Roberto Cipolla, and Jingjing Shen. 2023. DigiFace-1M: 1 Million Digital Face Images for Face Recognition. In 2023 IEEE Winter Confer- ence on Applications of Computer Vision (W ACV). IEEE

Show all 116 references
  1. [8]

    Chelsea Barabas, Colin Doyle, JB Rubinovitz, and Karthi k Dinakar. 2020. Study- ing up: reorienting the study of algorithmic fairness aroun d issues of power. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 167–176

  2. [9]

    Emily M Bender and Batya Friedman. 2018. Data statements for natural lan- guage processing: Toward mitigating system bias and enabli ng better science. Transactions of the Association for Computational Linguis tics 6 (2018), 587–604

  3. [10]

    Virginia Braun and Victoria Clarke. 2012. Thematic analysis. American Psy- chological Association

  4. [11]

    Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic anal- ysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597

  5. [12]

    Sarah Burkhardt and Bernhard Rieder. 2024. Foundation models are platform models: Prompting and the political economy of AI. Big Data & Society 11, 2 (2024), 20539517241247839

  6. [13]

    Cheng-Han Chiang and Hung-yi Lee. 2023. Can large langu age models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937 (2023)

  7. [14]

    Kranzinger, Michael Hind, Jen- nifer Wortman Vaughan, Margaret Mitchell, Julia Stoyanovi ch, Angelina McMillan-Major, Emily McReynolds, Kathleen Esfahany, Mary L

    Kasia Chmielinski, Sarah Newman, Chris N. Kranzinger, Michael Hind, Jen- nifer Wortman Vaughan, Margaret Mitchell, Julia Stoyanovi ch, Angelina McMillan-Major, Emily McReynolds, Kathleen Esfahany, Mary L. Gray, Audrey Chang, , and Maui Hudson. 2024. The CLeAR Documentation Fra...

  8. [15]

    Anamaria Crisan, Brittany Fiore-Gartland, and Melani e Tory. 2020. Passing the data baton: A retrospective analysis on data science work an d workers. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 1860–1870

  9. [16]

    Jessamyn Dahmen and Diane Cook. 2019. SynSys: A synthet ic data generation system for healthcare applications. Sensors 19, 5 (2019), 1181

  10. [17]

    Fernando Delgado, Stephen Yang, Michael Madaio, and Qi an Yang. 2023. The participatory turn in ai design: Theoretical foundations a nd the current state of practice. In Proceedings of the 3rd ACM Conference on Equity and Access in FAccT ’25, June 23–26, 2025, Athens, Greece K...

  11. [18]

    Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 20 22. Exploring how machine learning practitioners (try to) use fairness to olkits. In Proceed- ings of the 2022 ACM Conference on Fairness, Accounta...

  12. [19]

    Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahh are Eslami, Ken- neth Holstein, and Michael Madaio. 2023. Investigating Pra ctices and Opportu- nities for Cross-functional Collaboration around AI Fairn ess in Industry Prac- tice. In Proceedings of the 2023 ACM Conferenc...

  13. [20]

    Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang , Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Has himoto. 2024. Alpacafarm: A simulation framework for methods that learn f rom human feed- back. Advances in Neural Information Processing S...

  14. [21]

    Madeleine Clare Elish and Danah Boyd. 2018. Situating m ethods in the magic of Big Data and AI. Communication monographs 85, 1 (2018), 57–80

  15. [22]

    Alexander R Fabbri, Wojciech Kryściński, Bryan McCann , Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-eva luating summa- rization evaluation. Transactions of the Association for Computational Linguis - tics 9 (2021), 391–409

  16. [23]

    Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023. Factkb: Generalizable factuality evaluation using langua ge models enhanced with factual knowledge. arXiv preprint arXiv:2305.08281 (2023)

  17. [24]

    Andrew Fitzgerald. 2024. Why Synthetic Data Can Never B e Ethical: A Les- son from Media Ethics. Surveillance & Society 22, 4 (Dec. 2024), 477–482. https://doi.org/10.24908/ss.v22i4.18324

  18. [25]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Je nnifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (December 2021), 86–92

  19. [26]

    Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 202 3. ChatGPT outper- forms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences 120, 30 (2023), e2305016120

  20. [27]

    Raw Data

    Lisa Gitelman. 2013. “Raw Data” Is an Oxymoron

  21. [28]

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63, 11 (2020), 139–144

  22. [29]

    Maybank, and Dach eng Tao

    Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dach eng Tao. 2021. Knowledge Distillation: A Survey. International Journal of Computer Vision 129 (2021), 1789–1819

  23. [30]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab hinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur , Alan Schel- ten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)

  24. [31]

    Igor Grossmann, Matthew Feinberg, Dawn C Parker, Nicho las A Christakis, Philip E Tetlock, and William A Cunningham. 2023. AI and the t ransformation of social science research. Science 380, 6650 (2023), 1108–1109

  25. [32]

    Tom Gunter, Zirui Wang, Chong Wang, Ruoming Pang, Andy N arayanan, Aonan Zhang, Bowen Zhang, Chen Chen, Chung-Cheng Chiu, Davi d Qiu, et al. 2024. Apple intelligence foundation language models . arXiv preprint arXiv:2407.21075 (2024)

  26. [33]

    Alon Halevy, Peter Norvig, and Fernando Pereira. 2009. The unreasonable ef- fectiveness of data. IEEE intelligent systems 24, 2 (2009), 8–12

  27. [34]

    Lei Han, Tianwa Chen, Gianluca Demartini, Marta Induls ka, and Shazia Sadiq

  28. [35]

    Alex Hanna and Tina M Park. 2020. Against scale: Provoca tions and resistances to scale thinking. arXiv preprint arXiv:2010.08850 (2020)

  29. [36]

    Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maa rten Sap, Di- pankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Mach ine- Generated Dataset for Adversarial and Implicit Hate Speech Detection. http://arxiv.org/abs/2203.09509 arXiv:2203.09509 [cs]

  30. [37]

    Amy K Heger, Liz B Marquis, Mihaela Vorvoreanu, Hanna Wa llach, and Jen- nifer Wortman Vaughan. 2022. Understanding machine learni ng practitioners’ data documentation perceptions, needs, challenges, and desiderata. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2...

  31. [38]

    Paula Helm, Benjamin Lipp, and Roser Pujadas. 2024. Gen erating reality and silencing debate: Synthetic data as discursive device. Big Data & Society 11, 2 (June 2024), 20539517241249447. https://doi.org/10.1177/20539517241249447 Publisher: SAGE Publications Ltd

  32. [39]

    Hinton, Oriol Vinyals, and Jeffrey Dean

    Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 . Distilling the Knowl- edge in a Neural Network. arXiv preprint arXiv:1503.02531

  33. [40]

    Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learnin g systems: What do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–16

  34. [41]

    Qisheng Hu, Kaixin Li, Xu Zhao, Yuxi Xie, Tiedong Liu, Hu i Chen, Qizhe Xie, and Junxian He. 2023. Instructcoder: Empowering language m odels for code editing. arXiv preprint arXiv:2310.20329 (2023)

  35. [42]

    Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Lia ng, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebota r, et al. 2022. Inner monologue: Embodied reasoning through planning with language mod- els. arXiv preprint arXiv:2207.05608 (2022)

  36. [43]

    IBM. 2023. What is Synthetic Data? | IBM. https://www.ibm.com/think/topics/synthetic-data [Online; accessed 2025-01- 21]

  37. [44]

    Apple Inc. 2024. Introducing Apple’s On-Device and Ser ver Foundation Models - Apple Machine Learning Research. https://machinelearning.apple.com/research/introducing-apple-foundation-models [Online; accessed 2025-01-21]

  38. [45]

    Benjamin N Jacobsen. 2023. Machine learning and the pol itics of syn- thetic data. Big Data & Society 10, 1 (Jan. 2023), 20539517221145372. https://doi.org/10.1177/20539517221145372 Publisher: SAGE Publications Ltd

  39. [46]

    Nikita Jaipuria, Xianling Zhang, Rohan Bhasin, Mayar A rafa, Punarjay Chakravarty, Shubham Shrivastava, Sagar Manglani, and Vidya N Murali. 2020. Deflating dataset bias using synthetic data augmentation. I n Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...

  40. [47]

    Harry H Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan , Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Geb ru. 2023. AI Art and its Impact on Artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 363–374

  41. [48]

    Synthetic Data–what, why and how? arXiv preprint arXiv:2205.03257 (2022)

    James Jordon, Lukasz Szpruch, Florimond Houssiau, Mir ko Bottarelli, Giovanni Cherubin, Carsten Maple, Samuel N Cohen, and Adrian Weller.2022. Synthetic Data–what, why and how? arXiv preprint arXiv:2205.03257 (2022)

  42. [49]

    James Jordon, Jinsung Yoon, and Mihaela Van Der Schaar. 2018. PATE-GAN: Generating synthetic data with differential privacy guarantees. In International conference on learning representations

  43. [50]

    Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the Snark: An- notator Diversity in Data Practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Ma- chinery, New York, NY, USA, 1–15. https:...

  44. [51]

    Diederik P Kingma. 2013. Auto-encoding variational ba yes. arXiv preprint arXiv:1312.6114 (2013)

  45. [52]

    Laura Koesten, Elena Simperl, Tom Blount, Emilia Kacpr zak, and Jeni Tennison

  46. [53]

    Adam Kortylewski, Bernhard Egger, Andreas Schneider, Thomas Gerig, An- dreas Morel-Forster, and Thomas Vetter. 2019. Analyzing an d Reducing the Damage of Dataset Bias to Face Recognition With Synthetic Da ta. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  47. [54]

    Francis Lee, Saghi Hajisharif, and Ericka Johnson. 202 5. The ontological politics of synthetic data: Normalities, outliers, and intersectio nal hallucinations. Big Data & Society 12, 2 (2025), 20539517251318289

  48. [55]

    Peter Lee. 2024. Synthetic Data and the Future of AI. (20 24)

  49. [56]

    Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Haus man, Brian Ichter, Pete Florence, and Andy Zeng. 2023. Code as policies: Langua ge model pro- grams for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 9493–9500

  50. [57]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual in- struction tuning. Advances in neural information processing systems 36 (2024)

  51. [58]

    Ruibo Liu, Jerry Wei, Fangyu Liu, Chenglei Si, Yanzhe Zh ang, Jinmeng Rao, Steven Zheng, Daiyi Peng, Diyi Yang, Denny Zhou, and Andrew M . Dai. 2024. Best Practices and Lessons Learned on Synthetic Data for Lan guage Models. https://doi.org/10.48550/arXiv.2404.07503 arXiv:2404...

  52. [59]

    Dieuwertje Luitse and Wiebke Denkena. 2021. The great t ransformer: Examin- ing the role of large language models in the political econom y of AI. Big Data & Society 8, 2 (2021), 20539517211047734

  53. [60]

    Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lo u, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 20 23. Wizard- math: Empowering mathematical reasoning for large languag e models via re- inforced evol-instruct. arXiv preprint arXiv:2308.09583 (2023)

  54. [61]

    Michael Madaio, Lisa Egede, Hariharan Subramonyam, Je nnifer Wort- man Vaughan, and Hanna Wallach. 2022. Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Sup port. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022),...

  55. [62]

    Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan , and Hanna Wal- lach. 2020. Co-designing checklists to understand organiz ational challenges and opportunities around fairness in ai. In Proceedings of the 2020 CHI Confer- ence on Human Factors in Computing Systems . 1–14

  56. [63]

    Daniella Meeker, Crystal Kallem, Yan Heras, Stephanie Garcia, and Casey Thompson. 2022. Case report: evaluation of an open-source s ynthetic data platform for simulation studies. JAMIA open 5, 3 (2022), ooac067. Examining the Expanding Role of Synthetic Data Throughout th e AI...

  57. [64]

    Thom p- son, and Nico Grant

    Cade Metz, Cecilia Kang, Sheera Frenkel, Stuart A. Thom p- son, and Nico Grant. 2024. How Tech Giants Cut Cor- ners to Harvest Data for A.I. - The New York Times. https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html [Online; access...

  58. [65]

    Milagros Miceli, Martin Schuessler, and Tianling Yang . 2020. Between subjec- tivity and imposition: Power dynamics in data annotation fo r computer vision. Proceedings of the ACM on Human-Computer Interaction4, CSCW2 (2020), 1–25

  59. [66]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasser- man, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru

  60. [67]

    Michael Muller, Ingrid Lange, Dakuo Wang, David Piorko wski, Jason Tsay, Q Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How data science workers work with data: Discovery, capture, curation, desi gn, creation. In Pro- ceedings of the 2019 CHI conference on human factors ...

  61. [68]

    Laura Nader. 1972. Up the anthropologist: Perspective s gained from studying up. (1972)

  62. [69]

    Sergey I Nikolenko. 2021. Synthetic data for deep learning . Vol. 174. Springer

  63. [70]

    Code of Practice Working Groups. 2024. Second Draft of the General-Purpose AI Code of Practice. https://digital-strategy.ec.europa.eu/en/library/se cond-draft-general-purpose-ai-code-practice-publish ed-written-independent-experts [Online; accessed 2025-01-20]

  64. [71]

    Office of the Assistant Secretary for Planning and Evalua - tion (ASPE). 2022. A Synthetic Health Data Generation En- gine to Accelerate Patient-Centered Outcomes Research | AS PE. https://aspe.hhs.gov/synthetic-health-data-generati on-engine-accelerate-patient-centered-outcomes...

  65. [72]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carrol l Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Sl ama, Alex Ray, et al. 2022. Training language models to follow instruction s with human feed- back. Advances in neural information processing system...

  66. [73]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jin g Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Li nguistics. 311–318

  67. [74]

    Samir Passi and Steven Jackson. 2017. Data Vision: Lear ning to See Through Algorithmic Abstraction. In Proceedings of the 2017 ACM Confer- ence on Computer Supported Cooperative Work and Social Comp uting (CSCW ’17). Association for Computing Machinery, New York, NY, USA, 24 ...

  68. [75]

    Samir Passi and Steven J Jackson. 2018. Trust in data sci ence: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–28

  69. [76]

    Personal Data Protection Commission (PDPC). 2024. Pri vacy Enhanc- ing Technology (PET): Proposed Guide On Synthetic Data Gene ration. https://www.pdpc.gov.sg/-/media/files/pdpc/pdf-files/other-guides/proposed-guide-on-synthetic-data-gener ation.pdf [Online; accessed 2025-01-20]

  70. [77]

    Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Ro man Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving . 2022. Red Team- ing Language Models with Language Models. arXiv:2202.03286 [cs] (Feb. 2022). http://arxiv.org/abs/2202.03286 arXiv: 2202.03286

  71. [79]

    Bogdana Rakova, Jingying Yang, Henriette Cramer, and R umman Chowdhury

  72. [80]

    Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sag ot, Djamé Seddah, and Jacopo Staiano. 2020. Synthetic data augmentation for zero -shot cross-lingual question answering. arXiv preprint arXiv:2010.12643 (2020)

  73. [81]

    Kevin Roose. 2024. Data for A.I. Training Is Dis- appearing Fast, Study Shows - The New York Times. https://www.nytimes.com/2024/07/19/technology/ai-data-restrictions.html [Online; accessed 2025-01-21]

  74. [82]

    Roy A Ruddle, James Cheshire, and Sara Johansson Fernst ad. 2023. Tasks and visualizations used for data profiling: A survey and interview study. IEEE Trans- actions on Visualization and Computer Graphics (2023)

  75. [83]

    Everyone wants to do the mo del work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, D iana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the mo del work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Syst...

  76. [84]

    Siamak Shakeri, Noah Constant, Mihir Sanjay Kale, and L inting Xue. 2020. To- wards zero-shot multilingual synthetic question and answe r generation for cross-lingual reading comprehension. arXiv preprint arXiv:2010.12008 (2020)

  77. [85]

    Hoo-Chang Shin, Neil A Tenenholtz, Jameson K Rogers, Ch ristopher G Schwarz, Matthew L Senjem, Jeffrey L Gunter, Katherine P Andr iole, and Mark Michalski. 2018. Medical image synthesis for data augm entation and anonymization using generative adversarial networks. In Simulatio...

  78. [86]

    Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Ander- son, and Yarin Gal. 2024. AI models collapse when trained on r ecursively gen- erated data. Nature 631 (2024), 755–759

  79. [87]

    Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of conte nt moderators: the emotional labor of commercial moderation and avenues fo r improving sup- port. In Proceedings of the 2021 CHI conference on h...

  80. [88]

    Daniel Susser, Daniel S Schiff, Sara Gerke, Laura Y Cabre ra, I Glenn Co- hen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet, J asmine McNealy, Michelle N Meyer, et al. 2024. Synthetic Health Data: Real Et hical Promise and Peril. Hastings Center Report 54, 5 (2024), 8–13

  81. [89]

    Schiff, Sara Gerke, Laura Y

    Daniel Susser, Daniel S. Schiff, Sara Gerke, Laura Y. Cab rera, I. Glenn Cohen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet , Jasmine Mc- Nealy, Michelle N. Meyer, W. Nicholson Price II, and Jennife r K. Wagner

  82. [90]

    Daniel Susser and Jeremy Seeman. [n. d.]. Dialogue Crit ical Provocations for Synthetic Data. ([n. d.])

  83. [91]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Bap tiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Milli- can, et al. 2023. Gemini: a family of highly capable multimod al models. arXiv preprint arXiv:2312.11805 (2023)

  84. [92]

    Risto Uuk, Annemieke Brouwer, Tim Schreier, Noemi Drek sler, Valeria Pulig- nano, and Rishi Bommasani. 2024. Effective Mitigations for S ystemic Risks from General-Purpose AI. arXiv preprint arXiv:2412.02145 (2024)

  85. [93]

    Janet Vertesi and Paul Dourish. 2011. The value of data: considering the context of production in data economies. In Proceedings of the ACM 2011 conference on Computer supported cooperative work . 533–542

  86. [94]

    Dickerso n

    Angelina Wang, Jamie Morgenstern, and John P. Dickerso n. 2024. Large lan- guage models should not replace human participants because they can mis- portray and flatten identity groups. arXiv preprint arXiv:2 402.01908

  87. [95]

    Hastings Center Report 54, 5 (2024), 8–13

    Synthetic Health Data: Real Ethical Promise and Peril . Hastings Center Report 54, 5 (2024), 8–13. https://doi.org/10.1002/hast.4911 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/hast.4911

  88. [96]

    Cedric Deslandes Whitney and Justin Norman. 2024. Real Risks of Fake Data: Synthetic Data, Diversity-Washing and Consent Circumvent ion. In The 2024 ACM Conference on Fairness, Accountability, and Transpare ncy. ACM, Rio de Janeiro Brazil, 1733–1744. https://doi.org/10.1145/36...

  89. [97]

    AI supply chain

    David Gray Widder and Dawn Nafus. 2023. Dislocated acco untabilities in the “AI supply chain”: Modularity and developers’ notions of re sponsibility. Big Data & Society 10, 1 (2023), 20539517231177620

  90. [98]

    David Gray Widder, Sarah West, and Meredith Whittaker. 2023. Open (for busi- ness): Big tech, concentrated power, and the political economy of open AI. Con- centrated Power, and the Political Economy of Open AI (August 1 7, 2023) (2023)

  91. [99]

    David Gray Widder and Richmond Wong. 2023. Thinking ups tream: Ethics and policy opportunities in AI supply chains. arXiv preprint arXiv:2303.07529 (2023)

  92. [100]

    David Gray Widder, Derrick Zhen, Laura Dabbish, and Ja mes Herbsleb. 2023. It’s about power: What ethical concerns do software enginee rs have, and what do they (feel they can) do about them?. In Proceedings of the 2023 ACM Confer- ence on Fairness, Accountability, and Transpa...

  93. [101]

    Ding Wang, Shantanu Prabhat, and Nithya Sambasivan. 20 22. Whose AI Dream? In search of the aspiration in data annotation.. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–16

  94. [102]

    Matteo Wong. 2024. The GPT Era Is Already Ending - The At lantic. https://www.theatlantic.com/technology/archive/2024/12/openai-o1-reasoning-models/680906/?utm_source=c hatgpt.com [Online; accessed 2025-01-22]

  95. [103]

    Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019 . Goals, process, and challenges of exploratory data analysis: An interview stud y. arXiv preprint arXiv:1911.00568 (2019)

  96. [104]

    Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot. 2024. Fairness feedback loops: training on synthetic data amplifies bias. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2113–2147

  97. [105]

    Liyang Xie, Kaixiang Lin, Shu Wang, Fei Wang, and Jiayu Zhou. 2018. Differen- tially private generative adversarial network. arXiv preprint arXiv:1802.06739 (2018)

  98. [106]

    Zhenpei Yang, Yuning Chai, Dragomir Anguelov, Yin Zho u, Pei Sun, Dumitru Erhan, Sean Rafferty, and Henrik Kretzschmar. 2020. Surfelg an: Synthesizing realistic sensor data for autonomous driving. In Proceedings of the IEEE/CVF FAccT ’25, June 23–26, 2025, Athens, Greece Kapani...

  99. [107]

    Tanja Wiehn. 2024. Synthetic Data: From Data Scarcity to Data Pollution. Surveillance & Society 22, 4 (Dec. 2024), 472–476. https://doi.org/10.24908/ss.v22i4.18327

  100. [108]

    Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengy ing Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 202 3. Metamath: Bootstrap your own mathematical questions for large langua ge models. arXiv preprint arXiv:2309.12284 (2023)

  101. [109]

    Amy X Zhang, Michael Muller, and Dakuo Wang. 2020. How d o data science workers collaborate? Roles, workflows, and tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–23

  102. [110]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhu ang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, etal. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Informa- tion Processing Systems 36 (2023), 46595–46623. A ...

  103. [113]

    Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fer- nanda Viégas. 2023. Investigating How Practitioners Use Human-AI Guidelines: A Case Study on the People+ AI Guidebook. In Proceedings of the 2023 CHI Con- ference on Human Factors in Computing Systems . 1–13

  104. [2019]

    In Proceedings of the Conference on Fair- ness, Accountability, and Transparency

    Model cards for model reporting. In Proceedings of the Conference on Fair- ness, Accountability, and Transparency. 220–229

  105. [2020]

    International Journal of Human-Computer Studies 135 (March 2020), 102367

    Everything you always wanted to know about a dataset: S tudies in data summarisation. International Journal of Human-Computer Studies 135 (March 2020), 102367. https://doi.org/10.1016/j.ijhcs.2019.10.004

  106. [2021]

    Proceedings of the ACM on Human- Computer Interaction 5, CSCW1 (2021), 1–23

    Where responsible AI meets reality: Practitioner per spectives on en- ablers for shifting organizational practices. Proceedings of the ACM on Human- Computer Interaction 5, CSCW1 (2021), 1–23

  107. [2023]

    ACM Transactions on Information Systems 41, 3 (2023), 1–35

    A data-driven analysis of behaviors in data curation p rocesses. ACM Transactions on Information Systems 41, 3 (2023), 1–35

  108. [2024]

    arXiv preprint arXiv:2412.16089 (2024)

    The Evolution of LLM Adoption in Industry Data Curatio n Practices. arXiv preprint arXiv:2412.16089 (2024)

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.