REVIEW 3 major objections 5 minor 1 cited by
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Auxiliary generative models are now embedded in nearly all stages of the AI development pipeline, with synthetic data increasingly replacing human-generated data and labels — and validation is mostly informal spot-checking.
desk verdict A careful, honest qualitative study of synthetic data practices; the sample self-selects toward adoption, so read the 'ubiquitous' claims as hypothesis-generating. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the 'auxiliary model': a large, off-the-shelf generative model whose outputs — training examples, test cases, simulated user interactions, or numeric scores as in LLM-as-a-judge — are treated as synthetic data for a separate 'primary model.' The argument is carried methodologically by a two-phase interview study: 29 semi-structured interviews with 19 practitioners and 10 responsible-AI experts, analyzed with reflexive thematic analysis, using three vignettes in the second phase to surface ethical trade-offs. The conceptual move that organizes the findings is to count model-generated scores and labels as synthetic data, which reveals that evaluation itself has become a site where models judge models.
What would settle it
A systematic audit of production AI pipelines across many organizations — for example, reading public technical reports, model cards, or data documentation — that found most pipelines do not use model-generated data for training or evaluation would contradict the paper's central claim of near-universal embeddedness.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that one large generative model — the 'auxiliary model' — now routinely produces data, labels, and scores used to train or evaluate another 'primary model,' and this is happening across nearly every stage of the AI pipeline. Practitioners describe synthetic data as a way to overcome data scarcity, cost, and compliance friction, and as the only feasible way to evaluate generative systems at scale. But the promised control is often not realized: outputs are brittle to prompt changes, underrepresented groups are rendered as 'caricatures,' and most validation consists of manual spot-checking because teams cannot define what 'good' synthetic data means. Responsible-AI experts add structural concerns: chaining the same or similar models across stages creates feedback loops that can amplify bias and risk model collapse, and synthetic data removes pathways for data subjects to exercise agency.
Load-bearing premise
The interviews were with 29 people, all based in the United States and mostly working on text-based systems, and the paper's broad claims about synthetic data being ubiquitous across the industry depend on that sample being representative enough to generalize.
Editorial extensions
If this is right
- If auxiliary models are as widespread as claimed, then benchmark comparisons between systems increasingly compare one model's outputs through another model's judgment, so scores embed the scorer's biases.
- Evaluation at scale of generative systems would be impractical without auxiliary models, so policy should focus on transparent and validated use rather than simply removing synthetic data.
- Validation practices need to move from spot-checking to systematic methods, including rubrics, calibration experiments, and sampling that prioritizes edge cases.
- Documentation of auxiliary model choice, prompts, generation parameters, and rejected outputs becomes necessary for accountability.
- Chaining the same or similar models across training and evaluation risks feedback loops and model collapse, so model selection should be deliberate rather than defaulting to the newest model.
Reading between the lines
- If the ubiquity claim holds, then an independent 'audit' market could develop around auxiliary models as data generators, benchmarking how different generators perform at producing diverse, faithful data for a fixed downstream task.
- The reliance on visual spot-checking suggests a concrete, testable gap: one could measure how many defects in synthetic datasets survive manual inspection at typical sample sizes, which would quantify the risk the interviews describe.
- The finding that identity-related prompting produces 'caricatures' implies that generating synthetic data about marginalized groups is not a substitute for data from those groups, supporting a design rule that synthetic data should be validated against community input before use.
- A longitudinal extension of this study could track whether practitioners' 'full control' expectations erode as chaining becomes more common, for example by re-interviewing participants as their pipelines accumulate more synthetic data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a two-phase interview study (19 AI practitioners and 10 responsible-AI experts, all US-based and predominantly working with text data) on how synthetic data is used across the AI development pipeline. The authors define synthetic data to include labels and scores produced by auxiliary generative models, and they document use cases spanning training, fine-tuning, evaluation, red-teaming, and model scoring. The findings describe practitioners' motivations (cost, scale, controllability, competitive edge), generation practices and desiderata (diversity, naturalness), validation practices (mostly manual spot-checking), and broader ethical limitations raised by RAI experts (chaining, stereotyping, loss of agency, organizational pressure). The paper concludes with considerations for responsible use, documentation, model selection, validation, and community engagement.
Significance. If the empirical claims are accepted after appropriate recalibration, this is a valuable contribution for FAccT and the broader HCI/CSCW community. The study complements concurrent work such as Qian et al. [78] by spanning multiple organizations and explicitly covering evaluation-stage uses such as LLM-as-a-judge. The paper is unusually transparent methodologically: the appendix contains full interview protocols and vignettes, the analysis follows a named reflexive thematic analysis approach, and the authors disclose positionality and potential adverse impacts. The participant quotes are instructive, and the discussion of the AI supply chain (Section 5.1) is a useful framing for future work. The main weakness is that the breadth of the central prevalence claims—'ubiquitous,' 'embedded in nearly all stages,' 'increasingly substituted'—goes beyond what the sampling and elicitation design can support, as detailed in the major comments.
major comments (3)
- [Section 3 (Recruitment) and Section 1] The inference from the interview sample to the abstract's claim that auxiliary models are 'now widely used across the AI development pipeline' and 'embedded in nearly all stages' is underdetermined by the sampling design. Recruitment was conducted through advertisements on X and LinkedIn, direct professional emails, and snowball sampling for a study whose title and abstract foreground synthetic data. Practitioners who do not use synthetic data, or use it rarely, have little reason to volunteer, and snowball sampling compounds the bias by drawing from the networks of existing synthetic-data users. The Limitations subsection explicitly acknowledges the US-only and text-focused skew but does not discuss self-selection on the outcome. As written, the prevalence claims in Section 4.1 and the abstract report the experiences of a purposive sample of adopters, not an industry-wide trend. I ask the authors to rescope these claims to the interviewed sample, or to provide explicit evidence about non-adopters and discuss the likely direction and magnitude of self-selection bias.
- [Section 4.1.1 and Appendix A.1] The interview protocol primes the outcome before asking about pipeline practices. Appendix A.1 asks, 'Have you ever or do you currently use synthetic, simulated, model-generated, or augmented data in your work?' before asking about the participant's actual pipeline, and later asks, 'Do you ever use models to evaluate other models or systems?' In addition, Section 4.1.1 reports that several participants initially did not consider their use of auxiliary models for scoring model outputs to be synthetic data and came to identify these labels as 'synthetic evaluation data' only after the interviewer's framing. Because the paper's definition in Section 1 counts labels and scores produced by auxiliary models as synthetic data, the empirical claim that scoring is a synthetic-data use is partly definitional rather than a participant-generated finding. This is load-bearing for the claim that synthetic data is 'increasingly substituted for traditional human-generated data and labels.' Please report separately which uses were spontaneously described by participants and which were accepted after the interviewer's definition, and qualify the substitution claim accordingly.
- [Section 4.3.1 and Section 4.2.2] Several descriptive claims use frequency language that the qualitative analysis does not appear to support with counts or systematic comparison. For example, Section 4.3.1 says manual inspection 'emerged as the most common validation method,' and Section 4.2.2 says the most common approach to generating scores 'was to specify annotation guidelines.' In a sample of 19 practitioners with varied roles and projects, 'most common' implies a quantitative distribution that is not reported anywhere in the paper. If no counts were computed, the authors should either report the relevant frequencies or soften these statements to 'the most frequently described method in our interviews' or similar. This matters because the paper's contribution includes characterizing current practices, and the reader should be able to see how widely each practice was distributed among participants.
minor comments (5)
- [Abstract and Section 5.2] The abstract promises 'concrete steps towards the development of best practices,' but Section 5.2 explicitly says the authors 'focus on foregrounding key considerations' rather than offering prescriptive recommendations. Aligning these two statements would avoid overstating the practical guidance.
- [Section 4.1.1, P11 quote] The quote from P11 contains the hedge 'which I think could count as synthetic data insofar as it's making a judgment on the output of the [primary model].' The surrounding text should preserve this hedge and attribute the categorization to the interviewer's definitional frame, since P11's own framing is tentative.
- [Section 3 (Analysis)] The statement that the team 'surfaced 762 first-level codes' would be more useful if accompanied by a description of how saturation was assessed or a table of the final themes and their definitions, since the nine initial domain categories and four top-level categories are described only briefly.
- [Section 4.4.1] The discussion of 'chaining' auxiliary models is compelling but relies heavily on hypothetical examples (P10's GPT-4-as-scorer example) and expert concerns. Clarifying which observations are documented practices versus expert warnings would strengthen the analytical separation between findings and interpretation.
- [Table 1 and Section 3] The paper would benefit from reporting the distribution of participants by organization type (e.g., large technology companies, startups, academia) and by years of experience, since the paper's claims about organizational pressure and 'studying up' depend on the institutional positions of the participants.
Circularity Check
No significant circularity: central claims derive from new interviews; self-citations are minor and non-load-bearing.
full rationale
The paper's claimed derivation chain is empirical: 29 semi-structured interviews are coded via reflexive thematic analysis into findings and then into recommendations. I checked each circularity pattern against this chain. (1) No self-definitional step: the paper transparently scopes 'synthetic data' to include model-generated scores and labels (Section 2.1: 'We view this as a form of synthetic data as the scores produced by the auxiliary model take the place of labels...'), and then reports practitioners' practices under that scoping; the definition is an input framing, not something the findings are used to prove. (2) No fitted-input-called-prediction: the paper makes no quantitative predictions and fits no parameters; the 'ubiquitous integration' claim is a thematic summary of interview data. (3) Self-citations in related work (e.g., Heger et al. [37], Holstein et al. [40], Madaio et al. [61,62], Sambasivan et al. [83]) are contextual and not load-bearing; no uniqueness theorem or ansatz is imported from the authors' prior work. (4) No renaming of a known result as new organization: 'LLM-as-a-judge' is cited to external work and explicitly reframed as synthetic data for the paper's scope, which is transparent. The Limitations paragraph in Section 3 acknowledges the US-only, text-heavy sample; that is an external-validity limitation, not circularity. One participant's reported use of an auxiliary model's confidence to validate its own output (Section 4.3.1) is an object-level finding, not the paper's own reasoning. Thus the central claim rests on new interview evidence and is not circular; only minor non-load-bearing self-citations appear.
Assumptions & free parameters
assumptions (3)
- domain assumption Participant self-reports during interviews correspond to actual synthetic data practices.
- domain assumption The sample of 29 participants from 14 US organizations is sufficiently representative to support claims of widespread integration of synthetic data.
- ad hoc to paper Labels and scores produced by auxiliary models count as synthetic data.
Cite this review
Pith. "Pith review of Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline." pith.science (2026). https://pith.science/paper/4DOD7SRO
@misc{pith2026250118493,
author = {Pith},
title = {Pith review of: Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline},
year = {2026},
howpublished = {\url{https://pith.science/paper/4DOD7SRO}},
note = {Machine review of arXiv:2501.18493}
}
read the original abstract
Alongside the growth of generative AI, we are witnessing a surge in the use of synthetic data across all stages of the AI development pipeline. It is now common practice for researchers and practitioners to use one large generative model (which we refer to as an auxiliary model) to generate synthetic data that is used to train or evaluate another, reconfiguring AI workflows and reshaping the very nature of data. While scholars have raised concerns over the risks of synthetic data, policy guidance and best practices for its responsible use have not kept up with these rapidly evolving industry trends, in part because we lack a clear picture of current practices and challenges. Our work aims to address this gap. Through 29 interviews with AI practitioners and responsible AI experts, we examine the expanding role of synthetic data in AI development. Our findings reveal how auxiliary models are now widely used across the AI development pipeline. Practitioners describe synthetic data as crucial for addressing data scarcity and providing a competitive edge, noting that evaluation of generative AI systems at scale would be infeasible without auxiliary models. However, they face challenges controlling the outputs of auxiliary models, generating data that accurately depict underrepresented groups, and scaling data validation practices that are based primarily on manual inspection. We detail general limitations of and ethical considerations for synthetic data and conclude with a proposal of concrete steps towards the development of best practices for its responsible use.
Forward citations
Cited by 1 Pith paper
-
Synthetic CVs To Build and Test Fairness-Aware Hiring Tools
A new synthetic CV dataset, generated from donated real CVs, is proposed as a benchmark for fairness-aware algorithmic hiring research.
Reference graph
Works this paper leans on
-
[78]
Crystal Qian, Michael Xieyang Liu, Emily Reif, Grady Si mon, Nada Hussein, Nathan Clement, James Wexler, Carrie J Cai, Michael Terry, and Minsuk Kahng
-
[1]
Nazmiye Ceren Abay, Yan Zhou, Murat Kantarcioglu, Bhava ni Thuraisingham, and Latanya Sweeney. 2019. Privacy preserving synthetic da ta release using deep learning. In Machine Learning and Knowledge Discovery in Databases: Eu- ropean Conference, ECML PKDD 2018, Dublin, Ireland, Septemb er 10–14, 2018, Proceedings, Part I 18 . Springer, 510–526
2019
-
[2]
Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bube ck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. 2024. Phi-4 technical report. arXiv preprint arXiv:2412.08905 (2024)
arXiv 2024
-
[3]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad , Ilge Akkaya, Flo- rencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sa m Altman, Shya- mal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[4]
William Agnew, A Stevie Bergman, Jennifer Chien, Mark Dí az, Seliem El-Sayed, Jaylen Pittman, Shakir Mohamed, and Kevin R McKee. 2024. The illusion of artificial inclusion. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–12
2024
-
[5]
Samuel A Assefa, Danial Dervovic, Mahmoud Mahfouz, Robe rt E Tillman, Prashant Reddy, and Manuela Veloso. 2020. Generating synth etic data in fi- nance: opportunities, challenges and pitfalls. In Proceedings of the First ACM International Conference on AI in Finance . 1–8
2020
-
[6]
Pierre Azoulay, Joshua L Krieger, and Abhishek Nagaraj. 2024. Old moats for new models: Openness, control, and competition in generati ve ai. Technical Re- port. National Bureau of Economic Research
2024
-
[7]
Gwangbin Bae, Martin de La Gorce, Tadas Baltrušaitis, Ch arlie Hewitt, Dong Chen, Julien Valentin, Roberto Cipolla, and Jingjing Shen. 2023. DigiFace-1M: 1 Million Digital Face Images for Face Recognition. In 2023 IEEE Winter Confer- ence on Applications of Computer Vision (W ACV). IEEE
2023
Show all 116 references
-
[8]
Chelsea Barabas, Colin Doyle, JB Rubinovitz, and Karthi k Dinakar. 2020. Study- ing up: reorienting the study of algorithmic fairness aroun d issues of power. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 167–176
2020
-
[9]
Emily M Bender and Batya Friedman. 2018. Data statements for natural lan- guage processing: Toward mitigating system bias and enabli ng better science. Transactions of the Association for Computational Linguis tics 6 (2018), 587–604
2018
-
[10]
Virginia Braun and Victoria Clarke. 2012. Thematic analysis. American Psy- chological Association
2012
-
[11]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic anal- ysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597
2019
-
[12]
Sarah Burkhardt and Bernhard Rieder. 2024. Foundation models are platform models: Prompting and the political economy of AI. Big Data & Society 11, 2 (2024), 20539517241247839
2024
-
[13]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can large langu age models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937 (2023)
2023 arXiv
-
[14]
Kranzinger, Michael Hind, Jen- nifer Wortman Vaughan, Margaret Mitchell, Julia Stoyanovi ch, Angelina McMillan-Major, Emily McReynolds, Kathleen Esfahany, Mary L
Kasia Chmielinski, Sarah Newman, Chris N. Kranzinger, Michael Hind, Jen- nifer Wortman Vaughan, Margaret Mitchell, Julia Stoyanovi ch, Angelina McMillan-Major, Emily McReynolds, Kathleen Esfahany, Mary L. Gray, Audrey Chang, , and Maui Hudson. 2024. The CLeAR Documentation Fra...
2024
-
[15]
Anamaria Crisan, Brittany Fiore-Gartland, and Melani e Tory. 2020. Passing the data baton: A retrospective analysis on data science work an d workers. IEEE Transactions on Visualization and Computer Graphics 27, 2 (2020), 1860–1870
2020
-
[16]
Jessamyn Dahmen and Diane Cook. 2019. SynSys: A synthet ic data generation system for healthcare applications. Sensors 19, 5 (2019), 1181
2019
-
[17]
Fernando Delgado, Stephen Yang, Michael Madaio, and Qi an Yang. 2023. The participatory turn in ai design: Theoretical foundations a nd the current state of practice. In Proceedings of the 3rd ACM Conference on Equity and Access in FAccT ’25, June 23–26, 2025, Athens, Greece K...
2023
-
[18]
Wesley Hanwen Deng, Manish Nagireddy, Michelle Seng Ah Lee, Jatinder Singh, Zhiwei Steven Wu, Kenneth Holstein, and Haiyi Zhu. 20 22. Exploring how machine learning practitioners (try to) use fairness to olkits. In Proceed- ings of the 2022 ACM Conference on Fairness, Accounta...
2022
-
[19]
Wesley Hanwen Deng, Nur Yildirim, Monica Chang, Motahh are Eslami, Ken- neth Holstein, and Michael Madaio. 2023. Investigating Pra ctices and Opportu- nities for Cross-functional Collaboration around AI Fairn ess in Industry Prac- tice. In Proceedings of the 2023 ACM Conferenc...
2023
-
[20]
Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang , Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Has himoto. 2024. Alpacafarm: A simulation framework for methods that learn f rom human feed- back. Advances in Neural Information Processing S...
2024
-
[21]
Madeleine Clare Elish and Danah Boyd. 2018. Situating m ethods in the magic of Big Data and AI. Communication monographs 85, 1 (2018), 57–80
2018
-
[22]
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann , Caiming Xiong, Richard Socher, and Dragomir Radev. 2021. Summeval: Re-eva luating summa- rization evaluation. Transactions of the Association for Computational Linguis - tics 9 (2021), 391–409
2021
-
[23]
Shangbin Feng, Vidhisha Balachandran, Yuyang Bai, and Yulia Tsvetkov. 2023. Factkb: Generalizable factuality evaluation using langua ge models enhanced with factual knowledge. arXiv preprint arXiv:2305.08281 (2023)
2023 arXiv
-
[24]
Andrew Fitzgerald. 2024. Why Synthetic Data Can Never B e Ethical: A Les- son from Media Ethics. Surveillance & Society 22, 4 (Dec. 2024), 477–482. https://doi.org/10.24908/ss.v22i4.18324
2024 doi
-
[25]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Je nnifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM 64, 12 (December 2021), 86–92
2021
-
[26]
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 202 3. ChatGPT outper- forms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences 120, 30 (2023), e2305016120
2023
-
[27]
Raw Data
Lisa Gitelman. 2013. “Raw Data” Is an Oxymoron
2013
-
[28]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020. Generative adversarial networks. Commun. ACM 63, 11 (2020), 139–144
2020
-
[29]
Maybank, and Dach eng Tao
Jianping Gou, Baosheng Yu, Stephen J. Maybank, and Dach eng Tao. 2021. Knowledge Distillation: A Survey. International Journal of Computer Vision 129 (2021), 1789–1819
2021
-
[30]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab hinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur , Alan Schel- ten, Alex Vaughan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[31]
Igor Grossmann, Matthew Feinberg, Dawn C Parker, Nicho las A Christakis, Philip E Tetlock, and William A Cunningham. 2023. AI and the t ransformation of social science research. Science 380, 6650 (2023), 1108–1109
2023
-
[32]
Tom Gunter, Zirui Wang, Chong Wang, Ruoming Pang, Andy N arayanan, Aonan Zhang, Bowen Zhang, Chen Chen, Chung-Cheng Chiu, Davi d Qiu, et al. 2024. Apple intelligence foundation language models . arXiv preprint arXiv:2407.21075 (2024)
2024
-
[33]
Alon Halevy, Peter Norvig, and Fernando Pereira. 2009. The unreasonable ef- fectiveness of data. IEEE intelligent systems 24, 2 (2009), 8–12
2009
-
[34]
Lei Han, Tianwa Chen, Gianluca Demartini, Marta Induls ka, and Shazia Sadiq
-
[35]
Alex Hanna and Tina M Park. 2020. Against scale: Provoca tions and resistances to scale thinking. arXiv preprint arXiv:2010.08850 (2020)
2020 arXiv
-
[36]
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maa rten Sap, Di- pankar Ray, and Ece Kamar. 2022. ToxiGen: A Large-Scale Mach ine- Generated Dataset for Adversarial and Implicit Hate Speech Detection. http://arxiv.org/abs/2203.09509 arXiv:2203.09509 [cs]
2022 arXiv
-
[37]
Amy K Heger, Liz B Marquis, Mihaela Vorvoreanu, Hanna Wa llach, and Jen- nifer Wortman Vaughan. 2022. Understanding machine learni ng practitioners’ data documentation perceptions, needs, challenges, and desiderata. Proceedings of the ACM on Human-Computer Interaction 6, CSCW2...
2022
-
[38]
Paula Helm, Benjamin Lipp, and Roser Pujadas. 2024. Gen erating reality and silencing debate: Synthetic data as discursive device. Big Data & Society 11, 2 (June 2024), 20539517241249447. https://doi.org/10.1177/20539517241249447 Publisher: SAGE Publications Ltd
2024 doi
-
[39]
Hinton, Oriol Vinyals, and Jeffrey Dean
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 . Distilling the Knowl- edge in a Neural Network. arXiv preprint arXiv:1503.02531
2015 arXiv
-
[40]
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. 2019. Improving fairness in machine learnin g systems: What do industry practitioners need?. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–16
2019
-
[41]
Qisheng Hu, Kaixin Li, Xu Zhao, Yuxi Xie, Tiedong Liu, Hu i Chen, Qizhe Xie, and Junxian He. 2023. Instructcoder: Empowering language m odels for code editing. arXiv preprint arXiv:2310.20329 (2023)
2023 arXiv
-
[42]
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Lia ng, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebota r, et al. 2022. Inner monologue: Embodied reasoning through planning with language mod- els. arXiv preprint arXiv:2207.05608 (2022)
2022 arXiv
-
[43]
IBM. 2023. What is Synthetic Data? | IBM. https://www.ibm.com/think/topics/synthetic-data [Online; accessed 2025-01- 21]
2023
-
[44]
Apple Inc. 2024. Introducing Apple’s On-Device and Ser ver Foundation Models - Apple Machine Learning Research. https://machinelearning.apple.com/research/introducing-apple-foundation-models [Online; accessed 2025-01-21]
2024
-
[45]
Benjamin N Jacobsen. 2023. Machine learning and the pol itics of syn- thetic data. Big Data & Society 10, 1 (Jan. 2023), 20539517221145372. https://doi.org/10.1177/20539517221145372 Publisher: SAGE Publications Ltd
2023 doi
-
[46]
Nikita Jaipuria, Xianling Zhang, Rohan Bhasin, Mayar A rafa, Punarjay Chakravarty, Shubham Shrivastava, Sagar Manglani, and Vidya N Murali. 2020. Deflating dataset bias using synthetic data augmentation. I n Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2020
-
[47]
Harry H Jiang, Lauren Brown, Jessica Cheng, Mehtab Khan , Abhishek Gupta, Deja Workman, Alex Hanna, Johnathan Flowers, and Timnit Geb ru. 2023. AI Art and its Impact on Artists. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society . 363–374
2023
-
[48]
Synthetic Data–what, why and how? arXiv preprint arXiv:2205.03257 (2022)
James Jordon, Lukasz Szpruch, Florimond Houssiau, Mir ko Bottarelli, Giovanni Cherubin, Carsten Maple, Samuel N Cohen, and Adrian Weller.2022. Synthetic Data–what, why and how? arXiv preprint arXiv:2205.03257 (2022)
2022 arXiv
-
[49]
James Jordon, Jinsung Yoon, and Mihaela Van Der Schaar. 2018. PATE-GAN: Generating synthetic data with differential privacy guarantees. In International conference on learning representations
2018
-
[50]
Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A hunt for the Snark: An- notator Diversity in Data Practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). Association for Computing Ma- chinery, New York, NY, USA, 1–15. https:...
2023
-
[51]
Diederik P Kingma. 2013. Auto-encoding variational ba yes. arXiv preprint arXiv:1312.6114 (2013)
2013 arXiv
-
[52]
Laura Koesten, Elena Simperl, Tom Blount, Emilia Kacpr zak, and Jeni Tennison
-
[53]
Adam Kortylewski, Bernhard Egger, Andreas Schneider, Thomas Gerig, An- dreas Morel-Forster, and Thomas Vetter. 2019. Analyzing an d Reducing the Damage of Dataset Bias to Face Recognition With Synthetic Da ta. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2019
-
[54]
Francis Lee, Saghi Hajisharif, and Ericka Johnson. 202 5. The ontological politics of synthetic data: Normalities, outliers, and intersectio nal hallucinations. Big Data & Society 12, 2 (2025), 20539517251318289
2025
-
[55]
Peter Lee. 2024. Synthetic Data and the Future of AI. (20 24)
2024
-
[56]
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Haus man, Brian Ichter, Pete Florence, and Andy Zeng. 2023. Code as policies: Langua ge model pro- grams for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 9493–9500
2023
-
[57]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual in- struction tuning. Advances in neural information processing systems 36 (2024)
2024
- [58]
-
[59]
Dieuwertje Luitse and Wiebke Denkena. 2021. The great t ransformer: Examin- ing the role of large language models in the political econom y of AI. Big Data & Society 8, 2 (2021), 20539517211047734
2021
-
[60]
Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Jianguang Lo u, Chongyang Tao, Xiubo Geng, Qingwei Lin, Shifeng Chen, and Dongmei Zhang. 20 23. Wizard- math: Empowering mathematical reasoning for large languag e models via re- inforced evol-instruct. arXiv preprint arXiv:2308.09583 (2023)
2023 arXiv
-
[61]
Michael Madaio, Lisa Egede, Hariharan Subramonyam, Je nnifer Wort- man Vaughan, and Hanna Wallach. 2022. Assessing the Fairness of AI Systems: AI Practitioners’ Processes, Challenges, and Needs for Sup port. Proceedings of the ACM on Human-Computer Interaction 6, CSCW1 (2022),...
2022
-
[62]
Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan , and Hanna Wal- lach. 2020. Co-designing checklists to understand organiz ational challenges and opportunities around fairness in ai. In Proceedings of the 2020 CHI Confer- ence on Human Factors in Computing Systems . 1–14
2020
-
[63]
Daniella Meeker, Crystal Kallem, Yan Heras, Stephanie Garcia, and Casey Thompson. 2022. Case report: evaluation of an open-source s ynthetic data platform for simulation studies. JAMIA open 5, 3 (2022), ooac067. Examining the Expanding Role of Synthetic Data Throughout th e AI...
2022
-
[64]
Thom p- son, and Nico Grant
Cade Metz, Cecilia Kang, Sheera Frenkel, Stuart A. Thom p- son, and Nico Grant. 2024. How Tech Giants Cut Cor- ners to Harvest Data for A.I. - The New York Times. https://www.nytimes.com/2024/04/06/technology/tech-giants-harvest-data-artificial-intelligence.html [Online; access...
2024
-
[65]
Milagros Miceli, Martin Schuessler, and Tianling Yang . 2020. Between subjec- tivity and imposition: Power dynamics in data annotation fo r computer vision. Proceedings of the ACM on Human-Computer Interaction4, CSCW2 (2020), 1–25
2020
-
[66]
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasser- man, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru
-
[67]
Michael Muller, Ingrid Lange, Dakuo Wang, David Piorko wski, Jason Tsay, Q Vera Liao, Casey Dugan, and Thomas Erickson. 2019. How data science workers work with data: Discovery, capture, curation, desi gn, creation. In Pro- ceedings of the 2019 CHI conference on human factors ...
2019
-
[68]
Laura Nader. 1972. Up the anthropologist: Perspective s gained from studying up. (1972)
1972
-
[69]
Sergey I Nikolenko. 2021. Synthetic data for deep learning . Vol. 174. Springer
2021
-
[70]
Code of Practice Working Groups. 2024. Second Draft of the General-Purpose AI Code of Practice. https://digital-strategy.ec.europa.eu/en/library/se cond-draft-general-purpose-ai-code-practice-publish ed-written-independent-experts [Online; accessed 2025-01-20]
2024
-
[71]
Office of the Assistant Secretary for Planning and Evalua - tion (ASPE). 2022. A Synthetic Health Data Generation En- gine to Accelerate Patient-Centered Outcomes Research | AS PE. https://aspe.hhs.gov/synthetic-health-data-generati on-engine-accelerate-patient-centered-outcomes...
2022
-
[72]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carrol l Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Sl ama, Alex Ray, et al. 2022. Training language models to follow instruction s with human feed- back. Advances in neural information processing system...
2022
-
[73]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jin g Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Li nguistics. 311–318
2002
-
[74]
Samir Passi and Steven Jackson. 2017. Data Vision: Lear ning to See Through Algorithmic Abstraction. In Proceedings of the 2017 ACM Confer- ence on Computer Supported Cooperative Work and Social Comp uting (CSCW ’17). Association for Computing Machinery, New York, NY, USA, 24 ...
2017
-
[75]
Samir Passi and Steven J Jackson. 2018. Trust in data sci ence: Collaboration, translation, and accountability in corporate data science projects. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 1–28
2018
-
[76]
Personal Data Protection Commission (PDPC). 2024. Pri vacy Enhanc- ing Technology (PET): Proposed Guide On Synthetic Data Gene ration. https://www.pdpc.gov.sg/-/media/files/pdpc/pdf-files/other-guides/proposed-guide-on-synthetic-data-gener ation.pdf [Online; accessed 2025-01-20]
2024
-
[77]
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Ro man Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving . 2022. Red Team- ing Language Models with Language Models. arXiv:2202.03286 [cs] (Feb. 2022). http://arxiv.org/abs/2202.03286 arXiv: 2202.03286
2022 arXiv
-
[79]
Bogdana Rakova, Jingying Yang, Henriette Cramer, and R umman Chowdhury
-
[80]
Arij Riabi, Thomas Scialom, Rachel Keraron, Benoît Sag ot, Djamé Seddah, and Jacopo Staiano. 2020. Synthetic data augmentation for zero -shot cross-lingual question answering. arXiv preprint arXiv:2010.12643 (2020)
2020 arXiv
-
[81]
Kevin Roose. 2024. Data for A.I. Training Is Dis- appearing Fast, Study Shows - The New York Times. https://www.nytimes.com/2024/07/19/technology/ai-data-restrictions.html [Online; accessed 2025-01-21]
2024
-
[82]
Roy A Ruddle, James Cheshire, and Sara Johansson Fernst ad. 2023. Tasks and visualizations used for data profiling: A survey and interview study. IEEE Trans- actions on Visualization and Computer Graphics (2023)
2023
-
[83]
Everyone wants to do the mo del work, not the data work
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, D iana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the mo del work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Syst...
2021
-
[84]
Siamak Shakeri, Noah Constant, Mihir Sanjay Kale, and L inting Xue. 2020. To- wards zero-shot multilingual synthetic question and answe r generation for cross-lingual reading comprehension. arXiv preprint arXiv:2010.12008 (2020)
2020 arXiv
-
[85]
Hoo-Chang Shin, Neil A Tenenholtz, Jameson K Rogers, Ch ristopher G Schwarz, Matthew L Senjem, Jeffrey L Gunter, Katherine P Andr iole, and Mark Michalski. 2018. Medical image synthesis for data augm entation and anonymization using generative adversarial networks. In Simulatio...
2018
-
[86]
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Ander- son, and Yarin Gal. 2024. AI models collapse when trained on r ecursively gen- erated data. Nature 631 (2024), 755–759
2024
-
[87]
Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of conte nt moderators: the emotional labor of commercial moderation and avenues fo r improving sup- port. In Proceedings of the 2021 CHI conference on h...
2021
-
[88]
Daniel Susser, Daniel S Schiff, Sara Gerke, Laura Y Cabre ra, I Glenn Co- hen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet, J asmine McNealy, Michelle N Meyer, et al. 2024. Synthetic Health Data: Real Et hical Promise and Peril. Hastings Center Report 54, 5 (2024), 8–13
2024
-
[89]
Schiff, Sara Gerke, Laura Y
Daniel Susser, Daniel S. Schiff, Sara Gerke, Laura Y. Cab rera, I. Glenn Cohen, Megan Doerr, Jordan Harrod, Kristin Kostick-Quenet , Jasmine Mc- Nealy, Michelle N. Meyer, W. Nicholson Price II, and Jennife r K. Wagner
-
[90]
Daniel Susser and Jeremy Seeman. [n. d.]. Dialogue Crit ical Provocations for Synthetic Data. ([n. d.])
-
[91]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Bap tiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Milli- can, et al. 2023. Gemini: a family of highly capable multimod al models. arXiv preprint arXiv:2312.11805 (2023)
2023 arXiv
-
[92]
Risto Uuk, Annemieke Brouwer, Tim Schreier, Noemi Drek sler, Valeria Pulig- nano, and Rishi Bommasani. 2024. Effective Mitigations for S ystemic Risks from General-Purpose AI. arXiv preprint arXiv:2412.02145 (2024)
2024 arXiv
-
[93]
Janet Vertesi and Paul Dourish. 2011. The value of data: considering the context of production in data economies. In Proceedings of the ACM 2011 conference on Computer supported cooperative work . 533–542
2011
-
[94]
Dickerso n
Angelina Wang, Jamie Morgenstern, and John P. Dickerso n. 2024. Large lan- guage models should not replace human participants because they can mis- portray and flatten identity groups. arXiv preprint arXiv:2 402.01908
2024
-
[95]
Hastings Center Report 54, 5 (2024), 8–13
Synthetic Health Data: Real Ethical Promise and Peril . Hastings Center Report 54, 5 (2024), 8–13. https://doi.org/10.1002/hast.4911 _eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/hast.4911
2024 doi
-
[96]
Cedric Deslandes Whitney and Justin Norman. 2024. Real Risks of Fake Data: Synthetic Data, Diversity-Washing and Consent Circumvent ion. In The 2024 ACM Conference on Fairness, Accountability, and Transpare ncy. ACM, Rio de Janeiro Brazil, 1733–1744. https://doi.org/10.1145/36...
2024
-
[97]
AI supply chain
David Gray Widder and Dawn Nafus. 2023. Dislocated acco untabilities in the “AI supply chain”: Modularity and developers’ notions of re sponsibility. Big Data & Society 10, 1 (2023), 20539517231177620
2023
-
[98]
David Gray Widder, Sarah West, and Meredith Whittaker. 2023. Open (for busi- ness): Big tech, concentrated power, and the political economy of open AI. Con- centrated Power, and the Political Economy of Open AI (August 1 7, 2023) (2023)
2023
-
[99]
David Gray Widder and Richmond Wong. 2023. Thinking ups tream: Ethics and policy opportunities in AI supply chains. arXiv preprint arXiv:2303.07529 (2023)
2023 arXiv
-
[100]
David Gray Widder, Derrick Zhen, Laura Dabbish, and Ja mes Herbsleb. 2023. It’s about power: What ethical concerns do software enginee rs have, and what do they (feel they can) do about them?. In Proceedings of the 2023 ACM Confer- ence on Fairness, Accountability, and Transpa...
2023
-
[101]
Ding Wang, Shantanu Prabhat, and Nithya Sambasivan. 20 22. Whose AI Dream? In search of the aspiration in data annotation.. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–16
2022
-
[102]
Matteo Wong. 2024. The GPT Era Is Already Ending - The At lantic. https://www.theatlantic.com/technology/archive/2024/12/openai-o1-reasoning-models/680906/?utm_source=c hatgpt.com [Online; accessed 2025-01-22]
2024
-
[103]
Kanit Wongsuphasawat, Yang Liu, and Jeffrey Heer. 2019 . Goals, process, and challenges of exploratory data analysis: An interview stud y. arXiv preprint arXiv:1911.00568 (2019)
2019 arXiv
-
[104]
Sierra Wyllie, Ilia Shumailov, and Nicolas Papernot. 2024. Fairness feedback loops: training on synthetic data amplifies bias. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 2113–2147
2024
-
[105]
Liyang Xie, Kaixiang Lin, Shu Wang, Fei Wang, and Jiayu Zhou. 2018. Differen- tially private generative adversarial network. arXiv preprint arXiv:1802.06739 (2018)
2018 arXiv
-
[106]
Zhenpei Yang, Yuning Chai, Dragomir Anguelov, Yin Zho u, Pei Sun, Dumitru Erhan, Sean Rafferty, and Henrik Kretzschmar. 2020. Surfelg an: Synthesizing realistic sensor data for autonomous driving. In Proceedings of the IEEE/CVF FAccT ’25, June 23–26, 2025, Athens, Greece Kapani...
2020
-
[107]
Tanja Wiehn. 2024. Synthetic Data: From Data Scarcity to Data Pollution. Surveillance & Society 22, 4 (Dec. 2024), 472–476. https://doi.org/10.24908/ss.v22i4.18327
2024 doi
-
[108]
Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengy ing Liu, Yu Zhang, James T Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu. 202 3. Metamath: Bootstrap your own mathematical questions for large langua ge models. arXiv preprint arXiv:2309.12284 (2023)
2023 arXiv
-
[109]
Amy X Zhang, Michael Muller, and Dakuo Wang. 2020. How d o data science workers collaborate? Roles, workflows, and tools. Proceedings of the ACM on Human-Computer Interaction 4, CSCW1 (2020), 1–23
2020
-
[110]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhu ang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, etal. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Informa- tion Processing Systems 36 (2023), 46595–46623. A ...
2023
-
[113]
Nur Yildirim, Mahima Pushkarna, Nitesh Goyal, Martin Wattenberg, and Fer- nanda Viégas. 2023. Investigating How Practitioners Use Human-AI Guidelines: A Case Study on the People+ AI Guidebook. In Proceedings of the 2023 CHI Con- ference on Human Factors in Computing Systems . 1–13
2023
-
[2019]
In Proceedings of the Conference on Fair- ness, Accountability, and Transparency
Model cards for model reporting. In Proceedings of the Conference on Fair- ness, Accountability, and Transparency. 220–229
-
[2020]
International Journal of Human-Computer Studies 135 (March 2020), 102367
Everything you always wanted to know about a dataset: S tudies in data summarisation. International Journal of Human-Computer Studies 135 (March 2020), 102367. https://doi.org/10.1016/j.ijhcs.2019.10.004
2020 doi
-
[2021]
Proceedings of the ACM on Human- Computer Interaction 5, CSCW1 (2021), 1–23
Where responsible AI meets reality: Practitioner per spectives on en- ablers for shifting organizational practices. Proceedings of the ACM on Human- Computer Interaction 5, CSCW1 (2021), 1–23
2021
-
[2023]
ACM Transactions on Information Systems 41, 3 (2023), 1–35
A data-driven analysis of behaviors in data curation p rocesses. ACM Transactions on Information Systems 41, 3 (2023), 1–35
2023
-
[2024]
arXiv preprint arXiv:2412.16089 (2024)
The Evolution of LLM Adoption in Industry Data Curatio n Practices. arXiv preprint arXiv:2412.16089 (2024)
2024 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.