Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes

T0 review · 2 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read Community ambassadors, not web forms, made this food dataset possible.

desk verdict A candid, well-scoped retrospective on mediator labor in participatory dataset building; the causal 'necessity' claim outruns the self-reported evidence, but the descriptive account and five lessons are worth a serious read. read the letter →

arxiv 2502.05961 v3 pith:P4EXRL6R submitted 2025-02-09 cs.CY

classification cs.CY
keywords dataworkculturalrepresentationpositionalitycommunity-baseddesignparticipatoryAICommunityAmbassadorslong-tailedcrowdsourcing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that bottom-up, community-led dataset building succeeds only when trusted community insiders perform visible human labour that connects, supports, and interprets for the people contributing data. Drawing on a design retrospective of World Wide Dishes, a participatory image-and-text dataset of 765 dishes from 106 countries, the authors identify three functions of this labour by Community Ambassadors: building trust between communities and researchers, making participation genuinely accessible, and contextualising cultural practices so the data captured is meaningful. The paper treats this mediating labour as necessary infrastructure for participatory AI, not a side activity, and turns the experience into five lessons for building future data-collection projects. A reader should care because the result shifts attention from the dataset as a product to the social process that produces it, with concrete implications for how participatory AI is funded, staffed, and evaluated.

What carries the argument

The central object is the Community Ambassador, a participatory mediator who is simultaneously a community insider and a member of the research team. The paper defines this role in contrast to a simple community gatekeeper: ambassadors shaped the data collection instrument, hosted focus groups, translated and sometimes filled in submissions for contributors, reviewed regional entries, and fed community concerns back to the Core Organisers. Three named labour dimensions carry the argument: building trust, making participation accessible, and contextualising community values. The design retrospective, built from documentation artefacts and post-mortem reflections from the organisers, is the method that surfaces these dimensions, and the three dimensions are the mechanism claimed to convert community willingness into curated, granular data.

What would settle it

A concrete test would be a parallel participatory dataset effort in comparable communities where one group is onboarded through Community Ambassadors and a control group receives only the accessible web form; if the control group produces comparable submission rates and metadata quality, the claim that ambassador labour is necessary would be falsified. A simpler observational check would survey Community Contributors directly about why they participated and compare their answers with the ambassadors' retrospective accounts.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that Community Ambassadors make bottom-up, community-led data collection possible by performing three kinds of labour: they build trust with community members, they make participation accessible, and they contextualise community values so that data collection is meaningful. The authors show through their retrospective analysis of World Wide Dishes that data did not simply arrive once the web form existed; it was produced through social interaction, with contributors consulting parents and grandparents, ambassadors holding office hours and focus groups, and the team moving to the messaging platforms where contributors already were. In this account, the dataset's fine-grained regional metadata and the participation of communities that usually do not appear in web-scraped datasets are direct results of this mediating work. The authors conclude that this labour is necessary to make participatory AI efforts successful in action, and they derive five design lessons for building infrastructure that supports it.

Load-bearing premise

The account rests on retrospective self-reports from the nine Core Organisers, who were themselves the Community Ambassadors, with the coding done by two co-authors who were also participants; no systematic data from Community Contributors or outside observers checks whether these reported causes were the real reasons people contributed.

Editorial extensions

If this is right

  • Participatory dataset projects should budget for and formally support Community Ambassador roles, treating them as core infrastructure rather than volunteer extras.
  • Cultural data collection should support collaborative and social production, such as 'phone a friend' contributions, rather than assuming each entry comes from one isolated contributor.
  • Open-source and Creative Commons licensing impose significant downstream verification labour that can shrink a volunteer-built dataset, so future projects need funding and infrastructure for that work.
  • New participatory efforts should anchor themselves in existing community-led initiatives and communication channels rather than importing unfamiliar platforms.
  • The promises made to contributors must be honest: better representation in a dataset cannot by itself undo historical exclusion, so recruitment and consent processes should say so.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not test this, but if ambassador labour is the load-bearing input, then a testable extension is whether paid or formally recognised ambassador roles yield measurably higher contribution volume and metadata richness than volunteer roles in equivalent communities.
  • As an extension the authors do not draw, dataset documentation could record process metadata, such as who mediated, how trust was built, and what accessibility accommodations were made, because that context shapes what the data can legitimately be used for.
  • The three labour dimensions may generalise to other intangible cultural heritage domains, but the emotional and relational cost of asking friends and family for unpaid contributions may not scale the same way, so future work should track ambassador burden explicitly.
  • The strongest check on this account, beyond what the paper includes, would be direct interviews with Community Contributors about why they participated, a data source the present retrospective does not include.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This paper presents World Wide Dishes (WWD), a bottom-up, community-led dataset of culinary dishes and cultural customs. Through a design retrospective employing first-person methods, the authors analyze documentation and post-mortem reflections from nine Core Organisers (who also served as Community Ambassadors) to argue that Community Ambassadors performed three kinds of labour—building trust, making participation accessible, and contextualising community values—that made the data collection effort possible. The paper derives five lessons for infrastructure supporting participatory AI. The manuscript is transparent about its methods and positionality and includes detailed appendices.

Significance. If the account is accepted, the paper provides a rich empirical characterization of the hidden human labour in participatory dataset construction, extending CSCW work on data work and mediators to the participatory AI context. The manuscript's strengths include detailed documentation of the data collection system, explicit positionality statements, an appendix of reflection prompts, and recognition of all contributors. The descriptive account of Community Ambassador activities is well supported by concrete quotes. However, the significance of the broader claim that CA labour is necessary for participatory AI success depends on the strength of the causal inference, which is the main point of concern.

major comments (2)
  1. [Section 6 and Section 5.2] The central claim that CA labour 'made bottom-up, community-led data collection possible' and is 'necessary' (Section 1) is supported only by retrospective self-reports of nine Core Organisers who also served as Community Ambassadors. The study has no variation in CA presence, no counterfactual, and no data from Community Contributors on why they participated. The prompts in Appendix F, such as 'Can you reflect on why you think contributors engaged with you?', invite the very explanatory account that the analysis abstracts into themes, and the coding was done by participants. This does not invalidate the descriptive findings, but it cannot support the general necessity claim on which the five lessons in Section 7 rely. I recommend reframing the claims as accounts of the authors' experience ('in this project, CA labour was necessary') or adding contributor-side or observational evidence.
  2. [Section 5.1 vs. Table E.1 and Section 6] Quantitative details are inconsistent. Section 5.1 reports the final addition of COs in 'April 2025,' which is after the April 2024 data collection period and appears to be a typo for April 2024; Section 6 reports 'twelve COs' while Section 5.3 says nine COs produced post-mortems and the author list includes eleven; and Section 6 reports 'more than 170 CAs and CCs' while Table E.1 reports 162 Contributors and Community Ambassadors. These discrepancies need resolution because the scale of the effort is part of the paper's empirical grounding.
minor comments (5)
  1. [Section 5.1] The date 'April 2025' should presumably read 'April 2024'; the timeline otherwise places the final CO addition after the data collection period.
  2. [References [45] and [46]] The author field in both references begins with '∀,' which appears to be a placeholder or corrupt character; this should be corrected to the actual author names.
  3. [Appendix F] The text 'share a voicenote' should be 'share a voice note'.
  4. [Section 7, Lesson Five] The phrase 'The authors speculated' breaks the first-person voice used throughout the paper; 'we speculated' would be consistent with the paper's methodological framing.
  5. [Section 4.3.1] The claim that 'COs and CAs worked together to ensure the data collection system was easily shareable... to support multiple browsers' is vague; specify what compatibility testing was performed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central themes are grounded in documented reflections and concrete project outcomes, not in a self-referential equation or load-bearing self-citation.

full rationale

The paper's central claim—that Community Ambassadors' labour (building trust, making participation accessible, contextualising values) made bottom-up data collection possible—is a qualitative interpretive finding from a design retrospective, not a quantity fitted to data or a definitional tautology. The evidence chain runs from documentation artefacts and post-mortem reflections (Sections 5.2–5.3), through quoted CA accounts (Section 6), to lessons (Section 7); the themes are analytic labels for specific reported actions (e.g., O6 individually reaching out when social-media posts failed, the co-designed concise consent form, office hours explaining image licensing). Although the authors were also participants, Section 5.4 explicitly discloses this first-person, autoethnographic stance; that is a validity limitation, not the circularity pattern of 'finding = input by construction.' The self-citations present are peripheral: [76] is a nod to the authors' earlier WWD dataset paper and does not carry the CA-labour argument, and [32] supports a normative point in Lesson Five rather than the empirical derivation. No uniqueness theorem, ansatz, or fitted parameter is smuggled in through self-citation. The dataset's documented outcomes (765 dishes, 106 countries, 131 languages, Table E.2) and the quoted contributor anecdotes provide independent anchors for the retrospective narrative. The absence of contributor-side counterfactual data weakens the causal 'necessary' claim, but that is an evidentiary gap to be weighed under correctness, not evidence of circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on assumptions about dataset bias, the cultural significance of food, the value of bottom-up crowdsourcing, and the validity of retrospective self-analysis. No numerical free parameters or invented entities are introduced.

assumptions (4)
  • domain assumption Web-scraped Internet datasets disproportionately represent perspectives from the Global North and reflect existing power hierarchies.
    Invoked in Sections 1 and 3 as the motivating problem; no empirical verification is provided within this paper.
  • domain assumption Food is a salient cultural artifact deeply intertwined with history, geography, and religious symbolism.
    Stated in Sections 1 and 2.1 as the reason for choosing food as the data domain; treated as self-evident with references to UNESCO.
  • domain assumption Bottom-up community-based crowdsourcing can yield datasets that better represent marginalized communities compared to top-down approaches.
    Asserted in Section 2.3 with examples from AYA and Masakhane, but the present paper does not compare outcomes against a top-down baseline.
  • domain assumption First-person retrospective methods (collaborative autoethnography) provide valid evidence about design processes.
    This methodological premise is adopted in Section 5.2 and justified by citation to prior HCI work, not demonstrated within the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes." pith.science (2026). https://pith.science/paper/P4EXRL6R

@misc{pith2026250205961,
  author       = {Pith},
  title        = {Pith review of: The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P4EXRL6R}},
  note         = {Machine review of arXiv:2502.05961}
}
read the original abstract

This paper provides guidance for building and maintaining infrastructure for participatory AI efforts by sharing reflections on building World Wide Dishes (WWD), a bottom-up, community-led image and text dataset of culinary dishes and associated cultural customs. We present WWD as an example of participatory dataset creation, where community members both guide the design of the research process and contribute to the crowdsourced dataset. This approach incorporates localised expertise and knowledge to address the limitations of web-scraped Internet datasets acknowledged in the Participatory AI discourse. We show that our approach can result in curated, high-quality data that supports decentralised contributions from communities that do not typically contribute to datasets due to a variety of systemic factors. Our project demonstrates the importance of participatory mediators in supporting community engagement by identifying the kinds of labour they performed to make WWD possible. We surface three dimensions of labour performed by participatory mediators that are crucial for participatory dataset construction: building trust with community members, making participation accessible, and contextualising community values to support meaningful data collection. Drawing on our findings, we put forth five lessons for building infrastructure to support future participatory AI efforts.

Figures

Figures reproduced from arXiv: 2502.05961 by the authors.

Figure 1
Figure 1. Stakeholders in World Wide Dishes. There are three stakeholder groups in WWD: Community Contributors, Community Ambassadors, and Core Organisers. This figure illustrates the overlapping stake￾holder roles in the project. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW492. Publication date: November 2025 [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Overview of the World Wide Dishes flow. (A) A Community Contributor accesses WWD through a web browser. They consent to be a research participant and decide whether to create an account or proceed as a guest. They then fill out the data collection form with information about themselves and the dish they submit. (B) The submission is then stored in the WWD database. We store Community Contributors’ information separa… view at source ↗
Figure 3
Figure 3. Screenshot of the World Wide Dishes Leaderboard showing the top 10 countries by number of dishes contributed, the total number of dishes contributed, and the total number of contributors per country. (2) Translations were required in special cases. For submissions made by French-speaking CCs from the Democratic Republic of Congo, and to make concessions for accessibility, the COs accepted these specific entries and … view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

Reference graph

Works this paper leans on

122 extracted references · 47 canonical work pages · cited by 1 Pith paper

  1. [1]

    [n. d.]. AfriQA: Cross-lingual Open-Retrieval Question Answering for African Languages, author = Ogundepo, Odunayo and Gwadabe, Tajuddeen R and Rivera, Clara E and Clark, Jonathan H and Ruder, Sebastian and Adelani, David Ifeoluwa and Dossou, Bonaventure FP and Diop, Abdou Aziz and Sikasote, Claytone and Hacheme, Gilles and others, year = 2023, journal = ...

  2. [2]

    [n. d.]. The Alternative Epistemologies of Data Activism, author=Milan, Stefania and Velden, Lonneke van der, jour- nal=Digital culture & society, volume=2, number=2, pages=57–74, year=2016, publisher=transcript Verlag. ([n. d.])

  3. [3]

    [n. d.]. Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark Datasets, author = Blodgett, Su Lin and Lopez, Gilsinia and Olteanu, Alexandra and Sim, Robert and Wallach, Hanna, year = 2021, booktitle = Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conferenc...

  4. [4]

    David Ifeoluwa Adelani, Graham Neubig, Sebastian Ruder, Shruti Rijhwani, Michael Beukman, Chester Palen-Michel, Constantine Lignos, Jesujoba O Alabi, Shamsuddeen H Muhammad, Peter Nabende, et al. 2022. MasakhaNER 2.0: Africa-centric Transfer Learning for Named Entity Recognition. arXiv preprint arXiv:2210.12391 (2022)

  5. [5]

    Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Ashutosh Dwivedi, Al- ham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. 2024. Towards Measuring and Modeling “Culture” in LLMs: A Survey. arXiv preprint arXiv:2403.15412 (2024)

  6. [6]

    Peter M Asaro. 2000. Transforming society by transforming technology: the science and politics of participatory design. Accounting, Management and Information Technologies 10, 4 (2000), 257–290

  7. [7]

    Seyram Avle, Emmanuel Quartey, and David Hutchful. 2018. Research on Mobile Phone Data in the Global South: Opportunities and Challenges. (2018)

  8. [8]

    Liam Bannon, Jeffrey Bardzell, and Susanne Bødker. 2018. Reimagining Participatory Design. Interactions 26, 1 (2018), 26–32

Show all 122 references
  1. [9]

    Chelsea Barabas, Colin Doyle, JB Rubinovitz, and Karthik Dinakar. 2020. Studying Up: Reorienting the study of algorithmic fairness around issues of power. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain) (FAT* ’20). Associa...

  2. [10]

    Ruha Benjamin. 2016. Informed Refusal: Toward a Justicebased Bioethics. Science, Technology, & Human Values 41, 6 (2016), 967–990

  3. [11]

    R Benjamin. 2019. Race after Technology: Abolitionist Tools for the New Jim Code . Cambridge and Medford: Polity Press

  4. [12]

    Sebastian Benthall and Bruce D Haynes. 2019. Racial categories in machine learning. In Proceedings of the conference on Fairness, Accountability, and Transparency. 289–298

  5. [13]

    Hugo Berg, Siobhan Mackenzie Hall, Yash Bhalgat, Wonsuk Yang, Hannah Rose Kirk, Aleksandar Shtedritski, and Max Bain. 2022. A Prompt Array Keeps the Bias Away: Debiasing Vision-Language Models with Adversarial Learning. arXiv preprint arXiv:2203.11933 (2022)

  6. [14]

    Mukul Bhutani, Kevin Robinson, Vinodkumar Prabhakaran, Shachi Dave, and Sunipa Dev. 2024. SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes. arXiv preprint arXiv:2403.05696 (2024)

  7. [15]

    Abeba Birhane, William Isaac, Vinodkumar Prabhakaran, Mark Díaz, Madeleine Clare Elish, Iason Gabriel, and Shakir Mohamed. 2022. Power to the People? Opportunities and Challenges for Participatory AI. arXiv:2209.07572 [cs] doi:10.1145/3551624.3555290

  8. [16]

    Abeba Birhane and Vinay Uday Prabhu. 2021. Large image datasets: A pyrrhic win for computer vision?. In 2021 IEEE Winter Conference on Applications of Computer Vision (W ACV). IEEE, 1536–1546

  9. [17]

    Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021. Multimodal datasets: misogyny, pornography, and malignant stereotypes. arXiv preprint arXiv:2110.01963 (2021)

  10. [18]

    Janet Blake. 2000. On Defining the Cultural Heritage. International & Comparative Law Quarterly 49, 1 (2000), 61–85

  11. [19]

    Susanne Bødker, Christian Dindler, Ole S Iversen, and Rachel C Smith. 2022. What Can We Learn from the History of Participatory Design? In Participatory Design. Springer, 15–29. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW492. Publication date: November 2025. ...

  12. [20]

    Marcel Borowski, Bjarke V Fog, Carla F Griggio, James R Eagan, and Clemens N Klokmose. 2023. Between Principle and Pragmatism: Reflections on Prototyping Computational Media with Webstrates. ACM Transactions on Computer- Human Interaction 30, 4 (2023), 1–53

  13. [21]

    Geoffrey C Bowker. 2000. Sorting things out: Classification and its consequences . MIT press

  14. [22]

    Tone Bratteteig and Ina Wagner. 2016. Unpacking the Notion of Participation in Participatory Design. Computer Supported Cooperative Work (CSCW) 25 (2016), 425–475

  15. [23]

    Jed R Brubaker and Gillian R Hayes. 2011. SELECT* FROM USER: infrastructure and socio-technical representation. In Proceedings of the ACM 2011 conference on Computer supported cooperative work . 369–378

  16. [24]

    Joy Buolamwini and Timnit Gebru. 2018. Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research, Vol. 81), Sorelle A. Frie...

  17. [25]

    Cartografias da Internet. 2025. Cartografias da Internet. https://www.cartografiasdainternet.org/en/ [Accessed 13 January 2025]

  18. [26]

    Heewon Chang. 2021. Individual and Collaborative Autoethnography for Social Science Research. In Handbook of autoethnography. Routledge, 53–65

  19. [27]

    John Cheney-Lippold. 2017. We Are Data: Algorithms and the Making of our Digital Selves. In We Are Data. New York University Press

  20. [28]

    Marika Cifor and Patricia Garcia. 2019. Inscribing Gender: A Duoethnographic Examination of Gendered Values and Practices in Fitness Tracker Design. (2019)

  21. [29]

    Combahee River Collective. 1980. Eleven Black Women: Why Did They Die? The Collective

  22. [30]

    Eric Corbett, Emily Denton, and Sheena Erete. 2023. Power and Public Participation in AI. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization . 1–13

  23. [31]

    Geoffrey Currie, Johnathan Hewis, Elizabeth Hawk, and Eric Rohren. 2024. Gender and Ethnicity Bias of Text-to-Image Generative Artificial Intelligence in Medical Imaging, Part 1: Preliminary Evaluation. Journal of Nuclear Medicine Technology (2024)

  24. [32]

    inclusion

    Samantha Dalal, Siobhan Mackenzie Hall, and Nari Johnson. 2024. Provocation: Who benefits from" inclusion" in Generative AI? arXiv preprint arXiv:2411.09102 (2024)

  25. [33]

    Jol" or

    Dipto Das, Carsten Østerlund, and Bryan Semaan. 2021. "Jol" or "Pani"?: How Does Governance Shape a Platform’s Identity? Proc. ACM Hum.-Comput. Interact. 5, CSCW2, Article 473 (Oct. 2021), 25 pages. doi:10.1145/3479860

  26. [34]

    Dipto Das and Bryan Semaan. 2022. Collaborative Identity Decolonization as Reclaiming Narrative Agency: Identity Work of Bengali Communities on Quora. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association ...

  27. [35]

    Fernando Delgado, Stephen Yang, Michael Madaio, and Qian Yang. 2023. The Participatory Turn in AI Design: Theoretical Foundations and the Current State of Practice. In Proceedings of the 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization . 1–23

  28. [36]

    Remi Denton, Mark Díaz, Ian Kivlichan, Vinodkumar Prabhakaran, and Rachel Rosen. 2021. Whose Ground Truth? Accounting for Individual and Collective Identities Underlying Dataset Annotation. arXiv preprint arXiv:2112.04554 (2021)

  29. [37]

    Audrey Desjardins, Oscar Tomico, Andrés Lucero, Marta E Cecchinato, and Carman Neustaedter. 2021. Introduction to the special issue on first-person methods in HCI. ACM Transactions on Computer-Human Interaction (TOCHI) 28, 6, 1–12

  30. [38]

    Catherine D’Ignazio. 2024. Counting feminicide: Data feminism in action . MIT Press

  31. [39]

    Catherine D’ignazio and Lauren F Klein. 2023. Data feminism. MIT press

  32. [40]

    Emile Durkheim and Marcel Mauss. 2009. Primitive classification (Routledge revivals) . Routledge

  33. [41]

    Steven Epstein. 2008. The Rise of Recruitmentology’ Clinical Research, Racial Knowledge, and the Politics of Inclusion and Difference. Social Studies of Science 38, 5 (2008), 801–832

  34. [42]

    Rankin, and Jakita O

    Sheena Erete, Yolanda A. Rankin, and Jakita O. Thomas. 2021. I Can’t Breathe: Reflections from Black Women in CSCW and HCI. Proc. ACM Hum.-Comput. Interact. 4, CSCW3, Article 234 (Jan. 2021), 23 pages. doi:10.1145/3432933

  35. [43]

    Dipo Faloyin. 2022. Africa is Not a Country: Breaking Stereotypes of Modern Africa . Random House, New York

  36. [44]

    Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander Toshev, and Vaishaal Shankar. 2023. Data Filtering Networks. arXiv preprint arXiv:2309.17425 (2023)

  37. [45]

    ∀, Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Tajudeen Kolawole, Taiwo Fagbo- hungbe, Solomon Oluwole Akinola, Shamsuddeen Hassan Muhammad, Salomon Kabongo, Salomey Osei, et al. 2020. Participatory Research for Low-resourced Machine Translation: A C...

  38. [46]

    ∀, Iroro Orife, Julia Kreutzer, Blessing Sibanda, Daniel Whitenack, Kathleen Siminyu, Laura Martinus, Jamiil Toure Ali, Jade Abbott, Vukosi Marivate, Salomon Kabongo, et al. 2020. Masakhane–Machine Translation For Africa. arXiv preprint arXiv:2003.11529 (2020)

  39. [47]

    Andrew L Friedman and Dominic S Cornford. 1989. Computer Systems Development: History Organization and Implementation. John Wiley & Sons, Inc

  40. [48]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86–92

  41. [49]

    Clifford Geertz. 2008. Thick Description: Toward an Interpretive Theory of Culture. In The cultural geography reader . Routledge, 41–51

  42. [50]

    R Stuart Geiger, Kevin Yu, Yanlai Yang, Mindy Dai, Jie Qiu, Rebekah Tang, and Jenny Huang. 2020. Garbage In, Garbage Out? Do Machine Learning Application Papers in Social Computing Report Where Human-Labeled Training Data Comes From?. In Proceedings of the 2020 conference on F...

  43. [51]

    Tarleton Gillespie. 2014. The Relevance of Algorithms. Media technologies: Essays on communication, materiality, and society 167, 2014 (2014), 167

  44. [52]

    Tarleton Gillespie. 2024. Generative AI and the politics of visibility.Big Data & Society 11, 2 (2024), 20539517241252131

  45. [53]

    Jonathan Grudin. 1988. Why CSCW Applications Fail: Problems in the Design and Evaluation of Organizational Interfaces. In Proceedings of the 1988 ACM conference on Computer-supported cooperative work . 85–93

  46. [54]

    Akhil Gupta and James Ferguson. 2008. Beyond "Culture": Space, Identity, and the Politics of Difference. In The cultural geography reader. Routledge, 72–79

  47. [55]

    Siobhan Mackenzie Hall, Fernanda Gonçalves Abrantes, Hanwen Zhu, Grace Sodunke, Aleksandar Shtedritski, and Hannah Rose Kirk. 2024. VisoGender: A dataset for benchmarking gender bias in image-text pronoun resolution. Advances in Neural Information Processing Systems 36 (2024)

  48. [56]

    Donna Haraway. 2013. Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective

  49. [57]

    Routledge, 455–472

    In Women, science, and technology. Routledge, 455–472

  50. [58]

    Christina Harrington, Sheena Erete, and Anne Marie Piper. 2019. Deconstructing Community-Based Collaborative Design: Towards More Equitable Participatory Design Engagements. Proceedings of the ACM on human-computer interaction 3, CSCW (2019), 1–25

  51. [59]

    Tajanae Harris. 2024. Data for Whom, Data from Whom: How Social Movements Might Create Value for Their Community Data Practices. XRDS 30, 4 (June 2024), 31–35. doi:10.1145/3665597

  52. [60]

    Sarah Homewood. 2023. Self-Tracking to Do Less: An Autoethnography of Long COVID That Informs the Design of Pacing Technologies. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–14

  53. [61]

    Rachel Hong, William Agnew, Tadayoshi Kohno, and Jamie Morgenstern. 2024. Who’s in and who’s out? A case study of multimodal CLIP-filtering in DataComp. In Proceedings of the 4th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization . 1–17

  54. [62]

    Rachel Hong, Tadayoshi Kohno, and Jamie Morgenstern. 2023. Evaluation of targeted dataset collection on racial equity in face recognition. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society (Montréal, QC, Canada) (AIES ’23). Association for Computing Mac...

  55. [63]

    Noura Howell, Audrey Desjardins, and Sarah Fox. 2021. Cracks in the success narrative: Rethinking failure in design research through a retrospective trioethnography. ACM Transactions on Computer-Human Interaction (TOCHI) 28, 6 (2021), 1–31

  56. [64]

    Lilly C Irani and M Six Silberman. 2016. Stories We Tell About Labor: Turkopticon and the Trouble with" Design". In Proceedings of the 2016 CHI conference on human factors in computing systems . 4573–4586

  57. [65]

    Azra Ismail and Neha Kumar. 2018. Engaging Solidarity in Data Collection Practices for Community Health. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (Nov. 2018), 1–24. doi:10.1145/3274345

  58. [66]

    Akshita Jha, Vinodkumar Prabhakaran, Remi Denton, Sarah Laszlo, Shachi Dave, Rida Qadri, Chandan Reddy, and Sunipa Dev. 2024. ViSAGe: A Global-Scale Analysis of Visual Stereotypes in Text-to-Image Generation. InProceedings of the 62nd Annual Meeting of the Association for Comp...

  59. [67]

    Phil Jones. 2021. Refugees help power machine learning advances at Microsoft, Facebook, and Amazon. Rest of World 22 (2021)

  60. [68]

    Shivani Kapania, Alex S Taylor, and Ding Wang. 2023. A Hunt for the Snark: Annotator Diversity in Data Practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . ACM, Hamburg Germany, 1–15. doi:10.1145/3544548.3580645

  61. [69]

    Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. 2023. Taxonomizing and Measuring Representational Harms: A Look at Image Tagging. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 14277–14285

  62. [70]

    Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, et al. 2024. The PRISM Alignment Project: What Participatory, Representative Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, ...

  63. [71]

    Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton

    Aniket Kittur, Jeffrey V. Nickerson, Michael Bernstein, Elizabeth Gerber, Aaron Shaw, John Zimmerman, Matt Lease, and John Horton. 2013. The Future of Crowd Work. In Proceedings of the 2013 Conference on Computer Supported Cooperative Work (San Antonio, Texas, USA) (CSCW ’13)....

  64. [72]

    Christopher A Le Dantec and Sarah Fox. 2015. Strangers at the Gate: Gaining Access, Building Rapport, and Co- Constructing Community-Based Research. InProceedings of the 18th ACM conference on computer supported cooperative work & social computing . 1348–1358

  65. [73]

    Shayne Longpre, Robert Mahari, Anthony Chen, Naana Obeng-Marnu, Damien Sileo, William Brannon, Niklas Muennighoff, Nathan Khazam, Jad Kabbara, Kartik Perisetla, et al. 2024. A large-scale audit of dataset licensing and attribution in AI. Nature Machine Intelligence 6, 8 (2024)...

  66. [74]

    Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2024. Stable Bias: Evaluating Societal Representations in Diffusion Models. Advances in Neural Information Processing Systems 36 (2024)

  67. [75]

    Kelly Mack, Maitraye Das, Dhruv Jain, Danielle Bragg, John Tang, Andrew Begel, Erin Beneteau, Josh Urban Davis, Abraham Glasser, Joon Sung Park, and Venkatesh Potluri. 2021. Mixed Abilities and Varied Experiences: a group autoethnography of a virtual summer internship. In Proc...

  68. [76]

    Douglas MacMillan. 2023. Cameras, facial recognition watch over public housing. The Washington Post. https: //www.washingtonpost.com/business/2023/05/16/surveillance-cameras-public-housing/

  69. [77]

    Jabez Magomere, Shu Ishida, Tejumade Afonja, Aya Salama, Daniel Kochin, Yuehgoh Foutse, Imane Hamzaoui, Raesetje Sefala, Aisha Alaagib, Samantha Dalal, et al. 2025. The World Wide recipe: A community-centred framework for fine-grained data collection and regional bias operatio...

  70. [78]

    Daniela Massiceti, Luisa Zintgraf, John Bronskill, Lida Theodorou, Matthew Tobias Harris, Edward Cutrell, Cecily Morrison, Katja Hofmann, and Simone Stumpf. 2021. ORBIT: A real-world few-shot dataset for teachable object recognition. In Proceedings of the IEEE/CVF Internationa...

  71. [79]

    David W McMillan and David M Chavis. 1986. Sense of Community: A Definition and Theory. Journal of community psychology 14, 1 (1986), 6–23

  72. [80]

    Milagros Miceli and Julian Posada. 2022. The Data-Production Dispositif. Proc. ACM Hum.-Comput. Interact. 6, CSCW2, Article 460 (Nov. 2022), 37 pages. doi:10.1145/3555561

  73. [81]

    Milagros Miceli, Julian Posada, and Tianling Yang. 2022. Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power? Proc. ACM Hum.-Comput. Interact. 6, GROUP, Article 34 (Jan. 2022), 14 pages. doi:10.1145/ 3492853

  74. [82]

    Pine, Trine Rask Nielsen, and Gina Neff

    Naja Holten Møller, Claus Bossen, Kathleen H. Pine, Trine Rask Nielsen, and Gina Neff. 2020. Who Does the Work of Data? Interactions 27, 3 (April 2020), 52–55. doi:10.1145/3386389

  75. [83]

    Esther Mwema and Abeba Birhane. 2025. Undersea cables in Africa: The new frontiers of digital colonialism. First Monday 29, 4 (2025). doi:10.5210/fm.v29i4.13637 [Accessed 13 January 2025]

  76. [84]

    Safiya Umoja Noble. 2018. Algorithms of Oppression: How Search Engines Reinforce Racism . New York University Press, New York, USA. doi:doi:10.18574/nyu/9781479833641.001.0001

  77. [85]

    Victor Ojewale, Ryan Steed, Briana Vecchione, Abeba Birhane, and Inioluwa Deborah Raji. 2024. Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling. arXiv:2402.17861 (March 2024). arXiv:2402.17861 [cs]

  78. [86]

    Orlikowski, and Joanne Yates

    Kazuo Okamura, Masayo Fujimoto, Wanda J. Orlikowski, and Joanne Yates. 1995. Helping CSCW Applications Succeed: The Role of Mediators in the Context of Use. The Information Society 11, 3 (July 1995), 157–172. doi:10.1080/ 01972243.1995.9960190

  79. [87]

    Jennifer Pierre, Roderic Crooks, Morgan Currie, Britt Paris, and Irene Pasquetto. 2021. Getting Ourselves Together: Data-centered participatory design research & epistemic burden. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–11

  80. [88]

    Trevor J Pinch and Wiebe E Bijker. 1984. The Social Construction of Facts and Artefacts: Or How the Sociology of Science and the Sociology of Technology Might Benefit Each Other. Social studies of science 14, 3 (1984), 399–441

  81. [89]

    The Good Robot Podcast. 2025. Margaret Mitchell on Large Language Models and Misogyny in Tech. https://podcasts. apple.com/fr/podcast/margaret-mitchell-on-large-language-models-and/id1570237963?i=1000569683327 [Accessed 13 January 2025]

  82. [90]

    Julian Alberto Posada Gutierrez. 2022. The Coloniality of Data Work: Power and Inequality in Outsourced Data Production for Machine Learning . Ph. D. Dissertation. University of Toronto. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW492. Publication date: Novemb...

  83. [91]

    Vinodkumar Prabhakaran, Rida Qadri, and Ben Hutchinson. 2022. Cultural Incongruencies in Artificial Intelligence. arXiv preprint arXiv:2211.13069 (2022)

  84. [92]

    Puri and Sundeep Sahay

    Satish K. Puri and Sundeep Sahay. 2007. Role of ICTs in participatory development: An Indian experience.Information Technology for Development 13, 2 (2007), 133–160. doi:10.1002/itdj.20058

  85. [93]

    Rida Qadri, Mark Diaz, Ding Wang, and Michael Madaio. 2025. The Case for ‘Thick Evaluations’ of Cultural Representation in AI. arXiv preprint arXiv:2503.19075 (2025)

  86. [94]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al . 2021. Learning Transferable Visual Models From Natural Language Supervision. In International conference on machine learn...

  87. [95]

    Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron B

    Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao, Aaron B. Adcock, Laurens van der Maaten, Deepti Ghadi- yaram, and Olga Russakovsky. 2023. GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition. arXiv:2301.02560 [cs.CV] https://arxiv.org/abs/2301.02560

  88. [96]

    Peter Reason and Hilary Bradbury. 2008. The SAGE Handbook of Action Research . SAGE Publications Ltd, 1 Oliver’s Yard, 55 City Road, London England EC1Y 1SP United Kingdom. doi:10.4135/9781848607934

  89. [97]

    William A Gaviria Rojas, Sudnya Diamos, Keertan Ranjan Kini, David Kanter, Vijay Janapa Reddi, and Cody Coleman

  90. [98]

    Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. Gender Bias in Coreference Resolution. arXiv preprint arXiv:1804.09301 (2018)

  91. [99]

    David Romero, Chenyang Lyu, Haryo Akbarianto Wibowo, Teresa Lynn, Injy Hamed, Aditya Nanda Kishore, Aishik Mandal, Alina Dragonetti, Artem Abzaliev, Atnafu Lambebo Tonja, et al. 2024. CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark. arXiv preprint arX...

  92. [100]

    Brubaker

    Morgan Klaus Scheuerman and Jed R. Brubaker. 2024. Products of Positionality: How Tech Workers Shape Identity Concepts in Computer Vision. In Proceedings of the CHI Conference on Human Factors in Computing Systems . ACM, Honolulu HI USA, 1–18. doi:10.1145/3613904.3641890

  93. [101]

    Everyone wants to do the model work, not the data work

    Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen Paritosh, and Lora M Aroyo. 2021. “Everyone wants to do the model work, not the data work”: Data Cascades in High-Stakes AI. In proceedings of the 2021 CHI Conference on Human Factors in Computing Syste...

  94. [102]

    Brubaker

    Morgan Klaus Scheuerman, Jacob M Paul, and Jed R. Brubaker. 2019. How Computers See Gender: An Evaluation of Gender Classification in Commercial Facial Analysis and Image Labeling Services. Proc. ACM Hum.-Comput. Interact. 3, CSCW, Article 144 (2019), 33 pages. doi:10.1145/3359246

  95. [103]

    Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021. Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset Development. Proceedings of the ACM on Human-Computer Interaction 5, CSCW2 (Oct. 2021), 1–37. doi:10.1145/3476058

  96. [104]

    Shreya Shankar, Yoni Halpern, Eric Breck, James Atwood, Jimbo Wilson, and D Sculley. 2017. No Classification without Representation: Assessing Geodiversity Issues in Open Data Sets for the Developing World. arXiv preprint arXiv:1711.08536 (2017)

  97. [105]

    Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al. 2022. LAION-5B: An open large-scale dataset for training next generation image-text models.Advances in Neural In...

  98. [106]

    Susan Leigh Star. 1990. Power, technology and the phenomenology of conventions: on being allergic to onions. The Sociological Review 38, 1_suppl (1990), 26–56

  99. [107]

    Shivalika Singh, Freddie Vargus, Daniel D’souza, Börje F Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura O’Mahony, et al. 2024. Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning. arXiv preprint arXiv:24...

  100. [108]

    Cella M Sum, Anh-Ton Tran, Jessica Lin, Rachel Kuo, Cynthia L Bennett, Christina Harrington, and Sarah E Fox. 2023. Translation as (Re) mediation: How Ethnic Community-Based Organizations Negotiate Legitimacy. In Proceedings of the 2023 CHI Conference on Human Factors in Compu...

  101. [109]

    Susan Leigh Star and Anselm Strauss. 1999. Layers of Silence, Arenas of Voice: The Ecology of Visible and Invisible Work. Computer supported cooperative work (CSCW) 8 (1999), 9–30

  102. [110]

    Latanya Sweeney. 2013. Discrimination in Online Ad Delivery. Commun. ACM 56, 5 (2013), 44–54

  103. [111]

    Harini Suresh, Emily Tseng, Meg Young, Mary Gray, Emma Pierson, and Karen Levy. 2024. Participation in the Age of Foundation Models. In The 2024 ACM Conference on Fairness, Accountability, and Transparency . 1609–1621

  104. [112]

    UNESCO. 2003. Convention for the Safeguarding of the Intangible Cultural Heritage. https://ich.unesco.org/ en/convention Adopted by the General Conference of the United Nations Educational, Scientific and Cultural Organization, 32nd Session, Paris, 17 October 2003

  105. [113]

    Atnafu Lambebo Tonja, Bonaventure FP Dossou, Jessica Ojo, Jenalea Rajab, Fadel Thior, Eric Peter Wairagala, Aremu Anuoluwapo, Pelonomi Moiloa, Jade Abbott, Vukosi Marivate, et al. 2024. InkubaLM: A small language model for low-resource African languages. arXiv preprint arXiv:2...

  106. [114]

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, et al. 2023. Sociotechnical Safety Evaluation of Generative AI Systems. arXiv preprint arXiv:2310.11986 (2023)

  107. [115]

    Elizabeth Anne Watkins. 2023. Face Work: A Human-Centered Investigation into Facial Verification in Gig Work. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023), 1–24

  108. [116]

    Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Yutong Wang, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, et al. 2024. WorldCuisines: A Massive- Scale Benchmark for Multilingual and Multicultural Visual...

  109. [117]

    Joseph West, Meagan Kaufenberg-Lashua, Jamie Kelley, and Valeria Stepanova. 2024. A Field-specific Analysis of Gender and Racial Biases in Generative AI. (2024)

  110. [118]

    Meg Young, Uphol Ehsan, Ranjit Singh, Emnet Tafesse, Michele Gilman, Christina Harrington, and Jacob Metcalf

  111. [119]

    Jordan Wirfs-Brock, Alli Fam, Laura Devendorf, and Brian Keegan. 2021. Examining Narrative Sonification: Using First-Person Retrospection Methods to Translate Radio Production to Interaction Design. ACM Transactions on Computer-Human Interaction (TOCHI) 28, 6 (2021), 1–34

  112. [122]

    About You

    Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng. 2023. Deep Long-Tailed Learning: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10795–10816. Proc. ACM Hum.-Comput. Interact., Vol. 9, No. 7, Article CSCW492. Publicat...

  113. [2022]

    In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track

    The Dollar Street Dataset: Images Representing the Geographic and Socioeconomic Diversity of the World. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track . https://openreview. net/forum?id=qnfYsave0U4

  114. [2024]

    First Monday (2024)

    Participation versus Scale: Tensions in the Practical Demands on Participatory AI. First Monday (2024)

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.