Pith. sign in

REVIEW 3 major objections 4 minor 1 cited by

Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Automated moderation failures in low-resource languages stem from colonial-era inequities, the paper argues, not just missing data.

desk verdict A genuinely useful interview study of moderation barriers in four low-resource languages, wrapped in a causal claim about colonial suppression the data don't actually demonstrate. read the letter →

arxiv 2501.13836 v3 pith:GP3TJKOF submitted 2025-01-23 cs.CL cs.HC

classification cs.CLcs.HC
keywords automatedcontentmoderationlow-resourcelanguagescolonialitydatascarcitycode-mixingagglutinativemorphologyGlobalSouthqualitativeinterviews
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that automated moderation fails for Tamil, Swahili, Maghrebi Arabic, and Quechua not primarily because these languages lack data, but because the entire pipeline—data access, annotation, preprocessing, and model training—is built around English and Western market priorities. Based on semi-structured interviews with 22 AI experts working on harmful-content detection in these languages, it identifies how tech companies' monopoly on user data, weak financial incentives for Global South markets, and reliance on biased machine translation and outdated corpora reproduce historical colonial hierarchies. The paper treats these failures as structural rather than technical, and proposes multi-stakeholder fixes: funding local research capacity, democratizing data access, and adopting language-aware methods such as morphological tokenization. A sympathetic reader would take the central claim as a reframing of 'data scarcity' from a neutral technical condition into a consequence of political and economic choices.

What carries the argument

The carrying object is the automated moderation pipeline, broken into four stages: data curation, annotation, preprocessing, and model training. The load-bearing linguistic mechanism is agglutinative morphology combined with code-mixing: Tamil, Swahili, Maghrebi Arabic, and Quechua build thousands of words from a single root, so frequency-based tokenizers, stemming, and normalization designed for English systematically mangle the forms that carry harmful meaning. The interpretive machinery is the coloniality lens, which converts observed pipeline failures into symptoms of persistent power asymmetries—data monopolies, profit-driven neglect of Global South markets, and English-centric model design—rather than neutral technical gaps.

What would settle it

A controlled benchmark would settle the technical core: if, at matched data sizes, a linguistically motivated morphological tokenizer does not improve harmful-content detection over a frequency-based tokenizer for Tamil, Swahili, Maghrebi Arabic, or Quechua, the claim that English-centric preprocessing drives moderation failure in agglutinative languages would be undermined. Separately, if platforms granted vetted researchers full data access in one low-resource language and moderation accuracy did not improve, the data-monopoly link would weaken.

Watch

Extended reading notes

Core claim

The central discovery is that moderation failures in low-resource languages persist even when data volume is not the binding constraint; the binding constraints are who controls the data, who defines harm, and which linguistic features the tools are designed to see. The paper documents how English-centric frequency-based tokenizers split agglutinative words incorrectly—for example, the Tamil word Mulaicchu (meaning 'nipples') can be reduced to Mulai (meaning 'sprout')—so that sexually harassing language evades detection; how language-identification tools mangle code-mixed text; and how toxicity models trained on Western media associate Arabic phrases like 'Allahu Akbar' with terrorism. The authors then read these findings through coloniality: data monopolies, reliance on colonial-era texts for Quechua, underfunded annotation, and the flattening of annotator diversity into a single label are continuous with colonial suppression of non-Western languages. The paper's conclusion is that techno-solutionist fixes that only add more data will not yield equitable moderation; the pipeline itself must be redesigned around local linguistic knowledge and community self-determination.

Load-bearing premise

The central claim rests on treating the self-reported experiences of 22 purposively sampled AI experts—only three of them Quechua specialists and many based in Western institutions—as reliable evidence for how entire moderation pipelines behave across four language communities, and for the causal interpretation that these failures are rooted in colonial suppression.

Editorial extensions

If this is right

  • If the diagnosis is right, adding more labeled data alone will not fix moderation for these languages; data access, annotation incentives, and model architecture must change together.
  • Language-aware preprocessing—morphological segmenters, rule-based translation, and code-mixed language identification—should outperform generic frequency-based tokenizers and multilingual models on harmful-content detection for agglutinative languages.
  • Tech companies' API restrictions and shutdown of public research tools will continue to block academic study of evolving hate speech in the Global South unless policy such as the Digital Services Act grants vetted researchers data access.
  • Regulatory pressure requiring local moderators, locally defined benchmarks, and reporting of recall and precision for low-resource languages would surface failures that accuracy metrics hide.
  • Investment in grassroots research ecosystems and equitable data-sharing could transfer capacity to local researchers and reduce dependence on Western computational resources.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the paper's diagnosis is right, the same pipeline critique should apply to other morphologically rich low-resource languages beyond the four studied, and would predict that generic multilingual models underperform linguistically motivated tools there too.
  • The paper's account implies a testable asymmetry: for low-resource languages with simple morphology, data scarcity should dominate failure, whereas for agglutinative languages, architecture mismatches should dominate even at matched data sizes.
  • A concrete policy experiment follows: if platforms opened vetted data access to Global South researchers, hate-speech detection corpora and models should improve faster for those languages than equivalent private investment in more data alone would achieve.
  • A companion large-scale audit of platform transparency reports, comparing moderation accuracy across languages and regions, could test whether the structural inequities reported by the 22 experts hold beyond their experiences.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper reports a qualitative interview study of 22 AI researchers and practitioners who work on harmful content detection in four low-resource languages: Tamil, Swahili, Maghrebi Arabic, and Quechua. Using semi-structured interviews and reflexive thematic analysis, the authors document barriers across data curation, annotation, preprocessing, and model training, and argue that these barriers are not reducible to data scarcity. They interpret the findings through a coloniality lens, claiming that moderation failures are rooted in structural inequities and colonial suppression of non-Western languages, and they propose multi-stakeholder recommendations for improving moderation. The paper contributes an empirical account of systemic pipeline-level challenges and connects them to decolonial theory.

Significance. If the findings hold, the paper makes a useful contribution by moving the discussion of low-resource content moderation beyond a narrow data-scarcity framing. Its strengths include rich interview material with concrete examples across four typologically diverse languages, a reflexive positionality statement, and an explicit attempt to connect technical pipeline decisions to historical and political contexts. The recommendations in Section 5.2 are concrete and grounded in the participants' accounts. However, the central causal claim that the problems are 'rooted in colonial suppression' is asserted more strongly than the empirical evidence supports, and the sample is small and uneven across the four languages. The paper is publishable after the framing and evidentiary status of that causal claim are clarified.

major comments (3)
  1. [Abstract and §5.1] The abstract and Section 5.1 claim that moderation failures are 'rooted in colonial suppression of non-Western languages,' but the interview evidence in Section 4 documents contemporary structural barriers without any participant attributing those barriers to colonial history, and the coloniality framework is introduced as the authors' analytic lens in Section 5.1 rather than as a theme that emerged from the data. This is load-bearing because it is the paper's central contribution. Please either reframe the causal claim as an interpretive argument grounded in the cited decolonial literature, or provide explicit evidence linking participants' accounts to colonial legacies.
  2. [§3 (Participants) and §5.1] The systemic conclusions are drawn across four language communities from a sample that includes only three Quechua experts, with half of the participants affiliated with Western institutions, and the paper does not discuss how this imbalance constrains the cross-language claims. The Quechua-related findings in Sections 4.1 and 4.3 rest on very few voices. Please either narrow the scope of the cross-language conclusions, provide a per-language saturation assessment, or explicitly analyze how the uneven sample limits the generality of the claims.
  3. [§4.3 and Abstract] The 'even if more data were available' counterfactual, which is central to the paper's 'beyond data scarcity' thesis, is supported only by isolated examples such as the Tamil stemming error Mulaicchu→Mulai and by participants' beliefs, not by any controlled comparison of model performance at matched data sizes across languages. Please present this counterfactual explicitly as a hypothesis or as participant perception, or add comparative evidence, so that readers can distinguish observed barriers from the authors' interpretation.
minor comments (4)
  1. [§3 (Data Collection)] The paper does not include the interview protocol or a description of how saturation was determined; adding these would strengthen reproducibility and trustworthiness.
  2. [Table 2] Participants P7 and P8 are listed as specializing in 'Indic languages' rather than one of the four focal languages, which is inconsistent with the claim that all 22 participants specialize in Tamil, Swahili, Maghrebi Arabic, or Quechua; please clarify.
  3. [Abstract and §4.3] The Tamil example is spelled 'Mualichhu' in the abstract and 'Mulaicchu' in Section 4.3; please make the transliteration consistent throughout.
  4. [§3 (Data Collection)] Interviews were conducted in English even for participants whose native language is one of the focal languages; this should be acknowledged as a potential limitation for capturing in-group linguistic and cultural concepts.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the coloniality conclusion is an interpretive frame applied to independent interview evidence, not an input fitted into the output.

full rationale

This is a qualitative interview study rather than a formal derivation, so the circularity burden is low and no equation-level reduction exists. The paper's central empirical claims—data-access restrictions, underfunded annotation, tokenizer and stemming failures, and English-centric model design—are grounded in direct participant quotes in Section 4 (e.g., P14 on Swahili hate-speech evolution, P6 on Tamil overstemming of Mulaicchu to Mulai-, P9 on Quechua tokenization). The coloniality framing is explicitly introduced as an analytic lens in Section 2.2 and applied in Section 5.1, with supporting citations to Quijano, Kwet, Said, and others; it is not defined in terms of the outcome it is said to explain. The authors do cite their own prior work—Shahid and Vashistha (2023), Elswah (2024a), and Elswah (2024b)—but these citations are used as background evidence in Related Work and as corroborating references in the Discussion, not as the sole load-bearing justification for the central claim. The stronger assertion that moderation failures are 'rooted in colonial suppression' is an authorial interpretation that goes beyond what the interview data alone entail, and the 'even if more data were available' counterfactual is asserted rather than experimentally tested; however, interpretive overreach and untested counterfactuals are evidentiary weaknesses, not circularity. The paper even includes a positionality statement acknowledging the authors' situated perspective and the 'partial perspective' of participants, which further indicates that the coloniality reading is presented as an interpretive contribution rather than a self-justifying derivation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests on qualitative assumptions rather than mathematical postulates. No free parameters are fitted. The main axioms are the validity of self-reported interview data, the representativeness of the four selected languages, and the appropriateness of coloniality as an explanatory lens.

assumptions (3)
  • domain assumption Self-reports of 22 AI experts accurately reflect systemic barriers in moderation pipelines.
    The study relies on participants' accounts of company behavior, model failures, and data access without independent verification. This is standard for qualitative research but is a load-bearing premise.
  • domain assumption Tamil, Swahili, Maghrebi Arabic, and Quechua represent the diversity of low-resource languages in the Global South.
    The paper generalizes from four languages to the Global South; the selection is purposive and covers different families and regions, but representativeness is assumed.
  • domain assumption Coloniality is a valid explanatory frame for interpreting technical design choices and resource allocation.
    The paper applies a decolonial theoretical lens from prior literature (Quijano, Kwet, etc.) to interpret findings; the causal force of colonial history is assumed, not measured.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages." pith.science (2026). https://pith.science/paper/GP3TJKOF

@misc{pith2026250113836,
  author       = {Pith},
  title        = {Pith review of: Think Outside the Data: Colonial Biases and Systemic Issues in Automated Moderation Pipelines for Low-Resource Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GP3TJKOF}},
  note         = {Machine review of arXiv:2501.13836}
}
read the original abstract

Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews with 22 AI experts working on harmful content detection in four low-resource languages: Tamil (South Asia), Swahili (East Africa), Maghrebi Arabic (North Africa), and Quechua (South America)--we examine systemic issues in building automated moderation tools for these languages. Our findings reveal that beyond data scarcity, socio-political factors such as tech companies' monopoly on user data and lack of investment in moderation for low-profit Global South markets exacerbate historic inequities. Even if more data were available, the English-centric and data-intensive design of language models and preprocessing techniques overlooks the need to design for morphologically complex, linguistically diverse, and code-mixed languages. We argue these limitations are not just technical gaps caused by "data scarcity" but reflect structural inequities, rooted in colonial suppression of non-Western languages. We discuss multi-stakeholder approaches to strengthen local research capacity, democratize data access, and support language-aware solutions to improve automated moderation for low-resource languages.

Figures

Figures reproduced from arXiv: 2501.13836 by the authors.

Figure 1
Figure 1. Issues affecting different stages of automated moderation pipeline for low-resource languages. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. It is not enough to give your moderation rules to ChatGPT: Policy-as-Prompt Moderation and Its Potential Impacts on Community Governance

    cs.CY 2026-07 unverdicted novelty 5.0 of 10

    Writing a moderation policy as an LLM prompt cannot by itself ensure meaningful community governance.

Reference graph

Works this paper leans on

117 extracted references · 77 canonical work pages · cited by 1 Pith paper

  1. [1]

    Abdelali, A.; Hassan, S.; Mubarak, H.; Darwish, K.; and Samih, Y. 2021. Pre-training bert on arabic tweets: Practical considerations

  2. [2]

    Ahmad, S.; and Krzywdzinski, M. 2022. Moderating in obscurity: How Indian content moderators work in global content moderation value chains, chapter 5, 77--95. Cambridge, MA: The MIT Press

  3. [3]

    Ali, S. M. 2016. A brief introduction to decolonial computing. XRDS: Crossroads, The ACM Magazine for Students, 22(4): 16--21

  4. [4]

    Alimardani, M.; and Elswah, M. 2021. Digital orientalism:\# SaveSheikhJarrah and Arabic content moderation

  5. [5]

    Anderson, B. 2020. Imagined communities: Reflections on the origin and spread of nationalism. In The new social theory reader, 282--288. United Kingdom: Routledge

  6. [6]

    Arnett, C.; and Bergen, B. K. 2024. Why do language models perform worse for morphologically complex languages?

  7. [7]

    Arney, J. 2024. Data dump: Meta killed CrowdTangle. What does it mean for researchers, reporters?

  8. [8]

    Bank, W. 2014. Discriminated against for speaking their own language

Show all 117 references
  1. [9]

    Bellan, R. 2024. Meta axed CrowdTangle, a tool for tracking disinformation. Critics claim its replacement has just `1\

  2. [10]

    Bender, E. M. 2009. Linguistically Na \"i ve != Language Independent: Why NLP Needs Linguistic Typology. In Proceedings of the EACL 2009 Workshop on the Interaction between Linguistics and Computational Linguistics: Virtuous, Vicious or Vacuous? , 26--32. Athens, Greece: Assoc...

  3. [11]

    Benjamin, R. 2023. Race after technology. In Social Theory Re-Wired, 405--415. New York: Routledge

  4. [12]

    Bhabha, H. K. 2011. Our neighbours, ourselves: Contemporary reflections on survival. De Gruyter

  5. [13]

    Bhattacharyya, G. 2018. Rethinking racial capitalism: Questions of reproduction and survival. Maryland, USA: Rowman & Littlefield

  6. [14]

    Biddle, S. 2022. Facebook's Tamil Censorship Highlights Risks to Everyone

  7. [15]

    Bird, S. 2020. Decolonising speech and language technology. In 28th International Conference on Computational Linguistics, COLING 2020, 3504--3519. online: Association for Computational Linguistics (ACL)

  8. [16]

    Bird, S. 2022. Local languages, third spaces, and other high-resource scenarios. In 60th Annual Meeting of the Association for Computational Linguistics, ACL 2022, 7817--7829. Dublin: Association for Computational Linguistics (ACL)

  9. [17]

    Braun, V.; and Clarke, V. 2006. Using thematic analysis in psychology. Qualitative research in psychology, 3(2): 77

  10. [18]

    Centre, B. . H. R. R. 2024. Dismantling the facade: A global south perspective on the state of engagement with tech companies

  11. [19]

    A.; Arnett, C.; Tu, Z.; and Bergen, B

    Chang, T. A.; Arnett, C.; Tu, Z.; and Bergen, B. K. 2023. When is multilinguality a curse? language modeling for 250 high-and low-resource languages

  12. [20]

    Christodouloupoulos, C.; and Steedman, M. 2015. A massively parallel corpus: the bible in 100 languages. Language resources and evaluation, 49: 375--395

  13. [21]

    Coleman, D. 2018. Digital colonialism: The 21st century scramble for Africa through the extraction and control of user data and the limitations of data protection laws. Michigan Journal of Race and Law, 24: 417--439

  14. [22]

    Commission, E. 2025. Commission adopts delegated act on data access under the Digital Services Act

  15. [23]

    Couldry, N.; and Mejias, U. A. 2019. Data colonialism: Rethinking big data's relation to the contemporary subject. Television & New Media, 20(4): 336--349

  16. [24]

    R.; and Semaan, B

    Das, D.; Guha, S.; Brubaker, J. R.; and Semaan, B. 2024. The ``Colonial Impulse" of Natural Language Processing: An Audit of Bengali Sentiment Analysis Tools and Their Identity-based Biases. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. New Y...

  17. [25]

    Das, M.; Saha, P.; Mathew, B.; and Mukherjee, A. 2022. Hatecheckhin: Evaluating hindi hate speech detection models

  18. [26]

    De Gregorio, G.; and Stremlau, N. 2023. Inequalities and content moderation. Global Policy, 14(5): 870--879

  19. [27]

    S.; Kannimuthu, S.; and Madasamy, A

    Devi, V. S.; Kannimuthu, S.; and Madasamy, A. K. 2024. The Effect of Phrase Vector Embedding in Explainable Hierarchical Attention-Based Tamil Code-Mixed Hate Speech and Intent Detection. IEEE Access, 12(0): 11316--11329

  20. [28]

    Divon, T.; and Ong, J. C. 2025. Tech Bro Power Play: Zuckerberg vs. Global Tech Justice

  21. [29]

    Dourish, P.; and Mainwaring, S. D. 2012. Ubicomp's colonial impulse. In Proceedings of the 2012 ACM conference on ubiquitous computing, 133--142. New York, USA: ACM

  22. [30]

    W.; and Cabato, R

    Elizabeth Dwoskin, J. W.; and Cabato, R. 2019. Content moderators at YouTube, Facebook and Twitter see the worst of the web — and suffer silently

  23. [31]

    Elswah, M. 2024 a . Moderating Kiswahili Content on Social Media

  24. [32]

    Elswah, M. 2024 b . Moderating Maghrebi Arabic Content on Social Media

  25. [33]

    Errington, J. 2007. Linguistics in a colonial world: A story of language, meaning, and power. John Wiley & Sons

  26. [34]

    Fanon, F. 2023. Black skin, white masks. In Social theory re-wired, 355--361. United Kingdom: Routledge

  27. [35]

    Fine, M. 1994. Working the hyphens. Handbook of qualitative research, 2

  28. [36]

    Fishman, J. A. 1989. Language and ethnicity in minority sociolinguistic perspective. United Kingdom: Multilingual Matters

  29. [37]

    L.; and Bowen, C

    Garfinkel, S. L.; and Bowen, C. M. 2022. Preserving Privacy While Sharing Data

  30. [38]

    Garimella, K.; and Chauchard, S. 2024. WhatsApp Explorer: A Data Donation Tool To Facilitate Research on WhatsApp

  31. [39]

    S.; Yu, K.; Yang, Y.; Dai, M.; Qiu, J.; Tang, R.; and Huang, J

    Geiger, R. S.; Yu, K.; Yang, Y.; Dai, M.; Qiu, J.; Tang, R.; and Huang, J. 2020. Garbage in, garbage out? do machine learning application papers in social computing report where human-labeled training data comes from? In Proceedings of the 2020 Conference on Fairness, Accounta...

  32. [40]

    Ghosh, S.; and Caliskan, A. 2023. ChatGPT Perpetuates Gender Bias in Machine Translation and Ignores Non-Gendered Pronouns: Findings across Bengali and Five other Low-Resource Languages. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES '23, 901–9...

  33. [41]

    Gorwa, R. 2019. What is platform governance? Information, communication & society, 22(6): 854--871

  34. [42]

    Gramsci, A. 2020. Selections from the prison notebooks. In The applied theatre reader, 141--142. New York, USA: Routledge

  35. [43]

    Gupta, S. 2024. The AI arms race: Which LLMs are winning the enterprise battlefield?

  36. [44]

    Haraway, D. 2013. Situated knowledges: The science question in feminism and the privilege of partial perspective 1. Women, science, and technology, 1: 455--472

  37. [45]

    Held, W.; Harris, C.; Best, M.; and Yang, D. 2023. A material lens on coloniality in nlp

  38. [46]

    Heller, M.; and McElhinny, B. 2017. Language, capitalism, colonialism: Toward a critical history. Canada: University of Toronto Press

  39. [47]

    IDRC. 2024. Artificial Intelligence for Development

  40. [48]

    Irani, L.; Vertesi, J.; Dourish, P.; Philip, K.; and Grinter, R. E. 2010. Postcolonial computing: a lens on design and development. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '10, 1311–1320. New York, NY, USA: ACM

  41. [49]

    Iyer, P. 2025. What a New Study Reveals About Content Moderation in Tigray

  42. [50]

    Jim\' e nez, J. 2024. Worried About Meta Using Your Instagram to Train Its A.I.? Here's What to Know

  43. [51]

    Kak, A. 2020. ``The Global South is everywhere, but also always somewhere'': National Policy Narratives and AI Justice. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, AIES '20, 307–312. New York, NY, USA: ACM

  44. [52]

    Kapelke, C. 2020. Using differential privacy to harness big data and preserve privacy

  45. [53]

    M.; and Campos, D

    Kennedy, W. M.; and Campos, D. V. 2025. Vernacularizing Taxonomies of Harm is Essential for Operationalizing Holistic AI Safety. In Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society, AIES '25, 698–710. New York, NY, USA: ACM

  46. [54]

    S.; Yadav, K.; Alsharabi, N.; and Ahmad, A

    Khan, M.; Ullah, K.; Alharbi, Y.; Alferaidi, A.; Alharbi, T. S.; Yadav, K.; Alsharabi, N.; and Ahmad, A. 2023. Understanding the research challenges in low-resource language and linking bilingual news articles in multilingual news archive. Applied Sciences, 13(15): 8566

  47. [55]

    Kolli, V. 2024. Linguistic Colonialism: Moroccan Education and its Dark Past

  48. [56]

    Kreutzer, J.; Caswell, I.; Wang, L.; Wahab, A.; van Esch, D.; Ulzii-Orshikh, N.; Tapo, A.; Subramani, N.; Sokolov, A.; Sikasote, C.; et al. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, ...

  49. [57]

    Kwet, M. 2019. Digital colonialism: US empire and the new imperialism in the Global South. Race & Class, 60(4): 3--26

  50. [58]

    Legon, A.; and Alsalman, A. 2020. How Facebook can Flatten the Curve of the Coronavirus Infodemic. Technical report, Avaaz

  51. [59]

    E.; Abdilla, A.; Arista, N.; Baker, K.; Benesiinaabandan, S.; Brown, M.; Cheung, M.; Coleman, M.; Cordes, A.; Davison, J.; Duncan, K.; Garzon, S.; Harrell, D

    Lewis, J. E.; Abdilla, A.; Arista, N.; Baker, K.; Benesiinaabandan, S.; Brown, M.; Cheung, M.; Coleman, M.; Cordes, A.; Davison, J.; Duncan, K.; Garzon, S.; Harrell, D. F.; Jones, P.-L.; Kealiikanakaoleohaililani, K.; Kelleher, M.; Kite, S.; Lagon, O.; Leigh, J.; Levesque, M.;...

  52. [60]

    Malik, S. 2022. Global labor chains of the western AI

  53. [61]

    Mehta, I. 2023. X updates its terms to ban crawling and scraping

  54. [62]

    MeitY. 2023. The Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Rules, 2021

  55. [63]

    Milmo, D. 2021. Rohingya sue Facebook for 150bn over Myanmar genocide

  56. [64]

    Mohamed, S.; Png, M.-T.; and Isaac, W. 2020. Decolonial AI: Decolonial theory as sociotechnical foresight in artificial intelligence. Philosophy & Technology, 33: 659--684

  57. [65]

    Mufwene, S. S. 2004. The ecology of language evolution. United Kingdom: Cambridge University Press

  58. [66]

    Nicholas, G.; and Bhatia, A. 2023. Toward Better Automated Content Moderation in Low-Resource Languages. Journal of Online Trust and Safety, 2(1)

  59. [67]

    Nicholas, G.; and Thakur, D. 2022. Learning to Share: Lessons on Data-Sharing from Beyond Social Media

  60. [68]

    H.; and Raji, I

    Nigatu, H. H.; and Raji, I. D. 2024. ``I Searched for a Religious Song in Amharic and Got Sexual Content Instead'': Investigating Online Harm in Low-Resourced Languages on YouTube. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT '...

  61. [69]

    H.; Tonja, A

    Nigatu, H. H.; Tonja, A. L.; Rosman, B.; Solorio, T.; and Choudhury, M. 2024. The Zeno's Paradox ofLow-Resource'Languages

  62. [70]

    Obi-Young, O. 2018. Bantu's Swahili, or How to Steal a Language from Africa | Kamau Muiga

  63. [71]

    low-resource languages

    \`O g \'u nr\. e \` m \' , T.; Nekoto, W. O.; and Samuel, S. 2023. Decolonizing nlp for "low-resource languages": Applying abebe birhane's relational ethics

  64. [72]

    OpenAI. 2024. OpenAI and Reddit Partnership

  65. [73]

    Ovalle, A.; Subramonian, A.; Gautam, V.; Gee, G.; and Chang, K.-W. 2023. Factoring the Matrix of Domination: A Critical Review and Reimagination of Intersectionality in AI Fairness. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, AIES '23, 496–511. N...

  66. [74]

    M.; Serapio-Garc \'i a, G.; Taylor, A

    Parrish, A.; Prabhakaran, V.; Aroyo, L.; D \'i az, M.; Homan, C. M.; Serapio-Garc \'i a, G.; Taylor, A. S.; and Wang, D. 2024. Diversity-Aware Annotation for Conversational AI Safety. In Dinkar, T.; Attanasio, G.; Cercas Curry, A.; Konstas, I.; Hovy, D.; and Rieser, V., eds., ...

  67. [75]

    Perez, S. 2024. Reddit locks down its public data in new content policy, says use now requires a contract

  68. [76]

    Perrigo, B. 2023. The Workers Behind AI Rarely See Its Rewards. This Indian Startup Wants to Fix That

  69. [77]

    Popli, N. 2021. The 5 Most Important Revelations From the `Facebook Papers'

  70. [78]

    Posada, J. 2021. The Coloniality of Data Work in Latin America. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, AIES '21, 277–278. New York, NY, USA: ACM

  71. [79]

    Quijano, A. 2000. Coloniality of power and Eurocentrism in Latin America. International sociology, 15(2): 215--232

  72. [80]

    Quijano, A. 2007 a . Coloniality and modernity/rationality. Cultural studies, 21(2-3): 168--178

  73. [81]

    Quijano, A. 2007 b . Questioning ``race''. Socialism and democracy, 21(1): 45--53

  74. [82]

    Radiya-Dixit, E.; and Bogen, M. 2024. Beyond English-Centric AI Lessons on Community Participation from Non-English NLP Groups

  75. [83]

    Rananga, S.; Isong, B.; Modupe, A.; and Marivate, V. 2024. Misinformation Detection: A Review for High and Low-Resource Languages. Journal of Information Systems and Informatics, 6(4): 2892--2922

  76. [84]

    Rowe, J. 2022. Marginalised languages and the content moderation challenge

  77. [85]

    Said, E. W. 1977. Orientalism. The Georgia Review, 31(1): 162--206

  78. [86]

    Said, E. W. 2000. Out of Place—A Memoir. United Kingdom: Vintage Books

  79. [87]

    Everyone wants to do the model work, not the data work

    Sambasivan, N.; Kapania, S.; Highfill, H.; Akrong, D.; Paritosh, P.; and Aroyo, L. M. 2021. "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM

  80. [88]

    Samuels, E. 2020. How misinformation on WhatsApp led to a mob killing in India

  81. [89]

    Sap, M.; Card, D.; Gabriel, S.; Choi, Y.; and Smith, N. A. 2019. The Risk of Racial Bias in Hate Speech Detection. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 1668--1678. Florence, Italy: Association for Computational Linguistics

  82. [90]

    Scarcella, M. 2024. Elon Musk's X wins appeal to block part of California content moderation law

  83. [91]

    K.; and Brubaker, J

    Scheuerman, M. K.; and Brubaker, J. R. 2024. Products of Positionality: How Tech Workers Shape Identity Concepts in Computer Vision. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI '24. New York, NY, USA: ACM

  84. [92]

    Sch \"o pf, C. M. 2020. The Coloniality of Global Knowledge Production: Theorizing the Mechanisms of Academic Dependency. Social Transformations: Journal of the Global South, 8(2): 5--46

  85. [93]

    Schwartz, L. 2022. P rimum N on N ocere: B efore working with I ndigenous data, the ACL must confront ongoing colonialism. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 724--731. Dublin, Ireland: Associatio...

  86. [94]

    Shahid, F.; and Vashistha, A. 2023. Decolonizing Content Moderation: Does Uniform Global Community Standard Resemble Utopian Equality or Western Power Hegemony? In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. New York, USA: ACM

  87. [95]

    Siapera, E. 2022. AI Content Moderation, Racism and (de) Coloniality. International Journal of Bullying Prevention, 4(1): 55--65

  88. [96]

    Sreelekha, S.; Bhattacharyya, P.; and Malathi, D. 2018. Statistical vs. rule-based machine translation: A comparative study on indian languages. In International Conference on Intelligent Computing and Applications: ICICA 2016, 663--675. Australia: Springer

  89. [97]

    Stokel-Walker, C. 2024. Under Elon Musk, X is denying API access to academics who study misinformation

  90. [98]

    stream. 2025 a . Audio Moderation

  91. [99]

    stream. 2025 b . Video Moderation

  92. [100]

    Thakur, D. 2025. Moderating Quechua Content on Social Media. Technical report, Center for Democracy and Technology

  93. [101]

    Thiong'o, N. u. i. w. 1986. Decolonising the Mind: The Politics of Language in African Literature. East Africa: EAEP

  94. [102]

    TikTok. 2024. Supporting independent research

  95. [103]

    Udupa, S.; Maronikolakis, A.; and Wisiorek, A. 2023. Ethical scaling for content moderation: Extreme speech and the (in) significance of artificial intelligence. Big Data & Society, 10(1): 1--15

  96. [104]

    Union, U. G. 2025. Content moderators launch first-ever global alliance, demand safe working conditions and accountability from tech giants

  97. [105]

    van Esch, D.; Sarbar, E.; Lucassen, T.; O'Brien, J.; Breiner, T.; Prasad, M.; Crew, E.; Nguyen, C.; and Beaufays, F. 2019. Writing across the world's languages: Deep internationalization for Gboard, the Google keyboard

  98. [106]

    Varshney, K. R. 2024. Decolonial AI Alignment: Openness, Vi\' s esa-Dharma, and Including Excluded Knowledges. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7 of AIES '24, 1467--1481. New York, NY, USA: ACM

  99. [107]

    Verran, H.; and Christie, M. 2007. Using/designing digital technologies of representation in Aboriginal Australian knowledge practices. Human Technology, 3(2): 214--227

  100. [108]

    Villenas, S. 1996. The colonizer/colonized Chicana ethnographer: Identity, marginalization, and co-optation in the field. Harvard educational review, 66(4): 711--732

  101. [109]

    Accuracy

    Wei, J. T.-Z.; Zufall, F.; and Jia, R. 2025. Operationalizing Content Moderation "Accuracy" in the Digital Services Act. In Proceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society, AIES '25, 1527–1538. New York, NY, USA: ACM

  102. [110]

    Witness, G. 2022. Facebook unable to detect hate speech weeks away from tight Kenyan election

  103. [111]

    C.; and Ernst, J

    Wong, J. C.; and Ernst, J. 2021. Facebook knew of Honduran president’s manipulation campaign – and let it continue for 11 months

  104. [112]

    C.; and Harding, L

    Wong, J. C.; and Harding, L. 2021. `Facebook isn't interested in countries like ours': Azerbaijan troll network returns months after ban

  105. [113]

    Yibeltal, K.; and Muia, W. 2023. Facebook's algorithms `supercharged' hate speech in Ethiopia's Tigray conflict

  106. [114]

    Zevallos, R.; and Bel, N. 2023. Hints on the data for language modeling of synthetic languages with transformers. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 12508--12522. Toronto, Canada: Association for ...

  107. [115]

    Zhong, T.; Yang, Z.; Liu, Z.; Zhang, R.; Liu, Y.; Sun, H.; Pan, Y.; Li, Y.; Zhou, Y.; Jiang, H.; et al. 2024. Opportunities and challenges of large language models for low-resource languages in humanities research

  108. [116]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all...

  109. [117]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.