REVIEW 4 major objections 5 minor 53 references
Fantastic Biases (What are They) and Where to Find Them
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that bias is a deviation from a subjective norm rather than an inherent flaw, and that harmful biases in AI systems can be systematically detected and mitigated through data inspection, counterfactual testing, and…
desk verdict A readable, well-organized survey of AI bias that never quite pins down its own central definition of bias as deviation from a norm. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a general definition of bias as a deviation from a norm, together with a taxonomic 'zoology' of biases, a catalog of named biases that classifies them by their source: reporting, selection, representation, group attribution, implicit, annotation, cultural, linguistic, political, demographic, and temporal. This taxonomy carries the survey: each bias is located in a data pipeline or model component, which then determines which detection method (data inspection, counterfactual testing, per-subgroup performance analysis) and which mitigation method (sampling, weighting, augmentation, adversarial loss, pluralistic alignment) applies. The link between the mathematical bias term and cognitive bias is made concrete through the example of initializing a language model's final-layer bias with the log-unigram distribution of words, showing that a bias can encode a useful prior.
What would settle it
A concrete observation that would test the central claim is a case in which a model's output is judged unfair by every human stakeholder even though it exactly matches the chosen norm; such a case would show that fairness cannot be reduced to deviation from a norm. Alternatively, finding a reproducible bias that fits none of the paper's categories and escapes all of its listed detection methods would falsify the survey's coverage.
Extended reading notes
Core claim
The central claim is that biases are not fundamentally bad, just a deviation from a (subjective and defined) norm or value. The paper uses this definition to connect the mathematical bias term of a linear model, human cognitive biases such as the Dunning-Kruger effect, and social or cultural norms, showing that all are instances of the same phenomenon. It then argues that most harmful biases in machine learning originate from selection biases in data, annotations, and cultural or linguistic representation, and that they surface either in the data itself, in the model's robustness to counterfactual changes, or in uneven performance across demographic groups. On the mitigation side, the paper catalogs resampling, weighted loss functions, data augmentation, adversarial objectives, and alignment techniques as means to reduce harmful bias without abandoning the useful priors that bias supplies.
Load-bearing premise
The framework rests on the premise that a norm can be chosen and specified well enough to measure bias as a deviation from it, even though the paper only states that norms are subjective and context-dependent.
Editorial extensions
If this is right
- If bias is a deviation from a norm, then any debiasing effort must first make the norm explicit, turning fairness from a purely technical metric into a stated value choice.
- The taxonomy gives practitioners a shared vocabulary: naming a bias as selection or annotation bias points directly to the detection and mitigation techniques most likely to address it.
- Counterfactual testing and per-subgroup performance analysis become mandatory validation steps for any model that will be deployed on diverse populations.
- Because some biases encode useful priors, debiasing is not about removing all deviation but about removing deviations that are harmful for a chosen norm.
- Methods such as resampling, weighted loss, data augmentation, and adversarial debiasing are each suited to particular bias types, so a one-size-fits-all debiasing approach is unlikely to work.
Reading between the lines
- An implicit consequence of the paper's definition is that bias metrics are inherently normative: a model can only be judged unbiased relative to a stated norm, so audits and leaderboards should disclose which norm they assume.
- The taxonomy likely extends beyond NLP to vision-language models, which inherit the same selection and representation biases from web-scale training data; the paper's examples from image generation already hint at this.
- A testable prediction from the framework is that mitigation methods matched to the diagnosed bias type will outperform generic debiasing; a benchmarking study across the paper's bias categories could verify this.
- The definition also suggests that data feedback loops, where model outputs contaminate future training data, are a bias-amplification mechanism that fits naturally into the temporal-bias category, a connection the paper leaves implicit.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a conceptual survey that attempts to define 'bias' in a general sense, distinguish harmful from benign biases, present a taxonomy of common biases in machine learning, natural language processing, and large language models, and catalog methods for detecting and mitigating them. It argues that biases are pervasive and sometimes useful, and that negative biases arise from a deviation from a (subjective) norm or value, offering a 'zoology' of bias types (reporting, selection, representation, group attribution, implicit, cultural, linguistic, ideological, demographic, temporal, confirmation) and a set of detection and mitigation techniques (data inspection, counterfactual robustness, over/under-sampling, weighting, augmentation, adversarial loss, and human/machine perspective integration).
Significance. If its conceptual framework were rigorous, the paper would serve as an accessible introduction and reference for practitioners seeking to name and address bias in ML systems. It compiles a broad set of relevant literature and examples, and its emphasis on bias as a deviation from a chosen norm is a reasonable starting point. However, the framework is not operationalized, and the taxonomy shifts among incompatible norm types. This weakens the paper's central claim to provide a usable framework, but the survey content and reference collection still have pedagogical value. The paper does not provide machine-checked proofs, reproducible code, or falsifiable predictions; its contribution is purely conceptual and expository.
major comments (4)
- [Section 2.1] The definition of bias as 'a deviation from a (subjective and defined) norm or value' is never operationalized. The paper does not specify who chooses the norm, how it is elicited, or what distance function measures the deviation. This is load-bearing because the subsequent taxonomy and detection methods all depend on this definition; without a norm-selection rule, the framework cannot distinguish harmful bias from benign statistical variation.
- [Section 2.2] The taxonomy silently shifts among mutually incompatible norm types: reporting, selection, and automation biases are framed as deviations from truth, accuracy, or real-world prevalence; representation bias is framed as a deviation from a desired egalitarian target even when the data 'represents the reality' (the mostly-male CEO example); social and cultural biases are framed as deviations from group-dependent social norms. A single dataset can therefore be biased under one norm and unbiased under another, and the paper gives no procedure for selecting the relevant norm. This undermines the claim of providing a coherent 'zoology' of biases.
- [Section 3] The detection methods described in Sections 3.1 and 3.2 (data skewness, heterogeneous performance over target groups, counterfactual robustness) detect statistical skew, performance gaps, or decision changes, not deviation from an explicitly chosen norm. The connection between the Section 2.1 definition and these methods is never established. The paper should either weaken its claims about what is being measured or show how each method operationalizes a specific norm type, e.g., by defining the reference distribution or equality criterion.
- [Sections 1.2 and 4] The paper acknowledges in Section 1.2 that 'there is no universally accepted standard' for fairness and that fairness depends on individual values, yet the conclusion (Section 4) states that bias mitigation can 'tend to fairer IA models.' This tension is never resolved: if norms are subjective, then 'fairer' is undefined without a specified norm. The paper should discuss how norm selection could be grounded, for example through stakeholder participation, legal frameworks, or explicit value statements, rather than leaving it implicit.
minor comments (5)
- [Throughout] There are numerous typographical and grammatical errors: 'Duning-Krugger' should be 'Dunning-Kruger' (Figure 2), 'IAs' should be 'AIs', 'a such universal systems' should be 'such universal systems', and 'it exists a complete zoology' should be 'there exists a complete zoology.'
- [Section 3.3] The explanation of weighted sampling is garbled: 'you can weight the samples of group A by 1 0.9' should read 'by 1/0.9' (and similarly for group B). The formula is clear in intent but the typesetting is broken.
- [Section 2.3] The capitalization of 'Selection Biases' and 'Feature selection bias' is inconsistent; the latter appears with a line break in the heading 'F eature selection bias'. Please align heading styles throughout.
- [References] Some reference entries contain formatting artifacts, such as 'J's M.R.os' in the Sharma et al. entry, and several entries have non-standard line breaks. Please proofread the reference list against the original sources.
- [Figure 8] The caption cites 'Barriere et al. (2023)' as a source of the data-augmentation example, which is one of several self-citations. The reliance on the author's own prior work is acceptable, but the paper would benefit from a broader set of illustrative references for this particular technique.
Circularity Check
No significant circularity: the paper is a conceptual survey whose taxonomies and methods are supported by external literature, and its self-citations serve only as illustrative examples.
full rationale
This manuscript is a position/survey paper, not a derivation of predictions from first principles. Its central definition of bias as 'a deviation from a (subjective and defined) norm or value' (Section 2.1) is an explicitly stated stipulative definition, not a conclusion derived from prior results. The taxonomy in Section 2.2 and the detection/mitigation methods in Section 3 are presented as descriptions of existing work, with external citations such as Meister et al. (2022), Ribeiro et al. (2020), Santy et al. (2023), and Sharma et al. (2020) carrying the evidentiary weight. The author's own works (Barriere and Cifuentes 2024a,b; Barriere et al. 2023) appear only as illustrative examples of bias detection via name-based counterfactuals and of multimodal data augmentation, respectively; the paper's conceptual claims do not depend on these self-citations being true. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through self-citation. The skeptic's concern that the 'norm' in the definition is not operationalized is a genuine conceptual limitation of the framework, but it is not circularity: an under-specified definition does not make the survey's claims equivalent to their inputs by construction. The paper is self-contained as an opinionated overview, so the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Almost nothing in the world is pure randomness; structures are everywhere.
- domain assumption Bias is a deviation from a (subjective and defined) norm or value.
- domain assumption Deep learning models learn correlations from data and can amplify pre-existing biases.
Cite this review
Pith. "Pith review of Fantastic Biases (What are They) and Where to Find Them." pith.science (2026). https://pith.science/paper/P4DLU5AC
@misc{pith2026241115051,
author = {Pith},
title = {Pith review of: Fantastic Biases (What are They) and Where to Find Them},
year = {2026},
howpublished = {\url{https://pith.science/paper/P4DLU5AC}},
note = {Machine review of arXiv:2411.15051}
}
read the original abstract
Deep Learning models tend to learn correlations of patterns on huge datasets. The bigger these systems are, the more complex are the phenomena they can detect, and the more data they need for this. The use of Artificial Intelligence (AI) is becoming increasingly ubiquitous in our society, and its impact is growing everyday. The promises it holds strongly depend on their fair and universal use, such as access to information or education for all. In a world of inequalities, they can help to reach the most disadvantaged areas. However, such a universal systems must be able to represent society, without benefiting some at the expense of others. We must not reproduce the inequalities observed throughout the world, but educate these IAs to go beyond them. We have seen cases where these systems use gender, race, or even class information in ways that are not appropriate for resolving their tasks. Instead of real causal reasoning, they rely on spurious correlations, which is what we usually call a bias. In this paper, we first attempt to define what is a bias in general terms. It helps us to demystify the concept of bias, to understand why we can find them everywhere and why they are sometimes useful. Second, we focus over the notion of what is generally seen as negative bias, the one we want to avoid in machine learning, before presenting a general zoology containing the most common of these biases. We finally conclude by looking at classical methods to detect them, by means of specially crafted datasets of templates and specific algorithms, and also classical methods to mitigate them.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Gavin Abercrombie, Valerio Basile, Davide Bernadi, Shiran Dudy, Simona Frenda, Lucy Havens, and Sara Tonelli. 2024. Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 . In Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024
work page 2024
-
[4]
Gavin Abercrombie, Valerio Basile, Sara Tonelli, Verena Rieser, and Alexandra Uma. 2022. Proceedings of the 1st Workshop on Perspectivist Approaches to NLP@ LREC2022 . In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP@ LREC2022
work page 2022
-
[5]
Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate
Lisa P Argyle, Christopher A Bail, Ethan C Busby, Joshua R. Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate. 2023. https://doi.org/10.1073/pnas.2311627120 Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale . Proceedings of the National Academy of Sciences of the Unite...
-
[6]
Drake Baer. 2017. https://www.thecut.com/2017/01/kahneman-biases-act-like-optical-illusions.html Kahneman: Your Cognitive Biases Act Like Optical Illusions
work page 2017
-
[7]
Valentin Barriere. 2024. https://www.dcc.uchile.cl/media/bits/pdfs/bits26.2-sesgos-fantasticos.pdf Sesgos fant \' a sticos: Qu \' e son y d \' o nde encontrarlos . Bits de Ciencias, 26:02--13
work page 2024
-
[8]
Valentin Barriere and Sebastian Cifuentes. 2024 a . https://aclanthology.org/2024.emnlp-main.34 A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers . In Proceedings of EMNLP, Miami, Florida, USA. Association for Computational Linguistics
work page 2024
Show all 53 references
-
[9]
Valentin Barriere and Sebastian Cifuentes. 2024 b . https://aclanthology.org/2024.lrec-main.134 Are Text Classifiers Xenophobic? A Country-Oriented Bias Detection Method with Least Confounding Variables . In Proceedings of the 2024 Joint International Conference on Computation...
2024
-
[10]
Valentin Barriere, Felipe Del Rio, Andres Carvallo, Carlos Aspillaga, Eugenio Herrera-Berg, and Cristian Buc. 2023. https://aclanthology.org/2023.gem-1.21 Targeted Image Data Augmentation Increases Basic Skills Captioning Robustness . In Proceedings of the Third Workshop on Na...
2023
-
[11]
Yunjey Choi, Minje Choi, Munyoung Kim, Jung Woo Ha, Sunghun Kim, and Jaegul Choo. 2018. https://doi.org/10.1109/CVPR.2018.00916 StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation . In Proceedings of the IEEE Computer Society Conference...
2018
-
[12]
Alba Curry and Amanda Cercas Curry. 2023. https://doi.org/10.18653/v1/2023.findings-acl.515 Computer says “No”: The Case Against Empathetic Conversational AI . In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 8123--8130
2023 doi
-
[13]
Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Mohamed Bin Zayed, and Dirk Hovy. 2024. https://arxiv.org/abs/2403.04445v1 Classist Tools: Social Class Correlates with Performance in NLP . (1964)
1964 arXiv
-
[14]
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021. https://doi.org/10.1162/tacl \_ a \_ 00425 Quantifying social biases in nlp: A generalization and empirical comparison of extrinsic fairness metrics . Transactions of the Association for Computational Linguistics, 9:1249--1267
2021 doi
-
[15]
Jonathan Dunn, Benjamin Adams, and Harish Tayyar Madabushi. 2024. Pre-Trained Language Models Represent Some Geographic Populations Better than Others . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (L...
2024
-
[16]
Yanai Elazar and Yoav Goldberg. 2018. https://doi.org/10.18653/v1/d18-1002 Adversarial removal of demographic attributes from text data . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018, pages 11--21
2018 doi
-
[17]
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. http://arxiv.org/abs/2305.08283 From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models . In ACL, volume 1, pages 11737--11762
2023 arXiv
-
[18]
Nathan Godey, Éric de la Clergerie, and Benoît Sagot. 2024. https://arxiv.org/abs/2402.19406v1 On the Scaling Laws of Geographical Representation in Language Models . In LREC-COLING
2024 arXiv
-
[19]
Melissa Hall, Laurens van der Maaten, Laura Gustafson, Maxwell Jones, and Aaron Adcock. 2022. http://arxiv.org/abs/2201.11706 A Systematic Study of Bias Amplification . Trustworthy and Socially Responsible Machine Learning (TSRML) at Neurips, 1(1)
2022 arXiv
-
[20]
Rothkopf, Alexander Fraser, and Kristian Kersting
Katharina H \" a mmerl, Björn Deiseroth, Patrick Schramowski, Jindřich Libovick \' y , Constantin A. Rothkopf, Alexander Fraser, and Kristian Kersting. 2022. http://arxiv.org/abs/2211.07733 Speaking Multiple Languages Affects the Moral Bias of Language Models . In Findings of ...
2022 arXiv
-
[21]
Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. 2023. http://arxiv.org/abs/2311.09799 How Far Can We Extract Diverse Perspectives from Large Language Models?
2023 arXiv
-
[22]
Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, Scott Johnston, Ben Mann, Chris Olah, Catherine Olsson, Dario Amodei, Nicholas Joseph, Jared Kaplan, and Sam McCandlish. 2022. htt...
2022 arXiv
-
[23]
Yusuke Hirota, Yuta Nakashima, and Noa Garcia. 2022. https://doi.org/10.1109/CVPR52688.2022.01309 Quantifying Societal Bias Amplification in Image Captioning . Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2022-June:13440--13449
2022
-
[24]
Nicolas Jounin, Fatine Ahmadouchi, Aurélie Bachiri, Boubou Bakhayokho, Julien Bihet, Requia Bouali, Nedjma Cognasse, Sarah El Mellah, Camille Gicquel, Marie Josse, Yasmina Kettal, Nina Krumnow, Alice Mimoun, Laëtitia Mokrani, Jordan Mongongnon, Pierre Orsini, Camilla Otto, Luc...
2015 doi
-
[25]
Daniel Kahneman. 2011. Thinking, Fast and Slow
2011
-
[26]
Yelin Kim and Jeesun Kim. 2018. HUMAN-LIKE EMOTION RECOGNITION : MULTI-LABEL LEARNING FROM NOISY LABELED AUDIO-VISUAL EXPRESSIVE SPEECH . In ICASSP, pages 5104--5108
2018
-
[27]
Andy Liu, Mona Diab, and Daniel Fried. 2024. http://arxiv.org/abs/2405.20253 Evaluating Large Language Model Biases in Persona-Steered Generation
2024 arXiv
-
[28]
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. https://doi.org/10.1109/ICCV.2015.425 Deep learning face attributes in the wild . In Proceedings of the IEEE International Conference on Computer Vision, volume 2015 Inter, pages 3730--3738
2015 doi
-
[29]
Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, and Jose Camacho-Collados. 2022. http://arxiv.org/abs/2202.03829 TimeLMs: Diachronic Language Models from Twitter
2022 arXiv
-
[30]
Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020. https://doi.org/10.18653/v1/2020.acl-main.769 End-to-end bias mitigation by modelling biases in corpora . In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 8706--8716
2020 doi
-
[31]
Rohin Manvi, Samar Khanna, Marshall Burke, David Lobell, and Stefano Ermon. 2024. Large language models are geographically biased . arXiv preprint arXiv:2402.02680
2024 arXiv
-
[32]
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. https://doi.org/10.1145/3457607 A Survey on Bias and Fairness in Machine Learning
2021 doi
-
[33]
Clara Meister, Wojciech Stokowiec, Tiago Pimentel, Lei Yu, Laura Rimell, and Adhiguna Kuncoro. 2022. http://arxiv.org/abs/2212.09686 A Natural Bias for Language Generation Models . In ACL, volume 2, pages 243--255
2022 arXiv
-
[34]
Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Valentin Barriere, Doratossadat Dastgheib, Omid Ghahroodi, Mohammad Ali Sadraei, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, and Benno Stein. 2024. http://arxiv.org...
2024 arXiv
-
[35]
Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2024. http://arxiv.org/abs/2305.14456 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models . ACL
2024 arXiv
-
[36]
Observatoire des In \' e galit \' e s . 2021. https://inegalites.fr/Des-controles-de-police-tres-inegaux-selon-la-couleur-de-la-peau Des contr \^ o les de police tr \` e s in \' e gaux selon la couleur de la peau
2021
-
[37]
Hadas Orgad and Yonatan Belinkov. 2022. http://arxiv.org/abs/2212.10563 BLIND: Bias Removal With No Demographics . In ACL, volume 1, pages 8801--8821
2022 arXiv
-
[38]
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. 2023. Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions . In EACL 2023 - 17th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conferenc...
2023
-
[39]
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond Accuracy: Behavioral Testing of NLP Models . ACL
2020
-
[40]
Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap
Sebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. http://arxiv.org/abs/2306.01943 NLPositionality: Characterizing Design Biases of Datasets and Models . 1:9080--9102
2023 arXiv
-
[41]
Smith, and Yejin Choi
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language . Proceedings ofthe 58th Annual Meeting ofthe Association for Computational Linguistics, pages 5477--5490
2020
-
[42]
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. https://doi.org/10.18653/v1/2022.naacl-main.431 Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection . NAACL 2022 - 2022 Conference of the N...
2022 doi
-
[43]
Dominik Schlechtweg, Anna H \" a tty, Marco del Tredici, and Sabine Schulte im Walde. 2020. https://doi.org/10.18653/v1/p19-1072 A wind of change: Detecting and evaluating lexical semantic change across times and domains . In ACL 2019 - 57th Annual Meeting of the Association f...
2020 doi
-
[44]
Varshney
Shubham Sharma, Yunfeng Zhang, Jes's M.R.os Aliaga, Djallel Bouneffouf, Vinod Muthusamy, and Kush R. Varshney. 2020. https://doi.org/10.1145/3375627.3375865 Data augmentation for discrimination prevention and bias disambiguation . In AIES 2020 - Proceedings of the AAAI/ACM Con...
2020
-
[45]
Yaqian Shi and Lei Lei. 2020. The evolution of LGBT labelling words: Tracking 150 years of the interaction of semantics with social and cultural changes . English Today, 36(4):33--39
2020
-
[46]
Taylor Sorensen, Liwei Jiang, Jena Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi. 2023. http://arxiv.org/abs/2309.00779 Value Kaleidoscope: Engaging AI with Pluralistic H...
2023 arXiv
-
[47]
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. http://arxiv.org/abs/2402.05070 Position: A Roadmap to Pluralistic Alignment . ...
2024 arXiv
-
[48]
Shabnam Tafreshi, Orphée De Clercq, Valentin Barriere, João Sedoc, Sven Buechel, and Alexandra Balahur. 2021. WASSA 2021 Shared Task : Predicting Empathy and Emotion in Reaction to News Stories . In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivi...
2021
-
[49]
Rohan Taori and Tatsunori B Hashimoto. 2023. Data Feedback Loops: Model-driven Amplification of Dataset Biases . In Proceedings of Machine Learning Research, volume 202, pages 33883--33920
2023
-
[50]
Nitasha Tiku, Kevin Schaul, and Szu Yu Chen. 2023. https://www.washingtonpost.com/technology/interactive/2023/ai-generated-images-bias-racism-sexism-stereotypes/ This is how AI image generators see the world
2023
-
[51]
Mathieu Valette. 2024. What Does Perspectivism Mean? An Ethical and Methodological Countercriticism . In 3rd Workshop on Perspectivist Approaches to NLP, NLPerspectives 2024 at LREC-COLING 2024 - Workshop Proceedings, pages 111--115
2024
-
[52]
Liwen Wang, Yuanmeng Yan, Keqing He, Yanan Wu, and Weiran Xu. 2021. Dynamically Disentangling Social Bias from Task-Oriented Representations with Adversarial Attack . In Proceedings of NAACL-HLT, pages 3740--3750
2021
-
[53]
Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019. Detection of abusive language: The problem of biased datasets . Proceedings of NAACL HLT 2019, 1:602--608
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.