Pith. sign in

REVIEW 4 major objections 5 minor 53 references

Fantastic Biases (What are They) and Where to Find Them

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper argues that bias is a deviation from a subjective norm rather than an inherent flaw, and that harmful biases in AI systems can be systematically detected and mitigated through data inspection, counterfactual testing, and…

desk verdict A readable, well-organized survey of AI bias that never quite pins down its own central definition of bias as deviation from a norm. read the letter →

arxiv 2411.15051 v1 pith:P4DLU5AC submitted 2024-11-22 cs.CL cs.CVcs.CYcs.LG

classification cs.CLcs.CVcs.CYcs.LG
keywords biasfairnessmachinelearningnaturallanguageprocessinglargemodelsdetectionmitigationsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to demystify the concept of bias in artificial intelligence. It argues that bias is not fundamentally bad; it is simply a deviation from a subjective and context-defined norm, and such deviations are everywhere, from human perception and social norms to the structure of data and models. On this foundation, the paper develops a 'zoology' of common biases that appear in machine learning, natural language processing, and large language models, and then organizes the main detection and mitigation methods according to where the bias lives. A reader comes away with a shared vocabulary for naming a bias and a practical map of the tools available to find and reduce the harmful ones.

What carries the argument

The organizing device is a general definition of bias as a deviation from a norm, together with a taxonomic 'zoology' of biases, a catalog of named biases that classifies them by their source: reporting, selection, representation, group attribution, implicit, annotation, cultural, linguistic, political, demographic, and temporal. This taxonomy carries the survey: each bias is located in a data pipeline or model component, which then determines which detection method (data inspection, counterfactual testing, per-subgroup performance analysis) and which mitigation method (sampling, weighting, augmentation, adversarial loss, pluralistic alignment) applies. The link between the mathematical bias term and cognitive bias is made concrete through the example of initializing a language model's final-layer bias with the log-unigram distribution of words, showing that a bias can encode a useful prior.

What would settle it

A concrete observation that would test the central claim is a case in which a model's output is judged unfair by every human stakeholder even though it exactly matches the chosen norm; such a case would show that fairness cannot be reduced to deviation from a norm. Alternatively, finding a reproducible bias that fits none of the paper's categories and escapes all of its listed detection methods would falsify the survey's coverage.

Watch

Extended reading notes

Core claim

The central claim is that biases are not fundamentally bad, just a deviation from a (subjective and defined) norm or value. The paper uses this definition to connect the mathematical bias term of a linear model, human cognitive biases such as the Dunning-Kruger effect, and social or cultural norms, showing that all are instances of the same phenomenon. It then argues that most harmful biases in machine learning originate from selection biases in data, annotations, and cultural or linguistic representation, and that they surface either in the data itself, in the model's robustness to counterfactual changes, or in uneven performance across demographic groups. On the mitigation side, the paper catalogs resampling, weighted loss functions, data augmentation, adversarial objectives, and alignment techniques as means to reduce harmful bias without abandoning the useful priors that bias supplies.

Load-bearing premise

The framework rests on the premise that a norm can be chosen and specified well enough to measure bias as a deviation from it, even though the paper only states that norms are subjective and context-dependent.

Editorial extensions

If this is right

  • If bias is a deviation from a norm, then any debiasing effort must first make the norm explicit, turning fairness from a purely technical metric into a stated value choice.
  • The taxonomy gives practitioners a shared vocabulary: naming a bias as selection or annotation bias points directly to the detection and mitigation techniques most likely to address it.
  • Counterfactual testing and per-subgroup performance analysis become mandatory validation steps for any model that will be deployed on diverse populations.
  • Because some biases encode useful priors, debiasing is not about removing all deviation but about removing deviations that are harmful for a chosen norm.
  • Methods such as resampling, weighted loss, data augmentation, and adversarial debiasing are each suited to particular bias types, so a one-size-fits-all debiasing approach is unlikely to work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implicit consequence of the paper's definition is that bias metrics are inherently normative: a model can only be judged unbiased relative to a stated norm, so audits and leaderboards should disclose which norm they assume.
  • The taxonomy likely extends beyond NLP to vision-language models, which inherit the same selection and representation biases from web-scale training data; the paper's examples from image generation already hint at this.
  • A testable prediction from the framework is that mitigation methods matched to the diagnosed bias type will outperform generic debiasing; a benchmarking study across the paper's bias categories could verify this.
  • The definition also suggests that data feedback loops, where model outputs contaminate future training data, are a bias-amplification mechanism that fits naturally into the temporal-bias category, a connection the paper leaves implicit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper is a conceptual survey that attempts to define 'bias' in a general sense, distinguish harmful from benign biases, present a taxonomy of common biases in machine learning, natural language processing, and large language models, and catalog methods for detecting and mitigating them. It argues that biases are pervasive and sometimes useful, and that negative biases arise from a deviation from a (subjective) norm or value, offering a 'zoology' of bias types (reporting, selection, representation, group attribution, implicit, cultural, linguistic, ideological, demographic, temporal, confirmation) and a set of detection and mitigation techniques (data inspection, counterfactual robustness, over/under-sampling, weighting, augmentation, adversarial loss, and human/machine perspective integration).

Significance. If its conceptual framework were rigorous, the paper would serve as an accessible introduction and reference for practitioners seeking to name and address bias in ML systems. It compiles a broad set of relevant literature and examples, and its emphasis on bias as a deviation from a chosen norm is a reasonable starting point. However, the framework is not operationalized, and the taxonomy shifts among incompatible norm types. This weakens the paper's central claim to provide a usable framework, but the survey content and reference collection still have pedagogical value. The paper does not provide machine-checked proofs, reproducible code, or falsifiable predictions; its contribution is purely conceptual and expository.

major comments (4)
  1. [Section 2.1] The definition of bias as 'a deviation from a (subjective and defined) norm or value' is never operationalized. The paper does not specify who chooses the norm, how it is elicited, or what distance function measures the deviation. This is load-bearing because the subsequent taxonomy and detection methods all depend on this definition; without a norm-selection rule, the framework cannot distinguish harmful bias from benign statistical variation.
  2. [Section 2.2] The taxonomy silently shifts among mutually incompatible norm types: reporting, selection, and automation biases are framed as deviations from truth, accuracy, or real-world prevalence; representation bias is framed as a deviation from a desired egalitarian target even when the data 'represents the reality' (the mostly-male CEO example); social and cultural biases are framed as deviations from group-dependent social norms. A single dataset can therefore be biased under one norm and unbiased under another, and the paper gives no procedure for selecting the relevant norm. This undermines the claim of providing a coherent 'zoology' of biases.
  3. [Section 3] The detection methods described in Sections 3.1 and 3.2 (data skewness, heterogeneous performance over target groups, counterfactual robustness) detect statistical skew, performance gaps, or decision changes, not deviation from an explicitly chosen norm. The connection between the Section 2.1 definition and these methods is never established. The paper should either weaken its claims about what is being measured or show how each method operationalizes a specific norm type, e.g., by defining the reference distribution or equality criterion.
  4. [Sections 1.2 and 4] The paper acknowledges in Section 1.2 that 'there is no universally accepted standard' for fairness and that fairness depends on individual values, yet the conclusion (Section 4) states that bias mitigation can 'tend to fairer IA models.' This tension is never resolved: if norms are subjective, then 'fairer' is undefined without a specified norm. The paper should discuss how norm selection could be grounded, for example through stakeholder participation, legal frameworks, or explicit value statements, rather than leaving it implicit.
minor comments (5)
  1. [Throughout] There are numerous typographical and grammatical errors: 'Duning-Krugger' should be 'Dunning-Kruger' (Figure 2), 'IAs' should be 'AIs', 'a such universal systems' should be 'such universal systems', and 'it exists a complete zoology' should be 'there exists a complete zoology.'
  2. [Section 3.3] The explanation of weighted sampling is garbled: 'you can weight the samples of group A by 1 0.9' should read 'by 1/0.9' (and similarly for group B). The formula is clear in intent but the typesetting is broken.
  3. [Section 2.3] The capitalization of 'Selection Biases' and 'Feature selection bias' is inconsistent; the latter appears with a line break in the heading 'F eature selection bias'. Please align heading styles throughout.
  4. [References] Some reference entries contain formatting artifacts, such as 'J's M.R.os' in the Sharma et al. entry, and several entries have non-standard line breaks. Please proofread the reference list against the original sources.
  5. [Figure 8] The caption cites 'Barriere et al. (2023)' as a source of the data-augmentation example, which is one of several self-citations. The reliance on the author's own prior work is acceptable, but the paper would benefit from a broader set of illustrative references for this particular technique.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a conceptual survey whose taxonomies and methods are supported by external literature, and its self-citations serve only as illustrative examples.

full rationale

This manuscript is a position/survey paper, not a derivation of predictions from first principles. Its central definition of bias as 'a deviation from a (subjective and defined) norm or value' (Section 2.1) is an explicitly stated stipulative definition, not a conclusion derived from prior results. The taxonomy in Section 2.2 and the detection/mitigation methods in Section 3 are presented as descriptions of existing work, with external citations such as Meister et al. (2022), Ribeiro et al. (2020), Santy et al. (2023), and Sharma et al. (2020) carrying the evidentiary weight. The author's own works (Barriere and Cifuentes 2024a,b; Barriere et al. 2023) appear only as illustrative examples of bias detection via name-based counterfactuals and of multimodal data augmentation, respectively; the paper's conceptual claims do not depend on these self-citations being true. There is no fitted parameter renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz smuggled in through self-citation. The skeptic's concern that the 'norm' in the definition is not operationalized is a genuine conceptual limitation of the framework, but it is not circularity: an under-specified definition does not make the survey's claims equivalent to their inputs by construction. The paper is self-contained as an opinionated overview, so the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters, empirical fits, or new postulated entities. Its conceptual claims rest on domain assumptions about the structured nature of the world, the validity of defining bias as deviation from a norm, and the behavior of deep learning models on biased data.

assumptions (3)
  • domain assumption Almost nothing in the world is pure randomness; structures are everywhere.
    This premise justifies the claim that bias is universal and often useful. Stated in Section 2.1.
  • domain assumption Bias is a deviation from a (subjective and defined) norm or value.
    This is the paper's working definition; it assumes norms can be defined well enough to measure deviation. Section 2.1.
  • domain assumption Deep learning models learn correlations from data and can amplify pre-existing biases.
    Background assumption of the fairness literature, supported by cited works such as Hall et al. (2022).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Fantastic Biases (What are They) and Where to Find Them." pith.science (2026). https://pith.science/paper/P4DLU5AC

@misc{pith2026241115051,
  author       = {Pith},
  title        = {Pith review of: Fantastic Biases (What are They) and Where to Find Them},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P4DLU5AC}},
  note         = {Machine review of arXiv:2411.15051}
}
read the original abstract

Deep Learning models tend to learn correlations of patterns on huge datasets. The bigger these systems are, the more complex are the phenomena they can detect, and the more data they need for this. The use of Artificial Intelligence (AI) is becoming increasingly ubiquitous in our society, and its impact is growing everyday. The promises it holds strongly depend on their fair and universal use, such as access to information or education for all. In a world of inequalities, they can help to reach the most disadvantaged areas. However, such a universal systems must be able to represent society, without benefiting some at the expense of others. We must not reproduce the inequalities observed throughout the world, but educate these IAs to go beyond them. We have seen cases where these systems use gender, race, or even class information in ways that are not appropriate for resolving their tasks. Instead of real causal reasoning, they rely on spurious correlations, which is what we usually call a bias. In this paper, we first attempt to define what is a bias in general terms. It helps us to demystify the concept of bias, to understand why we can find them everywhere and why they are sometimes useful. Second, we focus over the notion of what is generally seen as negative bias, the one we want to avoid in machine learning, before presenting a general zoology containing the most common of these biases. We finally conclude by looking at classical methods to detect them, by means of specially crafted datasets of templates and specific algorithms, and also classical methods to mitigate them.

Figures

Figures reproduced from arXiv: 2411.15051 by the authors.

Figure 1
Figure 1. The chessboard shadow illusion are shaking up society in ways we never imagined, often reinforcing existing inequalities and discrim￾ination. Imagine an algorithm deciding who gets a loan, who gets hired, or even who gets bail. If these algorithms are biased, they can unfairly tar￾get certain groups, like minority communities or women, perpetuating injustice and discrimination. In criminal justice, biased algorithms… view at source ↗
Figure 2
Figure 2. Duning-Krugger effect regarding their bias is the famous gold and white vs black and blue dress.3 At higher-level of complexity, we can men￾tion among others the Dunning-Kruger effect, the availability bias, or the confirmation bias. The Dunning-Kruger effect, illustrated in [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. An example of coverage bias from the xkcd comics #2618 Selection bias It occurs if a dataset’s examples are chosen in a way that is not reflective of their real-world dis￾tribution.8 This bias can take various forms. The coverage bias arises when some groups are in￾adequately represented in the training data, like surveying in a Computer Science classroom in or￾der to poll about how much people know about programmin… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Evolution of meaning of the words gay and queer accross time (Shi and Lei, 2020) contains way more Irish/Scottish references than Chilean/Argentinean ones. Ideological and political biases In the training data, some political and ideological biases are more represented…
Figure 7
Figure 7. Figure 7: Bias detection with respect to coun￾tries, using counterfactual examples over a senti￾ment analysis system, which should be supposed to output the same predictions. Annotation not representative Ask the demographic of the annotators when col￾lecting them is now seeing …
Figure 8
Figure 8. Figure 8: Example of multimodal data-augmentation fostering diversity ( [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 27 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Gavin Abercrombie, Valerio Basile, Davide Bernadi, Shiran Dudy, Simona Frenda, Lucy Havens, and Sara Tonelli. 2024. Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024 . In Proceedings of the 3rd Workshop on Perspectivist Approaches to NLP (NLPerspectives)@ LREC-COLING 2024

  4. [4]

    Gavin Abercrombie, Valerio Basile, Sara Tonelli, Verena Rieser, and Alexandra Uma. 2022. Proceedings of the 1st Workshop on Perspectivist Approaches to NLP@ LREC2022 . In Proceedings of the 1st Workshop on Perspectivist Approaches to NLP@ LREC2022

  5. [5]

    Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate

    Lisa P Argyle, Christopher A Bail, Ethan C Busby, Joshua R. Gubler, Thomas Howe, Christopher Rytting, Taylor Sorensen, and David Wingate. 2023. https://doi.org/10.1073/pnas.2311627120 Leveraging AI for democratic discourse: Chat interventions can improve online political conversations at scale . Proceedings of the National Academy of Sciences of the Unite...

  6. [6]

    Drake Baer. 2017. https://www.thecut.com/2017/01/kahneman-biases-act-like-optical-illusions.html Kahneman: Your Cognitive Biases Act Like Optical Illusions

  7. [7]

    Valentin Barriere. 2024. https://www.dcc.uchile.cl/media/bits/pdfs/bits26.2-sesgos-fantasticos.pdf Sesgos fant \' a sticos: Qu \' e son y d \' o nde encontrarlos . Bits de Ciencias, 26:02--13

  8. [8]

    Valentin Barriere and Sebastian Cifuentes. 2024 a . https://aclanthology.org/2024.emnlp-main.34 A Study of Nationality Bias in Names and Perplexity using Off-the-Shelf Affect-related Tweet Classifiers . In Proceedings of EMNLP, Miami, Florida, USA. Association for Computational Linguistics

Show all 53 references
  1. [9]

    Valentin Barriere and Sebastian Cifuentes. 2024 b . https://aclanthology.org/2024.lrec-main.134 Are Text Classifiers Xenophobic? A Country-Oriented Bias Detection Method with Least Confounding Variables . In Proceedings of the 2024 Joint International Conference on Computation...

  2. [10]

    Valentin Barriere, Felipe Del Rio, Andres Carvallo, Carlos Aspillaga, Eugenio Herrera-Berg, and Cristian Buc. 2023. https://aclanthology.org/2023.gem-1.21 Targeted Image Data Augmentation Increases Basic Skills Captioning Robustness . In Proceedings of the Third Workshop on Na...

  3. [11]

    Yunjey Choi, Minje Choi, Munyoung Kim, Jung Woo Ha, Sunghun Kim, and Jaegul Choo. 2018. https://doi.org/10.1109/CVPR.2018.00916 StarGAN: Unified Generative Adversarial Networks for Multi-domain Image-to-Image Translation . In Proceedings of the IEEE Computer Society Conference...

  4. [12]

    Alba Curry and Amanda Cercas Curry. 2023. https://doi.org/10.18653/v1/2023.findings-acl.515 Computer says “No”: The Case Against Empathetic Conversational AI . In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 8123--8130

  5. [13]

    Amanda Cercas Curry, Giuseppe Attanasio, Zeerak Talat, Mohamed Bin Zayed, and Dirk Hovy. 2024. https://arxiv.org/abs/2403.04445v1 Classist Tools: Social Class Correlates with Performance in NLP . (1964)

  6. [14]

    Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021. https://doi.org/10.1162/tacl \_ a \_ 00425 Quantifying social biases in nlp: A generalization and empirical comparison of extrinsic fairness metrics . Transactions of the Association for Computational Linguistics, 9:1249--1267

  7. [15]

    Jonathan Dunn, Benjamin Adams, and Harish Tayyar Madabushi. 2024. Pre-Trained Language Models Represent Some Geographic Populations Better than Others . In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (L...

  8. [16]

    Yanai Elazar and Yoav Goldberg. 2018. https://doi.org/10.18653/v1/d18-1002 Adversarial removal of demographic attributes from text data . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018, pages 11--21

  9. [17]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. http://arxiv.org/abs/2305.08283 From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models . In ACL, volume 1, pages 11737--11762

  10. [18]

    Nathan Godey, Éric de la Clergerie, and Benoît Sagot. 2024. https://arxiv.org/abs/2402.19406v1 On the Scaling Laws of Geographical Representation in Language Models . In LREC-COLING

  11. [19]

    Melissa Hall, Laurens van der Maaten, Laura Gustafson, Maxwell Jones, and Aaron Adcock. 2022. http://arxiv.org/abs/2201.11706 A Systematic Study of Bias Amplification . Trustworthy and Socially Responsible Machine Learning (TSRML) at Neurips, 1(1)

  12. [20]

    Rothkopf, Alexander Fraser, and Kristian Kersting

    Katharina H \" a mmerl, Björn Deiseroth, Patrick Schramowski, Jindřich Libovick \' y , Constantin A. Rothkopf, Alexander Fraser, and Kristian Kersting. 2022. http://arxiv.org/abs/2211.07733 Speaking Multiple Languages Affects the Moral Bias of Language Models . In Findings of ...

  13. [21]

    Shirley Anugrah Hayati, Minhwa Lee, Dheeraj Rajagopal, and Dongyeop Kang. 2023. http://arxiv.org/abs/2311.09799 How Far Can We Extract Diverse Perspectives from Large Language Models?

  14. [22]

    Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, Scott Johnston, Ben Mann, Chris Olah, Catherine Olsson, Dario Amodei, Nicholas Joseph, Jared Kaplan, and Sam McCandlish. 2022. htt...

  15. [23]

    Yusuke Hirota, Yuta Nakashima, and Noa Garcia. 2022. https://doi.org/10.1109/CVPR52688.2022.01309 Quantifying Societal Bias Amplification in Image Captioning . Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2022-June:13440--13449

  16. [24]

    Nicolas Jounin, Fatine Ahmadouchi, Aurélie Bachiri, Boubou Bakhayokho, Julien Bihet, Requia Bouali, Nedjma Cognasse, Sarah El Mellah, Camille Gicquel, Marie Josse, Yasmina Kettal, Nina Krumnow, Alice Mimoun, Laëtitia Mokrani, Jordan Mongongnon, Pierre Orsini, Camilla Otto, Luc...

  17. [25]

    Daniel Kahneman. 2011. Thinking, Fast and Slow

  18. [26]

    Yelin Kim and Jeesun Kim. 2018. HUMAN-LIKE EMOTION RECOGNITION : MULTI-LABEL LEARNING FROM NOISY LABELED AUDIO-VISUAL EXPRESSIVE SPEECH . In ICASSP, pages 5104--5108

  19. [27]

    Andy Liu, Mona Diab, and Daniel Fried. 2024. http://arxiv.org/abs/2405.20253 Evaluating Large Language Model Biases in Persona-Steered Generation

  20. [28]

    Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang. 2015. https://doi.org/10.1109/ICCV.2015.425 Deep learning face attributes in the wild . In Proceedings of the IEEE International Conference on Computer Vision, volume 2015 Inter, pages 3730--3738

  21. [29]

    Daniel Loureiro, Francesco Barbieri, Leonardo Neves, Luis Espinosa Anke, and Jose Camacho-Collados. 2022. http://arxiv.org/abs/2202.03829 TimeLMs: Diachronic Language Models from Twitter

  22. [30]

    Rabeeh Karimi Mahabadi, Yonatan Belinkov, and James Henderson. 2020. https://doi.org/10.18653/v1/2020.acl-main.769 End-to-end bias mitigation by modelling biases in corpora . In Proceedings of the Annual Meeting of the Association for Computational Linguistics, pages 8706--8716

  23. [31]

    Rohin Manvi, Samar Khanna, Marshall Burke, David Lobell, and Stefano Ermon. 2024. Large language models are geographically biased . arXiv preprint arXiv:2402.02680

  24. [32]

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021. https://doi.org/10.1145/3457607 A Survey on Bias and Fairness in Machine Learning

  25. [33]

    Clara Meister, Wojciech Stokowiec, Tiago Pimentel, Lei Yu, Laura Rimell, and Adhiguna Kuncoro. 2022. http://arxiv.org/abs/2212.09686 A Natural Bias for Language Generation Models . In ACL, volume 2, pages 243--255

  26. [34]

    Nailia Mirzakhmedova, Johannes Kiesel, Milad Alshomary, Maximilian Heinrich, Nicolas Handke, Xiaoni Cai, Valentin Barriere, Doratossadat Dastgheib, Omid Ghahroodi, Mohammad Ali Sadraei, Ehsaneddin Asgari, Lea Kawaletz, Henning Wachsmuth, and Benno Stein. 2024. http://arxiv.org...

  27. [35]

    Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2024. http://arxiv.org/abs/2305.14456 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models . ACL

  28. [36]

    Observatoire des In \' e galit \' e s . 2021. https://inegalites.fr/Des-controles-de-police-tres-inegaux-selon-la-couleur-de-la-peau Des contr \^ o les de police tr \` e s in \' e gaux selon la couleur de la peau

  29. [37]

    Hadas Orgad and Yonatan Belinkov. 2022. http://arxiv.org/abs/2212.10563 BLIND: Bias Removal With No Demographics . In ACL, volume 1, pages 8801--8821

  30. [38]

    Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. 2023. Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions . In EACL 2023 - 17th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conferenc...

  31. [39]

    Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond Accuracy: Behavioral Testing of NLP Models . ACL

  32. [40]

    Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap

    Sebastin Santy, Jenny T. Liang, Ronan Le Bras, Katharina Reinecke, and Maarten Sap. 2023. http://arxiv.org/abs/2306.01943 NLPositionality: Characterizing Design Biases of Datasets and Models . 1:9080--9102

  33. [41]

    Smith, and Yejin Choi

    Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020. Social Bias Frames: Reasoning about Social and Power Implications of Language . Proceedings ofthe 58th Annual Meeting ofthe Association for Computational Linguistics, pages 5477--5490

  34. [42]

    Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A. Smith. 2022. https://doi.org/10.18653/v1/2022.naacl-main.431 Annotators with Attitudes: How Annotator Beliefs And Identities Bias Toxic Language Detection . NAACL 2022 - 2022 Conference of the N...

  35. [43]

    Dominik Schlechtweg, Anna H \" a tty, Marco del Tredici, and Sabine Schulte im Walde. 2020. https://doi.org/10.18653/v1/p19-1072 A wind of change: Detecting and evaluating lexical semantic change across times and domains . In ACL 2019 - 57th Annual Meeting of the Association f...

  36. [44]

    Varshney

    Shubham Sharma, Yunfeng Zhang, Jes's M.R.os Aliaga, Djallel Bouneffouf, Vinod Muthusamy, and Kush R. Varshney. 2020. https://doi.org/10.1145/3375627.3375865 Data augmentation for discrimination prevention and bias disambiguation . In AIES 2020 - Proceedings of the AAAI/ACM Con...

  37. [45]

    Yaqian Shi and Lei Lei. 2020. The evolution of LGBT labelling words: Tracking 150 years of the interaction of semantics with social and cultural changes . English Today, 36(4):33--39

  38. [46]

    Taylor Sorensen, Liwei Jiang, Jena Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, Maarten Sap, John Tasioulas, and Yejin Choi. 2023. http://arxiv.org/abs/2309.00779 Value Kaleidoscope: Engaging AI with Pluralistic H...

  39. [47]

    Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. http://arxiv.org/abs/2402.05070 Position: A Roadmap to Pluralistic Alignment . ...

  40. [48]

    Shabnam Tafreshi, Orphée De Clercq, Valentin Barriere, João Sedoc, Sven Buechel, and Alexandra Balahur. 2021. WASSA 2021 Shared Task : Predicting Empathy and Emotion in Reaction to News Stories . In Proceedings of the Eleventh Workshop on Computational Approaches to Subjectivi...

  41. [49]

    Rohan Taori and Tatsunori B Hashimoto. 2023. Data Feedback Loops: Model-driven Amplification of Dataset Biases . In Proceedings of Machine Learning Research, volume 202, pages 33883--33920

  42. [50]

    Nitasha Tiku, Kevin Schaul, and Szu Yu Chen. 2023. https://www.washingtonpost.com/technology/interactive/2023/ai-generated-images-bias-racism-sexism-stereotypes/ This is how AI image generators see the world

  43. [51]

    Mathieu Valette. 2024. What Does Perspectivism Mean? An Ethical and Methodological Countercriticism . In 3rd Workshop on Perspectivist Approaches to NLP, NLPerspectives 2024 at LREC-COLING 2024 - Workshop Proceedings, pages 111--115

  44. [52]

    Liwen Wang, Yuanmeng Yan, Keqing He, Yanan Wu, and Weiran Xu. 2021. Dynamically Disentangling Social Bias from Task-Oriented Representations with Adversarial Attack . In Proceedings of NAACL-HLT, pages 3740--3750

  45. [53]

    Michael Wiegand, Josef Ruppenhofer, and Thomas Kleinbauer. 2019. Detection of abusive language: The problem of biased datasets . Proceedings of NAACL HLT 2019, 1:602--608

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.