Pith. sign in

REVIEW 3 major objections 6 minor 44 references

Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection

T0 review · 3 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that large language models exhibit a systematic split between explicit and implicit bias, with explicit bias appearing mild and implicit bias strong, and that scaling and alignment training widen this split rather than…

desk verdict A broad and provocative empirical study, but the 'implicit vs. explicit' framing overclaims: the implicit probe is an unvalidated completion task, the 'self-reflection' step never uses the model's own outputs, and the scaling analysis conflates data size with model size. read the letter →

arxiv 2501.02295 v4 pith:BNBZ372L submitted 2025-01-04 cs.CL

classification cs.CL
keywords implicitbiasexplicitself-reflectionAssociationTestsocialinLLMsmodelscalingeffectsalignmenttrainingstereotypicalscore
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models, like people, may show a gap between what they openly state and what they reveal under indirect testing. The authors measure implicit bias by asking models to fill blanks in analogy statements that pair social groups with attributes, and explicit bias by asking the same models to rate whether such statements express stereotypes on a Likert scale. Across six social dimensions and six models, they report that explicit stereotyping is mild while implicit stereotyping is strong, and that this gap widens as models grow and as alignment training is applied.

What carries the argument

The self-reflection evaluation framework. Implicit bias is measured by adapting the Implicit Association Test to prompt templates of the form "<mask> is attrX as <mask> is attrY", where the model must choose two names or group stimuli to fill the blanks; explicit bias is measured by presenting the same association with explicit target groups and asking the model to rate agreement on a 5-point Likert scale, instructing it to reflect on the statement it had just completed. The pairing of the two measures on identical attribute pairs is what allows the explicit–implicit comparison.

What would settle it

If a probe that measures association strength without requiring a conscious choice — for example, comparing the model's token probabilities for stereotype-consistent versus stereotype-inconsistent completions — fails to reproduce the strong implicit scores, or if a model can be induced to fill the blanks stereotype-consistently while showing no corresponding probability difference, the claimed implicit–explicit distinction would be an artifact of the completion prompt.

Watch

Extended reading notes

Core claim

The paper's central discovery is a systematic explicit–implicit bias inconsistency in large language models: when a model is asked directly whether men are to CEOs as women are to secretaries, it strongly disagrees, but when the same association is probed through masked word-choice completions, it reliably selects the stereotypical pairing. The authors contend this mirrors the human pattern documented in social psychology, and that the two measures diverge with respect to scaling: more training data and parameters reduce self-reported stereotyping while increasing stereotype-consistent completions, and preference alignment reduces the former while leaving the latter largely unchanged.

Load-bearing premise

The results stand on the assumption that filling blanks in a masked analogy sentence captures the same automatic, unconscious associations that the Implicit Association Test captures in people, rather than merely reflecting surface word co-occurrence in training data.

Editorial extensions

If this is right

  • If the claim holds, bias evaluations that rely on self-reported or refusal-based tests underestimate the stereotypes that language models actually operate with.
  • Scaling up models and training data would make the explicit–implicit discrepancy worse, not better, so mitigation would need to target implicit associations directly rather than relying on scale.
  • Alignment methods such as DPO can train models to deny stereotypes while leaving the underlying statistical associations intact, implying that surface behavior can be made non-stereotypical without changing the model's internal associations.
  • The reported positive correlation between model capability ranking and implicit bias would mean that the most capable models are also the most implicitly biased.
  • Any bias benchmark that uses only explicit probes would be measuring a compliance behavior rather than the strength of the underlying association.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An untested reading is that the 'implicit bias' here is really task-format bias: masked completion may simply reflect corpus co-occurrence statistics rather than an attitude, while the explicit rating may reflect instruction-following; if so, the central gap is a difference between a cloze task and a refusal task, not between two kinds of cognition.
  • A testable extension would be to correlate the masked-completion score with an independent response-time or log-probability measure of association strength; if they do not track each other, the 'implicit' label is doing too much work.
  • The scaling results suggest a possible 'bias scaling law' opposite to the one for capabilities: implicit association strength rises with scale while explicit compliance rises, so any safety evaluation that only checks stated preferences would be blind to the trend.
  • Because the alignment experiments use only DPO on a single dataset, the conclusion that alignment cannot touch implicit bias should be read as specific to that method and data, not as a general property of all alignment techniques.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper proposes a two-phase, 'self-reflection-based' framework for measuring explicit and implicit social bias in LLMs. In the implicit phase, models fill masked analogy templates such as '<mask> is often to CEO, as <mask> is often to secretary' by choosing candidate names. In the explicit phase, models rate agreement with stereotype statements (e.g., 'Men are often to CEOs as women are to secretaries') on a 5-point Likert scale. The authors define a Stereotypical Score (SC) as the proportion of stereotype-consistent responses and report experiments on six prominent LLMs across six social dimensions (age, disability, gender-career, gender-occupation, race, science). They find high implicit SC and low explicit SC across all models. Scaling experiments on LLaMA families show that explicit bias decreases while implicit bias increases with model size and training data, and DPO-based alignment reduces explicit bias but leaves implicit bias stable. The paper concludes that LLMs exhibit an explicit-implicit bias inconsistency analogous to that found in humans.

Significance. The paper addresses a timely and important question: whether direct agreement probes and indirect association probes reveal different levels of stereotyping in LLMs, and how those levels change with model scaling and alignment. The use of multiple LLaMA model families with documented token counts is a strength, and the paper's core empirical observation—that indirect completion templates yield much higher stereotype-consistent rates than direct Likert items—is reproducible in principle. The framing is also valuable because it jointly studies explicit and implicit bias, which is rare in the LLM bias literature. However, the significance depends entirely on interpreting the indirect measure as 'implicit bias' in the psychological sense. That interpretation is not validated, and the explicit phase is not actually based on self-reflection as described. Without addressing these issues, the central conclusion is unsupported.

major comments (3)
  1. [Section 3.1] The implicit bias measure is not validated. The masked sentence-completion task is a deliberative choice among candidate names; no evidence is given that it measures automatic, unconscious associations. The psychological IAT relies on response-time differences under speeded conditions, which a word-choice task does not reproduce. The authors do not provide convergent validity (e.g., correlation with human IAT D-scores or established bias benchmarks), test-retest reliability, or per-item consistency checks. Without such evidence, the high 'implicit' SC may simply reflect corpus co-occurrence or instruction-following, and the observed explicit-implicit dissociation may be a task-format artifact. This concern directly undermines the central claim in Sections 5.1 and 6 that LLMs exhibit implicit bias that persists after alignment.
  2. [Section 3.2 and Figure 1] The explicit measure is not self-reflection as described. Section 3.2 states that the model evaluates 'its potential attitudes demonstrated during the implicit bias measurement phase,' but the actual prompt in Figure 1(b) is a generic Likert rating of a stereotype statement. The model is never shown its own implicit-phase completions. Moreover, the explicit task differs from the implicit task in multiple ways beyond direct versus indirect measurement: it uses group labels (men/women) instead of names (John/Lisa) and an agreement scale instead of a forced choice. These differences alone could produce the score gap. The paper therefore does not implement the claimed self-reflection methodology, and the comparison is not a clean test of explicit versus implicit attitude.
  3. [Section 5.2.1 and Figure 3] The scaling analysis cannot separate training-data scale from model size. In the LLaMA family, model size and pre-training token count are nearly perfectly confounded: LLaMA-2 7B/13B/70B all use 2T tokens, LLaMA-3 8B/70B and LLaMA-3.1 405B all use 15T, and LLaMA-3.2 1B/3B use 9T. The claim that 'increasing training data and parameters' drives the observed trends is therefore not supported by these data. Disentangling the two effects would require models trained with varying token counts at a fixed parameter count, or vice versa. This confound also limits the generality of the conclusion that implicit bias increases with scaling.
minor comments (6)
  1. [Abstract and Section 3.2] The abstract says explicit bias is evaluated by 'prompting LLMs to analyze their own generated content,' but the explicit prompt does not use content generated in the implicit phase; it uses fixed stereotype statements. Please align the abstract with the actual procedure.
  2. [Section 4.1] The claimed total of 14,400 experiments counts only the six models in Table 1, but the scaling and alignment experiments on additional LLaMA base and tuned models are not included in this arithmetic. Please clarify the total number of experiments actually run.
  3. [Figure 3] The figure should include a legend or point labels identifying model family and token count; without this, the confound between model size and training data is not visually transparent.
  4. [Table 3] The Stereotypical Scores are reported as point estimates without confidence intervals or standard errors, even though N=200 per cell; the table would benefit from uncertainty quantification to prevent over-reading of small differences.
  5. [Section 3.1 and Appendix A] The template phrase 'is often to' is ungrammatical; consider using a grammatical analogy format and reporting whether results are robust to template wording.
  6. [Limitations] The Limitations section does not acknowledge the absence of validity evidence for the implicit measure or the fact that the explicit phase is not self-reflection; these are central limitations and should be discussed.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the explicit/implicit dissociation is an empirical measurement outcome, not a formal consequence of the definitions; the only self-citation is a non-load-bearing related-work reference.

full rationale

The paper's chain is experimental rather than derivational. The Stereotypical Score (Eq. 1) is a frequency count of directly observed response behaviors, with no fitted parameters reused as predictions. The implicit probe (masked association, Section 3.1) and the explicit probe (Likert agreement, Section 3.2) are defined independently, and the reported explicit-implicit inconsistency is an empirical result: nothing in the definitions forces implicit scores to be high when explicit scores are low. The main vulnerabilities—whether the masked-completion task validly measures unconscious 'implicit' bias, and whether the explicit condition implements true self-reflection rather than a direct agreement probe—are construct-validity concerns, not circular reductions. The single citation to the authors' prior work (Zhao et al., 2024) appears only in a related-work list of implicit-bias studies and supplies no load-bearing premise, uniqueness theorem, or ansatz. The appended Limitations section candidly limits scope (single-category biases, three factors) but does not reveal any circular dependence. Therefore the central claims do not reduce to their inputs by construction; the low score reflects the non-load-bearing self-citation and the framing mismatch, not derivation circularity.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central claim depends on the assumption that the two prompt-based tasks measure the same psychological constructs (implicit and explicit bias) that they are named after. No new entities are introduced, and the only hand-chosen parameter is the Likert cutoff for explicit stereotypes.

free parameters (1)
  • Explicit stereotype threshold = agree or strongly agree (options 4 and 5 on a 5-point Likert scale)
    Section 4.2 defines a stereotype as present only when the model selects 'agree' or 'strongly agree'. Moving this cutoff would change explicit bias scores and the size of the explicit-implicit gap.
assumptions (3)
  • domain assumption The Implicit Association Test (IAT) paradigm, when adapted to LLMs by using a fill-in-the-blank choice, measures implicit bias.
    Section 3.1 uses the template '_ is attrX as _ is attrY' and treats stereotypical choices as implicit bias; this assumes the choice reveals unconscious association rather than lexical statistics.
  • domain assumption Self-Report Assessment using a 5-point Likert scale measures explicit bias in LLMs.
    Section 3.2 treats agreement with stereotype statements as explicit bias; this assumes the model's stated agreement reflects its deliberate attitude.
  • domain assumption The WEAT/IAT stimulus word lists are valid for probing these biases in LLMs.
    Section 4.1 says four categories use validated words from WEAT/IAT and two are from official IAT and BLS data; this assumes these human-oriented stimuli work for LLMs.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection." pith.science (2026). https://pith.science/paper/BNBZ372L

@misc{pith2026250102295,
  author       = {Pith},
  title        = {Pith review of: Explicit vs. Implicit: Investigating Social Bias in Large Language Models through Self-Reflection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BNBZ372L}},
  note         = {Machine review of arXiv:2501.02295}
}
read the original abstract

Large Language Models (LLMs) have been shown to exhibit various biases and stereotypes in their generated content. While extensive research has investigated biases in LLMs, prior work has predominantly focused on explicit bias, with minimal attention to implicit bias and the relation between these two forms of bias. This paper presents a systematic framework grounded in social psychology theories to investigate and compare explicit and implicit biases in LLMs. We propose a novel self-reflection-based evaluation framework that operates in two phases: first measuring implicit bias through simulated psychological assessment methods, then evaluating explicit bias by prompting LLMs to analyze their own generated content. Through extensive experiments on advanced LLMs across multiple social dimensions, we demonstrate that LLMs exhibit a substantial inconsistency between explicit and implicit biases: while explicit bias manifests as mild stereotypes, implicit bias exhibits strong stereotypes. We further investigate the underlying factors contributing to this explicit-implicit bias inconsistency, examining the effects of training data scale, model size, and alignment techniques. Experimental results indicate that while explicit bias declines with increased training data and model size, implicit bias exhibits a contrasting upward trend. Moreover, contemporary alignment methods effectively suppress explicit bias but show limited efficacy in mitigating implicit bias.

Figures

Figures reproduced from arXiv: 2501.02295 by the authors.

Figure 1
Figure 1. Our proposed assessment methodology based [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Average stereotypical scores: comparing ex [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 5
Figure 5. Impact of alignment training steps on LLaMA [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 12 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Abubakar Abid, Maheen Farooqi, and James Zou. 2021. Persistent anti-muslim bias in large language models. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, pages 298--306

  4. [4]

    Muhammad Ali, Swetasudha Panda, Qinlan Shen, Michael Wick, and Ari Kobren. 2024. Understanding the interplay of scale, data, and bias in language models: A case study with bert. arXiv preprint arXiv:2407.21058

  5. [5]

    Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L Griffiths. 2024. Measuring implicit bias in explicitly unbiased large language models. arXiv preprint arXiv:2402.04105

  6. [6]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022 a . Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862

  7. [7]

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. 2022 b . Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073

  8. [8]

    Andrew Scott Baron and Mahzarin R Banaji. 2006. The development of implicit attitudes: Evidence of race evaluations from ages 6 and 10 and adulthood. Psychological science, 17(1):53--58

Show all 44 references
  1. [9]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey W...

  2. [10]

    Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017. Semantics derived automatically from language corpora contain human-like biases. Science, 356(6334):183--186

  3. [11]

    Myra Cheng, Esin Durmus, and Dan Jurafsky. 2023. https://doi.org/10.18653/v1/2023.acl-long.84 Marked personas: Using natural language prompts to measure stereotypes in language models . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics ...

  4. [12]

    Gonzalez, and Ion Stoica

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. https://proceedings.mlr.press/v235/chiang24b.html Chatbot arena: An open platform for evaluating...

  5. [13]

    Christian S Crandall, Amy Eshleman, and Laurie O'brien. 2002. Social norms and the expression and suppression of prejudice: the struggle for internalization. Journal of personality and social psychology, 82(3):359

  6. [14]

    John F Dovidio. 2010. The SAGE handbook of prejudice, stereotyping and discrimination. Sage Publications

  7. [15]

    John F Dovidio, Kerry Kawakami, and Kelly R Beach. 2001. Implicit and explicit attitudes: Examination of the relationship between measures of intergroup bias. Blackwell handbook of social psychology: Intergroup processes, 4:175--197

  8. [16]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  9. [17]

    Deep Ganguli, Amanda Askell, Nicholas Schiefer, Thomas I Liao, Kamil \.e Luko s i \=u t \.e , Anna Chen, Anna Goldie, Azalia Mirhoseini, Catherine Olsson, Danny Hernandez, et al. 2023. The capacity for moral self-correction in large language models. arXiv preprint arXiv:2302.07459

  10. [18]

    Daniel T Gilbert and J Gregory Hixon. 1991. The trouble of thinking: Activation and application of stereotypic beliefs. Journal of Personality and social Psychology, 60(4):509

  11. [19]

    Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett. 2023. https://doi.org/10.18653/v1/2023.findings-acl.139 This prompt is measuring mask : evaluating bias evaluation in language models . In Findings of the Association for Computational Linguistics: AC...

  12. [20]

    Anthony G Greenwald and Mahzarin R Banaji. 1995. Implicit social cognition: attitudes, self-esteem, and stereotypes. Psychological review, 102(1):4

  13. [21]

    Anthony G Greenwald, Debbie E McGhee, and Jordan LK Schwartz. 1998. Measuring individual differences in implicit cognition: the implicit association test. Journal of personality and social psychology, 74(6):1464

  14. [22]

    Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 Lo RA : Low-rank adaptation of large language models . In International Conference on Learning Representations

  15. [23]

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276

  16. [24]

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361

  17. [25]

    Hannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal, Elias Benussi, Frederic Dreyer, Aleksandar Shtedritski, and Yuki Asano. 2021. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models. Advances in neural ...

  18. [26]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. Gender bias and stereotypes in large language models. In Proceedings of the ACM collective intelligence conference, pages 12--24

  19. [27]

    Timo Lajunen and Türker Özkan. 2011. https://doi.org/10.1016/B978-0-12-381984-0.10004-9 Chapter 4 - self-report instruments and methods . In Bryan E. Porter, editor, Handbook of Traffic Psychology, pages 43--59. Academic Press, San Diego

  20. [28]

    Rensis Likert. 1932. A technique for the measurement of attitudes. Archives of Psychology

  21. [29]

    Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023. https://proceed...

  22. [30]

    Corinne A Moss-Racusin, John F Dovidio, Victoria L Brescoll, Mark J Graham, and Jo Handelsman. 2012. Science faculty’s subtle gender biases favor male students. Proceedings of the national academy of sciences, 109(41):16474--16479

  23. [31]

    Brian A Nosek. 2007. Implicit--explicit relations. Current directions in psychological science, 16(2):65--69

  24. [32]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 3...

  25. [33]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2024. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36

  26. [34]

    Preethi Seshadri, Pouya Pezeshkpour, and Sameer Singh. 2022. https://openreview.net/forum?id=rIhzjia7SLa Quantifying social biases using templates is unreliable . In Workshop on Trustworthy and Socially Responsible Machine Learning, NeurIPS 2022

  27. [35]

    Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. https://proceedings.neurips.cc/paper_files/paper/2023/file/1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf Reflexion: language agents with verbal reinforcement learning . In Advances...

  28. [36]

    Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.625 `` I ' m sorry to hear that '' : Finding new biases in language models with a holistic descriptor dataset . In Proceedings of the 202...

  29. [37]

    Joanna Smith and Helen Noble. 2014. Bias in research. Evidence-based nursing, 17(4):100--101

  30. [38]

    Leanne S Son Hing, Greg A Chung-Yan, Leah K Hamilton, and Mark P Zanna. 2008. A two-dimensional model that employs explicit and implicit attitudes to characterize prejudice. Journal of Personality and Social Psychology, 94(6):971

  31. [39]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805

  32. [40]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  33. [41]

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, et al. 2023. Decodingtrust: A comprehensive assessment of trustworthiness in gpt models. In NeurIPS

  34. [42]

    Julia Watson, Barend Beekhuizen, and Suzanne Stevenson. 2023. https://doi.org/10.18653/v1/2023.acl-long.375 What social attitudes about gender does BERT encode? leveraging insights from psycholinguistics . In Proceedings of the 61st Annual Meeting of the Association for Comput...

  35. [43]

    Yixuan Weng, Minjun Zhu, Fei Xia, Bin Li, Shizhu He, Shengping Liu, Bin Sun, Kang Liu, and Jun Zhao. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.167 Large language models are better reasoners with self-verification . In Findings of the Association for Computational L...

  36. [44]

    Yachao Zhao, Bo Wang, Yan Wang, Dongming Zhao, Xiaojia Jin, Jijun Zhang, Ruifang He, and Yuexian Hou. 2024. https://aclanthology.org/2024.lrec-main.17/ A comparative study of explicit and implicit gender biases in large language models via self-evaluation . In Proceedings of t...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.