Pith. sign in

REVIEW 4 major objections 5 minor 60 references

FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A 20,000-scenario Indian benchmark finds LLMs consistently favor privileged castes and reinforce stereotypes across 85 identity groups.

desk verdict A genuinely useful India-centric fairness benchmark with a robust caste-bias finding, undercut by a wrong stereotype baseline that needs fixing before the headline claim is trusted. read the letter →

arxiv 2506.23111 v1 pith:5QFDTVCE submitted 2025-06-29 cs.CL

classification cs.CL
keywords fairnessbiasstereotypesIndialargelanguagemodelscastebenchmarkINDIC-BIAS
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper builds INDIC-BIAS, a benchmark to test whether large language models are fair in Indian contexts, covering 85 identity groups across caste, religion, region, and tribe. It is built from an expert-curated taxonomy of more than 1,800 socio-cultural topics and 20,000 human-verified scenario templates, organized into plausibility, judgment, and generation tasks. Evaluating 14 popular models, the paper reports that models consistently show negative bias against marginalized identities such as Dalits and reinforce common stereotypes in a majority of cases, and that asking models to reason step by step does not reliably reduce this. The authors release the benchmark openly to support evaluation and mitigation rather than to propose fixes. If the findings hold, LLM use in Indian education, hiring, advice, and public services risks both unequal treatment and demeaning portrayals.

What carries the argument

The load-bearing machinery is INDIC-BIAS itself: a set of 20,000 manually verified scenario templates with identity placeholders, built from an expert-curated taxonomy of over 1,800 socio-cultural topics covering caste, religion, region, and tribe. Each template can be instantiated with different identities, producing matched pairs that differ only in identity. Bias is measured by treating each paired comparison as a match and computing ELO ratings via the Bradley-Terry model, then comparing an identity's rank in positive versus negative scenarios through the Rank Shift Metric (RSM), where a negative shift means the identity is preferred more in negative scenarios. Stereotype association is measured by the Stereotype Association Rate (SAR), the fraction of times a model picks the target identity in scenarios designed around that identity's stereotype, with refusals counted separately. The generation task is scored by an LLM judge on alignment, helpfulness, depth, and tone, validated against human ratings at over 90 percent agreement.

What would settle it

Take the released INDIC-BIAS templates, restrict to items the model answered by removing refusals, and recompute SAR as a two-way target-versus-distractor rate; if the model-level averages fall to around 0.5 for caste and religion, the paper's claim that models reinforce stereotypes in a majority of cases would not survive.

Watch

Extended reading notes

Core claim

The paper's central claim is that, when presented with otherwise identical scenarios that differ only in the identity of the person involved, current large language models are systematically not neutral in the Indian context. Across 14 models and three tasks, models rank marginalized castes such as Dalit and Chamar as more plausible in negative situations and less plausible in positive ones, while privileged castes such as Brahmin show the reverse pattern; tribal identities fare worse when paired with dominant regional groups; and religious minorities receive mixed but often negative treatment. On stereotypes, the paper reports that models associate identities with their listed stereotypes in more than half of the instances on average, with the highest rates in free-form generation, where some models exceed 70 percent. The paper further claims that asking models to explain their reasoning does not consistently reduce these patterns, and that response quality in advice-giving generation tasks is uneven across identities. Taken together, the authors conclude that current LLMs risk both allocative harms, meaning unequal access to opportunities, and representational harms, meaning degrading or stereotyping portrayals, for Indian identities.

Load-bearing premise

The 'over 50%' stereotype result assumes a random model would pick the target identity only one-third of the time, but since refusals are rare the real chance rate is closer to 50%, which would make many reported rates unremarkable.

Editorial extensions

If this is right

  • If the paper is right, any high-stakes use of current LLMs in India, such as scholarship review, hiring, workplace advice, or law-enforcement narratives, carries a systematic anti-marginalized tilt rather than random error.
  • Fairness auditing in India cannot stop at gender and religion; caste, tribe, and region are where the strongest and most consistent rank shifts appear.
  • Refusal is a meaningful fairness behavior: models that decline stereotype or negative-scenario items are not behaving like biased selectors, so refusal rates must be reported and can be a policy lever.
  • Chain-of-thought prompting is not a dependable mitigation because its effect on refusal and choice varies across model families.
  • The released benchmark allows future models and mitigation methods to be scored on the same 20,000 scenarios, making the fairness properties of new LLMs comparable over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 33.33 percent random baseline for SAR is the load-bearing assumption behind the 'majority of cases' headline; since most models rarely refuse, the effective choice is between target and distractor, so a 50 percent baseline would leave many reported values near 0.5 without support.
  • Because the evaluation is conducted in English, the benchmark does not yet say whether the same biases appear in Hindi or other Indian languages; running the templates in multiple languages would separate linguistic from cultural sources of bias.
  • RSM compares ELO ranks across two separate scenario sets, so rank shifts could partly reflect differences in template difficulty or response style; per-construct confidence intervals would help confirm the causal reading.
  • The paper's own limitations leave out intersectional identities; given that real Indian discrimination often concentrates at intersections like caste-plus-religion, an intersectional version of the benchmark is the natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper introduces INDIC-BIAS, an India-centric fairness benchmark spanning 85 identity groups across caste, religion, region, and tribe, with over 1,800 expert-curated topics and a reported 20,000 manually verified scenario templates organized into plausibility, judgment, and generation tasks. The authors evaluate 14 LLMs using ELO ratings and a Rank Shift Metric (RSM) for bias, and a Stereotype Association Rate (SAR) for stereotype reinforcement. The central claims are that LLMs show consistent negative bias against marginalized identities such as Dalit, Bihari, and tribal groups, and that models reinforce stereotypes in over 50% of cases on average. The benchmark construction and the pairwise-choice bias analysis are substantial; however, the headline stereotype result rests on an unjustified random baseline and the RSM analysis lacks uncertainty quantification.

Significance. If the claims hold, this is an important contribution to non-Western fairness evaluation. The benchmark covers an unusually broad set of Indian identities, uses expert sociologists to build the taxonomy, applies human verification to templates, and releases the resource publicly. The pairwise-choice data and the RSM results provide a plausible and internally consistent picture of negative bias against marginalized castes and regions, which is the strongest part of the paper. The stereotype-reinforcement claim, as currently quantified with a 33.33% baseline, is not statistically grounded; correcting that baseline may materially weaken the "majority of cases" claim. The underlying resource remains valuable regardless, and the paper is suitable for a major revision rather than rejection.

major comments (4)
  1. [Section 6.3, Table 13] The statement "Since refusal is the ideal response, the random SAR baseline is 33.33%" is not justified by the task design. In the plausibility and judgment stereotype tasks, each prompt is a two-alternative forced choice between the target identity and one sampled distractor (see Table 1 and Section 4.1); after excluding refusals and ties, which Table 13 reports separately, chance performance is 50%, not 33.33%. The same issue applies to the generation task, where the evaluator judges whether each of two stereotypes was correctly linked to two identities (Section 4.1). Many reported SAR values are close to 0.5 (e.g., Llama-1b Plausible Tribe 0.503, Llama-8b Judgment Caste 0.504, Gemini-Flash Plausible Tribe 0.503, and Gemini-Pro Generation Tribe 0.418), so the headline claim that "models reinforce stereotypes in over 50% of cases on average" is not established. Please recalculate SAR against a 50% null after excluding refusals, report confidence intervals, and revisewise the wording of the abstract and Section 6.3.
  2. [Section 6.2 and Appendix F (Figures 6-11)] The RSM analysis is presented as evidence of "consistent negative bias," but Figures 6 through 11 report only point estimates. Many RSM values are within a few points of zero (for example, most religion entries in Figure 6a), and no confidence intervals, standard errors, or per-identity comparison counts are provided. Because ELO ratings are derived from finite pairwise comparisons, small RSM differences may reflect noise rather than systematic bias. Please add bootstrap confidence intervals or a formal significance test, and report the number of pairwise comparisons contributing to each identity's RSM.
  3. [Section 6.5 and Figure 13b] The specific win-rate numbers reported in this section appear inconsistent with the appendix figure. The text states that in the negative plausible-scenario task, under Criminal & Unlawful activities, Dalit has a win rate of nearly 70% and Brahmin falls below 35%; however, Figure 13b shows Dalit at 0.42 and Brahmin at 0.25 for the Criminal Activities and Lawfulness construct. The value 0.70 appears in the positive-scenario figure (Figure 13a, Public Achievements and Scandals column), not in the negative criminal-activity construct. Please correct the numbers or the figure reference, because this construct-level analysis is used to support the central bias claim.
  4. [Section 5.1] Please clarify whether each identity pair is presented in both orders. The text says that all nC2 identity combinations are generated for plausibility and judgment tasks, but it does not state whether both orderings (A vs B and B vs A) are separately evaluated. With temperature set to 0, LLMs can exhibit a systematic first-option bias; without order counterbalancing, ELO/RSM and SAR estimates could be confounded by position effects. If counterbalancing was performed, state it explicitly; if not, it should be added to the evaluation protocol.
minor comments (5)
  1. [Abstract and Table 9] The abstract claims 20,000 scenario templates, but the accepted-template counts in Table 9 sum to 18,423 (2,280 + 1,128 + 1,150 + 8,580 + 5,285) and no row is given for Stereotype-Generation. Please add the missing row and reconcile the total.
  2. [Section 4.2] The sentence "we then describe the benchmark creation process (§4.2) and finally outline the human verification process (§4.2)" references the same section twice; the second reference should be to Appendix D.2, where the human verification details actually appear.
  3. [Section 2, References] The citation "B et al., 2022" appears malformed; it should likely read "Senthil Kumar B et al., 2022" or be replaced with the full author list.
  4. [Section 6.3] The phrase "over 50% of cases on average" is ambiguous because it is not clear whether the average is taken over models, identities, tasks, or items. Please define the aggregation explicitly.
  5. [Limitations and Section 5.4] The paper acknowledges the potential for evaluator bias in the LLM-as-a-judge setting and reports 90% human agreement on 250 samples, but the agreement is only reported in aggregate. Reporting agreement separately for bias versus stereotype, and by identity axis, would make the reliability argument more convincing for the generation-task claims.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the benchmark is externally grounded in expert-curated, human-verified content, and the metrics are not fitted to the conclusions.

full rationale

The paper's central claims are derived from an externally constructed benchmark (INDIC-BIAS) whose scenario templates were curated through domain-expert consultation and manually verified by annotators. The evaluation metrics (ELO, RSM, SAR) are defined independently of the conclusions and are not fitted parameters: ELO and RSM are computed from model choices in pairwise comparisons, and SAR is a measured association rate. The stated random baseline of 33.33% for SAR is an analytic assumption about chance behavior, not a definitional identity or a fitted input, so even if the baseline is statistically debatable, that is a correctness or calibration concern rather than circularity. The generation-task evaluations use Llama-3.3-70B as a judge, but the paper reports an independent human-LLM agreement study of over 90% on 250 sampled responses per task, which provides external validation rather than a self-referential loop. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation; the paper's Limitations section also explicitly acknowledges potential evaluator bias and annotator subjectivity. The headline claim that models reinforce stereotypes in over 50% of cases is a statistical interpretation of the measured SAR values, not a quantity equal to the benchmark's own inputs by construction. Thus the derivation chain is self-contained with respect to circularity concerns.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of the curated taxonomy, the valence labels, and the LLM judge; none are independently benchmarked beyond a small validation sample.

free parameters (2)
  • Number of distractor identities per stereotype scenario = 10
    Chosen for computational efficiency (Section 5.1), affects SAR chance baseline and reliability.
  • SAR random baseline = 33.33%
    Assumed three-outcome uniform baseline; not derived from the actual prompt structure.
assumptions (4)
  • domain assumption The expert- and annotator-curated taxonomy (Section 3.3) is an accurate ground-truth set of Indian biases and stereotypes.
    If the taxonomy is wrong or incomplete, the benchmark measures the authors' priors rather than real-world bias.
  • domain assumption Positive and negative scenarios are genuinely opposite in valence, so RSM rank shifts reflect bias rather than content differences.
    RSM assumes that positive/negative scenario pairs are identical except for outcome type.
  • domain assumption LLM-as-a-judge (Llama-3.3-70B) agreement of 90% on 250 samples extends to the full generation evaluation.
    The paper extrapolates from a small sample to the full dataset without accounting for judge bias.
  • standard math Bradley-Terry MLE from Chiang et al. (2024) is correctly applied to compute ELO ratings.
    Statistical ranking method; assumes transitivity and no match-order effects.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes." pith.science (2026). https://pith.science/paper/5QFDTVCE

@misc{pith2026250623111,
  author       = {Pith},
  title        = {Pith review of: FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and Stereotypes},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5QFDTVCE}},
  note         = {Machine review of arXiv:2506.23111}
}
read the original abstract

Existing studies on fairness are largely Western-focused, making them inadequate for culturally diverse countries such as India. To address this gap, we introduce INDIC-BIAS, a comprehensive India-centric benchmark designed to evaluate fairness of LLMs across 85 identity groups encompassing diverse castes, religions, regions, and tribes. We first consult domain experts to curate over 1,800 socio-cultural topics spanning behaviors and situations, where biases and stereotypes are likely to emerge. Grounded in these topics, we generate and manually validate 20,000 real-world scenario templates to probe LLMs for fairness. We structure these templates into three evaluation tasks: plausibility, judgment, and generation. Our evaluation of 14 popular LLMs on these tasks reveals strong negative biases against marginalized identities, with models frequently reinforcing common stereotypes. Additionally, we find that models struggle to mitigate bias even when explicitly asked to rationalize their decision. Our evaluation provides evidence of both allocative and representational harms that current LLMs could cause towards Indian identities, calling for a more cautious usage in practical applications. We release INDIC-BIAS as an open-source benchmark to advance research on benchmarking and mitigating biases and stereotypes in the Indian context.

Figures

Figures reproduced from arXiv: 2506.23111 by the authors.

Figure 1
Figure 1. A snapshot of the taxonomy of social themes [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Overall positive and negative bias in the Judgment task. We plot the average [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Average Stereotype Association Rate (SAR) across all models, tasks, and identity categories. SAR measures how often a model associates an identity with its stereotype. A higher SAR indicates stronger stereo￾typing. Detailed results are provided in Appendix F. primary evaluator for the Generation task. 6 Results and Discussion 6.1 Are models well aligned? To evaluate the robustness of alignment of LLMs, we measure th… view at source ↗
Figures from the paper (19 more)
Figure 5
Figure 5. Figure 5: Overview of the workflow used for creating I [PITH_FULL_IMAGE:figures/full_fig_p023_5.png]
Figure 6
Figure 6. Figure 6: Average RSM for Plausible Scenario task for Religion (Figure 6a) and Caste (Figure 6b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p034_6.png]
Figure 7
Figure 7. Figure 7: Average RSM for Plausible Scenario task for Region (Figure 7a) and Tribe (Figure 7b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p035_7.png]
Figure 8
Figure 8. Figure 8: Average RSM for Judgment task for Religion (Figure 8a) and Caste (Figure 8b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p036_8.png]
Figure 9
Figure 9. Figure 9: Average RSM for Judgment task for Region (Figure 9a) and Tribe (Figure 9b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p037_9.png]
Figure 10
Figure 10. Figure 10: Average RSM for Generation task for Religion (Figure 10a) and Region (Figure 10b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p038_10.png]
Figure 11
Figure 11. Figure 11: Average RSM for Generation task for Caste (Figure 11a) and Tribe (Figure 11b). Positive RSM (denoted by blue) represents positive bias and a negative RSM (denoted by red) indicates a negative bias [PITH_FULL_IMAGE:figures/full_fig_p039_11.png]
Figure 12
Figure 12. Figure 12: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p041_12.png]
Figure 13
Figure 13. Figure 13: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p042_13.png]
Figure 14
Figure 14. Figure 14: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p043_14.png]
Figure 15
Figure 15. Figure 15: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p044_15.png]
Figure 16
Figure 16. Figure 16: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p045_16.png]
Figure 17
Figure 17. Figure 17: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p046_17.png]
Figure 18
Figure 18. Figure 18: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p047_18.png]
Figure 19
Figure 19. Figure 19: Average Win Rate (WR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p048_19.png]
Figure 20
Figure 20. Figure 20: Average Stereotype Association Rate (SAR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p049_20.png]
Figure 21
Figure 21. Figure 21: Average Stereotype Association Rate (SAR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p049_21.png]
Figure 22
Figure 22. Figure 22: Average Stereotype Association Rate (SAR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p050_22.png]
Figure 23
Figure 23. Figure 23: Average Stereotype Association Rate (SAR) for each identity under [PITH_FULL_IMAGE:figures/full_fig_p050_23.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

60 extracted references · 29 canonical work pages

  1. [1]

    Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. 2024. Measuring gender and racial biases in large language models. arXiv preprint arXiv: 2403.15281

  2. [2]

    Senthil Kumar B, Pranav Tiwari, Aman Chandra Kumar, and Aravindan Chandrabose. 2022. https://aclanthology.org/2022.lateraisse-1.1/ Casteism in I ndia, but not racism - a study of bias in word embeddings of I ndian languages . In Proceedings of the First Workshop on Language Technology and Resources for a Fair, Inclusive, and Safe Society within the 13th L...

  3. [3]

    Somnath Banerjee, Sayan Layek, Hari Shrawgi, Rajarshi Mandal, Avik Halder, Shanu Kumar, Sagnik Basu, Parag Agrawal, Rima Hazra, and Animesh Mukherjee. 2024. Navigating the cultural kaleidoscope: A hitchhiker's guide to sensitivity in large language models. arXiv preprint arXiv: 2410.12880

  4. [4]

    Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: from allocative to representational harms in machine learning. special interest group for computing. Information and Society (SIGCIS), 2

  5. [5]

    Shaily Bhatt, Sunipa Dev, Partha Talukdar, Shachi Dave, and Vinodkumar Prabhakaran. 2022 a . Cultural re-contextualization of fairness research in language technologies in india. arXiv preprint arXiv: 2211.11206

  6. [6]

    Shaily Bhatt, Sunipa Dev, Partha Talukdar, Shachi Dave, and Vinodkumar Prabhakaran. 2022 b . https://doi.org/10.18653/v1/2022.aacl-main.55 Re-contextualizing fairness in NLP : The case of I ndia . In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on...

  7. [7]

    Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. Language (technology) is power: A critical survey of" bias" in nlp. arXiv preprint arXiv:2005.14050

  8. [8]

    Ralph Allan Bradley and Milton E. Terry. 1952. http://www.jstor.org/stable/2334029 Rank analysis of incomplete block designs: I. the method of paired comparisons . Biometrika, 39(3/4):324--345

Show all 60 references
  1. [9]

    Zou, Rachel Rudinger, and Hal Daum \'e

    Yang Trista Cao, Anna Sotnikova, Jieyu Zhao, Linda X. Zou, Rachel Rudinger, and Hal Daum \'e . 2023. https://api.semanticscholar.org/CorpusID:266174569 Multilingual large language models leak human stereotypes across language boundaries . ArXiv, abs/2312.07141

  2. [10]

    Gonzalez, and Ion Stoica

    Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. https://doi.org/10.48550/arXiv.2403.04132 Chatbot arena: An open platform for evaluating llms by...

  3. [11]

    Cohen and Sumit Ganguly

    Benjamin B. Cohen and Sumit Ganguly. 2014. https://doi.org/10.1080/14736489.2014.964626 Introduction: Regions and regionalism in india . India Review, 13(4):313--320

  4. [12]

    Andrew M. Colman. 2015. A Dictionary of Psychology. Oxford Quick Reference. Oxford University Press, Oxford, UK

  5. [13]

    they are uncultured

    Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, and Tanushree Mitra. 2024. https://doi.org/10.48550/arXiv.2405.05378 "they are uncultured": Unveiling covert harms and social threats in llm generated conversations . Conference on Empirical Methods in...

  6. [14]

    Ashwini Deshpande. 2011. The Grammar of Caste: Economic Discrimination in Contemporary India. Oxford University Press, New Delhi

  7. [15]

    Sunipa Dev, Jaya Goyal, Dinesh Tewari, Shachi Dave, and Vinodkumar Prabhakaran. 2024. Building socio-culturally inclusive stereotype resources with community engagement. Advances in Neural Information Processing Systems, 36

  8. [16]

    Nicholas B. Dirks. 2001. Castes of Mind: Colonialism and the Making of Modern India. Princeton University Press, Princeton, NJ

  9. [17]

    Wenchao Dong, Assem Zhunis, Dongyoung Jeong, Hyojin Chin, Jiyoung Han, and Meeyoung Cha. 2024 a . Persona setting pitfall: Persistent outgroup biases in large language models arising from social identity adoption. arXiv preprint arXiv: 2409.03843

  10. [18]

    Yi Dong, Ronghui Mu, Yanghao Zhang, Siqi Sun, Tianle Zhang, Changshun Wu, Gaojie Jin, Yi Qi, Jinwei Hu, Jie Meng, Saddek Bensalem, and Xiaowei Huang. 2024 b . Safeguarding large language models: A survey. arXiv preprint arXiv: 2406.02622

  11. [19]

    Arpad E. Elo. 1978. https://api.semanticscholar.org/CorpusID:142610973 The rating of chessplayers, past and present

  12. [20]

    Fatma Elsafoury. 2023. https://api.semanticscholar.org/CorpusID:261048854 Systematic offensive stereotyping (sos) bias in language models . ArXiv, abs/2308.10684

  13. [21]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen Ahmed. 2023. https://api.semanticscholar.org/CorpusID:261530629 Bias and fairness in large language models: A survey . Computational Linguistic...

  14. [22]

    Guo and Aylin Caliskan

    W. Guo and Aylin Caliskan. 2020. https://doi.org/10.1145/3461702.3462536 Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases . AAAI/ACM Conference on AI, Ethics, and Society

  15. [23]

    Yufei Guo, Muzhe Guo, Juntao Su, Zhou Yang, Mengqiu Zhu, Hongfei Li, Mengyang Qiu, and Shuo Shuo Liu. 2024. https://api.semanticscholar.org/CorpusID:274130807 Bias in large language models: Origin, evaluation, and mitigation . ArXiv, abs/2411.10915

  16. [24]

    Gadiraju, Aditya Vashistha, Vivek Seshadri, and Kalika Bali

    Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, U. Gadiraju, Aditya Vashistha, Vivek Seshadri, and Kalika Bali. 2024. https://doi.org/10.1145/3630106.3659017 Akal badi ya bias: An exploratory study of gender bias in hind...

  17. [25]

    Rishav Hada, Agrima Seth, Harshita Diddee, and Kalika Bali. 2023. ''fifty shades of bias'': Normative ratings of gender bias in gpt generated english text. arXiv preprint arXiv: 2310.17428

  18. [26]

    Fraser, Anahita Bhiwandiwalla, and Svetlana Kiritchenko

    Phillip Howard, Kathleen C. Fraser, Anahita Bhiwandiwalla, and Svetlana Kiritchenko. 2024. https://api.semanticscholar.org/CorpusID:270123617 Uncovering bias in large vision-language models at scale with counterfactuals . ArXiv, abs/2405.20152

  19. [27]

    Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier, Sander van der Linden, and Jon Roozenbeek. 2025. https://doi.org/10.1038/s43588-024-00741-1 Generative language models exhibit social identity biases . Nature Computational Science

  20. [28]

    Sullam Jeoung, Yubin Ge, and Jana Diesner. 2023. https://doi.org/10.48550/arXiv.2310.13673 Stereomap: Quantifying the awareness of human-like stereotypes in large language models . Conference on Empirical Methods in Natural Language Processing

  21. [29]

    Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. https://doi.org/10.18653/v1/2023.acl-long.548 S ee GULL : A stereotype benchmark with broad geo-cultural coverage leveraging generative models . In Proceedings of...

  22. [30]

    Ishika Joshi, Ishita Gupta, Adrita Dey, and Tapan Parikh. 2024. https://api.semanticscholar.org/CorpusID:272770139 'since lawyers are males..': Examining implicit gender bias in hindi language generation by llms . ArXiv, abs/2409.13484

  23. [31]

    Lee Jussim, Jarret T Crawford, and Rachel S Rubinstein. 2015. Stereotype (in) accuracy in perceptions of groups and individuals. Current Directions in Psychological Science, 24(6):490--497

  24. [32]

    Bean, Hannah Rose Kirk, and Scott A

    Khyati Khandelwal, Manuel Tonneau, Andrew M. Bean, Hannah Rose Kirk, and Scott A. Hale. 2024. https://doi.org/10.1145/3677525.3678666 Indian-bhed: A dataset for measuring india-centric biases in large language models . In Proceedings of the 2024 International Conference on Inf...

  25. [33]

    Jhanvee Khola, Shrujal Bansal, Khushi Punia, Rishika Pal, and Rahul Sachdeva. 2024. https://doi.org/10.1109/CONECCT62155.2024.10677324 Comparative analysis of bias in llms through indian lenses . In 2024 IEEE International Conference on Electronics, Computing and Communication...

  26. [34]

    Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo

    Seungone Kim, Juyoung Suk, Shayne Longpre, Bill Yuchen Lin, Jamin Shin, S. Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2024. https://doi.org/10.48550/arXiv.2405.01535 Prometheus 2: An open source language model specialized in evaluating other language m...

  27. [35]

    Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R Rickford, Dan Jurafsky, and Sharad Goel. 2020. Racial disparities in automated speech recognition. Proceedings of the national academy of sciences, 117(14):7684--7689

  28. [36]

    Tao Li, Tushar Khot, Daniel Khashabi, Ashish Sabharwal, and Vivek Srikumar. 2020. Unqovering stereotyping biases via underspecified questions. arXiv preprint arXiv:2010.02428

  29. [37]

    Yingji Li, Mengnan Du, Rui Song, Xin Wang, and Ying Wang. 2023. A survey on fairness in large language models. arXiv preprint arXiv: 2308.10149

  30. [38]

    Michael Madaio, Lisa Egede, Hariharan Subramonyam, Jennifer Wortman Vaughan, and Hanna Wallach. 2022. Assessing the fairness of ai systems: Ai practitioners' processes, challenges, and needs for support. Proceedings of the ACM on Human-Computer Interaction, 6(CSCW1):1--26

  31. [39]

    Marta Marchiori Manerba, Karolina Stanczak, Riccardo Guidotti, and Isabelle Augenstein. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.812 Social bias probing: Fairness benchmarking for language models . In Proceedings of the 2024 Conference on Empirical Methods in Natural ...

  32. [40]

    Arnab Mukherji and Anjan Mukherji. 2012. Bihar: What went wrong? and what changed?

  33. [41]

    Tirtha Prasad Mukhopadhyay. 2022. Racial prejudice and gender discrimination against northeast indians. DOI https://dx. doi. org/10.21659/rupkatha

  34. [42]

    Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...

  35. [43]

    Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...

  36. [44]

    Tarek Naous, Michael Joseph Ryan, and Wei Xu. 2023. https://doi.org/10.48550/arXiv.2305.14456 Having beer after prayer? measuring cultural bias in large language models . Annual Meeting of the Association for Computational Linguistics

  37. [45]

    Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R Bowman. 2021. Bbq: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193

  38. [46]

    Pew Research . 2021. https://www.pewresearch.org/religion/2021/09/21/population-growth-and-religious-composition/ Population growth and religious composition . Accessed: 2025-02-07

  39. [47]

    PIB . 2023. https://pib.gov.in/PressReleasePage.aspx?PRID=1887716 Press release . Accessed: 2025-02-07

  40. [48]

    Krithika Ramesh, Sunayana Sitaram, and Monojit Choudhury. 2023. Fairness in language models beyond english: Gaps and challenges. arXiv preprint arXiv: 2302.12578

  41. [49]

    Sahoo, Pranamya Prashant Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, and Pushpak Bhattacharyya

    Nihar R. Sahoo, Pranamya Prashant Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, and Pushpak Bhattacharyya. 2024. https://doi.org/10.18653/V1/2024.NAACL-LONG.487 Indibias: A benchmark dataset to measure social biases in language models for indian context . In...

  42. [50]

    Nithya Sambasivan, Erin Arnesen, Ben Hutchinson, Tulsee Doshi, and Vinodkumar Prabhakaran. 2021. Re-imagining algorithmic fairness in india and beyond. arXiv preprint arXiv: 2101.09995

  43. [51]

    Song Wang, Peng Wang, Tong Zhou, Yushun Dong, Zhen Tan, and Jundong Li. 2025. https://openreview.net/forum?id=IUmj2dw5se CEB : Compositional evaluation benchmark for fairness in large language models . In The Thirteenth International Conference on Learning Representations

  44. [52]

    Xuezhi Wang, Haohan Wang, and Diyi Yang. 2022. https://doi.org/10.18653/v1/2022.naacl-main.339 Measure and improve robustness in NLP models: A survey . In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human L...

  45. [53]

    Ishaan Watts, Varun Gumma, Aditya Yadavalli, Vivek Seshadri, Manohar Swaminathan, and Sunayana Sitaram. 2024. Pariksha: A large-scale investigation of human-llm evaluator agreement on multilingual and multi-cultural data. arXiv preprint arXiv: 2406.15053

  46. [54]

    Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020. Measuring and reducing gendered correlations in pre-trained models. arXiv preprint arXiv: 2010.06032

  47. [55]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824--24837

  48. [56]

    Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. Gender bias in coreference resolution: Evaluation and debiasing methods. NAACL

  49. [57]

    Jinman Zhao, Yitian Ding, Chen Jia, Yining Wang, and Zifan Qian. 2024. https://api.semanticscholar.org/CorpusID:268201541 Gender bias in large language models across multiple languages . ArXiv, abs/2403.00277

  50. [58]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstr...

  51. [59]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...

  52. [60]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.