REVIEW 4 major objections 6 minor 18 references
Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs
T0 review · 4 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Western LLMs overestimate ableist harm, Indic LLMs underestimate it, and all eight models judge the same ableist comments as less harmful in Hindi than in English.
desk verdict A first cross-cultural ableism audit worth reviewing, but the 'all LLMs more tolerant in Hindi' claim needs translation-equivalence testing and a softer wording. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a comparative harm-assessment audit: a curated corpus of 300 social media comments—100 English, 100 formal Hindi, 100 casual Hindi—derived from an existing ableist speech dataset, scored on 0–10 toxicity and ableism scales with written justifications by both human raters (130 US-based and 45 India-based people with disabilities) and eight LLMs. Hindi's formality registers, marked by second-person pronouns (tū/casual versus āp/formal), are the probe that reveals how models misread sociolinguistic context. Statistical comparisons (Wilcoxon signed-rank, Kruskal-Wallis, Spearman) and open and deductive coding of explanations are used to locate where model ratings diverge from human ratings and where explanations omit culturally specific reasoning.
What would settle it
Have independent Hindi speakers, blind to the study hypothesis, back-translate the Hindi comments into English and rate them for ableism: if the back-translations score systematically lower than the originals, the Hindi leniency effect is at least partly a translation artifact. Separately, recruit a larger regionally stratified sample of Indian PwD and check whether the 45-person group's mean ratings and qualitative framings replicate; if they shift substantially, the ground-truth anchor changes.
Extended reading notes
Core claim
The central discovery is a twofold miscalibration. Western LLMs (GPT-4o, Gemini 2.0 Flash, Claude 3.7 Sonnet, Llama 3.1 70B) cluster at higher toxicity and ableism ratings than Indian PwD, flagging as "inspiration porn" or dehumanizing comments that Indian participants read as well-intentioned encouragement; Indic LLMs (Nanda, Krutrim, Gajendra, Airavata) cluster lower, missing harm in comments about faking disability, invisible conditions, and microaggressions. The paper further finds that all eight models rate ableist comments as less toxic and less ableist in Hindi than in English, with Western models showing large effect sizes, and that models misinterpret Hindi formality registers (ฤู vs. ฤฤรฤี) as rudeness in places where Indian participants read intimacy or care. Indian PwD's explanations centered intent, relationality, resilience, and intersectionality with gender, caste, and class—framings that the paper shows are absent from LLM explanations. The paper concludes that multilingual models are not multicultural and that local disabled people, not Western-centric benchmarks, should set the ground truth for ableist harm.
Load-bearing premise
The headline language effect depends on the manual Hindi translations preserving the same ableist weight and sociolinguistic nuance as the original English comments, and on 45 Indian PwD recruited via Prolific and snowballing giving a stable estimate of Indian cultural norms; neither assumption is back-translated or independently re-measured.
Editorial extensions
If this is right
- Content moderation using Western LLMs would over-remove disability advocacy and emotionally charged critique, because these models systematically over-flag comments that Indian disabled people do not experience as harmful.
- Moderation using Indic LLMs would under-remove ableist content, since these models miss harm in microaggressions and dismiss invisible or intellectual disabilities.
- Ableist comments in Hindi are less likely to be flagged than the same content in English by all eight models, leaving Hindi-speaking PwD more exposed.
- Fine-tuning a base model like Llama on Indian data does not by itself produce cultural adaptation: Nanda tracked Llama's ratings closely, and demographic prompting changed few models' scores.
- Fairness evaluations of harm detection should treat local disabled people's judgments as ground truth rather than relying on Western-centric benchmarks.
Reading between the lines
- A back-translation study with independent raters, blind to the English originals, could test whether the Hindi leniency effect survives translation quality control; absent that check, part of the language gap could reflect translation rather than model bias.
- The same audit design could be extended to other Indic languages and to generative images; the paper's "multilingual but not multicultural" conclusion would predict similar under-detection wherever training data skew Western.
- The 45-person Indian sample recruited via Prolific and snowballing may over-represent urban, English-literate, internet-connected disabled people; a larger regionally stratified sample is a natural next step before treating specific rating means as national norms.
- If the cultural-framing finding generalizes, harm-detection benchmarks themselves need to be pluralized: one universal ableism score is not a well-posed target.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a comparative audit of how eight large language models (four developed in the U.S.: GPT-4o, Gemini 2.0 Flash, Claude 3.7 Sonnet, Llama 3.1 70B; four developed in India: Nanda, Krutrim, Gajendra, Airavata) assess ableist speech, benchmarked against ratings and explanations from 175 people with disabilities (130 in the U.S., 45 in India). The authors translate a publicly available English ableist speech dataset (100 comments) into Hindi in two formality registers (casual and formal), collect toxicity and ableism ratings from models and PwD, and analyze both quantitative ratings and qualitative explanations. They report that Western LLMs overestimate ableist harm relative to Indian PwD, Indic LLMs underestimate it, and that LLMs are more tolerant of ableist speech in Hindi than in English. They also document culturally specific framings among Indian PwD, emphasizing intention, relationality, resilience, and intersections with gender, caste, and class, and argue for grounding AI harm detection in local disability experiences.
Significance. If the findings are robust, this is a valuable and timely contribution to cross-cultural AI fairness and human-centered content moderation. The paper offers the first Hindi ableist speech dataset, a first audit of Indic LLMs for disability bias, and a methodological template for incorporating PwD perspectives into LLM evaluation. The qualitative analysis of Indian PwD interpretations is rich and provides concrete examples of how Western-centric ableism framings miss local cultural context. The manuscript also benefits from a within-model, within-group design and reports effect sizes for many comparisons. However, the headline claim about language-based leniency depends on unvalidated translation equivalence and is broader than the reported statistics, so the central contribution requires additional evidential support before the main conclusions can be taken as established.
major comments (4)
- [Methodology, 'Hindi Translation'] The English-vs-Hindi comparison rests on the assumption that the Hindi translations carry the same ableist weight as the original English comments, but the paper provides no empirical validation of this equivalence. The text states that two native Hindi speakers translated the comments and two additional Hindi/Urdu speakers validated them, 'ensuring that the Hindi comments carried the same weight and intention as the original English texts,' yet no back-translation, no inter-rater agreement statistic, and no quantitative equivalence check (e.g., independent ratings of English and Hindi items by bilingual raters) are reported. This is load-bearing because the paper's most consequential claim—that LLMs are more tolerant of ableism in Hindi—is a direct comparison of scores assigned to English versus Hindi stimuli. If the Hindi versions are systematically milder in ableist force, the observed drop in LLM ratings would be a stimulus artifact rather than evidence of model bias. The concern is amplified by the deliberate manipulation of second-person pronouns (tu/aap) to create casual and formal registers, since the paper's own analyses (Table 2, Hin-C:Hin-F) show that register shifts significantly change both PwD and LLM ratings. I request that the authors provide equivalence evidence, for example by having a separate group of bilingual Hindi-speaking PwD rate the English and Hindi items for perceived ableist harm, or by performing back-translation with independent translators and reporting agreement.
- [Ableism Detection in Hindi (Table 2)] The abstract and Introduction assert that 'all LLMs were more tolerant of ableism when it was expressed in Hindi,' and the Discussion repeats that 'LLMs consistently rated Hindi ableist speech much lower than English speech.' These claims are contradicted by the paper's own within-model statistics in Table 2. For toxicity, Nanda shows no significant English-Hindi difference in either Eng:Hin-C or Eng:Hin-F, and Airavata also shows none. For ableism, Nanda shows no significant difference in either English-Hindi comparison, while Llama and Claude show no significant difference for Eng:Hin-C. The evidence supports a more qualified claim: most LLMs, particularly the four Western models, rated Hindi items lower, but 'all' is not supported by the reported tests. The authors should either revise the wording to reflect which models show significant effects, or report the non-significant directions and explain why they nonetheless regard the pattern as uniform.
- [Methodology, 'From LLMs'] The paper states that 'the prompt was provided in English for all models,' including for the Hindi dataset. This creates a potential confound in the English-vs-Hindi comparisons: lower LLM ratings for Hindi content could reflect degraded processing when the instruction language differs from the target text language, rather than language-specific tolerance of ableism. This is especially relevant for the smaller Indic models, which may be less robust to cross-lingual instruction. The authors should clarify whether the Hindi comments were also accompanied by an English prompt, and if so, test a subset of Hindi items with a matching Hindi prompt to determine whether the observed effect is robust to prompt language. If the prompt language was indeed English, this limitation should be acknowledged and addressed in the interpretation.
- [Findings, 'Ableism Detection in Hindi'] The paper reports many pairwise Wilcoxon signed-rank tests (8 models x 3 language/register comparisons x 2 rating types, plus PwD comparisons) without any correction for multiple comparisons. While the large effect sizes for several Western models make some conclusions robust, the uncorrected p-values inflate the risk of false positives, particularly for marginal entries such as PwD Toxicity Hin-C:Hin-F (Z = -1.86, p < 0.001) with a small effect size (r = 0.21). The authors should either apply a multiple-comparison correction (e.g., Holm-Bonferroni) or explicitly justify why the pattern of results is robust to this concern.
minor comments (6)
- [Abstract] The sentence 'Even more concerning, all LLMs were more tolerant of ableism when it was expressed in Hindi and asserted Western framings of ableist harm' is grammatically and conceptually confusing; 'asserted Western framings of ableist harm' does not clearly attach to the models or the finding. Please rephrase for clarity.
- [Table 2] The table reports absolute Z-values without indicating the direction of the difference; the text must be consulted to know whether a positive Z means higher or lower ratings in the second condition. Please add a note explaining the sign convention, and replace the '-' entries with 'ns' to avoid ambiguity with 'not applicable.'
- [Methodology, Qualitative Coding] The paper reports that two researchers independently analyzed 50 explanations and co-developed a codebook, but no inter-rater reliability statistic (e.g., Cohen's kappa) is reported for the final codebook application. Reporting this would strengthen confidence in the qualitative findings.
- [Related Work] There is a typo: 'and and studies show that' should read 'and studies show that.'
- [Discussion] There is a duplicated word in 'allowing allowing ableism to circulate unchecked'; please correct.
- [Throughout] The paper uses 'ground truth' to refer to the ratings of the 45 Indian PwD. Given the small, non-representative sample (acknowledged in Limitations), I recommend using a less definitive term such as 'reference judgments' or 'local PwD baseline' to avoid overstating the universality of this anchor.
Circularity Check
No significant circularity: LLM–human comparisons use newly collected ratings as ground truth, not fitted parameters or prior labels.
full rationale
The derivation chain is a measurement study: model outputs (zero-shot toxicity/ableism ratings and explanations) are compared with newly collected ratings from 130 US and 45 Indian PwD, with Indian PwD ratings explicitly designated as ground truth ("we consider ratings and explanations provided by Indian PwD as our 'ground truth' data"). No model parameter is fitted to human ratings, and no target quantity is defined in terms of model outputs, so the over/under-estimation and Hindi-tolerance findings are empirical comparisons rather than reductions. The only self-referential input is the reuse of the authors' prior ableist speech dataset (Phutane, Seelam, and Vashistha 2025) as stimuli; its prior labels are not used as the outcome, and the paper's claims are computed from new human and LLM ratings, so this is not load-bearing circularity. The methodology's Hindi translation step relies on the authors' and two validators' judgment that translations "carried the same weight and intention as the original English texts"; no back-translation or inter-rater agreement statistic is reported, and the Limitations section acknowledges the small Indian sample (n=45). These are validity and representativeness limitations, not circular reductions, because the English-vs-Hindi LLM ratings are independent outputs elicited after translation. No uniqueness theorem, ansatz-by-citation, fitted-input-called-prediction, or renamed-known-result pattern is present. Verdict: no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Numeric toxicity and ableism ratings on an 11-point scale are commensurable across human raters and LLMs.
- domain assumption The manually translated Hindi comments are semantically and pragmatically equivalent to the English source comments.
- ad hoc to paper Ratings from 45 Indian PwD are an adequate ground truth for Indian cultural norms of ableism.
Cite this review
Pith. "Pith review of Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs." pith.science (2026). https://pith.science/paper/6B4SQ3P5
@misc{pith2026250716130,
author = {Pith},
title = {Pith review of: Disability Across Cultures: A Human-Centered Audit of Ableism in Western and Indic LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/6B4SQ3P5}},
note = {Machine review of arXiv:2507.16130}
}
read the original abstract
People with disabilities (PwD) experience disproportionately high levels of discrimination and hate online, particularly in India, where entrenched stigma and limited resources intensify these challenges. Large language models (LLMs) are increasingly used to identify and mitigate online hate, yet most research on online ableism focuses on Western audiences with Western AI models. Are these models adequately equipped to recognize ableist harm in non-Western places like India? Do localized, Indic language models perform better? To investigate, we adopted and translated a publicly available ableist speech dataset to Hindi, and prompted eight LLMs--four developed in the U.S. (GPT-4, Gemini, Claude, Llama) and four in India (Krutrim, Nanda, Gajendra, Airavata)--to score and explain ableism. In parallel, we recruited 175 PwD from both the U.S. and India to perform the same task, revealing stark differences between groups. Western LLMs consistently overestimated ableist harm, while Indic LLMs underestimated it. Even more concerning, all LLMs were more tolerant of ableism when it was expressed in Hindi and asserted Western framings of ableist harm. In contrast, Indian PwD interpreted harm through intention, relationality, and resilience--emphasizing a desire to inform and educate perpetrators. This work provides groundwork for global, inclusive standards of ableism, demonstrating the need to center local disability experiences in the design and evaluation of AI systems.
Figures
Reference graph
Works this paper leans on
-
[8]
Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation.Proceedings of the ACM on Human-Computer Interaction, 6(CSCW2): 1–28. Haq, R.; Klarsfeld, A.; Kornau, A.; and Ngunjiri, F. W. 2020. Diversity in India: addressing caste, disability and gender. Equality, Diversity and Inclusion: An International Journal, 39(6): 585–...
work page Pith review arXiv 2020
-
[10]
In Proceedings of the CHI Conference on Human Fac- tors in Computing Systems, 1–15
Challenges to Online Disability Rights Advocacy in India. In Proceedings of the CHI Conference on Human Fac- tors in Computing Systems, 1–15. Honolulu HI USA: ACM. ISBN 9798400703300. Kelion, L. 2019. TikTok suppressed disabled users’ videos. BBC. Keller, R. M.; and Galgay, C. E. 2010. Microaggressive experiences of people with disabilities. In Microaggre...
work page 2019
-
[12]
When Being Unseen from mBERT is just the Be- ginning: Handling New Languages With Multilingual Lan- guage Models. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Compu- tational Linguistics: Human Language Technologies , 448–
work page 2021
-
[15]
In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–13
”I was really, really nervous posting it”: Communi- cating about Invisible Chronic Illnesses across Social Media Platforms. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–13. Glasgow Scot- land Uk: ACM. ISBN 978-1-4503-5970-2. Sannon, S.; Young, J.; Shusas, E.; and Forte, A. 2023. Disability Activism on Social Media: So...
work page 2019
-
[16]
ACM Transactions on Computer-Human Interaction, 26(6): 1–40
Agency of Autistic Children in Technology Re- search—A Critical Literature Review. ACM Transactions on Computer-Human Interaction, 26(6): 1–40. Suresh, V .; and Dyaram, L. 2020. Workplace disability in- clusion in India: review and directions. Management Re- search Review, 43(12). Tacheva, J.; and Ramasubramanian, S. 2023. AI Empire: Unraveling the interl...
arXiv 2020
-
[462]
Online: Association for Computational Linguistics. Naous, T.; Ryan, M. J.; Ritter, A.; and Xu, W. 2024. Having Beer after Prayer? Measuring Cultural Bias in Large Lan- guage Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 16366–16393. Bangkok, Thailand: Associa- tion for Computati...
arXiv 2024
-
[2005]
Fifty-Eight World Health Assembly
Disability, including prevention, management and re- habilitation. Fifty-Eight World Health Assembly
-
[2016]
Disability Divides in India: Evidence from the 2011 Census. PLOS ONE, 11(8): e0159809. Sambasivan, N.; Batool, A.; Ahmed, N.; Matthews, T.; Thomas, K.; Gayt ´an-Lugo, L. S.; Nemer, D.; Bursztein, E.; Churchill, E.; and Consolvo, S. 2019. ”They Don’t Leave Us Alone Anywhere We Go”: Gender and Digital Abuse in South Asia. In Proceedings of the 2019 CHI Conf...
work page 2011
Show all 18 references
-
[2019]
Proceedings of the ACM on Human-Computer Interaction , 3(CSCW): 1–33
”Did You Suspect the Post Would be Removed?”: Un- derstanding User Reactions to Content Removals on Reddit. Proceedings of the ACM on Human-Computer Interaction , 3(CSCW): 1–33. Kaur, S.; Swaminathan, M.; Bali, K.; and Vashistha, A
-
[2020]
In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics, 5454–
Language (Technology) is Power: A Critical Survey of “Bias” in NLP. In Proceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics, 5454–
-
[2021]
Office of Justice Programs Bureau of Justice Statistics
Crime Against Persons with Disabilities, 2009–2019 – Statistical Tables. Office of Justice Programs Bureau of Justice Statistics. Abbo, G. A.; Marchesi, S.; Wykowska, A.; and Belpaeme, T
2009
-
[2022]
In CHI Conference on Hu- man Factors in Computing Systems, 1–20
Trauma-Informed Computing: Towards Safer Tech- nology Experiences for All. In CHI Conference on Hu- man Factors in Computing Systems, 1–20. New Orleans LA USA: ACM. ISBN 978-1-4503-9157-3. Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V .; Wenzek, G.; Guzm ´an, F.; Grave,...
2020
-
[2023]
Version Number: 1
A Prompt Pattern Catalog to Enhance Prompt En- gineering with ChatGPT. Version Number: 1. Whittaker, M.; Alper, M.; Bennett, C. L.; Hendren, S.; Kaz- iunas, L.; Mills, M.; Morris, M. R.; Rankin, J.; Rogers, E.; Salas, M.; et al. 2019. Disability, bias, and AI. Wicks, D. 2017. ...
2019
-
[2024]
In Osman, N.; and Steels, L., eds., Value Engineering in Ar- tificial Intelligence, volume 14520, 83–97
Social Value Alignment in Large Language Models. In Osman, N.; and Steels, L., eds., Value Engineering in Ar- tificial Intelligence, volume 14520, 83–97. Cham: Springer Nature Switzerland. ISBN 978-3-031-58204-2 978-3-031- 58202-8. Series Title: Lecture Notes in Computer Scien...
2022
-
[2025]
Inter- national Journal of Human-Computer Studies, 198: 103468
A critical reflection on the use of toxicity detection algorithms in proactive content moderation systems. Inter- national Journal of Human-Computer Studies, 198: 103468. Watts, I.; Gumma, V .; Yadavalli, A.; Seshadri, V .; Swami- nathan, M.; and Sitaram, S. 2024. PARIKSHA: A ...
2024
-
[3207]
ISBN 978-1-4503- 9385-0
Washington DC USA: ACM. ISBN 978-1-4503- 9385-0. Linxen, S.; Sturm, C.; Br ¨uhlmann, F.; Cassau, V .; Opwis, K.; and Reinecke, K. 2021. How WEIRD is CHI? In Pro- ceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 1–14. Yokohama Japan: ACM. ISBN 978-1-45...
2021
-
[5476]
Blodgett, S
Online: Association for Computational Linguistics. Blodgett, S. L.; and O’Connor, B. 2017. Racial Disparity in Natural Language Processing: A Case Study of Social Me- dia African-American English. ArXiv:1707.00061 [cs]. Brosnan, K.; Gr¨un, B.; and Dolnicar, S. 2021. Cognitive ...
2017 arXiv
-
[6365]
Glazko, K.; Mohammed, Y .; Kosa, B.; Potluri, V .; and Mankoff, J
Barcelona, Spain (Online): International Committee on Computational Linguistics. Glazko, K.; Mohammed, Y .; Kosa, B.; Potluri, V .; and Mankoff, J. 2024. Identifying and Improving Disabil- ity Bias in GPT-Based Resume Screening. In The 2024 ACM Conference on Fairness, Accounta...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.