Pith. sign in

REVIEW 1 major objections 4 minor 19 references

Artificial Intelligence Fairness in the Context of Accessibility Research on Intelligent Systems for People who are Deaf or Hard of Hearing

T0 review · 1 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read AI fairness that overlooks people with disabilities will systematically disadvantage them once AI systems are deployed.

desk verdict Useful position paper that extends AI fairness discourse to disability; the training-data remedy rests on an unproven causal assumption, but the core argument and evaluation-metrics insight still make it worth publishing. read the letter →

arxiv 1908.10414 v2 pith:22GZBHRP submitted 2019-08-27 cs.HC

classification cs.HC
keywords artificialintelligencefairnesspeoplewithdisabilitiesdeaforhardofhearingautomaticspeechrecognitioncaptioninginterpretabilitytrainingdata
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a position statement arguing that AI fairness research has largely overlooked people with disabilities, and that this omission will have concrete costs as AI systems spread. Drawing on studies of speech recognition, automatic captioning, sign-language animation, and voice assistants, it claims that mainstream AI systems perform poorly for deaf and hard-of-hearing (DHH) users, that current evaluation metrics and black-box models hide these failures, and that researchers and companies have an ethical responsibility to include disabled users in data, design, and deployment decisions. The paper's central assertion is that if AI-based tools are deployed in critical or popular applications, some groups of people—including people with disabilities—will be disadvantaged. A sympathetic reader should care because the paper converts an abstract fairness debate into a specific, testable demand: disability should be a first-class axis of AI fairness, not an afterthought.

What carries the argument

The argument is carried by the automatic speech recognition (ASR) captioning pipeline and the research studies built around it, rather than by a formal mathematical model. The pipeline works as a recurring case: ASR trained mostly on hearing voices performs poorly for DHH users; adding a live captioning system changes how hearing users speak, which means deployment itself alters the speech distribution the system must handle; and DHH users' willingness to use imperfect captions depends on seeing confidence signals. The paper generalizes from these observations to a mechanism that applies across AI fairness for disability: training data determines who is served, model opacity blocks oversight, evaluation metrics steer research effort, and new AI interfaces create new disabling requirements that are harder to define and compensate for than earlier rule-based technologies.

What would settle it

Collect or assemble a large, diverse corpus of speech from deaf and hard-of-hearing speakers, train an ASR model on it together with hearing speech, and measure word error rate on held-out DHH speakers. If the error gap over hearing speakers remains large even with substantial DHH training data, the claim that missing training data is the central cause of the fairness gap would be falsified.

Watch

Extended reading notes

Core claim

The paper claims that AI systems inherit biases from the people left out of their training data, and that deaf and hard-of-hearing users are one such left-out group: automatic speech recognition works poorly on DHH voices, emotion-detection systems can misread facial expressions used in sign language as anger, and pedestrian detectors may fail to recognize people using wheelchairs. Beyond data inclusion, the paper argues that fairness for disabled users also requires interpretable models, evaluation metrics that reflect real user comprehension rather than convenient technical scores, honest public communication about system limits to prevent premature replacement of human interpreters, and research on how AI changes the behaviors that users must perform. The authors use their own work on ASR-based live captioning to ground each issue, showing that DHH users will accept imperfect automatic captions when they can judge the system's confidence, and that the presence of such systems changes how hearing speakers talk. The upshot is a research agenda in which disability is a standard category in AI fairness analysis and accessibility researchers take on explicit ethical duties when building or deploying AI-based access technologies.

Load-bearing premise

The load-bearing premise is that ASR's poor performance on deaf and hard-of-hearing voices comes mainly from those voices being absent from training data; if the gap were caused by something about how deaf and hard-of-hearing people naturally speak that no amount of extra training examples could fix, the paper's main remedy would not close the fairness gap.

Editorial extensions

If this is right

  • If ASR systems trained without DHH speech are deployed in voice assistants or automatic captions, deaf and hard-of-hearing users will receive systematically worse service than hearing users, which is itself an AI fairness harm.
  • Fairness evaluations of AI systems should report performance separately for people with disabilities, not only across race and gender, if the goal is to prevent this kind of harm.
  • Metrics such as word error rate can mislead when they do not track what DHH users actually understand, so user-centered metrics should be part of the standard evaluation of captioning systems.
  • Because an AI system can change the behavior of the people using it, training and evaluation data should be collected in real deployment contexts rather than assumed to match existing speech corpora.
  • Researchers and companies should state limits of AI-based access technologies honestly, so cost-saving decision makers do not replace human interpreters with systems that are not ready for critical settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's central remedy would be to train an ASR model on a large, diverse DHH speech corpus and measure whether the error gap on held-out DHH voices closes; if it does not, missing training data is not the whole story.
  • The same training-data fairness argument extends beyond ASR to any AI system whose input distribution excludes disabled users, such as image classifiers that must recognize wheelchairs, canes, or atypical body movements.
  • A policy consequence left implicit in the paper is that procurement rules and accessibility regulations could require disability-disaggregated fairness audits before AI-based access technologies are deployed in education, healthcare, or government.
  • The finding that captioning changes hearing speakers' behavior suggests a design direction the paper only hints at: automatic captioning could be deliberately used to encourage slower, clearer speech, turning a fairness concern into an accessibility intervention.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 4 minor

Summary. This position paper argues that AI fairness research and practice must explicitly include people with disabilities, with a focus on people who are Deaf or Hard of Hearing (DHH). Drawing on the authors' prior work in accessibility and human-computer interaction, the paper identifies five interrelated issues: the need to include data from people with disabilities in AI training sets, the lack of interpretability of AI systems, the ethical responsibilities of researchers and companies, the need for evaluation metrics that reflect the needs of DHH users, and the ways that AI systems alter human behavior and the skills valued in society. The central assertion is that, if current AI-based tools are deployed in critical or popular applications, they will systematically disadvantage some groups, including people with disabilities. The paper is framed as a commentary or position piece rather than an empirical study, using illustrative examples from the authors' own research on ASR-based captioning, sign-language animation, and evaluation methods.

Significance. If accepted, the paper makes a timely and important contribution by connecting AI fairness with accessibility research and by giving concrete examples of how AI systems can disadvantage DHH users. Its central message—that disability should be an explicit dimension in AI fairness debates—is well motivated and supported by external literature such as the Gender Shades study and by community statements like the WFD/WASLI declaration on signing avatars. The paper also proposes concrete responsibilities for researchers, such as avoiding overclaimed press releases and developing user-centered evaluation metrics. However, the evidence base is largely self-referential: most examples come from the authors' prior publications, and the key causal claim about training data is asserted rather than demonstrated. These limitations mean the paper is best read as a programmatic call to action rather than as a settled empirical analysis, and the strength of its central claim rests on the plausibility of the examples rather than on systematic evidence.

major comments (1)
  1. [Need for Inclusion in Training Data] The paper states that poor ASR performance on DHH voices is 'likely due to a lack of inclusion of speech from people who are DHH in the training data sets' and then presents greater data diversity as the corresponding remedy. This is a causal hypothesis that is not supported by any cited evidence or error analysis, and it is load-bearing for the paper's call for training-data inclusion. If poor ASR performance instead arises from inherent acoustic or articulatory characteristics of DHH speech that are not simply a matter of under-representation in the feature space, then adding DHH speech to corpora may not close the fairness gap; other remedies such as speaker adaptation, multi-modal interfaces, or alternative interaction modalities would be required. Because the paper's broader argument about disadvantage does not collapse, the issue is not fatal, but the authors should either cite prior empirical work testing this causal pathway, present their own error analyses comparing DHH and hearing speakers under controlled conditions, or explicitly reframe the claim as an open research question rather than a likely cause. This revision is necessary to make the recommended remedy credible.
minor comments (4)
  1. [Lack of Interpretability] There is a typo: 'However; more research is needed' should read 'However, more research is needed.'
  2. [Ethical Responsibility of Researchers and Experts] The narrative about the authors' decision to include the WFD/WASLI statement in their ASSETS'18 submission is informative but written in a personal voice that is slightly informal for a journal article; consider condensing it and focusing more on the generalizable lesson.
  3. [Introduction] The paper never defines what it means by AI 'fairness,' even though the term has multiple technical definitions in the literature; adding a sentence that clarifies whether the authors adopt a distributive, procedural, or capability-based notion would help position the arguments.
  4. [Need for Appropriate Evaluation Metrics] The discussion of the proposed evaluation metric in [12] would benefit from a brief explanation of what the metric measures and why it correlates better with DHH user opinions than WER, since the metric is not described in the text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: a position paper whose self-citations are illustrative, not load-bearing.

full rationale

This is a position paper rather than a formal derivation; it contains no equations, fitted parameters, or constructed predictions that could reduce to inputs. The central claim—that AI systems can disadvantage people with disabilities—is supported by external evidence (e.g., Gender Shades [7], ASR bias by gender/accent [4,19], the WFD/WASLI statement [20]) and by the authors' prior empirical studies. Those prior studies are used as examples and illustrations, not as premises that are then repackaged as conclusions. The 'Need for Inclusion in Training Data' section contains an explicitly hedged causal hypothesis ('likely due to a lack of inclusion of speech from people who are DHH in the training data sets') about why ASR performs poorly for DHH voices; this is an unverified assumption about causation and a potential correctness risk, but it is not a circular step because the paper's main normative argument does not reduce to that hypothesis. No uniqueness theorem, ansatz, or fitted parameter is imported from the authors' prior work in a way that forces the conclusions. Therefore, although the paper is self-citational in places, the self-citations are not load-bearing in the reductionist sense, and no significant circularity is present.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper makes no testable predictions and introduces no empirical quantities. Its argument rests on the representativeness of self-cited studies and on common assumptions about AI interpretability and training-data causality.

assumptions (3)
  • domain assumption The poor performance of ASR on DHH voices is primarily due to missing DHH speech in training data.
    Section 'Need for Inclusion in Training Data' states the poor performance is 'likely due to a lack of inclusion of speech from people who are DHH in the training data sets.' This assumption drives the paper's main recommendation.
  • domain assumption The authors' prior studies on ASR captioning and sign-language animation are valid and representative of AI-based access technology research.
    The paper generalizes from these self-cited examples to broad claims about AI fairness for people with disabilities, especially in sections on interpretability and evaluation metrics.
  • domain assumption Deep-learning AI systems are effectively black boxes that users, regulators, and even experts cannot easily interpret.
    Section 'Lack of Interpretability' assumes this widely held view; the paper's call for interpretability tools depends on it.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Artificial Intelligence Fairness in the Context of Accessibility Research on Intelligent Systems for People who are Deaf or Hard of Hearing." pith.science (2026). https://pith.science/paper/22GZBHRP

@misc{pith2026190810414,
  author       = {Pith},
  title        = {Pith review of: Artificial Intelligence Fairness in the Context of Accessibility Research on Intelligent Systems for People who are Deaf or Hard of Hearing},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/22GZBHRP}},
  note         = {Machine review of arXiv:1908.10414}
}
read the original abstract

We discuss issues of Artificial Intelligence (AI) fairness for people with disabilities, with examples drawn from our research on human-computer interaction (HCI) for AI-based systems for people who are Deaf or Hard of Hearing (DHH). In particular, we discuss the need for inclusion of data from people with disabilities in training sets, the lack of interpretability of AI systems, ethical responsibilities of access technology researchers and companies, the need for appropriate evaluation metrics for AI-based access technologies (to determine if they are ready to be deployed and if they can be trusted by users), and the ways in which AI systems influence human behavior and influence the set of abilities needed by users to successfully interact with computing systems.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

19 extracted references · 12 canonical work pages

  1. [1]

    Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz

    Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N. Bennett, Kori Inkpen, Jaime Teevan, Ruth Kikin-Gil, and Eric Horvitz. 2019. Guidelines for Human-AI Interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI '19). ACM, New York, NY, USA, Pape...

  2. [2]

    Association for Computing Machinery. 2019. ACM Code of Ethics and Professional Conduct. Retrieved from http://www.acm.org/code-of-ethics on July 1, 2019

  3. [4]

    Erbes, D

    Mohamed Faouzi BenZeghiba, Renato De Mori, Olivier Deroo, Stephane Dupont, T. Erbes, D. Jouvet, Luciano Fissore, Pietro Laface, Alfred Mertins, Christophe Ris, Richard Rose. 2007. Automatic speech recognition and speech variability: A review. Speech communication. 49(10-11). 763-786. DOI: https://doi.org/10.1016/j.specom.2007.02.006

  4. [6]

    Larwan Berke, Sushant Kafle, and Matt Huenerfauth

  5. [7]

    Joy Buolamwini and Timnit Gebru. 2018. Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the First Conference on Fairness, Accountability and Transparency. New York, NY, USA. http://proceedings.mlr.press/v81/buolamwini18a.html

  6. [8]

    Consortium for Citizens with Disabilities. 2018. CCD Transportation Task Force Autonomous Vehicle Principles. Retrieved from http://www.c-c- d.org/fichiers/CCD-Transp-TF-AV-Principles- 120318.pdf on July 1, 2019

  7. [9]

    Abraham Glasser. 2019. Automatic Speech Recognition Services: Deaf and Hard-of-Hearing Usability. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (CHI EA '19). ACM, New York, NY, USA, Paper SRC06, 6 pages. DOI: https://doi.org/10.1145/3290607.3308461

  8. [10]

    HireVue. 2019. HireVue Official Site. Retrieved from http://www.hirevue.com on July 2, 2019

Show all 19 references
  1. [11]

    Matt Huenerfauth and Vicki L. Hanson. 2009. Sign Language in the Interface: Access for Deaf Signers. In C. Stephanidis (Ed.), The Universal Access Handbook. Mahwah, NJ: Lawrence Erlbaum Associates, Inc

  2. [12]

    Sushant Kafle and Matt Huenerfauth. 2017. Evaluating the Usability of Automatically Generated Captions for People who are Deaf or Hard of Hearing. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS '17). ACM, New York, NY, ...

  3. [13]

    Rafal Kocielnik, Saleema Amershi, and Paul N. Bennett. 2019. Will You Accept an Imperfect AI?: Exploring Designs for Adjusting End-user Expectations of AI Systems. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI '19). ACM, New York, NY, USA...

  4. [14]

    Riccardo Miotto, Fei Wang, Shuang Wang, Xiaoqian Jiang, Joel T Dudley, Deep learning for healthcare: review, opportunities and challenges, Briefings in Bioinformatics, Volume 19, Issue 6, November 2018, 1236–1246, DOI: https://doi.org/10.1093/bib/bbx044

  5. [15]

    National Science Foundation. 2009. Women, Minorities, and Persons with Disabilities in Science and Engineering, Report No. 09-305. Arlington, VA: National Center for Science and Engineering Statistics

  6. [16]

    National Transportation Safety Board. 2018. Preliminary Report Highway HWY18MH010. Retrieved from https://www.ntsb.gov/investigations/AccidentReports/ Reports/HWY18MH010-prelim.pdf on July 1, 2019

  7. [17]

    Matthew Seita, Khaled Albusays, Sushant Kafle, Michael Stinson, and Matt Huenerfauth. 2018. Behavioral Changes in Speakers who are Automatically Captioned in Meetings with Deaf or Hard-of-Hearing Peers. In Proceedings of the 20th International ACM SIGACCESS Conference on Compu...

  8. [18]

    Irene Rogan Shaffer. 2018. Exploring the Performance of Facial Expression Recognition Technologies on Deaf Adults and Their Children. In Proceedings of the 20th International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS '18). ACM, New York, NY, USA, 474-476....

  9. [19]

    Rachel Tatman. 2017. Gender and dialect bias in YouTube’s automatic captions. In Proceedings of the First ACL Workshop on Ethics in Natural Language Processing, ACL, 53-59. http://www.ethicsinnlp.org/ workshop/pdf/EthNLP06.pdf

  10. [20]

    World Federation of the Deaf. 2018. WFD and WASLI Issue Statement on Signing Avatars. Retrieved from https://wfdeaf.org/news/wfd-wasli-issue-statement- signing-avatars/ on March 20, 2018

  11. [2018]

    In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI '18)

    Methods for Evaluation of Imperfect Captioning Tools by Deaf or Hard-of-Hearing Users at Different Reading Literacy Levels. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems (CHI '18). ACM, New York, NY, USA, Paper 91, 12 pages. DOI: https://doi.o...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.