REVIEW 5 cited by
Quantifying Bias in Automatic Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Automatic speech recognition (ASR) systems promise to deliver objective interpretation of human speech. Practice and recent evidence suggests that the state-of-the-art (SotA) ASRs struggle with the large variation in speech due to e.g., gender, age, speech impairment, race, and accents. Many factors can cause the bias of an ASR system. Our overarching goal is to uncover bias in ASR systems to work towards proactive bias mitigation in ASR. This paper is a first step towards this goal and systematically quantifies the bias of a Dutch SotA ASR system against gender, age, regional accents and non-native accents. Word error rates are compared, and an in-depth phoneme-level error analysis is conducted to understand where bias is occurring. We primarily focus on bias due to articulation differences in the dataset. Based on our findings, we suggest bias mitigation strategies for ASR development.
Forward citations
Cited by 5 Pith papers
-
Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
Automatic speech recognition errors on Newcastle English are driven primarily by regional dialect features such as glottalisation, monophthongisation, and local pronouns, rather than by speaker age or gender.
-
Addressing Pitfalls in Auditing Practices of Automatic Speech Recognition Technologies: A Case Study of People with Aphasia
Across six ASR services, speakers with aphasia receive worse transcriptions than controls, and standard audit methods mask within-group disparities and hallucination risks.
-
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
A distillation method that decays teacher loss then applies self-distillation yields a Whisper-derived ASR model with 5x lower latency and slightly better average WER only on in-domain noisy datasets.
-
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
On Dutch end-to-end ASR systems, average word error rate hides large performance gaps across speaker groups, and bias mitigation can lower the average while increasing bias.
-
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
FairASR pretrains a Conformer with InfoNCE plus a gradient-reversed supervised contrastive loss over demographic labels, reducing demographic WER gaps on FairSpeech with small overall WER cost.
Discussion (0). Continue with ORCID to comment.