Pith. sign in

REVIEW 3 cited by

Towards measuring fairness in speech recognition: Fair-Speech dataset

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.12734 v1 pith:3DM5KTBP submitted 2024-08-22 cs.AI cs.CYcs.SDeess.ASstat.ML

Towards measuring fairness in speech recognition: Fair-Speech dataset

classification cs.AI cs.CYcs.SDeess.ASstat.ML
keywords datasetmodelsspeechacrossdemographicfair-speechfairnessrecognition
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The current public datasets for speech recognition (ASR) tend not to focus specifically on the fairness aspect, such as performance across different demographic groups. This paper introduces a novel dataset, Fair-Speech, a publicly released corpus to help researchers evaluate their ASR models for accuracy across a diverse set of self-reported demographic information, such as age, gender, ethnicity, geographic variation and whether the participants consider themselves native English speakers. Our dataset includes approximately 26.5K utterances in recorded speech by 593 people in the United States, who were paid to record and submit audios of themselves saying voice commands. We also provide ASR baselines, including on models trained on transcribed and untranscribed social media videos and open source models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Evaluating Bias in Phoneme-Based Automatic Speech Recognition Systems: An Analysis of IPA Transcription Models

    cs.CL 2026-06 unverdicted novelty 5.0

    Evaluation of WhisperIPA and ZIPA reveals persistent performance gaps across languages, accents, gender, ethnicity, and age even after allowing for similar phoneme substitutions.

  2. Cross-Cultural Bias in Mel-Scale Representations: Evidence and Alternatives from Speech and Music

    cs.SD 2026-04 unverdicted novelty 5.0

    Mel-scale features exhibit measurable cultural bias with 12.5% higher WER on tonal languages and 15.7% F1 drop on non-Western music, while adaptive alternatives reduce these gaps substantially.

  3. "OK Aura, Be Fair With Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection

    cs.CL 2026-04 unverdicted novelty 4.0

    Demographics-agnostic training with augmentation and distillation reduces predictive disparity in wake-up word detection by 40-84% across demographic groups.