REVIEW 3 cited by
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Simple and Effective Zero-shot Cross-lingual Phoneme Recognition
read the original abstract
Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in many cases there is labeled data available for related languages which is not utilized by these methods. This paper extends previous work on zero-shot cross-lingual transfer learning by fine-tuning a multilingually pretrained wav2vec 2.0 model to transcribe unseen languages. This is done by mapping phonemes of the training languages to the target language using articulatory features. Experiments show that this simple method significantly outperforms prior work which introduced task-specific architectures and used only part of a monolingually pretrained model.
Forward citations
Cited by 3 Pith papers
-
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
PhonoQ-2.0 directly predicts structured phonological features from self-supervised models with a gating mechanism, outperforming phoneme baselines by 8+ macro-F1 points on average across in-domain, out-of-domain, and ...
-
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
Harf-Speech delivers phoneme-level Arabic pronunciation scores that correlate 0.79 with certified speech-language pathologists on 40 utterances.
-
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
A modular Arabic phoneme-level pronunciation scorer reaches Pearson 0.791 and ICC 0.659 with three SLPs on 40 utterances, with best speech-to-phoneme PER of 8.92%.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.