Pith. sign in

REVIEW 1 cited by

Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.04710 v1 pith:IMOICZUN submitted 2025-03-06 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords speechchildrecognitionmodelsbaseself-supervisedwavlmlearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Child speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Having explored various architectures for child speech recognition in previous work, in this article we tackle recent self-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM models adapted to phoneme recognition in French child speech, and continue our experiments with the best of them, WavLM base+. We then further adapt it by unfreezing its transformer blocks during fine-tuning on child speech, which greatly improves its performance and makes it significantly outperform our base model, a Transformer+CTC. Finally, we study in detail the behaviour of these two models under the real conditions of our application, and show that WavLM base+ is more robust to various reading tasks and noise levels. Index Terms: speech recognition, child speech, self-supervised learning

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings

    eess.AS 2025-09 conditional novelty 6.0 of 10

    Pre-training HuBERT on 13,164 hours of multilingual child-centered audio and fine-tuning for voice type classification yields 64.6% average F1, beating English-only and adult-speech baselines by 5.9 and 13.2 points.

Pith tools