Pith. sign in

REVIEW 1 cited by

Automatic recognition of suprasegmentals in speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.01122 v2 pith:DDSWI5X7 submitted 2021-08-02 cs.CL

classification cs.CL
keywords recognitionautomaticimprovesyllablestonalunitsfine-tuninghelpful
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study reports our efforts to improve automatic recognition of suprasegmentals by fine-tuning wav2vec 2.0 with CTC, a method that has been successful in automatic speech recognition. We demonstrate that the method can improve the state-of-the-art on automatic recognition of syllables, tones, and pitch accents. Utilizing segmental information, by employing tonal finals or tonal syllables as recognition units, can significantly improve Mandarin tone recognition. Language models are helpful when tonal syllables are used as recognition units, but not helpful when tones are recognition units. Finally, Mandarin tone recognition can benefit from English phoneme recognition by combining the two tasks in fine-tuning wav2vec 2.0.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.

Pith tools