Pith. sign in

REVIEW 2 cited by

Unified Mandarin TTS Front-end Based on Distilled BERT Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.15404 v1 pith:OJCSFRRB submitted 2020-12-31 cs.SD cs.CL

classification cs.SDcs.CL
keywords modelfront-endberttasksdistilledencodermandarinmodule
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The front-end module in a typical Mandarin text-to-speech system (TTS) is composed of a long pipeline of text processing components, which requires extensive efforts to build and is prone to large accumulative model size and cascade errors. In this paper, a pre-trained language model (PLM) based model is proposed to simultaneously tackle the two most important tasks in TTS front-end, i.e., prosodic structure prediction (PSP) and grapheme-to-phoneme (G2P) conversion. We use a pre-trained Chinese BERT[1] as the text encoder and employ multi-task learning technique to adapt it to the two TTS front-end tasks. Then, the BERT encoder is distilled into a smaller model by employing a knowledge distillation technique called TinyBERT[2], making the whole model size 25% of that of benchmark pipeline models while maintaining competitive performance on both tasks. With the proposed the methods, we are able to run the whole TTS front-end module in a light and unified manner, which is more friendly to deployment on mobile devices.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A transcript-prompted Whisper model with dictionary-based decoding automatically produces phonemic and prosodic annotations for Japanese audio-transcript pairs, improving Japanese TTS naturalness.

  2. Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Counterfactual gradient edits to a pretrained TTS model's encoder activations can control prosody and correct mispronunciations at inference time, at least on Tacotron 2.

Pith tools