Pith. sign in

REVIEW 2 cited by

Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.09271 v1 pith:LEONA3UK submitted 2020-05-19 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords accentnativereferenceconversionspeechqualityspeakersystem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving accent conversion applicability, as well as quality. First of all, we assume no reference speech is available at the conversion stage, and hence we employ an end-to-end text-to-speech system that is trained on native speech to generate native reference speech. To improve the quality and accent of the converted speech, we introduce reference encoders which make us capable of utilizing multi-source information. This is motivated by acoustic features extracted from native reference and linguistic information, which are complementary to conventional phonetic posteriorgrams (PPGs), so they can be concatenated as features to improve a baseline system based only on PPGs. Moreover, we optimize model architecture using GMM-based attention instead of windowed attention to elevate synthesized performance. Experimental results indicate when the proposed techniques are applied the integrated system significantly raises the scores of acoustic quality (30$\%$ relative increase in mean opinion score) and native accent (68$\%$ relative preference) while retaining the voice identity of the non-native speaker.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TokAN: Accent Normalization Using Self-Supervised Speech Tokens

    cs.SD 2026-07 accept novelty 6.0 of 10

    TokAN maps L2 speech tokens to L1-like tokens via an autoregressive converter plus GRPO rewards, cutting WER to 9.23% on seven English accents without natural parallel L1-L2 recordings.

  2. Controllable Accent Normalization via Discrete Diffusion

    eess.AS 2026-03 conditional novelty 6.0 of 10

    Masked discrete diffusion over SSL speech tokens plus a Common Token Predictor yields the lowest WER among compared accent-normalization systems and continuous accent-strength control via source-token reuse.

Pith tools