Achieving Human Parity in Conversational Speech Recognition

A. Stolcke; D. Yu; F. Seide; G. Zweig; J. Droppo; M. Seltzer; W. Xiong; X. Huang

arxiv: 1610.05256 · v2 · pith:2YE3LVDYnew · submitted 2016-10-17 · 💻 cs.CL · eess.AS

Achieving Human Parity in Conversational Speech Recognition

W. Xiong , J. Droppo , X. Huang , F. Seide , M. Seltzer , A. Stolcke , D. Yu , G. Zweig This is my paper

classification 💻 cs.CL eess.AS

keywords humansystemerrorrecognitionspeechachievingacousticautomated

0 comments

read the original abstract

Conversational speech recognition has served as a flagship speech recognition task since the release of the Switchboard corpus in the 1990s. In this paper, we measure the human error rate on the widely used NIST 2000 test set, and find that our latest automated system has reached human parity. The error rate of professional transcribers is 5.9% for the Switchboard portion of the data, in which newly acquainted pairs of people discuss an assigned topic, and 11.3% for the CallHome portion where friends and family members have open-ended conversations. In both cases, our automated system establishes a new state of the art, and edges past the human benchmark, achieving error rates of 5.8% and 11.0%, respectively. The key to our system's performance is the use of various convolutional and LSTM acoustic model architectures, combined with a novel spatial smoothing method and lattice-free MMI acoustic training, multiple recurrent neural network language modeling approaches, and a systematic use of system combination.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
cs.CR 2017-12 unverdicted novelty 7.0

Injecting around 50 poisoned samples with a stealthy trigger creates backdoors in deep learning models achieving over 90% attack success under a weak threat model with no model or data knowledge required.
Auxiliary Interference Speaker Loss for Target-Speaker Speech Recognition
cs.CL 2019-06 unverdicted novelty 6.0

Introduces auxiliary interference speaker loss for target-speaker ASR achieving 6.6% relative WER reduction from 18.06% to 16.87% on mixed speech.
Multilingual Bottleneck Features for Query by Example Spoken Term Detection
cs.CL 2019-06 unverdicted novelty 4.0

Multilingual bottleneck features extracted with residual networks outperform feedforward versions for QbE-STD on QUESST 2014 when trained on GlobalPhone.