REVIEW 4 cited by
Jasper: An End-to-End Convolutional Neural Acoustic Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we report state-of-the-art results on LibriSpeech among end-to-end speech recognition models without any external training data. Our model, Jasper, uses only 1D convolutions, batch normalization, ReLU, dropout, and residual connections. To improve training, we further introduce a new layer-wise optimizer called NovoGrad. Through experiments, we demonstrate that the proposed deep architecture performs as well or better than more complex choices. Our deepest Jasper variant uses 54 convolutional layers. With this architecture, we achieve 2.95% WER using a beam-search decoder with an external neural language model and 3.86% WER with a greedy decoder on LibriSpeech test-clean. We also report competitive results on the Wall Street Journal and the Hub5'00 conversational evaluation datasets.
Forward citations
Cited by 4 Pith papers
-
Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks
StateMixNN learns particle-filter transition and proposal densities as Gaussian mixtures parameterized by neural networks, trained only on the observation likelihood, and reports improved state recovery on Lorenz 96 a...
-
TPCNet: Representation learning for HI mapping
A CNN-Transformer hybrid with sinusoidal positional encoding predicts cold HI fraction and opacity correction from 21-cm emission, outperforming CNN baselines but biased at high column density.
-
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
The paper combines machine speech chain text-to-speech replay with gradient episodic memory to let an ASR model learn a noisy speech task without forgetting clean speech, reporting a 40% average CER reduction over fin...
-
GraphGrad: Efficient Estimation of Sparse Polynomial Representations for General State-Space Models
A differentiable particle filter with L1 proximal updates estimates sparse polynomial transition functions and interaction graphs for nonlinear state-space models.
Discussion (0). Continue with ORCID to comment.