Pith. sign in

REVIEW 4 cited by

Jasper: An End-to-End Convolutional Neural Acoustic Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.03288 v3 pith:6VKQKA67 submitted 2019-04-05 eess.AS cs.CLcs.LGcs.SD

classification eess.AScs.CLcs.LGcs.SD
keywords jaspermodelarchitectureconvolutionaldecoderend-to-endexternallibrispeech
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we report state-of-the-art results on LibriSpeech among end-to-end speech recognition models without any external training data. Our model, Jasper, uses only 1D convolutions, batch normalization, ReLU, dropout, and residual connections. To improve training, we further introduce a new layer-wise optimizer called NovoGrad. Through experiments, we demonstrate that the proposed deep architecture performs as well or better than more complex choices. Our deepest Jasper variant uses 54 convolutional layers. With this architecture, we achieve 2.95% WER using a beam-search decoder with an external neural language model and 3.86% WER with a greedy decoder on LibriSpeech test-clean. We also report competitive results on the Wall Street Journal and the Hub5'00 conversational evaluation datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning state and proposal dynamics in state-space models using differentiable particle filters and neural networks

    cs.LG 2024-11 conditional novelty 6.0 of 10

    StateMixNN learns particle-filter transition and proposal densities as Gaussian mixtures parameterized by neural networks, trained only on the observation likelihood, and reports improved state recovery on Lorenz 96 a...

  2. TPCNet: Representation learning for HI mapping

    astro-ph.GA 2024-11 conditional novelty 6.0 of 10

    A CNN-Transformer hybrid with sinusoidal positional encoding predicts cold HI fraction and opacity correction from 21-cm emission, outperforming CNN baselines but biased at high column density.

  3. Continual Learning in Machine Speech Chain Using Gradient Episodic Memory

    cs.CL 2024-11 reject novelty 5.0 of 10

    The paper combines machine speech chain text-to-speech replay with gradient episodic memory to let an ASR model learn a noisy speech task without forgetting clean speech, reporting a 40% average CER reduction over fin...

  4. GraphGrad: Efficient Estimation of Sparse Polynomial Representations for General State-Space Models

    stat.CO 2024-11 conditional novelty 5.0 of 10

    A differentiable particle filter with L1 proximal updates estimates sparse polynomial transition functions and interaction graphs for nonlinear state-space models.

Pith tools