Pith. sign in

REVIEW 5 cited by

Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1701.02720 v1 pith:BW5XSFWW submitted 2017-01-10 cs.CL cs.LGstat.ML

classification cs.CLcs.LGstat.ML
keywords cnnsrecognitionspeechend-to-endmodelsnetworksneuralapproach
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs with Hidden Markov Models/Gaussian Mixture Models (HMMs/GMMs) have achieved the state-of-the-art in various benchmarks. Meanwhile, Connectionist Temporal Classification (CTC) with Recurrent Neural Networks (RNNs), which is proposed for labeling unsegmented sequences, makes it feasible to train an end-to-end speech recognition system instead of hybrid settings. However, RNNs are computationally expensive and sometimes difficult to train. In this paper, inspired by the advantages of both CNNs and the CTC approach, we propose an end-to-end speech framework for sequence labeling, by combining hierarchical CNNs with CTC directly without recurrent connections. By evaluating the approach on the TIMIT phoneme recognition task, we show that the proposed model is not only computationally efficient, but also competitive with the existing baseline systems. Moreover, we argue that CNNs have the capability to model temporal correlations with appropriate context information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spatio-temporal Manifold Learning for Human Motions via Long-horizon Modeling

    cs.GR 2019-08 conditional novelty 6.0 of 10

    STRNN, a spatio-temporal recurrent network with part-based skeleton encoding and long-horizon prediction, generates stable human motions for up to 20,000 frames in open-loop mode.

  2. V2S attack: building DNN-based voice conversion from automatic speaker verification

    cs.SD 2019-08 conditional novelty 6.0 of 10

    A voice impersonation system is trained by deceiving a white-box automatic speaker verification model, using an ASR model to preserve content, and it performs comparably to voice conversion trained on only a few targe...

  3. Attention Control with Metric Learning Alignment for Image Set-based Recognition

    cs.CV 2019-08 conditional novelty 5.0 of 10

    An actor-critic reinforcement learning module that assigns dependency-aware weights to images in a set improves set-based and video-based face recognition over independent quality weighting.

  4. Core Placement Optimization of Many-core Brain-Inspired Near-Storage Systems for Spiking Neural Network Training

    cs.AR 2024-11 conditional novelty 4.0 of 10

    A core placement method using deep reinforcement learning is claimed to reduce communication cost and improve throughput for SNN training on many-core near-memory systems in unreleased simulator experiments.

  5. End-to-End ASR for Code-switched Hindi-English Speech

    eess.AS 2019-06 unverdicted novelty 4.0 of 10

    End-to-end ASR for code-switched Hindi-English with <50 hours of data shows gains from multi-task learning and corpus balancing but underperforms cascaded baselines.

Pith tools