REVIEW 5 cited by
Towards End-to-End Speech Recognition with Deep Convolutional Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Convolutional Neural Networks (CNNs) are effective models for reducing spectral variations and modeling spectral correlations in acoustic features for automatic speech recognition (ASR). Hybrid speech recognition systems incorporating CNNs with Hidden Markov Models/Gaussian Mixture Models (HMMs/GMMs) have achieved the state-of-the-art in various benchmarks. Meanwhile, Connectionist Temporal Classification (CTC) with Recurrent Neural Networks (RNNs), which is proposed for labeling unsegmented sequences, makes it feasible to train an end-to-end speech recognition system instead of hybrid settings. However, RNNs are computationally expensive and sometimes difficult to train. In this paper, inspired by the advantages of both CNNs and the CTC approach, we propose an end-to-end speech framework for sequence labeling, by combining hierarchical CNNs with CTC directly without recurrent connections. By evaluating the approach on the TIMIT phoneme recognition task, we show that the proposed model is not only computationally efficient, but also competitive with the existing baseline systems. Moreover, we argue that CNNs have the capability to model temporal correlations with appropriate context information.
Forward citations
Cited by 5 Pith papers
-
Spatio-temporal Manifold Learning for Human Motions via Long-horizon Modeling
STRNN, a spatio-temporal recurrent network with part-based skeleton encoding and long-horizon prediction, generates stable human motions for up to 20,000 frames in open-loop mode.
-
V2S attack: building DNN-based voice conversion from automatic speaker verification
A voice impersonation system is trained by deceiving a white-box automatic speaker verification model, using an ASR model to preserve content, and it performs comparably to voice conversion trained on only a few targe...
-
Attention Control with Metric Learning Alignment for Image Set-based Recognition
An actor-critic reinforcement learning module that assigns dependency-aware weights to images in a set improves set-based and video-based face recognition over independent quality weighting.
-
Core Placement Optimization of Many-core Brain-Inspired Near-Storage Systems for Spiking Neural Network Training
A core placement method using deep reinforcement learning is claimed to reduce communication cost and improve throughput for SNN training on many-core near-memory systems in unreleased simulator experiments.
-
End-to-End ASR for Code-switched Hindi-English Speech
End-to-end ASR for code-switched Hindi-English with <50 hours of data shows gains from multi-task learning and corpus balancing but underperforms cascaded baselines.
Discussion (0). Continue with ORCID to comment.