REVIEW 3 cited by
First-Pass Large Vocabulary Continuous Speech Recognition using Bi-Directional Recurrent DNNs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present a method to perform first-pass large vocabulary continuous speech recognition using only a neural network and language model. Deep neural network acoustic models are now commonplace in HMM-based speech recognition systems, but building such systems is a complex, domain-specific task. Recent work demonstrated the feasibility of discarding the HMM sequence modeling framework by directly predicting transcript text from audio. This paper extends this approach in two ways. First, we demonstrate that a straightforward recurrent neural network architecture can achieve a high level of accuracy. Second, we propose and evaluate a modified prefix-search decoding algorithm. This approach to decoding enables first-pass speech recognition with a language model, completely unaided by the cumbersome infrastructure of HMM-based systems. Experiments on the Wall Street Journal corpus demonstrate fairly competitive word error rates, and the importance of bi-directional network recurrence.
Forward citations
Cited by 3 Pith papers
-
CANDLE: CTC-based Arabic Noisy-character Deduplication using a Lightweight Encoder
CANDLE uses CTC on lightweight character encoders for Arabic noise deduplication, reporting 5.37% SER on benchmarks and up to 12.8% tokenizer fertility reduction.
-
Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
A streaming CTC keyword-spotting decoder with cross-layer cosine-similarity refinement reports a 6.8% absolute recall gain over graph-based decoding on the Hey Snips dataset.
-
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
Delayed fusion scores partial ASR hypotheses with a pre-trained LLM only after pruning and at word boundaries, giving lower word error rates than N-best rescoring without retraining the ASR model.
Discussion (0). Continue with ORCID to comment.