Pith. sign in

REVIEW 4 cited by

End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1412.1602 v1 pith:Z2PZKXBH submitted 2014-12-04 cs.NE cs.LGstat.ML

classification cs.NEcs.LGstat.ML
keywords recurrentattentioncontinuousdecoderemitsinputmechanismnetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We replace the Hidden Markov Model (HMM) which is traditionally used in in continuous speech recognition with a bi-directional recurrent neural network encoder coupled to a recurrent neural network decoder that directly emits a stream of phonemes. The alignment between the input and output sequences is established using an attention mechanism: the decoder emits each symbol based on a context created with a subset of input symbols elected by the attention mechanism. We report initial results demonstrating that this new approach achieves phoneme error rates that are comparable to the state-of-the-art HMM-based decoders, on the TIMIT dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Combining single-channel speech separation with end-to-end multi-talker ASR improves accuracy on heavily overlapped audio, and a new segment-based output ordering aids offline transcription readability.

  2. Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers

    cs.SD 2025-02 conditional novelty 6.0 of 10

    A transformer speech encoder can learn to internally rearrange audio information into text order, enabling a lightweight decoder trained with simple cross-entropy to nearly match RNN-Transducer accuracy with faster inference.

  3. Cross-Attention End-to-End ASR for Two-Party Conversations

    eess.AS 2019-07 unverdicted novelty 6.0 of 10

    End-to-end ASR model with speaker-specific cross-attention for two-party conversations outperforms standard models on the Switchboard corpus.

  4. End-to-End ASR for Code-switched Hindi-English Speech

    eess.AS 2019-06 unverdicted novelty 4.0 of 10

    End-to-end ASR for code-switched Hindi-English with <50 hours of data shows gains from multi-task learning and corpus balancing but underperforms cascaded baselines.

Pith tools