Pith. sign in

REVIEW 2 cited by

Relaxing the Conditional Independence Assumption of CTC-based ASR by Conditioning on Intermediate Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.02724 v2 pith:EWD36ORE submitted 2021-04-06 eess.AS cs.CLcs.SD

classification eess.AScs.CLcs.SD
keywords intermediatemodelcorpusctc-basedlayermethodassumptionconditional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper proposes a method to relax the conditional independence assumption of connectionist temporal classification (CTC)-based automatic speech recognition (ASR) models. We train a CTC-based ASR model with auxiliary CTC losses in intermediate layers in addition to the original CTC loss in the last layer. During both training and inference, each generated prediction in the intermediate layers is summed to the input of the next layer to condition the prediction of the last layer on those intermediate predictions. Our method is easy to implement and retains the merits of CTC-based ASR: a simple model architecture and fast decoding speed. We conduct experiments on three different ASR corpora. Our proposed method improves a standard CTC model significantly (e.g., more than 20 % relative word error rate reduction on the WSJ corpus) with a little computational overhead. Moreover, for the TEDLIUM2 corpus and the AISHELL-1 corpus, it achieves a comparable performance to a strong autoregressive model with beam search, but the decoding speed is at least 30 times faster.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    MTL for dual-output L2 ASR degrades surface transcription due to encoder-level representational entanglement, with stronger effects in English linked to surface-meaning divergence.

  2. DESign: Dynamic Context-Aware Convolution and Efficient Subnet Regularization for Continuous Sign Language Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sign language recognition model using context-aware dynamic convolutions and subnetwork CTC regularization reports new state-of-the-art word error rates on PHOENIX14, PHOENIX14-T, and CSL-Daily.

Pith tools