Pith. sign in

REVIEW 5 cited by

CR-CTC: Consistency regularization on CTC for improved speech recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05101 v4 pith:O6XJAAON submitted 2024-10-07 eess.AS cs.LGcs.SD

classification eess.AScs.LGcs.SD
keywords cr-ctcrecognitionspeechaugmentedconsistencydifferentdistributionsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Connectionist Temporal Classification (CTC) is a widely used method for automatic speech recognition (ASR), renowned for its simplicity and computational efficiency. However, it often falls short in recognition performance. In this work, we propose the Consistency-Regularized CTC (CR-CTC), which enforces consistency between two CTC distributions obtained from different augmented views of the input speech mel-spectrogram. We provide in-depth insights into its essential behaviors from three perspectives: 1) it conducts self-distillation between random pairs of sub-models that process different augmented views; 2) it learns contextual representation through masked prediction for positions within time-masked regions, especially when we increase the amount of time masking; 3) it suppresses the extremely peaky CTC distributions, thereby reducing overfitting and improving the generalization ability. Extensive experiments on LibriSpeech, Aishell-1, and GigaSpeech datasets demonstrate the effectiveness of our CR-CTC. It significantly improves the CTC performance, achieving state-of-the-art results comparable to those attained by transducer or systems combining CTC and attention-based encoder-decoder (CTC/AED). We release our code at https://github.com/k2-fsa/icefall.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hybrid Decoding: Rapid Pass and Selective Detailed Correction for Sequence Models

    eess.AS 2025-08 conditional novelty 6.0 of 10

    A hybrid decoding scheme with a fast TDT draft decoder and selective transformer patches matches baseline word error rate while cutting decoder latency roughly threefold.

  2. DESign: Dynamic Context-Aware Convolution and Efficient Subnet Regularization for Continuous Sign Language Recognition

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A sign language recognition model using context-aware dynamic convolutions and subnetwork CTC regularization reports new state-of-the-art word error rates on PHOENIX14, PHOENIX14-T, and CSL-Daily.

  3. HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A hierarchical transducer with self-distillation and a tuned blank penalty improves joint speech recognition and translation, matching offline attention models on conversational data.

  4. TVTA: Trajectory-Aware Viseme-Guided Temporal Aggregation for Event-Based Lip Reading

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Trajectory-aware local temporal modeling before spatial aggregation plus CTC viseme supervision and EMA consistency lifts DVS-Lip word accuracy to 77.49%.

  5. NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR

    eess.AS 2026-04 unverdicted novelty 4.0 of 10

    NIM4-ASR delivers SOTA ASR performance on public benchmarks using a 2.3B-parameter LLM with multi-stage training, real-time streaming, and million-scale hotword customization via RAG.

Pith tools