REVIEW 4 cited by
Why does CTC result in peaky behavior?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The peaky behavior of CTC models is well known experimentally. However, an understanding about why peaky behavior occurs is missing, and whether this is a good property. We provide a formal analysis of the peaky behavior and gradient descent convergence properties of the CTC loss and related training criteria. Our analysis provides a deep understanding why peaky behavior occurs and when it is suboptimal. On a simple example which should be trivial to learn for any model, we prove that a feed-forward neural network trained with CTC from uniform initialization converges towards peaky behavior with a 100% error rate. Our analysis further explains why CTC only works well together with the blank label. We further demonstrate that peaky behavior does not occur on other related losses including a label prior model, and that this improves convergence.
Forward citations
Cited by 4 Pith papers
-
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
LCS-CTC, a phoneme recognizer trained with similarity-aware LCS alignment masks constraining CTC, outperforms vanilla CTC on all reported PER, WPER, boundary-loss, and articulatory metrics.
-
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
A symmetric blank-selection method for CTC knowledge distillation lets a student model train without any CTC loss and with no loss in word error rate.
-
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
Wildcard CTC on intermediate encoder layers spots user-listed keywords at inference and biases later layers, improving unknown-word F1 by up to 29% relative without retraining or TTS modules.
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
Adding a language-identification block trained with non-peaky CTC and injecting the resulting language posteriors reduces mixed-error rate on Mandarin-English SEAME by about 0.5 to 0.8 percent absolute over the D-MoE ...
Discussion (0). Continue with ORCID to comment.