REVIEW 4 cited by
CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We demonstrate that carefully adjusting the tokenizer of the Whisper speech recognition model significantly improves the precision of word-level timestamps when applying dynamic time warping to the decoder's cross-attention scores. We fine-tune the model to produce more verbatim speech transcriptions and employ several techniques to increase robustness against multiple speakers and background noise. These adjustments achieve state-of-the-art performance on benchmarks for verbatim speech transcription, word segmentation, and the timed detection of filler events, and can further mitigate transcription hallucinations. The code is available open https://github.com/nyrahealth/CrisperWhisper.
Forward citations
Cited by 4 Pith papers
-
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.
-
Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset
Seamless Interaction provides over 4,000 hours of in-person dyadic video and dyadic audiovisual motion models that generate synchronized face and body behavior from conversational audio and visual inputs.
-
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.
-
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
An 8B LLaMa decoder with a Conformer audio encoder generates disfluency tokens and timestamps, and works even when the text hints come from imperfect phoneme or word aligners.
Discussion (0). Continue with ORCID to comment.