Pith. sign in

REVIEW 4 cited by

CrisperWhisper: Accurate Timestamps on Verbatim Speech Transcriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16589 v1 pith:Y4BLRL74 submitted 2024-08-29 cs.LG

classification cs.LG
keywords speechverbatimcrisperwhispermodeltimestampstranscriptiontranscriptionsaccurate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We demonstrate that carefully adjusting the tokenizer of the Whisper speech recognition model significantly improves the precision of word-level timestamps when applying dynamic time warping to the decoder's cross-attention scores. We fine-tune the model to produce more verbatim speech transcriptions and employ several techniques to increase robustness against multiple speakers and background noise. These adjustments achieve state-of-the-art performance on benchmarks for verbatim speech transcription, word segmentation, and the timed detection of filler events, and can further mitigate transcription hallucinations. The code is available open https://github.com/nyrahealth/CrisperWhisper.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing

    cs.CL 2026-07 conditional novelty 7.0 of 10

    REDDIT corrects non-speech-induced timestamp drift in autoregressive ASR by editing timestamp targets under cached replay context while anchoring non-timestamp behavior to the frozen base distribution.

  2. Seamless Interaction: Dyadic Audiovisual Motion Modeling and Large-Scale Dataset

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Seamless Interaction provides over 4,000 hours of in-person dyadic video and dyadic audiovisual motion models that generate synchronized face and body behavior from conversational audio and visual inputs.

  3. Word Level Timestamp Generation for Automatic Speech Recognition and Translation

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The paper teaches the Canary ASR and speech-translation model to output word-level start and end timestamps directly using forced-alignment teacher labels.

  4. Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts

    cs.SD 2025-06 conditional novelty 4.0 of 10

    An 8B LLaMa decoder with a Conformer audio encoder generates disfluency tokens and timestamps, and works even when the text hints come from imperfect phoneme or word aligners.

Pith tools