REVIEW 2 cited by
Reverb: Open-Source ASR and Diarization from Rev
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Today, we are open-sourcing our core speech recognition and diarization models for non-commercial use. We are releasing both a full production pipeline for developers as well as pared-down research models for experimentation. Rev hopes that these releases will spur research and innovation in the fast-moving domain of voice technology. The speech recognition models released today outperform all existing open source speech recognition models across a variety of long-form speech recognition domains.
Forward citations
Cited by 2 Pith papers
-
Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
Mode-tag conditioning on paired verbatim/intended data makes Whisper produce either verbatim or intended transcripts on demand, with cross-lingual disfluency control and improved word timestamps.
-
Smooth Operators: LLMs Translating Imperfect Hints into Disfluency-Rich Transcripts
An 8B LLaMa decoder with a Conformer audio encoder generates disfluency tokens and timestamps, and works even when the text hints come from imperfect phoneme or word aligners.
Discussion (0). Continue with ORCID to comment.