REVIEW 5 cited by
Open-Source Conversational AI with SpeechBrain 1.0
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete "recipes" of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.
Forward citations
Cited by 5 Pith papers
-
Autoregressive Speech Enhancement via Acoustic Tokens
Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.
-
LLAMAPIE: Proactive In-Ear Conversation Assistants
A two-model, on-device in-ear assistant that decides when to whisper one to three words of guidance matches a reactive chatbot's accuracy in live interviews while cutting response latency and perceived disruption by m...
-
A Benchmark of French ASR Systems Based on Error Severity
Four severity classes for ASR errors are defined and used to benchmark ten French ASR systems, with Kaldi plus RNNLM rescoring best overall and LeBenchmark 7k character models best at avoiding critical errors.
-
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data
A TSE dataset using clean LibriTTS targets, noisy VoxCeleb2 interference, synthetic speaker augmentation and curriculum learning reports iSDR gains of 1.39 dB and 0.78 dB on Libri2Talker and Libri2Vox test sets.
-
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
A systematic benchmark shows that standard saliency methods faithfully reflect a Parkinson's speech classifier's decisions, yet the resulting spectrogram highlights are not readily usable by domain experts.
Discussion (0). Continue with ORCID to comment.