Pith. sign in

REVIEW 5 cited by

Open-Source Conversational AI with SpeechBrain 1.0

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.00463 v5 pith:DGNEMUTJ submitted 2024-06-29 cs.LG cs.AIcs.CLcs.HCeess.AS

classification cs.LGcs.AIcs.CLcs.HCeess.AS
keywords modelsspeechspeechbraintasksconversationaldiverselanguagemodalities
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

SpeechBrain is an open-source Conversational AI toolkit based on PyTorch, focused particularly on speech processing tasks such as speech recognition, speech enhancement, speaker recognition, text-to-speech, and much more. It promotes transparency and replicability by releasing both the pre-trained models and the complete "recipes" of code and algorithms required for training them. This paper presents SpeechBrain 1.0, a significant milestone in the evolution of the toolkit, which now has over 200 recipes for speech, audio, and language processing tasks, and more than 100 models available on Hugging Face. SpeechBrain 1.0 introduces new technologies to support diverse learning modalities, Large Language Model (LLM) integration, and advanced decoding strategies, along with novel models, tasks, and modalities. It also includes a new benchmark repository, offering researchers a unified platform for evaluating models across diverse tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Autoregressive Speech Enhancement via Acoustic Tokens

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.

  2. LLAMAPIE: Proactive In-Ear Conversation Assistants

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A two-model, on-device in-ear assistant that decides when to whisper one to three words of guidance matches a reactive chatbot's accuracy in live interviews while cutting response latency and perceived disruption by m...

  3. A Benchmark of French ASR Systems Based on Error Severity

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Four severity classes for ASR errors are defined and used to benchmark ten French ASR systems, with Kaldi plus RNNLM rescoring best overall and LeBenchmark 7k character models best at avoiding critical errors.

  4. Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A TSE dataset using clean LibriTTS targets, noisy VoxCeleb2 interference, synthetic speaker augmentation and curriculum learning reports iSDR gains of 1.39 dB and 0.78 dB on Libri2Talker and Libri2Vox test sets.

  5. Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech

    cs.SD 2024-11 conditional novelty 5.0 of 10

    A systematic benchmark shows that standard saliency methods faithfully reflect a Parkinson's speech classifier's decisions, yet the resulting spectrogram highlights are not readily usable by domain experts.

Pith tools