REVIEW 9 cited by
VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/wavlab-speech/versa.
Forward citations
Cited by 9 Pith papers
-
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.
-
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.
-
OpusLM: A Family of Open Unified Speech Language Models
A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.
-
A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion
A sequence-to-sequence voice conversion model trained on a native rater's shadowing utterances can spot unintelligible segments in L2 speech, beating an ASR baseline on the native rater but not on all listeners.
-
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
A new benchmark, ERSB, shows that neural speech codecs in noisy environments degrade both reconstruction quality and downstream speech enhancement and recognition consistency.
-
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.
-
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
Staged ASR-to-text-response-to-TTS training makes open end-to-end spoken dialogue systems trainable on 300 hours of public human-human data and more coherent than one-step speech-to-speech models.
-
SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit
SHEET provides unified training and evaluation for MOS predictors, and its benchmark shows WavLM large and XLS-R 1b are the best SSL backbones for SSL-MOS on the tested datasets.
-
Uni-VERSA: Versatile Speech Assessment with a Unified Network
A single network predicts eleven speech quality metrics across five dimensions and shows strong correlations within speech enhancement, but not on out-of-domain data.
Discussion (0). Sign in to comment.