Pith. sign in

REVIEW 9 cited by

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.17667 v2 pith:Z7X3PEH2 submitted 2024-12-23 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords versatoolkitaudiospeechevaluationmusicincludingmetrics
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA offers 65 metrics with 729 metric variations based on different configurations. These metrics encompass evaluations utilizing diverse external resources, including matching and non-matching reference audio, text transcriptions, and text captions. As a lightweight yet comprehensive toolkit, VERSA is versatile to support the evaluation of a wide range of downstream scenarios. To demonstrate its capabilities, this work highlights example use cases for VERSA, including audio coding, speech synthesis, speech enhancement, singing synthesis, and music generation. The toolkit is available at https://github.com/wavlab-speech/versa.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.

  2. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

    cs.SD 2025-07 conditional novelty 6.0 of 10

    OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.

  3. OpusLM: A Family of Open Unified Speech Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new open family of speech-language models trained on public data matches or beats prior systems across ASR, TTS, and text benchmarks.

  4. A Perception-Based L2 Speech Intelligibility Indicator: Leveraging a Rater's Shadowing and Sequence-to-sequence Voice Conversion

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A sequence-to-sequence voice conversion model trained on a native rater's shadowing utterances can spot unintelligible segments in L2 speech, beating an ASR baseline on the native rater but not on all listeners.

  5. Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A new benchmark, ERSB, shows that neural speech codecs in noisy environments degrade both reconstruction quality and downstream speech enhancement and recognition consistency.

  6. CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.

  7. Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Staged ASR-to-text-response-to-TTS training makes open end-to-end spoken dialogue systems trainable on 300 hours of public human-human data and more coherent than one-step speech-to-speech models.

  8. SHEET: A Multi-purpose Open-source Speech Human Evaluation Estimation Toolkit

    cs.SD 2025-05 conditional novelty 5.0 of 10

    SHEET provides unified training and evaluation for MOS predictors, and its benchmark shows WavLM large and XLS-R 1b are the best SSL backbones for SSL-MOS on the tested datasets.

  9. Uni-VERSA: Versatile Speech Assessment with a Unified Network

    cs.SD 2025-05 conditional novelty 4.0 of 10

    A single network predicts eleven speech quality metrics across five dimensions and shows strong correlations within speech enhancement, but not on out-of-domain data.

Pith tools