Pith. sign in

REVIEW 1 cited by

TorchAudio-Squim: Reference-less Speech Quality and Intelligibility measures in TorchAudio

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01448 v1 pith:MICOZFFB submitted 2023-04-04 eess.AS

classification eess.AS
keywords speechintelligibilitymetricsmodelsqualityprocessingsubjectivetorchaudio-squim
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Measuring quality and intelligibility of a speech signal is usually a critical step in development of speech processing systems. To enable this, a variety of metrics to measure quality and intelligibility under different assumptions have been developed. Through this paper, we introduce tools and a set of models to estimate such known metrics using deep neural networks. These models are made available in the well-established TorchAudio library, the core audio and speech processing library within the PyTorch deep learning framework. We refer to it as TorchAudio-Squim, TorchAudio-Speech QUality and Intelligibility Measures. More specifically, in the current version of TorchAudio-squim, we establish and release models for estimating PESQ, STOI and SI-SDR among objective metrics and MOS among subjective metrics. We develop a novel approach for objective metric estimation and use a recently developed approach for subjective metric estimation. These models operate in a ``reference-less" manner, that is they do not require the corresponding clean speech as reference for speech assessment. Given the unavailability of clean speech and the effortful process of subjective evaluation in real-world situations, such easy-to-use tools would greatly benefit speech processing research and development.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A 2D patch-quantized VQ-VAE with a single 4096-entry codebook plus a HiFi-GAN vocoder reaches ~7.5 kbits/s for 16 kHz speech, with intelligibility below EnCodec and DAC but a simpler, non-residual architecture.

Pith tools