Pith. sign in

REVIEW 3 cited by

Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.05344 v2 pith:QXS724YC submitted 2018-08-16 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords speechqualityassessmentquality-netreferencecleanevaluationnon-intrusive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Nowadays, most of the objective speech quality assessment tools (e.g., perceptual evaluation of speech quality (PESQ)) are based on the comparison of the degraded/processed speech with its clean counterpart. The need of a "golden" reference considerably restricts the practicality of such assessment tools in real-world scenarios since the clean reference usually cannot be accessed. On the other hand, human beings can readily evaluate the speech quality without any reference (e.g., mean opinion score (MOS) tests), implying the existence of an objective and non-intrusive (no clean reference needed) quality assessment mechanism. In this study, we propose a novel end-to-end, non-intrusive speech quality evaluation model, termed Quality-Net, based on bidirectional long short-term memory. The evaluation of utterance-level quality in Quality-Net is based on the frame-level assessment. Frame constraints and sensible initializations of forget gate biases are applied to learn meaningful frame-level quality assessment from the utterance-level quality label. Experimental results show that Quality-Net can yield high correlation to PESQ (0.9 for the noisy speech and 0.84 for the speech processed by speech enhancement). We believe that Quality-Net has potential to be used in a wide variety of applications of speech signal processing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

    cs.SD 2025-02 unverdicted novelty 6.0 of 10

    Unified no-reference models assess audio aesthetics across speech, music, and sound via four perceptual axes and achieve performance comparable or superior to human mean opinion scores.

  2. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

  3. SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

    cs.SD 2025-06 reject novelty 4.0 of 10

    SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.

Pith tools