Pith. sign in

REVIEW 8 cited by

NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.09494 v1 pith:GVTB2DHF submitted 2021-04-19 eess.AS cs.AIcs.LGcs.SD

classification eess.AScs.AIcs.LGcs.SD
keywords modelspeechqualitydatasetsnisqaoverallpredictiontrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present an update to the NISQA speech quality prediction model that is focused on distortions that occur in communication networks. In contrast to the previous version, the model is trained end-to-end and the time-dependency modelling and time-pooling is achieved through a Self-Attention mechanism. Besides overall speech quality, the model also predicts the four speech quality dimensions Noisiness, Coloration, Discontinuity, and Loudness, and in this way gives more insight into the cause of a quality degradation. Furthermore, new datasets with over 13,000 speech files were created for training and validation of the model. The model was finally tested on a new, live-talking test dataset that contains recordings of real telephone calls. Overall, NISQA was trained and evaluated on 81 datasets from different sources and showed to provide reliable predictions also for unknown speech samples. The code, model weights, and datasets are open-sourced.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CallScreenBench: Benchmarking On-Device Models as Phone Secretaries

    cs.CR 2026-08 conditional novelty 7.0 of 10

    A new phone-secretary benchmark shows quality scaling with capability, no triage scaling once degenerate baselines are subtracted, and more capable models relaying scam callback numbers more often.

  2. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

    cs.CL 2025-07 conditional novelty 7.0 of 10

    With prompt engineering (audio concatenation plus in-context examples), large audio models rank speech synthesis systems in line with human preferences, reaching up to 0.91 Spearman correlation.

  3. LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Using LLM-generated text descriptions of enhanced speech converted to sentiment scores as PPO rewards improves PESQ, STOI, and neural quality scores over supervised and DNSMOS-reward baselines on AVSEC-4.

  4. Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

    cs.CL 2025-05 conditional novelty 6.0 of 10

    By appending optimized token sequences to harmful speech, the authors achieve up to 89% attack success rate on SpeechGPT across six forbidden categories.

  5. CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

    cs.SD 2026-07 reject novelty 5.0 of 10

    CS-ETS applies Lyapunov and detrended-fluctuation-analysis losses inside a Samba encoder, but its headline audio gains are confounded by a DTW alignment step not applied to baselines.

  6. GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model

    eess.AS 2025-12 conditional novelty 5.0 of 10

    A two-stage decoder-only language model with continuous embeddings and UTMOS-based preference fine-tuning reports improved target-speaker-extraction scores on Libri2Mix.

  7. Schr\"odinger Bridge Mamba for One-Step Speech Enhancement

    cs.SD 2025-10 conditional novelty 5.0 of 10

    A Mamba-based speech enhancer trained with Schrödinger Bridge objectives produces strong denoising and dereverberation in one inference step with a low real-time factor.

  8. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

Pith tools