Pith. sign in

Paper Citation Record · LEDGER

NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2104.09494.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.09494 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:05:11.436227Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:08:37.813294Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be99caf8-6851-409d-8e79-0bcec4553505 · inbound

Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model cites this paper.

Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T11:30:04.920547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:30:04.920547Z digest=sha256:53d586a71efca1039417b3f22a357da8d71ce7e81c9bb80c329421e831c4de24

Observation 772a9847-6895-4a12-a802-c361164763e7 · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.879183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.879183Z digest=sha256:7c502f39d23fce89d3468520162cd6cd0a4479e4b67bf2629c2b0b0e8b0310e1

Observation be93759f-f1ab-42ed-97ef-c47e7ef11280 · inbound

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators cites this paper.

Audio Large Language Models Can Be Descriptive Speech Quality Evaluators NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T12:30:52.089145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T12:30:52.089145Z digest=sha256:28cb4a4901ce301deebb52419f2134f409f1cf57fcd4217e82af9a79ad3207ca

Observation c12aaf43-7de6-45ee-adfd-7cfa2152ef9d · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.656523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.656523Z digest=sha256:2d054ad45195b0ebb2ac137902f7f70d2666e2fa5c6136b53100e25d998749af

Observation 0d180f06-89bc-49a4-8a9c-4eeb242cd7b7 · inbound

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget cites this paper.

Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:11.436227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:05:11.436227Z digest=sha256:8b2bb6cde84e75caf5ad2512677141cdb9193ee8f94e123b75874bd04f859a28

Observation 1926ba2a-b443-45ad-9567-ba7ca98d6dc9 · inbound

Towards Flow-Matching-based TTS without Classifier-Free Guidance cites this paper.

Towards Flow-Matching-based TTS without Classifier-Free Guidance NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:37:50.085103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:37:50.085103Z digest=sha256:48b0c254a4302bf2a1ebedada50e613a1dfeb5fed84c24a2a3310d1c26de6192

Observation 01cb78cd-3329-4f01-bac8-a992308224c2 · inbound

RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations cites this paper.

RoVo: Robust Voice Protection Against Unauthorized Speech Synthesis with Embedding-Level Perturbations NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:33:39.467645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:33:39.467645Z digest=sha256:fdb7b84e13ff2e86c5bb68ab47a45798eeb9533ca50d29f78cfb93b9eab54f77

Observation acce009c-0393-49ea-a7a9-d118aa8da34c · inbound

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework cites this paper.

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:07.110867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:07.110867Z digest=sha256:90563bded5efeb1dad813942f43cbbc96657aa2c71269c963658f46f42b57f68

Observation c9c1005d-6da5-4701-ad4b-a3d196cfb89f · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 270

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.176541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.176541Z digest=sha256:3794ae21b5775bb9b00339c7803b8b75eb541605df7a4453871a509d10741af9

Observation 07980a83-9295-4dc1-a55c-f5146b5b06e7 · inbound

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation cites this paper.

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:56.738747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:47:56.738747Z digest=sha256:8c28a72c452f8d9ddef4b1043245639138242cfe84d796f3e4dcf0266662ec7d

Observation 66d5267e-3512-4829-a45e-6f977e03a8cb · inbound

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement cites this paper.

Schr\"odinger Bridge Mamba for One-Step Speech Enhancement NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:12:59.994039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:12:59.994039Z digest=sha256:d501ca1be66712c3a3f4374412a12d2cc456bc36743c9fcc33dc138a9bddf77d

Observation 4bc17981-6d85-46b7-9710-5c344c87022d · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.842814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:f983e1ea27a4c8298889c8f458c0ee932720f6fb2997c5ee0c01a8bbd197c0e4

Observation b5c520e3-968d-4498-a4f7-828dc82af443 · inbound

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model cites this paper.

GenTSE: Enhancing Target Speaker Extraction via a Coarse-to-Fine Generative Language Model NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:20:49.194237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:20:49.194237Z digest=sha256:d8882f4b7832e3f63f2efa0697f4561c253dc8c284733cc8e23810279d05c3e5

Observation accabfca-7a05-4f46-97bc-7fe84ad890bc · inbound

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement cites this paper.

LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T18:15:19.385299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:15:19.385299Z digest=sha256:47f23c83127d80a26a9c0f920c7cb30ae7e2baffbe9bdc43c911ae14b765d12f

Observation 6fba3f46-deb1-4120-b68d-1e010ce83738 · inbound

Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models cites this paper.

Discrete Token Modeling for Multi-Stem Music Source Separation with Language Models NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.421331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T16:28:06.677963Z digest=sha256:7bd5b6f2962d80171677d4f9ef81a6e1f8fb13d5274ce048284b558931fae25c

Observation 286239c9-47a8-4c7a-ae16-bbadbbfbeb63 · inbound

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment cites this paper.

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:10:09.297512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T01:07:50.243903Z digest=sha256:fbda6fffa1ffa96b04ee237bda6c31c0b1fdec700f7c0472574b39b0bc70b5a2

Observation d1606140-ce0a-4f09-a6dc-4b181af79233 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:08.365461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:9d58ae049f3b5d7b022b343b9782b2dccc9ff73ecd29fd3c1e2327bfd37ad20d

Observation acb0c766-be82-4ead-aa38-b0a6b470f1f4 · inbound

A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models cites this paper.

A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:27:53.833743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T20:26:59.049472Z digest=sha256:6ceb3833e0729a2fb9391a07792aeeed8d9698c01a7bafd1329819d2d57050d5

Observation fd924b34-dd36-43b5-8c1a-a366a3914541 · inbound

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition cites this paper.

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:47:04.326382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T00:09:19.291564Z digest=sha256:3b33990ce1eb75aa9e77be25da2cd8dd00091e6222b1ce4d9249826d7277b67c

Observation b6454594-db1d-4567-bde7-bf1accec18b5 · inbound

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions cites this paper.

Feature-Aligned Speech Watermarking for Robustness to Reconstruction Distortions NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:08:08.739075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T08:27:00.881610Z digest=sha256:d3207c6f5d38ff3120321c28e8403b1fbbdd97741d21bd31f5750756d5413911

Observation c8922d62-d0fb-4c6a-990d-4c59296eed29 · inbound

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech cites this paper.

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:08:37.814833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T06:01:06.123832Z digest=sha256:326c55fd87ce0efba8e29373bab0538f5d03ad1e778e0850783bd6c1e640883f

Observation a3de0a8f-1d1c-4804-bd6a-984d304ecd90 · inbound

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses cites this paper.

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T14:51:46.710045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:51:46.710045Z digest=sha256:9425ad6d78194de0d798eb5ff34d9967687bb748634550bb2767e771e94ec123

Observation cb0b5c6e-2eb4-4a5e-9dac-db672dcc5553 · inbound

CallScreenBench: Benchmarking On-Device Models as Phone Secretaries cites this paper.

CallScreenBench: Benchmarking On-Device Models as Phone Secretaries NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:41:58.389986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:41:58.389986Z digest=sha256:d4a9526e6a52832935be43af28bdb61ec8176271e4fe21eb5bd0bf2a417873d6

Observation 19601c4e-067d-4dc4-ac06-f51e5edc8609 · inbound

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions cites this paper.

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T04:16:51.721966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:16:51.721966Z digest=sha256:4630990ccb2f5bfca3286c75fe76e6f94965c91d26925052168f3dbf541d5e97