Pith. sign in

Paper Citation Record · LEDGER

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

As of 8 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.02082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02082 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:17.796045Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact4
  • verified fuzzy10
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02691a0c-79b2-460b-9721-f69bf7e95692 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.618021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.618021Z digest=sha256:2f2de0134f69bc9ccc45239465a17d7c00585e7ee040351735e7598c380ef48d

Observation c2e23147-c707-4322-a1db-88b238e4d96c · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.625066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.625066Z digest=sha256:14d04c7d192165d8f0b57b90f6bd1d3ef55bb6cf2623c6480ad47b32e22f19b8

Observation 19fe8f58-2378-4a75-80be-20807df4b347 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.631566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.631566Z digest=sha256:fb82924bf7ab00f69af298c0212e0a6e05c3fc0e463664d722905f515059c490

Observation 0ea38d4e-bb61-43e4-974a-8cdc59922e95 · outbound

This paper cites TERA: Self-supervised learning of transformer encoder representation for speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction TERA: Self-supervised learning of transformer encoder representation for speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.390260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.638024Z digest=sha256:45171eee094f5bd77ab50b373123ed6f35901c84d447549a7b192386c55e75ca

Observation a1983489-b854-4836-aa1b-fdb7682379f2 · outbound

This paper cites AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.644998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.644998Z digest=sha256:81ddb477dac67a6ffc2dc9668579a47b38fe2dc95ed22cd43758b44f8eef154c

Observation 888f7599-bd55-4125-84bc-a9f43f7a0964 · outbound

This paper cites Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.651602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.651602Z digest=sha256:76eb80cfcb88c503f1d6f848be6526a0c3f533fdb84cdd41d33553bd99adb766

Observation d57acf9b-c798-4cf0-9b6a-35142091227b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.365044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.659966Z digest=sha256:bb75ff53eab0098fc735f2c494db4719cbfc8e5e9db06cef2f0c0bf48246b946

Observation b02c79b6-39e8-437a-9d8e-d2bff1aaa8b0 · outbound

This paper cites Evaluation of speech representations for mos prediction,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Evaluation of speech representations for mos prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.346302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.667199Z digest=sha256:5ac7da54fc56df24bfe770103586c36adce1de92a63b2ce95735b0b3103acb46

Observation 9148ec3c-421a-43bd-8625-5cc092c63f54 · outbound

This paper cites Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.324451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.672517Z digest=sha256:17fedb355c2f19344732ffcf48f20f37876fa3a72c2f6b5721311dcb4bff155b

Observation ed080602-d705-4148-958f-1e36bc095878 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.680497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.680497Z digest=sha256:6464d4555dd8774427bed53da5ac22d4ae2a132d93df8eacac92c0d6b7d554eb

Observation dab1aff8-dbb6-417d-aaa0-f98a13c6d6fc · outbound

This paper cites MBNET: Mos prediction for synthesized speech with mean-bias network,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MBNET: Mos prediction for synthesized speech with mean-bias network,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.302247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.688319Z digest=sha256:2bf83d364f7b55618fa6c9d5bea17eb9a960dccd7f5a7a870ce7d4a7e572f212

Observation f1bc8676-32ed-4d9a-ab5e-8020b98bab0c · outbound

This paper cites LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.274372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.695218Z digest=sha256:316f70056e382829d30701e105ed99bb3f7b00c55edbc3c3503bf2e522a99174

Observation 367dae72-feec-411a-8996-461f21409134 · outbound

This paper cites DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.048488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.702839Z digest=sha256:d95308dcca417bcf2424ff5dbc023298302e8ff48e539fbe75b5c89d115eae22

Observation a7e43f76-0e26-4106-a15a-d284420f2959 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.709392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.709392Z digest=sha256:59404751a761f48c9c6ab97318fdcce9c17e1a5614d45a605816ac7a6dec1591

Observation b8b67440-aadc-440a-adc2-e0d77a008300 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.715160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.715160Z digest=sha256:c82499e8e1984d1406edecda0ffde9ca4583d66e0edb2e2692e222269ee9cc99

Observation b5d7a4de-583e-489f-a575-3c0c73ee414a · outbound

This paper cites Fusion of Self-supervised Learned Models for MOS Prediction.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Fusion of Self-supervised Learned Models for MOS Prediction

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.004321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.720703Z digest=sha256:cd1708697b8219f170245c733b9be8b12c3d5685a2c5798fac3f43a0ff59ad03

Observation 41d06764-3c28-4b9a-9229-41f4a80baf47 · outbound

This paper cites MOSPC: MOS Prediction Based on Pairwise Comparison.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSPC: MOS Prediction Based on Pairwise Comparison

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.727766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.727766Z digest=sha256:a89314df48b48d26834dcc01c0dcfd5d6cdcdb2217127f479735491b71cbf908

Observation 220b25bd-0434-4023-9b39-cf319156dd00 · outbound

This paper cites Speech Quality Assessment through MOS using Non-Matching References.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Speech Quality Assessment through MOS using Non-Matching References

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.946189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.734793Z digest=sha256:ea9d09b98f043a74ab011fce803e878831b4ed0cda8588709a9911a9be21bfaf

Observation 2dfbb2a6-1e08-49e7-a198-02301099feb2 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction U-net: Convolutional networks for biomedical image segmentation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.741382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.741382Z digest=sha256:c817917d3c47ae84a64a0078ec1f2eb735ea9f8395a784f75c8a1a84a1537030

Observation e1f91c3c-ccd9-419b-9590-0feaf4e98c56 · outbound

This paper cites Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.223846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.748440Z digest=sha256:ba6f492c3bb983d39acec582054435c2d7791440d93291c8878107fd996bbce7

Observation 52aae80b-685e-455a-9b2f-bb90c42a645f · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.199291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.755047Z digest=sha256:33561ca7482454d7a22f654571581398ba74388314d688cd44212ddfae83bccb

Observation b03e58a3-6fd4-4615-98b8-157532293267 · outbound

This paper cites The VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The VoiceMOS Challenge 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.760799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.760799Z digest=sha256:e37404dc781d44faeb34b7d9699190b8ef55c0df1fab852f12ee20c3d0365809

Observation 69057cac-d56f-43b8-88c7-20aa16930c5c · outbound

This paper cites The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.767102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.767102Z digest=sha256:3e3e02774d11238d34d56605e1386b5c7d9415c83de53d19a57faf9c016c41a5

Observation 85f38dec-661a-4107-b374-415b0001405f · outbound

This paper cites SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.773360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.773360Z digest=sha256:c67011bcf5daa414e65a93ca68dd0d7e68cb42345e7687aa3775273ae615598e

Observation a6b79536-61c2-4870-96f6-7dc5100f0201 · outbound

This paper cites InQSS: a speech intelligibility and quality assessment model using a multi-task learning network.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction InQSS: a speech intelligibility and quality assessment model using a multi-task learning network

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.855886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.783207Z digest=sha256:ab2da0edc53ea5508c4eb86ea9d92555009c6f7ac1d6ad6eed47fc2284c9c1d4

Observation b16f79fd-15d8-4d3e-8122-3f08a56ba9a7 · outbound

This paper cites LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.180511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.790700Z digest=sha256:eaf3168afa43051f915693f3a8e9a5f26046366796d621e14e72a8fe664fa061

Observation 065f0d59-b72a-410d-8802-f5139d7fc6ee · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction X-vectors: Robust dnn embeddings for speaker recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.158851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:44:17.796045Z digest=sha256:e15ceace58333b037c4dbd10d1e73f552a28932d9ca8a5dead2a91825ddcf530

Pith citing papers

No inbound Pith citation observations are available.