Pith. sign in

Paper Citation Record · LEDGER

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.02082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02082 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:44:17.796045Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact4
  • verified fuzzy10
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02691a0c-79b2-460b-9721-f69bf7e95692 · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.618021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.618021Z digest=sha256:65cacd3e994f660dc9f183a0c05e22c3ca674d1fdc84106643aa4e7217ca3b6f

Observation c2e23147-c707-4322-a1db-88b238e4d96c · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.625066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.625066Z digest=sha256:14d04c7d192165d8f0b57b90f6bd1d3ef55bb6cf2623c6480ad47b32e22f19b8

Observation 19fe8f58-2378-4a75-80be-20807df4b347 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.631566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.631566Z digest=sha256:fb82924bf7ab00f69af298c0212e0a6e05c3fc0e463664d722905f515059c490

Observation 0ea38d4e-bb61-43e4-974a-8cdc59922e95 · outbound

This paper cites TERA: Self-supervised learning of transformer encoder representation for speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction TERA: Self-supervised learning of transformer encoder representation for speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.390260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.638024Z digest=sha256:1d12df7115b61b5b900ef9919419edf66b9e50ce0cd7d96c99ea7340e2241b60

Observation a1983489-b854-4836-aa1b-fdb7682379f2 · outbound

This paper cites AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.644998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.644998Z digest=sha256:68b8488413f1589cfa7ff226c5ab8097463b11ca07ca203a4bc57d32af31ce8d

Observation 888f7599-bd55-4125-84bc-a9f43f7a0964 · outbound

This paper cites Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Quality-Net: An End-to-End Non-intrusive Speech Quality Assessment Model based on BLSTM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.651602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.651602Z digest=sha256:3903d9b67c528c0beed998bf80f6eb0f62f50946679f1879fd9f12f4b62f8dbb

Observation d57acf9b-c798-4cf0-9b6a-35142091227b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.365044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.659966Z digest=sha256:2aa4fd4e5a85cd820edc4df0c5cfdd8caca523f310f5562b3c711b530f7923e1

Observation b02c79b6-39e8-437a-9d8e-d2bff1aaa8b0 · outbound

This paper cites Evaluation of speech representations for mos prediction,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Evaluation of speech representations for mos prediction,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.346302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.667199Z digest=sha256:395c9f8286a5e26e203e2d07fac06218f78a647b84f7f275c6f94ccaa9947eec

Observation 9148ec3c-421a-43bd-8625-5cc092c63f54 · outbound

This paper cites Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Deep learning-based non-intrusive multi-objective speech assessment model with cross-domain features,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.324451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.672517Z digest=sha256:a6f52b5f269eec38af5b31fa50dc05669d5cc32c97b92f24e9d5d87d9c3aa00c

Observation ed080602-d705-4148-958f-1e36bc095878 · outbound

This paper cites MOSNet: Deep Learning based Objective Assessment for Voice Conversion.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSNet: Deep Learning based Objective Assessment for Voice Conversion

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.680497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.680497Z digest=sha256:75e48628efdea63605f924b0ef2de1b0cfdfb5d3315360b1b34917318923daeb

Observation dab1aff8-dbb6-417d-aaa0-f98a13c6d6fc · outbound

This paper cites MBNET: Mos prediction for synthesized speech with mean-bias network,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MBNET: Mos prediction for synthesized speech with mean-bias network,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.302247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.688319Z digest=sha256:db587f723aa1f22f9bc8fcceb6200380c62dcfe5b0a2d10e1595fff43af5268a

Observation f1bc8676-32ed-4d9a-ab5e-8020b98bab0c · outbound

This paper cites LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LDNET: Unified listener dependent modeling in mos prediction for synthetic speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.274372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.695218Z digest=sha256:465b9595ab5b32b0bd2341665c58ed74cd7732666f82c63e6cf370724b923869

Observation 367dae72-feec-411a-8996-461f21409134 · outbound

This paper cites DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction DDOS: A MOS Prediction Framework utilizing Domain Adaptive Pre-training and Distribution of Opinion Scores

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.048488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.702839Z digest=sha256:7236bab4e5743d68323dded5caa6beaab69f454684a232e5eab8799efbedb475

Observation a7e43f76-0e26-4106-a15a-d284420f2959 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.709392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.709392Z digest=sha256:59404751a761f48c9c6ab97318fdcce9c17e1a5614d45a605816ac7a6dec1591

Observation b8b67440-aadc-440a-adc2-e0d77a008300 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.715160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.715160Z digest=sha256:ae46c556c31d677d908e33c5358b50ab96e168e791b32fd185755a6b19fb4bcb

Observation b5d7a4de-583e-489f-a575-3c0c73ee414a · outbound

This paper cites Fusion of Self-supervised Learned Models for MOS Prediction.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Fusion of Self-supervised Learned Models for MOS Prediction

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:18.004321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.720703Z digest=sha256:5e1ace94b16100b4fc683a77477e35d9ce6ad0feb9c1f41ed35630ee8cd002cb

Observation 41d06764-3c28-4b9a-9229-41f4a80baf47 · outbound

This paper cites MOSPC: MOS Prediction Based on Pairwise Comparison.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction MOSPC: MOS Prediction Based on Pairwise Comparison

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.727766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.727766Z digest=sha256:363118eb379e512ab48ad1e4b026f4bdde2ab66fd0b242c070abe9e993cf117f

Observation 220b25bd-0434-4023-9b39-cf319156dd00 · outbound

This paper cites Speech Quality Assessment through MOS using Non-Matching References.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Speech Quality Assessment through MOS using Non-Matching References

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.946189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.734793Z digest=sha256:e1e23ca6ec779dac2e5fc73822046b02c6d6023fd58209d2bb80c92ab3173f7f

Observation 2dfbb2a6-1e08-49e7-a198-02301099feb2 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction U-net: Convolutional networks for biomedical image segmentation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.741382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.741382Z digest=sha256:c817917d3c47ae84a64a0078ec1f2eb735ea9f8395a784f75c8a1a84a1537030

Observation e1f91c3c-ccd9-419b-9590-0feaf4e98c56 · outbound

This paper cites Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction Perceptual objective listening quality assess- ment (POLQA), the third generation itu-t standard for end-to-end speech quality measurement part i—temporal alignment,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.223846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.748440Z digest=sha256:e07a1cbd5394999d0f4d26c91f7f056120177927eaabe7074d00e8b1e8dd8119

Observation 52aae80b-685e-455a-9b2f-bb90c42a645f · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.199291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.755047Z digest=sha256:002e78948da3dd8e5f83fad4bf685f376cf7f6f73acc7322363949f2262c7572

Observation b03e58a3-6fd4-4615-98b8-157532293267 · outbound

This paper cites The VoiceMOS Challenge 2022.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The VoiceMOS Challenge 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.760799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.760799Z digest=sha256:1a0b575eb4011ba3b97a59697633770ef1f640a221be5090bd5d6d7c5881d34b

Observation 69057cac-d56f-43b8-88c7-20aa16930c5c · outbound

This paper cites The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction The Voice Conversion Challenge 2018: Promoting Development of Parallel and Nonparallel Methods

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.767102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.767102Z digest=sha256:627e773c1ff2be8de2952fd0dd54e239a03a1a14f678326caa513caced7e0e64

Observation 85f38dec-661a-4107-b374-415b0001405f · outbound

This paper cites SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:17.773360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:17.773360Z digest=sha256:763ccc273521ff0db260f8260f6ba0b86b316b0574b96566e9d331fdb92d8dd1

Observation a6b79536-61c2-4870-96f6-7dc5100f0201 · outbound

This paper cites InQSS: a speech intelligibility and quality assessment model using a multi-task learning network.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction InQSS: a speech intelligibility and quality assessment model using a multi-task learning network

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:44:17.855886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.783207Z digest=sha256:768e87a8fb6da02c2f363361774a607623fb2540fceef6ab7ec98c39a9b56a8c

Observation b16f79fd-15d8-4d3e-8122-3f08a56ba9a7 · outbound

This paper cites LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction LE-SSL-MOS: Self-supervised learning mos prediction with listener enhancement,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.180511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.790700Z digest=sha256:6f62bc3dee0be9dfc1df856cae60c4f3581f11c6c677e097af8e1ea93391850c

Observation 065f0d59-b72a-410d-8802-f5139d7fc6ee · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction X-vectors: Robust dnn embeddings for speaker recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:44:18.158851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T11:44:17.796045Z digest=sha256:6f8bbbd575b4cea0d09d1f194bc4cdc0ba9cb648e4553736e5c11126738c6a8f

Pith citing papers

No inbound Pith citation observations are available.