Pith. sign in

Paper Citation Record · LEDGER

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2411.11232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11232 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:50:08.405900Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:50:08.261278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T18:50:08.499324Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a10d66e1-013c-4a02-a300-2a91e0e2666b · outbound

This paper cites Common ob- jective evaluation metrics, such as mel-cepstral distance (MCD).

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Common ob- jective evaluation metrics, such as mel-cepstral distance (MCD)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.906628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.253206Z digest=sha256:00020ad2ab832892c965898f751e2402132e605e0d75f44b498bf30e288bfb1c

Observation 310222d8-cc72-43e2-8675-17cf92c5aab0 · outbound

This paper cites As a result, some ob- jective measures or models related to human perception have been proposed [3, 4, 5, 6].

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features As a result, some ob- jective measures or models related to human perception have been proposed [3, 4, 5, 6]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.893734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.257531Z digest=sha256:bea2e321d3bc5156deffa14171fc45cbcda810dc16c965c36ef2c6b4b279e243

Observation e87ed4fb-8977-4686-88e2-04c264b79dd2 · outbound

This paper cites Dataset In this paper, the experiments followed the same settings as the V oiceMOS Challenge 2022 [15].

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Dataset In this paper, the experiments followed the same settings as the V oiceMOS Challenge 2022 [15]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.857226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.273361Z digest=sha256:1e955c38cb5111e1f9ec21e2333c23831f82fa8e8753c3878c8ebead8704ef46

Observation 780f391c-0607-4a61-aa33-3478aefec9fb · outbound

This paper cites mean-listener.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features mean-listener

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.881238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.265344Z digest=sha256:98e7131bdce8fb4afbc79093b42d70716ce9460dbfa2fd7bf28c595b2f967180

Observation ee4f1d43-f25c-4638-bcb1-238c065bba6e · outbound

This paper cites When the rater ID is not the mean-listener, the label representing the sample is the score given by the individual rater (an integer i from 1 to 5).

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features When the rater ID is not the mean-listener, the label representing the sample is the score given by the individual rater (an integer i from 1 to 5)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.869423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.269386Z digest=sha256:abf62b75979f7d7f13e5749d80ec6e813f2a05f6afef310babbb057ce77fb18b

Observation 1f5c6476-c622-4219-8e09-7e6423a05b1d · outbound

This paper cites NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features NISQA: A deep CNN-self-attention model for multidimensional speech quality prediction with crowdsourced datasets,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.762518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.308516Z digest=sha256:3fa6435869ddd628d1c2d6c1b78d3fc164fb36c28b20a70af004a413a99121d0

Observation 67bc5e51-1fd6-407d-bb72-5fb44d2d6e40 · outbound

This paper cites Comparision with baseline methods We first compare the proposed SAMOS with the baselines.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Comparision with baseline methods We first compare the proposed SAMOS with the baselines

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.846329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.277306Z digest=sha256:966fdb00f18a27195da191a28f6ed8935400d1080c1edaeaae2fe25accd14f06

Observation 41c32c20-0186-47f1-a6b0-8cb3bec56c93 · outbound

This paper cites We can see that removing the semantic module resulted in the degradation of all the metrics on both datasets, indicating the importance of semantic repre- sentations from SSL model.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features We can see that removing the semantic module resulted in the degradation of all the metrics on both datasets, indicating the importance of semantic repre- sentations from SSL model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.834338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.281046Z digest=sha256:3926826a5c6b637aee399eb0a3f6b6a4c25239a02a53e69873d07eabe8ac66f2

Observation cdec2479-b405-4bf8-85b1-1d9413048786 · outbound

This paper cites To improve prediction accuracy, SAMOS employs parallel regression and classification heads, and finally outputs the final MOS score through an aggregation layer.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features To improve prediction accuracy, SAMOS employs parallel regression and classification heads, and finally outputs the final MOS score through an aggregation layer

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.821712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.285155Z digest=sha256:642dc7277f76dee36c69e3018bb27b1a4f1a52072bcc45889779c69a1102dcea

Observation a4a7314c-03f6-4e6b-9fbc-0c1eacc56cf3 · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assessment,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Mel-cepstral distance measure for objective speech quality assessment,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.808890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.288976Z digest=sha256:9b5b9ba67b3f98d5756b11a326fd2258dc281167b9fd5b62185192ccff3ac2e0

Observation 9029bd3b-988a-4b5d-a2ce-96aa0b219744 · outbound

This paper cites SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.503803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.261278Z digest=sha256:8cd9a3a63c13dc3bf6d80f18e84727a9bc63cab59acbfc6e20e5bd3da6356d01

Observation 30193a09-e986-462f-9cfc-2ca600386acb · outbound

This paper cites SDR– half-baked or well done?.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SDR– half-baked or well done?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.292704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.292704Z digest=sha256:c66c3508353cee8bc09d1b06e52669b937b8cb3cf3649b32619984f51732d578

Observation e07f6f7b-3792-4d31-b07a-1e0954f8979d · outbound

This paper cites Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.296480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.296480Z digest=sha256:0e3d949743698ece6e32a8e429e5ea8681a23200d10994ca32276fa448d2fc62

Observation e44b9181-7668-4052-8766-de5ac55a6734 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.300376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.300376Z digest=sha256:dcb1e13104322947b0dae058134c6df1f1ffc3ac62470c226f25e935c9804db2

Observation f594f578-5842-4e96-9dde-2ef897188c1c · outbound

This paper cites ViSQOL: An objective speech quality model,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ViSQOL: An objective speech quality model,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.774239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.304873Z digest=sha256:edd889c79062721c00dc5ade4b9e8eaf25b476df515208cc5371ef4d0cfcc309

Observation e699b29c-92a3-4aad-8d43-3aa357453f6f · outbound

This paper cites AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features AutoMOS: Learning a non-intrusive assessor of naturalness-of-speech

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.312306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.312306Z digest=sha256:b215bdd7eb74c9623ec72494be3835f046704c53447712171fb0effe944ada10

Observation 7b82dcd7-8716-4fbe-8279-c3b28f26b027 · outbound

This paper cites Quality-Net: An end-to-end non-intrusive speech quality assessment model based on blstm,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Quality-Net: An end-to-end non-intrusive speech quality assessment model based on blstm,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.751072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.316397Z digest=sha256:eb91d46674f3d2f4299c012c5d2a178e3df2d7ef358fdf4edf163416df840859

Observation 3e78365f-4946-4647-898c-9c39a313e118 · outbound

This paper cites MOSNet: Deep learning-based objec- tive assessment for voice conversion,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features MOSNet: Deep learning-based objec- tive assessment for voice conversion,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.320018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.320018Z digest=sha256:9da2e80ed7acf4e1dfbc6fd84f373fb01af268dba1c18a8fe90f8cc176d5f85d

Observation 3d77dae0-f23a-4757-844d-9fa4cf8e0cbe · outbound

This paper cites MBNet: MOS prediction for synthesized speech with mean-bias network,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features MBNet: MOS prediction for synthesized speech with mean-bias network,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.323526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.323526Z digest=sha256:524fe07738ef90e7da4574803728221f6c82b6aedfa6743640caac75d96c2fe1

Observation 794cebee-6b36-4bbb-b50d-89a86f18a8e7 · outbound

This paper cites LDNet: Unified listener dependent modeling in mos prediction for syn- thetic speech,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features LDNet: Unified listener dependent modeling in mos prediction for syn- thetic speech,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.724766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.326959Z digest=sha256:25a7d6a2c66bf548c885011754d4debb02198b43782bd9639c85800ae43afab4

Observation 52e2e90f-d821-48e3-befb-f3f91db3124b · outbound

This paper cites Generaliza- tion ability of mos prediction networks,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Generaliza- tion ability of mos prediction networks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.712368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.330731Z digest=sha256:3f88703dc0c4c37104ede8f2084219ea076944337de1df511c48116da36c1563

Observation fb553b37-1660-4ac6-83a8-6a2253627619 · outbound

This paper cites Deep learning-based non-intrusive multi- objective speech assessment model with cross-domain features,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Deep learning-based non-intrusive multi- objective speech assessment model with cross-domain features,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.700005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.334617Z digest=sha256:9582d6ba84d7cbdf1248c0dd727efadf3637d1bfe40edc57f3284f6a096c33a0

Observation 60af0c23-252a-4d64-97a0-0496ea529157 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.337864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.337864Z digest=sha256:2fd24fc64dceb070586dcc44e4380dd716aceea8b4c9809126bed47417d43bd3

Observation 54fe6ef4-8be8-4958-8218-b5dc94710d9f · outbound

This paper cites The V oiceMOS Challenge 2022,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features The V oiceMOS Challenge 2022,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.679439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.341380Z digest=sha256:261d8ec642a504f341742a2fb95e8da36e39efa27bffd374e89c3b44aa814fb9

Observation 1c22bcfd-ac74-4282-8180-efc23733ff3e · outbound

This paper cites A transfer and multi-task learning based approach for MOS predic- tion,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features A transfer and multi-task learning based approach for MOS predic- tion,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.667934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.344950Z digest=sha256:517b9145eed70163e379d7fe4d8cae049b3c3159e1c6bd75f61f8a4a2d2ca443

Observation 25448d39-b0e7-4b40-8abc-220ca5382bb6 · outbound

This paper cites DDOS: A MOS predic- tion framework utilizing domain adaptive pre-training and distri- bution of opinion scores,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features DDOS: A MOS predic- tion framework utilizing domain adaptive pre-training and distri- bution of opinion scores,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.656349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.348461Z digest=sha256:1ce611e655513c4a1f8417b88d79ee01b463d39c24f6fb302dd3986bcafbcf70

Observation b4ec5239-740e-4ec7-912c-06563030db64 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features UTMOS: UTokyo-SaruLab system for V oice- MOS Challenge 2022,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.643694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.351975Z digest=sha256:63956a8b4d170ed6a6d13032247656d9fe80839dfe4097b3cf284cf1233b45be

Observation c7776e9b-87ab-4c7b-9e29-10bdc810d981 · outbound

This paper cites Fusion of self-supervised learned models for MOS pre- diction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Fusion of self-supervised learned models for MOS pre- diction,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.631520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.355250Z digest=sha256:92a846f57f250fe0261b8a531f9d4bd98b4f1baba47b4de1a8cea201bf5a0bdd

Observation dded1169-beef-48c3-a257-46a1ba2d1be0 · outbound

This paper cites Ensem- ble of deep neural network models for MOS prediction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Ensem- ble of deep neural network models for MOS prediction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.618523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.359094Z digest=sha256:f820742003a3ee065279d8024dab8f1b60ac5ed7c41d5e13b3496bb78c3995af

Observation 447a59fb-38f9-4672-a89c-98c72831c85d · outbound

This paper cites RAMP: Retrieval- augmented MOS prediction via confidence-based dynamic weighting,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features RAMP: Retrieval- augmented MOS prediction via confidence-based dynamic weighting,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.606030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.362859Z digest=sha256:ef505e507f99097a36235f921137411f51933115f6e23f0b1353ea9b0717b87f

Observation 91352e5e-3eba-47d5-ab95-306699f19a0c · outbound

This paper cites Investigating content-aware neural text-to-speech MOS prediction using prosodic and linguistic fea- tures,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Investigating content-aware neural text-to-speech MOS prediction using prosodic and linguistic fea- tures,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.593262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.366284Z digest=sha256:fcc07a6789e905dca2fc58c177a0c4d80d89ad73fcd7d161e3948d0fed440bd0

Observation 1fee44e5-7b9f-45fb-8672-4163a1ba579f · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.580675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.369808Z digest=sha256:a5e8a8f5f8c72ee13d5ef292910cc6e392e83c7d97c53ef46107001329754ab4

Observation fffca9b7-87f0-4c05-9ff7-43ce8e77bfb5 · outbound

This paper cites BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.475207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.373433Z digest=sha256:ebdf047cd570768d569b36f50d9f905bef673795eb2d1b7d30f782436989b8cc

Observation 1027c81d-02da-479f-ac32-dcac8a1dbf17 · outbound

This paper cites SQAT-LD: Speech quality assessment transformer utilizing listener depen- dent modeling for zero-shot out-of-domain MOS prediction,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SQAT-LD: Speech quality assessment transformer utilizing listener depen- dent modeling for zero-shot out-of-domain MOS prediction,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.569287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.377396Z digest=sha256:e77ec47d8da748c2ecb39659ae31741b6897311bdfb06cd4d917e705b1a1d859

Observation 151e9e29-7922-43e7-982b-329479541b82 · outbound

This paper cites How do Voices from Past Speech Synthesis Challenges Compare Today?.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features How do Voices from Past Speech Synthesis Challenges Compare Today?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.381069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.381069Z digest=sha256:3f232141c7c34baaebddd9e710d5fd15f1e50f6bb432260cdfcd40781ade6de1

Observation 02825d8f-0937-4a6c-8255-0eaafae92b5f · outbound

This paper cites The Blizzard Challenge 2019,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features The Blizzard Challenge 2019,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.385075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.385075Z digest=sha256:46205c6fc04ea265fac2dcde6f87eccfd1710148dd9767e1c23120be3ec410d2

Observation 3d4e407a-af34-485f-aa7f-04e59493193e · outbound

This paper cites Conformer: Convolution- augmented transformer for speech recognition,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Conformer: Convolution- augmented transformer for speech recognition,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T18:50:08.389039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:50:08.389039Z digest=sha256:74613b46d1edabe4087f824cd2b4b161b386d70e40a6af56d5a0cbc12ef0c963

Observation 0fd6d85c-a83c-4c24-98d7-879b107e970d · outbound

This paper cites ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.543400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.393014Z digest=sha256:ecc58fa0e9e0aa4f1e8dc6610e4bbf51e49042f944599680b09c7d1d8e901165

Observation 632c41a8-ad18-440c-8953-02c144402482 · outbound

This paper cites Improving Self-Supervised Learning-based MOS Prediction Networks.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features Improving Self-Supervised Learning-based MOS Prediction Networks

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.445952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.397560Z digest=sha256:4c40d15109d7987a9e5f1a6c0ea56478c3c9a9d61b5f1fa204d1d2d717ea16a6

Observation 567d4e74-2d84-44c6-a656-ca38ed976c65 · outbound

This paper cites ESP- Net: End-to-end speech processing toolkit,.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features ESP- Net: End-to-end speech processing toolkit,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.531013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.402117Z digest=sha256:9dc1977e8ec24438c3486d1fbbcdde31848ba5308eb249ef4231cf45f5b90964

Observation b41c0c1c-fce8-4f73-bbbc-cc8666f5f234 · outbound

This paper cites CSTR VCTK cor- pus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features CSTR VCTK cor- pus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:50:08.518311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.405900Z digest=sha256:6c0a70f747a9812c33ff616d07e9674f67b11c9312f67c0541da9917c2be21a4

Pith citing papers

Observation 9029bd3b-988a-4b5d-a2ce-96aa0b219744 · inbound

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features cites this paper.

SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:50:08.503803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T18:50:08.261278Z digest=sha256:8cd9a3a63c13dc3bf6d80f18e84727a9bc63cab59acbfc6e20e5bd3da6356d01