Pith. sign in

Paper Citation Record · LEDGER

The AudioMOS Challenge 2025

As of 17 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 5 inbound Pith citation observations for arXiv:2509.01336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01336 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:41:28.171047Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T09:53:16.690620Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 71747265-4239-494d-9b26-0d3a57761db5 · outbound

This paper cites The V oiceMOS Challenge 2022,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2022,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:24.747365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:24.747365Z digest=sha256:09f962c86b86bd83e1a5d130d96fa5034ebbc0bf56ebd33ccfd1daaf276229a4

Observation 75c08112-cd56-493a-bcd9-95ed609c0cf4 · outbound

This paper cites The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2023: Zero-Shot Subjective Speech Quality Prediction for Multiple Domains,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.105692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.803002Z digest=sha256:84f7b95a698a4d18d0038bc735cd3882f71b092fef75ba29a3b329b4029893ba

Observation 9f15f70c-f59e-4775-97da-fec479570f35 · outbound

This paper cites The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,.

The AudioMOS Challenge 2025 The V oiceMOS Challenge 2024: Beyond Speech Quality Prediction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.090267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.838581Z digest=sha256:44c7aadd4b8177caa9593937a3d8ab5e7834c7dceadcef3ab13420ae0910c29b

Observation bfbc1959-0a0c-4d00-8ae3-09eb3a872dc6 · outbound

This paper cites How do voices from past speech synthesis challenges compare today?.

The AudioMOS Challenge 2025 How do voices from past speech synthesis challenges compare today?

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.076034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:24.911721Z digest=sha256:5cd180a6471c6bf399744672129ed982eb7c5db8c571af25d7f27efbd7a75e7e

Observation 9fee816d-8cca-43a4-89c3-636f263db927 · outbound

This paper cites Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,.

The AudioMOS Challenge 2025 Fr ´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.061380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.018687Z digest=sha256:674a54882206bad59266338a0f72f0be9677260b19fed178278092b16e746b67

Observation 7d788b0d-7c87-4918-9c0b-ce177cd587fc · outbound

This paper cites Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,.

The AudioMOS Challenge 2025 Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.044903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.099194Z digest=sha256:a860cda0f3df3a801d38d41d6c989792605d6943752cc9c10aca63c4588e2861

Observation 8cf02930-b725-4b01-84d8-8d5c8ef8c35a · outbound

This paper cites Evaluating generative audio systems and their metrics,.

The AudioMOS Challenge 2025 Evaluating generative audio systems and their metrics,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.029625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.176618Z digest=sha256:89e91268f2187a6ccbfa1abeaf02c8a406b35c30d3b78a53236dfa329cc2819f

Observation 444c3028-1fea-4c8b-ad67-517eb4e10b73 · outbound

This paper cites Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,.

The AudioMOS Challenge 2025 Correlation of Fr ´echet Audio Distance With Human Perception of Environmental Audio Is Embedding Dependent,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:29.014505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.264913Z digest=sha256:93b356952d3e4d3db3727db49977dafc40a604292c191d894dbdee1408302d62

Observation afbd36a3-ac50-4275-8770-a90feaf8348b · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

The AudioMOS Challenge 2025 Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:25.350124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:25.350124Z digest=sha256:14b83d1db05e5e5ff8a398d1e4dbc2e742022af1214759044fd949238623c28e

Observation 076c667e-3f55-40e0-b9fd-ef547b50bbd3 · outbound

This paper cites MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,.

The AudioMOS Challenge 2025 MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.995782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.426012Z digest=sha256:46f7838678204ee0ee488d5d979cada2a8f08ae470d17186a83957a24687421d

Observation 8a45dfc5-e190-4e62-8d36-03f44096efd7 · outbound

This paper cites LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,.

The AudioMOS Challenge 2025 LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.980192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.499732Z digest=sha256:a5bc6f643526bc0ca37fcdb2b944a7497ac5199178d97daddccea32ae0c61e55

Observation 3c74dd48-e844-46c3-a5f3-b34603c82a54 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

The AudioMOS Challenge 2025 AudioCaps: Generating captions for audios in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.964053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.614593Z digest=sha256:e92b640fdc238303df27312993963230acf8d299dfde00f00bb1c9d9df0a2240

Observation 6257b821-5018-4178-97e7-df8164efd9b9 · outbound

This paper cites MusicLM: Generating Music From Text.

The AudioMOS Challenge 2025 MusicLM: Generating Music From Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:25.791485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:25.791485Z digest=sha256:bf17a95a247dededf0386549f7e46eed1a9c48e31edae12796ad4f26babfb634

Observation 094207ec-0676-478b-b7b7-cc10dc9b3636 · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,.

The AudioMOS Challenge 2025 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.948461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.897729Z digest=sha256:a61a0db4dd18897a248076e4141b7c4391b95c414b63dc75c788b38dd6e703c9

Observation 8a3ff8ec-1567-4005-a81d-116bb2d934fc · outbound

This paper cites Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,.

The AudioMOS Challenge 2025 Hi-Fi-CAPTAIN: High-fidelity and high-capacity conversational speech synthesis corpus developed by NICT,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.932801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:25.973158Z digest=sha256:721df39daf0358b4a4f8157c7da754bfa06f2336c44cd729b2660a25b51d4a39

Observation 35552078-7a96-431f-94bc-d8529290e268 · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

The AudioMOS Challenge 2025 World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.915871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.096740Z digest=sha256:9d1690f65bb6fd963e08c9c9ca4a705fe1ce301651321ec0b1d658b6c1eb50dc

Observation c59f6e4b-80de-434f-b4ad-d173d91396c7 · outbound

This paper cites Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,.

The AudioMOS Challenge 2025 Fast Neural Speech Waveform Generative Models With Fully-Connected Layer-Based Upsampling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.898082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.238332Z digest=sha256:7d1de13a24de00245dd7e96caa77cf2c34d6218483b5e40b3b786171fca52e48

Observation 1716e873-33db-4514-9b46-f6a1c648661a · outbound

This paper cites Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,.

The AudioMOS Challenge 2025 Speech masking system based on spatially separated multiple TTS maskers with a compact circular loudspeaker array,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.879302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.300292Z digest=sha256:abeda9d720270fd1cfa49c76c1b65f52112fb6dc9e295938c11ef63aadcc51b9

Observation a4f16271-e27a-455a-bb07-aa84247e09b7 · outbound

This paper cites AudioSR: Versatile audio super-resolution at scale,.

The AudioMOS Challenge 2025 AudioSR: Versatile audio super-resolution at scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.860470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.410771Z digest=sha256:0c2dcb05cb871b940af1b1e758733d3a57b6a492a1cd750a04e656be19363467

Observation e0184f3c-dd9c-48ff-9dc4-9d0c847682df · outbound

This paper cites pyloudnorm: A simple yet flexible loudness meter in python,.

The AudioMOS Challenge 2025 pyloudnorm: A simple yet flexible loudness meter in python,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.844989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.520713Z digest=sha256:35a83afd5d6b2e2ce47fc89b60a26a1f1bfde29f7964304104bd390e9ad1e56b

Observation 7056e9ba-8ce3-49e2-a2c5-c4e50924387a · outbound

This paper cites Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,.

The AudioMOS Challenge 2025 Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.814314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.671368Z digest=sha256:ee8fe0f7953b89a534e34c459199c79957e49db3b93550909f866ada19647c77

Observation ed998dd8-45e3-4d88-a9e8-13eed8541ef6 · outbound

This paper cites HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,.

The AudioMOS Challenge 2025 HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.797594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.767113Z digest=sha256:1afd84725e07429b0189cfe9c3e4f500b208f5224fa58f34f5a530778c2ab215

Observation 0a6684ce-f2dc-4ee8-ad7d-d7ddcc3aa8c1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

The AudioMOS Challenge 2025 Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.859550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.859550Z digest=sha256:40179f5417bdd58a1ae13e574fba71eb0d1d17152c9f5a0ed095175997c5e932

Observation 172d95ca-1a10-4e60-866c-eebf56c26229 · outbound

This paper cites Attention is all you need,.

The AudioMOS Challenge 2025 Attention is all you need,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.890826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.890826Z digest=sha256:b1ebd4bf0a4ded89d54ae2d32326bbeb16cd1a35e856c9318ba71ec2ace2372a

Observation 17a12fbd-41de-4400-bc07-6938649d4dfb · outbound

This paper cites Layer Normalization.

The AudioMOS Challenge 2025 Layer Normalization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:26.917094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:26.917094Z digest=sha256:71fa13421c29bbe1919d8a77646ed262a2795df876742d0d42572d3ea51edc0b

Observation 9500da5a-ebd8-4855-b507-6863ad5ae1ad · outbound

This paper cites Gaussian Error Linear Units (GELUs).

The AudioMOS Challenge 2025 Gaussian Error Linear Units (GELUs)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.079074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.079074Z digest=sha256:12443c1f9b987dd80927d2f5825d4b07af2089ba4e393a72a20bc2885485a8c3

Observation a227ea61-bb4b-4bb2-ae57-34891ddc50c2 · outbound

This paper cites Generalization ability of MOS prediction networks,.

The AudioMOS Challenge 2025 Generalization ability of MOS prediction networks,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.140555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.140555Z digest=sha256:e70454e9cd4ab7512450283c1972e88c0c477aa0e7da392436f3919c965b09e7

Observation 3b1348ed-832e-45ef-9242-286d6553fe4c · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment,.

The AudioMOS Challenge 2025 PAM: Prompting Audio-Language Models for Audio Quality Assessment,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.748591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.250841Z digest=sha256:0fead42d153a52baf6c02b55e03edd8dc8295080d9e3d19e2bb1712b0a044e1a

Observation 818ff978-f69d-4b09-9641-8cef1b9b56b2 · outbound

This paper cites EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,.

The AudioMOS Challenge 2025 EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.733512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.363762Z digest=sha256:13da4c40e4f99135bfcc8271e61dfe6ab0d2bf7f7745bfd6b4b399d4f1fa0334

Observation 4cb3aca0-b399-434a-b226-15d309c40397 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

The AudioMOS Challenge 2025 Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:27.466107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:27.466107Z digest=sha256:98a9831f1992efd053d7b2c908faaa40019bcc3cfdbcb0e0f823609f2a89b138

Observation 26bc3c32-a67e-4203-96cb-347d0b747098 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers,.

The AudioMOS Challenge 2025 BEATs: Audio Pre-Training with Acoustic Tokenizers,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.718092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.569724Z digest=sha256:06518999ebe6285ef82088215d6bbc929a8483ae5ea8d922b489247dec35e724

Observation 67cf87db-4185-409f-a2b8-e18bf6b0b21e · outbound

This paper cites Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,.

The AudioMOS Challenge 2025 Masked Modeling Duo: Towards a Universal Audio Pre-training Frame- work,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.703862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.684183Z digest=sha256:c4bbd508541b074b131b909cfc034a3d60c5ed563cee88e7975b11be35435dcd

Observation 7eee63f3-1fd0-41d5-aff8-9dea27ca73bf · outbound

This paper cites High Fidelity Neural Audio Compression,.

The AudioMOS Challenge 2025 High Fidelity Neural Audio Compression,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.685458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.738336Z digest=sha256:2c29a7e64a1eb40f68f05d92ed76e893e88ab27fb30852f7f27554899faa3a5b

Observation fbd19bd8-0b75-4f1e-8863-3654bf650e33 · outbound

This paper cites Scaling up masked audio encoder learning for general audio classification,.

The AudioMOS Challenge 2025 Scaling up masked audio encoder learning for general audio classification,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.669001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.798487Z digest=sha256:8a3a453ee0482c1dcd6b7a01c66c87135fae93ed039db1fb9c02e728b2f04355

Observation 2c3b9206-cf74-418d-a760-716d816c152c · outbound

This paper cites MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,.

The AudioMOS Challenge 2025 MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.653507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:27.960529Z digest=sha256:76d18aca3e03be4e31e2369364539940b9ab0921749275f87e144a7dcef9cf15

Observation 9ea7b282-98af-4e89-a446-fe71f12d221b · outbound

This paper cites MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization.

The AudioMOS Challenge 2025 MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.086037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.086037Z digest=sha256:d73aa0dd7ac151efe555f17ffd9bd488259b45571c1521da4be050fc9ad2f56f

Observation 1808c671-6225-400b-8cb7-30eecc0d2c3a · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding,.

The AudioMOS Challenge 2025 BERT: Pre-training of deep bidirectional transformers for language understanding,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.638751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.091853Z digest=sha256:35a30578c8fee733a06119b23961d3bde702a7d4e7a11aa2f820e891bbbc84aa

Observation 371f0a5f-9676-4638-98b1-7420db71a21b · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

The AudioMOS Challenge 2025 RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.096665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.096665Z digest=sha256:246b81c93f3f8afc29f105d0c9b895320efff9aaeff03331da425cd8115fd35e

Observation 717265ef-56f9-4984-845e-912e8ccf3430 · outbound

This paper cites Qwen3 Technical Report.

The AudioMOS Challenge 2025 Qwen3 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.101543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.101543Z digest=sha256:63b8696dbbbef7f134cc6780725e8d9861e14edad8c5e2dcd55f78ad695bf6ae

Observation decdce1d-cc57-4be4-9034-0b74cec2acf4 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Super- vision,.

The AudioMOS Challenge 2025 Robust Speech Recognition via Large-Scale Weak Super- vision,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.624216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.106411Z digest=sha256:3f576f3ca764189dc8b745ed10e2d539901c0aad3510b8ef81b0838e6d2637b1

Observation 242f1acb-fb3b-437f-ac5d-02f21b625cce · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

The AudioMOS Challenge 2025 HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.608435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.111341Z digest=sha256:d73513dbba5ac270a6892dea746a4035114a34bf092e36650c1196c84206ee80

Observation 44ff0cea-57f8-4d50-ab61-f525cc04c564 · outbound

This paper cites Scaling Speech Technology to 1,000+ Languages,.

The AudioMOS Challenge 2025 Scaling Speech Technology to 1,000+ Languages,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.590327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.115839Z digest=sha256:8d7fc346451da74e87e6b05fcb0f063ecf9a8200b1d35c2f8a7537b34b523bf2

Observation 30883e1f-8453-4819-b884-c147b9987d9a · outbound

This paper cites EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,.

The AudioMOS Challenge 2025 EAT: Self- Supervised Pre-Training with Efficient Audio Transformer,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.571687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.121108Z digest=sha256:8fd95b366a953047e6c8f3a50e5e1015215ca83f8c157bfe1e25811be51e3d3b

Observation 66f7ec39-c620-4676-9128-c3b5ad480448 · outbound

This paper cites LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,.

The AudioMOS Challenge 2025 LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.549925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.126326Z digest=sha256:9a53f96cf4ca8f61e366fda6162b7336870ad264228ca6a8f4e1d5a48e37c4c9

Observation a147eebb-9fd7-4379-9024-91c8fc04aa03 · outbound

This paper cites KAN: Kolmogorov–Arnold networks,.

The AudioMOS Challenge 2025 KAN: Kolmogorov–Arnold networks,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.527216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.132640Z digest=sha256:8bac6079e118f5a5906da7f02ce6486ea94dca7bd887a1a87407fa1a1238d493

Observation 9688a5e9-93bd-485b-a80f-a456901d0779 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,.

The AudioMOS Challenge 2025 Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.505271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.138856Z digest=sha256:3d86beb457c1673aadeded287a8eef945d11bfeb89d9d08dac807e87e82973b0

Observation 5555346e-949e-49d5-b1bb-a279437c8bdf · outbound

This paper cites Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,.

The AudioMOS Challenge 2025 Sampling- Frequency-Independent Convolutional Layer and its Application to Audio Source Separation,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.484447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.144515Z digest=sha256:d73ad48bd66a8bcdedb1684896adc5379290efdb5d5f4bca08229290eedee103

Observation 68362bec-02a2-44b1-b7d7-fc29fbe516ec · outbound

This paper cites Kolmogorov-Arnold Transformer.

The AudioMOS Challenge 2025 Kolmogorov-Arnold Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T12:41:28.151523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:41:28.151523Z digest=sha256:223427091b6efe38258cdc63cd52fa99e03dda94527caa82c7d70c12580de158

Observation a58e257f-295f-4f00-9adf-156dd1676430 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,.

The AudioMOS Challenge 2025 VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.457972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.157430Z digest=sha256:063731f3f90b0c03aff704f09544ed381dae531076a70cd2cce9ed999bc4b85d

Observation fbbb3225-c21b-4a2a-afb7-d2c1f8254618 · outbound

This paper cites XGBoost: A Scalable Tree Boosting Sys- tem,.

The AudioMOS Challenge 2025 XGBoost: A Scalable Tree Boosting Sys- tem,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.436635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.165507Z digest=sha256:440698f50c491324f81d10ba6a43210a5398cb781d1139fab3a6da22302ae8d4

Observation 37341450-6e5d-4fa2-9674-b6724494b15c · outbound

This paper cites Pseudo Label Is Better Than Human Label,.

The AudioMOS Challenge 2025 Pseudo Label Is Better Than Human Label,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:41:28.419805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:28.171047Z digest=sha256:cf0247e9b61fb03ab064972f6f91144aada0581ae4568f23cd53f0a2a1a2fec5

Observation 635d2866-4a97-49a0-9603-f6d928cb71e9 · outbound

This paper cites an unresolved cited work.

The AudioMOS Challenge 2025 Unresolved cited work

Reference 150

Resolution
unresolved
raw_fallback, observed 2026-08-05T12:41:28.829801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:41:26.598556Z digest=sha256:4a41cae762dc1c5a8408072e40b091f64382131e9d88a1992dce0195daaf5007

Pith citing papers

Observation 08935185-338f-4cbd-a06d-1242a073ac96 · inbound

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling cites this paper.

Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling The AudioMOS Challenge 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T09:53:16.690620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T09:53:16.690620Z digest=sha256:103d46159429972bc7624ed2c5a31d8d88e90a092749ba50aae48512334731e6

Observation e76adabc-464e-4244-bdcb-bf4f0b2cb540 · inbound

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions cites this paper.

JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions The AudioMOS Challenge 2025

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:01:08.480283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T16:43:33.397158Z digest=sha256:6cc3ce32ddc27b607778003568b6e876c93685a4f69b5c6691283d1dd5f461cc

Observation a724a6a2-505c-4287-aab0-d21125a9daa1 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents The AudioMOS Challenge 2025

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.494952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:f9f1efb21eae247c3ec0652ae91de76086f6ea6fc131258a348ca6b5e43564e9

Observation 92cc6aa3-c70b-470b-8915-a73b1bfc21c0 · inbound

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation cites this paper.

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation The AudioMOS Challenge 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T13:57:36.846479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T13:57:36.846479Z digest=sha256:b0388b4e0f385d7e849c31562218625c408c63f7861b62f54fe9c91f95c35911

Observation e086a207-662d-4702-8c06-a55293b673a3 · inbound

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves cites this paper.

Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves The AudioMOS Challenge 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T20:33:36.029311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:33:36.029311Z digest=sha256:6c2351c3a14202adb4e3a3c5d73244c334dc598acc697aa6cacc0f1792a14ce7