Pith. sign in

Paper Citation Record · LEDGER

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.05899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05899 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bc31a3e-2706-46af-8798-ee2fc057ec8a · outbound

This paper cites MusicLM: Generating Music From Text.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.146491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.146491Z digest=sha256:2f963c9047df6521ffaad6467c34e46ab63f642f892bdea58f2a1fc8992101ec

Observation ad315d80-4b72-4627-83d5-86978052c1e7 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.194385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.194385Z digest=sha256:35e9d33411c5141db2b3d305152236e4d3532a5da5262944d92ab70002bc501f

Observation bcc3cde5-97b3-4695-8872-6c514b116744 · outbound

This paper cites Fast timing- conditioned latent audio diffusion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Fast timing- conditioned latent audio diffusion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.613505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.243298Z digest=sha256:f71d4453affe5b08faf5409e26a771599fc2ee6a92656f8d66df9a2920335433

Observation c2eda0e4-5e05-4345-8d6b-d655fcd3f419 · outbound

This paper cites Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.450106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.337412Z digest=sha256:ae7c1a39087c3fe94702c39ec6e666e9b16949c00ce44edcedcf8df1d786ed43

Observation faae7500-a287-40ef-b230-0f8e7653af5a · outbound

This paper cites SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.291059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.454075Z digest=sha256:0fb6261ada933a9fa9b3c07196b8327e2fa458b33fef060de8fa983dde4cfc1b

Observation c88a3958-4fb3-4d70-9463-13037c7bb64e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.161316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.513674Z digest=sha256:a367db6684808cc03c2ceebebb73789e9713d1b3933f23f67d3c153d5b5de3a7

Observation 009ebc16-5eae-417e-a12e-54ec420aee6e · outbound

This paper cites MOSNet: Deep learning based objective assessment for voice conver- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MOSNet: Deep learning based objective assessment for voice conver- sion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.028276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.583734Z digest=sha256:ca7f18ad90ea1998eef2f47c034743d7960822076c50a5dd618328c5ddcdd397

Observation 9d4e82e1-9659-4400-b67a-79047a6b82c5 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.648042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.648042Z digest=sha256:f9764afeed99bee77871185cbaa307e77a3005c13ddcbbab5cb032b4723d3754

Observation ab3abf6f-2794-4f18-af57-8b4cd7e9445c · outbound

This paper cites Qwen Technical Report.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.739123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.739123Z digest=sha256:591005b814375d0e5a320ed3bbda9b184d233a3df897bb69e1a1a67815a396bb

Observation f12e455d-4998-41ec-b50a-b64df0ffd3f0 · outbound

This paper cites Efficient optimal transport algorithm by accelerated gradient descent,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Efficient optimal transport algorithm by accelerated gradient descent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.868162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:54.811164Z digest=sha256:63321bf2bf5be0bf5bd552b8c04691de69eaf898563a310cc137b0d7f6b51cb8

Observation eceb604f-d319-4102-b005-320ad57c850d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.921831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.921831Z digest=sha256:1bb3ebc12237e39befa91c3e8f7214856cbb59395ac841d62acd276081533a9f

Observation 2fb0be64-6805-4edd-bb43-a0989960304e · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.010659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.010659Z digest=sha256:19d10f396547b0f10c1427cd4b0ab5056d4f07cfa0269a86287ca0cee222dac3

Observation 6cf873c1-3055-466d-8794-ed5457ae00cb · outbound

This paper cites Simple and controllable music generation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Simple and controllable music generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.722622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.126990Z digest=sha256:f2faeb6cd8d0e24362be1fd5d5f40fc4701abf14d32b2ecd9cbc4370ca9745ea

Observation b117bfbd-7da0-466e-a6db-321892dbb1ee · outbound

This paper cites Mo ˆusai: Efficient text-to-music diffusion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mo ˆusai: Efficient text-to-music diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.515398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.247126Z digest=sha256:febaee290dfefacdf1175aa5b39afcd55774323180494d0803f20ca563deb87f

Observation 110b80d9-f855-4f71-a97d-7964703822bf · outbound

This paper cites Musicmagus: Zero-shot text-to-music editing via diffu- sion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musicmagus: Zero-shot text-to-music editing via diffu- sion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.341182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.371848Z digest=sha256:a9a48776eff7d931ed61ba7e8aed41f886d3429a34876235393c9eb64db511e1

Observation 32b714ee-bd08-48cd-a7fa-e2067ef54485 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.490675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.490675Z digest=sha256:3eb7538939fbea14d8775a21e0364d0fe63adeca0b2c785472995d1bd42b2d36

Observation 948aa870-0d5d-40a3-8c25-d567d71de983 · outbound

This paper cites MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.607323Z digest=sha256:1ea2708d13200f5f6c56687c45d006bfca22ff910db6e4e73d59a42f9f1bae7b

Observation fcda9ac2-ff87-4e34-9b23-1215077a15cb · outbound

This paper cites Mospc: Mos prediction based on pairwise comparison,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mospc: Mos prediction based on pairwise comparison,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.727593Z digest=sha256:68786d2ff8a0f7e8a0ab652f48a34d31c8646c2bca253c552d2c3d286e59baca

Observation e78a8927-adc8-4697-9afa-7ca3b1b274bc · outbound

This paper cites Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.910254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.786881Z digest=sha256:5d7aa37d4c8fa4c67ade8a9d63a955eb6aa5d252ead7ede5b5ff324cffdda061

Observation 622b45c8-a352-4575-9197-7d1cc3bbd279 · outbound

This paper cites Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.645210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:55.905504Z digest=sha256:7216c8867c033d208d24c13856fb81fe68104ef20fbc8bdc13b1d3e519ef187d

Observation 83e39c97-ec97-486e-a77b-473dd56ec2ef · outbound

This paper cites APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:56.066230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:56.066230Z digest=sha256:c6e75a5836b90f22640aa3b87f3acfca01c96b17548baefd9054aed8af93e2d5

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:f4b7a2c183d36b2913209b3d07cf6c5985bc2fc08f76884a75edeaacf563388a

Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · outbound

This paper cites U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:56.749500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:56.294362Z digest=sha256:55883ef762af2c8e6afe098643a55a4129a0c58962fb5f43bc1e452ed5489f53

Observation 258d12b3-f868-4a2d-b655-afd4cdcc71fe · outbound

This paper cites Cmot: Cross-modal mixup via optimal transport for speech translation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Cmot: Cross-modal mixup via optimal transport for speech translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.395890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:16:56.451078Z digest=sha256:7461fbe2180803ca3e5dc6c788410abb0ba7edb01b9ca01b40f5c87db8d3da65

Pith citing papers

No inbound Pith citation observations are available.