Pith. sign in

Paper Citation Record · LEDGER

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.05899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05899 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:16:56.451078Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bc31a3e-2706-46af-8798-ee2fc057ec8a · outbound

This paper cites MusicLM: Generating Music From Text.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLM: Generating Music From Text

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.146491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.146491Z digest=sha256:21e391d56b49724a12d9eda73c8e78d9892306233bf52dfe790268038682c857

Observation ad315d80-4b72-4627-83d5-86978052c1e7 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.194385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.194385Z digest=sha256:99ed9cf30b8c8391a56df06033ffbc53781711f5d2875b197162a78e18ba5d57

Observation bcc3cde5-97b3-4695-8872-6c514b116744 · outbound

This paper cites Fast timing- conditioned latent audio diffusion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Fast timing- conditioned latent audio diffusion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.613505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.243298Z digest=sha256:3bc71cc90f1febca99d9153c802dbbd430d8a3ab2dc47f161473fd10e321b32c

Observation c2eda0e4-5e05-4345-8d6b-d655fcd3f419 · outbound

This paper cites Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musiceval: A generative music dataset with expert ratings for automatic text-to-music evaluation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.450106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.337412Z digest=sha256:3979f23fa773113725977fb788f2f2f8a9b4b24bbaf4647be8e6fde4077b920d

Observation faae7500-a287-40ef-b230-0f8e7653af5a · outbound

This paper cites SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction SSL-MOS: A Self-Supervised Learning Based Approach with A Transformer Target Model For MOS Pre- diction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.291059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.454075Z digest=sha256:8ec85459adc476b769b3d26d3974c831a031e4b8b851c2f7c386ac553bbfee1b

Observation c88a3958-4fb3-4d70-9463-13037c7bb64e · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.161316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.513674Z digest=sha256:4b74229ff09354e0825bd58dd806d2089b5ce8aa4500cc9dfbe3edefccbcb420

Observation 009ebc16-5eae-417e-a12e-54ec420aee6e · outbound

This paper cites MOSNet: Deep learning based objective assessment for voice conver- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MOSNet: Deep learning based objective assessment for voice conver- sion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:59.028276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.583734Z digest=sha256:e5c6b4bddc5c8188871d5ba92b73cbebf8526b97e6803c53ec45a63283c9ffa7

Observation 9d4e82e1-9659-4400-b67a-79047a6b82c5 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Robust speech recognition via large-scale weak supervi- sion,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.648042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.648042Z digest=sha256:df4832337c9ee338ac50230822190dbf618ff7330f880fd2dee866d3cf98c5c4

Observation ab3abf6f-2794-4f18-af57-8b4cd7e9445c · outbound

This paper cites Qwen Technical Report.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Qwen Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.739123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.739123Z digest=sha256:09187240261ee61725d1461fb7e01e5d85778e21c098ec4483651f71081ad080

Observation f12e455d-4998-41ec-b50a-b64df0ffd3f0 · outbound

This paper cites Efficient optimal transport algorithm by accelerated gradient descent,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Efficient optimal transport algorithm by accelerated gradient descent,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.868162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:54.811164Z digest=sha256:628b696bc21d71d40eb19524470f69a56771259a5b9a2c449c3844b981135bef

Observation eceb604f-d319-4102-b005-320ad57c850d · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:54.921831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:54.921831Z digest=sha256:4227bc8e20bf779deca6f3273efa9c6071a9bda7747bc33e7499ef18cd61e5e7

Observation 2fb0be64-6805-4edd-bb43-a0989960304e · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.010659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.010659Z digest=sha256:0ec62b28c3e089e081187f4df5c6310fadebdba68bfdd6ed6659afe2def5703f

Observation 6cf873c1-3055-466d-8794-ed5457ae00cb · outbound

This paper cites Simple and controllable music generation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Simple and controllable music generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.722622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.126990Z digest=sha256:770d2ba670c21f4d204fd53e45871c49b267b4ed5ec881b4c092770ad988fd6d

Observation b117bfbd-7da0-466e-a6db-321892dbb1ee · outbound

This paper cites Mo ˆusai: Efficient text-to-music diffusion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mo ˆusai: Efficient text-to-music diffusion models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.515398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.247126Z digest=sha256:481c03b7edc1522387f8f3563f40db9f79074a63bfed4f32482762f10842c015

Observation 110b80d9-f855-4f71-a97d-7964703822bf · outbound

This paper cites Musicmagus: Zero-shot text-to-music editing via diffu- sion models,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Musicmagus: Zero-shot text-to-music editing via diffu- sion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.341182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.371848Z digest=sha256:fdb76bab79814ef7b4a56e9dd2e302268c8d92fcefe60f3af5c512694b87c7bb

Observation 32b714ee-bd08-48cd-a7fa-e2067ef54485 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.490675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.490675Z digest=sha256:d774726a1a9d14c4502ab63748c5e489e05ffce73f4f517f8b9e95ae80c50861

Observation 948aa870-0d5d-40a3-8c25-d567d71de983 · outbound

This paper cites MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction MusicLDM: Enhancing Novelty in Text-to-Music Generation Using Beat-Synchronous Mixup Strategies

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:55.607323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:55.607323Z digest=sha256:d9ff67ea4be65b33f1a0bde90893a882ae8e9ea7099875d74c78da7a489e23dc

Observation fcda9ac2-ff87-4e34-9b23-1215077a15cb · outbound

This paper cites Mospc: Mos prediction based on pairwise comparison,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Mospc: Mos prediction based on pairwise comparison,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:58.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.727593Z digest=sha256:02ec4d819f81fa2b048b3d1f8339d0db21ee57bbe2366b44c3813ff18c299fa3

Observation e78a8927-adc8-4697-9afa-7ca3b1b274bc · outbound

This paper cites Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Resource-efficient fine- tuning strategies for automatic mos prediction in text-to-speech for low- resource languages,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.910254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.786881Z digest=sha256:4dde1f2bcfd1993a1e6d48658ebe1b99b3d2ee97cca76e4f7a9119fa051fd5ea

Observation 622b45c8-a352-4575-9197-7d1cc3bbd279 · outbound

This paper cites Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Zero- shot out-of-domain is no joke: Lessons learned in the voicemos 2023 challenge,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.645210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:55.905504Z digest=sha256:2752396eb5c152dd847ef9e856521220f9f9915017028f2fb601ddd85bfa0d32

Observation 83e39c97-ec97-486e-a77b-473dd56ec2ef · outbound

This paper cites APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:16:56.066230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:16:56.066230Z digest=sha256:293fe64407c71881f131d1537936bfc2062d1a34ab40043081d6fe096c0d97d6

Observation 0e99131c-60e1-43be-b897-8b27fd569fc6 · outbound

This paper cites LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:57.039584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:56.175634Z digest=sha256:ddff4f6dd1209fd6e34073687917a3bd6373a8e6ee46575fad196b531595de02

Observation 44d951d0-8696-4fcf-921d-c7575a48ecc7 · outbound

This paper cites U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:16:56.749500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:56.294362Z digest=sha256:88eaab91deeb9d446e3c10cf0f166ce8034aa970759b4485af58fbfc6b03f09b

Observation 258d12b3-f868-4a2d-b655-afd4cdcc71fe · outbound

This paper cites Cmot: Cross-modal mixup via optimal transport for speech translation,.

WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction Cmot: Cross-modal mixup via optimal transport for speech translation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:16:57.395890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:16:56.451078Z digest=sha256:e3a363480f251be78308c9642847c98c1c7b60e20c9ba918ef77c59dcf626006

Pith citing papers

No inbound Pith citation observations are available.