Pith. sign in

Paper Citation Record · LEDGER

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.03021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03021 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:20:42.798080Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c48ec5ef-93ba-458c-b48d-85d221fa485d · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Soundstream: An end-to-end neural audio codec,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.490899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.490899Z digest=sha256:ee1da27e0ca89187520822e6688f572dad342e03907a0d3cd0bd02c0b9c21203

Observation 7ee4344b-31a1-44a7-b153-63acbfd0edc8 · outbound

This paper cites High Fidelity Neural Audio Compression.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation High Fidelity Neural Audio Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.496573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.496573Z digest=sha256:ec2933d9e818973402c301c13a9bec520ebe61cbf6ff8f152d7bc55483bce645

Observation 06ed9330-9b54-4bae-8762-5cb022b7f751 · outbound

This paper cites Language models are few-shot learners,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Language models are few-shot learners,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.502395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.502395Z digest=sha256:b257a4051dd267330d68fac81c9420d186f797a6591c02181e26ec402d5945c2

Observation e3cb3fea-7c16-4ebc-85c5-df80d5f6a30c · outbound

This paper cites Zero-shot text-to-image generation,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Zero-shot text-to-image generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.633663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.507927Z digest=sha256:f6e5817ad752c4e7626d05dd8400cd11faa5e5c873d2bf898c469945332c3e80

Observation add13ae3-cab1-4af9-a984-38732bfc7505 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.513042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.513042Z digest=sha256:cbce56fdec4358da522dd7927e08d8c11b4cbdecfd7e63674f39558e9f029e9f

Observation 22b9c149-dacb-4662-bc31-3d8d9bab254d · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Audiolm: a language modeling approach to audio generation,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.519016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.519016Z digest=sha256:d29d0355e5672c70ea679adaa77052d34a9f94dc78e8f635797832260d1fcd61

Observation 6c8d2918-b001-460b-b8d6-aec5bcdcf6fa · outbound

This paper cites Simple and controllable music generation,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Simple and controllable music generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.605938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.524182Z digest=sha256:4baaaeb9b721fd13a07c8e8197655109550a73631b242aebf190f9d83fb9f1e0

Observation ef762dc9-c2ec-4904-b46b-de4311c0ca1f · outbound

This paper cites Codec does matter: Exploring the semantic shortcoming of codec for audio language model,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Codec does matter: Exploring the semantic shortcoming of codec for audio language model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.588572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.529427Z digest=sha256:abece41ebddc5685f0f87b87fb3a71b0dc4977bba64861ed2d80e865877c95b2

Observation 0ca0fea9-eb9e-4bb3-8543-ea1ff80a54cc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.533990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.533990Z digest=sha256:28f98fb07396ae282f0c0e53668b8b311a5b1491c6a68344a1067ff9e52143a1

Observation 9efab4d2-cd97-49f1-94fa-edeb1e746723 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.539324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.539324Z digest=sha256:9543ad80356ac15f2f0bbdee0e7377fab93ded9fa88110a745e81e68794d412f

Observation 960591ba-4569-4579-a413-d242cfde5e92 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.544907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.544907Z digest=sha256:4e58082c5d79b55314effc7298f195d7b0bd78001144ad418a541a240d8ce1e8

Observation a434ea2d-a5f3-4c61-b77b-df25f926ee55 · outbound

This paper cites Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:20:43.354150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.550446Z digest=sha256:6f9c935fd9a0f21a56b90d96a3e0260fe0bf39db5e09dd9761e41f33abf21b78

Observation 74272db4-2a6a-45bc-91b6-432dd1ab0927 · outbound

This paper cites Scaling under-resourced tts: A data-optimized framework with advanced acoustic modeling for thai,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Scaling under-resourced tts: A data-optimized framework with advanced acoustic modeling for thai,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.572568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.555937Z digest=sha256:083ad4286c6fab48326058185760147684fbe8fb1695903e3bf768d510da91f8

Observation e53bb86b-848b-4872-8df8-c59b8e3ffc94 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.560749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.560749Z digest=sha256:f34bec2cbfc8dd89c7e847ec1b5c88acf495b6c6cf84e0835fbb3c1d26bb1480

Observation c4a2efb6-a068-4e6c-b8f7-f812fa96a3ac · outbound

This paper cites Vevo2: Bridging controllable speech and singing voice generation via unified prosody learning,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Vevo2: Bridging controllable speech and singing voice generation via unified prosody learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.566266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.566266Z digest=sha256:c3d6cac9be3969a383f602d734a67f97badf4cb8fa710ee9a99f7bd5f54069a1

Observation 065e4e4a-a1e1-419a-803b-824cdb89d15e · outbound

This paper cites Accompanied Singing Voice Synthesis with Fully Text-controlled Melody.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Accompanied Singing Voice Synthesis with Fully Text-controlled Melody

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.571505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.571505Z digest=sha256:278f579379d5fb7f55c32793b4aae138ec809946556becb125efcb57ba485704

Observation bf76ce6f-7390-45bd-b71f-38ca88797097 · outbound

This paper cites Moee: Mixture of emotion experts for audio-driven portrait animation,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Moee: Mixture of emotion experts for audio-driven portrait animation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.556277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.577736Z digest=sha256:29cdc1b086c4650d35cd7309d9bc405fd6fc092ddcf50fcf248d5e82beeb3538

Observation 05eda9be-a5a9-481f-9728-d3cce7c63990 · outbound

This paper cites CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:20:43.188175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.582813Z digest=sha256:5b960520632d12ae838848d6556179d3dd31a23a7409bdaec707f3cb4b473feb

Observation b05783ca-fdf2-4451-a276-f9e3b75fffad · outbound

This paper cites Neural audio codecs for prompt- driven universal source separation,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Neural audio codecs for prompt- driven universal source separation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.539884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.588566Z digest=sha256:c469f16006d8b8db0707b6a3046c1df925f0ca7561400b1f271116197740e036

Observation 0da3e5b7-04ef-4702-b681-7569660fc2d6 · outbound

This paper cites VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation VaSAB: The variable size adaptive information bottleneck for disentanglement on speech and singing voice

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.593969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.593969Z digest=sha256:9a41a5c9d25874710989772d1e15b32de36fc495952e7976bbeea671c641d4dc

Observation 4ea00bab-c182-4192-b1b2-aba2e404c08a · outbound

This paper cites Generative de-quantization for neural speech codec via latent diffusion,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Generative de-quantization for neural speech codec via latent diffusion,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.523262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.612516Z digest=sha256:9fbb4456f6523babf13ac68b565d65d16f25d9c6656840ebe769859c5c12ee07

Observation 64724a5d-b20d-420d-af6c-84f4d10c480c · outbound

This paper cites Hq-svc: Towards high-quality zero- shot singing voice conversion in low-resource scenarios,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Hq-svc: Towards high-quality zero- shot singing voice conversion in low-resource scenarios,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.641478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.641478Z digest=sha256:b1cb3a38855303812d128bb590310028ab01c580f8603695308bc11b64dd85b0

Observation bff9680b-a80b-4765-805a-8c75a96ecb0d · outbound

This paper cites Muse: A multi-agent framework for unconstrained story envisioning via closed-loop cognitive orchestration,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Muse: A multi-agent framework for unconstrained story envisioning via closed-loop cognitive orchestration,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.660039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.660039Z digest=sha256:752364bb9237dc89f3bc0022c5b48d53c41692c622e14b3f48714cb1d0a872cf

Observation 91213f84-0abe-44fc-b610-c1219d93b31c · outbound

This paper cites Anyac- comp: Generalizable accompaniment generation via quantized melodic bottleneck,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Anyac- comp: Generalizable accompaniment generation via quantized melodic bottleneck,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.690555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.690555Z digest=sha256:510a33fc8b1fbd9690d5149a69538dc4d3c05674d5efa3fe887ae4ea501d3721

Observation 0972bb41-116a-44b7-8886-5d6d703f1f66 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation High-fidelity audio compression with improved rvqgan,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.706234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.706234Z digest=sha256:351f1445ea85f7e69a5869ddd1b235f440ce880c2636fe56488f6df8266a3c5c

Observation e527c669-83ca-4ce4-85fd-0eaeae17eb4e · outbound

This paper cites MuCodec: Ultra Low-Bitrate Music Codec.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation MuCodec: Ultra Low-Bitrate Music Codec

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.740593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.740593Z digest=sha256:e77d69f609e637a77ef2939c4b35d7647f919d1e1d803bc74359e596aebb30fa

Observation 5f5ae095-07c6-4950-b631-9e75e266b7d2 · outbound

This paper cites Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Opencpop: A High-Quality Open Source Chinese Popular Song Corpus for Singing Voice Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T04:20:42.775161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:20:42.775161Z digest=sha256:c6222f3e5a29aa49cc75bf86826acc6f2a1ede3ec9082bda3fe8fd5a94732125

Observation 75f099c2-92fa-4e00-b64f-72aeb23bdc29 · outbound

This paper cites Visqol: an objective speech quality model,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Visqol: an objective speech quality model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.497025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.780448Z digest=sha256:0f5b2d4ba21c64ca03ce85674b025be8f5bd72749ec9c16eb3cd941de9c4b145

Observation caa3e414-3588-4465-821a-7451fd51e340 · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.480460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.788441Z digest=sha256:929ec96e8ce61bebf20e2c46c175e70a246abedabd0ef9b9e95b7fd0582c4025

Observation 53c5d624-6bd3-475a-9800-5b664c821e41 · outbound

This paper cites Resnet 34,.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Resnet 34,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.465190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.793281Z digest=sha256:ad08b40916fabdf3159f15dcbd23b0bd40d8d16836827abb7b3b8ab548560d53

Observation 4dfe5f9e-2eb5-44e8-a35c-e4b7b20446d7 · outbound

This paper cites Towards the next generation of web-based experiments: A case study assessing basic audio quality following the itu-r recommendation bs. 1534 (mushra),.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Towards the next generation of web-based experiments: A case study assessing basic audio quality following the itu-r recommendation bs. 1534 (mushra),

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:20:43.450570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:20:42.798080Z digest=sha256:3f4043c986cd2906a348deb13d8d876cfd4bc73bceccaa2ae7a960fbafe4586d

Pith citing papers

No inbound Pith citation observations are available.