Pith. sign in

Paper Citation Record · LEDGER

Long-form music generation with latent diffusion

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2404.10301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.10301 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:26:56.478713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.451108Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 160c66ad-a6b7-4f0d-9c1c-fedb914090e1 · inbound

Compression of Higher Order Ambisonics with Multichannel RVQGAN cites this paper.

Compression of Higher Order Ambisonics with Multichannel RVQGAN Long-form music generation with latent diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T18:03:55.409469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:03:55.409469Z digest=sha256:47fa316b3c6cf472772b33776ec3a7ed671f38654340145b34eda70e62db5af8

Observation c37ee982-b157-47eb-8672-99a29fa3df47 · inbound

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models cites this paper.

Improving Controllability and Editability for Pretrained Text-to-Music Generation Models Long-form music generation with latent diffusion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T17:22:17.879005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:22:17.879005Z digest=sha256:a66013089102034a5f27d48f78717c3846dd39c148f223cb020b111067dfde79

Observation f7d6bd55-941c-47ff-aae2-fe5f5474c818 · inbound

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features cites this paper.

VidMusician: Video-to-Music Generation with Semantic-Rhythmic Alignment via Hierarchical Visual Features Long-form music generation with latent diffusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:53:26.954271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:53:26.954271Z digest=sha256:8c2b1416193ad90c277061398548682d2e4fcbab0b0ae70af93a76e988e45545

Observation cb3edcbf-530e-4901-98df-1f02fd63055c · inbound

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor cites this paper.

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor Long-form music generation with latent diffusion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:52:03.756797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:52:03.756797Z digest=sha256:48fb78404eaa9db85bf0ff728afed60ca9adeaf86cc6fdbae00999bf40447e20

Observation 9ff3bd98-15fe-40b3-aa88-26a6ca38b496 · inbound

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment cites this paper.

FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment Long-form music generation with latent diffusion

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:47:56.064401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:47:56.064401Z digest=sha256:92ed2e573ed823f3fd6b0c70b6650f421fd45817aa230733d6e7b32f9faf7714

Observation b4cf4c1b-afda-4ac7-bf37-dbac0f24533b · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Long-form music generation with latent diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.716105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.716105Z digest=sha256:00cc654489b96452f543077adfe1b5120a3cc8a7d8feb3f7517f071ae631398a

Observation bd53f9b4-b7ed-4ea9-ac89-ca6ec49087b2 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Long-form music generation with latent diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.508244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.508244Z digest=sha256:10f24e451efa00f44674dd42d92c8bd398b5f1caf7e7c279b68a6bc3a6cefa61

Observation a9e8101e-7f58-4e99-94e3-769e1ffca8f8 · inbound

Bridge-SR: Schr\"odinger Bridge for Efficient SR cites this paper.

Bridge-SR: Schr\"odinger Bridge for Efficient SR Long-form music generation with latent diffusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:36:57.433761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:36:57.433761Z digest=sha256:a6d826625cc6611f6d033d4ef10cf63db7a2822953ec2b26ff15031ad94caa20

Observation 977ec0a8-ce75-480f-8b9d-719d97fe4c4a · inbound

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation cites this paper.

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation Long-form music generation with latent diffusion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T20:12:29.405437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:12:29.405437Z digest=sha256:8b24a9de7e61a28df0e0ccc6b4a81cb1bdbabfc9f7c734631b0183e797053949

Observation f94b4d3f-73cd-4986-9216-18231a889e31 · inbound

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding cites this paper.

Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding Long-form music generation with latent diffusion

Reference 1940

Resolution
unresolved
no resolver link, observed 2026-08-15T22:26:56.478713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:26:56.478713Z digest=sha256:acda7488053db6e17a453d7e4d18fab9bb38b35cda555d725dbeab17731611df

Observation c55e1cf7-8ac6-413a-9d38-0f61d29607c3 · inbound

Fast Text-to-Audio Generation with Adversarial Post-Training cites this paper.

Fast Text-to-Audio Generation with Adversarial Post-Training Long-form music generation with latent diffusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:09:43.340083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:09:43.340083Z digest=sha256:8b6353bd4ffda0bd4636b11120d964ab985b7fad25d903cced8738d448b746aa

Observation 18e83ac3-3a05-4a89-a806-c1f937e01239 · inbound

Universal Semantic Disentangled Privacy-preserving Speech Representation Learning cites this paper.

Universal Semantic Disentangled Privacy-preserving Speech Representation Learning Long-form music generation with latent diffusion

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:19.426727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:19.426727Z digest=sha256:8317e882786fbe151e258fc8fb8e2b332f49b81ce2e63dbd1e72259882eb872e

Observation 351fe317-95a3-40ac-a656-48523bf94e57 · inbound

In-the-wild Audio Spatialization with Flexible Text-guided Localization cites this paper.

In-the-wild Audio Spatialization with Flexible Text-guided Localization Long-form music generation with latent diffusion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:46.716586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:46.716586Z digest=sha256:6c46f82559c09174b19b4fba8e767f0bbe5c184911341f1565e8f0a507a9cdd5

Observation 237bc281-e6e2-478d-99bc-f3e97e626b97 · inbound

WAKE: Watermarking Audio with Key Enrichment cites this paper.

WAKE: Watermarking Audio with Key Enrichment Long-form music generation with latent diffusion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:26.350553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:26.350553Z digest=sha256:9d09dc7f331e5e1a52f0c024d92c80047f01d63871ab4d6dd15ae50ba0de7bd9

Observation c272e73f-2a7a-487d-9aae-b8c315fdf398 · inbound

Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation cites this paper.

Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation Long-form music generation with latent diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:42:05.689145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:42:05.689145Z digest=sha256:57d9be94278d26c91cec3a2af464625f3934d3a26afb645d7aae3833b0660ebb

Observation 10d9b764-8e6b-474f-819d-61db03b33a6b · inbound

ASAudio: A Survey of Advanced Spatial Audio Research cites this paper.

ASAudio: A Survey of Advanced Spatial Audio Research Long-form music generation with latent diffusion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:54.920307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:54.920307Z digest=sha256:00eea7fa5a4d855710a408e7814f9381391ab4a023d58f6e429ff59b739a4d64

Observation d630949b-9041-49c4-b6a2-89e5cf734cab · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model Long-form music generation with latent diffusion

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.857234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:06cfdc6dfaa46d97d12cfa191f51c5aae41e6dbce759b89b7c0c6ba5f84aebed

Observation 1684b2a8-9e63-471a-bbc4-f3debbe0da5b · inbound

Seconds-Aligned PCA-DAC Latent Diffusion for Symbolic-to-Audio Drum Rendering cites this paper.

Seconds-Aligned PCA-DAC Latent Diffusion for Symbolic-to-Audio Drum Rendering Long-form music generation with latent diffusion

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:39:22.183244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T18:38:43.559539Z digest=sha256:8f7b98e32b9540afa7bf3ee179fbe2f1bbd09db59a1f3b21ee6ea42afbc05710

Observation 180e0950-93d0-4cab-9479-d8fe8e82756b · inbound

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators cites this paper.

Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators Long-form music generation with latent diffusion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T03:25:58.995820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-22T03:24:50.604019Z digest=sha256:fb4a75cc9fdc2b7f13ab1b8bc5e310be98d9dd4e9a59b724a3fbe9ebecbbfadf

Observation 71065932-b94a-463a-8a0a-5fdda690985a · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation Long-form music generation with latent diffusion

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.779223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:da6e43dbf9f99052e5b03effd1e54426bd3effb04b1bccf61307bab49ce2b589

Observation 1d69060d-9e2a-4a32-a8e5-3b446e87df76 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Long-form music generation with latent diffusion

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.452391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:a2410dd66cd97f1a6cb136502140ee042ae00d753c8bafbede7ca6a23ac703e4

Observation d6a3e05c-0ad7-42ab-9d30-eb6c2525024d · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Long-form music generation with latent diffusion

Reference 111

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:27fc31d2ad5a7f112eda4b62e0de475f3efe7a61b7a03109993f09d7a738e1a6

Observation 85dbf6a9-efbb-4786-8349-a07511050edd · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens Long-form music generation with latent diffusion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:48.723618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:48.723618Z digest=sha256:ab1e8dcc33680144a1aa1c3731a3bab75c6162635d038ea0f42eda65f71dd9b2

Observation be5684af-b5ae-4839-a628-2b5371d2a9eb · inbound

On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs cites this paper.

On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs Long-form music generation with latent diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:08:31.796045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:08:31.796045Z digest=sha256:0bbdbda8d185ab6b5d11c5d2d04619135aa79d64715a71a29d4c408f926aaa20

Observation 775a07d5-4098-423c-b0a6-ba1b3b538eeb · inbound

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities cites this paper.

AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities Long-form music generation with latent diffusion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:43:18.097916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:43:18.097916Z digest=sha256:15cdcf4659c8199d3a5ef9670c86127724021065aa95a4213ca372d2ad842592

Observation ebe00892-9970-4c75-9100-2afe2ff70979 · inbound

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation cites this paper.

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation Long-form music generation with latent diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:19:59.851959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:19:59.851959Z digest=sha256:98beaeeeb4ff071759fbe8cbdc11af9b9e3b123821e2af346382fee2927e8fa4