Pith. sign in

Paper Citation Record · LEDGER

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model

As of 9 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.05222.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05222 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:54:23.104469Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a6515c44-4ef2-4503-b7bd-b62dc5b0f1e6 · outbound

This paper cites MusicLM: Generating Music From Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.857355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.857355Z digest=sha256:9e83f82d03a8b08fa0f7451c3a186209532928a577a996e94cb4b7ab4e251f1a

Observation 4f43d365-518a-40f4-84b4-30d389fc9897 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.867787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.867787Z digest=sha256:9f4a04ac6665f9f94cc8e809a52f04d8a8e5407f1491a3edc4324c23a92852de

Observation 273693a8-f2cf-469e-9228-d62a582ed790 · outbound

This paper cites InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2024-2024 IEEE Interna- tional Conference on Acoustics, Speech and Signal Process- ing (ICASSP), 1206–1210

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.539916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:54:22.873016Z digest=sha256:90ed054fa5471aa9671f1ad0e1edfbea4a082c5ff0d287a074573b17272555c8

Observation e1d1ea82-7ed2-438b-92db-8fe2228288cb · outbound

This paper cites InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1–5

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:54:23.524635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:54:22.887371Z digest=sha256:ec23571f8ba4bd76a831d8619dda6b266bbf1e986b29dcf5030c5b78c7c2c2c9

Observation ae39f69c-a169-4759-b809-9e29a31f997f · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Fast Timing-Conditioned Latent Audio Diffusion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.891858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.891858Z digest=sha256:0d4e927a58c200174e2270f7920f13789edb1c41b61c226e79c36594f987c404

Observation 105e8c7f-b948-40ec-9d5a-68954b12baaa · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.896826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.896826Z digest=sha256:5bba160afe8495f1577f0601ad53c757c1492963aa5bd3d101d1ada3f78d94ab

Observation cc7317ee-6fd8-49ca-b85c-d9dc23da835f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Classifier-Free Diffusion Guidance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.901792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.901792Z digest=sha256:4478d3848f5acc5b782be1691000135f2a62ccfb6aaf8c348d1fa30bfcdcb8a0

Observation 4f81140f-afe0-49b6-8345-0031d08cab2b · outbound

This paper cites MuLan: A Joint Embedding of Music Audio and Natural Language.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuLan: A Joint Embedding of Music Audio and Natural Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.907977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.907977Z digest=sha256:0fc3b187210a13226f68b215c846c11d666b3f81c4083682a4ea285ed407433c

Observation 05672a3b-b744-4038-b8c4-c0a1684da041 · outbound

This paper cites Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Multi-view MidiVAE: Fusing Track- and Bar-view Representations for Long Multi-track Symbolic Music Generation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T17:54:23.318894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T17:54:22.922897Z digest=sha256:6e7eb4c13f2c5fd1e51acb92a7848c786feea10b0cd226a50f65f5c840d55117

Observation c39156a9-052c-4d01-a835-1c5b26ce463e · outbound

This paper cites Symphony Generation with Permutation Invariant Language Model.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symphony Generation with Permutation Invariant Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.927741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.927741Z digest=sha256:ecf71970be7e87909e212e1cec22a0680b6de21872de6aeeaf8ffb82d4a394d1

Observation af262a4d-1262-4886-a43d-acd0802b97e5 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model MuseCoco: Generating Symbolic Music from Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.950607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.950607Z digest=sha256:f3cd12c87d513b0aa21cd2b1d558a3cae223e8718a8bf17ec8ec0223737de4cd

Observation 8fc956f7-7bdc-42f8-9c5b-bb93f9cf7fbc · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mustango: Toward Controllable Text-to-Music Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.979881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.979881Z digest=sha256:7f23df47c7514711d705d1d1bc13f42de264f4e41c41133f2fa719e4e195bc58

Observation 50bc3d6e-8142-479c-bc00-ec77c11e1c30 · outbound

This paper cites Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Polyffusion: A Diffusion Model for Polyphonic Score Generation with Internal and External Controls

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.032198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.032198Z digest=sha256:e6c7fc4cb43af41d3b17a6c272e6ac06ec0e4b72857f33a7f40557f9fe4eb2af

Observation 7ae14d82-eb1b-4d0b-bd1e-92bb6d562f48 · outbound

This paper cites Symbolic Music Generation with Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Symbolic Music Generation with Diffusion Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.055772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.055772Z digest=sha256:31273040db940fd823ac8f647b88ca1420bcd516acbe1a979bbf7a1f8535abec

Observation 3177bb59-396d-4118-8ae3-b8ca441ec781 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.080284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.080284Z digest=sha256:3174eb058da07b0d9645ca113a99bdf15756c7becefd0a9fb5cd9f61ff25bbbc

Observation 81e39820-6994-451f-99b8-a8c7a942db3b · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.085526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.085526Z digest=sha256:2f4773ea799490d0a9bf14e6529ba7b0e1b724b68c94003b49c86addd0c65c8a

Observation 44d78cd7-78f1-4711-9823-c20a555e641b · outbound

This paper cites Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.094990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.094990Z digest=sha256:ea1182fc220b61e772da4a546d7d7c179e6ed5bde7e567e9fba3d5beaa1ab4c7

Observation 3571aefe-3462-4d1a-8a65-1b5a2c7edae0 · outbound

This paper cites Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Exploring the Efficacy of Pre-trained Checkpoints in Text-to-Music Generation Task

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.099837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.099837Z digest=sha256:d482d778b621cbdf31b02d55e2bd6ad7190344513e7fd6116dc98d62b25970f6

Observation 663438e4-e295-4e3d-851e-943b5c9e488a · outbound

This paper cites Masked Audio Generation using a Single Non-Autoregressive Transformer.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Masked Audio Generation using a Single Non-Autoregressive Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.104469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.104469Z digest=sha256:051c6ad249aef6d426974015f1e78c24695400969ccf19a7071e42a048b24b32

Observation 92bf661a-9c73-41c0-ab1e-0979609781d1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Auto-Encoding Variational Bayes

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.918152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.918152Z digest=sha256:9577ce5507fa0ece8c897ea807e2f82b33f4d3a3d0309ea05313a6ff15d94409

Observation 70069825-e6b9-4023-bce2-8e464c80c663 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.882276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.882276Z digest=sha256:8f58a99033a5a8f259a21af2af04f3b8d1d0e32a264b69fd20d30ebf9cde191d

Observation e8e6da20-9c3e-400e-8201-e5f354deb0db · outbound

This paper cites POP909: A Pop-song Dataset for Music Arrangement Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model POP909: A Pop-song Dataset for Music Arrangement Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:23.090051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:23.090051Z digest=sha256:b700ca1fb4086ad08d58d7df89f261249dbb124b750a57be719bae317bf9921d

Observation 6fec1bb8-bd79-4231-b913-453677663ad7 · outbound

This paper cites EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model EMOPIA: A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.913422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.913422Z digest=sha256:fe5464d70b1239409e5fbc95f077553e90dfece1257171332f4d6fc6e9faef3a

Observation 90a167dd-addb-4891-94d6-4db58d9af18d · outbound

This paper cites High Fidelity Neural Audio Compression.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.877436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.877436Z digest=sha256:75de08cfab808fb717baa8ea3ab49fccab4775d190cff3b14a348ad483806743

Observation 43418b24-ca96-4809-ab7f-55dae6a259ec · outbound

This paper cites GPT-4 Technical Report.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.839229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.839229Z digest=sha256:1a48744ac031e5f2b4504ac3636e2bfc258fd61b3d0231cfa696691ddb9586d7

Observation 16343409-1236-4d50-895e-c0f0a44e432c · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Diff-Symbo: Text-Controlled Long-Duration Symbolic Music Generation Using Autoregressive Latent Diffusion Model Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T17:54:22.862705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:54:22.862705Z digest=sha256:a740ef348610fc291984ebe04ff0cc62c719ca8a122824f15008a3c5d17f96df

Pith citing papers

No inbound Pith citation observations are available.