Pith. sign in

Paper Citation Record · LEDGER

Exploring State-Space-Model based Language Model in Music Generation

As of 17 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.06674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06674 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.400382Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.319848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:02:24.469579Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52ca52a8-ca71-4afe-bf4a-16e4c5b02055 · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.654861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.315970Z digest=sha256:cc963c50e16460001be1b399f57c7255a6a8247ebcd4cfe35cb6d6ba170ced17

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:e9ec8641d9281a7414f6ffc7e5424e733a522629fa38cf5a8b8a26d48ec13dcb

Observation 933e7d64-8e41-4ebb-b62c-3c85af3ace30 · outbound

This paper cites We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3].

Exploring State-Space-Model based Language Model in Music Generation We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.646549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.323872Z digest=sha256:28a9e92ba27bd612b16e8e6b88e589a7f9660702a08d9b2484149942ce21827a

Observation 92e4a2a2-8d09-43d2-842a-bfbf43b85354 · outbound

This paper cites Both FAD and KLD are lower the better, while CLAP is higher the better.

Exploring State-Space-Model based Language Model in Music Generation Both FAD and KLD are lower the better, while CLAP is higher the better

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.638280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.327623Z digest=sha256:5a27dfabc352a4d584f64d9f52dc09b0ddfb54ae6fc26a010d3d3283b8cc5662

Observation 67e548ea-689c-482e-8389-041b800fafdb · outbound

This paper cites Decoupled weight decay regularization,.

Exploring State-Space-Model based Language Model in Music Generation Decoupled weight decay regularization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.629458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.330769Z digest=sha256:69b9d142cb9482c1b6438f98fa564f652a7d1ecc945fbd48d92600f8e1776ab9

Observation c631407a-b575-4016-9edd-38cd09ccdec3 · outbound

This paper cites The MTG-Jamendo Dataset for Automatic Music Tagging,.

Exploring State-Space-Model based Language Model in Music Generation The MTG-Jamendo Dataset for Automatic Music Tagging,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.620895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.334460Z digest=sha256:10977eeb5d6f05e0dcf9b657bf5a54b68d5061782e66b30adc94a755a07bbbaf

Observation 00e27017-3956-4034-9a79-9e6083dc83e5 · outbound

This paper cites Hybrid Trans- formers for music source separation,.

Exploring State-Space-Model based Language Model in Music Generation Hybrid Trans- formers for music source separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.612619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.337857Z digest=sha256:f9311600b72ef9fe0cd4368b9f37b1c38dac54433c8389afe468ff4602cf591a

Observation 11f9d51f-5607-4a3d-842a-24024fa52584 · outbound

This paper cites LP-MusicCaps: LLM-based pseudo music captioning,.

Exploring State-Space-Model based Language Model in Music Generation LP-MusicCaps: LLM-based pseudo music captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.603686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.340676Z digest=sha256:384fa36d9b120474c179587bd4d20637825e4369cc5b477485a3a1b4367bc719

Observation 598b522d-7970-4ceb-9cde-8c1f1bf82d13 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring State-Space-Model based Language Model in Music Generation The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.343931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.343931Z digest=sha256:1b32d3665711946473a2bd93ec8724d0e50daebef96ca775f512b2536ecc771d

Observation b85630d5-8619-4ffd-9772-07fe6ea42ae6 · outbound

This paper cites The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,.

Exploring State-Space-Model based Language Model in Music Generation The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.594043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.347731Z digest=sha256:2972d953e2c390117eb7086f18bb87aa80a2ef0895417581c06c48a0681589dc

Observation 000d75ce-3637-4739-b292-b75a829d140c · outbound

This paper cites Scaling instruction-finetuned language mod- els,.

Exploring State-Space-Model based Language Model in Music Generation Scaling instruction-finetuned language mod- els,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.585432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.351332Z digest=sha256:69b0a6c840cfb85f2c5f0d08a27d3041b8b493baae003132f1f0994cea5bb501

Observation 7277841f-9344-40b2-a704-6fb3e73b614d · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Exploring State-Space-Model based Language Model in Music Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.354155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.354155Z digest=sha256:7cd08aaf9823a515357b54b6c02bc3932576bfbb8bb2e965e53234f5c8dbc87b

Observation 30aec164-c4aa-49f2-b7c3-cb9b7ef8d4fc · outbound

This paper cites Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,.

Exploring State-Space-Model based Language Model in Music Generation Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.576707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.357348Z digest=sha256:52bf4a7c0560a44e7f0fe321debbeb48329e62798f0b47baef042a21667be58c

Observation 2e692a93-aa63-4c43-8af3-e2c6e037539c · outbound

This paper cites SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series.

Exploring State-Space-Model based Language Model in Music Generation SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.360110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.360110Z digest=sha256:993f16f1f13cc08682a17a1d823d907659de864f4fc4a7a1630dd62296d42a92

Observation 8fb1814f-e275-41fa-a96a-c12fce8d4b72 · outbound

This paper cites High-fidelity audio compression with im- proved RVQGAN,.

Exploring State-Space-Model based Language Model in Music Generation High-fidelity audio compression with im- proved RVQGAN,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.568184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.363148Z digest=sha256:85fa81fa32ec8daed585eb78df29dd4127e19f7671b094ac7ea75c9a98d8a168

Observation 756330a9-bde3-40ae-bc9f-83851b512d1a · outbound

This paper cites Coarse-to-fine text-to- music latent diffusion,.

Exploring State-Space-Model based Language Model in Music Generation Coarse-to-fine text-to- music latent diffusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.559894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.365909Z digest=sha256:5863da5e059426461c8ab2f3eb35d41a258defb1f2b5cbf6e94652668614f20d

Observation f1861585-e6c7-4b67-8a64-3bd05e108397 · outbound

This paper cites MusicLM: Generating Music From Text.

Exploring State-Space-Model based Language Model in Music Generation MusicLM: Generating Music From Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.368575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.368575Z digest=sha256:8e101a77ca61fbcfa518facd130d52cf172798402363f14c14ad3fc07522a7db

Observation 712c6245-8d1c-458f-a56a-62f7ad753d21 · outbound

This paper cites Simple and control- lable music generation,.

Exploring State-Space-Model based Language Model in Music Generation Simple and control- lable music generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.551716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.371630Z digest=sha256:c5b9894b53c7eee092b95f93df71a02f41883c556ed00f9cc8fb2ad36dcf114f

Observation 4c80ae5d-f2e2-4948-9501-035de94da30d · outbound

This paper cites Stable audio open,.

Exploring State-Space-Model based Language Model in Music Generation Stable audio open,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.543061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.374380Z digest=sha256:5a71c22c7b5693a38ff05ee433db0a810945966fae6596492565123cb882d68b

Observation eb504c78-366e-48ec-9b61-4c94020b752b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,.

Exploring State-Space-Model based Language Model in Music Generation AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.533682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.377638Z digest=sha256:4b56d558050ef28e751066de50c528579ec7f61439755d797dbc98a60d725f51

Observation 6052973c-9722-4fbe-9ef5-59f9fd2b497e · outbound

This paper cites Mustango: Toward controllable text-to-music generation,.

Exploring State-Space-Model based Language Model in Music Generation Mustango: Toward controllable text-to-music generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.523416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.380443Z digest=sha256:32fdfdd89cbf81881e87d682905bfba4940514ad0b80d095f414fac67d3337e8

Observation 18188d7c-b941-47ea-abf2-a948124e77d0 · outbound

This paper cites JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,.

Exploring State-Space-Model based Language Model in Music Generation JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.514119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.384619Z digest=sha256:d6827736a51a1e5fb6f6de4df28b929d93b24e09ade53ac72954d13c874d9d2f

Observation 981006d2-d1a6-4647-8332-b80cd2f36eda · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Exploring State-Space-Model based Language Model in Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.387244Z digest=sha256:85577c7cfb22b4f8b1b41a69e03fda5ee8d604611010960424b5bb180032a3cc

Observation 826ee30f-3694-43f6-8d8d-1672cfda6509 · outbound

This paper cites Music ControlNet: Multiple time-varying controls for music generation,.

Exploring State-Space-Model based Language Model in Music Generation Music ControlNet: Multiple time-varying controls for music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.503756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.390173Z digest=sha256:1bd30fa579a95916af6d67eb18c9638b80097e149b23fd821999a59e2b479a6b

Observation 24806345-b78f-48b7-bf08-ab5630f53a63 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Exploring State-Space-Model based Language Model in Music Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.393797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.393797Z digest=sha256:d8e6deb633a602e1e71d250028015d3305152a8fa762b9284f560202c5958378

Observation 3892877c-48bd-44d7-9e39-bd73ceda8384 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

Exploring State-Space-Model based Language Model in Music Generation CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.493484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.396775Z digest=sha256:25980f8bdccb189ab970149757d12911da9671426fb55f91784a2d2ae8f33dad

Observation 90bba355-2469-4f7f-b1ed-0d607a9d9f5e · outbound

This paper cites MuseControlLite: Multifunctional music generation with lightweight conditioners,.

Exploring State-Space-Model based Language Model in Music Generation MuseControlLite: Multifunctional music generation with lightweight conditioners,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.484848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.400382Z digest=sha256:dadd45f045e697f645a9a3da01f9f1b971c95b4f6a7e47b68fef1d80c055bd40

Pith citing papers

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · inbound

Exploring State-Space-Model based Language Model in Music Generation cites this paper.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:e9ec8641d9281a7414f6ffc7e5424e733a522629fa38cf5a8b8a26d48ec13dcb