Pith. sign in

Paper Citation Record · LEDGER

Exploring State-Space-Model based Language Model in Music Generation

As of 9 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2507.06674.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06674 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.400382Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:02:24.319848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T19:02:24.469579Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy20
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52ca52a8-ca71-4afe-bf4a-16e4c5b02055 · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.654861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.315970Z digest=sha256:3b0dcf58ac8b2d08fd95a2790253d9a02a1125e028ad5626c93b02fa304c61f7

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · outbound

This paper cites Exploring State-Space-Model based Language Model in Music Generation.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:a343786fa07306396791087b12839bff04e486201107a73e38c45c1e66e8ec75

Observation 933e7d64-8e41-4ebb-b62c-3c85af3ace30 · outbound

This paper cites We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3].

Exploring State-Space-Model based Language Model in Music Generation We re-sample all the audio into 44.1kHz and convert them into mono audio, splitting the tracks into non-overlapping 30s clips with vo- cals removed by HTDemucs [3]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.646549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.323872Z digest=sha256:b441302ecd3f822ddb116d0bc8c329f6e64e13c61331943e4dfdd128e5fdcf0b

Observation 92e4a2a2-8d09-43d2-842a-bfbf43b85354 · outbound

This paper cites Both FAD and KLD are lower the better, while CLAP is higher the better.

Exploring State-Space-Model based Language Model in Music Generation Both FAD and KLD are lower the better, while CLAP is higher the better

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.638280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.327623Z digest=sha256:9b92fa78061ffc0dfb162a43831a9d21497d8e32da9ceb57c33eeaefb32a1225

Observation 67e548ea-689c-482e-8389-041b800fafdb · outbound

This paper cites Decoupled weight decay regularization,.

Exploring State-Space-Model based Language Model in Music Generation Decoupled weight decay regularization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.629458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.330769Z digest=sha256:efcebfe04fdb52a9a78560c2743e20f87a22d4804cc9860db3b02173d18b0bc4

Observation c631407a-b575-4016-9edd-38cd09ccdec3 · outbound

This paper cites The MTG-Jamendo Dataset for Automatic Music Tagging,.

Exploring State-Space-Model based Language Model in Music Generation The MTG-Jamendo Dataset for Automatic Music Tagging,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.620895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.334460Z digest=sha256:b8c1e48d565f7a4169cd2211f380faf7f59f60c4dda39c081696d6c2c17e7413

Observation 00e27017-3956-4034-9a79-9e6083dc83e5 · outbound

This paper cites Hybrid Trans- formers for music source separation,.

Exploring State-Space-Model based Language Model in Music Generation Hybrid Trans- formers for music source separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.612619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.337857Z digest=sha256:d16464c8b88d7ddf4acc5b949b4531475b3e4e480f529da1e3102a3018f81f24

Observation 11f9d51f-5607-4a3d-842a-24024fa52584 · outbound

This paper cites LP-MusicCaps: LLM-based pseudo music captioning,.

Exploring State-Space-Model based Language Model in Music Generation LP-MusicCaps: LLM-based pseudo music captioning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.603686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.340676Z digest=sha256:146529e342fed1fe0ac0adb180c2b0ee084402af64051ea3f7bcd4d49720012d

Observation 598b522d-7970-4ceb-9cde-8c1f1bf82d13 · outbound

This paper cites The Llama 3 Herd of Models.

Exploring State-Space-Model based Language Model in Music Generation The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.343931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.343931Z digest=sha256:82fd8c0285e1b240399745a683f66deb0733c872c3249f1224cd59a097c9487d

Observation b85630d5-8619-4ffd-9772-07fe6ea42ae6 · outbound

This paper cites The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,.

Exploring State-Space-Model based Language Model in Music Generation The Song De- scriber Dataset: a corpus of audio captions for music- and-language evaluation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.594043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.347731Z digest=sha256:2d177c3a39c3b159183d4da3ccc22fc404b89caf71fa097be75092a885010099

Observation 000d75ce-3637-4739-b292-b75a829d140c · outbound

This paper cites Scaling instruction-finetuned language mod- els,.

Exploring State-Space-Model based Language Model in Music Generation Scaling instruction-finetuned language mod- els,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.585432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.351332Z digest=sha256:b69a0f79f5deb34ae5b30c1af2a91b4cf6fa2275dc9e885c009866f60d74e3c8

Observation 7277841f-9344-40b2-a704-6fb3e73b614d · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Exploring State-Space-Model based Language Model in Music Generation Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.354155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.354155Z digest=sha256:9b0f55362855df82151076bfe3d3117929da334dae61585cedd7ceb4c59a184c

Observation 30aec164-c4aa-49f2-b7c3-cb9b7ef8d4fc · outbound

This paper cites Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,.

Exploring State-Space-Model based Language Model in Music Generation Transformers are SSMs: General- ized models and efficient algorithms through structured state space duality,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.576707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.357348Z digest=sha256:fd51a055e1515c189f187bd1e27cda1c3545b685acb9123e930be2b28049137d

Observation 2e692a93-aa63-4c43-8af3-e2c6e037539c · outbound

This paper cites SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series.

Exploring State-Space-Model based Language Model in Music Generation SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.360110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.360110Z digest=sha256:7728bd72ae73d71a5b318c9d7446bbe5dbab578692563fbdca634d99f7763287

Observation 8fb1814f-e275-41fa-a96a-c12fce8d4b72 · outbound

This paper cites High-fidelity audio compression with im- proved RVQGAN,.

Exploring State-Space-Model based Language Model in Music Generation High-fidelity audio compression with im- proved RVQGAN,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.568184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.363148Z digest=sha256:97460f730a51b5161854ab3e62a8bd7f087d23cb8b3fce462092e27e70d51405

Observation 756330a9-bde3-40ae-bc9f-83851b512d1a · outbound

This paper cites Coarse-to-fine text-to- music latent diffusion,.

Exploring State-Space-Model based Language Model in Music Generation Coarse-to-fine text-to- music latent diffusion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.559894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.365909Z digest=sha256:8101c3652097046f19feb601a545ce67dc5e07346782a1385c6319a193b082ca

Observation f1861585-e6c7-4b67-8a64-3bd05e108397 · outbound

This paper cites MusicLM: Generating Music From Text.

Exploring State-Space-Model based Language Model in Music Generation MusicLM: Generating Music From Text

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.368575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.368575Z digest=sha256:a48a224ee00fdc2ba86973758f3ff9328b622aa617de57305b4002c0e5ba36df

Observation 712c6245-8d1c-458f-a56a-62f7ad753d21 · outbound

This paper cites Simple and control- lable music generation,.

Exploring State-Space-Model based Language Model in Music Generation Simple and control- lable music generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.551716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.371630Z digest=sha256:3aa51c3151bd0979ecb9f13de2d9429eed9802474d4d283a7d1dce08df965cca

Observation 4c80ae5d-f2e2-4948-9501-035de94da30d · outbound

This paper cites Stable audio open,.

Exploring State-Space-Model based Language Model in Music Generation Stable audio open,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.543061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.374380Z digest=sha256:47f763c44b814470aa4a541dbbb1049353a12d548254bc4a4ebc7afad2595c99

Observation eb504c78-366e-48ec-9b61-4c94020b752b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,.

Exploring State-Space-Model based Language Model in Music Generation AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.533682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.377638Z digest=sha256:f6fea53f92746e633011fa9bac5cc43ec40f442f283bae8f13c61a63d08b50c1

Observation 6052973c-9722-4fbe-9ef5-59f9fd2b497e · outbound

This paper cites Mustango: Toward controllable text-to-music generation,.

Exploring State-Space-Model based Language Model in Music Generation Mustango: Toward controllable text-to-music generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.523416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.380443Z digest=sha256:d43ef9592a7b2a79654e7f08b79a9b8afc33f605b168bda6addda903dabf763f

Observation 18188d7c-b941-47ea-abf2-a948124e77d0 · outbound

This paper cites JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,.

Exploring State-Space-Model based Language Model in Music Generation JEN-1: Text-guided universal music gener- ation with omnidirectional diffusion models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.514119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.384619Z digest=sha256:1370384adff6214105be60e8f668b74b17a1bbdf08b3aeed0b16d352951b4adc

Observation 981006d2-d1a6-4647-8332-b80cd2f36eda · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Exploring State-Space-Model based Language Model in Music Generation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.387244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.387244Z digest=sha256:b90e5f1cc94482da7905231e9d8bbc019f9ced0a3ebd93c0e6899048c9bc5c28

Observation 826ee30f-3694-43f6-8d8d-1672cfda6509 · outbound

This paper cites Music ControlNet: Multiple time-varying controls for music generation,.

Exploring State-Space-Model based Language Model in Music Generation Music ControlNet: Multiple time-varying controls for music generation,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.503756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.390173Z digest=sha256:43debb55364a44f559e2e92054ebfd53cfc4d0ab1d347a43dee38b89ed2956e0

Observation 24806345-b78f-48b7-bf08-ab5630f53a63 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Exploring State-Space-Model based Language Model in Music Generation Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:02:24.393797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:02:24.393797Z digest=sha256:ae3bd7c3c3250f6be05f41b7198c664e74d5d32547333c109ca0ee104e8db63a

Observation 3892877c-48bd-44d7-9e39-bd73ceda8384 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

Exploring State-Space-Model based Language Model in Music Generation CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.493484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.396775Z digest=sha256:ec2f42a9f5cbceee6765cce78a52ae328c4efdeecadf9d66f8c76b1d0167c74d

Observation 90bba355-2469-4f7f-b1ed-0d607a9d9f5e · outbound

This paper cites MuseControlLite: Multifunctional music generation with lightweight conditioners,.

Exploring State-Space-Model based Language Model in Music Generation MuseControlLite: Multifunctional music generation with lightweight conditioners,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:02:24.484848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.400382Z digest=sha256:e598d5eeb99ec7f8cb1500c551dbe26c45d0aac1d1c38010b0690f2c7003d2d5

Pith citing papers

Observation 00fb6b57-fae5-4d8d-8ec0-16a53675b20c · inbound

Exploring State-Space-Model based Language Model in Music Generation cites this paper.

Exploring State-Space-Model based Language Model in Music Generation Exploring State-Space-Model based Language Model in Music Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:02:24.474768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:02:24.319848Z digest=sha256:a343786fa07306396791087b12839bff04e486201107a73e38c45c1e66e8ec75