Pith. sign in

Paper Citation Record · LEDGER

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.11899.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11899 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:28:27.726082Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact6
  • verified fuzzy27
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0872711-aa1b-48fd-8ba2-2c101cb7a09e · outbound

This paper cites Simple and Controllable Music Generation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Simple and Controllable Music Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.109837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.109837Z digest=sha256:1f73edaa4d84198852e1434e414949c22d57341644c289258de1db714d780523

Observation 66cbd473-cc96-4d28-9b23-decfb6e1de8f · outbound

This paper cites MusicLM: Generating Music From Text.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.177180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.177180Z digest=sha256:7bdf0a7c8e4d56af7d3d1dc464e7dfb06169b7354372f924049611cf9e703d5b

Observation 11c254ee-56bb-40d0-91c6-07e21c5ba28e · outbound

This paper cites ACE-Step: A Step Towards Music Generation Foundation Model.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.284368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.284368Z digest=sha256:32dd1d21c91fe996116e1b6d828adaea058a204ab939df89076c9d7d21a303c7

Observation 71a38278-fbe8-45d8-9349-73208e10f413 · outbound

This paper cites doi:10.48550/arXiv.2602.00744 , urldate =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping doi:10.48550/arXiv.2602.00744 , urldate =

Reference 4

Resolution
verified exact
doi, observed 2026-08-16T00:28:28.718546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.289360Z digest=sha256:2d09f8cfa25d3ff03c30d67480cb815942a7873b6cd387e595998a4bf594403c

Observation b09bc5a3-8721-495a-998c-841276ab5a9f · outbound

This paper cites MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.293852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.293852Z digest=sha256:31048579c8fa07e1e85bbdf4a579de03b589182a1f9dba2e3ec4b0abecada15f

Observation 10ae5514-10db-4268-85d7-dc4947ce3815 · outbound

This paper cites MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.314995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.314995Z digest=sha256:541a90145173b721a0a8142465940b26a1187f03db22370b64c19415960097fe

Observation 3bd89cb3-28a2-405c-bb5c-bc39a5f62d48 · outbound

This paper cites CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.320414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.320414Z digest=sha256:3b1ca1fcbf29a0365b652101847bc4efc2c2bff13f62baf9a64e04ed7186a54e

Observation 6397be86-bc80-45fb-90be-73f64e8da7fc · outbound

This paper cites CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.325260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.325260Z digest=sha256:28ec678e8bf4c76c2fb04cc26f5c694f481e72efab38e62e8f55a746433f813e

Observation 87473a7e-b1ba-4115-a49c-4d34f3dcfb7f · outbound

This paper cites doi:10.48550/arXiv.2503.08638 , urldate =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping doi:10.48550/arXiv.2503.08638 , urldate =

Reference 9

Resolution
verified exact
doi, observed 2026-08-16T00:28:28.546361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.331735Z digest=sha256:63f77722ee3edfb78eff8b1e2bf4f6821492828add565537f4362a56307b0a37

Observation 85a18ad6-5dbb-45b8-8e7b-288842d9022a · outbound

This paper cites doi:10.48550/arXiv.2509.23350 , urldate =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping doi:10.48550/arXiv.2509.23350 , urldate =

Reference 10

Resolution
verified exact
doi, observed 2026-08-16T00:28:28.414010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.336974Z digest=sha256:6c094365ab95532c4d958b9d5451adb04e2a7b0c5d177eb481ffb1c3f5b4bd9e

Observation d4c41290-afdf-46b1-8157-ba3285c4881b · outbound

This paper cites Music ControlNet: Multiple Time-varying Controls for Music Generation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Music ControlNet: Multiple Time-varying Controls for Music Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.427618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.427618Z digest=sha256:959f046594e90568c77655d07c21fac5150dcaa927a51981ee365f3242a0ecfa

Observation fa1a66d0-3cab-4bfa-a741-1b5ac83915cf · outbound

This paper cites doi:10.48550/arXiv.2510.22950 , urldate =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping doi:10.48550/arXiv.2510.22950 , urldate =

Reference 12

Resolution
verified exact
doi, observed 2026-08-16T00:28:28.305501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.520700Z digest=sha256:5d15b006100ea9d4f86b5c4bdc58b505f47fef11aa98e503cdbe399d90787144

Observation 1f31472f-ec84-423c-a039-ce74e65a3297 · outbound

This paper cites doi:10.48550/arXiv.2506.07520 , urldate =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping doi:10.48550/arXiv.2506.07520 , urldate =

Reference 13

Resolution
verified exact
doi, observed 2026-08-16T00:28:28.236350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.585183Z digest=sha256:e02ba6ea3753b0aa4cf939b1b610bb6cba90f7273b325ac7c30dff2d4942dec1

Observation b16d90f8-71d5-4d99-9097-6627734e9f41 · outbound

This paper cites an unresolved cited work.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:28:30.098546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.593641Z digest=sha256:049eacca9279c982ed0842aa5626e1f75a4dddeaeaff73f544f3ac10b564d3aa

Observation 90e40958-e696-4903-91ed-1eab3767a1a6 · outbound

This paper cites Stable Audio Open.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Stable Audio Open

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.599102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.599102Z digest=sha256:d1a76213fe68e4b6b92e6d01222e9ff4a16516636616dacb8cd159931c5c2b21

Observation 46811aac-ca2b-4b6d-8dc6-2de61b0836ac · outbound

This paper cites Stable Audio 3.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Stable Audio 3

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.604045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.604045Z digest=sha256:e267912b64e9e1cdecf8dc87f903affbb8f09271d9aaea9c9c61e1c6958abf82

Observation b301c622-6332-45e6-a3b7-1ad8f5731c5a · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Mustango: Toward Controllable Text-to-Music Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.609238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.609238Z digest=sha256:88a1373187cc5c2a5d1bb2fdf1ec355e33a53f7309d18e056f2ee98904c06f4e

Observation a58f21de-024f-442e-bf12-583c1fc44492 · outbound

This paper cites MuseCoco: Generating Symbolic Music from Text.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MuseCoco: Generating Symbolic Music from Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.614282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.614282Z digest=sha256:7870c092e01b8bb1c3a608414c47741999488be65b2ada85a4db1e0402d1a122

Observation 17d04a20-2b74-4785-8f1b-d196a622d6ba · outbound

This paper cites MuseMorphose: Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MuseMorphose: Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.618547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.618547Z digest=sha256:e56cba683a35e340def5b971debd68685e3365257928de6b7aadd22e92e51ba6

Observation c50d75e1-b465-4543-b85b-428ea73e466e · outbound

This paper cites SongEval: A Benchmark Dataset for Song Aesthetics Evaluation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.622928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.622928Z digest=sha256:44b18ec563f7db8cb8794fb46ef2137b38091175d17579bff2dbca84a2d6a3fa

Observation cffb460f-ea7f-46bd-8e1f-9d125e47fcf4 · outbound

This paper cites FIGARO: Generating Symbolic Music with Fine-Grained Artistic Control.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping FIGARO: Generating Symbolic Music with Fine-Grained Artistic Control

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.626971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.626971Z digest=sha256:32dedee9cfba3747a57ad27a7821e14d874341047ce2a3910d8898c17f798df5

Observation 8c045c9b-6d69-4982-8b47-5795996ca347 · outbound

This paper cites LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:28:28.000730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.658529Z digest=sha256:150421c84ba36d6d48cd39c5be259b8d35f3d8f8a25e55c0ce2646295ea65cc4

Observation 0188d9d8-2da6-4774-8558-1e227682ead6 · outbound

This paper cites Proceedings of the International Computer Music Conference , address =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Proceedings of the International Computer Music Conference , address =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:30.033586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.762197Z digest=sha256:1c7f7b8a7109e87ef90bb4a5bdeabfc78326ccbd984f2bdd44c1bea75bc3f9fa

Observation 7ee19cf0-4b67-44f1-96b9-b453146af944 · outbound

This paper cites Proceedings of the 7th International Conference on Music Information Retrieval , url =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Proceedings of the 7th International Conference on Music Information Retrieval , url =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:30.001789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.787481Z digest=sha256:3ffb8f91717a7193da7928b71f9133f1390a4b6d5d5b582730461e37f2e78f8c

Observation e911ecbe-eaf5-488e-9d50-484c98b5b0a7 · outbound

This paper cites and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.905805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.792385Z digest=sha256:8fcfc78fd648483c2a204e393223ebbd6f99c01557555c9c2c88c3b2dfd774c6

Observation bc17555b-aa7f-4d2d-bdcd-3b265a31f76f · outbound

This paper cites Proceedings of the 16th International Society for Music Information Retrieval Conference , address =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Proceedings of the 16th International Society for Music Information Retrieval Conference , address =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.855632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.983549Z digest=sha256:06aea20883d7033e9139f594ece73aa982a4b6209fe5b31aca8b7d18fc05c115

Observation 31915631-7fa8-4fa6-888b-c64fc3e92324 · outbound

This paper cites Proceedings of the AES 25th International Conference on Metadata for Audio , address =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Proceedings of the AES 25th International Conference on Metadata for Audio , address =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.841855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:26.997457Z digest=sha256:31bf16a2b257f384261a57dc6a14a6d47b85312565e596e7431c87987f363e50

Observation a6de812f-a8e4-40e7-8cee-2416cb5c12cf · outbound

This paper cites 2015 , address =.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping 2015 , address =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.826958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.003800Z digest=sha256:3551995c199e1286e0c3b5f4c20a60932b1b9216403cea4a1fe309ecdf14bd6e

Observation 51f13b55-ef89-4a52-bcac-da853f1150e2 · outbound

This paper cites Beyond Accuracy: Behavioral Testing of.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Beyond Accuracy: Behavioral Testing of

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.812752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.007770Z digest=sha256:46b4c6566edc4e6a04674a2959f80ffe1fa148db4af967989e33c50a8fde643c

Observation e840c0b4-05cf-4db5-9e83-d1e0275d23f4 · outbound

This paper cites Compared to What? Baselines and Metrics for Counterfactual Prompting.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Compared to What? Baselines and Metrics for Counterfactual Prompting

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:27.014667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:27.014667Z digest=sha256:368e9ba7d92221e57f53ee4ee6939b2f8db699ce6ed23978130c614e7143aa2b

Observation 0bd818e5-3bf9-4c68-8449-37400a9bcbbe · outbound

This paper cites an unresolved cited work.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:28:29.752173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.029047Z digest=sha256:773415595e0da09f8a0ee2156cef85cb386cf1174b444165afb4309ce4822c8c

Observation f00ba2f1-5739-4d87-9881-90740a8de78b · outbound

This paper cites Simple and Controllable Music Generation , January 2024.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Simple and Controllable Music Generation , January 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.636996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.033014Z digest=sha256:921bce2a0e9e594f075853918f09e80a4e4434270b81c9ebd1af20eb3935a005

Observation 4a49088f-a280-4a2b-a2a5-360ea31c9166 · outbound

This paper cites Parker, C.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Parker, C

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.543411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.036868Z digest=sha256:10d87e3dd750b135018dcdca92290667a30233043a8c89371fac31d28edd4363

Observation c0da051c-e40c-479f-aaea-9d7f073f4f71 · outbound

This paper cites Parker, Matthew Rice, C.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Parker, Matthew Rice, C

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.530241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.041167Z digest=sha256:af4b11d61b6609cbda27d5554439ab7b0e186fed6379ffe05c90275c0ef068dc

Observation 90534b1d-90a2-4609-8460-4b0fabb0c2a2 · outbound

This paper cites Beat this! Accurate beat tracking without DBN postprocessing.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Beat this! Accurate beat tracking without DBN postprocessing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:27.045584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:27.045584Z digest=sha256:0d06a2552161f4b5b3a31a54f60cabaaa9d2c6c33af7c27ded6287f2c59dc7de

Observation cbe7d1a3-d254-4060-a3f8-9110dd7b0f07 · outbound

This paper cites ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.516594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.049879Z digest=sha256:163cf2a611579d60010016838b0c23f3dad52a5713d7bb0f85a81b9cf8582c39

Observation 087ed2f8-ff0b-42c0-9621-87e777581fd1 · outbound

This paper cites S-KEY: Self-supervised Learning of Major and Minor Keys from Audio.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping S-KEY: Self-supervised Learning of Major and Minor Keys from Audio

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:28:27.873320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.054010Z digest=sha256:e9da40c2cc44c61b8effa68bcfea0077854c3b569194bfd10ea36552b0d3e2d0

Observation b2498062-495e-42dc-b035-261bd7866921 · outbound

This paper cites MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.501593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.057995Z digest=sha256:a940121755d99a44e61e9129fafcc9df43b1a1bdb0bcac45c5a18174f1af0593

Observation 373bf2f6-651c-477c-a17e-1b3e026d191b · outbound

This paper cites LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.397572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.062323Z digest=sha256:0587cec0a5552536977e44ca43426f33062218a1ccce70b4a175bdf7e71ccae4

Observation 193efec9-9939-4bbd-a7bd-1d156b5b252f · outbound

This paper cites MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.251069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.111878Z digest=sha256:eb184f4e5d1023610fba461ddf6ce3c8fbf46f5abac8334bc35589045acbb82d

Observation 041a9299-b036-4efc-9332-45bac6ca7dc1 · outbound

This paper cites MuseCoco : Generating Symbolic Music from Text , May 2023.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MuseCoco : Generating Symbolic Music from Text , May 2023

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.144617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.250537Z digest=sha256:95615e8b2305becd5d6097c3448046241ec9cde2620354dd1c723c8c70aab408

Observation 38fb4f0e-b9c7-44d1-8658-287f86c21660 · outbound

This paper cites CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.131106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.354980Z digest=sha256:2a903832c482b399b1f3693032434582eced995d416bbd7ce58c483c4ee4782c

Observation bc7d5d78-4b85-4108-b016-76c2555faf66 · outbound

This paper cites GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.117742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.367529Z digest=sha256:07dbb88c7724becdd897d918fb2969c597921861da287257d7fd237bf19273ae

Observation bae363d7-b273-4d8c-a480-5cae0820101e · outbound

This paper cites Mustango: Toward Controllable Text-to-Music Generation , June 2024.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Mustango: Toward Controllable Text-to-Music Generation , June 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.105018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.371944Z digest=sha256:a90d737fab88d6e4375c03d1d424ee242709a03fbc7a4b49880dbce3df9b58ff

Observation 158ddff2-d311-40c5-accb-5fe599f06154 · outbound

This paper cites Genre-specific key profiles.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Genre-specific key profiles

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.091795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.375932Z digest=sha256:d590721c0d4b8ea21279c4b6bae4e8406a8bbd41f4139be403adea487aea938f

Observation 388dd3f3-e260-40c1-9815-c7e6691c98f6 · outbound

This paper cites Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.076735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.381075Z digest=sha256:8ccfa6651306da9bbcdcfcc5494adc670066f0edf1dec9ba01f6d4cf43c586da

Observation 594cbda2-64ab-42f4-860a-1ac7983ea9f1 · outbound

This paper cites Beyond accuracy: Behavioral testing of NLP models with CheckList.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Beyond accuracy: Behavioral testing of NLP models with CheckList

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.062481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.385941Z digest=sha256:d6a33b27b8360a2676b4cb550e90fb6c12a5cf38bc7b214e0b78a106a2aa662c

Observation 31ad490c-df38-41ed-b1bf-e6ca8cebe65f · outbound

This paper cites The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:27.391018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:27.391018Z digest=sha256:1beb9dd68e8ba5a7b2cde8ce66e077db2d44190c8820a6bc6933dcee8fd26b2b

Observation 32059aa9-efa1-424d-9801-40201cd7a5a9 · outbound

This paper cites FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.048500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.395458Z digest=sha256:a82ea69ecffac62ab1d65c7f25859ca347e4d7f85a1829ae40debe45c6414da0

Observation db0ac36a-4f63-45cb-add7-7252d4743935 · outbound

This paper cites MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:29.033954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.399811Z digest=sha256:058633d082039aa9e83dcf8e16edd64bf16dbae99d1442231436deab172d9a82

Observation c8b97a66-e4bc-4ce3-b12b-eb5c19e5308f · outbound

This paper cites an unresolved cited work.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:28:28.922944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.404349Z digest=sha256:13630164b88de6019f4dd0177da0028ca623b07b2ac0fd26cb905026bcc1a6cb

Observation 6ff3f0ad-f5d8-4b08-9451-3405a5246e9c · outbound

This paper cites an unresolved cited work.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:28:28.867279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.409664Z digest=sha256:4dae8f1f25a1e16669e041f66442e79075df1e12faf57bca1cf48c4a0cb05650

Observation aa3307b1-d2b0-4a0d-8a6b-2d4a083e39b6 · outbound

This paper cites SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:28.833294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.462374Z digest=sha256:8d3d3e11a40edbf839f86ebf9980be4efad232b37ad770a61a03693ee71a29ab

Observation f8da1bb1-e90b-4a19-8230-023b91214fd1 · outbound

This paper cites YuE : Scaling Open Foundation Models for Long-Form Music Generation , September 2025.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping YuE : Scaling Open Foundation Models for Long-Form Music Generation , September 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:28.819312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.539810Z digest=sha256:8d1b0adc5aa589ab15962eb719be61fd2ceaebbc86008422f88be4cf87883ed7

Observation ba977370-faf3-420e-a08c-d37c6d8995f1 · outbound

This paper cites ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:28.804967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.656813Z digest=sha256:6da986144ef4a6928771decf8e7a2018b05a9ba895aba857f62bc4bddfdf4956

Observation 48b8c151-d3ac-4ca6-93c3-5196891bb1fb · outbound

This paper cites TPSMG : Text-Controllable Polyphonic Symbolic Music Generation.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping TPSMG : Text-Controllable Polyphonic Symbolic Music Generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:28:28.788155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-16T00:28:27.726082Z digest=sha256:689a87059a656bd8dbce4481d2c6bca16388a62366c460d6dfdb97c0b760999c

Pith citing papers

No inbound Pith citation observations are available.