Pith. sign in

Paper Citation Record · LEDGER

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2504.21815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21815 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.433651Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:57:47.249767Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T04:57:47.830028Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4cf1924a-2025-401e-a7b6-1796ad82b4e8 · outbound

This paper cites an unresolved cited work.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:57:48.215884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.244102Z digest=sha256:a6263d44f31baa7560a1548ab28d6e329ab23363a92d0ad00f28898632cdb112

Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · outbound

This paper cites From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.834783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.249767Z digest=sha256:41673f05ba43ca8b53fdac2b20abdbd9b2f80bae8cab3fbd680df96886be9eed

Observation a677cbcd-5180-4be5-a9bd-baa87ef75644 · outbound

This paper cites an unresolved cited work.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:57:48.202200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.255187Z digest=sha256:495ae17e97326042b9c5f4d70f4ff358973f0fac031dbfa930a5adaf654f928a

Observation cc880505-c91d-4713-ba40-49d21a8370b6 · outbound

This paper cites low quality.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems low quality

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.187932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.259974Z digest=sha256:33ff8d68707b78e984b35e28fbbd661bf87804d65d35db209248a16a90366246

Observation f5756c59-4e1f-441c-938f-2c2b27cdce4e · outbound

This paper cites We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems We examine reference-based evaluation metrics on the generated dataset computed in Section 4, offering a complementary perspective to subjective assessments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.174473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.265373Z digest=sha256:afb52d9a792c42c054eb7946aae9544777f007e20b2f2432d33954c6262370c0

Observation 748cb1b8-505f-46c1-8363-626873488520 · outbound

This paper cites Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Our results reveal substantial incon- sistencies between different evaluation perspectives, highlighting the challenges of fully capturing human judgment through automated proxies

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.160227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.270241Z digest=sha256:2214e103f4d5d097d220aaff4fc0f3f12d04a79e204c803ba6f11dfd295a985b

Observation 4cad6a93-1e79-487a-961c-8255a4a1b058 · outbound

This paper cites Acoustic scene generation with condi- tional SampleRNN,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Acoustic scene generation with condi- tional SampleRNN,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.146152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.274640Z digest=sha256:ad8a62677cf588c0152292cc8a1b4ac7c3a69fc0e0ed8fbb8614430c5a630ce8

Observation 1ec9efc9-90c6-45da-99da-6d7fb5062a42 · outbound

This paper cites AudioGen: Textually guided audio genera- tion,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioGen: Textually guided audio genera- tion,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.132578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.279168Z digest=sha256:338ce4ea877ae929d70e66c23cbe905c488d0aab519dd9fb1b4e255d337133ca

Observation 263790c9-df72-4cc2-b01e-e74fa0af038f · outbound

This paper cites DExter: Learning and Controlling Performance Expression with Diffusion Models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DExter: Learning and Controlling Performance Expression with Diffusion Models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.118563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.283916Z digest=sha256:faf72497a55c1d046e7e747d5ed6a041a35da2cefd7a9ad3c1249c2ccb1831d5

Observation e15067ec-328f-4ba1-a97f-963ca8c703f6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.288199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.288199Z digest=sha256:de56f454a1b247a75bebfe1f73a7cf62c832b829555b9d48d71af49e529340ae

Observation dea05df3-4aa9-463e-90b1-868cb15c4b7f · outbound

This paper cites Qwen2.5 Technical Report.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Qwen2.5 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.293130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.293130Z digest=sha256:d886714609fbb04e3a367e7ee1da8daa89cf85f2a47224aff201bec0916ad502

Observation 21984b2c-a8a6-4ea1-b1a0-66d8afc92516 · outbound

This paper cites DSPO: Direct score preference optimization for diffusion model align- ment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DSPO: Direct score preference optimization for diffusion model align- ment,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.104139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.298409Z digest=sha256:9d523749726db233a20fa9dd0251f78b2ec72647dd4a8cfe364d7270ee9d8f44

Observation 0a4454cb-d7b0-496d-835e-1e617e964cd5 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.302806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.302806Z digest=sha256:d01959597199c0f91600e418275336aa4dceb4ecb67c18774fd839c166559834

Observation 97b67597-354f-4ca0-a053-4efa3937514b · outbound

This paper cites 1, Association for Computing Machinery, 2024.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems 1, Association for Computing Machinery, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.090417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.307649Z digest=sha256:293058ec79e6feed73410a88f5aa5f04be8d44e71153af7de56d293b0ae2d776

Observation 04fec364-3565-464e-8d24-9d25be516a88 · outbound

This paper cites BATON: Aligning text-to-audio model using human preference feedback,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems BATON: Aligning text-to-audio model using human preference feedback,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.076969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.311697Z digest=sha256:412872106436bf64777ee8eba8d1fc7842215261abb6eb9cecdcbec13add4010

Observation 270de9c1-f454-41b2-9ba2-b54894bab1ae · outbound

This paper cites DRAGON: Distributional rewards optimize diffusion generative models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DRAGON: Distributional rewards optimize diffusion generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.315838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.315838Z digest=sha256:9f4e0caad18458934162e47b4e397142c9c3799f04cb11a7118038000d1db85f

Observation bf4d53e4-f123-46d0-83fa-e685f3268f87 · outbound

This paper cites SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.709388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.320069Z digest=sha256:b95ad0767ae4af2ecada647e1633b49e623b92784f016cbcca56487b15d1d485

Observation 9858a7fc-3339-4148-9bae-3beae96a5fad · outbound

This paper cites Aligning Text-to-Music Evaluation with Human Preferences.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Aligning Text-to-Music Evaluation with Human Preferences

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.324607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.324607Z digest=sha256:affe6f8a800eb9f24e78db7773f85f3121ff9054312333be7a87d17fbdf0d104

Observation a78cb8af-1d39-436a-8cd2-25fcb146a866 · outbound

This paper cites KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems KAD: No More FAD! An Effective and Efficient Evaluation Metric for Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.329053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.329053Z digest=sha256:6945ae4dcc3c1955535783a7fcb30bd1dc033df76bea9a459735aae9fc5d3613

Observation 3b40369f-1a92-466d-8fc3-8383255a9adf · outbound

This paper cites WavCraft: Audio editing and generation with large language models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavCraft: Audio editing and generation with large language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.063073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.333481Z digest=sha256:d1e80e2acf16d7e6876d6a25abc92cc03e57a34bd09ee9b07003c02c310a8042

Observation 6a11a3b2-b24f-4853-87a4-2747096b6ca4 · outbound

This paper cites WavJourney: Compositional audio creation with large language models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems WavJourney: Compositional audio creation with large language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.048898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.337921Z digest=sha256:86e518ffe40e48cbbe6f10e02f7bd48dfc0af1573053e7915006c483f62af7b9

Observation 1216f0f5-e14e-4400-9816-8a87becf1229 · outbound

This paper cites Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.342303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.342303Z digest=sha256:06baed072394413aaf152f0f4055281db4434a61d0224b9ac892582f4eeb0e25

Observation fe47a6b5-73ee-402c-b1a6-0c9ebede57bb · outbound

This paper cites Hierarchical Symbolic Pop Music Generation with Graph Neural Networks.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Hierarchical Symbolic Pop Music Generation with Graph Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.346856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.346856Z digest=sha256:8d646e68d6fb62ebeebb3f56e68cff25311e033f4a9ebcf7bfe5a98c48d618c2

Observation 2943ddd5-9d71-441b-bef7-e8db6ebeeb4b · outbound

This paper cites RenderBox: Expressive Performance Rendering with Text Control.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems RenderBox: Expressive Performance Rendering with Text Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.351237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.351237Z digest=sha256:42b099359d0849da36d39c7b0199b533f485308318d7685421e5b76be7521bae

Observation 1f301de7-9b12-4a39-abe0-7b2b434be3fc · outbound

This paper cites Leveraging pre-trained audioldm for sound generation: A benchmark study,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Leveraging pre-trained audioldm for sound generation: A benchmark study,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.034883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.355679Z digest=sha256:ce0a250dc19e37e3574f3c10fee10223d13b9d2ec4e448f845870ff152237fc0

Observation 4d29258e-06ab-4c71-87a2-4e204c684abf · outbound

This paper cites Diffsound: Discrete diffu- sion model for text-to-sound generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Diffsound: Discrete diffu- sion model for text-to-sound generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.021286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.360127Z digest=sha256:96ceec49b89b7571ef304cca95814052186ebc5c5ae3697aad9a5e5cabe337a0

Observation 6b059db3-0c0e-465d-9a3b-635fc6f088b5 · outbound

This paper cites Adapting frechet audio distance for generative music evaluation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Adapting frechet audio distance for generative music evaluation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:48.007803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.364505Z digest=sha256:1d24c4c7e5ed26ce8508d54fd94f370fc90ce989c8d64f5159f7dbfa663b8cd8

Observation 55d0523a-c8ad-4a58-8426-2b6794426a8a · outbound

This paper cites Zero-shot unsupervised and text-based audio editing using DDPM inversion,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Zero-shot unsupervised and text-based audio editing using DDPM inversion,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.993797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.368632Z digest=sha256:c2a2f5ef0714208d3ae715f22f305dc646d221119ac7b1bf4058a9a995be4090

Observation 09b81ee9-7b6f-4e12-9723-5b6787abab2b · outbound

This paper cites AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems AudioMorphix: Training-free audio editing with diffu- sion probabilistic models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.980370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.372707Z digest=sha256:8d00254f2b1b7dd4af77581d474da68e9160e0c9a97b025d7414f3751dfd2d49

Observation ad3c2866-b980-487e-ab89-467c7df63f34 · outbound

This paper cites A comparison of deep learning MOS pre- dictors for speech synthesis quality,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems A comparison of deep learning MOS pre- dictors for speech synthesis quality,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.965893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.377051Z digest=sha256:76d871f11083d9f30da9defaea609a982db0b8e84df1fc36a3b257edefa3955a

Observation 1558adc6-5a1b-4109-83b4-9dde11ff723e · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.381234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.381234Z digest=sha256:a096a2d75a8e7634c16f3c55a2c3654e9e807131696c1ae380f623bf4a2fc204

Observation e846aac5-a3a4-4cd6-8382-3a1863dbbf31 · outbound

This paper cites From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.952196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.385669Z digest=sha256:689bd5becc2347d0ec378b95232364e532e526169ec6e99fadb7858b714dc28e

Observation dda5b3b9-e940-4e24-916c-524783aad0f3 · outbound

This paper cites Piano Skills Assessment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Piano Skills Assessment,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.937513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.389788Z digest=sha256:72fa731c8b4454e6f8ea0c3018f620663582f213e041df9c2e97afb4e2b590a2

Observation 998adf1f-4074-453f-bba7-18b6777c6225 · outbound

This paper cites LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.923524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.394180Z digest=sha256:1618cf04d84da995e08aaeeb4328a0e4c3809d677c04e35186b06ee18c64b6d8

Observation 8e010328-1240-4842-8bea-849015f58b41 · outbound

This paper cites MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.398280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.398280Z digest=sha256:a87deb7df43083075149f8bbf21f2642ae1e45adf2a7ed0a10525c3b2e516e3d

Observation c6744b11-d13b-4c25-896e-f467dc0231f9 · outbound

This paper cites How does the teacher rate? Observations from the NeuroPiano dataset,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems How does the teacher rate? Observations from the NeuroPiano dataset,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.908177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.402747Z digest=sha256:1f3727b30a9fdccefbff2c3fe0dc3eebb53ead85dddeaf53bb5f91ed3c3a59c1

Observation 3f4fcfd6-91c2-40c6-9266-2dc96d8a1da0 · outbound

This paper cites Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.893180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.407376Z digest=sha256:cc2d9cf314c7e821c9db04d15291b31b18a7f037ff4af6fbaa015b039116a1ee

Observation 84a339d1-68ee-441f-86ed-c914ee47a8ff · outbound

This paper cites Stable Audio Open.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Stable Audio Open

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.411635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.411635Z digest=sha256:2c20ff4c5ea06c1a4ae26fbcf5e97e1ad342938fd15e119dc4a4e1b9304dfe17

Observation 19aed1bd-d670-4ec6-b619-7e7231e7a0aa · outbound

This paper cites Simple and controllable music generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Simple and controllable music generation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.878743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.416011Z digest=sha256:5c36256bfacff0d4c8bf0db4fb9ab47813afc2ca569e787e4a0ab49c191abe6c

Observation 576d0a2c-d804-4a90-8709-69a1647f8eed · outbound

This paper cites Yue: Scaling open foundation models for long-form music generation,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems Yue: Scaling open foundation models for long-form music generation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.420722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.420722Z digest=sha256:bd7a1f44467fd86a5a5f07614d870471952ad553d7aeecf398e7e0426b97fbc9

Observation 012bcc6c-7aa6-46ec-8959-cd14eaf12736 · outbound

This paper cites DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T04:57:47.425009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:57:47.425009Z digest=sha256:5ee806823e1b3ca37aeda20f786209a5e0e7147292cd2b78eef9e51248a3788b

Observation 64e66076-b27c-452f-81ff-b8209cd95143 · outbound

This paper cites LP-MusicCaps: LLM-based pseudo music captioning,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems LP-MusicCaps: LLM-based pseudo music captioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.864251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.429436Z digest=sha256:b7dad16a454ea384368a9a10408c1b83cf95c2000df897d8d1182d133541a3ce

Observation b11a7715-a5a8-4d16-a5a2-6f7b40aeebd0 · outbound

This paper cites PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems PANNs: Large-scale pre- trained audio neural networks for audio pattern recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:57:47.849442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.433651Z digest=sha256:89d476d5ea5ca5e7a94700cc5c2daacca2063f58e2de74f6be6da7fa3645ad43

Pith citing papers

Observation 6d90708e-c42e-408a-b0ac-d99b91e74eb1 · inbound

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems cites this paper.

From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:57:47.834783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-16T04:57:47.249767Z digest=sha256:41673f05ba43ca8b53fdac2b20abdbd9b2f80bae8cab3fbd680df96886be9eed