Pith. sign in

Paper Citation Record · LEDGER

ACE-Step: A Step Towards Music Generation Foundation Model

As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 38 inbound Pith citation observations for arXiv:2506.00045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00045 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:19.338517Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:28:26.284368Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:48:59.860115Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3000579-c655-4ccd-8eeb-6ad66dab2aef · outbound

This paper cites Yue: Scaling open foundation models for long-form music generation, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Yue: Scaling open foundation models for long-form music generation, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.668650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:12.788886Z digest=sha256:88c6ffeb96ccc64178a20567bfb172894524cfedb95c888e43f192f1db2a3161

Observation 4255b198-fae2-4325-a752-c7ad77d29d30 · outbound

This paper cites Songgen: A single stage auto-regressive transformer for text-to-song generation, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Songgen: A single stage auto-regressive transformer for text-to-song generation, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.446330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:12.857886Z digest=sha256:2d81b557976958cb2a2f5f755effc54532e20755f70e9148b3fde0323b3e144b

Observation d98a8a89-41db-4cec-936a-a6fd0dab535f · outbound

This paper cites DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion.

ACE-Step: A Step Towards Music Generation Foundation Model DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:12.977552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:12.977552Z digest=sha256:706ea6f4ee5fc7f676c029fb49f414275d24e559b08eed5fd0dc3150df401f91

Observation b6a86041-e35e-4705-83e3-54706ec5eb98 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.077727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.077727Z digest=sha256:9ee999a9c70064367f7074e4ed65c18c4740175a0dffd4b38b1642d70fbb0a0e

Observation 433aee22-c312-42d5-86a4-7cf845f2b03f · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

ACE-Step: A Step Towards Music Generation Foundation Model Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.142079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.142079Z digest=sha256:f76c03dbea4d7e2c128117d13013c5b56337edb2e64fdbb03dc497776998674b

Observation 2589ed82-51ac-4461-8b3c-f03ef7a4ae49 · outbound

This paper cites Mert: Acoustic music understanding model with large-scale self-supervised training, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Mert: Acoustic music understanding model with large-scale self-supervised training, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.180698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:13.234494Z digest=sha256:99611b22d29d72a1a543150ad302098e62c441465c66797f2e7c5a978e890db9

Observation 3c5ea20a-81e4-40f1-a753-405ba1c4e262 · outbound

This paper cites mhubert-147: A compact multilingual hubert model.

ACE-Step: A Step Towards Music Generation Foundation Model mhubert-147: A compact multilingual hubert model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.971652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:13.348470Z digest=sha256:8921d2a81641262fb57db629299ceba89703c2a6b3bc9264bea578e176ddc783

Observation cf59c7bb-f0ad-4a2a-9c97-9ebdb1aecd2e · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

ACE-Step: A Step Towards Music Generation Foundation Model Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.472359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.472359Z digest=sha256:7a93beaa71b5a0180b59dc17a20ee5c6919515293b86533f9e4859c1a0d5c514

Observation 3aa12378-f07a-4b68-b16b-689f0133f0bd · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

ACE-Step: A Step Towards Music Generation Foundation Model High-resolution image synthesis with latent diffusion models, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.580729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.580729Z digest=sha256:1539638f3afeb5f65a75dda0a079c5adec60c67f9d894000e0be2bf614743c68

Observation a665c522-449b-4a78-9fe9-ab3b64547272 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.698954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.698954Z digest=sha256:969d014e2058f84a9644dc70467ab01fecc898456a619a95e52a3c78012bfacb

Observation 2528ee31-d1c1-4366-b2e2-c0c7be8d4db0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ACE-Step: A Step Towards Music Generation Foundation Model LLaMA: Open and Efficient Foundation Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.846965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.846965Z digest=sha256:2dd711ac35488e0698349151f364766d23dfc31929f61bd9c39a749ca894c927

Observation 5dbff2c0-301d-4a87-acd6-9b7f09410249 · outbound

This paper cites Qwen Technical Report.

ACE-Step: A Step Towards Music Generation Foundation Model Qwen Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.964802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.964802Z digest=sha256:b0d2fa4498b5395d9bb143c7ae042ae72db2cb1c478989ad570873a41b759724

Observation d73db496-44ac-42ad-88d2-2a32b812dc3d · outbound

This paper cites Deepseek-v3 technical report, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Deepseek-v3 technical report, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.113728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.113728Z digest=sha256:ae8bc9919c9a66320e795883a6511a64a3fcd8f8c28476b7b5504f1d649b62fc

Observation 54a29bf8-ac64-44fc-87d2-774e64440123 · outbound

This paper cites Improving Image Generation with Better Captions.

ACE-Step: A Step Towards Music Generation Foundation Model Improving Image Generation with Better Captions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.769749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:14.253432Z digest=sha256:d74f69a120b495d3df75340b7d8e0abb028f119d9fb9e37f8789fd4e918db370

Observation 3db75307-76be-43f4-be64-410f63a112bc · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.370744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.370744Z digest=sha256:b6251a51373279b3f14b7234bea87141159c2e257e06069114763cc67d628e1e

Observation 032a81a8-48ef-4875-acc7-b16e737b0078 · outbound

This paper cites Imagen 3.

ACE-Step: A Step Towards Music Generation Foundation Model Imagen 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.542576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:14.498247Z digest=sha256:3d172efceb88a5be984ee264f1af572a44dde9fdd2dcc93aa0c40238836b1209

Observation b3d4face-d1e9-4fe9-8575-ff0f68040a27 · outbound

This paper cites Video generation models as world simulators.

ACE-Step: A Step Towards Music Generation Foundation Model Video generation models as world simulators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.306066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:14.590362Z digest=sha256:e7cd5cfc50b45dc99fb7b77630c613fc4f217239b04045ccd944c5f4071b6e60

Observation 36bfe483-0532-4eb5-a634-1a25ad3b1c5c · outbound

This paper cites Kling AI Video Generation Large Model.

ACE-Step: A Step Towards Music Generation Foundation Model Kling AI Video Generation Large Model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.129381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:14.758686Z digest=sha256:8a4b18dc2b91a43e61cbe89041f44cc968545984c25e40ab352ab323c69e2c8c

Observation 0eeb16a1-3e25-4afd-939d-d8ab09c85dcc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

ACE-Step: A Step Towards Music Generation Foundation Model Wan: Open and Advanced Large-Scale Video Generative Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.901360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.901360Z digest=sha256:1dfbf7c5341ad24ada6781d5838870a5a4090ef8f3b6781c2ec55e56df148441

Observation c1ab91cc-a288-4e78-94fd-1b9f9c61077f · outbound

This paper cites Jukebox: A Generative Model for Music.

ACE-Step: A Step Towards Music Generation Foundation Model Jukebox: A Generative Model for Music

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.015975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.015975Z digest=sha256:bc14bb7b3e68f26db3741e2bbd4fee0a14b1c1f389ea158bb49101588b9480a3

Observation fb8fcf73-5fa6-4dad-94bf-0430639983ac · outbound

This paper cites Simple and controllable music generation, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Simple and controllable music generation, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.969399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.135052Z digest=sha256:9c45681bd9f992c8888b1bdb528caab445021bbba598b52f4db54e28aa65a2dd

Observation cefd8126-934c-499e-86dd-8758f887d83b · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models.

ACE-Step: A Step Towards Music Generation Foundation Model AudioLDM: Text-to-audio generation with latent diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.277712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.277712Z digest=sha256:d39bcd3cd176563911d23c9c16db59965f9fabd08f475804d2a95d2fbc6b9400

Observation 2c80b67f-2374-4816-a3e5-48bee299dd49 · outbound

This paper cites Plumbley.

ACE-Step: A Step Towards Music Generation Foundation Model Plumbley

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.735523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.444744Z digest=sha256:b38de62cbe1ca746b13346a8a87b2c68b86ff98e541ea37cceacd9d99cda5bde

Observation be19aa28-23b5-4c92-b8ac-b9afdf88f0ea · outbound

This paper cites Suno: Ai music generation platform.

ACE-Step: A Step Towards Music Generation Foundation Model Suno: Ai music generation platform

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.529406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.551102Z digest=sha256:d4dbbd867a8b6419458a6e5995740b575e3ad3a48400df5c9030bf7721eacd27

Observation 77fd87bf-e74c-457d-b26c-12c7b13daf85 · outbound

This paper cites Udio: Ai music creation platform.

ACE-Step: A Step Towards Music Generation Foundation Model Udio: Ai music creation platform

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.260573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.666192Z digest=sha256:974ea7c0ad251a1701193369e3be741e89f0038d55a391c98d6aad54bd7d0424

Observation da6d2e8c-483f-4fd1-86da-7988f2c1a9f4 · outbound

This paper cites Riffusion: Stable diffusion for real-time music generation.

ACE-Step: A Step Towards Music Generation Foundation Model Riffusion: Stable diffusion for real-time music generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.004960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.731915Z digest=sha256:8eb4e25df5a84591e1b203d5534b79bb977e87ab5336a92a655846f739a15333

Observation 502e9e4b-5afd-432d-8ca9-7055c521d108 · outbound

This paper cites Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

ACE-Step: A Step Towards Music Generation Foundation Model Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.788346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.788346Z digest=sha256:1d12b452cb1bbc9ac9ed5b65795214386d2b733663bb8d591699c4c7ef1ca805

Observation 41b8446b-b645-45db-b479-d1ca6c03f281 · outbound

This paper cites Fast text-to-audio generation with adversarial post-training, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Fast text-to-audio generation with adversarial post-training, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.787291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.845585Z digest=sha256:486f4b8a5e6bdc33ec7f6d352a16c0332c4ff80e41073395ab3085f1fd16f081

Observation 5b8c6a90-b403-45be-91b9-c5ad4371ff17 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

ACE-Step: A Step Towards Music Generation Foundation Model Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.893630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.893630Z digest=sha256:39b0dfdc4f3b874d55800c1cb7afefd0b7034193910fc02d9447da4aeb24aaa4

Observation 67bd7ed7-129c-4f63-a0c9-86e7737876e3 · outbound

This paper cites Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.589340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:15.986885Z digest=sha256:148a0d019d40ba2cbadb61646a1e3673c2b6af62dbc8e330571ef02606bcd11c

Observation 6d5f9850-9028-4703-a449-c41822ad9d5a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

ACE-Step: A Step Towards Music Generation Foundation Model Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.081109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.081109Z digest=sha256:2a08824f04a3e221cfc12c27368e33cabfce0315900127662a1849ba23899ea3

Observation 63d68aac-c2fc-4035-8987-286191a8ca6c · outbound

This paper cites Adding conditional control to text-to-image diffusion models, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Adding conditional control to text-to-image diffusion models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.172758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.172758Z digest=sha256:8ffc100f4eeeb025cc1d9fcc5687836a4b3591a74a8fa696d3b23b47eeb0b432

Observation b9368db9-aa8b-432d-9efc-e8f9e1fadb3a · outbound

This paper cites Statistical parametric speech synthesis, 2007.

ACE-Step: A Step Towards Music Generation Foundation Model Statistical parametric speech synthesis, 2007

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.444048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:16.265993Z digest=sha256:e8970585adcdf6114362f93d618752c36efdf0429cb9679f95368218e76398ea

Observation 587953c3-962d-4b96-8ae5-507bfbef7791 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

ACE-Step: A Step Towards Music Generation Foundation Model MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.366580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.366580Z digest=sha256:6f4f7ade9924c7b2237597d58cabbf754e3cceb56d5f83ac47756737dbd46046

Observation f2543d06-48ff-44c9-b707-fe77133774a3 · outbound

This paper cites Neural discrete representation learning, 2018.

ACE-Step: A Step Towards Music Generation Foundation Model Neural discrete representation learning, 2018

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.460073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.460073Z digest=sha256:3415e42106f74e4e640d19a26a9d233073d8d7fc68cb1d16129aca06d1ff10b2

Observation 6dd4fbcc-2fb6-41c2-ae42-37249840015a · outbound

This paper cites Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank.

ACE-Step: A Step Towards Music Generation Foundation Model Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.263871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:16.534619Z digest=sha256:61625c98c05e998ed25891688b6469d585223831021841cdb93a5ac132a1712b

Observation 6e5baced-7c4b-4e37-96f8-8f2d22c21b12 · outbound

This paper cites High fidelity neural audio compression, 2022.

ACE-Step: A Step Towards Music Generation Foundation Model High fidelity neural audio compression, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.059052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:16.619063Z digest=sha256:d21a94fd1253c0a3572e9bca0e83151cebf357e3e423a9db68472b85d2e013d7

Observation b38abdfa-eddb-436a-96b3-1811361198a5 · outbound

This paper cites Large- scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

ACE-Step: A Step Towards Music Generation Foundation Model Large- scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.790631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:16.749152Z digest=sha256:6881a0af8eb68b0507e21d273c20cd8180fd933d7ec4079386a591b5dbc8d881

Observation a5bbe826-79cd-46da-9e6c-766a694cd5b4 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.842940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.842940Z digest=sha256:a6fdf36ad8e6ec67b6d6b18ce5073ccd3ab5b5d94476e8278306cea83b650319

Observation 0a48cfb5-87c1-4360-94a0-4a48302fc962 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.936678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.936678Z digest=sha256:909ed37fb3c2b03951e585f18715840680ec75e699d9c0e677156eafad8a1b18

Observation 55040861-a98a-43d6-87d5-39e4c0d1eed6 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Scalable diffusion models with transformers, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.004006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.004006Z digest=sha256:f11b3b6a33da593a3c1bee62d1ec7394462752c49e6c7c573de22d6c48b015bb

Observation 40dc43ad-6a7c-4351-b11e-9b8bd383284b · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

ACE-Step: A Step Towards Music Generation Foundation Model Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.092132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.092132Z digest=sha256:4679abeb95cd4a9f24ab76ccdd44337b66064412b0d1cb61abb06258ed5427b1

Observation f2465be3-9e95-4ce1-8488-3cc5f589b692 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation, 2015.

ACE-Step: A Step Towards Music Generation Foundation Model U-net: Convolutional networks for biomedical image segmentation, 2015

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.167427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.167427Z digest=sha256:38074e8a7299b4974cb29467a920cd84f83054bf156bcae810f07c7fe4b906f4

Observation 6b2dcfad-8129-4fce-96a4-18d49de674f5 · outbound

This paper cites Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound.

ACE-Step: A Step Towards Music Generation Foundation Model Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.253728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.253728Z digest=sha256:bb293edd271fb88b242cf21b8b1e574622385e70b6247c3ae4d7af20a030e124

Observation 62638e80-0a3d-4f37-8f4c-bc087b8beeee · outbound

This paper cites Qwen2.5-Omni Technical Report.

ACE-Step: A Step Towards Music Generation Foundation Model Qwen2.5-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.352668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.352668Z digest=sha256:e974e6a465b96e80ac36e6d1e4f362e2e9a61a47ac149df1eec0448710395297

Observation a724f43c-4e91-4834-95d8-45eda269d2b9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

ACE-Step: A Step Towards Music Generation Foundation Model Robust speech recognition via large-scale weak supervision, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.424130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.424130Z digest=sha256:c7841d2f07aaad414064b5a19902f5af33b24f1d52c47e3ce982747b0a771de3

Observation bfc12106-f8f9-4142-96a3-2b41058488de · outbound

This paper cites ekzhu/datasketch: v1.6.5, May 2024.

ACE-Step: A Step Towards Music Generation Foundation Model ekzhu/datasketch: v1.6.5, May 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.578319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:17.512771Z digest=sha256:4a4fa4caaeae2da7e1d2e27d138cf61e0a4bd84489ebf72c984e6cb3239eb7c0

Observation 35c29adf-ab98-49a5-b5b2-aa40a98aee3c · outbound

This paper cites Byt5 model for massively multilingual grapheme-to-phoneme conversion.

ACE-Step: A Step Towards Music Generation Foundation Model Byt5 model for massively multilingual grapheme-to-phoneme conversion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.373421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:17.590060Z digest=sha256:f9c9e3029e698f3ed5c1d4c2c2cd7fc843526e09d6168d77acdc0b9e8bb7c463

Observation b98c5d1f-f36e-4a92-99c6-a50be36cd851 · outbound

This paper cites All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio.

ACE-Step: A Step Towards Music Generation Foundation Model All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.189976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:17.671225Z digest=sha256:c6c27e84fc2acc7e709a0db496fc8230710379b7a12cb87a4c6fa0aa00e6bd22

Observation 68373a3d-a027-485c-9967-8950194a1811 · outbound

This paper cites Beat this! accurate beat tracking without DBN postpro- cessing.

ACE-Step: A Step Towards Music Generation Foundation Model Beat this! accurate beat tracking without DBN postpro- cessing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.948160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:17.799489Z digest=sha256:cdd1a06543237b50217e5c0862f64cce5b1bdf4bd3dd51f0e112f521b5b7901b

Observation cbdea552-a5c0-4afe-81ec-8c671937469c · outbound

This paper cites Essentia: An audio analysis library for music information retrieval.

ACE-Step: A Step Towards Music Generation Foundation Model Essentia: An audio analysis library for music information retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.776114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:17.920550Z digest=sha256:3fb2fa432abfa64a803bf5d73bab7fb5bcb10371ce871c1873d9cd19bf25338b

Observation 711bf726-379a-44c7-a7b0-2038a34f9ceb · outbound

This paper cites Xtts: a massively multilingual zero-shot text-to-speech model, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Xtts: a massively multilingual zero-shot text-to-speech model, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.580518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.015834Z digest=sha256:ebe58cd2e3f39506c0973cf0e57a0e06aa255c17829ed859167fc3181b98382f

Observation 876fc305-bd3e-4431-9b59-e4b551bf33c4 · outbound

This paper cites Fish-speech: Leveraging large language models for advanced multilingual text-to-speech synthesis, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Fish-speech: Leveraging large language models for advanced multilingual text-to-speech synthesis, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.426135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.106129Z digest=sha256:45511d010b7c0c0c2ebaab2a475850933cea88ec0a5c45faa4ccb58b6dcc0295

Observation e626520f-3c7f-4b8d-88fa-fefb94a03737 · outbound

This paper cites Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.231538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.231538Z digest=sha256:5f077d2df12ca6340ab9c7184fa6bb62f28385cd1f51ce830b45f06e6cba737c

Observation 9879f042-94ca-47cf-899d-307e3e426a08 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.277393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.345992Z digest=sha256:b9b797f484acf4d551354da89b4d00d1d0b66842a926a3bd00695cad0d4cd225

Observation 3f70d972-b2da-4ae0-9ce9-9a8f7eae6228 · outbound

This paper cites Learning diverse features with part-level resolution for person re-identification, 2020.

ACE-Step: A Step Towards Music Generation Foundation Model Learning diverse features with part-level resolution for person re-identification, 2020

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.133842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.495011Z digest=sha256:c1a498ac92094ec5c39570c65273cd94ff53998efbf184cb51d0bed23585279d

Observation 4ef08c98-252d-4b9d-8399-48c7620debb5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.572288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.572288Z digest=sha256:a41ad88f7609fcd08689c4ee8023425374f76928347d6158ce1e0e1917ced579

Observation 9e9a9e3c-20c2-44da-b4ca-9d0df6151958 · outbound

This paper cites Adapting frechet audio distance for generative music evaluation.

ACE-Step: A Step Towards Music Generation Foundation Model Adapting frechet audio distance for generative music evaluation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.979871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.701310Z digest=sha256:bec60b20c03458d661f039afe2e425525f4503cf001e3bcb7603748275077fcd

Observation 837ce3fc-4811-4b2a-81eb-be22b0b2d8ac · outbound

This paper cites Stable Audio Metrics, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Stable Audio Metrics, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.788623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.830656Z digest=sha256:7b0ffd361fd561b6c0cbf6e0cc2762ce958ebbc2782ed1d8c91440d2eceb40fa

Observation a56da72c-0086-41a0-85d5-db3e88f850a4 · outbound

This paper cites MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization.

ACE-Step: A Step Towards Music Generation Foundation Model MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.929184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.929184Z digest=sha256:856f41c610c7b6bf97de4b5b096870c22ac8a9b38b42d3c0df9063198400f167

Observation f76fba90-df7f-43ae-af5a-a196407df3fe · outbound

This paper cites stable-ts: Stabilizing Timestamps for Whisper, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model stable-ts: Stabilizing Timestamps for Whisper, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.581782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:16:18.996979Z digest=sha256:ee8b892495a4c550a649674c70e791816b55855744b238aea3de3c4e4510a5af

Observation 73391fcb-d9e8-4f78-a451-9e57ef435084 · outbound

This paper cites SongEval: A Benchmark Dataset for Song Aesthetics Evaluation.

ACE-Step: A Step Towards Music Generation Foundation Model SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.098611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.098611Z digest=sha256:c97e8714168400f1b7f126f2b6ee82b39c091eaf675953d533b75e17efcf5d56

Observation b5ccf2bb-3846-43b3-b876-238fa7d2be45 · outbound

This paper cites Simplifying, stabilizing and scaling continuous-time consistency models, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Simplifying, stabilizing and scaling continuous-time consistency models, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.207061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.207061Z digest=sha256:1f12b28a6624fb733fbe1ef5048965660edf0c81f5aa8e458357cb98d923447f

Observation 1a7b8550-da2e-4817-86e1-545e4bad6098 · outbound

This paper cites FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models.

ACE-Step: A Step Towards Music Generation Foundation Model FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.338517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.338517Z digest=sha256:56cf94986ffe8ea1f335d4cf4c1c6f08a5ef391b411898e757e42f11d20e7ecf

Pith citing papers

Observation 607794e3-305a-4413-8fa2-5d90040a27f4 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.507871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.507871Z digest=sha256:a01657ea6244869050fd5ddab3b46efc2e86a2a7cfe3968ac550f9a317f8d03f

Observation c7bb091b-edf2-421f-808d-1cf2109f21c1 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.880406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.880406Z digest=sha256:7ad68c5a70cebb8d6f9fea83558a501dca868a20f8fdf7f06dc93f33d813b7fc

Observation 0a61d2d4-36d4-439c-83c2-634dc4032319 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.692143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.692143Z digest=sha256:6556fc0eeae15c85bd1e8319ab9ac19e9b632eafeb3872eae3fd01c8900e493f

Observation 5110b3af-f508-44e5-b08f-763b95cffc20 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts ACE-Step: A Step Towards Music Generation Foundation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.596423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.596423Z digest=sha256:fdec3056b97cc10ab4b37c50bc3b3b6700ed1ec3d05e210749d8992b4063fca0

Observation af8bd234-643b-4344-aedd-ca8b67093cd6 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:03:15.219965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:11a0000d38f6c2e4850b25c7f073c8e74d014476348567a6519298f5345ffe40

Observation b6b2c2c2-c7a3-481f-8e9c-b03a6dd75d6a · inbound

SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision cites this paper.

SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:46:16.858659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T10:45:09.089710Z digest=sha256:97f1fd60222bade97809d658fdb166848683fe19d8dc2be5455cb84e09feb6f4

Observation 7ae5821c-dec4-4de1-bb54-c93df909c2ee · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing ACE-Step: A Step Towards Music Generation Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.887395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.887395Z digest=sha256:a72800c4b8c4af936818f1e828e86625c27ba7f788c9fe53f21cf06c238f81e8

Observation 9968680e-dce0-4bf3-84ac-c4ce297dd30f · inbound

TADA! Tuning Audio Diffusion Models through Activation Steering cites this paper.

TADA! Tuning Audio Diffusion Models through Activation Steering ACE-Step: A Step Towards Music Generation Foundation Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.551922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T13:40:47.112814Z digest=sha256:6cd105e315f40514fdb7df3cfba01836310e92ca829c5a6275c597639c9eb0bb

Observation 7c6cabd7-a8aa-43c8-be3d-6880c211ae87 · inbound

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline cites this paper.

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.918821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T20:17:45.920234Z digest=sha256:9e1e3ea06bb898080f1f0ad87b72563f47123b3fb133d4a17891ccb318f76a3f

Observation 310031d3-f4a0-4f8d-9b1a-2c3d248d3cb6 · inbound

Echoes: A semantically-aligned music deepfake detection dataset cites this paper.

Echoes: A semantically-aligned music deepfake detection dataset ACE-Step: A Step Towards Music Generation Foundation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T17:35:47.688979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:35:47.688979Z digest=sha256:dcfb89b87aa7d982174bbaf013015d193e6248cb7a3b25d98c1637bc553646d5

Observation ecbcf569-4432-4a67-8b22-40bd3d323ec8 · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:01.532365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:08:58.648967Z digest=sha256:280e3e26dc0d2dc45264ed188d3c726419522743f418504dc81d9b9f9a697adc

Observation 13a0087d-9539-4c4a-8519-843e847b4df9 · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:21.119692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:21.119692Z digest=sha256:5e21d8eb0c727f0f533ce639e9949663a2bc2e53a2e13abf2a319a23e3aeb742

Observation d0263534-865a-4bae-b7db-23f713ff310f · inbound

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment cites this paper.

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:22.334741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T10:14:41.499130Z digest=sha256:6b9ef3356746dcce9c548a5aa5f5aa96d501b7664709cea161a3c6eb777be385

Observation 86ef11ed-3261-4695-a0b5-79949d4777c1 · inbound

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation cites this paper.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.108044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T16:11:16.551910Z digest=sha256:f58b2a1a1feb818e5d54a3e7074213a28318286e9768cb4128ac534e4c0977ed

Observation be05e593-f1f9-4e55-abdc-b9638dcd151c · inbound

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music cites this paper.

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music ACE-Step: A Step Towards Music Generation Foundation Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:56:29.750313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T13:27:41.439049Z digest=sha256:fa082c871b329cf9d7c9e645611d2390f8bfb69446486e168c185054982e1603

Observation 36991cea-e9fa-440d-b263-15473293870c · inbound

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music cites this paper.

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music ACE-Step: A Step Towards Music Generation Foundation Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:45:12.341833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T00:35:42.697038Z digest=sha256:dd9888ae6937e9478b110ad7a884abce75e63be38e97a196d7e0b53bd513191d

Observation 8cef5aeb-d6b0-4e75-aee4-3ee51dfd1d14 · inbound

Cutting rules in strong field QED with application to trident pair production cites this paper.

Cutting rules in strong field QED with application to trident pair production ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T16:57:23.214654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:57:23.214654Z digest=sha256:e036140a28b7a12977b338488f9733260f1525e39f2ccf585fe8bf9515e0d452

Observation 12dd018e-0aad-4d75-9e9d-1619b69a25eb · inbound

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation cites this paper.

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:17:22.920096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T06:16:20.612291Z digest=sha256:26c98f3c2f53e6023ac1e4419a9f951b27c0c3fc9062aba0537df3e0f89fc97b

Observation 788b2689-12cb-414e-8b9a-0b20e04cbf89 · inbound

S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation cites this paper.

S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T22:52:50.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T22:50:08.321514Z digest=sha256:fe89d3bb21421950e020adb6ec42d91213ebda9e28bf20b8a26c8f5a2849962d

Observation ebc97580-98ed-49a8-a723-6ca14b67c2c3 · inbound

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches cites this paper.

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches ACE-Step: A Step Towards Music Generation Foundation Model

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:54.924541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T02:43:16.101544Z digest=sha256:ca8a650cdff236ff22fc7c0025880563a308532f0fed7a66f154824fa20fa43f

Observation 9d0fd53c-60e0-46d6-8c91-6b165351c0aa · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.662750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:0eb7e0ef9fb1964234fe78a582e06ed283aebbcd45cd32955c8507400052dfd8

Observation d47d9dbf-9398-49a1-b09c-e73fb9ac7a28 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:57:19.597322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:831c89acf79f09fca667273123583b6c3c1dd8a823141e9a86603fa40e033e03

Observation e3b93717-bbc5-4583-8be1-24859dd65462 · inbound

An Empirical Analysis of AI Slop in Music Streaming cites this paper.

An Empirical Analysis of AI Slop in Music Streaming ACE-Step: A Step Towards Music Generation Foundation Model

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:48:59.861457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T00:08:36.279506Z digest=sha256:4f91bf248ed468e94ddf4379bf411c2c58f7f0946a8d47c31f67fee6c9c8357e

Observation ddcc77e6-6770-4132-9653-1005ed10afa5 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training ACE-Step: A Step Towards Music Generation Foundation Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:15:47.551038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:6533bea1304084a776d5656cf950c7879670a234d416d8fa523617abbdf2d3b4

Observation 1db8f3b9-87af-4035-ab09-b0e72081ee04 · inbound

MusicMark: A Robust Generative Watermarking Framework for Music Generation cites this paper.

MusicMark: A Robust Generative Watermarking Framework for Music Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T06:56:53.384109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:56:53.384109Z digest=sha256:4094fcd4c0ea71fab4423bf7286f83aa3588e6a370a88a2e04ee6671d9875631

Observation f2752145-d739-4f3b-abbc-59bf5cc42406 · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T06:45:43.330341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:45:43.330341Z digest=sha256:70d3b736edd401795e60266e0100b5ab0eda6b93ca16ccc3cc1f765ff942416d

Observation 91482bf6-aca7-4c5f-9ea9-c288237debd6 · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:05:04.510458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:05:04.510458Z digest=sha256:515c0c6ec03cc1b2e92ce484f96aa5f3d0d7aaf9d700638ac01852d712fa0270

Observation 27e079cd-46e0-4357-a4b9-92661222b514 · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report ACE-Step: A Step Towards Music Generation Foundation Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T03:47:22.936776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:47:22.936776Z digest=sha256:3991b3b7e78216919b24998feeb9395a99e176b476e9fc5847247b9a7a4a2980

Observation fe00cef8-e6c0-4493-ab55-6566a33784d9 · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report ACE-Step: A Step Towards Music Generation Foundation Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:52:26.903347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:52:26.903347Z digest=sha256:4b25e0506d65c286534999fa5704553126de6a73e4b52209e9f62ba9495ae5a6

Observation db7d261c-a992-4881-837a-d860a877bf17 · inbound

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation cites this paper.

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:25:50.281276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:25:50.281276Z digest=sha256:c9c68e807c9359f95c121e8d52993b8cf1848b490ec7f59bd2692a26dc73cabb

Observation d949e0fb-134b-4f84-8d6c-0bedcbd70441 · inbound

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features cites this paper.

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features ACE-Step: A Step Towards Music Generation Foundation Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:01:57.134742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:01:57.134742Z digest=sha256:25d156671539189d94763002602146b43562f5ba32d51d66c4ea5b906f019e06

Observation c2a48299-6730-44a6-baaf-3ad500282caf · inbound

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering cites this paper.

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering ACE-Step: A Step Towards Music Generation Foundation Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:19.082295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:19.082295Z digest=sha256:f8936dd9e03d107e2393c6916c0019563b3cc0ef42124df26ef6c7c738cbba13

Observation 17568aa2-93fd-49dc-b420-c0f47a369205 · inbound

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation cites this paper.

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T23:40:07.801267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:40:07.801267Z digest=sha256:9da655eddddbc9c49d24bf438d96df7091a8ce6e0a2a1425c9757b2adb8ed22b

Observation 37152fb8-a22d-4c0f-b544-e5e048dca679 · inbound

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models cites this paper.

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models ACE-Step: A Step Towards Music Generation Foundation Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T01:28:28.022775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:28:28.022775Z digest=sha256:535ef3718e6ddb8d1e276d6d0cd86db8d32f42910f4c5e16ca66bf19ef4c1c81

Observation 4942d2d7-e8a2-4feb-bb5a-321663b226e2 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:20.556905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:20.556905Z digest=sha256:46681a2dc27a413f348e363f80977dd3e8d305b6c980b4e91c7e5478d2767977

Observation 0518d6f5-8d75-48c3-b78a-d633529e7370 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.128810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.128810Z digest=sha256:cb1492fcd3a9f66677987bc0c7baa8f530509694b5d1bcb87d75f23df923631c

Observation 49941bad-a2f8-4920-83bd-d304cf9244a5 · inbound

Beyond Reconstruction: Full-Context Generative DiT for Music Generation cites this paper.

Beyond Reconstruction: Full-Context Generative DiT for Music Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:30:19.892233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:30:19.892233Z digest=sha256:5aad81a95dd2b7b8b4c05ea47c6e7553f4bf2b816723f2f9a5896587d14e87f0

Observation 11c254ee-56bb-40d0-91c6-07e21c5ba28e · inbound

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping cites this paper.

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:28:26.284368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:28:26.284368Z digest=sha256:32dd1d21c91fe996116e1b6d828adaea058a204ab939df89076c9d7d21a303c7