Pith. sign in

Paper Citation Record · LEDGER

ACE-Step: A Step Towards Music Generation Foundation Model

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 36 inbound Pith citation observations for arXiv:2506.00045.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00045 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:16:19.338517Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:43.128810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T21:48:59.860115Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3000579-c655-4ccd-8eeb-6ad66dab2aef · outbound

This paper cites Yue: Scaling open foundation models for long-form music generation, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Yue: Scaling open foundation models for long-form music generation, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.668650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:12.788886Z digest=sha256:cbcae6f7dfa1e31fb3bcd7c93944f5d3cdc134b8d3424b3743c6fb9f2c50f8c0

Observation 4255b198-fae2-4325-a752-c7ad77d29d30 · outbound

This paper cites Songgen: A single stage auto-regressive transformer for text-to-song generation, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Songgen: A single stage auto-regressive transformer for text-to-song generation, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.446330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:12.857886Z digest=sha256:499ccdacf2b24bacaac10eecf4672185199b756eaf8e958648a19610365c69bb

Observation d98a8a89-41db-4cec-936a-a6fd0dab535f · outbound

This paper cites DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion.

ACE-Step: A Step Towards Music Generation Foundation Model DiffRhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:12.977552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:12.977552Z digest=sha256:352f7894d3e222387a7c171f80060fdd817b02f50ef734e8417f86b5182a71d4

Observation b6a86041-e35e-4705-83e3-54706ec5eb98 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Sana: Efficient high-resolution image synthesis with linear diffusion transformer, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.077727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.077727Z digest=sha256:25d179949b28183f4b82bfd7d64b7a4dd197bd1a6f6a5cdfc2603ec24b467a90

Observation 433aee22-c312-42d5-86a4-7cf845f2b03f · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

ACE-Step: A Step Towards Music Generation Foundation Model Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.142079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.142079Z digest=sha256:686f4d6187c1d03c17f43beec97dadd77a9f5da72395db567a57bac61f91dbea

Observation 2589ed82-51ac-4461-8b3c-f03ef7a4ae49 · outbound

This paper cites Mert: Acoustic music understanding model with large-scale self-supervised training, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Mert: Acoustic music understanding model with large-scale self-supervised training, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:25.180698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:13.234494Z digest=sha256:dba579dbc84ee9b3889bb7e8fd8b7e539c5a2e7b898e9f65302b57ee98e8ae72

Observation 3c5ea20a-81e4-40f1-a753-405ba1c4e262 · outbound

This paper cites mhubert-147: A compact multilingual hubert model.

ACE-Step: A Step Towards Music Generation Foundation Model mhubert-147: A compact multilingual hubert model

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.971652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:13.348470Z digest=sha256:8e8d0bffdb3fb3edad5f66b170c71675f5906b968b8050d2c9f0bcfa1d2b835b

Observation cf59c7bb-f0ad-4a2a-9c97-9ebdb1aecd2e · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

ACE-Step: A Step Towards Music Generation Foundation Model Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.472359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.472359Z digest=sha256:05163d769b78c5a0ea24728e045a37b7fd23abdf9152fbd7299daf90572e2824

Observation 3aa12378-f07a-4b68-b16b-689f0133f0bd · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

ACE-Step: A Step Towards Music Generation Foundation Model High-resolution image synthesis with latent diffusion models, 2021

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.580729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.580729Z digest=sha256:9f81abafdc1e6db97ab25cf5478ddd24e8ff0a173aa89fe3bb7951461c204835

Observation a665c522-449b-4a78-9fe9-ab3b64547272 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.698954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.698954Z digest=sha256:517a9c0bdaf4d24a038cc6acb79978bbf65023113461f3c150ef9a1dba9c6bfe

Observation 2528ee31-d1c1-4366-b2e2-c0c7be8d4db0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

ACE-Step: A Step Towards Music Generation Foundation Model LLaMA: Open and Efficient Foundation Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.846965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.846965Z digest=sha256:27f0e0c9a28d7e8f417822b8f5a020eb8278c81041adea6f6b2922e229e649e4

Observation 5dbff2c0-301d-4a87-acd6-9b7f09410249 · outbound

This paper cites Qwen Technical Report.

ACE-Step: A Step Towards Music Generation Foundation Model Qwen Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:13.964802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:13.964802Z digest=sha256:591d31b469a85206f8fc1ca6b4c3d251c991838d528dcfd4ddd89f97a559c35b

Observation d73db496-44ac-42ad-88d2-2a32b812dc3d · outbound

This paper cites Deepseek-v3 technical report, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Deepseek-v3 technical report, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.113728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.113728Z digest=sha256:7476beb2845139d59d0b5bcbcdd2bf7be484f40a6e834e88a101349a619110fb

Observation 54a29bf8-ac64-44fc-87d2-774e64440123 · outbound

This paper cites Improving Image Generation with Better Captions.

ACE-Step: A Step Towards Music Generation Foundation Model Improving Image Generation with Better Captions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.769749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:14.253432Z digest=sha256:cb44f8d9f99d208823b2fd3c3cd7362f0d6f1c8ef6c641e3a53c68f6e16e53bd

Observation 3db75307-76be-43f4-be64-410f63a112bc · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.370744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.370744Z digest=sha256:eec7275897f7a2ad7345cc20f4103a7b00f4a0cbdd704c7d7c5efc3be69ea6e6

Observation 032a81a8-48ef-4875-acc7-b16e737b0078 · outbound

This paper cites Imagen 3.

ACE-Step: A Step Towards Music Generation Foundation Model Imagen 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.542576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:14.498247Z digest=sha256:3e07be1260e3d8e99f63e3376c83b876b432cbbfc1a68536d0770d6ff5984b9d

Observation b3d4face-d1e9-4fe9-8575-ff0f68040a27 · outbound

This paper cites Video generation models as world simulators.

ACE-Step: A Step Towards Music Generation Foundation Model Video generation models as world simulators

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.306066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:14.590362Z digest=sha256:6cd47cbe679d9e1e7b342ab0095dfde1a8b8c6dc5a2ca47c48914f22ba3cfecf

Observation 36bfe483-0532-4eb5-a634-1a25ad3b1c5c · outbound

This paper cites Kling AI Video Generation Large Model.

ACE-Step: A Step Towards Music Generation Foundation Model Kling AI Video Generation Large Model

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:24.129381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:14.758686Z digest=sha256:a31ac949b2b252657281ea1d380faf382fb061b6da1f65f4ffe3df2e881c945b

Observation 0eeb16a1-3e25-4afd-939d-d8ab09c85dcc · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

ACE-Step: A Step Towards Music Generation Foundation Model Wan: Open and Advanced Large-Scale Video Generative Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:14.901360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:14.901360Z digest=sha256:985205ca9602c7b22c7c9809444eb38cd255f8fd8231066bbeec904ab6f652c3

Observation c1ab91cc-a288-4e78-94fd-1b9f9c61077f · outbound

This paper cites Jukebox: A Generative Model for Music.

ACE-Step: A Step Towards Music Generation Foundation Model Jukebox: A Generative Model for Music

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.015975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.015975Z digest=sha256:94ecf30825b0f0d8d8826e6d21bbddddf97fd5fe75c7a99912f0efd77eddc1d3

Observation fb8fcf73-5fa6-4dad-94bf-0430639983ac · outbound

This paper cites Simple and controllable music generation, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Simple and controllable music generation, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.969399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.135052Z digest=sha256:9d99f2b3aa1cb823066b4c21d25dce41ff45986bdd6cdd37f9ac1a42d8541e32

Observation cefd8126-934c-499e-86dd-8758f887d83b · outbound

This paper cites AudioLDM: Text-to-audio generation with latent diffusion models.

ACE-Step: A Step Towards Music Generation Foundation Model AudioLDM: Text-to-audio generation with latent diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.277712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.277712Z digest=sha256:a484b55bf53445bf36c4f5745c7e183995c7d142864c35696230664f2a0dada7

Observation 2c80b67f-2374-4816-a3e5-48bee299dd49 · outbound

This paper cites Plumbley.

ACE-Step: A Step Towards Music Generation Foundation Model Plumbley

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.735523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.444744Z digest=sha256:c31ebff998ff01740f969be369bc6bf39b04b4061e486ccf1cc5e63661760703

Observation be19aa28-23b5-4c92-b8ac-b9afdf88f0ea · outbound

This paper cites Suno: Ai music generation platform.

ACE-Step: A Step Towards Music Generation Foundation Model Suno: Ai music generation platform

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.529406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.551102Z digest=sha256:2665f5b0eeae33dfca0582582ba2b200cc2f9c9f25ae1880b526bfa426643b7c

Observation 77fd87bf-e74c-457d-b26c-12c7b13daf85 · outbound

This paper cites Udio: Ai music creation platform.

ACE-Step: A Step Towards Music Generation Foundation Model Udio: Ai music creation platform

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.260573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.666192Z digest=sha256:e1a48374837f71d2f52bef23d43ed5d549b8f74b353a3afbc3e45680eafe1f4d

Observation da6d2e8c-483f-4fd1-86da-7988f2c1a9f4 · outbound

This paper cites Riffusion: Stable diffusion for real-time music generation.

ACE-Step: A Step Towards Music Generation Foundation Model Riffusion: Stable diffusion for real-time music generation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:23.004960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.731915Z digest=sha256:d7411029bf87c9d1441e9705dc5c337fc474edd339c9b49472d2855087917f76

Observation 502e9e4b-5afd-432d-8ca9-7055c521d108 · outbound

This paper cites Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

ACE-Step: A Step Towards Music Generation Foundation Model Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.788346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.788346Z digest=sha256:1567c0fb5e839df65b6624ca308824be3e964ccf3e6d7757f007747c105d71c0

Observation 41b8446b-b645-45db-b479-d1ca6c03f281 · outbound

This paper cites Fast text-to-audio generation with adversarial post-training, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Fast text-to-audio generation with adversarial post-training, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.787291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.845585Z digest=sha256:98a391a1ca712140e67435d4bb35a5ebca976da835b2b86818e48244349c680b

Observation 5b8c6a90-b403-45be-91b9-c5ad4371ff17 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

ACE-Step: A Step Towards Music Generation Foundation Model Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:15.893630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:15.893630Z digest=sha256:239fed2e356b5ece6439b67b82d4f6070644cf910d44c92c47114d483fc27145

Observation 67bd7ed7-129c-4f63-a0c9-86e7737876e3 · outbound

This paper cites Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Efficientvit: Multi-scale linear attention for high-resolution dense prediction, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.589340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:15.986885Z digest=sha256:d6f12531932814db28e4433eea4f4e7b336db4a1ab5fad0413c47c3faca5dd43

Observation 6d5f9850-9028-4703-a449-c41822ad9d5a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

ACE-Step: A Step Towards Music Generation Foundation Model Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.081109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.081109Z digest=sha256:7fca79bceac594334b27271bc75bf63ac423968fd0e197e7601605ceb1dc6699

Observation 63d68aac-c2fc-4035-8987-286191a8ca6c · outbound

This paper cites Adding conditional control to text-to-image diffusion models, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Adding conditional control to text-to-image diffusion models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.172758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.172758Z digest=sha256:52ce72ad0ce8671d630eb1b08ab958508e2a28295d804a5512029e59d4d926f3

Observation b9368db9-aa8b-432d-9efc-e8f9e1fadb3a · outbound

This paper cites Statistical parametric speech synthesis, 2007.

ACE-Step: A Step Towards Music Generation Foundation Model Statistical parametric speech synthesis, 2007

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.444048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:16.265993Z digest=sha256:f7ef8e4dbd2a4a3d03d3556832f09be8bdf61a954cb866135606eb1382c46451

Observation 587953c3-962d-4b96-8ae5-507bfbef7791 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

ACE-Step: A Step Towards Music Generation Foundation Model MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.366580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.366580Z digest=sha256:898f4a2437aa780d11af2671420683c50cfc9b0ab299eaece64c461249b4cfd1

Observation f2543d06-48ff-44c9-b707-fe77133774a3 · outbound

This paper cites Neural discrete representation learning, 2018.

ACE-Step: A Step Towards Music Generation Foundation Model Neural discrete representation learning, 2018

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.460073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.460073Z digest=sha256:60f8d1f41ec1e1b4d080b76e7d1cffe4d77915e69cd2d7d86273cf45e8faa80f

Observation 6dd4fbcc-2fb6-41c2-ae42-37249840015a · outbound

This paper cites Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank.

ACE-Step: A Step Towards Music Generation Foundation Model Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.263871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:16.534619Z digest=sha256:cab291c32745d1fc4b4b0d6f87386d71cef7428fabfbfeae18dc275a360f54e3

Observation 6e5baced-7c4b-4e37-96f8-8f2d22c21b12 · outbound

This paper cites High fidelity neural audio compression, 2022.

ACE-Step: A Step Towards Music Generation Foundation Model High fidelity neural audio compression, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:22.059052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:16.619063Z digest=sha256:74334d496cf796e556545ad5941dd73ea0d85b9025b6e707de09c8ffe9a8d6c5

Observation b38abdfa-eddb-436a-96b3-1811361198a5 · outbound

This paper cites Large- scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

ACE-Step: A Step Towards Music Generation Foundation Model Large- scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.790631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:16.749152Z digest=sha256:1d710fd3c4dba9d1544403dff58d1315c236b21bbc0558d68585efca242133be

Observation a5bbe826-79cd-46da-9e6c-766a694cd5b4 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.842940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.842940Z digest=sha256:f79bc273885b58a06e3245d7bf9fa1bc2ee8889d8945d43aca99e52f1a975939

Observation 0a48cfb5-87c1-4360-94a0-4a48302fc962 · outbound

This paper cites an unresolved cited work.

ACE-Step: A Step Towards Music Generation Foundation Model Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:16.936678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:16.936678Z digest=sha256:efb7127fa4cf1023ab2b9a5582151e732466deb47c35b67b4455944e8eda1023

Observation 55040861-a98a-43d6-87d5-39e4c0d1eed6 · outbound

This paper cites Scalable diffusion models with transformers, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Scalable diffusion models with transformers, 2023

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.004006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.004006Z digest=sha256:c786b482079914bd65df123f6d89efacaa94edfea8ba0602763fcb4f85da6af6

Observation 40dc43ad-6a7c-4351-b11e-9b8bd383284b · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

ACE-Step: A Step Towards Music Generation Foundation Model Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.092132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.092132Z digest=sha256:319ae5a82d455787f88f727a5bb81b49b62fabab58c2edac5d09c867761cb3f4

Observation f2465be3-9e95-4ce1-8488-3cc5f589b692 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation, 2015.

ACE-Step: A Step Towards Music Generation Foundation Model U-net: Convolutional networks for biomedical image segmentation, 2015

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.167427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.167427Z digest=sha256:c0522ed1407d3bb88444374c2849bc6f4bb363cf3ead35bb299748874c805117

Observation 6b2dcfad-8129-4fce-96a4-18d49de674f5 · outbound

This paper cites Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound.

ACE-Step: A Step Towards Music Generation Foundation Model Meta audiobox aesthetics: Unified automatic quality assessment for speech, music, and sound

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.253728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.253728Z digest=sha256:23dea959a01e539029306acbb3e4aba4be973be74e9b926ccc91574ed6cbecc7

Observation 62638e80-0a3d-4f37-8f4c-bc087b8beeee · outbound

This paper cites Qwen2.5-Omni Technical Report.

ACE-Step: A Step Towards Music Generation Foundation Model Qwen2.5-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.352668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.352668Z digest=sha256:6412a8d883543f7a7bb9359a50557bc36c32423af5390cafbe889e2a13413386

Observation a724f43c-4e91-4834-95d8-45eda269d2b9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision, 2022.

ACE-Step: A Step Towards Music Generation Foundation Model Robust speech recognition via large-scale weak supervision, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:17.424130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:17.424130Z digest=sha256:545fc3b4f247afd8456a639e1b1be84cf68c2f92097cd376d6dfb962afd9fab0

Observation bfc12106-f8f9-4142-96a3-2b41058488de · outbound

This paper cites ekzhu/datasketch: v1.6.5, May 2024.

ACE-Step: A Step Towards Music Generation Foundation Model ekzhu/datasketch: v1.6.5, May 2024

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.578319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:17.512771Z digest=sha256:a65b358062c56231977f640614176cc21e1340c4026dcb0425100fb105bb9c76

Observation 35c29adf-ab98-49a5-b5b2-aa40a98aee3c · outbound

This paper cites Byt5 model for massively multilingual grapheme-to-phoneme conversion.

ACE-Step: A Step Towards Music Generation Foundation Model Byt5 model for massively multilingual grapheme-to-phoneme conversion

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.373421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:17.590060Z digest=sha256:d2c973b0139bdfe6d6eedac973139181002e1652c510824a9b1120af50c32d72

Observation b98c5d1f-f36e-4a92-99c6-a50be36cd851 · outbound

This paper cites All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio.

ACE-Step: A Step Towards Music Generation Foundation Model All-in-one metrical and functional structure analysis with neighborhood attentions on demixed audio

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:21.189976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:17.671225Z digest=sha256:3abfc1c6bcf868683e2724bf5c6dee54e75ca9e55b383fcda6dbd08f5f61a348

Observation 68373a3d-a027-485c-9967-8950194a1811 · outbound

This paper cites Beat this! accurate beat tracking without DBN postpro- cessing.

ACE-Step: A Step Towards Music Generation Foundation Model Beat this! accurate beat tracking without DBN postpro- cessing

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.948160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:17.799489Z digest=sha256:adcb37f57b2fdd45bf6a66b1497bc4b90b419ce853e690d1221b8726e37564b8

Observation cbdea552-a5c0-4afe-81ec-8c671937469c · outbound

This paper cites Essentia: An audio analysis library for music information retrieval.

ACE-Step: A Step Towards Music Generation Foundation Model Essentia: An audio analysis library for music information retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.776114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:17.920550Z digest=sha256:03d6924f5bbfc1d8129714569f647ced9d5194c5b3550d3729ed5c006a370fcc

Observation 711bf726-379a-44c7-a7b0-2038a34f9ceb · outbound

This paper cites Xtts: a massively multilingual zero-shot text-to-speech model, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Xtts: a massively multilingual zero-shot text-to-speech model, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.580518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.015834Z digest=sha256:94d915f1de9942edf72770e8b4e01230406622a193deaa8c3501a45784dfa7a1

Observation 876fc305-bd3e-4431-9b59-e4b551bf33c4 · outbound

This paper cites Fish-speech: Leveraging large language models for advanced multilingual text-to-speech synthesis, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Fish-speech: Leveraging large language models for advanced multilingual text-to-speech synthesis, 2024

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.426135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.106129Z digest=sha256:043428a38538e92c575e652f48a472f2b4024cd22db4a6f1643cd40a0c8bc286

Observation e626520f-3c7f-4b8d-88fa-fefb94a03737 · outbound

This paper cites Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis, 2023

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.231538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.231538Z digest=sha256:0eb0b36123c230b108433a26023554687752ab5ed1a90cd6f3f3989346b19e4e

Observation 9879f042-94ca-47cf-899d-307e3e426a08 · outbound

This paper cites Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023.

ACE-Step: A Step Towards Music Generation Foundation Model Unimax: Fairer and more effective language sampling for large-scale multilingual pretraining, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.277393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.345992Z digest=sha256:6c4649e28e091961c33d4d18f25b03fd4abd946498279fac8120ed55adba9e7a

Observation 3f70d972-b2da-4ae0-9ce9-9a8f7eae6228 · outbound

This paper cites Learning diverse features with part-level resolution for person re-identification, 2020.

ACE-Step: A Step Towards Music Generation Foundation Model Learning diverse features with part-level resolution for person re-identification, 2020

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:20.133842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.495011Z digest=sha256:f4b693c613282949ec7300414b07d78d8a41e3517c5034b3d84df76fbbd017e1

Observation 4ef08c98-252d-4b9d-8399-48c7620debb5 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.572288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.572288Z digest=sha256:0069ec41df9e4265f25500cb3ad1058c9491af7be979653233bcee028ec82d1f

Observation 9e9a9e3c-20c2-44da-b4ca-9d0df6151958 · outbound

This paper cites Adapting frechet audio distance for generative music evaluation.

ACE-Step: A Step Towards Music Generation Foundation Model Adapting frechet audio distance for generative music evaluation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.979871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.701310Z digest=sha256:8599d6fd4e5ab447c375236ed5b629157a615f0db3dfadfce68b12c552b7914d

Observation 837ce3fc-4811-4b2a-81eb-be22b0b2d8ac · outbound

This paper cites Stable Audio Metrics, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model Stable Audio Metrics, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.788623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.830656Z digest=sha256:2aa2c430a317f8abc866d6aafdf77ebcc718a68f6e82d91cfffcf5d5e93072d1

Observation a56da72c-0086-41a0-85d5-db3e88f850a4 · outbound

This paper cites MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization.

ACE-Step: A Step Towards Music Generation Foundation Model MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:18.929184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:18.929184Z digest=sha256:1d4f8ff2855aa15c409c186d37150cd04dc94f532e3a1af4d5a42de867dd0029

Observation f76fba90-df7f-43ae-af5a-a196407df3fe · outbound

This paper cites stable-ts: Stabilizing Timestamps for Whisper, 2024.

ACE-Step: A Step Towards Music Generation Foundation Model stable-ts: Stabilizing Timestamps for Whisper, 2024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:16:19.581782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:16:18.996979Z digest=sha256:0a95d7b9944a49607a83b11523851d038d871ec66417e37fa273f8baa619c589

Observation 73391fcb-d9e8-4f78-a451-9e57ef435084 · outbound

This paper cites SongEval: A Benchmark Dataset for Song Aesthetics Evaluation.

ACE-Step: A Step Towards Music Generation Foundation Model SongEval: A Benchmark Dataset for Song Aesthetics Evaluation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.098611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.098611Z digest=sha256:cefe3fe2728518e6eac1423eb3b08596cd5f28e8c321448e6405e662c09013e1

Observation b5ccf2bb-3846-43b3-b876-238fa7d2be45 · outbound

This paper cites Simplifying, stabilizing and scaling continuous-time consistency models, 2025.

ACE-Step: A Step Towards Music Generation Foundation Model Simplifying, stabilizing and scaling continuous-time consistency models, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.207061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.207061Z digest=sha256:8dd229599860b8f86c2a018f809363d30f16c9ebe0c7a2a292301cf593db6063

Observation 1a7b8550-da2e-4817-86e1-545e4bad6098 · outbound

This paper cites FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models.

ACE-Step: A Step Towards Music Generation Foundation Model FlowEdit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:19.338517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:19.338517Z digest=sha256:cc52afc8352ec93b20adfaee087f630b036ff5f3869388670ac1b99bd12bdc94

Pith citing papers

Observation 607794e3-305a-4413-8fa2-5d90040a27f4 · inbound

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment cites this paper.

JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:14:37.507871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:14:37.507871Z digest=sha256:1c274a68659c2a8e1c0b6143fe0d69e61142676f8f3124a74e3aafefb857f3a5

Observation c7bb091b-edf2-421f-808d-1cf2109f21c1 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.880406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.880406Z digest=sha256:c96d56280323bf8d947840f1e9ce42d1dfad2829bf42bebbb8dd34b7dd03437d

Observation 0a61d2d4-36d4-439c-83c2-634dc4032319 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.692143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.692143Z digest=sha256:1c05065fa915aff10019fb46192faa3ac59df48424f98b5c6ad41e8eab723687

Observation 5110b3af-f508-44e5-b08f-763b95cffc20 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts ACE-Step: A Step Towards Music Generation Foundation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.596423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.596423Z digest=sha256:bdffdd44aa91322696cad54a5260dc5a11810710f5567ca01c1cd84c7bb8f15c

Observation af8bd234-643b-4344-aedd-ca8b67093cd6 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:03:15.219965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:3c817c732976be45d9198bc9e6adc064927a9189293562e503313b7deaa2b315

Observation b6b2c2c2-c7a3-481f-8e9c-b03a6dd75d6a · inbound

SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision cites this paper.

SongFormer: Scaling Music Structure Analysis with Heterogeneous Supervision ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T10:46:16.858659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:45:09.089710Z digest=sha256:a9d6d969ad45c16a60ad647dc1c7c0e70209f5b143f20b71c1df4035e4881464

Observation 7ae5821c-dec4-4de1-bb54-c93df909c2ee · inbound

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing cites this paper.

JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing ACE-Step: A Step Towards Music Generation Foundation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:25:55.887395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:25:55.887395Z digest=sha256:5eb48598e529515653eec451b4f22f4edf75e4f7c2be2c5e21b3f88cbee24efc

Observation 9968680e-dce0-4bf3-84ac-c4ce297dd30f · inbound

TADA! Tuning Audio Diffusion Models through Activation Steering cites this paper.

TADA! Tuning Audio Diffusion Models through Activation Steering ACE-Step: A Step Towards Music Generation Foundation Model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:44:11.551922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T13:40:47.112814Z digest=sha256:3dfd97bcd6f2ad9f231629f4f0b11aec64812ff923c3f2c36d3339782f154706

Observation 7c6cabd7-a8aa-43c8-be3d-6880c211ae87 · inbound

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline cites this paper.

MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline ACE-Step: A Step Towards Music Generation Foundation Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:20:17.918821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:17:45.920234Z digest=sha256:c68b1d4738d5ed8208e42f7cd3f2dffe6d8fa606df8491dfc899691a4381b837

Observation 310031d3-f4a0-4f8d-9b1a-2c3d248d3cb6 · inbound

Echoes: A semantically-aligned music deepfake detection dataset cites this paper.

Echoes: A semantically-aligned music deepfake detection dataset ACE-Step: A Step Towards Music Generation Foundation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T17:35:47.688979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:35:47.688979Z digest=sha256:085e152eb64dd8b12f152151a8a858f4d74ac7ce99449f283388e3689c04a853

Observation ecbcf569-4432-4a67-8b22-40bd3d323ec8 · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:01.532365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:08:58.648967Z digest=sha256:2104d9b892c5cf748d01868b0296de581022f3ac2ee8488ae41a1fce4fedc8cb

Observation 13a0087d-9539-4c4a-8519-843e847b4df9 · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:21.119692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:21.119692Z digest=sha256:db8001ad6d378d7a047991b562b2c747e4664a5586188ebfa1740afa7ccd67ad

Observation d0263534-865a-4bae-b7db-23f713ff310f · inbound

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment cites this paper.

SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:24:22.334741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T10:14:41.499130Z digest=sha256:22b68cd0999032a6c6d40fe8382fc228b7d730aff9bbb66fe90a2eacef9819fd

Observation 86ef11ed-3261-4695-a0b5-79949d4777c1 · inbound

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation cites this paper.

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:06.108044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T16:11:16.551910Z digest=sha256:d7e2837702f15cf2cbeef00b7963e6515e739f5a521e04a0bd1564772d78166b

Observation be05e593-f1f9-4e55-abdc-b9638dcd151c · inbound

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music cites this paper.

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music ACE-Step: A Step Towards Music Generation Foundation Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:56:29.750313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T13:27:41.439049Z digest=sha256:2339fcebd194b7a6fb130277dd2be585199463207d6825b95d7d27641f316a7f

Observation 36991cea-e9fa-440d-b263-15473293870c · inbound

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music cites this paper.

APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music ACE-Step: A Step Towards Music Generation Foundation Model

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T00:45:12.341833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T00:35:42.697038Z digest=sha256:55638f6210f776fa0061b389e2fe975473142b5a82d92de796fcc8ee44e92b2f

Observation 8cef5aeb-d6b0-4e75-aee4-3ee51dfd1d14 · inbound

Cutting rules in strong field QED with application to trident pair production cites this paper.

Cutting rules in strong field QED with application to trident pair production ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T16:57:23.214654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:57:23.214654Z digest=sha256:e8a9baabe6dcf76743e2a2f9885ecad5c413460bb7c65e9f5e56adae5b71992f

Observation 12dd018e-0aad-4d75-9e9d-1619b69a25eb · inbound

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation cites this paper.

OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:17:22.920096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:16:20.612291Z digest=sha256:b9b7c82e15ac028fe64a0b6b76da42f878ed37b54cb36f3892b97b494fb383b5

Observation 788b2689-12cb-414e-8b9a-0b20e04cbf89 · inbound

S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation cites this paper.

S2Accompanist: A Semantic-Aware and Structure-Guided Diffusion Model for Music Accompaniment Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T22:52:50.459830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T22:50:08.321514Z digest=sha256:d64c5e62307e55925b6db72d92ca480c28d8836b3f13e99d0a6b054de9b8bc72

Observation ebc97580-98ed-49a8-a723-6ca14b67c2c3 · inbound

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches cites this paper.

Instrumental Text-to-Music Generation with Auxiliary Conditioning Branches ACE-Step: A Step Towards Music Generation Foundation Model

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T02:43:54.924541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T02:43:16.101544Z digest=sha256:f3596ec95e12bc6403311ca266d35e81c29ac2686fe7398a864c6af33110bd93

Observation 9d0fd53c-60e0-46d6-8c91-6b165351c0aa · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment ACE-Step: A Step Towards Music Generation Foundation Model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.662750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:3b11d4e79ffe05ff199f5f58328abda04beff842402e8705edb96d67f49193e2

Observation d47d9dbf-9398-49a1-b09c-e73fb9ac7a28 · inbound

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation cites this paper.

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:57:19.597322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:14:08.243894Z digest=sha256:afe22237fd5c6110b126430158821066aa6af7c67f6a32624d5395b79f93ccee

Observation e3b93717-bbc5-4583-8be1-24859dd65462 · inbound

An Empirical Analysis of AI Slop in Music Streaming cites this paper.

An Empirical Analysis of AI Slop in Music Streaming ACE-Step: A Step Towards Music Generation Foundation Model

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:48:59.861457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T00:08:36.279506Z digest=sha256:a220b084c0d3ff19c7f32e9504f543cc9744f4224862c3cf980448bc90f93478

Observation ddcc77e6-6770-4132-9653-1005ed10afa5 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training ACE-Step: A Step Towards Music Generation Foundation Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:15:47.551038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:de4a549ba3161fb8bb40b8f69c555f028a2b7195482cca7849b838085cc57065

Observation 1db8f3b9-87af-4035-ab09-b0e72081ee04 · inbound

MusicMark: A Robust Generative Watermarking Framework for Music Generation cites this paper.

MusicMark: A Robust Generative Watermarking Framework for Music Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T06:56:53.384109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:56:53.384109Z digest=sha256:035065eee80395ee6a8d0a88d336303a244e78c857314ade846cfc0bfe103225

Observation f2752145-d739-4f3b-abbc-59bf5cc42406 · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T06:45:43.330341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:45:43.330341Z digest=sha256:aad67ea9edff9fda3aa56b2d7e7a015e5f5258d46da190c229e7e3007e7857bd

Observation 91482bf6-aca7-4c5f-9ea9-c288237debd6 · inbound

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance cites this paper.

Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:05:04.510458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:05:04.510458Z digest=sha256:d4a7b3f3b85fe9ed569d64b057d5939e4aaf58a4576c8b0362de3cea5b9c296a

Observation 27e079cd-46e0-4357-a4b9-92661222b514 · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report ACE-Step: A Step Towards Music Generation Foundation Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T03:47:22.936776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:47:22.936776Z digest=sha256:7fdbc84b449e94984b1626c170dac5de62c7594df673843adc6a983617b12018

Observation fe00cef8-e6c0-4493-ab55-6566a33784d9 · inbound

Qwen-Music Technical Report cites this paper.

Qwen-Music Technical Report ACE-Step: A Step Towards Music Generation Foundation Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:52:26.903347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:52:26.903347Z digest=sha256:8d5a26babab779d4b0de7c93796c039f51eebfb29cc2bca7db4588ac306115a0

Observation db7d261c-a992-4881-837a-d860a877bf17 · inbound

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation cites this paper.

Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T03:25:50.281276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:25:50.281276Z digest=sha256:fba0b089a5238a71d70a6b0cb42dfdaa78e7d8b550c2b7c7b09c09851763c602

Observation d949e0fb-134b-4f84-8d6c-0bedcbd70441 · inbound

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features cites this paper.

A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features ACE-Step: A Step Towards Music Generation Foundation Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T12:01:57.134742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:01:57.134742Z digest=sha256:44ad110d2b59314b5c59206cd3a49281a9794a862ebcbb0fca806a17eb5d20fe

Observation c2a48299-6730-44a6-baaf-3ad500282caf · inbound

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering cites this paper.

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering ACE-Step: A Step Towards Music Generation Foundation Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T10:24:19.082295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:24:19.082295Z digest=sha256:7bc5506f449abb1033600573b859e0531c86df937c3a0099c5af2fbe5ecb975c

Observation 17568aa2-93fd-49dc-b420-c0f47a369205 · inbound

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation cites this paper.

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation ACE-Step: A Step Towards Music Generation Foundation Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-30T23:40:07.801267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:40:07.801267Z digest=sha256:b0998111092461336ceb68c7dfd1729c83f7a117990cd26739b1246a444cc453

Observation 37152fb8-a22d-4c0f-b544-e5e048dca679 · inbound

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models cites this paper.

TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models ACE-Step: A Step Towards Music Generation Foundation Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T01:28:28.022775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:28:28.022775Z digest=sha256:739fbdcecdc4ba687c384252d39ce9b67a033528741d522670e366620642ed4a

Observation 4942d2d7-e8a2-4feb-bb5a-321663b226e2 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:20.556905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:20.556905Z digest=sha256:054a8904cdb28f35bb6a23a2cc4e672549ce045b3adb78838adf57898a8fd8ad

Observation 0518d6f5-8d75-48c3-b78a-d633529e7370 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ACE-Step: A Step Towards Music Generation Foundation Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:43.128810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:43.128810Z digest=sha256:ee88cc201bc11d974bea81d5d3f7dc2b7145160eb06ce6cfc633bddbc968b4f5