Pith. sign in

Paper Citation Record · LEDGER

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

As of 9 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2507.11096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11096 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.637571Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.422666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.492700Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy36
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1836cb7-56b4-40b5-8a95-1dd7398a3814 · outbound

This paper cites EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.371790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.410643Z digest=sha256:94348977083201cee38319a9c1616496c4820d3bb3d03a6af595d50288c3cada

Observation 39fb741f-b162-410a-847e-7ee8294c203f · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.362693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.418748Z digest=sha256:a96f4188f94722a55ed2ed42be3b2ad131ae5760156cd0d3ea6d319a5981734a

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · outbound

This paper cites EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2667cc41be133a822258569ba3062d1ca92c1fa27f12cef2feca2b98b0cb0105

Observation 62c952b9-1fb4-4883-81ca-5f8d0c33861e · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.352291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.427018Z digest=sha256:b10442c34aee38f138cf45567d8593c43ec6a526a9baa3861097359005c2b65c

Observation cd8834d9-09b7-4c01-aebb-3e43448982ef · outbound

This paper cites Yang et al.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Yang et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.342500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.438096Z digest=sha256:3599883fdf4c53d3cf0cb1b90c530f888fc099a7e81dffeae5d1d4ede8f10cb0

Observation 14ee1a2e-467a-4ef6-b21b-050c29796121 · outbound

This paper cites We began with Auffusion, leveraging its existing capabilities for prompt- based editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing We began with Auffusion, leveraging its existing capabilities for prompt- based editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.269463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.468785Z digest=sha256:c7e397ff6ddff03e8a10979b312a0df7967cf8ac3a9f81441462950affec0510

Observation 74e85710-6cef-4bc2-8fa4-66c6473c4f52 · outbound

This paper cites Masked autoencoders that listen,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Masked autoencoders that listen,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.197286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.499750Z digest=sha256:ab9379165c70d4c288452be241bc72fe4ba3c38a8052fd8a34d4f5b3a83f19ec

Observation 8f6c65de-b84c-44fb-be5b-9f5ef05bc44e · outbound

This paper cites The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs).

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.309637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.455289Z digest=sha256:004a39dbe156f6d70e69d78490519fa48206e6a6a63b065db30634dca7cc4d10

Observation d2e3bb99-8b24-4daa-9251-504a96eea1a2 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.289769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.459793Z digest=sha256:f1d125486b9840efec2bcce360d89293ca63fd82d26272042365e9633d6aa613

Observation 7ffbaf05-ae58-446b-8e7f-cfb9a91a1885 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.279470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.463949Z digest=sha256:baea9c16f4c581f8ef7f1c27a1d01f8a10eeb96b6e659ba125bcd8cdf4a74a6a

Observation 5ea41673-d7c3-4a12-95d4-bf4abbf531ea · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.175498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.520183Z digest=sha256:1c78b2697f457daddac8539933228954f264cf441f633bcbe87b7eeb2dff3f7b

Observation dfbc75aa-9c63-4334-925d-661392fa34f9 · outbound

This paper cites Prompt-to- prompt image editing with cross-attention con- trol,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Prompt-to- prompt image editing with cross-attention con- trol,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.259659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.473034Z digest=sha256:b4a3b590d32d510c8c57f5b7f31995e99102ef3d035a77949f4830b74f90bf59

Observation 51391606-a24e-44bd-a28a-3088e7a2d9f7 · outbound

This paper cites Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.333086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.444796Z digest=sha256:03d1bd30ba62114b0417324a6942799b088aa5a32719bd7f319da3e82d8c0ba0

Observation 13c1e6ef-09d2-4070-974d-b5c4e9ff79cb · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.249345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.477277Z digest=sha256:01ec8fa9c49c8ac531ba3c4c5387019010676d10b2286f8183495ca02f3251d8

Observation d163a200-2080-4189-81c4-ec27564fde37 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.239226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.481096Z digest=sha256:cb0ad796a9e7bc9abdee00ada364cecabc7664a19cd6990d58161ea4748edfb5

Observation 0c747d61-fac2-41a2-9f55-67d3490b89d9 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.228419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.484451Z digest=sha256:f25f8169cc260091e683de2c6dc70c3865de995880d7a1a17532b9e413b6c956

Observation 05a24ad4-193e-4e78-a85c-39471fa3a39e · outbound

This paper cites Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.323341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.450753Z digest=sha256:aa869b221101803966add552c9a5b86a56b11ca2bb8d1400756e2833ce757fe7

Observation 094ae7f7-c06b-4836-970a-c16f3a7d3432 · outbound

This paper cites AudioLDM: Text- to-audio generation with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM: Text- to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.213796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.491844Z digest=sha256:79e0d841fea621f810a056418ca11097be3fbf7c037bbc8ba99b4f389104c28a

Observation 641e40c4-6c03-4697-812a-0c7803250b6b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.495779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.495779Z digest=sha256:6041a5cfb4e9f09729f5406e1821ee9cf95d568b1458331bfa107a27954ce332

Observation 6f653922-2083-43b9-a3de-f852195c57c6 · outbound

This paper cites AUDIT: Audio editing by following in- structions with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AUDIT: Audio editing by following in- structions with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.186979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.506419Z digest=sha256:06843f9c323869ff3a7099d1a10af6706777aac265053445d05ac8eecc6db101

Observation b58c729e-353c-4979-b850-2ae52784f25a · outbound

This paper cites Text-to-audio generation using instruction guided latent diffusion model,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Text-to-audio generation using instruction guided latent diffusion model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.510770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.510770Z digest=sha256:f114bc86a1c3cca9fa953cf20326df5623b6004c0e1e0542d6162e6b6890a909

Observation cb1355dd-1f1a-4735-9bf3-4289fa82414b · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.516295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.516295Z digest=sha256:460de44db2a1ec9e9795e86f83af474beda78785440ded5164f9a71f3939b5d1

Observation 41414649-1fb4-454d-9aa2-7b702a3ff99c · outbound

This paper cites MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.165442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.524168Z digest=sha256:97103b1a308673ea0adc30c52bd18e65ae382844372e7e4193da050a12bdb514

Observation 96f0f70d-1ed0-4b70-9f00-2f55e8cd2515 · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.780086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.527735Z digest=sha256:95d5cc22fc1f0e13a50d0f13d26ce6928c374ddb3be6955420aae184e49845a5

Observation 27980d42-36ee-4c10-96a4-2befd82c6f23 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.531204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.531204Z digest=sha256:21ef18b64102a0da208e2de6fc04526c3be2bd8301bf92b0c2193b9c26d3693e

Observation 1b32fa68-7f95-4f25-8b3f-9db7c204b828 · outbound

This paper cites Audiogen: Textually guided audio genera- tion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audiogen: Textually guided audio genera- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.154915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.535045Z digest=sha256:5919f54295fde1f979a96ac67f8770ac0ebe9892e31ec081e33325bd7e96b627

Observation e8084bfe-9e05-4bb0-86ea-71ba8fef237c · outbound

This paper cites Jukebox: A Generative Model for Music.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Jukebox: A Generative Model for Music

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.540076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.540076Z digest=sha256:458494d11c3cca381d9fd8faf0e3e34ea912f14af92ec12bb16fbe27118c533c

Observation 05075a94-0224-4dae-8a53-2acfa6dd4680 · outbound

This paper cites Neural discrete representation learning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Neural discrete representation learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.140555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.543739Z digest=sha256:2a899b39d91e408b422a0fef29fe0e7853bf87c4c7bd744ec29346feb0592bb4

Observation 0f0f49c6-7193-4135-8835-6f6a1965dd3e · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLM: a Language Modeling Approach to Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.547479Z digest=sha256:4e7e3475336fad8658f465a4fb719858eed0beb557a7b1d65ea913cbeaf0d9b8

Observation 757c56d3-8f1f-4fe2-a449-14bc158deee4 · outbound

This paper cites Soundstream: An end-to-end neu- ral audio codec,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Soundstream: An end-to-end neu- ral audio codec,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.123953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.551187Z digest=sha256:9d8338129503f92fc046562be5df087c05581aece8582e8ecf13e3f13e9554fd

Observation 430323d8-6cab-48ee-aeb1-1e38f82c512c · outbound

This paper cites End-to-end optimized speech cod- ing with deep neural networks,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing End-to-end optimized speech cod- ing with deep neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.107018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.554641Z digest=sha256:99af6342ade13fbdd40f6cd3ab370a1e257f626abf09ebde05cb89fd8f162a5b

Observation df02e93d-76c6-4a72-a3b7-d9874ff8e559 · outbound

This paper cites Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.096933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.558163Z digest=sha256:c9849fe519f181fe32de6583a6b4ac75c6088b682f608b9046314a0081afe138

Observation 7dcc803c-55e9-4d6b-9ea3-99ce39c413ab · outbound

This paper cites MusicLM: Generating Music From Text.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLM: Generating Music From Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.562352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.562352Z digest=sha256:1a69e3de54a5a3a37951023742ecf64f6b0fc544cd8be4dbc2b3b910e1d36ee6

Observation 8e362c39-afec-48c0-98cf-0b86c0f3a3d8 · outbound

This paper cites Simple and control- lable music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Simple and control- lable music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.085641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.566201Z digest=sha256:88ea0a2b34a5d51ca3f904662827ea01ebbebfae16fddb5c30670d82ca23d050

Observation 45e2670d-ef47-403c-b9a5-c0e5e8c579ad · outbound

This paper cites High fidelity neural audio compression,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing High fidelity neural audio compression,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.569796Z digest=sha256:98558e2dc395f2b135738ed080d64dea0e09c16be07219c4ecc96b7f487cc214

Observation efb8e301-1609-4cef-8345-274ad718c106 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Null-text inversion for editing real im- ages using guided diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.054968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.573681Z digest=sha256:e9245bdd4e2053b80624580f4f1865d30c194ba9f13c0c5cc402633d34923a9f

Observation baaa9fc4-4c3f-49cc-ba3c-acb8cd9b5641 · outbound

This paper cites Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.041200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.578547Z digest=sha256:534c7043c38f0713a8dc95be4261c2278f25b05ecabb0dec1562d366fee9061d

Observation 769f635d-1cac-4d9e-8b60-21c208ca9497 · outbound

This paper cites Multi-concept customization of text-to-image diffusion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Multi-concept customization of text-to-image diffusion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.010492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.581943Z digest=sha256:bc1ebbab2d13655d3d2e8fbd5e87a01bb98adb4396fc1dde087247afcd55d2b0

Observation 3235656a-5e43-4ab0-b6fa-d678989ffe1c · outbound

This paper cites SVDiff: Compact parameter space for dif- fusion fine-tuning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing SVDiff: Compact parameter space for dif- fusion fine-tuning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.992132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.585641Z digest=sha256:0756cabea48544199e8bbd541856192fcf0d7bbe56e9ca112cc0d9cf779dbef3

Observation 7f52f4b6-2a27-4684-9c18-f4a4979c02b9 · outbound

This paper cites Countering language drift via visual grounding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Countering language drift via visual grounding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.979777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.591831Z digest=sha256:abba01db65cfa4457e9062e50253d8770779cb61fa4dbf12e6238e806fa6ca48

Observation e7739275-b8fb-4c94-bd53-a58cdba6eccd · outbound

This paper cites Likert scale: Explored and explained,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Likert scale: Explored and explained,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.899844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.637571Z digest=sha256:1efd77077c77357b812592206edfad894a16b0a610d08f2a3a77811265e03e62

Observation 6f6943c1-c394-47c9-a915-fc26e93bb154 · outbound

This paper cites Investigating personaliza- tion methods in text to music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Investigating personaliza- tion methods in text to music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.958897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.599929Z digest=sha256:418e71ab1ff68db1a48064f2ae51bc42583c0035df27e54e70f7a2876c904366

Observation e13327b7-eb7e-4d5a-8ffa-013c22004306 · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.603502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.603502Z digest=sha256:404c6833510c7429f4d381975af49d48a640654a286e30e6b2655dc2a31c76a5

Observation 79b0cdd2-eeb8-4192-88d7-48621f19d2f2 · outbound

This paper cites An Edit Friendly DDPM Noise Space: Inversion and Manipulations.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An Edit Friendly DDPM Noise Space: Inversion and Manipulations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.607581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.607581Z digest=sha256:69d852d0c26957aa68d2a14b247fc0a4ae13f03feefd1f608202b94cf4eb8bdc

Observation e87ace66-a28a-4ddf-a7f4-86419d50b3ce · outbound

This paper cites Photorealistic text-to-image diffu- sion models with deep language understanding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Photorealistic text-to-image diffu- sion models with deep language understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.948459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.611446Z digest=sha256:8c88d2b2a422a7c47ead818ff56f479b625a273b9915c17d17e0a7e962b98af5

Observation 41af34df-31d2-47da-984d-967b6956125f · outbound

This paper cites Music ControlNet: Multiple Time-varying Controls for Music Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Music ControlNet: Multiple Time-varying Controls for Music Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.697050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.615013Z digest=sha256:e8d4f36bdc4249507780da432b3c15fed59ed58de1c858f7cef320f9135aee29

Observation d1855218-0518-461c-ba1a-da8cc5f7a7b3 · outbound

This paper cites Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T17:21:31.671027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.618387Z digest=sha256:54bc618f7e4e78db86c1561fd455e89debdd08e1f2a3fb1556aa22e820831b04

Observation 1e2e7b21-c867-484f-a509-356a3ea75b37 · outbound

This paper cites mir_eval: A transparent implementation of common mir metrics,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing mir_eval: A transparent implementation of common mir metrics,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.938083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.621895Z digest=sha256:33054837bc1048711575b451ad7f6303972d118a0eb5eed955125ab610d9dac6

Observation 149c2df2-8efe-4795-8e0c-12f93b27d828 · outbound

This paper cites An efficient state- space model for joint tempo and meter tracking.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An efficient state- space model for joint tempo and meter tracking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.625142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.625142Z digest=sha256:30f6d6882bcdbede888ee4672b85b55d0cd3ae38f7ea66d028406adca4a207b9

Observation 7928f099-baf0-4901-8ba7-de10090931ff · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.921391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.628136Z digest=sha256:6b8aa537f4e9b7b7ff0ee71054738ffee6e9cf82b07c3d263e7e53a20a040d22

Observation 983330f6-1d59-42aa-94ef-5bd46daeb1dc · outbound

This paper cites HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.911114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.631335Z digest=sha256:16dfedb9951e8ca4112886e2ae4915aa98f6488a0d43f2ccdc940d648b55f7a2

Observation 8c76b7c9-5998-4a40-a9ba-417eed918ec9 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Representation Learning with Contrastive Predictive Coding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.634601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.634601Z digest=sha256:eaaf6e821b9ebf9e6a5744b0fdb52d3bf1b76fc3c4aad22de50c922d727f380c

Observation 9f26091d-4751-422e-9db3-f3d6d8799954 · outbound

This paper cites Available: https://aclanthology.org/ D19-1447.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Available: https://aclanthology.org/ D19-1447

Reference 4395

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.969764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:21:31.596119Z digest=sha256:c1a0251c0dca2cbdc8382bd3e9daf69d083a8ef08fed77dcd0d784266e6865ea

Pith citing papers

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · inbound

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing cites this paper.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2667cc41be133a822258569ba3062d1ca92c1fa27f12cef2feca2b98b0cb0105

Observation 0615c238-6ed3-4616-b1d8-e7d03f23a4a6 · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.494215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:21d179e2c3b23445f7657fd64cfcbcc9a676dbe44eeff2b52df2c08fb0f84a53