Pith. sign in

Paper Citation Record · LEDGER

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

As of 15 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 2 inbound Pith citation observations for arXiv:2507.11096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.11096 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.637571Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:21:31.422666Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:39:38.492700Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy36
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1836cb7-56b4-40b5-8a95-1dd7398a3814 · outbound

This paper cites EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harness- ing Cross Attention Control for Instruction-Based auto-regressive Audio Editing

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.371790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.410643Z digest=sha256:e71599c7ee6feca5855c5b4ecaea36b5a025b2d492576078cc240233476766b4

Observation 39fb741f-b162-410a-847e-7ee8294c203f · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.362693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.418748Z digest=sha256:10789ad0460f701ad654fc72db63740c52e926614bf57e6f77bfa98edc295e0f

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · outbound

This paper cites EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2295887fdee14bb952f42c47226bf220f239d793eb26b317ddca16ae02fd2bc9

Observation 62c952b9-1fb4-4883-81ca-5f8d0c33861e · outbound

This paper cites an unresolved cited work.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:21:32.352291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.427018Z digest=sha256:f77a2e50d4041a4e474c1485a54cc5efe8a22b5d4870593be047bdbffb792ef3

Observation cd8834d9-09b7-4c01-aebb-3e43448982ef · outbound

This paper cites Yang et al.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Yang et al

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.342500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.438096Z digest=sha256:30511549acc1ed44cc523f6cbbd2426fa522f2b78112cd71a6b9bb89e8230516

Observation 14ee1a2e-467a-4ef6-b21b-050c29796121 · outbound

This paper cites We began with Auffusion, leveraging its existing capabilities for prompt- based editing.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing We began with Auffusion, leveraging its existing capabilities for prompt- based editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.269463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.468785Z digest=sha256:e683b0fa1dfaa899077461d3f37dd85fed4063be18bbcce7011fbdea45edcb82

Observation 74e85710-6cef-4bc2-8fa4-66c6473c4f52 · outbound

This paper cites Masked autoencoders that listen,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Masked autoencoders that listen,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.197286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.499750Z digest=sha256:78df7efff40b33aa68df07b17d2b7d2935c472608189e7b853ae1a12482daace

Observation 8f6c65de-b84c-44fb-be5b-9f5ef05bc44e · outbound

This paper cites The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs).

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing The input, a reference audio random variableX, is encoded into a continuous ten- sor with a lower frame rate ( fr) compared to the sample rate (fs)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.309637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.455289Z digest=sha256:ab93f5438cc550ec20f8c9b761b5bbc581ce342540321cef0deb2f7095985388

Observation d2e3bb99-8b24-4daa-9251-504a96eea1a2 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.289769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.459793Z digest=sha256:26ca9fab143dc82a268d76ac5e24bf4b2f6cd452f53735d5b1a65543b8dd0c74

Observation 7ffbaf05-ae58-446b-8e7f-cfb9a91a1885 · outbound

This paper cites prompt strength.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing prompt strength

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.279470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.463949Z digest=sha256:148cb75ff7b422d633dc97dad27f565dea4d67ab6a8ac01918afecb20ec41a42

Observation 5ea41673-d7c3-4a12-95d4-bf4abbf531ea · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.175498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.520183Z digest=sha256:9d8599196fee9425af9857baeeeb50b59f36f98fd1b373591bcb9dcb6a4b008e

Observation dfbc75aa-9c63-4334-925d-661392fa34f9 · outbound

This paper cites Prompt-to- prompt image editing with cross-attention con- trol,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Prompt-to- prompt image editing with cross-attention con- trol,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.259659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.473034Z digest=sha256:ea4976b1e31b2634669b1f9c9ffaa842a925236fb2b87ac286e7bc135d1d5501

Observation 51391606-a24e-44bd-a28a-3088e7a2d9f7 · outbound

This paper cites Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auto-regressive audio generation: As an alternative to diffusion-based models for audio and music generation, autoregressive models have shown promise in recent years

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.333086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.444796Z digest=sha256:5e0cdc8a43436f00391e3a9bc47fe4df503a64bb15f589cc5c0fe0f3a024ef82

Observation 13c1e6ef-09d2-4070-974d-b5c4e9ff79cb · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.249345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.477277Z digest=sha256:7d83e0933d3b074032529e53461c5cef9a3b207777c61cab417dbfd82f15c487

Observation d163a200-2080-4189-81c4-ec27564fde37 · outbound

This paper cites Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Make- an-audio: Text-to-audio generation with prompt- enhanced diffusion models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.239226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.481096Z digest=sha256:d8f763d644ebd0fd1f8e9b31146b4b5ea921a021a63b7835a49a1e3770d5c169

Observation 0c747d61-fac2-41a2-9f55-67d3490b89d9 · outbound

This paper cites CLAP: Learning audio concepts from natural lan- guage supervision,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing CLAP: Learning audio concepts from natural lan- guage supervision,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.228419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.484451Z digest=sha256:4278b0a51281e26ee94d523f04da8295c2221a47aa9cacee2fd999191243171c

Observation 05a24ad4-193e-4e78-a85c-39471fa3a39e · outbound

This paper cites Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audi- oLM [18] utilizes tokens generated by a SoundStream [19] neural codec [20, 21] as targets for a sequence modeling task

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.323341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.450753Z digest=sha256:66d98e9ec2f10fb29ef83e6a800cb47937bf711d12859380923f5ed8e0411975

Observation 094ae7f7-c06b-4836-970a-c16f3a7d3432 · outbound

This paper cites AudioLDM: Text- to-audio generation with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM: Text- to-audio generation with latent diffusion models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.213796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.491844Z digest=sha256:fddf6c50d04630c2ca4cf6d2dfb9b594280c85165294b71b3ef07e26c4ff4267

Observation 641e40c4-6c03-4697-812a-0c7803250b6b · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.495779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.495779Z digest=sha256:c608b53e6c31a9d278d8d7bda3c13c661cb6784131d58d387db72ca7bc2fe5d9

Observation 6f653922-2083-43b9-a3de-f852195c57c6 · outbound

This paper cites AUDIT: Audio editing by following in- structions with latent diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AUDIT: Audio editing by following in- structions with latent diffusion models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.186979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.506419Z digest=sha256:d0597c27d3cff0744fce61196c8ee245921f372127c18b5d226491adb541caba

Observation b58c729e-353c-4979-b850-2ae52784f25a · outbound

This paper cites Text-to-audio generation using instruction guided latent diffusion model,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Text-to-audio generation using instruction guided latent diffusion model,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.510770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.510770Z digest=sha256:58cf5f65b53c44e097bd6c8bb21f7406672ddac487a1b0bc7b2039c1da2bc50d

Observation cb1355dd-1f1a-4735-9bf3-4289fa82414b · outbound

This paper cites Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.516295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.516295Z digest=sha256:4f850c901a68ff65453e7b327f29771b4154fd5b6566dd7c03aaddbb6c0db99c

Observation 41414649-1fb4-454d-9aa2-7b702a3ff99c · outbound

This paper cites MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLDM: Enhanc- ing novelty in text-to-music generation using beat- synchronous mixup strategies,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.165442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.524168Z digest=sha256:961bb4dd2a17a173f690fe289c1b05b9ef2309adc84515de3b21a18a4e6ac10b

Observation 96f0f70d-1ed0-4b70-9f00-2f55e8cd2515 · outbound

This paper cites InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing InstructME: An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.780086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.527735Z digest=sha256:cee150ebb46e7a185cdc58ae2d4c2c44512d9ed9322d19fcb0f78552b8e4d3a2

Observation 27980d42-36ee-4c10-96a4-2befd82c6f23 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing WaveNet: A Generative Model for Raw Audio

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.531204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.531204Z digest=sha256:cb9203f0117b9e000bc8de02380e9594ebc0c76fc574f9672cb997f03676483f

Observation 1b32fa68-7f95-4f25-8b3f-9db7c204b828 · outbound

This paper cites Audiogen: Textually guided audio genera- tion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Audiogen: Textually guided audio genera- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.154915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.535045Z digest=sha256:b86c6c1638f292f3207705e14180ec4ad85550c1d0ac04d1a15674d83d0d94c2

Observation e8084bfe-9e05-4bb0-86ea-71ba8fef237c · outbound

This paper cites Jukebox: A Generative Model for Music.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Jukebox: A Generative Model for Music

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.540076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.540076Z digest=sha256:7d17f0eebe1289a9bfba5d7226c061c13b8309768e90f489db6fad7543bc55cb

Observation 05075a94-0224-4dae-8a53-2acfa6dd4680 · outbound

This paper cites Neural discrete representation learning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Neural discrete representation learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.140555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.543739Z digest=sha256:00a64bc7bac44c7fe33e17fc83ecb583124520953cd6ebee003248e0ae119c2a

Observation 0f0f49c6-7193-4135-8835-6f6a1965dd3e · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing AudioLM: a Language Modeling Approach to Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.547479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.547479Z digest=sha256:94329b76208670f1dd0ddefa72b4385d628c9affb7cd31a346f026dd7a05278b

Observation 757c56d3-8f1f-4fe2-a449-14bc158deee4 · outbound

This paper cites Soundstream: An end-to-end neu- ral audio codec,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Soundstream: An end-to-end neu- ral audio codec,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.123953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.551187Z digest=sha256:e8c6296f8068192c839db59fe7833f1d0075ad481d27d58bcbb82be942eacb6b

Observation 430323d8-6cab-48ee-aeb1-1e38f82c512c · outbound

This paper cites End-to-end optimized speech cod- ing with deep neural networks,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing End-to-end optimized speech cod- ing with deep neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.107018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.554641Z digest=sha256:d401bffcd4ccbe1770e8ba3b00d1a82f04a4e0160928229028d893019d0e9b5d

Observation df02e93d-76c6-4a72-a3b7-d9874ff8e559 · outbound

This paper cites Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Harp- net: Hyper-autoencoded reconstruction propagation for scalable neural audio coding,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.096933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.558163Z digest=sha256:4ad2be9acfba903eed10eab6567c2ede18df59263a397934430ff31b2652c825

Observation 7dcc803c-55e9-4d6b-9ea3-99ce39c413ab · outbound

This paper cites MusicLM: Generating Music From Text.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing MusicLM: Generating Music From Text

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.562352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.562352Z digest=sha256:519a8b2581d6bad79531837a694be01613a363e8a59ec599bf2bed501e1ee694

Observation 8e362c39-afec-48c0-98cf-0b86c0f3a3d8 · outbound

This paper cites Simple and control- lable music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Simple and control- lable music generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.085641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.566201Z digest=sha256:2542541131442c2110e210d10ba73f4df8aff76b0bddc82475998e4c40289974

Observation 45e2670d-ef47-403c-b9a5-c0e5e8c579ad · outbound

This paper cites High fidelity neural audio compression,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing High fidelity neural audio compression,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.569796Z digest=sha256:4ffe40105b8fd4424ae0919f3cd2540c2a6b1b13cebafe5ab8b000b7d379fbe5

Observation efb8e301-1609-4cef-8345-274ad718c106 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Null-text inversion for editing real im- ages using guided diffusion models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.054968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.573681Z digest=sha256:567ee4d556bd1a152c58aebc3b8c5ec9b85a4be7ea30650942f392c16e53724c

Observation baaa9fc4-4c3f-49cc-ba3c-acb8cd9b5641 · outbound

This paper cites Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Dreambooth: Fine tuning text-to- image diffusion models for subject-driven generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.041200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.578547Z digest=sha256:8cfb364a8604c8dcd920eeb31ef6324636f5065f741cbad78949a65b0d29711d

Observation 769f635d-1cac-4d9e-8b60-21c208ca9497 · outbound

This paper cites Multi-concept customization of text-to-image diffusion,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Multi-concept customization of text-to-image diffusion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:32.010492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.581943Z digest=sha256:061340ad2e4f30232a355a6cc6637ae3cbdcf14c83b2ec28b932c5881137f7e5

Observation 3235656a-5e43-4ab0-b6fa-d678989ffe1c · outbound

This paper cites SVDiff: Compact parameter space for dif- fusion fine-tuning,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing SVDiff: Compact parameter space for dif- fusion fine-tuning,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.992132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.585641Z digest=sha256:8d3fc6a5386b88ea6792f53caf9d431fae3474addce66714922de3819e5ceb8f

Observation 7f52f4b6-2a27-4684-9c18-f4a4979c02b9 · outbound

This paper cites Countering language drift via visual grounding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Countering language drift via visual grounding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.979777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.591831Z digest=sha256:3f10a4a4349f0b491c92e74010630cc7705e179129eb664d721c2d006a45427e

Observation e7739275-b8fb-4c94-bd53-a58cdba6eccd · outbound

This paper cites Likert scale: Explored and explained,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Likert scale: Explored and explained,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.899844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.637571Z digest=sha256:8a0d6076d74958d53490f6da5c0dc4982295888492ebf809a86d7ed8cb57f625

Observation 6f6943c1-c394-47c9-a915-fc26e93bb154 · outbound

This paper cites Investigating personaliza- tion methods in text to music generation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Investigating personaliza- tion methods in text to music generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.958897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.599929Z digest=sha256:3cd642419f1e38aa80fa5e9e2c12f07b61d9838eb8e0db395a1a43405aad44e7

Observation e13327b7-eb7e-4d5a-8ffa-013c22004306 · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.603502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.603502Z digest=sha256:d685e6ac7dbdebcdba139fc9a1c719072e88a55b74c539aa0062c2567167d9de

Observation 79b0cdd2-eeb8-4192-88d7-48621f19d2f2 · outbound

This paper cites An Edit Friendly DDPM Noise Space: Inversion and Manipulations.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An Edit Friendly DDPM Noise Space: Inversion and Manipulations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.607581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.607581Z digest=sha256:7588624293d3c1b0fc28fa6cdf71d29f3360f30bab4df36087d3beacecadf26a

Observation e87ace66-a28a-4ddf-a7f4-86419d50b3ce · outbound

This paper cites Photorealistic text-to-image diffu- sion models with deep language understanding,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Photorealistic text-to-image diffu- sion models with deep language understanding,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.948459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.611446Z digest=sha256:cae80fd5796a4d37b407c8c548ca3d47f99b544cb5a6c214c695ad7cc0df0945

Observation 41af34df-31d2-47da-984d-967b6956125f · outbound

This paper cites Music ControlNet: Multiple Time-varying Controls for Music Generation.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Music ControlNet: Multiple Time-varying Controls for Music Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:21:31.697050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.615013Z digest=sha256:7ed67f6fb38f31b5c50ba353813a707742d473548513c7c54706c412b9efb0d1

Observation d1855218-0518-461c-ba1a-da8cc5f7a7b3 · outbound

This paper cites Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Evaluation of Audio Beat Tracking and Music Tempo Extraction Algorithms,

Reference 47

Resolution
verified exact
doi, observed 2026-08-06T17:21:31.671027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.618387Z digest=sha256:5733534b90d821fb746afba6a5a4bcba6ec9866dde36a1e1524119fa53b37fce

Observation 1e2e7b21-c867-484f-a509-356a3ea75b37 · outbound

This paper cites mir_eval: A transparent implementation of common mir metrics,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing mir_eval: A transparent implementation of common mir metrics,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.938083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.621895Z digest=sha256:1c70cc018712a9227efcc5bca4668b40e798130d9caf8fe3d52892c161bdde42

Observation 149c2df2-8efe-4795-8e0c-12f93b27d828 · outbound

This paper cites An efficient state- space model for joint tempo and meter tracking.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing An efficient state- space model for joint tempo and meter tracking

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.625142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.625142Z digest=sha256:dd54eae4ac329571b2b2bac4f03bfbcfe0941f417a51bfa0458e45405cfbaf69

Observation 7928f099-baf0-4901-8ba7-de10090931ff · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.921391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.628136Z digest=sha256:df121ba7674932f0c0467e4b1cf1af1b8b3b7a139e3aa26994226de58de920d8

Observation 983330f6-1d59-42aa-94ef-5bd46daeb1dc · outbound

This paper cites HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing HTS-AT: A hierarchical token- semantic audio transformer for sound classification and detection,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.911114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.631335Z digest=sha256:758208436fe97821ff7b558a8a75d1c73e10793e0210c1d6d63f33e1c19f88cd

Observation 8c76b7c9-5998-4a40-a9ba-417eed918ec9 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Representation Learning with Contrastive Predictive Coding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.634601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.634601Z digest=sha256:2eb6431d759269e18fbacd5c9a6d7bb16fa0bf573fc0e79dd658b765f31504a0

Observation 9f26091d-4751-422e-9db3-f3d6d8799954 · outbound

This paper cites Available: https://aclanthology.org/ D19-1447.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing Available: https://aclanthology.org/ D19-1447

Reference 4395

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:21:31.969764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:21:31.596119Z digest=sha256:b403244a9fbdadd46da14d9dea6fbc062c7e01b7ec99cf255b0dae07bd608750

Pith citing papers

Observation 1af8fecf-7da0-47ff-a5d9-9cb2350948dd · inbound

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing cites this paper.

EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:21:31.422666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:21:31.422666Z digest=sha256:2295887fdee14bb952f42c47226bf220f239d793eb26b317ddca16ae02fd2bc9

Observation 0615c238-6ed3-4616-b1d8-e7d03f23a4a6 · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption EditGen: Harnessing Cross-Attention Control for Instruction-Based Auto-Regressive Audio Editing

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.494215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:5ad4c8b9a384cc9f0fae4d88bb5b5ac72f9b0422aae9ceef7c38fc0924a33940