Pith. sign in

Paper Citation Record · LEDGER

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

As of 21 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2508.20379.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20379 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:31.849774Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3920ab5b-1b92-456c-babc-9857e328da00 · outbound

This paper cites GPT-4 Technical Report.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.658133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.658133Z digest=sha256:bf8cf9811fc0d25c5a7009370bfc792a6bbc8771eafbfdb6bdd1e2bfe0232fd3

Observation adc71368-bcf9-480b-aff4-9589c8e14b6e · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2live: Text-driven layered image and video editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.469198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.662463Z digest=sha256:7c54cea5a7ddb6a711e2a268fc3f70781a42ba8831677df6d5f08c3c6ebe7945

Observation c1b863f0-6252-49cf-8419-27a579b8d80f · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.202014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.667945Z digest=sha256:809aa950ba5497dda9a8a9bd124ca01237f923bdbc566e484d0adcf960ca79fc

Observation 47e2c4cb-2364-43fa-8d26-fad3ea7f3196 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align your latents: High-resolution video synthesis with latent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.671833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.671833Z digest=sha256:a2204afe5ca814bea268578897512c42002068fe5b91798d27aa43607c51b402

Observation cb1a46e1-9ea0-42b4-9a49-88786c7f2087 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ledits++: Limitless image editing using text-to-image models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.454984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.674911Z digest=sha256:0a17dc5dc4166d631d6e51df5fe872b6579721bad8363bb5cf95f0c258a028f0

Observation 8038bc95-56ab-4c82-a49a-426e5953eb8f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.444636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.678420Z digest=sha256:1fa99a37929bd85876c703abcf6f1631b638b2c23277eb2e6448d50c0c5ba101

Observation 15255ef9-bc95-4305-ae2d-92fa964d1eab · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.682377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.682377Z digest=sha256:3a3592408f34ab8620c92ed96814b727e6184ff7055379bad78c62de4a6b53c7

Observation 3a9ba256-3b8c-407b-9a4f-47b4036a63ea · outbound

This paper cites Con- ditional generation of audio from video via foley analogies.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Con- ditional generation of audio from video via foley analogies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.434962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.685950Z digest=sha256:316675f9cd7ad2576aa79446d3c7f9b253970fe85f35f577926c6b352f4d27b9

Observation 70e339d7-4417-4a3b-8b9f-20a65d53e31f · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.425812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.689192Z digest=sha256:a3abbe02c20fb0687e0be9889f2102dd6f844604cc2ffff8328abd76391c96ae

Observation c1a0ee9a-b796-402c-a5d3-19025e5b4ed8 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.692936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.692936Z digest=sha256:095626ce4743da3f0a85fad7d4c71ef6e4776a2cf703969cca7ffd4f993f2aa2

Observation 017fc962-2b1e-4a23-b0a5-9fed696c05ea · outbound

This paper cites Imagebind: One embedding space to bind them all.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Imagebind: One embedding space to bind them all

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.416553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.696749Z digest=sha256:0fb3e761c615bdc839ba26071e1c4f3a119eb84e1532209cf04a3f920ba42a9d

Observation a38220d8-905a-45b9-a58f-de8655d3d5f0 · outbound

This paper cites AtomoVideo: High Fidelity Image-to-Video Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AtomoVideo: High Fidelity Image-to-Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.700664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.700664Z digest=sha256:34eee2ef887cbfcabd60c4df70ddb92f5b8558d01d0345e58c3e1c00f3660314

Observation b3099f05-64d3-4e41-bf32-662f496e67f9 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.704083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.704083Z digest=sha256:84990cb8ea27356c5a9658b9414da4c6cb52f0e8b261dc262aca0926d0683e80

Observation c1de15e2-4ec3-43b2-9a43-f8d8bd6fb07e · outbound

This paper cites FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.154458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.708990Z digest=sha256:56235a39212ed7c4289d63bcb9e59fbc1706a2891da2046a0ce2ac2f7a4a0a6d

Observation 6c938676-cf99-4b22-b3d5-29caf78618bd · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.713176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.713176Z digest=sha256:8f34c3b7eb7855f5132fc756ba777c40b64615d7ff7a4149083b4f1a22f866cb

Observation d9c1f186-073e-48e5-a586-4874717b2784 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.716262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.716262Z digest=sha256:477d0da24a5c5602d459584946e2d5b4db3e5b40cddcfbd0819c011c1aeab98e

Observation c4435c43-05ee-4b21-9678-726f78a5d31f · outbound

This paper cites Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.719815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.719815Z digest=sha256:734243813014f379ba0aede5bc7342ed04dcba4e177dc4ca6dfed2d3faf9257e

Observation 38c81fa2-f69a-4611-8626-a787974b7d05 · outbound

This paper cites Text2video-zero: Text-to- image diffusion models are zero-shot video generators.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2video-zero: Text-to- image diffusion models are zero-shot video generators

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.407619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.722751Z digest=sha256:ed6352701c44cb9b7090fef800272b12ff4042e541b5d314b87a79dcd3ef9e2a

Observation d69db800-7323-4db2-a592-e0ed9afcac39 · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.725752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.725752Z digest=sha256:c2f4164d7a1a10b662b5b4c19df8677ef8e71cee88d8b708587edaa747853783

Observation 6ebd06e6-9de3-4d2c-8045-5942e4032e18 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Multi-concept customization of text-to-image diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.728658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.728658Z digest=sha256:0032c63cf0a2d2bc26b0a51f56e66c5f997e601f85997086c602abc9713899bd

Observation e6b2b2f6-8177-49e4-9f46-f74b2ee39ef4 · outbound

This paper cites Sound-guided semantic image manipulation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Sound-guided semantic image manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.392936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.731687Z digest=sha256:625da6fdfdf755c2c0ec99ee3ef5a5ed7f11970ef9dd7182790a631b1cbc21d0

Observation 44e2fce0-b5e8-460a-89a0-53bb27120f97 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.734970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.734970Z digest=sha256:c7eecb530588e02bd419c9cce0f554f1afd23901acf150b1cd262d0a932df78e

Observation 2b58a244-9615-45e9-87b1-0f98f98c73d5 · outbound

This paper cites Generating real- istic images from in-the-wild sounds.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Generating real- istic images from in-the-wild sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.383588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.737985Z digest=sha256:221af67849493505e7eceba7dd21cb342964d2a931cb9a4d1583c02f3a956d95

Observation dbb0e571-384b-4b43-9fc2-58e48bf14f32 · outbound

This paper cites Learning visual styles from audio-visual associations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Learning visual styles from audio-visual associations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.374499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.741316Z digest=sha256:06aad7b6700d5563831e45fc4ab1e91121d3949f2cf69d2224a4b5d14e96c325

Observation d0ef518f-84ba-4c12-bfae-ff480a868876 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.365099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.744280Z digest=sha256:a21607b4d3b0d76e7919bfad2d0ad387d80f2b00f68492dac1deeef22d9d1484

Observation ff760cde-a948-48cb-b605-60bd5298135c · outbound

This paper cites Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.747156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.747156Z digest=sha256:edd240b4b858403796a7591000edde0e4c78bebeca142bd8457f40028522f3e0

Observation d6bfd0e8-3685-41e7-ab70-c02f60c6012f · outbound

This paper cites Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.750629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.750629Z digest=sha256:8e34917f5ec978f18cde736e4cf64b641b9f87944ca0313501356eff7ac2dac9

Observation b0ef9092-7b6a-45a5-a957-900a8070585c · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.753877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.753877Z digest=sha256:92e053ed7aace93c7c5cfed5dea3976433e4753fc426b2141bfc4c2d525fa348

Observation 2af4216c-7cf4-4b16-8b8f-ab8664569220 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.355091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.757006Z digest=sha256:3cf6bfcdd5efee8ee34cd668d394d1a8710d33617508b801a7f9ac98ea153162

Observation 79bccee3-4715-4b5f-baf5-07bff6f8656c · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Videofusion: Decomposed diffusion models for high-quality video generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.345669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.760023Z digest=sha256:8c1ce3c4a56551fd29da43386bd51da3047ef289cd764b46d3644c00b51b8e68

Observation 75333614-6bef-46d9-9f16-2e05dbf39fcb · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.763240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.763240Z digest=sha256:ea3c92ec0778a52c4d72930b6d8442427c52038d97625c5b4e09eba4214d7d53

Observation 81065768-7ebe-4cfd-8d4a-63f846f602dc · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.766335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.766335Z digest=sha256:3769d4542d227722981fa30eb6d4025b6bb70295819ed011c01bda0d5c2da9c5

Observation b15750d2-18e5-4fd5-9900-ea7ac919ef72 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.336618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.769728Z digest=sha256:21adbdaddfb367983ed27e8bb097cc5824afa992808dd5f587535437e583be5f

Observation f717f9cb-22d1-47f5-97fb-41227fd3e1b8 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Null-text inversion for editing real images using guided diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.327476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.773612Z digest=sha256:21fa5a0e2f19b68dbcd14eceb486dfaf89709311613bd1f79d876b29d0ba2db6

Observation 705322b1-435b-438a-83ee-9b6922827448 · outbound

This paper cites Conditional image-to-video generation with latent flow diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Conditional image-to-video generation with latent flow diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.318131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.776779Z digest=sha256:009eeb931d9b04a792e5632cfaec8d84abffbc87984ad0e094d07aa1fc849b61

Observation 0d3e2c24-9ba4-496b-8b26-cda1590b0b48 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.780021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.780021Z digest=sha256:a4d6c401b77afc99a84af77851028133012ef25e668634579f075dc559e27d06

Observation c37a2fe5-dc43-484b-a07a-55102189a072 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The 2017 DAVIS Challenge on Video Object Segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.784026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.784026Z digest=sha256:d3685df14b00dcab442a6b3f00af6e6bae09b2368a3a53800aea9d06a04a4afd

Observation ed40fd61-e4ed-44a3-9e05-3eaaa2dddf45 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Grad-tts: A diffusion probabilistic model for text-to-speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.788198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.788198Z digest=sha256:e6df2c85a17cc14152829b03af5a061414cda4a505abb2e3f0ef1dd52c8477ab

Observation 8d8de396-afe3-4b6b-93c1-4d7b4d040fc0 · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Fatezero: Fusing attentions for zero-shot text-based video editing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.301659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.791205Z digest=sha256:fa55d543a0b7e7f092932e346917c922dec6c4c5daca487bc25358ea5f4d77a9

Observation a225ef8d-4329-4cee-8f46-ac398c6ce10a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.794473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.794473Z digest=sha256:3c6b0bfbde5055c4de39fcbdbf323c7bc843731a403e6cdc9e50809b858c786d

Observation d343c0ec-5f2a-4eb8-9491-4e2c3ba36eb8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.292428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.797335Z digest=sha256:67f1187c21971abae8c8595b8e8264db5be547863b952193158dc8fbe77411b3

Observation cc1637b2-aefb-40aa-a17b-b07656bcda18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Photorealistic text-to-image diffusion models with deep language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.283106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.801313Z digest=sha256:0eed8b230f7ee1406ffb677ac0ad8a18e9d9c865b6c147a4187fcd3598ce7acb

Observation 6562b558-0a3e-403b-9f40-55481a3dceab · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.804516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.804516Z digest=sha256:2d43ee3bd3d1adaf3758faf8c95d83a91aa28573d8866ad112eacf26f49be000

Observation 0277b3f2-5a4f-4538-8813-931e2d73d77c · outbound

This paper cites Denoising Diffusion Implicit Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Denoising Diffusion Implicit Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.807657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.807657Z digest=sha256:0e1bdf4f362f93c0d4243fbcae73e2c1a6fa65b98f95125e085a714737221c4e

Observation 68203758-5db6-4225-9d4d-4f63e58a8ec0 · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.811712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.811712Z digest=sha256:0619b9d82866d4232c6bfa47a96532f0d3c6507812d09e7ef2d90c5a5c115c5f

Observation 8002fae2-0e9c-43c9-b91c-38834b9edba0 · outbound

This paper cites Any-to-Any Generation via Composable Diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Any-to-Any Generation via Composable Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.814837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.814837Z digest=sha256:348fc8333b9b8b8d27e4c6190fcc4ac220865c4136a0ef01d95a99ea73bfa849

Observation 8390fd36-fdb4-4726-9136-619a2a8f4708 · outbound

This paper cites Splicing vit features for semantic appearance transfer.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Splicing vit features for semantic appearance transfer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.273359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.818212Z digest=sha256:6ff8bce8547dad3c83e56e977f7ef874f8e25fa6f3587c8a16edd4ed7c3be403

Observation 99c62098-9257-42b2-8a37-27ae0cebcfaa · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Plug-and-play diffusion features for text-driven image-to-image translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.821082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.821082Z digest=sha256:a986f8502bf4dec34e5277c58b4310c745de02d72de346db116ecb75cf0f83dc

Observation fa7fff4f-95cf-4e2e-8053-f3d43caeab9a · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Audit: Audio editing by following instructions with latent diffusion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.258243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.824068Z digest=sha256:5e53b630d53674a43c445bcd1cb8d73646c738539158992ad4081d691e369233

Observation 509d7ec8-fc05-4c4a-a237-72817d40639c · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.248945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.827209Z digest=sha256:fdb19a5b19f5c201aeab5938bff36cfc11f2edb4f16d824ef830cbadca61f2b2

Observation 12fef8f2-af64-4a1a-b178-bf8f1f08b6cc · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CVPR 2023 Text Guided Video Editing Competition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.830361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.830361Z digest=sha256:469a6dc9a562a6693edabcadd3f8a4fe4969419d4d5a937710c8a198da297808

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:8b572e214cfae3b0865b42780a95e0b2370d9222cb68316dd1bbc3e4ca9cd4ae

Observation b239d8f0-2453-48f3-a5b9-0c93d6930919 · outbound

This paper cites Ar-diffusion: Auto-regressive diffusion model for text generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ar-diffusion: Auto-regressive diffusion model for text generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.239486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.837124Z digest=sha256:deafff688ec965b1aa276c482b5f62572df32f670c073b72787ac574df293c00

Observation 6617b490-5b00-4988-a50e-b1cc63fb1f43 · outbound

This paper cites Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.902187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.840253Z digest=sha256:fef2f06b81dece43dcc8cd515ef760316fee9d73b4d420fb08b7c2ec6d301f80

Observation 3de73937-bc91-4434-b7a8-97c435a9ecdc · outbound

This paper cites Align, Adapt and Inject: Sound-guided Unified Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align, Adapt and Inject: Sound-guided Unified Image Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.886034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.843363Z digest=sha256:aea580345544ca59215ff0d0bbe9dbe6ae65bf2b5571906459c53bf6662b9987

Observation 91f0718d-73b0-4c14-a3c9-8a09db9e19bf · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The unreasonable effectiveness of deep features as a perceptual metric

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.230376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.846869Z digest=sha256:f8c387b459d6edc2cd3b4a780575092d5346bf08d3c3ec683c35d8f19930b045

Observation 40491051-3946-400f-9b6c-7283985938a4 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.221511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T15:10:31.849774Z digest=sha256:d31238b24f00fca3351d3f4282a0adc577901b55e4b3835e795960967feafa8a

Pith citing papers

No inbound Pith citation observations are available.