Pith. sign in

Paper Citation Record · LEDGER

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2508.20379.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20379 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:10:31.849774Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact4
  • verified fuzzy25
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3920ab5b-1b92-456c-babc-9857e328da00 · outbound

This paper cites GPT-4 Technical Report.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.658133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.658133Z digest=sha256:a182b0f062cfe7d34c55b93af701012d59e603d295a29df74285270c8ed0b7c8

Observation adc71368-bcf9-480b-aff4-9589c8e14b6e · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2live: Text-driven layered image and video editing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.469198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.662463Z digest=sha256:46bc01a0ae55215e3b58384829c6f5eeb643e0f96fe4c519015b25ff7b60bc86

Observation c1b863f0-6252-49cf-8419-27a579b8d80f · outbound

This paper cites SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.202014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.667945Z digest=sha256:97cb0e79736833edb340a6457792d5cd58e814730dd173cbc8ccf22b73b37292

Observation 47e2c4cb-2364-43fa-8d26-fad3ea7f3196 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align your latents: High-resolution video synthesis with latent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.671833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.671833Z digest=sha256:87407f32439df685fb041fa5bc345414fce27efaa4a31ed79bcf1a9b39422071

Observation cb1a46e1-9ea0-42b4-9a49-88786c7f2087 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ledits++: Limitless image editing using text-to-image models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.454984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.674911Z digest=sha256:9e7bbf38fe58c249d7b22240636e7b0521cc8cb988975ca6f1022474cbd4162b

Observation 8038bc95-56ab-4c82-a49a-426e5953eb8f · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.444636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.678420Z digest=sha256:e63cab60a8687c6b1358b18ce3c7affcd303b54dba0b1e66f02be9f1af766808

Observation 15255ef9-bc95-4305-ae2d-92fa964d1eab · outbound

This paper cites DiffEdit: Diffusion-based semantic image editing with mask guidance.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffEdit: Diffusion-based semantic image editing with mask guidance

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.682377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.682377Z digest=sha256:60316fe2a06e4975e5f06a2d259c64a0e359bb5ecf65e3cb0dbdfca64809c2d0

Observation 3a9ba256-3b8c-407b-9a4f-47b4036a63ea · outbound

This paper cites Con- ditional generation of audio from video via foley analogies.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Con- ditional generation of audio from video via foley analogies

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.434962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.685950Z digest=sha256:a797a2b30c503088c8472d895f2d9775d5bf867f57e2dffc877c69154e777d49

Observation 70e339d7-4417-4a3b-8b9f-20a65d53e31f · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.425812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.689192Z digest=sha256:1446208090028afaf3c1facb131474b7821fb3ff154036c0cb3686871aa766d4

Observation c1a0ee9a-b796-402c-a5d3-19025e5b4ed8 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.692936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.692936Z digest=sha256:66b49f60e6059a71922f5c981f6fd3607c85ffe34c82ae40dd3ca0e3e8acb439

Observation 017fc962-2b1e-4a23-b0a5-9fed696c05ea · outbound

This paper cites Imagebind: One embedding space to bind them all.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Imagebind: One embedding space to bind them all

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.416553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.696749Z digest=sha256:907ffd206098b4872b1666a1b91d87450ab3afbfe550e297debe6c5f367ce947

Observation a38220d8-905a-45b9-a58f-de8655d3d5f0 · outbound

This paper cites AtomoVideo: High Fidelity Image-to-Video Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AtomoVideo: High Fidelity Image-to-Video Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.700664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.700664Z digest=sha256:b69a96a8988909b7a6293af776ce56fa8a96a40786985b2ed707a27def567bf5

Observation b3099f05-64d3-4e41-bf32-662f496e67f9 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.704083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.704083Z digest=sha256:79c2f44f66f00500a07cc00c4a2ae7c01543a6192d33c192e7b7d7c297c2ef60

Observation c1de15e2-4ec3-43b2-9a43-f8d8bd6fb07e · outbound

This paper cites FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:32.154458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.708990Z digest=sha256:d8e428f5e97edc3f47a3126539cf8b8d981cef6e1e66a2858c79826ab610d198

Observation 6c938676-cf99-4b22-b3d5-29caf78618bd · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.713176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.713176Z digest=sha256:e705604eb2d76a174fb2cb35a0947c6c94cf54cfd2db0d66b6ccb8e44d3ae864

Observation d9c1f186-073e-48e5-a586-4874717b2784 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.716262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.716262Z digest=sha256:3ebf022f07220beb1d686f73d8a34f61ff67bba776d4ea94c357e20b311456b3

Observation c4435c43-05ee-4b21-9678-726f78a5d31f · outbound

This paper cites Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.719815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.719815Z digest=sha256:78fb98a03a9f40c28483c47c27c25afdb7c644b3c97f461cd82798d018fab6d3

Observation 38c81fa2-f69a-4611-8626-a787974b7d05 · outbound

This paper cites Text2video-zero: Text-to- image diffusion models are zero-shot video generators.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Text2video-zero: Text-to- image diffusion models are zero-shot video generators

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.407619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.722751Z digest=sha256:6462bedaa3ce7bfa4ddc5d19a516cb180e77141276761336ab41000cec20546a

Observation d69db800-7323-4db2-a592-e0ed9afcac39 · outbound

This paper cites Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Enclap: Combining neural audio codec and audio-text joint embedding for automated audio captioning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.725752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.725752Z digest=sha256:38cdc180f53d781a37f12ebea502b75287daa0b9ffbf2b81cc92baae717d9c70

Observation 6ebd06e6-9de3-4d2c-8045-5942e4032e18 · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Multi-concept customization of text-to-image diffusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.728658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.728658Z digest=sha256:ca1de99331199e70e5c966d3ff9b5cb824dd3846b385606bd44cff156410aab6

Observation e6b2b2f6-8177-49e4-9f46-f74b2ee39ef4 · outbound

This paper cites Sound-guided semantic image manipulation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Sound-guided semantic image manipulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.392936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.731687Z digest=sha256:8371f6a47e6a6e7128588e0dba9e9ccc530c47cda3613cc870056ddc76155388

Observation 44e2fce0-b5e8-460a-89a0-53bb27120f97 · outbound

This paper cites Soundini: Sound-Guided Diffusion for Natural Video Editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Soundini: Sound-Guided Diffusion for Natural Video Editing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.734970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.734970Z digest=sha256:253380bbcfdec0750e7a60b40f427df144be86611cc6bb021a4a1c967f0c243f

Observation 2b58a244-9615-45e9-87b1-0f98f98c73d5 · outbound

This paper cites Generating real- istic images from in-the-wild sounds.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Generating real- istic images from in-the-wild sounds

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.383588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.737985Z digest=sha256:14426e0617fd4ee3ab1f96ca1df5f74ae6825573b24867cb0154bfbb26dc86b0

Observation dbb0e571-384b-4b43-9fc2-58e48bf14f32 · outbound

This paper cites Learning visual styles from audio-visual associations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Learning visual styles from audio-visual associations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.374499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.741316Z digest=sha256:85e1c2f959bac401217c5efd7bbf8eefc24a5bff965629eae6fb6173665283d8

Observation d0ef518f-84ba-4c12-bfae-ff480a868876 · outbound

This paper cites Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.365099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.744280Z digest=sha256:a15eeb0d6aaaabaa98e25b1e3ddbf265fac2b4934f301f9432dff1244fccf203

Observation ff760cde-a948-48cb-b605-60bd5298135c · outbound

This paper cites Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.747156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.747156Z digest=sha256:7bfc9224ef1167b1a2314847adc431f10156b249f5e825d4a1cf7efc6da3365b

Observation d6bfd0e8-3685-41e7-ab70-c02f60c6012f · outbound

This paper cites Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.750629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.750629Z digest=sha256:28d4d2d932384b7b530458824e69a26abd7a703b92b911e6dcd5fb48dfffe828

Observation b0ef9092-7b6a-45a5-a957-900a8070585c · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.753877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.753877Z digest=sha256:2b273ca11d0813ce217817fd58101b2a24f17fd2f9fd179bdc2ff58808cbf848

Observation 2af4216c-7cf4-4b16-8b8f-ab8664569220 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.355091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.757006Z digest=sha256:67f30eea7c375ed9614b71d6831f61e026839617221640b0f199a2c89c9e621f

Observation 79bccee3-4715-4b5f-baf5-07bff6f8656c · outbound

This paper cites Videofusion: Decomposed diffusion models for high-quality video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Videofusion: Decomposed diffusion models for high-quality video generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.345669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.760023Z digest=sha256:e2bbc08caf2c7f68b4b1abfa7ad113a34e30de4956b9f7940ac3f0a4d3c944e6

Observation 75333614-6bef-46d9-9f16-2e05dbf39fcb · outbound

This paper cites Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Zero-Shot Unsupervised and Text-Based Audio Editing Using DDPM Inversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.763240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.763240Z digest=sha256:dfb34b5b60543584590e99420a2b5917224a3e3237a4f8e2da5ecbc366385831

Observation 81065768-7ebe-4cfd-8d4a-63f846f602dc · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.766335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.766335Z digest=sha256:c5cc62be0bab35cfc112d4e1ed1fbc51fb22aaf5c6e4f15bbfa6632733a51959

Observation b15750d2-18e5-4fd5-9900-ea7ac919ef72 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.336618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.769728Z digest=sha256:5ef194f4091356191440391bf8014a13886c64032984d9ddc8f87d32c7052fc1

Observation f717f9cb-22d1-47f5-97fb-41227fd3e1b8 · outbound

This paper cites Null-text inversion for editing real images using guided diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Null-text inversion for editing real images using guided diffusion models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.327476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.773612Z digest=sha256:bf3fdd65a88a0eec9436f11ab7572f9725e43a830184f1e5cb399529815bf861

Observation 705322b1-435b-438a-83ee-9b6922827448 · outbound

This paper cites Conditional image-to-video generation with latent flow diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Conditional image-to-video generation with latent flow diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.318131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.776779Z digest=sha256:780d11e89250b83fc9c450e02c72e84760f815da38aa2b4c20d5e1be76e80d30

Observation 0d3e2c24-9ba4-496b-8b26-cda1590b0b48 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.780021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.780021Z digest=sha256:bc288b33b1af4b287ba1f2b54d5854aeb26f8d91be526d0da40b0ea13f7a15db

Observation c37a2fe5-dc43-484b-a07a-55102189a072 · outbound

This paper cites The 2017 DAVIS Challenge on Video Object Segmentation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The 2017 DAVIS Challenge on Video Object Segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.784026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.784026Z digest=sha256:b3a3bd5a11402fc037ed1b72afef1d412b32720ee9cf4180293592605a651f1f

Observation ed40fd61-e4ed-44a3-9e05-3eaaa2dddf45 · outbound

This paper cites Grad-tts: A diffusion probabilistic model for text-to-speech.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Grad-tts: A diffusion probabilistic model for text-to-speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.788198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.788198Z digest=sha256:f71217a5b99451d3fa1f4deed02d00c8938c7d486f5888f7d72ff3528b9c2bc8

Observation 8d8de396-afe3-4b6b-93c1-4d7b4d040fc0 · outbound

This paper cites Fatezero: Fusing attentions for zero-shot text-based video editing.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Fatezero: Fusing attentions for zero-shot text-based video editing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.301659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.791205Z digest=sha256:ee1495dfc16e23a6e3615e9f98d48ad7495108b4a2a43b6454bf717db68b07aa

Observation a225ef8d-4329-4cee-8f46-ac398c6ce10a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.794473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.794473Z digest=sha256:795ccefd674751b22607360b1803628d3c458e5d610826a030f12ee8bcf54ae9

Observation d343c0ec-5f2a-4eb8-9491-4e2c3ba36eb8 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts High-resolution image synthesis with latent diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.292428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.797335Z digest=sha256:a91d8bd30628c60e001ff6110c89f2c215d22d03c2fcc137fee5c1cc3e1c8a39

Observation cc1637b2-aefb-40aa-a17b-b07656bcda18 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Photorealistic text-to-image diffusion models with deep language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.283106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.801313Z digest=sha256:4f4fa27485b290196a329a58099f834d3eaa3795f28088125cc356388e5d75ea

Observation 6562b558-0a3e-403b-9f40-55481a3dceab · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.804516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.804516Z digest=sha256:e9ed09f979954ff9a0d6e743532a58e36968f564417e095a02a4b342479f9f4e

Observation 0277b3f2-5a4f-4538-8813-931e2d73d77c · outbound

This paper cites Denoising Diffusion Implicit Models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Denoising Diffusion Implicit Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.807657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.807657Z digest=sha256:7f1bc0faa1ddb312dbebabcd612d5f689c5d77f0fdc969cbb30d56878ffcd0c7

Observation 68203758-5db6-4225-9d4d-4f63e58a8ec0 · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.811712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.811712Z digest=sha256:1643f331ffa5bc42b041ee0616c293d42747a96aee2ccc29cbf692f1b26bbafa

Observation 8002fae2-0e9c-43c9-b91c-38834b9edba0 · outbound

This paper cites Any-to-Any Generation via Composable Diffusion.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Any-to-Any Generation via Composable Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.814837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.814837Z digest=sha256:95869f47a5a97bf7f8b3638fbf8d7e8052dde2c9f683fb16a43a6b7123bd7e26

Observation 8390fd36-fdb4-4726-9136-619a2a8f4708 · outbound

This paper cites Splicing vit features for semantic appearance transfer.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Splicing vit features for semantic appearance transfer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.273359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.818212Z digest=sha256:29fa350ca6f39eee3f26475224e8584cfc4d89cb29a00b05a2c3bd6e2a274f63

Observation 99c62098-9257-42b2-8a37-27ae0cebcfaa · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Plug-and-play diffusion features for text-driven image-to-image translation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.821082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.821082Z digest=sha256:517ef7f4aeaaa0c9164f65367d9a3411f74fb554b3a66b13fef231d7ce2946a7

Observation fa7fff4f-95cf-4e2e-8053-f3d43caeab9a · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Audit: Audio editing by following instructions with latent diffusion models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.258243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.824068Z digest=sha256:96b7c5c1cfd1f75012ab5204f3025b0bde42ecb9e41d6c58c3231f86b98bb517

Observation 509d7ec8-fc05-4c4a-a237-72817d40639c · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.248945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.827209Z digest=sha256:33fa851ad4e91a13d83006fd0622bf74f3c706442edd029460cc6c1884144c58

Observation 12fef8f2-af64-4a1a-b178-bf8f1f08b6cc · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts CVPR 2023 Text Guided Video Editing Competition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.830361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.830361Z digest=sha256:3ccbcdb17ffde90dcfebf7c67f19559217048905a4c6d4e109e80e3bcd057db7

Observation 0d76b3bd-d2d5-474b-bdc9-8f2c483b45cd · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts NExT-GPT: Any-to-Any Multimodal LLM

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T15:10:31.833708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:10:31.833708Z digest=sha256:3c2243601d790c0d6af3d5ce1d94411c1b520c9803798a557fd9f855aecb036d

Observation b239d8f0-2453-48f3-a5b9-0c93d6930919 · outbound

This paper cites Ar-diffusion: Auto-regressive diffusion model for text generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Ar-diffusion: Auto-regressive diffusion model for text generation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.239486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.837124Z digest=sha256:9973ccf50d25c8721351c4e1a0e3f49a4360b8fc4fd48b726da4599da90775b4

Observation 6617b490-5b00-4988-a50e-b1cc63fb1f43 · outbound

This paper cites Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.902187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.840253Z digest=sha256:c93c6d9b2793c5b4d1d04b3d193919619aee74b72ab16cda2f08afb1410534f1

Observation 3de73937-bc91-4434-b7a8-97c435a9ecdc · outbound

This paper cites Align, Adapt and Inject: Sound-guided Unified Image Generation.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Align, Adapt and Inject: Sound-guided Unified Image Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:10:31.886034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.843363Z digest=sha256:4ded2e061028a3da4336e97af055409394d03ce7880db6a0d27b9ed7b0e05f6e

Observation 91f0718d-73b0-4c14-a3c9-8a09db9e19bf · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts The unreasonable effectiveness of deep features as a perceptual metric

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.230376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.846869Z digest=sha256:2af29ca0f5272a3e8a393c778ac1b56833b2dc778f0a6c847663f4402ff0e4b5

Observation 40491051-3946-400f-9b6c-7283985938a4 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

Audio-Guided Visual Editing with Complex Multi-Modal Prompts Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:10:32.221511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T15:10:31.849774Z digest=sha256:21c662c1f06784cc73b80705596782b528b08df7f96d0efd700324410c49d31d

Pith citing papers

No inbound Pith citation observations are available.