Pith. sign in

Paper Citation Record · LEDGER

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance

As of 17 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.20995.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20995 v4

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:45:08.841789Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2efef86-b747-425b-86c4-8cc021f95666 · outbound

This paper cites The Foley Grail: The Art of Performing Sound for Film, Games, and Animation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance The Foley Grail: The Art of Performing Sound for Film, Games, and Animation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.351577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:04.880574Z digest=sha256:55d12ea4e036b4c62462172265b8885748a79c795198df88e158c071a1bf15a3

Observation bf501168-58c8-491e-88bf-c0087732d0c7 · outbound

This paper cites Qwen2.5-VL Technical Report.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:04.943204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:04.943204Z digest=sha256:ab398d3a9aee3bc3ca038703c260d9cdfa444d44a79e9b7e5d1efa18e2fd407b

Observation 138ac0b1-663d-4891-ae36-377b6e6f1b38 · outbound

This paper cites Erasedraw: Learning to insert objects by erasing them from images.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Erasedraw: Learning to insert objects by erasing them from images

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.340230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.024306Z digest=sha256:480ebd04c306add342f764a6d3d7bd852f12c1c7afab896d9ed9cafa746434f3

Observation 9cc5593b-09cc-4f8e-9f1e-3ef031053a27 · outbound

This paper cites Action2sound: Ambient-aware generation of action sounds from egocentric videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Action2sound: Ambient-aware generation of action sounds from egocentric videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.327905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.113307Z digest=sha256:cfa9d6661ce8536a7f5fbd1c35191302ab8aa46b3dada94a6ad44d87aeb1dc1b

Observation f19cb447-815d-43e3-b6e9-057ba58a388e · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.316890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.234622Z digest=sha256:8e2e99fc301d0fd6283c6314e18ea5bb28ca54cc060518a7126b3b258b947a50

Observation 50d1729f-30ce-4e2c-9a55-874ca54bf8db · outbound

This paper cites Generating visually aligned sound from videos.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generating visually aligned sound from videos

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.305399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.353789Z digest=sha256:2675ca9c6baaa71815074e76c9cd2c603ff4f50585c4f48977c5adff14f9d7e0

Observation 83671cc3-1f8c-46ab-adbb-f180111af7f3 · outbound

This paper cites Video-Guided Foley Sound Generation with Multimodal Controls.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Video-Guided Foley Sound Generation with Multimodal Controls

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:05.422297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:05.422297Z digest=sha256:1f56ceed706bfb59e89baaa71a706b27693bd7c512eae3e201bb3c7da6b0d63e

Observation 465ce467-2a71-45bb-94bb-5bd783ab6dd0 · outbound

This paper cites Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Mmaudio: Taming multimodal joint training for high-quality video-to-audio synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.293192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.510328Z digest=sha256:0445823ea6cc00850e2c912c5a56540306aab8f1efe66e97d914452dfba61c01

Observation f6f79a6e-9c2b-497d-9ccb-aa341c3f8f5e · outbound

This paper cites Clotho: An audio captioning dataset.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Clotho: An audio captioning dataset

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.281452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.606265Z digest=sha256:4b91f046f9eddb260a4aa862b9db82bc39b196878f033b4671010aa843b09e6d

Observation c48835b7-ca7e-4973-9b99-a865742c2709 · outbound

This paper cites Compositional visual generation with energy based models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with energy based models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.269734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.722732Z digest=sha256:5eff5d24eac34982ebf6dff474fca6b96e9a44a96dfba4b4b55ce335d622d68a

Observation 8fcf95f3-9723-4d9e-8778-5fdb8415bf20 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Scaling rectified flow transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.258064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.808806Z digest=sha256:e82ddad4b6d9b467aed7e8a2dbeb3c9ca60eea94b9e4bd3e899295e4fc7e94f4

Observation 5a745c0a-b954-408e-bd54-993eacde0289 · outbound

This paper cites Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Sketch2sound: Controllable audio generation via time-varying signals and sonic imitations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.245835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.904797Z digest=sha256:ba93599d375e6fe8b2cab8da6702fea79923fe30b15dc9e7da3689793a6cbb48

Observation e2701f7c-992a-43d0-bba1-362698e35d76 · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio set: An ontology and human-labeled dataset for audio events

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.233666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:05.968972Z digest=sha256:82901ea29d361080331ea1047a95012034c749960c9e17a06e9a49517d81b1db

Observation 8ab04e99-e18d-484c-9a57-60fc2b2b7602 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Imagebind: One embedding space to bind them all

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.222069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.021697Z digest=sha256:06e9f5efee63b1b7dcec04adaada923bfaa892a174e7135d48dbe46066ee7378

Observation d2359448-9249-44de-8680-c523e1dad7d8 · outbound

This paper cites Instructme: An instruction guided music edit and remix framework with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Instructme: An instruction guided music edit and remix framework with latent diffusion models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.209188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.075730Z digest=sha256:80f002df2fe636a91dab92cefb2293e3c0db91040a9041bb4af2a6fd8d59ecaa

Observation 784a3cab-fd12-4243-b191-259de7ea9baf · outbound

This paper cites Classifier-free diffusion guidance.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Classifier-free diffusion guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.151487Z digest=sha256:73dc364ef1e284b6fbc8786deead185f95beed03f6cc36a2d98bccda108914b8

Observation 791d89c6-44a6-45a4-9721-f24578894e16 · outbound

This paper cites Taming visually guided sound generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Taming visually guided sound generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.186932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.203790Z digest=sha256:5c0b3481d88e5ea97f2d4c7b8d5b339d932a46e5244b9f04f8ffd843e0b00a6e

Observation 3f00ebf4-1299-48dc-809f-7d36449e9505 · outbound

This paper cites Synchformer: Efficient synchronization from sparse cues.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Synchformer: Efficient synchronization from sparse cues

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.174622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.275002Z digest=sha256:51105b68aac64f6a4ae12ae0ac4261987d1bd36e3a6f361e47bfa96018404b91

Observation 76d83d63-5816-4a20-bcb7-456bbf0e3d3b · outbound

This paper cites Audioeditor: A training-free diffusion-based audio editing framework.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audioeditor: A training-free diffusion-based audio editing framework

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.163649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.323102Z digest=sha256:4d71e2eababe6f7160f79234343b9d21c3037619fa621c46c5898f7f2549847b

Observation a6d67126-3422-4ee6-b220-50bc93246462 · outbound

This paper cites Simultaneous music separation and generation using multi-track latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Simultaneous music separation and generation using multi-track latent diffusion models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.151760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.376824Z digest=sha256:87e6aa4f7f9521c42ac2179ed400288a16ac071df01f9ae2baaaf40eb934b619

Observation 543f135c-bde9-4a08-907e-32953996f07c · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Analyzing and improving the training dynamics of diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.140505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.432024Z digest=sha256:02ae3b97d1131392174002377936f666c3589040476238dbe68c0f24d6e6cf8e

Observation f0c3b862-20cb-440b-8185-4c5cc2213f44 · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audiocaps: Generating captions for audios in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.128711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.488469Z digest=sha256:f2d07016d70583b191cea049d747622ca8a24bcd825e0a8d6aa5afc0df908880

Observation ee2e0e48-327f-416e-93d1-9878a97a85f3 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Panns: Large-scale pretrained audio neural networks for audio pattern recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.116398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.539100Z digest=sha256:440bdf8cfaa20c3ec94a985177c7fbc8fd4bde2c342548addf6517d05ab84828

Observation 2285d8a5-0e8d-4396-94bf-2d20c51f43b8 · outbound

This paper cites Vintage: Joint video and text conditioning for holistic audio generation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Vintage: Joint video and text conditioning for holistic audio generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.104521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.608604Z digest=sha256:8e4f1aabff7caae6e85da20ca804dd035009ccfd59a555d1f3a033ed513071c5

Observation 3b64cfc8-9b49-4187-b718-f8eddc76d437 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Gonzalez, Hao Zhang, and Ion Stoica

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.091741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.675323Z digest=sha256:0012b4e32f51d00cb68fc28bf4a16665019f296c65f65b2838d77a547548f52f

Observation 8d5aed84-561b-4ff3-b3da-80fc48acb6f0 · outbound

This paper cites Flow matching for generative modeling.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Flow matching for generative modeling

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.080943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.722269Z digest=sha256:8ff5fe584c3a4d7b279a242b8fa4dbbe6274e12c81fd2cad55d6e6bc7a23bdba

Observation 72f18168-e4f9-43bd-bd23-35e8754133ee · outbound

This paper cites Compositional visual generation with composable diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Compositional visual generation with composable diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.069743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.789909Z digest=sha256:1922a170b8cdd2e1b66a803487d8b5a6e700b495965dd5e7b8e1de34b4ea4ef4

Observation fbac445a-9263-4075-8796-713d4175d007 · outbound

This paper cites Tell what you hear from what you see-video to audio generation through text.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Tell what you hear from what you see-video to audio generation through text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.058915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.849457Z digest=sha256:6b7d05be00738d9eb69200f893433959dae487c4f04096f24fc33bc16d949696

Observation 5c27eb4f-97b7-4233-b673-0f8f00d8fd96 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.047856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.916636Z digest=sha256:fcf34415dc89dc657aa702cbf886ce0dc9f1f5dbaeedfa230f74f18db6c0a2ba

Observation 79b5cffe-d872-4a6b-a594-afe71466b990 · outbound

This paper cites Multi-source diffusion models for simultaneous music generation and separation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Multi-source diffusion models for simultaneous music generation and separation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.037038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:06.972516Z digest=sha256:2ea7020836a77271ef224482678404fcb2a35d7a5725eabf6b20694f7df1176a

Observation 1c51d7c4-9601-4119-a8e5-7b6e238e7efa · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.024863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.046130Z digest=sha256:1d1fe983e173d8b01597ef4cb45bdebb6220ef3beaafc593444c713b94ebeb5a

Observation 20fca138-03ed-415c-958c-0173fb57af3a · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audio-visual scene analysis with self-supervised multisensory features

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.012958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.096950Z digest=sha256:2ffb56874534abfb028ba0f83173cd1d7d48c681df3135899c2375f788b83b33

Observation 0f77170d-dce8-464e-84cd-b1b0bb67c57f · outbound

This paper cites Stemgen: A music generation model that listens.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stemgen: A music generation model that listens

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:12.001536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.138164Z digest=sha256:d39e9dc04e8f713d5a86e726cba4773a7cea7110f7e83e450aa92584e7224995

Observation 186907ed-1502-45d5-82a0-79d9c0153b15 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Movie Gen: A Cast of Media Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:07.206934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:07.206934Z digest=sha256:ef5c0db0656c16089d91b6743e006f55d48d93d42783938bf91a20d4df200978

Observation 463dd510-0502-4450-9692-436cc26a425b · outbound

This paper cites Generalized multi-source inference for text conditioned music diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Generalized multi-source inference for text conditioned music diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.990280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.274186Z digest=sha256:98a3d323e226619dc710f2f6920d97127cd8b76fdabcb2148d6e81f5f271253f

Observation b96150f7-b8f9-4748-a4de-37565fc32655 · outbound

This paper cites Improved techniques for training gans.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Improved techniques for training gans

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.978310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.347645Z digest=sha256:3165c213b8bd3f3816100021fa164a3c7e9ca2287a938171759d4a7cadcd5247

Observation e65fd7f1-2874-4163-a217-7e3bb9e57f15 · outbound

This paper cites Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Smartmask: context aware high-fidelity mask generation for fine-grained object insertion and layout control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.960730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.473056Z digest=sha256:5b74e03f03fe5a26ebcf2835062ed09aacab547bbf364cb2f35cc61bf5e9a96a

Observation a4f55f75-e5ed-4ed4-bb7c-d95b5b065c55 · outbound

This paper cites Visually guided sound source separation with audio-visual predictive coding.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation with audio-visual predictive coding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.832404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.585707Z digest=sha256:7fedc42bb26f7e54d2e493ba1c2d784202d50d96a54a7bd51558c080c8bb58d0

Observation 4e31c034-6b9a-4f01-93ab-ae7ff2779e8c · outbound

This paper cites sd3.5, 2024.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance sd3.5, 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.502321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.682202Z digest=sha256:5a3b42d915e12537dd202d6a618be1d213679f851a8d476c74b215ffb5d0d255

Observation 37145041-e787-421d-b6f3-a1e17e010e53 · outbound

This paper cites Steinmetz and Joshua D.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Steinmetz and Joshua D

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.318638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.779691Z digest=sha256:af2f440aa013031be372bace7d0f16e98968e3fbf940f81b9f0b98b90f0171a7

Observation 03c38e60-10be-4a19-ae61-8e74395b6276 · outbound

This paper cites Add-it: Training-free object insertion in images with pretrained diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Add-it: Training-free object insertion in images with pretrained diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:11.045041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:07.883778Z digest=sha256:202d41fab130b489a67ff2236f3b074700e2f3a726ec0608850da93936b197ff

Observation 0bc16ca5-f94f-4c75-ba2a-1b4004b179b2 · outbound

This paper cites Liu, Kevin J.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Liu, Kevin J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.900034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.011468Z digest=sha256:b5b4786217768ee236c8eaad6b4f8ae035c5eccb30d03117dcea5270c919418a

Observation 9777209a-2c83-403a-82c9-35f0b3cdcafe · outbound

This paper cites Temporally aligned audio for video with autoregression.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Temporally aligned audio for video with autoregression

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.758908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.098102Z digest=sha256:83dce3c23bb9270da865b26692d9b1d76faf2b393305eab54c65378778d271db

Observation f6e8b40c-660b-4d33-9552-06e095b8af5c · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.554435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.178693Z digest=sha256:4624df7e264a4dd0aac6f6b74f5b7265538e52ac769737127e31362d7d913644

Observation 5ba41fe0-c0b5-41dc-88ec-625b962f1694 · outbound

This paper cites Frieren: Efficient video-to-audio generation network with rectified flow matching.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Frieren: Efficient video-to-audio generation network with rectified flow matching

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.395064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.285868Z digest=sha256:574285f575ea6dc0e1598533d9842ef6e791ddc11b941fea42ff69583990ce3d

Observation e0dde940-a687-4083-9f75-0920aa92541e · outbound

This paper cites Audit: Audio editing by following instructions with latent diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Audit: Audio editing by following instructions with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:10.226759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.365573Z digest=sha256:f2c4eaa114577309a1d70f1885dac20e6258c61292537e08a5fa0d68288fe241

Observation d4d331b0-7589-46d7-a134-6f921ba9eae9 · outbound

This paper cites Stable diffusion 2.0 and the importance of negative prompts for good results, 2022.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Stable diffusion 2.0 and the importance of negative prompts for good results, 2022

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.978231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.474046Z digest=sha256:ba2cb6842c1f11e7f701aa2afb9e9f67e07b85e637d765d50f363f73e9156c67

Observation fbf2a7b1-9d47-4e00-a5ab-eed923cef885 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.732295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.535917Z digest=sha256:3dc53e5614e6f179d138bb0d5fd6924b98c545069c9b8ddb199c30dcf2cbd3cb

Observation 0102a54a-4938-4296-8fd9-88cd98f822af · outbound

This paper cites Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.528185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.546056Z digest=sha256:c036bb3e42c2d7d54efa924676783c9e47a38c1ca7e628536d5fec830a1ca121

Observation 6d026ab3-f7b1-4d29-9af6-1d3023bc86fd · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.304523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.621106Z digest=sha256:2c2cd15f10f1d254a19a17067fd1849fef5e04d2a655dfdc44e5ab21ed007f02

Observation f86c4cf0-2d82-4547-8a55-783902287549 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:08.730625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:08.730625Z digest=sha256:ab5b5f96b1e5698167e30a634fe100bd27222b34f7b48489c673f864174063d2

Observation b754503b-876d-4787-b826-07e6136fc464 · outbound

This paper cites Visually guided sound source separation using cascaded opponent filter network.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Visually guided sound source separation using cascaded opponent filter network

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:45:09.086947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T22:45:08.841789Z digest=sha256:7baf1304774e4448de1ed94dd957ac8caba75662d538ab11613849a3d075348c

Pith citing papers

No inbound Pith citation observations are available.