Pith. sign in

REVIEW 1 cited by

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.17203 v1 pith:P2TDLSYZ submitted 2023-06-29 cs.SD cs.CVcs.LGeess.AS

classification cs.SDcs.CVcs.LGeess.AS
keywords diff-foleyaudio-visualfeatureslatentvideo-to-audioaudiocavp-aligneddiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation quality in terms of temporal synchronization and audio-visual relevance. We present Diff-Foley, a synchronized Video-to-Audio synthesis method with a latent diffusion model (LDM) that generates high-quality audio with improved synchronization and audio-visual relevance. We adopt contrastive audio-visual pretraining (CAVP) to learn more temporally and semantically aligned features, then train an LDM with CAVP-aligned visual features on spectrogram latent space. The CAVP-aligned features enable LDM to capture the subtler audio-visual correlation via a cross-attention module. We further significantly improve sample quality with `double guidance'. Diff-Foley achieves state-of-the-art V2A performance on current large scale V2A dataset. Furthermore, we demonstrate Diff-Foley practical applicability and generalization capabilities via downstream finetuning. Project Page: see https://diff-foley.github.io/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sounding that Object: Interactive Object-Aware Image to Audio Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A latent diffusion audio model is trained to ground sound in image patches, then uses SAM segmentation masks at test time so users can generate audio for selected objects in a scene.

Pith tools