Pith. sign in

REVIEW 2 cited by

Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.08628 v1 pith:WUP76UNY submitted 2024-09-13 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords audioalignmentaudio-visualframeworksemanticadapterbeatseamless
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our research introduces an innovative framework for video-to-audio synthesis, which solves the problems of audio-video desynchronization and semantic loss in the audio. By incorporating a semantic alignment adapter and a temporal synchronization adapter, our method significantly improves semantic integrity and the precision of beat point synchronization, particularly in fast-paced action sequences. Utilizing a contrastive audio-visual pre-trained encoder, our model is trained with video and high-quality audio data, improving the quality of the generated audio. This dual-adapter approach empowers users with enhanced control over audio semantics and beat effects, allowing the adjustment of the controller to achieve better results. Extensive experiments substantiate the effectiveness of our framework in achieving seamless audio-visual alignment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

    cs.CV 2024-12 conditional novelty 7.0 of 10

    AV-Link unifies video-to-audio and audio-to-video generation by aligning frozen diffusion-model activations with temporally matched rotary position embeddings in a shared Fusion Block.

  2. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.

Pith tools