Pith. sign in

REVIEW 3 cited by

Text-to-Audio Generation Synchronized with Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07938 v1 pith:Z3Z6LZAD submitted 2024-03-08 cs.SD cs.AIcs.CVcs.LGcs.MMeess.AS

classification cs.SDcs.AIcs.CVcs.LGcs.MMeess.AS
keywords generationtemporalaudioembeddingstextt2avtext-to-audiovisual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent times, the focus on text-to-audio (TTA) generation has intensified, as researchers strive to synthesize audio from textual descriptions. However, most existing methods, though leveraging latent diffusion models to learn the correlation between audio and text embeddings, fall short when it comes to maintaining a seamless synchronization between the produced audio and its video. This often results in discernible audio-visual mismatches. To bridge this gap, we introduce a groundbreaking benchmark for Text-to-Audio generation that aligns with Videos, named T2AV-Bench. This benchmark distinguishes itself with three novel metrics dedicated to evaluating visual alignment and temporal consistency. To complement this, we also present a simple yet effective video-aligned TTA generation model, namely T2AV. Moving beyond traditional methods, T2AV refines the latent diffusion approach by integrating visual-aligned text embeddings as its conditional foundation. It employs a temporal multi-head attention transformer to extract and understand temporal nuances from video data, a feat amplified by our Audio-Visual ControlNet that adeptly merges temporal visual representations with text embeddings. Further enhancing this integration, we weave in a contrastive learning objective, designed to ensure that the visual-aligned text embeddings resonate closely with the audio features. Extensive evaluations on the AudioCaps and T2AV-Bench demonstrate that our T2AV sets a new standard for video-aligned TTA generation in ensuring visual alignment and temporal consistency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A single flow-matching transformer trained jointly on audio-video and audio-text data produces state-of-the-art public video-to-audio synthesis with a frame-level synchronization module.

  2. VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A text-and-video conditioned flow transformer that generates onscreen plus offscreen audio, evaluated on a new curated benchmark and on VGGSound.

  3. SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

    cs.MM 2024-12 conditional novelty 6.0 of 10

    SyncFlow jointly generates temporally aligned 16 FPS video and 48kHz audio from text using a dual-diffusion-transformer with modality adaptors, and reports better audio-video alignment than cascaded and contrastive baselines.

Pith tools