Pith. sign in

REVIEW 2 cited by

Efficient Video to Audio Mapper with Visual Scene Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09823 v1 pith:3HHETYAF submitted 2024-09-15 cs.SD cs.MMeess.AS

classification cs.SDcs.MMeess.AS
keywords videoaudiomodelmultiplescenesbaselinegenerationhandle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have made significant progress in bridging the domain gap between video and audio, generating audio that is semantically aligned with the video content. However, a critical limitation of these approaches is their inability to effectively recognize and handle multiple scenes within a video, often leading to suboptimal audio generation in such cases. In this paper, we first reimplement a state-of-the-art V2A model with a slightly modified light-weight architecture, achieving results that outperform the baseline. We then propose an improved V2A model that incorporates a scene detector to address the challenge of switching between multiple visual scenes. Results on VGGSound show that our model can recognize and handle multiple scenes within a video and achieve superior performance against the baseline for both fidelity and relevance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

    cs.MM 2025-01 conditional novelty 6.0 of 10

    AGAV-Rater, an LMM fine-tuned in two stages, achieves state-of-the-art quality scores for AI-generated audio-visual content, text-to-audio, and text-to-music.

  2. FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment

    cs.SD 2024-12 conditional novelty 6.0 of 10

    FolAI predicts an editable RMS envelope from silent video and uses it, with semantic embeddings, to condition a Stable Audio diffusion model for 44.1 kHz stereo foley generation.

Pith tools