REVIEW 5 cited by
Audio-Visual Instance Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, containing over 90K instance masks from 26 semantic categories in 926 long videos. Additionally, we propose a strong baseline model for this task. Our model first localizes sound source within each frame, and condenses object-specific contexts into concise tokens. Then it builds long-range audio-visual dependencies between these tokens using window-based attention, and tracks sounding objects among the entire video sequences. Extensive experiments reveal that our method performs best on AVISeg, surpassing the existing methods from related tasks. We further conduct the evaluation on several multi-modal large models. Unfortunately, they exhibits subpar performance on instance-level sound source localization and temporal perception. We expect that AVIS will inspire the community towards a more comprehensive multi-modal understanding. Dataset and code is available at https://github.com/ruohaoguo/avis.
Forward citations
Cited by 5 Pith papers
-
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
A new patch-level sounding object tracking method with motion-, sound-, and question-driven graph modules achieves 78.42% average accuracy on MUSIC-AVQA, competitive with large-scale pretraining approaches.
-
CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization
CLASP localizes dense audio-visual events under weak supervision by identifying cross-modal salient anchors and propagating their semantics along the timeline.
-
Towards Open-Vocabulary Audio-Visual Event Localization
An ImageBind-based fine-tuned model outperforms a training-free zero-shot baseline on the new OV-AVEBench, which spans 67 event classes with 21 unseen at test time.
-
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
CCNet combines cross-modal consistency and multi-temporal granularity modules to achieve state-of-the-art dense audio-visual event localization on UnAV-100.
-
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
A class-aware feature decoupling module with a background class plus co-occurrence and local-global fusion blocks improves weakly-supervised audio-visual video parsing.
Discussion (0). Continue with ORCID to comment.