Pith. sign in

REVIEW 5 cited by

Audio-Visual Instance Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.18709 v4 pith:GFOYPZZ7 submitted 2023-10-28 cs.CV cs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.LGcs.MMcs.SDeess.AS
keywords audio-visualavisinstancemulti-modalavisegmodelproposesegmentation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To facilitate this research, we introduce a high-quality benchmark named AVISeg, containing over 90K instance masks from 26 semantic categories in 926 long videos. Additionally, we propose a strong baseline model for this task. Our model first localizes sound source within each frame, and condenses object-specific contexts into concise tokens. Then it builds long-range audio-visual dependencies between these tokens using window-based attention, and tracks sounding objects among the entire video sequences. Extensive experiments reveal that our method performs best on AVISeg, surpassing the existing methods from related tasks. We further conduct the evaluation on several multi-modal large models. Unfortunately, they exhibits subpar performance on instance-level sound source localization and temporal perception. We expect that AVIS will inspire the community towards a more comprehensive multi-modal understanding. Dataset and code is available at https://github.com/ruohaoguo/avis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Patch-level Sounding Object Tracking for Audio-Visual Question Answering

    cs.MM 2024-12 conditional novelty 7.0 of 10

    A new patch-level sounding object tracking method with motion-, sound-, and question-driven graph modules achieves 78.42% average accuracy on MUSIC-AVQA, competitive with large-scale pretraining approaches.

  2. CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CLASP localizes dense audio-visual events under weak supervision by identifying cross-modal salient anchors and propagating their semantics along the timeline.

  3. Towards Open-Vocabulary Audio-Visual Event Localization

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An ImageBind-based fine-tuned model outperforms a training-free zero-shot baseline on the new OV-AVEBench, which spans 67 event classes with 21 unseen at test time.

  4. Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration

    cs.CV 2024-12 conditional novelty 5.0 of 10

    CCNet combines cross-modal consistency and multi-temporal granularity modules to achieve state-of-the-art dense audio-visual event localization on UnAV-100.

  5. Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A class-aware feature decoupling module with a background class plus co-occurrence and local-global fusion blocks improves weakly-supervised audio-visual video parsing.

Pith tools