Pith. sign in

REVIEW 5 major objections 4 minor 86 references

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper introduces OmniAVS, a benchmark of 61,095 omnimodal referring expressions for audio-visual segmentation, and OISA, a multimodal LLM that reaches 41.1% J&F on it, beating prior best by 5.0 points.

desk verdict New omnimodal referring dataset is a real contribution, but the missing audio-ablation control leaves its central audio-understanding claim unproven. read the letter →

arxiv 2507.22886 v2 pith:ODETY2BB submitted 2025-07-30 cs.CV

classification cs.CV
keywords referringaudio-visualsegmentationomnimodalexpressionsreasoningmultimodallargelanguagemodelinterleavingquerypropagationdatasetbenchmarkexplanationgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that referring audio-visual segmentation should move beyond detecting obvious sound sources toward understanding the content of sounds and performing reasoning over multimodal cues. To make this concrete, it introduces OmniAVS, a dataset of 2,104 videos with 61,095 expressions that flexibly combine text, speech, sound, and images in eight modality types, requiring skills like inferring illness from coughing or locating an object by matching a sound and an image. It also presents OISA, a multimodal large language model that handles these omnimodal expressions and outputs both a segmentation mask and a natural-language explanation. On OmniAVS, OISA-1B achieves 41.1% average J&F, outperforming the strongest prior method LISA-13B by 5.0 points, and it also reports competitive results on existing referring segmentation benchmarks.

What carries the argument

The load-bearing mechanisms are Audio-Visual Interleaving and Query Propagation. Audio-Visual Interleaving divides the audio token sequence into clips and places each clip immediately after its corresponding frame's vision tokens, forming a synchronized sequence of the form $\{v_1, a_1, v_2, a_2, \ldots, v_N, a_N\}$ without adding parameters. Query Propagation updates the [SEG] token frame-by-frame in the mask decoder, rather than using one fixed token for all frames, so the query tracks object motion and avoids identity switches. The [SEG] token, produced by the MLLM from the interleaved multimodal context, is fed into the mask head for segmentation.

What would settle it

Re-annotate a random 5% of OmniAVS test expressions with an independent annotation team and measure inter-annotator agreement on the referred-object masks; if agreement falls near the 5.0-point J&F gap between OISA-1B and LISA-13B, the ranking would not be trustworthy. Alternatively, rerun OISA-1B with the audio and video tokens interleaved in random order instead of the frame-aligned order; if J&F does not drop substantially, the claimed benefit of Audio-Visual Interleaving is not load-bearing.

Watch

Extended reading notes

Core claim

The central claim is that a new benchmark, OmniAVS, can push referring audio-visual segmentation from surface-level acoustic attributes to semantic content understanding and reasoning, and that a multimodal LLM-based model can be adapted to this harder task. OISA-1B accomplishes this by interleaving audio and visual tokens for temporal synchronization and by propagating a single segmentation query across frames, achieving state-of-the-art results on OmniAVS and strong transfer to related referring and reasoning segmentation tasks.

Load-bearing premise

The benchmark's reliability rests on the assumption that its 61,095 expressions and the associated mask annotations are clean and unambiguous; the paper reports no inter-annotator agreement or quality checks, so if those labels are noisy, the reported rankings and difficulty comparisons could change.

Editorial extensions

If this is right

  • If OmniAVS becomes a standard benchmark, referring segmentation evaluation will include audio-content reasoning and explanation quality, not just acoustic event detection.
  • The eight-expression-type interface could push development of models that accept arbitrary combinations of text, speech, sound, and image as referring input, closer to human interaction.
  • Audio-Visual Interleaving and Query Propagation are architecture-agnostic enough to be incorporated into other MLLM-based segmentation systems, potentially improving video-level referring segmentation broadly.
  • The explanations provided for reasoning expressions enable quantifying a model's interpretability, a dimension absent from prior referring audio-visual benchmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because speech expressions are generated by converting text via TTS, there may be a gap between these utterances and natural spontaneous speech; a follow-up could test whether OISA's gains persist with human-spoken expressions.
  • The full-temporal mask convention, inherited from MeViS, may be ill-matched to expressions that are only valid in part of a video; a temporally-gated evaluation variant could reveal whether models locate objects during the relevant segment.
  • The 5.0-point gain over LISA-13B may partly stem from the audio-text alignment pretraining stage rather than the interleaving mechanism itself; ablating that stage would clarify which contribution is decisive.
  • The benchmark's emphasis on sound-content reasoning connects naturally to audio-visual question answering, so a joint model fine-tuned on both OmniAVS and A-VQA may improve both tasks through shared reasoning supervision.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper proposes OmniAVS, a new referring audio-visual segmentation dataset with 2,104 videos and 61,095 referring expressions spanning eight modality combinations (text/speech with sound and/or image). The authors argue that existing RAVS datasets such as Ref-AVS rely on surface-level acoustic properties, whereas OmniAVS expressions require understanding audio content and performing reasoning, such as inferring illness from coughing sounds. The paper also introduces OISA-1B, an MLLM-based baseline with two technical components: audio-visual interleaving for temporal alignment and query propagation for mask decoding. Experiments report state-of-the-art results on OmniAVS (41.1% J&F, outperforming LISA-13B by 5.0 points) and competitive results on Ref-AVS, referring image/video segmentation, and ReVOS.

Significance. If the dataset annotations are reliable, OmniAVS is a useful resource that moves referring audio-visual segmentation from sound-presence cues toward audio-content understanding and reasoning. The OISA design choices, audio-visual interleaving and query propagation, are simple, parameter-free, and show consistent gains in the in-model ablations. The evaluation is performed on a held-out test split and does not rely on fitted constants, so the main benchmark claim is not circular. However, the absence of an audio-ablated control leaves the core claim that OmniAVS expressions demand audio-content understanding empirically unsupported, and the benchmark's reliability is not yet established because annotation quality is not measured. For these reasons the contribution is promising but not fully established.

major comments (5)
  1. [§5.2–5.3 (Tables 3 and 5)] No experiment removes or masks the video audio stream. All fusion ablations in Table 3 vary only how real audio tokens are combined, and all benchmark conditions in Table 5 feed audio into the model. Because many OmniAVS expressions are semantically redundant with text and vision (e.g., 'Who is most likely to be sick?' is inferable from visible coughing, and 'The dog warning' names the sound in the text), a vision-language model without audio understanding could plausibly achieve much of the reported 41.1% J&F. An audio-ablated control (e.g., replacing audio tokens with silence or zeros, evaluated for both OISA and adapted LISA) is needed to support the paper's central claim that OmniAVS demands audio-content understanding beyond sound-presence detection.
  2. [§3.2] The reliability of OmniAVS as a benchmark depends on annotation quality, but Section 3.2 reports no inter-annotator agreement, no double-annotation rate, and no quality-control metric for the 61,095 expressions or the 206k mask labels. The expression rules (e.g., 'emphasize the sound's content') leave room for subjective judgment, and annotator use of SAM2 assistance does not by itself validate mask correctness. Please report agreement statistics (e.g., mask IoU between annotators and expression-validity agreement) on a sample, and state how ambiguous or failing annotations were resolved.
  3. [§5.3 (Table 5)] The LISA baseline is enhanced only with the same audio-text alignment as OISA, while OISA additionally uses audio-visual interleaving and query propagation. The 5.0-point advantage over LISA-13B therefore conflates the proposed architectural components with the ability to reason about audio content. To support the claim that OISA outperforms existing methods on omnimodal reasoning, the comparison should give LISA the same interleaving and query-propagation benefits, or isolate each component's contribution on the OmniAVS test set.
  4. [§5.3–5.4 (Tables 5–8)] All reported metrics are single-run point estimates with no variance or significance testing. This matters for claims such as the 5.0-point gain over LISA-13B and the 0.2-point improvement over VISA on ReVOS; without error bars or multiple seeds, small differences may not be reproducible. Please report standard deviations over multiple runs, or at least provide evidence that the main conclusions are stable under different random seeds.
  5. [§5.4 (Table 6)] The discussion of the Ref-AVS Null split is speculative: the claim that EEMC's 0.7% S score 'likely' reflects overfitting to audio patterns rather than genuine null understanding is not tested. Because OISA-1B's S=9.8% is substantially worse on this split, the paper should either provide an error analysis supporting the overfitting explanation or qualify the claim of 'greatly surpassing' EEMC on Ref-AVS.
minor comments (4)
  1. [§5.3] The sentence 'splits VII (text+speech+image) and VIII (text+sound+image)' mislabels the expression types: according to Section 3.2, type VII is text+sound+image and type VIII is speech+sound+image.
  2. [§5.1] The terms 'dense frames' and 'sparse frames' are used without definition; please clarify how the dense and sparse frame subsets are selected and how they differ during training.
  3. [§5 (Evaluation Metrics)] For no-target expressions, J&F is set to 1 when the prediction is empty and 0 otherwise; this gives full credit for predicting an empty mask, which could inflate scores. Please report the proportion of no-target expressions in OmniAVS and the sensitivity of the overall J&F to this convention.
  4. [§4.2] Audio tokens are divided uniformly across the N sampled frames, but OmniAVS videos have annotation frame rates of 3–15 FPS; please clarify whether the audio segmentation uses actual timestamps or a simplified uniform division, and discuss the effect on audio-visual alignment for variable-FPS videos.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: benchmark and method are evaluated on a held-out test split and against external benchmarks.

full rationale

The paper has two contributions: the OmniAVS dataset and the OISA method. The central numerical claims (OISA-1B reaching 41.1% J&F on OmniAVS, outperforming LISA-13B by 5.0 points, and 58.0% J&F on Ref-AVS) are produced by training on training splits and evaluating on held-out test splits; no parameter is fitted to the test set and no reported metric is a renamed training objective. The audio-content emphasis of OmniAVS is built through annotation rules in Section 3.2, not derived from OISA's outputs, and OISA's results do not define the benchmark's content. Citations to prior work by the same authors, such as MeViS for the full-temporal mask protocol or MOVE, MOSEv2, and MMT-Bench as related baselines, are methodological or bibliographic and do not carry the derivation of any reported result. The absence of an audio-ablated control in Section 5 is a valid experimental concern about whether OISA's scores prove audio-content understanding, but it is not circularity: a missing control does not make the result equivalent to its input. The paper is also self-contained against external benchmarks (Ref-AVS, MeViS, RefCOCO, ReVOS), further supporting the independence of the method's assessment. No load-bearing step reduces, by construction or by self-citation, to the paper's own inputs.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central contribution is a dataset and a trained baseline; there are no mathematical derivations. Free parameters are training choices rather than fitted constants. The load-bearing assumptions concern annotation quality, evaluation conventions, and the audio-visual alignment model.

free parameters (1)
  • Frame sampling counts = 10 training frames, 32 inference frames, 4 dense frames
    Chosen in Section 5.1; the audio-visual interleaving splits audio into a matching number of clips, so these values drive the alignment assumption and evaluation cost.
assumptions (4)
  • domain assumption Uniform frame sampling and interleaved audio clips preserve audio-visual synchronization.
    Section 4.2: audio is split into N clips and aligned with N sampled frames; this assumes frame-synchronous alignment is sufficient for reasoning, which may fail for off-screen sound sources.
  • domain assumption Full-temporal masks are valid when an object only partially matches an expression.
    Section 3.2 footnote: follows MeViS to segment all frames even when objects only partially match expressions; this shapes evaluation labels.
  • domain assumption The [SEG] token output by the LLM can be used directly as a mask query.
    Section 4.3: after LISA, the [SEG] token is passed to the mask decoder; assumes the LLM embeds the target identity in a single token.
  • ad hoc to paper Videos selected for informative audio and complex scenes represent the target task.
    Section 3.1: selection criteria favor informative audio and complex scenes, so statistics may not generalize to arbitrary audio-visual content.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation." pith.science (2026). https://pith.science/paper/ODETY2BB

@misc{pith2026250722886,
  author       = {Pith},
  title        = {Pith review of: Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ODETY2BB}},
  note         = {Machine review of arXiv:2507.22886}
}
read the original abstract

Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propose Omnimodal Referring Audio-Visual Segmentation (OmniAVS), a new dataset containing 2,104 videos and 61,095 multimodal referring expressions. OmniAVS stands out with three key innovations: (1) 8 types of multimodal expressions that flexibly combine text, speech, sound, and visual cues; (2) an emphasis on understanding audio content beyond just detecting their presence; and (3) the inclusion of complex reasoning and world knowledge in expressions. Furthermore, we introduce Omnimodal Instructed Segmentation Assistant (OISA), to address the challenges of multimodal reasoning and fine-grained understanding of audiovisual content in OmniAVS. OISA uses MLLM to comprehend complex cues and perform reasoning-based segmentation. Extensive experiments show that OISA outperforms existing methods on OmniAVS and achieves competitive results on other related tasks.

Figures

Figures reproduced from arXiv: 2507.22886 by the authors.

Figure 1
Figure 1. Examples of the proposed benchmark Omnimodal Referring Audio-Visual Segmentation (OmniAVS) to show its nature and flexibility. OmniAVS introduces 3 key features: 1) It supports diverse multimodal referring expressions that flexibly combine text , speech , sound , and image for referring audio-visual segmentation; 2) It emphasizes understanding the content of audio rather than merely hearing them; 3) It incorporates … view at source ↗
Figure 2
Figure 2. Comparison of (a) Ref-AVS [68] and (b) our proposed OmniAVS. In Ref-AVS, expressions mainly focus on surface-level sound properties, while OmniAVS demands deeper understanding and reasoning about sound content, enabling complex reasoning like identifying potential illness from coughing sounds. inputs can greatly enhance flexible referring and human￾machine interactions. However, existing datasets [17, 68] are limite… view at source ↗
Figure 3
Figure 3. Distribution of expression types. SI: sound+image. Text 25,474 Sound 1,738 Image 1,731 SI 1,553 Speech 25,472 Sound 1,755 SI 1,561 Text Speech Image 1,811 OmniAVS 61,095 the annotation team is required to carefully confirm the corresponding object in the video, then track and anno￾tate the object’s mask in all frames † . We designed an interactive annotation tool to automatically load the video and corresponding obj… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Overview of the proposed method OISA. Vision/audio encoders and text embedding are omitted for clarity. integrated into the corresponding text tokens to produce the final Omnimodal Expression Tokens. All types of tokens are subsequently fed into the MLLM. The MLLM gene…
Figure 5
Figure 5. Figure 5: Comparision of different mask decoder forms. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Success and failure cases of OISA-1B [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

86 extracted references · 71 canonical work pages

  1. [1]

    Qwen Technical Report

    Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen Technical Report. arXiv, 2023. 6

  2. [2]

    Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities. arXiv preprint arXiv:2308.12966 ,

  3. [3]

    One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

    Zechen Bai, Tong He, Haiyang Mei, Pichao W ANG, Ziteng Gao, Joya Chen, liulei, Zheng Zhang, and Mike Zheng Shou. One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos. In Adv. Neural Inform. Process. Syst., 2024. 2, 3, 6, 7, 8

  4. [4]

    METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

    Satanjeev Banerjee and Alon Lavie. METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments. In Assoc. Comput. Linguist. Worksh., 2005. 6

  5. [5]

    End-to-End Referring Video Object Segmentation with Mul- timodal Transformers

    Adam Botach, Evgenii Zheltonozhskii, and Chaim Baskin. End-to-End Referring Video Object Segmentation with Mul- timodal Transformers. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 8

  6. [6]

    Auditory Scene Analysis: The Perceptual Organization of Sound

    Albert S Bregman. Auditory Scene Analysis: The Perceptual Organization of Sound. MIT press, 1994. 2

  7. [7]

    COCO- Stuff: Thing and Stuff Classes in Context

    Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. COCO- Stuff: Thing and Stuff Classes in Context. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 6

  8. [8]

    TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition

    Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, and Dima Damen. TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 3

Show all 86 references
  1. [9]

    GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

    Guoguo Chen, Shuzhou Chai, et al. GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio. In Proc. Interspeech 2021, 2021. 6

  2. [10]

    VGGSound: A Large-scale Audio-Visual Dataset

    Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zis- serman. VGGSound: A Large-scale Audio-Visual Dataset. In IEEE Int. Conf. Acoust. Speech Signal Process., 2020. 3

  3. [11]

    Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts

    Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille. Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts. In IEEE Conf. Comput. Vis. Pattern Recog.,

  4. [12]

    Vision Transformer Adapter for Dense Predictions

    Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision Transformer Adapter for Dense Predictions. In Int. Conf. Learn. Represent., 2023. 5, 6

  5. [13]

    How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open- Source Suites

    Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open- Source Suites. arXiv preprint arXiv:2404.16821 , 2024. 2, 3, 6, 7

  6. [14]

    Masked-attention Mask Transformer for Universal Image Segmentation

    Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention Mask Transformer for Universal Image Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 5, 6

  7. [15]

    VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs

    Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs. arXiv preprint arXiv:2406.07476, 2024. 5, 7

  8. [16]

    Qwen2-Audio Technical Report

    Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. Qwen2-Audio Technical Report. arXiv preprint arXiv:2407.10759, 2024. 3

  9. [17]

    MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

    Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy. MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions. In Int. Conf. Comput. Vis., 2023. 2, 3, 4, 6, 7, 8

  10. [18]

    MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

    Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Philip HS Torr, and Song Bai. MOSE: A New Dataset for Video Object Segmentation in Complex Scenes. InInt. Conf. Comput. Vis., 2023. 6

  11. [19]

    Multimodal referring segmentation: A survey

    Henghui Ding, Song Tang, Shuting He, Chang Liu, Zuxuan Wu, and Yu-Gang Jiang. Multimodal referring segmentation: A survey. arXiv, 2025. 2

  12. [20]

    MOSEv2: A more challenging dataset for video object segmentation in complex scenes

    Henghui Ding, Kaining Ying, Chang Liu, Shuting He, Yu- Gang Jiang, Philip HS Torr, and Song Bai. MOSEv2: A more challenging dataset for video object segmentation in complex scenes. arXiv, 2025. 6

  13. [21]

    VITA: Towards Open-Source Interactive Omni Multimodal LLM

    Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Xiong Wang, Di Yin, Long Ma, Xiawu Zheng, et al. VITA: Towards Open-Source Interactive Omni Multimodal LLM. arXiv, 2024. 3, 4, 6

  14. [22]

    A VSegFormer: Audio-Visual Segmentation with Transformer

    Shengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang, and Tong Lu. A VSegFormer: Audio-Visual Segmentation with Transformer. In AAAI, 2024. 8

  15. [23]

    https://github.com/RVC- Boss/ GPT-SoVITS, 2024

    GPT-SoVITS. https://github.com/RVC- Boss/ GPT-SoVITS, 2024. 4

  16. [24]

    Open- V ocabulary Audio-Visual Semantic Segmentation

    Ruohao Guo, Liao Qu, Dantong Niu, Yanyu Qi, Wenzhen Yue, Ji Shi, Bowei Xing, and Xianghua Ying. Open- V ocabulary Audio-Visual Semantic Segmentation. In ACM Int. Conf. Multimedia, 2024. 2

  17. [25]

    Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception

    Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu, Jun- Yan He, Jin-Peng Lan, Bin Luo, and Xuansong Xie. Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception. In IEEE Conf. Comput. Vis. Pattern Recog. ,

  18. [26]

    Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation

    Shuting He and Henghui Ding. Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2

  19. [27]

    A Generalized Framework for Video Instance Segmentation

    Miran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim, Seoung Wug Oh, Joon-Young Lee, and Seon Joo Kim. A Generalized Framework for Video Instance Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 6

  20. [28]

    Deep clustering: Discriminative embeddings for segmentation and separation

    John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe. Deep clustering: Discriminative embeddings for segmentation and separation. In IEEE Int. Conf. Acoust. Speech Signal Process., 2016. 7

  21. [29]

    LoRA: Low-Rank Adaptation of Large Language Models

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In Int. Conf. Learn. Represent., 2022. 6 9

  22. [30]

    Egocentric Audio-Visual Object Localization

    Chao Huang, Yapeng Tian, Anurag Kumar, and Chenliang Xu. Egocentric Audio-Visual Object Localization. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 3

  23. [31]

    Video Object Segmentation with Language Referring Expressions

    Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video Object Segmentation with Language Referring Expressions. In ACCV, 2019. 2, 6, 8

  24. [32]

    Segment Anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment Anything. In Int. Conf. Comput. Vis., 2023. 3, 5

  25. [33]

    LISA: Reasoning Segmen- tation via Large Language Model

    Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmen- tation via Large Language Model. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2, 3, 5, 6, 7, 8

  26. [34]

    TVQA: Localized, Compositional Video Question Answer- ing

    Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. TVQA: Localized, Compositional Video Question Answer- ing. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2018. 3, 7

  27. [35]

    Learning to Answer Questions in Dynamic Audio-Visual Scenarios

    Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji- Rong Wen, and Di Hu. Learning to Answer Questions in Dynamic Audio-Visual Scenarios. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 3

  28. [36]

    Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

    Guangyao Li, Henghui Du, and Di Hu. Boosting Audio Visual Question Answering via Key Semantic-Aware Cues. In ACM Int. Conf. Multimedia, pages 5997–6005, 2024

  29. [37]

    Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025

    Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo, and Yapeng Tian. Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025. 3

  30. [38]

    Robust Referring Video Object Segmentation with Cyclic Structural Consensus

    Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, and Yan Lu. Robust Referring Video Object Segmentation with Cyclic Structural Consensus. In Int. Conf. Comput. Vis.,

  31. [39]

    Baichuan-Omni Technical Report

    Yadong Li, Haoze Sun, Mingan Lin, Tianpeng Li, Guosheng Dong, Tao Zhang, Bowen Ding, Wei Song, Zhenglin Cheng, Yuqi Huo, et al. Baichuan-Omni Technical Report. arXiv preprint arXiv:2410.08565, 2024. 3

  32. [40]

    Losh: Long-short text joint prediction network for referring video object segmentation

    Linfeng Yuan and Miaojing Shi and Zijie Yue and Qijun Chen. Losh: Long-short text joint prediction network for referring video object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2

  33. [41]

    GRES: Generalized Referring Expression Segmentation

    Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized Referring Expression Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 5, 6, 8

  34. [42]

    Primitivenet: decomposing the global constraints for referring segmenta- tion

    Chang Liu, Xudong Jiang, and Henghui Ding. Primitivenet: decomposing the global constraints for referring segmenta- tion. Visual Intelligence, 2024. 2

  35. [43]

    Visual Instruction Tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In Adv. Neural Inform. Process. Syst., 2023. 3

  36. [44]

    Improved Baselines with Visual Instruction Tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved Baselines with Visual Instruction Tuning. InIEEE Conf. Comput. Vis. Pattern Recog., 2024. 7

  37. [45]

    ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

    Shuo Liu, Kaining Ying, Hao Zhang, Yue Yang, Yuqi Lin, Tianle Zhang, Chuanhao Li, Yu Qiao, Ping Luo, Wenqi Shao, et al. ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models. In Adv. Neural Inform. Proc...

  38. [46]

    Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration

    Yi Luo and Nima Mesgarani. Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration. IEEE/ACM Trans. Audio Speech Lang. Process., 27 (8), 2019. 7

  39. [47]

    Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation

    Juncheng Ma, Peiwen Sun, Yaoting Wang, and Di Hu. Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation. In Eur. Conf. Comput. Vis.,

  40. [48]

    Generation and Comprehension of Unambiguous Object Descriptions

    Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and Comprehension of Unambiguous Object Descriptions. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 6, 8

  41. [49]

    V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation

    Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation. In IEEE Int. Conf. 3D Vis. ,

  42. [50]

    https://platform.openai.com/docs/ guides/text-to-speech, 2023

    OpenAI. https://platform.openai.com/docs/ guides/text-to-speech, 2023. 4

  43. [51]

    https://openai.com/index/hello- gpt-4o, 2024

    OpenAI. https://openai.com/index/hello- gpt-4o, 2024. 2, 3

  44. [52]

    Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

    Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu, Xi- uqiang He, Zhigeng Pan, Lianli Gao, Jun Yu, Fei Wu, and Qi Tian. Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 4

  45. [53]

    DetGPT: Detect What You Need via Reasoning

    Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, et al. DetGPT: Detect What You Need via Reasoning. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2023. 2, 3

  46. [54]

    Robust Speech Recognition via Large-Scale Weak Supervision

    Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust Speech Recognition via Large-Scale Weak Supervision. InInt. Conf. Mach. Learn., 2023. 6

  47. [55]

    PACO: Parts and Attributes of Common Objects

    Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen, Baixue Zheng, Baishan Guo, Rui Wang, Aaron Mar- quez, Rama Kovvuri, Abhishek Kadian, et al. PACO: Parts and Attributes of Common Objects. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 6

  48. [56]

    SAM 2: Segment Anything in Images and Videos

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. SAM 2: Segment Anything in Images and Videos. arXiv preprint arXiv:2408.00714, 2024. 4

  49. [57]

    URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

    Seonguk Seo, Joon-Young Lee, and Bohyung Han. URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark. In Eur. Conf. Comput. Vis., 2020. 2, 4, 6, 8

  50. [58]

    video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

    Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, and Chao Zhang. video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models. In Int. Conf. Mach. Learn., 2024. 3, 5

  51. [59]

    Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning

    Luoyi Sun, Xuenan Xu, Mengyue Wu, and Weidi Xie. Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning. In ACM Int. Conf. Multimedia, 2024. 6 10

  52. [60]

    Unveiling and Mitigating Bias in Audio Visual Segmentation

    Peiwen Sun, Honggang Zhang, and Di Hu. Unveiling and Mitigating Bias in Audio Visual Segmentation. In ACM Int. Conf. Multimedia, 2024. 3

  53. [61]

    SALMONN: Towards Generic Hearing Abilities for Large Language Models

    Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. SALMONN: Towards Generic Hearing Abilities for Large Language Models. In Int. Conf. Learn. Represent., 2023. 3

  54. [62]

    Audio-Visual Event Localization in Unconstrained Videos

    Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chen- liang Xu. Audio-Visual Event Localization in Unconstrained Videos. In Eur. Conf. Comput. Vis., 2018. 3

  55. [63]

    Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv preprint arXiv:2409.12191, 2024. 3

  56. [64]

    Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation

    Yefei Wang, Kaili Wang, Yi Wang, Di Guo, Huaping Liu, and Fuchun Sun. Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation. InIEEE Int. Conf. Robot. Autom., 2022. 3

  57. [66]

    Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

    Yaoting Wang, Weisong Liu, Guangyao Li, Jian Ding, Di Hu, and Xi Li. Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer. In AAAI,

  58. [67]

    Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur

    Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang, and Di Hu. Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur. Conf. Comput. Vis., 2024

  59. [68]

    Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes

    Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li, Honggang Zhang, and Di Hu. Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes. In Eur. Conf. Comput. Vis.,

  60. [69]

    Language as Queries for Referring Video Object Segmentation

    Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, and Ping Luo. Language as Queries for Referring Video Object Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. ,

  61. [70]

    VISA: Reasoning Video Object Segmentation via Large Language Models

    Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, and Efstratios Gavves. VISA: Reasoning Video Object Segmentation via Large Language Models. In Eur. Conf. Comput. Vis., 2024. 2, 3, 4, 6, 8

  62. [71]

    Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

    Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen, Wei Zhang, Hongyang Li, Yu Qiao, Hao Dong, Zhongjiang He, and Peng Gao. Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation. In AAAI, 2024. 3, 7

  63. [72]

    A VQA: A Dataset for Audio-Visual Question Answering on Videos

    Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. A VQA: A Dataset for Audio-Visual Question Answering on Videos. In ACM Int. Conf. Multimedia, 2022. 3

  64. [73]

    LA VT: Language-Aware Vision Transformer for Referring Image Segmentation

    Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Heng- shuang Zhao, and Philip HS Torr. LA VT: Language-Aware Vision Transformer for Referring Image Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 8

  65. [74]

    Isda: Position-aware instance segmentation with deformable attention

    Kaining Ying, Zhenhua Wang, Cong Bai, and Pengfei Zhou. Isda: Position-aware instance segmentation with deformable attention. In IEEE Int. Conf. Acoust. Speech Signal Process.,

  66. [75]

    CTVIS: Consistent Training for Online Video Instance Segmentation

    Kaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang, Hao Chen, Lin Yuanbo Wu, Yifan Liu, Chengxiang Fan, Yunzhi Zhuge, and Chunhua Shen. CTVIS: Consistent Training for Online Video Instance Segmentation. In Int. Conf. Comput. Vis., 2023. 6

  67. [76]

    MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

    Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao. MMT-Bench: A Compr...

  68. [77]

    MOVE: Motion-guided few-shot video object segmentation

    Kaining Ying, Hengrui Hu, and Henghui Ding. MOVE: Motion-guided few-shot video object segmentation. In Int. Conf. Comput. Vis., 2025. 6

  69. [78]

    Modeling Context in Referring Expressions

    Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling Context in Referring Expressions. In Eur. Conf. Comput. Vis., 2016. 6, 8

  70. [79]

    MOTR: End-to-End Multiple-Object Tracking with Transformer

    Fangao Zeng, Bin Dong, Yuang Zhang, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. MOTR: End-to-End Multiple-Object Tracking with Transformer. In Eur. Conf. Comput. Vis., 2022. 6

  71. [80]

    Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

    Hang Zhang, Xin Li, and Lidong Bing. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2023. 5

  72. [81]

    LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention

    Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention. In Int. Conf. Learn. Represent., 2024. 3

  73. [82]

    GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

    Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Yu Liu, Kai Chen, and Ping Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. arXiv preprint arXiv:2307.03601, 2023. 3

  74. [83]

    DVIS: Decoupled Video Instance Segmentation Framework

    Tao Zhang, Xingye Tian, Yu Wu, Shunping Ji, Xuebo Wang, Yuan Zhang, and Pengfei Wan. DVIS: Decoupled Video Instance Segmentation Framework. In Int. Conf. Comput. Vis., 2023. 6

  75. [84]

    ViLLa: Video Reasoning Segmentation with Large Language Model

    Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang, Kun Wang, Yu Qiao, and Hengshuang Zhao. ViLLa: Video Reasoning Segmentation with Large Language Model. arXiv preprint arXiv:2407.14500, 2024. 2, 6

  76. [85]

    Scene Parsing through ADE20K Dataset

    Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene Parsing through ADE20K Dataset. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 6

  77. [86]

    Audio-Visual Segmentation

    Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong. Audio-Visual Segmentation. In Eur. Conf. Comput. Vis., 2022. 2, 3, 8

  78. [87]

    Tracking with Human-Intent Reasoning

    Jiawen Zhu, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li, Bin Luo, Huchuan Lu, Yifeng Geng, and Xuansong Xie. Tracking with Human-Intent Reasoning. arXiv, 2023. 8 11

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.