Pith. sign in

Paper Citation Record · LEDGER

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

As of 18 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2507.22886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22886 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:10.475961Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:08.927215Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ce8bf37-9ae5-49f7-b3e4-78a3253ac6ea · outbound

This paper cites Qwen Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.019397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.019397Z digest=sha256:a1fd467f17030fd544553473f5a3ae19b01b055ef1d2a22648315d298bfa024a

Observation 2f00d2a4-c04b-4525-9a71-c51d555eb531 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.094700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.094700Z digest=sha256:335b8b213e476656283bdc69fa0de52f8c89f14d6ee830d9f1ca0deb896c28fd

Observation a5525f58-491f-4f40-93f7-0b2218537b1b · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.158565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.158565Z digest=sha256:12266cdaa1331972a88bc38721fe421bfa5a421dbcfba5901b36acaf30ae0bf0

Observation 022d88cb-68fc-4714-ac2f-f3c71fb547e8 · outbound

This paper cites METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.221846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.221846Z digest=sha256:1385323c34d2b155cd0c86552fecb988016d60f0e27a3cdf92b5ddd42b08cfc2

Observation 777041a2-00ee-4a1a-9dc6-3c07deffeaeb · outbound

This paper cites End-to-End Referring Video Object Segmentation with Mul- timodal Transformers.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation End-to-End Referring Video Object Segmentation with Mul- timodal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.274355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.274355Z digest=sha256:70ab82e8f5f5574b1d9a6326631bf70e9242b5647bea1111678781835c2cd122

Observation e8e3eb20-6961-4062-a322-8376d74a004c · outbound

This paper cites Auditory Scene Analysis: The Perceptual Organization of Sound.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auditory Scene Analysis: The Perceptual Organization of Sound

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.346121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.346121Z digest=sha256:49a5a31eee1aea45369775c55b8be1f76fb8d961718068b5fe42f9cbbc514fbb

Observation 9784a756-a69b-4dc8-8bc6-cbd6e11ed1b5 · outbound

This paper cites COCO- Stuff: Thing and Stuff Classes in Context.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation COCO- Stuff: Thing and Stuff Classes in Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.422372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.422372Z digest=sha256:b9efde8a8f6246d332067baf325f8bcc6b0dd76d1c4657910620c54cf29fb350

Observation e53caa16-5819-4119-b0bd-d8051b52809f · outbound

This paper cites TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.485367Z digest=sha256:e34f75c51002420a04dd9f82b6f7651fd14ee8702b0555a618f9cf79410efa39

Observation c505fc3d-73be-4808-9367-b907a3ce0dc7 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.252941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.596129Z digest=sha256:0fc11fffe06ee1680a2fe26c411524d854c086c562cf6f0270ccc86c09cfeb8e

Observation 1ae1bae4-45c0-424a-8d2f-ee1031106421 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VGGSound: A Large-scale Audio-Visual Dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.232394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.636040Z digest=sha256:99da37282125306ddae2b969b3ce87b845f48effe4a47d8319937792d3affbcf

Observation 042946cb-fc4c-4e03-a6e0-73db51bb935a · outbound

This paper cites Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.210671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.696935Z digest=sha256:b3e467759865f4761efef665853a72b35c7f631675aed0459825927c5ba8e6b8

Observation 646c8340-d9cb-4d95-8eeb-802e3e114d5e · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Vision Transformer Adapter for Dense Predictions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.190592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.774702Z digest=sha256:92ba5b89790d1f652661910a1f01873790e46ecde442f5fdf181e1a6c340231b

Observation 76f515e9-2c19-402f-a1c4-383e0f99abb1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.856376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.856376Z digest=sha256:2be9fb4e4e98ad34ffa350e60b187fdee9fe3b7ca6dad4b08ac56696f70fb7e7

Observation 1e824bbf-bee8-4cc9-9ad0-a5e34c2278c8 · outbound

This paper cites Masked-attention Mask Transformer for Universal Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Masked-attention Mask Transformer for Universal Image Segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.171004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:06.974243Z digest=sha256:2217e6e11310549e6b3854d98342f5a8ae348f1d0bdab48f0505a63dbb1cffb7

Observation e2e4208a-25f8-4a3c-bd42-3fac05f3a85d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.045271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.045271Z digest=sha256:96e2e01a7aadb34732caf32954fc2327a9c42479318144dc40ab0dba4b8d9213

Observation b10fd106-834c-4e6d-8b7b-ad7c8fdc633b · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-Audio Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.101084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.101084Z digest=sha256:3ea6710472a156b6e7f8bad95b574308d0e7d1efebafc83f23e39fc35d0d87eb

Observation f6258804-80df-43ff-9a14-46e88b13db52 · outbound

This paper cites MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.286988Z digest=sha256:90148fb66ed0286eeeefbaac732e9350fe5d611d10e935da3030afcd9e7b5fa0

Observation 01f7c39d-536f-4319-b801-04d6ef0370b8 · outbound

This paper cites MOSE: A New Dataset for Video Object Segmentation in Complex Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.128980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.373963Z digest=sha256:6455b9f9fcffee85618b224120f38c074e18f40edf6b5a5228571a96c13dac2d

Observation 4f1abf28-9b01-48f0-80fc-2db69cc2c347 · outbound

This paper cites Multimodal referring segmentation: A survey.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multimodal referring segmentation: A survey

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.110276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.468288Z digest=sha256:26ee1779e6c64cb2db1e6c06a63385d99cf5fd7251bbae17ff602a51926a7570

Observation 6c58ef9c-39ad-43f1-91f7-91a7b233dcde · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.089024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.555289Z digest=sha256:66dc58abc5c04cec689d8b932344cc1cd17f95ce0758b4e100732f4ab9ac98ef

Observation a47ca59b-9d8f-4b35-bf3b-54d1f65feaf9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.070253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.614968Z digest=sha256:21fcee06711e9ac11f9bb5aca45865bfb441f16d4dc54478d12b38ee97318f47

Observation 48a5fa52-7c80-4b18-8b04-ac3a66fa6e90 · outbound

This paper cites A VSegFormer: Audio-Visual Segmentation with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VSegFormer: Audio-Visual Segmentation with Transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.052365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.702971Z digest=sha256:64b87df1b42c3bd332b91ad877e312f7872a0499ef07a5372f65833b7f04ea36

Observation 4dd43122-d005-4916-bbd3-36838b48a4d8 · outbound

This paper cites https://github.com/RVC- Boss/ GPT-SoVITS, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://github.com/RVC- Boss/ GPT-SoVITS, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.032935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.890555Z digest=sha256:35920dbedd2b427f49ec3f2f577076f93a42c965f6a45c7798f9e2215cf0e6dd

Observation 405c891e-2575-4736-bdea-30422574e22c · outbound

This paper cites Open- V ocabulary Audio-Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Open- V ocabulary Audio-Visual Semantic Segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.016780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:07.963845Z digest=sha256:5e3e09a8ea42aed998ad1031438aad943a14fe8aa25c205c11b7615eaaede59f

Observation 2807b96b-cd94-4504-a8de-d82e0c18893f · outbound

This paper cites Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.996594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.053343Z digest=sha256:a47ca4977e1c6de4b1e7f544cdfa79f7c30c24cac5fecb74bd0a176e6591ca2c

Observation 1371932c-7134-4480-abe1-c9c9ad800365 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.977821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.146958Z digest=sha256:23fec64255887958f6843fba80512751360526f347f96427b0c9c8e8117d3f4f

Observation f6c70001-9a2d-426c-ad4f-06a8182e72a6 · outbound

This paper cites A Generalized Framework for Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A Generalized Framework for Video Instance Segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.959155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.229361Z digest=sha256:b1293c10d2370f12ac896994563cd0a969a98dbaca1d63aed478b51b6c85844b

Observation cd2dc9d9-8f31-40c9-a458-3c397f3283c0 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Deep clustering: Discriminative embeddings for segmentation and separation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.942633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.279887Z digest=sha256:cedcbace25a59258f98364c25765f191351695eb593dc63b22b49e7d2d95bd7d

Observation db461c65-20e1-4067-a559-4ffe04388662 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.922008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.397805Z digest=sha256:da06c01da09aea8683c8f3e13b3dbb6d5e41074bf6cd729b86f1d07847c2151b

Observation d1af3745-296f-45f2-8d5f-9cc30daa77ac · outbound

This paper cites Egocentric Audio-Visual Object Localization.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Egocentric Audio-Visual Object Localization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.905791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.609708Z digest=sha256:1d3695081f80813e76555888723f0fdc82378bc7e27e51ec8b8c591682428418

Observation 50a7f8a0-b7c0-4f1b-9ca8-a92f7895044c · outbound

This paper cites Video Object Segmentation with Language Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video Object Segmentation with Language Referring Expressions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.888913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.667631Z digest=sha256:53adf4ff5d24d85425125928e46d9e8ca4412d998608269485d0b01b627c03a9

Observation 672b7acd-ee7b-4197-8a0b-e39f025dd6dd · outbound

This paper cites Segment Anything.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Segment Anything

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.869618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.752455Z digest=sha256:1cfa6f324074089a025d2e5843eb6e3795880d1458d625869d0d7cdab63ee360

Observation e8dc9a84-77ca-4455-8bb8-8e76388a25f0 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LISA: Reasoning Segmen- tation via Large Language Model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.853787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.844260Z digest=sha256:3c6feb3bb00d57481d1f013e6185291cefb5a95a45e3f0aa1cc4007870abf73a

Observation 64d9a684-9cda-4be5-9d2a-d1b0722f00db · outbound

This paper cites TVQA: Localized, Compositional Video Question Answer- ing.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TVQA: Localized, Compositional Video Question Answer- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.837725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:08.989571Z digest=sha256:bfbf416ff1a7a95c293df536512d76ecdb094ab215ef827c88badc6f83e08226

Observation f176f74b-fda8-43ac-901b-0ec06487a5e0 · outbound

This paper cites Learning to Answer Questions in Dynamic Audio-Visual Scenarios.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.819872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.133635Z digest=sha256:5f67eb8ab0341fba249b594181d7e5e3c3277a4da3ab7b1909adff8df5c9b204

Observation f8a8afd9-e3f0-4389-9471-37c213797fff · outbound

This paper cites Boosting Audio Visual Question Answering via Key Semantic-Aware Cues.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.261515Z digest=sha256:fe266e2b359f31823101747a18116be0b5a87f799019bc12ad5d6e57f6c83186

Observation fcccefce-30db-4420-a1eb-d1a5452bf3e4 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.782524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.423301Z digest=sha256:cf4b5aa0fe5bf2159ac669cfe3068e1ba0fad2c87bc76f33d40e98da78324cb9

Observation 645602e0-6408-4407-97dc-4093b00fe8fb · outbound

This paper cites Robust Referring Video Object Segmentation with Cyclic Structural Consensus.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Referring Video Object Segmentation with Cyclic Structural Consensus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.764123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.565012Z digest=sha256:e3ca08c118de064cefe1e49155fbba13594d0a93dbd89effda8384c3dcee95d1

Observation c8b02ec2-3a3d-4696-8be3-73d21c9be3ea · outbound

This paper cites Baichuan-Omni Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Baichuan-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:09.667789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:09.667789Z digest=sha256:c2160d868c79ad281bd5c35cdc679ed61d102246cb4d1341653af056ef20a857

Observation d5aae2dc-9bb8-440b-9bc7-5ff61ddce81d · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.743164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.765828Z digest=sha256:b371229c77c8b2f71086813cf57b8667d04243d3254474c61c43e96ec8e23417

Observation 50923f62-0614-498b-9011-7e18564c6710 · outbound

This paper cites GRES: Generalized Referring Expression Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GRES: Generalized Referring Expression Segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.725625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.891803Z digest=sha256:5b425463fe3d2a786dc69b1fa428ab4f524a714e8d191f6459ef53e219598d3b

Observation 891931f1-a7d3-4608-bc87-9c6f7b102264 · outbound

This paper cites Primitivenet: decomposing the global constraints for referring segmenta- tion.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Primitivenet: decomposing the global constraints for referring segmenta- tion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.705305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:09.973694Z digest=sha256:ff76950d4c1bb4f452e7a37212b73a13b322d62b02e1c329d7df3dc85f1313f1

Observation bdfc6c6f-d5ba-4b18-93c0-8dfbcddcee92 · outbound

This paper cites Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Visual Instruction Tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.688159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.155107Z digest=sha256:935c21c32e4495b449bdb0a3accbd43b91cf24fc2137315667672b8152983e3a

Observation 1590031a-698e-41ca-b72f-fa229f64fbbc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Improved Baselines with Visual Instruction Tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.672272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.204533Z digest=sha256:9e216d765b060a67d64b39a71a13c274697ee0bda10b21eb838653c53b4daa61

Observation 10ab099c-43fb-47a7-a078-5903e288deae · outbound

This paper cites ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.654463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.210803Z digest=sha256:708b4532ac664717512b651653e5ce6b646a4bd828d212ad10b6a745de2569f8

Observation 8abb241e-01d6-4491-9e08-383ddf8654f1 · outbound

This paper cites Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.635223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.217190Z digest=sha256:02200dfe7f7ea4957e9133b82e353317e103eb7ba4187f630fe1cabfaad451da

Observation 015e1808-5f00-4085-98c3-22a9a3013810 · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.610896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.222759Z digest=sha256:6cc4c0790e4ba0df51a8c5dca73f1aa69c976971cbcef0bde66de1367e4608ec

Observation 9a25fee9-c7ce-46fd-a42f-0e562430d94a · outbound

This paper cites Generation and Comprehension of Unambiguous Object Descriptions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Generation and Comprehension of Unambiguous Object Descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.588517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.228825Z digest=sha256:bf769adb9ba30bf0e2280b50bb96035fee5824cf82b2367e9b4b5c280f42505a

Observation 0501c317-ebcf-452a-b993-8c1f8f8b4d2c · outbound

This paper cites V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.570411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.234547Z digest=sha256:f91b1da0f6fa8382ea34d5aa76c969eed108644b244c1496914dd96e3ced8cdb

Observation 6cbd765e-9f01-4245-a1ce-0c1c298c72f3 · outbound

This paper cites https://platform.openai.com/docs/ guides/text-to-speech, 2023.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://platform.openai.com/docs/ guides/text-to-speech, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.552583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.247239Z digest=sha256:8a31a6b60cacd1b5a3e48197fd8bef863cf21e25ca2169753ceb491e108de6ac

Observation 82d8109d-93ff-4236-b5d7-e40bc2dcf588 · outbound

This paper cites https://openai.com/index/hello- gpt-4o, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://openai.com/index/hello- gpt-4o, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.534588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.252563Z digest=sha256:8f4de53fca7abc0c241671047ef948a70871430c6613f2342ae5a9d249c02639

Observation 6a185d90-b484-4179-baa4-a6d681f6e7b2 · outbound

This paper cites Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.516684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.260334Z digest=sha256:da91ceaef2d68806c02cf1e811975cc92673ad06b723b2a2b3746f5a3670d769

Observation 1b467605-e0aa-40c9-bc9c-023ef0f3f312 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DetGPT: Detect What You Need via Reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.498976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.269021Z digest=sha256:d272eb8d3244e6651b03ecc770c23b108ac69e63536e254ba5d7d5440f097e44

Observation cbcbf9cf-d2f6-449e-bbdc-68cfc1239e0c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.480708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.278238Z digest=sha256:7b19bb4e8dcd75e91d0f647b204e04a0889c21aebe908d62fceb0d7be0f19159

Observation c563c37b-4b53-45a7-be18-589aece25f6f · outbound

This paper cites PACO: Parts and Attributes of Common Objects.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation PACO: Parts and Attributes of Common Objects

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.444625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.285742Z digest=sha256:a9221340cc0304facde7e23bf26b30fb3a4b1c59a10cfc0821e0f20db54f0cfb

Observation 6e4bfd0c-8c92-4dff-ba5d-033188d7a457 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SAM 2: Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.292503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.292503Z digest=sha256:0160b4790c34e9e0ad62c5aaaf2b9e2c2492a0f52622ce904f1de328523c345f

Observation 00c96345-f537-42d2-996d-8d48ef230711 · outbound

This paper cites URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.405620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.298834Z digest=sha256:8e4277a9adff46fee2783bf561c813f426b5e027b96e710cc12e3a123add2332

Observation ec9d5370-ca5d-47a2-a75d-17e77bc45517 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.364909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.304319Z digest=sha256:9c7abecd68528174dfc809d19f7db78793ebd2c59fc85b9d7aab69c45c475106

Observation 43f66742-91dd-46ec-8af0-56562b948142 · outbound

This paper cites Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.347843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.312650Z digest=sha256:6c1782597c4c32776277753111443a0d1e28bc5ec2f7d31c5bd0a5e1ed421dfa

Observation e9f020c3-53c3-4b7e-9352-6abb95a70ee2 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.330373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.319486Z digest=sha256:c432f17529f963a52ba8f708a9b15c70a8188d0c2b657c1ed22040fe69996ecc

Observation 10441679-27b3-40d4-ae73-4bb067d5a7d8 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.311783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.324756Z digest=sha256:827de1a945ccdf92714ac0a4f90d16dfb4a35a436daca5329687b7c637529c83

Observation d62a2be0-8151-4f94-8b91-b354b07c882c · outbound

This paper cites Audio-Visual Event Localization in Unconstrained Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Event Localization in Unconstrained Videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.290558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.330368Z digest=sha256:c6665efb911c3b964dc9b797875be2d1aea636958df5352148d41e47d953d551

Observation 2d7d6602-7a14-4756-bee0-279c1d915fb5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.336496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.336496Z digest=sha256:ec97ebe0d8db933cf99c0090da7ee736d399cb5ff12213ce487b65b8154c72eb

Observation 9dc3b0b3-a9e6-4331-bf5a-a7eaa1d6b5cb · outbound

This paper cites Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.271327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.341868Z digest=sha256:6317a0e9d6656a280593c0d2166761d373c47bae16255d60397ebbaebc0d356a

Observation 244a95f7-f82a-40e5-9e11-73152644fc18 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.253156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.357613Z digest=sha256:53876718a33b149d9e18c2ad8ddfc979ba9f2cad0e0c7376cc244430b1c530b5

Observation 977b99c5-f402-4b6c-a997-dce4e653e001 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.234846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.363089Z digest=sha256:257a9bc0733a50561f2c153ad228dcd6a9d5708d345165202d88aa17561e96bf

Observation 631e99e6-604d-4d89-ad36-31efa91f39f7 · outbound

This paper cites Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.369265Z digest=sha256:6265166d6a68e8912632fbfa29e647aeb80d0ef7b2ba0f1cd40359270037aa72

Observation 04473527-a83e-40a6-811e-0b61852742e8 · outbound

This paper cites Language as Queries for Referring Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Language as Queries for Referring Video Object Segmentation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.198186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.374194Z digest=sha256:acd8be95b1249f019c69b18e40e950c72a4d638ddc7454c793ea751a76e8dc6e

Observation 777ad773-26d4-453f-8f81-38b7847d8705 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.174877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.378825Z digest=sha256:835d5999d7462811c314045576d3e1b7e05a9f83a8279d6eeba4c1f65d76d51b

Observation 8116845e-6924-4180-b5be-c901cef114e4 · outbound

This paper cites Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.146807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.387330Z digest=sha256:1ff008642c72865ccb750ccfef1822f9c68d296b7dbac9251dded1a138b2c8e1

Observation 46080188-6cf1-486a-8db9-15479764098e · outbound

This paper cites A VQA: A Dataset for Audio-Visual Question Answering on Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VQA: A Dataset for Audio-Visual Question Answering on Videos

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.121804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.394479Z digest=sha256:11c75c253d50644d317879c59092bef69bad17d17a0c69664a5e197d76e16ed5

Observation d1978e52-dd29-4c52-b280-80f4a309769b · outbound

This paper cites LA VT: Language-Aware Vision Transformer for Referring Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LA VT: Language-Aware Vision Transformer for Referring Image Segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.095686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.399323Z digest=sha256:6310df145ecce610a4d51e72509c4d6f9ffa81a91da3316048409717fffbaab5

Observation 95031c31-a3e3-417d-93f3-0937d2a67d77 · outbound

This paper cites Isda: Position-aware instance segmentation with deformable attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Isda: Position-aware instance segmentation with deformable attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.064986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.409682Z digest=sha256:dca8203367e79989e302881342deed1d736f5c20d61d3bda41533721b54d5a67

Observation 585b8fae-a251-4094-b1d6-05e743ac25d6 · outbound

This paper cites CTVIS: Consistent Training for Online Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation CTVIS: Consistent Training for Online Video Instance Segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.044795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.414531Z digest=sha256:ddca3b80fa6db85faf25455d231254765e6baa8c22e1c24fae555f1887a38fbd

Observation 4478e543-28aa-49f9-8f7f-66ef2796a7c2 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.017438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.420373Z digest=sha256:d1d4a8b4a3c01fc2d72dd8fa3177601683c76253f39a17e457178ac863fd1a94

Observation 4abbf53c-1812-4652-8e2e-20332cfcfee0 · outbound

This paper cites MOVE: Motion-guided few-shot video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOVE: Motion-guided few-shot video object segmentation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.995052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.425968Z digest=sha256:26db4a15fe0f761a88d02ea3e2e19f1261d5553f5a8b70bb6eee0312108927d4

Observation e239833f-6bb0-4853-bbd8-114595adcdeb · outbound

This paper cites Modeling Context in Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Modeling Context in Referring Expressions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.964609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.430262Z digest=sha256:8c2a46cfcdea012f2e517aa995a60654065f1ca2b668003430fe4231f7866e06

Observation 3f0d94cf-bbe8-4685-8d81-ca5875f7e677 · outbound

This paper cites MOTR: End-to-End Multiple-Object Tracking with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOTR: End-to-End Multiple-Object Tracking with Transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.945246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.435156Z digest=sha256:c68429b60f55fb00feafd3d9c815c7f55cca91a07b2e9eb1a65fe5e7bd1c2d91

Observation ae1f22c0-d178-4c3b-ad7c-2f9fcd04f24e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.921392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.439995Z digest=sha256:d2afc65f57edad543f308e6d18260b62930fd1ecd6d4e735908b03e3e56531d8

Observation 60db074e-da48-4dee-ac7e-11891f91c7ef · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.902176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.444754Z digest=sha256:3c6047b85aaae4c8424846a42adbaa95d5e85f9c88caf2f31a2d09548c14f66a

Observation a24455e6-9922-4a95-9d94-7a0c5ac81fa1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.449340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.449340Z digest=sha256:97f6b58ddf8521c21b94bdfb3d405a45754eeec178f75755ae2524185b32a578

Observation 3dd6e515-9ec8-43cd-987f-fa1c1d08a13f · outbound

This paper cites DVIS: Decoupled Video Instance Segmentation Framework.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DVIS: Decoupled Video Instance Segmentation Framework

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.885373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.454902Z digest=sha256:43bfe302445bb48ec411461a3c022b40d546a8ce337573b7b20fc8f384e77dc2

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:0b2f1988a7fba2245009cfa3f4d40cf3dce14270502d17e531bef32d85c16032

Observation 77a47579-e561-4dec-99a8-67aeadf588f5 · outbound

This paper cites Scene Parsing through ADE20K Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Scene Parsing through ADE20K Dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.868533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.465447Z digest=sha256:60a42e33e20db4248121aba1a72a42ada7d37d9b7fc863eae9e2fa7be9c123e9

Observation b4440455-5ce6-4e3d-bcb1-fdbf7fc961df · outbound

This paper cites Audio-Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Segmentation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.852538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.470620Z digest=sha256:f90773998bb5b6a6d53fc113da473fc0fd3528e0c288948854f40cd298ca7ca5

Observation 68aa03c1-fccc-480d-9709-0703fb587fc1 · outbound

This paper cites Tracking with Human-Intent Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Tracking with Human-Intent Reasoning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.832815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T11:16:10.475961Z digest=sha256:6f9e31d7ec0b32cda00ecfc53f5c33bd55135aee934fd18ffeb0852c98bfe40c

Pith citing papers

Observation f4afe86d-fbd3-4785-a992-0802dc92b695 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.958133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:4198c9487ef9a6e01d90db6c48a9c0e90029c64bd578c414ff7374b9088d3174