Pith. sign in

Paper Citation Record · LEDGER

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

As of 9 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2507.22886.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22886 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:16:10.475961Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:35:37.095627Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:11:08.927215Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy71
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1ce8bf37-9ae5-49f7-b3e4-78a3253ac6ea · outbound

This paper cites Qwen Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.019397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.019397Z digest=sha256:a71b1d6a2b3f686897b3e55cddf8ee06fb6aaf7caa05941ae76ad2525af10ffb

Observation 2f00d2a4-c04b-4525-9a71-c51d555eb531 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.094700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.094700Z digest=sha256:3b936114258e2878c7e6a6b6b6740cb9b9caa1d3403dffec771be30a1b7e8533

Observation a5525f58-491f-4f40-93f7-0b2218537b1b · outbound

This paper cites One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.158565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.158565Z digest=sha256:8e8d1bfded2ea8ea122350e96fe2e7cc070f6b311b7f46fd894046989a44bc04

Observation 022d88cb-68fc-4714-ac2f-f3c71fb547e8 · outbound

This paper cites METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.221846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.221846Z digest=sha256:b6cf68890d42fc0133d00edbc3ead0179a9c94bf8f87d26a6a6c4aa14868ff38

Observation 777041a2-00ee-4a1a-9dc6-3c07deffeaeb · outbound

This paper cites End-to-End Referring Video Object Segmentation with Mul- timodal Transformers.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation End-to-End Referring Video Object Segmentation with Mul- timodal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.274355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.274355Z digest=sha256:f9103deb8cb9be12c5d732dda7994dab90e07da3cc8c43d56a1d1ca82089f8dc

Observation e8e3eb20-6961-4062-a322-8376d74a004c · outbound

This paper cites Auditory Scene Analysis: The Perceptual Organization of Sound.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auditory Scene Analysis: The Perceptual Organization of Sound

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.346121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.346121Z digest=sha256:dc0d9a9416c4f54fe4ceafa004172b1824acf4c773ab028e100c2f9d5257083d

Observation 9784a756-a69b-4dc8-8bc6-cbd6e11ed1b5 · outbound

This paper cites COCO- Stuff: Thing and Stuff Classes in Context.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation COCO- Stuff: Thing and Stuff Classes in Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.422372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.422372Z digest=sha256:7003a40305af5f149df11803d1634b6e7c33d75d05562d6eff8fe9b1f633a942

Observation e53caa16-5819-4119-b0bd-d8051b52809f · outbound

This paper cites TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.269880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.485367Z digest=sha256:d1ef11b5773e450903ae7d5d923799073a0f15a107c06bb129cea75143935921

Observation c505fc3d-73be-4808-9367-b907a3ce0dc7 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.252941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.596129Z digest=sha256:c101580c8bc74e0658fc1decfc86a2a18997a43fd507a92d33f0d9c9bf5f1201

Observation 1ae1bae4-45c0-424a-8d2f-ee1031106421 · outbound

This paper cites VGGSound: A Large-scale Audio-Visual Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VGGSound: A Large-scale Audio-Visual Dataset

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.232394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.636040Z digest=sha256:8e6456aa7cc9935e16b212389faa617cdfe48c2f4b88d96f112f8302b5b8c7bd

Observation 042946cb-fc4c-4e03-a6e0-73db51bb935a · outbound

This paper cites Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.210671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.696935Z digest=sha256:de46ea5131d09fdfde20b98bbc24f99b131465a5aa51013434a07157f3568a95

Observation 646c8340-d9cb-4d95-8eeb-802e3e114d5e · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Vision Transformer Adapter for Dense Predictions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.190592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.774702Z digest=sha256:2fe1bc2a03ac554c734553765995d6793ca994c772410b974cdfed661217a606

Observation 76f515e9-2c19-402f-a1c4-383e0f99abb1 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:06.856376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:06.856376Z digest=sha256:744945c1835212af5f6e30656a6356b864d8411a22422bbe3abbb97adf24d914

Observation 1e824bbf-bee8-4cc9-9ad0-a5e34c2278c8 · outbound

This paper cites Masked-attention Mask Transformer for Universal Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Masked-attention Mask Transformer for Universal Image Segmentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.171004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:06.974243Z digest=sha256:fe6e161870f6a03ead062c27de2cc840233c887101dd23cf0165c3fbeaba610e

Observation e2e4208a-25f8-4a3c-bd42-3fac05f3a85d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.045271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.045271Z digest=sha256:09caab5b006e7220763e024c5290b65bc4418ab686cafe6b4cfdd4a880fc418b

Observation b10fd106-834c-4e6d-8b7b-ad7c8fdc633b · outbound

This paper cites Qwen2-Audio Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-Audio Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:07.101084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:07.101084Z digest=sha256:1d95d64de773b92cf7ee900e46116fe71d7d6002de0e9834d4886d551edbeaa5

Observation f6258804-80df-43ff-9a14-46e88b13db52 · outbound

This paper cites MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.149630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.286988Z digest=sha256:604f926cea87d1b06dfc8bbced45dacbadabb7047c511f62b49917262992068a

Observation 01f7c39d-536f-4319-b801-04d6ef0370b8 · outbound

This paper cites MOSE: A New Dataset for Video Object Segmentation in Complex Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSE: A New Dataset for Video Object Segmentation in Complex Scenes

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.128980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.373963Z digest=sha256:1a850b909fdb940dd9dab06499b6f3cbc5e3f0d24254701bb57eed98a9c15008

Observation 4f1abf28-9b01-48f0-80fc-2db69cc2c347 · outbound

This paper cites Multimodal referring segmentation: A survey.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multimodal referring segmentation: A survey

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.110276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.468288Z digest=sha256:cf5fae27884dd619d0015ec90bb7938ef0f404c77d1c13c16f3f5d516b457389

Observation 6c58ef9c-39ad-43f1-91f7-91a7b233dcde · outbound

This paper cites MOSEv2: A more challenging dataset for video object segmentation in complex scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOSEv2: A more challenging dataset for video object segmentation in complex scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.089024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.555289Z digest=sha256:edc68ceb765ad84be234560b2c31c2654465c68a677ae3f76bb5ef5e1ea267b5

Observation a47ca59b-9d8f-4b35-bf3b-54d1f65feaf9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.070253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.614968Z digest=sha256:106f8ddd410c75b23f3acc676ad11ad3df37c8c8fd33975361d63d8f171590e8

Observation 48a5fa52-7c80-4b18-8b04-ac3a66fa6e90 · outbound

This paper cites A VSegFormer: Audio-Visual Segmentation with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VSegFormer: Audio-Visual Segmentation with Transformer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.052365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.702971Z digest=sha256:21573b902bb5306dcb3eed972bab25536077b4b330c2d052f2450f6955dc6264

Observation 4dd43122-d005-4916-bbd3-36838b48a4d8 · outbound

This paper cites https://github.com/RVC- Boss/ GPT-SoVITS, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://github.com/RVC- Boss/ GPT-SoVITS, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.032935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.890555Z digest=sha256:192f1ee9e2da6d58671f7ba17d0d5faa0830316fff4af247dc970264d910a1e3

Observation 405c891e-2575-4736-bdea-30422574e22c · outbound

This paper cites Open- V ocabulary Audio-Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Open- V ocabulary Audio-Visual Semantic Segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:12.016780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:07.963845Z digest=sha256:44af5e058a06c7490d95dcd85f45d5fc1f4679c07afa3ccbbeede41972808647

Observation 2807b96b-cd94-4504-a8de-d82e0c18893f · outbound

This paper cites Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.996594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.053343Z digest=sha256:fdc581b2ad387e34aeee57f035656c3bdcb01afe0448ff5caf37b34b358c2fd1

Observation 1371932c-7134-4480-abe1-c9c9ad800365 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.977821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.146958Z digest=sha256:12d6593ce6d265c5b38d228a6afbc11f68b7b5799e3b54c3e603d0bb60f19cf4

Observation f6c70001-9a2d-426c-ad4f-06a8182e72a6 · outbound

This paper cites A Generalized Framework for Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A Generalized Framework for Video Instance Segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.959155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.229361Z digest=sha256:9099878ac53eb71e4a511193fd79e6c696d281a0177fb65bdc8154f5e8df32ec

Observation cd2dc9d9-8f31-40c9-a458-3c397f3283c0 · outbound

This paper cites Deep clustering: Discriminative embeddings for segmentation and separation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Deep clustering: Discriminative embeddings for segmentation and separation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.942633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.279887Z digest=sha256:9f0d23c4b57c7391bd1ddb5b30d1ade8c12800f278487054579b16b774628f1d

Observation db461c65-20e1-4067-a559-4ffe04388662 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.922008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.397805Z digest=sha256:09f27d8a1e2b8241853303b15f6645533dfa7bfe020fa1f7b7bcd254f6d05e3a

Observation d1af3745-296f-45f2-8d5f-9cc30daa77ac · outbound

This paper cites Egocentric Audio-Visual Object Localization.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Egocentric Audio-Visual Object Localization

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.905791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.609708Z digest=sha256:e766f515b883851f930a416f0885d25916685a8e44a0da8e64ea1e22f4228ed5

Observation 50a7f8a0-b7c0-4f1b-9ca8-a92f7895044c · outbound

This paper cites Video Object Segmentation with Language Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video Object Segmentation with Language Referring Expressions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.888913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.667631Z digest=sha256:39fd74cdf6a7b95025a7e5d4fd363ad226902d94b58e3d07930ba692f449b6e9

Observation 672b7acd-ee7b-4197-8a0b-e39f025dd6dd · outbound

This paper cites Segment Anything.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Segment Anything

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.869618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.752455Z digest=sha256:12a694e0c402968e6c8c74a8f3e17d520f274a9ec517434a5dcd3c23e96e481c

Observation e8dc9a84-77ca-4455-8bb8-8e76388a25f0 · outbound

This paper cites LISA: Reasoning Segmen- tation via Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LISA: Reasoning Segmen- tation via Large Language Model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.853787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.844260Z digest=sha256:fcf9ebb28b9a3124f203b8d230710df44d95cfc21d2dc23c9c5ab485abd7fc05

Observation 64d9a684-9cda-4be5-9d2a-d1b0722f00db · outbound

This paper cites TVQA: Localized, Compositional Video Question Answer- ing.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation TVQA: Localized, Compositional Video Question Answer- ing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.837725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:08.989571Z digest=sha256:ecdde22243aa38e8f89cbf5d3899548cf47f216a0fc806a265af681a966d2ac0

Observation f176f74b-fda8-43ac-901b-0ec06487a5e0 · outbound

This paper cites Learning to Answer Questions in Dynamic Audio-Visual Scenarios.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Learning to Answer Questions in Dynamic Audio-Visual Scenarios

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.819872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.133635Z digest=sha256:71f585fb1eefa5edede4f0e48ef3bb535d94d1fd4551d9d8821fb031a3308c09

Observation f8a8afd9-e3f0-4389-9471-37c213797fff · outbound

This paper cites Boosting Audio Visual Question Answering via Key Semantic-Aware Cues.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Boosting Audio Visual Question Answering via Key Semantic-Aware Cues

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.800716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.261515Z digest=sha256:018dbf6b213517b91cedb1ef3077a96441be4cc6c46c170684e51a996efd1267

Observation fcccefce-30db-4420-a1eb-d1a5452bf3e4 · outbound

This paper cites Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.782524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.423301Z digest=sha256:0c28fd9e74d67c8375f2070a6628bc0f90da1bf70ab4d4cbe2c2413adf28db96

Observation 645602e0-6408-4407-97dc-4093b00fe8fb · outbound

This paper cites Robust Referring Video Object Segmentation with Cyclic Structural Consensus.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Referring Video Object Segmentation with Cyclic Structural Consensus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.764123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.565012Z digest=sha256:b620ecdaac7a53cc93d638db20edc877990ec7d09b0b1b33c4825efa937ee19d

Observation c8b02ec2-3a3d-4696-8be3-73d21c9be3ea · outbound

This paper cites Baichuan-Omni Technical Report.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Baichuan-Omni Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:09.667789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:09.667789Z digest=sha256:6df77a51d76ebcf890b6e3812eab393227edc32dc4f8b33dd126664155dc357b

Observation d5aae2dc-9bb8-440b-9bc7-5ff61ddce81d · outbound

This paper cites Losh: Long-short text joint prediction network for referring video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Losh: Long-short text joint prediction network for referring video object segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.743164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.765828Z digest=sha256:2f44b979a8c3ae9d96bc2bf8c5cc38418be438dc0b84eabd94aae12db1c9a3b1

Observation 50923f62-0614-498b-9011-7e18564c6710 · outbound

This paper cites GRES: Generalized Referring Expression Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GRES: Generalized Referring Expression Segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.725625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.891803Z digest=sha256:928d861a130d82f729c0b69e6ea976b2b9e61f7b62524575016d8c846d2274d1

Observation 891931f1-a7d3-4608-bc87-9c6f7b102264 · outbound

This paper cites Primitivenet: decomposing the global constraints for referring segmenta- tion.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Primitivenet: decomposing the global constraints for referring segmenta- tion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.705305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:09.973694Z digest=sha256:71337717cb2f308d48c98137bbb731b970e38441622493f4c64240e9188b15ff

Observation bdfc6c6f-d5ba-4b18-93c0-8dfbcddcee92 · outbound

This paper cites Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Visual Instruction Tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.688159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.155107Z digest=sha256:513c1f34d7986c3bf061fd1fde2e919e980c23f362e68fe48e178a194d6fe40c

Observation 1590031a-698e-41ca-b72f-fa229f64fbbc · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Improved Baselines with Visual Instruction Tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.672272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.204533Z digest=sha256:8a7a0580b56228973b8cc2a721a4f88f480d71a04eaefbfd7fdd072c0f53e1c1

Observation 10ab099c-43fb-47a7-a078-5903e288deae · outbound

This paper cites ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.654463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.210803Z digest=sha256:b0f68b4bb7024ced7e5cd24533431462e279df7ba8d8b8ca745a1c6b43be5424

Observation 8abb241e-01d6-4491-9e08-383ddf8654f1 · outbound

This paper cites Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.635223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.217190Z digest=sha256:a0de8400807335ff51230e42a4e330188d1a4c1b84f352bdac16572ec8c32252

Observation 015e1808-5f00-4085-98c3-22a9a3013810 · outbound

This paper cites Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.610896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.222759Z digest=sha256:62df6aba3e4f37df9905a247a35c3082e4fd3830f5bb010b5ba385882b78ed23

Observation 9a25fee9-c7ce-46fd-a42f-0e562430d94a · outbound

This paper cites Generation and Comprehension of Unambiguous Object Descriptions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Generation and Comprehension of Unambiguous Object Descriptions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.588517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.228825Z digest=sha256:89c3b4db5580d5463f3d59029b2ebc8b88ebf540a87bf5adb8e857f377507299

Observation 0501c317-ebcf-452a-b993-8c1f8f8b4d2c · outbound

This paper cites V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.570411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.234547Z digest=sha256:355904da7440f981cdc56c9a324fd5ae080ca98ed8456f7e9aec4ad538dd7aaf

Observation 6cbd765e-9f01-4245-a1ce-0c1c298c72f3 · outbound

This paper cites https://platform.openai.com/docs/ guides/text-to-speech, 2023.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://platform.openai.com/docs/ guides/text-to-speech, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.552583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.247239Z digest=sha256:1f08054c2838f5b9c86fd4719e0df7f0eba4dc3764d2dc0bc64ea0f28fdccbf2

Observation 82d8109d-93ff-4236-b5d7-e40bc2dcf588 · outbound

This paper cites https://openai.com/index/hello- gpt-4o, 2024.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation https://openai.com/index/hello- gpt-4o, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.534588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.252563Z digest=sha256:8cc756397d1f12f5e4e0496f06d0253f8c917b6f847fcceba5374f5f321b4883

Observation 6a185d90-b484-4179-baa4-a6d681f6e7b2 · outbound

This paper cites Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.516684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.260334Z digest=sha256:5b0877b478a7908853acf0384e6900ccc4f0b6932c30e40cd62a08f69ace3f49

Observation 1b467605-e0aa-40c9-bc9c-023ef0f3f312 · outbound

This paper cites DetGPT: Detect What You Need via Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DetGPT: Detect What You Need via Reasoning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.498976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.269021Z digest=sha256:18cbde9bbfd5419cf8ad4be2f19959eee5349c20d68fadbc9d653a03e47ab600

Observation cbcbf9cf-d2f6-449e-bbdc-68cfc1239e0c · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.480708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.278238Z digest=sha256:da1388c8d8b13f0e8586315465060e8a91bb4fa4622963a897ec114f10a8674e

Observation c563c37b-4b53-45a7-be18-589aece25f6f · outbound

This paper cites PACO: Parts and Attributes of Common Objects.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation PACO: Parts and Attributes of Common Objects

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.444625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.285742Z digest=sha256:eba35349b334785ed77003ab1c5fc5d9dbd321d6c729fd759a3fd2a16b7b5dc8

Observation 6e4bfd0c-8c92-4dff-ba5d-033188d7a457 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SAM 2: Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.292503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.292503Z digest=sha256:e6d4e6f1b4723df510661311c9bbe28faac1f9fa9b43dada78c0196f483df501

Observation 00c96345-f537-42d2-996d-8d48ef230711 · outbound

This paper cites URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.405620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.298834Z digest=sha256:2fc94f32e54a649309460bf50f6cc1a5f8e2aaac6acbbf2123862593d640cc5a

Observation ec9d5370-ca5d-47a2-a75d-17e77bc45517 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.364909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.304319Z digest=sha256:0c3e95680904a6d27431aa5e1e3e9298ce54315e4ed60c7c9ae56574f4405a24

Observation 43f66742-91dd-46ec-8af0-56562b948142 · outbound

This paper cites Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.347843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.312650Z digest=sha256:af541740ee01abd22d5a6bacbe26f3ad4761c8c418febacaaf1f0cb12932debb

Observation e9f020c3-53c3-4b7e-9352-6abb95a70ee2 · outbound

This paper cites Unveiling and Mitigating Bias in Audio Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Unveiling and Mitigating Bias in Audio Visual Segmentation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.330373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.319486Z digest=sha256:ffe710c11bc32cc7d6599baf45b04a7ca0b4a437b456dad0979d9194138eb0f8

Observation 10441679-27b3-40d4-ae73-4bb067d5a7d8 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.311783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.324756Z digest=sha256:385d75e0854807ba7fd2bede32086bee04e8d79e19d12a520a5a0ed3610c997c

Observation d62a2be0-8151-4f94-8b91-b354b07c882c · outbound

This paper cites Audio-Visual Event Localization in Unconstrained Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Event Localization in Unconstrained Videos

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.290558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.330368Z digest=sha256:6c19a4a5ec3d2d006bc37df1756cd22eaa40b5b3902b59ab04d787f01239cc1c

Observation 2d7d6602-7a14-4756-bee0-279c1d915fb5 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.336496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.336496Z digest=sha256:5cc456e1d68dfc7664aaa0da5a2abf6162ab829d710265097e32100265eb1270

Observation 9dc3b0b3-a9e6-4331-bf5a-a7eaa1d6b5cb · outbound

This paper cites Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.271327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.341868Z digest=sha256:6b1b7b122e287358ff24d2ea119801910ab0c5b4d84fb5d39ab4ebbaabc63282

Observation 244a95f7-f82a-40e5-9e11-73152644fc18 · outbound

This paper cites Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.253156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.357613Z digest=sha256:1d2e1c08d9cb2968b12211b95a6b7edb1463953db760ad956b5b1e5045d78f05

Observation 977b99c5-f402-4b6c-a997-dce4e653e001 · outbound

This paper cites Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.234846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.363089Z digest=sha256:311d2a585735c530a3f48175c45ebb95575d2f014b86f7b065121546c4a5c9e9

Observation 631e99e6-604d-4d89-ad36-31efa91f39f7 · outbound

This paper cites Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.218188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.369265Z digest=sha256:35def3e9b159bce5f2c31d4f49ecb5f1fe31c475cd8c3f3dfee1dc43e6ea57bb

Observation 04473527-a83e-40a6-811e-0b61852742e8 · outbound

This paper cites Language as Queries for Referring Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Language as Queries for Referring Video Object Segmentation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.198186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.374194Z digest=sha256:1e3aaafc521a3cd75e15456790377e61cc91ad950571187d62514e00b49a5372

Observation 777ad773-26d4-453f-8f81-38b7847d8705 · outbound

This paper cites VISA: Reasoning Video Object Segmentation via Large Language Models.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation VISA: Reasoning Video Object Segmentation via Large Language Models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.174877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.378825Z digest=sha256:3d7f7e9d00dde45e46663fce8e427dcd6c4d77c7aa26b08d9063f9d722ab3107

Observation 8116845e-6924-4180-b5be-c901cef114e4 · outbound

This paper cites Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.146807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.387330Z digest=sha256:7978510922cc07cda95debb2fe1269b6653e732825cee008b90166a89f6e04cc

Observation 46080188-6cf1-486a-8db9-15479764098e · outbound

This paper cites A VQA: A Dataset for Audio-Visual Question Answering on Videos.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation A VQA: A Dataset for Audio-Visual Question Answering on Videos

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.121804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.394479Z digest=sha256:213fb5378b88af2a46fc2cc6dcccb37b9868a75bac9b2abfb65dab2916aa53c2

Observation d1978e52-dd29-4c52-b280-80f4a309769b · outbound

This paper cites LA VT: Language-Aware Vision Transformer for Referring Image Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LA VT: Language-Aware Vision Transformer for Referring Image Segmentation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.095686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.399323Z digest=sha256:1bfc6789d34532b93df4811fab12f162d061524c1bca3e10d9cf6cbf38ad5c5a

Observation 95031c31-a3e3-417d-93f3-0937d2a67d77 · outbound

This paper cites Isda: Position-aware instance segmentation with deformable attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Isda: Position-aware instance segmentation with deformable attention

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.064986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.409682Z digest=sha256:65cd24fce737028042c83f38d5d0e07205bc01af1ecb5c94696224765b8d4c8d

Observation 585b8fae-a251-4094-b1d6-05e743ac25d6 · outbound

This paper cites CTVIS: Consistent Training for Online Video Instance Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation CTVIS: Consistent Training for Online Video Instance Segmentation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.044795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.414531Z digest=sha256:ef1c41252ed421d965f400c1af14f3cf0695dce0ef0a7a647997848c47e3d207

Observation 4478e543-28aa-49f9-8f7f-66ef2796a7c2 · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:11.017438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.420373Z digest=sha256:1e9a63ba788ff0b4601c009f8d25ad525c9c0ef5a98ab9b8a36cbf6b4927705a

Observation 4abbf53c-1812-4652-8e2e-20332cfcfee0 · outbound

This paper cites MOVE: Motion-guided few-shot video object segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOVE: Motion-guided few-shot video object segmentation

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.995052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.425968Z digest=sha256:95beed898af554f6dba4ae7c0e46e9d11b69a14f85b647d741a7fbd2df207c5c

Observation e239833f-6bb0-4853-bbd8-114595adcdeb · outbound

This paper cites Modeling Context in Referring Expressions.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Modeling Context in Referring Expressions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.964609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.430262Z digest=sha256:ec7a00c782cc2a34e6006f3b0c5653845c0ac149c93eb2a35809ade2e1a53d83

Observation 3f0d94cf-bbe8-4685-8d81-ca5875f7e677 · outbound

This paper cites MOTR: End-to-End Multiple-Object Tracking with Transformer.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation MOTR: End-to-End Multiple-Object Tracking with Transformer

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.945246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.435156Z digest=sha256:b5c225a78d71a14beb303b250b5364b168a8c04df4d4276e4278943750d7089f

Observation ae1f22c0-d178-4c3b-ad7c-2f9fcd04f24e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.921392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.439995Z digest=sha256:2d1a949b0133884d2d1a1a113837daca264091e0c427fd14740cc9d4a9d0eb63

Observation 60db074e-da48-4dee-ac7e-11891f91c7ef · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.902176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.444754Z digest=sha256:13f785a8fb474504999c783d45b326878615d74a6f377303e897568db7ee5bac

Observation a24455e6-9922-4a95-9d94-7a0c5ac81fa1 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.449340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.449340Z digest=sha256:fc45ab3e3f5b524b5b2a09c107d2749be345b59dcd6d024004fca7c7a5d287ec

Observation 3dd6e515-9ec8-43cd-987f-fa1c1d08a13f · outbound

This paper cites DVIS: Decoupled Video Instance Segmentation Framework.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation DVIS: Decoupled Video Instance Segmentation Framework

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.885373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.454902Z digest=sha256:7e07e66a79dc0f270a08de2e29099f72970b603c5106e98f269473238033fa8e

Observation f35ef5b4-7b2b-4406-9312-0d74b021a834 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T11:16:10.460552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:16:10.460552Z digest=sha256:31cfb040ba876d4e91e05a12c1f83cc4ebfc33931b6f374af0d5bcc888d5d00a

Observation 77a47579-e561-4dec-99a8-67aeadf588f5 · outbound

This paper cites Scene Parsing through ADE20K Dataset.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Scene Parsing through ADE20K Dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.868533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.465447Z digest=sha256:bc700e57824dfecfcf01e3bc22668987b49179da987cf943b889c446c9a835d2

Observation b4440455-5ce6-4e3d-bcb1-fdbf7fc961df · outbound

This paper cites Audio-Visual Segmentation.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Audio-Visual Segmentation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.852538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.470620Z digest=sha256:25e4c4d8ea7d964988926e2d1d9aba8190ac3a156d30a96403c1bb55624f016c

Observation 68aa03c1-fccc-480d-9709-0703fb587fc1 · outbound

This paper cites Tracking with Human-Intent Reasoning.

Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation Tracking with Human-Intent Reasoning

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:16:10.832815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T11:16:10.475961Z digest=sha256:8c118aae2f821e2695049e73433e48a3b65066feaaeb132ad3247d8230de1a15

Pith citing papers

Observation f4afe86d-fbd3-4785-a992-0802dc92b695 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation

Reference 210

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.958133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:96b5383fcf4d290b36ac942315c2b2191161a08738c59ef1856d2180a9c57b6e