Pith. sign in

Paper Citation Record · LEDGER

InterRVOS: Interaction-aware Referring Video Object Segmentation

As of 20 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2506.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02356 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:29.724751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.213452Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee763a0-162e-4fdf-803b-272cdbdf7042 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation One token to seg them all: Language instructed reasoning segmentation in videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.253558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.623640Z digest=sha256:56dfc3281a904695a483cf8168f07d176fbab110303f854528dc212cdafb5314

Observation 85184742-8ac3-4641-a574-aecbcd50b8c5 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

InterRVOS: Interaction-aware Referring Video Object Segmentation End-to-end referring video object segmentation with multimodal transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.057605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.674228Z digest=sha256:e3354bb233785339886887416b84643e801767b49b445d61931b94a9633e52af

Observation 5a381820-99cb-45bc-9192-783b00a298f8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InterRVOS: Interaction-aware Referring Video Object Segmentation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:27.777419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:27.777419Z digest=sha256:5d194df8c4e01a6a0d2eec2a364bb6e02153c973f4695aed5f147c2076849753

Observation 9368ec74-88b7-4f2f-a2c8-0c690f01fbda · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Vision-language transformer and query generation for referring segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.834762Z digest=sha256:21a32e0135af557edab58db95644aa77d6ee3f2c4deff352d3e6f5e6426629bb

Observation baae64bd-df07-4a70-b212-5abaec036c7b · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.640172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.917308Z digest=sha256:40b84909b5f2dfa81421f4d56eb60901030d507348b2f239dade7cbac31a513a

Observation b91dbc1a-cbb9-4d1f-aa8b-a8e0ac7e1dd8 · outbound

This paper cites Moma: A multi-object multi-action dataset for understanding human activities.

InterRVOS: Interaction-aware Referring Video Object Segmentation Moma: A multi-object multi-action dataset for understanding human activities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.410861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.989653Z digest=sha256:835382d0083246bbde3a15c373696f169f239a6529fe70ff08569990f71c45e7

Observation bf4a6e33-31c5-4bfb-ad76-b6f4018096bb · outbound

This paper cites Actor and action video segmentation from a sentence.

InterRVOS: Interaction-aware Referring Video Object Segmentation Actor and action video segmentation from a sentence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.306269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.076797Z digest=sha256:a4d7e2282739e98a4d142104a1419b16d96f909fb9275dcf11d3e346c02b8921

Observation 52848177-7ddd-4880-9b09-00a850dcaec7 · outbound

This paper cites The llama 3 herd of models.

InterRVOS: Interaction-aware Referring Video Object Segmentation The llama 3 herd of models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.101988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.133567Z digest=sha256:32c15e8f7b23e3cda471fc162957351b500f7a896097ab21130585c2a30584b4

Observation f51f17a5-0d00-409d-b20a-94e1e8dfbb97 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Lora: Low-rank adaptation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.190322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.190322Z digest=sha256:5b3283148b6956b944045c8dbe09978ea284e5d9937951fc68ace8bb573286db

Observation 82a39cc5-de55-4213-afa2-c3f0483ebaad · outbound

This paper cites GPT-4o System Card.

InterRVOS: Interaction-aware Referring Video Object Segmentation GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.254083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.254083Z digest=sha256:f4a9b63d22426adcdad88673a0c3a50a1de1df4e061027fd6895d0462d3c2176

Observation 5ad1d67c-d91c-4f6d-a940-9f0cc036ade9 · outbound

This paper cites Action genome: Actions as compositions of spatiotemporal scene graphs.

InterRVOS: Interaction-aware Referring Video Object Segmentation Action genome: Actions as compositions of spatiotemporal scene graphs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.908278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.343745Z digest=sha256:dc198aff57c7e463aff97f252627d52fc2245f8ff01878d71605f06e5538ee5d

Observation 86063e0b-4cee-4535-af7a-e12a2afdbae7 · outbound

This paper cites Video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video object segmentation with language referring expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.412310Z digest=sha256:e87e178303ebeeb9de1ed534639a8a0dd9a812ae855031ce53fe57d81f94631c

Observation 3c42c50f-28d6-4f2a-a0c1-5d2179bed05c · outbound

This paper cites VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.478864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.478864Z digest=sha256:8d37c5112fcd5a2cca9b69923b5b7b33f9b38fa0dc9d9b4d1c20b43368f16d35

Observation 9dd12239-b4e5-43e3-95f1-eab0bf2524d0 · outbound

This paper cites Visual instruction tuning.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.562199Z digest=sha256:0d7daa0204ef60db93836438ea63e13e6a2d924302b88f4a7c2aadec3aea1b55

Observation b8e90b91-2bcd-4c5a-b288-d8694652df00 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.416144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.655247Z digest=sha256:5561aa4b3e5e2e9762c3c808c05950f5a0d8831557495c88bcfd2d9772c5bb2f

Observation 945be114-8625-4e57-9bbe-9fc57529e925 · outbound

This paper cites Refer-youtube-vos: A dataset for video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Refer-youtube-vos: A dataset for video object segmentation with language referring expressions

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.178119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.718185Z digest=sha256:95e16780bb3a306d33bed4c9744229c88cdc7b81642d66da7b8698f46a5c6316

Observation 80d7bdab-9f61-4f66-a496-b15b3a3355e5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation SAM 2: Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.763180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.763180Z digest=sha256:82d99b882978bb6d12884cfc085408ba730a1221522b3a352de1bf11fa960dec

Observation db77b5ba-1be2-4a7a-9791-269ea4c563db · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

InterRVOS: Interaction-aware Referring Video Object Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.861439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.841080Z digest=sha256:368534a20bed3b3086dd29be85800a91130c4209a815574c324d7471752571af

Observation 2aa89ca2-753c-40bd-b9e7-699f0531a945 · outbound

This paper cites Video relationship detection.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video relationship detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.474955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.914632Z digest=sha256:29f7dbfe1cd4263db7fd1edaffd7344cc39dc1b1e299b01dfb95b3594162f658

Observation 71050332-79e4-42da-8df2-789af37553ba · outbound

This paper cites Annotating objects and relations in user-generated videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Annotating objects and relations in user-generated videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.094806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.981295Z digest=sha256:9c8d9a990637be6538a81fee1f723ba165194c6267f3261f180c26f7f4789a7c

Observation e30af10b-5310-4e50-a02a-6535d18e2ab7 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

InterRVOS: Interaction-aware Referring Video Object Segmentation Yfcc100m: The new data in multimedia research

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.865552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.039012Z digest=sha256:4e7d643afb52e2d7947356a78e9d480e333e2ec94fd1dac0f5b519b0457dfca2

Observation 5cb9c047-76ec-408a-8c66-0f6c3378604a · outbound

This paper cites SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.102921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.102921Z digest=sha256:b4651b6b8b8a66e5f637d16097955922a465afbaed4696715e5df21175013c9c

Observation 51c3e5c4-7511-4120-a3b2-65b765c8c9b2 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

InterRVOS: Interaction-aware Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.149297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.149297Z digest=sha256:ee590e32199c98a0347314805852d07b7544fc02261f57f355bbd30fd433b688

Observation e701675b-9556-4572-b972-89f1ac135800 · outbound

This paper cites Language as queries for referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Language as queries for referring video object segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.644265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.205824Z digest=sha256:7a15707c4487bbfeb9ab8ebcd12cb1750670dfeef7085e279e1053c18deaffd8

Observation 759cfd64-7b93-4d6e-9e2e-1b9fa7556250 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.288871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.288871Z digest=sha256:b437d0977f6a2e072b1e5e7ac70818e6a84bd8bb3e6dc5943631a93f3ff1a006

Observation 3efc5d46-5ba5-4ef7-8314-11c0529b260c · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visa: Reasoning video object segmentation via large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.477138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.323770Z digest=sha256:34e1d10842e118ca1edb98eb11b2c2c9eac5704edfbed35b7a1d262d8d3b61b2

Observation 651c05f3-96ee-4b49-9b2e-fae4fe859f67 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.406573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.406573Z digest=sha256:46aa6d45f8b4acbfe5066280f628cf2eb814ce079187a2a3de2be28277834f25

Observation f2776946-a890-4711-adbb-6403f95f64e9 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:30.142164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.474932Z digest=sha256:b099e0d41285d8bf266a2d4804a0998b2641257aecb337cbd99f9b8d803a29df

Observation 9034104e-4cb4-4125-b42f-18ec53fd65aa · outbound

This paper cites write newline.

InterRVOS: Interaction-aware Referring Video Object Segmentation write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.534945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.534945Z digest=sha256:260b039c09f9207f88a2533e82bf9f673c9578cab9dd992d2e2903efe3f42131

Observation 1d30c372-044f-4464-b690-037358a9bc37 · outbound

This paper cites @esa (Ref.

InterRVOS: Interaction-aware Referring Video Object Segmentation @esa (Ref

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.594298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.594298Z digest=sha256:c7440ce9c2d1479ca9abbca2743843474cade0538b1716baa0324a762ba2af06

Observation c13ea854-4ba4-4081-8f51-25305ab4c2b2 · outbound

This paper cites an unresolved cited work.

InterRVOS: Interaction-aware Referring Video Object Segmentation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.642842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.642842Z digest=sha256:be76cbf8f1ad67ff53aed83b68272074a70c4db518e490624bad0dd8c1f07d7c

Observation ce1d1cc0-77ce-4bad-b412-6a3faf66d064 · outbound

This paper cites A child helping another child with a backpack.

InterRVOS: Interaction-aware Referring Video Object Segmentation A child helping another child with a backpack

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.724751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.724751Z digest=sha256:a4f195d7e45db21daa138957eb8bc9a46d3d2d894f19c6092d8abbfe89bb4556

Pith citing papers

Observation 66f5d8ba-b7bd-4887-b54e-b3e839842af9 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models InterRVOS: Interaction-aware Referring Video Object Segmentation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.214944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:c9adca432f56e13cf13757ea4a19c448a07eef6c56fb6f0399d5e2fc9f472f7b