Pith. sign in

Paper Citation Record · LEDGER

InterRVOS: Interaction-aware Referring Video Object Segmentation

As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2506.02356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02356 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:29:29.724751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T01:50:54.242508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.213452Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dee763a0-162e-4fdf-803b-272cdbdf7042 · outbound

This paper cites One token to seg them all: Language instructed reasoning segmentation in videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation One token to seg them all: Language instructed reasoning segmentation in videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.253558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.623640Z digest=sha256:e7887efbb722dad79e0112c42c8a42130922091205106efbf9383fa55c1c8964

Observation 85184742-8ac3-4641-a574-aecbcd50b8c5 · outbound

This paper cites End-to-end referring video object segmentation with multimodal transformers.

InterRVOS: Interaction-aware Referring Video Object Segmentation End-to-end referring video object segmentation with multimodal transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:34.057605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.674228Z digest=sha256:a8c40816afa07bbe09ebf9178518439175608bb4d1aa1de3e87e539a91fa613e

Observation 5a381820-99cb-45bc-9192-783b00a298f8 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

InterRVOS: Interaction-aware Referring Video Object Segmentation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:27.777419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:27.777419Z digest=sha256:2fcf8511481514a420f508a787432dccc012d474dcf8ee1130bab1b5673c0d37

Observation 9368ec74-88b7-4f2f-a2c8-0c690f01fbda · outbound

This paper cites Vision-language transformer and query generation for referring segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Vision-language transformer and query generation for referring segmentation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.862292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.834762Z digest=sha256:06bb6cfeb35970976fc7ace0bc77a97e81e4de0d22c06145be20846c425f684b

Observation baae64bd-df07-4a70-b212-5abaec036c7b · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.640172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.917308Z digest=sha256:b669a3daec92c5fbbd2baccd76e7115cd0140caf1234f2a1d2565222b3a56750

Observation b91dbc1a-cbb9-4d1f-aa8b-a8e0ac7e1dd8 · outbound

This paper cites Moma: A multi-object multi-action dataset for understanding human activities.

InterRVOS: Interaction-aware Referring Video Object Segmentation Moma: A multi-object multi-action dataset for understanding human activities

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.410861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:27.989653Z digest=sha256:d5c800916236106d58b9046d1162831b7be582a2feeb099d6d17cca916a2a24a

Observation bf4a6e33-31c5-4bfb-ad76-b6f4018096bb · outbound

This paper cites Actor and action video segmentation from a sentence.

InterRVOS: Interaction-aware Referring Video Object Segmentation Actor and action video segmentation from a sentence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.306269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.076797Z digest=sha256:3ccd1e38a888dc4381ee6305f56ef7ecdf9d1b0372d5ba128301afa7e36e99ab

Observation 52848177-7ddd-4880-9b09-00a850dcaec7 · outbound

This paper cites The llama 3 herd of models.

InterRVOS: Interaction-aware Referring Video Object Segmentation The llama 3 herd of models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:33.101988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.133567Z digest=sha256:7981c1b9367862ac03830d7bb8313503f89b77897c3dee30fe149c964e536b0f

Observation f51f17a5-0d00-409d-b20a-94e1e8dfbb97 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Lora: Low-rank adaptation of large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.190322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.190322Z digest=sha256:68457185e79a474fa8ab452259e150b01183c90fe490513d0edefecb5f7f8c7b

Observation 82a39cc5-de55-4213-afa2-c3f0483ebaad · outbound

This paper cites GPT-4o System Card.

InterRVOS: Interaction-aware Referring Video Object Segmentation GPT-4o System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.254083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.254083Z digest=sha256:e58d58cfe5ed13f7444a6b133f5fbc646a495724cc02ac4ba47a04943823cfc4

Observation 5ad1d67c-d91c-4f6d-a940-9f0cc036ade9 · outbound

This paper cites Action genome: Actions as compositions of spatiotemporal scene graphs.

InterRVOS: Interaction-aware Referring Video Object Segmentation Action genome: Actions as compositions of spatiotemporal scene graphs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.908278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.343745Z digest=sha256:7bd59ab3477c444557877c44c97095485142aa02dacc4683b6e4e7ecc1a990d2

Observation 86063e0b-4cee-4535-af7a-e12a2afdbae7 · outbound

This paper cites Video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video object segmentation with language referring expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.412310Z digest=sha256:0de574aa6e91f2f77bdea987c0a0b9f85c8f57026edf191005ae6bc2c776331f

Observation 3c42c50f-28d6-4f2a-a0c1-5d2179bed05c · outbound

This paper cites VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.478864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.478864Z digest=sha256:49de397c5115cbba002534ca42b4f8cf146b709d622487168790e09f3ccc158f

Observation 9dd12239-b4e5-43e3-95f1-eab0bf2524d0 · outbound

This paper cites Visual instruction tuning.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visual instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.562199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.562199Z digest=sha256:1906386fe2baf383b80ea8958d9e606c29c9656eba1aa4ea98e9b819f87d7601

Observation b8e90b91-2bcd-4c5a-b288-d8694652df00 · outbound

This paper cites Spectrum-guided multi-granularity referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Spectrum-guided multi-granularity referring video object segmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.416144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.655247Z digest=sha256:2a78307c643ee66b6969ad295fba9f6fd96043f76f36f389498a378aa45f37c2

Observation 945be114-8625-4e57-9bbe-9fc57529e925 · outbound

This paper cites Refer-youtube-vos: A dataset for video object segmentation with language referring expressions.

InterRVOS: Interaction-aware Referring Video Object Segmentation Refer-youtube-vos: A dataset for video object segmentation with language referring expressions

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:32.178119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.718185Z digest=sha256:edc700ae2405debdbfb0e9d6132d301ed50049761e9af267e0c49483a8f6bf74

Observation 80d7bdab-9f61-4f66-a496-b15b3a3355e5 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation SAM 2: Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:28.763180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:28.763180Z digest=sha256:76f35f0391ff6d1e8f34e2bebfc8c252b595f2a4c2acc82ffa3bad75616f31c4

Observation db77b5ba-1be2-4a7a-9791-269ea4c563db · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

InterRVOS: Interaction-aware Referring Video Object Segmentation Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.861439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.841080Z digest=sha256:6288fb10530c1581937c5adc4a9bdc63db27b38051433b51e125d2828028a214

Observation 2aa89ca2-753c-40bd-b9e7-699f0531a945 · outbound

This paper cites Video relationship detection.

InterRVOS: Interaction-aware Referring Video Object Segmentation Video relationship detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.474955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.914632Z digest=sha256:1c1a8dec52379ac2d6b44edbc82bfc720bd7b9f2817ea513443dfb736b1073dc

Observation 71050332-79e4-42da-8df2-789af37553ba · outbound

This paper cites Annotating objects and relations in user-generated videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Annotating objects and relations in user-generated videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:31.094806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:28.981295Z digest=sha256:199d3cfc44c79059222d1064dedb8e175ca5647df8b12e5a380f3ca853aaf50b

Observation e30af10b-5310-4e50-a02a-6535d18e2ab7 · outbound

This paper cites Yfcc100m: The new data in multimedia research.

InterRVOS: Interaction-aware Referring Video Object Segmentation Yfcc100m: The new data in multimedia research

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.865552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.039012Z digest=sha256:428808b7a52bbbb442ec731dc6624d49287289a1729f945fbc65f39b4abb99bf

Observation 5cb9c047-76ec-408a-8c66-0f6c3378604a · outbound

This paper cites SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.102921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.102921Z digest=sha256:ba6d4f633ab4e96695b6564e4968dca39fd5e4ccd831df46dd38f582c3ba9320

Observation 51c3e5c4-7511-4120-a3b2-65b765c8c9b2 · outbound

This paper cites ViLLa: Video Reasoning Segmentation with Large Language Model.

InterRVOS: Interaction-aware Referring Video Object Segmentation ViLLa: Video Reasoning Segmentation with Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.149297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.149297Z digest=sha256:d90c09a696a95a932fe761520f8f1310e1c7fda7c134901d344da323a00ab912

Observation e701675b-9556-4572-b972-89f1ac135800 · outbound

This paper cites Language as queries for referring video object segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Language as queries for referring video object segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.644265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.205824Z digest=sha256:01e0d0530860ed343d731062ade2fb1e0267d5b55d375f483751faaecd7dc953

Observation 759cfd64-7b93-4d6e-9e2e-1b9fa7556250 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.288871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.288871Z digest=sha256:e123ac52287e34019e7e7cae13ce9c80d3c076512003e4d07352446d6cccdd74

Observation 3efc5d46-5ba5-4ef7-8314-11c0529b260c · outbound

This paper cites Visa: Reasoning video object segmentation via large language models.

InterRVOS: Interaction-aware Referring Video Object Segmentation Visa: Reasoning video object segmentation via large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:29:30.477138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.323770Z digest=sha256:eb991f6dff967dd92ad7db40045dc36679f3b126ba1342b95e124e8a0bcbb5da

Observation 651c05f3-96ee-4b49-9b2e-fae4fe859f67 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

InterRVOS: Interaction-aware Referring Video Object Segmentation Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.406573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.406573Z digest=sha256:399d293986010add3b9dd3773a75fb977d7a3d6578dfd06513af8f13ed5359f6

Observation f2776946-a890-4711-adbb-6403f95f64e9 · outbound

This paper cites Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation.

InterRVOS: Interaction-aware Referring Video Object Segmentation Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:30.142164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:29:29.474932Z digest=sha256:783e02265c0f8dfef523ad8c878d7d2494d5d83eb10416d82a98f5bcc6b87a5c

Observation 9034104e-4cb4-4125-b42f-18ec53fd65aa · outbound

This paper cites write newline.

InterRVOS: Interaction-aware Referring Video Object Segmentation write newline

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.534945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.534945Z digest=sha256:753ddf49d472b9a6b2b4b979ad2af9abd4fd5bed5ad609b1fdd0e072d9dd12c3

Observation 1d30c372-044f-4464-b690-037358a9bc37 · outbound

This paper cites @esa (Ref.

InterRVOS: Interaction-aware Referring Video Object Segmentation @esa (Ref

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.594298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.594298Z digest=sha256:84fed820ede1864c8dff55613c402871b4c39034464844962350f364da22e702

Observation c13ea854-4ba4-4081-8f51-25305ab4c2b2 · outbound

This paper cites an unresolved cited work.

InterRVOS: Interaction-aware Referring Video Object Segmentation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.642842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.642842Z digest=sha256:4e4664dc3a8cae5cd58dd6ad671832fc968ba28afd164c28e23e245516546143

Observation ce1d1cc0-77ce-4bad-b412-6a3faf66d064 · outbound

This paper cites A child helping another child with a backpack.

InterRVOS: Interaction-aware Referring Video Object Segmentation A child helping another child with a backpack

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:29:29.724751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:29:29.724751Z digest=sha256:e8e2311b8b0b848932755f3c555e83e368a0a5473dc0479fcbf340bf0fb88fba

Pith citing papers

Observation 66f5d8ba-b7bd-4887-b54e-b3e839842af9 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models InterRVOS: Interaction-aware Referring Video Object Segmentation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.214944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:991bb30eec4e32c23fbc384af986ec086229ea50e3417c82180810ca2a3b78dd