Pith. sign in

Paper Citation Record · LEDGER

Referring Video Object Segmentation via Language-aligned Track Selection

As of 15 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2412.01136.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01136 v2

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:42:30.341132Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:08:16.065856Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T05:09:31.997736Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72b3f13b-d469-49cc-af3f-03b302024ed0 · outbound

This paper cites GPT-4 Technical Report.

Referring Video Object Segmentation via Language-aligned Track Selection GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.187102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.187102Z digest=sha256:2174ffc92d4e63319f8e5edcbf6b12e927c2dbffdab03dddec1b15debf914f6f

Observation 7aedcaaa-2530-4ba2-8bfd-3eca2b38a316 · outbound

This paper cites End-to-end referring video object segmentation with mul- timodal transformers.

Referring Video Object Segmentation via Language-aligned Track Selection End-to-end referring video object segmentation with mul- timodal transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.055324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.192902Z digest=sha256:0832e59d701e66f2ec87d50cd240272e3b2d426aee0c54182492e31f03dac6c3

Observation 2e5a6799-73e8-4236-97bb-42b1e1f4627f · outbound

This paper cites Vision-language transformer and query generation for refer- ring segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Vision-language transformer and query generation for refer- ring segmentation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.039986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.198030Z digest=sha256:994de2680036e16e8fc3afe7040a354b78ea0754398cb2d71bb2881eb5e4e656

Observation 021f977b-b978-4745-96c1-a0bf0c45470d · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Mevis: A large-scale benchmark for video segmentation with motion expressions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.023971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.203375Z digest=sha256:ed26e2a9711f7dbb773295fafab5c4dd49af8e4a86bab69eeb4b7a368e07274a

Observation e8af9f4e-02b6-4310-813e-d1c9fe1ff8ca · outbound

This paper cites Language-bridged spatial-temporal interaction for referring video object segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Language-bridged spatial-temporal interaction for referring video object segmentation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:31.003957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.208440Z digest=sha256:f3012e7c62b455ce6151cad127d04074529525f4b2e5cbed519a467b4c8babb2

Observation ee750bb6-3a96-4573-bdf1-84b2f0fed86c · outbound

This paper cites The pascal visual object classes (voc) challenge.

Referring Video Object Segmentation via Language-aligned Track Selection The pascal visual object classes (voc) challenge

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.985638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.213445Z digest=sha256:9ab798ab6340f17a3d8655531ad06aff5ee28aa5e775dc503c547eef0ef4fb7d

Observation 2e25aaf5-ba70-495e-9386-7bcf1ceb2269 · outbound

This paper cites Actor and action video segmentation from a sen- tence.

Referring Video Object Segmentation via Language-aligned Track Selection Actor and action video segmentation from a sen- tence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.969164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.220559Z digest=sha256:95f403f60530119cad52ba755525db6bf386fd27d38f476ecfc6967bdeb9d8ed

Observation ce2302a8-491a-4ef8-8d50-72b4a0c643a7 · outbound

This paper cites Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation.

Referring Video Object Segmentation via Language-aligned Track Selection Html: Hybrid temporal-scale mul- timodal learning framework for referring video object seg- mentation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.951065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.228708Z digest=sha256:75fcbb4e15c00d133a58df0243e35e06713d464d0bae66af1f0974146e7d85cf

Observation 5ba59491-4c49-46a1-8788-8e9f46a22f65 · outbound

This paper cites Decoupling static and hier- archical motion perception for referring video segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Decoupling static and hier- archical motion perception for referring video segmentation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.934513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.233906Z digest=sha256:40864dd817116aa001947e0b901d8969649b03b9c4ddf244cabf7731918b5335

Observation 72a94c27-7bf6-4408-8487-2d83cf546cdf · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Referring Video Object Segmentation via Language-aligned Track Selection Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.238720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.238720Z digest=sha256:49562d6b66a9ccf21eab527dce6e4d1a1e981c9b4ab55fade36aedce66489129

Observation 41c97bf8-dcb4-4d16-ac06-e2152538a441 · outbound

This paper cites Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.243704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.243704Z digest=sha256:8e0e3cc9e69fc0e2c7d7668a8ac32fe46bd4dffadf270fa7005ca03c4a654246

Observation 4b34fb5a-31c7-4a10-a5f1-1511cd520211 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Referring Video Object Segmentation via Language-aligned Track Selection Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.915946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.249752Z digest=sha256:113a2991ad6adc59dc7a7353a360aecb79ae49033e19c6b0856bfde5e5ba98a1

Observation 19fee968-32f3-468c-a432-a56361293f7d · outbound

This paper cites Video object segmentation with language referring expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.899243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.254626Z digest=sha256:a4e33ffccd1add52dcaafb93f260f3ca7a5e5d598e8f54df6000a978d2398d51

Observation 5c37430b-569c-4528-b8ca-035f1ecc81c4 · outbound

This paper cites Video object segmentation with language referring expressions.

Referring Video Object Segmentation via Language-aligned Track Selection Video object segmentation with language referring expressions

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.881871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.259230Z digest=sha256:6db481a932bf11f6f2d8c930f24ed23a10178b1fbab055889e7acdf768c68d6c

Observation a9ab9048-f549-47f7-aca3-3a4c9a07bc0a · outbound

This paper cites Segment any- thing.

Referring Video Object Segmentation via Language-aligned Track Selection Segment any- thing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.264270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.264270Z digest=sha256:edc8ef5936c00e6a9ab368e35d1425a8b4146546e99d371bcd388b6c7b1cba7b

Observation 2102bee9-3c31-493a-a017-c3540bfab65e · outbound

This paper cites Referring image seg- mentation via recurrent refinement networks.

Referring Video Object Segmentation via Language-aligned Track Selection Referring image seg- mentation via recurrent refinement networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.855992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.268630Z digest=sha256:0a06c06f8d0d601faf05388c99e6db0a6943de38e11480ebe7935d85a79d110a

Observation c5ebdb6e-1be1-4721-acbe-1f04f1af376b · outbound

This paper cites Robust referring video object segmentation with cyclic structural consensus.

Referring Video Object Segmentation via Language-aligned Track Selection Robust referring video object segmentation with cyclic structural consensus

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.838830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.273139Z digest=sha256:6e718cbe36e849280dda9960a05ff6ae7e93a33f21629b57c1df42914c720c04

Observation c5366fd6-a2d9-4667-8ec4-f8cfbac51261 · outbound

This paper cites RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.278910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.278910Z digest=sha256:12ff1f6c05dd2b580402ebdb9d6d5b1f87f1f1e1e334abbc071e836c63e1dc2d

Observation c146b9a3-2a06-4a09-b7ce-677443afd363 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Referring Video Object Segmentation via Language-aligned Track Selection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.285147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.285147Z digest=sha256:fe1cacbd5c0a4649fed6227b8a8e9e9f7d3d5792aa6841bb019d3701821e1b75

Observation 70c4a044-c2b8-473a-85f6-7cd0364320c3 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Referring Video Object Segmentation via Language-aligned Track Selection RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.292818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.292818Z digest=sha256:5d16e9cecf8c9b30bfc2d47b0db1da4c04ef19803770958d3089fbf1e7edf53c

Observation af574789-eae5-4493-ab62-1d0c2fe1bb50 · outbound

This paper cites Temporally consistent referring video object segmentation with hybrid memory.

Referring Video Object Segmentation via Language-aligned Track Selection Temporally consistent referring video object segmentation with hybrid memory

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.821262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.298412Z digest=sha256:64b76004253b86eae77bf615f3f6386fd628f7e8053ce2905d4fab284a69d130

Observation 74a8702f-3397-492b-9705-167cb268bcbe · outbound

This paper cites Efficient non- maximum suppression.

Referring Video Object Segmentation via Language-aligned Track Selection Efficient non- maximum suppression

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.801716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.304055Z digest=sha256:daa80d6dde3fedc0a368cf10c6d55903592018e4689a5ce9c465e7afa3b3eb1d

Observation 6f332e0b-e8e8-422c-8616-623c1c5c75c7 · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

Referring Video Object Segmentation via Language-aligned Track Selection A benchmark dataset and evaluation methodology for video object segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.783350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.309190Z digest=sha256:0c6a9aad152cee913f5e13bcb10b02a948ca0228a9ae3f47ec74c6fbbb48f710

Observation deaf41a7-9c1d-4c4d-90c4-d639b97bbee4 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Referring Video Object Segmentation via Language-aligned Track Selection SAM 2: Segment Anything in Images and Videos

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.313812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.313812Z digest=sha256:6201482382ca20bf8bceb717cbe25b2a4dc1883fceae054070c3c1dd0783864f

Observation 4a8dcfc2-c299-46a1-9f6d-86faa7caf827 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Referring Video Object Segmentation via Language-aligned Track Selection Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.763510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.320392Z digest=sha256:fd91846885b5584ca8a37764b4c67132b2549b41744d757b712c1f5b0f8842be

Observation 1617197a-0fdd-4668-afc2-5ef4222d579f · outbound

This paper cites Attention is all you need.

Referring Video Object Segmentation via Language-aligned Track Selection Attention is all you need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:42:30.325968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:42:30.325968Z digest=sha256:c7a1229fa992041bfc30329b17dea28ca8da1ffc2cee361e4131c02058198330

Observation 77670a91-7112-4f11-89a8-e3b7ef0d0497 · outbound

This paper cites Data-efficient mul- timodal fusion on a single gpu.

Referring Video Object Segmentation via Language-aligned Track Selection Data-efficient mul- timodal fusion on a single gpu

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.735719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.330343Z digest=sha256:edb579d3d7e180fb9ca8e2eb4929b0deedad82cbb5032484f6584062367ba477

Observation 4a000abf-1e12-427a-b2a3-5943585bfade · outbound

This paper cites Language as queries for referring video object segmen- tation.

Referring Video Object Segmentation via Language-aligned Track Selection Language as queries for referring video object segmen- tation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.719106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.336155Z digest=sha256:9edd7ceaf3e70c372615d0318631360d65d97ece5ce01fb21c66485ad56799fe

Observation c50724b1-ee13-4332-949a-76b6a6232ef8 · outbound

This paper cites Going right.

Referring Video Object Segmentation via Language-aligned Track Selection Going right

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:42:30.695296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T04:42:30.341132Z digest=sha256:8b007caccb8c2dabbeea4102dc0ed4c89f97b28f80f21809a726b2ffd514852c

Pith citing papers

Observation ec8a01c8-6981-4e8b-8e64-4b8fa22f65bb · inbound

MOVE: Motion-Guided Few-Shot Video Object Segmentation cites this paper.

MOVE: Motion-Guided Few-Shot Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:08:16.065856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:08:16.065856Z digest=sha256:617ce2968ad6d22bb1d0f4c8f895bc7b7f271f4fcc88b0d9c9858a817c324201

Observation cff9ce97-b9ba-401e-82fe-fa7684435c61 · inbound

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation cites this paper.

Unleashing Hierarchical Reasoning: An LLM-Driven Framework for Training-Free Referring Video Object Segmentation Referring Video Object Segmentation via Language-aligned Track Selection

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:09:32.003729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-05T05:09:31.671155Z digest=sha256:6997e0842878d7c04e186567cbe5016066e807090e1cca30e7df4fcf14426e8b