Pith. sign in

Paper Citation Record · LEDGER

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation

As of 10 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.05948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05948 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:17:34.548221Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44cec9f3-2f55-4dd1-946f-2731467497a5 · outbound

This paper cites Stem-seg: Spatio-temporal em- beddings for instance segmentation in videos.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stem-seg: Spatio-temporal em- beddings for instance segmentation in videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.225576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:30.801393Z digest=sha256:24069486118903940f61a78b72eac276e9dd2a98e9fe29b236cd53708245b556

Observation 4aa3b6fc-8dfd-4296-950a-de684d2ad91e · outbound

This paper cites Is space-time attention all you need for video understanding? In International Conference on Machine Learning, 2021.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Is space-time attention all you need for video understanding? In International Conference on Machine Learning, 2021

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.215551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:30.900956Z digest=sha256:130023b7f64c172a977bbc1e13bf46276db04a33175dface716e907b1d5e0e2b

Observation 4fd57a2b-e37a-4bd8-a81b-f11ca7fe49d5 · outbound

This paper cites MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.044998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.044998Z digest=sha256:b7f1bf509bfef42402c5311bc51f6b03cd20f8bc668954f00c458001a4e43ee5

Observation f42aa1f6-bfb6-4184-b5cd-a5166450ac14 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.260099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.260099Z digest=sha256:28cf845042ccd548e619ab85cc37290984c21e8ef294522b328f9cd9d91f98bf

Observation 40888369-7740-4196-8a40-0ca91b09051e · outbound

This paper cites End-to- end object detection with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to- end object detection with transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.205943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.370901Z digest=sha256:4316a6377254e79765346815200e2ecaa0ce3ce74a5bcaf47c954c799492115d

Observation de672815-1933-4925-b478-5ca43aa28665 · outbound

This paper cites Liu, Yen-Cheng Liu, and Yu- Chiang Frank Wang.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Liu, Yen-Cheng Liu, and Yu- Chiang Frank Wang

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.195868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.495860Z digest=sha256:96c6d0fc0c4c30949121410faa9a3aeffd9ddd59241e060e31fc2c2a4356cadf

Observation 883b2de5-6e7e-4ce7-82b9-f04d83ff57dc · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vision Transformer Adapter for Dense Predictions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.571904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.571904Z digest=sha256:a67839a8dad82c3e0669e41825ea03fea6c701b443504950d1aaef0f8243bf69

Observation bee4c805-d516-4b99-a360-31bdf9138c91 · outbound

This paper cites Collins, Yukun Zhu, Ting Liu, Thomas S.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Collins, Yukun Zhu, Ting Liu, Thomas S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.185091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.648706Z digest=sha256:7d10862377443de7c23a6086efd69234a6eba09b3459eab399d4bca4a700072d

Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · outbound

This paper cites Mask2Former for Video Instance Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:31.737828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:31.737828Z digest=sha256:41cb55a6fafbb23cd71ad1943782a82879af0f784dee1964e253195223f190c9

Observation f2b25763-c8f1-4016-af5e-0fa2158affd1 · outbound

This paper cites Per- pixel classification is not all you need for semantic segmen- tation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Per- pixel classification is not all you need for semantic segmen- tation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.174848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.802513Z digest=sha256:0cb26456e9968e614dc50f4e0d77005079fcb3b110580751a9fe0adcca89bd0e

Observation f6b9670c-78a4-47e4-9fc6-780bc041e9df · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Masked-attention mask transformer for universal image segmentation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.164477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.862573Z digest=sha256:3e19aeda27001d205f8c79d6535d953082dff12a189e1786a1966b71e4d27572

Observation eecf82d2-8b0c-41c0-bc50-4a1879bd9af6 · outbound

This paper cites MeViS: A large-scale benchmark for video segmentation with motion expressions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.153426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:31.953003Z digest=sha256:effd0349cd4c123c209dbd81916754386d80e1a3794fa677c6e37e2f1a4f1678

Observation f81a655b-2269-4813-b6dd-62845d006d2c · outbound

This paper cites MOSE: A new dataset for video object segmentation in complex scenes.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MOSE: A new dataset for video object segmentation in complex scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.143479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.053673Z digest=sha256:33ca326d379af957691808f2148646e9d358f29d7b2689808cf1045f27c3d19a

Observation 4ea2f2aa-c283-4462-aa95-184f0563ba7c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation An image is worth 16x16 words: Transformers for image recognition at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.132546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.120327Z digest=sha256:e32e0aaccd9ece56f6049cbf4fa764b084f80f2bfee4fe8af83ffcd5ac1fc273

Observation 2037d9f0-e08e-4e28-a9a3-0c6ebf7a7c0c · outbound

This paper cites Depth map prediction from a single image using a multi-scale deep net- work.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth map prediction from a single image using a multi-scale deep net- work

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.121778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.186887Z digest=sha256:eab6b187db2c29eecd96987313ef98b1e4a3f4f8a3144544e54d2ab9ed74a9ab

Observation 19f79190-0953-41c4-9daa-0e0f6c04b649 · outbound

This paper cites Deep Ordinal Regression Network for Monocular Depth Estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep Ordinal Regression Network for Monocular Depth Estimation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.111325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.259930Z digest=sha256:fd92e578fbd327ad0d5f8abc6abdfb26457988898c261afcb911dea576214137

Observation e1676818-0491-41a7-bc6c-4f4c2cc0a5ea · outbound

This paper cites Gratt-vis: Gated residual atten- tion for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gratt-vis: Gated residual atten- tion for video instance segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.100906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.350163Z digest=sha256:cebbdbece79a39913c817fa65d0326f33d48e39dbec4771350c91d9b8a4edaae

Observation dfe9f3e3-c457-4e73-a771-9c6023ce1132 · outbound

This paper cites Deep residual learning for image recognition.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep residual learning for image recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.090612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.414701Z digest=sha256:fdf82a069f0eb05404e1a892cd85cb050ce17d36d3fb7269504f7cbc63a61244

Observation b898921b-6756-44f9-b53b-4806e2b439ab · outbound

This paper cites Vita: Video instance segmentation via object token association.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vita: Video instance segmentation via object token association

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.080147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.485519Z digest=sha256:242537683875495ee59efade7dfd1338ccea65ba7b5036256b4226bc388cd0bc

Observation d1b04145-c257-4b6e-8936-eb3f465a6564 · outbound

This paper cites A generalized framework for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A generalized framework for video instance segmentation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.067855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.555234Z digest=sha256:3dc9a0f3f2ba2aa88b9a6ecb6b30332e4c83b164ecdda2299691acfd6c14f3aa

Observation 47c16f76-4c88-48f2-8444-e17c89f1fdba · outbound

This paper cites Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.056740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.638086Z digest=sha256:a78b4284c9cfafa137688c739a5133935938e9b5080cadebab697350cc17a69a

Observation a80c305d-816f-4cb4-92ea-bdd5366268d3 · outbound

This paper cites Min- vis: A minimal video instance segmentation framework without video-based training.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Min- vis: A minimal video instance segmentation framework without video-based training

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.046211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.704845Z digest=sha256:0a86dcd4c5978a0cea26bb1fb8f59f45c0355d3468f3ba5a38c2a85e03cae7fb

Observation 92e1425d-cafc-45f8-89f8-662caf0af0ba · outbound

This paper cites Video object segmentation with language referring expressions.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video object segmentation with language referring expressions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.035711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.777215Z digest=sha256:b58fc5ac055ea5dae479ccad810abd550c910a709dcda91fc8a80b3c9b21e66d

Observation f48d6ef3-5bee-4933-8c66-12578057479c · outbound

This paper cites Self-Supervised Monocular Depth Es- timation: Solving the Dynamic Object Problem by Seman- tic Guidance.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Self-Supervised Monocular Depth Es- timation: Solving the Dynamic Object Problem by Seman- tic Guidance

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.023162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:32.849148Z digest=sha256:c570e7fd0a8a427a5b9b276f5adc1a231b9a1685e7f26d92864e43e4696943a7

Observation a105ac78-8cbf-4cd0-b8fd-1e2449f60562 · outbound

This paper cites CAVIS: Context-Aware Video Instance Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation CAVIS: Context-Aware Video Instance Segmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:32.925664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:32.925664Z digest=sha256:e6660316aa759b7dc29b35e23e664486ee043d4a210bfbc626a6aeac37ff4331

Observation 3e00813c-e74e-4afb-85a3-4eeab842981a · outbound

This paper cites Tcovis: Temporally consistent online video instance seg- mentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Tcovis: Temporally consistent online video instance seg- mentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:35.010380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.011814Z digest=sha256:ceccb2396f32f428a5924e74f0c5fec9a190e33b7a60fdc75937bc518d57e205

Observation 97a9959c-2aba-4451-9256-df2fac6f7eab · outbound

This paper cites Video k-net: A simple, strong, and unified baseline for video segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video k-net: A simple, strong, and unified baseline for video segmentation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.998688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.092555Z digest=sha256:1bc1dd7f0317569b4da170b28aeadb4dc8c618251de347c1d087cfd683e7bc3b

Observation 2f402b5d-1fa7-45e1-acea-a5cfffc5d4e9 · outbound

This paper cites Transformer-based visual segmenta- tion: A survey.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Transformer-based visual segmenta- tion: A survey

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.984574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.214189Z digest=sha256:09f89f04e1aa6f5f411a4fe4e228e21b635685d8e812b7ad0b15845b4f62ecec

Observation 1fb97ea8-58a2-4419-a407-c80f8ccb521d · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? In Conference on Computer Vision and Pattern Recognition,.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Omg-seg: Is one model good enough for all segmentation? In Conference on Computer Vision and Pattern Recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.972284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.279285Z digest=sha256:d4823a76cb25d1d928cc28fa675b14486d772ecc80c27f800fc5c060bc8904b0

Observation 68151916-4bf7-463c-89ba-f2551fba9e48 · outbound

This paper cites Lawrence Zitnick.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Lawrence Zitnick

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:33.336373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:33.336373Z digest=sha256:1ad9264289e456e2da420fec99a16765e9618bf60a09b9323e02abc5ce1a6709

Observation fb27e04d-cf28-4f62-9267-6313a358a888 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.956085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.406701Z digest=sha256:0d75d63c113c3d3ced2b3cdbdbde1b9a37baa91263156f560834a4d6e8da5552

Observation 54b62558-b8a6-4d0a-938d-3233084b5af9 · outbound

This paper cites Decoupled weight de- cay regularization.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Decoupled weight de- cay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:33.456111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:33.456111Z digest=sha256:3140f359ea84c6156999a600c9fbb3404b75269638829bdc0155385cdca2ab5d

Observation 8abc1bd3-3b74-4663-b440-c98b2466d8cb · outbound

This paper cites an unresolved cited work.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:17:34.937844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.525177Z digest=sha256:72f8b0996bb3e2e2a4014028fe4cc70b1cbd5a0eb95b90b30a41ff290ef2a6db

Observation b76ac901-acca-46a2-a6b2-c304a2344ab0 · outbound

This paper cites Keeping your eye on the ball: Trajec- tory attention in video transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Keeping your eye on the ball: Trajec- tory attention in video transformers

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.926944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.620382Z digest=sha256:1a22e870df804ba1932a863f98d390f2ac32fb849e9f82899582c11d8698c9e4

Observation 4561aba0-e6f1-4af7-aea8-2c6015f83e70 · outbound

This paper cites A benchmark dataset and evaluation methodology for video object segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A benchmark dataset and evaluation methodology for video object segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.913748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.639897Z digest=sha256:a3afbe194c6ee29e234522391c00606b6d7f8cd62855c57f9abd92f59f62f724

Observation 98ad0c44-f3dc-4db5-96f9-c420c1c8def2 · outbound

This paper cites Occluded video instance segmentation: A bench- mark.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Occluded video instance segmentation: A bench- mark

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.902520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.718789Z digest=sha256:9ad6f42fe685eb9edd9aa9a60c79d878222f3fc53f4abb00294384fe16996edf

Observation 096fb90e-f66d-4bc0-aaff-31b9fe73e979 · outbound

This paper cites ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:17:34.592459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.822660Z digest=sha256:22fa784e78f196d4d9d6ea1cc4fdab459f1a406d9bf5289779500bba607ef66a

Observation b8ff83cf-20d8-4991-94ad-7925dcb7cca5 · outbound

This paper cites Vi- sion transformers for dense prediction.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vi- sion transformers for dense prediction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.890407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:33.965005Z digest=sha256:e61bafff042675d1897cb46375f7b8a3ac29a81d56f2955f47ec43317dcaab67

Observation 069c13b9-6382-42de-8880-d43d1eddbc4b · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.880084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.075884Z digest=sha256:716e866e36503b11368e562a5f5d3edc176e480952ea0b1c705f352f31ee96bb

Observation b6ed3fb2-ff4f-49d4-b10b-fec2342cbc41 · outbound

This paper cites Boosting monocular depth with panoptic segmentation maps.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Boosting monocular depth with panoptic segmentation maps

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.868877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.280261Z digest=sha256:d8e27c89a8589218af8a12436d7e7f6c7dc3a6ca1d3e97cb0561e96fc93f2055

Observation faf64410-40a8-4a62-a17c-9a8c78694bee · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.856873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.460870Z digest=sha256:a484e3fbad5f64c3acf44c33d5b9a9f868e5eca3e89ce63e5c454ada14732c8f

Observation 2b611cf6-04e3-4099-900f-d70730844b99 · outbound

This paper cites Sigma: Siamese mamba network for multi-modal semantic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sigma: Siamese mamba network for multi-modal semantic segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.846739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.492979Z digest=sha256:0d94242cf8095c1574f722646cdf7e0fb80425b36dce07775b05022d77893e87

Observation f244ce5c-8d5f-48ac-8843-81dac7de53ce · outbound

This paper cites Sdc-depth: Semantic divide-and-conquer net- work for monocular depth estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sdc-depth: Semantic divide-and-conquer net- work for monocular depth estimation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.835266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.496984Z digest=sha256:9fc4314b9368fe94c0fcf9eb87a6bfbbb0bc056892f3fad0b33f2ca8e0d1dfe5

Observation 85d59161-665f-4a1c-ac11-b17da801093c · outbound

This paper cites End-to-end video instance segmentation with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to-end video instance segmentation with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.822760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.499751Z digest=sha256:838a6dd00ac85d6c88e0fc1ef1b083eea7cdc962bfc911ec5c6512718da106be

Observation a4cb9d1e-976b-410b-9635-cc681dc350b4 · outbound

This paper cites Metric3d: Towards zero-shot metric 3d prediction from a single image.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d: Towards zero-shot metric 3d prediction from a single image

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.812977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.503601Z digest=sha256:c8f654aaca3be2b2177b011c4862b24b8f4864919cd290c2cd02912ceb10b70e

Observation 44fad11c-bfcd-4554-bbf2-cabbca2bbdf3 · outbound

This paper cites Seqformer: Sequential transformer for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Seqformer: Sequential transformer for video instance segmentation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.803309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.507477Z digest=sha256:250f150e28a6bb2702bc830f5cc413e33d241e7c083bf76eb142316390b8bda7

Observation 0986d5df-f94b-4fe6-beb2-80d01711e7f5 · outbound

This paper cites In defense of online models for video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation In defense of online models for video instance segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.792784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.510411Z digest=sha256:d06955baa7c581faae3e225eeb019d049ab8e8ad97a3304a396dcdfd4da4c9c3

Observation a4435c75-3b7d-47ec-a1ab-a37ebf32d005 · outbound

This paper cites Depth Any Video with Scalable Synthetic Data.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth Any Video with Scalable Synthetic Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:17:34.513382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:17:34.513382Z digest=sha256:67651fe3b934b2396c7b27830f37f00d3700f4cf4c2855b6b9da82a7330130e2

Observation b603f057-5f35-41a2-9a6e-ca11f6a59264 · outbound

This paper cites Video instance seg- mentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video instance seg- mentation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.781260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.516735Z digest=sha256:6c3ea5b33e03ec4e24a10aa90dabe2953aeb08964a22949a51bf88c5e13311c0

Observation 5ce503cb-9725-4c3e-bc4e-48fdd850bb6d · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth anything: Unleashing the power of large-scale unlabeled data

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.769290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.519533Z digest=sha256:3d6ebb2bb00502bfdd5edd76c1aebf1f3e8b097dd21fee352d269f6aa1ddb4aa

Observation 249a832b-6022-41b9-bfe0-e82b34dad44e · outbound

This paper cites Depth any- thing v2.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth any- thing v2

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.757631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.522356Z digest=sha256:7ff9212b8de6ab18277c3657c4eac56047a8673dfcad2c67a328af84f96221c0

Observation 43c01758-bda8-406c-9a97-f7f7c2c6ee56 · outbound

This paper cites Ctvis: Consistent train- ing for online video instance segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Ctvis: Consistent train- ing for online video instance segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.745864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.525132Z digest=sha256:075e7db359ead2da3f2f0eb8d9a213f0c3afb22e49b31f4fccb6f6ad5c5002c0

Observation 9448895f-a749-442b-907e-7e7bed8179af · outbound

This paper cites Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.733631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.528509Z digest=sha256:d352e4213e55f6c4a1bb43749f927e662888bc42c4ede3a44e688f4e1ae146e1

Observation ee2c0955-4265-45f8-bb61-2b8b621a1eae · outbound

This paper cites Geometry meets semantic for semi-supervised monocular depth estimation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Geometry meets semantic for semi-supervised monocular depth estimation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.721040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.531650Z digest=sha256:2ac2daece2c0197e6d1a90d1f33e75a3e314bbe39152ad407ce58ed009c5e981

Observation cf7c2cbb-991d-4625-a445-749a2e551d55 · outbound

This paper cites Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.710318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.534259Z digest=sha256:95fc0cb5af15da76ebd5bfc0f2f3dc6eef6aa9837c4edc7d5b41e427bb89c440

Observation 33a3296b-c367-409b-b49a-db6338d5efd5 · outbound

This paper cites Delivering arbitrary-modal semantic segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Delivering arbitrary-modal semantic segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.698098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.536874Z digest=sha256:a1e4d56f618a3eaf793217352b9c6f4fa51d203e9acc1acff0bce441e33f1a6d

Observation 4ba3d615-4146-414e-975e-0459af2efcc2 · outbound

This paper cites Dvis: Decoupled video in- stance segmentation framework.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis: Decoupled video in- stance segmentation framework

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.688108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.539546Z digest=sha256:bbeb6dde8bc5ff0dcb315871582e63e908d014364f78bd9a15e0ca5947e1ad0f

Observation 648c8882-4c46-458c-8834-57f0771046d8 · outbound

This paper cites Dvis++: Improved decoupled frame- work for universal video segmentation.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis++: Improved decoupled frame- work for universal video segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.677425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.542429Z digest=sha256:fefda779e8fb7d49103943637dc598222f50cbc7c4c25938abd0086853ba9231

Observation fc0301be-19a4-4b39-a3f9-c4686cf35dbe · outbound

This paper cites Dvis-daq: Improving video segmentation via dynamic anchor queries.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis-daq: Improving video segmentation via dynamic anchor queries

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.666323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.545623Z digest=sha256:5faffb4bff07fd4e87c61c1ac429f8f66a05562315e7c186970240f2c2bc4322

Observation f903c52f-5a6a-4f45-9b73-c3fdb9922883 · outbound

This paper cites Deformable detr: Deformable transformers for end-to-end object detection.

Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deformable detr: Deformable transformers for end-to-end object detection

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:17:34.655837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:17:34.548221Z digest=sha256:7c6016aeb256149d931cffd010f82a1b12e843ed93746e1d2b9da6a4dcb956d6

Pith citing papers

No inbound Pith citation observations are available.