Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:17:34.548221Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2507.05948.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:17:34.548221Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44cec9f3-2f55-4dd1-946f-2731467497a5 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stem-seg: Spatio-temporal em- beddings for instance segmentation in videos
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4aa3b6fc-8dfd-4296-950a-de684d2ad91e · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Is space-time attention all you need for video understanding? In International Conference on Machine Learning, 2021
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fd57a2b-e37a-4bd8-a81b-f11ca7fe49d5 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MiDaS v3.1 -- A Model Zoo for Robust Monocular Relative Depth Estimation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f42aa1f6-bfb6-4184-b5cd-a5166450ac14 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40888369-7740-4196-8a40-0ca91b09051e · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to- end object detection with transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de672815-1933-4925-b478-5ca43aa28665 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Liu, Yen-Cheng Liu, and Yu- Chiang Frank Wang
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 883b2de5-6e7e-4ce7-82b9-f04d83ff57dc · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vision Transformer Adapter for Dense Predictions
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bee4c805-d516-4b99-a360-31bdf9138c91 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Collins, Yukun Zhu, Ting Liu, Thomas S
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd7d29b2-c6db-4d8c-b6fc-7cec9ca92856 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Mask2Former for Video Instance Segmentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2b25763-c8f1-4016-af5e-0fa2158affd1 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Per- pixel classification is not all you need for semantic segmen- tation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6b9670c-78a4-47e4-9fc6-780bc041e9df · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Masked-attention mask transformer for universal image segmentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eecf82d2-8b0c-41c0-bc50-4a1879bd9af6 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MeViS: A large-scale benchmark for video segmentation with motion expressions
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f81a655b-2269-4813-b6dd-62845d006d2c · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation MOSE: A new dataset for video object segmentation in complex scenes
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ea2f2aa-c283-4462-aa95-184f0563ba7c · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation An image is worth 16x16 words: Transformers for image recognition at scale
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2037d9f0-e08e-4e28-a9a3-0c6ebf7a7c0c · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth map prediction from a single image using a multi-scale deep net- work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19f79190-0953-41c4-9daa-0e0f6c04b649 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep Ordinal Regression Network for Monocular Depth Estimation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1676818-0491-41a7-bc6c-4f4c2cc0a5ea · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gratt-vis: Gated residual atten- tion for video instance segmentation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfe9f3e3-c457-4e73-a771-9c6023ce1132 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deep residual learning for image recognition
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b898921b-6756-44f9-b53b-4806e2b439ab · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vita: Video instance segmentation via object token association
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1b04145-c257-4b6e-8936-eb3f465a6564 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A generalized framework for video instance segmentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47c16f76-4c88-48f2-8444-e17c89f1fdba · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d v2: A versatile monocular geomet- ric foundation model for zero-shot metric depth and surface normal estimation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a80c305d-816f-4cb4-92ea-bdd5366268d3 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Min- vis: A minimal video instance segmentation framework without video-based training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92e1425d-cafc-45f8-89f8-662caf0af0ba · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video object segmentation with language referring expressions
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f48d6ef3-5bee-4933-8c66-12578057479c · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Self-Supervised Monocular Depth Es- timation: Solving the Dynamic Object Problem by Seman- tic Guidance
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a105ac78-8cbf-4cd0-b8fd-1e2449f60562 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation CAVIS: Context-Aware Video Instance Segmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e00813c-e74e-4afb-85a3-4eeab842981a · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Tcovis: Temporally consistent online video instance seg- mentation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97a9959c-2aba-4451-9256-df2fac6f7eab · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video k-net: A simple, strong, and unified baseline for video segmentation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f402b5d-1fa7-45e1-acea-a5cfffc5d4e9 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Transformer-based visual segmenta- tion: A survey
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fb97ea8-58a2-4419-a407-c80f8ccb521d · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Omg-seg: Is one model good enough for all segmentation? In Conference on Computer Vision and Pattern Recognition,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68151916-4bf7-463c-89ba-f2551fba9e48 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Lawrence Zitnick
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb27e04d-cf28-4f62-9267-6313a358a888 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Swin transformer: Hierarchical vision transformer using shifted windows
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54b62558-b8a6-4d0a-938d-3233084b5af9 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Decoupled weight de- cay regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8abc1bd3-3b74-4663-b440-c98b2466d8cb · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b76ac901-acca-46a2-a6b2-c304a2344ab0 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Keeping your eye on the ball: Trajec- tory attention in video transformers
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4561aba0-e6f1-4af7-aea8-2c6015f83e70 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation A benchmark dataset and evaluation methodology for video object segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 98ad0c44-f3dc-4db5-96f9-c420c1c8def2 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Occluded video instance segmentation: A bench- mark
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 096fb90e-f66d-4bc0-aaff-31b9fe73e979 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation ViP-DeepLab: Learning Visual Perception with Depth-aware Video Panoptic Segmentation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8ff83cf-20d8-4991-94ad-7925dcb7cca5 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Vi- sion transformers for dense prediction
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 069c13b9-6382-42de-8880-d43d1eddbc4b · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6ed3fb2-ff4f-49d4-b10b-fec2342cbc41 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Boosting monocular depth with panoptic segmentation maps
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation faf64410-40a8-4a62-a17c-9a8c78694bee · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b611cf6-04e3-4099-900f-d70730844b99 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sigma: Siamese mamba network for multi-modal semantic segmentation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f244ce5c-8d5f-48ac-8843-81dac7de53ce · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Sdc-depth: Semantic divide-and-conquer net- work for monocular depth estimation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85d59161-665f-4a1c-ac11-b17da801093c · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation End-to-end video instance segmentation with transformers
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4cb9d1e-976b-410b-9635-cc681dc350b4 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Metric3d: Towards zero-shot metric 3d prediction from a single image
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44fad11c-bfcd-4554-bbf2-cabbca2bbdf3 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Seqformer: Sequential transformer for video instance segmentation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0986d5df-f94b-4fe6-beb2-80d01711e7f5 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation In defense of online models for video instance segmentation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4435c75-3b7d-47ec-a1ab-a37ebf32d005 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth Any Video with Scalable Synthetic Data
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b603f057-5f35-41a2-9a6e-ca11f6a59264 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Video instance seg- mentation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ce503cb-9725-4c3e-bc4e-48fdd850bb6d · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth anything: Unleashing the power of large-scale unlabeled data
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 249a832b-6022-41b9-bfe0-e82b34dad44e · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Depth any- thing v2
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43c01758-bda8-406c-9a97-f7f7c2c6ee56 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Ctvis: Consistent train- ing for online video instance segmentation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9448895f-a749-442b-907e-7e7bed8179af · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Polyphonicformer: Unified query learning for depth-aware video panoptic segmentation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee2c0955-4265-45f8-bb61-2b8b621a1eae · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Geometry meets semantic for semi-supervised monocular depth estimation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf7c2cbb-991d-4625-a445-749a2e551d55 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Cmx: Cross-modal fusion for rgb-x semantic segmentation with transformers
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 33a3296b-c367-409b-b49a-db6338d5efd5 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Delivering arbitrary-modal semantic segmentation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ba3d615-4146-414e-975e-0459af2efcc2 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis: Decoupled video in- stance segmentation framework
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 648c8882-4c46-458c-8834-57f0771046d8 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis++: Improved decoupled frame- work for universal video segmentation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc0301be-19a4-4b39-a3f9-c4686cf35dbe · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Dvis-daq: Improving video segmentation via dynamic anchor queries
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f903c52f-5a6a-4f45-9b73-c3fdb9922883 · outbound
Beyond Appearance: Geometric Cues for Robust Video Instance Segmentation Deformable detr: Deformable transformers for end-to-end object detection
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.