Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:01:18.047778Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 100 of 110 outbound references and 2 inbound Pith citation observations for arXiv:2508.21809.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T14:01:18.047778Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T00:56:18.867382Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:27:15.698319Z
100 of 110 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9177577e-76bf-47c2-b879-04cb54ee956f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b28fbf4-cae5-49e4-b130-0ecb18ca97cd · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e08dbdd-2edc-4f41-bce3-00a182780fd5 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Context r-cnn: Long term temporal context for per-camera object detection
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2874a56a-9b25-4054-b70e-13bcec7da245 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt JAX: composable transformations of Python+NumPy programs, 2018
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f439b7b-b054-4b52-8a43-5bb14697f2f5 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt The 2019 DAVIS Challenge on VOS: Unsupervised Multi-Object Segmentation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d125a6c-7184-413c-8e2a-93cdfe3cb89e · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt nuscenes: A multimodal dataset for autonomous driving
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df8eea38-b958-46b0-88d4-aedcbe49966f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Stablevideo: Text-driven consistency-aware diffusion video editing
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b95b0c0-f4fe-4c59-9daf-220cedfc8e0e · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt PaLI: A Jointly-Scaled Multilingual Language-Image Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58a7584-aef2-4cc0-8aa1-b51c705e54b3 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Per-pixel classification is not all you need for semantic segmentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e9961d-3d4b-4c70-bf9a-e5c47ce592f4 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd00fb5a-7009-4953-9119-15a1db801fea · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Putting the object back into video object segmentation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d34e96-2a6a-4207-a531-14fd3f85fee9 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Segment and Track Anything
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ef5bd6-536b-4f9c-a433-c3cbf1968cd2 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Find first, track next: Decoupling identification and propagation in referring video object segmentation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6eb5932-be1c-44b8-927a-9f6c338c64f2 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Ow-viscaptor: Abstractors for open-world video instance segmentation and captioning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 721f38f1-5d3f-41d8-88f2-94a18394aa17 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Scenic: A jax library for computer vision research and beyond
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a7ea64-fec1-4e92-9025-369d3f0ecf66 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Memsam: Taming segment anything model for echocardiography video segmentation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 508e5e38-71f8-4802-a7f4-7bda04a4eca4 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Bert: Pre-training of deep bidirectional transformers for language understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45a4fb90-4416-4353-bf32-56a713abf84c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Mevis: A large-scale benchmark for video segmentation with motion expressions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00dd6e78-bbe7-4fe8-9136-235ab55110a5 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt MOSE: A new dataset for video object segmentation in complex scenes
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67cbf05-cbd4-482f-920d-67c153d78ab6 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt An image is worth 16x16 words: Transformers for image recognition at scale
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6747a0f9-37f7-4fcf-9395-2915513161f3 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt EVA-02: A Visual Representation for Neon Genesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b629c41-94ce-44c9-8d93-a6c4abe62451 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8176bb-ec41-460f-a59b-d8711ef0de9e · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt VideoSAM: Open-World Video Segmentation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6d7c230e-35c5-4568-906a-76a94457ba27 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Masked autoencoders are scalable vision learners
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6702ecee-8c0c-42ba-a0f7-6d2634eee613 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Decoupling static and hierarchical motion perception for referring video segmentation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c6717fd-d5db-4a7e-a5e3-b13941ee6882 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Instruct-imagen: Image generation with multi-modal instruction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259459b1-63b5-4526-a464-f708dd6b1677 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Segment and caption anything
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d9da2d-f6c1-4ba2-b1bf-fd554fb63942 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt A better use of audio-visual cues: Dense video captioning with bi-modal transformer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b991e6-60d9-40a5-ba43-cdb91bcaa6b7 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Perceiver: General perception with iterative attention
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 075a4d78-04a8-40e2-b57e-23de063ff4f5 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Scaling up visual and vision-language representation learning with noisy text supervision
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c3529bb-aa1e-468e-8f57-7a1d95d60664 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Densecap: Fully convolutional localization networks for dense captioning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c77ef26a-3202-4983-88b7-8ebf93db0f22 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Kanani, Sriparna Saha, and Pushpak Bhattacharyya
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2069cbfc-56f4-47ec-9832-55555d20f88a · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Video object segmentation with language referring expressions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c63d63-bd18-4096-8acf-48f1bf76e627 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Video panoptic segmentation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0e447a3-09f0-47d3-88bc-0a9aaaa470be · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Segment anything
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ff2a8b-c79d-417b-8f09-65ed0a8c227c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Dense-captioning events in videos
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac9522b-363c-4045-8bff-59058700cd60 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02709fa6-6806-4d60-91e0-406cdbac60d9 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Bidirectional Correlation-Driven Inter-Frame Interaction Transformer for Referring Video Object Segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e5c2ef69-ed6d-4bd2-a648-46efa7d3c36b · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88920aa7-0cda-42af-9c7d-dc9954f24fbd · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0aad7dbe-5ff1-4662-80ab-fe44c90c0184 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Learning object context for dense captioning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d712224b-a72a-4e4a-8f3e-a0bfed0eaa41 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Exploring plain vision transformer backbones for object detection
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0a7bcacd-eeb5-4fb0-93e9-c08967baec2c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Beyond mot: Semantic multi-object tracking
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b819030b-abb2-474b-a43c-1fdb5303f7b7 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e05588-bf31-435b-ac20-27d2e85eb79f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Visual instruction tuning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 209ca39c-956b-4d91-a462-2f5cd2b9b926 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Image segmentation using text and image prompts
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ac956e77-cd96-456a-bf9d-d075ed517381 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Soc: Semantic-assisted object cluster for referring video object segmentation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1b214c00-92d9-4c56-820f-076d44c8e8f8 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Generation and comprehension of unambiguous object descriptions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97fb6b6c-8501-423f-9c11-619ec0520100 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Scaling open-vocabulary object detection
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c0e9d8b2-f9da-4f7d-a238-3d5f21ac534c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Pivot: Iterative visual prompting elicits actionable knowledge for vlms
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e8613517-5df2-446e-b494-aa657ac9f9a4 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Gpt-4v(ision) technical work and authors
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c99daecb-c971-4c8e-b2cf-4124212af177 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a349ecd-bd05-47ae-9458-a62320139450 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Perazzi, J
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 047b2d75-2e1b-4597-8ec3-1352b02d854f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation daf5929b-7733-4dd7-b7d2-2f9781d1fcf8 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Connecting vision and language with localized narratives
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f433843-de19-4feb-9c59-54571dcbd76e · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 00934094-f590-4efc-8065-b9b4418ece4c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Learning transferable visual models from natural language supervision
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 67b91da0-785f-4eaf-a0f7-61732255d9dd · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 015ba149-0ab4-47dd-a96f-453410770f6f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt SAM 2: Segment Anything in Images and Videos
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298b0a5c-55b7-446e-9eec-1bfdfe473786 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Youtube- boundingboxes: A large high-precision human-annotated data set for object detection in video
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d1f04666-ef6f-4fd3-b188-1f2075ce5503 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt U-net: Convolutional networks for biomedical image segmentation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c89d1ec-60f0-4ec5-8062-292d5eefe26a · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Berg, and Li Fei-Fei
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9c714a6c-cc32-4d0d-8fb1-c6e0d7e68ec9 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Hiera: A hierarchical vision transformer without the bells-and-whistles
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c36ae280-65e0-471a-a2e6-738d8c47fe3d · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Tokenlearner: Adaptive space-time tokenization for videos
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation de29efa0-952a-47d5-94d1-8310d0897d9c · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Urvos: Unified referring video object segmentation network with a large-scale benchmark
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eb82c4d2-4bf1-49f2-9550-83c922099e37 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Annotating objects and relations in user-generated videos
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 317edd14-e39c-4dbb-984f-57f9c1f2ccab · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Region-object relation-aware dense captioning via transformer
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b4551fa-fdd1-4b54-a5a1-245f1020df62 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt What does clip know about a red circle? visual prompt engineering for vlms
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 41ad1629-a574-43b7-84d9-88bfbd4b6591 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Video foundation models for animal behavior analysis
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8e79880d-d477-4a52-9edc-08113a32497b · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Scalability in perception for autonomous driving: Waymo open dataset
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66638174-134c-448a-8f47-8dd4c160250f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Gemma: Open Models Based on Gemini Research and Technology
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494a617f-2afe-4654-8bf5-8521bb3e8b92 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Gemma 2: Improving Open Language Models at a Practical Size
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b3fd1a2-1124-4050-8d20-b132447d8328 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt LLaMA: Open and Efficient Foundation Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee97b73b-be38-47f1-8c9a-70004761ad0f · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Attention is all you need
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fd80b34-72bd-4cc3-9c36-b1359b36f6c2 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Cider: Consensus-based image description evaluation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 04161a33-0846-43ff-a739-6bf328814047 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Phenaki: Variable length video generation from open domain textual descriptions
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 93bb1291-efe9-4fea-a86c-4c3458880499 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Connecting vision and language with video localized narratives
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 76f7b4b0-3460-45d5-a143-cdc5a9dbe52a · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Git: A generative image-to-text transformer for vision and language
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ddd0f84e-71ef-41e2-b5b5-b85091610457 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt End-to-end dense video captioning with parallel decoding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d78b6e5f-978c-4106-b90a-5f503598f5ac · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Caption Anything: Interactive Image Description with Diverse Multimodal Controls
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edfe104c-d0a3-4856-b530-fe9e64d72b31 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Unidentified video objects: A benchmark for dense, open-world segmentation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5dccc15b-8b05-4d5b-9b86-998459bd7ad9 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt The all-seeing project: Towards panoptic visual recognition and understanding of the open world
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1daf37a9-7708-467f-9ccb-7631b853dd2b · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt The All-Seeing Project V2: Towards General Relation Comprehension of the Open World
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a59b876c-4e5d-41bb-aa62-b96cdfca49c2 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Instancediffusion: Instance-level control for image generation
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d24b68d0-16ec-4dfa-bc4b-bb4a07d1beb5 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Language as queries for referring video object segmentation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2e6c14ce-75b9-49c9-96ef-bcefd03f0a57 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b41afe1f-22a7-4ab7-9a9e-29c1d14a7a23 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt General object foundation model for images and videos at scale
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6ba209f7-85dc-46f5-be63-4b677a35aee7 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Grit: A generative region-to-text transformer for object understanding
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bd63d38f-194e-4ba3-9319-3ee30bf19a55 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Dettoolchain: A new prompting paradigm to unleash detection ability of mllm
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a952e8fd-2d3c-4193-abab-4eb57748c530 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Pixel-aligned language model
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 10886054-f88b-4cdd-8a12-d32a13b39de9 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Youtube-vos: Sequence-to-sequence video object segmentation
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0ea80731-72c7-4620-b19b-ed2f23f557f6 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt xgen-mm (blip-3): A family of open large multimodal models
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 524bf4cc-24fa-4109-9c75-fe402ae84dab · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dc07d712-d243-49bb-9167-63d0ab8cde11 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Track Anything: Segment Anything Meets Videos
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0deecbdc-995a-4edd-b1b3-a9fed35a2141 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d442d00-cc26-4cd2-ab3a-5d1ef693dfdc · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Decoupling features in hierarchical propagation for video object segmentation
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3d09fbf5-b2e4-47fe-925c-160612d99570 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Associating objects with transformers for video object segmentation
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4f0f111-19a8-4746-a4f8-cd040d2366b8 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Scalable video object segmentation with identification mechanism
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1c1d4855-d106-4023-abbd-8fadeba161e6 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Describing videos by exploiting temporal structure
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 37ff9346-5dfa-40b8-b6c7-454f4ad94255 · outbound
VoCap: Video Object Captioning and Segmentation from Any Prompt Modeling context in referring expressions
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 501080c7-8714-494f-9325-48e7a4702264 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs VoCap: Video Object Captioning and Segmentation from Any Prompt
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c1f47cb2-a1bc-4129-a261-406d0b614598 · inbound
Video Generation Models are General-Purpose Vision Learners VoCap: Video Object Captioning and Segmentation from Any Prompt
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.