Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:34.495197Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2411.10922.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:34.495197Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
91 of 91 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation db3981e3-74ce-482e-bbb4-dd3ac75ca216 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Bridg- ing the gap between object and image-level representations for open-vocabulary detection
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17fc7182-20d2-41c2-95e3-017c4828fb4d · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection End- to-end object detection with transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03136e5-d99e-44cd-bf34-03704abbffe5 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Emerg- ing properties in self-supervised vision transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b891a697-38a8-4291-87b8-3a0063cc9ca9 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaa2407a-400c-4507-8cbb-083426c58af3 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e5f1f7ba-2173-4c7a-a937-a29ce4570943 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Efficient video action detection with token dropout and context refinement
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751bcd4b-75e4-4952-9bb9-4b5230a32a7a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Watch only once: An end-to-end video action detection framework
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e35feb40-0671-4a36-b17c-bee968bc33d5 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Gabriellav2: Towards better generalization in surveillance videos for ac- tion detection
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899f6d60-3d9a-423f-95ef-a1f52c97fe31 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection An image is worth 16x16 words: Trans- formers for image recognition at scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98a2ede4-d1b3-4ed4-89fc-1cc8556ffbad · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for open-vocabulary ob- ject detection with vision-language model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd0b121-fc8f-4d39-82ed-b4a83da6c90e · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Holistic interaction transformer network for action detection
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbb3167-c1e4-4cb3-9bdd-c19cd89a3440 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection X3d: Expanding architectures for efficient video recognition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d8fe21-be1c-4d83-bd36-40bd9b5e3e62 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Slowfast networks for video recognition
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 39c3202b-f348-4ef9-9528-cdfa635e7663 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-adapter: Better vision-language models with feature adapters
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6c8f9fde-1643-4644-a992-a20489ee6010 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Adamixer: A fast-converging query-based object detector
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3647d439-e593-4e2e-9808-f6f413784297 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Video action transformer network
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ada785ec-8665-41e8-a8f3-868448bcf185 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ava: A video dataset of spatio-temporally localized atomic visual actions
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0948440f-5a6f-41f8-b49b-dbc1d9aef6cd · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8dcb99de-b9f4-4ade-afba-f62d7bb1d01e · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b7ff74fc-5739-40ef-81bb-4a05f78aa929 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Interaction-aware prompting for zero-shot spatio-temporal action detection
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c0f167db-0d6f-4c01-b12b-11a15a5744f7 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards understanding ac- tion recognition
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 48bfafee-2a01-41bd-a998-cac1e8fb3254 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f2e8739-d94a-49ee-8651-719104ac68a4 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58b2af96-fcf7-41c8-b199-885dd85a5e02 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Action tubelet detector for spatio- temporal action localization
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f4cf0649-040f-417d-80a7-279a2b3bd1cd · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Region- aware pretraining for open-vocabulary object detection with vision transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 13de2aab-4ba6-495f-9b9c-b9dac35bac06 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627737ca-7693-4dc1-aec4-a29fc32d4c5d · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection F-vlm: Open-vocabulary object detection upon frozen vision and language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a6de3308-f534-4333-adba-57ddb94b64a6 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Multisports: A multi-person video dataset of spatio-temporally localized sports actions
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 992132a9-a5a3-44d3-a65c-db8ca27af360 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring plain vision transformer backbones for object de- tection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6060094-1d26-4859-9fa5-e6c7837ca89c · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection A Closer Look at the Explainability of Contrastive Language-Image Pre-training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5d97d2-4307-47ec-bd8c-6a10c6b79641 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring Visual Interpretability for Contrastive Language-Image Pre-training
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d102cf0-b0a6-4361-bb6a-81b2b6873817 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary semantic segmentation with mask-adapted clip
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64abda54-12b0-469d-9f22-61e2b7c3dd2a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning object-language alignments for open-vocabulary object de- tection
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 782e6c26-1364-4b98-9f88-2eaa67143b93 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Frozen clip models are efficient video learners
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fdf89e7f-b712-4165-84cc-0524b34cb94b · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 83cf222f-01f6-4fc0-b64d-7d5bc2a43bf3 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5e553e47-f158-4342-b759-54babcbf0257 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c08765bf-906b-45d3-8e29-d86035509d44 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Decoupled weight decay regularization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e3dd92-e866-4dcc-8f9f-485a8a455d83 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Verbs in action: Improv- ing verb understanding in video-language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation aa9df3ab-1c93-4e5c-af9e-c2f424e757b2 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0db54a28-09cc-447f-80f8-dedade732fd3 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9944dc8b-70e2-4561-9129-f828df311ed2 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ebb2f9c1-454b-48a6-9593-05de32ece069 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d4e87b73-854e-4aee-b8e7-adb107377104 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection GPT-4 Technical Report
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da5c551-d16e-4ac1-ae8c-614c9c120058 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-context-actor relation net- work for spatio-temporal action localization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 81dffbc5-8af5-4945-94d0-0e49a45e975f · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection St-adapter: Parameter-efficient image-to-video transfer learning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b5874fa3-658e-4643-a83d-ff134bca9207 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d2e88c0-1bb0-42e1-9891-0f9add2444a0 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2aae92bd-2bf9-4752-8b8f-0f0205d3967c · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Fine-tuned clip models are efficient video learners
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e253c8d0-066f-41b0-9685-f289a970ef83 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary temporal action detection with off-the-shelf image-text features
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 083eb89f-447a-48a8-8054-c1df10550bad · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 724675cb-153f-4780-b4b2-a6798463402a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generalized in- tersection over union: A metric and a loss for bounding box regression
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59106dfc-b53e-4ec7-88b5-be3eeeb0d5b6 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse DETR: Efficient end-to-end object detec- tion with learnable sparsity
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f93e6906-a7cc-4ba2-9da8-4af4973beca9 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Hi- era: A hierarchical vision transformer without the bells-and- whistles
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 91d9c389-3853-478b-9aa6-f4b0d78c39f5 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grad-cam: Visual explanations from deep networks via gradient-based localization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5bdd13fe-1598-41bc-b9a9-e7c8978c548c · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Road: The road event awareness dataset for autonomous driving
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c60eda70-48eb-4399-bea5-d0e95325d959 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ucf101: A dataset of 101 human actions classes from videos in the wild
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9e8d9b78-66a6-4176-b7c2-12f4eb1c56f1 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-centric relation network
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e6691d63-bfe3-4569-b800-587c8c19d26e · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Relational action forecasting
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d2becb4-77a4-4488-9e45-fb5cf60f330a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse r-cnn: End-to-end object detec- tion with learnable proposals
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3b7da0b-fc8a-4876-8656-6f785bd2d0ec · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Asynchronous interaction aggregation for action detection
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9b3ac7e4-8a5e-4709-a9a6-a0f99214bebb · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mlp-mixer: An all-mlp architecture for vision
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d4ded328-8ca6-40d1-8e5d-ea4b27506309 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Attention is all you need
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2862efa0-b746-4ef8-8848-e4580b94e97a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection ActionCLIP: A New Paradigm for Video Action Recognition
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef25c1c-96cf-4f48-a3cf-a785f982f251 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vita-clip: Video and text adaptive clip via multimodal prompting
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 49b02c60-0397-4add-bb39-60f46e93c4b0 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 22930374-aaaa-4cb5-bc82-56616bdefd0b · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Long-term feature banks for detailed video understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 81150c78-644c-49be-82a4-af4bbc0f6241 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Context-aware rcnn: A baseline for ac- tion detection in videos
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fa01248c-5077-4de8-8e0e-ca3bf8273f72 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards Open Vocabulary Learning: A Survey
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 512e006d-b5fa-4039-8c0f-5ce3411677e8 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6ed96869-fb21-4314-ad30-5ee6828fb25e · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6228390b-7bcb-49f7-b9e7-eb4aeaf4a8f0 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-Vocabulary Spatio-Temporal Action Detection
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a450ad62-ef08-4428-a427-86f9125e5f3a · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4a15141a-cc08-4019-9d1f-55152793a02b · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5576a9a8-ab22-4929-9de9-141a61b7f618 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1dde812-30d6-474c-a4db-8f8112fc3d92 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Unloc: A unified framework for video localization tasks
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6a60e58f-1d3a-480f-ad50-8620f40144e3 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Detecting human actions in surveillance videos
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee6e3520-08f2-4373-9558-c55e4b34417c · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Contextualized spatio-temporal contrastive learning with self-supervision
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1a35debf-220f-4151-9dcf-e3d2c372e435 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vision-based garbage dumping action detection for real-world surveillance platform.ETRI Journal, 41(4):494–505, 2019
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba00c246-766c-4df0-9b77-2e412d624b1c · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary detr with conditional matching
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 778c0bbe-b2f3-4369-b6e5-55093f7d3d82 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c3229a69-53c5-4c6a-b995-cd6e3a770985 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f94dff5c-4afc-40ef-b6f8-8bf85f6e1a47 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection From recognition to cognition: Visual commonsense reason- ing
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c2e45b16-72d3-4443-9478-a1e903d3ec23 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vinvl: Revisiting visual representations in vision-language models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2316d72b-c42a-4441-a29f-faeed5df7917 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Tuber: Tubelet transformer for video action detection
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fb604aaf-58c5-40a6-8b26-0a72517c0866 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection MRSN: Multi-Relation Support Network for Video Action Detection
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf65d351-a680-4c7c-a476-f51acb3f9e45 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Regionclip: Region-based language-image pretraining
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f0230f2-e4ce-4cdc-98e5-cc72b0593fee · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning deep features for discrimi- native localization
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2a8ecf9d-e5b1-4f1f-9cea-074b3d180b78 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for vision-language models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d972e5f-770b-48e3-93a5-930a182a96a3 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zegclip: Towards adapting clip for zero-shot se- mantic segmentation
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6347dfc0-d29f-4afe-915d-746364a3c506 · outbound
Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection For the action type {CLS}, what are the visual descriptions? Please respond with a list of 16 short sentences
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.