Pith. sign in

Paper Citation Record · LEDGER

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection

As of 20 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2411.10922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10922 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:34.495197Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact2
  • verified fuzzy65
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db3981e3-74ce-482e-bbb4-dd3ac75ca216 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Bridg- ing the gap between object and image-level representations for open-vocabulary detection

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.897362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.897362Z digest=sha256:1ac70d661e18a6bf7aac15dc28567d74c310ef1fc2fd1433d2eb8f7f30e0caf8

Observation 17fc7182-20d2-41c2-95e3-017c4828fb4d · outbound

This paper cites End- to-end object detection with transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection End- to-end object detection with transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.902105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.902105Z digest=sha256:0f4e6c49aa276db36df7958c038a2cbe504d9e3dbec0fd5fcd41a40a589dc50d

Observation b03136e5-d99e-44cd-bf34-03704abbffe5 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Emerg- ing properties in self-supervised vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.908058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.908058Z digest=sha256:92f596d3fd0d79ace0b6422e1702ebb8c169ce71685b39032ca9e885bfd42266

Observation b891a697-38a8-4291-87b8-3a0063cc9ca9 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.913872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.913872Z digest=sha256:f1e9d17c653a18964630ac18c32cd4566381b7cd34e47e749f8e584764743187

Observation eaa2407a-400c-4507-8cbb-083426c58af3 · outbound

This paper cites CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:14:34.741882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.920517Z digest=sha256:2f0e928b6ab49897c9aadb0642ed3a4d74b43f3d3cd21308978de9ca8e50df2f

Observation e5f1f7ba-2173-4c7a-a937-a29ce4570943 · outbound

This paper cites Efficient video action detection with token dropout and context refinement.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Efficient video action detection with token dropout and context refinement

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.926112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.926112Z digest=sha256:53eb6244e35d62ea20a51fc313cb96804e69fc79fd05b0b22d3e33fad475e488

Observation 751bcd4b-75e4-4952-9bb9-4b5230a32a7a · outbound

This paper cites Watch only once: An end-to-end video action detection framework.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Watch only once: An end-to-end video action detection framework

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.931019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.931019Z digest=sha256:cce2fe3cba142c4639a86f7dd075b24ebead9ea598b4715978faf936acaace16

Observation e35feb40-0671-4a36-b17c-bee968bc33d5 · outbound

This paper cites Gabriellav2: Towards better generalization in surveillance videos for ac- tion detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Gabriellav2: Towards better generalization in surveillance videos for ac- tion detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.938524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.938524Z digest=sha256:deb0a1c8b749364eef778543d1059eae66e91106797592a0183f91b2866c994a

Observation 899f6d60-3d9a-423f-95ef-a1f52c97fe31 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.942737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.942737Z digest=sha256:37e31ea60b9c5e341f6715dd18413c108547372e49aef7f734fc419ed7f33be0

Observation 98a2ede4-d1b3-4ed4-89fc-1cc8556ffbad · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.947581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.947581Z digest=sha256:76c21afe5db7389581de8f2e53d6ba4784bf54ad79bef4d5a752e4f0d5ccc40e

Observation fbd0b121-fc8f-4d39-82ed-b4a83da6c90e · outbound

This paper cites Holistic interaction transformer network for action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Holistic interaction transformer network for action detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.953166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.953166Z digest=sha256:bb05d8ed2ec28b58be9c164100028104cdcbd21c5991b2322115f4d71f19b487

Observation bbbb3167-c1e4-4cb3-9bdd-c19cd89a3440 · outbound

This paper cites X3d: Expanding architectures for efficient video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection X3d: Expanding architectures for efficient video recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.958245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.958245Z digest=sha256:f80f11ff773efe3ae34cdaa68d6e6f594f118be30b8526e2bd4ab7324e2bb3ac

Observation e1d8fe21-be1c-4d83-bd36-40bd9b5e3e62 · outbound

This paper cites Slowfast networks for video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Slowfast networks for video recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.305254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.962864Z digest=sha256:94464bf38e2a133d587d0de4d2cf1dac7055ed7a75ef439dc7466f43e1dc0f2b

Observation 39c3202b-f348-4ef9-9528-cdfa635e7663 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-adapter: Better vision-language models with feature adapters

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.289572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.968625Z digest=sha256:8097e90c57cf3fecc620172f5c48b4ecc009da530f80ce7636a117636224e040

Observation 6c8f9fde-1643-4644-a992-a20489ee6010 · outbound

This paper cites Adamixer: A fast-converging query-based object detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Adamixer: A fast-converging query-based object detector

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.275566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.972844Z digest=sha256:c173357654d188b047ecff207a13da71d575029adb00a799317b1126e9757167

Observation 3647d439-e593-4e2e-9808-f6f413784297 · outbound

This paper cites Video action transformer network.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Video action transformer network

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.261882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.978719Z digest=sha256:e264269fb6fca74f45fb3d80301c11cf15629009d1c7301007950dde5edb90b2

Observation ada785ec-8665-41e8-a8f3-868448bcf185 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.245927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.984004Z digest=sha256:b139bc0331e554b101dd871a26fd19f549c95b97e801324ceeb412b1f5534548

Observation 0948440f-5a6f-41f8-b49b-dbc1d9aef6cd · outbound

This paper cites Mask r-cnn.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.233579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.989696Z digest=sha256:ee79c589709a276a3d96ee5f07d64d47d48c2c11be1d0ddf9263c47ede70dde7

Observation 8dcb99de-b9f4-4ade-afba-f62d7bb1d01e · outbound

This paper cites Mask r-cnn.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.219321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:33.997303Z digest=sha256:38d6fa5fdbb98fbef777d8a86df7371a2749b78b60a8878acd4188c96dcdc145

Observation b7ff74fc-5739-40ef-81bb-4a05f78aa929 · outbound

This paper cites Interaction-aware prompting for zero-shot spatio-temporal action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Interaction-aware prompting for zero-shot spatio-temporal action detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.204945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.002986Z digest=sha256:5b6dd9b1c2b431fa9c08ca6489ddd9732145ee7f9309be3c073fe633789169dc

Observation c0f167db-0d6f-4c01-b12b-11a15a5744f7 · outbound

This paper cites Towards understanding ac- tion recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards understanding ac- tion recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.191381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.008057Z digest=sha256:2391379376db3dfc023e0ed0751066996dc5688d2dec2bb4d172964209c2662b

Observation 48bfafee-2a01-41bd-a998-cac1e8fb3254 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.174641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.014997Z digest=sha256:cae9eb4d28401a1a808e11e0de97a1e4520f61f4af81e54ba20e3bdd3a200e1a

Observation 3f2e8739-d94a-49ee-8651-719104ac68a4 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.162136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.021273Z digest=sha256:d5e6a762a30c7564b47645c166881ad5cf3e2be311e59280386ca8fe920965b6

Observation 58b2af96-fcf7-41c8-b199-885dd85a5e02 · outbound

This paper cites Action tubelet detector for spatio- temporal action localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Action tubelet detector for spatio- temporal action localization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.147873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.026541Z digest=sha256:dbd4e0b537c35bf9728a3020e40185c722fde8ed3f9b0fd9e635d7d3817ce75a

Observation f4cf0649-040f-417d-80a7-279a2b3bd1cd · outbound

This paper cites Region- aware pretraining for open-vocabulary object detection with vision transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Region- aware pretraining for open-vocabulary object detection with vision transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.133135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.032374Z digest=sha256:6a48813f5afa53bbf9da5130431416b7348b275f3465dc31adde6112a13dbecc

Observation 13de2aab-4ba6-495f-9b9c-b9dac35bac06 · outbound

This paper cites You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.042842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.042842Z digest=sha256:3fd6011683bf140c51e7b03d3e3611be922958e26f5dde687262dd0a94f9659e

Observation 627737ca-7693-4dc1-aec4-a29fc32d4c5d · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.115041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.053125Z digest=sha256:6f093f00e83ef16426df94a57b03bf3edef0ff04c586a7a7f145c1b4c461644b

Observation a6de3308-f534-4333-adba-57ddb94b64a6 · outbound

This paper cites Multisports: A multi-person video dataset of spatio-temporally localized sports actions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Multisports: A multi-person video dataset of spatio-temporally localized sports actions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.099488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.057827Z digest=sha256:d686bacb3b66e1ba17e87e9448f75dc6fffca3d13bd4b5aaf395c6e106ef9f86

Observation 992132a9-a5a3-44d3-a65c-db8ca27af360 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring plain vision transformer backbones for object de- tection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.082400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.062023Z digest=sha256:0c1b579c1e51e9dfe7859ef18a3e1ff41882f325cbcd9a99fb4c95d37460eb23

Observation d6060094-1d26-4859-9fa5-e6c7837ca89c · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.072739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.072739Z digest=sha256:04526e3ea5f0903deaee0858b00989f6099235bd5f0542b423790cc3c0833955

Observation 6d5d97d2-4307-47ec-bd8c-6a10c6b79641 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.090090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.090090Z digest=sha256:b1f24594eec8fbf19ee43fcf7b1c99ac95ba9c40c6e6905d2391714c58e8cc3c

Observation 4d102cf0-b0a6-4361-bb6a-81b2b6873817 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary semantic segmentation with mask-adapted clip

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.064805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.095188Z digest=sha256:becaac0e35b42a428e47416ab87b9d56d1e404167d1ce5e8ce985b3658e746f6

Observation 64abda54-12b0-469d-9f22-61e2b7c3dd2a · outbound

This paper cites Learning object-language alignments for open-vocabulary object de- tection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning object-language alignments for open-vocabulary object de- tection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.050592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.106098Z digest=sha256:046d52f3b94a3c25a8b725781de57a19b7b733c2526f16ccd8809c55facdf779

Observation 782e6c26-1364-4b98-9f88-2eaa67143b93 · outbound

This paper cites Frozen clip models are efficient video learners.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Frozen clip models are efficient video learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.119730Z digest=sha256:ba04610ed73fd04287569b3f4dcad30f773b797584c3475a25b87b6baabfcbe0

Observation fdf89e7f-b712-4165-84cc-0524b34cb94b · outbound

This paper cites Revisiting temporal modeling for clip-based image-to-video knowledge transferring.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.125878Z digest=sha256:9ad3ad936b988378bbc7e3e3017b4e537c0475b0615882f2cc1a3393e4c85815

Observation 83cf222f-01f6-4fc0-b64d-7d5bc2a43bf3 · outbound

This paper cites Revisiting temporal modeling for clip-based image-to-video knowledge transferring.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.002043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.131528Z digest=sha256:1bf948c61716db08036f1d20620068c4086fd3d8f557c8a8c0fbb7cf57bb9358

Observation 5e553e47-f158-4342-b759-54babcbf0257 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.142444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.142444Z digest=sha256:cf757a1df519847c22ba460724717ee8f5912cab975190762a8ebe10d520ecef

Observation c08765bf-906b-45d3-8e29-d86035509d44 · outbound

This paper cites Decoupled weight decay regularization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Decoupled weight decay regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.147093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.147093Z digest=sha256:6e00456fa96b8bbb78ef5c65e5a6eb679b1739503f7cd363274222690ae280a4

Observation 79e3dd92-e866-4dcc-8f9f-485a8a455d83 · outbound

This paper cites Verbs in action: Improv- ing verb understanding in video-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Verbs in action: Improv- ing verb understanding in video-language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.967268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.152280Z digest=sha256:a3d299a779c98a0fbdcb58291c0edcd5ef815182a0420e0bcfb8ac0a3d69d135

Observation aa9df3ab-1c93-4e5c-af9e-c2f424e757b2 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.953181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.167226Z digest=sha256:94affde82666efb2060065429bfd68e618aac8eac3d581134442cf6c3ec507e9

Observation 0db54a28-09cc-447f-80f8-dedade732fd3 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.925219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.173557Z digest=sha256:db838ddec122f70985cce88ae1b6acebc5d44d1d4fc04ad2693f30f88e28c252

Observation 9944dc8b-70e2-4561-9129-f828df311ed2 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.905040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.179499Z digest=sha256:833b356af78fc457e8a193ccb474c90e612b8fc864832042e51a520e011c233b

Observation ebb2f9c1-454b-48a6-9593-05de32ece069 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.886681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.185592Z digest=sha256:0dfab10a625fab20ef0221750b364cb842074b372aa4cf483619ab0d6e4d8e27

Observation d4e87b73-854e-4aee-b8e7-adb107377104 · outbound

This paper cites GPT-4 Technical Report.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.199421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.199421Z digest=sha256:c418ca6227ae1ba31a2de3ee2ed28fc8c830382659ff1ae1ed4a9d3cbbe0d668

Observation 3da5c551-d16e-4ac1-ae8c-614c9c120058 · outbound

This paper cites Actor-context-actor relation net- work for spatio-temporal action localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-context-actor relation net- work for spatio-temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.870041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.203815Z digest=sha256:9c20199385460e9869ae227163923968d5d0c2afb5615bae5c23aa05b2ac6950

Observation 81dffbc5-8af5-4945-94d0-0e49a45e975f · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection St-adapter: Parameter-efficient image-to-video transfer learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.854822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.210893Z digest=sha256:3fe657eb8cc8984053437567a7311d72f48e8be1d9311f51653d50bfd69560f8

Observation b5874fa3-658e-4643-a83d-ff134bca9207 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.836146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.221147Z digest=sha256:1e8f317929e916655e1fd4448c8502d0316f6f840d9cfb9f1344d92ee31d4a1a

Observation 4d2e88c0-1bb0-42e1-9891-0f9add2444a0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.813548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.226553Z digest=sha256:302f7e7f78e8d85dc6666a01cfe22db490eb03135172cee2a21755479d99312d

Observation 2aae92bd-2bf9-4752-8b8f-0f0205d3967c · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Fine-tuned clip models are efficient video learners

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.796054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.235051Z digest=sha256:e5b21c8487a522706b5c02ca7d25119a2410700b7a48dea77c0120daaacd9922

Observation e253c8d0-066f-41b0-9685-f289a970ef83 · outbound

This paper cites Open-vocabulary temporal action detection with off-the-shelf image-text features.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary temporal action detection with off-the-shelf image-text features

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.780311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.248648Z digest=sha256:d80c97d104a931fa809b88df9e241914617f94bb622119b142f4ab531f5f3fc1

Observation 083eb89f-447a-48a8-8054-c1df10550bad · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.253502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.253502Z digest=sha256:7098c2b9fc35666b241113bba66adf71a2b251c5ff377549b3717af3c038ca02

Observation 724675cb-153f-4780-b4b2-a6798463402a · outbound

This paper cites Generalized in- tersection over union: A metric and a loss for bounding box regression.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generalized in- tersection over union: A metric and a loss for bounding box regression

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.257602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.257602Z digest=sha256:0be63b74f947feb13aace84c602a9df1055aa3def41badb159a9d5616563d481

Observation 59106dfc-b53e-4ec7-88b5-be3eeeb0d5b6 · outbound

This paper cites Sparse DETR: Efficient end-to-end object detec- tion with learnable sparsity.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse DETR: Efficient end-to-end object detec- tion with learnable sparsity

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.743383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.261894Z digest=sha256:32597b7b7e48f827f6d9a321aa7aa3de33014ec48c20cdc33a13da508df73183

Observation f93e6906-a7cc-4ba2-9da8-4af4973beca9 · outbound

This paper cites Hi- era: A hierarchical vision transformer without the bells-and- whistles.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Hi- era: A hierarchical vision transformer without the bells-and- whistles

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.726724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.267147Z digest=sha256:d8c8f77ad3beecc95e813d5e1a25aa4a8466738f4f3a8a569b760c1d3a642b68

Observation 91d9c389-3853-478b-9aa6-f4b0d78c39f5 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.703196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.271341Z digest=sha256:b719ceada3d0f9a28a189f765bd9a7b87cc94921fa92b6f9c6cbcf8b8e2f1cf3

Observation 5bdd13fe-1598-41bc-b9a9-e7c8978c548c · outbound

This paper cites Road: The road event awareness dataset for autonomous driving.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Road: The road event awareness dataset for autonomous driving

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.688783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.275944Z digest=sha256:ab2c2606c22e9f8fc8ad3a4d14c992db94664861e443fc297af1062cd740f1cd

Observation c60eda70-48eb-4399-bea5-d0e95325d959 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ucf101: A dataset of 101 human actions classes from videos in the wild

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.670972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.280753Z digest=sha256:df99859181f4b4aa1142814d29f30686498f65edd311959b87e18ae484850c7e

Observation 9e8d9b78-66a6-4176-b7c2-12f4eb1c56f1 · outbound

This paper cites Actor-centric relation network.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-centric relation network

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.654926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.286480Z digest=sha256:ba22d4981ca65fb37748c88806b2a6f76edf5569665146ab54e0833890345c35

Observation e6691d63-bfe3-4569-b800-587c8c19d26e · outbound

This paper cites Relational action forecasting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Relational action forecasting

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.620142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.292836Z digest=sha256:09435f82d9bdd47bb7f5b9170dfbee25a837370b40104c4a5ac49860861fd82c

Observation 2d2becb4-77a4-4488-9e45-fb5cf60f330a · outbound

This paper cites Sparse r-cnn: End-to-end object detec- tion with learnable proposals.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse r-cnn: End-to-end object detec- tion with learnable proposals

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.596591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.300260Z digest=sha256:278982c779c6579309282d31e355bc78e79c35cfce4696e40bc5d1260185f757

Observation f3b7da0b-fc8a-4876-8656-6f785bd2d0ec · outbound

This paper cites Asynchronous interaction aggregation for action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Asynchronous interaction aggregation for action detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.576476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.308022Z digest=sha256:d3f538e5990fd7f6da742add8a060a8d8e988179458f1d7c1ef1beb26a1670c1

Observation 9b3ac7e4-8a5e-4709-a9a6-a0f99214bebb · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mlp-mixer: An all-mlp architecture for vision

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.558998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.317360Z digest=sha256:2be6f2ee388a5e4e0772c86e7edfa64fcbf85bb4c2e88f78fc99b01faf38985d

Observation d4ded328-8ca6-40d1-8e5d-ea4b27506309 · outbound

This paper cites Attention is all you need.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Attention is all you need

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.540707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.321857Z digest=sha256:45053f31a4bafd464d20f621d89073cf3f8a76a87ed03304e7eed3127b58d5e6

Observation 2862efa0-b746-4ef8-8848-e4580b94e97a · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection ActionCLIP: A New Paradigm for Video Action Recognition

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.333164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.333164Z digest=sha256:98bb6eadd28c0c60a5ea19e7da4f6a55b32b09d16de19cacae0994f1faa5b253

Observation cef25c1c-96cf-4f48-a3cf-a785f982f251 · outbound

This paper cites Vita-clip: Video and text adaptive clip via multimodal prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vita-clip: Video and text adaptive clip via multimodal prompting

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.501357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.337888Z digest=sha256:8bebdb146be248d7632bf745c9b2836104712389320734bb44b3b0f3990f9854

Observation 49b02c60-0397-4add-bb39-60f46e93c4b0 · outbound

This paper cites Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.477256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.342182Z digest=sha256:604208fe43a47e6d2fbaddb67a7bff193d7f8bc84bfe6998f65336d307eba855

Observation 22930374-aaaa-4cb5-bc82-56616bdefd0b · outbound

This paper cites Long-term feature banks for detailed video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Long-term feature banks for detailed video understanding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.460235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.346537Z digest=sha256:cec05e322cc0fef40e87c786d068a281d6dbc7eda104b17330a843382467c9f2

Observation 81150c78-644c-49be-82a4-af4bbc0f6241 · outbound

This paper cites Context-aware rcnn: A baseline for ac- tion detection in videos.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Context-aware rcnn: A baseline for ac- tion detection in videos

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.442768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.351318Z digest=sha256:740199f5adce257335afdceac944572999b3e0065d186cd0189f69b47d32dff8

Observation fa01248c-5077-4de8-8e0e-ca3bf8273f72 · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards Open Vocabulary Learning: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.355336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.355336Z digest=sha256:3a93f3de05d6039f5bb3c99fd2e9f854742daa489c2e7a891e4c94ad3251b8d9

Observation 512e006d-b5fa-4039-8c0f-5ce3411677e8 · outbound

This paper cites Stmixer: A one-stage sparse action detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.425422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.360234Z digest=sha256:8bf432ea1c513c3cb75199f4eb32e264a8a54cb4b3b94f5a6bdca7072b21c3a6

Observation 6ed96869-fb21-4314-ad30-5ee6828fb25e · outbound

This paper cites Stmixer: A one-stage sparse action detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.405882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.364807Z digest=sha256:dcd3f27e05a2e66f585a2f834ad73898be8877ef543af9f230e0cf3e6e30bed7

Observation 6228390b-7bcb-49f7-b9e7-eb4aeaf4a8f0 · outbound

This paper cites Open-Vocabulary Spatio-Temporal Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-Vocabulary Spatio-Temporal Action Detection

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.369317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.369317Z digest=sha256:5fe435522796d3e3aaa62adaa6a9c080533601835a508cb26fae32bc3ac3c662

Observation a450ad62-ef08-4428-a427-86f9125e5f3a · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.390028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.374056Z digest=sha256:ff199b581d49ea0634f5a94e05a06a83b4fd715fa8d177c6c4f03870656e47f8

Observation 4a15141a-cc08-4019-9d1f-55152793a02b · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.375890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.378372Z digest=sha256:0439906805649bb95d169f7038f1833d5dace733860ba6b1a74b9cae56aa0f11

Observation 5576a9a8-ab22-4929-9de9-141a61b7f618 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.359394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.382792Z digest=sha256:729df486ab4c1e87db293e3ef9f4664a0c6965e9cfee339bbd93934b144c0649

Observation c1dde812-30d6-474c-a4db-8f8112fc3d92 · outbound

This paper cites Unloc: A unified framework for video localization tasks.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Unloc: A unified framework for video localization tasks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.342125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.389183Z digest=sha256:bdd9d05c13d2a0de6033f4bb5e1da4c09645c5a6358eb0480173d532c3b58c81

Observation 6a60e58f-1d3a-480f-ad50-8620f40144e3 · outbound

This paper cites Detecting human actions in surveillance videos.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Detecting human actions in surveillance videos

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.323223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.398600Z digest=sha256:e12d85f0b92676a2f12da84fae271c18d590b399d1988e54b4e597340b3e5016

Observation ee6e3520-08f2-4373-9558-c55e4b34417c · outbound

This paper cites Contextualized spatio-temporal contrastive learning with self-supervision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Contextualized spatio-temporal contrastive learning with self-supervision

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.286810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.404552Z digest=sha256:dd31addc8affbbdb3e025fd5256e7b375c7709ba13527904dade144fc32778df

Observation 1a35debf-220f-4151-9dcf-e3d2c372e435 · outbound

This paper cites Vision-based garbage dumping action detection for real-world surveillance platform.ETRI Journal, 41(4):494–505, 2019.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vision-based garbage dumping action detection for real-world surveillance platform.ETRI Journal, 41(4):494–505, 2019

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.259768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.412125Z digest=sha256:1ee756166d49ef2c6c2665941304f949e5ec41baee1a48aaafbb99aa05da25fa

Observation ba00c246-766c-4df0-9b77-2e412d624b1c · outbound

This paper cites Open-vocabulary detr with conditional matching.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary detr with conditional matching

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.227399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.418542Z digest=sha256:6e2cd1aa8128c57cbe39ba83f6a61b9c8d870597dead909b451cc317917737b4

Observation 778c0bbe-b2f3-4369-b6e5-55093f7d3d82 · outbound

This paper cites Open-vocabulary object detection using captions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.195512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.428209Z digest=sha256:9388e3acc89910b84382699a206e8eaa7f762dbb729949a05e8c5a1be06d64a2

Observation c3229a69-53c5-4c6a-b995-cd6e3a770985 · outbound

This paper cites Open-vocabulary object detection using captions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.168723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.432539Z digest=sha256:5f7a36d6ec35799fb6b1633becb2820fe71f67e93458308a311b3e30f7506702

Observation f94dff5c-4afc-40ef-b6f8-8bf85f6e1a47 · outbound

This paper cites From recognition to cognition: Visual commonsense reason- ing.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection From recognition to cognition: Visual commonsense reason- ing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.139140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.443045Z digest=sha256:387fe519b2224d4d3ab940ff6274ee022fee017a136186181d72d06bba3f5a6c

Observation c2e45b16-72d3-4443-9478-a1e903d3ec23 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vinvl: Revisiting visual representations in vision-language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.089742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.449259Z digest=sha256:e1689cd52f9f2a509f782717ac0cf68ada47a72d70fc1a09abe8dcc1f66c9f38

Observation 2316d72b-c42a-4441-a29f-faeed5df7917 · outbound

This paper cites Tuber: Tubelet transformer for video action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Tuber: Tubelet transformer for video action detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.065261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.453982Z digest=sha256:13168e240b8ef4b53b4fc8b6cb0ef9377a4ded1183dfe4a1cba225f9e8263ee0

Observation fb604aaf-58c5-40a6-8b26-0a72517c0866 · outbound

This paper cites MRSN: Multi-Relation Support Network for Video Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection MRSN: Multi-Relation Support Network for Video Action Detection

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:14:34.558149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.460035Z digest=sha256:89753577c61bbf587010dda3904ef8a49eb9e2887d9851b168d3f5df5e06f0ea

Observation bf65d351-a680-4c7c-a476-f51acb3f9e45 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Regionclip: Region-based language-image pretraining

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.040683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.472500Z digest=sha256:9658e1945ae52a74123c13c586ca12b31d5fe30d560e051b504109d8e27da359

Observation 8f0230f2-e4ce-4cdc-98e5-cc72b0593fee · outbound

This paper cites Learning deep features for discrimi- native localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning deep features for discrimi- native localization

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.005673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.477338Z digest=sha256:30f751d742e92b17165bb4da945a3309246f2126f09497aa5eccf13eb72f3d48

Observation 2a8ecf9d-e5b1-4f1f-9cea-074b3d180b78 · outbound

This paper cites Learning to prompt for vision-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for vision-language models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.482606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.482606Z digest=sha256:fffeb3f4dffe9659e59ab2cda5cc49e30b639f03f3e3ed86d68863381795e447

Observation 7d972e5f-770b-48e3-93a5-930a182a96a3 · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot se- mantic segmentation.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.487802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.487802Z digest=sha256:1b06f5e8cd3241b33893f2fc82a5df8a3ae92d1313a40b1bab02a54a3e2bd2ea

Observation 6347dfc0-d29f-4afe-915d-746364a3c506 · outbound

This paper cites For the action type {CLS}, what are the visual descriptions? Please respond with a list of 16 short sentences.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection For the action type {CLS}, what are the visual descriptions? Please respond with a list of 16 short sentences

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:34.950581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:14:34.495197Z digest=sha256:4854314a4da893c3de5ff30a52d4fbf95eaf5d35782551dedc9b36ea06a3582c

Pith citing papers

No inbound Pith citation observations are available.