Pith. sign in

Paper Citation Record · LEDGER

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection

As of 20 August 2026, this Paper Citation Record lists 91 of 91 outbound references and 0 inbound Pith citation observations for arXiv:2411.10922.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10922 v1

Coverage vector

measured 91 of 91 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:34.495197Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

91 of 91 outbound references displayed

  • verified exact2
  • verified fuzzy65
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db3981e3-74ce-482e-bbb4-dd3ac75ca216 · outbound

This paper cites Bridg- ing the gap between object and image-level representations for open-vocabulary detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Bridg- ing the gap between object and image-level representations for open-vocabulary detection

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.897362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.897362Z digest=sha256:1ac70d661e18a6bf7aac15dc28567d74c310ef1fc2fd1433d2eb8f7f30e0caf8

Observation 17fc7182-20d2-41c2-95e3-017c4828fb4d · outbound

This paper cites End- to-end object detection with transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection End- to-end object detection with transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.902105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.902105Z digest=sha256:0f4e6c49aa276db36df7958c038a2cbe504d9e3dbec0fd5fcd41a40a589dc50d

Observation b03136e5-d99e-44cd-bf34-03704abbffe5 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Emerg- ing properties in self-supervised vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.908058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.908058Z digest=sha256:92f596d3fd0d79ace0b6422e1702ebb8c169ce71685b39032ca9e885bfd42266

Observation b891a697-38a8-4291-87b8-3a0063cc9ca9 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.913872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.913872Z digest=sha256:f1e9d17c653a18964630ac18c32cd4566381b7cd34e47e749f8e584764743187

Observation eaa2407a-400c-4507-8cbb-083426c58af3 · outbound

This paper cites CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection CycleACR: Cycle Modeling of Actor-Context Relations for Video Action Detection

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:14:34.741882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.920517Z digest=sha256:ece3866324045af1b368df2688c1c1f1f0cac368654bbc3853dcaa327fce959a

Observation e5f1f7ba-2173-4c7a-a937-a29ce4570943 · outbound

This paper cites Efficient video action detection with token dropout and context refinement.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Efficient video action detection with token dropout and context refinement

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.926112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.926112Z digest=sha256:53eb6244e35d62ea20a51fc313cb96804e69fc79fd05b0b22d3e33fad475e488

Observation 751bcd4b-75e4-4952-9bb9-4b5230a32a7a · outbound

This paper cites Watch only once: An end-to-end video action detection framework.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Watch only once: An end-to-end video action detection framework

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.931019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.931019Z digest=sha256:cce2fe3cba142c4639a86f7dd075b24ebead9ea598b4715978faf936acaace16

Observation e35feb40-0671-4a36-b17c-bee968bc33d5 · outbound

This paper cites Gabriellav2: Towards better generalization in surveillance videos for ac- tion detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Gabriellav2: Towards better generalization in surveillance videos for ac- tion detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.938524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.938524Z digest=sha256:deb0a1c8b749364eef778543d1059eae66e91106797592a0183f91b2866c994a

Observation 899f6d60-3d9a-423f-95ef-a1f52c97fe31 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.942737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.942737Z digest=sha256:37e31ea60b9c5e341f6715dd18413c108547372e49aef7f734fc419ed7f33be0

Observation 98a2ede4-d1b3-4ed4-89fc-1cc8556ffbad · outbound

This paper cites Learning to prompt for open-vocabulary ob- ject detection with vision-language model.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for open-vocabulary ob- ject detection with vision-language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.947581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.947581Z digest=sha256:76c21afe5db7389581de8f2e53d6ba4784bf54ad79bef4d5a752e4f0d5ccc40e

Observation fbd0b121-fc8f-4d39-82ed-b4a83da6c90e · outbound

This paper cites Holistic interaction transformer network for action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Holistic interaction transformer network for action detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.953166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.953166Z digest=sha256:bb05d8ed2ec28b58be9c164100028104cdcbd21c5991b2322115f4d71f19b487

Observation bbbb3167-c1e4-4cb3-9bdd-c19cd89a3440 · outbound

This paper cites X3d: Expanding architectures for efficient video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection X3d: Expanding architectures for efficient video recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:33.958245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:33.958245Z digest=sha256:f80f11ff773efe3ae34cdaa68d6e6f594f118be30b8526e2bd4ab7324e2bb3ac

Observation e1d8fe21-be1c-4d83-bd36-40bd9b5e3e62 · outbound

This paper cites Slowfast networks for video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Slowfast networks for video recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.305254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.962864Z digest=sha256:a5ea9e1746f7f06b7f500076a11c4e0e65524d9d00722fc56047e2329ae05e85

Observation 39c3202b-f348-4ef9-9528-cdfa635e7663 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-adapter: Better vision-language models with feature adapters

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.289572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.968625Z digest=sha256:9f818b1d7e6a276b30229c3b24c0f8b992c867d9df0ba2d88dd4b2a23a5900a7

Observation 6c8f9fde-1643-4644-a992-a20489ee6010 · outbound

This paper cites Adamixer: A fast-converging query-based object detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Adamixer: A fast-converging query-based object detector

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.275566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.972844Z digest=sha256:d729051bbbe9fa3b601700593516703b295f057aab91d82e872235872332f989

Observation 3647d439-e593-4e2e-9808-f6f413784297 · outbound

This paper cites Video action transformer network.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Video action transformer network

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.261882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.978719Z digest=sha256:7462443c902b33c7deb2e2409a74bddd8e56e4b803143096f1cc12aaa648e6bd

Observation ada785ec-8665-41e8-a8f3-868448bcf185 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.245927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.984004Z digest=sha256:f800b842691ed19e5036ef548df77dd0cd7a005c584c73cbee2ead37d40125cf

Observation 0948440f-5a6f-41f8-b49b-dbc1d9aef6cd · outbound

This paper cites Mask r-cnn.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.233579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.989696Z digest=sha256:271133af0fe026e5893c51540afc3a5ce9ba69f056fc908fc2edaa689a7c23ba

Observation 8dcb99de-b9f4-4ade-afba-f62d7bb1d01e · outbound

This paper cites Mask r-cnn.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mask r-cnn

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.219321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:33.997303Z digest=sha256:ab2d8c03e26f25fb7a37c8f2d671203ca02fb2dd9379be8e6eaa52c0beb4151a

Observation b7ff74fc-5739-40ef-81bb-4a05f78aa929 · outbound

This paper cites Interaction-aware prompting for zero-shot spatio-temporal action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Interaction-aware prompting for zero-shot spatio-temporal action detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.204945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.002986Z digest=sha256:659adbce59dba6377cce47845a5521e37f24595492f032d54790564371bc6589

Observation c0f167db-0d6f-4c01-b12b-11a15a5744f7 · outbound

This paper cites Towards understanding ac- tion recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards understanding ac- tion recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.191381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.008057Z digest=sha256:2f0b9901dae3f43caddbfd3a9489885b04e59171bb7eab020c897d4e7ed958bc

Observation 48bfafee-2a01-41bd-a998-cac1e8fb3254 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.174641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.014997Z digest=sha256:331adb158d6422f4f571ed7386b37c15e71fbf6ce305e1e2181e660758e000b1

Observation 3f2e8739-d94a-49ee-8651-719104ac68a4 · outbound

This paper cites Prompting visual-language models for efficient video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Prompting visual-language models for efficient video understanding

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.162136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.021273Z digest=sha256:40a4f7668addaf792e86a75cbaac1494bd668c1eeacf3534d7154ee61363cddf

Observation 58b2af96-fcf7-41c8-b199-885dd85a5e02 · outbound

This paper cites Action tubelet detector for spatio- temporal action localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Action tubelet detector for spatio- temporal action localization

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.147873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.026541Z digest=sha256:588e7442eb3ccbf6c722890e3aabd91d746ff23788332c6d6a796b6e9cf41c8d

Observation f4cf0649-040f-417d-80a7-279a2b3bd1cd · outbound

This paper cites Region- aware pretraining for open-vocabulary object detection with vision transformers.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Region- aware pretraining for open-vocabulary object detection with vision transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.133135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.032374Z digest=sha256:32d4666b84cc797a6c9627dd5d6b5a0e0e2dd01eea679c45c562bf12bd14d91c

Observation 13de2aab-4ba6-495f-9b9c-b9dac35bac06 · outbound

This paper cites You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.042842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.042842Z digest=sha256:3fd6011683bf140c51e7b03d3e3611be922958e26f5dde687262dd0a94f9659e

Observation 627737ca-7693-4dc1-aec4-a29fc32d4c5d · outbound

This paper cites F-vlm: Open-vocabulary object detection upon frozen vision and language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection F-vlm: Open-vocabulary object detection upon frozen vision and language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.115041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.053125Z digest=sha256:2be82906ef63eb71cee16d9ce00b48d8188145044be7379df083a5f2833a3da3

Observation a6de3308-f534-4333-adba-57ddb94b64a6 · outbound

This paper cites Multisports: A multi-person video dataset of spatio-temporally localized sports actions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Multisports: A multi-person video dataset of spatio-temporally localized sports actions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.099488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.057827Z digest=sha256:8eed1669457e11d4282661047f9ac858c9dd92d7567bc57379aa9da12a636e16

Observation 992132a9-a5a3-44d3-a65c-db8ca27af360 · outbound

This paper cites Exploring plain vision transformer backbones for object de- tection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring plain vision transformer backbones for object de- tection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.082400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.062023Z digest=sha256:d9082ec808c6275b5143dee5d6156efe4e010083b2ab014bc87ce46db08e3480

Observation d6060094-1d26-4859-9fa5-e6c7837ca89c · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.072739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.072739Z digest=sha256:04526e3ea5f0903deaee0858b00989f6099235bd5f0542b423790cc3c0833955

Observation 6d5d97d2-4307-47ec-bd8c-6a10c6b79641 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.090090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.090090Z digest=sha256:b1f24594eec8fbf19ee43fcf7b1c99ac95ba9c40c6e6905d2391714c58e8cc3c

Observation 4d102cf0-b0a6-4361-bb6a-81b2b6873817 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary semantic segmentation with mask-adapted clip

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.064805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.095188Z digest=sha256:07de1ef5c76e1c591aefe2963488233bf1f1197788777150396695cbc0facb33

Observation 64abda54-12b0-469d-9f22-61e2b7c3dd2a · outbound

This paper cites Learning object-language alignments for open-vocabulary object de- tection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning object-language alignments for open-vocabulary object de- tection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.050592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.106098Z digest=sha256:2c7137871c4b8268f844bc6076b2c702aad6b4f570582ae2cf69bf0f712de5a9

Observation 782e6c26-1364-4b98-9f88-2eaa67143b93 · outbound

This paper cites Frozen clip models are efficient video learners.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Frozen clip models are efficient video learners

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.036110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.119730Z digest=sha256:ae08f3c93d76331acb19620b29d49fe093db9262c5972fc486a27ccb398b465e

Observation fdf89e7f-b712-4165-84cc-0524b34cb94b · outbound

This paper cites Revisiting temporal modeling for clip-based image-to-video knowledge transferring.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.020826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.125878Z digest=sha256:6d6620fbfb8c0ffec912a297a0120ecf07bb9d514a4998af8bdc3d538352fd7d

Observation 83cf222f-01f6-4fc0-b64d-7d5bc2a43bf3 · outbound

This paper cites Revisiting temporal modeling for clip-based image-to-video knowledge transferring.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Revisiting temporal modeling for clip-based image-to-video knowledge transferring

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:36.002043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.131528Z digest=sha256:2b3b4c2539132d8fb9b66a52e41c877a2f70f1476aa03bf8410b76e397022733

Observation 5e553e47-f158-4342-b759-54babcbf0257 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.142444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.142444Z digest=sha256:cf757a1df519847c22ba460724717ee8f5912cab975190762a8ebe10d520ecef

Observation c08765bf-906b-45d3-8e29-d86035509d44 · outbound

This paper cites Decoupled weight decay regularization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Decoupled weight decay regularization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.147093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.147093Z digest=sha256:6e00456fa96b8bbb78ef5c65e5a6eb679b1739503f7cd363274222690ae280a4

Observation 79e3dd92-e866-4dcc-8f9f-485a8a455d83 · outbound

This paper cites Verbs in action: Improv- ing verb understanding in video-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Verbs in action: Improv- ing verb understanding in video-language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.967268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.152280Z digest=sha256:d996f2466d2521bc4f49c4c802e655d6fbab9634a3584c77e29681fbfca02c76

Observation aa9df3ab-1c93-4e5c-af9e-c2f424e757b2 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.953181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.167226Z digest=sha256:b99eda8353f7b75b91c0dfb1a99c4f3440dece0af8acc77dd94d4c76b41c2e06

Observation 0db54a28-09cc-447f-80f8-dedade732fd3 · outbound

This paper cites Zero-shot temporal action detection via vision-language prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zero-shot temporal action detection via vision-language prompting

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.925219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.173557Z digest=sha256:86590de8d8e9d2ae28fa79395c65d29d6415cc00269e93db2ab6e6af6b0f3254

Observation 9944dc8b-70e2-4561-9129-f828df311ed2 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.905040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.179499Z digest=sha256:daf347fa624c3b6602955a6ba13a5ac022d2e2aada20fafb5fc88b5fc5eb968d

Observation ebb2f9c1-454b-48a6-9593-05de32ece069 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Expanding language-image pretrained models for gen- eral video recognition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.886681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.185592Z digest=sha256:d4306d498e2066e614a9f308e2755c8b4f4a685efbd652ef96867dc486d77121

Observation d4e87b73-854e-4aee-b8e7-adb107377104 · outbound

This paper cites GPT-4 Technical Report.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.199421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.199421Z digest=sha256:c418ca6227ae1ba31a2de3ee2ed28fc8c830382659ff1ae1ed4a9d3cbbe0d668

Observation 3da5c551-d16e-4ac1-ae8c-614c9c120058 · outbound

This paper cites Actor-context-actor relation net- work for spatio-temporal action localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-context-actor relation net- work for spatio-temporal action localization

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.870041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.203815Z digest=sha256:fd61c09b268df464903710f98beda3faf8156591e253a6f12142b2d45536ea2f

Observation 81dffbc5-8af5-4945-94d0-0e49a45e975f · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection St-adapter: Parameter-efficient image-to-video transfer learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.854822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.210893Z digest=sha256:24ef1d0d9964b7bf68cc3ef476c0130a9c75cad4898e895bd8fb1fdf0d650cfc

Observation b5874fa3-658e-4643-a83d-ff134bca9207 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.836146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.221147Z digest=sha256:4200eb0b2f6f882a0e8ee09874a14a4c479da2e8002ca6cf0800a6a0ad3d2f03

Observation 4d2e88c0-1bb0-42e1-9891-0f9add2444a0 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learn- ing transferable visual models from natural language super- vision

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.813548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.226553Z digest=sha256:fdcb0ec066a5b716a7615b86f262c9bd6cd2e3669c18ac652a64496fbc1f215a

Observation 2aae92bd-2bf9-4752-8b8f-0f0205d3967c · outbound

This paper cites Fine-tuned clip models are efficient video learners.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Fine-tuned clip models are efficient video learners

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.796054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.235051Z digest=sha256:6dc6cc902f122cd01126f2ed2da592a722c2e6d4504eeb28d57219f0c1599eb7

Observation e253c8d0-066f-41b0-9685-f289a970ef83 · outbound

This paper cites Open-vocabulary temporal action detection with off-the-shelf image-text features.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary temporal action detection with off-the-shelf image-text features

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.780311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.248648Z digest=sha256:d463c54ee278c18b712ca9fd07867b1d3825e40e90987f8aaaa0b4b6971de1f9

Observation 083eb89f-447a-48a8-8054-c1df10550bad · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.253502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.253502Z digest=sha256:7098c2b9fc35666b241113bba66adf71a2b251c5ff377549b3717af3c038ca02

Observation 724675cb-153f-4780-b4b2-a6798463402a · outbound

This paper cites Generalized in- tersection over union: A metric and a loss for bounding box regression.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Generalized in- tersection over union: A metric and a loss for bounding box regression

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.257602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.257602Z digest=sha256:0be63b74f947feb13aace84c602a9df1055aa3def41badb159a9d5616563d481

Observation 59106dfc-b53e-4ec7-88b5-be3eeeb0d5b6 · outbound

This paper cites Sparse DETR: Efficient end-to-end object detec- tion with learnable sparsity.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse DETR: Efficient end-to-end object detec- tion with learnable sparsity

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.743383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.261894Z digest=sha256:09c10de63b2dea61930feee9ec176c79b01a72a2143bdfdd4ed7849b9210e4b6

Observation f93e6906-a7cc-4ba2-9da8-4af4973beca9 · outbound

This paper cites Hi- era: A hierarchical vision transformer without the bells-and- whistles.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Hi- era: A hierarchical vision transformer without the bells-and- whistles

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.726724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.267147Z digest=sha256:fa24731ee11af9d816478a28aa17c97fcd29162a65884b6480092a83eebd5058

Observation 91d9c389-3853-478b-9aa6-f4b0d78c39f5 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.703196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.271341Z digest=sha256:52889556561d72fe88b6417515329785cfc8cbe37a480c3cc61c3ffb9777ac3c

Observation 5bdd13fe-1598-41bc-b9a9-e7c8978c548c · outbound

This paper cites Road: The road event awareness dataset for autonomous driving.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Road: The road event awareness dataset for autonomous driving

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.688783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.275944Z digest=sha256:490bfab5ee2e6aeb1726f17a7b5f27fb747f83171808d70823fb0741ea897a73

Observation c60eda70-48eb-4399-bea5-d0e95325d959 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Ucf101: A dataset of 101 human actions classes from videos in the wild

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.670972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.280753Z digest=sha256:751017069d8c5ee38042a6a4a36ce22747032a4e773c1841d98c7626b139dd9f

Observation 9e8d9b78-66a6-4176-b7c2-12f4eb1c56f1 · outbound

This paper cites Actor-centric relation network.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Actor-centric relation network

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.654926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.286480Z digest=sha256:9cb0d6ea8882cabc7ea206051b18b8f09491a11fa1ffe85cb159b318b800cf6b

Observation e6691d63-bfe3-4569-b800-587c8c19d26e · outbound

This paper cites Relational action forecasting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Relational action forecasting

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.620142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.292836Z digest=sha256:cc3922f3e67ef378fe53e0bbcc21688b9fd22278ea27ff40da88825b6a46122a

Observation 2d2becb4-77a4-4488-9e45-fb5cf60f330a · outbound

This paper cites Sparse r-cnn: End-to-end object detec- tion with learnable proposals.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Sparse r-cnn: End-to-end object detec- tion with learnable proposals

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.596591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.300260Z digest=sha256:f16a760791512f46560d1f1b297fd1a1271d6f0576b28089a75fe56fff9c5f05

Observation f3b7da0b-fc8a-4876-8656-6f785bd2d0ec · outbound

This paper cites Asynchronous interaction aggregation for action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Asynchronous interaction aggregation for action detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.576476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.308022Z digest=sha256:3a759affd8b6899279e408f0f55c1fa2420059baea2b68315ec4a3062c984a93

Observation 9b3ac7e4-8a5e-4709-a9a6-a0f99214bebb · outbound

This paper cites Mlp-mixer: An all-mlp architecture for vision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Mlp-mixer: An all-mlp architecture for vision

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.558998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.317360Z digest=sha256:58540620026de4183afa3bc9973e51398159f1c0dbd0542d60b1e44f78edef0b

Observation d4ded328-8ca6-40d1-8e5d-ea4b27506309 · outbound

This paper cites Attention is all you need.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Attention is all you need

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.540707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.321857Z digest=sha256:fde7cf244abc954e4f4e77374c350cd7cd028d6a6951b58394ff5a560e31f66d

Observation 2862efa0-b746-4ef8-8848-e4580b94e97a · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection ActionCLIP: A New Paradigm for Video Action Recognition

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.333164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.333164Z digest=sha256:98bb6eadd28c0c60a5ea19e7da4f6a55b32b09d16de19cacae0994f1faa5b253

Observation cef25c1c-96cf-4f48-a3cf-a785f982f251 · outbound

This paper cites Vita-clip: Video and text adaptive clip via multimodal prompting.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vita-clip: Video and text adaptive clip via multimodal prompting

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.501357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.337888Z digest=sha256:ec207d5ac8d7f2252616a1d51ff4ee9020133f5d0bff205bbe53ec84606b854e

Observation 49b02c60-0397-4add-bb39-60f46e93c4b0 · outbound

This paper cites Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vclip: Transforming clip to an open-vocabulary video model via interpolated weight optimization

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.477256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.342182Z digest=sha256:9f532d14a3d56e8bb3d12d7f78def45b8a57c5728fe7696f14886b37299c8ae0

Observation 22930374-aaaa-4cb5-bc82-56616bdefd0b · outbound

This paper cites Long-term feature banks for detailed video understanding.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Long-term feature banks for detailed video understanding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.460235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.346537Z digest=sha256:dd5375476672f4b57a6b63cf5b68ddacfa0bea3e6ee97e57a47a558477babcd8

Observation 81150c78-644c-49be-82a4-af4bbc0f6241 · outbound

This paper cites Context-aware rcnn: A baseline for ac- tion detection in videos.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Context-aware rcnn: A baseline for ac- tion detection in videos

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.442768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.351318Z digest=sha256:54617918665009f0d6bd3aadff33f7a95ff0e60412e340ee64ddc36155bae8ec

Observation fa01248c-5077-4de8-8e0e-ca3bf8273f72 · outbound

This paper cites Towards Open Vocabulary Learning: A Survey.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Towards Open Vocabulary Learning: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.355336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.355336Z digest=sha256:3a93f3de05d6039f5bb3c99fd2e9f854742daa489c2e7a891e4c94ad3251b8d9

Observation 512e006d-b5fa-4039-8c0f-5ce3411677e8 · outbound

This paper cites Stmixer: A one-stage sparse action detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.425422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.360234Z digest=sha256:e62dd4a107e6cf5598bb48050528713f33a6f102f08dc293dc0f0c624d500f8b

Observation 6ed96869-fb21-4314-ad30-5ee6828fb25e · outbound

This paper cites Stmixer: A one-stage sparse action detector.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Stmixer: A one-stage sparse action detector

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.405882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.364807Z digest=sha256:a3b56461738cc407277c0ab0781446e44442064ede44698ef16417f8b51aa264

Observation 6228390b-7bcb-49f7-b9e7-eb4aeaf4a8f0 · outbound

This paper cites Open-Vocabulary Spatio-Temporal Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-Vocabulary Spatio-Temporal Action Detection

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.369317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.369317Z digest=sha256:5fe435522796d3e3aaa62adaa6a9c080533601835a508cb26fae32bc3ac3c662

Observation a450ad62-ef08-4428-a427-86f9125e5f3a · outbound

This paper cites Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.390028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.374056Z digest=sha256:3545ed276abfcc07bb186ecaab61c95b6569d3626c49d033a692d2a24d835d82

Observation 4a15141a-cc08-4019-9d1f-55152793a02b · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.375890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.378372Z digest=sha256:3b791eb577b75aacd11e5313d7d72c81c04d56f0802f9bf6d2ee53c62fdfcbf9

Observation 5576a9a8-ab22-4929-9de9-141a61b7f618 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language representation alignment.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Clip-vip: Adapting pre- trained image-text model to video-language representation alignment

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.359394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.382792Z digest=sha256:503a164a3b8425eeb2153c4ceeb3aaae359270c839b91660fa79e21aefcb9c05

Observation c1dde812-30d6-474c-a4db-8f8112fc3d92 · outbound

This paper cites Unloc: A unified framework for video localization tasks.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Unloc: A unified framework for video localization tasks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.342125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.389183Z digest=sha256:2ff578774d76f340b9dcf7dfb483ad7897bdd1735850fe53ed242c5297ad976a

Observation 6a60e58f-1d3a-480f-ad50-8620f40144e3 · outbound

This paper cites Detecting human actions in surveillance videos.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Detecting human actions in surveillance videos

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.323223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.398600Z digest=sha256:acbbbbc5ab4f4d86a70be0dc0a7a1168cee4aeeae617e42ff2f6253b79cd92f8

Observation ee6e3520-08f2-4373-9558-c55e4b34417c · outbound

This paper cites Contextualized spatio-temporal contrastive learning with self-supervision.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Contextualized spatio-temporal contrastive learning with self-supervision

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.286810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.404552Z digest=sha256:04b0209dc0e9d710d31c09848954259aad642b314857705f5c9f4ebe4ddf34c4

Observation 1a35debf-220f-4151-9dcf-e3d2c372e435 · outbound

This paper cites Vision-based garbage dumping action detection for real-world surveillance platform.ETRI Journal, 41(4):494–505, 2019.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vision-based garbage dumping action detection for real-world surveillance platform.ETRI Journal, 41(4):494–505, 2019

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.259768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.412125Z digest=sha256:5c63451f7a3cfa9817bf3929f87f435ce9790d7958be9bc25bf3aba8ab806ca4

Observation ba00c246-766c-4df0-9b77-2e412d624b1c · outbound

This paper cites Open-vocabulary detr with conditional matching.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary detr with conditional matching

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.227399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.418542Z digest=sha256:630c3e628683cc61afa72dfcc26be049180ead767651e5db200f19d11c134773

Observation 778c0bbe-b2f3-4369-b6e5-55093f7d3d82 · outbound

This paper cites Open-vocabulary object detection using captions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.195512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.428209Z digest=sha256:c34e8338eec331e1dab1d939a98efc466849a1627d0fb4491c79e4310b7d3c8e

Observation c3229a69-53c5-4c6a-b995-cd6e3a770985 · outbound

This paper cites Open-vocabulary object detection using captions.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Open-vocabulary object detection using captions

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.168723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.432539Z digest=sha256:0acc799d35a09a4e28cb2f4afb96726ef84d58baa36fd5cabc1d869d725a0f32

Observation f94dff5c-4afc-40ef-b6f8-8bf85f6e1a47 · outbound

This paper cites From recognition to cognition: Visual commonsense reason- ing.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection From recognition to cognition: Visual commonsense reason- ing

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.139140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.443045Z digest=sha256:a47f242dc5c0135ef6d27a532ad5d69214e942e412a59a5848a72724538dbef0

Observation c2e45b16-72d3-4443-9478-a1e903d3ec23 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Vinvl: Revisiting visual representations in vision-language models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.089742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.449259Z digest=sha256:af2aa0fbff1e9fae4424c51af1597af4375a45d9d8d1ab17ebb0e46bd7c50326

Observation 2316d72b-c42a-4441-a29f-faeed5df7917 · outbound

This paper cites Tuber: Tubelet transformer for video action detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Tuber: Tubelet transformer for video action detection

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.065261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.453982Z digest=sha256:553bd3477740fcd00ef78ffb87424d40ca7195e068d59120b0cfde34eff9d257

Observation fb604aaf-58c5-40a6-8b26-0a72517c0866 · outbound

This paper cites MRSN: Multi-Relation Support Network for Video Action Detection.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection MRSN: Multi-Relation Support Network for Video Action Detection

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:14:34.558149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.460035Z digest=sha256:f3a209d3132aaf345c44e693a039935a805eca2a5c2c8790c49e64c2af50cbb1

Observation bf65d351-a680-4c7c-a476-f51acb3f9e45 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Regionclip: Region-based language-image pretraining

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.040683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.472500Z digest=sha256:07a59e7c2411352154acde4c82174826d3fda54c17ca95eda1160fe9e0077589

Observation 8f0230f2-e4ce-4cdc-98e5-cc72b0593fee · outbound

This paper cites Learning deep features for discrimi- native localization.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning deep features for discrimi- native localization

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:35.005673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.477338Z digest=sha256:74459c0c167ec45a8cac04db7358b6ec89491a6c74e4dde87a5a3c696d9d4254

Observation 2a8ecf9d-e5b1-4f1f-9cea-074b3d180b78 · outbound

This paper cites Learning to prompt for vision-language models.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Learning to prompt for vision-language models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.482606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.482606Z digest=sha256:fffeb3f4dffe9659e59ab2cda5cc49e30b639f03f3e3ed86d68863381795e447

Observation 7d972e5f-770b-48e3-93a5-930a182a96a3 · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot se- mantic segmentation.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection Zegclip: Towards adapting clip for zero-shot se- mantic segmentation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:34.487802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:34.487802Z digest=sha256:1b06f5e8cd3241b33893f2fc82a5df8a3ae92d1313a40b1bab02a54a3e2bd2ea

Observation 6347dfc0-d29f-4afe-915d-746364a3c506 · outbound

This paper cites For the action type {CLS}, what are the visual descriptions? Please respond with a list of 16 short sentences.

Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection For the action type {CLS}, what are the visual descriptions? Please respond with a list of 16 short sentences

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:34.950581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T19:14:34.495197Z digest=sha256:d438aea47d5fc32409615ccea7c06dfd73655343f957f95bda739b29c767e2c4

Pith citing papers

No inbound Pith citation observations are available.