Pith. sign in

Paper Citation Record · LEDGER

ActionCLIP: A New Paradigm for Video Action Recognition

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2109.08472.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.08472 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:31:23.569839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.259998Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5f97fecb-928b-4e39-a1ce-cd7485922d02 · inbound

InternVideo: General Video Foundation Models via Generative and Discriminative Learning cites this paper.

InternVideo: General Video Foundation Models via Generative and Discriminative Learning ActionCLIP: A New Paradigm for Video Action Recognition

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:36:53.345981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:36:53.235740Z digest=sha256:783917b4fdf7cba02650d91c5fed55a50fbc3f1317e0b3013882cca509b70d2a

Observation 63c284db-2302-4cfa-87e9-dda14f932cb2 · inbound

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels cites this paper.

SRL-CLIP: Efficient CLIP Video Adaptation via Structured Semantic Role Labels ActionCLIP: A New Paradigm for Video Action Recognition

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:48:54.278709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-24T04:46:18.542521Z digest=sha256:bca7f24e25bd6b6ec4c54278184b802153d8e5e6cdf00ddf3545eab14092fc5a

Observation 1a11ae7e-743f-4e0d-b04e-64b1d6bc76f4 · inbound

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models cites this paper.

LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:31:23.569839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:31:23.569839Z digest=sha256:c45db22e7ac4d20ebc304a94c089fe576bb77b22f9b805bdaab49ac6bf01d60b

Observation 717a7f45-29f6-4f8d-a1c0-9a0d749f5d34 · inbound

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living cites this paper.

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living ActionCLIP: A New Paradigm for Video Action Recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T04:44:58.877082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:44:58.877082Z digest=sha256:46c493be4d947812f336850ea0e60d51b1e5a934544ea0e31ad62e950d6817f6

Observation 04503e34-8e11-4846-8e67-8a304247919c · inbound

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners cites this paper.

Kronecker Mask and Interpretive Prompts are Language-Action Video Learners ActionCLIP: A New Paradigm for Video Action Recognition

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-09T04:37:09.819109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:37:09.819109Z digest=sha256:a7a9694c821ee86bc02907d40f612d87e68a1c1948434c558ac06e069df30e9b

Observation db75da33-556d-40dd-8625-aa56000599e1 · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.562855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.562855Z digest=sha256:5e7c29d70dbdec4fae7be54508e8fa8d1ca7ae7f0f524ce00131b99043f38fc8

Observation 64738afd-5232-4775-aa1e-ca4dfe9670b1 · inbound

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation cites this paper.

M2R2: MultiModal Robotic Representation for Temporal Action Segmentation ActionCLIP: A New Paradigm for Video Action Recognition

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.079251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T17:39:23.051951Z digest=sha256:af18cecb1a87562f0b1f1280d9bb81f06d54cf52488c5c89579137c155f681c4

Observation ca03ad4d-a30e-4554-a856-6b3bd56dea65 · inbound

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation cites this paper.

From Data to Modeling: Fully Open-vocabulary Scene Graph Generation ActionCLIP: A New Paradigm for Video Action Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:06.655155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:06:06.655155Z digest=sha256:835e0067a7ea34b64788689bb4c8ecc7ce0df7dc10809ffe2a4732f2dd929e48

Observation 9324a207-f7ac-4fba-aa21-6728672fc579 · inbound

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control cites this paper.

From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control ActionCLIP: A New Paradigm for Video Action Recognition

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:08.610513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:08.610513Z digest=sha256:75a5220c806c56d0825aca823d6d7ac680bce651d1ed5e16193c44c876cfdd05

Observation f133875f-465d-4055-a12b-a1b2bf6f951e · inbound

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory cites this paper.

FRAME: Pre-Training Video Feature Representations via Anticipation and Memory ActionCLIP: A New Paradigm for Video Action Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:26:59.796244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:26:59.796244Z digest=sha256:547756730cea07cd4f49ff77663bde8c2f9fbfcc6cfb7bfaefe3d82a9fa54624

Observation db80e471-c663-44e1-999d-d3c94c8a6fbd · inbound

Feature Hallucination for Self-supervised Action Recognition cites this paper.

Feature Hallucination for Self-supervised Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 168

Resolution
unresolved
no resolver link, observed 2026-08-06T22:55:24.152416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:55:24.152416Z digest=sha256:5e3a80f8162647807cfe270063c664028430b1091cf2ba373df85636b0b20cbc

Observation fefb174e-7ad2-4264-838b-62483d589634 · inbound

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition cites this paper.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.921969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.921969Z digest=sha256:634ceb5e8931bafce9ed27df2119c9c92ede8ca0caa521a3419b93e809cca4c1

Observation 4e85078f-a7d3-4f7d-9755-7bc0e018c68a · inbound

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching cites this paper.

LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:01.269123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:46:01.269123Z digest=sha256:80f6ac1c643e578cca5201d21e3a74786f70650f6368c47dbacd5ae76994e651

Observation 7064ddaa-7c79-407f-98fb-91c5b4832711 · inbound

MReg: A Novel Regression Model with MoE-based Video Feature Mining for Mitral Regurgitation Diagnosis cites this paper.

MReg: A Novel Regression Model with MoE-based Video Feature Mining for Mitral Regurgitation Diagnosis ActionCLIP: A New Paradigm for Video Action Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:44.379725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:44.379725Z digest=sha256:7dad859a6a0807582ecd6d7663d21399bb87d3cebbe0e1de9a29706ccce57efb

Observation 5e0caaa7-7186-43ea-a47c-e988f75f6f33 · inbound

Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives cites this paper.

Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives ActionCLIP: A New Paradigm for Video Action Recognition

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T21:29:05.004317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:29:05.004317Z digest=sha256:1c3bea1ca9a734fa3b1b5ae8d31c701ae74d7b8101e0af1fb7fb6b5ad58b0e8c

Observation 82048a79-c77a-45fa-830c-fd89501bfcc4 · inbound

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People cites this paper.

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People ActionCLIP: A New Paradigm for Video Action Recognition

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:47.752536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:47.752536Z digest=sha256:ea9d71bca048433ad6f8a739a457734b4e781101f5a46a7f09a2658e84cfd6da

Observation cbe83651-a236-407f-91e9-f4fb52d25874 · inbound

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition cites this paper.

Cross-Modal Dual-Causal Learning for Long-Term Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:06:10.178430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:06:10.178430Z digest=sha256:609fb66c20aa6968e63bb5fb6ac849d846b8dceb5c67139198d4911e77089793

Observation 0e899e0e-25bf-4e3e-ab1b-d6b312477551 · inbound

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space cites this paper.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space ActionCLIP: A New Paradigm for Video Action Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.168158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.168158Z digest=sha256:0ddc3155cf4c88b226cb3eb3f8d47c208fa21a3cf3afed8f2c72dc2a35876d94

Observation 332235fb-b1dd-4aa6-915b-f080b3dd2d37 · inbound

MoExDA: Domain Adaptation for Edge-based Action Recognition cites this paper.

MoExDA: Domain Adaptation for Edge-based Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T04:53:54.848055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:53:54.848055Z digest=sha256:d6b896b886d5a31a04f35f56f43a03d8b87b32a970491df9e545649ab90cdf7e

Observation 8793f05a-a5f2-4c71-9e9d-648154d60476 · inbound

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models cites this paper.

Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T16:56:27.442977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:56:27.442977Z digest=sha256:d7bb06423905e1ef40ff45843a26e727ea09c1e234549257f649d80e487ba2dc

Observation 3d76a287-af6c-469c-bde0-b99ccc3cfd3f · inbound

What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos cites this paper.

What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos ActionCLIP: A New Paradigm for Video Action Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T14:02:04.935174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:02:04.935174Z digest=sha256:4a83920acf3e8f27f38bd618554947bb3252feb5dc6cc2005b8132c06e2776f3

Observation d81fe392-c880-4688-9e6e-3676c10a83e5 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:41.696663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:41.696663Z digest=sha256:06a5bcad577f930b28d9645352b9985017148a669cac09b83d80e84197245249

Observation c2e75ed0-5b86-48b1-9aaa-87c5ba2f0735 · inbound

Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection cites this paper.

Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:16:40.083368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T17:14:25.533181Z digest=sha256:c2221ec669dbe0a417cacbefa3df0bbbe691a861b9bf2772cd0b8f3273d5d321

Observation 0984b6a3-f540-4e40-80e9-ffdbb441b3ec · inbound

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation cites this paper.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation ActionCLIP: A New Paradigm for Video Action Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:21.923280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:21.923280Z digest=sha256:d7b8b0fe9974a03eec8e8a173b169fc3d9ab1eedba7c5f095cefaa6239e1461b

Observation 20e27860-9ce1-44c0-8ec3-f1d915dd61d7 · inbound

Adapting MLLMs for Nuanced Video Retrieval cites this paper.

Adapting MLLMs for Nuanced Video Retrieval ActionCLIP: A New Paradigm for Video Action Recognition

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.876851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:20:09.051957Z digest=sha256:1abc28b5d4dccb215f11e95f301749ccac12bcee1cf20b37d15d9a45a5e95e41

Observation de5a1c17-d768-477b-ad83-b51daef083d0 · inbound

Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification cites this paper.

Context Matters: Peer-Aware Student Behavioral Engagement Measurement via VLM Action Parsing and LLM Sequence Classification ActionCLIP: A New Paradigm for Video Action Recognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:48:03.803999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:47:08.925519Z digest=sha256:3c7abced65117d92e38b63f7f9011810afecf53da5558d9ff4f781e07add8913

Observation 87ef0e5d-2b09-4f9c-93cd-d1c08ccddb87 · inbound

TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition cites this paper.

TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:40:59.528093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:33:28.632640Z digest=sha256:fe2a27ff371f6596d27b5aecbdf618a86152f7b40bd650da917141c1a683e511

Observation c9d0742d-7ef4-4a97-9a98-793f70546892 · inbound

EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges cites this paper.

EV-CLIP: Efficient Visual Prompt Adaptation for CLIP in Few-shot Action Recognition under Visual Challenges ActionCLIP: A New Paradigm for Video Action Recognition

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:07.304347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T12:29:32.897198Z digest=sha256:a366a5a37746ec35ad702c149d805ace0257a9af90deae04dfd2b2e56d0a17da

Observation 2441715e-1b69-4569-845d-2751f516b4fe · inbound

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition cites this paper.

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:15:22.369888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:13:20.840430Z digest=sha256:968cd377b8581f040a3c5061b6483b7b4df57bcad54df57da9a9814d3cc745c1

Observation 82f2df7e-a194-4303-a9bc-8dd02a225af3 · inbound

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer cites this paper.

VidPrism: Heterogeneous Mixture of Experts for Image-to-Video Transfer ActionCLIP: A New Paradigm for Video Action Recognition

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:03:25.922776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T13:01:42.880738Z digest=sha256:7e803ec46cb0ee997b13bcba86c0741e818b2f5a2e8badfed622b05293e5c733

Observation fbb1ef4b-4342-42aa-b0d0-84eb9d5bcc46 · inbound

Demystifying the Optimal Fair Classifier in Multi-Class Classification cites this paper.

Demystifying the Optimal Fair Classifier in Multi-Class Classification ActionCLIP: A New Paradigm for Video Action Recognition

Reference 184

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:52:35.633070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T18:49:29.377237Z digest=sha256:214f1c90e441845b8c8e3a3044c91353ed8027bdd20961e5ebe9ae8ed0e238cf

Observation b89ac15d-cda5-433e-ba95-01e276df524b · inbound

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding cites this paper.

Toward Low-Latency Vision-Language Models with Doubly-Correct Predictions in Egocentric Visual Understanding ActionCLIP: A New Paradigm for Video Action Recognition

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:20:00.470725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:47:12.032082Z digest=sha256:8631551bce19e9788d77e67bb28382462a113ee8ef2d7748093eb39f9674c97c

Observation 5105fa1e-95b7-4df3-9dbb-c4dc0fe454fa · inbound

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition cites this paper.

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:00:08.261733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:54:39.254414Z digest=sha256:3606d5d649e40d4f1a6b209afb42aeecf91f5480e71e17737528d054beb12e10

Observation f7dcf529-7636-4356-acf1-0530476bda84 · inbound

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition cites this paper.

TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.715522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T09:39:12.683562Z digest=sha256:508c251b88fc2ff9dbbdf78d0ebb2afa12967450b61fba195b567cfb0a36b50b

Observation 7b0ae27a-24ac-4b29-8281-5227163d1908 · inbound

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics cites this paper.

In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics ActionCLIP: A New Paradigm for Video Action Recognition

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:50.802179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:15:22.022274Z digest=sha256:896ca71ac999d5cccfdcdd0fa4e6bd1b96d8e8ad917994fd6991a358305af718

Observation 980d0c9c-a66e-42b1-944c-ae7f1d73f149 · inbound

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning cites this paper.

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning ActionCLIP: A New Paradigm for Video Action Recognition

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.208396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T05:08:02.242341Z digest=sha256:347493912e7cb788c92110fc52ca4016f473c5a1db4d61f585b8ca29a9bb0c19

Observation b1de98bf-1973-4d4e-9c6d-69ee5aaf6fbe · inbound

Breaking the 15% Barrier: A Real-World Data-Driven System for Proactive Social Robot Triggered by User Nonverbal Cues cites this paper.

Breaking the 15% Barrier: A Real-World Data-Driven System for Proactive Social Robot Triggered by User Nonverbal Cues ActionCLIP: A New Paradigm for Video Action Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T04:15:21.548043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:15:21.548043Z digest=sha256:b26d1489ef383dc3f7b81873d6150761e8fedcaf3700604769d80fded2ccb7cb

Observation 71a45313-f943-4d80-943f-5f4f0e7ccf5b · inbound

GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning cites this paper.

GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning ActionCLIP: A New Paradigm for Video Action Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T04:51:54.256318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T04:51:54.256318Z digest=sha256:a6ba4f430b181a5a6a63b6aef1e6021f94de88e73dc2c80de64725131cce2359

Observation 37ad474c-adf1-4974-8d68-3e0b5f0c6071 · inbound

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment cites this paper.

Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment ActionCLIP: A New Paradigm for Video Action Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T03:19:52.676974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:19:52.676974Z digest=sha256:f972cf10e7c677d3d32640f2a447a1ffc3e973cb691f877060d398981ef505cb

Observation 0cc73b54-d097-46e3-afed-9aa40a8047c3 · inbound

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition cites this paper.

Knowledge-guided Disentanglement with Atomic Actions for Action Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:41.727505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:02:41.727505Z digest=sha256:8b07fb4c40cdb7eb381eda0ce543c0086bfb5f141e45458c5066ebd6cf619a3d