Pith. sign in

Paper Citation Record · LEDGER

AIM: Adapting Image Models for Efficient Video Action Recognition

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2302.03024.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.03024 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:02:21.096816Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:57:09.595862Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2944d945-23c6-4222-b44a-1bae9b2964e5 · inbound

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment cites this paper.

LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 211

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:27:59.144394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-17T03:27:58.952076Z digest=sha256:20358cafacb9356275f9fa9674a3ced8b784120675e27d4269d958720195ad65

Observation 8cd07702-9560-475b-835c-0737c4af8c0d · inbound

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling cites this paper.

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:16:01.083925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-24T05:14:41.833172Z digest=sha256:a703b7806a01c1ce6c3470685a0635cb04a1c2da217a5856dbdeec671e52b158

Observation 8278e180-6b12-4f8f-bfee-e260a5605e70 · inbound

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey cites this paper.

Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:32:36.964213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T11:32:36.738536Z digest=sha256:c09fa3645813277aba6b8bbb94ad4d76e656749f8a04451969567d49cfd34f6d

Observation cd240235-3b9b-41ee-b783-60e85f662059 · inbound

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models cites this paper.

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:21.096816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:02:21.096816Z digest=sha256:02bec571ddb3a65ae61a03b9da971a59a6076483188bfba55769bf728192490d

Observation 40b61d70-3788-421a-ba08-c6f1f372a2a6 · inbound

Structured Relational Reasoning for Group Activity Assessment cites this paper.

Structured Relational Reasoning for Group Activity Assessment AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T21:49:41.229532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:49:41.229532Z digest=sha256:0605956553c8aed5971d0504054c85efedb1d39aecc8ff8d0f5f947d367cdec4

Observation 82f17ee0-9989-4a68-a2c4-e9a6126c7f2a · inbound

MPT: Motion Prompt Tuning for Micro-Expression Recognition cites this paper.

MPT: Motion Prompt Tuning for Micro-Expression Recognition AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T21:05:11.716434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:05:11.716434Z digest=sha256:a279c9756839dcc2bada1160665feea76b2e6a2df0d78bed65ee95e581701e0f

Observation 0c37717b-7435-4463-b43d-0bd8d978ebeb · inbound

GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models cites this paper.

GA2-CLIP: Generic Attribute Anchor for Efficient Prompt Tuningin Video-Language Models AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:19:04.966588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T05:16:47.253028Z digest=sha256:855419137ff15b0fbf95a953d85c20cf2f6f7d924262783439d4e3530936d0d4

Observation a6da16a7-e77c-4ffc-924f-621a0219e0f3 · inbound

Exploring High-Order Self-Similarity for Video Understanding cites this paper.

Exploring High-Order Self-Similarity for Video Understanding AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 92

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.439564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:59:03.890135Z digest=sha256:ad2b108a8b2abeb53bcc9c288b95a1b5cac50000e328f210d9d03c17169bd2e0

Observation 119103ea-5fcf-46b9-94a5-4b662a588ead · inbound

Seeing Through Fog: Towards Fog-Invariant Action Recognition cites this paper.

Seeing Through Fog: Towards Fog-Invariant Action Recognition AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:03:59.194902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T06:03:42.648937Z digest=sha256:5163dc0d17c7fc4fe7b2756eed2149cb9f926918c94d6b28cfae2ceb95a114d1

Observation a854cb2d-7d7a-4438-9889-11ba87f876b2 · inbound

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition cites this paper.

Spatio-Temporal Similarity Volume Aggregation for Open-Vocabulary Action Recognition AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:15:22.375170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:13:20.840430Z digest=sha256:bcb19154a6ef1a5b10d29c3a2ba157905e613352379a7936e542e55dce52c412

Observation 3b88dcbd-b944-4248-901d-70b1612e6fa3 · inbound

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition cites this paper.

Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:57:09.597195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:17:49.723802Z digest=sha256:49f526b3a67a34e4e0da6df26fe6d26220b6aeb85f846b8b81cf73f05a6b977c

Observation e3b70cbd-acb8-4db6-92a5-14ac9c9393ff · inbound

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning cites this paper.

TRUST: Efficient Abdominal Trauma Recognition via Image-to-Ultrasound-Video Transfer Learning AIM: Adapting Image Models for Efficient Video Action Recognition

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:51.198175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T05:08:02.242341Z digest=sha256:c598f466ad9bee3a7420ec6bca45b729c78f9f91b6aac5f2a74ae53d6ec82788