Pith. sign in

Paper Citation Record · LEDGER

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition

As of 8 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2507.12426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12426 v2

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:52:55.019874Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact2
  • verified fuzzy70
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d6fa9a7-95b3-4e9b-b0c8-bfc6deb3b292 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.543670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.543670Z digest=sha256:a78ef456bd5be686c64e554e251a1a5a4941ae1b1e9e54c1bae013187f1c2ecb

Observation f6dfecc4-6bac-46de-8600-043905dbef6d · outbound

This paper cites Spatiotemporal residual networks for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal residual networks for video action recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.603264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.603264Z digest=sha256:b828cee742fd765f20bf9c7fd07a36cb1fa07748608a64f9aea92966a21bd3b5

Observation da7f7c26-20f2-4582-829d-229133615bd3 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.762590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.762590Z digest=sha256:eadf9790a9259f0258eafb939ef86efff135139939c01a3adda37bdd77f9431d

Observation 1139d00a-d086-4239-9ee1-f7fbbe824386 · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Large-scale video classification with convolutional neural networks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.834049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.834049Z digest=sha256:6a707f36d5585b8a3cfc8a0b9c6639289affca1c2e2d3dfea068a2d5e667caad

Observation 81df5fdd-3704-4da7-864d-36ff4ef50b15 · outbound

This paper cites Beyond short snippets: Deep networks for video classification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Beyond short snippets: Deep networks for video classification,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.912945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.912945Z digest=sha256:4fffb9a973078a07d9934fff3c08a5dee5457f04d6ef73b06c3a3458f3187320

Observation 30e07942-23e6-46a0-9685-6d2274f38ca8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Two-stream convolutional networks for action recognition in videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.999396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.999396Z digest=sha256:93dca2750c81c38e8177768eb978eff037b39e7bfb73a04dabcd4c9ce9de8884

Observation 2ab53b6f-85c6-446e-b41c-2ad3340312f1 · outbound

This paper cites Action recognition using deep 3d cnns with sequential feature aggregation and attention,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition using deep 3d cnns with sequential feature aggregation and attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.061463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.061463Z digest=sha256:ad221a6560a79d8bbbab2cdf51db35d54b64b1c56759d6a81caf50a2592bef6a

Observation befd3be9-d7eb-4f72-a1a3-e92c65ab6b2a · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A closer look at spatiotemporal convolutions for action recognition,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.139621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.139621Z digest=sha256:a793117bbb769a24d89a1e418a4597d1465fe908cf838feda965ceaeb683a2d6

Observation 89171f78-0d81-4a0f-8b44-9d16f4498847 · outbound

This paper cites Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.215422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.215422Z digest=sha256:b624b068795f08423469ac9609f265a652480ea06c7de2c466522ecebd9ae9c4

Observation ea61037a-f432-42f1-8c1e-5bc5d0ed4871 · outbound

This paper cites Action recognition in videos using pre-trained 2d convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition in videos using pre-trained 2d convolutional neural networks,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.293116Z digest=sha256:76f071c1245604b0c1d7dcf8ef1b2da0eb8a0e5d17b41d2cd013f734857f9970

Observation 57809d8d-af21-4dad-a9c4-caa9ce0ce57b · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video-focalnets: Spatio-temporal focal modulation for video action recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.365973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.365973Z digest=sha256:1a267fab3363c838b5f36aa6a456db28fc17a774924a34f5481063afc2dbc4a9

Observation cbd63ea4-ac98-4630-b0d5-2f34a2a50e21 · outbound

This paper cites Attention is all you need,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Attention is all you need,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.458192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.458192Z digest=sha256:081a86e147d74163034ed256dacfac110cdc7dd2dcd083aa0da160ee383ba6c3

Observation 94573c9d-c9d9-405f-a809-205fd9385d3b · outbound

This paper cites Morph: flexible acceleration for 3d cnn-based video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Morph: flexible acceleration for 3d cnn-based video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.549508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.549508Z digest=sha256:9eaa73fbbbf5541a4c8e9fabdbe0f1813a77f94f8b6ccb0e689c0259bc446799

Observation 7dacace1-b6ab-4d7f-a2f6-134e5c83ccb9 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.646960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.646960Z digest=sha256:aac01132818716b9240327d30483201bf0169850cead72368b2f866d8c2646cf

Observation f0cf805b-b70b-45ea-9f7b-4c0d38c27a59 · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.720373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.720373Z digest=sha256:a79f433a61f16b1cb526391a93c587e3f243ba9920f2b6b2b55506acb206ae69

Observation 3eb7b69d-ddd7-49e0-9c4f-fe022bf11c61 · outbound

This paper cites Vivit: A video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.347041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:47.802678Z digest=sha256:1d6dcf2d5cc4a002f7a930808abc45f660538a1caa38750a3a7c49d6821b853b

Observation 82e5e3f7-b6c3-41a8-9fc9-f6b66ec19b23 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:47.915356Z digest=sha256:696a0bb43bf482ff2070e0e576f3fbc906a4aa4e6e2509dd8d50efceecd8ba16

Observation 05571d0b-4696-43a4-a4c0-f3ef2b565420 · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.027968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.002422Z digest=sha256:9acae904a486d5001514611c2517896b43d05a2324c32c22090c545f06cf6505

Observation d19db6ac-64a3-4d5b-8bc4-38a75d42c01d · outbound

This paper cites Vivit: a video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: a video vision transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.838377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.113860Z digest=sha256:598f41723abbe3f10b961693a600d82f502d41a03ac82c3dbdb619cddbba8e9a

Observation b1bcdfe6-4cdc-44f8-aad7-61da12e91435 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A Short Note on the Kinetics-700 Human Action Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.257823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.257823Z digest=sha256:7b4aa1b61e953dda178e2f824c6eb1d4d5216a8f89f349044e83cf37254692ba

Observation a818e4a1-4b5e-4a94-b1fa-ad39061bc3c9 · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.619678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.310977Z digest=sha256:c14a14ab01203cbb08d20d7906bc7ad59812fd5dd484551c35d11b9dd8713d7d

Observation 619217a0-41a1-4d5f-bb13-e281e3ab75a4 · outbound

This paper cites Tsnet: token sparsification for efficient video transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsnet: token sparsification for efficient video transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.374484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.397443Z digest=sha256:18ccea8f101c02e1981af827203a6de7ea6a540b33b39b329f0cf23cc61fc632

Observation 14426831-847c-42f1-a119-ab7de1c633fd · outbound

This paper cites Dualformer: local- global stratified transformer for efficient video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dualformer: local- global stratified transformer for efficient video recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.198144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.495776Z digest=sha256:9c70c5392f0e14b1d3a29c30fe64ef30489bb69f2ac0a58caf67e4194d291bca

Observation 9e639f13-90b9-4e10-890f-1c9f41288763 · outbound

This paper cites Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.032271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.645264Z digest=sha256:0683658a1ebc0409079bc7f32153af46cc14df80b1c075d9327fa533722f61c7

Observation ecd963af-485a-4dc1-8fed-b62f79566714 · outbound

This paper cites Tsm: temporal shift module for efficient video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsm: temporal shift module for efficient video understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.810289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.745524Z digest=sha256:34e17bab256c73099e0ead520bd7aee915ba9ed7d8ed5346169ccfdd123fbf49

Observation 02dadde9-ea1b-43a4-9339-e281a66574ed · outbound

This paper cites Multi-stream interaction networks for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multi-stream interaction networks for human action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.903466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.903466Z digest=sha256:d7c6a77a53cd001ad08d62f3a154dab45b28c1bd17e26fca7a4d31ad3aa44b6e

Observation f4125143-b5a1-47c0-ad04-f863252f597f · outbound

This paper cites Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.577825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:48.988513Z digest=sha256:35995442ab77cbd12f95cd70aa45d05340219cb758476cf17c6897281a2431c7

Observation 7239b067-039d-4a18-b7ed-b876d7c032a5 · outbound

This paper cites Agpn: Action granularity pyramid network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Agpn: Action granularity pyramid network for video action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.413558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.052816Z digest=sha256:c34575f962c0f9144bdef0827fc6e7b0c6bdd705d7b43d8dd4758be58033b749

Observation 10739f60-4c9f-4362-bc5d-71672cacce9c · outbound

This paper cites Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.180393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.130816Z digest=sha256:287fbf8680ec75d1c37709624e29bca04629b3dc46c1170b1cd1949d8dc12450

Observation a3b87581-b7aa-4cce-8d76-78178a971912 · outbound

This paper cites Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.004396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.203358Z digest=sha256:98bfa2ace937531a1ef9a6f2a326f5b7683abd85232250a016f5d572f88a3d03

Observation dc449c7c-7aab-4408-9730-7df66158556e · outbound

This paper cites Decoupled knowledge distillation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Decoupled knowledge distillation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.788026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.291727Z digest=sha256:def46c90789cc0a094294799ff5328067fdb95f40c4574d382c2660404404776

Observation 6c54a01c-3a35-469e-8b2d-214a8e7009c1 · outbound

This paper cites Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.616547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.359338Z digest=sha256:96384bf640fc90aa98cb538ecae7a0e543456dac9b34fa4faa20496f4bb4b69e

Observation 0c4740cb-6e71-425d-b39b-8bb395344d67 · outbound

This paper cites Tomato leaf disease recognition based on multi-task distillation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tomato leaf disease recognition based on multi-task distillation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.441174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.445207Z digest=sha256:a4a6f31decb95fa4d9d890d4a5d52a881ff542e25433be2d66c374d2ed14d78c

Observation abc9a7cf-cff3-428b-86d2-4a8ec455593e · outbound

This paper cites Videoadviser: video knowledge distillation for multimodal transfer learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videoadviser: video knowledge distillation for multimodal transfer learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.264063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.510135Z digest=sha256:4fd01ac259e67af57e0ebe81a897a7af75bf95820fc47c7aed1853821b690706

Observation f0e5042d-7bdc-44c7-ad1e-eb7de09209f5 · outbound

This paper cites Generative model- based feature knowledge distillation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Generative model- based feature knowledge distillation for action recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.099709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.627318Z digest=sha256:73fc82bf0556f2307fcba9cd29e1c83fb342e3a24ed2ed75485d4808f42ab63f

Observation 7ac67c90-6540-4d09-9d3f-49fb1a495643 · outbound

This paper cites Distillation of human-object interaction contexts for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Distillation of human-object interaction contexts for action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.880266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.680694Z digest=sha256:cd22018b441a99ed838a7d47f2900db6d161fbb73774c439aa2bed1c636955eb

Observation 51a66055-0d5c-4e05-ae76-efc69f911c32 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Gaussian Error Linear Units (GELUs)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:49.779462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:49.779462Z digest=sha256:5e95cfc4b037d44bd8fc09ac6676eca6524e786d57f8bfdee7a8f4eaa4af376c

Observation 42c9bee4-121f-49f0-bb40-e03ca6c20299 · outbound

This paper cites Recognizing 50 human action categories of web videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Recognizing 50 human action categories of web videos,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.753222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:49.860074Z digest=sha256:a61499233e2fb10c6744fd454ad214edf169a3881ddd38b981f4bd06ac1ba25e

Observation 9a6faf9b-3c84-4091-854b-7615a44fba10 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.009362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.009362Z digest=sha256:aa0e77cec42f98aa6b8c34cd955fc21aca218a7b5be44fcaf3e99851e166d403

Observation 3f17c9e8-b6e2-4c9c-bddf-97a6c44e5f88 · outbound

This paper cites Hmdb: a large video database for human motion recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Hmdb: a large video database for human motion recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.650847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.106813Z digest=sha256:a92342f8abdd344c49c50e12f0b4dc7b88d441e4eee197c955bf427928f69cfd

Observation c75510fe-4399-4f99-b634-5292bf4a285f · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.556260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.174304Z digest=sha256:8e988cddbe8948d485b491b3a9a98f2c129df0d538d7b89ece32ac164033f46b

Observation b3a55fc2-d47e-44ba-bc91-dbe64eb2618d · outbound

This paper cites The Kinetics Human Action Video Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The Kinetics Human Action Video Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.256104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.256104Z digest=sha256:327de68814a15fdf1c44cb4c8e8d37810f4410656dc365288d34d45dc2751833

Observation 1c01ee1e-7001-44a3-a8f4-b282d3eed1db · outbound

This paper cites UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.390492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.390492Z digest=sha256:f4837fd8f5df69a906693371a143c72c7a5aa9689a3fc137ae2e8955330abdb3

Observation 771bb7c5-6e0d-495d-8627-a7fe6aec94b5 · outbound

This paper cites Going deeper with convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Going deeper with convolutions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.389286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.471227Z digest=sha256:172c71d7fd37a17a114a9ee86d9b33f63a4e618965ed358970b3faabfe1aab69

Observation 24febacd-fbf9-4fa0-95f4-43b6defeb682 · outbound

This paper cites Making sense of neuromorphic event data for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Making sense of neuromorphic event data for human action recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.263934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.520535Z digest=sha256:26b22629253a6c96ff0becb7a311665126b8286cb8b2ea2267d0f2a9f2fbc93d

Observation d670189b-3b5f-453a-9bb4-193e9c814505 · outbound

This paper cites Human action recognition using dis- tance transform and entropy based features,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using dis- tance transform and entropy based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.200658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.601096Z digest=sha256:4f3d5fd48566af02f4a7ea8cfe45db6c23cd67cf80148d8fe834d1ec37fe0291

Observation 121473c8-a8a4-402a-bb27-b0ab77aab99d · outbound

This paper cites Human action recognition using hybrid deep evolving neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using hybrid deep evolving neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.116040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.719001Z digest=sha256:23efd9414aece686f37a97bba8e8abc3d5dc940cf0729da060d5865c8e585174

Observation 36e68b1f-41af-4cf8-b42b-2f5eebc74e94 · outbound

This paper cites Simple-action-guided dictionary learning for complex action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Simple-action-guided dictionary learning for complex action recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.035273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.830577Z digest=sha256:0aa36e1e247a14d50f5bd3b62c30ad9403c65cc4a767d46584c29d49bb415911

Observation 6a77b057-5bb1-4b19-9eeb-6b4c9707ba87 · outbound

This paper cites Human activity classification using the 3dcnn architecture,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human activity classification using the 3dcnn architecture,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.921818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:50.937287Z digest=sha256:59581b2049aaaa6be8bc9ee10ee4e58eb9820cd9e59b8c2e2e69d8ce7221f5da

Observation 46e3ecc0-c91d-4f86-8982-338625a5b2c6 · outbound

This paper cites Fast classification and action recognition with event-based imaging,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Fast classification and action recognition with event-based imaging,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.770390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.096512Z digest=sha256:3d937115521307a9c0efdf13667b9ceed841b9faca2d6acb76a135397c1c0b7c

Observation 04678e70-a280-4266-b2bf-1cbc78009620 · outbound

This paper cites Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.610341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.214457Z digest=sha256:c3cd19f7b94753bb2fcab804f0fcf9d60b8446780b349bdb44a7e9cd9c4dfaaa

Observation 26c54948-d8b3-443e-b9b7-c1fb8ddcacac · outbound

This paper cites Human action recognition using multi-stream fusion and hybrid deep neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using multi-stream fusion and hybrid deep neural networks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.327664Z digest=sha256:e3e13a8096d4e163824434fe6c3ec609eb86fb15c9512f6b58f214cb7d23e3ab

Observation 51710b76-8eed-4762-b895-c9825289873f · outbound

This paper cites Self-supervised video representation learning by uncovering spatio-temporal statistics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video representation learning by uncovering spatio-temporal statistics,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.420896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.425880Z digest=sha256:624ce55c786b90da60fc9cb28f3f63d915c28d017c3eb7ed0d66873b13a4c4e4

Observation 7057970c-2ce6-4cfe-a130-d6867f5205b4 · outbound

This paper cites Enhancing self-supervised video representation learning via multi-level feature optimization,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Enhancing self-supervised video representation learning via multi-level feature optimization,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.332744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.505885Z digest=sha256:c6817c2216093c4d625e3181a868d7d98947387233fe4311b071a268bccebc1f

Observation d59234f2-5c29-47a9-83f8-a177742866da · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videomoco: Contrastive video representation learning with temporally adversarial examples,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.169505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.608989Z digest=sha256:0d71fa33400257534a5cc3de6691791c5df988c5d803d9030ceb3084a9c19217

Observation ef665412-a01e-4b14-ae2b-f6664f0c0d74 · outbound

This paper cites Action recognition from a single coded image,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition from a single coded image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.708625Z digest=sha256:535240e97fcfa528b2d0e61c0ceac6d8afcd8d911d2d0c5e726a40744e9e793b

Observation 7c3b9fb0-aafb-4894-bebe-e9dd1baf0f02 · outbound

This paper cites Tclr: Temporal contrastive learning for video representation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tclr: Temporal contrastive learning for video representation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.949242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.782435Z digest=sha256:3540c1fa1315cb56b2388c7cec9b7d05eab10a1b6c53470e16a913ad5e655688

Observation 8e4ccfc8-44d5-4f05-b48a-8474961d34c8 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learn2augment: learning to composite videos for data augmentation in action recognition,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.857703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.887584Z digest=sha256:b3afbdd0826291b1b517e9fddba8e55386407a54be3de768c6bf819165ea9b44

Observation 5e96c946-27dc-4cae-b725-e151ef84c116 · outbound

This paper cites Learning from temporal gradient for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning from temporal gradient for semi-supervised action recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.657387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:51.993613Z digest=sha256:dbdb3f6dbe89d89374ecbc4673309372fa3b1fca3547cd16290f78230d07e490

Observation 627c3103-b39a-40f0-aac2-395d19dc5380 · outbound

This paper cites Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.326931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.098718Z digest=sha256:caab44e8325ea7dade611f19461b2e44d4cf89e54c2784ed986ff03e45ea976f

Observation 64a00eee-4644-447b-b2c3-22bcf3d2d01a · outbound

This paper cites Extreme low- resolution action recognition with confident spatial-temporal attention transfer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Extreme low- resolution action recognition with confident spatial-temporal attention transfer,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.478916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.195968Z digest=sha256:c09646fb2faa6018be6e210b392f6074bc38549e6572f74418bbe9fa0e4ab922

Observation 0a6f9449-9838-4d17-a28b-95f3b333cdee · outbound

This paper cites Self-supervised video-based action recognition with disturbances,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video-based action recognition with disturbances,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.286153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.315336Z digest=sha256:d2d4ae48716ca0c8620608b33c6e4f57942616c8f51ef7da58313c682420cfd0

Observation 24aaf547-bc02-420f-8f6f-ae968884f75f · outbound

This paper cites Spatial-temporal exclusive capsule network for open set action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal exclusive capsule network for open set action recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.146160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.428855Z digest=sha256:0ba3aacaf1a0c8e1b3007b5ba0f4e121a10e7b543da108a25143f0cd8328264b

Observation 7cff12ac-649e-46be-b717-016115aaf030 · outbound

This paper cites Sv- former: Semi-supervised video transformer for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sv- former: Semi-supervised video transformer for action recognition,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.986980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.509199Z digest=sha256:30e181405eda16f099d64f0453074a800845d2a92099ef0dd51b545627900c03

Observation 6e74282b-f332-4229-bd4c-a2008662ed7a · outbound

This paper cites ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.174880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.604016Z digest=sha256:29f8420017ffcd90f38e2d87d14c5ee018237af88ccef41d7b04090a239edb98

Observation 3ac35da3-01a0-4f25-8917-be0cde4a5013 · outbound

This paper cites Self-supervised learning via multi-transformation classification for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised learning via multi-transformation classification for action recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.802805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.730793Z digest=sha256:de80257f0aaec9238d7af37e8069cca9a3f1156146920220600f197f6724124a

Observation 48d46bd4-9904-490f-9eaa-43495d8f3c7f · outbound

This paper cites Semi-supervised action recog- nition with dynamic temporal information fusion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Semi-supervised action recog- nition with dynamic temporal information fusion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.610844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.788124Z digest=sha256:1cd993e16a0cf9aa24b0707b053b447ea96ea7bb2f19e8381eb1e9951b8af581

Observation fa98add2-c079-44e4-aac9-8e0bc6c79754 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal contrastive video representation learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.852312Z digest=sha256:5f804b44d7347279eb7f42be367f7566208fd91c42f1a8f3176bd394cc5d19c1

Observation 9dd942c3-2298-4b33-bc85-24eb0ffdef8a · outbound

This paper cites Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.292874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.921579Z digest=sha256:547c10402cfbcff7f070466883524c34fd56999604fd0e11c99053daff12b5ec

Observation 001c1faf-2447-4562-8b39-aec022990205 · outbound

This paper cites Motion-driven visual tempo learn- ing for video-based action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Motion-driven visual tempo learn- ing for video-based action recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.149388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:52.991457Z digest=sha256:6a36f7a0c42c5d6e7051387b1f7fd29592877b5a72e9536d750182cf6d342fbd

Observation 1fea8405-db0e-418c-a4ed-095cc6fcfc9b · outbound

This paper cites Learning spatiotemporal and motion features in a unified 2d network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal and motion features in a unified 2d network for action recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.958930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.056544Z digest=sha256:fd950525ca5b41db0742e9102c2592da9a0f8dca43373036dd6d473ab3dc5586

Observation 72f9b3c9-d38f-4a75-b885-c9606e8b1c55 · outbound

This paper cites Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.805940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.175022Z digest=sha256:8d1122e218b584778a1941e09927a9f4f79ffb653cfbe22445d6f78dcf2b8df8

Observation 6274f1a9-a5e3-4101-ace2-d631abf33154 · outbound

This paper cites Spatial-temporal interleaved net- work for efficient action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal interleaved net- work for efficient action recognition,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.668315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.283537Z digest=sha256:07dbb03b7aee876694b19cd798731acdbc6e9fb11ac415b7faa6483a0fa4ee3b

Observation 35e04df3-1d2c-4336-90ce-67622692a76b · outbound

This paper cites A hybrid transformer framework for efficient activity recog- nition using consumer electronics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A hybrid transformer framework for efficient activity recog- nition using consumer electronics,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.469065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.396459Z digest=sha256:decca9d1bf858504bd68b90f00d25f58cd7cc93a3f87aa142d56495b929bc263

Observation affa9eb6-feaf-49ce-ac9c-4e00473fd572 · outbound

This paper cites A knowledge-based hierarchical causal inference network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A knowledge-based hierarchical causal inference network for video action recognition,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.287590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.500055Z digest=sha256:472968d390466cfa6d32eee093dd9049af4b9ea90dd16547f95bebcefbb0e066

Observation 158182ec-519a-4c5b-9884-f96baf7d8dfb · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:53.559512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:53.559512Z digest=sha256:a6a0718002d8f243761d79f0e9481eb7ce7b8494d5c9938da8971f4eabf23898

Observation 070b335d-9239-4dfb-b0cb-f10f38c64ee1 · outbound

This paper cites Vidtr: Video transformer without convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vidtr: Video transformer without convolutions,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.141635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.632514Z digest=sha256:2b19ed4dd8f93a8330b297bda423b797587941d4de2c6b63c7afd2e3dd79713f

Observation cc5a6118-9795-47ed-af13-2a4b5f8d00de · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.961367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.714960Z digest=sha256:f9acd1f2c01238ec4079dcc7a1815bf5c1e61ec5eeb5b137c2857536662c0f04

Observation 7be14330-1fd3-4783-b88b-c3567693fad5 · outbound

This paper cites Multiscale vision transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiscale vision transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.813054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.837040Z digest=sha256:17c8cd87c420cd52fe0b25333293539a720026b53c06f88e43d6d26a282631c0

Observation 30bfd305-887a-4872-9621-edd153208baf · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.677629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:53.978259Z digest=sha256:5dd9e7c0a9f7d3bc50ef84f7d7a877709d68f09073f16fc292a400bda1c89e84

Observation f796a309-6d28-429e-ab81-1870fb80b295 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.567092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.063447Z digest=sha256:00ff613dc4be5459d5f678499077c625b9fe3f9aa1432821c4382381f65634ca

Observation 7af3d5bd-983f-4ee8-a0c6-9cd773fde5e9 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mvitv2: Improved multiscale vision transformers for classification and detection,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.452633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.121446Z digest=sha256:f2a5aa4b52611ae176c16099ef05f0dd28365c39607e3dc8c3a2fb145200a4cf

Observation 1ac19103-74cc-4d7a-b7f4-eef175cdbd75 · outbound

This paper cites A novel spatio-temporal-wise network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A novel spatio-temporal-wise network for action recognition,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.285877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.207067Z digest=sha256:1168b57be803ec016d6ca493f85656fbc5ae882c93425638dde604b977daf352

Observation bf3879a6-426f-4218-82e6-3ee66704101c · outbound

This paper cites D-tsm: Discriminative temporal shift module for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition D-tsm: Discriminative temporal shift module for action recognition,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.163489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.282421Z digest=sha256:c0018a4afd983981888e9f97faab2b29e163f6c1d6b03023a9e517370e1f8c74

Observation e352fb0d-f249-40ef-b2c8-c164288a4a29 · outbound

This paper cites Scene adaptive mechanism for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Scene adaptive mechanism for action recognition,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.021801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.344367Z digest=sha256:7036d0b04b088117b32fc4151dd312da7136e6be4c606e6288518c63581c47e4

Observation 8affa414-54cd-4ff1-b837-4f01b33aaef0 · outbound

This paper cites Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.855411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.396120Z digest=sha256:edb65340b1f3e760cf61ba0ff18c6820461c50465fe65a25d78ad5d50b370350

Observation bab42be6-c522-4b76-95fa-665b474a8812 · outbound

This paper cites Short-term action learning for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Short-term action learning for video action recognition,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.720156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.487275Z digest=sha256:bd9ef5f0130788aa1f86f82acb2f9de4072ab4d2587d927672f4862f1427c7da

Observation 75ed9bf1-ed5b-4bae-a3e7-d754c807b0ba · outbound

This paper cites Tea: Temporal excitation and aggregation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tea: Temporal excitation and aggregation for action recognition,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.585739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.566351Z digest=sha256:4d0a1feb947c3948d4d5a4e891cf71b8c2b19779efd43d0cf8f32f41e114e584

Observation 1c59ee0d-d2d0-4368-8eab-645b74f76a7e · outbound

This paper cites Movinets: Mobile video networks for efficient video recogni- tion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Movinets: Mobile video networks for efficient video recogni- tion,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.394162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.657092Z digest=sha256:f36bdc5819f428f1148070d5d8c8b378e29113f8cf97959739c8e7a4afe1bd4f

Observation 3901d6cc-435a-4bb1-87dc-5176f6e73f41 · outbound

This paper cites Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.180372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.746931Z digest=sha256:24b45d02ef74d6bf1c71cd838821f5ab2a02e24e534c0f9f6f927fa9be14eb27

Observation 89d04c6f-4153-4f0e-b987-06a169856d17 · outbound

This paper cites Dilated multi-temporal modeling for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dilated multi-temporal modeling for action recognition,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.035863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.862860Z digest=sha256:902ec9f4247b72f9c12b6929b52dbb8686a61e902b4d374973c7d28fd17e3d93

Observation 2d9693e9-22a7-4efb-9789-44932768db2c · outbound

This paper cites Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:54.894113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:54.894113Z digest=sha256:364b764baf034ef3a62bff73374527a5f917ec4bd6f35978f5b87bebe8c2276a

Observation a4bdf994-b313-41e5-a2f3-029c2ad88fd7 · outbound

This paper cites Discrimina- tive segment focus network for fine-grained video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Discrimina- tive segment focus network for fine-grained video action recognition,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.869769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.944965Z digest=sha256:2a47aa75b76099537b4eb9db6b0dd27c33cfa4ff7225942daa2ef1b2e9b2b95a

Observation 37f17f06-1d35-4c0e-ad1f-0514af327267 · outbound

This paper cites Temporal difference attention for action recog- nition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Temporal difference attention for action recog- nition,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.666237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:54.978639Z digest=sha256:6cec33fa847e513f1ef1cbd5e08429b07d06f68d6b8c5a986a7d0c246b68e46e

Observation 62891965-1003-479c-9435-0d893005c4e6 · outbound

This paper cites An efficient motion visual learning method for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition An efficient motion visual learning method for video action recognition,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.515616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:52:55.019874Z digest=sha256:a4408f67f7fc0b0f55d9d2aefd084fd7b75bfb0c97df2b422dd93d80f51afdb7

Pith citing papers

No inbound Pith citation observations are available.