Pith. sign in

Paper Citation Record · LEDGER

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition

As of 11 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2507.12426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12426 v2

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:52:55.019874Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact2
  • verified fuzzy70
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d6fa9a7-95b3-4e9b-b0c8-bfc6deb3b292 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.543670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.543670Z digest=sha256:a3b087860909383001e4e483eb8058546cb05e763a6175b4b561a37a41f84385

Observation f6dfecc4-6bac-46de-8600-043905dbef6d · outbound

This paper cites Spatiotemporal residual networks for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal residual networks for video action recognition,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.603264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.603264Z digest=sha256:f2fd213275abcdb751572ee51980d91a896c8e407df3d88697d8a2ab16f1fbdd

Observation da7f7c26-20f2-4582-829d-229133615bd3 · outbound

This paper cites Learning spatiotemporal features with 3d convolutional networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal features with 3d convolutional networks,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.762590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.762590Z digest=sha256:e21689149610fbd86c91eb167e6355ba7956951f76aecac89ecdf4b3cc917c1d

Observation 1139d00a-d086-4239-9ee1-f7fbbe824386 · outbound

This paper cites Large-scale video classification with convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Large-scale video classification with convolutional neural networks,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.834049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.834049Z digest=sha256:b2964a0ad0b4655af0085df6de68e2da5406d0a1bce5377723fd5625e315c864

Observation 81df5fdd-3704-4da7-864d-36ff4ef50b15 · outbound

This paper cites Beyond short snippets: Deep networks for video classification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Beyond short snippets: Deep networks for video classification,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.912945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.912945Z digest=sha256:0d4869470f8649c46f43468daabf33e9a538802f6dd2a9484d3b6b0f4f729f3c

Observation 30e07942-23e6-46a0-9685-6d2274f38ca8 · outbound

This paper cites Two-stream convolutional networks for action recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Two-stream convolutional networks for action recognition in videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:46.999396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:46.999396Z digest=sha256:df3762ad8654110e84b4c988b4aa4d92db098b5ede0ecd6b27aff1f705a455fc

Observation 2ab53b6f-85c6-446e-b41c-2ad3340312f1 · outbound

This paper cites Action recognition using deep 3d cnns with sequential feature aggregation and attention,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition using deep 3d cnns with sequential feature aggregation and attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.061463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.061463Z digest=sha256:7ff4c278ccd6548524a10410e44a3be7b020d78c6ffe87fc309a12d3be5f1803

Observation befd3be9-d7eb-4f72-a1a3-e92c65ab6b2a · outbound

This paper cites A closer look at spatiotemporal convolutions for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A closer look at spatiotemporal convolutions for action recognition,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.139621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.139621Z digest=sha256:4d6fb0b87a0c9eb138d06424785d9e97356bc51ddc03e77d525bcea5a22ef58c

Observation 89171f78-0d81-4a0f-8b44-9d16f4498847 · outbound

This paper cites Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition in videos using convolution long short-term memory network with spatio-temporal networks,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.215422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.215422Z digest=sha256:8e53a275b4f2ee728dbe1f0ad3fe99f0a1908d37387a287451380803ec28d843

Observation ea61037a-f432-42f1-8c1e-5bc5d0ed4871 · outbound

This paper cites Action recognition in videos using pre-trained 2d convolutional neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition in videos using pre-trained 2d convolutional neural networks,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.293116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.293116Z digest=sha256:f073de163b497f7167b8ff203d984dc336c8e8bc24069beb3a50fac3a1caad6f

Observation 57809d8d-af21-4dad-a9c4-caa9ce0ce57b · outbound

This paper cites Video-focalnets: Spatio-temporal focal modulation for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video-focalnets: Spatio-temporal focal modulation for video action recognition,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.365973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.365973Z digest=sha256:6667d4f9a77b71d5438c79373e17b72f5e1177f20776d8a85f045022b2bc2bb4

Observation cbd63ea4-ac98-4630-b0d5-2f34a2a50e21 · outbound

This paper cites Attention is all you need,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Attention is all you need,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.458192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.458192Z digest=sha256:cafe574366d96a20e2fce582f63c180c5b0e22697d4891c128569c7ce94c6bf9

Observation 94573c9d-c9d9-405f-a809-205fd9385d3b · outbound

This paper cites Morph: flexible acceleration for 3d cnn-based video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Morph: flexible acceleration for 3d cnn-based video understanding,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.549508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.549508Z digest=sha256:7168b00f831bae34e76a9958c42c94c0e5c8210baa209bddda8f49d11d744e78

Observation 7dacace1-b6ab-4d7f-a2f6-134e5c83ccb9 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.646960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.646960Z digest=sha256:09c4d633715e0fc75fa06dba8a0d3c09e9eba75b4885b347414038bf9738939b

Observation f0cf805b-b70b-45ea-9f7b-4c0d38c27a59 · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:47.720373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:47.720373Z digest=sha256:b9b44c830d0df0cba6082d630ba74d98d3fdcdf56e3d6302a96c910a619eb388

Observation 3eb7b69d-ddd7-49e0-9c4f-fe022bf11c61 · outbound

This paper cites Vivit: A video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.347041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:47.802678Z digest=sha256:474bc39d517d885d75cb819a61b6ac227c5343ac89b475e26b4fb6d630a89506

Observation 82e5e3f7-b6c3-41a8-9fc9-f6b66ec19b23 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.256262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:47.915356Z digest=sha256:9b273ab47db45ab2b23d58b24f9e5eeaa5e69cf0ea38f43fc6b5af2e459d6446

Observation 05571d0b-4696-43a4-a4c0-f3ef2b565420 · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:06.027968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.002422Z digest=sha256:ee82ec5940f01016812edfa4328c2964278e8165a843aafd05f9e827513eee75

Observation d19db6ac-64a3-4d5b-8bc4-38a75d42c01d · outbound

This paper cites Vivit: a video vision transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vivit: a video vision transformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.838377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.113860Z digest=sha256:7c63ddd517c491e308905b83d33cdd8c57b080efde04926c024bb31d0001be94

Observation b1bcdfe6-4cdc-44f8-aad7-61da12e91435 · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A Short Note on the Kinetics-700 Human Action Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.257823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.257823Z digest=sha256:1af1467ac3f4bfe0b555d6428a5caa94fa7a0e8141d64df29125d693a9b3b038

Observation a818e4a1-4b5e-4a94-b1fa-ad39061bc3c9 · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.619678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.310977Z digest=sha256:330a24762aef90d65c241bac36aca3860f77a20e2770fd2ec151b65cfc5ce614

Observation 619217a0-41a1-4d5f-bb13-e281e3ab75a4 · outbound

This paper cites Tsnet: token sparsification for efficient video transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsnet: token sparsification for efficient video transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.374484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.397443Z digest=sha256:2797b235e34fae925e87d013edbce2ee7eb8fda819f5e8a647ce8a8532422e7a

Observation 14426831-847c-42f1-a119-ab7de1c633fd · outbound

This paper cites Dualformer: local- global stratified transformer for efficient video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dualformer: local- global stratified transformer for efficient video recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.198144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.495776Z digest=sha256:c2d623ca5afa57df85d121867585f37133f3a1a08c5d0f789d94de1c394a499a

Observation 9e639f13-90b9-4e10-890f-1c9f41288763 · outbound

This paper cites Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Aerobics action recognition algorithm based on three-dimensional convolutional neural network and multilabel clas- sification,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:05.032271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.645264Z digest=sha256:9dd97e7b2300283b88046496f626403afcd17e9ac850860a07f001aa24763c92

Observation ecd963af-485a-4dc1-8fed-b62f79566714 · outbound

This paper cites Tsm: temporal shift module for efficient video understanding,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tsm: temporal shift module for efficient video understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.810289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.745524Z digest=sha256:c98e777fc7eab66c600de6871859a3e6f1df2f23a30434eb9c0568e3489d482d

Observation 02dadde9-ea1b-43a4-9339-e281a66574ed · outbound

This paper cites Multi-stream interaction networks for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multi-stream interaction networks for human action recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:48.903466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:48.903466Z digest=sha256:3b46f66b20f480c60b47708366a78d0cfbc1c2e9af188bf61df79e73179f1e46

Observation f4125143-b5a1-47c0-ad04-f863252f597f · outbound

This paper cites Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.577825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:48.988513Z digest=sha256:b9f48b065967597a338f9e327d8fd4b1f85c0dd65a3d9fd69884ddece225500d

Observation 7239b067-039d-4a18-b7ed-b876d7c032a5 · outbound

This paper cites Agpn: Action granularity pyramid network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Agpn: Action granularity pyramid network for video action recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.413558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.052816Z digest=sha256:bba2b86937d929d2df0a28e95e966bb7a4b8b87876f26fe0d1d247a3fbff8d92

Observation 10739f60-4c9f-4362-bc5d-71672cacce9c · outbound

This paper cites Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mawkdn: A multimodal fusion wavelet knowledge distillation approach based on cross-view attention for action recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.180393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.130816Z digest=sha256:5d19aad472ed5528ce1f42c7e207b3581de8d0ecf3463eea3a3ab0d04a4625b4

Observation a3b87581-b7aa-4cce-8d76-78178a971912 · outbound

This paper cites Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Convolutional neural networks or vision transformers: who will win the race for action recognitions in visual data?

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:04.004396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.203358Z digest=sha256:f60059005e414b4c0c6705f6b81fb25098456053af5f0aa4ec38e7d89a528450

Observation dc449c7c-7aab-4408-9730-7df66158556e · outbound

This paper cites Decoupled knowledge distillation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Decoupled knowledge distillation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.788026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.291727Z digest=sha256:949e01a7afb64d15c81359dd56ea6395c9e17cb3c636e5ce67b300b136c59310

Observation 6c54a01c-3a35-469e-8b2d-214a8e7009c1 · outbound

This paper cites Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Knowledge distil- lation in video-based human action recognition: an intuitive approach to efficient and flexible model training,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.616547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.359338Z digest=sha256:e0397631a8aef23082653dc6a3ea7058d532a441a526ddfa36f2f1aeb7e7aaee

Observation 0c4740cb-6e71-425d-b39b-8bb395344d67 · outbound

This paper cites Tomato leaf disease recognition based on multi-task distillation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tomato leaf disease recognition based on multi-task distillation learning,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.441174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.445207Z digest=sha256:84c50633f1e199ae5341e5cc24d967c873c366d54023aa02449ae8525bfa4c13

Observation abc9a7cf-cff3-428b-86d2-4a8ec455593e · outbound

This paper cites Videoadviser: video knowledge distillation for multimodal transfer learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videoadviser: video knowledge distillation for multimodal transfer learning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.264063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.510135Z digest=sha256:9053c2f0cdc59f83b26a804bfd5bcf586987db6c0eb1ff3aa03950dbc7f90719

Observation f0e5042d-7bdc-44c7-ad1e-eb7de09209f5 · outbound

This paper cites Generative model- based feature knowledge distillation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Generative model- based feature knowledge distillation for action recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:03.099709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.627318Z digest=sha256:f3cf96ac663bcfcf2b64e13d096fbb8195b67d1a1806c13a2c505ec81deec6dc

Observation 7ac67c90-6540-4d09-9d3f-49fb1a495643 · outbound

This paper cites Distillation of human-object interaction contexts for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Distillation of human-object interaction contexts for action recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.880266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.680694Z digest=sha256:f0177cd0cf5fbe4b4b06716f7aaf2ff2877b349a57ac77d441ce10ba4ac1eb18

Observation 51a66055-0d5c-4e05-ae76-efc69f911c32 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Gaussian Error Linear Units (GELUs)

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:49.779462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:49.779462Z digest=sha256:cdc5d335c8c70796c3e986881bcaae1d8ba79f762e9ff031e8e3f53c48dbc016

Observation 42c9bee4-121f-49f0-bb40-e03ca6c20299 · outbound

This paper cites Recognizing 50 human action categories of web videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Recognizing 50 human action categories of web videos,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.753222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:49.860074Z digest=sha256:b1d50453228c648457ec51c312818087ee7694a6c66d02e5d99c7484ef05e8e1

Observation 9a6faf9b-3c84-4091-854b-7615a44fba10 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.009362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.009362Z digest=sha256:efe87ede17fab88b130d42cddb42fc4d29e4a22c6ab470739f6e7d6cbd6d6f3f

Observation 3f17c9e8-b6e2-4c9c-bddf-97a6c44e5f88 · outbound

This paper cites Hmdb: a large video database for human motion recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Hmdb: a large video database for human motion recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.650847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.106813Z digest=sha256:564b9bec109ec1755a7c3dd34e182f02f47c1bf8b7f7a11eda9528b802209cf5

Observation c75510fe-4399-4f99-b634-5292bf4a285f · outbound

This paper cites The" something something.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The" something something

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.556260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.174304Z digest=sha256:571e3f36e092b59ccdf56df976b77f478fb48ac2891d317648ee1615572267ac

Observation b3a55fc2-d47e-44ba-bc91-dbe64eb2618d · outbound

This paper cites The Kinetics Human Action Video Dataset.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition The Kinetics Human Action Video Dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.256104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.256104Z digest=sha256:f1061f39f1d027cb0fb1a088440e177d152f43b71532aa2eb0a56b3aef6f3204

Observation 1c01ee1e-7001-44a3-a8f4-b282d3eed1db · outbound

This paper cites UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition UniFormer: Unified Transformer for Efficient Spatiotemporal Representation Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:50.390492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:50.390492Z digest=sha256:d42a541c786be489ab766b0ab04aba25f9036e42c0eb9f34a4ab9fa7935adca7

Observation 771bb7c5-6e0d-495d-8627-a7fe6aec94b5 · outbound

This paper cites Going deeper with convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Going deeper with convolutions,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.389286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.471227Z digest=sha256:31938572265cc888ef8d427c22ac9100e1c73c1a65436a4197e84b055ce1bdd3

Observation 24febacd-fbf9-4fa0-95f4-43b6defeb682 · outbound

This paper cites Making sense of neuromorphic event data for human action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Making sense of neuromorphic event data for human action recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.263934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.520535Z digest=sha256:61ade682a92c8953d7fe3e147401d14e8ca93342c37c8749038f97f8bd5a0294

Observation d670189b-3b5f-453a-9bb4-193e9c814505 · outbound

This paper cites Human action recognition using dis- tance transform and entropy based features,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using dis- tance transform and entropy based features,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.200658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.601096Z digest=sha256:d284263443fd28c4f1b696b1b0420a1855c22dc7bcad42d88cff54cf2ab66924

Observation 121473c8-a8a4-402a-bb27-b0ab77aab99d · outbound

This paper cites Human action recognition using hybrid deep evolving neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using hybrid deep evolving neural networks,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.116040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.719001Z digest=sha256:e2b7377576932fdd4bec92e96c534dc1e9a3010fd261cc03226c6849c94c360b

Observation 36e68b1f-41af-4cf8-b42b-2f5eebc74e94 · outbound

This paper cites Simple-action-guided dictionary learning for complex action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Simple-action-guided dictionary learning for complex action recognition,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:02.035273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.830577Z digest=sha256:f5a32431ede586f42849555f9020e979d8155593f97545f4402a50bf4ab6d0c6

Observation 6a77b057-5bb1-4b19-9eeb-6b4c9707ba87 · outbound

This paper cites Human activity classification using the 3dcnn architecture,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human activity classification using the 3dcnn architecture,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.921818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:50.937287Z digest=sha256:0e983d501f769e2f4375e34cc29190e6511e912bf5818c8790c1e6cc02ab4305

Observation 46e3ecc0-c91d-4f86-8982-338625a5b2c6 · outbound

This paper cites Fast classification and action recognition with event-based imaging,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Fast classification and action recognition with event-based imaging,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.770390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.096512Z digest=sha256:7bdbc1ec114d237fee2860be7f42b65c2d1c783dde07e27f052726e9da515901

Observation 04678e70-a280-4266-b2bf-1cbc78009620 · outbound

This paper cites Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatio-temporal features based human action recognition using convolutional long short-term deep neural network,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.610341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.214457Z digest=sha256:6a53dd55b6ed9b514904aa13564555ba44d618c6948d25db7d40966583e91e5a

Observation 26c54948-d8b3-443e-b9b7-c1fb8ddcacac · outbound

This paper cites Human action recognition using multi-stream fusion and hybrid deep neural networks,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Human action recognition using multi-stream fusion and hybrid deep neural networks,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.479218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.327664Z digest=sha256:6eb9907e64a240a665cd72de6fe3cebbda5ed3165978c9b2d979b9a0bfd37a40

Observation 51710b76-8eed-4762-b895-c9825289873f · outbound

This paper cites Self-supervised video representation learning by uncovering spatio-temporal statistics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video representation learning by uncovering spatio-temporal statistics,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.420896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.425880Z digest=sha256:4d203d71ef9f9cd176fac46544ef93eb86dd60c2a059d46798bda5d5be02e4e9

Observation 7057970c-2ce6-4cfe-a130-d6867f5205b4 · outbound

This paper cites Enhancing self-supervised video representation learning via multi-level feature optimization,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Enhancing self-supervised video representation learning via multi-level feature optimization,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.332744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.505885Z digest=sha256:384e43a9becbd50509fec2d61a8b50d98321927c2ee6a66e9011227afbb5853c

Observation d59234f2-5c29-47a9-83f8-a177742866da · outbound

This paper cites Videomoco: Contrastive video representation learning with temporally adversarial examples,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Videomoco: Contrastive video representation learning with temporally adversarial examples,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.169505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.608989Z digest=sha256:45c4d7f325a28371753bbf36db8f15078c7d431e2d803c8f3e0294d71c4adfee

Observation ef665412-a01e-4b14-ae2b-f6664f0c0d74 · outbound

This paper cites Action recognition from a single coded image,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Action recognition from a single coded image,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:01.043969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.708625Z digest=sha256:0962e2780ef9af2a5116c58648961e67109cc1891e4ce10e3e776e7edcfedefa

Observation 7c3b9fb0-aafb-4894-bebe-e9dd1baf0f02 · outbound

This paper cites Tclr: Temporal contrastive learning for video representation,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tclr: Temporal contrastive learning for video representation,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.949242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.782435Z digest=sha256:8bf94295aa46bf62f8a07ecf961f0866a08c13cc1fa674d713ca84374ae27586

Observation 8e4ccfc8-44d5-4f05-b48a-8474961d34c8 · outbound

This paper cites Learn2augment: learning to composite videos for data augmentation in action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learn2augment: learning to composite videos for data augmentation in action recognition,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.857703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.887584Z digest=sha256:ddb0d97373bd73c7d9a560ed2ea89d40a69d674e6dd9dfed9472726f5f96bbfc

Observation 5e96c946-27dc-4cae-b725-e151ef84c116 · outbound

This paper cites Learning from temporal gradient for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning from temporal gradient for semi-supervised action recognition,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.657387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:51.993613Z digest=sha256:edb421d6a7ac84c29f030c5fbbba48624ab91de7898dd9bf4f31f9ce7ebf387e

Observation 627c3103-b39a-40f0-aac2-395d19dc5380 · outbound

This paper cites Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Preserve Pre-trained Knowledge: Transfer Learning With Self-Distillation For Action Recognition

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.326931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.098718Z digest=sha256:466140331158b777bbda52cc09142cfbbb474dcfe33f976e8c5bd3220276f5c2

Observation 64a00eee-4644-447b-b2c3-22bcf3d2d01a · outbound

This paper cites Extreme low- resolution action recognition with confident spatial-temporal attention transfer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Extreme low- resolution action recognition with confident spatial-temporal attention transfer,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.478916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.195968Z digest=sha256:bd17a93755711c1747e1555e77e5d7bc350e1685fe73c6da357bb67f3d48400f

Observation 0a6f9449-9838-4d17-a28b-95f3b333cdee · outbound

This paper cites Self-supervised video-based action recognition with disturbances,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised video-based action recognition with disturbances,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.286153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.315336Z digest=sha256:e85312a86451ad69e8f95bc0cf6b00c5600e7e540f6b53b26709e056225616f1

Observation 24aaf547-bc02-420f-8f6f-ae968884f75f · outbound

This paper cites Spatial-temporal exclusive capsule network for open set action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal exclusive capsule network for open set action recognition,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:53:00.146160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.428855Z digest=sha256:02023e0dd505295562f5b75ee933999225130328e423fabf36f35d48e506c37b

Observation 7cff12ac-649e-46be-b717-016115aaf030 · outbound

This paper cites Sv- former: Semi-supervised video transformer for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sv- former: Semi-supervised video transformer for action recognition,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.986980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.509199Z digest=sha256:d3e6cbfbb35727613e6c581805b71ccb8dece290172975ef2947ce21609dad0b

Observation 6e74282b-f332-4229-bd4c-a2008662ed7a · outbound

This paper cites ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:52:55.174880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.604016Z digest=sha256:d48208d0f107ed548b022c9676745515832d8ad7116ebc4e62fb5de7474ff59d

Observation 3ac35da3-01a0-4f25-8917-be0cde4a5013 · outbound

This paper cites Self-supervised learning via multi-transformation classification for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Self-supervised learning via multi-transformation classification for action recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.802805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.730793Z digest=sha256:e47b9d1627d4191876a5c94704c6ca524e1ddf3ac3afe3ebcbec86d8cf7608eb

Observation 48d46bd4-9904-490f-9eaa-43495d8f3c7f · outbound

This paper cites Semi-supervised action recog- nition with dynamic temporal information fusion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Semi-supervised action recog- nition with dynamic temporal information fusion,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.610844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.788124Z digest=sha256:89deb5d0fc98bb15783f84d6251ebbf1eceaa580bb70a976cf97423367f39df3

Observation fa98add2-c079-44e4-aac9-8e0bc6c79754 · outbound

This paper cites Spatiotemporal contrastive video representation learning,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatiotemporal contrastive video representation learning,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.422918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.852312Z digest=sha256:5da57944bb9a0f4a284a651ce7ce35e5d675702090a7f98ece99d9c609bb6176

Observation 9dd942c3-2298-4b33-bc85-24eb0ffdef8a · outbound

This paper cites Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Representation learning for compressed video action recognition via attentive cross- modal interaction with motion enhancement,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.292874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.921579Z digest=sha256:56291e1c9b53e389f6df88e4e8b04bcfde0d7bae6f0bd60aad142b10b1ccfacc

Observation 001c1faf-2447-4562-8b39-aec022990205 · outbound

This paper cites Motion-driven visual tempo learn- ing for video-based action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Motion-driven visual tempo learn- ing for video-based action recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:59.149388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:52.991457Z digest=sha256:b80413ee16c760f5547cf32ef306281d92300cc082b06d40ee53ba5707d8a6b7

Observation 1fea8405-db0e-418c-a4ed-095cc6fcfc9b · outbound

This paper cites Learning spatiotemporal and motion features in a unified 2d network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning spatiotemporal and motion features in a unified 2d network for action recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.958930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.056544Z digest=sha256:89d20e7258c53c7960745e4dfdabc308b478fafd6255beb9b30bfb87978c2553

Observation 72f9b3c9-d38f-4a75-b885-c9606e8b1c55 · outbound

This paper cites Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vit-ret: Vision and recurrent transformer neural networks for human activity recognition in videos,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.805940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.175022Z digest=sha256:ee215356a96042217352a9ebe185a08425c06424d979954868985f91ac004da8

Observation 6274f1a9-a5e3-4101-ace2-d631abf33154 · outbound

This paper cites Spatial-temporal interleaved net- work for efficient action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Spatial-temporal interleaved net- work for efficient action recognition,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.668315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.283537Z digest=sha256:328ffd6ee2ccb2ca8f134cc93fa4f815f9f26fd465d76c2cd227105ff512ce07

Observation 35e04df3-1d2c-4336-90ce-67622692a76b · outbound

This paper cites A hybrid transformer framework for efficient activity recog- nition using consumer electronics,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A hybrid transformer framework for efficient activity recog- nition using consumer electronics,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.469065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.396459Z digest=sha256:27d60ead11abc6cfed1b30dffda319f780945f8708904447474dc7b1da2dec5c

Observation affa9eb6-feaf-49ce-ac9c-4e00473fd572 · outbound

This paper cites A knowledge-based hierarchical causal inference network for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A knowledge-based hierarchical causal inference network for video action recognition,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.287590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.500055Z digest=sha256:4490680c564ed43d436f3eb22051618e544d496bc229cf2b6e44c33d48ea7ce0

Observation 158182ec-519a-4c5b-9884-f96baf7d8dfb · outbound

This paper cites Is space-time attention all you need for video understanding?.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Is space-time attention all you need for video understanding?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:53.559512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:53.559512Z digest=sha256:6b3759455aaaeb53b50dc0ebcd8faf7526ad91e7081d3c28279225cb9cc412d1

Observation 070b335d-9239-4dfb-b0cb-f10f38c64ee1 · outbound

This paper cites Vidtr: Video transformer without convolutions,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Vidtr: Video transformer without convolutions,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:58.141635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.632514Z digest=sha256:c76380832e7643fdfa85c3ae299cf0465931cee8d5a7493aa93b00fad135a1a4

Observation cc5a6118-9795-47ed-af13-2a4b5f8d00de · outbound

This paper cites Keeping your eye on the ball: Tra- jectory attention in video transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Keeping your eye on the ball: Tra- jectory attention in video transformers,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.961367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.714960Z digest=sha256:6e18dc35bd84dc832ec18f8d8db5d86de0048d6871112c40131abf7cfcefed46

Observation 7be14330-1fd3-4783-b88b-c3567693fad5 · outbound

This paper cites Multiscale vision transformers,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiscale vision transformers,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.813054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.837040Z digest=sha256:d5f936ece426320253ea1a22b519a02a999c53e22f484b3bf4933ea0ff4e5600

Observation 30bfd305-887a-4872-9621-edd153208baf · outbound

This paper cites Multiview transformers for video recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Multiview transformers for video recognition,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.677629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:53.978259Z digest=sha256:ee91aa153ee65eca389bd2852113f6f1228c62fed8c3c9839353514be0e4961b

Observation f796a309-6d28-429e-ab81-1870fb80b295 · outbound

This paper cites Video swin transformer,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Video swin transformer,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.567092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.063447Z digest=sha256:8ceb214a0bc1508b997ad777bd927ec09f44b360f694973cff0a8f7d55e41da9

Observation 7af3d5bd-983f-4ee8-a0c6-9cd773fde5e9 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Mvitv2: Improved multiscale vision transformers for classification and detection,

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.452633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.121446Z digest=sha256:3c7896047187277f502e43cdb87bc4e88d5f0f00888931b453a9f225ba702c42

Observation 1ac19103-74cc-4d7a-b7f4-eef175cdbd75 · outbound

This paper cites A novel spatio-temporal-wise network for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition A novel spatio-temporal-wise network for action recognition,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.285877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.207067Z digest=sha256:02e53664666d8023b6769de03b3c594c333c6fe66b5c7161e7e7fa449f21dbdf

Observation bf3879a6-426f-4218-82e6-3ee66704101c · outbound

This paper cites D-tsm: Discriminative temporal shift module for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition D-tsm: Discriminative temporal shift module for action recognition,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.163489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.282421Z digest=sha256:f37fba3be7dfefbe1f84598a24ff32570b07ac5efd8d15ea2d65bf08187ac882

Observation e352fb0d-f249-40ef-b2c8-c164288a4a29 · outbound

This paper cites Scene adaptive mechanism for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Scene adaptive mechanism for action recognition,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:57.021801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.344367Z digest=sha256:9bcd709a0c7c5ba662e1953ec67f8d8c902c7162d2e2d66c610655a89b933c13

Observation 8affa414-54cd-4ff1-b837-4f01b33aaef0 · outbound

This paper cites Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Sta+: Spatiotemporal adaptation with adaptive model selection for video action recognition,

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.855411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.396120Z digest=sha256:4e7562f863a1423ccaaf94c6a2eac69cca32aea9b8d52bff06d3135d89a93a65

Observation bab42be6-c522-4b76-95fa-665b474a8812 · outbound

This paper cites Short-term action learning for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Short-term action learning for video action recognition,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.720156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.487275Z digest=sha256:e850c6b32a0b19207dec5b77f817ef904865859a918bfdb00c9cf06effa2607f

Observation 75ed9bf1-ed5b-4bae-a3e7-d754c807b0ba · outbound

This paper cites Tea: Temporal excitation and aggregation for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Tea: Temporal excitation and aggregation for action recognition,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.585739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.566351Z digest=sha256:c2483c739297585f5e42eb1f214d83e4dc20ef54323e719dc41100066f50fe97

Observation 1c59ee0d-d2d0-4368-8eab-645b74f76a7e · outbound

This paper cites Movinets: Mobile video networks for efficient video recogni- tion,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Movinets: Mobile video networks for efficient video recogni- tion,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.394162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.657092Z digest=sha256:6248d160d273a9335ea822e52e182032a0e7f95e958bb75ee809f2163b314307

Observation 3901d6cc-435a-4bb1-87dc-5176f6e73f41 · outbound

This paper cites Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Timebal- ance: Temporally-invariant and temporally-distinctive video represen- tations for semi-supervised action recognition,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.180372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.746931Z digest=sha256:7e9ce50abf10b52da18ecfdc22057910744c800aa4332d716b114899803dfc2d

Observation 89d04c6f-4153-4f0e-b987-06a169856d17 · outbound

This paper cites Dilated multi-temporal modeling for action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Dilated multi-temporal modeling for action recognition,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:56.035863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.862860Z digest=sha256:6bb4e346c54044e2226d3281635752d0f42ecc0bc486f09d15fd0449274d68b4

Observation 2d9693e9-22a7-4efb-9789-44932768db2c · outbound

This paper cites Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T16:52:54.894113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:52:54.894113Z digest=sha256:bd11e5b6aef25b0d25fa48ff96171235deef13f12873f1ce7c450a99a34acc25

Observation a4bdf994-b313-41e5-a2f3-029c2ad88fd7 · outbound

This paper cites Discrimina- tive segment focus network for fine-grained video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Discrimina- tive segment focus network for fine-grained video action recognition,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.869769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.944965Z digest=sha256:4d309ae76fa16c0edc20e1c81c311f6f7ca56954b6b507cde56dd07eb9db333f

Observation 37f17f06-1d35-4c0e-ad1f-0514af327267 · outbound

This paper cites Temporal difference attention for action recog- nition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition Temporal difference attention for action recog- nition,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.666237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:54.978639Z digest=sha256:dfc4b5d0f9a91b7c963f5e184a1d4ce0c79ac64f5f8a1fc17c1da3fc97823b5a

Observation 62891965-1003-479c-9435-0d893005c4e6 · outbound

This paper cites An efficient motion visual learning method for video action recognition,.

DVFL-Net: A Lightweight Distilled Video Focal Modulation Network for Spatio-Temporal Action Recognition An efficient motion visual learning method for video action recognition,

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:52:55.515616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T16:52:55.019874Z digest=sha256:1d3566dded5a839689bf55a4dd9630d7253a68fc0d4cba10b5a20a0f26d0fa63

Pith citing papers

No inbound Pith citation observations are available.