Pith. sign in

REVIEW 4 major objections 5 minor 176 references

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis

T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This survey organizes deep video-understanding models by how they handle time — extracting spatial features per frame, splitting space and motion into separate streams, or learning spatiotemporal features directly.

desk verdict A broad but careless survey: the narrative is usable, but Table 3's source attributions are impossible and the internal dataset numbers conflict, so the survey can't be trusted as a reference without major revision. read the letter →

arxiv 2502.07277 v1 pith:5CASGN2C submitted 2025-02-11 cs.CV cs.AI

classification cs.CVcs.AI
keywords VideounderstandingActionrecognitionSpatiotemporalfeatures3DconvolutionalnetworksTwo-streamVisiontransformersbenchmarksDeepneural
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a narrative survey of deep neural networks for video understanding, organized around spatiotemporal feature extraction. It claims that the field's models fall into three methodologies: extracting spatial features frame-by-frame, extracting spatial and temporal features in separate streams, and extracting spatiotemporal features directly from the video volume. On that basis, the paper reviews the main model families, from two-stream CNNs to transformers, surveys major video and action-recognition datasets, compares reported model results, and catalogs open problems such as computational cost, input variance, and long-term dependencies. A reader who wants a structured orientation to the field would find a usable map of the model landscape and its benchmarks.

What carries the argument

The organizing device is the spatiotemporal feature, defined as information about both location and time that is most relevant for video understanding. The survey's taxonomy of three extraction methodologies, spatial-only, separate spatial-plus-temporal streams, and direct spatiotemporal extraction, carries the whole argument, with supporting concepts including the early, late, and slow temporal fusion types from [69], the two-stream architecture inspired by the brain's dorsal and ventral pathways, attention and shifting mechanisms, and aggregation methods such as NetVLAD. Each model family is presented as a way of instantiating one of the three extraction schemes.

What would settle it

Check the numbers in Table 3 against the original papers — for example, the Charades mAP for SlowFast and the Something-Something V1 accuracy for I3D — using the exact evaluation protocols reported there. If the numbers differ from the sources, or if the sources used incompatible protocols that the survey does not flag, the overview's comparisons would mislead.

Watch

Extended reading notes

Core claim

The paper's central claim is that video understanding models are best understood by how they treat the temporal dimension. It distinguishes three feature-extraction schemes: purely spatial extraction from individual frames, separate spatial and temporal streams, typically with optical flow or a trainable motion stream, and direct spatiotemporal extraction using 3D convolutions or other volume-level operations. The survey then uses this taxonomy to organize a broad review of structural designs, including temporal frame fusion, pooling and aggregation, attention and shifting, memory-based and recursive networks, multi-stream networks, and transformer models, and to frame the central challenges of the field. The paper does not propose a new model; its contribution is the organizing overview and the comparative table of reported results.

Load-bearing premise

The survey's reliability depends on the reported benchmark numbers in Table 3 being accurate and directly comparable, even though the models use different backbones, pretraining datasets, clip lengths, and evaluation protocols.

Editorial extensions

If this is right

  • A reader can use the three-way taxonomy to locate any video-understanding model and its structural design.
  • The survey shows a clear progression from image-extension models to two-stream networks to transformer-based models, with each step responding to the temporal dimension.
  • The comparison table gives a snapshot of reported performance on major benchmarks, even if direct comparability is limited.
  • The review identifies the open problems that would need solving for practical deployment: computational cost, dataset scale, input invariance, and online or hardware-efficient processing.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy likely generalizes beyond the surveyed period: newer video models such as masked autoencoders and video diffusion models still either pool frame features, split space and motion, or learn joint spatiotemporal representations.
  • A unified evaluation protocol, with fixed pretraining, clip length, and sampling, would turn the comparative table from a collection of reported numbers into a trustworthy ranking.
  • The emphasis on long-term dependencies suggests that clip-based benchmarks may underestimate model performance on long, real-world videos.
  • The dorsal and ventral stream inspiration behind two-stream architectures suggests that neuroscience-grounded multi-stream designs will remain a productive line of research.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper is a narrative survey of deep neural networks for video understanding, organized around a three-way taxonomy of spatial, separate spatial/temporal, and joint spatiotemporal feature extraction. It reviews preprocessing, temporal fusion, aggregation/pooling, attention and shifting, and then surveys model families including structural-search models, memory/recursive models, multi-stream networks, and transformer networks. It also lists major video datasets in Table 2 and gives a cross-model performance comparison in Table 3, followed by a discussion of open challenges such as computational cost, dataset scale, and input variance. The paper's stated goal, repeated in the introduction and conclusion, is to provide a holistic overview of video understanding models focused on spatiotemporal feature detection.

Significance. If the factual problems in the comparative material were corrected, the paper would be a serviceable entry-level narrative review: it covers the main model families and datasets, and a reader could use it to identify landmark architectures such as C3D, two-stream networks, I3D, SlowFast, and video transformers. The manuscript is honest in its overall style that reported numbers come from external sources, and it does not introduce new results or fitted parameters. However, the central value of a survey is accurate curation, and the current Table 3 contains several source attributions that are impossible given the reference dates, while the surrounding text presents the table as a comparison. Because the survey's utility rests on coverage and accuracy, these errors are load-bearing rather than cosmetic.

major comments (4)
  1. [Section 6, Table 3] The caption states "In all cases, the source literature provided the results," but this is contradicted by the reference list's own dates for at least three rows. Two-Stream [122] is Simonyan and Zisserman 2014, while Charades [121] is 2016, so the original two-stream paper cannot report a Charades mAP of 22.4. C3D [135] is from 2015, while Kinetics400 [70] is from 2017, and the original C3D paper does not use a VGG16 backbone or report 59.5% Kinetics accuracy. TSN [144] is ECCV 2016, before Kinetics400 existed, yet it is assigned a Kinetics400 number. These rows must be replaced with values from the actual later re-evaluations that reported them, with the proper citations, or deleted; otherwise the table's comparative claim is unsupported and actively misleading.
  2. [Section 6, Table 3] Even where a number could plausibly come from the cited paper, the table mixes results obtained under different protocols: backbones vary from AlexNet to ViT, pretraining includes ImageNet, Sports1M, Kinetics400, and Kinetics600, and the input modalities include RGB, optical flow, and compressed video. The text introduces the table as "comparison of main introduced structures," but no caveat states that the numbers are not directly comparable across rows. The authors should either add an explicit statement that each entry is the result reported under its source paper's protocol and should not be read as an apples-to-apples comparison, or restrict the table to a single standardized evaluation setting.
  3. [Section 6, text vs. Table 2] The paragraph above Table 2 says that dataset scale grew from 7K videos and 51 classes in HMDB51 to "over 6M videos and 3862 classes in Youtube-8M," while Table 2 itself lists YouTube-8M as 8M videos and 4800 classes. Reference [1] is the original YouTube-8M paper, which reports 8 million videos and 4800 classes. The text and table cannot both be right. This is not a minor typo because the paragraph uses the numbers as evidence of the field's scaling trend, and a reader relying on the survey for dataset facts will be misled.
  4. [Sections 5 and 6] The paper gives no methodology or inclusion criteria for selecting the models and datasets it reviews. The introduction promises "a holistic overview of the video understanding models," but the choice of which architectures appear in Section 5 and which rows appear in Table 3 is not justified; for example, several recent state-of-the-art video transformers are omitted while less central variants are included. For a survey, the authors should either state the scope and selection criteria explicitly (e.g., covering historically influential architectures rather than all recent methods) or temper the claim of holism. Without this, the selection appears arbitrary and the comparative value of Table 3 is further weakened.
minor comments (5)
  1. [References] The reference list heading is misspelled as "REFRENCES."
  2. [Abstract and Introduction] The prose in the abstract and opening paragraph is informal for a survey, including "It's no secret" and the incomplete sentence "It's a trend going to continue as video continues to dominate the digital landscape." A careful copyedit would improve readability.
  3. [Section 4.2] The assertion that "slow fusion closely resembles natural vision" is presented without supporting evidence beyond citation [12], which is about temporal image fusion in human vision but does not directly establish the claimed analogy to slow fusion in convolutional networks.
  4. [Reference [37]] The self-citation "In our recent research [37]" has no venue, arXiv identifier, or publication year in the reference list; this should be completed or the sentence removed.
  5. [Section 6, Table 2] The YouTube-8M row reports an average duration of 229.6 s, which is the total video length scale for that dataset; if this figure is the median or mean duration, the column heading "Ave. Duration" should be clarified so that readers do not confuse dataset-scale statistics with typical clip lengths used in evaluation.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a survey of externally published results, and its only self-citation is a non-load-bearing pointer.

full rationale

This is a narrative survey with no original derivations, fitted parameters, or predictive claims, so the standard circularity patterns do not apply. The one self-citation, [37] in Section 7.3, is a pointer to the authors' own earlier comparative study of multi-channel architectures and is used only to say that the authors have previously examined input variance; it is not load-bearing for any central claim, and no uniqueness theorem or ansatz is imported from it. The paper's comparative Table 3 attributes results to the source literature, and while the skeptic's concern that some rows cite papers that predate the datasets (e.g., Two-Stream [122] from 2014 cannot report a Charades mAP because Charades is from 2016, and C3D [135] from 2015 cannot report Kinetics400 accuracy) is a serious external-attribution and correctness problem, it is not circularity: the numbers are presented as coming from external benchmarks, not as outputs of a derivation defined in terms of the survey's own inputs. No equation is equated to another by construction, and no prediction is statistically forced by a fitted parameter. The inconsistency between the text's 'over 6M videos and 3862 classes' for YouTube-8M and Table 2's '8M videos and 4800 classes' is likewise an editing/accuracy issue, not a circular reduction. Under the hard rule that circularity requires quoting a specific reduction or a fitted parameter renamed as a prediction, none is present, so the appropriate finding is no significant circularity with score 0.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters are fitted in this review. The paper's tables copy published results, so the central 'derivation' (such as it is) is a selection and summarization of prior work. The assumptions are the trustworthiness and comparability of the cited results and the validity of the three-way feature-extraction taxonomy.

assumptions (2)
  • domain assumption Reported benchmark numbers in the cited literature are accurate and comparable.
    Table 3 reproduces results from prior papers and asserts 'the source literature provided the results'; the review's comparisons are only as good as those reports.
  • ad hoc to paper The taxonomy of approaches into spatial, temporal, and spatiotemporal feature extraction is a meaningful partition of the field.
    Section 3 presents this three-way partition as the organizing frame for the survey; it is the authors' framing rather than an externally established classification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis." pith.science (2026). https://pith.science/paper/5CASGN2C

@misc{pith2026250207277,
  author       = {Pith},
  title        = {Pith review of: Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/5CASGN2C}},
  note         = {Machine review of arXiv:2502.07277}
}
read the original abstract

It's no secret that video has become the primary way we share information online. That's why there's been a surge in demand for algorithms that can analyze and understand video content. It's a trend going to continue as video continues to dominate the digital landscape. These algorithms will extract and classify related features from the video and will use them to describe the events and objects in the video. Deep neural networks have displayed encouraging outcomes in the realm of feature extraction and video description. This paper will explore the spatiotemporal features found in videos and recent advancements in deep neural networks in video understanding. We will review some of the main trends in video understanding models and their structural design, the main problems, and some offered solutions in this topic. We will also review and compare significant video understanding and action recognition datasets.

Figures

Figures reproduced from arXiv: 2502.07277 by the authors.

Figure 5
Figure 5. 3-D convolution (a) and (2+1)-D convolution (b) model. Spatial dimensions, including width and height, denoted as x and y, and the temporal extent represented by t, characterize our filter. Figure adapted from [137] [135] introduced C3D as a generic video descriptor founded on 3-D convolutional networks. Through empirical evidence, the authors demonstrated that employing homogeneous 3x3x3 filters consistently outper… view at source ↗
Figure 8
Figure 8. The above figure illustrates pyramidal Fusion with Temporal Gaussian attention filters. In this fusion, multiple le [PITH_FULL_IMAGE:figures/full_fig_p009_8.png] view at source ↗
Figure 10
Figure 10. NetVLAD structure. The VLAD core will augment the input vectors and produce a vector representing them all. NetVLA [PITH_FULL_IMAGE:figures/full_fig_p011_10.png] view at source ↗
Figures from the paper (14 more)
Figure 11
Figure 11. Figure 11: Attention is applied to focus the network on the action. Based on Optical flow [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: The quantized temporal shift. Images (a) and (b) show the attention window. We simplify the attention on each chan [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: This figure illustrates the spatiotemporal shift in RubicksNet. The shift happens in all dimensions. Figure adapte [PITH_FULL_IMAGE:figures/full_fig_p012_13.png]
Figure 14
Figure 14. Figure 14: An example of context gating for a video about skiing. Snow and Skiing are more significant for action detection t [PITH_FULL_IMAGE:figures/full_fig_p013_14.png]
Figure 15
Figure 15. Figure 15: The WILLOW structure includes multiple streams for video and audio feature extractions. The augmented data with fu [PITH_FULL_IMAGE:figures/full_fig_p013_15.png]
Figure 16
Figure 16. Figure 16: The above image shows a 5-frame window that does not include any clues about the main action that is taking place. With five more frames, we can distinguish between the different sports of running and jumping and their variations. Memory-based network models, which ut…
Figure 17
Figure 17. Figure 17: General Recurrent Neural Network Structure for video understanding with memory blocks. Red blocks show convolution [PITH_FULL_IMAGE:figures/full_fig_p014_17.png]
Figure 18
Figure 18. Figure 18: Using feature banks alongside the spatiotemporal stream for better classification. Feature banks will map long [PITH_FULL_IMAGE:figures/full_fig_p015_18.png]
Figure 19
Figure 19. Figure 19: Dorsal and Ventral visual streams are two parallel pathways for natural vision. The visual dorsal is tasked with m [PITH_FULL_IMAGE:figures/full_fig_p015_19.png]
Figure 20
Figure 20. Figure 20: The Slowfast structure with slow and fast streams. The red arrow shows feedback from slow to fast stream. The blue [PITH_FULL_IMAGE:figures/full_fig_p016_20.png]
Figure 21
Figure 21. Figure 21: A Mixture of expert classifiers with an attention [PITH_FULL_IMAGE:figures/full_fig_p017_21.png]
Figure 22
Figure 22. Figure 22: Self-Attention Block. The triplet (Key, Query, Value) applies to input and will rewrite the values. Finally, Output projection (W) applies to the data to transform it. Figure adapted from [71] [PITH_FULL_IMAGE:figures/full_fig_p017_22.png]
Figure 23
Figure 23. Figure 23: Parallel self-attention blocks create a multi-head attention block. The results concat and are sent to the next layer. Figure adapted from [71] Research has established that multi-head self-attention when equipped with an adequate number of parameters, offers greater …
Figure 25
Figure 25. Figure 25: NetVLAD can classify correctly with variant inputs. The images are from the same location. However, the perspectiv [PITH_FULL_IMAGE:figures/full_fig_p023_25.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

176 extracted references · 52 canonical work pages

  1. [37]

    AmirHosein Fadaei and Mohammad-Reza A Dehaqani. 2023. Beyond Still Images: Robust Multi -Stream Spatiotemporal Networks

  2. [122]

    Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. Adv Neural Inf Process Syst 27, (2014)

  3. [121]

    Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11 –14, 2016, Proceedings, Part I 14, Springer, 510–526

  4. [135]

    Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3d c onvolutional networks. In Proceedings of the IEEE international conference on computer vision , 4489–4497

  5. [70]

    Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Tre vor Back, and Paul Natsev. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)

  6. [144]

    Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision, Springer, 20–36

  7. [1]

    Sami Abu -El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675 (2016)

  8. [2]

    Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. 2016. NetVLAD: CNN architecture for weakly supervised place recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 5297–5307

Show all 176 references
  1. [3]

    Relja Arandjelovic and Andrew Zisserman. 2013. All about VLAD. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 1578–1585

  2. [4]

    Farshid Arman, Arding Hsu, and Ming-Yee Chiu. 1997. Method for representing contents of a single video shot using frames

  3. [5]

    Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transf ormer. In Proceedings of the IEEE/CVF international conference on computer vision , 6836–6846

  4. [6]

    Moez Baccouche, Franck Mamalet, Christian Wolf, Christophe Garcia, and Atilla Baskurt. 2011. Sequential deep learning for hum an action recognition. In Human Behavior Understanding: Second International Workshop, HBU 2011, Amsterdam, The Netherlands, November 16, 2011. Proceed...

  5. [7]

    Tara Baldacchino, Elizabeth J Cross, Keith Worden, and Jennifer Rowson. 2016. Variational Bayesian mixture of experts models and sensitivity analysis for nonlinear dynamical systems. Mech Syst Signal Process 66, (2016), 178–200

  6. [8]

    Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. 2008. Speeded-up robust features (SURF). Computer vision and image understanding 110, 3 (2008), 346–359

  7. [9]

    Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le. 2018. Understanding and simplifying one -shot architecture search. In International conference on machine learning, PMLR, 550–559

  8. [10]

    Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In ICML, 4

  9. [11]

    Shweta Bhardwaj, Mukundhan Srinivasan, and Mitesh M Khapra. 2019. Efficient video classification using fewer frames. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 354–363

  10. [12]

    Hans Brettel, Lei Shi, and Hans Strasburger. 2006. Temporal image fusion in human vision. Vision Res 46, 6–7 (2006), 774–781

  11. [13]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, G irish Sastry, and Amanda Askell. 2020. Language models are few -shot learners. Adv Neural Inf Process Syst 33, (2020), 1877–1901

  12. [14]

    Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large -scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition , 961–970

  13. [15]

    John Canny. 1986. A computational approach to edge detection. IEEE Trans Pattern Anal Mach Intell 6 (1986), 679–698

  14. [16]

    Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu. 2019. Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE/CVF international conference on computer vision workshops , 0

  15. [17]

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End -to-end object detection with transformers. In European conference on computer vision, Springer, 213–229

  16. [18]

    Joao Carreira, Eric Noland, Andras Banki -Horvath, Chloe Hillier, and Andrew Zisserman. 2018. A short note about kinetics -600. arXiv preprint arXiv:1808.01340 (2018)

  17. [19]

    Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019. A short note on the kinetics -700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)

  18. [20]

    Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6299–6308

  19. [21]

    Yunpeng Chang, Zhigang Tu, Wei Xie, and Junsong Yuan. 2020. Clustering driven deep autoencoder for video anomaly detection. I n Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XV 16, Springer, 329–345

  20. [22]

    Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2021. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 12299–12310

  21. [23]

    Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. 2021. Pix2seq: A language modeling framework for object detection. arXiv preprint arXiv:2109.10852 (2021)

  22. [24]

    Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014)

  23. [25]

    Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Dav is, Afroz Mohiuddin, and Lukasz Kaiser. 2020. Rethinking attention with performers. arXiv preprint arXiv:2009.14794 (2020)

  24. [26]

    Cisco Visual Networking. 2019. Forecast and Trends, 2017–2022, White Paper c11-741490-00

  25. [27]

    Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2019. On the relationship between self -attention and convolutional layers. arXiv preprint arXiv:1911.03584 (2019)

  26. [28]

    Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. 2017. Deformable convolutional networks. In Proceedings of the IEEE international conference on computer vision, 764–773

  27. [29]

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, Ieee, 248–255

  28. [30]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)

  29. [31]

    Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jurgen Gall, Rainer Stiefelhagen, and Luc Van Gool. 2019. Holistic lar ge scale video understanding. arXiv preprint arXiv:1904.11451 38, 39 (2019), 9

  30. [32]

    Ali Diba, Vivek Sharma, and Luc Van Gool. 2017. Deep temporal linear encoding networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2329–2338

  31. [33]

    Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Dar rell. 2015. Long - term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer visi...

  32. [34]

    Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. 2022. Cswin transfor mer: A general vision transformer backbone with cross-shaped windows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  33. [35]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly. 2020. An image is worth 16x16 words: Transformers for image recognition a t scale. arXiv preprint a...

  34. [36]

    Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high -resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12873–12883

  35. [38]

    Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021. Multiscale vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision , 6824–6835

  36. [39]

    Linxi Fan, Shyamal Buch, Guanzhi Wang, Ryan Cao, Yuke Zhu, Juan Carlos Niebles, and Li Fei-Fei. 2020. Rubiksnet: Learnable 3d-shift for efficient video action recognition. In European Conference on Computer Vision, Springer, 505–521

  37. [40]

    Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. 2021. You only look at one sequence: Rethinking transformer in vision through object detection. Adv Neural Inf Process Syst 34, (2021), 26183–26197

  38. [41]

    William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research 23, 1 (2022), 5232–5270

  39. [42]

    Christoph Feichtenhofer. 2020. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 203–213

  40. [43]

    Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 6202–6211

  41. [44]

    Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes. 2017. Spatiotemporal multiplier networks for video action recogniti on. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4768–4777

  42. [45]

    Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional two -stream network fusion for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 1933–1941

  43. [46]

    Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell. 2017. Actionvlad: Learning spatio -temporal aggregation for action classification. In Proceedings of the IEEE conference on computer vision and pattern recognition , 971–980

  44. [47]

    Adam Golinski, Reza Pourreza, Yang Yang, Guillaume Sautiere, and Taco S Cohen. 2020. Feedback recurrent autoencoder for video compression. In Proceedings of the Asian Conference on Computer Vision

  45. [48]

    Melvyn A Goodale and A David Milner. 1992. Separate visual pathways for perception and action. Trends Neurosci 15, 1 (1992), 20–25

  46. [49]

    something something

    Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ing o Fruend, Peter Yianilos, and Moritz Mueller -Freitag. 2017. The" something something" video database for learning and evaluati ng visual common sense....

  47. [50]

    Alex Graves, Abdel -rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing , Ieee, 6645–6649

  48. [51]

    Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2016. LSTM: A search space odyssey. IEEE Trans Neural Netw Learn Syst 28, 10 (2016), 2222–2232

  49. [52]

    Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, and Rahul Sukthankar. 2018. Ava: A video dataset of spatio -temporally localized atomic visual actions. In Proceedings of the IEEE con...

  50. [53]

    Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi -Min Hu. 2021. Pct: Point cloud transformer. Comput Vis Media (Beijing) 7, (2021), 187–199

  51. [54]

    Zichao Guo, Xiangyu Zhang, Haoyuan Mu, Wen Heng, Zechun Liu, Yichen Wei, and Jian Sun. 2020. Single path one -shot neural architecture search with uniform sampling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XVI ...

  52. [55]

    Ghouthi Boukli Hacene, Carlos Lassance, Vincent Gripon, Matthieu Courbariaux, and Yoshua Bengio. 2021. Attention based prunin g for shift networks. In 2020 25th International Conference on Pattern Recognition (ICPR) , IEEE, 4054–4061

  53. [56]

    Chris Harris and Mike Stephens. 1988. A combined corner and edge detector. In Alvey vision conference, Citeseer, 10–5244

  54. [57]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778

  55. [58]

    Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019. Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180 (2019)

  56. [59]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short -term memory. Neural Comput 9, 8 (1997), 1735–1780

  57. [60]

    Karen Hollingsworth, Tanya Peters, Kevin W Bowyer, and Patrick J Flynn. 2009. Iris recognition using signal -level fusion of frames from video. IEEE Transactions on Information Forensics and Security 4, 4 (2009), 837–848

  58. [61]

    Berthold K P Horn and Brian G Schunck. 1981. Determining optical flow. Artif Intell 17, 1–3 (1981), 185–203

  59. [62]

    Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7132–7141

  60. [63]

    Wenbing Huang, Fuchun Sun, Lele Cao, Deli Zhao, Huaping Liu, and Mehrtash Harandi. 2016. Sparse coding and dictionary learning with linear dynamical systems. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 3938–3947

  61. [64]

    Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. 2019. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision , 603–612

  62. [65]

    Seong Jae Hwang, Joonseok Lee, Balakrishnan Varadarajan, Ariel Gordon, Zheng Xu, and Apostol Natsev. 2019. Large -scale training framework for video annotation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2394–2402

  63. [66]

    Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez. 2010. Aggregating local descriptors into a compact image rep resentation. In 2010 IEEE computer society conference on computer vision and pattern recognition , IEEE, 3304–3311

  64. [67]

    Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2012. 3D convolutional neural networks for human action recognition. IEEE Trans Pattern Anal Mach Intell 35, 1 (2012), 221–231

  65. [68]

    Samira Ebrahimi Kahou, Christopher Pal, Xavier Bouthillier, Pierre Froumenty, Çaglar Gülçehre, Roland Memisevic, Pascal Vince nt, Aaron Courville, Yoshua Bengio, and Raul Chandias Ferrari. 2013. Combining modality specific deep neural networks for emot ion recognition in video...

  66. [69]

    Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei -Fei. 2014. Large -scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 1725–1732

  67. [71]

    Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2022. Transformers in vi sion: A survey. ACM computing surveys (CSUR) 54, 10s (2022), 1–41

  68. [72]

    Saeed Reza Kheradpisheh, Masoud Ghodrati, Mohammad Ganjtabesh, and Timothée Masquelier. 2016. Deep networks can resemble huma n feed-forward vision in invariant object recognition. Sci Rep 6, 1 (2016), 32672

  69. [73]

    Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451 (2020)

  70. [74]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks . Adv Neural Inf Process Syst 25, (2012)

  71. [75]

    Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: a large video database for human motion recognition. In 2011 International conference on computer vision , IEEE, 2556–2563

  72. [76]

    Manoj Kumar, Dirk Weissenborn, and Nal Kalchbrenner. 2021. Colorization transformer. arXiv preprint arXiv:2102.04432 (2021)

  73. [77]

    Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2020. Hierarchical conditional relation networks for video questio n answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 9972–9981

  74. [78]

    Joonseok Lee, Walter Reade, Rahul Sukthankar, and George Toderici. 2018. The 2nd youtube-8m large-scale video understanding challenge. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0

  75. [79]

    Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for at tention-based permutation-invariant neural networks. In International conference on machine learning, PMLR, 3744–3753

  76. [80]

    Juho Lee, Yoonho Lee, and Yee Whye Teh. 2019. Deep amortized clustering. arXiv preprint arXiv:1909.13433 (2019)

  77. [81]

    Ang Li, Meghana Thotakuri , David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman. 2020. The ava -kinetics localized human actions video dataset. arXiv preprint arXiv:2005.00214 (2020)

  78. [82]

    Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. 2020. Tea: Temporal excitation and aggregation for action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 909–918

  79. [83]

    Yong Li, Yang Fu, Hui Li, and Si-Wen Zhang. 2009. The improved training algorithm of back propagation neural network with self -adaptive learning rate. In 2009 international conference on computational intelligence and natural computing , IEEE, 73–76

  80. [84]

    Kirt Lillywhite, Dah-Jye Lee, Beau Tippetts, and James Archibald. 2013. A feature construction method for general object recognition. Pattern Recognit 46, 12 (2013), 3300–3314

  81. [85]

    Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF international conference on computer vision, 7083–7093

  82. [86]

    Rongcheng Lin, Jing Xiao, and Jianping Fan. 2018. Nextvlad: An efficient neural network to aggregate frame -level features for large -scale video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0

  83. [87]

    Oskar Linde and Tony Lindeberg. 2004. Object recognition using composed receptive field histograms of higher dimensionality. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. , IEEE, 1–6

  84. [88]

    Tony Lindeberg. 2012. Scale invariant feature transform. (2012)

  85. [89]

    Tianqi Liu and Qizhan Shao. 2019. BERT for large-scale video segment classification with test-time augmentation. arXiv preprint arXiv:1912.01127 (2019)

  86. [90]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)

  87. [91]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchi cal vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision , 10012–10022

  88. [92]

    David G Lowe. 1999. Object recognition from local scale -invariant features. In Proceedings of the seventh IEEE international conference on computer vision, Ieee, 1150–1157

  89. [93]

    David G Lowe. 2004. Distinctive image features from scale -invariant keypoints. Int J Comput Vis 60, (2004), 91–110

  90. [94]

    Antoine Miech, Ivan Laptev, and Josef Sivic. 2017. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905 (2017)

  91. [95]

    Melanie Mitchell. 1998. An introduction to genetic algorithms. MIT press

  92. [96]

    Rakesh Mohan and Ramakant Nevatia. 1992. Perceptual organization for scene segmentation and description. IEEE Trans Pattern Anal Mach Intell 14, 06 (1992), 616–635

  93. [97]

    Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfr eund, and Carl Vondrick. 2019. Moments in time dataset: one million videos for event understanding. IEEE Trans Pattern Anal Mach Intell 42, 2 (2019), 502–508

  94. [98]

    Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. 2021. Video transformer network. In Proceedings of the IEEE/CVF international conference on computer vision, 3163–3172

  95. [99]

    Shinji Nishimoto, An T Vu, Thomas Naselaris, Yuval Benjamini, Bin Yu, and Jack L Gallant. 2011. Reconstructing visual experie nces from brain activity evoked by natural movies. Current biology 21, 19 (2011), 1641–1646

  96. [100]

    Bo Pang, Gao Peng, Yizhuo Li, and Cewu Lu. 2021. Pgt: A progressive method for training models on long videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11379–11389

  97. [101]

    Viorica Patraucean, Ankur Handa, and Roberto Cipolla. 2015. Spatio -temporal video autoencoder with differentiable memory. arXiv preprint arXiv:1511.06309 (2015)

  98. [102]

    Karl Pearson. 1901. LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin philosophical magazine and journal of science 2, 11 (1901), 559–572

  99. [103]

    Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. 2021. Random feature attention. arXiv preprint arXiv:2103.02143 (2021)

  100. [104]

    Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. 2018. Efficient neural architecture search via parameters sharin g. In International conference on machine learning, PMLR, 4095–4104

  101. [105]

    A Piergiovanni, Chenyou Fan, and Michael Ryoo. 2017. Learning latent subevents in activity videos using temporal attention fi lters. In Proceedings of the AAAI Conference on Artificial Intelligence

  102. [106]

    A J Piergiovanni and Michael Ryoo. 2020. Avid dataset: Anonymized videos from diverse countries. Adv Neural Inf Process Syst 33, (2020), 16711–16721

  103. [107]

    Zhaofan Qiu, Ting Yao, and Tao Mei. 2017. Learning spatio -temporal representation with pseudo -3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, 5533–5541

  104. [108]

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre -training. (2018)

  105. [109]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21, 1 (2020), 5485–5551

  106. [110]

    Amir M Rahimi, Kevin Lee, Amit Agarwal, Hyukseong Kwon, and Rajan Bhattacharyya. 2021. Toward Improving The Visual Characteri zation of Sport Activities With Abstracted Scene Graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 4500–4507

  107. [111]

    Michael S Ryoo, A J Piergiovanni, Juhana Kangaspunta, and Anelia Angelova. 2020. Assemblenet++: Assembling modality represent ations via attention connections. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XX 16 , S...

  108. [112]

    Michael S Ryoo, A J Piergiovanni, Mingxing Tan, and Anelia Angelova. 2019. Assemblenet: Searching for multi -stream neural connectivity in video architectures. arXiv preprint arXiv:1905.13209 (2019)

  109. [113]

    Mohammad Sabokrou, Mahmood Fathy, and Mojtaba Hoseini. 2016. Video anomaly detection and localisation based on the sparsity and reconstruction error of auto‐encoder. Electron Lett 52, 13 (2016), 1122–1124

  110. [114]

    Seyed Morteza Safdarnejad, Xiaoming Liu, Lalita Udpa, Brooks Andrus, John Wood, and Dean Craven. 2015. Sports videos in the w ild (svw): A video dataset for sports analysis. In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)...

  111. [115]

    Allah Bux Sargano, Plamen Angelov, and Zulfiqar Habib. 2017. A comprehensive review on handcrafted and learning -based action representation approaches for human activity recognition. applied sciences 7, 1 (2017), 110

  112. [116]

    Bernt Schiele and James L Crowley. 2000. Recognition without correspondence using multidimensional receptive field histograms . Int J Comput Vis 36, (2000), 31–50

  113. [117]

    Fadime Sener, Dipika Singhania, and Angela Yao. 2020. Temporal aggregate representations for long -range video understanding. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XVI 16, Springer, 154–171

  114. [118]

    Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. 2020. Finegym: A hierarchical video dataset for fine -grained action understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2616–2625

  115. [119]

    Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor. 2021. An image is worth 16x16 words, what is a video worth? arXiv preprint arXiv:2103.13915 (2021)

  116. [120]

    Shikhar Sharma, Ryan Kiros, and Ruslan Salakhutdinov. 2015. Action recognition using visual attention. arXiv preprint arXiv:1511.04119 (2015)

  117. [123]

    Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)

  118. [124]

    Jingkuan Song, Hanwang Zhang, Xiangpeng Li, Lianli Gao, Meng Wang, and Richang Hong. 2018. Self -supervised video hashing with hierarchical binary auto-encoder. IEEE Transactions on Image Processing 27, 7 (2018), 3210–3221

  119. [125]

    Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)

  120. [126]

    Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. 2021. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision , 7262–7272

  121. [127]

    Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. Videobert: A joint model for video and langua ge representation learning. In Proceedings of the IEEE/CVF international conference on computer vision , 7464–7473

  122. [128]

    Peter Sykora, Patrik Kamencay, Robert Hudec, Miroslav Benco, and Martin Sinko. 2018. Comparison of Feature Extraction Methods and Deep Learning Framework for Depth Map Recognition. In 2018 New Trends in Signal Processing (NTSP), IEEE, 1–7

  123. [129]

    Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015. Improved semantic representations from tree -structured long short -term memory networks. arXiv preprint arXiv:1503.00075 (2015)

  124. [130]

    Yi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei, and Xun Yang. 2021. Selective dependency aggregation for action classification. In Proceedings of the 29th ACM International Conference on Multimedia, 592–601

  125. [131]

    Yongyi Tang, Lin Ma, and Lianqiang Zhou. 2019. Hallucinating optical flow features for video classification. arXiv preprint arXiv:1905.11799 (2019)

  126. [132]

    Yongyi Tang, Xing Zhang, Lin Ma, Jingwen Wang, Shaoxiang Chen, and Yu -Gang Jiang. 2018. Non -local netvlad encoding for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0

  127. [133]

    Yi Tay, Dara Bahri, Donald Metzler, Da -Cheng Juan, Zhe Zhao, and Che Zheng. 2021. Synthesizer: Rethinking self -attention for transformer models. In International conference on machine learning, PMLR, 10183–10192

  128. [134]

    Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da -Cheng Juan. 2020. Sparse sinkhorn attention. In International Conference on Machine Learning , PMLR, 9438–9447

  129. [136]

    Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. 2019. Video classification with channel-separated convolutional networks. In Proceedings of the IEEE/CVF international conference on computer vision , 5552–5561

  130. [137]

    Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convo lutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 6450–6459

  131. [138]

    Gül Varol, Ivan Laptev, and Cordelia Schmid. 2017. Long -term temporal convolutions for action recognition. IEEE Trans Pattern Anal Mach Intell 40, 6 (2017), 1510–1517

  132. [139]

    Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. 2021. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 12894–12904

  133. [140]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin . 2017. Attention is all you need. Adv Neural Inf Process Syst 30, (2017)

  134. [141]

    Chunyu Wang, Yizhou Wang, and Alan L Yuille. 2013. An approach to pose -based action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 915–922

  135. [142]

    Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, and Song Han. 2020. Hat: Hardware -aware transformers for efficient natural language processing. arXiv preprint arXiv:2005.14187 (2020)

  136. [143]

    Limin Wang, Yu Qiao, and Xiaoou Tang. 2015. Action recognition with trajectory -pooled deep -convolutional descriptors. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4305–4314

  137. [145]

    Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self -attention with linear complexity. arXiv preprint arXiv:2006.04768 (2020)

  138. [146]

    Wenhai Wang, Enze Xie, Xiang Li, Deng -Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2022. Pvt v2: Improved baselines with pyramid vision transformer. Comput Vis Media (Beijing) 8, 3 (2022), 415–424

  139. [147]

    Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018. Non -local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7794–7803

  140. [148]

    Xinpeng Wang, Chandan Yeshwanth, and Matthias Nießner. 2021. Sceneformer: Indoor scene generation with transformers. In 2021 International Conference on 3D Vision (3DV), IEEE, 106–115

  141. [149]

    Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. 2022. Anchor detr: Query design for transformer -based detector. In Proceedings of the AAAI conference on artificial intelligence, 2567–2575

  142. [150]

    Thomas A Woolsey, Joseph Hanaway, and Mokhtar H Gado. 2017. The brain atlas: A visual guide to the human central nervous system. John Wiley & Sons

  143. [151]

    Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. 2019. Long-term feature banks for detailed video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 284–293

  144. [152]

    Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl. 2018. Compressed video action reco gnition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 6026–6035

  145. [153]

    Wenhao Wu, Dongliang He, Tianwei Lin, Fu Li, Chuang Gan, and Errui Ding. 2021. Mvfnet: Multi -view fusion network for efficient video recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , 2943–2951

  146. [154]

    Yu-Huan Wu, Yun Liu, Xin Zhan, and Ming-Ming Cheng. 2022. P2T: Pyramid pooling transformer for scene understanding. IEEE Trans Pattern Anal Mach Intell (2022)

  147. [155]

    Zuxuan Wu, Yu -Gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue. 2016. Multi -stream multi-class fusion of deep networks for video classification. In Proceedings of the 24th ACM international conference on Multimedia , 791–800

  148. [156]

    Bruce Xiaohan Nie, Caiming Xiong, and Song -Chun Zhu. 2015. Joint action recognition and pose estimation from video. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1293–1301

  149. [157]

    Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. 2021. SegFormer: Simple and efficient desi gn for semantic segmentation with transformers. Adv Neural Inf Process Syst 34, (2021), 12077–12090

  150. [158]

    Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018. Rethinking spatiotemporal feature learning: Speed -accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV) , 305–321

  151. [159]

    Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021. Nyströmformer: A nyström-based algorithm for approximating self-attention. In Proceedings of the AAAI Conference on Artificial Intelligence , 14138–14148

  152. [160]

    Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto. 2021. Long short-term transformer for online action detection. Adv Neural Inf Process Syst 34, (2021), 1086–1099

  153. [161]

    Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. 2020. Learning texture transformer network for image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 5791–5800

  154. [162]

    Ren Yang, Fabian Mentzer, Luc Van Gool, and Radu Timofte. 2020. Learning for video compression with recurrent auto -encoder and recurrent probability model. IEEE J Sel Top Signal Process 15, 2 (2020), 388–401

  155. [163]

    Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition , 4694–4702

  156. [164]

    Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12104–12113

  157. [165]

    Chuhan Zhang, Ankush Gupta, and Andrew Zisserman. 2021. Temporal query networks for fine -grained video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 4486–4496

  158. [166]

    Hao Zhang, Yanbin Hao, and Chong -Wah Ngo. 2021. Token shift transformer for video classification. In Proceedings of the 29th ACM International Conference on Multimedia, 917–925

  159. [167]

    Yujia Zhang, Xiaodan Liang, Dingwen Zhang, Min Tan, and Eric P Xing. 2020. Unsupervised object -level video summarization with online motion auto - encoder. Pattern Recognit Lett 130, (2020), 376–385

  160. [168]

    Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. 2019. Hacs: Human action clips and segments dataset for rec ognition and temporal localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision , 8668–8678

  161. [169]

    Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H S Torr, and Vladlen Koltun. 2021. Point transformer. In Proceedings of the IEEE/CVF international conference on computer vision, 16259–16268

  162. [170]

    Yiru Zhao, Bing Deng, Chen Shen, Yao Liu, Hongtao Lu, and Xian -Sheng Hua. 2017. Spatio -temporal autoencoder for video anomaly detection. In Proceedings of the 25th ACM international conference on Multimedia , 1933–1941

  163. [171]

    Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, and Phil ip H S Torr. 2021. Rethinking semantic segmentation from a sequence -to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conferen...

  164. [172]

    Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV), 803–818

  165. [173]

    Linchao Zhu, Du Tran, Laura Sevilla -Lara, Yi Yang, Matt Feiszli, and Heng Wang. 2020. Faster recurrent networks for efficient video classification. In Proceedings of the AAAI conference on artificial intelligence , 13098–13105

  166. [174]

    Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159 (2020)

  167. [175]

    Yi Zhu, Zhenzhong Lan, Shawn Newsam, and Alexander Hauptmann. 2019. Hidden two-stream convolutional networks for action recognition. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2 –6, 2018, Revised Selected Papers, Part III...

  168. [176]

    Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox. 2018. Eco: Efficient convolutional network for online video understanding. In Proceedings of the European conference on computer vision (ECCV), 695–712

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.