REVIEW 4 major objections 5 minor 176 references
Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis
T0 review · 4 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This survey organizes deep video-understanding models by how they handle time — extracting spatial features per frame, splitting space and motion into separate streams, or learning spatiotemporal features directly.
desk verdict A broad but careless survey: the narrative is usable, but Table 3's source attributions are impossible and the internal dataset numbers conflict, so the survey can't be trusted as a reference without major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the spatiotemporal feature, defined as information about both location and time that is most relevant for video understanding. The survey's taxonomy of three extraction methodologies, spatial-only, separate spatial-plus-temporal streams, and direct spatiotemporal extraction, carries the whole argument, with supporting concepts including the early, late, and slow temporal fusion types from [69], the two-stream architecture inspired by the brain's dorsal and ventral pathways, attention and shifting mechanisms, and aggregation methods such as NetVLAD. Each model family is presented as a way of instantiating one of the three extraction schemes.
What would settle it
Check the numbers in Table 3 against the original papers — for example, the Charades mAP for SlowFast and the Something-Something V1 accuracy for I3D — using the exact evaluation protocols reported there. If the numbers differ from the sources, or if the sources used incompatible protocols that the survey does not flag, the overview's comparisons would mislead.
Extended reading notes
Core claim
The paper's central claim is that video understanding models are best understood by how they treat the temporal dimension. It distinguishes three feature-extraction schemes: purely spatial extraction from individual frames, separate spatial and temporal streams, typically with optical flow or a trainable motion stream, and direct spatiotemporal extraction using 3D convolutions or other volume-level operations. The survey then uses this taxonomy to organize a broad review of structural designs, including temporal frame fusion, pooling and aggregation, attention and shifting, memory-based and recursive networks, multi-stream networks, and transformer models, and to frame the central challenges of the field. The paper does not propose a new model; its contribution is the organizing overview and the comparative table of reported results.
Load-bearing premise
The survey's reliability depends on the reported benchmark numbers in Table 3 being accurate and directly comparable, even though the models use different backbones, pretraining datasets, clip lengths, and evaluation protocols.
Editorial extensions
If this is right
- A reader can use the three-way taxonomy to locate any video-understanding model and its structural design.
- The survey shows a clear progression from image-extension models to two-stream networks to transformer-based models, with each step responding to the temporal dimension.
- The comparison table gives a snapshot of reported performance on major benchmarks, even if direct comparability is limited.
- The review identifies the open problems that would need solving for practical deployment: computational cost, dataset scale, input invariance, and online or hardware-efficient processing.
Reading between the lines
- The taxonomy likely generalizes beyond the surveyed period: newer video models such as masked autoencoders and video diffusion models still either pool frame features, split space and motion, or learn joint spatiotemporal representations.
- A unified evaluation protocol, with fixed pretraining, clip length, and sampling, would turn the comparative table from a collection of reported numbers into a trustworthy ranking.
- The emphasis on long-term dependencies suggests that clip-based benchmarks may underestimate model performance on long, real-world videos.
- The dorsal and ventral stream inspiration behind two-stream architectures suggests that neuroscience-grounded multi-stream designs will remain a productive line of research.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is a narrative survey of deep neural networks for video understanding, organized around a three-way taxonomy of spatial, separate spatial/temporal, and joint spatiotemporal feature extraction. It reviews preprocessing, temporal fusion, aggregation/pooling, attention and shifting, and then surveys model families including structural-search models, memory/recursive models, multi-stream networks, and transformer networks. It also lists major video datasets in Table 2 and gives a cross-model performance comparison in Table 3, followed by a discussion of open challenges such as computational cost, dataset scale, and input variance. The paper's stated goal, repeated in the introduction and conclusion, is to provide a holistic overview of video understanding models focused on spatiotemporal feature detection.
Significance. If the factual problems in the comparative material were corrected, the paper would be a serviceable entry-level narrative review: it covers the main model families and datasets, and a reader could use it to identify landmark architectures such as C3D, two-stream networks, I3D, SlowFast, and video transformers. The manuscript is honest in its overall style that reported numbers come from external sources, and it does not introduce new results or fitted parameters. However, the central value of a survey is accurate curation, and the current Table 3 contains several source attributions that are impossible given the reference dates, while the surrounding text presents the table as a comparison. Because the survey's utility rests on coverage and accuracy, these errors are load-bearing rather than cosmetic.
major comments (4)
- [Section 6, Table 3] The caption states "In all cases, the source literature provided the results," but this is contradicted by the reference list's own dates for at least three rows. Two-Stream [122] is Simonyan and Zisserman 2014, while Charades [121] is 2016, so the original two-stream paper cannot report a Charades mAP of 22.4. C3D [135] is from 2015, while Kinetics400 [70] is from 2017, and the original C3D paper does not use a VGG16 backbone or report 59.5% Kinetics accuracy. TSN [144] is ECCV 2016, before Kinetics400 existed, yet it is assigned a Kinetics400 number. These rows must be replaced with values from the actual later re-evaluations that reported them, with the proper citations, or deleted; otherwise the table's comparative claim is unsupported and actively misleading.
- [Section 6, Table 3] Even where a number could plausibly come from the cited paper, the table mixes results obtained under different protocols: backbones vary from AlexNet to ViT, pretraining includes ImageNet, Sports1M, Kinetics400, and Kinetics600, and the input modalities include RGB, optical flow, and compressed video. The text introduces the table as "comparison of main introduced structures," but no caveat states that the numbers are not directly comparable across rows. The authors should either add an explicit statement that each entry is the result reported under its source paper's protocol and should not be read as an apples-to-apples comparison, or restrict the table to a single standardized evaluation setting.
- [Section 6, text vs. Table 2] The paragraph above Table 2 says that dataset scale grew from 7K videos and 51 classes in HMDB51 to "over 6M videos and 3862 classes in Youtube-8M," while Table 2 itself lists YouTube-8M as 8M videos and 4800 classes. Reference [1] is the original YouTube-8M paper, which reports 8 million videos and 4800 classes. The text and table cannot both be right. This is not a minor typo because the paragraph uses the numbers as evidence of the field's scaling trend, and a reader relying on the survey for dataset facts will be misled.
- [Sections 5 and 6] The paper gives no methodology or inclusion criteria for selecting the models and datasets it reviews. The introduction promises "a holistic overview of the video understanding models," but the choice of which architectures appear in Section 5 and which rows appear in Table 3 is not justified; for example, several recent state-of-the-art video transformers are omitted while less central variants are included. For a survey, the authors should either state the scope and selection criteria explicitly (e.g., covering historically influential architectures rather than all recent methods) or temper the claim of holism. Without this, the selection appears arbitrary and the comparative value of Table 3 is further weakened.
minor comments (5)
- [References] The reference list heading is misspelled as "REFRENCES."
- [Abstract and Introduction] The prose in the abstract and opening paragraph is informal for a survey, including "It's no secret" and the incomplete sentence "It's a trend going to continue as video continues to dominate the digital landscape." A careful copyedit would improve readability.
- [Section 4.2] The assertion that "slow fusion closely resembles natural vision" is presented without supporting evidence beyond citation [12], which is about temporal image fusion in human vision but does not directly establish the claimed analogy to slow fusion in convolutional networks.
- [Reference [37]] The self-citation "In our recent research [37]" has no venue, arXiv identifier, or publication year in the reference list; this should be completed or the sentence removed.
- [Section 6, Table 2] The YouTube-8M row reports an average duration of 229.6 s, which is the total video length scale for that dataset; if this figure is the median or mean duration, the column heading "Ave. Duration" should be clarified so that readers do not confuse dataset-scale statistics with typical clip lengths used in evaluation.
Circularity Check
No significant circularity: the paper is a survey of externally published results, and its only self-citation is a non-load-bearing pointer.
full rationale
This is a narrative survey with no original derivations, fitted parameters, or predictive claims, so the standard circularity patterns do not apply. The one self-citation, [37] in Section 7.3, is a pointer to the authors' own earlier comparative study of multi-channel architectures and is used only to say that the authors have previously examined input variance; it is not load-bearing for any central claim, and no uniqueness theorem or ansatz is imported from it. The paper's comparative Table 3 attributes results to the source literature, and while the skeptic's concern that some rows cite papers that predate the datasets (e.g., Two-Stream [122] from 2014 cannot report a Charades mAP because Charades is from 2016, and C3D [135] from 2015 cannot report Kinetics400 accuracy) is a serious external-attribution and correctness problem, it is not circularity: the numbers are presented as coming from external benchmarks, not as outputs of a derivation defined in terms of the survey's own inputs. No equation is equated to another by construction, and no prediction is statistically forced by a fitted parameter. The inconsistency between the text's 'over 6M videos and 3862 classes' for YouTube-8M and Table 2's '8M videos and 4800 classes' is likewise an editing/accuracy issue, not a circular reduction. Under the hard rule that circularity requires quoting a specific reduction or a fitted parameter renamed as a prediction, none is present, so the appropriate finding is no significant circularity with score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption Reported benchmark numbers in the cited literature are accurate and comparable.
- ad hoc to paper The taxonomy of approaches into spatial, temporal, and spatiotemporal feature extraction is a meaningful partition of the field.
Cite this review
Pith. "Pith review of Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis." pith.science (2026). https://pith.science/paper/5CASGN2C
@misc{pith2026250207277,
author = {Pith},
title = {Pith review of: Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis},
year = {2026},
howpublished = {\url{https://pith.science/paper/5CASGN2C}},
note = {Machine review of arXiv:2502.07277}
}
read the original abstract
It's no secret that video has become the primary way we share information online. That's why there's been a surge in demand for algorithms that can analyze and understand video content. It's a trend going to continue as video continues to dominate the digital landscape. These algorithms will extract and classify related features from the video and will use them to describe the events and objects in the video. Deep neural networks have displayed encouraging outcomes in the realm of feature extraction and video description. This paper will explore the spatiotemporal features found in videos and recent advancements in deep neural networks in video understanding. We will review some of the main trends in video understanding models and their structural design, the main problems, and some offered solutions in this topic. We will also review and compare significant video understanding and action recognition datasets.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[37]
AmirHosein Fadaei and Mohammad-Reza A Dehaqani. 2023. Beyond Still Images: Robust Multi -Stream Spatiotemporal Networks
2023
-
[122]
Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. Adv Neural Inf Process Syst 27, (2014)
work page 2014
-
[121]
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11 –14, 2016, Proceedings, Part I 14, Springer, 510–526
work page 2016
-
[135]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning spatiotemporal features with 3d c onvolutional networks. In Proceedings of the IEEE international conference on computer vision , 4489–4497
work page 2015
-
[70]
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Tre vor Back, and Paul Natsev. 2017. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950 (2017)
arXiv 2017
-
[144]
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In European conference on computer vision, Springer, 20–36
work page 2016
-
[1]
Sami Abu -El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016. Youtube-8m: A large-scale video classification benchmark. arXiv preprint arXiv:1609.08675 (2016)
arXiv 2016
-
[2]
Relja Arandjelovic, Petr Gronat, Akihiko Torii, Tomas Pajdla, and Josef Sivic. 2016. NetVLAD: CNN architecture for weakly supervised place recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 5297–5307
2016
Show all 176 references
-
[3]
Relja Arandjelovic and Andrew Zisserman. 2013. All about VLAD. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 1578–1585
2013
-
[4]
Farshid Arman, Arding Hsu, and Ming-Yee Chiu. 1997. Method for representing contents of a single video shot using frames
1997
-
[5]
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transf ormer. In Proceedings of the IEEE/CVF international conference on computer vision , 6836–6846
2021
-
[6]
Moez Baccouche, Franck Mamalet, Christian Wolf, Christophe Garcia, and Atilla Baskurt. 2011. Sequential deep learning for hum an action recognition. In Human Behavior Understanding: Second International Workshop, HBU 2011, Amsterdam, The Netherlands, November 16, 2011. Proceed...
2011
-
[7]
Tara Baldacchino, Elizabeth J Cross, Keith Worden, and Jennifer Rowson. 2016. Variational Bayesian mixture of experts models and sensitivity analysis for nonlinear dynamical systems. Mech Syst Signal Process 66, (2016), 178–200
2016
-
[8]
Herbert Bay, Andreas Ess, Tinne Tuytelaars, and Luc Van Gool. 2008. Speeded-up robust features (SURF). Computer vision and image understanding 110, 3 (2008), 346–359
2008
-
[9]
Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le. 2018. Understanding and simplifying one -shot architecture search. In International conference on machine learning, PMLR, 550–559
2018
-
[10]
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding? In ICML, 4
2021
-
[11]
Shweta Bhardwaj, Mukundhan Srinivasan, and Mitesh M Khapra. 2019. Efficient video classification using fewer frames. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 354–363
2019
-
[12]
Hans Brettel, Lei Shi, and Hans Strasburger. 2006. Temporal image fusion in human vision. Vision Res 46, 6–7 (2006), 774–781
2006
-
[13]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, G irish Sastry, and Amanda Askell. 2020. Language models are few -shot learners. Adv Neural Inf Process Syst 33, (2020), 1877–1901
2020
-
[14]
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. Activitynet: A large -scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition , 961–970
2015
-
[15]
John Canny. 1986. A computational approach to edge detection. IEEE Trans Pattern Anal Mach Intell 6 (1986), 679–698
1986
-
[16]
Yue Cao, Jiarui Xu, Stephen Lin, Fangyun Wei, and Han Hu. 2019. Gcnet: Non-local networks meet squeeze-excitation networks and beyond. In Proceedings of the IEEE/CVF international conference on computer vision workshops , 0
2019
-
[17]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End -to-end object detection with transformers. In European conference on computer vision, Springer, 213–229
2020
-
[18]
Joao Carreira, Eric Noland, Andras Banki -Horvath, Chloe Hillier, and Andrew Zisserman. 2018. A short note about kinetics -600. arXiv preprint arXiv:1808.01340 (2018)
2018 arXiv
-
[19]
Joao Carreira, Eric Noland, Chloe Hillier, and Andrew Zisserman. 2019. A short note on the kinetics -700 human action dataset. arXiv preprint arXiv:1907.06987 (2019)
2019 arXiv
-
[20]
Joao Carreira and Andrew Zisserman. 2017. Quo vadis, action recognition? a new model and the kinetics dataset. In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6299–6308
2017
-
[21]
Yunpeng Chang, Zhigang Tu, Wei Xie, and Junsong Yuan. 2020. Clustering driven deep autoencoder for video anomaly detection. I n Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XV 16, Springer, 329–345
2020
-
[22]
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2021. Pre-trained image processing transformer. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 12299–12310
2021
-
[23]
Ting Chen, Saurabh Saxena, Lala Li, David J Fleet, and Geoffrey Hinton. 2021. Pix2seq: A language modeling framework for object detection. arXiv preprint arXiv:2109.10852 (2021)
2021 arXiv
-
[24]
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv preprint arXiv:1406.1078 (2014)
2014 arXiv
-
[25]
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Dav is, Afroz Mohiuddin, and Lukasz Kaiser. 2020. Rethinking attention with performers. arXiv preprint arXiv:2009.14794 (2020)
2020 arXiv
-
[26]
Cisco Visual Networking. 2019. Forecast and Trends, 2017–2022, White Paper c11-741490-00
2019
-
[27]
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2019. On the relationship between self -attention and convolutional layers. arXiv preprint arXiv:1911.03584 (2019)
2019 arXiv
-
[28]
Jifeng Dai, Haozhi Qi, Yuwen Xiong, Yi Li, Guodong Zhang, Han Hu, and Yichen Wei. 2017. Deformable convolutional networks. In Proceedings of the IEEE international conference on computer vision, 764–773
2017
-
[29]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, Ieee, 248–255
2009
-
[30]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[31]
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jurgen Gall, Rainer Stiefelhagen, and Luc Van Gool. 2019. Holistic lar ge scale video understanding. arXiv preprint arXiv:1904.11451 38, 39 (2019), 9
2019 arXiv
-
[32]
Ali Diba, Vivek Sharma, and Luc Van Gool. 2017. Deep temporal linear encoding networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, 2329–2338
2017
-
[33]
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Dar rell. 2015. Long - term recurrent convolutional networks for visual recognition and description. In Proceedings of the IEEE conference on computer visi...
2015
-
[34]
Xiaoyi Dong, Jianmin Bao, Dongdong Chen, Weiming Zhang, Nenghai Yu, Lu Yuan, Dong Chen, and Baining Guo. 2022. Cswin transfor mer: A general vision transformer backbone with cross-shaped windows. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2022
-
[35]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, and Sylvain Gelly. 2020. An image is worth 16x16 words: Transformers for image recognition a t scale. arXiv preprint a...
2020 arXiv
-
[36]
Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high -resolution image synthesis. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12873–12883
2021
-
[38]
Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021. Multiscale vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision , 6824–6835
2021
-
[39]
Linxi Fan, Shyamal Buch, Guanzhi Wang, Ryan Cao, Yuke Zhu, Juan Carlos Niebles, and Li Fei-Fei. 2020. Rubiksnet: Learnable 3d-shift for efficient video action recognition. In European Conference on Computer Vision, Springer, 505–521
2020
-
[40]
Yuxin Fang, Bencheng Liao, Xinggang Wang, Jiemin Fang, Jiyang Qi, Rui Wu, Jianwei Niu, and Wenyu Liu. 2021. You only look at one sequence: Rethinking transformer in vision through object detection. Adv Neural Inf Process Syst 34, (2021), 26183–26197
2021
-
[41]
William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. The Journal of Machine Learning Research 23, 1 (2022), 5232–5270
2022
-
[42]
Christoph Feichtenhofer. 2020. X3d: Expanding architectures for efficient video recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 203–213
2020
-
[43]
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 6202–6211
2019
-
[44]
Christoph Feichtenhofer, Axel Pinz, and Richard P Wildes. 2017. Spatiotemporal multiplier networks for video action recogniti on. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4768–4777
2017
-
[45]
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional two -stream network fusion for video action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 1933–1941
2016
-
[46]
Rohit Girdhar, Deva Ramanan, Abhinav Gupta, Josef Sivic, and Bryan Russell. 2017. Actionvlad: Learning spatio -temporal aggregation for action classification. In Proceedings of the IEEE conference on computer vision and pattern recognition , 971–980
2017
-
[47]
Adam Golinski, Reza Pourreza, Yang Yang, Guillaume Sautiere, and Taco S Cohen. 2020. Feedback recurrent autoencoder for video compression. In Proceedings of the Asian Conference on Computer Vision
2020
-
[48]
Melvyn A Goodale and A David Milner. 1992. Separate visual pathways for perception and action. Trends Neurosci 15, 1 (1992), 20–25
1992
-
[49]
something something
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ing o Fruend, Peter Yianilos, and Moritz Mueller -Freitag. 2017. The" something something" video database for learning and evaluati ng visual common sense....
2017
-
[50]
Alex Graves, Abdel -rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In 2013 IEEE international conference on acoustics, speech and signal processing , Ieee, 6645–6649
2013
-
[51]
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2016. LSTM: A search space odyssey. IEEE Trans Neural Netw Learn Syst 28, 10 (2016), 2222–2232
2016
-
[52]
Chunhui Gu, Chen Sun, David A Ross, Carl Vondrick, Caroline Pantofaru, Yeqing Li, Sudheendra Vijayanarasimhan, George Toderici, Susanna Ricco, and Rahul Sukthankar. 2018. Ava: A video dataset of spatio -temporally localized atomic visual actions. In Proceedings of the IEEE con...
2018
-
[53]
Meng-Hao Guo, Jun-Xiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R Martin, and Shi -Min Hu. 2021. Pct: Point cloud transformer. Comput Vis Media (Beijing) 7, (2021), 187–199
2021
-
[54]
Zichao Guo, Xiangyu Zhang, Haoyuan Mu, Wen Heng, Zechun Liu, Yichen Wei, and Jian Sun. 2020. Single path one -shot neural architecture search with uniform sampling. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XVI ...
2020
-
[55]
Ghouthi Boukli Hacene, Carlos Lassance, Vincent Gripon, Matthieu Courbariaux, and Yoshua Bengio. 2021. Attention based prunin g for shift networks. In 2020 25th International Conference on Pattern Recognition (ICPR) , IEEE, 4054–4061
2021
-
[56]
Chris Harris and Mike Stephens. 1988. A combined corner and edge detector. In Alvey vision conference, Citeseer, 10–5244
1988
-
[57]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 770–778
2016
-
[58]
Jonathan Ho, Nal Kalchbrenner, Dirk Weissenborn, and Tim Salimans. 2019. Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180 (2019)
2019 arXiv
-
[59]
Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short -term memory. Neural Comput 9, 8 (1997), 1735–1780
1997
-
[60]
Karen Hollingsworth, Tanya Peters, Kevin W Bowyer, and Patrick J Flynn. 2009. Iris recognition using signal -level fusion of frames from video. IEEE Transactions on Information Forensics and Security 4, 4 (2009), 837–848
2009
-
[61]
Berthold K P Horn and Brian G Schunck. 1981. Determining optical flow. Artif Intell 17, 1–3 (1981), 185–203
1981
-
[62]
Jie Hu, Li Shen, and Gang Sun. 2018. Squeeze-and-excitation networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7132–7141
2018
-
[63]
Wenbing Huang, Fuchun Sun, Lele Cao, Deli Zhao, Huaping Liu, and Mehrtash Harandi. 2016. Sparse coding and dictionary learning with linear dynamical systems. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 3938–3947
2016
-
[64]
Zilong Huang, Xinggang Wang, Lichao Huang, Chang Huang, Yunchao Wei, and Wenyu Liu. 2019. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision , 603–612
2019
-
[65]
Seong Jae Hwang, Joonseok Lee, Balakrishnan Varadarajan, Ariel Gordon, Zheng Xu, and Apostol Natsev. 2019. Large -scale training framework for video annotation. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2394–2402
2019
-
[66]
Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez. 2010. Aggregating local descriptors into a compact image rep resentation. In 2010 IEEE computer society conference on computer vision and pattern recognition , IEEE, 3304–3311
2010
-
[67]
Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2012. 3D convolutional neural networks for human action recognition. IEEE Trans Pattern Anal Mach Intell 35, 1 (2012), 221–231
2012
-
[68]
Samira Ebrahimi Kahou, Christopher Pal, Xavier Bouthillier, Pierre Froumenty, Çaglar Gülçehre, Roland Memisevic, Pascal Vince nt, Aaron Courville, Yoshua Bengio, and Raul Chandias Ferrari. 2013. Combining modality specific deep neural networks for emot ion recognition in video...
2013
-
[69]
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei -Fei. 2014. Large -scale video classification with convolutional neural networks. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 1725–1732
2014
-
[71]
Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, and Mubarak Shah. 2022. Transformers in vi sion: A survey. ACM computing surveys (CSUR) 54, 10s (2022), 1–41
2022
-
[72]
Saeed Reza Kheradpisheh, Masoud Ghodrati, Mohammad Ganjtabesh, and Timothée Masquelier. 2016. Deep networks can resemble huma n feed-forward vision in invariant object recognition. Sci Rep 6, 1 (2016), 32672
2016
-
[73]
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. arXiv preprint arXiv:2001.04451 (2020)
2020 arXiv
-
[74]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks . Adv Neural Inf Process Syst 25, (2012)
2012
-
[75]
Hildegard Kuehne, Hueihan Jhuang, Estíbaliz Garrote, Tomaso Poggio, and Thomas Serre. 2011. HMDB: a large video database for human motion recognition. In 2011 International conference on computer vision , IEEE, 2556–2563
2011
-
[76]
Manoj Kumar, Dirk Weissenborn, and Nal Kalchbrenner. 2021. Colorization transformer. arXiv preprint arXiv:2102.04432 (2021)
2021 arXiv
-
[77]
Thao Minh Le, Vuong Le, Svetha Venkatesh, and Truyen Tran. 2020. Hierarchical conditional relation networks for video questio n answering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 9972–9981
2020
-
[78]
Joonseok Lee, Walter Reade, Rahul Sukthankar, and George Toderici. 2018. The 2nd youtube-8m large-scale video understanding challenge. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0
2018
-
[79]
Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, and Yee Whye Teh. 2019. Set transformer: A framework for at tention-based permutation-invariant neural networks. In International conference on machine learning, PMLR, 3744–3753
2019
-
[80]
Juho Lee, Yoonho Lee, and Yee Whye Teh. 2019. Deep amortized clustering. arXiv preprint arXiv:1909.13433 (2019)
2019 arXiv
-
[81]
Ang Li, Meghana Thotakuri , David A Ross, João Carreira, Alexander Vostrikov, and Andrew Zisserman. 2020. The ava -kinetics localized human actions video dataset. arXiv preprint arXiv:2005.00214 (2020)
2020 arXiv
-
[82]
Yan Li, Bin Ji, Xintian Shi, Jianguo Zhang, Bin Kang, and Limin Wang. 2020. Tea: Temporal excitation and aggregation for action recognition. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 909–918
2020
-
[83]
Yong Li, Yang Fu, Hui Li, and Si-Wen Zhang. 2009. The improved training algorithm of back propagation neural network with self -adaptive learning rate. In 2009 international conference on computational intelligence and natural computing , IEEE, 73–76
2009
-
[84]
Kirt Lillywhite, Dah-Jye Lee, Beau Tippetts, and James Archibald. 2013. A feature construction method for general object recognition. Pattern Recognit 46, 12 (2013), 3300–3314
2013
-
[85]
Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. In Proceedings of the IEEE/CVF international conference on computer vision, 7083–7093
2019
-
[86]
Rongcheng Lin, Jing Xiao, and Jianping Fan. 2018. Nextvlad: An efficient neural network to aggregate frame -level features for large -scale video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0
2018
-
[87]
Oskar Linde and Tony Lindeberg. 2004. Object recognition using composed receptive field histograms of higher dimensionality. In Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. , IEEE, 1–6
2004
-
[88]
Tony Lindeberg. 2012. Scale invariant feature transform. (2012)
2012
-
[89]
Tianqi Liu and Qizhan Shao. 2019. BERT for large-scale video segment classification with test-time augmentation. arXiv preprint arXiv:1912.01127 (2019)
2019 arXiv
-
[90]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692 (2019)
2019 arXiv
-
[91]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchi cal vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision , 10012–10022
2021
-
[92]
David G Lowe. 1999. Object recognition from local scale -invariant features. In Proceedings of the seventh IEEE international conference on computer vision, Ieee, 1150–1157
1999
-
[93]
David G Lowe. 2004. Distinctive image features from scale -invariant keypoints. Int J Comput Vis 60, (2004), 91–110
2004
-
[94]
Antoine Miech, Ivan Laptev, and Josef Sivic. 2017. Learnable pooling with context gating for video classification. arXiv preprint arXiv:1706.06905 (2017)
2017 arXiv
-
[95]
Melanie Mitchell. 1998. An introduction to genetic algorithms. MIT press
1998
-
[96]
Rakesh Mohan and Ramakant Nevatia. 1992. Perceptual organization for scene segmentation and description. IEEE Trans Pattern Anal Mach Intell 14, 06 (1992), 616–635
1992
-
[97]
Mathew Monfort, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa Brown, Quanfu Fan, Dan Gutfr eund, and Carl Vondrick. 2019. Moments in time dataset: one million videos for event understanding. IEEE Trans Pattern Anal Mach Intell 42, 2 (2019), 502–508
2019
-
[98]
Daniel Neimark, Omri Bar, Maya Zohar, and Dotan Asselmann. 2021. Video transformer network. In Proceedings of the IEEE/CVF international conference on computer vision, 3163–3172
2021
-
[99]
Shinji Nishimoto, An T Vu, Thomas Naselaris, Yuval Benjamini, Bin Yu, and Jack L Gallant. 2011. Reconstructing visual experie nces from brain activity evoked by natural movies. Current biology 21, 19 (2011), 1641–1646
2011
-
[100]
Bo Pang, Gao Peng, Yizhuo Li, and Cewu Lu. 2021. Pgt: A progressive method for training models on long videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 11379–11389
2021
-
[101]
Viorica Patraucean, Ankur Handa, and Roberto Cipolla. 2015. Spatio -temporal video autoencoder with differentiable memory. arXiv preprint arXiv:1511.06309 (2015)
2015 arXiv
-
[102]
Karl Pearson. 1901. LIII. On lines and planes of closest fit to systems of points in space. The London, Edinburgh, and Dublin philosophical magazine and journal of science 2, 11 (1901), 559–572
1901
-
[103]
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A Smith, and Lingpeng Kong. 2021. Random feature attention. arXiv preprint arXiv:2103.02143 (2021)
2021 arXiv
-
[104]
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. 2018. Efficient neural architecture search via parameters sharin g. In International conference on machine learning, PMLR, 4095–4104
2018
-
[105]
A Piergiovanni, Chenyou Fan, and Michael Ryoo. 2017. Learning latent subevents in activity videos using temporal attention fi lters. In Proceedings of the AAAI Conference on Artificial Intelligence
2017
-
[106]
A J Piergiovanni and Michael Ryoo. 2020. Avid dataset: Anonymized videos from diverse countries. Adv Neural Inf Process Syst 33, (2020), 16711–16721
2020
-
[107]
Zhaofan Qiu, Ting Yao, and Tao Mei. 2017. Learning spatio -temporal representation with pseudo -3d residual networks. In proceedings of the IEEE International Conference on Computer Vision, 5533–5541
2017
-
[108]
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018. Improving language understanding by generative pre -training. (2018)
2018
-
[109]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research 21, 1 (2020), 5485–5551
2020
-
[110]
Amir M Rahimi, Kevin Lee, Amit Agarwal, Hyukseong Kwon, and Rajan Bhattacharyya. 2021. Toward Improving The Visual Characteri zation of Sport Activities With Abstracted Scene Graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 4500–4507
2021
-
[111]
Michael S Ryoo, A J Piergiovanni, Juhana Kangaspunta, and Anelia Angelova. 2020. Assemblenet++: Assembling modality represent ations via attention connections. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XX 16 , S...
2020
-
[112]
Michael S Ryoo, A J Piergiovanni, Mingxing Tan, and Anelia Angelova. 2019. Assemblenet: Searching for multi -stream neural connectivity in video architectures. arXiv preprint arXiv:1905.13209 (2019)
2019 arXiv
-
[113]
Mohammad Sabokrou, Mahmood Fathy, and Mojtaba Hoseini. 2016. Video anomaly detection and localisation based on the sparsity and reconstruction error of auto‐encoder. Electron Lett 52, 13 (2016), 1122–1124
2016
-
[114]
Seyed Morteza Safdarnejad, Xiaoming Liu, Lalita Udpa, Brooks Andrus, John Wood, and Dean Craven. 2015. Sports videos in the w ild (svw): A video dataset for sports analysis. In 2015 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)...
2015
-
[115]
Allah Bux Sargano, Plamen Angelov, and Zulfiqar Habib. 2017. A comprehensive review on handcrafted and learning -based action representation approaches for human activity recognition. applied sciences 7, 1 (2017), 110
2017
-
[116]
Bernt Schiele and James L Crowley. 2000. Recognition without correspondence using multidimensional receptive field histograms . Int J Comput Vis 36, (2000), 31–50
2000
-
[117]
Fadime Sener, Dipika Singhania, and Angela Yao. 2020. Temporal aggregate representations for long -range video understanding. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23 –28, 2020, Proceedings, Part XVI 16, Springer, 154–171
2020
-
[118]
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. 2020. Finegym: A hierarchical video dataset for fine -grained action understanding. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2616–2625
2020
-
[119]
Gilad Sharir, Asaf Noy, and Lihi Zelnik-Manor. 2021. An image is worth 16x16 words, what is a video worth? arXiv preprint arXiv:2103.13915 (2021)
2021 arXiv
-
[120]
Shikhar Sharma, Ryan Kiros, and Ruslan Salakhutdinov. 2015. Action recognition using visual attention. arXiv preprint arXiv:1511.04119 (2015)
2015 arXiv
-
[123]
Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)
2014 arXiv
-
[124]
Jingkuan Song, Hanwang Zhang, Xiangpeng Li, Lianli Gao, Meng Wang, and Richang Hong. 2018. Self -supervised video hashing with hierarchical binary auto-encoder. IEEE Transactions on Image Processing 27, 7 (2018), 3210–3221
2018
-
[125]
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
2012 arXiv
-
[126]
Robin Strudel, Ricardo Garcia, Ivan Laptev, and Cordelia Schmid. 2021. Segmenter: Transformer for semantic segmentation. In Proceedings of the IEEE/CVF international conference on computer vision , 7262–7272
2021
-
[127]
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019. Videobert: A joint model for video and langua ge representation learning. In Proceedings of the IEEE/CVF international conference on computer vision , 7464–7473
2019
-
[128]
Peter Sykora, Patrik Kamencay, Robert Hudec, Miroslav Benco, and Martin Sinko. 2018. Comparison of Feature Extraction Methods and Deep Learning Framework for Depth Map Recognition. In 2018 New Trends in Signal Processing (NTSP), IEEE, 1–7
2018
-
[129]
Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015. Improved semantic representations from tree -structured long short -term memory networks. arXiv preprint arXiv:1503.00075 (2015)
2015 arXiv
-
[130]
Yi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei, and Xun Yang. 2021. Selective dependency aggregation for action classification. In Proceedings of the 29th ACM International Conference on Multimedia, 592–601
2021
-
[131]
Yongyi Tang, Lin Ma, and Lianqiang Zhou. 2019. Hallucinating optical flow features for video classification. arXiv preprint arXiv:1905.11799 (2019)
2019 arXiv
-
[132]
Yongyi Tang, Xing Zhang, Lin Ma, Jingwen Wang, Shaoxiang Chen, and Yu -Gang Jiang. 2018. Non -local netvlad encoding for video classification. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops , 0
2018
-
[133]
Yi Tay, Dara Bahri, Donald Metzler, Da -Cheng Juan, Zhe Zhao, and Che Zheng. 2021. Synthesizer: Rethinking self -attention for transformer models. In International conference on machine learning, PMLR, 10183–10192
2021
-
[134]
Yi Tay, Dara Bahri, Liu Yang, Donald Metzler, and Da -Cheng Juan. 2020. Sparse sinkhorn attention. In International Conference on Machine Learning , PMLR, 9438–9447
2020
-
[136]
Du Tran, Heng Wang, Lorenzo Torresani, and Matt Feiszli. 2019. Video classification with channel-separated convolutional networks. In Proceedings of the IEEE/CVF international conference on computer vision , 5552–5561
2019
-
[137]
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A closer look at spatiotemporal convo lutions for action recognition. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition , 6450–6459
2018
-
[138]
Gül Varol, Ivan Laptev, and Cordelia Schmid. 2017. Long -term temporal convolutions for action recognition. IEEE Trans Pattern Anal Mach Intell 40, 6 (2017), 1510–1517
2017
-
[139]
Ashish Vaswani, Prajit Ramachandran, Aravind Srinivas, Niki Parmar, Blake Hechtman, and Jonathon Shlens. 2021. Scaling local self-attention for parameter efficient visual backbones. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 12894–12904
2021
-
[140]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin . 2017. Attention is all you need. Adv Neural Inf Process Syst 30, (2017)
2017
-
[141]
Chunyu Wang, Yizhou Wang, and Alan L Yuille. 2013. An approach to pose -based action recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, 915–922
2013
-
[142]
Hanrui Wang, Zhanghao Wu, Zhijian Liu, Han Cai, Ligeng Zhu, Chuang Gan, and Song Han. 2020. Hat: Hardware -aware transformers for efficient natural language processing. arXiv preprint arXiv:2005.14187 (2020)
2020 arXiv
-
[143]
Limin Wang, Yu Qiao, and Xiaoou Tang. 2015. Action recognition with trajectory -pooled deep -convolutional descriptors. In Proceedings of the IEEE conference on computer vision and pattern recognition, 4305–4314
2015
-
[145]
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. 2020. Linformer: Self -attention with linear complexity. arXiv preprint arXiv:2006.04768 (2020)
2020 arXiv
-
[146]
Wenhai Wang, Enze Xie, Xiang Li, Deng -Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao. 2022. Pvt v2: Improved baselines with pyramid vision transformer. Comput Vis Media (Beijing) 8, 3 (2022), 415–424
2022
-
[147]
Xiaolong Wang, Ross Girshick, Abhinav Gupta, and Kaiming He. 2018. Non -local neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, 7794–7803
2018
-
[148]
Xinpeng Wang, Chandan Yeshwanth, and Matthias Nießner. 2021. Sceneformer: Indoor scene generation with transformers. In 2021 International Conference on 3D Vision (3DV), IEEE, 106–115
2021
-
[149]
Yingming Wang, Xiangyu Zhang, Tong Yang, and Jian Sun. 2022. Anchor detr: Query design for transformer -based detector. In Proceedings of the AAAI conference on artificial intelligence, 2567–2575
2022
-
[150]
Thomas A Woolsey, Joseph Hanaway, and Mokhtar H Gado. 2017. The brain atlas: A visual guide to the human central nervous system. John Wiley & Sons
2017
-
[151]
Chao-Yuan Wu, Christoph Feichtenhofer, Haoqi Fan, Kaiming He, Philipp Krahenbuhl, and Ross Girshick. 2019. Long-term feature banks for detailed video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 284–293
2019
-
[152]
Chao-Yuan Wu, Manzil Zaheer, Hexiang Hu, R Manmatha, Alexander J Smola, and Philipp Krähenbühl. 2018. Compressed video action reco gnition. In Proceedings of the IEEE conference on computer vision and pattern recognition , 6026–6035
2018
-
[153]
Wenhao Wu, Dongliang He, Tianwei Lin, Fu Li, Chuang Gan, and Errui Ding. 2021. Mvfnet: Multi -view fusion network for efficient video recognition. In Proceedings of the AAAI Conference on Artificial Intelligence , 2943–2951
2021
-
[154]
Yu-Huan Wu, Yun Liu, Xin Zhan, and Ming-Ming Cheng. 2022. P2T: Pyramid pooling transformer for scene understanding. IEEE Trans Pattern Anal Mach Intell (2022)
2022
-
[155]
Zuxuan Wu, Yu -Gang Jiang, Xi Wang, Hao Ye, and Xiangyang Xue. 2016. Multi -stream multi-class fusion of deep networks for video classification. In Proceedings of the 24th ACM international conference on Multimedia , 791–800
2016
-
[156]
Bruce Xiaohan Nie, Caiming Xiong, and Song -Chun Zhu. 2015. Joint action recognition and pose estimation from video. In Proceedings of the IEEE conference on computer vision and pattern recognition, 1293–1301
2015
-
[157]
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M Alvarez, and Ping Luo. 2021. SegFormer: Simple and efficient desi gn for semantic segmentation with transformers. Adv Neural Inf Process Syst 34, (2021), 12077–12090
2021
-
[158]
Saining Xie, Chen Sun, Jonathan Huang, Zhuowen Tu, and Kevin Murphy. 2018. Rethinking spatiotemporal feature learning: Speed -accuracy trade-offs in video classification. In Proceedings of the European conference on computer vision (ECCV) , 305–321
2018
-
[159]
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. 2021. Nyströmformer: A nyström-based algorithm for approximating self-attention. In Proceedings of the AAAI Conference on Artificial Intelligence , 14138–14148
2021
-
[160]
Mingze Xu, Yuanjun Xiong, Hao Chen, Xinyu Li, Wei Xia, Zhuowen Tu, and Stefano Soatto. 2021. Long short-term transformer for online action detection. Adv Neural Inf Process Syst 34, (2021), 1086–1099
2021
-
[161]
Fuzhi Yang, Huan Yang, Jianlong Fu, Hongtao Lu, and Baining Guo. 2020. Learning texture transformer network for image super-resolution. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 5791–5800
2020
-
[162]
Ren Yang, Fabian Mentzer, Luc Van Gool, and Radu Timofte. 2020. Learning for video compression with recurrent auto -encoder and recurrent probability model. IEEE J Sel Top Signal Process 15, 2 (2020), 388–401
2020
-
[163]
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015. Beyond short snippets: Deep networks for video classification. In Proceedings of the IEEE conference on computer vision and pattern recognition , 4694–4702
2015
-
[164]
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. 2022. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12104–12113
2022
-
[165]
Chuhan Zhang, Ankush Gupta, and Andrew Zisserman. 2021. Temporal query networks for fine -grained video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 4486–4496
2021
-
[166]
Hao Zhang, Yanbin Hao, and Chong -Wah Ngo. 2021. Token shift transformer for video classification. In Proceedings of the 29th ACM International Conference on Multimedia, 917–925
2021
-
[167]
Yujia Zhang, Xiaodan Liang, Dingwen Zhang, Min Tan, and Eric P Xing. 2020. Unsupervised object -level video summarization with online motion auto - encoder. Pattern Recognit Lett 130, (2020), 376–385
2020
-
[168]
Hang Zhao, Antonio Torralba, Lorenzo Torresani, and Zhicheng Yan. 2019. Hacs: Human action clips and segments dataset for rec ognition and temporal localization. In Proceedings of the IEEE/CVF International Conference on Computer Vision , 8668–8678
2019
-
[169]
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip H S Torr, and Vladlen Koltun. 2021. Point transformer. In Proceedings of the IEEE/CVF international conference on computer vision, 16259–16268
2021
-
[170]
Yiru Zhao, Bing Deng, Chen Shen, Yao Liu, Hongtao Lu, and Xian -Sheng Hua. 2017. Spatio -temporal autoencoder for video anomaly detection. In Proceedings of the 25th ACM international conference on Multimedia , 1933–1941
2017
-
[171]
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, and Phil ip H S Torr. 2021. Rethinking semantic segmentation from a sequence -to-sequence perspective with transformers. In Proceedings of the IEEE/CVF conferen...
2021
-
[172]
Bolei Zhou, Alex Andonian, Aude Oliva, and Antonio Torralba. 2018. Temporal relational reasoning in videos. In Proceedings of the European conference on computer vision (ECCV), 803–818
2018
-
[173]
Linchao Zhu, Du Tran, Laura Sevilla -Lara, Yi Yang, Matt Feiszli, and Heng Wang. 2020. Faster recurrent networks for efficient video classification. In Proceedings of the AAAI conference on artificial intelligence , 13098–13105
2020
-
[174]
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159 (2020)
2020 arXiv
-
[175]
Yi Zhu, Zhenzhong Lan, Shawn Newsam, and Alexander Hauptmann. 2019. Hidden two-stream convolutional networks for action recognition. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2 –6, 2018, Revised Selected Papers, Part III...
2019
-
[176]
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox. 2018. Eco: Efficient convolutional network for online video understanding. In Proceedings of the European conference on computer vision (ECCV), 695–712
2018
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.