REVIEW 4 major objections 4 minor 52 references
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read VidEvent introduces the first large-scale video dataset for extracting event scripts and analyzing long-term event evolution, built from 1,110 movie recaps with 23,989 events and 17,525 relations.
desk verdict Useful new dataset and task, but the load-bearing reliability claim is unverified — no annotation quality metrics, and text-only baselines nearly match multimodal on the core benchmarks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The event script is the load-bearing object: a set of structured events, each with a trigger verb and a role-filled argument set, bundled with binary relations among those events so that a whole video becomes a chain of causally and temporally linked events. The annotation scheme around this object uses PropBank as the event vocabulary, segments videos into event clips whose boundaries follow the events rather than fixed timescales, gives each person-related entity a stable identity number across the video, and labels relations following an extended scheme of temporal, causal, conditional, coreference, and subevent types. The script object carries the argument because it converts understanding a video into a concrete prediction problem: given a partial script, choose the next event among candidates.
What would settle it
Annotate a random sample of VidEvent videos a second time with new annotators who have not seen the original labels, using the same guidelines, and measure agreement on event boundaries, triggers, arguments, and relation types; if agreement is low, the claim that the annotations reliably represent event evolution in the recaps is falsified.
Extended reading notes
Core claim
The central claim is that VidEvent is the first large-scale video dataset that supports both the extraction of highly concluded events and the analysis of long-term event evolution. Each event is represented as a structure with a PropBank verb as trigger and a set of semantic roles as arguments; a unique identifier is assigned to each character so role fillers stay co-referenced across an entire video. Events are linked by binary relations of five types — temporal, causal, conditional, coreference, and subevent — and the resulting scripts have an average chain length of 10.57. The paper supplies an annotation protocol built on movie recap videos, four baseline tasks with evaluation metrics, and experiments showing that semantic aggregation improves event localization and that combining visual and textual features helps relation classification.
Load-bearing premise
The whole dataset rests on the assumption that movie recap videos are faithful, complete records of event structure and logic; if recaps omit or distort events, the ground truth is unreliable.
Editorial extensions
If this is right
- Video event understanding gives a common task definition and evaluation metrics for event-level reasoning, so future systems can be compared on localization, extraction, relation classification, and script induction.
- Because event chains average 10.57 links, the dataset supports studying long-range logical structure rather than only pairwise actions.
- The use of PropBank-style triggers and co-referenced entities connects video event understanding directly to NLP event extraction, potentially allowing transfer between text and video.
- Baseline results indicate that adding semantic aggregation over atomic action boundaries improves event localization, and that visual and textual modalities complement each other in relation classification.
- With all annotations made public through video URLs, VidEvent offers a reproducible training and test bed for the proposed tasks.
Reading between the lines
- Movie recap videos are condensations selected by their narrators, so VidEvent's event distributions may over-represent dramatic, dialogue-driven plot points and under-represent visual or ambient events; this could be tested by comparing event chains against full-length source films.
- The single-choice, five-candidate induction setup sidesteps open-ended generation; an extension would evaluate whether a model can generate the next event without an existing candidate set.
- The unique character identifiers could support character-centric downstream tasks, such as tracking who caused what across a long script and grounding coreference in visual appearance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a new task, video event understanding, which comprises four progressive subtasks: video event localization, video event extraction, event relation classification, and script event induction. To support this task, the authors introduce VidEvent, a dataset built from 1,110 movie recap videos containing 23,989 events, 80,822 event arguments, and 17,525 relations, annotated with PropBank-style event structures and relations adapted from MAVEN-ERE with an additional conditional relation type. The paper provides formal definitions of event clips, event structures, relations, and event scripts, and it reports a baseline framework with vision and text encoders for each subtask. The central claim is that VidEvent is the first large-scale dataset that supports extracting highly-concluded events and analyzing long-term event evolution in videos, and the paper presents statistical comparisons with existing datasets to support this claim. The dataset is advertised as publicly available at www.videvent.top.
Significance. If the dataset is reliable and accessible, it would fill a genuine gap: most video datasets focus on atomic actions, sentence-level captions, or situation structures without long, typed event chains. The four subtasks are clearly defined, and the baseline framework is a reasonable starting point for future work. The paper is not circular: the baselines are evaluated on held-out data, and no fitted parameters are used to generate the annotations. The statistical comparisons of event density, argument distribution, relation variety, and coreference chain length are informative. However, the value of the contribution depends critically on annotation quality and on the extent to which the event labels are visually grounded. The manuscript does not yet provide evidence for either, and the near-equality of text-only and multimodal baseline results raises a substantive concern that the task may be largely solvable from subtitles without genuine video understanding.
major comments (4)
- [Dataset: VidEvent / Data Composition (Table 2)] The central claim that VidEvent is a 'well-labeled' and reliable dataset is not backed by any annotation-quality measurement. The manuscript reports no inter-annotator agreement, no number of annotators, no instruction-manual summary, and no definition of the 96% 'Annotation Coverage' entry in Table 2. Since the dataset is the main contribution, these omissions are load-bearing. Please report agreement statistics (e.g., boundary agreement and Cohen's kappa or Krippendorff's alpha for trigger, argument, and relation labels), define the coverage metric, and describe the annotation review process in enough detail for readers to assess label reliability.
- [Data Composition, 'Movie Recap Videos' paragraph; Tables 5 and 6] The statement that 'each sentence of the subtitles in the movie recap videos corresponds to one event' suggests that event boundaries and possibly event structures are derived largely from transcript text rather than from visual evidence. The baseline results are consistent with this risk: in Table 5, text-only RoBERTa reaches 0.55 accuracy in script event induction, essentially matching TF+RoBERTa at 0.56; in Table 6, text-only RoBERTa achieves F1=0.41 versus SF+RoBERTa at 0.48. The paper should provide quantitative evidence that event triggers, arguments, and relations are visually grounded, for example by having annotators label events from video without subtitles and measuring agreement with the released annotations, or by reporting the fraction of events whose arguments are not explicitly named in subtitles. Without such evidence, the claim that VidEvent supports 'video event understanding' rather than 'subtitle-based script induction' is not established.
- [Experiments, Tables 3-6] All baseline results are reported as single point estimates without variance, error bars, number of runs, or significance tests. Because the baselines are intended as a benchmark for future comparisons, the absence of uncertainty estimates makes it impossible to judge whether small differences, such as 0.55 versus 0.56 in Table 5, are meaningful. Please report means and standard deviations over multiple seeds or runs, and state whether the differences between text-only and multimodal models are statistically significant.
- [Script Event Induction in Vision] The construction of the five candidate events is underspecified. The text says that 'two events are selected from a unified candidate pool, one event outside the event chain within the same video, and one event serves as a distractor by randomly replacing an event argument,' but it does not define the candidate pool, the sampling distribution, the difficulty of the distractors, or the chance level of the resulting task. Consequently, the text-only accuracy of 0.55 cannot be interpreted. Please specify the candidate-generation procedure, report the distribution of candidate types, and consider reporting results separately for different distractor difficulties.
minor comments (4)
- [Table 1] The symbols in the rows of Table 1 (for example 'verb ! % %' and 'verb ! % !') are not explained and appear to be rendering artifacts of check or cross marks; please use clearly defined symbols and add a legend.
- [Event Relations and Experiments] The Event Relations section describes five relation types (temporal, causal, conditional, coreference, subevent), but the Experiments section states that relation classification is a '6-way classification problem.' Please clarify whether a 'no relation' class is included and how the taxonomy maps to the six classes.
- [Throughout] There are several typos and ungrammatical sentences: 'event ralation classification' in the Figure 3 caption, 'occurence' in the Event Relations paragraph, 'correference' in the section heading, 'analyzing the dynamic complex events ... depict the extensive research' in the Task section, and 'induct with scripts' in the Introduction.
- [Data Composition] Because VidEvent provides only URLs rather than video files, please state the licensing terms and the expected stability of the URLs, and include a complete sample annotation in the supplementary material so that the JSON format and argument/relation conventions can be inspected without downloading the full dataset.
Circularity Check
No significant circularity: VidEvent is a dataset-and-benchmark paper whose claims are supported by held-out evaluations; no prediction reduces to an input by construction.
full rationale
The paper's central artifacts are a newly constructed dataset and baseline evaluations. The derivation chain is: (i) define event structures and relations; (ii) annotate movie recap videos under PropBank and MAVEN-ERE conventions; (iii) propose four subtasks; (iv) train standard video and text models and report held-out metrics. None of these steps fits a parameter to an output and then reports that output as a prediction. The baseline numbers in Tables 3-6 are computed on test sets using standard metrics (mAP, P@5/R@5, F1, METEOR/CIDEr/SPICE, top-1 accuracy), and text-only and multimodal comparisons are reported, so the results are not forced by construction. The assumption that recap videos provide complete event structure is a data-quality premise, not a circular reduction; it could weaken validity but does not make any equation or result equivalent to its input. Citations to PropBank, MAVEN-ERE, VidSitu, and the recap-story work are external sources, not self-citations that carry the paper's central claim. Consequently, no circular step meeting the paper-quoting standard can be identified.
Assumptions & free parameters
assumptions (3)
- domain assumption PropBank verb frames and argument roles are an appropriate vocabulary for visual events.
- domain assumption Movie recap videos provide complete and coherent event structures suitable for annotation.
- domain assumption Pairwise binary relations are sufficient to capture dynamic event evolution.
Cite this review
Pith. "Pith review of VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos." pith.science (2026). https://pith.science/paper/VFR5LDI3
@misc{pith2026250602448,
author = {Pith},
title = {Pith review of: VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/VFR5LDI3}},
note = {Machine review of arXiv:2506.02448}
}
read the original abstract
Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose the task of video event understanding that extracts event scripts and makes predictions with these scripts from videos. To support this task, we introduce VidEvent, a large-scale dataset containing over 23,000 well-labeled events, featuring detailed event structures, broad hierarchies, and logical relations extracted from movie recap videos. The dataset was created through a meticulous annotation process, ensuring high-quality and reliable event data. We also provide comprehensive baseline models offering detailed descriptions of their architecture and performance metrics. These models serve as benchmarks for future research, facilitating comparisons and improvements. Our analysis of VidEvent and the baseline models highlights the dataset's potential to advance video event understanding and encourages the exploration of innovative algorithms and models. The dataset and related resources are publicly available at www.videvent.top.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
P.; Toderici, G.; Varadarajan, B.; and Vijayanarasimhan, S
Abu-El-Haija, S.; Kothari, N.; Lee, J.; Natsev, A. P.; Toderici, G.; Varadarajan, B.; and Vijayanarasimhan, S. 2016. YouTube-8M: A Large-Scale Video Classification Benchmark. In arXiv:1609.08675
arXiv 2016
-
[4]
Anderson, P.; Fernando, B.; Johnson, M.; and Gould, S. 2016. SPICE: Semantic Propositional Image Caption Evaluation. In ECCV
work page 2016
-
[5]
Banerjee, S.; and Lavie, A. 2005. METEOR : An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization , 65--72. Association for Computational Linguistics
work page 2005
-
[6]
Bertasius, G.; Wang, H.; and Torresani, L. 2021. Is space-time attention all you need for video understanding? In ICML, volume 2, 4
work page 2021
-
[7]
Chen, Z.; Ma, L.; Luo, W.; and Wong, K.-Y. K. 2019. Weakly-Supervised Spatio-Temporally Grounding Natural Sentence in Video. In ACL
work page 2019
-
[8]
Cheng, F.; and Bertasius, G. 2022. Tallformer: Temporal action localization with a long-memory transformer. In ECCV, 503--521. Springer
work page 2022
Show all 52 references
-
[9]
M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M
Damen, D.; Doughty, H.; Farinella, G. M.; Fidler, S.; Furnari, A.; Kazakos, E.; Moltisanti, D.; Munro, J.; Perrett, T.; Price, W.; and Wray, M. 2018. Scaling Egocentric Vision: The EPIC-KITCHENS Dataset. arXiv:1804.02748
2018 arXiv
-
[10]
Diba, A.; Fayyaz, M.; Sharma, V.; Paluri, M.; Gall, J.; Stiefelhagen, R.; and Van Gool, L. 2020. Large Scale Holistic Video Understanding. In ECCV 2020, 593–610. Springer-Verlag
2020
-
[11]
Duan, H.; Zhao, Y.; Chen, K.; Lin, D.; and Dai, B. 2022. Revisiting skeleton-based action recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2969--2978
2022
-
[12]
G., Victor Escorcia; and Niebles, J
Fabian Caba Heilbron, B. G., Victor Escorcia; and Niebles, J. C. 2015. ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 961--970
2015
-
[13]
Feichtenhofer, C.; Fan, H.; Malik, J.; and He, K. 2019. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, 6202--6211
2019
-
[14]
Gao, J.; Sun, C.; Yang, Z.; and Nevatia, R. 2017. TALL: Temporal Activity Localization via Language Query. arXiv:1705.02101
2017 arXiv
-
[15]
Something Something
Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; Hoppe, F.; Thurau, C.; Bax, I.; and Memisevic, R. 2017. The "Something Something" Video Database for Learning and Evaluating Visual ...
2017
-
[16]
A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; Schmid, C.; and Malik, J
Gu, C.; Sun, C.; Ross, D. A.; Vondrick, C.; Pantofaru, C.; Li, Y.; Vijayanarasimhan, S.; Toderici, G.; Ricco, S.; Sukthankar, R.; Schmid, C.; and Malik, J. 2018. AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions. arXiv:1705.08421
2018 arXiv
-
[17]
A.; Toderici, G.; Li, Y.; Ricco, S.; Sukthankar, R.; Schmid, C.; and Malik, J
Gu, C.; Sun, C.; Vijayanarasimhan, S.; Pantofaru, C.; Ross, D. A.; Toderici, G.; Li, Y.; Ricco, S.; Sukthankar, R.; Schmid, C.; and Malik, J. 2017. AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions. 2018 IEEE/CVF Conference on Computer Vision and Patter...
2017
-
[18]
A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B
Hendricks, L. A.; Wang, O.; Shechtman, E.; Sivic, J.; Darrell, T.; and Russell, B. 2017. Localizing Moments in Video with Natural Language. arXiv:1708.01641
2017 arXiv
-
[19]
in the wild
Idrees, H.; Zamir, A. R.; Jiang, Y.-G.; Gorban, A.; Laptev, I.; Sukthankar, R.; and Shah, M. 2017. The THUMOS challenge on action recognition for videos “in the wild”. Computer Vision and Image Understanding, 155: 1--23
2017
-
[20]
Jhuang, H.; Gall, J.; Zuffi, S.; Schmid, C.; and Black, M. J. 2013. Towards understanding action recognition. In International Conf. on Computer Vision (ICCV), 3192--3199
2013
-
[21]
Khan, Z.; Jawahar, C.; and Tapaswi, M. 2022. Grounded Video Situation Recognition. Advances in Neural Information Processing Systems, 35: 8199--8210
2022
-
[22]
Kingsbury, P.; and Palmer, M. 2003. Propbank: the next level of treebank. In Proceedings of Treebanks and lexical Theories, volume 3. Citeseer
2003
-
[23]
K.; On, K.-W.; Roh, B.; and Kim, H
Ko, D.; Choi, J.; Choi, H. K.; On, K.-W.; Roh, B.; and Kim, H. J. 2023. MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20105--20115
2023
-
[24]
Krishna, R.; Hata, K.; Ren, F.; Fei-Fei, L.; and Carlos Niebles, J. 2017. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, 706--715
2017
-
[25]
Lei, J.; Yu, L.; Bansal, M.; and Berg, T. L. 2019. TVQA: Localized, Compositional Video Question Answering. arXiv:1809.01696
2019 arXiv
-
[26]
Li, M.; Xu, R.; Wang, S.; Zhou, L.; Lin, X.; Zhu, C.; Zeng, M.; Ji, H.; and Chang, S.-F. 2022. Clip-event: Connecting text and images with event structures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 16420--16429
2022
-
[27]
Liu, X.; Wang, Q.; Hu, Y.; Tang, X.; Zhang, S.; Bai, S.; and Bai, X. 2022. End-to-end temporal action detection with transformer. IEEE Transactions on Image Processing, 31: 5427--5441
2022
-
[28]
Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; and Stoyanov, V. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692
2019 arXiv
-
[29]
Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; and Sivic, J. 2019. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. arXiv:1906.03327
2019 arXiv
-
[30]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In Meila, M.; and Zhang, T., eds., Proceedings of...
2021
-
[31]
A.; and Zacks, J
Radvansky, G. A.; and Zacks, J. M. 2017. Event boundaries in memory and cognition. Current opinion in behavioral sciences, 17: 133--140
2017
-
[32]
Rohrbach, A.; Torabi, A.; Rohrbach, M.; Tandon, N.; Pal, C.; Larochelle, H.; Courville, A.; and Schiele, B. 2016. Movie Description. arXiv:1605.03705
2016 arXiv
-
[33]
Sadhu, A.; Gupta, T.; Yatskar, M.; Nevatia, R.; and Kembhavi, A. 2021 a . Visual Semantic Role Labeling for Video Understanding. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2021
-
[34]
Sadhu, A.; Gupta, T.; Yatskar, M.; Nevatia, R.; and Kembhavi, A. 2021 b . Visual semantic role labeling for video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5589--5600
2021
-
[35]
Schroff, F.; Kalenichenko, D.; and Philbin, J. 2015. FaceNet: A unified embedding for face recognition and clustering. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 815--823
2015
-
[36]
Shi, D.; Zhong, Y.; Cao, Q.; Ma, L.; Li, J.; and Tao, D. 2023. Tridet: Temporal action detection with relative boundary modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 18857--18866
2023
-
[37]
A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A
Sigurdsson, G. A.; Varol, G.; Wang, X.; Farhadi, A.; Laptev, I.; and Gupta, A. 2016. Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding. arXiv:1604.01753
2016 arXiv
-
[38]
Previously on
Singh, A. K.; Srivastava, D.; and Tapaswi, M. 2024. "Previously on..." from Recaps to Story Summarization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024 , 13635--13646. IEEE
2024
-
[39]
L.; and Parikh, D
Vedantam, R.; Zitnick, C. L.; and Parikh, D. 2015. CIDEr: Consensus-based image description evaluation. In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 4566--4575
2015
-
[40]
Wang, R.; Chen, D.; Wu, Z.; Chen, Y.; Dai, X.; Liu, M.; Yuan, L.; and Jiang, Y.-G. 2023 a . Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern...
2023
-
[41]
Wang, X.; Chen, Y.; Ding, N.; Peng, H.; Wang, Z.; Lin, Y.; Han, X.; Hou, L.; Li, J.; Liu, Z.; et al. 2022. MAVEN-ERE: A Unified Large-scale Dataset for Event Coreference, Temporal, Causal, and Subevent Relation Extraction. In Proceedings of the 2022 Conference on Empirical Met...
2022
-
[42]
Wang, X.; Li, J.; Zhu, L.; Zhang, Z.; Chen, Z.; Li, X.; Wang, Y.; Tian, Y.; and Wu, F. 2023 b . Visevent: Reliable object tracking via collaboration of frame and event flows. IEEE Transactions on Cybernetics
2023
-
[43]
Wang, X.; Wu, J.; Chen, J.; Li, L.; Wang, Y.-F.; and Wang, W. Y. 2019. Vatex: A large-scale, high-quality multilingual dataset for video-and-language research. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 4581--4591
2019
-
[44]
Wang, X.; Wu, J.; Chen, J.; Li, L.; Wang, Y.-F.; and Wang, W. Y. 2020. VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research. arXiv:1904.03493
2020 arXiv
-
[45]
Wei, X.; Bai, Y.; Zheng, Y.; Shi, D.; and Gong, Y. 2023. Autoregressive Visual Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9697--9706
2023
-
[46]
Xu, D.; Zhao, Z.; Xiao, J.; Wu, F.; Zhang, H.; He, X.; and Zhuang, Y. 2017. Video Question Answering via Gradually Refined Attention over Appearance and Motion. In Proceedings of the 25th ACM International Conference on Multimedia, 1645–1653. Association for Computing Machinery
2017
-
[47]
Xu, J.; Mei, T.; Yao, T.; and Rui, Y. 2016. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, 5288--5296
2016
-
[48]
Zhang, C.-L.; Wu, J.; and Li, Y. 2022. Actionformer: Localizing moments of actions with transformers. In European Conference on Computer Vision, 492--510. Springer
2022
-
[49]
Zhang, Z.; Zhao, Z.; Zhao, Y.; Wang, Q.; Liu, H.; and Gao, L. 2020. Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences. In CVPR
2020
-
[50]
Zhao, H.; Torralba, A.; Torresani, L.; and Yan, Z. 2019. HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization. arXiv:1712.09374
2019 arXiv
-
[51]
Zhong, Y.; Xiao, J.; Ji, W.; Li, Y.; Deng, W.; and Chua, T.-S. 2022. Video Question Answering: Datasets, Algorithms and Challenges. arXiv:2203.01225
2022 arXiv
-
[52]
Zisserman, A.; Carreira, J.; Simonyan, K.; Kay, W.; Zhang, B.; Hillier, C.; Vijayanarasimhan, S.; Viola, F.; Green, T.; Back, T.; Natsev, P.; and Suleyman, M. 2017. In The Kinetics Human Action Video Dataset
2017
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.