REVIEW 4 major objections 5 minor 2 cited by
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read CineTechBench, an expert-annotated benchmark spanning seven cinematography dimensions, finds that 15+ multimodal models and 5+ video generators all fall short at film technique, both in understanding and in reproduction.
desk verdict A useful expert-annotated benchmark for cinematographic understanding, with a generation-evaluation section whose quantitative claims currently rest on an unvalidated trajectory estimator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the seven-dimension cinematography taxonomy (scale, angle, composition, movement, lighting, color, focal length) built from professional filmmaking sources and refined with expert cinematographers, together with the annotation pipeline behind it: trained annotators label each image and clip, GPT-4o drafts question-answer pairs and descriptions, and annotators refine them. On the understanding side the benchmark measures per-dimension accuracy on QA pairs and adapts the CAPability caption metrics (hit rate, precision, recall, F1) with GPT-4.1-nano scoring each description as Miss, Positive, or Negative. On the generation side it uses MonST3R to estimate camera trajectories from both the original clips and the generated videos, then compares them with rotation error, translation error, and CamMC, plus a CLIP-based frame similarity score.
What would settle it
Re-annotate a random subset — about 100 images and 30 clips — with two independent professional cinematographers who never see the original labels, and compute per-dimension agreement; then have human judges re-score a sample of 50 model captions with the same Miss/Positive/Negative rubric used by the automated judge. If agreement on lighting direction or focal length falls below roughly 0.6 kappa, or if the automated and human caption scores diverge on a meaningful share of samples, the reported model rankings would not be stable.
Extended reading notes
Core claim
The paper's central claim is that CineTechBench is the first benchmark to cover all seven core dimensions of cinematography at once and to pair them with expert labels, and that measuring against it reveals a clear, consistent shortfall. On static image QA, the best commercial model reaches 70.16% overall accuracy while open-source models trail by about 15 points; on camera movement QA the best model reaches only 56.69%, with rotation the weakest movement type. On description generation, models mention the right dimensions often (hit rates above 80% for static dimensions) but describe them precisely only about a third to a half of the time (F1 between 30% and 50%), and camera-movement descriptions have a hit rate near 30%. For generation, image-to-video models conditioned on first and last frames produce lower rotation and translation errors than first-frame-only models, yet all tested generators struggle with high angular velocity, sometimes producing a roll in the wrong direction or no roll at all.
Load-bearing premise
The benchmark's conclusions assume the expert labels are consistent (no inter-annotator agreement is reported), the GPT-4.1-nano judge scores descriptions the way a human would, and MonST3R's estimated camera trajectories are accurate.
Editorial extensions
If this is right
- Any future multimodal model can be plugged into the same QA and description protocols and ranked on the same seven dimensions, making cinematographic understanding a trackable metric rather than an anecdotal one.
- The documented gaps — lighting direction, focal length, and camera rotation direction — give a concrete agenda for model builders, since these are the dimensions where visual encoders and training data are weakest.
- For video generation, the results imply that camera control, especially angular motion such as high-amplitude rolls, is the main bottleneck separating current generators from cinema-quality output, and that first-frame-plus-last-frame conditioning is only a partial fix.
- The near-30% hit rate for camera-movement descriptions shows that dynamic cinematography is disproportionately harder for multimodal models than static dimensions, pointing to video-specific perception as a distinct weakness.
- Because the benchmark is public, its numbers can serve as a baseline that later models are expected to exceed, giving the community a shared yardstick for film-level understanding and generation.
Reading between the lines
- The paper does not report inter-annotator agreement, so a reader extending it should first check how stable the labels are: a small re-annotation study by independent cinematographers would tell whether the per-dimension accuracy ordering is trustworthy.
- The description rankings depend on GPT-4.1-nano agreeing with human judgment about whether a caption is correct; a human side-by-side study on a sample of captions would reveal whether the automated judge's Miss/Positive/Negative calls skew the reported F1 scores.
- The generation numbers inherit any error in MonST3R's trajectory estimates, which the paper acknowledges; obtaining clips with true camera motion metadata, for example from virtual production or on-set motion capture, would sharpen the rotation and translation error measurements.
- A natural next experiment suggested by the results is to train or fine-tune a video generator on prompts that explicitly name rotation direction and amplitude, such as 'roll clockwise', and test on the same 120 clips to see whether the documented rotation failures shrink.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CineTechBench is a new benchmark for evaluating multimodal LLMs and video generation models on seven cinematography dimensions: shot scale, shot angle, composition, camera movement, lighting, color, and focal length. It comprises over 600 expert-annotated movie stills and over 120 movie clips, with QA pairs and captions that were initially generated by GPT-4o and then manually reviewed. The understanding task reports per-dimension QA accuracy for 15+ MLLMs and CAPability-style description metrics; the generation task reports RotErr, TransErr, CamMC, and CLIP-IS for 5+ image-to-video models, where trajectory metrics are computed from MonST3R-estimated camera paths. The central findings are that MLLMs are weakest on camera rotation and lighting direction, and that video generation models perform worst under high-amplitude rotation.
Significance. The benchmark addresses a real gap: existing movie benchmarks focus on high-level semantics, while professional cinematography labels are scarce. The seven-dimension taxonomy is a reasonable synthesis of standard film terminology, the model coverage is broad, and the paper includes qualitative examples that support the broad qualitative conclusions (e.g., rotation direction errors in Figure 4 and failure to produce roll in Figure 5). The authors are transparent about limitations in Appendix F and release code and dataset links, which is a clear strength. The quantitative generation claims and part of the description claims, however, rest on unvalidated pseudo-ground-truth: MonST3R trajectories for camera motion and GPT-4.1-nano judgments for caption scoring. These issues are load-bearing for the paper's quantitative conclusions, though they appear fixable within a revision; the relative finding that models struggle with fine-grained rotation and lighting direction is internally consistent and plausible.
major comments (4)
- [Section 4.2, Equations (1)-(2), Table 4, Figure 6] The generation conclusion that models 'struggle with camera movement with intense rotation amplitude' and the Table 4 rankings rest on RotErr, TransErr, and CamMC computed from MonST3R trajectories, and Appendix F.1 concedes that such pose estimates often introduce inaccuracies for dynamic scenes, motion blur, and non-rigid object motion. No validation of MonST3R against known camera paths is reported, and the highlighted failure regime (fast rotations) is precisely where rotation estimation is most fragile. Because the original and generated clips differ in content, estimator errors need not cancel and could systematically inflate errors for particular models or motion types; the quantitative ordering in Table 4 and the angular-velocity trend in Figure 6 may therefore reflect estimator sensitivity rather than generation quality. To support these claims, please validate the estimator on sequences with known trajectories (including fast rolls) or add a human evaluation of rotation direction and amplitude on a subset of generations. The qualitative examples in Figure 5 support a broad claim of poor rotation generation but not the amplitude-dependent conclusion.
- [Section 3.2 and Appendix B] The annotation pipeline is described as expert-annotated with ambiguous cases discarded, but no inter-annotator agreement statistics are reported. Since every QA accuracy number in Tables 1 and 2 is measured against these labels, label noise could affect the relative ordering of models, particularly in boundary-prone dimensions such as shot scale (e.g., Medium Close-Up versus Close-Up) and lighting direction (Side versus Back versus Top). Please report per-dimension agreement (e.g., Cohen's kappa or Krippendorff's alpha) on a sample of the annotations, or quantitatively describe the adjudication protocol, to substantiate the claim of 'precise, manual annotation' in the abstract.
- [Appendix C.1, Tables 3 and 6] The description metrics (HR, AP, AR, F1) are computed by a GPT-4.1-nano judge that classifies each caption as Miss, Positive, or Negative, but the judge's agreement with human judgments is not measured. The unusually low Movement hit rate (29.83%) and the F1 differences that determine model rankings could in principle be artifacts of judge bias. Please validate the judge on a human-rated sample (e.g., 50-100 captions per dimension), report the agreement, and consider providing the per-caption judge outputs alongside the released benchmark.
- [Section 3.2] The understanding test items were generated by GPT-4o and then manually reviewed, but the paper does not quantify how much of the GPT-4o output was changed by human reviewers. Because GPT-4o is itself among the evaluated models, the benchmark may inadvertently reward agreement with GPT-4o's annotation style rather than with expert cinematography. Please report the modification or rewrite rate, the number of items substantively revised, and the criteria reviewers used to override or confirm GPT-4o's content; this information is also needed to support the description of the benchmark as based on manual expert annotation.
minor comments (5)
- [Throughout the manuscript] There are numerous typographical errors and inconsistent naming conventions (e.g., 'flim clips' in the Introduction, 'ineatographic' in Section 2.1, 'LLaV A' in tables and text, and 'Monst3r' in the Figure 5 caption); a careful proofreading pass is needed.
- [Table 4] The table header 'Rel. Abs. Rel. Abs.' is ambiguous; please explicitly state that relative and absolute variants apply to RotError and TransError, and define the normalization used for each variant in the caption or in Appendix C.2.
- [Equation (1)] The definition of RotErr sums per-frame angular errors but does not specify n or how the frames of the original and generated clips are temporally aligned after downsampling; please state the alignment procedure and the frame count used.
- [Figure 6] The x-axis quantities 'translation speed' and 'angular velocity' are not defined; please provide units, binning details, and the number of test clips per bin so that the trend can be interpreted.
- [Section 4.2, Equation (3)] The CLIP-IS metric compares original and generated frames, so it partly reflects content similarity rather than camera-movement fidelity; the text should explicitly caveat that CLIP-IS is not a pure motion-quality measure.
Circularity Check
No load-bearing circularity; evaluation-validity caveats (GPT-4o-generated items, LLM judge, MonST3R trajectories) are not by-construction reductions.
full rationale
CineTechBench is a benchmark-and-evaluation paper, not a derivation. The central claims—that current MLLMs underperform on cinematographic understanding and that video generators struggle with rotation—are empirical measurements obtained by scoring models against human-reviewed labels and external trajectory estimates. The strongest candidates for circularity do not survive scrutiny. (1) Test items were originally drafted by GPT-4o, one of the evaluated models, but Section 3.2 states 'All generated content was manually reviewed and refined by trained annotators to ensure accuracy, clarity, and alignment with professional cinematography standards,' so the labels are not GPT-4o's outputs by construction. (2) Description scoring uses GPT-4.1-nano as an automated judge (Appendix C.1); the judge is not the model being scored, and the classification categories (Miss/Positive/Negative) are anchored to human ground truth, so no equation reduces to the evaluated model's output. (3) Camera-motion metrics use MonST3R [51] to estimate trajectories of both original and generated clips (Eq. 1–2). Appendix F.1 concedes such estimates 'often introduce inaccuracies due to complex cinematographic factors such as dynamic scenes, motion blur, and non-rigid object motion'; this is a real measurement-validity limitation that could affect model rankings in Table 4 and Figure 6, but it is not circular because the estimator is external to the evaluated models and no fitted parameter is renamed as a prediction. The only overlapping-author citation, [10] (Conmo, which includes co-authors Liang and Ma), appears in related work as an example of video-generation advances and is not load-bearing; no uniqueness theorem from the authors is invoked. No pattern of self-definitional labels, fitted inputs called predictions, or ansatz smuggled via citation is present. Accordingly, the circularity score is minimal; the flagged limitations belong to correctness risk rather than circular reasoning.
Assumptions & free parameters
assumptions (3)
- domain assumption Expert annotations and GPT-4o-generated, human-reviewed QA and descriptions are accurate ground truth for cinematographic techniques.
- domain assumption MonST3R-estimated camera trajectories from movie clips are accurate enough to serve as pseudo ground truth for evaluating generated camera motion.
- domain assumption GPT-4.1-nano prompt classifications (Miss, Positive, Negative) agree with human judgments for CAPability-style metrics.
Cite this review
Pith. "Pith review of CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation." pith.science (2026). https://pith.science/paper/CJZLYI3F
@misc{pith2026250515145,
author = {Pith},
title = {Pith review of: CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/CJZLYI3F}},
note = {Machine review of arXiv:2505.15145}
}
read the original abstract
Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects-shot scale, shot angle, composition, camera movement, lighting, color, and focal length-and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question answer pairs and annotated descriptions to assess MLLMs' ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatically film production and appreciation. The code and benchmark can be accessed at https://github.com/PRIS-CV/CineTechBench.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 2 Pith papers
-
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
A film-academy cinematic taxonomy and reverse-engineered multi-shot prompts expose large gaps in leading video generators that web-style benchmarks miss.
-
Natural Language Camera Movement Understanding
A cinematographic taxonomy, atomic real+synthetic benchmark (ACaM), and targeted-augmentation SFT let an 8B VLM outperform Gemini-3.1-Pro by 10-11% on camera-movement recognition, yet a large human gap remains.
Reference graph
Works this paper leans on
-
[1]
Qwen2.5-vl technical report, 2025
Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025
2025
-
[2]
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare V oss, editors,Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan,...
work page 2005
-
[3]
Align your latents: High-resolution video synthesis with latent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023
2023
-
[4]
Movieclip: Visual scene recognition in movies
Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Haoyang Zhang, Yin Cui, Kree Cole- McLaughlin, Huisheng Wang, and Shrikanth Narayanan. Movieclip: Visual scene recognition in movies. In 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2082–2091, 2023
work page 2023
-
[5]
Skyreels-v2: Infinite-length film generative model, 2025
Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengcheng Ma, Weiming Xiong, Wei Wang, Nuo Pang, Kang Kang, Zhiheng Xu, Yuzhe Jin, Yupeng Liang, Yubing Song, Peng Zhao, Boyuan Xu, Di Qiu, Debang Li, Zhengcong Fei, Yang Li, and Yahui Zhou. Skyreels-v2: Infinite-length film generative model, 2025
2025
-
[6]
Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025
Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jia...
2025
-
[7]
Lmdeploy: A toolkit for compressing, deploying, and serving llm
LMDeploy Contributors. Lmdeploy: A toolkit for compressing, deploying, and serving llm. https://github.com/InternLM/lmdeploy, 2023
2023
-
[8]
High-level features for movie style understanding
Robin Courant, Christophe Lino, Marc Christie, and Vicky Kalogeiton. High-level features for movie style understanding. Le Centre pour la Communication Scientifique Directe - HAL - Diderot,Le Centre pour la Communication Scientifique Directe - HAL - Diderot, Oct 2021
work page 2021
Show all 64 references
-
[9]
One-minute video generation with test-time training, 2025
Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, and Xiaolong Wang. One-minute video generation with test-time training, 2025
2025
-
[10]
Conmo: Controllable motion disentanglement and recomposition for zero-shot motion transfer, 2025
Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng, Kongming Liang, Zhanyu Ma, Jun Guo, and Yang Liu. Conmo: Controllable motion disentanglement and recomposition for zero-shot motion transfer, 2025
2025
-
[11]
ImageInWords: Unlocking Hyper-Detailed Image Descriptions
Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Ya- sumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Michael Baldridge, and Radu Soricut. ImageInWords: Unlocking Hyper-Detailed Image Descriptions. In Proc. EMNLP 2024, pages 93–127, Miami, Flo...
2024
-
[12]
Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Team GLM, :, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Jingyu Sun, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zh...
2024
-
[13]
Animatediff: Animate your personalized text-to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. In ICLR, 2024
2024
-
[14]
Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025
Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025
2025
-
[15]
Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model, 2025
Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, Xianfang Zeng, Xinhao Zhang, Gang Yu, Yuhe Yin, Qiling Wu, Wen Sun, Kang An, Xin Han, Deshan Sun, Wei Ji, Bizhu Huang, Brian Li, Chenfei Wu, Guanzhe Huang, Hu...
2025
-
[16]
Movienet: A holistic dataset for movie understanding
Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. Movienet: A holistic dataset for movie understanding. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan- Michael Frahm, editors, Computer Vision – ECCV 2020, pages 709–727, Cham, 2020. Springer International Publishing
2020
-
[17]
Vbench: Comprehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat...
2024
-
[18]
Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024
Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench++: Comprehensive and versatile benchmark suite for v...
2024
-
[19]
Hunyuanvideo: A systematic framework for large video generative models, 2025
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...
2025
-
[20]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention, 2023
2023
-
[21]
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024
2024
-
[22]
LLaV A-onevision: Easy visual task transfer
Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. LLaV A-onevision: Easy visual task transfer. Transactions on Machine Learning Research, 2025
2025
-
[23]
Realcam-i2v: Real-world image-to-video generation with interactive complex camera control, 2025
Teng Li, Guangcong Zheng, Rui Jiang, Shuigenzhan, Tao Wu, Yehao Lu, Yining Lin, and Xi Li. Realcam-i2v: Real-world image-to-video generation with interactive complex camera control, 2025. 11
2025
-
[24]
Evaluation of text-to-video generation models: A dynamics perspective
Mingxiang Liao, Hannan Lu, Xinyu Zhang, Fang Wan, Tianyu Wang, Yuzhong Zhao, Wang- meng Zuo, Qixiang Ye, and Jingdong Wang. Evaluation of text-to-video generation models: A dynamics perspective. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zha...
2024
-
[25]
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summariza- tion Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics
2004
-
[26]
Towards understanding camera motions in any video, 2025
Zhiqiu Lin, Siyuan Cen, Daniel Jiang, Jay Karhade, Hewei Wang, Chancharik Mitra, Tiffany Ling, Yuhan Huang, Sifan Liu, Mingyu Chen, Rushikesh Zawar, Xue Bai, Yilun Du, Chuang Gan, and Deva Ramanan. Towards understanding camera motions in any video, 2025
2025
-
[27]
What is a good caption? a comprehensive visual caption benchmark for evaluating both correctness and thoroughness, 2025
Zhihang Liu, Chen-Wei Xie, Bin Wen, Feiwu Yu, Jixuan Chen, Boqiang Zhang, Nianzu Yang, Pandeng Li, Yinglu Li, Zuan Gao, Yun Zheng, and Hongtao Xie. What is a good caption? a comprehensive visual caption benchmark for evaluating both correctness and thoroughness, 2025
2025
-
[28]
T2vsafetybench: Evaluating the safety of text-to-video generative models
Yibo Miao, Yifan Zhu, Lijia Yu, Jun Zhu, Xiao-Shan Gao, and Yinpeng Dong. T2vsafetybench: Evaluating the safety of text-to-video generative models. In A. Globerson, L. Mackey, D. Bel- grave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Pr...
2024
-
[29]
Gpt-4 technical report, 2024
OpenAI. Gpt-4 technical report, 2024
2024
-
[30]
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , ...
2002
-
[31]
Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Ja- gadeesh, Kunpeng Li, ...
2025
-
[32]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...
2021
-
[33]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila an...
2021
-
[34]
A unified framework for shot type classification based on subject centric lens
Anyi Rao, Jiaze Wang, Linning Xu, Xuekun Jiang, Qingqiu Huang, Bolei Zhou, and Dahua Lin. A unified framework for shot type classification based on subject centric lens. In Andrea 12 Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 20...
2020
-
[35]
Seaweed-7b: Cost-effective training of video generation foundation model, 2025
Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, Feng Cheng, Feilong Zuo Xuejiao Zeng, Ziyan Yang, Fangyuan Kong, Zhiwu Qing, Fei Xiao, Meng Wei, Tuyen Hoang, Siyu Zhang, Peihao Zhu, Qi Zhao, Jiangqiao Yan, Lia...
2025
-
[36]
Simões, Jônatas Wehrmann, Rodrigo C
Gabriel S. Simões, Jônatas Wehrmann, Rodrigo C. Barros, and Duncan D. Ruiz. Movie genre classification with convolutional neural networks. In 2016 International Joint Conference on Neural Networks (IJCNN), pages 259–266, 2016
2016
-
[37]
Moviechat: From dense token to sparse memory for long video understanding
Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt...
2024
-
[38]
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. Movieqa: Understanding stories in movies through question-answering. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4631–4640, 2016
2016
-
[39]
Gemma 3 technical report, 2025
Gemma Team. Gemma 3 technical report, 2025
2025
-
[40]
Kimi-vl technical report, 2025
Kimi Team. Kimi-vl technical report, 2025
2025
-
[41]
The llama 3 herd of models, 2024
Llama3 Team. The llama 3 herd of models, 2024
2024
-
[42]
Phi-3 technical report: A highly capable language model locally on your phone, 2024
Phi-3 Team. Phi-3 technical report: A highly capable language model locally on your phone, 2024
2024
-
[43]
Moviegraphs: Towards understanding human-centric situations from videos
Paul Vicol, Makarand Tapaswi, Lluis Castrejon, and Sanja Fidler. Moviegraphs: Towards understanding human-centric situations from videos. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[44]
Motionctrl: A unified and flexible motion controller for video generation
Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers, SIGGRAPH ’24, New York, NY , USA, 2024. Association for Co...
2024
-
[45]
Wan: Open and advanced large-scale video generative models, 2025
WanTeam, :, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pan...
2025
-
[46]
Cinematography, April 2025
Wikipedia. Cinematography, April 2025. Page Version ID: 1286012571
2025
-
[47]
Qwen2.5-omni technical report, 2025
Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. Qwen2.5-omni technical report, 2025. 13
2025
-
[48]
Minicpm-v: A gpt-4v level mllm on your phone, 2024
Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun. M...
2024
-
[49]
NUW A-XL: Diffusion over diffusion for eXtremely long video generation
Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, and Nan Duan. NUW A-XL: Diffusion over diffusion for eXtremely long video generatio...
2023
-
[50]
Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation
Shenghai Yuan, Jinfa Huang, Yongqi Xu, YaoYang Liu, Shaofeng Zhang, Yujun Shi, Rui- Jie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation. In The Thirty-eight Conference on Neural Informa...
2024
-
[51]
MonST3r: A simple approach for estimating geometry in the presence of motion
Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. MonST3r: A simple approach for estimating geometry in the presence of motion. In The Thirteenth International Conference on Learning Representa- tions, 2025
2025
-
[52]
Packing input frame context in next-frame prediction models for video generation, 2025
Lvmin Zhang and Maneesh Agrawala. Packing input frame context in next-frame prediction models for video generation, 2025
2025
-
[53]
Generative ai for film creation: A survey of recent advances, 2025
Ruihan Zhang, Borou Yu, Jiajian Min, Yetong Xin, Zheng Wei, Juncheng Nemo Shi, Mingzhen Huang, Xianghao Kong, Nix Liu Xin, Shanshan Jiang, Praagya Bahuguna, Mark Chan, Khushi Hora, Lijian Yang, Yongqi Liang, Runhe Bian, Yunlei Liu, Isabela Campillo Valencia, Patri- cia Morales...
2025
-
[54]
Cami2v: Camera- controlled image-to-video diffusion model, 2024
Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera- controlled image-to-video diffusion model, 2024
2024
-
[55]
Gonzalez, Clark Barrett, and Ying Sheng
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. Sglang: Efficient execution of structured language model programs, 2024
2024
-
[56]
Open-sora: Democratizing efficient video production for all, 2024
Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, 2024
2024
-
[57]
Magicvideo: Efficient video generation with latent diffusion models, 2023
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models, 2023
2023
-
[58]
Karandikar, and James M
Howard Zhou, Tucker Hermans, Asmita V . Karandikar, and James M. Rehg. Movie genre classification via scene categorization. InProceedings of the 18th ACM International Conference on Multimedia, MM ’10, page 747–750, New York, NY , USA, 2010. Association for Computing Machinery
2010
-
[59]
Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...
2025
-
[60]
• Atmosphere Description: A brief description of the mood or feeling conveyed
Description Structure: These generated descriptions generally follow a standard structure: • Scene Description: A general depiction of the visual scene. • Atmosphere Description: A brief description of the mood or feeling conveyed. • Cinematographic Technique Analysis: An anal...
-
[61]
• Cross-reference with the context or plot summary of the film to ensure accuracy
Scene and Atmosphere Verification: • Review the scene and atmosphere descriptions. • Cross-reference with the context or plot summary of the film to ensure accuracy. • Make necessary corrections for clarity, factual accuracy, and alignment with the scene
-
[62]
• Remove any unnecessary or inaccurate techniques
Technique Analysis Refinement: • Verify that the description covers all relevant cinematographic techniques. • Remove any unnecessary or inaccurate techniques. • Ensure that all technical terms align with the predefined standardized taxonomy
-
[63]
• Cross-check with film critique websites to ensure the effects are consistent with expert interpretations
Effect Explanation Correction: • Refine the explanation of the effects generated by the identified techniques. • Cross-check with film critique websites to ensure the effects are consistent with expert interpretations
-
[64]
{caption}
Final Review: • Ensure the description is coherent, grammatically correct, and accurately represents the visual content. • Submit the refined description. Quality Control • Each refined description will be reviewed by a senior annotator for quality assurance. • Descriptions fa...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.