Pith. sign in

REVIEW 4 major objections 5 minor 2 cited by

CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read CineTechBench, an expert-annotated benchmark spanning seven cinematography dimensions, finds that 15+ multimodal models and 5+ video generators all fall short at film technique, both in understanding and in reproduction.

desk verdict A useful expert-annotated benchmark for cinematographic understanding, with a generation-evaluation section whose quantitative claims currently rest on an unvalidated trajectory estimator. read the letter →

arxiv 2505.15145 v1 pith:CJZLYI3F submitted 2025-05-21 cs.CV

classification cs.CV
keywords cinematographymultimodallargelanguagemodelsvideogenerationcameramovementbenchmarkfilmunderstandingshotscalelighting
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

CineTechBench is a new benchmark built to test whether AI systems actually grasp the craft of filmmaking, not just the plot. It spans seven cinematography dimensions — shot scale, angle, composition, camera movement, lighting, color, and focal length — using over 600 expert-annotated movie stills and 120 film clips. The paper evaluates 15+ multimodal large language models on recognizing and describing these techniques, and 5+ video generation models on reconstructing cinematic camera movements. Its results show every model class has clear gaps: the best models still mix up camera rotation direction and lighting angle, and video generators fail to reproduce fast, intense rotations. The benchmark is released as a reusable testbed for tracking progress in automated film analysis and production.

What carries the argument

The load-bearing machinery is the seven-dimension cinematography taxonomy (scale, angle, composition, movement, lighting, color, focal length) built from professional filmmaking sources and refined with expert cinematographers, together with the annotation pipeline behind it: trained annotators label each image and clip, GPT-4o drafts question-answer pairs and descriptions, and annotators refine them. On the understanding side the benchmark measures per-dimension accuracy on QA pairs and adapts the CAPability caption metrics (hit rate, precision, recall, F1) with GPT-4.1-nano scoring each description as Miss, Positive, or Negative. On the generation side it uses MonST3R to estimate camera trajectories from both the original clips and the generated videos, then compares them with rotation error, translation error, and CamMC, plus a CLIP-based frame similarity score.

What would settle it

Re-annotate a random subset — about 100 images and 30 clips — with two independent professional cinematographers who never see the original labels, and compute per-dimension agreement; then have human judges re-score a sample of 50 model captions with the same Miss/Positive/Negative rubric used by the automated judge. If agreement on lighting direction or focal length falls below roughly 0.6 kappa, or if the automated and human caption scores diverge on a meaningful share of samples, the reported model rankings would not be stable.

Watch

Extended reading notes

Core claim

The paper's central claim is that CineTechBench is the first benchmark to cover all seven core dimensions of cinematography at once and to pair them with expert labels, and that measuring against it reveals a clear, consistent shortfall. On static image QA, the best commercial model reaches 70.16% overall accuracy while open-source models trail by about 15 points; on camera movement QA the best model reaches only 56.69%, with rotation the weakest movement type. On description generation, models mention the right dimensions often (hit rates above 80% for static dimensions) but describe them precisely only about a third to a half of the time (F1 between 30% and 50%), and camera-movement descriptions have a hit rate near 30%. For generation, image-to-video models conditioned on first and last frames produce lower rotation and translation errors than first-frame-only models, yet all tested generators struggle with high angular velocity, sometimes producing a roll in the wrong direction or no roll at all.

Load-bearing premise

The benchmark's conclusions assume the expert labels are consistent (no inter-annotator agreement is reported), the GPT-4.1-nano judge scores descriptions the way a human would, and MonST3R's estimated camera trajectories are accurate.

Editorial extensions

If this is right

  • Any future multimodal model can be plugged into the same QA and description protocols and ranked on the same seven dimensions, making cinematographic understanding a trackable metric rather than an anecdotal one.
  • The documented gaps — lighting direction, focal length, and camera rotation direction — give a concrete agenda for model builders, since these are the dimensions where visual encoders and training data are weakest.
  • For video generation, the results imply that camera control, especially angular motion such as high-amplitude rolls, is the main bottleneck separating current generators from cinema-quality output, and that first-frame-plus-last-frame conditioning is only a partial fix.
  • The near-30% hit rate for camera-movement descriptions shows that dynamic cinematography is disproportionately harder for multimodal models than static dimensions, pointing to video-specific perception as a distinct weakness.
  • Because the benchmark is public, its numbers can serve as a baseline that later models are expected to exceed, giving the community a shared yardstick for film-level understanding and generation.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report inter-annotator agreement, so a reader extending it should first check how stable the labels are: a small re-annotation study by independent cinematographers would tell whether the per-dimension accuracy ordering is trustworthy.
  • The description rankings depend on GPT-4.1-nano agreeing with human judgment about whether a caption is correct; a human side-by-side study on a sample of captions would reveal whether the automated judge's Miss/Positive/Negative calls skew the reported F1 scores.
  • The generation numbers inherit any error in MonST3R's trajectory estimates, which the paper acknowledges; obtaining clips with true camera motion metadata, for example from virtual production or on-set motion capture, would sharpen the rotation and translation error measurements.
  • A natural next experiment suggested by the results is to train or fine-tune a video generator on prompts that explicitly name rotation direction and amplitude, such as 'roll clockwise', and test on the same 120 clips to see whether the documented rotation failures shrink.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. CineTechBench is a new benchmark for evaluating multimodal LLMs and video generation models on seven cinematography dimensions: shot scale, shot angle, composition, camera movement, lighting, color, and focal length. It comprises over 600 expert-annotated movie stills and over 120 movie clips, with QA pairs and captions that were initially generated by GPT-4o and then manually reviewed. The understanding task reports per-dimension QA accuracy for 15+ MLLMs and CAPability-style description metrics; the generation task reports RotErr, TransErr, CamMC, and CLIP-IS for 5+ image-to-video models, where trajectory metrics are computed from MonST3R-estimated camera paths. The central findings are that MLLMs are weakest on camera rotation and lighting direction, and that video generation models perform worst under high-amplitude rotation.

Significance. The benchmark addresses a real gap: existing movie benchmarks focus on high-level semantics, while professional cinematography labels are scarce. The seven-dimension taxonomy is a reasonable synthesis of standard film terminology, the model coverage is broad, and the paper includes qualitative examples that support the broad qualitative conclusions (e.g., rotation direction errors in Figure 4 and failure to produce roll in Figure 5). The authors are transparent about limitations in Appendix F and release code and dataset links, which is a clear strength. The quantitative generation claims and part of the description claims, however, rest on unvalidated pseudo-ground-truth: MonST3R trajectories for camera motion and GPT-4.1-nano judgments for caption scoring. These issues are load-bearing for the paper's quantitative conclusions, though they appear fixable within a revision; the relative finding that models struggle with fine-grained rotation and lighting direction is internally consistent and plausible.

major comments (4)
  1. [Section 4.2, Equations (1)-(2), Table 4, Figure 6] The generation conclusion that models 'struggle with camera movement with intense rotation amplitude' and the Table 4 rankings rest on RotErr, TransErr, and CamMC computed from MonST3R trajectories, and Appendix F.1 concedes that such pose estimates often introduce inaccuracies for dynamic scenes, motion blur, and non-rigid object motion. No validation of MonST3R against known camera paths is reported, and the highlighted failure regime (fast rotations) is precisely where rotation estimation is most fragile. Because the original and generated clips differ in content, estimator errors need not cancel and could systematically inflate errors for particular models or motion types; the quantitative ordering in Table 4 and the angular-velocity trend in Figure 6 may therefore reflect estimator sensitivity rather than generation quality. To support these claims, please validate the estimator on sequences with known trajectories (including fast rolls) or add a human evaluation of rotation direction and amplitude on a subset of generations. The qualitative examples in Figure 5 support a broad claim of poor rotation generation but not the amplitude-dependent conclusion.
  2. [Section 3.2 and Appendix B] The annotation pipeline is described as expert-annotated with ambiguous cases discarded, but no inter-annotator agreement statistics are reported. Since every QA accuracy number in Tables 1 and 2 is measured against these labels, label noise could affect the relative ordering of models, particularly in boundary-prone dimensions such as shot scale (e.g., Medium Close-Up versus Close-Up) and lighting direction (Side versus Back versus Top). Please report per-dimension agreement (e.g., Cohen's kappa or Krippendorff's alpha) on a sample of the annotations, or quantitatively describe the adjudication protocol, to substantiate the claim of 'precise, manual annotation' in the abstract.
  3. [Appendix C.1, Tables 3 and 6] The description metrics (HR, AP, AR, F1) are computed by a GPT-4.1-nano judge that classifies each caption as Miss, Positive, or Negative, but the judge's agreement with human judgments is not measured. The unusually low Movement hit rate (29.83%) and the F1 differences that determine model rankings could in principle be artifacts of judge bias. Please validate the judge on a human-rated sample (e.g., 50-100 captions per dimension), report the agreement, and consider providing the per-caption judge outputs alongside the released benchmark.
  4. [Section 3.2] The understanding test items were generated by GPT-4o and then manually reviewed, but the paper does not quantify how much of the GPT-4o output was changed by human reviewers. Because GPT-4o is itself among the evaluated models, the benchmark may inadvertently reward agreement with GPT-4o's annotation style rather than with expert cinematography. Please report the modification or rewrite rate, the number of items substantively revised, and the criteria reviewers used to override or confirm GPT-4o's content; this information is also needed to support the description of the benchmark as based on manual expert annotation.
minor comments (5)
  1. [Throughout the manuscript] There are numerous typographical errors and inconsistent naming conventions (e.g., 'flim clips' in the Introduction, 'ineatographic' in Section 2.1, 'LLaV A' in tables and text, and 'Monst3r' in the Figure 5 caption); a careful proofreading pass is needed.
  2. [Table 4] The table header 'Rel. Abs. Rel. Abs.' is ambiguous; please explicitly state that relative and absolute variants apply to RotError and TransError, and define the normalization used for each variant in the caption or in Appendix C.2.
  3. [Equation (1)] The definition of RotErr sums per-frame angular errors but does not specify n or how the frames of the original and generated clips are temporally aligned after downsampling; please state the alignment procedure and the frame count used.
  4. [Figure 6] The x-axis quantities 'translation speed' and 'angular velocity' are not defined; please provide units, binning details, and the number of test clips per bin so that the trend can be interpreted.
  5. [Section 4.2, Equation (3)] The CLIP-IS metric compares original and generated frames, so it partly reflects content similarity rather than camera-movement fidelity; the text should explicitly caveat that CLIP-IS is not a pure motion-quality measure.

Circularity Check

0 steps flagged · score 1.0 of 10

No load-bearing circularity; evaluation-validity caveats (GPT-4o-generated items, LLM judge, MonST3R trajectories) are not by-construction reductions.

full rationale

CineTechBench is a benchmark-and-evaluation paper, not a derivation. The central claims—that current MLLMs underperform on cinematographic understanding and that video generators struggle with rotation—are empirical measurements obtained by scoring models against human-reviewed labels and external trajectory estimates. The strongest candidates for circularity do not survive scrutiny. (1) Test items were originally drafted by GPT-4o, one of the evaluated models, but Section 3.2 states 'All generated content was manually reviewed and refined by trained annotators to ensure accuracy, clarity, and alignment with professional cinematography standards,' so the labels are not GPT-4o's outputs by construction. (2) Description scoring uses GPT-4.1-nano as an automated judge (Appendix C.1); the judge is not the model being scored, and the classification categories (Miss/Positive/Negative) are anchored to human ground truth, so no equation reduces to the evaluated model's output. (3) Camera-motion metrics use MonST3R [51] to estimate trajectories of both original and generated clips (Eq. 1–2). Appendix F.1 concedes such estimates 'often introduce inaccuracies due to complex cinematographic factors such as dynamic scenes, motion blur, and non-rigid object motion'; this is a real measurement-validity limitation that could affect model rankings in Table 4 and Figure 6, but it is not circular because the estimator is external to the evaluated models and no fitted parameter is renamed as a prediction. The only overlapping-author citation, [10] (Conmo, which includes co-authors Liang and Ma), appears in related work as an example of video-generation advances and is not load-bearing; no uniqueness theorem from the authors is invoked. No pattern of self-definitional labels, fitted inputs called predictions, or ansatz smuggled via citation is present. Accordingly, the circularity score is minimal; the flagged limitations belong to correctness risk rather than circular reasoning.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The benchmark contribution is a dataset and evaluation protocol rather than a derived physical result. No numeric free parameters are fit. The main assumptions are about annotation reliability, the pseudo-ground-truth camera trajectories, and the automated LLM judge.

assumptions (3)
  • domain assumption Expert annotations and GPT-4o-generated, human-reviewed QA and descriptions are accurate ground truth for cinematographic techniques.
    Invoked in Section 3.2 and used for all accuracy and CAPability metrics; no inter-annotator agreement or expert double-check rate is reported.
  • domain assumption MonST3R-estimated camera trajectories from movie clips are accurate enough to serve as pseudo ground truth for evaluating generated camera motion.
    Used in Section 4.2 and Appendix C.2; Appendix F.1 concedes these estimates introduce inaccuracies from motion blur, dynamic scenes, and non-rigid motion.
  • domain assumption GPT-4.1-nano prompt classifications (Miss, Positive, Negative) agree with human judgments for CAPability-style metrics.
    Appendix C.1 and Figure 11 define the automated judging protocol; no human agreement study for the adapted cinematography dimensions is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation." pith.science (2026). https://pith.science/paper/CJZLYI3F

@misc{pith2026250515145,
  author       = {Pith},
  title        = {Pith review of: CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CJZLYI3F}},
  note         = {Machine review of arXiv:2505.15145}
}
read the original abstract

Cinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects-shot scale, shot angle, composition, camera movement, lighting, color, and focal length-and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question answer pairs and annotated descriptions to assess MLLMs' ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatically film production and appreciation. The code and benchmark can be accessed at https://github.com/PRIS-CV/CineTechBench.

Figures

Figures reproduced from arXiv: 2505.15145 by the authors.

Figure 1
Figure 1. Cinematography taxonomy and data examples in our CineTechBench. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our benchmark focus on the cinematographic techniques in film production and ap￾preciation. Compared with similar benchmarks, our benchmark include more core dimensions in cinematography. of benchmarks tailored specifically for evaluating cinematographic techniques, particularly camera movement, in generated video content. Our benchmark fills this gap by focusing on the assessment of cinema-level camera movement gen… view at source ↗
Figure 3
Figure 3. Overview of our benchmark building process. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Visualization of MLLMs’ answers on cinematographic technique question answering task. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Generated movie clips by different video generation models and the corresponding camera [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Average TransError and RotError on different translation speed and angular velocity. Results The overall results are shown in Ta￾ble 4. Among the commercial video genera￾tion model, Klingv1.6 with first and last frame control achieve the best performance on both RotErr…
Figure 7
Figure 7. Figure 7: Statistical and semantic overview of the CineTechBench. [PITH_FULL_IMAGE:figures/full_fig_p015_7.png]
Figure 8
Figure 8. Figure 8: Illustration of categories in the angle and scale dimension. [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: An annotation refine example for MLLM generated description. [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: An example label interface. C Evaluation Metrics C.1 Description Evaluation Metrics Inspired by the CAPability benchmark [27], which proposes a comprehensive framework to evaluate the correctness and thoroughness of visual captions, we adopt a similar metric design to…
Figure 11
Figure 11. Figure 11: Prompt template used for static dimension evaluation (e.g., [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Average hit rate (HR) and F1 score of all MLLMs on seven dimensions of cinematographic [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Visualization of MLLMs’ answers on image cinematographic technique question answer [PITH_FULL_IMAGE:figures/full_fig_p023_13.png]
Figure 14
Figure 14. Figure 14: Visualization of MLLMs’ answers on video question-answering task and generated [PITH_FULL_IMAGE:figures/full_fig_p024_14.png]
Figure 15
Figure 15. Figure 15: Generated movie clips by different video generation models and the corresponding camera [PITH_FULL_IMAGE:figures/full_fig_p025_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A film-academy cinematic taxonomy and reverse-engineered multi-shot prompts expose large gaps in leading video generators that web-style benchmarks miss.

  2. Natural Language Camera Movement Understanding

    cs.CV 2026-07 accept novelty 6.5 of 10

    A cinematographic taxonomy, atomic real+synthetic benchmark (ACaM), and targeted-augmentation SFT let an 8B VLM outperform Gemini-3.1-Pro by 10-11% on camera-movement recognition, yet a large human gap remains.

Reference graph

Works this paper leans on

64 extracted references · 41 canonical work pages · cited by 2 Pith papers

  1. [1]

    Qwen2.5-vl technical report, 2025

    Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025

  2. [2]

    METEOR: An automatic metric for MT evaluation with improved correlation with human judgments

    Satanjeev Banerjee and Alon Lavie. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Jade Goldstein, Alon Lavie, Chin-Yew Lin, and Clare V oss, editors,Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pages 65–72, Ann Arbor, Michigan,...

  3. [3]

    Align your latents: High-resolution video synthesis with latent diffusion models

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 22563–22575, 2023

  4. [4]

    Movieclip: Visual scene recognition in movies

    Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Haoyang Zhang, Yin Cui, Kree Cole- McLaughlin, Huisheng Wang, and Shrikanth Narayanan. Movieclip: Visual scene recognition in movies. In 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 2082–2091, 2023

  5. [5]

    Skyreels-v2: Infinite-length film generative model, 2025

    Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengcheng Ma, Weiming Xiong, Wei Wang, Nuo Pang, Kang Kang, Zhiheng Xu, Yuzhe Jin, Yupeng Liang, Yubing Song, Peng Zhao, Boyuan Xu, Di Qiu, Debang Li, Zhengcong Fei, Yang Li, and Yahui Zhou. Skyreels-v2: Infinite-length film generative model, 2025

  6. [6]

    Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

    Zhe Chen, Weiyun Wang, Yue Cao, Yangzhou Liu, Zhangwei Gao, Erfei Cui, Jinguo Zhu, Shenglong Ye, Hao Tian, Zhaoyang Liu, Lixin Gu, Xuehui Wang, Qingyun Li, Yimin Ren, Zixuan Chen, Jiapeng Luo, Jiahao Wang, Tan Jiang, Bo Wang, Conghui He, Botian Shi, Xingcheng Zhang, Han Lv, Yi Wang, Wenqi Shao, Pei Chu, Zhongying Tu, Tong He, Zhiyong Wu, Huipeng Deng, Jia...

  7. [7]

    Lmdeploy: A toolkit for compressing, deploying, and serving llm

    LMDeploy Contributors. Lmdeploy: A toolkit for compressing, deploying, and serving llm. https://github.com/InternLM/lmdeploy, 2023

  8. [8]

    High-level features for movie style understanding

    Robin Courant, Christophe Lino, Marc Christie, and Vicky Kalogeiton. High-level features for movie style understanding. Le Centre pour la Communication Scientifique Directe - HAL - Diderot,Le Centre pour la Communication Scientifique Directe - HAL - Diderot, Oct 2021

Show all 64 references
  1. [9]

    One-minute video generation with test-time training, 2025

    Karan Dalal, Daniel Koceja, Gashon Hussein, Jiarui Xu, Yue Zhao, Youjin Song, Shihao Han, Ka Chun Cheung, Jan Kautz, Carlos Guestrin, Tatsunori Hashimoto, Sanmi Koyejo, Yejin Choi, Yu Sun, and Xiaolong Wang. One-minute video generation with test-time training, 2025

  2. [10]

    Conmo: Controllable motion disentanglement and recomposition for zero-shot motion transfer, 2025

    Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng, Kongming Liang, Zhanyu Ma, Jun Guo, and Yang Liu. Conmo: Controllable motion disentanglement and recomposition for zero-shot motion transfer, 2025

  3. [11]

    ImageInWords: Unlocking Hyper-Detailed Image Descriptions

    Roopal Garg, Andrea Burns, Burcu Karagol Ayan, Yonatan Bitton, Ceslee Montgomery, Ya- sumasa Onoe, Andrew Bunner, Ranjay Krishna, Jason Michael Baldridge, and Radu Soricut. ImageInWords: Unlocking Hyper-Detailed Image Descriptions. In Proc. EMNLP 2024, pages 93–127, Miami, Flo...

  4. [12]

    Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

    Team GLM, :, Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Dan Zhang, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Jingyu Sun, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zh...

  5. [13]

    Animatediff: Animate your personalized text-to-image diffusion models without specific tuning

    Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text-to-image diffusion models without specific tuning. In ICLR, 2024

  6. [14]

    Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025

    Teng Hu, Zhentao Yu, Zhengguang Zhou, Sen Liang, Yuan Zhou, Qin Lin, and Qinglin Lu. Hunyuancustom: A multimodal-driven architecture for customized video generation, 2025

  7. [15]

    Step-video-ti2v technical report: A state-of-the-art text-driven image-to-video generation model, 2025

    Haoyang Huang, Guoqing Ma, Nan Duan, Xing Chen, Changyi Wan, Ranchen Ming, Tianyu Wang, Bo Wang, Zhiying Lu, Aojie Li, Xianfang Zeng, Xinhao Zhang, Gang Yu, Yuhe Yin, Qiling Wu, Wen Sun, Kang An, Xin Han, Deshan Sun, Wei Ji, Bizhu Huang, Brian Li, Chenfei Wu, Guanzhe Huang, Hu...

  8. [16]

    Movienet: A holistic dataset for movie understanding

    Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. Movienet: A holistic dataset for movie understanding. In Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan- Michael Frahm, editors, Computer Vision – ECCV 2020, pages 709–727, Cham, 2020. Springer International Publishing

  9. [17]

    Vbench: Comprehensive benchmark suite for video generative models

    Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat...

  10. [18]

    Vbench++: Comprehensive and versatile benchmark suite for video generative models, 2024

    Ziqi Huang, Fan Zhang, Xiaojie Xu, Yinan He, Jiashuo Yu, Ziyue Dong, Qianli Ma, Nattapol Chanpaisit, Chenyang Si, Yuming Jiang, Yaohui Wang, Xinyuan Chen, Ying-Cong Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench++: Comprehensive and versatile benchmark suite for v...

  11. [19]

    Hunyuanvideo: A systematic framework for large video generative models, 2025

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, J...

  12. [20]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with pagedattention, 2023

  13. [21]

    Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

    Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, and Chunyuan Li. Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024

  14. [22]

    LLaV A-onevision: Easy visual task transfer

    Bo Li, Yuanhan Zhang, Dong Guo, Renrui Zhang, Feng Li, Hao Zhang, Kaichen Zhang, Peiyuan Zhang, Yanwei Li, Ziwei Liu, and Chunyuan Li. LLaV A-onevision: Easy visual task transfer. Transactions on Machine Learning Research, 2025

  15. [23]

    Realcam-i2v: Real-world image-to-video generation with interactive complex camera control, 2025

    Teng Li, Guangcong Zheng, Rui Jiang, Shuigenzhan, Tao Wu, Yehao Lu, Yining Lin, and Xi Li. Realcam-i2v: Real-world image-to-video generation with interactive complex camera control, 2025. 11

  16. [24]

    Evaluation of text-to-video generation models: A dynamics perspective

    Mingxiang Liao, Hannan Lu, Xinyu Zhang, Fang Wan, Tianyu Wang, Yuzhong Zhao, Wang- meng Zuo, Qixiang Ye, and Jingdong Wang. Evaluation of text-to-video generation models: A dynamics perspective. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zha...

  17. [25]

    ROUGE: A package for automatic evaluation of summaries

    Chin-Yew Lin. ROUGE: A package for automatic evaluation of summaries. In Text Summariza- tion Branches Out, pages 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics

  18. [26]

    Towards understanding camera motions in any video, 2025

    Zhiqiu Lin, Siyuan Cen, Daniel Jiang, Jay Karhade, Hewei Wang, Chancharik Mitra, Tiffany Ling, Yuhan Huang, Sifan Liu, Mingyu Chen, Rushikesh Zawar, Xue Bai, Yilun Du, Chuang Gan, and Deva Ramanan. Towards understanding camera motions in any video, 2025

  19. [27]

    What is a good caption? a comprehensive visual caption benchmark for evaluating both correctness and thoroughness, 2025

    Zhihang Liu, Chen-Wei Xie, Bin Wen, Feiwu Yu, Jixuan Chen, Boqiang Zhang, Nianzu Yang, Pandeng Li, Yinglu Li, Zuan Gao, Yun Zheng, and Hongtao Xie. What is a good caption? a comprehensive visual caption benchmark for evaluating both correctness and thoroughness, 2025

  20. [28]

    T2vsafetybench: Evaluating the safety of text-to-video generative models

    Yibo Miao, Yifan Zhu, Lijia Yu, Jun Zhu, Xiao-Shan Gao, and Yinpeng Dong. T2vsafetybench: Evaluating the safety of text-to-video generative models. In A. Globerson, L. Mackey, D. Bel- grave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors,Advances in Neural Information Pr...

  21. [29]

    Gpt-4 technical report, 2024

    OpenAI. Gpt-4 technical report, 2024

  22. [30]

    Bleu: a method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for automatic evaluation of machine translation. In Pierre Isabelle, Eugene Charniak, and Dekang Lin, editors, Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , ...

  23. [31]

    Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

    Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Ja- gadeesh, Kunpeng Li, ...

  24. [32]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pa...

  25. [33]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila an...

  26. [34]

    A unified framework for shot type classification based on subject centric lens

    Anyi Rao, Jiaze Wang, Linning Xu, Xuekun Jiang, Qingqiu Huang, Bolei Zhou, and Dahua Lin. A unified framework for shot type classification based on subject centric lens. In Andrea 12 Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm, editors, Computer Vision – ECCV 20...

  27. [35]

    Seaweed-7b: Cost-effective training of video generation foundation model, 2025

    Team Seawead, Ceyuan Yang, Zhijie Lin, Yang Zhao, Shanchuan Lin, Zhibei Ma, Haoyuan Guo, Hao Chen, Lu Qi, Sen Wang, Feng Cheng, Feilong Zuo Xuejiao Zeng, Ziyan Yang, Fangyuan Kong, Zhiwu Qing, Fei Xiao, Meng Wei, Tuyen Hoang, Siyu Zhang, Peihao Zhu, Qi Zhao, Jiangqiao Yan, Lia...

  28. [36]

    Simões, Jônatas Wehrmann, Rodrigo C

    Gabriel S. Simões, Jônatas Wehrmann, Rodrigo C. Barros, and Duncan D. Ruiz. Movie genre classification with convolutional neural networks. In 2016 International Joint Conference on Neural Networks (IJCNN), pages 259–266, 2016

  29. [37]

    Moviechat: From dense token to sparse memory for long video understanding

    Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, et al. Moviechat: From dense token to sparse memory for long video understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patt...

  30. [38]

    Movieqa: Understanding stories in movies through question-answering

    Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. Movieqa: Understanding stories in movies through question-answering. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 4631–4640, 2016

  31. [39]

    Gemma 3 technical report, 2025

    Gemma Team. Gemma 3 technical report, 2025

  32. [40]

    Kimi-vl technical report, 2025

    Kimi Team. Kimi-vl technical report, 2025

  33. [41]

    The llama 3 herd of models, 2024

    Llama3 Team. The llama 3 herd of models, 2024

  34. [42]

    Phi-3 technical report: A highly capable language model locally on your phone, 2024

    Phi-3 Team. Phi-3 technical report: A highly capable language model locally on your phone, 2024

  35. [43]

    Moviegraphs: Towards understanding human-centric situations from videos

    Paul Vicol, Makarand Tapaswi, Lluis Castrejon, and Sanja Fidler. Moviegraphs: Towards understanding human-centric situations from videos. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018

  36. [44]

    Motionctrl: A unified and flexible motion controller for video generation

    Zhouxia Wang, Ziyang Yuan, Xintao Wang, Yaowei Li, Tianshui Chen, Menghan Xia, Ping Luo, and Ying Shan. Motionctrl: A unified and flexible motion controller for video generation. In ACM SIGGRAPH 2024 Conference Papers, SIGGRAPH ’24, New York, NY , USA, 2024. Association for Co...

  37. [45]

    Wan: Open and advanced large-scale video generative models, 2025

    WanTeam, :, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, Jianyuan Zeng, Jiayu Wang, Jingfeng Zhang, Jingren Zhou, Jinkai Wang, Jixuan Chen, Kai Zhu, Kang Zhao, Keyu Yan, Lianghua Huang, Mengyang Feng, Ningyi Zhang, Pan...

  38. [46]

    Cinematography, April 2025

    Wikipedia. Cinematography, April 2025. Page Version ID: 1286012571

  39. [47]

    Qwen2.5-omni technical report, 2025

    Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, Bin Zhang, Xiong Wang, Yunfei Chu, and Junyang Lin. Qwen2.5-omni technical report, 2025. 13

  40. [48]

    Minicpm-v: A gpt-4v level mllm on your phone, 2024

    Yuan Yao, Tianyu Yu, Ao Zhang, Chongyi Wang, Junbo Cui, Hongji Zhu, Tianchi Cai, Haoyu Li, Weilin Zhao, Zhihui He, Qianyu Chen, Huarong Zhou, Zhensheng Zou, Haoye Zhang, Shengding Hu, Zhi Zheng, Jie Zhou, Jie Cai, Xu Han, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun. M...

  41. [49]

    NUW A-XL: Diffusion over diffusion for eXtremely long video generation

    Shengming Yin, Chenfei Wu, Huan Yang, Jianfeng Wang, Xiaodong Wang, Minheng Ni, Zhengyuan Yang, Linjie Li, Shuguang Liu, Fan Yang, Jianlong Fu, Ming Gong, Lijuan Wang, Zicheng Liu, Houqiang Li, and Nan Duan. NUW A-XL: Diffusion over diffusion for eXtremely long video generatio...

  42. [50]

    Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation

    Shenghai Yuan, Jinfa Huang, Yongqi Xu, YaoYang Liu, Shaofeng Zhang, Yujun Shi, Rui- Jie Zhu, Xinhua Cheng, Jiebo Luo, and Li Yuan. Chronomagic-bench: A benchmark for metamorphic evaluation of text-to-time-lapse video generation. In The Thirty-eight Conference on Neural Informa...

  43. [51]

    MonST3r: A simple approach for estimating geometry in the presence of motion

    Junyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, and Ming-Hsuan Yang. MonST3r: A simple approach for estimating geometry in the presence of motion. In The Thirteenth International Conference on Learning Representa- tions, 2025

  44. [52]

    Packing input frame context in next-frame prediction models for video generation, 2025

    Lvmin Zhang and Maneesh Agrawala. Packing input frame context in next-frame prediction models for video generation, 2025

  45. [53]

    Generative ai for film creation: A survey of recent advances, 2025

    Ruihan Zhang, Borou Yu, Jiajian Min, Yetong Xin, Zheng Wei, Juncheng Nemo Shi, Mingzhen Huang, Xianghao Kong, Nix Liu Xin, Shanshan Jiang, Praagya Bahuguna, Mark Chan, Khushi Hora, Lijian Yang, Yongqi Liang, Runhe Bian, Yunlei Liu, Isabela Campillo Valencia, Patri- cia Morales...

  46. [54]

    Cami2v: Camera- controlled image-to-video diffusion model, 2024

    Guangcong Zheng, Teng Li, Rui Jiang, Yehao Lu, Tao Wu, and Xi Li. Cami2v: Camera- controlled image-to-video diffusion model, 2024

  47. [55]

    Gonzalez, Clark Barrett, and Ying Sheng

    Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. Sglang: Efficient execution of structured language model programs, 2024

  48. [56]

    Open-sora: Democratizing efficient video production for all, 2024

    Zangwei Zheng, Xiangyu Peng, Tianji Yang, Chenhui Shen, Shenggui Li, Hongxin Liu, Yukun Zhou, Tianyi Li, and Yang You. Open-sora: Democratizing efficient video production for all, 2024

  49. [57]

    Magicvideo: Efficient video generation with latent diffusion models, 2023

    Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models, 2023

  50. [58]

    Karandikar, and James M

    Howard Zhou, Tucker Hermans, Asmita V . Karandikar, and James M. Rehg. Movie genre classification via scene categorization. InProceedings of the 18th ACM International Conference on Multimedia, MM ’10, page 747–750, New York, NY , USA, 2010. Association for Computing Machinery

  51. [59]

    Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

    Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, Zhangwei Gao, Erfei Cui, Xuehui Wang, Yue Cao, Yangzhou Liu, Xingguang Wei, Hongjie Zhang, Haomin Wang, Weiye Xu, Hao Li, Jiahao Wang, Nianchen Deng, Songze Li,...

  52. [60]

    • Atmosphere Description: A brief description of the mood or feeling conveyed

    Description Structure: These generated descriptions generally follow a standard structure: • Scene Description: A general depiction of the visual scene. • Atmosphere Description: A brief description of the mood or feeling conveyed. • Cinematographic Technique Analysis: An anal...

  53. [61]

    • Cross-reference with the context or plot summary of the film to ensure accuracy

    Scene and Atmosphere Verification: • Review the scene and atmosphere descriptions. • Cross-reference with the context or plot summary of the film to ensure accuracy. • Make necessary corrections for clarity, factual accuracy, and alignment with the scene

  54. [62]

    • Remove any unnecessary or inaccurate techniques

    Technique Analysis Refinement: • Verify that the description covers all relevant cinematographic techniques. • Remove any unnecessary or inaccurate techniques. • Ensure that all technical terms align with the predefined standardized taxonomy

  55. [63]

    • Cross-check with film critique websites to ensure the effects are consistent with expert interpretations

    Effect Explanation Correction: • Refine the explanation of the effects generated by the identified techniques. • Cross-check with film critique websites to ensure the effects are consistent with expert interpretations

  56. [64]

    {caption}

    Final Review: • Ensure the description is coherent, grammatically correct, and accurately represents the visual content. • Submit the refined description. Quality Control • Each refined description will be reviewed by a senior annotator for quality assurance. • Descriptions fa...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.