REVIEW 3 major objections 5 minor 99 references
CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read CineVision, an AI storyboard tool that merges scriptwriting with real-time lighting, style, and character control, claims to lower filmmaker workload and raise perceived mutual understanding versus hand-drawn boards and prompt-only image…
desk verdict A genuinely useful integration and a plausible subjective-user study, but the abstract's task-time claim is unsupported by any measured data—fix or drop it before this can be accepted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a two-tier hierarchical input with weighted prompt conditioning: first-tier categories (environment, time of day, light direction, director style, actor basics) are weighted above second-tier details (facial features, hairstyle, clothing), and the total conditioning input is $W_{\mathrm{total}} = \sum_{i=1}^{n} w_i x_i$, so menu choices rather than free prose drive generation. Three learned components hang off this input structure: a diffusion image model fine-tuned on 1,000 curated film stills and given relighting ability so light can be changed without redrawing; a second fine-tune on 2,000 portraits that carries character and costume identity across shots; and an SQL-style retrieval over a film-still database, scored by visual-text similarity, that anchors each script line to a coherent reference shot group. The weighted hierarchy is what the paper credits with lowering the creation threshold and reducing prompt ambiguity, while the relighting and character modules are what it credits with visual continuity.
What would settle it
Run the same six-scene dialogue task under the three conditions while recording completion time per dyad; if hand-drawn or text-to-image pairs finish as fast or faster, the abstract's efficiency claim fails. Separately, replace the mutual self-rating of understanding with observer-coded clarification rounds or a match between final boards and an expert reading of the script, and check whether the CineVision group still leads.
Extended reading notes
Core claim
The paper's central claim is that integrating scriptwriting with generative pre-visualization, instead of piping a script through a separate image generator, changes how directors and cinematographers spend their collaborative time. On the authors' account, CineVision's combination of a hierarchical menu (which decomposes a scene into environment, lighting, time of day, director style, actor basics, then facial and costume details) with weighted prompt conditioning produces storyboards that keep character identity, lighting, and background consistent across shots. Its relighting module, built by fine-tuning a diffusion image model on 1,000 curated film stills and then adding relighting ability, lets a dyad test light direction, time of day, and director-specific looks in real time; its character module, fine-tuned on 2,000 portraits, carries facial features and costumes from shot to shot; and its film-still retrieval supplies scene anchors matched to the script. In the reported lab study, the CineVision group diverged from the baselines on subjective workload and usability, and posted the highest mutual-understanding score (mean 6.5 on a 7-point scale). The authors frame the result as a preliminary indication of easier early-stage communication, explicitly noting that perceived alignment is not yet demonstrated on-set efficiency.
Load-bearing premise
The paper's headline promise is that CineVision makes storyboarding faster, but the study never measures or reports task time, so the speed advantage stated in the abstract rests on an unverified assertion.
Editorial extensions
If this is right
- Dyads using CineVision reportedly spent their time discussing atmosphere, emotion, and style rather than redrawing or re-generating images, so the workflow shifts effort from technical adjustment to narrative discussion.
- Visual continuity across shots becomes a built-in property rather than a manual discipline: character identity and lighting set once carry forward, which the text-to-image baseline lacked and hand-drawing required repeated reference sheets to approximate.
- A shared menu vocabulary of lighting terms (soft, hard, key light; light direction; time of day) could serve as a common reference language that reduces re-explanation between directors and cinematographers.
- The authors position the system as the first step of an incremental ladder: future production-mode tiers with lens-accurate metadata, industry-standard exports, and multi-role participation would extend the same integration to the rest of the crew.
- The paper argues that AI pre-visualization will redistribute creative authority between director and cinematographer, and that the cinematographer's role may narrow to technical execution unless the tool is treated as enhancing rather than replacing their interpretive work.
Reading between the lines
- A natural next experiment would isolate the communication benefit from the generation benefit by comparing CineVision against its own components—for instance, the same weighted menu paired with a plain text-to-image model—to see how much of the workload reduction comes from the shared vocabulary and how much from image quality.
- The per-condition sample of eight participants, with only one professional director-cinematographer pair per group, means the reported differences could shift substantially with more professionals; a replication using only experienced filmmaking pairs would test whether the benefits persist where communication norms are already established.
- Because the system only handles two-character, well-lit live-action dialogue scenes drawn from a Hollywood film corpus, the generalizability question is as much about cinema as about users: period, fantasy, action, and crowd scenes would stress the lighting-coherence and costume-novelty modules that this study never exercised.
- If the paper's workload results replicate, the next bottleneck to measure is decision quality rather than effort: the collaboration score is a mutual self-rating, and an objective measure—such as whether the final storyboard matches an expert-coded reading of the script—would test whether perceived understanding tracks actual understanding.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents CineVision, an interactive pre-visualization system that combines script input, film-database frame retrieval, diffusion-based relighting and stylization, and character/costume customization, with the aim of supporting director–cinematographer collaboration during storyboarding. The authors describe a formative interview with two senior practitioners, the system's design and implementation, and a user study of 24 participants split into three groups: CineVision, DALL·E 3, and hand-drawn storyboards. The study reports NASA-TLX workload measures, a three-item UEQ, a mutual collaboration score, and qualitative interviews. The abstract claims that CineVision yielded shorter task times and higher usability ratings than the two baselines; the usability claims are supported for several self-report subscales, but no task-completion time is measured or reported anywhere in the paper.
Significance. If the usability and workload findings hold, CineVision is a meaningful contribution to AI-mediated creative collaboration in film pre-production: it integrates previously separate functions (script-to-frame retrieval, lighting control, style emulation, character design) into one system, and the paper provides an open-source implementation. The statistical approach for a small-n study (Shapiro–Wilk, Kruskal–Wallis, Dunn with Bonferroni) is appropriate, and several significant effects are reported on NASA-TLX and UEQ dimensions. The qualitative material, especially the participant quotes about reduced discussion time and the changing director–cinematographer division of labor, is valuable. However, the paper's headline efficiency claim (shorter task times) is not operationalized or measured, all quantitative outcomes are self-report, and storyboard quality is never scored objectively; the strength of the contribution is therefore restricted to perceived workload, perceived usefulness, and perceived mutual understanding until these gaps are addressed.
major comments (3)
- [Abstract; Section 5.1; Section 5.3; Section 6; Section 7.1; Section 8] The abstract's central claim that "CineVision yielded shorter task times" is unsupported by any task-completion time measurement. Section 5.1 defines dependent variables as NASA-TLX, UEQ, and observation/interviews, with no time-to-completion variable; Section 5.3 describes a fixed 50-minute task with no early-stop rule or time recording; and Section 6 reports no time measure. Nevertheless, Section 7.1 states that "participants using CineVision completed storyboarding tasks faster and with fewer modifications" and Section 8 repeats "reduced task time," both without data. The only quantitative proxy is the UEQ item "This workflow approach is a time saver" (Section 6.2), which measures perceived time savings, not elapsed time. The authors should either report the actual task-completion time data, if collected, or revise the abstract and discussion to claim only perceived time savings.
- [Section 5.1; Section 5.3; Section 6.3; Section 7.1] Every quantitative outcome reported in the evaluation is self-report: NASA-TLX, UEQ, and the Collaboration Score, which is a mutual 7-point subjective rating of perceived understanding (Section 5.3). There is no objective measure of storyboard quality, communication quality, or efficiency of the resulting storyboard. Consequently, the discussion's stronger claims—such as "real-time visualization not only accelerates pre-production but also preserves creative momentum" (Section 7.1)—go beyond what the measured variables can support. The authors should temper these claims or add objective measures (e.g., storyboard rating by blind judges, counts of modification iterations, or logged interaction data).
- [Section 5.2; Section 6.3] The group assignment yields very limited statistical power for professional-level conclusions: each of the three groups contains only one professional director–cinematographer pair (D1&C1, D2&C2, D3&C3 in Table 3), with the remaining six participants per group being amateurs. The significant collaboration-score difference (Section 6.3, χ²(2)=9.16, p<0.05) is driven by groups of eight participants each, and the paper does not report whether the one professional pair in Group A is the source of the highest score. The claim that results "particularly for new collaborators" (abstract) is speculative; the authors should either present a professional-only analysis or explicitly limit the collaboration conclusion to novice and mixed-expertise dyads.
minor comments (5)
- [Section 6.1; Section 6.2] For the significant Kruskal–Wallis tests, the post-hoc Dunn's test results are indicated only by asterisks, plus signs, and hash symbols in Table 4; exact p-values or adjusted significance levels for the pairwise comparisons should be reported.
- [Section 5.1; Section 5.3] The Collaboration Score is first introduced in Section 5.3 but is not listed among the dependent variables in Section 5.1; it should be added to the Dependent Variables subsection for consistency.
- [Section 4.1.4] Equation (1) for W_total is not numbered and the weight values w_i are never specified; please clarify how the weights are initialized and whether they were tuned or held fixed in the user study.
- [Section 6.4] The observational findings (gender differences, communication dynamics) are reported narratively without a protocol, coding scheme, or inter-rater reliability; if these observations are intended as results, they should be presented as qualitative analysis with a defined method.
- [Section 4.1.1; Section 4.1.2] Minor copy-editing issues include "chatGPT4o" inconsistencies and the phrase "role and costume creation" (Section 4.1.2) which seems to say "role" where "costume" is meant; please proofread for these and similar typos.
Circularity Check
No significant circularity: the user study is an independent evaluation; the unmeasured 'shorter task times' claim is a support gap, not a circular derivation.
full rationale
The paper's central results come from a 24-participant between-group user study comparing CineVision against DALL·E 3 and hand-drawn storyboards. The NASA-TLX, UEQ, and Collaboration Score outcomes are measured questionnaire responses, not quantities derived from system parameters, so there is no equation or fitted parameter that reduces to the target conclusion. The only equation in the paper, W_total = sum w_i x_i, is a weighted input aggregation for prompt construction, not a predictive model fitted to the study outcomes. The main self-citation surface is ScriptViz [59] and IC-Light [88], both by overlapping authors; the paper says it 'adopted the strategy proposed by Rao et al.' and 'Building upon Scriptviz [59]', and it 'follow[s] the IC-Light [88]' for relighting. These are legitimate system-building citations rather than load-bearing justifications for the paper's empirical claims, and the user-study results are not inputs to those prior systems. No uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation to prove the evaluation's conclusion. One significant non-circularity concern does appear: the abstract claims CineVision 'yielded shorter task times', but Section 5.1 lists only questionnaires and interviews as dependent variables, Section 5.3 describes a fixed 50-minute session with no time-to-completion recording, and Section 6 reports no task-time statistic. Section 7.1 then asserts that 'participants using CineVision completed storyboarding tasks faster and with fewer modifications' without citing any measured data. This is an unsupported empirical claim and a correctness/verification issue, not a definitional circularity: no measured task-time variable is being renamed or derived from its own definition. Overall, the derivation chain is self-contained and the evaluation is independent; the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (4)
- Category weights w_i in W_total =
not disclosed
- Curated dataset sizes for finetuning =
1,000 relighting images; 2,000 character images; 5 epochs
- Scene-complexity filter threshold =
6 visible characters
- Director style prompt set =
10 directors, 5 key shots each, 3 ChatGPT optimization rounds
assumptions (4)
- domain assumption Two expert interviews are sufficient to generalize design needs C1-C4 to the director-cinematographer population.
- domain assumption MovieNet dialogue frames provide adequate training coverage for cinematic relighting and character generation.
- domain assumption ChatGPT4o-generated captions are accurate enough to serve as ground truth for finetuning.
- domain assumption IC-Light relighting on the finetuned SD1.5 preserves character identity and scene coherence.
Cite this review
Pith. "Pith review of CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration." pith.science (2026). https://pith.science/paper/SYUIDF5E
@misc{pith2026250720355,
author = {Pith},
title = {Pith review of: CineVision: An Interactive Pre-visualization Storyboard System for Director-Cinematographer Collaboration},
year = {2026},
howpublished = {\url{https://pith.science/paper/SYUIDF5E}},
note = {Machine review of arXiv:2507.20355}
}
read the original abstract
Effective communication between directors and cinematographers is fundamental in film production, yet traditional approaches relying on visual references and hand-drawn storyboards often lack the efficiency and precision necessary during pre-production. We present CineVision, an AI-driven platform that integrates scriptwriting with real-time visual pre-visualization to bridge this communication gap. By offering dynamic lighting control, style emulation based on renowned filmmakers, and customizable character design, CineVision enables directors to convey their creative vision with heightened clarity and rapidly iterate on scene composition. In a 24-participant lab study, CineVision yielded shorter task times and higher usability ratings than two baseline methods, suggesting a potential to ease early-stage communication and accelerate storyboard drafts under controlled conditions. These findings underscore CineVision's potential to streamline pre-production processes and foster deeper creative synergy among filmmaking teams, particularly for new collaborators. Our code and demo are available at https://github.com/TonyHongtaoWu/CineVision.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Abubakar Abid, Ali Abdalla, Ali Abid, Dawood Khan, Abdulrahman Alfozan, and James Zou. 2019. Gradio: Hassle-free sharing and testing of ML models in the wild. arXiv preprint arXiv:1906.02569 (2019)
arXiv 2019
-
[2]
John Alton. 2013. Painting with light. University of California Press
2013
-
[3]
Tom Bartindale, Alia Sheikh, Nick Taylor, Peter Wright, and Patrick Olivier. 2012. StoryCrate: tabletop storyboarding for live film production. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . 169–178
2012
-
[4]
Staphord Bengesi, Hoda El-Sayed, Md Kamruzzaman Sarker, Yao Houkpati, John Irungu, and Timothy Oladunni. 2024. Advancements in Generative AI: A Compre- hensive Review of GANs, GPT, Autoencoders, Diffusion Model, and Transformers. IEEE Access (2024)
2024
-
[5]
Oloff C Biermann, Ning F Ma, and Dongwook Yoon. 2022. From tool to compan- ion: Storywriters want AI writers to respect their personal values and writing strategies. In Proceedings of the ACM Designing Interactive Systems Conference . 1209–1227
2022
-
[6]
Bruce Block. 2020. The visual story: Creating the visual structure of film, TV, and digital media. Routledge
2020
-
[7]
Fadi Boutros, Jonas Henry Grebe, Arjan Kuijper, and Naser Damer. 2023. Idiff-face: Synthetic-based face recognition through fizzy identity-conditioned diffusion model. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 19650–19661
2023
-
[8]
Karen Brewster and Melissa Shafer. 2011. Fundamentals of theatrical design: A guide to the basics of scenic, costume, and lighting design . Skyhorse Publishing Inc
2011
Show all 99 references
-
[9]
Blain Brown. 2016. Cinematography: theory and practice: image making for cinematographers and directors. Routledge
2016
-
[10]
Hung-Jen Chen, Hong-Han Shuai, and Wen-Huang Cheng. 2023. A survey of artificial intelligence in fashion. IEEE Signal Processing Magazine 40, 3 (2023), 64–73. CineVision: An Interactive Pre-visualization Storyboard System for Director–Cinematographer Collaboration UIST ’25, Se...
2023
-
[11]
John Joon Young Chung and Eytan Adar. 2023. Artinter: AI-powered Boundary Objects for Commissioning Visual Arts. InProceedings of the 2023 ACM Designing Interactive Systems Conference. 1997–2018
2023
-
[12]
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A Smith. 2018. Creative writing with a machine in the loop: Case studies on slogans and stories. In Proceedings of the 23rd International Conference on Intelligent User Interfaces. 329–340
2018
-
[13]
Rafael Concepcion. 2022. Adobe Photoshop and Lightroom Classic for Photogra- phers Classroom in a Book . Adobe Press
2022
-
[14]
Florinel-Alin Croitoru, Vlad Hondru, Radu Tudor Ionescu, and Mubarak Shah
-
[15]
Zheng Ding, Xuaner Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, and Xiuming Zhang. 2023. Diffusionrig: Learning personalized priors for facial appearance editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12736–12746
2023
-
[16]
Alexis Dinno. 2015. Nonparametric pairwise multiple comparisons in indepen- dent groups using Dunn’s test. The Stata Journal 15, 1 (2015), 292–300
2015
-
[17]
Tiffany D Do, Camille Isabella Protko, and Ryan P McMahan. 2024. Stepping into the Right Shoes: The Effects of User-Matched Avatar Ethnicity and Gender on Sense of Embodiment in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics (2024)
2024
-
[18]
Elham Doust-Haghighi. 2023. Laboratory Animation Productions: Strategies to Produce Customizable Animations to Be Used as Materials for Experimental Research on Media. The University of Texas at Dallas
2023
-
[19]
Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. arXiv preprint arXiv:1805.04833 (2018)
2018 arXiv
-
[20]
Rinon Gal, Moab Arar, Yuval Atzmon, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2023. Encoder-based domain tuning for fast personalization of text- to-image models. ACM Transactions on Graphics 42, 4 (2023), 1–13
2023
-
[21]
2012.Directing the story: professional storytelling and storyboarding techniques for live action and animation
Francis Glebas. 2012.Directing the story: professional storytelling and storyboarding techniques for live action and animation . Routledge
2012
-
[22]
Katherine Graham. 2022. The play of light: rethinking mood lighting in perfor- mance. Studies in Theatre and Performance 42, 2 (2022), 139–155
2022
-
[23]
Ziyue Guo, Zongyang Zhu, Yizhi Li, Shidong Cao, Hangyue Chen, and Gaoang Wang. 2023. AI assisted fashion design: A review. IEEE Access 11 (2023), 88403– 88415
2023
-
[24]
John Hart. 2013. The Art of the Storyboard: A filmmaker’s introduction. Routledge
2013
-
[25]
Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. 904–908
2006
-
[26]
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. 2024. Cameractrl: Enabling camera control for text-to-video generation. arXiv preprint arXiv:2404.02101 (2024)
2024 arXiv
-
[27]
Kai He, Kaixin Yao, Qixuan Zhang, Jingyi Yu, Lingjie Liu, and Lan Xu. 2024. Dresscode: Autoregressively sewing and generating garments from text guidance. ACM Transactions on Graphics 43, 4 (2024), 1–13
2024
-
[28]
Xudong Hong, Asad Sayeed, Khushboo Mehra, Vera Demberg, and Bernt Schiele
-
[29]
Donnesh Dustin Hosseini. 2024. Generative AI: a problematic illustration of the intersections of racialized gender, race, ethnicity
2024
-
[30]
Transactions of the Association for Computational Linguistics 11 (2023), 565–581
Visual writing prompts: Character-grounded story generation with curated image sequences. Transactions of the Association for Computational Linguistics 11 (2023), 565–581
2023
-
[31]
Li Hu. 2024. Animate anyone: Consistent and controllable image-to-video syn- thesis for character animation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8153–8163
2024
-
[32]
Andrew Hou, Ze Zhang, Michel Sarkis, Ning Bi, Yiying Tong, and Xiaoming Liu
-
[33]
Daphne Ippolito, David Grangier, Chris Callison-Burch, and Douglas Eck. 2019. Unsupervised hierarchical story infilling. In Proceedings of the First Workshop on Narrative Understanding. 37–43
2019
-
[34]
D Curtis Jamison. 2003. Structured query language (SQL) fundamentals. Current protocols in bioinformatics 1 (2003), 9–2
2003
-
[35]
Qingqiu Huang, Yu Xiong, Anyi Rao, Jiaze Wang, and Dahua Lin. 2020. Movienet: A holistic dataset for movie understanding. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part IV 16 . Springer, 709–727
2020
-
[36]
Patrick Keating. 2009. Hollywood Lighting from the Silent Era to Film Noir . Columbia University Press
2009
-
[37]
Minchul Kim, Feng Liu, Anil Jain, and Xiaoming Liu. 2023. Dcface: Synthetic face generation with dual condition diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12715–12725
2023
-
[38]
Steven Douglas Katz. 1991. Film directing shot by shot: visualizing from concept to screen. Gulf Professional Publishing
1991
-
[39]
Hyung-Kwon Ko, Subin An, Gwanmo Park, Seung Kwon Kim, Daesik Kim, Bo- hyoung Kim, Jaemin Jo, and Jinwook Seo. 2022. We-toon: A Communication Support System between Writers and Artists in Collaborative Webtoon Sketch Revision. In Proceedings of the 35th Annual ACM Symposium on ...
2022
-
[40]
Danrui Li, Samuel S Sohn, Sen Zhang, Che-Jui Chang, and Mubbasir Kapadia
-
[41]
Taeksoo Kim, Byungjun Kim, Shunsuke Saito, and Hanbyul Joo. 2024. Gala: Generating animatable layered assets from a single scan. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1535–1545
2024
-
[42]
Zhengqin Li, Jia Shi, Sai Bi, Rui Zhu, Kalyan Sunkavalli, Miloš Hašan, Zexiang Xu, Ravi Ramamoorthi, and Manmohan Chandraker. 2022. Physically-based editing of indoor scene lighting from a single image. InEuropean Conference on Computer Vision. Springer, 555–572
2022
-
[43]
Ming-Yu Liu, Xun Huang, Jiahui Yu, Ting-Chun Wang, and Arun Mallya. 2021. Generative adversarial networks for image and video synthesis: Algorithms and applications. Proc. IEEE 109, 5 (2021), 839–862
2021
-
[44]
Xingchao Liu, Xiwen Zhang, and Jianzhu et al. Ma. 2023. Instaflow: One step is enough for high-quality diffusion-based text-to-image generation. In The Twelfth International Conference on Learning Representations
2023
-
[45]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning. PMLR, 12888–12900
2022
-
[46]
Shilin Lu, Yanzhu Liu, and Adams Wai-Kin Kong. 2023. Tf-icon: Diffusion-based training-free cross-domain image composition. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 2294–2305
2023
-
[47]
Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. 2024. Mace: Mass concept erasure in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6430–6440
2024
-
[48]
Shilin Lu, Zihan Zhou, Jiayou Lu, Yuanzhi Zhu, and Adams Wai-Kin Kong. 2024. Robust watermarking using generative priors against image editing: From bench- marking to advances. arXiv preprint arXiv:2410.18775 (2024)
2024 arXiv
-
[49]
Vincent LoBrutto. 2002. The filmmaker’s guide to production design . Simon and Schuster
2002
-
[50]
John Mateer. 2017. Directing for cinematic virtual reality: how the traditional film director’s craft applies to immersive environments and notions of presence. Journal of Media Practice 18, 1 (2017), 14–25
2017
-
[51]
Patrick E McKight and Julius Najab. 2010. Kruskal-wallis test. The Corsini Encyclopedia of Psychology (2010), 1
2010
-
[52]
Henry Melki, Ian Montgomery, and Greg Maguire. 2019. An Investigation into the Creative Processes in Generating Believable Photorealistic Film Characters . Ph. D. Dissertation. Ulster University
2019
-
[53]
Mustafa Yousry Matbouly. 2022. Quantifying the unquantifiable: the color of cin- ematic lighting and its effect on audience’s impressions towards the appearance of film characters. Current Psychology 41, 6 (2022), 3694–3715
2022
-
[54]
G Painguzhali, Viswanath Ananth, Kavitha, et al. 2025. How artificial intelligence and generative AI is revolutionizing the fashion industry. In Generative AI for Business Analytics and Strategic Decision Making in Service Industry . IGI Global Scientific Publishing, 281–316
2025
-
[55]
Arjama Pal, Sakalya Mitra, and D Lakshmi. 2025. Illuminating the path from script to screen using lights, camera, and AI. In Transforming Cinema with Artificial Intelligence. IGI Global Scientific Publishing, 97–142
2025
-
[56]
Rohit Pandey, Sergio Orts-Escolano, Chloe Legendre, Christian Haene, Sofien Bouaziz, Christoph Rhemann, Paul E Debevec, and Sean Ryan Fanello. 2021. Total relighting: learning to relight portraits for background replacement. ACM Transactions on Graphics 40, 4 (2021), 43–1
2021
-
[57]
Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–34
2023
-
[58]
Stephen Prince. 2011. Digital visual effects in cinema: The seduction of reality . Rutgers University Press
2011
-
[59]
Anyi Rao, Jean-Peïc Chou, and Maneesh Agrawala. 2024. ScriptViz: A Visualiza- tion Tool to Aid Scriptwriting based on a Large Movie Database. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology . 1–13
2024
-
[60]
Anyi Rao, Xuekun Jiang, Yuwei Guo, Linning Xu, Lei Yang, Libiao Jin, Dahua Lin, and Bo Dai. 2023. Dynamic storyboard generation in an engine-based virtual environment for video production. In ACM SIGGRAPH 2023
2023
-
[61]
Puntawat Ponglertnapakorn, Nontawat Tritrong, and Supasorn Suwajanakorn
-
[62]
In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision
DiFaReli: Diffusion face relighting. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision . 22646–22657
-
[63]
Tassneam M Samy, Beshoy I Asham, Salwa O Slim, and Amr A Abohany. 2025. Revolutionizing online shopping with FITMI: a realistic virtual try-on solution. Neural Computing and Applications (2025), 1–20
2025
-
[64]
Christine Sciortino. 2020. Pre-Production. In Makeup Artistry for Film and Television. Routledge, 18–59
2020
-
[65]
Yang Shi, Nan Cao, Xiaojuan Ma, Siji Chen, and Pei Liu. 2020. EmoG: supporting the sketching of emotional expressions for storyboarding. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–12
2020
-
[66]
Nornadiah Mohd Razali, Yap Bee Wah, et al. 2011. Power comparisons of shapiro- wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests. Journal of statis- tical modeling and analytics 2, 1 (2011), 21–33
2011
-
[67]
Adrian Rusu and Amalia Rusu. 2024. Script-to-Storyboard-to-Story Reel Frame- work. In 2024 28th International Conference Information Visualisation (IV) . IEEE, 350–355. UIST ’25, September 28-October 1, 2025, Busan, Republic of Korea Wei et al
2024
-
[68]
Shun-Yu Wang, Wei-Chung Su, Serena Chen, Ching-Yi Tsai, Marta Misztal, Kather- ine M Cheng, Alwena Lin, Yu Chen, and Mike Y Chen. 2024. Roomdreaming: Generative-AI approach to facilitating iterative, preliminary interior design explo- ration. In Proceedings of the 2024 CHI Con...
2024
-
[69]
Zheng Wei, Yuzheng Chen, Wai Tong, Xuan Zong, Huamin Qu, Xian Xu, and Lik- Hang Lee. 2024. Hearing the Moment with MetaEcho! From Physical to Virtual in Synchronized Sound Recording. In Proceedings of the 32nd ACM International Conference on Multimedia. 6520–6529
2024
-
[70]
Zheng Wei, Shan Jin, Wai Tong, David Kei Man Yip, Pan Hui, and Xian Xu. 2024. Multi-Role VR Training System for Film Production: Enhancing Collaboration with MetaCrew. In ACM SIGGRAPH 2024 Posters. 1–2
2024
-
[71]
Mark Simon. 2012. Storyboards: motion in art . Routledge
2012
-
[72]
Ed S Tan. 2018. A psychology of the film. Palgrave Communications 4, 1 (2018)
2018
-
[73]
Hongtao Wu, Yijun Yang, Angelica I Aviles-Rivero, Jingjing Ren, Sixiang Chen, Haoyu Chen, and Lei Zhu. 2024. Semi-supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regular- ization. In European Conference on Computer Vision ...
2024
-
[74]
Hongtao Wu, Yijun Yang, Haoyu Chen, Jingjing Ren, and Lei Zhu. 2023. Mask- guided progressive network for joint raindrop and rain streak removal in videos. In Proceedings of the 31st ACM International Conference on Multimedia. 7216–7225
2023
-
[75]
Hongtao Wu, Yijun Yang, Huihui Xu, Weiming Wang, Jinni Zhou, and Lei Zhu
-
[76]
Zheng Wei, Jia Sun, Junxiang Liao, Lik-Hang Lee, Chan In Sio, Pan Hui, Huamin Qu, Wai Tong, and Xian Xu. 2025. Illuminating the Scene: How Virtual Environ- ments and Learning Modes Shape Film Lighting Mastery in Virtual Reality. IEEE Transactions on Visualization and Computer ...
2025
-
[77]
Zheng Wei, Xian Xu, Lik-Hang Lee, Wai Tong, Huamin Qu, and Pan Hui. 2023. Feeling Present! From Physical to Virtual Cinematography Lighting Education with Metashadow. In Proceedings of the 31st ACM International Conference on Multimedia. 1127–1136
2023
-
[78]
Xian Xu, Wai Tong, Zheng Wei, Meng Xia, Lik-Hang Lee, and Huamin Qu
-
[79]
Zheer Xu, Shanqing Cai, Mukund Varma T, Subhashini Venugopalan, and Shumin Zhai. 2024. SkipWriter: LLM-Powered Abbreviated Writing on Tablets. InProceed- ings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–13
2024
-
[80]
Jingyuan Yang, Jiawei Feng, and Hui Huang. 2024. EmoGen: Emotional image content generation with text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 6358–6368
2024
-
[81]
In Proceedings of the 32nd ACM International Conference on Multimedia
Rainmamba: Enhanced locality learning with state space models for video deraining. In Proceedings of the 32nd ACM International Conference on Multimedia . 7881–7890
-
[82]
Liwenhan Xie, Zhaoyu Zhou, Kerun Yu, Yun Wang, Huamin Qu, and Siming Chen. 2023. Wakey-wakey: Animate text by mimicking characters in a gif. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–14
2023
-
[83]
Peng Xu, Mostofa Patwary, Mohammad Shoeybi, Raul Puri, Pascale Fung, Anima Anandkumar, and Bryan Catanzaro. 2020. MEGATRON-CNTRL: Controllable story generation with external knowledge using large-scale language models. arXiv preprint arXiv:2010.00840 (2020)
2020 arXiv
-
[84]
Jun Yue, Leyuan Fang, Shaobo Xia, Yue Deng, and Jiayi Ma. 2023. Dif-fusion: Toward high color fidelity in infrared and visible image fusion with diffusion models. IEEE Transactions on Image Processing 32 (2023), 5705–5720
2023
-
[85]
Visual Informatics (2024)
Transforming cinematography lighting education in the metaverse. Visual Informatics (2024)
2024
-
[86]
Lingjun Zhang, Xinyuan Chen, Yaohui Wang, Yue Lu, and Yu Qiao. 2024. Brush your text: Synthesize any scene text on images via diffusion model. InProceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 7215–7223
2024
-
[87]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2023. Adding conditional con- trol to text-to-image diffusion models. InProceedings of the IEEE/CVF international conference on computer vision . 3836–3847
2023
-
[88]
Yijun Yang, Hongtao Wu, Angelica I Aviles-Rivero, Yulun Zhang, Jing Qin, and Lei Zhu. 2024. Genuine knowledge from practice: Diffusion test-time adaptation for video adverse weather removal. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, ...
2024
-
[89]
Yuyang Yin, Dejia Xu, Chuangchuang Tan, Ping Liu, Yao Zhao, and Yunchao Wei. 2023. Cle diffusion: Controllable light enhancement diffusion model. In Proceedings of the 31st ACM International Conference on Multimedia . 8145–8156
2023
-
[90]
Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: story writing with large language models. In Proceedings of the 27th International Conference on Intelligent User Interfaces . 841–852
2022
-
[91]
Xin Zhao. 2023. Leveraging artificial intelligence (AI) technology for English writing: Introducing wordtune as a digital writing assistant for EFL writers.RELC Journal 54, 3 (2023), 890–894
2023
-
[92]
Chong Zeng, Yue Dong, Pieter Peers, Youkang Kong, Hongzhi Wu, and Xin Tong. 2024. Dilightnet: Fine-grained lighting control for diffusion-based image generation. In ACM SIGGRAPH 2024 Conference Papers . 1–12
2024
-
[95]
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. 2024. Ic-light github page. (2024)
2024
-
[96]
Ruihan Zhang, Borou Yu, Jiajian Min, Yetong Xin, Zheng Wei, Juncheng Nemo Shi, Mingzhen Huang, Xianghao Kong, Nix Liu Xin, Shanshan Jiang, et al. 2025. Generative AI for Film Creation: A Survey of Recent Advances. In Proceedings of the Computer Vision and Pattern Recognition C...
2025
-
[97]
Yanbo Zhang and Chuanlan Liu. 2024. Unlocking the potential of artificial intel- ligence in fashion design and e-commerce applications: The case of Midjourney. Journal of Theoretical and Applied Electronic Commerce Research 19, 1 (2024), 654–670
2024
-
[99]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023. Minigpt-4: Enhancing vision-language understanding with advanced large lan- guage models. arXiv preprint arXiv:2304.10592 (2023)
2023 arXiv
-
[2021]
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Towards high fidelity face relighting with realistic shadows. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14719– 14728
-
[2023]
IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10850–10869
Diffusion models in vision: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 9 (2023), 10850–10869
2023
-
[2024]
In Proceedings of the 17th ACM SIGGRAPH Conference on Motion, Interaction, and Games
From Words to Worlds: Transforming One-line Prompts into Multi-modal Digital Stories with LLM Agents. In Proceedings of the 17th ACM SIGGRAPH Conference on Motion, Interaction, and Games . 1–12
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.