REVIEW 4 major objections 7 minor 1 cited by
VideoDiff: Human-AI Video Co-Creation with Alternatives
T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Aligned timelines and transcripts let video creators compare AI-generated edits in about half the time.
desk verdict A well-designed video co-creation tool with genuinely useful comparison views, but the quantitative claims are shakier than the p-values suggest. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the aligned comparison view: a color-coded section timeline and a synchronized transcript, both anchored to the source footage so that all variations share the same coordinate system. Sections are extracted from the source, not from each edit, which means every rough cut is shown against identical section boundaries, and the edited-vs-source toggle reveals which parts of the original each version keeps. This alignment converts the open-ended task of watching several videos and remembering what differed into a glanceable difference-highlighting task, and it is what the user study attributes the speed and accuracy gains to.
What would settle it
Show participants comparison questions about videos whose timelines and transcripts have been deliberately corrupted, for instance a section labeled 'grocery items' that actually contains different footage, and measure whether accuracy and speed fall to baseline levels; if accuracy collapses, the measured gains depend on the pipeline being truthful, and if it does not, the interface benefit is independent of representational fidelity.
Extended reading notes
Core claim
The central claim is that aligning multiple edited versions of the same source video against a shared timeline and transcript makes it practical for creators to work with many AI-generated alternatives at once. VideoDiff segments the source footage into sections, applies the same section structure consistently across every variation, and lets users toggle between the edited timeline and the source timeline, so a user can see which sections each rough cut includes or excludes and where B-rolls and text effects land. The paper reports that this design let participants compare faster and with less effort than a baseline that simply lists generated videos, and that creators used the freed time to refine, recombine, and regenerate suggestions, producing final videos they were more satisfied with.
Load-bearing premise
The evaluation assumes the AI pipeline that transcribes, segments, and edits the videos is accurate enough that the aligned timelines and transcripts describe what is actually in each edited video, since the paper's own limitations note transcription errors and hallucinations.
Editorial extensions
If this is right
- VideoDiff's results imply that AI video tools should treat comparison as a first-class design problem, not an afterthought of generation.
- Creators can productively start from ten alternatives when differences are visible at a glance, so reducing the number of suggestions is not the only remedy for overload.
- Because users made an average of 4.33 edits per video in the VideoDiff conditions, comparison support appears to shift effort from reviewing to customizing.
- The measured drop in mental demand, temporal demand, effort, and frustration suggests aligned views could make AI editing accessible to novices and to creators who find transcript-heavy review tiring.
Reading between the lines
- A testable extension of the paper's argument is that comparison cost, not generation quality, is the binding constraint on how many alternatives a creator can meaningfully consider; if so, better alignment should allow scaling to 20-30 variations without proportional time increases.
- The same alignment strategy may transfer to other temporal creative media, such as podcast editing, narrated slides, or multi-camera footage, where multiple AI-generated versions share a common source timeline.
- The study's measured benefits would likely shrink if the underlying pipeline misrepresents the videos; a direct test would inject transcription or segmentation errors and re-run the comparison task.
- The paper leaves verification of AI suggestions as future work, implying that a version of VideoDiff that flags likely hallucinations or jump cuts could further improve trust and accuracy.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents VideoDiff, a web-based human-AI video co-creation tool that generates and displays multiple AI-generated alternatives for three editing stages: rough cuts, B-roll insertion, and text effects. Its central innovation is a set of aligned timeline and transcript views that let creators compare alternatives side-by-side, with synchronized color coding, source-versus-edited toggles, and difference-highlighting thumbnails. The authors report a formative study with 8 professional editors that yields six design goals, a within-subjects user study with 12 participants comparing VideoDiff against a baseline interface, and three exploratory case studies with creators editing their own footage. The evaluation measures comparison time, answer accuracy, NASA-TLX workload, creativity support index, satisfaction, and usefulness, and reports significant benefits for VideoDiff on several time, workload, satisfaction, and creativity measures, along with qualitative evidence about how users compare and customize alternatives.
Significance. If the quantitative findings are robust, the paper makes a valuable contribution by demonstrating that the burden of comparison—not just generation speed—is a central factor in human-AI video co-creation. The design goals D1–D6 are well grounded in the formative study, and the aligned timeline/transcript visualization is a thoughtful response to the temporal and multimodal nature of video comparison. The study is carefully structured with counterbalancing, standardized instruments (NASA-TLX, CSI), and realistic materials, and the qualitative analysis offers concrete insights into user strategies and workflow impact. The authors also explicitly acknowledge limitations of their LLM-based generation pipeline. The main weakness is statistical: the quantitative claims depend on a large number of uncorrected pairwise tests with N=12, and no effect sizes or confidence intervals are reported, which leaves the headline quantitative results insecure.
major comments (4)
- [§5.2, Figure 9, Table 3] The paper's quantitative claims rest on many uncorrected Wilcoxon tests with N=12. In the comparison task alone there are 10 per-question time tests, 10 per-question accuracy tests, 5 NASA-TLX subscales, and several usefulness and satisfaction ratings. Most reported p-values cluster between 0.01 and 0.05 (e.g., mental demand Z=-2.08, temporal demand Z=-2.36, effort Z=-2.62, frustration Z=-1.78, satisfaction Z=2.63). With N=12, the smallest achievable two-sided Wilcoxon p-value is about 0.0005, and under a Benjamini-Hochberg correction at q=0.05 none of the reported p-values would remain significant. The expected number of false positives among the roughly 20 nominal rejections is close to one even if all null hypotheses were true. The paper also reports no effect sizes or confidence intervals. This does not invalidate the qualitative findings, but it means the quantitative foundation for H1–H7 is not currently established. The authors should report corrected p-values, effect sizes (e.g., matched-rank-biserial correlation), or a more appropriate repeated-measures model, and adjust the strength of their claims accordingly.
- [§5.2, Figure 8, H1] The headline "roughly half the time" (38s vs 74s) is presented without an omnibus statistical test; only five of the ten per-question time comparisons are reported as significant. H1 states that "VideoDiff significantly decreases the time in video comparison," but the evidence is a post-hoc-looking subset of questions. Please specify which comparisons were pre-specified, report the overall test result (e.g., a paired test on mean per-question completion time), or rephrase H1 to refer to the subset of questions that were significantly faster.
- [§5.1, RQ1/H2, Table 3] Accuracy differences in Table 3 are presented without any inferential test. Seven of ten questions show numerically higher mean accuracy for VideoDiff (e.g., Q3: 0.42 vs 1.00; Q7: 0.36 vs 0.90), but with N=12 and large standard deviations, these differences are likely not statistically significant; no test statistics or confidence intervals are reported. Consequently, H2 ("improves comprehension and accuracy in video comparison") is unsupported as stated. Please either provide significance tests for the accuracy comparisons or explicitly state that the accuracy differences were not statistically significant and remove H2 from the confirmed hypotheses.
- [§4.3, §7 Limitations] The system's internal representations—sections, B-roll placements, text-effect annotations, and even the transcripts—are generated by a pipeline that the paper itself acknowledges is "prone to transcription errors and LLM hallucinations," does not consider visual input, and "cannot follow visual edit requests." The evaluation, however, treats these representations as ground truth: comparison accuracy is scored against the actual video content, while the interface displays possibly incorrect segmentations and descriptions. The paper does not report any measure of pipeline accuracy (e.g., transcription WER, section-segmentation agreement, or B-roll placement precision). If the pipeline frequently misrepresents the video, the comparison task may be measuring the interface's ability to communicate inaccurate information, and the observed benefits could be tied to the specific failure modes of the generation pipeline. The authors should either add a pipeline accuracy analysis or explicitly separate the claims about interface-level comparison support from the quality of the underlying generation, ideally by reporting accuracy on a per-task basis and discussing how representation errors affect the validity of the user-study measurements.
minor comments (7)
- [§1, §2.2, Figure 9 caption] There are several typos: "asembling" in §1, "Boreczkly" in §2.2 should be "Boreczky", and the Figure 9 caption says "Wilcoxon text" instead of "Wilcoxon test". Please proofread the manuscript.
- [§4.3 footnote] The footnote "We used GPT-o to identify visually concrete keywords" is inconsistent with the rest of the paper, which uses GPT-4o; please correct the model name.
- [§5.1, Baseline] The baseline interface is described only by reference to Figure 7 and a short sentence. Please provide a precise list of its features (e.g., does it include a transcript view, keyword search, sorting, generation/regeneration?) so readers can assess whether the comparison isolates the aligned timeline/transcript contribution or conflates it with other UI differences.
- [§5.1, Procedure] The sentence "participants reviewed 10 different videos" is ambiguous; there were 10 variations generated per source video and each participant was assigned one of V1/V2 per condition. Please clarify the exact number of variations shown and how assignments were balanced.
- [§5.2, Figure 8] The x-axis labels in Figure 8 are confusing (the categories repeat and there is a duplicate "Q9"). Please clarify the mapping between the four column groups and the ten numbered questions, and label the significance marks individually.
- [§5.2, Time results] The phrase "In 5 questions that required checking multiple parts of the video" is vague; please list the specific question numbers that were significantly faster and specify whether these correspond to the multi-select questions or a different criterion.
- [§5.3, Authoring task] The report that VideoDiff users made 4.33 edits (SD=2.42) is not contextualized against the baseline, where most participants made no edits. A descriptive comparison of edit counts (even if not inferential) would help readers interpret the scale of the difference.
Circularity Check
No significant circularity: the central claim is an empirical interface comparison with no fitted parameters, derived equations, or load-bearing self-citation chain.
full rationale
VideoDiff makes an HCI systems claim: that aligned timeline and transcript views help creators compare and customize AI-generated video alternatives. The paper contains no mathematical derivation or fitted model; the central claim rests on a within-subjects user study (Section 5) comparing VideoDiff to a baseline interface. None of the measured outcomes (completion time, accuracy, NASA-TLX, CSI, self-rated usefulness/satisfaction) are defined in terms of VideoDiff's own outputs; they are independent user responses to two interface conditions. The formative study (Section 3) motivates design goals D1-D6, but those goals are presented as design rationale and are then tested empirically, not asserted as predictions derived from the system itself. The AI pipeline limitations acknowledged in Section 7 (transcription errors, LLM hallucinations, no visual input) are a real external-validity threat, but they do not make the comparison circular because the same generation pipeline feeds both conditions; the study isolates the interface as the manipulated variable. Self-citations to prior work (e.g., AVscript, ReelFramer, PodReels, B-Script) appear as related work and design inspiration, not as justification of the empirical results; no uniqueness theorem or load-bearing result is imported from the authors' prior work. No equation is reused as its own input, no fitted parameter is relabeled as a prediction, and no known result is merely renamed. Statistical concerns about uncorrected multiple comparisons, raised in the skeptic analysis, are a correctness risk rather than a circularity risk. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The baseline interface is representative of existing AI video editing tools such as OpusClip and CapCut.
- domain assumption GPT-4o and Whisper, with the authors' prompts, generate edit suggestions and transcript segmentations accurate enough for the study tasks.
- domain assumption The 12 study participants and 8 formative participants are representative of video creators who would use AI editing tools.
Cite this review
Pith. "Pith review of VideoDiff: Human-AI Video Co-Creation with Alternatives." pith.science (2026). https://pith.science/paper/TVUBWDL4
@misc{pith2026250210190,
author = {Pith},
title = {Pith review of: VideoDiff: Human-AI Video Co-Creation with Alternatives},
year = {2026},
howpublished = {\url{https://pith.science/paper/TVUBWDL4}},
note = {Machine review of arXiv:2502.10190}
}
read the original abstract
To make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and automation. A key strength of generative models is their ability to quickly generate multiple variations, but when provided with many alternatives, creators struggle to compare them to find the best fit. We propose VideoDiff, an AI video editing tool designed for editing with alternatives. With VideoDiff, creators can generate and review multiple AI recommendations for each editing process: creating a rough cut, inserting B-rolls, and adding text effects. VideoDiff simplifies comparisons by aligning videos and highlighting differences through timelines, transcripts, and video previews. Creators have the flexibility to regenerate and refine AI suggestions as they compare alternatives. Our study participants (N=12) could easily compare and customize alternatives, creating more satisfying results.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.
Reference graph
Works this paper leans on
-
[1]
Columbia Undergraduate Admissions. 2023. Columbia University Campus Tour. https://www.youtube.com/watch?v=OkxRkNjPHIk
2023
-
[2]
Adobe. [n. d.]. Adobe Premiere Pro. https://www.adobe.com/products/premiere. html
-
[3]
Shm Garanganao Almeda, JD Zamfirescu-Pereira, Kyu Won Kim, Pradeep Mani Rathnam, and Bjoern Hartmann. 2024. Prompting for Discovery: Flex- ible Sense-Making for AI Art-Making with Dreamsheets. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17
2024
-
[4]
Apple. [n. d.]. Apple iMovie. https://support.apple.com/imovie
-
[5]
Narges Ashtari, Andrea Bunt, Joanna McGrenere, Michael Nebeling, and Parmit K Chilana. 2020. Creating augmented and virtual reality applications: Current practices, challenges, and opportunities. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–13
2020
-
[6]
Adam Baker, Carl Gutwin, Justin Matejka, and Ian Stavness. 2024. Interaction Techniques for Comparing Video. In Graphics Interface 2024 Second Deadline
2024
-
[7]
Guha Balakrishnan, Frédo Durand, and John Guttag. 2015. Video diff: Highlight- ing differences between similar actions in videos. ACM Transactions on Graphics (TOG) 34, 6 (2015), 1–10
2015
-
[8]
Connelly Barnes, Dan B Goldman, Eli Shechtman, and Adam Finkelstein. 2010. Video tapestries with continuous temporal zoom. In ACM SIGGRAPH 2010 papers. 1–9
2010
Show all 95 references
-
[9]
Elena Benedetto, Gabriele Romano, Ilaria Torre, Mario Vallarino, and Gi- anni Viardo Vercelli. 2024. A Visual Comparison interface for educational videos. In Proceedings of the 2024 International Conference on Advanced Visual Interfaces . 1–5
2024
-
[10]
John Boreczky, Andreas Girgensohn, Gene Golovchinsky, and Shingo Uchihashi
-
[11]
Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Gross- man. 2023. Promptify: Text-to-image generation through interactive prompt exploration with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–14
2023
-
[12]
Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The impact of multiple parallel phrase suggestions on email input and composition behaviour of native and non-native english writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–13
2021
-
[13]
CapCut. [n. d.]. CapCut. https://www.capcut.com/
-
[14]
Minsuk Chang, Mina Huh, and Juho Kim. 2021. Rubyslippers: Supporting content- based voice navigation for how-to videos. InProceedings of the 2021 CHI conference on human factors in computing systems . 1–14
2021
-
[15]
Shizhe Chen, Bei Liu, Jianlong Fu, Ruihua Song, Qin Jin, Pingping Lin, Xiaoyu Qi, Chunting Wang, and Jin Zhou. 2019. Neural storyboard artist: Visualizing stories with coherent image sequences. In Proceedings of the 27th ACM International Conference on Multimedia. 2236–2244
2019
-
[16]
Erin Cherry and Celine Latulipe. 2014. Quantifying the creativity support of digital tools through the creativity support index.ACM Transactions on Computer- Human Interaction (TOCHI) 21, 4 (2014), 1–25
2014
-
[17]
Peggy Chi, Tao Dong, Christian Frueh, Brian Colonna, Vivek Kwatra, and Irfan Essa. 2022. Synthesis-Assisted Video Prototyping From a Document. In Proceed- ings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–10
2022
-
[18]
Peggy Chi, Nathan Frey, Katrina Panovich, and Irfan Essa. 2021. Automatic instructional video creation from a markdown-formatted tutorial. In The 34th Annual ACM Symposium on User Interface Software and Technology . 677–690
2021
-
[19]
Peggy Chi, Zheng Sun, Katrina Panovich, and Irfan Essa. 2020. Automatic video creation from a web page. In Proceedings of the 33rd annual ACM symposium on user interface software and technology . 279–292
2020
-
[20]
Pei-Yu Chi, Joyce Liu, Jason Linder, Mira Dontcheva, Wilmot Li, and Bjoern Hartmann. 2013. Democut: generating concise instructional videos for physi- cal demonstrations. In Proceedings of the 26th annual ACM symposium on User interface software and technology . 141–150
2013
-
[21]
John Joon Young Chung, Shiqing He, and Eytan Adar. 2021. The intersection of users, roles, interactions, and technologies in creativity support tools. In Proceedings of the 2021 ACM Designing Interactive Systems Conference. 1817–1833
2021
-
[22]
Opus Clip. [n. d.]. Opus Clip. https://www.opus.pro/
-
[23]
Descript. [n. d.]. Descript: Edit Videos & Podcasts Like a Doc | AI Video Editor. https://www.descript.com/
-
[24]
Steven P Dow, Alana Glassco, Jonathan Kass, Melissa Schwarz, Daniel L Schwartz, and Scott R Klemmer. 2010. Parallel prototyping leads to better design results, more divergence, and increased self-efficacy. ACM Transactions on Computer- Human Interaction (TOCHI) 17, 4 (2010), 1–24
2010
-
[25]
Steven M Drucker, Georg Petschnigg, and Maneesh Agrawala. 2006. Comparing and managing multiple versions of slide presentations. In Proceedings of the 19th annual ACM symposium on User interface software and technology . 47–56
2006
-
[26]
Fabio Duarte. [n. d.]. YouTube Content Creator Statistics (2024). https:// explodingtopics.com/blog/youtube-creator-stats CHI ’25, April 26-May 1, 2025, Yokohama, Japan Mina Huh, Dingzeyu Li, Kim Pimmel, Hijung Valentina Shin, Amy Pavel, and Mira Dontcheva
2024
-
[27]
Kasra Ferdowsi, Ruanqianqian Huang, Michael B James, Nadia Polikarpova, and Sorin Lerner. 2024. Validating AI-Generated Code with Live Programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–8
2024
-
[28]
C Ailie Fraser, Joy O Kim, Hijung Valentina Shin, Joel Brandt, and Mira Dontcheva
-
[29]
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shecht- man, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Transactions on Graphics (TOG) 38, 4 (2019), 1–14
2019
-
[30]
Simret Araya Gebreegziabher, Yukun Yang, Elena L Glassman, and Toby Jia-Jun Li. 2024. Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories. arXiv preprint arXiv:2409.16561 (2024)
2024 arXiv
-
[31]
Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K Kummerfeld, and Elena L Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21
2024
-
[32]
Michael Gleicher, Danielle Albers, Rick Walker, Ilir Jusufi, Charles D Hansen, and Jonathan C Roberts. 2011. Visual comparison for information visualization. Information Visualization 10, 4 (2011), 289–309
2011
-
[33]
Han L Han, Miguel A Renom, Wendy E Mackay, and Michel Beaudouin-Lafon
-
[34]
Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904–908
2006
-
[35]
Bernd Huber, Hijung Valentina Shin, Bryan Russell, Oliver Wang, and Gautham J Mysore. 2019. B-script: Transcript-based b-roll video editing with recommenda- tions. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–11
2019
-
[36]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Textlets: Supporting constraints and consistency in text documents. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
2020
-
[37]
Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17
2023
-
[38]
invideo AI. [n. d.]. invideo AI Script Generator. https://invideo.io/tools/script- generator/
-
[39]
Mina Huh, Yi-Hao Peng, and Amy Pavel. 2023. GenAssist: Making image gen- eration accessible. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–17
2023
-
[40]
Joy Kim, Mira Dontcheva, Wilmot Li, Michael S Bernstein, and Daniela Steinsapir
-
[41]
Janin Koch, Nicolas Taffin, Michel Beaudouin-Lafon, Markku Laine, Andrés Lucero, and Wendy E Mackay. 2020. Imagesense: An intelligent collaborative ideation tool to support diverse human-computer partnerships. Proceedings of the ACM on human-computer interaction 4, CSCW1 (2020), 1–27
2020
-
[42]
Haojian Jin, Yale Song, and Koji Yatani. 2017. Elasticplay: Interactive video summarization with dynamic time budgets. In Proceedings of the 25th ACM international conference on Multimedia . 1164–1172
2017
-
[43]
Philippe Laban, Jesse Vig, Marti Hearst, Caiming Xiong, and Chien-Sheng Wu
-
[44]
Mackenzie Leake, Abe Davis, Anh Truong, and Maneesh Agrawala. 2017. Com- putational video editing for dialogue-driven scenes. ACM Trans. Graph. 36, 4 (2017), 130–1
2017
-
[45]
Mackenzie Leake, Hijung Valentina Shin, Joy O Kim, and Maneesh Agrawala. 2020. Generating audio-visual slideshows from text articles using word concreteness. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–11
2020
-
[46]
Steffen Koch, Markus John, Michael Wörner, Andreas Müller, and Thomas Ertl
-
[47]
Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2024. Selenite: Scaffolding Online Sensemak- ing with Comprehensive Overviews Elicited from Large Language Models. In Proceedings of the CHI Conference on Human Factors in ...
2024
-
[48]
Vivian Liu, Tao Long, Nathan Raw, and Lydia Chilton. 2023. Generative disco: Text-to-video generation for music visualization. arXiv preprint arXiv:2304.08551 (2023)
2023 arXiv
-
[49]
Ference Marton. 2014. Necessary conditions of learning . Routledge
2014
-
[50]
Justin Matejka, Michael Glueck, Erin Bradner, Ali Hashemi, Tovi Grossman, and George Fitzmaurice. 2018. Dream lens: Exploration and visualization of large- scale generative design datasets. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12
2018
-
[51]
Justin Matejka, Tovi Grossman, and George Fitzmaurice. 2014. Video lens: rapid playback and exploration of large video collections and associated metadata. In Proceedings of the 27th annual ACM symposium on User interface software and technology. 541–550
2014
-
[52]
Ching Liu, Juho Kim, and Hao-Chuan Wang. 2018. ConceptScape: Collaborative concept mapping for video learning. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12
2018
-
[53]
Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–34
2023
-
[54]
Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the CHI Conference on Hu...
2024
-
[55]
OpenAI. [n. d.]. DALL-E 3. https://openai.com/index/dall-e-3/
-
[56]
OpenAI. 2024. Sora OpenAI video generation model. https://openai.com/sora
2024
-
[57]
2024 (accessed Sept 11, 2024)
Remotion.dev. 2024 (accessed Sept 11, 2024). Remotion: Make videos program- matically in react. https://www.remotion.dev/
2024
-
[58]
Midjourney. [n. d.]. Midjourney. https://www.midjourney.com/
-
[59]
Brittany Rose. 2024. Huge ALDI Grocery Haul | Shop with me and haul!! https: //www.youtube.com/watch?v=ti43mILvSug
2024
-
[60]
Brittany Rose. 2024. WEEKLY GROCERY HAUL AND MEAL PLAN | Shop with me!! https://www.youtube.com/watch?v=IxPofzudXxM
2024
-
[61]
Steve Rubin and Maneesh Agrawala. 2014. Generating emotionally relevant musical scores for audio stories. InProceedings of the 27th annual ACM symposium on User interface software and technology . 439–448
2014
-
[62]
Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–26
2024
-
[63]
Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling multilevel exploration and sensemaking with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18
2023
-
[64]
Mohi Reza, Nathan Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, Joseph Jay Williams, et al. 2023. ABScribe: Rapid Exploration of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Langu...
2023 arXiv
-
[65]
TED. [n. d.]. Does Working Hard Really Make You a Good Person? | Azim Shariff | TED. https://www.youtube.com/watch?v=z_ca983qJcI
-
[66]
Atima Tharatipyakul and Hyowon Lee. 2018. Towards a better video comparison: comparison as a way of browsing the video contents. In Proceedings of the 30th Australian Conference on Computer-Human Interaction . 349–353
2018
-
[67]
Bekzat Tilekbay, Saelyne Yang, Michal Adam Lewkowicz, Alex Suryapranata, and Juho Kim. 2024. ExpressEdit: Video Editing with Natural Language and Sketching. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 515–536
2024
-
[68]
Anh Truong and Maneesh Agrawala. 2019. A Tool for Navigating and Editing 360 Video of Social Conversations into Shareable Highlights.. In Graphics Interface. 14–1
2019
-
[69]
Anh Truong, Floraine Berthouzoz, Wilmot Li, and Maneesh Agrawala. 2016. Quickcut: An interactive tool for editing narrated video. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology . 497–507
2016
-
[70]
Amanda Swearngin, Chenglong Wang, Alannah Oleson, James Fogarty, and Amy J Ko. 2020. Scout: Rapid exploration of interface layout alternatives through high-level design constraints. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–13
2020
-
[71]
Vimeo. [n. d.]. Vimeo AI-Powered Video Platform. https://vimeo.com/
-
[72]
Vizard. [n. d.]. Vizard. https://vizard.ai/
-
[73]
Samangi Wadinambiarachchi, Ryan M Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso. 2024. The Effects of Generative AI on Design Fixation and Divergent Thinking. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18
2024
-
[74]
It Felt Like Having a Second Mind
Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2024. " It Felt Like Having a Second Mind": Investigating Human-AI Co-creativity in Prewriting with Large Language Models. Proceedings of the ACM on Human- Computer Interaction 8, CSCW1 (2024), 1–26
2024
-
[75]
Bryan Wang, Zeyu Jin, and Gautham Mysore. 2022. Record Once, Post Every- where: Automatic Shortening of Audio Stories for Social Media. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–11. VideoDiff: Human-AI Video Co-Creation with ...
2022
-
[76]
Babish Culinary Universe. [n. d.]. Eggs Benedict | Basics with Babish Live. https: //www.youtube.com/watch?v=ROAmMJeTO6Y
-
[77]
Miao Wang, Guo-Wei Yang, Shi-Min Hu, Shing-Tung Yau, Ariel Shamir, et al
-
[78]
Sitong Wang, Samia Menon, Tao Long, Keren Henderson, Dingzeyu Li, Kevin Crowston, Mark Hansen, Jeffrey V Nickerson, and Lydia B Chilton. 2024. Reel- Framer: Human-AI co-creation for news-to-video translation. In Proceedings of the CHI Conference on Human Factors in Computing S...
2024
-
[79]
Sitong Wang, Zheng Ning, Anh Truong, Mira Dontcheva, Dingzeyu Li, and Lydia B Chilton. 2024. PodReels: Human-AI Co-Creation of Video Podcast Teasers. In Proceedings of the 2024 ACM Designing Interactive Systems Conference . 958–974
2024
-
[80]
Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al
-
[81]
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22
2022
-
[82]
Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 699–714
2024
-
[83]
Liwenhan Xie, Zhaoyu Zhou, Kerun Yu, Yun Wang, Huamin Qu, and Siming Chen. 2023. Wakey-Wakey: Animate Text by Mimicking Characters in a GIF. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–14
2023
-
[84]
Saelyne Yang, Jisu Yim, Aitolkyn Baigutanova, Seoyoung Kim, Minsuk Chang, and Juho Kim. 2022. SoftVideo: Improving the Learning Experience of Soft- ware Tutorial Videos with Collective Interaction Data. In Proceedings of the 27th International Conference on Intelligent User In...
2022
-
[85]
Catherine Yeh, Gonzalo Ramos, Rachel Ng, Andy Huntington, and Richard Banks
-
[86]
Zheng Zhang, Zheng Ning, Chenliang Xu, Yapeng Tian, and Toby Jia-Jun Li. 2023. PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual Data. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18. CHI ’25, April 26-May 1, 2025...
2023
-
[90]
Haijun Xia, Jennifer Jacobs, and Maneesh Agrawala. 2020. Crosscast: adding visuals to audio travel podcasts. InProceedings of the 33rd annual ACM symposium on user interface software and technology . 735–746
2020
-
[94]
arXiv preprint arXiv:2402.08855 (2024)
GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency. arXiv preprint arXiv:2402.08855 (2024)
2024 arXiv
-
[2000]
In Proceedings of the SIGCHI conference on Human factors in computing systems
An interactive comic book presentation for exploring video. In Proceedings of the SIGCHI conference on Human factors in computing systems . 185–192
-
[2014]
IEEE transactions on visualization and computer graphics 20, 12 (2014), 1723–1732
VarifocalReader—in-depth visual analysis of large text documents. IEEE transactions on visualization and computer graphics 20, 12 (2014), 1723–1732
2014
-
[2015]
In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems
Motif: Supporting novice creativity through expert patterns. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems . 1211–1220
-
[2019]
ACM Trans
Write-a-video: computational video montage from themed text. ACM Trans. Graph. 38, 6 (2019), 177–1
2019
-
[2020]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Temporal segmentation of creative live streams. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12
2020
-
[2022]
In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229
2022
-
[2024]
In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology
Beyond the chat: Executable and verifiable text-editing with llms. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–23
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.