Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

VideoDiff: Human-AI Video Co-Creation with Alternatives

T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Aligned timelines and transcripts let video creators compare AI-generated edits in about half the time.

desk verdict A well-designed video co-creation tool with genuinely useful comparison views, but the quantitative claims are shakier than the p-values suggest. read the letter →

arxiv 2502.10190 v1 pith:TVUBWDL4 submitted 2025-02-14 cs.HC

classification cs.HC
keywords videoeditinghuman-AIco-creationgenerativeAIalternativescomparisontimelinevisualizationtranscriptviewuserstudycreativitysupport
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes VideoDiff, an AI video-editing tool built around the idea that the hard part of AI-assisted editing is not generating edits but comparing the many versions AI can produce. VideoDiff generates multiple alternative rough cuts, B-roll placements, and text effects, then aligns those alternatives on shared, color-coded timelines and synchronized transcripts so creators can see at a glance what differs. In a within-subjects study with 12 video creators, participants answered comparison questions in about half the time with VideoDiff than with a baseline interface (38 vs. 74 seconds), were more accurate on most questions, reported lower mental demand, effort, and frustration, and rated their final videos as more satisfying. The authors argue that comparison support, not generation speed alone, is what makes editing with AI alternatives usable.

What carries the argument

The load-bearing mechanism is the aligned comparison view: a color-coded section timeline and a synchronized transcript, both anchored to the source footage so that all variations share the same coordinate system. Sections are extracted from the source, not from each edit, which means every rough cut is shown against identical section boundaries, and the edited-vs-source toggle reveals which parts of the original each version keeps. This alignment converts the open-ended task of watching several videos and remembering what differed into a glanceable difference-highlighting task, and it is what the user study attributes the speed and accuracy gains to.

What would settle it

Show participants comparison questions about videos whose timelines and transcripts have been deliberately corrupted, for instance a section labeled 'grocery items' that actually contains different footage, and measure whether accuracy and speed fall to baseline levels; if accuracy collapses, the measured gains depend on the pipeline being truthful, and if it does not, the interface benefit is independent of representational fidelity.

Watch

Extended reading notes

Core claim

The central claim is that aligning multiple edited versions of the same source video against a shared timeline and transcript makes it practical for creators to work with many AI-generated alternatives at once. VideoDiff segments the source footage into sections, applies the same section structure consistently across every variation, and lets users toggle between the edited timeline and the source timeline, so a user can see which sections each rough cut includes or excludes and where B-rolls and text effects land. The paper reports that this design let participants compare faster and with less effort than a baseline that simply lists generated videos, and that creators used the freed time to refine, recombine, and regenerate suggestions, producing final videos they were more satisfied with.

Load-bearing premise

The evaluation assumes the AI pipeline that transcribes, segments, and edits the videos is accurate enough that the aligned timelines and transcripts describe what is actually in each edited video, since the paper's own limitations note transcription errors and hallucinations.

Editorial extensions

If this is right

  • VideoDiff's results imply that AI video tools should treat comparison as a first-class design problem, not an afterthought of generation.
  • Creators can productively start from ten alternatives when differences are visible at a glance, so reducing the number of suggestions is not the only remedy for overload.
  • Because users made an average of 4.33 edits per video in the VideoDiff conditions, comparison support appears to shift effort from reviewing to customizing.
  • The measured drop in mental demand, temporal demand, effort, and frustration suggests aligned views could make AI editing accessible to novices and to creators who find transcript-heavy review tiring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension of the paper's argument is that comparison cost, not generation quality, is the binding constraint on how many alternatives a creator can meaningfully consider; if so, better alignment should allow scaling to 20-30 variations without proportional time increases.
  • The same alignment strategy may transfer to other temporal creative media, such as podcast editing, narrated slides, or multi-camera footage, where multiple AI-generated versions share a common source timeline.
  • The study's measured benefits would likely shrink if the underlying pipeline misrepresents the videos; a direct test would inject transcription or segmentation errors and re-run the comparison task.
  • The paper leaves verification of AI suggestions as future work, implying that a version of VideoDiff that flags likely hallucinations or jump cuts could further improve trust and accuracy.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper presents VideoDiff, a web-based human-AI video co-creation tool that generates and displays multiple AI-generated alternatives for three editing stages: rough cuts, B-roll insertion, and text effects. Its central innovation is a set of aligned timeline and transcript views that let creators compare alternatives side-by-side, with synchronized color coding, source-versus-edited toggles, and difference-highlighting thumbnails. The authors report a formative study with 8 professional editors that yields six design goals, a within-subjects user study with 12 participants comparing VideoDiff against a baseline interface, and three exploratory case studies with creators editing their own footage. The evaluation measures comparison time, answer accuracy, NASA-TLX workload, creativity support index, satisfaction, and usefulness, and reports significant benefits for VideoDiff on several time, workload, satisfaction, and creativity measures, along with qualitative evidence about how users compare and customize alternatives.

Significance. If the quantitative findings are robust, the paper makes a valuable contribution by demonstrating that the burden of comparison—not just generation speed—is a central factor in human-AI video co-creation. The design goals D1–D6 are well grounded in the formative study, and the aligned timeline/transcript visualization is a thoughtful response to the temporal and multimodal nature of video comparison. The study is carefully structured with counterbalancing, standardized instruments (NASA-TLX, CSI), and realistic materials, and the qualitative analysis offers concrete insights into user strategies and workflow impact. The authors also explicitly acknowledge limitations of their LLM-based generation pipeline. The main weakness is statistical: the quantitative claims depend on a large number of uncorrected pairwise tests with N=12, and no effect sizes or confidence intervals are reported, which leaves the headline quantitative results insecure.

major comments (4)
  1. [§5.2, Figure 9, Table 3] The paper's quantitative claims rest on many uncorrected Wilcoxon tests with N=12. In the comparison task alone there are 10 per-question time tests, 10 per-question accuracy tests, 5 NASA-TLX subscales, and several usefulness and satisfaction ratings. Most reported p-values cluster between 0.01 and 0.05 (e.g., mental demand Z=-2.08, temporal demand Z=-2.36, effort Z=-2.62, frustration Z=-1.78, satisfaction Z=2.63). With N=12, the smallest achievable two-sided Wilcoxon p-value is about 0.0005, and under a Benjamini-Hochberg correction at q=0.05 none of the reported p-values would remain significant. The expected number of false positives among the roughly 20 nominal rejections is close to one even if all null hypotheses were true. The paper also reports no effect sizes or confidence intervals. This does not invalidate the qualitative findings, but it means the quantitative foundation for H1–H7 is not currently established. The authors should report corrected p-values, effect sizes (e.g., matched-rank-biserial correlation), or a more appropriate repeated-measures model, and adjust the strength of their claims accordingly.
  2. [§5.2, Figure 8, H1] The headline "roughly half the time" (38s vs 74s) is presented without an omnibus statistical test; only five of the ten per-question time comparisons are reported as significant. H1 states that "VideoDiff significantly decreases the time in video comparison," but the evidence is a post-hoc-looking subset of questions. Please specify which comparisons were pre-specified, report the overall test result (e.g., a paired test on mean per-question completion time), or rephrase H1 to refer to the subset of questions that were significantly faster.
  3. [§5.1, RQ1/H2, Table 3] Accuracy differences in Table 3 are presented without any inferential test. Seven of ten questions show numerically higher mean accuracy for VideoDiff (e.g., Q3: 0.42 vs 1.00; Q7: 0.36 vs 0.90), but with N=12 and large standard deviations, these differences are likely not statistically significant; no test statistics or confidence intervals are reported. Consequently, H2 ("improves comprehension and accuracy in video comparison") is unsupported as stated. Please either provide significance tests for the accuracy comparisons or explicitly state that the accuracy differences were not statistically significant and remove H2 from the confirmed hypotheses.
  4. [§4.3, §7 Limitations] The system's internal representations—sections, B-roll placements, text-effect annotations, and even the transcripts—are generated by a pipeline that the paper itself acknowledges is "prone to transcription errors and LLM hallucinations," does not consider visual input, and "cannot follow visual edit requests." The evaluation, however, treats these representations as ground truth: comparison accuracy is scored against the actual video content, while the interface displays possibly incorrect segmentations and descriptions. The paper does not report any measure of pipeline accuracy (e.g., transcription WER, section-segmentation agreement, or B-roll placement precision). If the pipeline frequently misrepresents the video, the comparison task may be measuring the interface's ability to communicate inaccurate information, and the observed benefits could be tied to the specific failure modes of the generation pipeline. The authors should either add a pipeline accuracy analysis or explicitly separate the claims about interface-level comparison support from the quality of the underlying generation, ideally by reporting accuracy on a per-task basis and discussing how representation errors affect the validity of the user-study measurements.
minor comments (7)
  1. [§1, §2.2, Figure 9 caption] There are several typos: "asembling" in §1, "Boreczkly" in §2.2 should be "Boreczky", and the Figure 9 caption says "Wilcoxon text" instead of "Wilcoxon test". Please proofread the manuscript.
  2. [§4.3 footnote] The footnote "We used GPT-o to identify visually concrete keywords" is inconsistent with the rest of the paper, which uses GPT-4o; please correct the model name.
  3. [§5.1, Baseline] The baseline interface is described only by reference to Figure 7 and a short sentence. Please provide a precise list of its features (e.g., does it include a transcript view, keyword search, sorting, generation/regeneration?) so readers can assess whether the comparison isolates the aligned timeline/transcript contribution or conflates it with other UI differences.
  4. [§5.1, Procedure] The sentence "participants reviewed 10 different videos" is ambiguous; there were 10 variations generated per source video and each participant was assigned one of V1/V2 per condition. Please clarify the exact number of variations shown and how assignments were balanced.
  5. [§5.2, Figure 8] The x-axis labels in Figure 8 are confusing (the categories repeat and there is a duplicate "Q9"). Please clarify the mapping between the four column groups and the ten numbered questions, and label the significance marks individually.
  6. [§5.2, Time results] The phrase "In 5 questions that required checking multiple parts of the video" is vague; please list the specific question numbers that were significantly faster and specify whether these correspond to the multi-select questions or a different criterion.
  7. [§5.3, Authoring task] The report that VideoDiff users made 4.33 edits (SD=2.42) is not contextualized against the baseline, where most participants made no edits. A descriptive comparison of edit counts (even if not inferential) would help readers interpret the scale of the difference.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claim is an empirical interface comparison with no fitted parameters, derived equations, or load-bearing self-citation chain.

full rationale

VideoDiff makes an HCI systems claim: that aligned timeline and transcript views help creators compare and customize AI-generated video alternatives. The paper contains no mathematical derivation or fitted model; the central claim rests on a within-subjects user study (Section 5) comparing VideoDiff to a baseline interface. None of the measured outcomes (completion time, accuracy, NASA-TLX, CSI, self-rated usefulness/satisfaction) are defined in terms of VideoDiff's own outputs; they are independent user responses to two interface conditions. The formative study (Section 3) motivates design goals D1-D6, but those goals are presented as design rationale and are then tested empirically, not asserted as predictions derived from the system itself. The AI pipeline limitations acknowledged in Section 7 (transcription errors, LLM hallucinations, no visual input) are a real external-validity threat, but they do not make the comparison circular because the same generation pipeline feeds both conditions; the study isolates the interface as the manipulated variable. Self-citations to prior work (e.g., AVscript, ReelFramer, PodReels, B-Script) appear as related work and design inspiration, not as justification of the empirical results; no uniqueness theorem or load-bearing result is imported from the authors' prior work. No equation is reused as its own input, no fitted parameter is relabeled as a prediction, and no known result is merely renamed. Statistical concerns about uncorrected multiple comparisons, raised in the skeptic analysis, are a correctness risk rather than a circularity risk. Therefore the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are present. The central claims rest on three domain assumptions about baseline validity, LLM pipeline accuracy, and participant representativeness.

assumptions (3)
  • domain assumption The baseline interface is representative of existing AI video editing tools such as OpusClip and CapCut.
    The within-subjects comparison attributes differences to VideoDiff's features; if the baseline is weaker than commercial tools, the effect sizes are inflated. Section 5.1 Baseline.
  • domain assumption GPT-4o and Whisper, with the authors' prompts, generate edit suggestions and transcript segmentations accurate enough for the study tasks.
    The timeline and transcript views are built on LLM output; the Limitations section acknowledges hallucinations, transcription errors, and no visual input, which could make representations unreliable. Section 4.3, Limitations.
  • domain assumption The 12 study participants and 8 formative participants are representative of video creators who would use AI editing tools.
    Claims about creators generalize from a small convenience sample recruited via mailing lists and Upwork. Section 5.1 Participants.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VideoDiff: Human-AI Video Co-Creation with Alternatives." pith.science (2026). https://pith.science/paper/TVUBWDL4

@misc{pith2026250210190,
  author       = {Pith},
  title        = {Pith review of: VideoDiff: Human-AI Video Co-Creation with Alternatives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TVUBWDL4}},
  note         = {Machine review of arXiv:2502.10190}
}
read the original abstract

To make an engaging video, people sequence interesting moments and add visuals such as B-rolls or text. While video editing requires time and effort, AI has recently shown strong potential to make editing easier through suggestions and automation. A key strength of generative models is their ability to quickly generate multiple variations, but when provided with many alternatives, creators struggle to compare them to find the best fit. We propose VideoDiff, an AI video editing tool designed for editing with alternatives. With VideoDiff, creators can generate and review multiple AI recommendations for each editing process: creating a rough cut, inserting B-rolls, and adding text effects. VideoDiff simplifies comparisons by aligning videos and highlighting differences through timelines, transcripts, and video previews. Creators have the flexibility to regenerate and refine AI suggestions as they compare alternatives. Our study participants (N=12) could easily compare and customize alternatives, creating more satisfying results.

Figures

Figures reproduced from arXiv: 2502.10190 by the authors.

Figure 1
Figure 1. VideoDiff is a Human-AI co-creative system that supports video creators to explore multiple variations. 1) VideoDiff [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of VideoDiff: Users can view an outline of the variations in the current editing stage (a). In this figure, we [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. At each editing stage, VideoDiff provides glanceable [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Users can switch between the edited and original source timeline in the transcript (a) and timeline (b) views. This [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: At each editing stage, VideoDiff provides glance [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Using VideoDiff, users can edit a variation, recom [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: The baseline interface shares a similar UI design [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Average task completion time to answer comparison [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Distribution of the rating scores for the Baseline [PITH_FULL_IMAGE:figures/full_fig_p012_9.png]
Figure 10
Figure 10. Figure 10: In exploratory case studies, creators edited their [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: Examples of user edits with VideoDiff [PITH_FULL_IMAGE:figures/full_fig_p019_11.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design

    cs.HC 2025-08 conditional novelty 5.0 of 10

    GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.

Reference graph

Works this paper leans on

95 extracted references · 47 canonical work pages · cited by 1 Pith paper

  1. [1]

    Columbia Undergraduate Admissions. 2023. Columbia University Campus Tour. https://www.youtube.com/watch?v=OkxRkNjPHIk

  2. [2]

    Adobe. [n. d.]. Adobe Premiere Pro. https://www.adobe.com/products/premiere. html

  3. [3]

    Shm Garanganao Almeda, JD Zamfirescu-Pereira, Kyu Won Kim, Pradeep Mani Rathnam, and Bjoern Hartmann. 2024. Prompting for Discovery: Flex- ible Sense-Making for AI Art-Making with Dreamsheets. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  4. [4]

    Apple. [n. d.]. Apple iMovie. https://support.apple.com/imovie

  5. [5]

    Narges Ashtari, Andrea Bunt, Joanna McGrenere, Michael Nebeling, and Parmit K Chilana. 2020. Creating augmented and virtual reality applications: Current practices, challenges, and opportunities. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–13

  6. [6]

    Adam Baker, Carl Gutwin, Justin Matejka, and Ian Stavness. 2024. Interaction Techniques for Comparing Video. In Graphics Interface 2024 Second Deadline

  7. [7]

    Guha Balakrishnan, Frédo Durand, and John Guttag. 2015. Video diff: Highlight- ing differences between similar actions in videos. ACM Transactions on Graphics (TOG) 34, 6 (2015), 1–10

  8. [8]

    Connelly Barnes, Dan B Goldman, Eli Shechtman, and Adam Finkelstein. 2010. Video tapestries with continuous temporal zoom. In ACM SIGGRAPH 2010 papers. 1–9

Show all 95 references
  1. [9]

    Elena Benedetto, Gabriele Romano, Ilaria Torre, Mario Vallarino, and Gi- anni Viardo Vercelli. 2024. A Visual Comparison interface for educational videos. In Proceedings of the 2024 International Conference on Advanced Visual Interfaces . 1–5

  2. [10]

    John Boreczky, Andreas Girgensohn, Gene Golovchinsky, and Shingo Uchihashi

  3. [11]

    Stephen Brade, Bryan Wang, Mauricio Sousa, Sageev Oore, and Tovi Gross- man. 2023. Promptify: Text-to-image generation through interactive prompt exploration with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–14

  4. [12]

    Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The impact of multiple parallel phrase suggestions on email input and composition behaviour of native and non-native english writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–13

  5. [13]

    CapCut. [n. d.]. CapCut. https://www.capcut.com/

  6. [14]

    Minsuk Chang, Mina Huh, and Juho Kim. 2021. Rubyslippers: Supporting content- based voice navigation for how-to videos. InProceedings of the 2021 CHI conference on human factors in computing systems . 1–14

  7. [15]

    Shizhe Chen, Bei Liu, Jianlong Fu, Ruihua Song, Qin Jin, Pingping Lin, Xiaoyu Qi, Chunting Wang, and Jin Zhou. 2019. Neural storyboard artist: Visualizing stories with coherent image sequences. In Proceedings of the 27th ACM International Conference on Multimedia. 2236–2244

  8. [16]

    Erin Cherry and Celine Latulipe. 2014. Quantifying the creativity support of digital tools through the creativity support index.ACM Transactions on Computer- Human Interaction (TOCHI) 21, 4 (2014), 1–25

  9. [17]

    Peggy Chi, Tao Dong, Christian Frueh, Brian Colonna, Vivek Kwatra, and Irfan Essa. 2022. Synthesis-Assisted Video Prototyping From a Document. In Proceed- ings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–10

  10. [18]

    Peggy Chi, Nathan Frey, Katrina Panovich, and Irfan Essa. 2021. Automatic instructional video creation from a markdown-formatted tutorial. In The 34th Annual ACM Symposium on User Interface Software and Technology . 677–690

  11. [19]

    Peggy Chi, Zheng Sun, Katrina Panovich, and Irfan Essa. 2020. Automatic video creation from a web page. In Proceedings of the 33rd annual ACM symposium on user interface software and technology . 279–292

  12. [20]

    Pei-Yu Chi, Joyce Liu, Jason Linder, Mira Dontcheva, Wilmot Li, and Bjoern Hartmann. 2013. Democut: generating concise instructional videos for physi- cal demonstrations. In Proceedings of the 26th annual ACM symposium on User interface software and technology . 141–150

  13. [21]

    John Joon Young Chung, Shiqing He, and Eytan Adar. 2021. The intersection of users, roles, interactions, and technologies in creativity support tools. In Proceedings of the 2021 ACM Designing Interactive Systems Conference. 1817–1833

  14. [22]

    Opus Clip. [n. d.]. Opus Clip. https://www.opus.pro/

  15. [23]

    Descript. [n. d.]. Descript: Edit Videos & Podcasts Like a Doc | AI Video Editor. https://www.descript.com/

  16. [24]

    Steven P Dow, Alana Glassco, Jonathan Kass, Melissa Schwarz, Daniel L Schwartz, and Scott R Klemmer. 2010. Parallel prototyping leads to better design results, more divergence, and increased self-efficacy. ACM Transactions on Computer- Human Interaction (TOCHI) 17, 4 (2010), 1–24

  17. [25]

    Steven M Drucker, Georg Petschnigg, and Maneesh Agrawala. 2006. Comparing and managing multiple versions of slide presentations. In Proceedings of the 19th annual ACM symposium on User interface software and technology . 47–56

  18. [26]

    Fabio Duarte. [n. d.]. YouTube Content Creator Statistics (2024). https:// explodingtopics.com/blog/youtube-creator-stats CHI ’25, April 26-May 1, 2025, Yokohama, Japan Mina Huh, Dingzeyu Li, Kim Pimmel, Hijung Valentina Shin, Amy Pavel, and Mira Dontcheva

  19. [27]

    Kasra Ferdowsi, Ruanqianqian Huang, Michael B James, Nadia Polikarpova, and Sorin Lerner. 2024. Validating AI-Generated Code with Live Programming. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–8

  20. [28]

    C Ailie Fraser, Joy O Kim, Hijung Valentina Shin, Joel Brandt, and Mira Dontcheva

  21. [29]

    Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shecht- man, Dan B Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, and Maneesh Agrawala. 2019. Text-based editing of talking-head video. ACM Transactions on Graphics (TOG) 38, 4 (2019), 1–14

  22. [30]

    Simret Araya Gebreegziabher, Yukun Yang, Elena L Glassman, and Toby Jia-Jun Li. 2024. Supporting Co-Adaptive Machine Teaching through Human Concept Learning and Cognitive Theories. arXiv preprint arXiv:2409.16561 (2024)

  23. [31]

    Katy Ilonka Gero, Chelse Swoopes, Ziwei Gu, Jonathan K Kummerfeld, and Elena L Glassman. 2024. Supporting Sensemaking of Large Language Model Outputs at Scale. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21

  24. [32]

    Michael Gleicher, Danielle Albers, Rick Walker, Ilir Jusufi, Charles D Hansen, and Jonathan C Roberts. 2011. Visual comparison for information visualization. Information Visualization 10, 4 (2011), 289–309

  25. [33]

    Han L Han, Miguel A Renom, Wendy E Mackay, and Michel Beaudouin-Lafon

  26. [34]

    Sandra G Hart. 2006. NASA-task load index (NASA-TLX); 20 years later. In Proceedings of the human factors and ergonomics society annual meeting , Vol. 50. Sage publications Sage CA: Los Angeles, CA, 904–908

  27. [35]

    Bernd Huber, Hijung Valentina Shin, Bryan Russell, Oliver Wang, and Gautham J Mysore. 2019. B-script: Transcript-based b-roll video editing with recommenda- tions. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 1–11

  28. [36]

    In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems

    Textlets: Supporting constraints and consistency in text documents. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13

  29. [37]

    Mina Huh, Saelyne Yang, Yi-Hao Peng, Xiang’Anthony’ Chen, Young-Ho Kim, and Amy Pavel. 2023. AVscript: Accessible Video Editing with Audio-Visual Scripts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17

  30. [38]

    invideo AI. [n. d.]. invideo AI Script Generator. https://invideo.io/tools/script- generator/

  31. [39]

    Mina Huh, Yi-Hao Peng, and Amy Pavel. 2023. GenAssist: Making image gen- eration accessible. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–17

  32. [40]

    Joy Kim, Mira Dontcheva, Wilmot Li, Michael S Bernstein, and Daniela Steinsapir

  33. [41]

    Janin Koch, Nicolas Taffin, Michel Beaudouin-Lafon, Markku Laine, Andrés Lucero, and Wendy E Mackay. 2020. Imagesense: An intelligent collaborative ideation tool to support diverse human-computer partnerships. Proceedings of the ACM on human-computer interaction 4, CSCW1 (2020), 1–27

  34. [42]

    Haojian Jin, Yale Song, and Koji Yatani. 2017. Elasticplay: Interactive video summarization with dynamic time budgets. In Proceedings of the 25th ACM international conference on Multimedia . 1164–1172

  35. [43]

    Philippe Laban, Jesse Vig, Marti Hearst, Caiming Xiong, and Chien-Sheng Wu

  36. [44]

    Mackenzie Leake, Abe Davis, Anh Truong, and Maneesh Agrawala. 2017. Com- putational video editing for dialogue-driven scenes. ACM Trans. Graph. 36, 4 (2017), 130–1

  37. [45]

    Mackenzie Leake, Hijung Valentina Shin, Joy O Kim, and Maneesh Agrawala. 2020. Generating audio-visual slideshows from text articles using word concreteness. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–11

  38. [46]

    Steffen Koch, Markus John, Michael Wörner, Andreas Müller, and Thomas Ertl

  39. [47]

    Michael Xieyang Liu, Tongshuang Wu, Tianying Chen, Franklin Mingzhe Li, Aniket Kittur, and Brad A Myers. 2024. Selenite: Scaffolding Online Sensemak- ing with Comprehensive Overviews Elicited from Large Language Models. In Proceedings of the CHI Conference on Human Factors in ...

  40. [48]

    Vivian Liu, Tao Long, Nathan Raw, and Lydia Chilton. 2023. Generative disco: Text-to-video generation for music visualization. arXiv preprint arXiv:2304.08551 (2023)

  41. [49]

    Ference Marton. 2014. Necessary conditions of learning . Routledge

  42. [50]

    Justin Matejka, Michael Glueck, Erin Bradner, Ali Hashemi, Tovi Grossman, and George Fitzmaurice. 2018. Dream lens: Exploration and visualization of large- scale generative design datasets. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12

  43. [51]

    Justin Matejka, Tovi Grossman, and George Fitzmaurice. 2014. Video lens: rapid playback and exploration of large video collections and associated metadata. In Proceedings of the 27th annual ACM symposium on User interface software and technology. 541–550

  44. [52]

    Ching Liu, Juho Kim, and Hao-Chuan Wang. 2018. ConceptScape: Collaborative concept mapping for video learning. In Proceedings of the 2018 CHI conference on human factors in computing systems . 1–12

  45. [53]

    Piotr Mirowski, Kory W Mathewson, Jaylen Pittman, and Richard Evans. 2023. Co-writing screenplays and theatre scripts with language models: Evaluation by industry professionals. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–34

  46. [54]

    Zheng Ning, Brianna L Wimer, Kaiwen Jiang, Keyi Chen, Jerrick Ban, Yapeng Tian, Yuhang Zhao, and Toby Jia-Jun Li. 2024. SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers. In Proceedings of the CHI Conference on Hu...

  47. [55]

    OpenAI. [n. d.]. DALL-E 3. https://openai.com/index/dall-e-3/

  48. [56]

    OpenAI. 2024. Sora OpenAI video generation model. https://openai.com/sora

  49. [57]

    2024 (accessed Sept 11, 2024)

    Remotion.dev. 2024 (accessed Sept 11, 2024). Remotion: Make videos program- matically in react. https://www.remotion.dev/

  50. [58]

    Midjourney. [n. d.]. Midjourney. https://www.midjourney.com/

  51. [59]

    Brittany Rose. 2024. Huge ALDI Grocery Haul | Shop with me and haul!! https: //www.youtube.com/watch?v=ti43mILvSug

  52. [60]

    Brittany Rose. 2024. WEEKLY GROCERY HAUL AND MEAL PLAN | Shop with me!! https://www.youtube.com/watch?v=IxPofzudXxM

  53. [61]

    Steve Rubin and Maneesh Agrawala. 2014. Generating emotionally relevant musical scores for audio stories. InProceedings of the 27th annual ACM symposium on User interface software and technology . 439–448

  54. [62]

    Sangho Suh, Meng Chen, Bryan Min, Toby Jia-Jun Li, and Haijun Xia. 2024. Luminate: Structured Generation and Exploration of Design Space with Large Language Models for Human-AI Co-Creation. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–26

  55. [63]

    Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling multilevel exploration and sensemaking with large language models. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18

  56. [64]

    Mohi Reza, Nathan Laundry, Ilya Musabirov, Peter Dushniku, Zhi Yuan Yu, Kashish Mittal, Tovi Grossman, Michael Liut, Anastasia Kuzminykh, Joseph Jay Williams, et al. 2023. ABScribe: Rapid Exploration of Multiple Writing Variations in Human-AI Co-Writing Tasks using Large Langu...

  57. [65]

    TED. [n. d.]. Does Working Hard Really Make You a Good Person? | Azim Shariff | TED. https://www.youtube.com/watch?v=z_ca983qJcI

  58. [66]

    Atima Tharatipyakul and Hyowon Lee. 2018. Towards a better video comparison: comparison as a way of browsing the video contents. In Proceedings of the 30th Australian Conference on Computer-Human Interaction . 349–353

  59. [67]

    Bekzat Tilekbay, Saelyne Yang, Michal Adam Lewkowicz, Alex Suryapranata, and Juho Kim. 2024. ExpressEdit: Video Editing with Natural Language and Sketching. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 515–536

  60. [68]

    Anh Truong and Maneesh Agrawala. 2019. A Tool for Navigating and Editing 360 Video of Social Conversations into Shareable Highlights.. In Graphics Interface. 14–1

  61. [69]

    Anh Truong, Floraine Berthouzoz, Wilmot Li, and Maneesh Agrawala. 2016. Quickcut: An interactive tool for editing narrated video. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology . 497–507

  62. [70]

    Amanda Swearngin, Chenglong Wang, Alannah Oleson, James Fogarty, and Amy J Ko. 2020. Scout: Rapid exploration of interface layout alternatives through high-level design constraints. In Proceedings of the 2020 CHI conference on human factors in computing systems . 1–13

  63. [71]

    Vimeo. [n. d.]. Vimeo AI-Powered Video Platform. https://vimeo.com/

  64. [72]

    Vizard. [n. d.]. Vizard. https://vizard.ai/

  65. [73]

    Samangi Wadinambiarachchi, Ryan M Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso. 2024. The Effects of Generative AI on Design Fixation and Divergent Thinking. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–18

  66. [74]

    It Felt Like Having a Second Mind

    Qian Wan, Siying Hu, Yu Zhang, Piaohong Wang, Bo Wen, and Zhicong Lu. 2024. " It Felt Like Having a Second Mind": Investigating Human-AI Co-creativity in Prewriting with Large Language Models. Proceedings of the ACM on Human- Computer Interaction 8, CSCW1 (2024), 1–26

  67. [75]

    Bryan Wang, Zeyu Jin, and Gautham Mysore. 2022. Record Once, Post Every- where: Automatic Shortening of Audio Stories for Social Media. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology . 1–11. VideoDiff: Human-AI Video Co-Creation with ...

  68. [76]

    Babish Culinary Universe. [n. d.]. Eggs Benedict | Basics with Babish Live. https: //www.youtube.com/watch?v=ROAmMJeTO6Y

  69. [77]

    Miao Wang, Guo-Wei Yang, Shi-Min Hu, Shing-Tung Yau, Ariel Shamir, et al

  70. [78]

    Sitong Wang, Samia Menon, Tao Long, Keren Henderson, Dingzeyu Li, Kevin Crowston, Mark Hansen, Jeffrey V Nickerson, and Lydia B Chilton. 2024. Reel- Framer: Human-AI co-creation for news-to-video translation. In Proceedings of the CHI Conference on Human Factors in Computing S...

  71. [79]

    Sitong Wang, Zheng Ning, Anh Truong, Mira Dontcheva, Dingzeyu Li, and Lydia B Chilton. 2024. PodReels: Human-AI Co-Creation of Video Podcast Teasers. In Proceedings of the 2024 ACM Designing Interactive Systems Conference . 958–974

  72. [80]

    Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, et al

  73. [81]

    Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. In Proceedings of the 2022 CHI conference on human factors in computing systems . 1–22

  74. [82]

    Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces. 699–714

  75. [83]

    Liwenhan Xie, Zhaoyu Zhou, Kerun Yu, Yun Wang, Huamin Qu, and Siming Chen. 2023. Wakey-Wakey: Animate Text by Mimicking Characters in a GIF. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–14

  76. [84]

    Saelyne Yang, Jisu Yim, Aitolkyn Baigutanova, Seoyoung Kim, Minsuk Chang, and Juho Kim. 2022. SoftVideo: Improving the Learning Experience of Soft- ware Tutorial Videos with Collective Interaction Data. In Proceedings of the 27th International Conference on Intelligent User In...

  77. [85]

    Catherine Yeh, Gonzalo Ramos, Rachel Ng, Andy Huntington, and Richard Banks

  78. [86]

    Zheng Zhang, Zheng Ning, Chenliang Xu, Yapeng Tian, and Toby Jia-Jun Li. 2023. PEANUT: A Human-AI Collaborative Tool for Annotating Audio-Visual Data. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 1–18. CHI ’25, April 26-May 1, 2025...

  79. [90]

    Haijun Xia, Jennifer Jacobs, and Maneesh Agrawala. 2020. Crosscast: adding visuals to audio travel podcasts. InProceedings of the 33rd annual ACM symposium on user interface software and technology . 735–746

  80. [94]

    arXiv preprint arXiv:2402.08855 (2024)

    GhostWriter: Augmenting Collaborative Human-AI Writing Experiences Through Personalization and Agency. arXiv preprint arXiv:2402.08855 (2024)

  81. [2000]

    In Proceedings of the SIGCHI conference on Human factors in computing systems

    An interactive comic book presentation for exploring video. In Proceedings of the SIGCHI conference on Human factors in computing systems . 185–192

  82. [2014]

    IEEE transactions on visualization and computer graphics 20, 12 (2014), 1723–1732

    VarifocalReader—in-depth visual analysis of large text documents. IEEE transactions on visualization and computer graphics 20, 12 (2014), 1723–1732

  83. [2015]

    In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems

    Motif: Supporting novice creativity through expert patterns. In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems . 1211–1220

  84. [2019]

    ACM Trans

    Write-a-video: computational video montage from themed text. ACM Trans. Graph. 38, 6 (2019), 177–1

  85. [2020]

    In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems

    Temporal segmentation of creative live streams. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12

  86. [2022]

    In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency

    Taxonomy of risks posed by language models. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency . 214–229

  87. [2024]

    In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology

    Beyond the chat: Executable and verifiable text-editing with llms. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–23

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.