REVIEW 3 major objections 6 minor 1 cited by
An Exploratory Study on Multi-modal Generative AI in AR Storytelling
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A study of 223 AR videos maps which AI-generated media fits which story element, and 30 storytellers confirm the mapping in practice.
desk verdict Useful exploratory mapping of modality preferences for AIGC in AR storytelling, but the headline map is tangled up with the uneven quality of the specific generative models used. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a two-dimensional design space: five Modalities (Text, Audio, Image, Video, 3D) crossed with four atomic Elements (Character, Background, Sentiment, Development). The authors derive it by open-coding 223 YouTube AR storytelling videos, then use it to structure a testbed in which a storyteller selects a sentence, chooses a modality, and receives generator output from Stable Diffusion (image), MusicGen (audio), a motion-diffusion plus Text2Video-Zero pipeline (video), and a text-to-3D model; the AR interface then triggers the saved content by speech while the narrator gestures with hand-tracked interaction. The design space organizes both the corpus analysis and the study tasks, so the preferences reported are preferences over cells of this five-by-four grid.
What would settle it
A systematically sampled corpus of AR storytelling videos that reveals a common modality outside the five (for example haptic feedback) or a common atomic element outside the four (for example audience interaction) would falsify the taxonomy's completeness; likewise, a replication with a state-of-the-art text-to-video model that erases the current video-for-development-only preference would falsify the modality-element mapping as a stable property of AR storytelling.
Extended reading notes
Core claim
The central claim is that multi-modal AIGC is suitable for AR storytelling, provided the modality is matched to the story element. From the 223-video analysis the authors derive a design space of five modalities and four atomic elements (Character, Background, Sentiment, Development). Their two studies, each with 15 participants, show that images are strongly preferred for characters (51%), backgrounds (47%), and sentiment (47%); video is preferred for development (40%), with text close behind (30%); and 3D is a secondary choice for characters. They also find that participants rate generated text highest in quality (4.43/5) and video lowest (2.8/5), and that while co-creating with AI feels fast and enjoyable, guiding the generation to match intention is the hardest part. The authors further report that participants could mostly tell AIGC apart from human-made content but said it did not hurt the storytelling, and that cross-modal inconsistencies and literal misinterpretation of metaphors are the main blockers to wider use.
Load-bearing premise
The findings stand on the assumption that the 223 manually collected YouTube videos represent the space of AR storytelling well enough for the five-by-four taxonomy to be the right frame for the testbed and the study tasks; the authors say plainly that the corpus was not collected by systematic search.
Editorial extensions
If this is right
- Future AR storytelling authoring tools can present images as the default augmentation for characters, backgrounds, and emotional tone, and reserve video for temporal development.
- Because participants found text the clearest and most reliable output, tools should keep text as a fallback or complement even when visuals are preferred.
- The low video-quality ratings imply that improvements in text-to-video generation could enlarge the role of video beyond development into other elements.
- The repeated difficulty in prompting suggests authoring systems need example-based, iterative, or in-AR prompting beyond plain text.
- The observed cross-modal misalignment argues for generating each story entity once and sharing its representation across modalities.
Reading between the lines
- The modality-element preferences are partly confounded by current model quality: participants avoided video for small details because outputs were poor, so a replication with a stronger text-to-video model might shift the mapping.
- The four-element taxonomy may be reusable beyond AR, as a general vocabulary for choosing generative media in slideware, virtual reality, or interactive fiction.
- The non-systematic YouTube sampling means the taxonomy should be treated as a starting hypothesis; a systematic corpus study could add elements such as interactivity or user choice that this work deliberately excludes.
- If cross-modal alignment is solved, the same testbed pattern could support live 'when to augment' decisions during presentations rather than pre-authored augmentation only.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an exploratory study of multi-modal generative AI (GenAI) for AR storytelling. The authors analyze 223 YouTube videos to derive a design space with two dimensions: five modalities (text, audio, image, video, 3D) and four atomic elements (Character, Background, Sentiment, Development). They implement a testbed that generates content in these modalities from textual narratives and displays it through an AR interface, then conduct two studies with 30 experienced storytellers and presenters (split into two groups of 15). The reported findings are a modality-to-element preference mapping (image preferred for Character, Background, and Sentiment; video preferred for Development), quality ratings of generated content, Likert-scale evaluations of co-creation interactions and suitability, and qualitative insights about alignment, selective augmentation, and context awareness. The paper concludes with design considerations for future AR storytelling systems with GenAI.
Significance. If the claims are properly qualified, the paper makes a useful exploratory contribution: it provides a taxonomy of modalities and story elements, a working testbed for studying GenAI-based AR authoring, and an empirical snapshot of author preferences. Strengths include the explicit acknowledgement of corpus-selection limitations, the reporting of inter-rater agreement for video filtering (κ=0.76), and the inclusion of the full story stimuli in an appendix. However, the central preference mapping is coupled to the specific generative models used in the testbed, and the claimed 'empirical comparison' with human-generated content is not supported by the data. Because the load-bearing conclusions overreach beyond what the design can establish, the paper needs revision before its claims can be accepted as stated.
major comments (3)
- [§1, Contribution bullet 3; §5.5.2] The contribution list claims an 'empirical comparison of AIGC with human-generated content,' but the manuscript reports no human-generated baseline and no systematic comparison. Section 5.5.2 describes only participants' subjective impressions ('Some of the participants felt the content was as good as human-generated content') and the authors' argument that video quality reflects algorithmic limitations rather than modality suitability. This is not an empirical comparison; the claim should be removed or replaced with a condition in which participants rate matched human-generated content.
- [§5.1, Fig. 5; §5.5.1; Fig. 6] The central modality-to-element preference mapping is confounded with per-modality AIGC quality. Preferences in Figure 5 were expressed after previewing outputs from specific models (Stable Diffusion for images, MDM plus Text2Video-Zero for video, MusicGen for audio, DreamFusion for 3D), and Figure 6 reports that video quality was rated lowest (AVG=2.8, SD=1.13). Section 5.5.1 states that 'the quality of the video was not good enough' and that participants used video only where 'details are safe to ignore.' The rebuttal in §5.5.2, that this is 'due to limitations in algorithmic development rather than the suitability of the video as a modality,' is not testable from the current data because each modality is instantiated by a single model. The paper should either frame the conclusions as preferences under current AIGC quality or include a controlled comparison with quality matched across modalities.
- [§3.1.1, §3.1.3] The design-space taxonomy is load-bearing for the study: it determines the elements and modalities used in Study 1's pre-highlighted elements, the testbed capabilities, and the interpretation of participant preferences. However, the corpus was assembled through a manual, non-systematic search (stated explicitly in §3.1.1), and inter-rater reliability is reported only for the filtering step (κ=0.76), not for the open coding of modalities/elements or for the final placement of videos in the design space. The study therefore provides no independent check that the four atomic elements are complete or reliably identifiable, which limits the generality of the subsequent preference findings. Please report coding reliability for the design-space dimensions or validate the taxonomy on an independent sample.
minor comments (6)
- [Abstract; §4.2.1] The abstract reports 'N=30' without clarifying that this number was split into two studies of 15 participants each; please state this split explicitly for accuracy.
- [Fig. 6] The x-axis label contains a typo: 'Accepable' should be 'Acceptable.'
- [§5.5.2] The heading 'Empirical comparison to human-generated content' overstates what is presented, because no human-generated baseline is included; consider renaming the subsection to 'Participants' perceptions of AIGC versus human-generated content.'
- [§8 Conclusion] The first sentence reads 'Mutli-modal Gen-AI'; this should be corrected to 'Multi-modal Gen-AI.'
- [§5.3, §5.4] The Mann-Whitney U tests are used to support the statement that there is 'no significant difference' between the two study conditions; however, with N=15 per group, non-significance does not establish equivalence, and no effect sizes or confidence intervals are reported. Please temper these interpretations.
- [§6.1.1, Fig. 11] The caption states 'All images are generated by ChatGPT,' but ChatGPT alone is not an image-generation model; please specify the image model used (e.g., DALL-E through ChatGPT) for reproducibility.
Circularity Check
No significant circularity: the design space is an empirical input, the preference mapping is directly measured, and same-author citations are not load-bearing.
full rationale
The paper's central chain is empirical rather than derivational: a 223-video corpus is open-coded into a modality/element design space (Section 3), a testbed is implemented from that space (Section 4.1), and two user studies elicit preferences and quality ratings (Section 5). No result is computed from the design space by construction. The Fig. 5 modality-to-element mapping is a direct tally of participant selections in Study 1 ('For each element previously highlighted in the interface, the participants are asked to select at least one modality of augmentation'), not a quantity fitted or assumed from the taxonomy. The Fig. 6 quality ratings are likewise direct Likert responses. The paper explicitly flags its corpus as non-systematic ('We do not claim that our corpus was collected by a systematic search'), which is an honest sampling limitation rather than a circular step. The video-quality confound (video rated AVG=2.8 and the authors arguing in Section 5.5.2 that this reflects 'limitations in the algorithmic development ... rather than the suitability of the video') is an internal-validity concern, not a circular reduction, because the suitability claim is not derived from the quality rating. Same-author references (e.g., [92]) appear only in related-work or discussion context and do not carry the central claim. Hence no circularity.
Assumptions & free parameters
assumptions (5)
- domain assumption A non-systematic YouTube corpus of 223 videos sufficiently represents the space of AR storytelling.
- ad hoc to paper The four atomic elements Character, Background, Sentiment, and Development decompose storytelling content in a way that covers the study stories.
- domain assumption Self-reported Likert-scale ratings and semi-structured interviews measure the suitability of AIGC for AR storytelling.
- domain assumption Desktop webcam AR is a sufficient testbed to reveal the impact of AIGC on AR storytelling.
- domain assumption The bundled generative models are representative of the current quality of each modality.
invented entities (2)
-
Atomic elements taxonomy (Character, Background, Sentiment, Development)
-
Modality-element preference mapping
Cite this review
Pith. "Pith review of An Exploratory Study on Multi-modal Generative AI in AR Storytelling." pith.science (2026). https://pith.science/paper/FXMCDF2G
@misc{pith2026250515973,
author = {Pith},
title = {Pith review of: An Exploratory Study on Multi-modal Generative AI in AR Storytelling},
year = {2026},
howpublished = {\url{https://pith.science/paper/FXMCDF2G}},
note = {Machine review of arXiv:2505.15973}
}
read the original abstract
Storytelling in AR has gained attention due to its multi-modality and interactivity. However, generating multi-modal content for AR storytelling requires expertise and efforts for high-quality conveyance of the narrator's intention. Recently, Generative-AI (GenAI) has shown promising applications in multi-modal content generation. Despite the potential benefit, current research calls for validating the effect of AI-generated content (AIGC) in AR Storytelling. Therefore, we conducted an exploratory study to investigate the utilization of GenAI. Analyzing 223 AR videos, we identified a design space for multi-modal AR Storytelling. Based on the design space, we developed a testbed facilitating multi-modal content generation and atomic elements in AR Storytelling. Through two studies with N=30 experienced storytellers and live presenters, we 1. revealed participants' preferences for modalities, 2. evaluated the interactions with AI to generate content, and 3. assessed the quality of the AIGC for AR Storytelling. We further discussed design considerations for future AR Storytelling with GenAI.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs
An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.
Reference graph
Works this paper leans on
-
[1]
Rameen Abdal, Peihao Zhu, John Femiani, Niloy Mitra, and Peter Wonka. 2022. Clip2stylegan: Unsupervised extraction of stylegan edit directions. InACM SIGGRAPH 2022 conference proceedings. 1–9
2022
-
[2]
Chaitanya Ahuja and Louis-Philippe Morency. 2019. Language2pose: Natural language grounded pose forecasting. In 2019 International Conference on 3D Vision (3DV). IEEE, 719–728
2019
-
[3]
Victor Nikhil Antony and Chien-Ming Huang. 2023. ID. 8: Co-Creating Visual Stories with Generative AI.arXiv preprint arXiv:2309.14228(2023)
arXiv 2023
-
[4]
National Storytelling Association et al. 2012. What Storytelling is. An attempt at defining the art form
2012
-
[5]
Paulo Bala, Stuart James, Alessio Del Bue, and Valentina Nisi. 2022. Writing with (Digital) Scissors: Designing a Text Editing Tool for Assisted Storytelling Using Crowd-Generated Content. InInternational Conference on Interactive Digital Storytelling. Springer, 139–158
2022
-
[6]
Valentin Bauer, Anna Nagele, Chris Baume, Tim Cowlishaw, Henry Cooke, Chris Pike, and Patrick GT Healey. 2019. Designing an interactive and collaborative experience in audio augmented reality. InVirtual Reality and Augmented Reality: 16th EuroVR International Conference, EuroVR 2019, Tallinn, Estonia, October 23–25, 2019, Proceedings 16. Springer, 305–311
2019
-
[7]
2010.Storytelling for social justice: Connecting narrative and the arts in antiracist teaching
Lee Anne Bell. 2010.Storytelling for social justice: Connecting narrative and the arts in antiracist teaching. Routledge
2010
-
[8]
Eden Bensaid, Mauro Martino, Benjamin Hoover, and Hendrik Strobelt. 2021. Fairytailor: A multimodal generative framework for storytelling.arXiv preprint arXiv:2108.04324(2021)
arXiv 2021
Show all 118 references
-
[9]
Sukanya Bhattacharjee and Parag Chaudhuri. 2020. A survey on sketch based content creation: from the desktop to virtual and augmented reality. InComputer Graphics Forum, Vol. 39. Wiley Online Library, 757–780
2020
-
[10]
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan. 2019. High fidelity speech synthesis with adversarial networks.arXiv preprint arXiv:1909.11646 (2019)
2019 arXiv
-
[11]
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis
-
[12]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258(2021)
2021 arXiv
-
[13]
I address race because race addresses me
Anjuli Joshi Brekke, Ralina Joseph, and Naheed Gina Aaftaab. 2021. “I address race because race addresses me”: women of color show receipts through digital storytelling.Review of Communication21, 1 (2021), 44–57
2021
-
[14]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural , Vol. 1, No. 1, Article . Publication date: Septembe...
2020
-
[15]
Licia Calvi. 2020. What do we know about AR Storytelling?. InProceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good. 278–280
2020
-
[16]
Jorge Camba, Manuel Contero, and Gustavo Salvador-Herranz. 2014. Desktop vs. mobile: A comparative study of augmented reality systems for engineering visualizations in education. In2014 IEEE Frontiers in Education Conference (FIE) Proceedings. IEEE, 1–8
2014
-
[17]
Weifeng Chen, Jie Wu, Pan Xie, Hefeng Wu, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. 2023. Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models.arXiv preprint arXiv:2305.13840(2023)
2023 arXiv
-
[18]
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: visual sketching of story generation with pretrained language models. InCHI Conference on Human Factors in Computing Systems Extended Abstracts. 1–4
2022
-
[19]
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023. Simple and Controllable Music Generation.arXiv preprint arXiv:2306.05284(2023)
2023 arXiv
-
[20]
Delneshin Danaei, Hamid R Jamali, Yazdan Mansourian, and Hassan Rastegarpour. 2020. Comparing reading comprehension between children reading augmented reality and print storybooks.Computers & Education153 (2020), 103900
2020
-
[21]
Dimitrios Darzentas, Martin Flintham, and Steve Benford. 2018. Object-focused mixed reality storytelling: technology- driven content creation and dissemination for engaging user experiences. InProceedings of the 22nd Pan-Hellenic Conference on Informatics. 278–281
2018
-
[22]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805(2018)
2018 arXiv
-
[23]
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022. Cogview2: Faster and better text-to-image generation via hierarchical transformers.Advances in Neural Information Processing Systems35 (2022), 16890–16902
2022
-
[24]
Runlin Duan, Shao-Kang Hsia, Yuzhao Chen, Yichen Hu, Ming Yin, and Karthik Ramani. 2025. Investigating Creativity in Humans and Generative AI Through Circles Exercises.arXiv preprint arXiv:2502.07292(2025)
2025 arXiv
-
[25]
Runlin Duan, Nachiketh Karthik, Jingyu Shi, Rahul Jain, Maria C Yang, and Karthik Ramani. 2024. ConceptVis: Generating and Exploring Design Concepts for Early-Stage Ideation Using Large Language Model. InInternational Design Engineering Technical Conferences and Computers and ...
2024
-
[26]
Runlin Duan, Chenfei Zhu, Yuzhao Chen, Yichen Hu, Jingyu Shi, and Karthik Ramani. 2025. DesignFromX: Empower- ing Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products.arXiv preprint arXiv:2505.11666(2025)
2025 arXiv
-
[27]
FastAPI. [n. d.]. FastAPI. https://fastapi.tiangolo.com
-
[28]
2005.Storytelling
Klaus Fog, Christian Budtz, and Baris Yakaboylu. 2005.Storytelling. Springer
2005
-
[29]
Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022. StyleGAN-NADA: CLIP-guided domain adaptation of image generators.ACM Transactions on Graphics (TOG)41, 4 (2022), 1–13
2022
-
[30]
Ze Gao, Anqi Wang, Pan Hui, and Tristan Braud. 2022. Bridging curatorial intent and visiting experience: Using ar guidance as a storytelling tool. InProceedings of the 18th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and its Applications in Industry. 1–10
2022
-
[31]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets.Advances in neural information processing systems27 (2014)
2014
-
[32]
Ariel Han and Zhenyao Cai. 2023. Design implications of generative AI systems for visual storytelling for young learners. InProceedings of the 22nd Annual ACM Interaction Design and Children Conference. 470–474
2023
-
[33]
Linda Hirsch, Robin Welsch, Beat Rossmy, and Andreas Butz. 2022. Embedded AR Storytelling Supports Active Indexing at Historical Places. InSixteenth International Conference on Tangible, Embedded, and Embodied Interaction. 1–12
2022
-
[34]
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models.arXiv preprint arXiv:2210.02303(2022)
2022 arXiv
-
[35]
Lissa Holloway-Attaway and Lars Vipsjö. 2020. Using augmented reality, gaming technologies, and transmedial storytelling to develop and co-design local cultural heritage experiences.Visual Computing for Cultural Heritage (2020), 177–204
2020
-
[36]
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al. 2020. spaCy: Industrial-strength natural language processing in python. (2020)
2020
-
[37]
Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu, and Karthik Ramani. 2025. GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality.arXiv preprint arXiv:2505.05441(2025). , Vol. 1, No. 1, Article . Public...
2025 arXiv
-
[38]
Yongquan Hu, Mingyue Yuan, Kaiqi Xian, Don Samitha Elvitigala, and Aaron Quigley. 2023. Exploring the Design Space of Employing AI-Generated Content for Augmented Reality Display.arXiv preprint arXiv:2303.16593(2023)
2023 arXiv
-
[39]
Reinis Indans, Eva Hauthal, and Dirk Burghardt. 2019. Towards an audio-locative mobile application for immersive storytelling.KN-Journal of Cartography and Geographic Information69 (2019), 41–50
2019
-
[40]
Myunggeun Ji and Junchul Chun. 2020. A Sketch-based 3D Object Retrieval Approach for Augmented Reality Models Using Deep Learning.Journal of Korean Society for Internet Information21, 1 (2020)
2020
-
[41]
BRENNAN JONES, YAN XU, MARY ANNE HOOD, MOHAMMAD SHAHIDUL KADER, and HAMID EGHBALZADEH. [n. d.]. Using Generative AI to Produce Situated Action Recommendations in Augmented Reality for High-Level Goals. ([n. d.])
-
[42]
Kwanghee Jung, Vinh T Nguyen, and Jaehoon Lee. 2021. Blocklyxr: An interactive extended reality toolkit for digital storytelling.Applied Sciences11, 3 (2021), 1073
2021
-
[43]
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018. Efficient neural audio synthesis. InInternational Conference on Machine Learning. PMLR, 2410–2419
2018
-
[44]
Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks.Advances in Neural Information Processing Systems34 (2021), 852–863
2021
-
[45]
Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4401–4410
2019
-
[46]
Sarah Ketchell, Winyu Chinthammit, and Ulrich Engelke. 2019. Situated storytelling with SLAM enabled augmented reality. InProceedings of the 17th International Conference on Virtual-Reality Continuum and Its Applications in Industry. 1–9
2019
-
[47]
Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. 2023. Text2video-zero: Text-to-image diffusion models are zero-shot video genera- tors.arXiv preprint arXiv:2303.13439(2023)
2023 arXiv
-
[48]
Doyeon Kim, Donggyu Joo, and Junmo Kim. 2020. Tivgan: Text to image to video generation with step-by-step evolutionary generator.IEEE Access8 (2020), 153113–153122
2020
-
[49]
Jihoon Kim, Jiseob Kim, and Sungjoon Choi. 2023. Flame: Free-form language-based motion synthesis & editing. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 8255–8263
2023
-
[50]
Kiyoung Kim, Noh-young Park, and Woontack Woo. 2014. Vision-based all-in-one solution for augmented reality and its storytelling applications.The Visual Computer30 (2014), 417–429
2014
-
[51]
Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114(2013)
2013 arXiv
-
[52]
Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2020. Diffwave: A versatile diffusion model for audio synthesis.arXiv preprint arXiv:2009.09761(2020)
2020 arXiv
-
[53]
Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebis- son, Yoshua Bengio, and Aaron C Courville. 2019. Melgan: Generative adversarial networks for conditional waveform synthesis.Advances in neural information process...
2019
-
[54]
Tomas Lawton, Francisco J Ibarrola, Dan Ventura, and Kazjon Grace. 2023. Drawing with Reframer: Emergence and Control in Co-Creative AI. InProceedings of the 28th International Conference on Intelligent User Interfaces. 264–277
2023
-
[55]
Kyungjun Lee, Hong Li, Muhammad Rizky Wellyanto, Yu Jiang Tham, Andrés Monroy-Hernández, Fannie Liu, Brian A Smith, and Rajan Vaish. 2023. Exploring Immersive Interpersonal Communication via AR.Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (2023), 1–25
2023
-
[56]
Jian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi, and Ryo Suzuki. 2022. RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–12
2022
-
[57]
Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang
Xingyu" Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang" Anthony" Chen, and Ruofei Du
-
[58]
Ziyi Liu, Zhengzhe Zhu, Lijun Zhu, Enze Jiang, Xiyun Hu, Kylie A Peppler, and Karthik Ramani. 2024. Classmeta: Designing interactive virtual classmate to promote VR classroom participation. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–17
2024
-
[59]
InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Visual Captions: Augmenting Verbal Communication With On-the-Fly Visuals. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20
2023
-
[60]
Tiago Madeira, Bernardo Marques, Pedro Neves, Paulo Dias, and Beatriz Sousa Santos. 2022. Comparing desktop vs. Mobile interaction for the creation of pervasive augmented reality experiences.Journal of Imaging8, 3 (2022), 79
2022
-
[61]
Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Ying Shan, Xiu Li, and Qifeng Chen. 2023. Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos.arXiv preprint arXiv:2304.01186(2023)
2023 arXiv
-
[62]
Dimitrios Markouzis and Georgios Fessakis. 2015. Interactive storytelling and mobile augmented reality applications for learning and entertainment—a rapid prototyping perspective. In2015 International Conference on Interactive Mobile , Vol. 1, No. 1, Article . Publication date...
2015
-
[63]
Jenny Mandelbaum. 2012. Storytelling in conversation.The handbook of conversation analysis(2012), 492–507
2012
-
[64]
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. 2016. SampleRNN: An unconditional end-to-end neural audio generation model.arXiv preprint arXiv:1612.07837(2016)
2016 arXiv
-
[65]
MediaPipe. [n. d.]. MediaPipe. https://mediapipe.dev/
-
[66]
Ron Mokady, Omer Tov, Michal Yarom, Oran Lang, Inbar Mosseri, Tali Dekel, Daniel Cohen-Or, and Michal Irani. 2022. Self-distilled stylegan: Towards generation from internet photos. InACM SIGGRAPH 2022 Conference Proceedings. 1–9
2022
-
[67]
2019.Digital Storytelling 4e: A creator’s guide to interactive entertainment
Carolyn Handler Miller. 2019.Digital Storytelling 4e: A creator’s guide to interactive entertainment. CRC Press
2019
-
[68]
Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023. Conditional Image-to-Video Generation with Latent Flow Diffusion Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18444–18455
2023
-
[69]
mozilla. [n. d.]. Web Speech API. https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API
-
[70]
Jennifer O’Meara and Kata Szita. 2021. AR cinema: Visual storytelling and embodied experiences with augmented reality filters and backgrounds.PRESENCE: Virtual and Augmented Reality30 (2021), 99–123
2021
-
[71]
Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741(2021)
2021 arXiv
-
[72]
Augusto Palombini. 2017. Storytelling and telling history. Towards a grammar of narratives for Cultural Heritage dissemination in the Digital Era.Journal of cultural heritage24 (2017), 134–139
2017
-
[73]
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbren- ner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio.arXiv preprint arXiv:1609.03499(2016)
2016 arXiv
-
[74]
Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021. Styleclip: Text-driven manipulation of stylegan imagery. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2085–2094
2021
-
[75]
Seung-Bo Park, Jason J Jung, and EunSoon You. 2015. Storytelling of collaborative learning system on augmented reality.New Trends in Computational Collective Intelligence(2015), 139–147
2015
-
[76]
Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao. 2020. Non-autoregressive neural text-to-speech. InInternational conference on machine learning. PMLR, 7586–7598
2020
-
[77]
John V Pavlik and Frank Bridges. 2013. The emergence of augmented reality (AR) as a storytelling medium in journalism.Journalism & Communication Monographs15, 1 (2013), 4–59
2013
-
[78]
Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. Anglekindling: Supporting journalistic angle ideation with large language models. InProceedings of the 2023 CHI Conference on ...
2023
-
[79]
Eric E Peterson and Kristin M Langellier. 2006. Communication as storytelling.Communication as... Perspectives on Theory(2006), 123–131
2006
-
[80]
Matthias Plappert, Christian Mandery, and Tamim Asfour. 2016. The KIT motion-language dataset.Big data4, 4 (2016), 236–252
2016
-
[81]
Judith Pintar. 2023. Invisible, aesthetic, and enrolled listeners across storytelling modalities: Immersive preference as situated player type.Convergence(2023), 13548565231206505
2023
-
[82]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...
2021
-
[83]
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:2209.14988(2022)
2022 arXiv
-
[84]
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents.arXiv preprint arXiv:2204.061251, 2 (2022), 3
2022 arXiv
-
[85]
Gideon Raeburn, Martin Welton, and Laurissa Tokarchuk. 2022. Developing a play-anywhere handheld AR storytelling app using remote data collection.Frontiers in Computer Science4 (2022), 927177
2022
-
[86]
Rosalie Rolón-Dow. 2011. Race (ing) stories: Digital storytelling as a tool for critical race scholarship.Race Ethnicity and Education14, 2 (2011), 159–173
2011
-
[87]
resfulapi. [n. d.]. RESTFul API. https://restfulapi.net/
-
[88]
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22500–22510
2023
-
[89]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695. , Vol. 1, No. 1, Article . Pu...
2022
-
[90]
Kari Salo, Diana Giova, and Tommi Mikkonen. 2016. Backend infrastructure supporting audio augmented reality and storytelling. InHuman Interface and the Management of Information: Applications and Services: 18th International Conference, HCI International 2016 Toronto, Canada, ...
2016
-
[91]
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding.Advances in Neural Inform...
2022
-
[92]
Jingyu Shi, Rahul Jain, Seunggeun Chi, Hyungjun Doh, Hyung-gun Chi, Alexander J Quinn, and Karthik Ramani
-
[93]
Narrative analysis
Emanuel A Schegloff. 1997. " Narrative analysis" thirty years later.Journal of narrative and life history7, 1-4 (1997), 97–106
1997
-
[94]
Jingyu Shi, Rahul Jain, Runlin Duan, and Karthik Ramani. 2023. Understanding Generative AI in Art: An Interview Study with Artists on G-AI from an HCI Perspective.arXiv preprint arXiv:2310.13149(2023)
2023 arXiv
-
[95]
Jae-eun Shin, Boram Yoon, Dooyoung Kim, and Woontack Woo. 2022. The Effects of Spatial Complexity on Narrative Experience in Space-Adaptive AR Storytelling.IEEE Transactions on Visualization and Computer Graphics(2022)
2022
-
[96]
Jingyu Shi, Rahul Jain, Hyungjun Doh, Ryo Suzuki, and Karthik Ramani. 2023. An HCI-centric survey and taxonomy of human-generative-AI interactions.arXiv preprint arXiv:2310.07127(2023)
2023 arXiv
-
[97]
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. 2022. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792 (2022)
2022 arXiv
-
[98]
Abbey Singh, Ramanpreet Kaur, Peter Haltner, Matthew Peachey, Mar Gonzalez-Franco, Joseph Malloch, and Derek Reilly. 2021. Story creatar: a toolkit for spatially-adaptive augmented reality storytelling. In2021 IEEE Virtual Reality and 3D User Interfaces (VR). IEEE, 713–722
2021
-
[99]
Bilal Şimşek and Bekir Direkçi. 2023. The effects of augmented reality storybooks on student’s reading comprehension. British Journal of Educational Technology54, 3 (2023), 754–772
2023
-
[100]
Murray Taylor, Mauricio Marrone, Mark Tayar, and Beate Mueller. 2018. Digital storytelling and visual metaphor in lectures: a study of student engagement.Accounting Education27, 6 (2018), 552–569
2018
-
[101]
Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. 2022. Motionclip: Exposing human motion generation to clip space. InEuropean Conference on Computer Vision. Springer, 358–374
2022
-
[102]
Helping Nemo!
Nayia Stylianidou, Angelos Sofianidis, Elpiniki Manoli, and Maria Meletiou-Mavrotheris. 2020. “Helping Nemo!”—Using augmented reality and alternate reality games in the context of universal design for learning.Education Sciences10, 4 (2020), 95
2020
-
[103]
Anastasia Tyurina. 2023. Leveraging AR-Driven Visual Storytelling to Enhance Communication of Complex Social Issues: Principles, Strategies, and Multifaceted Roles of Visuals. InSIGGRAPH Asia 2023 Educator’s Forum. 1–2
2023
-
[104]
Tom Van Laer, Stephanie Feiereisen, and Luca M Visconti. 2019. Storytelling in the digital era: A meta-analysis of relevant moderators of the narrative transportation effect.Journal of Business Research96 (2019), 135–146
2019
-
[105]
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. 2022. Human motion diffusion model.arXiv preprint arXiv:2209.14916(2022)
2022 arXiv
-
[106]
Lennart Wachowiak and Dagmar Gromann. 2023. Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1018–1032
2023
-
[107]
Josephine Walwema. 2015. The Art of Storytelling.Writing & Pedagogy7, 1 (2015)
2015
-
[108]
Fernando Vera and J Alfredo Sánchez. 2016. A model for in-situ augmented reality content creation based on storytelling and gamification. InProceedings of the 6th Mexican Conference on Human-Computer Interaction. 39–42
2016
-
[109]
Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. InICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICA...
2020
-
[110]
Hui Ye, Kin Chung Kwan, Wanchao Su, and Hongbo Fu. 2020. ARAnimator: In-situ character animation in mobile AR with user-defined motion gestures.ACM Transactions on Graphics (TOG)39, 4 (2020), 83–1
2020
-
[111]
Tianyi Wang, Xun Qian, Fengming He, Xiyun Hu, Yuanzhi Cao, and Karthik Ramani. 2021. Gesturar: An authoring system for creating freehand interactive augmented reality applications. InThe 34th Annual ACM Symposium on User Interface Software and Technology. 552–567
2021
-
[112]
Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. 2022. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.107892, 3 (2022), 5
2022 arXiv
-
[113]
Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022. Motiondif- fuse: Text-driven human motion generation with diffusion model.arXiv preprint arXiv:2208.15001(2022)
2022 arXiv
-
[114]
Rabia Meryem Yilmaz and Yuksel Goktas. 2017. Using augmented reality technology in storytelling activities: examining elementary students’ narrative skill and creativity.Virtual Reality21 (2017), 75–89. , Vol. 1, No. 1, Article . Publication date: September 2018. 24 Trovato an...
2017
-
[115]
ZhiYing Zhou, Adrian David Cheok, JiunHorng Pan, and Yu Li. 2004. An interactive 3D exploration narrative interface for storytelling. InProceedings of the 2004 conference on Interaction design and children: building a community. 155–156. A STORIES In this section, we manifest ...
2004
-
[117]
Zhenjie Zhao and Xiaojuan Ma. 2018. A compensation method of two-stage image generation for human-ai collabo- rated in-situ fashion design in augmented reality environment. In2018 IEEE International Conference on Artificial Intelligence and Virtual Reality (AIVR). IEEE, 76–83
2018
-
[2023]
InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Align your latents: High-resolution video synthesis with latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22563–22575
-
[2025]
InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems
CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–23
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.