Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

An Exploratory Study on Multi-modal Generative AI in AR Storytelling

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A study of 223 AR videos maps which AI-generated media fits which story element, and 30 storytellers confirm the mapping in practice.

desk verdict Useful exploratory mapping of modality preferences for AIGC in AR storytelling, but the headline map is tangled up with the uneven quality of the specific generative models used. read the letter →

arxiv 2505.15973 v1 pith:FXMCDF2G submitted 2025-05-21 cs.HC

classification cs.HC
keywords augmentedrealitystorytellinggenerativeAImulti-modalcontentgenerationhuman-AIinteractiondesignspaceauthoringtoolsuserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether AI-generated content can carry the multimodal load of augmented-reality storytelling, and answers with a qualified yes backed by a design space and two user studies. Analyzing 223 AR storytelling videos on YouTube, the authors identify five content modalities—text, audio, image, video, and 3D—and four atomic story elements: character, background, sentiment, and development. In studies with 30 experienced storytellers and presenters using a GenAI-powered authoring testbed, participants preferred images for characters, backgrounds, and sentiment, and video for plot development, while rating generated text highest in quality and generated video lowest. The result matters because it gives future AR authoring systems a reusable mapping of which AI modality to offer for which narrative element, and it pinpoints the open problems: aligning outputs across modalities, making prompts controllable, and deciding what to augment at all.

What carries the argument

The load-bearing object is a two-dimensional design space: five Modalities (Text, Audio, Image, Video, 3D) crossed with four atomic Elements (Character, Background, Sentiment, Development). The authors derive it by open-coding 223 YouTube AR storytelling videos, then use it to structure a testbed in which a storyteller selects a sentence, chooses a modality, and receives generator output from Stable Diffusion (image), MusicGen (audio), a motion-diffusion plus Text2Video-Zero pipeline (video), and a text-to-3D model; the AR interface then triggers the saved content by speech while the narrator gestures with hand-tracked interaction. The design space organizes both the corpus analysis and the study tasks, so the preferences reported are preferences over cells of this five-by-four grid.

What would settle it

A systematically sampled corpus of AR storytelling videos that reveals a common modality outside the five (for example haptic feedback) or a common atomic element outside the four (for example audience interaction) would falsify the taxonomy's completeness; likewise, a replication with a state-of-the-art text-to-video model that erases the current video-for-development-only preference would falsify the modality-element mapping as a stable property of AR storytelling.

Watch

Extended reading notes

Core claim

The central claim is that multi-modal AIGC is suitable for AR storytelling, provided the modality is matched to the story element. From the 223-video analysis the authors derive a design space of five modalities and four atomic elements (Character, Background, Sentiment, Development). Their two studies, each with 15 participants, show that images are strongly preferred for characters (51%), backgrounds (47%), and sentiment (47%); video is preferred for development (40%), with text close behind (30%); and 3D is a secondary choice for characters. They also find that participants rate generated text highest in quality (4.43/5) and video lowest (2.8/5), and that while co-creating with AI feels fast and enjoyable, guiding the generation to match intention is the hardest part. The authors further report that participants could mostly tell AIGC apart from human-made content but said it did not hurt the storytelling, and that cross-modal inconsistencies and literal misinterpretation of metaphors are the main blockers to wider use.

Load-bearing premise

The findings stand on the assumption that the 223 manually collected YouTube videos represent the space of AR storytelling well enough for the five-by-four taxonomy to be the right frame for the testbed and the study tasks; the authors say plainly that the corpus was not collected by systematic search.

Editorial extensions

If this is right

  • Future AR storytelling authoring tools can present images as the default augmentation for characters, backgrounds, and emotional tone, and reserve video for temporal development.
  • Because participants found text the clearest and most reliable output, tools should keep text as a fallback or complement even when visuals are preferred.
  • The low video-quality ratings imply that improvements in text-to-video generation could enlarge the role of video beyond development into other elements.
  • The repeated difficulty in prompting suggests authoring systems need example-based, iterative, or in-AR prompting beyond plain text.
  • The observed cross-modal misalignment argues for generating each story entity once and sharing its representation across modalities.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The modality-element preferences are partly confounded by current model quality: participants avoided video for small details because outputs were poor, so a replication with a stronger text-to-video model might shift the mapping.
  • The four-element taxonomy may be reusable beyond AR, as a general vocabulary for choosing generative media in slideware, virtual reality, or interactive fiction.
  • The non-systematic YouTube sampling means the taxonomy should be treated as a starting hypothesis; a systematic corpus study could add elements such as interactivity or user choice that this work deliberately excludes.
  • If cross-modal alignment is solved, the same testbed pattern could support live 'when to augment' decisions during presentations rather than pre-authored augmentation only.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents an exploratory study of multi-modal generative AI (GenAI) for AR storytelling. The authors analyze 223 YouTube videos to derive a design space with two dimensions: five modalities (text, audio, image, video, 3D) and four atomic elements (Character, Background, Sentiment, Development). They implement a testbed that generates content in these modalities from textual narratives and displays it through an AR interface, then conduct two studies with 30 experienced storytellers and presenters (split into two groups of 15). The reported findings are a modality-to-element preference mapping (image preferred for Character, Background, and Sentiment; video preferred for Development), quality ratings of generated content, Likert-scale evaluations of co-creation interactions and suitability, and qualitative insights about alignment, selective augmentation, and context awareness. The paper concludes with design considerations for future AR storytelling systems with GenAI.

Significance. If the claims are properly qualified, the paper makes a useful exploratory contribution: it provides a taxonomy of modalities and story elements, a working testbed for studying GenAI-based AR authoring, and an empirical snapshot of author preferences. Strengths include the explicit acknowledgement of corpus-selection limitations, the reporting of inter-rater agreement for video filtering (κ=0.76), and the inclusion of the full story stimuli in an appendix. However, the central preference mapping is coupled to the specific generative models used in the testbed, and the claimed 'empirical comparison' with human-generated content is not supported by the data. Because the load-bearing conclusions overreach beyond what the design can establish, the paper needs revision before its claims can be accepted as stated.

major comments (3)
  1. [§1, Contribution bullet 3; §5.5.2] The contribution list claims an 'empirical comparison of AIGC with human-generated content,' but the manuscript reports no human-generated baseline and no systematic comparison. Section 5.5.2 describes only participants' subjective impressions ('Some of the participants felt the content was as good as human-generated content') and the authors' argument that video quality reflects algorithmic limitations rather than modality suitability. This is not an empirical comparison; the claim should be removed or replaced with a condition in which participants rate matched human-generated content.
  2. [§5.1, Fig. 5; §5.5.1; Fig. 6] The central modality-to-element preference mapping is confounded with per-modality AIGC quality. Preferences in Figure 5 were expressed after previewing outputs from specific models (Stable Diffusion for images, MDM plus Text2Video-Zero for video, MusicGen for audio, DreamFusion for 3D), and Figure 6 reports that video quality was rated lowest (AVG=2.8, SD=1.13). Section 5.5.1 states that 'the quality of the video was not good enough' and that participants used video only where 'details are safe to ignore.' The rebuttal in §5.5.2, that this is 'due to limitations in algorithmic development rather than the suitability of the video as a modality,' is not testable from the current data because each modality is instantiated by a single model. The paper should either frame the conclusions as preferences under current AIGC quality or include a controlled comparison with quality matched across modalities.
  3. [§3.1.1, §3.1.3] The design-space taxonomy is load-bearing for the study: it determines the elements and modalities used in Study 1's pre-highlighted elements, the testbed capabilities, and the interpretation of participant preferences. However, the corpus was assembled through a manual, non-systematic search (stated explicitly in §3.1.1), and inter-rater reliability is reported only for the filtering step (κ=0.76), not for the open coding of modalities/elements or for the final placement of videos in the design space. The study therefore provides no independent check that the four atomic elements are complete or reliably identifiable, which limits the generality of the subsequent preference findings. Please report coding reliability for the design-space dimensions or validate the taxonomy on an independent sample.
minor comments (6)
  1. [Abstract; §4.2.1] The abstract reports 'N=30' without clarifying that this number was split into two studies of 15 participants each; please state this split explicitly for accuracy.
  2. [Fig. 6] The x-axis label contains a typo: 'Accepable' should be 'Acceptable.'
  3. [§5.5.2] The heading 'Empirical comparison to human-generated content' overstates what is presented, because no human-generated baseline is included; consider renaming the subsection to 'Participants' perceptions of AIGC versus human-generated content.'
  4. [§8 Conclusion] The first sentence reads 'Mutli-modal Gen-AI'; this should be corrected to 'Multi-modal Gen-AI.'
  5. [§5.3, §5.4] The Mann-Whitney U tests are used to support the statement that there is 'no significant difference' between the two study conditions; however, with N=15 per group, non-significance does not establish equivalence, and no effect sizes or confidence intervals are reported. Please temper these interpretations.
  6. [§6.1.1, Fig. 11] The caption states 'All images are generated by ChatGPT,' but ChatGPT alone is not an image-generation model; please specify the image model used (e.g., DALL-E through ChatGPT) for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the design space is an empirical input, the preference mapping is directly measured, and same-author citations are not load-bearing.

full rationale

The paper's central chain is empirical rather than derivational: a 223-video corpus is open-coded into a modality/element design space (Section 3), a testbed is implemented from that space (Section 4.1), and two user studies elicit preferences and quality ratings (Section 5). No result is computed from the design space by construction. The Fig. 5 modality-to-element mapping is a direct tally of participant selections in Study 1 ('For each element previously highlighted in the interface, the participants are asked to select at least one modality of augmentation'), not a quantity fitted or assumed from the taxonomy. The Fig. 6 quality ratings are likewise direct Likert responses. The paper explicitly flags its corpus as non-systematic ('We do not claim that our corpus was collected by a systematic search'), which is an honest sampling limitation rather than a circular step. The video-quality confound (video rated AVG=2.8 and the authors arguing in Section 5.5.2 that this reflects 'limitations in the algorithmic development ... rather than the suitability of the video') is an internal-validity concern, not a circular reduction, because the suitability claim is not derived from the quality rating. Same-author references (e.g., [92]) appear only in related-work or discussion context and do not carry the central claim. Hence no circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 2 invented entities

No numeric free parameters were fitted in this study; the results are descriptive statistics and non-parametric group comparisons. The load-bearing premises are conceptual: the corpus is treated as representative, the atomic-element decomposition is treated as complete, self-reports are treated as evidence of suitability, desktop webcam AR is treated as a sufficient environment, and specific GenAI models are treated as representative of their modalities. The atomic elements and modality-element mapping are author-introduced constructs with no external falsifiable validation.

assumptions (5)
  • domain assumption A non-systematic YouTube corpus of 223 videos sufficiently represents the space of AR storytelling.
    Invoked in Section 3.1.1, where the authors state they do not claim the corpus was collected by a systematic search, yet they use it to derive the design space.
  • ad hoc to paper The four atomic elements Character, Background, Sentiment, and Development decompose storytelling content in a way that covers the study stories.
    Introduced in Section 3.2.2 and then used to pre-highlight elements in Study 1, making the taxonomy part of the evaluation rather than something the evaluation tests.
  • domain assumption Self-reported Likert-scale ratings and semi-structured interviews measure the suitability of AIGC for AR storytelling.
    The study's main quantitative evidence is subjective preference data, as described in Sections 4.2.4 and 5; no behavioral or audience-based outcome is measured.
  • domain assumption Desktop webcam AR is a sufficient testbed to reveal the impact of AIGC on AR storytelling.
    Acknowledged in Section 7 as a limitation: the authors note that AR headsets offer hands-free and more immersive experiences, but they still generalize design considerations from the desktop testbed.
  • domain assumption The bundled generative models are representative of the current quality of each modality.
    The quality ratings in Figure 6 are tied to specific models (Stable Diffusion, MusicGen, MDM, Text2Video-Zero, text-to-3D), so the modality comparison conflates model quality with modality suitability, a point the authors partly acknowledge in Section 5.5.1.
invented entities (2)
  • Atomic elements taxonomy (Character, Background, Sentiment, Development)
    purpose: Used to code the 223-video corpus, define the design space, and assign augmentation targets in Study 1.
    The taxonomy is the authors' own open-coding product. It shapes the testbed and study conditions, so it is not independently validated by the same study, and no external benchmark or falsifiable prediction is offered.
  • Modality-element preference mapping
    purpose: Summarizes which of the five modalities participants preferred for each atomic element.
    The mapping is an author-generated summary of self-report data from the testbed sessions; it is not compared against an external ground truth or a human-authored content baseline.

how reviews work

0 comments
Cite this review

Pith. "Pith review of An Exploratory Study on Multi-modal Generative AI in AR Storytelling." pith.science (2026). https://pith.science/paper/FXMCDF2G

@misc{pith2026250515973,
  author       = {Pith},
  title        = {Pith review of: An Exploratory Study on Multi-modal Generative AI in AR Storytelling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FXMCDF2G}},
  note         = {Machine review of arXiv:2505.15973}
}
read the original abstract

Storytelling in AR has gained attention due to its multi-modality and interactivity. However, generating multi-modal content for AR storytelling requires expertise and efforts for high-quality conveyance of the narrator's intention. Recently, Generative-AI (GenAI) has shown promising applications in multi-modal content generation. Despite the potential benefit, current research calls for validating the effect of AI-generated content (AIGC) in AR Storytelling. Therefore, we conducted an exploratory study to investigate the utilization of GenAI. Analyzing 223 AR videos, we identified a design space for multi-modal AR Storytelling. Based on the design space, we developed a testbed facilitating multi-modal content generation and atomic elements in AR Storytelling. Through two studies with N=30 experienced storytellers and live presenters, we 1. revealed participants' preferences for modalities, 2. evaluated the interactions with AI to generate content, and 3. assessed the quality of the AIGC for AR Storytelling. We further discussed design considerations for future AR Storytelling with GenAI.

Figures

Figures reproduced from arXiv: 2505.15973 by the authors.

Figure 1
Figure 1. Our Design Space for Multi-modal AR Storytelling: Modalities and Elements. These examples illustrate [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. The testbed workflow. 1) Content Generator interface: The user employs the content generator to create AIGC for AR Storytelling, supporting five modalities for the selected sentence. 2) AR interface: The user can view the text corresponding to spoken words. Based on the transferred speech text, the user can interact with AIGC using hand. The testbed contains two separate interfaces. 1) Content Generator interface: t… view at source ↗
Figure 3
Figure 3. The testbed. (a) The Multi-Modal Content Generator interface. (a-i) The textual input section. The user can import a story text file for storytelling by entering the story title. (a-ii) The loaded story. Where the user can see the loaded story and select a sentence for augmentation by dragging it with the mouse cursor. (a-iii) The highlight (top) and save (bottom) buttons. The top button highlights the selected port… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: The distribution of the four elements that emerge in the stories we used for the study [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 6
Figure 6. Figure 6: The participants’ ratings of the overall quality of each modality. [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: Results of the Likert-scale questionnaires [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 10
Figure 10. Figure 10: Results of the Likert-scale questionnaires [PITH_FULL_IMAGE:figures/full_fig_p013_10.png]
Figure 11
Figure 11. Figure 11: An illustration of the Alignment among modalities in AR Storytelling. All images are generated by [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: An illustration of the concept of Selective Augmentation. In (a), the narrative depicts the motion [PITH_FULL_IMAGE:figures/full_fig_p016_12.png]
Figure 13
Figure 13. Figure 13: An illustration of context-awareness AIGC in AR Storytelling. Given a narrative, Gen-AI can generate [PITH_FULL_IMAGE:figures/full_fig_p017_13.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An integrated storytelling framework that generates text, scene graphs, images, and sound together reports higher expert-rated coherence than a sequential baseline, but the evaluation is small and qualitative.

Reference graph

Works this paper leans on

118 extracted references · 54 canonical work pages · cited by 1 Pith paper

  1. [1]

    Rameen Abdal, Peihao Zhu, John Femiani, Niloy Mitra, and Peter Wonka. 2022. Clip2stylegan: Unsupervised extraction of stylegan edit directions. InACM SIGGRAPH 2022 conference proceedings. 1–9

  2. [2]

    Chaitanya Ahuja and Louis-Philippe Morency. 2019. Language2pose: Natural language grounded pose forecasting. In 2019 International Conference on 3D Vision (3DV). IEEE, 719–728

  3. [3]

    Victor Nikhil Antony and Chien-Ming Huang. 2023. ID. 8: Co-Creating Visual Stories with Generative AI.arXiv preprint arXiv:2309.14228(2023)

  4. [4]

    National Storytelling Association et al. 2012. What Storytelling is. An attempt at defining the art form

  5. [5]

    Paulo Bala, Stuart James, Alessio Del Bue, and Valentina Nisi. 2022. Writing with (Digital) Scissors: Designing a Text Editing Tool for Assisted Storytelling Using Crowd-Generated Content. InInternational Conference on Interactive Digital Storytelling. Springer, 139–158

  6. [6]

    Valentin Bauer, Anna Nagele, Chris Baume, Tim Cowlishaw, Henry Cooke, Chris Pike, and Patrick GT Healey. 2019. Designing an interactive and collaborative experience in audio augmented reality. InVirtual Reality and Augmented Reality: 16th EuroVR International Conference, EuroVR 2019, Tallinn, Estonia, October 23–25, 2019, Proceedings 16. Springer, 305–311

  7. [7]

    2010.Storytelling for social justice: Connecting narrative and the arts in antiracist teaching

    Lee Anne Bell. 2010.Storytelling for social justice: Connecting narrative and the arts in antiracist teaching. Routledge

  8. [8]

    Eden Bensaid, Mauro Martino, Benjamin Hoover, and Hendrik Strobelt. 2021. Fairytailor: A multimodal generative framework for storytelling.arXiv preprint arXiv:2108.04324(2021)

Show all 118 references
  1. [9]

    Sukanya Bhattacharjee and Parag Chaudhuri. 2020. A survey on sketch based content creation: from the desktop to virtual and augmented reality. InComputer Graphics Forum, Vol. 39. Wiley Online Library, 757–780

  2. [10]

    Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman, Aidan Clark, Erich Elsen, Norman Casagrande, Luis C Cobo, and Karen Simonyan. 2019. High fidelity speech synthesis with adversarial networks.arXiv preprint arXiv:1909.11646 (2019)

  3. [11]

    Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dockhorn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis

  4. [12]

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258(2021)

  5. [13]

    I address race because race addresses me

    Anjuli Joshi Brekke, Ralina Joseph, and Naheed Gina Aaftaab. 2021. “I address race because race addresses me”: women of color show receipts through digital storytelling.Review of Communication21, 1 (2021), 44–57

  6. [14]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural , Vol. 1, No. 1, Article . Publication date: Septembe...

  7. [15]

    Licia Calvi. 2020. What do we know about AR Storytelling?. InProceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good. 278–280

  8. [16]

    Jorge Camba, Manuel Contero, and Gustavo Salvador-Herranz. 2014. Desktop vs. mobile: A comparative study of augmented reality systems for engineering visualizations in education. In2014 IEEE Frontiers in Education Conference (FIE) Proceedings. IEEE, 1–8

  9. [17]

    Weifeng Chen, Jie Wu, Pan Xie, Hefeng Wu, Jiashi Li, Xin Xia, Xuefeng Xiao, and Liang Lin. 2023. Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models.arXiv preprint arXiv:2305.13840(2023)

  10. [18]

    John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: visual sketching of story generation with pretrained language models. InCHI Conference on Human Factors in Computing Systems Extended Abstracts. 1–4

  11. [19]

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2023. Simple and Controllable Music Generation.arXiv preprint arXiv:2306.05284(2023)

  12. [20]

    Delneshin Danaei, Hamid R Jamali, Yazdan Mansourian, and Hassan Rastegarpour. 2020. Comparing reading comprehension between children reading augmented reality and print storybooks.Computers & Education153 (2020), 103900

  13. [21]

    Dimitrios Darzentas, Martin Flintham, and Steve Benford. 2018. Object-focused mixed reality storytelling: technology- driven content creation and dissemination for engaging user experiences. InProceedings of the 22nd Pan-Hellenic Conference on Informatics. 278–281

  14. [22]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding.arXiv preprint arXiv:1810.04805(2018)

  15. [23]

    Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang. 2022. Cogview2: Faster and better text-to-image generation via hierarchical transformers.Advances in Neural Information Processing Systems35 (2022), 16890–16902

  16. [24]

    Runlin Duan, Shao-Kang Hsia, Yuzhao Chen, Yichen Hu, Ming Yin, and Karthik Ramani. 2025. Investigating Creativity in Humans and Generative AI Through Circles Exercises.arXiv preprint arXiv:2502.07292(2025)

  17. [25]

    Runlin Duan, Nachiketh Karthik, Jingyu Shi, Rahul Jain, Maria C Yang, and Karthik Ramani. 2024. ConceptVis: Generating and Exploring Design Concepts for Early-Stage Ideation Using Large Language Model. InInternational Design Engineering Technical Conferences and Computers and ...

  18. [26]

    Runlin Duan, Chenfei Zhu, Yuzhao Chen, Yichen Hu, Jingyu Shi, and Karthik Ramani. 2025. DesignFromX: Empower- ing Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products.arXiv preprint arXiv:2505.11666(2025)

  19. [27]

    FastAPI. [n. d.]. FastAPI. https://fastapi.tiangolo.com

  20. [28]

    2005.Storytelling

    Klaus Fog, Christian Budtz, and Baris Yakaboylu. 2005.Storytelling. Springer

  21. [29]

    Rinon Gal, Or Patashnik, Haggai Maron, Amit H Bermano, Gal Chechik, and Daniel Cohen-Or. 2022. StyleGAN-NADA: CLIP-guided domain adaptation of image generators.ACM Transactions on Graphics (TOG)41, 4 (2022), 1–13

  22. [30]

    Ze Gao, Anqi Wang, Pan Hui, and Tristan Braud. 2022. Bridging curatorial intent and visiting experience: Using ar guidance as a storytelling tool. InProceedings of the 18th ACM SIGGRAPH International Conference on Virtual-Reality Continuum and its Applications in Industry. 1–10

  23. [31]

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial nets.Advances in neural information processing systems27 (2014)

  24. [32]

    Ariel Han and Zhenyao Cai. 2023. Design implications of generative AI systems for visual storytelling for young learners. InProceedings of the 22nd Annual ACM Interaction Design and Children Conference. 470–474

  25. [33]

    Linda Hirsch, Robin Welsch, Beat Rossmy, and Andreas Butz. 2022. Embedded AR Storytelling Supports Active Indexing at Historical Places. InSixteenth International Conference on Tangible, Embedded, and Embodied Interaction. 1–12

  26. [34]

    Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. 2022. Imagen video: High definition video generation with diffusion models.arXiv preprint arXiv:2210.02303(2022)

  27. [35]

    Lissa Holloway-Attaway and Lars Vipsjö. 2020. Using augmented reality, gaming technologies, and transmedial storytelling to develop and co-design local cultural heritage experiences.Visual Computing for Cultural Heritage (2020), 177–204

  28. [36]

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, Adriane Boyd, et al. 2020. spaCy: Industrial-strength natural language processing in python. (2020)

  29. [37]

    Xiyun Hu, Dizhi Ma, Fengming He, Zhengzhe Zhu, Shao-Kang Hsia, Chenfei Zhu, Ziyi Liu, and Karthik Ramani. 2025. GesPrompt: Leveraging Co-Speech Gestures to Augment LLM-Based Interaction in Virtual Reality.arXiv preprint arXiv:2505.05441(2025). , Vol. 1, No. 1, Article . Public...

  30. [38]

    Yongquan Hu, Mingyue Yuan, Kaiqi Xian, Don Samitha Elvitigala, and Aaron Quigley. 2023. Exploring the Design Space of Employing AI-Generated Content for Augmented Reality Display.arXiv preprint arXiv:2303.16593(2023)

  31. [39]

    Reinis Indans, Eva Hauthal, and Dirk Burghardt. 2019. Towards an audio-locative mobile application for immersive storytelling.KN-Journal of Cartography and Geographic Information69 (2019), 41–50

  32. [40]

    Myunggeun Ji and Junchul Chun. 2020. A Sketch-based 3D Object Retrieval Approach for Augmented Reality Models Using Deep Learning.Journal of Korean Society for Internet Information21, 1 (2020)

  33. [41]

    BRENNAN JONES, YAN XU, MARY ANNE HOOD, MOHAMMAD SHAHIDUL KADER, and HAMID EGHBALZADEH. [n. d.]. Using Generative AI to Produce Situated Action Recommendations in Augmented Reality for High-Level Goals. ([n. d.])

  34. [42]

    Kwanghee Jung, Vinh T Nguyen, and Jaehoon Lee. 2021. Blocklyxr: An interactive extended reality toolkit for digital storytelling.Applied Sciences11, 3 (2021), 1073

  35. [43]

    Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu. 2018. Efficient neural audio synthesis. InInternational Conference on Machine Learning. PMLR, 2410–2419

  36. [44]

    Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2021. Alias-free generative adversarial networks.Advances in Neural Information Processing Systems34 (2021), 852–863

  37. [45]

    Tero Karras, Samuli Laine, and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4401–4410

  38. [46]

    Sarah Ketchell, Winyu Chinthammit, and Ulrich Engelke. 2019. Situated storytelling with SLAM enabled augmented reality. InProceedings of the 17th International Conference on Virtual-Reality Continuum and Its Applications in Industry. 1–9

  39. [47]

    Levon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel, Zhangyang Wang, Shant Navasardyan, and Humphrey Shi. 2023. Text2video-zero: Text-to-image diffusion models are zero-shot video genera- tors.arXiv preprint arXiv:2303.13439(2023)

  40. [48]

    Doyeon Kim, Donggyu Joo, and Junmo Kim. 2020. Tivgan: Text to image to video generation with step-by-step evolutionary generator.IEEE Access8 (2020), 153113–153122

  41. [49]

    Jihoon Kim, Jiseob Kim, and Sungjoon Choi. 2023. Flame: Free-form language-based motion synthesis & editing. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37. 8255–8263

  42. [50]

    Kiyoung Kim, Noh-young Park, and Woontack Woo. 2014. Vision-based all-in-one solution for augmented reality and its storytelling applications.The Visual Computer30 (2014), 417–429

  43. [51]

    Diederik P Kingma and Max Welling. 2013. Auto-encoding variational bayes.arXiv preprint arXiv:1312.6114(2013)

  44. [52]

    Zhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao, and Bryan Catanzaro. 2020. Diffwave: A versatile diffusion model for audio synthesis.arXiv preprint arXiv:2009.09761(2020)

  45. [53]

    Kundan Kumar, Rithesh Kumar, Thibault De Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre De Brebis- son, Yoshua Bengio, and Aaron C Courville. 2019. Melgan: Generative adversarial networks for conditional waveform synthesis.Advances in neural information process...

  46. [54]

    Tomas Lawton, Francisco J Ibarrola, Dan Ventura, and Kazjon Grace. 2023. Drawing with Reframer: Emergence and Control in Co-Creative AI. InProceedings of the 28th International Conference on Intelligent User Interfaces. 264–277

  47. [55]

    Kyungjun Lee, Hong Li, Muhammad Rizky Wellyanto, Yu Jiang Tham, Andrés Monroy-Hernández, Fannie Liu, Brian A Smith, and Rajan Vaish. 2023. Exploring Immersive Interpersonal Communication via AR.Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (2023), 1–25

  48. [56]

    Jian Liao, Adnan Karim, Shivesh Singh Jadon, Rubaiat Habib Kazi, and Ryo Suzuki. 2022. RealityTalk: Real-Time Speech-Driven Augmented Presentation for AR Live Storytelling. InProceedings of the 35th Annual ACM Symposium on User Interface Software and Technology. 1–12

  49. [57]

    Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang

    Xingyu" Bruce" Liu, Vladimir Kirilyuk, Xiuxiu Yuan, Alex Olwal, Peggy Chi, Xiang" Anthony" Chen, and Ruofei Du

  50. [58]

    Ziyi Liu, Zhengzhe Zhu, Lijun Zhu, Enze Jiang, Xiyun Hu, Kylie A Peppler, and Karthik Ramani. 2024. Classmeta: Designing interactive virtual classmate to promote VR classroom participation. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–17

  51. [59]

    InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Visual Captions: Augmenting Verbal Communication With On-the-Fly Visuals. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20

  52. [60]

    Tiago Madeira, Bernardo Marques, Pedro Neves, Paulo Dias, and Beatriz Sousa Santos. 2022. Comparing desktop vs. Mobile interaction for the creation of pervasive augmented reality experiences.Journal of Imaging8, 3 (2022), 79

  53. [61]

    Yue Ma, Yingqing He, Xiaodong Cun, Xintao Wang, Ying Shan, Xiu Li, and Qifeng Chen. 2023. Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos.arXiv preprint arXiv:2304.01186(2023)

  54. [62]

    Dimitrios Markouzis and Georgios Fessakis. 2015. Interactive storytelling and mobile augmented reality applications for learning and entertainment—a rapid prototyping perspective. In2015 International Conference on Interactive Mobile , Vol. 1, No. 1, Article . Publication date...

  55. [63]

    Jenny Mandelbaum. 2012. Storytelling in conversation.The handbook of conversation analysis(2012), 492–507

  56. [64]

    Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. 2016. SampleRNN: An unconditional end-to-end neural audio generation model.arXiv preprint arXiv:1612.07837(2016)

  57. [65]

    MediaPipe. [n. d.]. MediaPipe. https://mediapipe.dev/

  58. [66]

    Ron Mokady, Omer Tov, Michal Yarom, Oran Lang, Inbar Mosseri, Tali Dekel, Daniel Cohen-Or, and Michal Irani. 2022. Self-distilled stylegan: Towards generation from internet photos. InACM SIGGRAPH 2022 Conference Proceedings. 1–9

  59. [67]

    2019.Digital Storytelling 4e: A creator’s guide to interactive entertainment

    Carolyn Handler Miller. 2019.Digital Storytelling 4e: A creator’s guide to interactive entertainment. CRC Press

  60. [68]

    Haomiao Ni, Changhao Shi, Kai Li, Sharon X Huang, and Martin Renqiang Min. 2023. Conditional Image-to-Video Generation with Latent Flow Diffusion Models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 18444–18455

  61. [69]

    mozilla. [n. d.]. Web Speech API. https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_API

  62. [70]

    Jennifer O’Meara and Kata Szita. 2021. AR cinema: Visual storytelling and embodied experiences with augmented reality filters and backgrounds.PRESENCE: Virtual and Augmented Reality30 (2021), 99–123

  63. [71]

    Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, and Mark Chen. 2021. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv preprint arXiv:2112.10741(2021)

  64. [72]

    Augusto Palombini. 2017. Storytelling and telling history. Towards a grammar of narratives for Cultural Heritage dissemination in the Digital Era.Journal of cultural heritage24 (2017), 134–139

  65. [73]

    Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbren- ner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio.arXiv preprint arXiv:1609.03499(2016)

  66. [74]

    Or Patashnik, Zongze Wu, Eli Shechtman, Daniel Cohen-Or, and Dani Lischinski. 2021. Styleclip: Text-driven manipulation of stylegan imagery. InProceedings of the IEEE/CVF International Conference on Computer Vision. 2085–2094

  67. [75]

    Seung-Bo Park, Jason J Jung, and EunSoon You. 2015. Storytelling of collaborative learning system on augmented reality.New Trends in Computational Collective Intelligence(2015), 139–147

  68. [76]

    Kainan Peng, Wei Ping, Zhao Song, and Kexin Zhao. 2020. Non-autoregressive neural text-to-speech. InInternational conference on machine learning. PMLR, 7586–7598

  69. [77]

    John V Pavlik and Frank Bridges. 2013. The emergence of augmented reality (AR) as a storytelling medium in journalism.Journalism & Communication Monographs15, 1 (2013), 4–59

  70. [78]

    Savvas Petridis, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V Nickerson, and Lydia B Chilton. 2023. Anglekindling: Supporting journalistic angle ideation with large language models. InProceedings of the 2023 CHI Conference on ...

  71. [79]

    Eric E Peterson and Kristin M Langellier. 2006. Communication as storytelling.Communication as... Perspectives on Theory(2006), 123–131

  72. [80]

    Matthias Plappert, Christian Mandery, and Tamim Asfour. 2016. The KIT motion-language dataset.Big data4, 4 (2016), 236–252

  73. [81]

    Judith Pintar. 2023. Invisible, aesthetic, and enrolled listeners across storytelling modalities: Immersive preference as situated player type.Convergence(2023), 13548565231206505

  74. [82]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learnin...

  75. [83]

    Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:2209.14988(2022)

  76. [84]

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents.arXiv preprint arXiv:2204.061251, 2 (2022), 3

  77. [85]

    Gideon Raeburn, Martin Welton, and Laurissa Tokarchuk. 2022. Developing a play-anywhere handheld AR storytelling app using remote data collection.Frontiers in Computer Science4 (2022), 927177

  78. [86]

    Rosalie Rolón-Dow. 2011. Race (ing) stories: Digital storytelling as a tool for critical race scholarship.Race Ethnicity and Education14, 2 (2011), 159–173

  79. [87]

    resfulapi. [n. d.]. RESTFul API. https://restfulapi.net/

  80. [88]

    Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman. 2023. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22500–22510

  81. [89]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684–10695. , Vol. 1, No. 1, Article . Pu...

  82. [90]

    Kari Salo, Diana Giova, and Tommi Mikkonen. 2016. Backend infrastructure supporting audio augmented reality and storytelling. InHuman Interface and the Management of Information: Applications and Services: 18th International Conference, HCI International 2016 Toronto, Canada, ...

  83. [91]

    Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. 2022. Photorealistic text-to-image diffusion models with deep language understanding.Advances in Neural Inform...

  84. [92]

    Jingyu Shi, Rahul Jain, Seunggeun Chi, Hyungjun Doh, Hyung-gun Chi, Alexander J Quinn, and Karthik Ramani

  85. [93]

    Narrative analysis

    Emanuel A Schegloff. 1997. " Narrative analysis" thirty years later.Journal of narrative and life history7, 1-4 (1997), 97–106

  86. [94]

    Jingyu Shi, Rahul Jain, Runlin Duan, and Karthik Ramani. 2023. Understanding Generative AI in Art: An Interview Study with Artists on G-AI from an HCI Perspective.arXiv preprint arXiv:2310.13149(2023)

  87. [95]

    Jae-eun Shin, Boram Yoon, Dooyoung Kim, and Woontack Woo. 2022. The Effects of Spatial Complexity on Narrative Experience in Space-Adaptive AR Storytelling.IEEE Transactions on Visualization and Computer Graphics(2022)

  88. [96]

    Jingyu Shi, Rahul Jain, Hyungjun Doh, Ryo Suzuki, and Karthik Ramani. 2023. An HCI-centric survey and taxonomy of human-generative-AI interactions.arXiv preprint arXiv:2310.07127(2023)

  89. [97]

    Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. 2022. Make-a-video: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792 (2022)

  90. [98]

    Abbey Singh, Ramanpreet Kaur, Peter Haltner, Matthew Peachey, Mar Gonzalez-Franco, Joseph Malloch, and Derek Reilly. 2021. Story creatar: a toolkit for spatially-adaptive augmented reality storytelling. In2021 IEEE Virtual Reality and 3D User Interfaces (VR). IEEE, 713–722

  91. [99]

    Bilal Şimşek and Bekir Direkçi. 2023. The effects of augmented reality storybooks on student’s reading comprehension. British Journal of Educational Technology54, 3 (2023), 754–772

  92. [100]

    Murray Taylor, Mauricio Marrone, Mark Tayar, and Beate Mueller. 2018. Digital storytelling and visual metaphor in lectures: a study of student engagement.Accounting Education27, 6 (2018), 552–569

  93. [101]

    Guy Tevet, Brian Gordon, Amir Hertz, Amit H Bermano, and Daniel Cohen-Or. 2022. Motionclip: Exposing human motion generation to clip space. InEuropean Conference on Computer Vision. Springer, 358–374

  94. [102]

    Helping Nemo!

    Nayia Stylianidou, Angelos Sofianidis, Elpiniki Manoli, and Maria Meletiou-Mavrotheris. 2020. “Helping Nemo!”—Using augmented reality and alternate reality games in the context of universal design for learning.Education Sciences10, 4 (2020), 95

  95. [103]

    Anastasia Tyurina. 2023. Leveraging AR-Driven Visual Storytelling to Enhance Communication of Complex Social Issues: Principles, Strategies, and Multifaceted Roles of Visuals. InSIGGRAPH Asia 2023 Educator’s Forum. 1–2

  96. [104]

    Tom Van Laer, Stephanie Feiereisen, and Luca M Visconti. 2019. Storytelling in the digital era: A meta-analysis of relevant moderators of the narrative transportation effect.Journal of Business Research96 (2019), 135–146

  97. [105]

    Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. 2022. Human motion diffusion model.arXiv preprint arXiv:2209.14916(2022)

  98. [106]

    Lennart Wachowiak and Dagmar Gromann. 2023. Does GPT-3 Grasp Metaphors? Identifying Metaphor Mappings with Generative Language Models. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1018–1032

  99. [107]

    Josephine Walwema. 2015. The Art of Storytelling.Writing & Pedagogy7, 1 (2015)

  100. [108]

    Fernando Vera and J Alfredo Sánchez. 2016. A model for in-situ augmented reality content creation based on storytelling and gamification. InProceedings of the 6th Mexican Conference on Human-Computer Interaction. 39–42

  101. [109]

    Ryuichi Yamamoto, Eunwoo Song, and Jae-Min Kim. 2020. Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram. InICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICA...

  102. [110]

    Hui Ye, Kin Chung Kwan, Wanchao Su, and Hongbo Fu. 2020. ARAnimator: In-situ character animation in mobile AR with user-defined motion gestures.ACM Transactions on Graphics (TOG)39, 4 (2020), 83–1

  103. [111]

    Tianyi Wang, Xun Qian, Fengming He, Xiyun Hu, Yuanzhi Cao, and Karthik Ramani. 2021. Gesturar: An authoring system for creating freehand interactive augmented reality applications. InThe 34th Annual ACM Symposium on User Interface Software and Technology. 552–567

  104. [112]

    Jiahui Yu, Yuanzhong Xu, Jing Yu Koh, Thang Luong, Gunjan Baid, Zirui Wang, Vijay Vasudevan, Alexander Ku, Yinfei Yang, Burcu Karagol Ayan, et al. 2022. Scaling autoregressive models for content-rich text-to-image generation. arXiv preprint arXiv:2206.107892, 3 (2022), 5

  105. [113]

    Mingyuan Zhang, Zhongang Cai, Liang Pan, Fangzhou Hong, Xinying Guo, Lei Yang, and Ziwei Liu. 2022. Motiondif- fuse: Text-driven human motion generation with diffusion model.arXiv preprint arXiv:2208.15001(2022)

  106. [114]

    Rabia Meryem Yilmaz and Yuksel Goktas. 2017. Using augmented reality technology in storytelling activities: examining elementary students’ narrative skill and creativity.Virtual Reality21 (2017), 75–89. , Vol. 1, No. 1, Article . Publication date: September 2018. 24 Trovato an...

  107. [115]

    ZhiYing Zhou, Adrian David Cheok, JiunHorng Pan, and Yu Li. 2004. An interactive 3D exploration narrative interface for storytelling. InProceedings of the 2004 conference on Interaction design and children: building a community. 155–156. A STORIES In this section, we manifest ...

  108. [117]

    Zhenjie Zhao and Xiaojuan Ma. 2018. A compensation method of two-stage image generation for human-ai collabo- rated in-situ fashion design in augmented reality environment. In2018 IEEE International Conference on Artificial Intelligence and Virtual Reality (AIVR). IEEE, 76–83

  109. [2023]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Align your latents: High-resolution video synthesis with latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 22563–22575

  110. [2025]

    InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems

    CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–23

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.