REVIEW 5 minor 27 references
Everyday AR through AI-in-the-Loop
T0 review · 0 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Everyday AR needs AI in the loop to become feasible
desk verdict A well-organized workshop call for participation, not a research paper; fine as a vision statement but nothing to referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the concept of everyday AR, defined as AR that is always available and integrated into users' daily environments. The mechanism that carries the argument is the AI-in-the-loop model: an AR system in which an AI component continuously senses context, including room geometry, object affordances, user activity, and user state, and uses that understanding to adapt content and interactions without explicit user commands. This mechanism is what the paper says will move AR beyond monolithic applications, because it makes both scene understanding and content generation dynamic. The paper identifies six research thrusts that instantiate the mechanism: adaptive and context-aware AR, LLM-powered always-on assistants, AI-assisted task guidance, generative on-demand content creation, AI-driven accessible design, and real-world-oriented AI agents.
What would settle it
A controlled field deployment would settle the claim: give one group an always-on, context-adapting AR assistant and another group a static AR interface for the same everyday tasks; if the adaptive system shows no improvement in completion time, error rate, or perceived load, or if participants disable it because it misreads their context, the premise that AI-in-the-loop is the route to everyday AR is not supported.
Extended reading notes
Core claim
The paper's central claim is that the combination of recent AI advances and maturing AR hardware enables a new class of experience the authors call everyday AR: always-available, seamlessly integrated digital content that could replace or augment smartphones and desktop computers for many interactions. To get there, the paper asserts, AR must adopt an AI-in-the-loop model, in which digital interactions and content continuously anticipate and adapt to users' changing needs and context. In this model, scene understanding and content generation both become dynamic: large language models support natural, always-on assistant interactions, generative models create on-demand content in real time, and context-aware systems adapt interfaces based on room geometry, object affordances, and user activity. The paper's contribution is a shared research agenda and a call for the community to define the requirements and limitations of this vision rather than a completed system or evaluation.
Load-bearing premise
The whole vision rests on the assumption that current AI, especially large language models and generative models, can be integrated into AR systems well enough to understand and respond to unpredictable real-world contexts in real time, and that users will accept the always-on sensing this requires.
Editorial extensions
If this is right
- AR would shift from niche applications such as productivity and maintenance to a general-purpose platform for everyday computing, potentially displacing smartphone and desktop interaction for many tasks.
- Always-on LLM-based assistants embedded in AR could support users implicitly, understanding context from gaze, gestures, and activities rather than requiring typed questions.
- Generative AI could produce interactive AR content on demand, including 3D objects, scenes, and code-generated interactions, removing the need for manual content authoring.
- Context-aware adaptation informed by room geometry, object affordances, and user activity would keep augmented interfaces from overwhelming users with irrelevant content.
- AI-driven input modalities such as voice, gaze, and facial expression recognition could open AR authoring and experiences to people with motor impairments, who are currently excluded by keyboard-mouse and controller interactions.
Reading between the lines
- [Editorial inference] If the AI-in-the-loop hypothesis is correct, evaluation of AR systems will need to move from lab-based rendering-quality measures to longitudinal, in-situ measures of context recognition and task benefit, because the central claim is about everyday use.
- [Editorial inference] The vision shifts control from the user to the system: users would delegate some interface decisions to AI, which makes trust, privacy, and the ability to override the AI central design problems rather than peripheral ones.
- [Editorial inference] A concrete testable consequence is that adaptive AR interfaces using live context models should outperform static, manually configured interfaces on real-world everyday tasks; a field deployment comparing the two would provide evidence for or against the paper's premise.
- [Editorial inference] The paper's emphasis on accessible design suggests AI-driven adaptation could make AR more equitable, but that holds only if the underlying context models do not carry the same biases found in their training data.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a workshop proposal for CHI EA 2025, authored by Suzuki, Gonzalez-Franco, Sra, and Lindlbauer. It argues that recent advances in AR hardware and AI/ML make 'everyday AR'—always-available, seamlessly integrated augmented reality—increasingly feasible, and that achieving it requires an AI-in-the-loop approach in which digital content and interactions continuously anticipate and adapt to user context. The paper identifies six topics of interest (adaptive and context-aware AR, LLM-powered always-on assistants, AI-assisted task guidance, generative AR content creation, accessible AR design, and real-world-oriented AI agents), describes the planned one-day workshop structure and activities, provides a call for participation, and includes organizer biographies and references. The manuscript contains no experiments, derivations, or datasets; its central feasibility claims are explicitly framed as beliefs ('we believe').
Significance. As a workshop proposal, this paper serves a community-building and agenda-setting purpose rather than making a technical contribution. If the vision of everyday AR were realized, it could indeed constitute a major shift in human-computer interaction, and the paper usefully organizes current research threads such as context-aware AR, generative content creation, and always-on AI assistance. It also explicitly acknowledges open challenges including accessibility, privacy, and explainability, and it draws on the organizers' prior workshop experience at UIST 2023 [24], which lends credibility to the organizational plan. However, the manuscript makes no testable claims and provides no evidence for feasibility; its value depends on the workshop's ability to generate and refine community discourse. For a reader expecting a standard research paper, the contribution is limited; for a workshop proposal, it is appropriate. The stress-test concern about whether AI can handle real-world unpredictability is genuine, but it is not load-bearing for this document because the paper's stated purpose is to invite discussion of that very question.
minor comments (5)
- [Section 1, first paragraph] The phrase 'make it possible to results in a paradigm shift' is ungrammatical; rewrite it as 'make possible a paradigm shift' or 'result in a paradigm shift.'
- [Section 1, first paragraph] The phrase 'users every-changing needs' contains a typo; it should be 'users' ever-changing needs.'
- [Section 3.2, Introductions and Lightning Talks] The text reads 'participant's lighting talks'; this should be 'participants' lightning talks.'
- [Figure 1 caption] The caption says 'Around 50 participants,' but Section 3.2 reports 40 participants for the same UIST 2023 workshop; please reconcile these numbers.
- [Section 3.2, Theme Organization and Discussion] The phrase 'assigned a ‘table. ’' contains stray quotation marks; it should simply read 'assigned a table.'
Circularity Check
No significant circularity: the paper is a workshop call for participation with no derivation chain or predicted result to reduce to its inputs.
full rationale
This manuscript is a CHI 2025 workshop proposal. It contains no formal derivation, no fitted parameters, no predictive model, and no empirical claim that could be reduced to its inputs by construction. The central statement, 'we believe that everyday AR is increasingly feasible' and that 'we need to adopt an AI-in-the-loop approach,' is an explicitly aspirational vision rather than a derived result. The paper's own text frames the document as a call for participation: 'make a call for participation in which we ask participants to share how they imagine the future of AR as we move to a world where AR is to AI what screens have been to computers.' The only self-citation is reference [24], the authors' prior UIST 2023 workshop, which is used as contextual background ('an evolution of our previous XR and AI Workshop at ACM UIST 2023') and as evidence that a similar workshop format was successfully run ('similar to our successful workshop at ACM UIST 2023'). That citation is not load-bearing for any technical claim, prediction, or uniqueness argument; it merely supports the workshop logistics claim that the format has been used before. All other citations point to external prior work and are used as examples of existing research directions, not as premises that force a conclusion. There is no self-definitional step, no fitted input renamed as a prediction, no uniqueness theorem imported from the authors, and no ansatz smuggled in via citation. The paper is self-contained relative to its purpose: it proposes discussion topics and an event structure, and it explicitly leaves open questions such as privacy, explainability, and sustainability. Because the manuscript makes no verifiable technical claim, no circular reasoning can attach to its central content. Score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Everyday AR is becoming increasingly feasible due to advances in AR hardware (smaller form factors, battery life, connectivity) and AI.
- domain assumption An AI-in-the-loop approach is necessary to make everyday AR usable.
- domain assumption LLMs and generative AI can effectively drive adaptive, context-aware AR experiences.
Cite this review
Pith. "Pith review of Everyday AR through AI-in-the-Loop." pith.science (2026). https://pith.science/paper/3U7WR4CO
@misc{pith2026241212681,
author = {Pith},
title = {Pith review of: Everyday AR through AI-in-the-Loop},
year = {2026},
howpublished = {\url{https://pith.science/paper/3U7WR4CO}},
note = {Machine review of arXiv:2412.12681}
}
read the original abstract
This workshop brings together experts and practitioners from augmented reality (AR) and artificial intelligence (AI) to shape the future of AI-in-the-loop everyday AR experiences. With recent advancements in both AR hardware and AI capabilities, we envision that everyday AR -- always-available and seamlessly integrated into users' daily environments -- is becoming increasingly feasible. This workshop will explore how AI can drive such everyday AR experiences. We discuss a range of topics, including adaptive and context-aware AR, generative AR content creation, always-on AI assistants, AI-driven accessible design, and real-world-oriented AI agents. Our goal is to identify the opportunities and challenges in AI-enabled AR, focusing on creating novel AR experiences that seamlessly blend the digital and physical worlds. Through the workshop, we aim to foster collaboration, inspire future research, and build a community to advance the research field of AI-enhanced AR.
Figures
Reference graph
Works this paper leans on
-
[24]
Ryo Suzuki, Mar Gonzalez-Franco, Misha Sra, and David Lindlbauer. 2023. XR and AI: AI-Enabled Virtual, Augmented, and Mixed Reality. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23 Adjunct). Association for Computing Machinery, New York, NY, USA, Article 108, 3 pages. htt...
arXiv 2023
-
[1]
Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, and Anthony Tang. 2024. How People Prompt Generative AI to Create Interactive VR Scenes. InProceedings of the 2024 ACM Designing Interactive Systems Conference. 2319–2340
work page 2024
-
[2]
Karan Ahuja, Eyal Ofek, Mar Gonzalez-Franco, Christian Holz, and Andrew D Wilson. 2021. Coolmoves: User motion accentuation in virtual reality.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies5, 2 (2021), 1–23
work page 2021
-
[3]
Riku Arakawa, Hiromu Yakura, and Mayank Goel. 2024. PrISM-Observer: In- tervention Agent to Help Users Perform Everyday Procedures Sensed using a Smartwatch. arXiv preprint arXiv:2407.16785 (2024)
work page Pith review arXiv 2024
-
[4]
Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J Gonzalez, Li-Te Cheng, and Mar Gonzalez-Franco. 2024. EmBARDiment: an Embodied AI Agent for Productivity in XR. arXiv preprint arXiv:2408.08158 (2024)
arXiv 2024
-
[5]
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, ...
arXiv 2020
-
[6]
Sonia Castelo, Joao Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Iran Roman, Roque Lopez, Ethan Brewer, Chen Zhao, et al. 2023. Argus: Visual- ization of ai-assisted task guidance in ar. IEEE Transactions on Visualization and Computer Graphics (2023)
2023
-
[7]
Neil Chulpongsatorn, Mille Skovhus Lunding, Nishan Soni, and Ryo Suzuki. 2023. Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math Textbooks. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–16
work page 2023
Show all 27 references
-
[8]
Fernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski- Fahey, Judith Amores Fernandez, and Jaron Lanier. 2024. Llmr: Real-time prompt- ing of interactive worlds using large language models. In Proceedings of the CHI Conference on Human Factors in Computing Sy...
2024
-
[9]
Mustafa Doga Dogan, Eric J Gonzalez, Andrea Colaco, Karan Ahuja, Ruofei Du, Johnny Lee, Mar Gonzalez-Franco, and David Kim. 2024. Augmented Object Intelligence: Making the Analog World Interactable with XR-Objects. arXiv preprint arXiv:2404.13274 (2024)
2024 arXiv
-
[10]
Ran Gal, Lior Shapira, Eyal Ofek, and Pushmeet Kohli. 2014. FLARE: Fast layout for augmented reality applications. In 2014 IEEE international symposium on mixed and augmented reality (ISMAR) . IEEE, 207–212
2014
-
[11]
Mar Gonzalez-Franco, Julio Cermeron, Katie Li, Rodrigo Pizarro, Jacob Thorn, Windo Hutabarat, Ashutosh Tiwari, and Pablo Bermell-Garcia. 2016. Immersive augmented reality training for complex manufacturing scenarios. arXiv preprint arXiv:1602.01944 (2016)
2016 arXiv
-
[12]
Mar Gonzalez-Franco and Andrea Colaco. 2024. Guidelines for Productivity in Virtual Reality. Interactions 31, 3 (2024), 46–53
2024
-
[13]
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Gener- ative Adversarial Nets. Advances in Neural Information Processing Sys- tems 27 (2014). https://proceedings.neurips.cc/paper_files/pape...
2014
-
[14]
Jens Grubert, Tobias Langlotz, Stefanie Zollmann, and Holger Regenbrecht. 2017. Towards Pervasive Augmented Reality: Context-Awareness in Augmented Reality. IEEE Transactions on Visualization and Computer Graphics 23, 6 (2017), 1706–1724. https://doi.org/10.1109/TVCG.2016.2543720
2017
-
[15]
Aditya Gunturu, Shivesh Jadon, Nandi Zhang, Morteza Faraji, Jarin Thundathil, Tafreed Ahmad, Wesley Willett, and Ryo Suzuki. 2024. RealitySummary: Explor- ing On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models. arXiv preprint arXiv:24...
2024
-
[16]
Violet Yinuo Han, Hyunsung Cho, Kiyosu Maeda, Alexandra Ion, and David Lindlbauer. 2023. BlendMR: A Computational Method to Create Ambient Mixed Reality Interfaces. Proceedings of the ACM on Human-Computer Interaction 7, ISS (2023), 217–241
2023
-
[17]
Fengming He, Xiyun Hu, Jingyu Shi, Xun Qian, Tianyi Wang, and Karthik Ramani
-
[18]
Teresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz, Pascal Knierim, and Kasper Hornbæk. 2023. When XR and AI Meet-A Scoping Review on Extended Reality and Artificial Intelligence. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–45
2023
-
[19]
Rahul Jain, Jingyu Shi, Runlin Duan, Zhengzhe Zhu, Xun Qian, and Karthik Ramani. 2023. Ubi-TOUCH: Ubiquitous Tangible Object Utilization through Consistent Hand-object interaction in Augmented Reality. In Proceedings of the 36th Annual ACM Symposium on User Interface Software ...
2023
-
[20]
Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. arXiv preprint arXiv:2407.13515 (2024)
2024
-
[21]
Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S Rodriguez, and Jon E Froehlich. 2024. GazePointAR: A Context-Aware Multimodal Voice Assis- tant for Pronoun Disambiguation in Wearable Augmented Reality. InProceedings of the CHI Conference on Human Factors in Compu...
2024
-
[22]
David Lindlbauer. 2022. The future of mixed reality is adaptive.XRDS: Crossroads, The ACM Magazine for Students 29, 1 (2022), 26–31
2022
-
[23]
David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-aware online adaptation of mixed reality interfaces. In Proceedings of the 32nd annual ACM symposium on user interface software and technology . 147–160
2019
-
[25]
Atieh Taheri, Ziv Weissman, and Misha Sra. 2021. Exploratory design of a hands- free video game controller for a quadriplegic individual. In Proceedings of the Augmented Humans International Conference 2021 . 131–140
2021
-
[26]
Santawat Thanyadit, Matthias Heintz, and Effie LC Law. 2023. Tutor In-sight: Guiding and Visualizing Students’ Attention with Mixed Reality Avatar Pre- sentation Tools. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20
2023
-
[2023]
In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
Ubi Edge: Authoring Edge-Based Opportunistic Tangible User Interfaces in Augmented Reality. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–14
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.