Pith. sign in

REVIEW 5 minor 27 references

Everyday AR through AI-in-the-Loop

T0 review · 0 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Everyday AR needs AI in the loop to become feasible

desk verdict A well-organized workshop call for participation, not a research paper; fine as a vision statement but nothing to referee. read the letter →

arxiv 2412.12681 v1 pith:3U7WR4CO submitted 2024-12-17 cs.HC cs.AIcs.LG

classification cs.HCcs.AIcs.LG
keywords AugmentedRealityMixedGenerativeAILargeLanguageModelsHuman-AIInteractionContext-awareARAI-in-the-loopEveryday
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This workshop paper argues that advances in AR hardware and in AI, especially large language models and generative models, make everyday augmented reality increasingly feasible: AR that is always available and seamlessly integrated into daily life, potentially as a general-purpose computing platform. The paper's core proposal is an AI-in-the-loop approach, in which the AR system continuously senses a user's context and adapts content and interactions in real time. A sympathetic reader would care because this reframes AR from a set of specialised productivity or maintenance applications into an infrastructure for daily life, and it gives AI a spatially grounded role: not just answering questions but deciding what the user sees and can do. The paper lays out six research areas that would need to develop together for this vision to hold, from context-aware adaptation to generative content creation to accessible design.

What carries the argument

The central object is the concept of everyday AR, defined as AR that is always available and integrated into users' daily environments. The mechanism that carries the argument is the AI-in-the-loop model: an AR system in which an AI component continuously senses context, including room geometry, object affordances, user activity, and user state, and uses that understanding to adapt content and interactions without explicit user commands. This mechanism is what the paper says will move AR beyond monolithic applications, because it makes both scene understanding and content generation dynamic. The paper identifies six research thrusts that instantiate the mechanism: adaptive and context-aware AR, LLM-powered always-on assistants, AI-assisted task guidance, generative on-demand content creation, AI-driven accessible design, and real-world-oriented AI agents.

What would settle it

A controlled field deployment would settle the claim: give one group an always-on, context-adapting AR assistant and another group a static AR interface for the same everyday tasks; if the adaptive system shows no improvement in completion time, error rate, or perceived load, or if participants disable it because it misreads their context, the premise that AI-in-the-loop is the route to everyday AR is not supported.

Watch

Extended reading notes

Core claim

The paper's central claim is that the combination of recent AI advances and maturing AR hardware enables a new class of experience the authors call everyday AR: always-available, seamlessly integrated digital content that could replace or augment smartphones and desktop computers for many interactions. To get there, the paper asserts, AR must adopt an AI-in-the-loop model, in which digital interactions and content continuously anticipate and adapt to users' changing needs and context. In this model, scene understanding and content generation both become dynamic: large language models support natural, always-on assistant interactions, generative models create on-demand content in real time, and context-aware systems adapt interfaces based on room geometry, object affordances, and user activity. The paper's contribution is a shared research agenda and a call for the community to define the requirements and limitations of this vision rather than a completed system or evaluation.

Load-bearing premise

The whole vision rests on the assumption that current AI, especially large language models and generative models, can be integrated into AR systems well enough to understand and respond to unpredictable real-world contexts in real time, and that users will accept the always-on sensing this requires.

Editorial extensions

If this is right

  • AR would shift from niche applications such as productivity and maintenance to a general-purpose platform for everyday computing, potentially displacing smartphone and desktop interaction for many tasks.
  • Always-on LLM-based assistants embedded in AR could support users implicitly, understanding context from gaze, gestures, and activities rather than requiring typed questions.
  • Generative AI could produce interactive AR content on demand, including 3D objects, scenes, and code-generated interactions, removing the need for manual content authoring.
  • Context-aware adaptation informed by room geometry, object affordances, and user activity would keep augmented interfaces from overwhelming users with irrelevant content.
  • AI-driven input modalities such as voice, gaze, and facial expression recognition could open AR authoring and experiences to people with motor impairments, who are currently excluded by keyboard-mouse and controller interactions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • [Editorial inference] If the AI-in-the-loop hypothesis is correct, evaluation of AR systems will need to move from lab-based rendering-quality measures to longitudinal, in-situ measures of context recognition and task benefit, because the central claim is about everyday use.
  • [Editorial inference] The vision shifts control from the user to the system: users would delegate some interface decisions to AI, which makes trust, privacy, and the ability to override the AI central design problems rather than peripheral ones.
  • [Editorial inference] A concrete testable consequence is that adaptive AR interfaces using live context models should outperform static, manually configured interfaces on real-world everyday tasks; a field deployment comparing the two would provide evidence for or against the paper's premise.
  • [Editorial inference] The paper's emphasis on accessible design suggests AI-driven adaptation could make AR more equitable, but that holds only if the underlying context models do not carry the same biases found in their training data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 5 minor

Summary. This manuscript is a workshop proposal for CHI EA 2025, authored by Suzuki, Gonzalez-Franco, Sra, and Lindlbauer. It argues that recent advances in AR hardware and AI/ML make 'everyday AR'—always-available, seamlessly integrated augmented reality—increasingly feasible, and that achieving it requires an AI-in-the-loop approach in which digital content and interactions continuously anticipate and adapt to user context. The paper identifies six topics of interest (adaptive and context-aware AR, LLM-powered always-on assistants, AI-assisted task guidance, generative AR content creation, accessible AR design, and real-world-oriented AI agents), describes the planned one-day workshop structure and activities, provides a call for participation, and includes organizer biographies and references. The manuscript contains no experiments, derivations, or datasets; its central feasibility claims are explicitly framed as beliefs ('we believe').

Significance. As a workshop proposal, this paper serves a community-building and agenda-setting purpose rather than making a technical contribution. If the vision of everyday AR were realized, it could indeed constitute a major shift in human-computer interaction, and the paper usefully organizes current research threads such as context-aware AR, generative content creation, and always-on AI assistance. It also explicitly acknowledges open challenges including accessibility, privacy, and explainability, and it draws on the organizers' prior workshop experience at UIST 2023 [24], which lends credibility to the organizational plan. However, the manuscript makes no testable claims and provides no evidence for feasibility; its value depends on the workshop's ability to generate and refine community discourse. For a reader expecting a standard research paper, the contribution is limited; for a workshop proposal, it is appropriate. The stress-test concern about whether AI can handle real-world unpredictability is genuine, but it is not load-bearing for this document because the paper's stated purpose is to invite discussion of that very question.

minor comments (5)
  1. [Section 1, first paragraph] The phrase 'make it possible to results in a paradigm shift' is ungrammatical; rewrite it as 'make possible a paradigm shift' or 'result in a paradigm shift.'
  2. [Section 1, first paragraph] The phrase 'users every-changing needs' contains a typo; it should be 'users' ever-changing needs.'
  3. [Section 3.2, Introductions and Lightning Talks] The text reads 'participant's lighting talks'; this should be 'participants' lightning talks.'
  4. [Figure 1 caption] The caption says 'Around 50 participants,' but Section 3.2 reports 40 participants for the same UIST 2023 workshop; please reconcile these numbers.
  5. [Section 3.2, Theme Organization and Discussion] The phrase 'assigned a ‘table. ’' contains stray quotation marks; it should simply read 'assigned a table.'

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a workshop call for participation with no derivation chain or predicted result to reduce to its inputs.

full rationale

This manuscript is a CHI 2025 workshop proposal. It contains no formal derivation, no fitted parameters, no predictive model, and no empirical claim that could be reduced to its inputs by construction. The central statement, 'we believe that everyday AR is increasingly feasible' and that 'we need to adopt an AI-in-the-loop approach,' is an explicitly aspirational vision rather than a derived result. The paper's own text frames the document as a call for participation: 'make a call for participation in which we ask participants to share how they imagine the future of AR as we move to a world where AR is to AI what screens have been to computers.' The only self-citation is reference [24], the authors' prior UIST 2023 workshop, which is used as contextual background ('an evolution of our previous XR and AI Workshop at ACM UIST 2023') and as evidence that a similar workshop format was successfully run ('similar to our successful workshop at ACM UIST 2023'). That citation is not load-bearing for any technical claim, prediction, or uniqueness argument; it merely supports the workshop logistics claim that the format has been used before. All other citations point to external prior work and are used as examples of existing research directions, not as premises that force a conclusion. There is no self-definitional step, no fitted input renamed as a prediction, no uniqueness theorem imported from the authors, and no ansatz smuggled in via citation. The paper is self-contained relative to its purpose: it proposes discussion topics and an event structure, and it explicitly leaves open questions such as privacy, explainability, and sustainability. Because the manuscript makes no verifiable technical claim, no circular reasoning can attach to its central content. Score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper is a workshop proposal and rests entirely on field-level assumptions about AI and AR capabilities. No free parameters or invented entities appear.

assumptions (3)
  • domain assumption Everyday AR is becoming increasingly feasible due to advances in AR hardware (smaller form factors, battery life, connectivity) and AI.
    Stated in Section 1 as a belief, not demonstrated.
  • domain assumption An AI-in-the-loop approach is necessary to make everyday AR usable.
    The workshop's premise; asserted without empirical support in Section 1.
  • domain assumption LLMs and generative AI can effectively drive adaptive, context-aware AR experiences.
    Implicit in the listed topics of interest, e.g., always-on AI assistants and generative AR content creation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Everyday AR through AI-in-the-Loop." pith.science (2026). https://pith.science/paper/3U7WR4CO

@misc{pith2026241212681,
  author       = {Pith},
  title        = {Pith review of: Everyday AR through AI-in-the-Loop},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3U7WR4CO}},
  note         = {Machine review of arXiv:2412.12681}
}
read the original abstract

This workshop brings together experts and practitioners from augmented reality (AR) and artificial intelligence (AI) to shape the future of AI-in-the-loop everyday AR experiences. With recent advancements in both AR hardware and AI capabilities, we envision that everyday AR -- always-available and seamlessly integrated into users' daily environments -- is becoming increasingly feasible. This workshop will explore how AI can drive such everyday AR experiences. We discuss a range of topics, including adaptive and context-aware AR, generative AR content creation, always-on AI assistants, AI-driven accessible design, and real-world-oriented AI agents. Our goal is to identify the opportunities and challenges in AI-enabled AR, focusing on creating novel AR experiences that seamlessly blend the digital and physical worlds. Through the workshop, we aim to foster collaboration, inspire future research, and build a community to advance the research field of AI-enhanced AR.

Figures

Figures reproduced from arXiv: 2412.12681 by the authors.

Figure 1
Figure 1. Photos of the XR+AI workshop at ACM UIST 2023. Around 50 participants engaged in a variety of activities such as [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 17 canonical work pages

  1. [24]

    Ryo Suzuki, Mar Gonzalez-Franco, Misha Sra, and David Lindlbauer. 2023. XR and AI: AI-Enabled Virtual, Augmented, and Mixed Reality. In Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23 Adjunct). Association for Computing Machinery, New York, NY, USA, Article 108, 3 pages. htt...

  2. [1]

    Setareh Aghel Manesh, Tianyi Zhang, Yuki Onishi, Kotaro Hara, Scott Bateman, Jiannan Li, and Anthony Tang. 2024. How People Prompt Generative AI to Create Interactive VR Scenes. InProceedings of the 2024 ACM Designing Interactive Systems Conference. 2319–2340

  3. [2]

    Karan Ahuja, Eyal Ofek, Mar Gonzalez-Franco, Christian Holz, and Andrew D Wilson. 2021. Coolmoves: User motion accentuation in virtual reality.Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies5, 2 (2021), 1–23

  4. [3]

    Riku Arakawa, Hiromu Yakura, and Mayank Goel. 2024. PrISM-Observer: In- tervention Agent to Help Users Perform Everyday Procedures Sensed using a Smartwatch. arXiv preprint arXiv:2407.16785 (2024)

  5. [4]

    Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J Gonzalez, Li-Te Cheng, and Mar Gonzalez-Franco. 2024. EmBARDiment: an Embodied AI Agent for Productivity in XR. arXiv preprint arXiv:2408.08158 (2024)

  6. [5]

    Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, ...

  7. [6]

    Sonia Castelo, Joao Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Iran Roman, Roque Lopez, Ethan Brewer, Chen Zhao, et al. 2023. Argus: Visual- ization of ai-assisted task guidance in ar. IEEE Transactions on Visualization and Computer Graphics (2023)

  8. [7]

    Neil Chulpongsatorn, Mille Skovhus Lunding, Nishan Soni, and Ryo Suzuki. 2023. Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math Textbooks. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–16

Show all 27 references
  1. [8]

    Fernanda De La Torre, Cathy Mengying Fang, Han Huang, Andrzej Banburski- Fahey, Judith Amores Fernandez, and Jaron Lanier. 2024. Llmr: Real-time prompt- ing of interactive worlds using large language models. In Proceedings of the CHI Conference on Human Factors in Computing Sy...

  2. [9]

    Mustafa Doga Dogan, Eric J Gonzalez, Andrea Colaco, Karan Ahuja, Ruofei Du, Johnny Lee, Mar Gonzalez-Franco, and David Kim. 2024. Augmented Object Intelligence: Making the Analog World Interactable with XR-Objects. arXiv preprint arXiv:2404.13274 (2024)

  3. [10]

    Ran Gal, Lior Shapira, Eyal Ofek, and Pushmeet Kohli. 2014. FLARE: Fast layout for augmented reality applications. In 2014 IEEE international symposium on mixed and augmented reality (ISMAR) . IEEE, 207–212

  4. [11]

    Mar Gonzalez-Franco, Julio Cermeron, Katie Li, Rodrigo Pizarro, Jacob Thorn, Windo Hutabarat, Ashutosh Tiwari, and Pablo Bermell-Garcia. 2016. Immersive augmented reality training for complex manufacturing scenarios. arXiv preprint arXiv:1602.01944 (2016)

  5. [12]

    Mar Gonzalez-Franco and Andrea Colaco. 2024. Guidelines for Productivity in Virtual Reality. Interactions 31, 3 (2024), 46–53

  6. [13]

    Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde- Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Gener- ative Adversarial Nets. Advances in Neural Information Processing Sys- tems 27 (2014). https://proceedings.neurips.cc/paper_files/pape...

  7. [14]

    Jens Grubert, Tobias Langlotz, Stefanie Zollmann, and Holger Regenbrecht. 2017. Towards Pervasive Augmented Reality: Context-Awareness in Augmented Reality. IEEE Transactions on Visualization and Computer Graphics 23, 6 (2017), 1706–1724. https://doi.org/10.1109/TVCG.2016.2543720

  8. [15]

    Aditya Gunturu, Shivesh Jadon, Nandi Zhang, Morteza Faraji, Jarin Thundathil, Tafreed Ahmad, Wesley Willett, and Ryo Suzuki. 2024. RealitySummary: Explor- ing On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models. arXiv preprint arXiv:24...

  9. [16]

    Violet Yinuo Han, Hyunsung Cho, Kiyosu Maeda, Alexandra Ion, and David Lindlbauer. 2023. BlendMR: A Computational Method to Create Ambient Mixed Reality Interfaces. Proceedings of the ACM on Human-Computer Interaction 7, ISS (2023), 217–241

  10. [17]

    Fengming He, Xiyun Hu, Jingyu Shi, Xun Qian, Tianyi Wang, and Karthik Ramani

  11. [18]

    Teresa Hirzle, Florian Müller, Fiona Draxler, Martin Schmitz, Pascal Knierim, and Kasper Hornbæk. 2023. When XR and AI Meet-A Scoping Review on Extended Reality and Artificial Intelligence. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–45

  12. [19]

    Rahul Jain, Jingyu Shi, Runlin Duan, Zhengzhe Zhu, Xun Qian, and Karthik Ramani. 2023. Ubi-TOUCH: Ubiquitous Tangible Object Utilization through Consistent Hand-object interaction in Augmented Reality. In Proceedings of the 36th Annual ACM Symposium on User Interface Software ...

  13. [20]

    Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. arXiv preprint arXiv:2407.13515 (2024)

  14. [21]

    Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S Rodriguez, and Jon E Froehlich. 2024. GazePointAR: A Context-Aware Multimodal Voice Assis- tant for Pronoun Disambiguation in Wearable Augmented Reality. InProceedings of the CHI Conference on Human Factors in Compu...

  15. [22]

    David Lindlbauer. 2022. The future of mixed reality is adaptive.XRDS: Crossroads, The ACM Magazine for Students 29, 1 (2022), 26–31

  16. [23]

    David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-aware online adaptation of mixed reality interfaces. In Proceedings of the 32nd annual ACM symposium on user interface software and technology . 147–160

  17. [25]

    Atieh Taheri, Ziv Weissman, and Misha Sra. 2021. Exploratory design of a hands- free video game controller for a quadriplegic individual. In Proceedings of the Augmented Humans International Conference 2021 . 131–140

  18. [26]

    Santawat Thanyadit, Matthias Heintz, and Effie LC Law. 2023. Tutor In-sight: Guiding and Visualizing Students’ Attention with Mixed Reality Avatar Pre- sentation Tools. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–20

  19. [2023]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Ubi Edge: Authoring Edge-Based Opportunistic Tangible User Interfaces in Augmented Reality. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–14

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.