REVIEW 3 major objections 7 minor 1 cited by
Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations
T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Interacting directly with physical objects is what makes them hard to select in mixed reality; shifting interaction to an abstract proxy that is functionally equivalent makes distant, crowded, or occluded targets as manageable as nearby…
desk verdict A promising, honestly-scoped systems contribution whose 'generalizable paradigm' claim needs more evidence than 10 experts and no baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the proxy: a fixed-size, rectangular, manipulable digital stand-in for a physical object, kept persistent near the user's hand and treated as functionally equivalent to the object it represents. Its role is to carry all interaction so that gestures designed for nearby virtual objects—skim, brush, pinch-and-hold, two-hand zoom—apply to real objects regardless of distance, size, or occlusion. Three mechanisms support this: an AI pipeline that recursively detects objects into a spatial hierarchy and attaches semantic attributes; a constraint-based layout optimization that rescales object positions while preserving relative ordering and a minimum spacing, solved with an SMT solver; and a lazy-follow placement that keeps proxies within comfortable reach without overreacting to hand jitter.
What would settle it
A controlled selection experiment on a crowded shelf or rack, comparing direct gaze-plus-pinch selection against proxy selection while systematically compressing the proxy layout, would settle the claim: the paper predicts accuracy and speed stay high under compression, while a sharp rise in selection errors or visible proxy-to-object disambiguation behavior would falsify the layout-fidelity premise.
Extended reading notes
Core claim
In the paper's own terms, a proxy is a manipulable object, and selecting a proxy is functionally equivalent to selecting the actual object. The central claim is that decoupling interaction from the physical constraints of real objects creates a more general interaction paradigm: users can perform set-level operations—brushing across proxies to multi-select, filtering by semantic attribute, grouping by shared properties, and zooming between spatial hierarchy levels—on distant, crowded, or partially occluded objects as easily as on nearby virtual ones. The system realizes this through a three-step process of activating, generating, and interacting with proxies, where AI extracts a hierarchical, semantic scene representation and a constraint-based layout preserves relative spatial order among fixed-size proxies placed near the user's hand. An expert evaluation found the approach useful, easy to learn, and applicable to scenarios ranging from large-scale building navigation to selecting tiny or inaccessible objects, while also revealing alignment and layout-fidelity issues.
Load-bearing premise
The approach depends on the assumption that preserving only the relative spatial order of objects (plus a 0.5 cm minimum spacing) in a compressed, fixed-size proxy layout is enough for users to map each proxy back to its physical object and interact accurately; the expert evaluation showed this mapping can fail, with participants reporting misalignment and needing to look down at proxies to select correctly.
Editorial extensions
If this is right
- If proxy selection is functionally equivalent to object selection, users can select distant or partially occluded objects as reliably as nearby ones, without new gestures or menu systems.
- AI-extracted attributes and spatial hierarchies turn single-object targeting into set-level operations: filter all books by topic, group rooms by department, or brush-select multiple drones.
- Because proxies are persistent and manipulable, they can serve as anchors for collaborative or asynchronous work and for selecting virtual groupings that have no physical existence.
- The approach generalizes beyond AI scene parsing: with digital twins or embedded trackers, the same proxy layer supports building navigation and control of dynamic objects such as drones.
- At the system level, basic proxy interactions such as skim and brush could supplement or refine default raycasting and gaze interactions on commercial MR devices.
Reading between the lines
- Beyond the paper, the proxy principle suggests that physical constraints—distance, size, and occlusion—should be treated as a design variable in MR interaction rather than a given; any real object with a tracked or detected representation could receive a proxy.
- A testable extension is dynamic scenes: with real-time tracking, as the paper sketches for basketball players and drones, proxies would update continuously, allowing selection by attribute or position to be evaluated against raycasting on moving targets.
- The paper's abstraction trade-off implies a design space between fidelity and manipulability; one could measure how much spatial compression users tolerate before the proxy-to-object mapping breaks, which would directly inform layout generation.
- The system effectively turns AI scene understanding into a manipulable interface, suggesting a general pattern for human-AI collaboration in MR: rather than passive annotations, users explore and correct the AI's structure through proxies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Reality Proxy, a Mixed Reality interaction technique that replaces direct selection of physical objects with selection of abstract 'proxies'—fixed-size rectangular representations placed near the user's hand. The system uses a DINO-X + GPT-4o pipeline to extract hierarchical and semantic scene structure, a Z3-based layout optimizer that preserves only relative order constraints and distances, and a lazy-follow hand placement. It claims this shift makes distant, crowded, or occluded objects as easy to interact with as nearby virtual objects, and demonstrates advanced interactions (skimming, brushing multi-select, attribute filtering, semantic grouping, spatial zooming, custom groups) across three application scenarios. The paper reports an expert evaluation with 10 experienced XR users, with mostly positive Likert ratings and qualitative feedback, and concludes that proxy-based abstractions are a 'powerful and generalizable interaction paradigm.'
Significance. If the central claim holds, Reality Proxy would be a genuinely useful contribution to MR interaction, since it addresses real occlusion/distance problems with a conceptually clean decoupling of interaction target from physical object. The paper's strengths are concrete: it ships an open-source prototype, describes a complete AI-to-interaction pipeline, shows diverse application scenarios, and is unusually candid about its limitations. The expert feedback is encouraging and the design space exploration is valuable. However, the significance is currently constrained: the 'generalizable paradigm' claim is supported only by qualitative expert opinion from 10 experienced XR developers in a controlled lab setting, with no quantitative performance data, no comparison baseline, and no direct measurement of the proxy-to-object mapping reliability that the entire concept depends on. The paper itself documents mapping failures. As a technology-probe and design exploration, the work is solid; as a validated general interaction paradigm, the evidence is incomplete.
major comments (3)
- [§3.2, §6.3.2] The paper's central premise—'selecting a proxy is functionally equivalent to selecting the actual object' (§3)—rests on the user reliably recovering the identity and location of the physical object from its proxy. The layout generation in §3.2 deliberately discards shape and metric fidelity: proxies are fixed-size rectangles, and only relative order constraints (e.g., x_A < x_B) plus distance-deviation minimization with 0.5 cm minimum spacing are preserved. In a dense shelf of similar rectangles, the user can only disambiguate if the compressed order maps cleanly back to the physical scene. The paper's own evaluation shows this premise can fail: E5 and E7 reported having to look down at proxies to identify the correct one, and E4 reported brushing boxes that did not cover the intended proxies (§6.3.2). The authors acknowledge in the future-work paragraph of §3.2 that the layout 'can sometimes lead to inaccuracies during interaction.' Since all advanced features (filtering, grouping, multi-select) operate on proxies, any systematic mapping error propagates to every downstream interaction. The claim of functional equivalence should either be reframed as a design goal, or the paper should provide quantitative evidence (e.g., target-identification accuracy and time in dense scenes) that the mapping is reliable enough to support the headline generalization.
- [§6, §7] The expert evaluation is the only empirical support for the 'generalizable interaction paradigm' claim, and it is an entirely internal validation: 10 experienced XR experts (with 2 additional pilots excluded), no baseline condition, and the same authors designed the system, the interaction features, and the Likert instrument that rated those features. The paper itself states 'there is currently no direct comparison baseline' (§6) and later 'we did not conduct a direct quantitative comparison against a baseline, as no fair or appropriate baseline exists' (§7). This is a legitimate choice for a technology probe, but it does not license the abstract's strong claim that the paradigm is 'powerful and generalizable.' The seven-point Likert ratings and thematic analysis demonstrate that experts found the concepts promising, not that proxies outperform or even match existing raycasting or gaze+pinch on objective task performance. I recommend either adding a small comparative or within-subjects task-performance measure (even with acknowledgment of its exploratory nature), or substantially softening the generalization claims throughout the abstract, introduction, and conclusion to scope them to a design exploration.
- [§3.1, §7] The AI scene-understanding pipeline is a load-bearing component for the office/kitchen scenarios, but its accuracy is not reported at all. The paper states (in §3.1) that the pipeline 'cannot handle objects that are fully occluded' and that 'the quality of the scene representation directly determines the user's interaction space'; §7 likewise acknowledges that the AI 'may misclassify or overlook objects.' Yet the abstract and §1 claim that proxies make 'distant, crowded, or partially occluded' objects easy to interact with. Partial occlusion and mis-detection are precisely the conditions that break the proxy generation, so the claimed benefit is asserted for the failure regime of its own core pipeline. At minimum, the paper should report detection and attribute-extraction accuracy on the studied scenes (or state that no accuracy data were collected), and should restrict the headline claim to the digital-twin and sensor-tracked scenarios (§5.2, §5.3) where scene data are known to be reliable.
minor comments (7)
- [Appendix A.2, Table 1] The column labeled 'Z-value' in Table 1 appears to report the Wilcoxon test statistic T (values like 45.0), not a z-score; please relabel and report effect sizes for the post-task questionnaire.
- [§6.4] The word 'declouping' in the sentence 'users' declouping could become problematic' should be 'decoupling.'
- [§4.7] Typo: 'Users can creat a cube-shape group container' should be 'create.'
- [Related Work, §2.2] Reference [78] is described as 'also refereed to as a proxy'; the intended word is 'referred.'
- [§7] The phrase 'explodable proxy layers' appears to be a typo for 'explorable proxy layers.' Additionally, the sentence 'As the players move, the position of each proxy would be to continuously updated' contains a grammar error ('would be to continuously updated').
- [References] Reference [40] (RealitySummary) appears twice in the bibliography; please deduplicate.
- [Figure 11] Consider reporting the exact Likert distributions (e.g., stacked bar counts) rather than only aggregate positive percentages, since the sample is only n=10 and the p-values for Zoom and Reconfigure ease-of-use are marginal (p = .059, p = .08).
Circularity Check
No circularity found: the paper makes no quantitative derivation, and its central proxy-equivalence claim is a stated design definition rather than a result derived from its own inputs.
full rationale
The paper contains no formal derivation chain to cycle. The central claim, "A proxy is a manipulable object, and selecting a proxy is functionally equivalent to selecting the actual object" (Section 3), is a design definition establishing the intended semantics of the proxy concept, not a conclusion inferred from fitted parameters or from a self-citation. The proxy layout step (Section 3.2) is an explicit constraint-based optimization: spatial constraints such as xA < xB are generated from the objects' 3D positions and solved with Z3, so the layout output is openly constructed from stated inputs rather than being presented as a prediction. No parameter is fitted to a subset of data and then renamed as a prediction; the evaluation is a qualitative expert study with Likert ratings and no numeric predictability claim. The paper includes many self-citations involving coauthor Mar Gonzalez-Franco (e.g., references 1, 14, 15, 24, 35, 36, 37, 52, 64, 70, 73, 87, 88), but these support background statements about gaze+pinch challenges, reach extensions, and prior systems; none is invoked as a uniqueness theorem or as the sole justification for forbidding alternative designs. The related-work discussion explicitly credits Poros and World-in-Miniature as conceptually close predecessors, so the proxy idea is not presented as a renamed known result. The fact that validation is entirely internal, with no external baseline, is a study-validity limitation that the authors themselves acknowledge in Section 7 ("we conducted only qualitative evaluations with a small example of experienced XR users ... limits the generalizability"), but under the stated rubric such a limitation is a correctness and evidence concern, not circularity. Accordingly, no circular step can be exhibited with quote and specific reduction.
Assumptions & free parameters
free parameters (4)
- Gaze region extension threshold =
20 cm
- IoU deduplication threshold =
0.75
- Minimum proxy spacing =
0.5 cm
- Proxy rectangle size =
fixed-size rectangle (exact size not specified)
assumptions (4)
- domain assumption The scene mesh provided by the headset operating system is accurate enough for raycast-based 3D localization.
- domain assumption Detected objects are convex and do not intersect, allowing a fixed-size rectangular proxy representation.
- ad hoc to paper Relative order constraints plus distance minimization in the Z3 layout are sufficient to preserve the user's mental model of the scene.
- domain assumption DINO-X and GPT-4o provide correct, complete object detection, segmentation, and semantic attribute extraction.
invented entities (1)
-
Proxy (abstract interaction representation for a physical object)
Cite this review
Pith. "Pith review of Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations." pith.science (2026). https://pith.science/paper/2TR542JH
@misc{pith2026250717248,
author = {Pith},
title = {Pith review of: Reality Proxy: Fluid Interactions with Real-World Objects in MR via Abstract Representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/2TR542JH}},
note = {Machine review of arXiv:2507.17248}
}
read the original abstract
Interacting with real-world objects in Mixed Reality (MR) often proves difficult when they are crowded, distant, or partially occluded, hindering straightforward selection and manipulation. We observe that these difficulties stem from performing interaction directly on physical objects, where input is tightly coupled to their physical constraints. Our key insight is to decouple interaction from these constraints by introducing proxies-abstract representations of real-world objects. We embody this concept in Reality Proxy, a system that seamlessly shifts interaction targets from physical objects to their proxies during selection. Beyond facilitating basic selection, Reality Proxy uses AI to enrich proxies with semantic attributes and hierarchical spatial relationships of their corresponding physical objects, enabling novel and previously cumbersome interactions in MR - such as skimming, attribute-based filtering, navigating nested groups, and complex multi object selections - all without requiring new gestures or menu systems. We demonstrate Reality Proxy's versatility across diverse scenarios, including office information retrieval, large-scale spatial navigation, and multi-drone control. An expert evaluation suggests the system's utility and usability, suggesting that proxy-based abstractions offer a powerful and generalizable interaction paradigm for future MR systems.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
Can AR Embedded Visualizations Foster Appropriate Reliance on AI in Spatial Decision-Making? A Comparative Study of AR X-Ray vs. 2D Minimap
In a time-pressured spatial selection task, an AR X-ray view caused more inappropriate over-reliance on AI than a 2D minimap did, despite improving spatial mapping.
Reference graph
Works this paper leans on
-
[1]
Parastoo Abtahi, Mar Gonzalez-Franco, Eyal Ofek, and Anthony Steed. 2019. I’m a giant: Walking in large virtual environments at high speed gains. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–13
2019
-
[2]
Karan Ahuja, Sujeath Pareddy, Robert Xiao, Mayank Goel, and Chris Harrison
-
[3]
Christoph Anthes, Rubén Jesús García-Hernández, Markus Wiedemann, and Dieter Kranzlmüller. 2016. State of the art of virtual reality technology. In 2016 IEEE aerospace conference. IEEE, 1–19
2016
-
[4]
Apple. 2025. VisionOS. https://developer.apple.com/visionos/. Accessed: 2024-09-13
2025
-
[5]
Argelaguet and C
F. Argelaguet and C. Andujar. 2013. A survey of 3D object selection techniques for virtual environments. Computers & Graphics 37, 3 (2013), 121–136
2013
-
[6]
Iro Armeni, Zhi-Yang He, JunYoung Gwak, Amir R Zamir, Martin Fischer, Jiten- dra Malik, and Silvio Savarese. 2019. 3d scene graph: A structure for unified semantics, 3d space, and camera. In Proceedings of the IEEE/CVF international conference on computer vision . 5664–5673
2019
-
[7]
Rowel Atienza, Ryan Blonna, Maria Isabel Saludares, Joel Casimiro, and Vivencio Fuentes. 2016. Interaction techniques using head gaze for virtual reality. In 2016 IEEE Region 10 Symposium (TENSYMP) . IEEE, 110–114
2016
-
[8]
Marc Baloup, Thomas Pietrzak, and Géry Casiez. 2019. Raycursor: A 3d point- ing facilitation technique based on raycasting. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems . 1–12
2019
Show all 114 references
-
[9]
Michel Beaudouin-Lafon. 2000. Instrumental interaction: an interaction model for designing post-WIMP user interfaces. InProceedings of the SIGCHI conference on Human factors in computing systems . 446–453
2000
-
[10]
Benjamin B Bederson, James D Hollan, Allison Druin, Jason Stewart, David Rogers, and David Proft. 1996. Local tools: An alternative to tool palettes. In Proceedings of the 9th annual ACM symposium on User interface software and technology. 169–170
1996
-
[11]
Lonni Besançon, Anders Ynnerman, Daniel F Keefe, Lingyun Yu, and Tobias Isenberg. 2021. The state of the art of spatial interfaces for 3D visualization. In Computer Graphics Forum, Vol. 40. Wiley Online Library, 293–326
2021
-
[12]
Bier, Maureen C
Eric A. Bier, Maureen C. Stone, Kenneth A. Pier, Kenneth P. Fishkin, Thomas Baudel, Matthew Conway, William Buxton, and Tony DeRose. 1994. Toolglass and magic lenses: the see-through interface. In Conference on Human Factors in Computing Systems, CHI 1994, Boston, Massachusett...
1994
-
[13]
Andrew Bluff and Andrew Johnston. 2019. Don’t panic: Recursive interactions in a miniature metaworld. In Proceedings of the 17th International Conference on Virtual-Reality Continuum and its Applications in Industry . 1–9
2019
-
[14]
Riccardo Bovo, Steven Abreu, Karan Ahuja, Eric J Gonzalez, Li-Te Cheng, and Mar Gonzalez-Franco. 2024. EmBARDiment: an Embodied AI Agent for Pro- ductivity in XR. arXiv preprint arXiv:2408.08158 (2024)
2024 arXiv
-
[15]
Riccardo Bovo, Karan Ahuja, Ryo Suzuki, Mustafa Doga Dogan, and Mar Gonzalez-Franco. 2025. Symbiotic AI: Augmenting Human Cognition from PCs to Cars. arXiv preprint arXiv:2504.03105 (2025)
2025 arXiv
-
[16]
Doug Bowman, Chadwick Wingrave, Joshua Campbell, and Vinh Ly. 2001. Using pinch gloves (tm) for both natural and abstract interaction techniques in virtual environments. (2001)
2001
-
[17]
Virginia Braun and Victoria Clarke. 2019. Reflecting on reflexive thematic analysis. Qualitative research in sport, exercise and health 11, 4 (2019), 589–597
2019
-
[18]
Han Joo Chae, Jeong-in Hwang, and Jinwook Seo. 2018. Wall-based space manipulation technique for efficient placement of distant objects in augmented reality. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology. 45–52
2018
-
[19]
Fang Chen, Natalie Ruiz, Eric Choi, Julien Epps, M Asif Khawaja, Ronnie Taib, Bo Yin, and Yang Wang. 2013. Multimodal behavior and interaction as indicators of cognitive load. ACM Transactions on Interactive Intelligent Systems (TiiS) 2, 4 (2013), 1–36
2013
-
[20]
Ruijia Chen, Junru Jiang, Pragati Maheshwary, Brianna R Cochran, and Yuhang Zhao. 2025. VisiMark: Characterizing and Augmenting Landmarks for People with Low Vision in Augmented Reality to Support Indoor Navigation. arXiv preprint arXiv:2502.10561 (2025)
2025 arXiv
-
[21]
Neil Chulpongsatorn, Mille Skovhus Lunding, Nishan Soni, and Ryo Suzuki. 2023. Augmented Math: Authoring AR-Based Explorable Explanations by Augmenting Static Math Textbooks. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology . 1–16
2023
-
[22]
Marianela Ciolfi Felice, Nolwenn Maudet, Wendy E Mackay, and Michel Beaudouin-Lafon. 2016. Beyond snapping: Persistent, tweakable alignment and distribution with StickyLines. In Proceedings of the 29th Annual Symposium on User Interface Software and Technology . 133–144
2016
-
[23]
Rory MS Clifford, Nikita Mae B Tuanquin, and Robert W Lindeman. 2017. Jedi forceextension: Telekinesis as a virtual reality interaction metaphor. In 2017 IEEE Symposium on 3D User Interfaces (3DUI) . IEEE, 239–240. UIST ’25, September 28-October 1, 2025, Busan, Republic of Kor...
2017
-
[24]
Brian A Cohn, Antonella Maselli, Eyal Ofek, and Mar Gonzalez-Franco. 2020. Snapmove: Movement projection mapping in virtual reality. In 2020 IEEE in- ternational conference on artificial intelligence and virtual reality (AIVR) . IEEE, 74–81
2020
-
[25]
Leonardo De Moura and Nikolaj Bjørner. 2008. Z3: An efficient SMT solver. In International conference on Tools and Algorithms for the Construction and Analysis of Systems. Springer, 337–340
2008
-
[26]
Shujie Deng, Nan Jiang, Jian Chang, Shihui Guo, and Jian J Zhang. 2017. Un- derstanding the impact of multimodal interaction using gaze informed mid-air gesture control in 3D virtual objects manipulation. International Journal of Human-Computer Studies 105 (2017), 68–80
2017
-
[27]
Mustafa Doga Dogan, Eric J Gonzalez, Andrea Colaco, Karan Ahuja, Ruofei Du, Johnny Lee, Mar Gonzalez-Franco, and David Kim. 2024. Augmented Object Intelligence: Making the Analog World Interactable with XR-Objects. arXiv preprint arXiv:2404.13274 (2024)
2024 arXiv
-
[28]
Ruofei Du, Alex Olwal, Mathieu Le Goc, Shengzhi Wu, Danhang Tang, Yinda Zhang, Jun Zhang, David Joseph Tan, Federico Tombari, and David Kim. 2022. Opportunistic interfaces for augmented reality: Transforming everyday objects into tangible 6dof interfaces using ad hoc ui. In CH...
2022
-
[29]
Ruofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso, Ivo Duarte, Jason Dour- garian, Joao Afonso, Jose Pascoal, Josh Gladstone, Nuno Cruces, et al . 2020. DepthLab: Real-time 3D interaction with depth maps for mobile augmented reality. In Proceedings of the 33rd Annual ACM Sy...
2020
-
[30]
Niklas Elmqvist. 2005. BalloonProbe: Reducing occlusion in 3D using interactive space distortion. In Proceedings of the ACM symposium on Virtual reality software and technology. 134–137
2005
-
[31]
Tiare Feuchtner and Jörg Müller. 2018. Ownershift: Facilitating overhead in- teraction in virtual reality with an ownership-preserving hand space shift. In Proceedings of the 31st Annual ACM Symposium on User Interface Software and Technology. 31–43
2018
-
[32]
J Randall Flanagan and Roland S Johansson. 2003. Action plans used in action observation. Nature 424, 6950 (2003), 769–771
2003
-
[33]
Christian Freksa. 1991. Qualitative spatial reasoning. In Cognitive and linguistic aspects of geographic space . Springer, 361–372
1991
-
[34]
Froehlich, Chu Li, Maryam Hosseini, Fabio Miranda, Andres Sevtsuk, and Yochai Eisenberg
Jon E. Froehlich, Chu Li, Maryam Hosseini, Fabio Miranda, Andres Sevtsuk, and Yochai Eisenberg. 2024. The Future of Urban Accessibility: The Role of AI. In Extended Abstract Proceedings of the 26th International ACM SIGACCESS Confer- ence on Computers and Accessibility (St. Jo...
2024
-
[35]
Mar Gonzalez-Franco and Andrea Colaco. 2024. Guidelines for Productivity in Virtual Reality. Interactions 31, 3 (2024), 46–53
2024
-
[36]
Mar Gonzalez-Franco, Zelia Egan, Matthew Peachey, Angus Antley, Tanmay Randhavane, Payod Panda, Yaying Zhang, Cheng Yao Wang, Derek F Reilly, Tabitha C Peck, et al. 2020. Movebox: Democratizing mocap for the microsoft rocketbox avatar library. In 2020 IEEE International Confer...
2020
-
[37]
Mar Gonzalez-Franco and Jaron Lanier. 2017. Model of illusions and virtual reality. Frontiers in psychology 8 (2017), 1125
2017
-
[38]
Tovi Grossman and Ravin Balakrishnan. 2006. The design and evaluation of selection techniques for 3D volumetric displays. InProceedings of the 19th annual ACM symposium on User interface software and technology . 3–12
2006
-
[40]
Aditya Gunturu, Shivesh Jadon, Nandi Zhang, Morteza Faraji, Jarin Thun- dathil, Tafreed Ahmad, Wesley Willett, and Ryo Suzuki. 2024. RealitySum- mary: Exploring On-Demand Mixed Reality Text Summarization and Ques- tion Answering using Large Language Models. arXiv:2405.18620 [c...
2024
-
[41]
Anhong Guo, Xiang’Anthony’ Chen, Haoran Qi, Samuel White, Suman Ghosh, Chieko Asakawa, and Jeffrey P Bigham. 2016. Vizlens: A robust and interactive screen reader for interfaces in the real world. In Proceedings of the 29th annual symposium on user interface software and techn...
2016
-
[42]
Anhong Guo, Jeeeun Kim, Xiang’Anthony’ Chen, Tom Yeh, Scott E Hudson, Jennifer Mankoff, and Jeffrey P Bigham. 2017. Facade: Auto-generating tactile interfaces to appliances. In Proc. CHI. 5826–5838
2017
-
[43]
Anhong Guo, Junhan Kong, Michael Rivera, Frank F Xu, and Jeffrey P Bigham
-
[44]
Aakar Gupta, Bo Rui Lin, Siyi Ji, Arjav Patel, and Daniel Vogel. 2020. Replicate and reuse: Tangible interaction design for digitally-augmented physical media objects. In Proc. CHI. 1–12
2020
-
[45]
Statelens: A reverse engineering solution for making existing dynamic touchscreens accessible. In Proc. UIST. 371–385
-
[46]
Han L Han, Miguel A Renom, Wendy E Mackay, and Michel Beaudouin-Lafon
-
[47]
Taejin Ha, Steven Feiner, and Woontack Woo. 2014. WeARHand: Head-worn, RGB-D camera-based, bare-hand user interface with visually enhanced depth perception. In 2014 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). IEEE, 219–228
2014
-
[48]
Anuruddha Hettiarachchi and Daniel Wigdor. 2016. Annexing reality: Enabling opportunistic use of everyday objects as tangible proxies in augmented reality. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems . 1957–1967
2016
-
[49]
Fernando Hitt. 1998. Difficulties in the articulation of different representations linked to the concept of function. The Journal of Mathematical Behavior 17, 1 (1998), 123–134
1998
-
[50]
Devamardeep Hayatpur, Seongkook Heo, Haijun Xia, Wolfgang Stuerzlinger, and Daniel Wigdor. 2019. Plane, ray, and point: Enabling precise spatial ma- nipulations with shape constraints. In Proceedings of the 32nd annual ACM symposium on user interface software and technology . ...
2019
-
[51]
Tai-Wei Kan, Chin-Hung Teng, and Wen-Shou Chou. 2009. Applying QR code in augmented reality applications. In Proceedings of the 8th international conference on virtual reality continuum and its applications in industry . 253–257
2009
-
[52]
Jinwook Kim, Sangmin Park Zhou, Mar Gonzalez-Franco, Jeongmi Lee, Ken Pfeuffer, et al. 2025. PinchCatcher: Enabling Multi-selection for Gaze+ Pinch. ACM CHI (2025)
2025
-
[53]
Ke Huo, Yuanzhi Cao, Sang Ho Yoon, Zhuangying Xu, Guiming Chen, and Karthik Ramani. 2018. Scenariot: Spatially mapping smart things within aug- mented reality scenes. In Proceedings of the 2018 CHI Conference on human factors in computing systems . 1–13
2018
-
[54]
Ioannis Kotziampasis, Nathan Sidwell, and Alan Chalmers. 2003. Portals: in- creasing visibility in virtual worlds. In Proceedings of the 19th Spring Conference on Computer Graphics. 257–261
2003
-
[55]
André Kunert, Alexander Kulik, Stephan Beck, and Bernd Froehlich. 2014. Pho- toportals: shared references in space and time. In Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing . 1388– 1399
2014
-
[56]
Regis Kopper, Felipe Bacim, and Doug A Bowman. 2011. Rapid and accurate 3D selection by progressive refinement. In 2011 IEEE symposium on 3D user interfaces (3DUI). IEEE, 67–74
2011
-
[57]
Joseph J LaViola Jr, Ernst Kruijff, Ryan P McMahan, Doug Bowman, and Ivan P Poupyrev. 2017. 3D user interfaces: theory and practice . Addison-Wesley Profes- sional
2017
-
[58]
Mathieu Le Goc, Lawrence H Kim, Ali Parsaei, Jean-Daniel Fekete, Pierre Drag- icevic, and Sean Follmer. 2016. Zooids: Building blocks for swarm user interfaces. In Proc. UIST. 97–109
2016
-
[59]
Mikko Kytö, Barrett Ens, Thammathip Piumsomboon, Gun A Lee, and Mark Billinghurst. 2018. Pinpointing: Precise head-and eye-based target selection for augmented reality. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1–14
2018
-
[60]
Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. In Proceedings of the 37th Annual...
2024
-
[61]
Jaewook Lee, Jun Wang, Elizabeth Brown, Liam Chu, Sebastian S Rodriguez, and Jon E Froehlich. 2024. GazePointAR: A Context-Aware Multimodal Voice Assis- tant for Pronoun Disambiguation in Wearable Augmented Reality. InProceedings of the CHI Conference on Human Factors in Compu...
2024
-
[62]
Gun A Lee, Gerard J Kim, and Mark Billinghurst. 2007. Interaction design for tangible augmented reality applications. In Emerging Technologies of Augmented Reality: Interfaces and Design . IGI Global, 261–282
2007
-
[63]
James Liu, Hirav Parekh, Majed Al-Zayer, and Eelke Folmer. 2018. Increasing walking in VR using redirected teleportation. In Proceedings of the 31st annual ACM symposium on user interface software and technology . 521–529
2018
-
[64]
Sebastian Marwecki, Andrew D Wilson, Eyal Ofek, Mar Gonzalez Franco, and Christian Holz. 2019. Mise-unseen: Using eye tracking to hide virtual reality scene changes in plain sight. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology . 777–789
2019
-
[65]
David Lindlbauer and Andy D Wilson. 2018. Remixed reality: Manipulating space and time in augmented reality. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems . 1–13
2018
-
[66]
Mark R Mine. 1995. Virtual environment interaction techniques. UNC Chapel Hill CS Dept (1995)
1995
-
[67]
Mark R Mine, Frederick P Brooks Jr, and Carlo H Sequin. 1997. Moving ob- jects in space: exploiting proprioception in virtual-environment interaction. In Proceedings of the 24th annual conference on Computer graphics and interactive techniques. 19–26. Reality Proxy: Fluid Inte...
1997
-
[68]
Daniel Mendes, Fabio Marco Caputo, Andrea Giachetti, Alfredo Ferreira, and Joaquim Jorge. 2019. A survey on 3d virtual object manipulation: From the desktop to immersive virtual environments. InComputer graphics forum, Vol. 38. Wiley Online Library, 21–45
2019
-
[69]
Kyzyl Monteiro, Ritik Vatsal, Neil Chulpongsatorn, Aman Parnami, and Ryo Suzuki. 2023. Teachable reality: Prototyping tangible augmented reality with everyday objects by leveraging interactive machine teaching. InProc. CHI. 1–15
2023
-
[70]
Gonçalo Padrao, Mar Gonzalez-Franco, Maria V Sanchez-Vives, Mel Slater, and Antoni Rodriguez-Fornells. 2016. Violating body movement semantics: Neural signatures of self-generated and external-generated errors. Neuroimage 124 (2016), 147–156
2016
-
[71]
Roberto A Montano Murillo, Sriram Subramanian, and Diego Martinez Plasencia
-
[72]
Siyou Pei, Alexander Chen, Jaewook Lee, and Yang Zhang. 2022. Hand interfaces: Using hands to imitate objects in ar/vr for expressive interactions. InProceedings of the 2022 CHI conference on human factors in computing systems . 1–16
2022
-
[73]
Ken Pfeuffer, Hans Gellersen, and Mar Gonzalez-Franco. 2024. Design principles and challenges for gaze+ pinch interaction in xr. IEEE Computer Graphics and Applications 44, 3 (2024), 74–81
2024
-
[74]
Ken Pfeuffer, Benedikt Mayer, Diako Mardanbegi, and Hans Gellersen. 2017. Gaze+ pinch interaction in virtual reality. In Proceedings of the 5th symposium on spatial user interaction . 99–108
2017
-
[75]
Randy Pausch, Tommy Burnette, Dan Brockway, and Michael E Weiblen. 1995. Navigation and locomotion in virtual worlds via flight into hand-held miniatures. In Proceedings of the 22nd annual conference on Computer graphics and interactive techniques. 399–400
1995
-
[76]
Jeffrey S Pierce, Brian C Stearns, and Randy Pausch. 1999. Voodoo dolls: seamless interaction at multiple scales in virtual environments. In Proceedings of the 1999 symposium on Interactive 3D graphics . 141–145
1999
-
[77]
Thammathip Piumsomboon, Gun A Lee, Andrew Irlitti, Barrett Ens, Bruce H Thomas, and Mark Billinghurst. 2019. On the shoulder of the giant: A multi-scale mixed reality collaboration with 360 video sharing and tangible interaction. In Proceedings of the 2019 CHI conference on hu...
2019
-
[78]
Henning Pohl, Klemen Lilija, Jess McIntosh, and Kasper Hornbæk. 2021. Poros: configurable proxies for distant interactions in VR. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–12
2021
-
[79]
Jeffrey S Pierce and Randy Pausch. 2004. Navigation with place representations and visible landmarks. In IEEE Virtual Reality 2004 . IEEE, 173–288
2004
-
[80]
Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng, Yuda Xiong, Wenlong Liu, Zhengyu Ma, Junyi Shen, Yuan Gao, Xiaoke Jiang, Xingyu Chen, Zhuheng Song, Yuhong Zhang, Hongjie Huang, Han Gao, Shilong Liu, Hao Zhang, Feng Li, Kent Yu, and Lei Zhang. 2024. DINO-X: A Unified Vision ...
2024 arXiv
-
[81]
Christian Sandor, Arindam Dey, Andrew Cunningham, Sebastien Barbier, Ulrich Eck, Donald Urquhart, Michael R Marner, Graeme Jarvis, and Sang Rhee. 2010. Egocentric space-distorting visualizations for rapid environment exploration in mobile mixed reality. In 2010 IEEE Virtual Re...
2010
-
[82]
Ben Shneiderman. 1983. Direct manipulation: A step beyond programming languages. Computer 16, 08 (1983), 57–69
1983
-
[83]
Ivan Poupyrev, Mark Billinghurst, Suzanne Weghorst, and Tadao Ichikawa. 1996. The go-go interaction technique: non-linear mapping for direct manipulation in VR. In Proceedings of the 9th annual ACM symposium on User interface software and technology. 79–80
1996
-
[84]
Conway, and Randy Pausch
Richard Stoakley, Matthew J. Conway, and Randy Pausch. 1995. Virtual reality on a WIM: interactive worlds in miniature. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Denver, Colorado, USA) (CHI ’95). ACM Press/Addison-Wesley Publishing Co., USA...
1995
-
[85]
Stanislav L Stoev and Dieter Schmalstieg. 2002. Application and taxonomy of through-the-lens techniques. In Proceedings of the ACM symposium on Virtual reality software and technology . 57–64
2002
-
[86]
Hemant Bhaskar Surale, Fabrice Matulic, and Daniel Vogel. 2019. Experimental analysis of barehand mid-air mode-switching techniques in virtual reality. In Proceedings of the 2019 CHI conference on human factors in computing systems . 1–14
2019
-
[87]
Frank Steinicke, Timo Ropinski, and Klaus Hinrichs. 2006. Object selection in virtual environments using an improved virtual pointer metaphor. In Computer Vision and Graphics: International Conference, ICCVG 2004, Warsaw, Poland, September 2004, Proceedings. Springer, 320–326
2006
-
[88]
Junjiao Tian, Lavisha Aggarwal, Andrea Colaco, Zsolt Kira, and Mar Gonzalez- Franco. 2024. Diffuse attend and segment: Unsupervised zero-shot segmentation using stable diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 3554–3563
2024
-
[89]
Wai Tong, Chen Zhu-Tian, Meng Xia, Leo Yu-Ho Lo, Linping Yuan, Benjamin Bach, and Huamin Qu. 2022. Exploring interactions with printed data visual- izations in augmented reality. IEEE transactions on visualization and computer graphics 29, 1 (2022), 418–428
2022
-
[90]
Barbara Tversky. 2008. Spatial Cognition: Embodied and Situated . Cambridge University Press, 201–216
2008
-
[91]
Ryo Suzuki, Eyal Ofek, Mike Sinclair, Daniel Leithinger, and Mar Gonzalez- Franco. 2021. Hapticbots: Distributed encountered-type haptics for vr with multiple shape-changing mobile robots. In The 34th Annual ACM Symposium on User Interface Software and Technology . 1269–1281
2021
-
[92]
Unity Technologies. 2024. Lazy Follow. https://docs.unity3d.com/Packages/com. unity.xr.interaction.toolkit@2.4/manual/lazy-follow.html Accessed: 2024-09-13
2024
-
[93]
Athanasios Voulodimos, Nikolaos Doulamis, Anastasios Doulamis, and Efty- chios Protopapadakis. 2018. Deep learning for computer vision: A brief review. Computational intelligence and neuroscience 2018, 1 (2018), 7068349
2018
-
[94]
Johann Wentzel, Matthew Lakier, Jeremy Hartmann, Falah Shazib, Géry Casiez, and Daniel Vogel. 2024. A Comparison of Virtual Reality Menu Archetypes: Raycasting, Direct Input, and Marking Menus.IEEE Transactions on Visualization and Computer Graphics (2024)
2024
-
[95]
Unity. 2025. PolySpatial. https://docs.unity3d.com/Packages/com.unity. polyspatial.visionos@0.1/manual/index.html. Accessed: 2024-09-13
2025
-
[96]
Dennis Wolf, Jan Gugenheimer, Marco Combosch, and Enrico Rukzio. 2020. Understanding the heisenberg effect of spatial interaction: A selection induced error for spatially tracked input devices. InProceedings of the 2020 CHI conference on human factors in computing systems . 1–10
2020
-
[97]
Haijun Xia, Bruno Araujo, Tovi Grossman, and Daniel Wigdor. 2016. Object- oriented drawing. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems. 4610–4621
2016
-
[98]
Haijun Xia, Bruno Araujo, and Daniel Wigdor. 2017. Collection objects: Enabling fluid formation and manipulation of aggregate selections. In Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems . 5592–5604
2017
-
[99]
Gerald Westheimer. 1954. Mechanism of saccadic eye movements.AMA Archives of Ophthalmology 52, 5 (1954), 710–724
1954
-
[100]
Robert Xiao, Julia Schwarz, Nick Throm, Andrew D Wilson, and Hrvoje Benko
-
[101]
Hui Ye and Hongbo Fu. 2022. ProGesAR: Mobile AR Prototyping for Proxemic and Gestural Interactions with Real-world IoT Enhanced Spaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems . 1–14
2022
-
[102]
Tianwei Yin, Xingyi Zhou, and Philipp Krahenbuhl. 2021. Center-based 3d object detection and tracking. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 11784–11793
2021
-
[103]
Haijun Xia, Sebastian Herscher, Ken Perlin, and Daniel Wigdor. 2018. Space- time: Enabling fluid individual and collaborative editing in virtual reality. In Proceedings of the 31st annual ACM symposium on user interface software and technology. 853–866
2018
-
[104]
Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR visualizations to facilitate stair navigation for people with low vision. In Proceedings of the 32nd annual ACM symposium on user interface software and technology ...
2019
-
[105]
Chen Zhu-Tian, Wai Tong, Qianwen Wang, Benjamin Bach, and Huamin Qu
-
[106]
Chen Zhu-Tian and Haijun Xia. 2022. CrossData: Leveraging Text-Data Connec- tions for Authoring Data Documents. In CHI Conference on Human Factors in Computing Systems. ACM, 95:1–95:15. https://doi.org/10.1145/3491102.3517485
2022
-
[107]
Chen Zhu-Tian, Qisen Yang, Jiarui Shan, Tica Lin, Johanna Beyer, Haijun Xia, and Hanspeter Pfister. 2023. iBall: Augmenting Basketball Videos with Gaze- moderated Embedded Visualizations. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI 2023...
2023
-
[108]
Yifu Zhang, Peize Sun, Yi Jiang, Dongdong Yu, Fucheng Weng, Zehuan Yuan, Ping Luo, Wenyu Liu, and Xinggang Wang. 2022. ByteTrack: Multi-Object Tracking by Associating Every Detection Box. (2022)
2022
-
[109]
Chen Zhu-Tian, Shuainan Ye, Xiangtong Chu, Haijun Xia, Hui Zhang, Huamin Qu, and Yingcai Wu. 2022. Augmenting Sports Videos with VisCommentator. IEEE Transactions on Visualization and Computer Graphics 28, 1 (Jan. 2022), 824–834. https://doi.org/10.1109/TVCG.2021.3114806 UIST ...
2022
-
[111]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Augmenting static visualizations with paparvis designer. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–12
2020
-
[114]
Chen Zhu-Tian, Qisen Yang, Xiao Xie, Johanna Beyer, Haijun Xia, Yingcai Wu, and Hanspeter Pfister. 2022. Sporthesia: Augmenting Sports Videos Using Natural Language. IEEE Transactions on Visualization and Computer Graphics (2022), 1–11. https://doi.org/10.1109/TVCG.2022.320949...
2022
-
[2017]
Erg-O: Ergonomic optimization of immersive virtual environments. In Proc. of UIST. 759–771
-
[2018]
IEEE transactions on visualization and computer graphics 24, 4 (2018), 1653–1660
MRTouch: Adding touch input to head-mounted mixed reality. IEEE transactions on visualization and computer graphics 24, 4 (2018), 1653–1660
2018
-
[2019]
In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology
Lightanchors: Appropriating point lights for spatially-anchored aug- mented reality interfaces. In Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology . 189–196
-
[2020]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
Textlets: Supporting constraints and consistency in text documents. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
2020
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.