Pith. sign in

REVIEW 3 major objections 4 minor 135 references

The paper argues that visually distinguishing objects by importance level in a wearable AR display shifts people with low vision toward high-importance objects, at the cost of recalling fewer objects overall.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 07:04 UTC pith:GVZV2IQN

load-bearing objection The design exploration is genuinely useful, but the headline quantitative claim about importance-based attention shift doesn't survive contact with the paper's own thresholds or its AR baseline. the 3 major comments →

arxiv 2607.10902 v2 pith:GVZV2IQN submitted 2026-07-12 cs.HC

What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision

classification cs.HC
keywords low visionaugmented realitycomplex scenesobject importanceattentionscene recallAR distinctionwearable computing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper proposes that to help people with low vision perceive cluttered real-world scenes, AR systems should do more than highlight objects—they should visually rank objects by importance and render each level differently. To test this, the authors built SceneGlance, a wearable AR system that detects kitchen and street objects and augments them with importance-differentiated outlines, overlays, or icons. In a lab study with 12 people with low vision, the importance-differentiated display shifted first glances toward high-importance objects (79% vs 21% without AR) but reduced overall object recall, an attention-recall tradeoff. The authors also identify new failure modes of multi-object AR, such as adjacent highlights merging into misleading shapes and floating icons misaligning with objects. The paper's contribution is the empirical demonstration that AR distinction steers attention in complex scenes, together with design guidance for making such augmentation work outdoors and under variable lighting.

Core claim

On its own terms, the paper establishes that AR distinction—rendering objects of different importance with different visual treatment—is a workable strategy for guiding attention in complex scenes for people with low vision. SceneGlance detects and segments up to 11 kitchen and 21 street object classes, assigns them to primary- or secondary-importance, and augments them accordingly. In a controlled mock-kitchen study, participants first noticed primary-important objects in 79.2% of SceneGlance trials versus 20.8% with only their own vision (p < 0.001), and recall ratios for primary-important objects rose significantly, while overall object recall fell by about 8 points in both AR conditions.

What carries the argument

SceneGlance: a HoloLens-based wearable system that streams video to a server running fine-tuned RTMDet segmentation models (fine-tuned on kitchen and street datasets), raycasts detections onto the environmental mesh for 3D placement, and renders three base augmentations (static outline, solid overlay, icon label) combined into three distinction methods (by form, by color, by additional visual information). The load-bearing idea is the importance ranking itself, derived from a six-participant formative study that categorized objects as safety-related, visually challenging, or frequently used, and rated them by risk severity and visual difficulty into primary versus secondary importance; this

Load-bearing premise

The claim that AR distinction shifted attention assumes the object detector's live performance was accurate enough that the effect comes from the importance-based rendering, not from which objects happened to get augmented; the reported offline false-negative rates (29.3% kitchen, 42.3% street) mean undetected objects—especially transparent glasses and curb cuts—could be driving both the attention shift and the recall drop.

What would settle it

Re-run Study I with per-trial logging of which objects were actually augmented; restrict the analysis to trials where every primary-important object in the layout was detected and augmented. If the first-noticed rate no longer differs from the Reality baseline, the attention effect is an artifact of detection, not of importance-based distinction.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Importance-differentiated AR augmentation reliably shifts first attention toward higher-importance objects for people with low vision in cluttered scenes.
  • Multi-object augmentation imposes an attention-recall tradeoff: gains in high-importance recall come with reduced recall of un-augmented non-important objects.
  • Adjacent augmentations of the same importance merge into misleading shapes, so distinction systems must also consider spatial relations, not only per-object importance.
  • In outdoor scenes, continuous surfaces and dynamic objects need different treatments: outlines for boundaries, and importance based on collision likelihood rather than fixed object type.
  • Color-based distinction degrades under variable outdoor lighting, while form-based distinction (overlay vs outline, outline thickness) remains recognizable.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the attention-recall tradeoff generalizes, importance-based AR acts like a spotlight that narrows the visual field; adaptive granularity (group outlines with counts) could recover breadth without losing the spotlight.
  • The reported effect may partly reflect detection asymmetries: with false negatives highest for transparent objects and curb cuts, the attention shift could be an artifact of which objects were actually augmented, so per-trial augmentation logging is needed to separate design from detection.
  • A testable extension: combining importance distinction with predicted trajectory (e.g., a car at a crosswalk rising to primary-importance as it approaches) should strengthen perceived safety benefits relative to static importance ranking.
  • The merging of adjacent augmentations suggests a concrete design heuristic: vary color or texture across adjacent same-importance objects, or use outline thickness, to break uniform connectedness—something the paper implies but does not implement.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents SceneGlance, a wearable AR system that detects important objects in complex scenes and visually distinguishes them by importance level (primary vs. secondary) using forms, colors, and/or additional information. The work is grounded in a formative study with six PLV that characterized important object categories and derived design guidelines. The system is evaluated in two studies: a controlled lab study with 12 PLV performing scene-perception tasks in a mock kitchen under three conditions (Reality baseline, AR baseline, SceneGlance), and a free-form outdoor think-aloud study with 13 PLV. The paper claims that AR distinction shifted attention toward primary-important objects, supported new perception strategies, and produced an attention-recall tradeoff, while also surfacing AR-specific challenges and design implications.

Significance. If the central claims were fully supported, this would be a valuable contribution to AR-based vision enhancement for low vision in complex scenes, an underexplored area. The paper's strengths include: a formative study that directly informs the system design; a fully implemented, near-real-time wearable system with a technical evaluation; a controlled three-condition lab study with PLV; and a complementary outdoor study that identifies realistic deployment challenges. The qualitative findings about augmentation interference, occlusion, and spatial misalignment are useful for future system design. However, the quantitative evidence does not currently support the abstract's strong claim that the importance-distinction design, rather than generic augmentation, drove the attention shift and recall tradeoff. The paper's own Bonferroni-corrected analyses show no significant difference between SceneGlance and the AR baseline on attention or recall measures, and two of the reported 'significant' pairwise comparisons exceed the stated threshold. These issues are fixable through revised analyses and reporting, but they are load-bearing for the paper's central message.

major comments (3)
  1. [§5.3.2, §5.3.3, Abstract] The abstract's claim that 'AR distinction on object importance shifted PLV's attention toward objects of higher importance' is not supported by the reported statistics. For primary-important recall ratio, SceneGlance does not differ from the AR baseline (p=0.927); for first-noticed objects, SceneGlance also does not differ from the AR baseline (p=0.149), and the AR baseline falls in between without significant separation from Reality (p=0.426). The only significant contrasts are versus the no-AR Reality baseline, which is consistent with a generic augmentation effect rather than a distinction-specific effect. The conclusion in Section 5.3.2 should be tempered, and the abstract should not attribute the effects to 'AR distinction' without significant SceneGlance-vs-AR-baseline evidence or additional analysis that isolates the distinction manipulation.
  2. [§5.3.3, §5.3.2] The reported pairwise p-values are inconsistent with the stated Bonferroni threshold (α=0.0056). In §5.3.3, overall recall is described as significantly lower under SceneGlance (diff=-0.079, p=0.013), but 0.013 > 0.0056, so this contrast is not significant by the paper's own criterion. Likewise, in §5.3.2 the primary-important recall ratio contrasts for SceneGlance (p=0.006) and the AR baseline (p=0.018) are labeled 'significantly higher,' yet both exceed 0.0056. This is an internal inconsistency that affects the abstract and the discussion of the attention-recall tradeoff. The authors should re-run or re-report the post-hoc comparisons with the stated correction and revise the claims accordingly.
  3. [§4.3.3, §5.3] The attention-shift and recall-reduction results relative to Reality could be confounded by differential detection coverage across importance levels. Offline false-negative rates are high and vary strongly by class (e.g., glasses 39.4%, jars 37.7%, curb cut 89.0%, crosswalk 79.1%), and no per-trial online augmentation coverage is reported. If undetected objects were systematically clustered in primary-important or secondary-important classes, the Reality-vs-SceneGlance differences could reflect which objects received augmentations rather than the importance-distinction design. The AR baseline partially controls for the detection pipeline, but because SceneGlance and the AR baseline are not significantly separated, the distinction-specific interpretation is not established. Please report per-condition augmentation coverage, and/or re-analyze attention and recall restricted to objects that
minor comments (4)
  1. [§4.3.2] The text says the models 'demonstrated strong robustness' and 'high accuracy,' but the reported mAP values (0.435 kitchen, 0.324 street) and false-negative rates (29.3% and 42.3%) suggest modest performance. Please soften this characterization to match the data.
  2. [§5.3.2] The phrase '(p=0.034 > 0.0056 with correction)' is awkward; the paper uses the threshold both as a significance level and as a post-hoc correction. Please clarify that no significance is claimed when p > 0.0056.
  3. [Table 2 and §5.3.2] The first-noticed analysis treats the two trials per condition as independent responses; this should be acknowledged or handled with a repeated-measures model. Also, the 22 vs 24 valid responses across conditions are not tested for the effect of missingness.
  4. [§5.1.3] In the AR baseline, participants chose their preferred single base augmentation, but the choice is not recorded as a factor. Differences in the chosen augmentation (outline vs overlay vs icon) could affect attention and recall; at least a sensitivity analysis or descriptive summary would help.

Circularity Check

0 steps flagged

No significant circularity: SceneGlance is an empirical design-and-evaluation study; importance labels are inputs, and attention/recall outcomes are measured independently against Reality and AR baselines.

full rationale

The paper makes no formal derivation or predictive claim that reduces to its inputs. The formative study (Section 3) elicited important objects and importance levels from six PLV; these labels were used as design inputs for SceneGlance (Section 4). The evaluation (Section 5) then measured attention allocation via recall ratios and first-noticed-object distribution under three conditions (Reality, AR baseline, SceneGlance). The outcome measures are behaviorally coded and statistically compared, not computed from the importance labels or from any fitted parameter. The AR baseline condition explicitly controls for generic augmentation, so the SceneGlance-vs-AR-baseline comparison is the right contrast for the distinction-specific claim; the fact that those comparisons were not significant (primary-important recall ratio p=0.927; first-noticed distribution p=0.149, Section 5.3.2) is a validity/interpretation concern, not a circularity. Self-citations (e.g., [123] for yellow color preference, [13,58,126] for augmentation designs) are used only to motivate design choices and describe related work; no load-bearing argument reduces to a self-cited theorem. The limitations stated in Section 7.2 (small sample, carryover, non-controlled Study II) are acknowledged threats to generalizability, not evidence of circular derivation. Accordingly, no circular step can be quoted with a specific reduction; the correct finding is no significant circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

No free parameters are fitted to a predictive mathematical model; the entries above are hand-chosen system and study settings that the central claims depend on. The main analytic assumptions are representative sampling, task realism, fixed importance labels, and detector reliability.

free parameters (4)
  • Importance category assignment (primary vs secondary) = Not disclosed
    The central manipulation depends on these hand-assigned categories, but the mapping from five-point importance ratings to the two levels is not reported.
  • Detection confidence threshold = 0.45
    Used to compute reported recognition/error rates (Tables 6-7) and to decide which objects get augmented; no sensitivity analysis is given.
  • IoU threshold for detection-level metrics = 0.3
    Chosen to match deployment parameters; small changes would alter false negative and false discovery rates.
  • Scene complexity parameters (mock kitchen) = 30-33 objects; 12 categories; 150cm x 60cm counter
    Selected to mimic a complex scene; no validation that this matches real kitchen clutter or that the outdoor route generalizes to other complex streets.
axioms (5)
  • domain assumption Participants are representative of the broader PLV population
    Small convenience samples (6, 12, 13) from local clinics, with three participants overlapping studies; claims about PLV in general rest on this.
  • domain assumption The mock kitchen and outdoor route approximate real-world complexity
    Objects arranged on a 150x60 cm counter and a preplanned 335-meter route; dynamic objects and lighting variability are only partially represented.
  • ad hoc to paper Object importance can be treated as a stable two-level property
    SceneGlance assigns fixed primary/secondary labels per object class from the formative study, but Section 7.1 states importance is context-dependent; this assumption is necessary for the central manipulation yet is partially contradicted by the paper's own discussion.
  • domain assumption Recognition errors did not systematically bias attention measures
    Offline false-negative rates are 29.3% (kitchen) and 42.3% (street), but per-trial online augmentation coverage is not reported; the analysis treats augmentations as a stable experimental condition.
  • standard math Standard statistical assumptions (normality, independence, chi-square approximation)
    Used LME/ANOVA and Pearson chi-square tests; non-normal measures were handled with ART ANOVA, but small-sample multiplicity remains a concern.

pith-pipeline@v1.3.0-alltime-deepseek · 35703 in / 15082 out tokens · 141315 ms · 2026-08-02T07:04:42.511565+00:00 · methodology

0 comments
read the original abstract

People with low vision (PLV) struggle to perceive complex scenes like busy kitchens and crowded streets, which contain many objects, visual clutter, and dynamic elements. Prior AR systems for low vision either enhance low-level visual features or augment task-relevant objects for single tasks in simple settings, leaving multi-object augmentation in complex scenes underexplored. Informed by a formative study characterizing important objects and their perceived importance for PLV, we built SceneGlance, a wearable AR system that recognizes important objects and visually distinguishes them by importance level. Through a controlled lab study with 12 PLV in a mock-up kitchen scene and a free-form think-aloud study with 13 PLV navigating an outdoor route, we found that AR distinction on object importance shifted PLV's attention toward objects of higher importance, and supported perception strategies such as building mental snapshots from the augmentation distribution and hierarchical scanning by importance. However, this attention shift came with a tradeoff of reduced overall scene recall. The studies also surfaced challenges posed by AR augmentations in complex scenes, such as adjacent augmentations blending or interfering with each other, yielding design implications for more practical AR vision enhancement systems in the complex real world.

Figures

Figures reproduced from arXiv: 2607.10902 by Jaewook Lee, Jia Li, Jon E. Froehlich, Kexin Zhang, Mengfong Lio, Ruijia Chen, Sanbrita Mondal, Weibing Wang, Yapeng Tian, Yuhang Zhao, Yuheng Wu.

Figure 1
Figure 1. Figure 1: We explore the opportunities and design challenges of augmenting and distinguishing multiple objects in complex [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Research method overview. A formative study with six PLV characterized important objects, their perceived importance [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 2
Figure 2. Figure 2: Research method overview. A formative study with six PLV characterized important objects, their perceived importance [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: An illustration of the three AR distinction methods in SceneGlance. (A) [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: System pipeline of SceneGlance: the HoloLens (frontend) streams video to the backend; the backend runs the fine-tuned [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Example inference results of the two fine-tuned models on test images in the kitchen (A) and outdoor environment (B, [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Kitchen Countertop Perception Task. (A) Participants sat at a mock-up kitchen table with 30–33 kitchen objects, [PITH_FULL_IMAGE:figures/full_fig_p009_6.png] view at source ↗
Figure 6
Figure 6. Figure 6: Kitchen Countertop Perception Task. (A) Participants sat at a mock-up kitchen table with 30–33 kitchen objects, [PITH_FULL_IMAGE:figures/full_fig_p010_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Examples of perception challenges in Study I. (A) [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: The route for the outdoor navigation study, which [PITH_FULL_IMAGE:figures/full_fig_p014_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Examples of perception challenges identified in Study II. (A) Outlines of the sidewalk and a railway track crossed the [PITH_FULL_IMAGE:figures/full_fig_p016_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Example scenarios and augmentation designs in the formative study design probe. (A)-(B) Two example kitchen [PITH_FULL_IMAGE:figures/full_fig_p021_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Example distinction methods in the formative study design probe. (A) The [PITH_FULL_IMAGE:figures/full_fig_p022_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

135 extracted references · 6 linked inside Pith

  1. [1]

    Walid I Al-Atabany, Muhammad A Memon, Susan M Downes, and Patrick A Degenaar. 2010. Designing and testing scene enhancement algorithms for patients with retina degenerative disorders.Biomedical engineering online9, 1 (2010), 27. ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal Wu et al

  2. [2]

    Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. 2016. Social lstm: Human trajectory prediction in crowded spaces. InProceedings of the IEEE conference on computer vision and pattern recognition. 961–971

  3. [3]

    George A Alvarez and Patrick Cavanagh. 2004. The capacity of visual short- term memory is set both by visual information load and by number of objects. Psychological science15, 2 (2004), 106–111

  4. [4]

    American Optometric Association. 2023. Legal blindness in America — aoa.org. https://www.aoa.org/news/clinical-eye-care/diseases-and-conditions/ legal-blindness-in-america?sso=y. [Accessed 12-03-2025]

  5. [5]

    Anastasios Nikolas Angelopoulos, Hossein Ameri, Debbie Mitra, and Mark Humayun. 2019. Enhanced depth navigation through augmented reality depth mapping in patients with low vision.Scientific reports9, 1 (2019), 11230

  6. [6]

    Richard A Armstrong. 2014. When to use the Bonferroni correction.Ophthalmic and physiological optics34, 5 (2014), 502–508

  7. [7]

    Marie Claire Bilyk, Jessica M Sontrop, Gwen E Chapman, Susan I Barr, and Linda Mamer. 2009. Food experiences and eating patterns of visually impaired and blind people.Canadian Journal of Dietetic practice and research70, 1 (2009), 13–18

  8. [8]

    Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology3, 2 (2006), 77–101

  9. [9]

    Virginia Braun and Victoria Clarke. 2024. Thematic analysis. InEncyclopedia of quality of life and well-being research. Springer, 7187–7193

  10. [10]

    Xiaoqian J Chai, Noa Ofen, Lucia F Jacobs, and John DE Gabrieli. 2010. Scene complexity: influence on perception, memory, and development in the medial temporal lobe.Frontiers in human neuroscience4 (2010), 1021

  11. [11]

    Ruei-Che Chang, Yuxuan Liu, and Anhong Guo. 2024. Worldscribe: Towards context-aware live visual descriptions. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–18

  12. [12]

    Xiaojun Chang, Pengzhen Ren, Pengfei Xu, Zhihui Li, Xiaojiang Chen, and Alex Hauptmann. 2021. A comprehensive survey of scene graphs: Generation and application.IEEE Transactions on Pattern Analysis and Machine Intelligence45, 1 (2021), 1–26

  13. [13]

    Ruijia Chen, Junru Jiang, Pragati Maheshwary, Brianna R Cochran, and Yuhang Zhao. 2025. Visimark: Characterizing and augmenting landmarks for people with low vision in augmented reality to support indoor navigation. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–20

  14. [14]

    Ruijia Chen, Yuheng Wu, Charlie Houseago, Filipe Gaspar, Filippo Ale- otti, Dorian Gálvez-López, Oliver Johnston, Diego Mazala, Guillermo Garcia- Hernando, Maryam Bandukda, Gabriel Brostow, and Jessica Van Brummelen

  15. [15]

    2013.Statistical power analysis for the behavioral sciences

    Jacob Cohen. 2013.Statistical power analysis for the behavioral sciences. Rout- ledge

  16. [16]

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray. 2018. Scaling Egocentric Vision: The EPIC- KITCHENS Dataset. InEuropean Conference on Computer Vision (ECCV)

  17. [17]

    Raghavendra Singh Dasila, Meet Trivedi, Shubham Soni, M Senthil, and M Narendran. 2017. Real time environment perception for visually impaired. In 2017 IEEE Technological Innovations in ICT for Agriculture and Rural Development (TIAR). IEEE, 168–172

  18. [18]

    Johanne Desrosiers, Marie-Chantal Wanet-Defalque, Khatoune Témisjian, Jacques Gresset, Marie-France Dubois, Judith Renaud, Claude Vincent, Jacque- line Rousseau, Mathieu Carignan, and Olga Overbury. 2009. Participation in daily activities and social roles of older adults with visual impairment.Disability and rehabilitation31, 15 (2009), 1227–1234

  19. [19]

    Juan C Dibene and Enrique Dunn. 2022. HoloLens 2 Sensor Streaming.arXiv preprint arXiv:2211.02648(2022)

  20. [20]

    Olive Jean Dunn. 1961. Multiple comparisons among means.Journal of the American statistical association56, 293 (1961), 52–64

  21. [21]

    Martin Eckert, Matthias Blex, Christoph M Friedrich, et al. 2018. Object detection featuring 3D audio localization for Microsoft HoloLens. InProc. 11th Int. Joint Conf. on Biomedical Engineering Systems and Technologies, Vol. 5. 555–561

  22. [22]

    Elkin, Matthew Kay, James J

    Lisa A. Elkin, Matthew Kay, James J. Higgins, and Jacob O. Wobbrock. 2021. An Aligned Rank Transform Procedure for Multifactor Contrast Tests. InThe 34th Annual ACM Symposium on User Interface Software and Technology(Virtual Event, USA)(UIST ’21). Association for Computing Machinery, New York, NY, USA, 754–768. doi:10.1145/3472749.3474784

  23. [23]

    Niklas Elmqvist and Philippas Tsigas. 2008. A taxonomy of 3d occlusion manage- ment for visualization.IEEE transactions on visualization and computer graphics 14, 5 (2008), 1095–1109

  24. [24]

    MR Everingham, BT Thomas, T Troscianko, et al. 1999. Head-mounted mobility aid for low vision using scene classification techniques.The International Journal of Virtual Reality3, 4 (1999), 3

  25. [25]

    Mark Everingham, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. 2010. The pascal visual object classes (voc) challenge. International journal of computer vision88 (2010), 303–338

  26. [26]

    Fox, Ahmad Ahmadzada, Clara T

    Dylan R. Fox, Ahmad Ahmadzada, Clara T. Friedman, Shiri Azenkot, Marlena A. Chu, Roberto Manduchi, and Emily A. Cooper. 2023. Using augmented reality to cue obstacles for people with low vision.Opt. Express31, 4 (Feb 2023), 6827–6848. doi:10.1364/OE.479258

  27. [27]

    Bhanuka Gamage, Nicola McDowell, Dijana Kovacic, Leona Holloway, Thanh- Toan Do, Arthur James Lowery, Nicholas Price, and Kim Marriott. 2025. Smart Glasses for CVI: Co-Designing Extended Reality Solutions to Support Environ- mental Perception by People with Cerebral Visual Impairment. InProceedings of the 27th International ACM SIGACCESS Conference on Com...

  28. [28]

    Yun Gao, Dan Wu, Jie Song, Xueyi Zhang, Bangbang Hou, Hengfa Liu, Junqi Liao, and Liang Zhou. 2025. A wearable obstacle avoidance device for visually impaired individuals with cross-modal learning.Nature Communications16, 1 (2025), 2857

  29. [29]

    Google. 2025. Protocol Buffers. https://developers.google.com/protocol-buffers. Accessed: 2025-03-30

  30. [30]

    Tenen- baum, Antonio Torralba, Florian Shkurti, and Liam Paull

    Qiao Gu, Alihusein Kuwajerwala, Sacha Morin, Krishna Murthy Jatavallab- hula, Bipasha Sen, Aditya Agarwal, Corban Rivera, William Paul, Kirsty El- lis, Rama Chellappa, Chuang Gan, Celso Miguel de Melo, Joshua B. Tenen- baum, Antonio Torralba, Florian Shkurti, and Liam Paull. 2023. Concept- Graphs: Open-Vocabulary 3D Scene Graphs for Perception and Plannin...

  31. [31]

    Fangli Guan, Zhixiang Fang, Lubin Wang, Xucai Zhang, Haoyu Zhong, and Haosheng Huang. 2022. Modelling people’s perceived scene complexity of real-world environments using street-view panoramas and open geodata.ISPRS Journal of Photogrammetry and Remote Sensing186 (2022), 315–331

  32. [32]

    Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi

  33. [33]

    Yu Hao, Alexey Magay, Hao Huang, Shuaihang Yuan, Congcong Wen, and Yi Fang. 2024. ChatMap: A Wearable Platform Based on the Multi-modal Foun- dation Model to Augment Spatial Cognition for People with Blindness and Low Vision. In2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 129–134

  34. [34]

    Yu Hao, Fan Yang, Hao Huang, Shuaihang Yuan, Sundeep Rangan, John-Ross Rizzo, Yao Wang, and Yi Fang. 2024. A multi-modal foundation model to assist people with blindness and low vision in environmental interaction.Journal of Imaging10, 5 (2024), 103

  35. [35]

    Sharon A Haymes, Alan W Johnston, and Anthony D Heyes. 2002. Relationship between vision impairment and ability to perform activities of daily living. Ophthalmic and Physiological Optics22, 2 (2002), 79–91

  36. [36]

    John M Henderson, Myriam Chanceaux, and Tim J Smith. 2009. The influence of clutter on real-world scene search: Evidence from search efficiency and eye movements.Journal of vision9, 1 (2009), 32–32

  37. [37]

    John M Henderson, Taylor R Hayes, Candace E Peacock, and Gwendolyn Rehrig

  38. [38]

    Marion Hersh. 2022. Wearable travel aids for blind and partially sighted people: A review with a focus on design issues.Sensors22, 14 (2022), 5454

  39. [39]

    Stephen L Hicks, Iain Wilson, Louwai Muhammed, John Worsfold, Susan M Downes, and Christopher Kennard. 2013. A depth-based head-mounted visual display to aid navigation in partially sighted individuals.PLOS ONE8, 7 (2013), e67695

  40. [40]

    Jonathan Huang, Max Kinateder, Matt J Dunn, Wojciech Jarosz, Xing-Dong Yang, and Emily A Cooper. 2019. An augmented reality sign-reading assistant for users with reduced vision.PloS one14, 1 (2019), e0210630

  41. [41]

    Alex D Hwang and Eli Peli. 2014. An augmented-reality edge enhancement application for Google Glass.Optometry and vision science91, 8 (2014), 1021– 1030

  42. [42]

    Md Touhidul Islam, Imran Kabir, Elena Ariel Pearce, Md Alimoor Reza, and Syed Masum Billah. 2024. Identifying Crucial Objects in Blind and Low-Vision Individuals’ Navigation. InProceedings of the 26th International ACM SIGACCESS Conference on Computers and Accessibility. 1–8

  43. [43]

    Watthanasak Jeamwatthanachai, Mike Wald, and Gary Wills. 2019. Indoor navigation by blind people: Behaviors and challenges in unfamiliar spaces and buildings.British Journal of Visual Impairment37, 2 (2019), 140–153. doi:10. 1177/0264619619833723

  44. [44]

    Nabila Jones, Hannah Elizabeth Bartlett, and Richard Cooke. 2019. An analysis of the impact of visual impairment on activities of daily living and vision-related quality of life in a visually impaired adult population.British Journal of Visual Impairment37, 1 (2019), 50–63

  45. [45]

    Amy A Kalia, Gordon E Legge, and Nicholas A Giudice. 2008. Learning building layouts with non-geometric visual information: The effects of visual impairment and age.Perception37, 11 (2008), 1677–1699. SceneGlance ASSETS ’26, October 25–28, 2026, Vila Nova de Gaia, Portugal

  46. [46]

    Avyay Ravi Kashyap. 2020. Behaviors, Problems and Strategies of Visually Impaired Persons During Meal Preparation in the Indian Context: Challenges and Opportunities for Design. InProceedings of the 22nd International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’20). Association for Computing Machinery, New York, NY, USA, Article 101, ...

  47. [47]

    Liam Kettle and Yi-Ching Lee. 2022. Augmented reality for vehicle-driver communication: a systematic review.Safety8, 4 (2022), 84

  48. [48]

    Doaa Khattab, Julie Buelow, and Donna Saccuteli. 2015. Understanding the barriers: Grocery stores and visually impaired shoppers.Journal of accessibility and design for all: JACCES5, 2 (2015), 157–173

  49. [49]

    Daniel Killough, Justin Feng, Zheng Xue Ching, Daniel Wang, Rithvik Dyava, Yapeng Tian, and Yuhang Zhao. 2025. VRSight: An AI-Driven Scene Description System to Improve Virtual Reality Accessibility for Blind People. InProceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST ’25). Association for Computing Machinery, Ne...

  50. [50]

    Benjamin Kommey, Kumbong Herrman, and Ernest Ofosu Addo. 2019. A smart vision based navigation aid for the visually impaired.Asian Journal of Research in Computer Science4, 3 (2019), 1–8

  51. [51]

    Masaki Kuribayashi, Kohei Uehara, Allan Wang, Shigeo Morishima, and Chieko Asakawa. 2025. Wanderguide: Indoor map-less robotic guide for exploration by blind people. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–21

  52. [52]

    Alexandra Kuznetsova, Per B Brockhoff, and Rune HB Christensen. 2017. lmerTest package: tests in linear mixed effects models.Journal of statistical software82 (2017), 1–26

  53. [53]

    MiYoung Kwon, Chaithanya Ramachandra, PremNandhini Satgunam, Bartlett W Mel, Eli Peli, and Bosco S Tjan. 2012. Contour enhancement benefits older adults with simulated central field loss.Optometry and vision science89, 9 (2012), 1374– 1384

  54. [54]

    Cameron Kyle-Davidson, Elizabeth Yue Zhou, Dirk B Walther, Adrian G Bors, and Karla K Evans. 2023. Characterising and dissecting human perception of scene complexity.Cognition231 (2023), 105319

  55. [55]

    Mikko Kytö, Barrett Ens, Thammathip Piumsomboon, Gun A Lee, and Mark Billinghurst. 2018. Pinpointing: Precise head-and eye-based target selection for augmented reality. InProceedings of the 2018 CHI conference on human factors in computing systems. 1–14

  56. [56]

    Florian Lang and Tonja Machulla. 2021. Pressing a button you cannot see: evaluating visual designs to assist persons with low vision through augmented reality. InProceedings of the 27th ACM Symposium on Virtual Reality Software and Technology. 1–10

  57. [57]

    Jaewook Lee, Yang Li, Dylan Bunarto, Eujean Lee, Olivia H Wang, Adrian Rodriguez, Yuhang Zhao, Yapeng Tian, and Jon E Froehlich. 2024. Towards ai-powered ar for enhancing sports playability for people with low vision: An exploration of arsports. In2024 IEEE International Symposium on Mixed and Augmented Reality Adjunct (ISMAR-Adjunct). IEEE, 228–233

  58. [58]

    Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–16

  59. [59]

    Meesung Lee, Heerim Lee, Sungjoo Hwang, and Minji Choi. 2021. Under- standing the impact of the walking environment on pedestrian perception and comprehension of the situation.Journal of Transport & Health23 (2021), 101267

  60. [60]

    Franklin Mingzhe Li, Jamie Dorst, Peter Cederberg, and Patrick Carrington. 2021. Non-Visual Cooking: Exploring Practices and Challenges of Meal Preparation by People with Visual Impairments. InProceedings of the 23rd International ACM SIGACCESS Conference on Computers and Accessibility (ASSETS ’21). Association for Computing Machinery, New York, NY, USA, ...

  61. [61]

    Ke Li and Jitendra Malik. 2016. Amodal instance segmentation. InEuropean Conference on Computer Vision. Springer, 677–693

  62. [62]

    Zhipeng Li, Christoph Gebhardt, Yves Inglin, Nicolas Steck, Paul Streli, and Christian Holz. 2024. Situationadapt: Contextual ui optimization in mixed reality with situation awareness via llm reasoning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–13

  63. [63]

    Haozhe Lin, Jiangtao Gong, Yu Wang, Jinsong Zhang, Bing Bai, Yan Zhang, Luyao Wang, Chenyu Wei, Yancheng Cao, Kun Li, et al . 2025. AI system facilitates people with blindness and low vision in interpreting and experiencing unfamiliar environments.npj Artificial Intelligence1, 1 (2025), 7

  64. [64]

    Lawrence Zitnick, and Piotr Dollár

    Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár

  65. [65]

    David Lindlbauer, Anna Maria Feit, and Otmar Hilliges. 2019. Context-aware online adaptation of mixed reality interfaces. InProceedings of the 32nd annual ACM symposium on user interface software and technology. 147–160

  66. [66]

    Alice Lo Valvo, Daniele Croce, Domenico Garlisi, Fabrizio Giuliano, Laura Giarré, and Ilenia Tinnirello. 2021. A navigation and augmented reality system for visually impaired people.Sensors21, 9 (2021), 3061

  67. [67]

    Gang Luo and Eli Peli. 2006. Use of an augmented-vision device for visual search by patients with tunnel vision.Investigative ophthalmology & visual science47, 9 (2006), 4152–4159

  68. [68]

    Chengqi Lyu, Wenwei Zhang, Haian Huang, Yue Zhou, Yudong Wang, Yanyi Liu, Shilong Zhang, and Kai Chen. 2022. Rtmdet: An empirical study of designing real-time object detectors.arXiv preprint arXiv:2212.07784(2022)

  69. [69]

    Márcio CF Macedo and Antonio L Apolinario. 2021. Occlusion handling in augmented reality: Past, present and future.IEEE Transactions on Visualization and Computer Graphics29, 2 (2021), 1590–1609

  70. [70]

    Sean P MacEvoy and Russell A Epstein. 2011. Constructing scenes from objects in human occipitotemporal cortex.Nature neuroscience14, 10 (2011), 1323–1329

  71. [71]

    Alexey Magay, Dhurba Tripathi, Yu Hao, and Yi Fang. 2024. A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision. InEuropean Conference on Computer Vision. Springer, 323–339

  72. [72]

    Florian Mathis and Johannes Schöning. 2025. LifeInsight: Design and Evaluation of an AI-Powered Assistive Wearable for Blind and Low Vision People Across Multiple Everyday Life Scenarios. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–25

  73. [73]

    2017.Designing experiments and analyzing data: A model comparison perspective

    Scott E Maxwell, Harold D Delaney, and Ken Kelley. 2017.Designing experiments and analyzing data: A model comparison perspective. Routledge

  74. [74]

    Microsoft. 2017. Seeing AI - Talking Camera for the Blind — seeingai.com. https://www.seeingai.com/. [Accessed 09-09-2025]

  75. [75]

    Margrain, Yu-Kun Lai, and Parisa Eslambolchilar

    Hein Min Htike, Tom H. Margrain, Yu-Kun Lai, and Parisa Eslambolchilar. 2021. Augmented reality glasses as an orientation and mobility aid for people with low vision: a feasibility study of experiences and requirements. InProceedings of the 2021 CHI Conference on Human Factors in Computing Systems. 1–15

  76. [76]

    Aliaksei Miniukovich and Antonella De Angeli. 2014. Quantification of interface visual complexity. InProceedings of the 2014 international working conference on advanced visual interfaces. 153–160

  77. [77]

    Pilar Montero. 2005. Nutritional assessment and diet quality of visually impaired Spanish children.Annals of Human Biology32, 4 (2005), 498–512

  78. [78]

    Karin Müller, Christin Engel, Claudia Loitsch, Rainer Stiefelhagen, and Gerhard Weber. 2022. Traveling more independently: a study on the diverse needs and challenges of people with visual or mobility impairments in unfamiliar indoor environments.ACM Transactions on Accessible Computing (TACCESS)15, 2 (2022), 1–44

  79. [79]

    National Eye Institute. 2025. Low Vision | National Eye Institute — nei.nih.gov. https://www.nei.nih.gov/learn-about-eye-health/eye-conditions- and-diseases/low-vision. [Accessed 19-06-2024]

  80. [80]

    Mark B Neider and Gregory J Zelinsky. 2011. Cutting through the clutter: Searching for targets in evolving complex scenes.Journal of Vision11, 14 (2011), 7–7

Showing first 80 references.