Pith. sign in

REVIEW 4 major objections 5 minor 279 references

Characterizing Visual Intents for People with Low Vision through Eye Tracking

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that low vision and sighted viewers share five in-the-moment visual intents—searching, observing, traversing, comparing, exploring—and that each leaves a distinct gaze signature modulated by visual acuity and peripheral…

desk verdict A solid bottom-up taxonomy paper; the qualitative core is new and useful, but the quantitative gaze contrasts are partly confounded by segment length and the labeling procedure. read the letter →

arxiv 2501.14327 v2 pith:O5L5HMUP submitted 2025-01-24 cs.HC

classification cs.HC
keywords eyetrackinglowvisionvisualintentgazebehaviorretrospectivethink-aloudacuityperipherallossassistivetechnology
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that during static image viewing, people's in-the-moment visual goals fall into five shared categories—searching, observing, traversing, comparing, and exploring—and that each category carries a measurable gaze signature. The authors derive this taxonomy from the bottom up using eye tracking with retrospective think-aloud on 20 low vision and 20 sighted participants, rather than starting from pre-defined intents. They show that gaze behavior differs across the five intents, that low vision participants scan more broadly than sighted participants during traversing and exploring, and that lower visual acuity and peripheral vision loss shorten saccades and increase the number of objects visited. If the classification is right, it gives intent-aware assistive technology a first cross-task vocabulary for recognizing what a low vision user is trying to see, and it says plainly that such recognition should combine gaze with visual ability and image context.

What carries the argument

The central object is the visual intent, defined as a distinct gaze pattern that reflects an in-the-moment, meta-level objective and is agnostic to the overall visual task. The procedure that carries the argument is the eye-tracking-based retrospective think-aloud protocol: participants viewed images while their gaze was tracked, then watched a playback of their own gaze trajectory and described what they were doing, and researchers segmented and labeled each trajectory into intent segments, discarding unresolved segments so the final set had complete labeling agreement. On those segments the authors computed fixation duration and rate, saccade amplitude, stationary entropy $H_s = -\sum_i \pi_i \log_2 \pi_i$ over an $8\times5$ grid of areas of interest, number of distinct objects visited, object attention variability, and foreground attention ratio, and compared them across intents, between low vision and sighted groups, and across visual abilities. These comparisons are what bind each intent label to a measurable gaze signature and what connect visual acuity and peripheral vision to those signatures.

What would settle it

Fresh coders blind to the original labels could re-segment the same gaze recordings from the retrospective think-aloud sessions; if their five-intent labels do not agree with the paper's segments at better than chance, the gaze differences attributed to intents are an artifact of the labeling procedure rather than a stable property of visual intent.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a five-category visual intent taxonomy shared by low vision and sighted viewers. Searching is a sequence of fixations aimed at locating a target; Observing concentrates fixations on one object or person to identify identity, details, or activity; Traversing moves fixations across adjacent objects, typically for counting or reading; Comparing shifts fixations back and forth between two or more objects to judge relationships; Exploring spreads fixations widely to gather overall context. The quantitative analyses tie these categories to distinct gaze signatures: observing shows longer fixations, shorter saccades, lower stationary entropy, and more uneven object attention; searching and exploring spend more time on background; comparing has longer saccades than traversing and more uneven attention. The paper further reports that low vision participants scanned more broadly during traversing and exploring, and that low visual acuity and peripheral vision loss were associated with shorter saccades and a larger number of objects visited, which the authors read as compensatory scanning. It also documents low-vision-specific gaze behaviors outside the five intents, such as briefly fixating a high-contrast object to restore confidence in color perception and confusion-driven irregular revisits after misidentifying an object.

Load-bearing premise

The load-bearing premise is that the intent labels placed on gaze segments from the retrospective think-aloud procedure are valid ground truth; if those labels do not correspond to stable mental categories, the gaze differences reported between intents could be an artifact of labeling and segmentation.

Editorial extensions

If this is right

  • An assistive system could use the taxonomy to choose support on the fly: magnify the fixated object during observing, highlight the relevant objects during comparing, and preview content beyond the visual field when a user with peripheral vision loss scans toward the boundary.
  • Visual intent recognition models for low vision users should be trained with visual ability and image context as inputs, because searching, traversing, and exploring are not reliably separated by gaze metrics alone.
  • The retrospective think-aloud protocol is feasible with low vision users but is too time-intensive and produces too variable segment lengths to build large training sets; standardized single-intent trials will be needed for model training.
  • The taxonomy is a starting point for static 2D viewing without assistive tools; dynamic content, real-world 3D scenes, screen magnifiers, and contrast enhancements are likely to introduce new intents and gaze behaviors such as smooth pursuit.
  • Gaze-based support must be personalized to the user's visual profile: peripheral vision loss and low visual acuity change how many objects a person visits and how far their saccades travel, so sighted gaze norms are not a safe default.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next experiment is to recruit balanced groups with acuity loss only and peripheral field loss only; because the paper's low vision groups overlap on both conditions, the separate effect of each visual ability remains provisional.
  • The 'palette cleanser' behavior—briefly fixating a high-contrast object to restore confidence in color perception—suggests a testable assistive design: detect gaze returning to a high-contrast region during a low-contrast judgment and offer color or contrast enhancement at that moment.
  • If the taxonomy is stable, a held-out classifier using gaze plus image context should label new low vision users' gaze segments with the five intent labels; the paper does not build such a model, so an accurate classifier would be the strongest confirmation of the taxonomy's usefulness.
  • The difficulty of separating searching, traversing, and exploring on low-level gaze measures alone suggests these categories may be graded or context-dependent rather than discrete, and a future recognition system might score them as soft labels instead of exclusive states.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a retrospective think-aloud eye-tracking study with 20 low vision and 20 sighted participants who viewed static images while answering questions at three information levels. From participants' verbal reflections and gaze replays, the authors derive a taxonomy of five visual intents—searching, observing, traversing, comparing, and exploring—shared by both groups, and they report low-vision-specific behaviors such as color-perception recalibration and visual confusion. Using fixation, saccade, spatial-entropy, and object-level gaze measures, the manuscript compares gaze behavior across intents, between low vision and sighted groups, and across visual acuity and peripheral vision subgroups. The central claim is that these five intents are associated with distinct, measurable gaze patterns and that visual abilities modulate those patterns, laying a foundation for intent-aware low vision assistive technology.

Significance. If the central claim holds, this is the first bottom-up, cross-task visual intent taxonomy for low vision users, and the paper's mixed-methods design is a real strength: the qualitative taxonomy is grounded in participants' own descriptions rather than imposed solely from gaze statistics, the sample is diverse in visual conditions, and the authors take care with calibration, gaze-data quality, and object-level annotations. The paper also makes concrete, falsifiable predictions about which gaze measures differ across intents and between ability groups. However, the quantitative corroboration in Sections 4.3–4.5 currently rests on segment labels whose construction and selection are not fully reported, and on gaze measures that may be sensitive to segment duration; those issues are fixable and are the main barrier to accepting the quantitative claims as they stand.

major comments (4)
  1. [Sections 3.3, 3.4.2, 4.3–4.5] The headline quantitative contrasts—especially stationary entropy and number of objects visited—are potentially confounded by segment duration. In the ETRTA procedure, segment boundaries are set by the researcher while listening to retrospective descriptions and refined with the participant, and Section 5.1 states that there is "no control over how many segments they produce or how long each segment is." If observing segments are systematically shorter (for example, single-object identification episodes), they will trivially have lower stationary entropy and fewer visited objects regardless of intent. None of the LME/ART models in Sections 4.3–4.5 includes segment duration or fixation count as a covariate. Please report the distribution of segment duration and fixation count per intent, and re-run the entropy and object-count analyses with duration control or duration-matched segment comparison. If the observing effect is reduced or disappears, the qualitative taxonomy can still stand, but the quantitative corroboration would need to be reframed accordingly.
  2. [Section 3.4.1] The validity of the intent labels is asserted through full agreement, but unresolved gaze segments were discarded to achieve complete agreement, and the paper does not report how many segments were discarded, whether discarding was systematic with respect to intent or participant group, or inter-coder agreement on the final segment labels. The reported Cohen's Kappa of 0.73 is for the initial codebook on three sample transcripts, not for the final labeling of all 510 segments. Because the quantitative analyses in Sections 4.3–4.5 inherit these labels, selection bias in segment labeling could produce the reported gaze differences. Please report the number and proportion of discarded segments, a reliability analysis of the final segment labels, and a sensitivity analysis that includes or models the unresolved segments.
  3. [Sections 4.3–4.5] The statistical inference is based on many correlated gaze measures and numerous post-hoc contrasts, but the paper does not state the total number of tests performed or apply a family-wise or false-discovery-rate correction across measures. At least one headline result is only a trend (fixation rate, p=0.067), and the effect-size confidence intervals reported for the LME results (e.g., eta-squared CI [0.30, 1.0]) are implausible as printed and need correction. Please clarify how many hypotheses were tested, apply an appropriate correction at the measure level, and correct the effect-size reporting.
  4. [Sections 4.5 and 5.5] The visual ability analysis is weakened by overlap between the low-acuity and peripheral-vision-loss groups: Section 5.5 states that six participants had both conditions. Section 4.5 nevertheless reports main effects of visual acuity and peripheral vision without reporting the cell sizes of the 2x2 ability grouping or a sensitivity analysis that excludes the overlapping group. Given the modest sample size (20 low vision participants), these effects should be interpreted with more caution, and the paper should either report the 2x2 cell sizes or explicitly model the overlap to show that the saccade-amplitude and object-count effects are not driven by the same six participants.
minor comments (5)
  1. [Section 2.3] "Stardgart's disease" should be "Stargardt's disease."
  2. [Section 3.4.4] The 20/100 visual acuity threshold is introduced without justification; since it is used to split the low vision group, please cite the prior source more clearly and, if possible, report sensitivity to the threshold.
  3. [Section 4.2, Figure 4] The figure caption and the text describe examples with participant IDs, but several panels (e.g., 4j) are not explicitly referenced in the prose; please ensure every panel is mentioned or state that some panels illustrate an intent collectively.
  4. [Overall] No data or analysis-code availability statement is included. Given the central quantitative claims, making at least aggregate data and analysis scripts available would substantially strengthen reproducibility.
  5. [Section 4.1] The phrase "The mean angular errors from the 5-dot validation was..." should be "The mean angular error... was..." for subject-verb agreement.

Circularity Check

2 steps flagged · score 6.0 of 10

Partial self-definitional circularity: several headline gaze contrasts restate the taxonomy's own definitions, and the method's validity is supported by measures derived from that same method.

  1. self definitional [Section 4.2 (Visual Intent Taxonomy) and Section 4.3 (Gaze Behavior under Different Visual Intents)]
    "Observing is defined as a sequence of fixations primarily concentrated on a single object... Exploring is characterized by widely distributed fixations across the entire image... observing had significantly lower entropy than searching, traversing, comparing, and exploring. This suggested that fixations during observing were more concentrated... participants visited significantly fewer objects during observing than during searching, traversing, comparing, and exploring."

    The taxonomy's definitions already assert the spatial gaze properties that Section 4.3 reports as empirical findings. Observing is defined by concentrated fixations on one object, and Exploring is defined by widely distributed fixations across the whole image, so the stationary-entropy and number-of-objects contrasts follow from the coding scheme by construction. Similarly, Comparing is defined as fixations shifting back and forth between objects, which entails longer saccades, and Traversing as fixations across adjacent objects, which entails broad object coverage. These quantitative results are therefore partly restatements of the definitions rather than independent confirmations that the intents have distinct gaze patterns.

  2. other [Section 5.1, Gaze Data Collection for Model Training]
    "The validity of our method was further supported by our quantitative findings, where visual intent had significant effects on multiple gaze measures."

    This sentence validates the ETRTA method using quantitative findings that were themselves produced by ETRTA-based segmentation and labeling. Because several of those quantitative effects are partly entailed by the intent definitions used during labeling, the validation loop closes on the method's own output rather than on an external benchmark. The qualitative observation that participants could verbalize their gaze trajectories is independent evidence for feasibility, but the cited quantitative support is not independent of the method being validated.

full rationale

The paper's qualitative taxonomy is genuinely bottom-up: the five intents were derived from participants' retrospective think-aloud descriptions and open coding, not from the quantitative gaze measures. However, the definitions of the intents are themselves stated in gaze terms (concentrated vs. distributed fixations, back-and-forth shifts, adjacent-object scanning), so the headline contrasts in stationary entropy, number of objects visited, and saccade amplitude partially reduce to those definitions. The between-group and visual-ability findings (low vision vs. sighted, visual acuity and peripheral vision effects) are not circular because they compare groups under the same intent labels rather than re-deriving the labels. The self-citations to the authors' prior work [94, 95] are used for calibration procedures, a visual-acuity threshold, and consistency with prior reading findings; they are not load-bearing for the taxonomy itself. The paper's own statement in Section 5.1 that the method's validity is supported by those quantitative findings adds a circular validation step, but it is partial: participant verbal reports and the qualitative codebook provide independent grounding. Overall, the core taxonomy is not forced by its inputs, but several quantitative 'characterizations' are definitional restatements, meriting a partial circularity score.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

No numbers are fitted to data. The central analyses depend on subjective intent labels, visual ability grouping choices, the 20/100 acuity split, and stimulus design assumptions; these are listed as axioms rather than free parameters.

free parameters (1)
  • Visual acuity split at 20/100 = 20/100 in better eye
    Grouping cutoff for low vs high visual acuity, adopted from prior work [94] and applied in Section 3.4.4; not fitted to the current data, but a hand-chosen threshold affecting all acuity comparisons.
assumptions (4)
  • domain assumption Participants' retrospective verbal reports accurately reflect their visual intents during image viewing.
    The taxonomy in Section 4.2 rests on participants' reflections during gaze playback; no independent behavioral validation is provided.
  • domain assumption Each gaze trajectory segment has a single dominant visual intent.
    Section 3.4.1 labels each segment with the closest intent and discards unresolved segments to force complete agreement.
  • domain assumption The simplified visual field test and self-report accurately classify peripheral vision as limited or intact.
    Sections 3.2.1 and 3.4.4 use these to define the PeripheralVision factor for all low vision analyses.
  • domain assumption The 117-image stimulus set and three-level information questions elicit gaze behaviors representative of everyday image viewing.
    Section 3.2.2 selects images from five daily contexts and designs questions by information level; this is assumed to approximate real-world visual tasks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Characterizing Visual Intents for People with Low Vision through Eye Tracking." pith.science (2026). https://pith.science/paper/O5L5HMUP

@misc{pith2026250114327,
  author       = {Pith},
  title        = {Pith review of: Characterizing Visual Intents for People with Low Vision through Eye Tracking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O5L5HMUP}},
  note         = {Machine review of arXiv:2501.14327}
}
read the original abstract

Accessing visual information is crucial yet challenging for people with low vision due to visual conditions like low visual acuity and limited visual fields. However, unlike blind people, low vision people have and prefer using their functional vision in daily tasks. Gaze patterns thus become an important indicator to uncover their visual challenges and intents, inspiring more adaptive visual support. We seek to deeply understand low vision users' gaze behaviors in different image-viewing tasks, characterizing typical visual intents and the unique gaze patterns exhibited by people with different low vision conditions. We conducted a retrospective think-aloud study using eye tracking with 20 low vision participants and 20 sighted controls. Participants completed various image-viewing tasks and watched the playback of their gaze trajectories to reflect on their visual experiences. Based on the study, we derived a visual intent taxonomy with five visual intents characterized by participants' gaze behaviors. We demonstrated the difference between low vision and sighted participants' gaze behaviors and how visual ability affected low vision participants' gaze patterns across visual intents. Our findings underscore the importance of combining visual ability information, visual context, and eye tracking data in visual intent recognition, setting up a foundation for intent-aware assistive technologies for low vision people.

Figures

Figures reproduced from arXiv: 2501.14327 by the authors.

Figure 1
Figure 1. Overview of five visual intents—Searching, Observing, Traversing, Comparing, and Exploring—identified through retrospective think-aloud study with both low vision and sighted participants using eye tracking. Each panel shows an illustration of the same example image, overlaid with a representative gaze trajectory for one visual intent. Gaze trajectories are color-coded to show progression (from red to yellow), and c… view at source ↗
Figure 2
Figure 2. Our study set up and gaze data collection & annotation system: (a) The positioning of participant screen S1 and [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Our gaze calibration & validation interface and visual field test interface: (a) 14-dot calibration. (b) 5-dot validation. (c) [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Examples of different visual intents exhibited by low vision participants. Gaze trajectory overlay on the images are [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Examples of gaze behaviors demonstrated by our participants, influenced by visual intent, visual condition (low vision [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

279 extracted references · 75 canonical work pages

  1. [1]

    Sonia J Ahn and Gordon E Ledge. 1995. Psychophysics of reading—XIII. Pre- dictors of magnifier-aided reading speed in low vision. Vision Research 35, 13 (1995), 1931–1938

  2. [2]

    Concetta F Alberti and Peter Bex. 2017. Do oculomotor adaptations to a vol- ume scotoma provide functional benefits for binocular vision? Investigative Ophthalmology & Visual Science 58, 8 (2017), 4695–4695

  3. [3]

    Apple. 2022. How to zoom in or out on Mac. Available online at: https: //support.apple.com/en-us/HT210978, last accessed on 8/23/2022

  4. [4]

    Apple. 2022. How to zoom in or out on Mac. Available online at: https://support. apple.com/guide/iphone/zoom-iph3e2e367e/ios, last accessed on 8/24/2023

  5. [5]

    Daniel S Asfaw, Pete R Jones, Vera M Mönter, Nicholas D Smith, and David P Crabb. 2018. Does glaucoma alter eye movements when viewing images of natural scenes? A between-eye study. Investigative ophthalmology & visual science 59, 8 (2018), 3189–3198

  6. [6]

    Roman Bednarik, Hana Vrzakova, and Michal Hradis. 2012. What do you want to do next: a novel approach for intent prediction in gaze-based interaction. In Proceedings of the symposium on eye tracking research and applications . 83–90

  7. [7]

    Kenan Bektaş, Jannis Strecker, Simon Mayer, Dr Kimberly Garcia, Jonas Her- mann, Kay Erik Jenß, Yasmine Sheila Antille, and Marc Solèr. 2023. GEAR: Gaze-enabled augmented reality for human activity recognition. In Proceedings of the 2023 symposium on eye tracking research and applications . 1–9

  8. [8]

    Dmitrii Boiarshinov, Jaime Guajardo, and Gabi Lanning. 2024. Providing LLM- generated Point of Interest Description Based on Gaze Tracking. (2024)

Show all 279 references
  1. [9]

    Avishai Ceder. 1977. Drivers’ eye movements as related to attention in simulated traffic flow conditions. Human factors 19, 6 (1977), 571–581

  2. [10]

    Ruijia Chen, Junru Jiang, Pragati Maheshwary, Brianna R Cochran, and Yuhang Zhao. 2025. VisiMark: Characterizing and Augmenting Landmarks for People with Low Vision in Augmented Reality to Support Indoor Navigation. arXiv preprint arXiv:2502.10561 (2025)

  3. [11]

    Allen MY Cheong, Gordon E Legge, Mary G Lawrence, Sing-Hang Cheung, and Mary A Ruff. 2007. Relationship between slow visual processing and reading speed in people with macular degeneration. Vision research 47, 23 (2007), 2943– 2955

  4. [12]

    Hwayoung Cho, Dakota Powell, Adrienne Pichon, Lisa M Kuhns, Robert Garo- falo, and Rebecca Schnall. 2019. Eye-tracking retrospective think-aloud as a novel approach for a usability evaluation. International journal of medical informatics 129 (2019), 366–373

  5. [13]

    Anustup Choudhury and Gérard Medioni. 2010. Color contrast enhancement for visually impaired people. In 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition-Workshops. IEEE, 33–40

  6. [14]

    Jacob Cohen. 2013. Statistical power analysis for the behavioral sciences . rout- ledge

  7. [15]

    Hélio Clemente Cuve, Jelka Stojanov, Xavier Roberts-Gaal, Caroline Catmur, and Geoffrey Bird. 2022. Validation of Gazepoint low-cost eye-tracking and psychophysiology bundle. Behavior research methods 54, 2 (2022), 1027–1049

  8. [16]

    Asim H Dar, Adina S Wagner, and Michael Hanke. 2021. REMoDNaV: robust eye-movement classification for dynamic stimulation.Behavior research methods 53, 1 (2021), 399–414

  9. [17]

    Brendan David-John, Candace Peacock, Ting Zhang, T Scott Murdison, Hrvoje Benko, and Tanya R Jonker. 2021. Towards gaze-based prediction of the intent to interact in virtual reality. In ACM symposium on eye tracking research and applications. 1–7

  10. [18]

    Ashley D Deemer, Christopher K Bradley, Nicole C Ross, Danielle M Natale, Rath Itthipanichpong, Frank S Werblin, and Robert W Massof. 2018. Low vision enhancement with head-mounted video display systems: are we there yet? Optometry and Vision Science 95, 9 (2018), 694–703. Cha...

  11. [19]

    Brad Dwyer, Joseph Nelson, J Solawetz, J Warner, A Smith, M Johnson, K Lee, R Brown, L Clark, T Harris, et al. 2024. Roboflow (Version 1.0)[Software](2022). Computer vision platform (2024)

  12. [20]

    Lisa A Elkin, Matthew Kay, James J Higgins, and Jacob O Wobbrock. 2021. An aligned rank transform procedure for multifactor contrast tests. In The 34th annual ACM symposium on user interface software and technology . 754–768

  13. [21]

    Frederick L Ferris III, Aaron Kassoff, George H Bresnick, and Ian Bailey. 1982. New visual acuity charts for clinical research.American journal of ophthalmology 94, 1 (1982), 91–96

  14. [22]

    Dylan R Fox, Ahmad Ahmadzada, Clara T Friedman, Shiri Azenkot, Marlena A Chu, Roberto Manduchi, and Emily A Cooper. 2023. Using augmented reality to cue obstacles for people with low vision. Optics Express 31, 4 (2023), 6827–6848

  15. [23]

    Shonraj Ballae Ganeshrao, Amina Jaleel, Srija Madicharla, Vanga Kavya Sri, Juwariah Zakir, Chandra S Garudadri, and Sirisha Senthil. 2021. Comparison of saccadic eye movements among the high-tension glaucoma, primary angle- closure glaucoma, and normal-tension glaucoma. Journa...

  16. [24]

    Clémentine Garric, Jean-François Rouland, and Quentin Lenoble. 2021. Glau- coma and computer use: do contrast and color enhancements improve visual comfort in patients? Ophthalmology Glaucoma 4, 5 (2021), 531–540

  17. [25]

    Wilson S Geisler and Lawrence K Cormack. 2011. Models of overt attention. The Oxford handbook of eye movements (2011), 439–454

  18. [26]

    Giovanni Giacomelli, Alessandro Farini, Ilaria Baldini, Marco Raffaelli, Giulia Bigagli, Alessandro Fossetti, and Gianni Virgili. 2021. Saccadic movements assessment in eccentric fixation: A study in patients with Stargardt disease. European Journal of Ophthalmology 31, 5 (202...

  19. [27]

    Ricardo E Gonzalez Penuela, Jazmin Collins, Cynthia Bennett, and Shiri Azenkot

  20. [28]

    Howard Greisdorf and Brian O’Connor. 2002. Modelling what users see when they look at images: a cognitive viewpoint.Journal of documentation 58, 1 (2002), 6–29

  21. [29]

    Miguel Grinberg. 2018. Flask-SocketIO documentation. Available online at: https://flask-socketio.readthedocs.io/en/latest/, last accessed on 9/10/2023

  22. [30]

    Zhiwei Guan, Shirley Lee, Elisabeth Cuddihy, and Judith Ramey. 2006. The validity of the stimulated retrospective think-aloud method as measured by eye tracking. In Proceedings of the SIGCHI conference on Human Factors in computing systems. 1253–1262

  23. [31]

    Elyse C Hallett. 2015. Reading without bounds: How different magnification methods affect the performance of students with low vision . California State University, Long Beach

  24. [32]

    Elyse C Hallett, Wayne Dick, Tom Jewett, and Kim-Phuong L Vu. 2018. How screen magnification with and without word-wrapping affects the user expe- rience of adults with low vision. In Advances in Usability and User Experience: Proceedings of the AHFE 2017 International Confere...

  25. [33]

    Seongsil Heo, Roberto Manduchi, and Suzana Chung. 2024. Reading with Screen Magnification: Eye Movement Analysis Using Compensated Gaze Tracks. In Proceedings of the 2024 Symposium on Eye Tracking Research and Applications . 1–6

  26. [34]

    Jutta Hild, Michael Voit, Christian Kühnle, and Jürgen Beyerer. 2018. Predicting observer’s task from eye movement patterns during motion image analysis. In Proceedings of the 2018 ACM symposium on eye tracking research & applications . 1–5

  27. [35]

    Bill Holton. 2014. A review of iOS access for all: Your comprehensive guide to accessibility for iPad, iPhone, and iPod touch, by Shelly Brisbin. AccessWorld Magazine 15, 7 (2014)

  28. [36]

    Zhiming Hu, Andreas Bulling, Sheng Li, and Guoping Wang. 2021. Ehtask: Recognizing user tasks from eye and head movements in immersive virtual reality. IEEE Transactions on Visualization and Computer Graphics 29, 4 (2021), 1992–2004

  29. [37]

    Chien-Ming Huang, Sean Andrist, Allison Sauppé, and Bilge Mutlu. 2015. Using gaze patterns to predict task intent in collaboration. Frontiers in psychology 6 (2015), 1049

  30. [38]

    Jonathan Huang, Max Kinateder, Matt J Dunn, Wojciech Jarosz, Xing-Dong Yang, and Emily A Cooper. 2019. An augmented reality sign-reading assistant for users with reduced vision. PloS one 14, 1 (2019), e0210630

  31. [39]

    Taewoo Jo, Dohyeon Yeo, Gwangbin Kim, Seokhyun Hwang, and SeungJun Kim

  32. [40]

    Melanie Kellar, Carolyn Watters, and Michael Shepherd. 2006. A goal-based classification of web information tasks. Proceedings of the American Society for Information Science and Technology 43, 1 (2006), 1–22

  33. [41]

    Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–32

    WatchCap: Improving Scanning Efficiency in People with Low Vision through Compensatory Head Movement Stimulation. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 8, 2 (2024), 1–32

  34. [42]

    Max Kinateder, Justin Gualtieri, Matt J Dunn, Wojciech Jarosz, Xing-Dong Yang, and Emily A Cooper. 2018. Using an Augmented Reality Device as a Distance- based Vision Aid—Promise and Limitations. Optometry and Vision Science 95, 9 (2018), 727

  35. [43]

    Peter Kiefer, Ioannis Giannopoulos, and Martin Raubal. 2013. Using eye move- ments to recognize activities on cartographic maps. In Proceedings of the 21st ACM SIGSPATIAL International Conference on Advances in Geographic Informa- tion Systems. 488–491

  36. [44]

    Alexandra Kuznetsova, Per B Brockhoff, and Rune Haubo Bojesen Christensen

  37. [45]

    Manu Kumar, Jeff Klingner, Rohan Puranik, Terry Winograd, and Andreas Paepcke. 2008. Improving the accuracy of gaze input for interaction. In Proceed- ings of the 2008 symposium on Eye tracking research & applications . 65–68

  38. [46]

    Florian Lang and Tonja Machulla. 2021. Pressing a button you cannot see: evaluating visual designs to assist persons with low vision through augmented reality. In Proceedings of the 27th ACM Symposium on Virtual Reality Software and Technology. 1–10

  39. [47]

    Tobias Langlotz, Jonathan Sutton, Stefanie Zollmann, Yuta Itoh, and Holger Regenbrecht. 2018. ChromaGlasses: Computational glasses for compensating colour blindness. In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1–12

  40. [48]

    Guohao Lan, Tim Scargill, and Maria Gorlatova. 2022. Eyesyn: Psychology- inspired eye movement synthesis for gaze-based activity recognition. In2022 21st ACM/IEEE international conference on information processing in sensor networks (IPSN). IEEE, 233–246

  41. [49]

    Susan J Leat, Gordon E Legge, and Mark A Bullimore. 1999. What is low vision? A re-evaluation of definitions. Optometry and Vision Science 76, 4 (1999), 198–211

  42. [50]

    Jaewook Lee, Andrew D Tjahjadi, Jiho Kim, Junpu Yu, Minji Park, Jiawen Zhang, Jon E Froehlich, Yapeng Tian, and Yuhang Zhao. 2024. CookAR: Affordance Augmentations in Wearable AR to Support Kitchen Tool Interactions for People with Low Vision. In Proceedings of the 37th Annual...

  43. [51]

    Augustinus Laude, Damon WK Wong, Ai Ping Yow, Muthu Mookiah, and Tock H Lim. 2018. Eye gaze tracking and its relationship with visual acuity, central visual field and age-related macular degeneration features.Investigative Ophthalmology & Visual Science 59, 9 (2018), 1264–1264

  44. [52]

    Samantha Sze-Yee Lee, Alex A Black, and Joanne M Wood. 2017. Effect of glaucoma on eye movement patterns and laboratory-based hazard detection ability. PloS one 12, 6 (2017), e0178876

  45. [53]

    Samantha Sze-Yee Lee, Alex A Black, and Joanne M Wood. 2019. Eye movements of drivers with glaucoma on a visual recognition slide test.Optometry and Vision Science 96, 7 (2019), 484–491

  46. [54]

    Susanna MY Lee and Joseph CW Cho. 2007. Low vision devices for children. Community Eye Health 20, 62 (2007), 28

  47. [55]

    Firas Lethaus, Martin RK Baumann, Frank Köster, and Karsten Lemmer. 2013. A comparison of selected simple supervised learning algorithms to predict driver intent based on gaze data. Neurocomputing 121 (2013), 108–130

  48. [56]

    Franklin Mingzhe Li, Jamie Dorst, Peter Cederberg, and Patrick Carrington. 2021. Non-visual cooking: exploring practices and challenges of meal preparation by people with visual impairments. In Proceedings of the 23rd International ACM SIGACCESS Conference on Computers and Acc...

  49. [57]

    Gordon E Legge, Gary S Rubin, Denis G Pelli, and Mary M Schleske. 1985. Psychophysics of reading—II. Low vision. Vision research 25, 2 (1985), 253–265

  50. [58]

    Gang Luo and Eli Peli. 2006. Use of an augmented-vision device for visual search by patients with Tunnel Vision. Investigative Opthalmology & Visual Science 47, 9 (2006), 4152. doi:10.1167/iovs.05-1672

  51. [59]

    Bhanuka Mahanama, Yasith Jayawardana, Sundararaman Rengarajan, Gavindya Jayawardena, Leanne Chukoskie, Joseph Snider, and Sampath Jayarathna. 2022. Eye movement and pupil measures: A review. frontiers in Computer Science 3 (2022), 733531

  52. [60]

    Franklin Mingzhe Li, Michael Xieyang Liu, Shaun K Kane, and Patrick Carring- ton. 2024. A Contextual Inquiry of People with Vision Impairments in Cooking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Sys- tems. 1–14

  53. [61]

    Sandra Malpica, Daniel Martin, Ana Serrano, Diego Gutierrez, and Belen Masia

  54. [62]

    Sina Masnadi, Brian Williamson, Andrés N Vargas González, and Joseph J LaViola. 2020. Vriassist: An eye-tracked virtual reality low vision assistance tool. In 2020 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). IEEE, 808–809

  55. [63]

    Bertram F Malle and Joshua Knobe. 1997. The folk concept of intentionality. Journal of experimental social psychology 33, 2 (1997), 101–121

  56. [64]

    James G May, Robert S Kennedy, Mary C Williams, William P Dunlap, and Julie R Brannan. 1990. Eye movement indices of mental workload. Acta psychologica 75, 1 (1990), 75–89

  57. [65]

    Meta. 2022. React - A JavaScript library for building user interfaces. Available online at: https://reactjs.org, last accessed on 9/7/2022

  58. [66]

    Microsoft. 2022. Use Magnifier to make things on the screen easier to see. Available online at: https://support.microsoft.com/en-us/windows/use- magnifier-to-make-things-on-the-screen-easier-to-see-414948ba-8b1c-d3bd- 8615-0e5e32204198#WindowsVersion=Windows_11, last accessed ...

  59. [67]

    Natalie Maus, Dalton Rutledge, Sedeeq Al-Khazraji, Reynold Bailey, Cecilia Oves- dotter Alm, and Kristen Shinohara. 2020. Gaze-guided magnification for individ- uals with vision impairments. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems...

  60. [68]

    NIH. 2020. Low Vision - National Eye Institute. Available on- line at: https://www.nei.nih.gov/learn-about-eye-health/eye-conditions-and- diseases/low-vision, last accessed on 1/10/2025

  61. [69]

    OpenAI. 2022. Introducing Whisper. Available online at: https://openai.com/ index/whisper/, last accessed on 1/17/2024

  62. [70]

    Delfina Sol Martinez Pandiani and Valentina Presutti. 2023. Seeing the Intangible: Survey of Image Classification into High-Level and Abstract Categories. arXiv preprint arXiv:2308.10562 (2023)

  63. [71]

    Mala D Naraine and Peter H Lindsay. 2011. Social inclusion of employees who are blind or low vision. Disability & Society 26, 4 (2011), 389–403

  64. [72]

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenh...

  65. [73]

    Keith Rayner. 1998. Eye movements in reading and information processing: 20 years of research. Psychological bulletin 124, 3 (1998), 372

  66. [74]

    ReBokeh. 2025. ReBokeh: Home. Available online at: https://https://www. rebokeh.com/, last accessed on 4/16/2025

  67. [75]

    Pradeep Y Ramulu, Bonnielin K Swenor, Joan L Jefferys, David S Friedman, and Gary S Rubin. 2013. Difficulty with out-loud and silent reading in glaucoma. Investigative ophthalmology & visual science 54, 1 (2013), 666–672

  68. [76]

    Jun Rekimoto. 2025. GazeLLM: Multimodal LLMs incorporating Human Visual Attention. arXiv preprint arXiv:2504.00221 (2025)

  69. [77]

    Ligao Ruan, Giles Hamilton-Fletcher, Mahya Beheshti, Todd E Hudson, Maurizio Porfiri, and JR Rizzo. 2024. Multi-faceted Sensory Substitution for Curb Alerting: A Pilot Investigation in Persons with Blindness and Low Vision. arXiv preprint arXiv:2408.14578 (2024)

  70. [78]

    Johnny Saldaña. 2021. The coding manual for qualitative researchers. (2021)

  71. [79]

    G Rees, CL Saw, EL Lamoureux, and JE Keeffe. 2007. Self-management programs for adults with low vision: needs and challenges.Patient education and counseling 69, 1-3 (2007), 39–46

  72. [80]

    Thorsten Schwarz, Arsalan Akbarioroumieh, Giuseppe Melfi, and Rainer Stiefel- hagen. 2020. Developing a magnification prototype based on head and eye- tracking for persons with low vision. In Computers Helping People with Special Needs: 17th International Conference, ICCHP 202...

  73. [81]

    Cassia Senger, Mirella Aparecida Oliveira, C Gustavo De Moraes, Andre Messias, Jayter Silva Paula, and Raquel pantojo Souza. 2020. Saccadic movements during an exploratory visual search task in patients with glaucomatous visual field loss. Investigative Ophthalmology & Visual ...

  74. [82]

    Abigale Stangl, Nitin Verma, Kenneth R Fleischmann, Meredith Ringel Morris, and Danna Gurari. 2021. Going beyond one-size-fits-all image descriptions to satisfy the information wants of people who are blind or have low vision. In Proceedings of the 23rd International ACM SIGAC...

  75. [83]

    Hosnieh Sattar, Mario Fritz, and Andreas Bulling. 2020. Deep gaze pooling: Inferring and visually decoding search intents from human gaze fixations. Neu- rocomputing 387 (2020), 369–382

  76. [84]

    Sarit Szpiro, Yuhang Zhao, and Shiri Azenkot. 2016. Finding a store, searching for a product: a study of daily challenges of low vision people. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing. 61–72

  77. [85]

    Sarit Felicia Anais Szpiro, Shafeka Hashash, Yuhang Zhao, and Shiri Azenkot

  78. [86]

    Meini Tang, Roberto Manduchi, Susana Chung, and Raquel Prado. 2023. Screen Magnification for Readers with Low Vision: A Study on Usability and Perfor- mance. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–15

  79. [87]

    Lee Stearns, Leah Findlater, and Jon E Froehlich. 2018. Design of an augmented reality magnification aid for low vision users. In Proceedings of the 20th interna- tional ACM SIGACCESS conference on computers and accessibility . 28–39

  80. [88]

    Rachel Thomas, Lucy Barker, Gary Rubin, and Annegret Dahlmann-Noor. 2015. Assistive technology for children and young people with low vision. Cochrane database of systematic reviews 6 (2015)

  81. [89]

    Tobii. 2022. Tobii Pro SDK. Available online at: https://www.tobiipro.com/ product-listing/tobii-pro-sdk/, last accessed on 9/7/2022

  82. [90]

    Bruce Tsuji1, Gitte Lindgaard1, and Avi Parush1. 2005. Landmarks for navigators who are visually impaired. (2005)

  83. [91]

    Samuel Tuhkanen, Jami Pekkanen, Richard M Wilkie, and Otto Lappi. 2021. Visual anticipation of the future path: Predictive gaze and steering. Journal of Vision 21, 8 (2021), 25–25

  84. [92]

    Enrico Tanuwidjaja, Derek Huynh, Kirsten Koa, Calvin Nguyen, Churen Shao, Patrick Torbett, Colleen Emmenegger, and Nadir Weibel. 2014. Chroma: a wearable augmented-reality solution for color blindness. In Proceedings of the 2014 ACM international joint conference on pervasive ...

  85. [93]

    Gianni Virgili, Ruthy Acosta, Sharon A Bentley, Giovanni Giacomelli, Claire Allcock, and Jennifer R Evans. 2018. Reading aids for adults with low vision. Cochrane Database of Systematic Reviews 4 (2018)

  86. [94]

    Ru Wang, Zach Potter, Yun Ho, Daniel Killough, Linxiu Zeng, Sanbrita Mon- dal, and Yuhang Zhao. 2024. GazePrompt: Enhancing Low Vision People’s Reading Experience with Gaze-Aware Augmentations. In Proceedings of the CHI Conference on Human Factors in Computing Systems . 1–17

  87. [95]

    Ru Wang, Linxiu Zeng, Xinyong Zhang, Sanbrita Mondal, and Yuhang Zhao

  88. [96]

    Ru Wang, Nihan Zhou, Tam Nguyen, Sanbrita Mondal, Bilge Mutlu, and Yuhang Zhao. 2023. Characterizing barriers and technology needs in the kitchen for blind and low vision people. arXiv preprint arXiv:2310.05396 (2023)

  89. [97]

    Preeti Verghese, Cécile Vullings, and Natela Shanidze. 2021. Eye movements in macular degeneration. Annual review of vision science 7, 1 (2021), 773–791

  90. [98]

    Zhimin Wang and Feng Lu. 2024. Tasks Reflected in the Eyes: Egocentric Gaze- Aware Visual Task Type Recognition in Virtual Reality. IEEE Transactions on Visualization and Computer Graphics (2024)

  91. [99]

    WHO. 2023. Blindness and vision impairment. Available online at: https://www. who.int/news-room/fact-sheets/detail/blindness-and-visual-impairment, last accessed on 1/10/2025

  92. [100]

    Alfred L Yarbus. 2013. Eye movements and vision . Springer

  93. [101]

    In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Understanding how low vision people read using eye tracking. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems . 1–17

  94. [102]

    Yuhang Zhao, Edward Cutrell, Christian Holz, Meredith Ringel Morris, Eyal Ofek, and Andrew D. Wilson. 2019. SeeingVR: A Set of Tools to Make Virtual Reality More Accessible to People with Low Vision. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Syste...

  95. [103]

    Ru Wang, Nihan Zhou, Tam Nguyen, Sanbrita Mondal, Bilge Mutlu, and Yuhang Zhao. 2023. Practices and Barriers of Cooking Training for Blind and Low Vision People. In Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility. 1–5

  96. [104]

    Yuhang Zhao, Elizabeth Kupferstein, Hathaitorn Rojnirun, Leah Findlater, and Shiri Azenkot. 2020. The effectiveness of visual and audio wayfinding guidance on smartglasses for people with low vision. In Proceedings of the 2020 CHI conference on human factors in computing syste...

  97. [105]

    It Looks Beautiful but Scary

    Yuhang Zhao, Elizabeth Kupferstein, Doron Tal, and Shiri Azenkot. 2018. " It Looks Beautiful but Scary" How Low Vision People Navigate Stairs and Other Surface Level Changes. In Proceedings of the 20th International ACM SIGACCESS Conference on Computers and Accessibility . 307–320

  98. [106]

    Yuhang Zhao, Sarit Szpiro, and Shiri Azenkot. 2015. Foresee: A customizable head-mounted vision enhancement system for people with low vision. In Pro- ceedings of the 17th international ACM SIGACCESS conference on computers & accessibility. 239–249

  99. [107]

    Zejia Zhang, Bo Yang, Xinxing Chen, Weizhuang Shi, Haoyuan Wang, Wei Luo, and Jian Huang. 2025. MindEye-OmniAssist: A Gaze-Driven LLM-Enhanced Assistive Robot System for Implicit Intention Recognition and Task Execution. arXiv preprint arXiv:2503.13250 (2025)

  100. [109]

    Yuhang Zhao, Elizabeth Kupferstein, Brenda Veronica Castro, Steven Feiner, and Shiri Azenkot. 2019. Designing AR visualizations to facilitate stair navigation for people with low vision. In Proceedings of the 32nd annual ACM symposium on user interface software and technology ...

  101. [113]

    Yuhang Zhao, Sarit Szpiro, Jonathan Knighten, and Shiri Azenkot. 2016. CueSee: exploring visual cues for people with low vision to facilitate a visual search task. In Proceedings of the 2016 ACM International Joint Conference on Pervasive and Ubiquitous Computing. 73–84. A App...

  102. [114]

    How to get to the T1 Departures?

  103. [115]

    Is there a person wearing a green plaid shirt?

  104. [116]

    What accessories does the woman wear?

  105. [117]

    What are the students holding?

  106. [118]

    What color are the woman’s gloves?

  107. [119]

    What color is the backpack worn by the boy on the right?

  108. [120]

    What color is the big truck?

  109. [121]

    What color is the boy’s hair?

  110. [122]

    What color is the cap of the tallest boy?

  111. [123]

    What color is the cap the man in the middle wears?

  112. [124]

    What color is the car on the left of the image?

  113. [125]

    What color is the car on the right?

  114. [126]

    What color is the chair on the left?

  115. [127]

    What color is the coat the mannequin wears on the right of the image?

  116. [128]

    What color is the crossbody bag on the left?

  117. [129]

    What color is the dog?

  118. [130]

    What color is the dress the girl is wearing in the center?

  119. [131]

    What color is the egg tart?

  120. [132]

    What color is the girl’s backpack?

  121. [133]

    What color is the girl’s clothes?

  122. [134]

    What color is the helmet of the biggest person?

  123. [135]

    What color is the ice cream?

  124. [136]

    What color is the kid’s t-shirt?

  125. [137]

    What color is the man’s beard?

  126. [138]

    What color is the man’s headscarf?

  127. [139]

    What color is the man’s shirt?

  128. [140]

    What color is the mask worn by the woman on the right in this image?

  129. [141]

    What color is the sofa in the center of the image?

  130. [142]

    What color is the speaker’s t-shirt?

  131. [143]

    What color is the suit of the person in front of the microphone?

  132. [144]

    What color is the swimming cap on the top of the image?

  133. [145]

    What color is the t-shirt of the jumping man?

  134. [146]

    What color is the t-shirt of the person standing on the right?

  135. [147]

    What color is the t-shirt the man wears on the left of this photo?

  136. [148]

    What color is the tablet?

  137. [149]

    What color is the tent?

  138. [150]

    What color is the towel hanging on the wall?

  139. [151]

    What color is the umbrella?

  140. [152]

    What color is the washer on the left?

  141. [153]

    What color is the woman in the center wearing?

  142. [154]

    What color is the woman’s backpack?

  143. [155]

    What color is the woman’s clothes?

  144. [156]

    What color is the woman’s dress?

  145. [157]

    What color is the woman’s pants?

  146. [158]

    What color is the woman’s t-shirt?

  147. [159]

    What color of clothes does the person on the left wear?

  148. [160]

    What color t-shirt does the biggest girl wear?

  149. [161]

    What food is in the center of the table?

  150. [162]

    What food is inside the bowl?

  151. [164]

    What ingredient is on the biggest donut?

  152. [165]

    What is his facial expression?

  153. [166]

    What is in the bowl on the right?

  154. [167]

    What is in the middle of the image?

  155. [168]

    What is in the pan on the left?

  156. [169]

    What is next to the woman’s laptop?

  157. [170]

    What is on the blue door on the left?

  158. [171]

    What is on the face of the person on the left?

  159. [172]

    What is on the plate?

  160. [174]

    What is on the wall?

  161. [175]

    What is the biggest person holding?

  162. [176]

    What is the color of the banner?

  163. [177]

    What is the color of the woman’s nails?

  164. [178]

    What is the facial expression of the man leaning on the couch?

  165. [179]

    What is the facial expression of the man on the left?

  166. [180]

    What is the facial expression of the person on the right?

  167. [181]

    What is the facial expression of the person wearing pink?

  168. [182]

    What is the facial expression of the woman on the left?

  169. [183]

    What is the facial expression of the woman?

  170. [184]

    What is the food on the plate next to the mug?

  171. [185]

    What is the hair color of the woman in the center?

  172. [186]

    What is the hairstyle of the girl on the left?

  173. [187]

    What is the hand gesture of the person in the gray t-shirt on the left?

  174. [188]

    What is the man holding?

  175. [189]

    What is the man wearing on his head?

  176. [190]

    What is the man’s facial expression?

  177. [191]

    What is the number at the bottom right of the image?

  178. [192]

    What is the pattern of the carpet?

  179. [193]

    What is the person doing?

  180. [194]

    What is the person in the red shirt on the right doing?

  181. [195]

    What is the person on the left holding? ASSETS ’25, October 26–29, 2025, Denver, CO, USA Wang et al

  182. [196]

    What is the posture of the girl on the right?

  183. [197]

    What is the posture of the man on the left?

  184. [198]

    What is the posture of the woman?

  185. [199]

    What is the title of the document on this screenshot?

  186. [200]

    What is the woman holding?

  187. [201]

    What is the woman in yellow holding?

  188. [202]

    What is the woman’s facial expression?

  189. [203]

    What is the woman’s hair color?

  190. [204]

    What kind of car is in the image?

  191. [205]

    What kind of food is on the plate?

  192. [206]

    What kind of instrument is the man on the front playing?

  193. [207]

    What kind of instrument is the person on the left playing?

  194. [208]

    Where is the calendar application?

  195. [209]

    Where is the search bar?

  196. [210]

    Where is the snack with pink packaging?

  197. [211]

    Which folder is selected on this google drive page in the screen- shot?

  198. [212]

    Who is wearing eye glasses?

  199. [213]

    Who is wearing sneakers?

  200. [214]

    Who is wearing the blue sneakers? A.1.2 Level 2: Cross-Object Information

  201. [215]

    How many boats are there in this photo?

  202. [216]

    How many buses are shown in this photo?

  203. [217]

    How many cars are there in this photo?

  204. [218]

    How many chairs are there?

  205. [219]

    How many cups are there?

  206. [220]

    How many donuts are there?

  207. [221]

    How many drawers are there?

  208. [222]

    How many egg tarts are there?

  209. [223]

    How many flavors of ice cream are there?

  210. [224]

    How many icons are there in this screenshot?

  211. [225]

    How many mannequins are there?

  212. [226]

    How many open drawers are in this photo?

  213. [227]

    How many pans are there on the stovetop?

  214. [228]

    How many people are crossing the street?

  215. [229]

    How many people are facing the sea?

  216. [230]

    How many people are in the hallway?

  217. [231]

    How many people are in this photo?

  218. [232]

    How many people are not standing?

  219. [233]

    How many people are raising their arms?

  220. [234]

    How many people are receiving the awards?

  221. [235]

    How many people are seated?

  222. [236]

    How many people are standing next to the speaker?

  223. [237]

    How many people are there in the second row?

  224. [238]

    How many people are wearing sunglasses?

  225. [239]

    How many people wear the green apron?

  226. [240]

    How many pillows are on the bed?

  227. [241]

    How many pineapples are on the shelf?

  228. [242]

    How many plates are there on the table?

  229. [243]

    How many police cars are there?

  230. [244]

    How many policemen are there?

  231. [245]

    How many pots are there on the stovetop?

  232. [246]

    How many sinks are there in this photo?

  233. [247]

    How many snacks are there?

  234. [248]

    How many suitcases are in this photo?

  235. [249]

    How many tabs are open in the chrome browser?

  236. [250]

    How many throw pillows are there?

  237. [251]

    How many toy cars are there?

  238. [252]

    What are people on the right of this photo doing?

  239. [253]

    What are the two boys in the middle of the photo doing?

  240. [254]

    What are the two girls doing?

  241. [255]

    What are the two people doing?

  242. [256]

    What are the two people in the middle doing?

  243. [257]

    What are the two women on the left doing?

  244. [258]

    What are these three people doing?

  245. [259]

    What are people on the image doing?

  246. [260]

    What food is the person cooking?

  247. [261]

    What is the man doing?

  248. [262]

    What is the man looking at?

  249. [263]

    What is the person doing with their hands?

  250. [264]

    What is the woman and the little kid doing?

  251. [265]

    What is the woman doing?

  252. [266]

    What is this person holding?

  253. [267]

    What is on the table?

  254. [268]

    What is the relationship between the people in this image?

  255. [269]

    Who is the man on the left looking at? A.1.3 Level 3: Overall Interpretation Information

  256. [270]

    What activity does this image describe?

  257. [271]

    What are the characteristics of this alley?

  258. [272]

    What emotion does the image convey?

  259. [273]

    What event does the image describe?

  260. [274]

    What event might have happened before this picture was taken?

  261. [275]

    What is the atmosphere of this image?

  262. [276]

    What is the purpose of taking this photo? Characterizing Visual Intents for People with Low Vision through Eye Tracking ASSETS ’25, October 26–29, 2025, Denver, CO, USA

  263. [277]

    What is the purpose of the application shown in the image?

  264. [278]

    What is the purpose of showing this image?

  265. [279]

    What kind of store is this?

  266. [280]

    Where are the egg tarts?

  267. [281]

    Where is this photo taken?

  268. [2016]

    In Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility

    How people with low vision access computing devices: Understanding challenges and opportunities. In Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility . 171–180

  269. [2017]

    lmerTest package: tests in linear mixed effects models.Journal of statistical software 82, 13 (2017)

  270. [2023]

    IEEE transactions on visualization and computer graphics (2023)

    Task-Dependent Visual Behavior in Immersive Environments: A Compar- ative Study of Free Exploration, Memory and Visual Search. IEEE transactions on visualization and computer graphics (2023)

  271. [2024]

    InProceedings of the CHI Conference on Human Factors in Computing Systems

    Investigating Use Cases of AI-Powered Scene Description Applications for Blind and Low Vision People. InProceedings of the CHI Conference on Human Factors in Computing Systems . 1–21

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.