REVIEW 2 cited by
Spotlight: Mobile UI Understanding using Vision-Language Models with a Focus
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Mobile UI understanding is important for enabling various interaction tasks such as UI automation and accessibility. Previous mobile UI modeling often depends on the view hierarchy information of a screen, which directly provides the structural data of the UI, with the hope to bypass challenging tasks of visual modeling from screen pixels. However, view hierarchies are not always available, and are often corrupted with missing object descriptions or misaligned structure information. As a result, despite the use of view hierarchies could offer short-term gains, it may ultimately hinder the applicability and performance of the model. In this paper, we propose Spotlight, a vision-only approach for mobile UI understanding. Specifically, we enhance a vision-language model that only takes the screenshot of the UI and a region of interest on the screen -- the focus -- as the input. This general architecture of Spotlight is easily scalable and capable of performing a range of UI modeling tasks. Our experiments show that our model establishes SoTA results on several representative UI tasks and outperforms previous methods that use both screenshots and view hierarchies as inputs. Furthermore, we explore multi-task learning and few-shot prompting capacities of the proposed models, demonstrating promising results in the multi-task learning direction.
Forward citations
Cited by 2 Pith papers
-
Scaling Mobile Chaos Testing with AI-Driven Test Execution
An integrated LLM-based mobile testing system and service-level fault injector ran 180,000+ chaos tests at Uber, finding 23 resilience defects that manual and backend-only testing missed.
-
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
GTArena is a unified benchmark showing that current multimodal LLMs perform poorly across all three subtasks of automated GUI testing.
Discussion (0). Continue with ORCID to comment.