Pith. sign in

REVIEW 2 cited by

Improved GUI Grounding via Iterative Narrowing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.13591 v7 pith:GSKR7OZD submitted 2024-11-18 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords groundingperformancegeneraliterativemodelsnarrowingvariousacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Graphical User Interface (GUI) grounding plays a crucial role in enhancing the capabilities of Vision-Language Model (VLM) agents. While general VLMs, such as GPT-4V, demonstrate strong performance across various tasks, their proficiency in GUI grounding remains suboptimal. Recent studies have focused on fine-tuning these models specifically for zero-shot GUI grounding, yielding significant improvements over baseline performance. We introduce a visual prompting framework that employs an iterative narrowing mechanism to further improve the performance of both general and fine-tuned models in GUI grounding. For evaluation, we tested our method on a comprehensive benchmark comprising various UI platforms and provided the code to reproduce our results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Auxiliary Reasoning Unleashes GUI Grounding in VLMs

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Overlaying labeled grids and axes on screenshots substantially improves zero-shot GUI grounding in most VLMs, with the best variant zooming into grid cells.

  2. DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning

    cs.AI 2025-06 conditional novelty 4.0 of 10

    Separating text and icon grounding with iterative zooming improves GUI-element localization accuracy of existing vision-language models without retraining.

Pith tools