Pith. sign in

REVIEW 3 cited by

Falcon-UI: Understanding GUI Before Following User Instructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09362 v1 pith:CASGOQGF submitted 2024-12-12 cs.CL

classification cs.CL
keywords datasetcontextandroidfalcon-uiinsight-uiunderstandinguseragent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pursuing human-like interaction for Graphical User Interface (GUI) agents requires understanding the GUI context and following user instructions. However, existing works typically couple these two aspects and focus more on instruct-following abilities, while ignoring the importance of understanding the GUI context. In this paper, we introduce an instruction-free GUI navigation dataset, termed Insight-UI Dataset, to enhance model comprehension of GUI environments. Insight-UI Dataset is automatically generated from the Common Crawl corpus, simulating various platforms -- including iOS, Android, Windows, and Linux -- across multiple resolutions on 312K domains. Although GUI interactions vary by context, diverse interfaces share common internal patterns, such as clicking an item to view its details. It implies the feasibility of independent GUI operation learning, followed by joint optimization with instruction tuning. Thereby, we develop the GUI agent model Falcon-UI, which is initially pretrained on Insight-UI Dataset and subsequently fine-tuned on Android and Web GUI datasets, including AITW, AITZ, Android Control, and Mind2Web. With 7 billion parameters, Falcon-UI achieves accuracy comparable to the 72 billion-parameter Qwen2VL on AITZ, validating the alignment between GUI context comprehension and agent performance. Our code and dataset will be open-sourced.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents

    cs.AI 2026-08 conditional novelty 6.0 of 10

    AppDeltaWorld predicts mobile GUI transitions as code updates retrieved under action constraints, and its generated trajectories improve an 8B mobile agent on several benchmarks.

  2. Scaling GUI Agents with Visual State Transitions

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A joint inverse-forward pretraining stage on visual screen transitions improves GUI-agent fine-tuning by 0.6–6.2 percentage points across three benchmarks.

  3. The Devil is in Fine-tuning and Long-tailed Problems:A New Benchmark for Scene Text Detection

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Scene text detectors overfit under dataset-specific fine-tuning; the paper proposes Joint-Dataset Learning and a 13-category Long-Tailed Benchmark (LTB) with a self-supervised MAEDet baseline.

Pith tools