Pith. sign in

REVIEW 3 major objections 4 minor 285 references

Software Engineering for and with GUI Agent

T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read GUI agents are closed-loop systems that benchmarks alone cannot validate

desk verdict First SE-lifecycle map of GUI agents that reads the literature as systems work; the gap counts are directionally right but the absence-as-evidence rule makes the magnitudes unverifiable. read the letter →

arxiv 2608.09278 v1 pith:G5VJYGW4 submitted 2026-08-10 cs.SE cs.AI

classification cs.SEcs.AI
keywords GUIagentssoftwareengineeringclosed-loopsystemsbenchmarksevaluationhumanoversightlifecyclelargelanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review of 336 papers argues that GUI agents have outgrown the framing of a model that clicks: they are now closed-loop software systems coupling perception, reasoning, execution, feedback, recovery, and human oversight. Because of that, the paper claims, capability improvements alone cannot guarantee deployment readiness. It finds the field has expanded sharply since 2024 while recovery, human escalation, safety enforcement, auditability, maintainability, observability, and privacy controls remain underdeveloped. The conclusion is that progress depends on lifecycle engineering, meaning explicit execution contracts, trajectory-based testing, reproducible evaluation, and governed human oversight, rather than on benchmark scores alone.

What carries the argument

The analytical object is the closed-loop perceive-decide-execute architecture treated as a software system, coded across 327 SE-parseable papers for signals such as verification, memory, reflection, recovery, execution contracts, observability, and governance. The central proposed mechanism is the execution contract: a specification of valid actions, acceptable state changes, retry and stop rules, rollback conditions, and permission boundaries that would make recovery and human oversight testable. The coding scheme's rule that absent signals count as lack of explicit evidence carries the survey's gap statistics.

What would settle it

Re-code a random sample of papers coded as lacking maintenance and observability signals, searching for implicit versioning, log, replay, escalation, or permission mechanisms; if a substantial fraction describes such mechanisms without the coded keywords, the survey's gap statistics would not hold.

Watch

Extended reading notes

Core claim

The paper establishes that GUI agents should be analyzed as closed-loop software systems whose behavior emerges from the interaction between a foundation model, an orchestration layer, the target application, and the user. Reviewing 336 papers from 2018 to April 2026, it finds that modular perceive-decide-act architectures dominate but that recovery, human escalation, safety enforcement, and auditability are far less developed than perception and planning. On evaluation, task success dominates as a metric while protocol comparability, reproducibility, and risk-awareness lag. Across the lifecycle, testing beyond benchmarks, maintainability, observability, privacy controls, and systematic human oversight are sparse. The paper concludes that capability improvements alone cannot ensure deployment readiness and that future work must connect dependable execution with lifecycle-centered testing, reproducible evaluation, and cost-aware, human-centered governance.

Load-bearing premise

The survey treats the absence of an explicitly coded engineering signal as evidence that the underlying system lacks that support, so if many agents implement recovery, maintenance, or human oversight under different terminology, the reported gaps would be overstated.

Editorial extensions

If this is right

  • Benchmark task success should no longer be treated as sufficient evidence of deployment readiness; a reported score describes the whole agent stack and protocol, not just the model.
  • Evaluation protocols must disclose observation access, action spaces, environment resets, success oracles, retry policies, and intervention rules before results can be meaningfully compared across benchmarks.
  • Research should invest in explicit execution contracts that define valid actions, stop conditions, escalation points, and reversible versus irreversible operations.
  • A GUI-agent test pyramid spanning component tests, integration tests, trajectory regression, and adversarial system-level stress tests is needed to replace benchmark-only validation.
  • Permissions, privacy controls, audit logs, and risk-adaptive human oversight must be designed as runtime components of the system rather than as post-hoc safeguards.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the coding rule is correct, published task-success rates likely overstate operational maturity, since many systems implement recovery or oversight implicitly without reporting it as an engineering artifact.
  • The execution-contract proposal could be made directly testable by expressing contracts as preconditions and postconditions over UI states and measuring violation rates under interface perturbations.
  • A natural extension is to treat benchmarks themselves as versioned software artifacts with reproducibility budgets, an idea the paper opens but does not fully formalize.
  • Merging early reinforcement-learning web agents with post-2023 LLM agents in one corpus may smooth over generational differences; separating those cohorts could sharpen the reported architectural and evaluation trends.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This paper surveys 336 GUI-agent papers published or posted between January 2018 and April 2026 from a software-engineering perspective. It poses five research questions covering the research landscape, system architectures, evaluation protocols, software-lifecycle concerns, and future opportunities. The central claim is that GUI agents are closed-loop software systems whose deployment readiness depends on lifecycle engineering, and that capability improvements alone cannot ensure dependable, maintainable, secure, and deployable systems. The survey reports multi-label, functionally interpreted coding of the corpus and finds that recovery, human escalation, safety enforcement, maintainability, observability, and oversight are underdeveloped relative to perception, planning, and task-capability work, and that evaluation remains centered on task success with limited cross-protocol comparability. The paper closes with a research agenda built around execution contracts, trajectory-based testing, reproducible evaluation, and governed human oversight.

Significance. If its findings are accepted, this survey provides a valuable corrective to benchmark-centric assessments of GUI agents and a useful map for software-engineering research on these systems. The paper has clear strengths: a large and explicitly bounded corpus; a transparent RQ structure; consistent reporting of per-RQ denominators; multi-label coding and an explicit functional-interpretation rule for mechanisms such as verifiers, critics, and reward models; and a candid threats-to-validity section. The synthesis connecting architecture, evaluation, and lifecycle concerns is genuinely useful and goes beyond existing capability-oriented surveys. However, the headline negative claims are built on counts that treat the absence of explicit textual signals as absence of engineering support, and the paper does not currently demonstrate the reliability of that coding. The significance of the survey is therefore conditional on coding transparency and validation, which are not yet provided.

major comments (3)
  1. [§11, §3.3, §7.3, §7.4] The central negative findings are load-bearing on the coding rule stated in §11: 'we treat absent signals as a lack of explicit evidence under the coding scheme.' For example, §7.3 reports that 279/327 papers lack a maintenance-process signal and 251/327 lack a monitoring or audit mechanism, and §7.4 reports runtime human oversight in only 47/327 papers. These counts are then used to conclude that lifecycle concerns are 'underdeveloped' and that 'capability improvements alone cannot ensure deployment readiness.' But a paper can implement recovery, confirmation, or oversight without using the survey's terminology; the functional-coding step in §3.3 mitigates this, yet §11 itself acknowledges that such concerns 'may remain implicit in system descriptions.' Without a sensitivity analysis that re-codes absent signals as unknown, or a validation sample checked against author statements or system repositories, the size of the reported gaps may be inflated. I would like to see either (a) a demonstration that absent signals correlate with absent mechanisms in a sample, or (b) rewording of the statistics as 'explicitly discussed under this coding scheme' rather than 'lacking support.'
  2. [§3.3, §11] Section 3.3 states that the coding scheme was developed iteratively and that records were checked for internal coherence, but Section 11 concedes that these 'internal-coherence checks... do not provide independently measured inter-rater agreement from multiple coders.' Because the RQ2 and RQ4 statistics are the quantitative backbone of the paper, the absence of any inter-rater agreement measure for the interpretive dimensions is a significant gap. The paper should report agreement per coding dimension—especially for human oversight, maintainability, and observability—on a sample of papers, and describe how disagreements were resolved. Without this, counts such as 95/327 versus 86/327 for related runtime-control signals are not auditable, and the precision of the gap percentages cannot be assessed.
  3. [§3.3, Table 6, §7] The RQ4 denominator of 327 out of 336 records is never explained. Section 3.3 says 'RQ4 uses the 327 records with complete software-engineering coding,' and Table 6 uses '327 SE-parseable papers,' but the reader is not told which nine papers are excluded or what 'SE-parseable' means. This matters because all RQ4 shares use the 327 denominator, and a reader cannot reproduce the subsets or check for selection effects. Please state the exclusion rule, provide the identifiers or at least the contribution types and years of the excluded records, and release the coding instrument with the paper so that the reported counts can be independently verified.
minor comments (4)
  1. [Table 7] Table 7 mixes denominators within one table: the first three rows use the 327-paper SE-parseable set, while the last row ('Explicit human-in-the-loop flag') uses the 145-paper framework set. The caption notes this, but it would be clearer to split the table or to add a denominator column to each row.
  2. [Figure 3] The right panel of Figure 3 reports contribution types as multi-label counts, and the text correctly notes that counts do not sum to the corpus size; the figure caption would benefit from stating this explicitly to avoid reader confusion.
  3. [§3.2, §4.1] Table 1 shows that 60.7% of the corpus consists of preprints. The paper would be strengthened by a brief robustness discussion of whether the RQ1 growth trends and the RQ4 gap counts change when the analysis is restricted to peer-reviewed records.
  4. [§3.3] The term 'SE-parseable' is used at first in Section 3.3 and then repeatedly in the RQ4 analysis, but it is only loosely defined. Please provide a precise definition at first use, even if the full list of excluded records appears in an appendix or supplementary artifact.

Circularity Check

1 steps flagged · score 2.0 of 10

Survey's gap statistics are partly self-definitional via the absence-as-evidence coding rule, but conclusions retain independent empirical content.

  1. self definitional [Section 11, Threats to Validity (second paragraph); operationalized in RQ4 (Section 7) and Table 13]
    "We use multi-label and functional coding to reduce this loss and treat absent signals as a lack of explicit evidence under the coding scheme."

    The RQ4 conclusions that maintainability, observability, privacy, audit, and human oversight are 'underdeveloped' are derived from counting papers that lack explicit coded signals for these concerns. Under the stated coding rule, 'underdeveloped' is defined as 'absent signal,' so the reported gap percentages (e.g., 279/327 missing a maintenance-process signal, 251/327 missing monitoring/audit) are equivalent to the coding input by construction. The paper itself acknowledges that lifecycle concerns 'may remain implicit in system descriptions,' meaning the coding rule is an assumption that makes the central negative claim about field immaturity partly self-referential.

full rationale

This is a survey paper that describes and synthesizes 336 external publications; it does not derive predictions or first-principles results from its own model. The only point where a conclusion reduces to an input by construction is the operationalization of 'underdeveloped' in RQ4: the coding rule in Section 11 equates absence of an explicit signal with lack of explicit evidence, and the paper then reports low counts as evidence of underdevelopment. This is a genuine self-definitional step for those specific descriptive findings (e.g., maintainability, observability, human oversight). Nevertheless, the paper discloses this rule in Threats to Validity, uses functional coding to mitigate terminology differences, and supports its broader conclusions with a large external corpus and with independent benchmark and safety studies. No self-citation chain, no fitted parameter masquerading as a prediction, and no uniqueness theorem imported from the authors' prior work were found. Given the transparency and the independent empirical base, a score of 2 is appropriate: one minor self-referential step that does not undermine the central claim's independent content.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No free parameters or invented entities. The survey's conclusions rest on four domain assumptions: corpus representativeness, functional coding reliability, absence-as-evidence, and the choice of SE quality attributes. These are stated or acknowledged in Sections 3.3 and 11, but they are not independently benchmarked.

assumptions (4)
  • domain assumption The 336-paper corpus is representative of GUI-agent research.
    Corpus built via keyword search, backward/forward snowballing, and manual consolidation; the paper admits it 'cannot constitute an exhaustive census' (Section 11).
  • domain assumption Functional coding of modules and engineering signals reflects actual system behavior.
    The paper codes by function when authors use different names, e.g., 'a verifier, critic, evaluator, reward model, or post-action checker may provide verification' (Section 3.3). This requires interpretation and could misclassify systems.
  • domain assumption Absence of an explicit signal means absence of explicit engineering support.
    The paper states it 'treat[s] absent signals as a lack of explicit evidence under the coding scheme' (Section 11). The RQ2 and RQ4 statistics rely on this rule.
  • ad hoc to paper The software-engineering quality attributes chosen (recovery, maintainability, observability, privacy, human oversight) are the right lens for deployment readiness.
    The paper selects these dimensions a priori from SE practice; other lenses (e.g., economic, organizational, or HCI-centric) could yield different priorities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Software Engineering for and with GUI Agent." pith.science (2026). https://pith.science/paper/G5VJYGW4

@misc{pith2026260809278,
  author       = {Pith},
  title        = {Pith review of: Software Engineering for and with GUI Agent},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/G5VJYGW4}},
  note         = {Machine review of arXiv:2608.09278}
}
read the original abstract

GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.

Figures

Figures reproduced from arXiv: 2608.09278 by the authors.

Figure 1
Figure 1. GUI-agent interaction loop. often required application-specific adaptation. Moving to a new application commonly required additional demonstrations, handcrafted selectors, or a new representation of target interface. Large language models and multimodal foundation models expanded this paradigm. Language models improved instruction interpretation and task decomposition, while vision-language mod￾els made screenshots … view at source ↗
Figure 2
Figure 2. GUI-agent research landscape. 4 RQ1: Research Landscape and Development Trends We use RQ1 to establish the empirical landscape for the later analyses. The corpus contains 336 papers from January 2018 to April 2026. Only 19 appeared before 2024, while 317 were published or posted during 2024–2026. This increase coincides with stronger multimodal models, interactive benchmarks across major platforms, and a shift from … view at source ↗
Figure 3
Figure 3. Publication years and contribution types. [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Platform and application-domain distributions. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Observation-medium distribution. 4.3 Screenshot-Centered Observation and Model Adaptation The enabling technology behind recent GUI agents is also changing. We summarize the observation media used to interpret interfaces in [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Modular GUI-agent architecture. screenshots and Android view hierarchies, and desktop agents that must coordinate across windows, applications, and files. We use [PITH_FULL_IMAGE:figures/full_fig_p015_6.png]
Figure 7
Figure 7. Figure 7: GUI-agent evaluation pipeline. 6 RQ3: Evaluation and Benchmarking We use RQ3 to examine whether current evaluations provide credible evidence for the systems characterized in RQ2. The analysis uses the 252 papers coded as frameworks, models, or evaluation studies. GUI-…
Figure 8
Figure 8. Figure 8: GUI-agent benchmark and evaluation families. [PITH_FULL_IMAGE:figures/full_fig_p023_8.png]
Figure 9
Figure 9. Figure 9: Software-engineering lifecycle of GUI agents. [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: Research roadmap for dependable GUI agents. [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

285 extracted references · 56 canonical work pages

  1. [1]

    Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan, Aditya Vempaty, and Ravi Kokku. 2024. Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems. arXiv:2407.13032 [cs.AI] doi:10.48550/arXiv.2407.13032

  2. [2]

    Saaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Eric Wang. 2024. Agent S: An Open Agentic Framework that Uses Computers Like a Human. arXiv:2410.08164 [cs.AI] doi:10.48550/arXiv.2410.08164

  3. [3]

    Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. 2025. Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents. arXiv:2504.00906 [cs.AI] doi:10.48550/arXiv.2504.00906

  4. [4]

    Pranjal Aggarwal and Sean Welleck. 2025. Programming with Pixels: Computer-Use Meets Software Engineering. arXiv:2502.18525 [cs.SE] doi:10.48550/arXiv.2502.18525

  5. [5]

    Jaewoo Ahn, Junseo Kim, Heeseung Yun, Jaehyeon Son, Dongmin Park, Jaewoong Cho, and Gunhee Kim. 2025. FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 23365–23395. doi:10.18653/v1/2025.emnlp- main.1192

  6. [6]

    Neeraj Anand, Rishabh Jain, Sohan Patnaik, Balaji Krishnamurthy, and Mausoom Sarkar. 2025. AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent. arXiv:2512.00846 [cs.CV] doi:10.48550/ arXiv.2512.00846

  7. [7]

    Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024. ScreenAI: A Vision-Language Model for UI and Infographics Understanding. arXiv:2402.04615 [cs.CV] doi:10.48550/arXiv.2402.04615

  8. [8]

    Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Agüera y Arcas

Show all 285 references
  1. [9]

    Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar. 2024. DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning. InAdvances in Neural Information Processing Systems 37. 12461–12495. doi:10.52202/079017-0397

  2. [10]

    Hao Bai, Yifei Zhou, Li Li, Sergey Levine, and Aviral Kumar. 2025. Digi-Q: Learning VLM Q-Value Functions for Training Device-Control Agents. InInternational Conference on Learning Representations. https://proceedings.iclr.cc/ paper_files/paper/2025/hash/519abe71ee55aac4fe821b...

  3. [11]

    Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks. InAdvances in Neural ...

  4. [12]

    Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Jang, and Zack Hui. 2024. Windows Agent Arena: Evaluating Multi- Modal OS Agents at Scale. arXiv:2409.08264 [cs.AI] doi:10....

  5. [13]

    Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. 2022. A Dataset for Interactive Vision-Language Navigation with Unknown Command Feasibility.Lecture Notes in Computer Science (2022), 312–328. doi:10.1007/978-3-031-20074-8_18

  6. [15]

    Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Yuchen Mao, Wenjing Hu, Tianbao Xie, Hongshen Xu, Danyang Zhang, Sida Wang, Ruoxi Sun, Pengcheng Yin, Caiming Xiong, Ansong Ni, Qian Liu, Victor Zhong, Lu Chen, Kai Yu...

  7. [16]

    Hyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak, Gwanwoo Song, Jihoon Kim, Sunghwan Kim, Dongha Lee, and Jinyoung Yeo. 2024. Web Agents with World Models: Learning and Leveraging Environment Software Engineering for and with GUI Agent 1:39 Dynamics in Web Navigation....

  8. [17]

    Yuxiang Chai, Siyuan Huang, Yazhe Niu, Han Xiao, Liang Liu, Guozhi Wang, Dingyu Zhang, Shuai Ren, and Hongsheng Li. 2025. AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 2138–2156. doi:10...

  9. [18]

    Yuxiang Chai, Shunye Tang, Han Xiao, Weifeng Lin, Hanhao Li, Jiayu Zhang, Liang Liu, Pengxiang Zhao, Guangyi Liu, Guozhi Wang, Shuai Ren, Rongduo Han, Haining Zhang, Siyuan Huang, and Hongsheng Li. 2025. A3: Android Agent Arena for Mobile GUI Agents with Essential-State Proced...

  10. [19]

    Iason Chaimalas, Arnas Vyšniauskas, and Gabriel Brostow. 2025. Explorer: Robust Collection of Interactable GUI Elements. arXiv:2504.09352 [cs.HC] doi:10.48550/arXiv.2504.09352

  11. [20]

    Rajat Chawla, Adarsh Jha, Muskaan Kumar, Mukunda NS, and Ishaan Bhola. 2024. GUIDE: Graphical User Interface Data for Execution. arXiv:2404.16048 [cs.HC] doi:10.48550/arXiv.2404.16048

  12. [21]

    Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu, Anzhou Li, Guo Gan, Ziyuan Huang, Cheng Zou, Jiajia Liu, Jingdong Chen, Hao Chen, and Chunhua Shen. 2025. GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks. arXiv:2509.23738 [cs.AI] doi:10.48550/arX...

  13. [22]

    Chen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu, Xiangcheng Liu, Hantao Yao, and Wu Liu. 2026. GUI-Eyes: Tool- Augmented Perception for Visual Grounding in GUI Agents.Proceedings of the AAAI Conference on Artificial Intelligence40, 35 (2026), 29350–29358. doi:10.1609/aaai.v40i35.40175

  14. [23]

    Chiyu Chen, Xinhao Song, Yunkai Chai, Yang Yao, Haodong Zhao, Lijun Li, Jie Li, Yan Teng, Gongshen Liu, and Yingchun Wang. 2025. GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On- Device Environments? arXiv:2510.20333 [cs.CR] doi:10.48550/arXi...

  15. [24]

    Chaoran Chen, Zhiping Zhang, Bingcan Guo, Shang Ma, Ibrahim Khalilov, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. The Obvious Invisible Threat: LLM-Powered GUI Agents’ Vulnerability to Fine-Print Injections. arXiv:2504.11...

  16. [25]

    Chaoran Chen, Zhiping Zhang, Ibrahim Khalilov, Bingcan Guo, Simret A Gebreegziabher, Yanfang Ye, Ziang Xiao, Yaxing Yao, Tianshi Li, and Toby Jia-Jun Li. 2025. Toward a Human-Centered Evaluation Framework for Trustworthy LLM-Powered GUI Agents. arXiv:2504.17934 [cs.HC] doi:10....

  17. [26]

    Dongping Chen, Yue Huang, Siyuan Wu, Jingyu Tang, Liuyi Chen, Yilin Bai, Zhigang He, Chenlong Wang, Huichi Zhou, Yiqiang Li, Tianshuo Zhou, Yue Yu, Chujie Gao, Qihui Zhang, Yi Gui, Zhen Li, Yao Wan, Pan Zhou, Jianfeng Gao, and Lichao Sun. 2024. GUI-World: A Video Benchmark and...

  18. [27]

    Gongwei Chen, Lirong Jie, Lexiao Zou, Weili Guan, Miao Zhang, and Liqiang Nie. 2025. Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator. InAdvances in Neural Information Processing Systems

  19. [28]

    Gongwei Chen, Xurui Zhou, Rui Shao, Yibo Lyu, Kaiwen Zhou, Shuai Wang, Wentao Li, Yinchuan Li, Zhongang Qi, and Liqiang Nie. 2025. Less is More: Empowering GUI Agent with Context-Aware Simplification. (2025), 5901–5911. doi:10.1109/iccv51701.2025.00558

  20. [29]

    Jingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang, Gongwei Chen, Zhihao Wu, Li Yixing, Xurui Zhou, Weiwen Liu, Shuai Wang, Kaiwen Zhou, Rui Shao, Liqiang Nie, Yasheng Wang, Jianye Hao, Jun Wang, and Kun Shao

  21. [30]

    Liang Chen, Haozhe Zhao, Yinzhen Huang, Yang Luo, Tsekai Lin, Weichu Xie, Ruoyu Wu, Peiyi Wang, Runxin Xu, Ming Wu, and Baobao Chang. 2025. CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3. InProceedings of the 34th ACM International Confer...

  22. [31]

    Qi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou, Hsiang-Ting Chen, and Qi Wu. 2024. WebVLN: Vision- and-Language Navigation on Websites.Proceedings of the AAAI Conference on Artificial Intelligence38, 2 (2024), 1165–1173. doi:10.1609/aaai.v38i2.27878

  23. [32]

    Ruihan Chen, Qiming Li, Xiaocheng Feng, Weihong Zhong, Xiaoliang Yang, Yuxuan Gu, Zekun Zhou, Yunfei Lu, Haoyu Ren, Kun Chen, Dandan Tu, and Bing Qin. 2025. MPR-GUI: Benchmarking and Enhancing Multilingual Perception and Reasoning in GUI Agents. arXiv:2512.00756 [cs.AI] doi:10...

  24. [33]

    Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Yue Zhao, Chongyi Wang, Jun Liu, Guirong Chen, Yupeng Huo, Yuan Yao, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2025. GUICourse: From General Vision Language Model to Versatile GUI Agent. (2025), 21936–21959. doi:10.186...

  25. [34]

    Wei Chen and Zhiyuan Li. 2024. Octopus v2: On-device language model for super agent. arXiv:2404.01744 [cs.CL] doi:10.48550/arXiv.2404.01744 1:40 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, and Zhenyu Chen

  26. [35]

    Wei Chen, Zhiyuan Li, Zhen Guo, and Yikang Shen. 2026. Octo-Planner: On-Device Language Model for Planner- Action Agents.Lecture Notes in Computer Science(2026), 141–156. doi:10.1007/978-3-032-18011-7_9

  27. [36]

    Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, and Wei Jiang. 2025. PG-Agent: An Agent Powered by Page Graph. InProceedings of the 33rd ACM International Conference on Multimedia. 6878–6887. doi:10.1145/3746027.3755189

  28. [38]

    https://proceedings.neurips.cc/paper_files/paper/2025/hash/d067d16e3e5fe8fa8a3e62909907659a-Abstract- Conference.html

  29. [39]

    Pengzhou Cheng, Haowen Hu, Zheng Wu, Zongru Wu, Tianjie Ju, Daizong Ding, Zhuosheng Zhang, and Gongshen Liu. 2025. Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents. InFindings of the Association for Computational Linguistics: EMNLP 2025. ...

  30. [40]

    Pengzhou Cheng, Zheng Wu, Zongru Wu, Tianjie Ju, Aston Zhang, Zhuosheng Zhang, and Gongshen Liu. 2025. OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 6701–6725. doi:10.18653/v1/2025.findings-acl.348

  31. [41]

    Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024. SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents. (2024), 9313–9332. doi:10.18653/v1/2024.acl-long.505

  32. [42]

    Ziming Cheng, Zhiyuan Huang, Junting Pan, Zhaohui Hou, and Mingjie Zhan. 2025. Navi-plus: Managing Ambiguous GUI Navigation Tasks with Follow-up Questions. arXiv:2503.24180 [cs.CV] doi:10.48550/arXiv.2503.24180

  33. [43]

    Xu, Siva Reddy, Quentin Cappart, Graham Neubig, Ruslan Salakhutdinov, Nicolas Chapados, and Alexandre Lacoste

    Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin, Massimo Caccia, Léo Boisvert, Megh Thakkar, Tom Marty, Rim Assouel, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, Dehan Kong, Frank F. Xu, Siva Reddy, Quentin Cappart, Graham Neubig, Ruslan...

  34. [44]

    Weihua Cheng, Junming Liu, Yifei Sun, Botian Shi, Yirong Chen, and Ding Wang. 2025. MGA: Memory-Driven GUI Agent for Observation-Centric Interaction. arXiv:2510.24168 [cs.AI] doi:10.48550/arXiv.2510.24168

  35. [45]

    Gaole Dai, Shiqi Jiang, Ting Cao, Yuanchun Li, Yuqing Yang, Rui Tan, Mo Li, and Lili Qiu. 2025. Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment. arXiv:2503.15937 [cs.AI] doi:10.48550/arXiv.2503.15937

  36. [46]

    Preetam Prabhu Srikar Dammu. 2025. Towards Ethical and Personalized Web Navigation Agents: A Framework for User-Aligned Task Execution. InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining. 1074–1076. doi:10.1145/3701551.3707420

  37. [47]

    arXiv:2412.05467 [cs.LG] doi:10.48550/arXiv.2412.05467

    The BrowserGym Ecosystem for Web Agent Research. arXiv:2412.05467 [cs.LG] doi:10.48550/arXiv.2412.05467

  38. [48]

    Filippos Christianos, Georgios Papoudakis, Thomas Coste, Jianye Hao, Jun Wang, and Kun Shao. 2024. Lightweight Neural App Control. arXiv:2410.17883 [cs.AI] doi:10.48550/arXiv.2410.17883

  39. [49]

    Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, and Tat-Seng Chua. 2024. On the Multi-turn Instruction Following for Conversational Web Agents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 87...

  40. [50]

    Lei Ding, Jeshwanth Bheemanpally, and Yi Zhang. 2024. Enhancing Mobile "How-to" Queries with Automated Search Results Verification and Reranking. arXiv:2404.08860 [cs.IR] doi:10.48550/arXiv.2404.08860

  41. [51]

    Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Liujianfeng Liujianfeng, Ang Li, Jian Luan, Bin Wang, Rui Yan, and Shuo Shang. 2024. Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents. InProceedings of the 62nd Annual Meeting of the Association for Computa...

  42. [52]

    Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2Web: Towards a Generalist Agent for the Web. InAdvances in Neural Information Processing Systems 36. 28091–28114. doi:10.52202/075280-1220

  43. [53]

    Lingzhong Dong, Ziqi Zhou, Shuaibo Yang, Haiyue Sheng, Pengzhou Cheng, Zongru Wu, Zheng Wu, Gongshen Liu, and Zhuosheng Zhang. 2025. Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents. arXiv:2510.02204 [cs.CL] doi:10.48550/arXiv.2510.02204

  44. [54]

    Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste

    Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. 2024. WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Task...

  45. [55]

    Tinghe Ding. 2024. MobileAgent: enhancing mobile control via human-machine interaction and SOP integration. arXiv:2401.04124 [cs.HC] doi:10.48550/arXiv.2401.04124

  46. [56]

    Jinhan Dong, Lei Jin, Zhihong Zhang, Wei Tang, Runqing Zhang, Liqiang Xu, and Junliang Xing. 2025. MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided Fusion.IEEE Internet of Things Journal(2025),

  47. [57]

    doi:10.1109/jiot.2025.3600573

  48. [58]

    Yue Fan, Lei Ding, Ching-Chen Kuo, Shan Jiang, Yang Zhao, Xinze Guan, Jie Yang, Yi Zhang, and Xin Eric Wang. 2024. Read Anywhere Pointed: Layout-aware GUI Screen Reading with Tree-of-Lens Grounding. arXiv:2406.19263 [cs.CL] doi:10.48550/arXiv.2406.19263

  49. [59]

    Yue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen, Xin Eric Wang, and Gang Wu. 2025. GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 33249–33266. doi:10.18...

  50. [60]

    Yong Du, Yuchen Yan, Fei Tang, Zhengxi Lu, Chang Zong, Weiming Lu, Shengpei Jiang, and Yongliang Shen. 2026. Test-Time Reinforcement Learning for GUI Grounding via Region Consistency.Proceedings of the AAAI Conference Software Engineering for and with GUI Agent 1:41 on Artific...

  51. [61]

    Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim, Suhong Moon, Hiroki Furuta, Gopala Anumanchipalli, Kurt Keutzer, and Amir Gholami. 2025. Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks. arXiv:2503.09572 [cs.CL] doi:10.48550/arXiv.2503.09572

  52. [62]

    Ivan Evtimov, Arman Zharmagambetov, Aaron Grattafiori, Chuan Guo, and Kamalika Chaudhuri. 2025. WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks. arXiv:2504.18575 [cs.CR] doi:10.48550/arXiv. 2504.18575

  53. [63]

    Difei Gao, Siyuan Hu, Zechen Bai, Qinghong Lin, and Mike Zheng Shou. 2024. AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation. InProceedings of the 32nd ACM International Conference on Multimedia. 11255–11257. doi:10.1145/3664647.3684998

  54. [65]

    Moghis Fereidouni, Adib Mosharrof, and A. B. Siddique. 2024. Grounded Language Agent for Product Search via Intelligent Web Interactions. arXiv:2404.10887 [cs.CL] doi:10.48550/arXiv.2404.10887

  55. [66]

    Hiroki Furuta, Kuang-Huei Lee, Ofir Nachum, Yutaka Matsuo, Aleksandra Faust, Shixiang Shane Gu, and Izzeddin Gur. 2023. Multimodal Web Navigation with Instruction-Finetuned Foundation Models. arXiv:2305.11854 [cs.LG] doi:10.48550/arXiv.2305.11854

  56. [67]

    Hiroki Furuta, Yutaka Matsuo, Aleksandra Faust, and Izzeddin Gur. 2023. Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web. arXiv:2311.18751 [cs.LG] doi:10.48550/arXiv.2311.18751

  57. [68]

    Yu Gu, Kai Zhang, Yuting Ning, Boyuan Zheng, Boyu Gou, Tianci Xue, Cheng Chang, Sanjari Srivastava, Yanan Xie, Peng Qi, Huan Sun, and Yu Su. 2024. Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents. arXiv:2411.06559 [cs.AI] doi:10.48550/arX...

  58. [70]

    Longxi Gao, Li Zhang, Shihe Wang, Pengzhi Gao, Wei Liu, Jian Luan, Shangguang Wang, Yuanchun Li, and Mengwei Xu. 2024. MobileViews: A Million-scale and Diverse Mobile GUI Dataset. arXiv:2409.14337 [cs.HC] doi:10.48550/ arXiv.2409.14337

  59. [71]

    Divyansh Garg, Shaun VanWeelden, Diego Caples, Andis Draguns, Nikil Ravi, Pranav Putta, Naman Garg, Tomas Abraham, Michael Lara, Federico Lopez, James Liu, Atharva Gundawar, Prannay Hebbar, Youngchul Joo, Jindong Gu, Charles London, Christian Schroeder de Witt, and Sumeet Motw...

  60. [72]

    Boyu Gou, Ruohan Wang, Boyuan Zheng, Yanan Xie, Cheng Chang, Yiheng Shu, Huan Sun, and Yu Su. 2024. Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents. arXiv:2410.05243 [cs.AI] doi:10.48550/arXiv.2410.05243

  61. [73]

    Wenkang Han, Zhixiong Zeng, Jing Huang, Shu Jiang, Liming Zheng, Longrong Yang, Haibo Qiu, Chang Yao, Jingyuan Chen, and Lin Ma. 2025. UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions. arXiv:2506.11127 [cs.CL] doi:10.48550/arXiv.2506.11127

  62. [74]

    Chao Hao, Shuai Wang, and Kaiwen Zhou. 2025. Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement. arXiv:2508.04025 [cs.AI] doi:10.48550/arXiv.2508. 1:42 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong...

  63. [75]

    Ziyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, and Ngai Wong. 2025. KG- RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven ...

  64. [76]

    Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent. arXiv:2511.06417 [cs.AI] doi:10.48550/arXiv.2511.06417

  65. [77]

    Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023. A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis. arXiv:2307.12856 [cs.LG] doi:10.48550/arXiv.2307.12856

  66. [78]

    Fung, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu

    Zhitao He, Zijun Liu, Peng Li, Yi R. Fung, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu. 2025. Advanc- ing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization. arXiv:2502.14496 [cs.CL] doi:10.48550/arXiv.2502.14496

  67. [79]

    Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. 2024. CogAgent: A Visual Language Model for GUI Agents. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14281–142...

  68. [80]

    Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...

  69. [81]

    Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, and Dong Yu. 2025. OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization. InProceedings of the 63rd Annual Meeting of the Asso...

  70. [82]

    Yanheng He, Jiahe Jin, Shijie Xia, Jiadi Su, Runze Fan, Haoyang Zou, Xiangkun Hu, and Pengfei Liu. 2024. PC Agent: While You Sleep, AI Works – A Cognitive Journey into Digital World. arXiv:2412.17589 [cs.AI] doi:10.48550/arXiv. 2412.17589

  71. [83]

    Zhiyuan Hu, Shiyun Xiong, Yifan Zhang, See-Kiong Ng, Anh Tuan Luu, Bo An, Shuicheng Yan, and Bryan Hooi

  72. [84]

    Jing Huang, Zhixiong Zeng, Wenkang Han, Yufeng Zhong, Liming Zheng, Shuai Fu, Jingyuan Chen, and Lin Ma. 2025. ScaleTrack: Scaling and back-tracking Automated GUI Agents. arXiv:2505.00416 [cs.AI] doi:10.48550/arXiv.2505.00416

  73. [85]

    Jakub Hoscilowicz, Bartosz Maj, Bartosz Kozakiewicz, Oleksii Tymoshchuk, and Artur Janicki. 2024. ClickAgent: Enhancing UI Location Capabilities of Autonomous Agents. arXiv:2410.11872 [cs.HC] doi:10.48550/arXiv.2410.11872

  74. [86]

    Siyuan Hu, Mingyu Ouyang, Difei Gao, and Mike Zheng Shou. 2024. The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use. arXiv:2411.10323 [cs.AI] doi:10.48550/arXiv.2411.10323

  75. [87]

    Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shawn Wang, Xinchen Xu, Shuofei Qiao, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wan...

  76. [88]

    Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang, Weiqi Sun, and Yuanchun Shi. 2025. Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts.ACM Transactions on Computer-Human Interaction32, 3 (2025), 1–45. doi:10.1145/3716132

  77. [90]

    Zheng Hui, Yinheng Li, Dan zhao, Tianyi Chen, Colby Banbury, and Kazuhito Koishida. 2025. WinClick: GUI Grounding with Multimodal Large Language Models. arXiv:2503.04730 [cs.CL] doi:10.48550/arXiv.2503.04730

  78. [91]

    Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai, Caiming Xiong, and Chien-Sheng Wu. 2025. GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness. arXiv:2510.00536 [cs.CL] doi:10.48550/arXiv.2510.00536

  79. [92]

    Tenghao Huang, Kinjal Basu, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, and Muhao Chen. 2025. R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory. InProceedings of the 63rd Annual Meeting of the Association for Computational Lingui...

  80. [93]

    Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang, Weiqi Sun, and Yuanchun Shi. 2024. PromptRPA: Generating Robotic Process Automation on Smartphones from Textual Prompts. arXiv:2404.02475 [cs.HC] doi:10. 48550/arXiv.2404.02475

  81. [94]

    Wenjia Jiang, Yangyang Zhuang, Chenxi Song, Xu Yang, Joey Tianyi Zhou, and Chi Zhang. 2025. AppAgentX: Evolving GUI Agents as Proficient Smartphone Users. arXiv:2503.02268 [cs.AI] doi:10.48550/arXiv.2503.02268

  82. [95]

    Yiqiao Jin, Stefano Petrangeli, Yu Shen, and Gang Wu. 2025. <scp>ScreenLLM:</scp> Stateful Screen Schema for Efficient Action Understanding and Prediction. InCompanion Proceedings of the ACM on Web Conference 2025. 2008–2013. doi:10.1145/3701716.3718379

  83. [96]

    Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, and Ruslan Salakhutdinov. 2024. OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.Lecture Notes in Computer Science(2024), 161–17...

  84. [97]

    Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang

  85. [98]

    InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)

    OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 72–81. doi:10.18653/v1/2024.acl-demos.8

  86. [99]

    Lawrence Jang, Yinheng Li, Dan Zhao, Charles Ding, Justin Lin, Paul Pu Liang, Rogerio Bonatti, and Kazuhito Koishida. 2024. VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks. arXiv:2410.19100 [cs.CV] doi:10.48550/arXiv.2410.19100 Softw...

  87. [100]

    Chengyou Jia, Minnan Luo, Zhuohang Dang, Qiushi Sun, Fangzhi Xu, Junlin Hu, Tianbao Xie, and Zhiyong Wu. 2025. AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer Assistant. InFindings of the Association for Computational Linguistics: AC...

  88. [101]

    SeokJoo Kwak, Jihoon Kim, Boyoun Kim, Jung Jae Yoon, Wooseok Jang, Jeonghoon Hong, Jaeho Yang, and Yeong-Dae Kwon. 2025. MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements. arXiv:2511.13087 [cs.AI] doi:10.48550/arXiv.2511.13087

  89. [103]

    Bradley Knox, and Kimin Lee

    Juyong Lee, Dongyoon Hahm, June Suk Choi, W. Bradley Knox, and Kimin Lee. 2026. MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control.Proceedings of the AAAI Conference on Artificial Intelligence 40, 44 (2026), 37565–37573. doi:10.1609/aaai.v40i44.41090

  90. [104]

    Su Kara, Fazle Faisal, and Suman Nath. 2025. WABER: Evaluating Reliability and Efficiency of Web Agents with Existing Benchmarks. InICLR Workshop on Foundation Models in the Wild. https://www.microsoft.com/en-us/ research/publication/waber-evaluating-reliability-and-efficiency...

  91. [105]

    Jihyung Kil, Chan Hee Song, Boyuan Zheng, Xiang Deng, Yu Su, and Wei-Lun Chao. 2024. Dual-View Visual Contextualization for Web Navigation. arXiv:2402.04476 [cs.CV] doi:10.48550/arXiv.2402.04476

  92. [106]

    Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks. InProceedings of the 62nd Annual Meeting of the Asso...

  93. [107]

    Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. 2024. Tree Search for Language Model Agents. arXiv:2407.01476 [cs.AI] doi:10.48550/arXiv.2407.01476

  94. [108]

    Hongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen, Li Qing, and Zhaoxiang Zhang. 2025. AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  95. [109]

    Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li, and Zhaoxiang Zhang. 2025. UIPro: Unleashing Superior Interaction Capability For GUI Agents. arXiv:2509.17328 [cs.CV] doi:10.48550/arXiv.2509.17328

  96. [110]

    Jiahao Li and Kaer Huang. 2025. A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement Learning. arXiv:2504.20464 [cs.AI] doi:10.48550/arXiv.2504.20464

  97. [111]

    Jungjae Lee, Dongjae Lee, Chihun Choi, Youngmin Im, Jaeyoung Wi, Kihong Heo, Sangeun Oh, Sunjae Lee, and Insik Shin. 2025. VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification. InProceedings of the 31st Annual International Conference on Mobile Com...

  98. [112]

    Juyong Lee, Taywon Min, Minyong An, Dongyoon Hahm, Haeone Lee, Changyeon Kim, and Kimin Lee. 2024. Benchmarking Mobile Device Control Agents across Diverse Configurations. arXiv:2404.16660 [cs.HC] doi:10.48550/ arXiv.2404.16660

  99. [113]

    Sunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi, Hojun Choi, Steve Ko, Sangeun Oh, and Insik Shin. 2024. MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation. InProceedings of the 30th Annual International Conference on Mobile Computing and ...

  100. [114]

    Ido Levy, Ben Wiesel, Sami Marreed, Alon Oved, Avi Yaeli, and Segev Shlomov. 2024. ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents. arXiv:2410.06703 [cs.AI] doi:10.48550/arXiv. 2410.06703

  101. [115]

    Wei Li, Fu-Lin Hsu, William Bishop, Folawiyo Campbell-Ajala, Max Lin, and Oriana Riva. 2024. UINav: A Practical Approach to Train On-Device Automation Agents. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: H...

  102. [116]

    Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020. Mapping Natural Language Instructions to Mobile UI Action Sequences. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 8198–8210. doi:10.18653/v1/2020.acl-main.729

  103. [117]

    Yanda Li, Chi Zhang, Wenjia Jiang, Wanqi Yang, Bin Fu, Pei Cheng, Xin Chen, Ling Chen, and Yunchao Wei. 2024. AppAgent v2: Advanced Agent for Flexible Mobile Interactions. arXiv:2408.11824 [cs.HC] doi:10.48550/arXiv.2408. 11824

  104. [118]

    Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua

  105. [119]

    InProceedings of the 33rd ACM International Conference on Multimedia

    ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. InProceedings of the 33rd ACM International Conference on Multimedia. 8778–8786. doi:10.1145/3746027.3755688

  106. [120]

    Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li. 2023. A Zero-Shot Language Agent for Computer Control with Structured Reflection. InFindings of the Association for Computational Linguistics: EMNLP 2023. 11261–11274. doi:10.18653/v1/2023.findings-emnlp.753 1:44 Shengchen...

  107. [123]

    Wei Li, William Bishop, Alice Li, Chris Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva

  108. [124]

    InAdvances in Neural Information Processing Systems 37

    On the Effects of Data Scale on UI Control Agents. InAdvances in Neural Information Processing Systems 37. 92130–92154. doi:10.52202/079017-2925

  109. [125]

    Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, WenHao Wang, Tianze Wu, Zhengxi Lu, Siheng Chen, LiLinghao, Hao Wang, Guanjing Xiong, Yong Liu, and Hongsheng Li. 2025. LLM-Power...

  110. [126]

    Guangyi Liu, Pengxiang Zhao, Liang Liu, Zhiming Chen, Yuxiang Chai, Shuai Ren, Hao Wang, Shibo He, and Wenchao Meng. 2025. LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark. arXiv:2504.13805 [cs.HC] doi:10.48550/arXiv.2504.13805

  111. [127]

    Haowei Liu, Xi Zhang, Haiyang Xu, Yuyang Wanyan, Junyang Wang, Ming Yan, Ji Zhang, Chunfeng Yuan, Changsheng Xu, Weiming Hu, and Fei Huang. 2025. PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC. arXiv:2502.14282 [cs.CV] doi:10.485...

  112. [128]

    Zhangheng Li, Keen You, Haotian Zhang, Di Feng, Harsh Agrawal, Xiujun Li, Mohana Prasad Sathya Moorthy, Jeff Nichols, Yinfei Yang, and Zhe Gan. 2024. Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms. arXiv:2410.18967 [cs.CV] doi:10.48550/arXiv.2410.18967

  113. [129]

    Shuquan Lian, Yuhang Wu, Jia Ma, Yifan Ding, Zihan Song, Bingqi Chen, Xiawu Zheng, Hui Li, and Rongrong Ji. 2025. UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Grounding. arXiv:2507.22025 [cs.AI] doi:10.48550/arXiv.2507.22025

  114. [130]

    Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun

  115. [131]

    arXiv:2409.11295 [cs.CR] doi:10.48550/arXiv.2409.11295

    EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage. arXiv:2409.11295 [cs.CR] doi:10.48550/arXiv.2409.11295

  116. [132]

    Kevin Lin, Linjie Li, Difei Gao, Qinchen Wu, Mingyi Yan, Zhengyuan Yang, Lijuan Wang, and Mike Shou. 2024. VideoGUI: A Benchmark for GUI Automation from Instructional Videos. InAdvances in Neural Information Processing Systems 37. 69329–69360. doi:10.52202/079017-2214

  117. [134]

    Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. 2018. Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration. arXiv:1802.08802 [cs.AI] doi:10.48550/arXiv.1802.08802

  118. [135]

    Guohong Liu, Jialei Ye, Jiacheng Liu, Yuanchun Li, Wei Liu, Pengzhi Gao, Jian Luan, and Yunxin Liu. 2025. Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties. InProceedings of the 2nd International Workshop on Edge and Mobile Foundation Models. ...

  119. [136]

    Ziwei Liu, Borui Kang, Hangjie Yuan, Zixiang Zhao, Wei Li, Yifan Zhu, and Tao Feng. 2026. Continual GUI Agents. arXiv:2601.20732 [cs.LG] doi:10.48550/arXiv.2601.20732

  120. [137]

    Fanbin Lu, Zhisheng Zhong, Shu Liu, Chi-Wing Fu, and Jiaya Jia. 2025. ARPO: End-to-End Policy Optimization for GUI Agents with Experience Replay. arXiv:2505.16282 [cs.CV] doi:10.48550/arXiv.2505.16282

  121. [138]

    Fanbin Lu, Zhisheng Zhong, Ziqin Wei, Shu Liu, Chi-Wing Fu, and Jiaya Jia. 2025. STEVE: A Step Verification Pipeline for Computer-use Agent Training. arXiv:2503.12532 [cs.CV] doi:10.48550/arXiv.2503.12532

  122. [139]

    Jiarun Liu, Jia Hao, Chunhong Zhang, and Zheng Hu. 2025. WEPO: Web Element Preference Optimization for LLM-based Web Navigation.Proceedings of the AAAI Conference on Artificial Intelligence39, 25 (2025), 26614–26622. doi:10.1609/aaai.v39i25.34863

  123. [140]

    Junpeng Liu, Yifan Song, Bill Yuchen Lin, Wai Lam, Graham Neubig, Yuanzhi Li, and Xiang Yue. 2024. VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding? arXiv:2404.05955 [cs.CL] doi:10.48550/arXiv.2404.05955

  124. [141]

    Xiao Liu, Bo Qin, Dongzhu Liang, Guang Dong, Hanyu Lai, Hanchen Zhang, Hanlin Zhao, Iat Long Iong, Jiadai Sun, Jiaqi Wang, Junjie Gao, Junjun Shan, Kangning Liu, Shudan Zhang, Shuntian Yao, Siyi Cheng, Wentao Yao, Wenyi Zhao, Xinghan Liu, Xinyi Liu, Xinying Chen, Xinyue Yang, ...

  125. [142]

    Xiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong, Yifan Xu, Xixuan Song, Shudan Zhang, Hanyu Lai, Xinyi Liu, Hanlin Zhao, Jiadai Sun, Xinyue Yang, Yu Yang, Zehan Qi, Shuntian Yao, Xueqiao Sun, Siyi Cheng, Qinkai Zheng, Hao Yu, Hanchen Zhang, Wenyi Hong, Ming Ding, Lihang Pan, Xia...

  126. [143]

    Xinyi Liu, Xiaoyi Zhang, Ziyun Zhang, and Yan Lu. 2025. UI-E2I-Synth: Advancing GUI Grounding with Large- Scale Instruction Synthesis. InFindings of the Association for Computational Linguistics: ACL 2025. 15668–15684. doi:10.18653/v1/2025.findings-acl.809

  127. [144]

    Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, and Fei Wu. 2026. InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection. In Proceedings of the 19th Conference of the European Chap...

  128. [145]

    Yuhang Liu, Pengxiang Li, Congkai Xie, Xavier Hu, Xiaotian Han, Shengyu Zhang, Hongxia Yang, and Fei Wu. 2025. InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners. arXiv:2504.14239 [cs.AI] doi:10.48550/arXiv.2504.14239

  129. [146]

    Yuxuan Liu, Hongda Sun, Wei Liu, Jian Luan, Bo Du, and Rui Yan. 2025. MobileSteward: Integrating Multiple App- Oriented Agents with Self-Evolution to Automate Cross-App Instructions. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1. 88...

  130. [147]

    Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue. arXiv:2402.05930 [cs.CL] doi:10.48550/arXiv.2402.05930

  131. [148]

    Pal, and Siva Reddy

    Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, and Siva Reddy. 2025. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. arXiv:2504.08942 [cs.LG] ...

  132. [149]

    Yadong Lü, Jianwei Yang, Yelong Shen, and Ahmed Hassan Awadallah. 2024. OmniParser for Pure Vision Based GUI Agent. arXiv:2408.00203 [cs.CV] doi:10.48550/arXiv.2408.00203 1:46 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, and Zhenyu Chen

  133. [150]

    Quanfeng Lu, Wenqi Shao, Zitao Liu, Lingxiao Du, Fanqing Meng, Boxuan Li, Botong Chen, Siyuan Huang, Kaipeng Zhang, and Ping Luo. 2024. GUIOdyssey: A Comprehensive Dataset for Cross-App GUI Navigation on Mobile Devices. arXiv:2406.08451 [cs.CV] doi:10.48550/arXiv.2406.08451

  134. [151]

    Yijie Lu, Tianjie Ju, Manman Zhao, Xinbei Ma, Yuan Guo, and Zhuosheng Zhang. 2025. EVA: Red-Teaming GUI Agents via Evolving Indirect Prompt Injection. arXiv:2505.14289 [cs.AI] doi:10.48550/arXiv.2505.14289

  135. [152]

    Yuheng Lu, Qian Yu, Hongru Wang, Zeming Liu, Wei Su, Yanping Liu, Yuhang Guo, Maocheng Liang, Yunhong Wang, and Haifeng Wang. 2025. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments. InFindings of the Association for...

  136. [154]

    Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Pengxiang Zhao, Guangyi Liu, Guanjing Xiong, and Hongsheng Li. 2026. UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.Proceedings of the AAAI Conference ...

  137. [155]

    Dezhao Luo, Bohan Tang, Kang Li, Georgios Papoudakis, Jifei Song, Shaogang Gong, Jianye Hao, Jun Wang, and Kun Shao. 2025. ViMo: A Generative Visual GUI World Model for App Agents. arXiv:2504.13936 [cs.HC] doi:10.48550/ arXiv.2504.13936

  138. [156]

    Run Luo, Lu Wang, Wanwei He, Longze Chen, Jiaming Li, and Xiaobo Xia. 2025. GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents. arXiv:2504.10458 [cs.CV] doi:10.48550/arXiv.2504.10458

  139. [157]

    Yibo Lyu, Gongwei Chen, Rui Shao, Weili Guan, and Liqiang Nie. 2026. PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records. arXiv:2601.09636 [cs.AI] doi:10.48550/arXiv.2601.09636

  140. [158]

    Runliang Niu, Jindong Li, Shiqi Wang, Yali Fu, Xiyu Hu, Xueyuan Leng, He Kong, Yi Chang, and Qi Wang. 2024. ScreenAgent: A Vision Language Model-driven Computer Control Agent. arXiv:2402.07945 [cs.HC] doi:10.48550/ arXiv.2402.07945

  141. [159]

    Songqin Nong, Xiaoxuan Tang, Jingxuan Xu, Sheng Zhou, Jianfeng Chen, Tao Jiang, and Wenhao Xu. 2025. CRAFT- GUI: Curriculum-Reinforced Agent For GUI Tasks. arXiv:2508.11360 [cs.AI] doi:10.48550/arXiv.2508.11360

  142. [160]

    Songqin Nong, Jiali Zhu, Rui Wu, Jiongchao Jin, Shuo Shan, Xiutian Huang, and Wenhao Xu. 2024. MobileFlow: A Multimodal LLM For Mobile GUI Agent. arXiv:2407.04346 [cs.CV] doi:10.48550/arXiv.2407.04346

  143. [161]

    Kaixin Ma, Hongming Zhang, Hongwei Wang, Xiaoman Pan, Wenhao Yu, and Dong Yu. 2023. LASER: LLM Agent with State-Space Exploration for Web Navigation. arXiv:2309.08172 [cs.CL] doi:10.48550/arXiv.2309.08172

  144. [162]

    Longhui Ma, Di Zhao, Siwei Wang, Zhao Lv, and Miao Wang. 2026. Beyond element-level understanding: Explicit relational understanding for GUI agents.Pattern Recognition176 (2026), 113262. doi:10.1016/j.patcog.2026.113262

  145. [164]

    Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, and Hai Zhao. 2025. Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions. InProceedings of the 63rd Annual Meeting of the Association for Computational Ling...

  146. [165]

    Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024. CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation. InFindings of the Association for Computational Linguistics ACL 2024. 9097–9110. doi:10.18653/v1/2024.findings-acl.539

  147. [166]

    Shikhar Murty, Hao Zhu, Dzmitry Bahdanau, and Christopher D. Manning. 2024. NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild. arXiv:2410.02907 [cs.CL] doi:10.48550/arXiv.2410. 02907

  148. [167]

    Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapa- dos, M

    Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapa- dos, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI-Vision: A Desktop-centric GUI Bench...

  149. [168]

    Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A

    Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, ...

  150. [169]

    Zehan Qi, Xiao Liu, Iat Long Iong, Hanyu Lai, Xueqiao Sun, Wenyi Zhao, Yu Yang, Xinyue Yang, Jiadai Sun, Shuntian Yao, Tianjie Zhang, Wei Xu, Jie Tang, and Yuxiao Dong. 2024. WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning. arXiv:2411....

  151. [170]

    Yijun Qian, Yujie Lu, Alexander Hauptmann, and Oriana Riva. 2024. Visual Grounding for User Interfaces. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track)....

  152. [171]

    Yujia Qin, Yining Ye, Junjie Fang, Haoming Wang, Shihao Liang, Shizuo Tian, Junda Zhang, Jiahao Li, Yunxin Li, Shijue Huang, Wanjun Zhong, Kuanye Li, Jiale Yang, Yu Miao, Woyu Lin, Longxiang Liu, Xu Jiang, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, Yaowei Zheng, C...

  153. [172]

    Vardaan Pahuja, Yadong Lu, Corby Rosset, Boyu Gou, Arindam Mitra, Spencer Whitehead, Yu Su, and Ahmed Hassan Awadallah. 2025. Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents. In Findings of the Association for Computational Linguistics: ...

  154. [173]

    Lihang Pan, Bowen Wang, Chun Yu, Yuxuan Chen, Xiangyu Zhang, and Yuanchun Shi. 2023. AutoTask: Executing Arbitrary Voice Commands by Exploring and Learning from Mobile GUI. arXiv:2312.16062 [cs.HC] doi:10.48550/ arXiv.2312.16062

  155. [174]

    Yichen Pan, Dehan Kong, Sida Zhou, Cheng Cui, Yifei Leng, Bing Jiang, Hangyu Liu, Yanyi Shang, Shuyan Zhou, Tongshuang Wu, and Zhengyang Wu. 2024. WebCanvas: Benchmarking Web Agents in Online Environments. arXiv:2406.12373 [cs.CL] doi:10.48550/arXiv.2406.12373

  156. [175]

    Georgios Papoudakis, Thomas Coste, Zhihao Wu, Jianye Hao, Jun Wang, and Kun Shao. 2025. AppVLM: A Lightweight Vision Language Model for Online App Control. arXiv:2502.06395 [cs.AI] doi:10.48550/arXiv.2502.06395

  157. [176]

    Manmatha, and Shabnam Ghadar

    Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju, Kunwar Yashraj Singh, R. Manmatha, and Shabnam Ghadar. 2025. R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding. InFindings of the Association for Computational Linguistics: ACL 2025. 9669–9685. doi:10....

  158. [177]

    Pawel Pawlowski, Krystian Zawistowski, Wojciech Lapacz, Adam Wiacek, Marcin Skorupa, Sebastien Postansque, and Jakub Hoscilowicz. 2025. TinyClick: Single-Turn Agent for Empowering GUI Automation. InInterspeech 2025. 3035–3039. doi:10.21437/interspeech.2025-176

  159. [178]

    Bigham, and Amy Pavel

    Yi-Hao Peng, Faria Huq, Yue Jiang, Jason Wu, Xin Yue Li, Jeffrey P. Bigham, and Amy Pavel. 2024. DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation.Lecture Notes in Computer Science(2024), 466–485. doi:10.1007/978-3-031-72691-0_26

  160. [179]

    Pranav Putta, Edmund Mills, Naman Garg, Sumeet Motwani, Chelsea Finn, Divyansh Garg, and Rafael Rafailov. 2024. Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents. arXiv:2408.07199 [cs.AI] doi:10.48550/arXiv. 2408.07199 Software Engineering for and with GUI Agent 1:47

  161. [180]

    Huawen Shen, Chang Liu, Gengluo Li, Xinlong Wang, Yu Zhou, Can Ma, and Xiangyang Ji. 2024. Falcon-UI: Understanding GUI Before Following User Instructions. arXiv:2412.09362 [cs.CL] doi:10.48550/arXiv.2412.09362

  162. [181]

    Junhong Shen, Atishay Jain, Zedian Xiao, Ishan Amlekar, Mouad Hadji, Aaron Podolny, and Ameet Talwalkar. 2024. ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data. arXiv:2411.15004 [cs.CL] doi:10.48550/arXiv.2411.15004

  163. [182]

    Yucheng Shi, Wenhao Yu, Jingyuan Huang, Wenlin Yao, Wenhu Chen, and Ninghao Liu. 2025. Towards Trustworthy GUI Agents: A Survey. arXiv:2503.23434 [cs.LG] doi:10.48550/arXiv.2503.23434

  164. [183]

    Haoyi Qiu, Alexander Fabbri, Divyansh Agarwal, Kung-Hsiang Huang, Sarah Tan, Nanyun Peng, and Chien-Sheng Wu

  165. [184]

    InFindings of the Association for Computational Linguistics: NAACL 2025

    Evaluating Cultural and Social Awareness of LLM Web Agents. InFindings of the Association for Computational Linguistics: NAACL 2025. 3978–4005. doi:10.18653/v1/2025.findings-naacl.222

  166. [185]

    Abdur Rahman, Rajat Chawla, Muskaan Kumar, Arkajit Datta, Adarsh Jha, Mukunda NS, and Ishaan Bhola. 2024. V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM. arXiv:2405.15341 [cs.AI] doi:10.48550/arXiv.2405.15341

  167. [186]

    Dezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu, Yuan Cao, Ying Zhang, Wei Yang, and Tao Xie. 2024. Guardian: A Runtime Framework for LLM-Based UI Exploration. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 958–970. doi:10.1145/3650...

  168. [187]

    Dezhi Ran, Mengzhou Wu, Hao Yu, Yuetong Li, Jun Ren, Yuan Cao, Xia Zeng, Haochuan Lu, Zexin Xu, Mengqian Xu, Ting Su, Liangchao Yao, Ting Xiong, Wei Yang, Yuetang Deng, Assaf Marron, David Harel, and Tao Xie. 2025. Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundat...

  169. [188]

    Christopher Rawles, Sarah Clinckemaillie, Yifan Chang, Jonathan Waltz, Gabrielle Lau, Marybeth Fair, Alice Li, William Bishop, Wei Li, Folawiyo Campbell-Ajala, Daniel Toyama, Robert Berry, Divya Tyamagundlu, Timothy Lillicrap, and Oriana Riva. 2024. AndroidWorld: A Dynamic Ben...

  170. [189]

    Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lillicrap. 2023. Android in the Wild: A Large-Scale Dataset for Android Device Control. arXiv:2307.10088 [cs.LG] doi:10.48550/arXiv.2307.10088

  171. [190]

    Mobina Shahbandeh, Parsa Alian, Noor Nashid, and Ali Mesbah. 2024. NaviQAte: Functionality-Guided Web Application Navigation. arXiv:2409.10741 [cs.SE] doi:10.48550/arXiv.2409.10741

  172. [191]

    Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces. InAdvances in Neural Information Proc...

  173. [192]

    Jiahui Sun, Zhichao Hua, and Yubin Xia. 2025. AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents. arXiv:2503.02403 [cs.AI] doi:10.48550/arXiv.2503.02403

  174. [193]

    Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu. 2022. META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI. (2022), 6699–6712. doi:10.18653/v1/2022.emnlp-main.449

  175. [194]

    Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2024. OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis...

  176. [195]

    Yucheng Shi, Wenhao Yu, Zaitang Li, Yong-Lin Wang, Hongming Zhang, Ninghao Liu, Haitao Mi, and Dong Yu

  177. [196]

    arXiv:2507.05720 [cs.LG] doi:10.48550/arXiv.2507.05720

    MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment. arXiv:2507.05720 [cs.LG] doi:10.48550/arXiv.2507.05720

  178. [197]

    Kunal Singh, Shreyas Singh, and Mukund Khanna. 2025. Trishul: Towards Region Identification and Screen Hierarchy Understanding for Large VLM Based GUI Agents. In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 170–179. doi:10.1109/cvprw673...

  179. [198]

    Paloma Sodhi, S. R. K. Branavan, Yoav Artzi, and Ryan McDonald. 2023. SteP: Stacked LLM Policies for Web Actions. arXiv:2310.03720 [cs.LG] doi:10.48550/arXiv.2310.03720

  180. [199]

    Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. VisionTasker: Mobile Task Au- tomation Using Vision Based UI Understanding and LLM Task Planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17. doi:10.1...

  181. [200]

    Yixiao Song, Katherine Thai, Chau Minh Pham, Yapei Chang, Mazin Nadaf, and Mohit Iyyer. 2025. BEARCUBS: A benchmark for computer-using web agents. arXiv:2503.07919 [cs.AI] doi:10.48550/arXiv.2503.07919

  182. [201]

    Yueqi Song, Frank Xu, Shuyan Zhou, and Graham Neubig. 2024. Beyond Browsing: API-Based Web Agents. arXiv:2410.16464 [cs.CL] doi:10.48550/arXiv.2410.16464 1:48 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, and Zhenyu Chen

  183. [202]

    Zirui Song, Yaohang Li, Meng Fang, Zhenhao Chen, Zecheng Shi, Yuan Huang, and Ling Chen. 2024. MMAC-Copilot: Multi-modal Agent Collaboration Operating System Copilot. arXiv:2404.18074 [cs.AI] doi:10.48550/arXiv.2404.18074

  184. [203]

    Trisanth Srinivasan and Santosh Patapati. 2025. WebNav: An Intelligent Agent for Voice-Controlled Web Navigation. arXiv:2503.13843 [cs.AI] doi:10.48550/arXiv.2503.13843

  185. [204]

    Hongjin Su, Ruoxi Sun, Jinsung Yoon, Pengcheng Yin, Tao Yu, and Sercan Ö. Arık. 2025. Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments. arXiv:2501.10893 [cs.LG] doi:10.48550/ arXiv.2501.10893

  186. [205]

    Xingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang, and Jing Tang. 2025. Understanding GUI Agent Localization Biases through Logit Sharpness. InFindings of the Association for Computational Linguistics: EMNLP 2025. 23361–23374. doi:10.18653/v1/2025.findings-emnlp.1268

  187. [206]

    Lucas-Andrei Thil, Mirela Popa, and Gerasimos Spanakis. 2024. Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning. InProceedings of the 39th ACM/SIGAPP Symposium on Applied Computing. 866–874. doi:10.1145/3605098.3635903

  188. [207]

    Chan, Jikun Kang, Wenqi Wu, Filippos Christianos, Fraser Greenlee, Andy Toulis, and Marvin Purtorab

    George Thomas, Alex J. Chan, Jikun Kang, Wenqi Wu, Filippos Christianos, Fraser Greenlee, Andy Toulis, and Marvin Purtorab. 2025. WebGames: Challenging General-Purpose Web-Browsing AI Agents. arXiv:2502.18356 [cs.LG] doi:10.48550/arXiv.2502.18356 Software Engineering for and w...

  189. [208]

    Qiushi Sun, Zhoumianze Liu, Chang Ma, Zichen Ding, Fangzhi Xu, Zhangyue Yin, Haiteng Zhao, Zhenyu Wu, Kanzhi Cheng, Zhaoyang Liu, Jianing Wang, Qintong Li, Xiangru Tang, Tianbao Xie, Xiachong Feng, Xiang Li, Ben Kao, Wenhai Wang, Biqing Qi, Lingpeng Kong, and Zhiyong Wu. 2025....

  190. [210]

    Srinivas Sunkara, Maria Wang, Lijuan Liu, Gilles Baechler, Yu-Chung Hsiao, Jindong, Chen, Abhanshu Sharma, and James Stout. 2022. Towards Better Semantic Understanding of Mobile Interfaces. arXiv:2210.02663 [cs.HC] doi:10.48550/arXiv.2210.02663

  191. [211]

    Karlsson, Bo An, Shuicheng Yan, and Zongqing Lu

    Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li, Bohan Zhou, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Yujie Wu, Xiaoqiang Chai, Yifei Bi, Tian- bao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Long Tian, Chaojie ...

  192. [212]

    Brian Tang and Kang G. Shin. 2024. Steward: Natural Language Web Automation. arXiv:2409.15441 [cs.AI] doi:10. 48550/arXiv.2409.15441

  193. [213]

    Fei Tang, Yongliang Shen, Hang Zhang, Siqi Chen, Guiyang Hou, Wenqi Zhang, Wenqiao Zhang, Kaitao Song, Weiming Lu, and Yueting Zhuang. 2025. Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems. arXiv:2503.06470 [cs.AI] doi:10.48550/arXiv.2503.06470

  194. [214]

    Fei Tang, Haolei Xu, Hang Zhang, Siqi Chen, Xingyu Wu, Yongliang Shen, Wenqi Zhang, Guiyang Hou, Zeqi Tan, Yuchen Yan, Kaitao Song, Jian Shao, Weiming Lu, Jun Xiao, and Yueting Zhuang. 2025. A Survey on (M)LLM-Based GUI Agents. arXiv:2504.13865 [cs.HC] doi:10.48550/arXiv.2504.13865

  195. [215]

    Jiaqi Tang, Yu Xia, Yi-Feng Wu, Yuwei Hu, Yuhui Chen, Qing-Guo Chen, Xiaogang Xu, Xiangyu Wu, Hao Lu, Yanqing Ma, Shiyin Lu, and Qifeng Chen. 2025. LPO: Towards Accurate GUI Agent Interaction via Location Preference Optimization. arXiv:2506.09373 [cs.LG] doi:10.48550/arXiv.2506.09373

  196. [216]

    Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, Zhihui Cao, Hailiang Pang, Heng Kong, He Yang, Mingxu Chai, Zhilin Gao, Xingyu Liu, Yingnan Fu, Jiaming Liu, Xuanjing Huang, Yu-Gang Jiang, Tao Gui, Qi Zhang, Kang Wang, Yunke Z...

  197. [217]

    Sizhe Tang, Rongqian Chen, and Tian Lan. 2026. Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents. arXiv:2602.02995 [cs.AI] doi:10.48550/arXiv.2602.02995

  198. [218]

    Ke Wang, Tianyu Xia, Zhangxuan Gu, Yi Zhao, Shuheng Shen, Changhua Meng, Weiqiang Wang, and Ke Xu. 2024. E-ANT: A Large-Scale Dataset for Efficient Automatic GUI NavigaTion. arXiv:2406.14250 [cs.CV] doi:10.48550/arXiv. 2406.14250

  199. [219]

    Luyuan Wang, Yongyu Deng, Yiwei Zha, Guodong Mao, Qinmin Wang, Tianchen Min, Wei Chen, and Shoufa Chen

  200. [220]

    Shuai Wang, Weiwen Liu, J. C. Chen, Yuqi Zhou, Weinan Gan, X. Zeng, Yuhan Che, Shicheng Yu, Xinlong Hao, Shao Kun, Bin Wang, Chuhan Wu, Y.S. Wang, Ruiming Tang, and Jianye Hao. 2024. GUI Agents with Foundation Models: A Comprehensive Survey. arXiv:2411.04890 [cs.AI] doi:10.485...

  201. [221]

    Shizuo Tian, Hao Wen, Yuxuan Chen, Jiacheng Liu, Shanhui Zhao, Guohong Liu, Ju Ren, Yunxin Liu, and Yuanchun Li. 2025. AgentProg: Empowering Long-Horizon GUI Agents with Program-Guided Context Management. arXiv:2512.10371 [cs.AI] doi:10.48550/arXiv.2512.10371

  202. [223]

    Brandon Trabucco, Gunnar Sigurdsson, Robinson Piramuthu, and Ruslan Salakhutdinov. 2025. InSTA: Towards Internet-Scale Training For Agents. arXiv:2502.06776 [cs.LG] doi:10.48550/arXiv.2502.06776

  203. [224]

    Ada Defne Tur, Nicholas Meade, Xing Han Lù, Alejandra Zambrano, Arkil Patel, Esin Durmus, Spandana Gella, Karolina Stańczak, and Siva Reddy. 2025. SafeArena: Evaluating the Safety of Autonomous Web Agents. arXiv:2503.04957 [cs.LG] doi:10.48550/arXiv.2503.04957

  204. [225]

    Minh Duc Vu, Han Wang, Zhuang Li, Jieshan Chen, Shengdong Zhao, Zhenchang Xing, and Chunyang Chen. 2024. GPTVoiceTasker: LLM-Powered Virtual Assistant for Smartphone. arXiv:2401.14268 [cs.HC] doi:10.48550/arXiv.2401. 14268

  205. [226]

    Bryan Wang, Gang Li, and Yang Li. 2023. Enabling Conversational Interaction with Mobile UI using Large Language Models. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17. doi:10.1145/3544548. 3580895

  206. [227]

    Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu. 2026. Computer Agent Arena: Toward Human-Centri...

  207. [228]

    Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang

  208. [229]

    arXiv:2406.01014 [cs.CL] doi:10.48550/arXiv.2406.01014

    Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration. arXiv:2406.01014 [cs.CL] doi:10.48550/arXiv.2406.01014

  209. [230]

    Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile- Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception. arXiv:2401.16158 [cs.CL] doi:10. 48550/arXiv.2401.16158

  210. [231]

    Junyang Wang, Haiyang Xu, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Jitao Sang. 2025. Mobile-Agent-V: Learning Mobile Device Operation Through Video-Guided Multi-Agent Collaboration. arXiv:2502.17110 [cs.CL] doi:10.48550/arXiv.2502.17110

  211. [232]

    Zhenhailong Wang, Haiyang Xu, Junyang Wang, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Heng Ji. 2025. Mobile- Agent-E: Self-Evolving Mobile Assistant for Complex Tasks. arXiv:2501.11733 [cs.CL] doi:10.48550/arXiv.2501.11733

  212. [233]

    Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig, and Daniel Fried. 2025. Inducing Programmatic Skills for Agentic Tasks. arXiv:2504.06821 [cs.CL] doi:10.48550/arXiv.2504.06821

  213. [234]

    arXiv:2406.08184 [cs.AI] doi:10.48550/arXiv.2406.08184

    MobileAgentBench: An Efficient and User-Friendly Benchmark for Mobile LLM Agents. arXiv:2406.08184 [cs.AI] doi:10.48550/arXiv.2406.08184

  214. [235]

    Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. AutoDroid: LLM-powered Task Automation in Android. InProceedings of the 30th Annual International Conference on Mobile Computing and Networking...

  215. [236]

    Taiyi Wang, Zhihao Wu, Jianheng Liu, Jianye Hao, Jun Wang, and Kun Shao. 2024. DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents. arXiv:2410.14803 [cs.LG] doi:10. 48550/arXiv.2410.14803

  216. [237]

    Wenhao Wang, Zijie Yu, Rui Ye, Jianqing Zhang, Siheng Chen, and Yanfeng Wang. 2025. FedMABench: Benchmarking Mobile Agents on Decentralized Heterogeneous User Data. arXiv:2503.05143 [cs.AI] doi:10.48550/arXiv.2503.05143

  217. [238]

    WenHao Wang, Zijie Yu, Rui Ye, Jianqing Zhang, Guangyi Liu, Liang Liu, Siheng Chen, and Yanfeng Wang. 2025. FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Proces...

  218. [239]

    WenHao Wang, Mengying Yuan, Zijie Yu, Guangyi Liu, Rui Ye, Tian Jin, Siheng Chen, and Yanfeng Wang. 2025. MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users. InProceedings of the Fourth Workshop on Bridging Human-Computer Interact...

  219. [240]

    Xiaoqiang Wang and Bang Liu. 2024. OSCAR: Operating System Control via State-Aware Reasoning and Re-Planning. arXiv:2410.18963 [cs.AI] doi:10.48550/arXiv.2410.18963

  220. [241]

    Xuehui Wang, Zhenyu Wu, JingJing Xie, Zichen Ding, Bowen Yang, Zehao Li, Zhaoyang Liu, Qingyun Li, Xuan Dong, Zhe Chen, Weiyun Wang, Xiangyu Zhao, Jixuan Chen, Haodong Duan, Tianbao Xie, Chenyu Yang, Shiqian Su, Yue Yu, Yuan Huang, Yiqian Liu, Xiao Zhang, Yanting Zhang, Xiangy...

  221. [242]

    Yiqin Wang, Haoji Zhang, Jingqi Tian, and Yansong Tang. 2025. Ponder &amp; Press: Advancing Visual GUI Agent towards General Computer Control. InFindings of the Association for Computational Linguistics: ACL 2025. 1461–1473. doi:10.18653/v1/2025.findings-acl.76

  222. [243]

    Yuanlei Wang, Liuzhou Zhang, Haohao Luo, and Ying Shen. 2025. INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent. InFindings of the Association for Computational Linguistics: EMNLP 2025. 9148–9160. doi:10.18653/v1/2025.findings-emnlp.486

  223. [244]

    Yanxi Wang, Zhiling Zhang, Wenbo Zhou, Weiming Zhang, Jie Zhang, Qiannan Zhu, Yu Shi, Shuxin Zheng, and Jiyan He. 2026. GUIGuard: Toward a General Framework for Privacy-Preserving GUI Agents. arXiv:2601.18842 [cs.CR] doi:10.48550/arXiv.2601.18842

  224. [246]

    Zilong Wang, Yuedong Cui, Li Zhong, Zimin Zhang, Da Yin, Bill Yuchen Lin, and Jingbo Shang. 2024. OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation. arXiv:2407.19056 [cs.CL] doi:10.48550/arXiv.2407.19056

  225. [247]

    Qingyuan Wu, Jianheng Liu, Jianye Hao, Jun Wang, and Kun Shao. 2025. Advancing Autonomous VLM Agents via Variational Subgoal-Conditioned Reinforcement Learning. arXiv:2502.07949 [cs.LG] doi:10.48550/arXiv.2502.07949

  226. [248]

    Qinzhuo Wu, Wei Liu, Jian Luan, and Bin Wang. 2025. ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Vo...

  227. [249]

    Yuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Jiabo Ye, Yutong Kou, Ming Yan, Fei Huang, Xiaoshan Yang, Weiming Dong, and Changsheng Xu. 2025. Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation. arXiv:2506.04614 [cs...

  228. [250]

    Wenyi Wu, Kun Zhou, Ruoxin Yuan, Vivian Yu, Stephen Wang, Zhiting Hu, and Biwei Huang. 2025. Auto-scaling Continuous Memory for GUI Agent. arXiv:2510.09038 [cs.AI] doi:10.48550/arXiv.2510.09038

  229. [251]

    Hao Wen, Shizuo Tian, Borislav Pavlov, Wenjie Du, Yixuan Li, Ge Chang, Shanhui Zhao, Jiacheng Liu, Yunxin Liu, Ya-Qin Zhang, and Yuanchun Li. 2025. AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation. InProceedings of the 23rd Annual International Conference on Mob...

  230. [252]

    Hao Wen, Hongming Wang, Jiaxuan Liu, and Yuanchun Li. 2023. DroidBot-GPT: GPT-powered UI Automation for Android. arXiv:2304.07061 [cs.SE] doi:10.48550/arXiv.2304.07061

  231. [253]

    Michael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre, Nigam Shah, and Christopher Ré. 2024. Automating the Enterprise with Foundation Models.Proceedings of the VLDB Endowment17, 11 (2024), 2805–2812. doi:10.14778/3681954.3681964

  232. [254]

    Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan Khare, Tathagat Verma, Tibor Thompson, Miguel Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez, Vardhan Agrawal, Althea Hudson, Nigam Shah, and Christopher R...

  233. [255]

    Biao Wu, Yanda Li, Zhiwei Zhang, Yunchao Wei, Meng Fang, and Ling Chen. 2024. Foundations and Recent Trends in Multimodal Mobile Agents: A Survey. arXiv:2411.02006 [cs.AI] doi:10.48550/arXiv.2411.02006

  234. [256]

    Benlong Wu, Yuang Qi, Xiuwei Shang, Weiming Zhang, Nenghai Yu, and Kejiang Chen. 2025. MMPro: A Decoupled Perception-Thinking-Execution Framework for Secure GUI Agent. InProceedings of the 33rd ACM International Conference on Multimedia. 4679–4687. doi:10.1145/3746027.3755553

  235. [257]

    Fangzhou Wu, Shutong Wu, Yulong Cao, and Chaowei Xiao. 2024. WIPI: A New Web Threat for LLM-Driven Web Agents. arXiv:2402.16965 [cs.CR] doi:10.48550/arXiv.2402.16965

  236. [258]

    Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Zekun Xi, Runnan Fang, Linhai Zhang, Yulan He, Deyu Zhou, Pengjun Xie, and Fei Huang. 2025. WebWalker: Benchmarking LLMs in Web Traversal. arXiv:2501.07572 [cs.CL] doi:10.48550/arXiv.2501.07572

  237. [259]

    Qianhui Wu, Kanzhi Cheng, Rui Yang, Chaoyun Zhang, Jianwei Yang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao, Reuben Tan, Si Qin, Lars Liden, Qingwei Lin, Huan Zhang, Tong Zhang, Jianbing Zhang, Dongmei Zhang, and Jianfeng Gao. 2025. GUI-Actor: Coordinate-Free Visual Groundi...

  238. [260]

    Qinchen Wu, Difei Gao, Kevin Qinghong Lin, Zhuoyu Wu, Xiangwu Guo, Peiran Li, Weichen Zhang, Hengxu Wang, and Mike Zheng Shou. 2024. GUI Action Narrator: Where and When Did That Action Take Place? arXiv:2406.13719 [cs.CV] doi:10.48550/arXiv.2406.13719 Software Engineering for ...

  239. [261]

    Qinzhuo Wu, Pengzhi Gao, Wei Liu, and Jian Luan. 2025. BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 4250–4272. doi:10.18653/v1/2025.emnlp-main.212

  240. [262]

    Yuquan Xie, Zaijing Li, Rui Shao, Gongwei Chen, Kaiwen Zhou, Yinchuan Li, Dongmei Jiang, and Liqiang Nie. 2025. Mirage-1: Augmenting and Updating GUI Agent with Hierarchical Multimodal Skills. arXiv:2506.10387 [cs.AI] doi:10.48550/arXiv.2506.10387

  241. [263]

    Mingzhe Xing, Rongkai Zhang, Hui Xue, Qi Chen, Fan Yang, and Zhen Xiao. 2024. Understanding the Weakness of Large Language Model Agents within a Complex Android Environment. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6061–6072. doi:...

  242. [264]

    Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Liujian Liujianfeng, Ang Li, Jian Luan, Bin Wang, and Shuo Shang. 2024. MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding. InFindings of the Association for Computational Linguistics: EMNLP 2024. 10231–1025...

  243. [265]

    Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li. 2024. AdvWeb: Controllable Black-box Attacks on VLM-powered Web Agents. arXiv:2410.17401 [cs.CR] doi:10.48550/arXiv.2410.17401

  244. [266]

    Zongru Wu, Pengzhou Cheng, Zheng Wu, Tianjie Ju, Zhuosheng Zhang, and Gongshen Liu. 2025. Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks. arXiv:2503.00401 [cs.CL] doi:10.48550/arXiv.2503.00401

  245. [267]

    Zhiyong Wu, Chengcheng Han, Zichen Ding, Zhenmin Weng, Zhoumianze Liu, Shunyu Yao, Tao Yu, and Lingpeng Kong. 2024. OS-Copilot: Towards Generalist Computer Agents with Self-Improvement. arXiv:2402.07456 [cs.AI] doi:10.48550/arXiv.2402.07456

  246. [268]

    Zheng Wu, Heyuan Huang, Xingyu Lou, Xiangmou Qu, Pengzhou Cheng, Zongru Wu, Weiwen Liu, Weinan Zhang, Jun Wang, Zhaoxiang Wang, and Zhuosheng Zhang. 2025. VeriOS: Query-Driven Proactive Human-Agent-GUI Interaction for Trustworthy OS Agents. arXiv:2509.07553 [cs.CL] doi:10.4855...

  247. [269]

    Zongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng, Tianjie Ju, Zheng Wu, Lingzhong Dong, Haiyue Sheng, Zhuosheng Zhang, and Gongshen Liu. 2025. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles. arXiv:2509.13615 [cs.AI] doi:10.4...

  248. [270]

    Zhiyong Wu, Zhenyu Wu, Fangzhi Xu, Yian Wang, Qiushi Sun, Chengyou Jia, Kanzhi Cheng, Zichen Ding, Liheng Chen, Paul Pu Liang, and Yu Qiao. 2024. OS-ATLAS: A Foundation Action Model for Generalist GUI Agents. arXiv:2410.23218 [cs.CL] doi:10.48550/arXiv.2410.23218

  249. [272]

    Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, Dawn Song, and Bo Li. 2025. GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning. In International Conference on Machine Learning. https://openre...

  250. [273]

    Han Xiao, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Hao He, Lue Fan, Liuyang Bian, Rui Hu, Liang Liu, Shuai Ren, Yafei Wen, Xiaoxin Chen, Aojun Zhou, and Hongsheng Li. 2025. UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents. arXiv...

  251. [274]

    Bin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou, Yinchuan Li, Jie Liu, Min Zhang, and Liqiang Nie. 2025. GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent. arXiv:2505.16827 [cs.AI] doi:10.48550/arXiv.2505.16827

  252. [275]

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for O...

  253. [276]

    Liu, Yiheng Xu, Hongjin Su, Dongchan Shin, Caiming Xiong, and Tao Yu

    Tianbao Xie, Fan Zhou, Zhoujun Cheng, Peng Shi, Luoxuan Weng, Yitao Liu, Toh Jing Hua, Junning Zhao, Qian Liu, Che Liu, Leo Z. Liu, Yiheng Xu, Hongjin Su, Dongchan Shin, Caiming Xiong, and Tao Yu. 2023. OpenAgents: An Open Platform for Language Agents in the Wild. arXiv:2310.1...

  254. [277]

    An Yan, Zhengyuan Yang, Wanrong Zhu, Kevin Lin, Linjie Li, Jianfeng Wang, Jianwei Yang, Yiwu Zhong, Julian McAuley, Jianfeng Gao, Zicheng Liu, and Lijuan Wang. 2023. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation. arXiv:2311.07562 [cs.CV]...

  255. [278]

    Zihe Yan, Jiaping Gui, Zhuosheng Zhang, and Gongshen Liu. 2025. LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents. arXiv:2507.10610 [cs.CR] doi:10.48550/arXiv.2507.10610

  256. [279]

    Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, and Shengyu Zhang. 2025. GUI-PRA: Process Reward Agent for GUI Tasks. arXiv:2509.23263 [cs.AI] doi:10.48550/arXiv.2509.23263 1:52 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, C...

  257. [280]

    Jingqi Yang, Zeng Song, Jiawei Chen, Mingli Song, Zhou Sheng, linjun sun, Xiaogang Ouyang, Chun Chen, and Can Wang. 2025. GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies. arXiv:2506.14477 [cs.AI] doi:10.48550/arXiv.2506.14477

  258. [281]

    Hai-Ming Xu, Qi Chen, Lei Wang, and Lingqiao Liu. 2025. Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning.Proceedings of the AAAI Conference on Artificial Intelligence 39, 8 (2025), 8851–8859. doi:10.1609/aaai.v39i8.32957

  259. [282]

    Kevin Xu, Yeganeh Kordi, Tanay Nayak, Adi Asija, Yizhong Wang, Kate Sanders, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, and Daniel Khashabi. 2025. TurkingBench: A Challenge Benchmark for Web Agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapt...

  260. [283]

    Nancy Xu, Sam Masling, Michael Du, Giovanni Campagna, Larry Heck, James Landay, and Monica S Lam. 2021. Grounding Open-Domain Instructions to Automate Web Support Tasks. arXiv:2103.16057 [cs.CL] doi:10.48550/arXiv. 2103.16057

  261. [284]

    Ho, Carl Yang, and Dong Yu

    Ran Xu, Kaixin Ma, Wenhao Yu, Hongming Zhang, Joyce C. Ho, Carl Yang, and Dong Yu. 2025. Retrieval-augmented GUI Agents with Generative Guidelines. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 17877–17886. doi:10.18653/v1/2025.emnlp...

  262. [285]

    Tianqi Xu, Linyao Chen, Dai-Jie Wu, Yanjun Chen, Zecheng Zhang, Xiang Yao, Zhiqiang Xie, Yongchao Chen, Shilong Liu, Bochen Qian, Anjie Yang, Zhaoxuan Jin, Jianbo Deng, Philip Torr, Bernard Ghanem, and Guohao Li. 2025. CRAB: Cross-environment Agent Benchmark for Multimodal Lan...

  263. [286]

    Yifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu, Hanchen Zhang, Bohao Jing, Shudan Zhang, Yuting Wang, Wenyi Zhao, and Yuxiao Dong. 2025. MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents. arXiv:2509.18119 [cs.LG] doi:10.48550/arXiv.2509.18119

  264. [287]

    Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong. 2025. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents. InProceedings of the 63rd Annual Meeting of the Association for Computationa...

  265. [288]

    Yiheng Xu, Dunjie Lu, Zhennan Shen, Junli Wang, Zekun Wang, Yuchen Mao, Caiming Xiong, and Tao Yu. 2024. AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials. arXiv:2412.09605 [cs.CL] doi:10. 48550/arXiv.2412.09605

  266. [289]

    Yiheng Xu, Zekun Wang, Junli Wang, Dunjie Lu, Tianbao Xie, Amrita Saha, Doyen Sahoo, Tao Yu, and Caiming Xiong. 2024. Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction. arXiv:2412.04454 [cs.CL] doi:10.48550/arXiv.2412.04454

  267. [290]

    Yibin Xu, Liang Yang, Hao Chen, Hua Wang, Zhi Chen, and Yaohua Tang. 2025. DeskVision: Large Scale Desktop Region Captioning for Advanced GUI Agents. arXiv:2503.11170 [cs.CL] doi:10.48550/arXiv.2503.11170

  268. [291]

    Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. 2025. An Illusion of Progress? Assessing the Current State of Web Agents. arXiv:2504.01382 [cs.AI] doi:10.48550/arXiv.2504.01382

  269. [294]

    Jiaxi Yang and Haowen Hou. 2025. RWKV-UI: UI Understanding with Enhanced Perception and Reasoning. In2025 IEEE International Conference on Multimedia and Expo (ICME). 1–6. doi:10.1109/icme59968.2025.11210007

  270. [296]

    Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, and Jianfeng Gao. 2025. Magma: A Foundation Model for Multimodal AI Agents. In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  271. [297]

    Ke Yang, Yao Liu, Sapana Chaudhary, Rasool Fakoor, Pratik Chaudhari, George Karypis, and Huzefa Rangwala. 2024. AgentOccam: A Simple Yet Strong Baseline for LLM-Based Web Agents. arXiv:2410.13825 [cs.AI] doi:10.48550/arXiv. 2410.13825

  272. [298]

    Pei Yang, Hai Ci, and Mike Zheng Shou. 2025. macOSWorld: A Multilingual Interactive Benchmark for GUI Agents. arXiv:2506.04135 [cs.AI] doi:10.48550/arXiv.2506.04135

  273. [299]

    Qi Yang, Weichen Bi, Haiyang Shen, Yaoqi Guo, and Yun Ma. 2025. PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels. arXiv:2504.16419 [cs.CV] doi:10.48550/arXiv.2504.16419 Software Engineering for and with GUI Agent 1:53

  274. [300]

    Rui Yang, Qianhui Wu, Zhaoyang Wang, Hanyang Chen, Ke Yang, Hao Cheng, Huaxiu Yao, Baolin Peng, Huan Zhang, Jianfeng Gao, and Tong Zhang. 2026. GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL. arXiv:2602.22190 [...

  275. [2021]

    InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence

    UIBert: Learning Generic Multimodal Representations for UI Understanding. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence. 1705–1712. doi:10.24963/ijcai.2021/235

  276. [2024]

    arXiv:2410.15164 [cs.AI] doi:10.48550/arXiv.2410.15164

    SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation. arXiv:2410.15164 [cs.AI] doi:10.48550/arXiv.2410.15164

  277. [2025]

    arXiv:2504.16073 [cs.CL] doi:10.48550/arXiv.2504.16073

    Guiding VLM Agents with Process Rewards at Inference Time for GUI Navigation. arXiv:2504.16073 [cs.CL] doi:10.48550/arXiv.2504.16073

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.