REVIEW 3 major objections 4 minor 285 references
Software Engineering for and with GUI Agent
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read GUI agents are closed-loop systems that benchmarks alone cannot validate
desk verdict First SE-lifecycle map of GUI agents that reads the literature as systems work; the gap counts are directionally right but the absence-as-evidence rule makes the magnitudes unverifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analytical object is the closed-loop perceive-decide-execute architecture treated as a software system, coded across 327 SE-parseable papers for signals such as verification, memory, reflection, recovery, execution contracts, observability, and governance. The central proposed mechanism is the execution contract: a specification of valid actions, acceptable state changes, retry and stop rules, rollback conditions, and permission boundaries that would make recovery and human oversight testable. The coding scheme's rule that absent signals count as lack of explicit evidence carries the survey's gap statistics.
What would settle it
Re-code a random sample of papers coded as lacking maintenance and observability signals, searching for implicit versioning, log, replay, escalation, or permission mechanisms; if a substantial fraction describes such mechanisms without the coded keywords, the survey's gap statistics would not hold.
Extended reading notes
Core claim
The paper establishes that GUI agents should be analyzed as closed-loop software systems whose behavior emerges from the interaction between a foundation model, an orchestration layer, the target application, and the user. Reviewing 336 papers from 2018 to April 2026, it finds that modular perceive-decide-act architectures dominate but that recovery, human escalation, safety enforcement, and auditability are far less developed than perception and planning. On evaluation, task success dominates as a metric while protocol comparability, reproducibility, and risk-awareness lag. Across the lifecycle, testing beyond benchmarks, maintainability, observability, privacy controls, and systematic human oversight are sparse. The paper concludes that capability improvements alone cannot ensure deployment readiness and that future work must connect dependable execution with lifecycle-centered testing, reproducible evaluation, and cost-aware, human-centered governance.
Load-bearing premise
The survey treats the absence of an explicitly coded engineering signal as evidence that the underlying system lacks that support, so if many agents implement recovery, maintenance, or human oversight under different terminology, the reported gaps would be overstated.
Editorial extensions
If this is right
- Benchmark task success should no longer be treated as sufficient evidence of deployment readiness; a reported score describes the whole agent stack and protocol, not just the model.
- Evaluation protocols must disclose observation access, action spaces, environment resets, success oracles, retry policies, and intervention rules before results can be meaningfully compared across benchmarks.
- Research should invest in explicit execution contracts that define valid actions, stop conditions, escalation points, and reversible versus irreversible operations.
- A GUI-agent test pyramid spanning component tests, integration tests, trajectory regression, and adversarial system-level stress tests is needed to replace benchmark-only validation.
- Permissions, privacy controls, audit logs, and risk-adaptive human oversight must be designed as runtime components of the system rather than as post-hoc safeguards.
Reading between the lines
- If the coding rule is correct, published task-success rates likely overstate operational maturity, since many systems implement recovery or oversight implicitly without reporting it as an engineering artifact.
- The execution-contract proposal could be made directly testable by expressing contracts as preconditions and postconditions over UI states and measuring violation rates under interface perturbations.
- A natural extension is to treat benchmarks themselves as versioned software artifacts with reproducibility budgets, an idea the paper opens but does not fully formalize.
- Merging early reinforcement-learning web agents with post-2023 LLM agents in one corpus may smooth over generational differences; separating those cohorts could sharpen the reported architectural and evaluation trends.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper surveys 336 GUI-agent papers published or posted between January 2018 and April 2026 from a software-engineering perspective. It poses five research questions covering the research landscape, system architectures, evaluation protocols, software-lifecycle concerns, and future opportunities. The central claim is that GUI agents are closed-loop software systems whose deployment readiness depends on lifecycle engineering, and that capability improvements alone cannot ensure dependable, maintainable, secure, and deployable systems. The survey reports multi-label, functionally interpreted coding of the corpus and finds that recovery, human escalation, safety enforcement, maintainability, observability, and oversight are underdeveloped relative to perception, planning, and task-capability work, and that evaluation remains centered on task success with limited cross-protocol comparability. The paper closes with a research agenda built around execution contracts, trajectory-based testing, reproducible evaluation, and governed human oversight.
Significance. If its findings are accepted, this survey provides a valuable corrective to benchmark-centric assessments of GUI agents and a useful map for software-engineering research on these systems. The paper has clear strengths: a large and explicitly bounded corpus; a transparent RQ structure; consistent reporting of per-RQ denominators; multi-label coding and an explicit functional-interpretation rule for mechanisms such as verifiers, critics, and reward models; and a candid threats-to-validity section. The synthesis connecting architecture, evaluation, and lifecycle concerns is genuinely useful and goes beyond existing capability-oriented surveys. However, the headline negative claims are built on counts that treat the absence of explicit textual signals as absence of engineering support, and the paper does not currently demonstrate the reliability of that coding. The significance of the survey is therefore conditional on coding transparency and validation, which are not yet provided.
major comments (3)
- [§11, §3.3, §7.3, §7.4] The central negative findings are load-bearing on the coding rule stated in §11: 'we treat absent signals as a lack of explicit evidence under the coding scheme.' For example, §7.3 reports that 279/327 papers lack a maintenance-process signal and 251/327 lack a monitoring or audit mechanism, and §7.4 reports runtime human oversight in only 47/327 papers. These counts are then used to conclude that lifecycle concerns are 'underdeveloped' and that 'capability improvements alone cannot ensure deployment readiness.' But a paper can implement recovery, confirmation, or oversight without using the survey's terminology; the functional-coding step in §3.3 mitigates this, yet §11 itself acknowledges that such concerns 'may remain implicit in system descriptions.' Without a sensitivity analysis that re-codes absent signals as unknown, or a validation sample checked against author statements or system repositories, the size of the reported gaps may be inflated. I would like to see either (a) a demonstration that absent signals correlate with absent mechanisms in a sample, or (b) rewording of the statistics as 'explicitly discussed under this coding scheme' rather than 'lacking support.'
- [§3.3, §11] Section 3.3 states that the coding scheme was developed iteratively and that records were checked for internal coherence, but Section 11 concedes that these 'internal-coherence checks... do not provide independently measured inter-rater agreement from multiple coders.' Because the RQ2 and RQ4 statistics are the quantitative backbone of the paper, the absence of any inter-rater agreement measure for the interpretive dimensions is a significant gap. The paper should report agreement per coding dimension—especially for human oversight, maintainability, and observability—on a sample of papers, and describe how disagreements were resolved. Without this, counts such as 95/327 versus 86/327 for related runtime-control signals are not auditable, and the precision of the gap percentages cannot be assessed.
- [§3.3, Table 6, §7] The RQ4 denominator of 327 out of 336 records is never explained. Section 3.3 says 'RQ4 uses the 327 records with complete software-engineering coding,' and Table 6 uses '327 SE-parseable papers,' but the reader is not told which nine papers are excluded or what 'SE-parseable' means. This matters because all RQ4 shares use the 327 denominator, and a reader cannot reproduce the subsets or check for selection effects. Please state the exclusion rule, provide the identifiers or at least the contribution types and years of the excluded records, and release the coding instrument with the paper so that the reported counts can be independently verified.
minor comments (4)
- [Table 7] Table 7 mixes denominators within one table: the first three rows use the 327-paper SE-parseable set, while the last row ('Explicit human-in-the-loop flag') uses the 145-paper framework set. The caption notes this, but it would be clearer to split the table or to add a denominator column to each row.
- [Figure 3] The right panel of Figure 3 reports contribution types as multi-label counts, and the text correctly notes that counts do not sum to the corpus size; the figure caption would benefit from stating this explicitly to avoid reader confusion.
- [§3.2, §4.1] Table 1 shows that 60.7% of the corpus consists of preprints. The paper would be strengthened by a brief robustness discussion of whether the RQ1 growth trends and the RQ4 gap counts change when the analysis is restricted to peer-reviewed records.
- [§3.3] The term 'SE-parseable' is used at first in Section 3.3 and then repeatedly in the RQ4 analysis, but it is only loosely defined. Please provide a precise definition at first use, even if the full list of excluded records appears in an appendix or supplementary artifact.
Circularity Check
Survey's gap statistics are partly self-definitional via the absence-as-evidence coding rule, but conclusions retain independent empirical content.
-
self definitional
[Section 11, Threats to Validity (second paragraph); operationalized in RQ4 (Section 7) and Table 13]
"We use multi-label and functional coding to reduce this loss and treat absent signals as a lack of explicit evidence under the coding scheme."
The RQ4 conclusions that maintainability, observability, privacy, audit, and human oversight are 'underdeveloped' are derived from counting papers that lack explicit coded signals for these concerns. Under the stated coding rule, 'underdeveloped' is defined as 'absent signal,' so the reported gap percentages (e.g., 279/327 missing a maintenance-process signal, 251/327 missing monitoring/audit) are equivalent to the coding input by construction. The paper itself acknowledges that lifecycle concerns 'may remain implicit in system descriptions,' meaning the coding rule is an assumption that makes the central negative claim about field immaturity partly self-referential.
full rationale
This is a survey paper that describes and synthesizes 336 external publications; it does not derive predictions or first-principles results from its own model. The only point where a conclusion reduces to an input by construction is the operationalization of 'underdeveloped' in RQ4: the coding rule in Section 11 equates absence of an explicit signal with lack of explicit evidence, and the paper then reports low counts as evidence of underdevelopment. This is a genuine self-definitional step for those specific descriptive findings (e.g., maintainability, observability, human oversight). Nevertheless, the paper discloses this rule in Threats to Validity, uses functional coding to mitigate terminology differences, and supports its broader conclusions with a large external corpus and with independent benchmark and safety studies. No self-citation chain, no fitted parameter masquerading as a prediction, and no uniqueness theorem imported from the authors' prior work were found. Given the transparency and the independent empirical base, a score of 2 is appropriate: one minor self-referential step that does not undermine the central claim's independent content.
Assumptions & free parameters
assumptions (4)
- domain assumption The 336-paper corpus is representative of GUI-agent research.
- domain assumption Functional coding of modules and engineering signals reflects actual system behavior.
- domain assumption Absence of an explicit signal means absence of explicit engineering support.
- ad hoc to paper The software-engineering quality attributes chosen (recovery, maintainability, observability, privacy, human oversight) are the right lens for deployment readiness.
Cite this review
Pith. "Pith review of Software Engineering for and with GUI Agent." pith.science (2026). https://pith.science/paper/G5VJYGW4
@misc{pith2026260809278,
author = {Pith},
title = {Pith review of: Software Engineering for and with GUI Agent},
year = {2026},
howpublished = {\url{https://pith.science/paper/G5VJYGW4}},
note = {Machine review of arXiv:2608.09278}
}
read the original abstract
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
Tamer Abuelsaad, Deepak Akkil, Prasenjit Dey, Ashish Jagmohan, Aditya Vempaty, and Ravi Kokku. 2024. Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic Systems. arXiv:2407.13032 [cs.AI] doi:10.48550/arXiv.2407.13032
-
[2]
Saaket Agashe, Jiuzhou Han, Shuyu Gan, Jiachen Yang, Ang Li, and Xin Eric Wang. 2024. Agent S: An Open Agentic Framework that Uses Computers Like a Human. arXiv:2410.08164 [cs.AI] doi:10.48550/arXiv.2410.08164
-
[3]
Saaket Agashe, Kyle Wong, Vincent Tu, Jiachen Yang, Ang Li, and Xin Eric Wang. 2025. Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents. arXiv:2504.00906 [cs.AI] doi:10.48550/arXiv.2504.00906
-
[4]
Pranjal Aggarwal and Sean Welleck. 2025. Programming with Pixels: Computer-Use Meets Software Engineering. arXiv:2502.18525 [cs.SE] doi:10.48550/arXiv.2502.18525
-
[5]
Jaewoo Ahn, Junseo Kim, Heeseung Yun, Jaehyeon Son, Dongmin Park, Jaewoong Cho, and Gunhee Kim. 2025. FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 23365–23395. doi:10.18653/v1/2025.emnlp- main.1192
-
[6]
Neeraj Anand, Rishabh Jain, Sohan Patnaik, Balaji Krishnamurthy, and Mausoom Sarkar. 2025. AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent. arXiv:2512.00846 [cs.CV] doi:10.48550/ arXiv.2512.00846
-
[7]
Gilles Baechler, Srinivas Sunkara, Maria Wang, Fedir Zubach, Hassan Mansoor, Vincent Etter, Victor Cărbune, Jason Lin, Jindong Chen, and Abhanshu Sharma. 2024. ScreenAI: A Vision-Language Model for UI and Infographics Understanding. arXiv:2402.04615 [cs.CV] doi:10.48550/arXiv.2402.04615
-
[8]
Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, and Blaise Agüera y Arcas
Show all 285 references
-
[9]
Hao Bai, Yifei Zhou, Mert Cemri, Jiayi Pan, Alane Suhr, Sergey Levine, and Aviral Kumar. 2024. DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning. InAdvances in Neural Information Processing Systems 37. 12461–12495. doi:10.52202/079017-0397
2024 doi
-
[10]
Hao Bai, Yifei Zhou, Li Li, Sergey Levine, and Aviral Kumar. 2025. Digi-Q: Learning VLM Q-Value Functions for Training Device-Control Agents. InInternational Conference on Learning Representations. https://proceedings.iclr.cc/ paper_files/paper/2025/hash/519abe71ee55aac4fe821b...
2025
-
[11]
Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks. InAdvances in Neural ...
2024 doi
- [12]
-
[13]
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A. Plummer. 2022. A Dataset for Interactive Vision-Language Navigation with Unknown Command Feasibility.Lecture Notes in Computer Science (2022), 312–328. doi:10.1007/978-3-031-20074-8_18
2022 doi
-
[15]
Ruisheng Cao, Fangyu Lei, Haoyuan Wu, Jixuan Chen, Yeqiao Fu, Hongcheng Gao, Xinzhuang Xiong, Hanchong Zhang, Yuchen Mao, Wenjing Hu, Tianbao Xie, Hongshen Xu, Danyang Zhang, Sida Wang, Ruoxi Sun, Pengcheng Yin, Caiming Xiong, Ansong Ni, Qian Liu, Victor Zhong, Lu Chen, Kai Yu...
2024 doi
- [16]
-
[17]
Yuxiang Chai, Siyuan Huang, Yazhe Niu, Han Xiao, Liang Liu, Guozhi Wang, Dingyu Zhang, Shuai Ren, and Hongsheng Li. 2025. AMEX: Android Multi-annotation Expo Dataset for Mobile GUI Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 2138–2156. doi:10...
2025 doi
-
[18]
Yuxiang Chai, Shunye Tang, Han Xiao, Weifeng Lin, Hanhao Li, Jiayu Zhang, Liang Liu, Pengxiang Zhao, Guangyi Liu, Guozhi Wang, Shuai Ren, Rongduo Han, Haining Zhang, Siyuan Huang, and Hongsheng Li. 2025. A3: Android Agent Arena for Mobile GUI Agents with Essential-State Proced...
2025 doi
- [19]
- [20]
-
[21]
Cong Chen, Kaixiang Ji, Hao Zhong, Muzhi Zhu, Anzhou Li, Guo Gan, Ziyuan Huang, Cheng Zou, Jiajia Liu, Jingdong Chen, Hao Chen, and Chunhua Shen. 2025. GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks. arXiv:2509.23738 [cs.AI] doi:10.48550/arX...
2025 doi
-
[22]
Chen Chen, Jiawei Shao, Dakuan Lu, Haoyi Hu, Xiangcheng Liu, Hantao Yao, and Wu Liu. 2026. GUI-Eyes: Tool- Augmented Perception for Visual Grounding in GUI Agents.Proceedings of the AAAI Conference on Artificial Intelligence40, 35 (2026), 29350–29358. doi:10.1609/aaai.v40i35.40175
2026 doi
-
[23]
Chiyu Chen, Xinhao Song, Yunkai Chai, Yang Yao, Haodong Zhao, Lijun Li, Jie Li, Yan Teng, Gongshen Liu, and Yingchun Wang. 2025. GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On- Device Environments? arXiv:2510.20333 [cs.CR] doi:10.48550/arXi...
2025 doi
- [24]
- [25]
- [26]
-
[27]
Gongwei Chen, Lirong Jie, Lexiao Zou, Weili Guan, Miao Zhang, and Liqiang Nie. 2025. Enhancing GUI Agent with Uncertainty-Aware Self-Trained Evaluator. InAdvances in Neural Information Processing Systems
2025
-
[28]
Gongwei Chen, Xurui Zhou, Rui Shao, Yibo Lyu, Kaiwen Zhou, Shuai Wang, Wentao Li, Yinchuan Li, Zhongang Qi, and Liqiang Nie. 2025. Less is More: Empowering GUI Agent with Context-Aware Simplification. (2025), 5901–5911. doi:10.1109/iccv51701.2025.00558
2025
-
[29]
Jingxuan Chen, Derek Yuen, Bin Xie, Yuhao Yang, Gongwei Chen, Zhihao Wu, Li Yixing, Xurui Zhou, Weiwen Liu, Shuai Wang, Kaiwen Zhou, Rui Shao, Liqiang Nie, Yasheng Wang, Jianye Hao, Jun Wang, and Kun Shao
-
[30]
Liang Chen, Haozhe Zhao, Yinzhen Huang, Yang Luo, Tsekai Lin, Weichu Xie, Ruoyu Wu, Peiyi Wang, Runxin Xu, Ming Wu, and Baobao Chang. 2025. CCAgent: Coordinating Collaborative Data Scaling for Operating System Agents via Web3. InProceedings of the 34th ACM International Confer...
2025
-
[31]
Qi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou, Hsiang-Ting Chen, and Qi Wu. 2024. WebVLN: Vision- and-Language Navigation on Websites.Proceedings of the AAAI Conference on Artificial Intelligence38, 2 (2024), 1165–1173. doi:10.1609/aaai.v38i2.27878
2024 doi
- [32]
-
[33]
Wentong Chen, Junbo Cui, Jinyi Hu, Yujia Qin, Junjie Fang, Yue Zhao, Chongyi Wang, Jun Liu, Guirong Chen, Yupeng Huo, Yuan Yao, Yankai Lin, Zhiyuan Liu, and Maosong Sun. 2025. GUICourse: From General Vision Language Model to Versatile GUI Agent. (2025), 21936–21959. doi:10.186...
2025 doi
- [34]
-
[35]
Wei Chen, Zhiyuan Li, Zhen Guo, and Yikang Shen. 2026. Octo-Planner: On-Device Language Model for Planner- Action Agents.Lecture Notes in Computer Science(2026), 141–156. doi:10.1007/978-3-032-18011-7_9
2026 doi
-
[36]
Weizhi Chen, Ziwei Wang, Leyang Yang, Sheng Zhou, Xiaoxuan Tang, Jiajun Bu, Yong Li, and Wei Jiang. 2025. PG-Agent: An Agent Powered by Page Graph. InProceedings of the 33rd ACM International Conference on Multimedia. 6878–6887. doi:10.1145/3746027.3755189
2025
-
[38]
https://proceedings.neurips.cc/paper_files/paper/2025/hash/d067d16e3e5fe8fa8a3e62909907659a-Abstract- Conference.html
2025
-
[39]
Pengzhou Cheng, Haowen Hu, Zheng Wu, Zongru Wu, Tianjie Ju, Daizong Ding, Zhuosheng Zhang, and Gongshen Liu. 2025. Hidden Ghost Hand: Unveiling Backdoor Vulnerabilities in MLLM-Powered Mobile GUI Agents. InFindings of the Association for Computational Linguistics: EMNLP 2025. ...
2025 doi
-
[40]
Pengzhou Cheng, Zheng Wu, Zongru Wu, Tianjie Ju, Aston Zhang, Zhuosheng Zhang, and Gongshen Liu. 2025. OS-Kairos: Adaptive Interaction for MLLM-Powered GUI Agents. InFindings of the Association for Computational Linguistics: ACL 2025. 6701–6725. doi:10.18653/v1/2025.findings-acl.348
2025 doi
-
[41]
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024. SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents. (2024), 9313–9332. doi:10.18653/v1/2024.acl-long.505
2024 doi
- [42]
-
[43]
Xu, Siva Reddy, Quentin Cappart, Graham Neubig, Ruslan Salakhutdinov, Nicolas Chapados, and Alexandre Lacoste
Thibault Le Sellier De Chezelles, Maxime Gasse, Alexandre Drouin, Massimo Caccia, Léo Boisvert, Megh Thakkar, Tom Marty, Rim Assouel, Sahar Omidi Shayegan, Lawrence Keunho Jang, Xing Han Lù, Ori Yoran, Dehan Kong, Frank F. Xu, Siva Reddy, Quentin Cappart, Graham Neubig, Ruslan...
- [44]
-
[45]
Gaole Dai, Shiqi Jiang, Ting Cao, Yuanchun Li, Yuqing Yang, Rui Tan, Mo Li, and Lili Qiu. 2025. Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment. arXiv:2503.15937 [cs.AI] doi:10.48550/arXiv.2503.15937
2025 doi
-
[46]
Preetam Prabhu Srikar Dammu. 2025. Towards Ethical and Personalized Web Navigation Agents: A Framework for User-Aligned Task Execution. InProceedings of the Eighteenth ACM International Conference on Web Search and Data Mining. 1074–1076. doi:10.1145/3701551.3707420
2025
- [47]
- [48]
-
[49]
Yang Deng, Xuan Zhang, Wenxuan Zhang, Yifei Yuan, See-Kiong Ng, and Tat-Seng Chua. 2024. On the Multi-turn Instruction Following for Conversational Web Agents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 87...
2024 doi
- [50]
-
[51]
Shihan Deng, Weikai Xu, Hongda Sun, Wei Liu, Tao Tan, Liujianfeng Liujianfeng, Ang Li, Jian Luan, Bin Wang, Rui Yan, and Shuo Shang. 2024. Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents. InProceedings of the 62nd Annual Meeting of the Association for Computa...
2024 doi
-
[52]
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2Web: Towards a Generalist Agent for the Web. InAdvances in Neural Information Processing Systems 36. 28091–28114. doi:10.52202/075280-1220
2023 doi
-
[53]
Lingzhong Dong, Ziqi Zhou, Shuaibo Yang, Haiyue Sheng, Pengzhou Cheng, Zongru Wu, Zheng Wu, Gongshen Liu, and Zhuosheng Zhang. 2025. Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents. arXiv:2510.02204 [cs.CL] doi:10.48550/arXiv.2510.02204
2025 doi
-
[54]
Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, Léo Boisvert, Megh Thakkar, Quentin Cappart, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. 2024. WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Task...
- [55]
-
[56]
Jinhan Dong, Lei Jin, Zhihong Zhang, Wei Tang, Runqing Zhang, Liqiang Xu, and Junliang Xing. 2025. MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided Fusion.IEEE Internet of Things Journal(2025),
2025
-
[57]
doi:10.1109/jiot.2025.3600573
2025
- [58]
-
[59]
Yue Fan, Handong Zhao, Ruiyi Zhang, Yu Shen, Xin Eric Wang, and Gang Wu. 2025. GUI-Bee: Align GUI Action Grounding to Novel Environments via Autonomous Exploration. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 33249–33266. doi:10.18...
2025 doi
-
[60]
Yong Du, Yuchen Yan, Fei Tang, Zhengxi Lu, Chang Zong, Weiming Lu, Shengpei Jiang, and Yongliang Shen. 2026. Test-Time Reinforcement Learning for GUI Grounding via Region Consistency.Proceedings of the AAAI Conference Software Engineering for and with GUI Agent 1:41 on Artific...
2026 doi
- [61]
- [62]
-
[63]
Difei Gao, Siyuan Hu, Zechen Bai, Qinghong Lin, and Mike Zheng Shou. 2024. AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation. InProceedings of the 32nd ACM International Conference on Multimedia. 11255–11257. doi:10.1145/3664647.3684998
2024
- [65]
- [66]
- [67]
- [68]
-
[70]
Longxi Gao, Li Zhang, Shihe Wang, Pengzhi Gao, Wei Liu, Jian Luan, Shangguang Wang, Yuanchun Li, and Mengwei Xu. 2024. MobileViews: A Million-scale and Diverse Mobile GUI Dataset. arXiv:2409.14337 [cs.HC] doi:10.48550/ arXiv.2409.14337
2024 doi
- [71]
- [72]
-
[73]
Wenkang Han, Zhixiong Zeng, Jing Huang, Shu Jiang, Liming Zheng, Longrong Yang, Haibo Qiu, Chang Yao, Jingyuan Chen, and Lin Ma. 2025. UITron-Speech: Towards Automated GUI Agents Based on Speech Instructions. arXiv:2506.11127 [cs.CL] doi:10.48550/arXiv.2506.11127
2025 doi
- [74]
-
[75]
Ziyi Guan, Jason Chun Lok Li, Zhijian Hou, Pingping Zhang, Donglai Xu, Yuzhi Zhao, Mengyang Wu, Jinpeng Chen, Thanh-Toan Nguyen, Pengfei Xian, Wenao Ma, Shengchao Qin, Graziano Chesi, and Ngai Wong. 2025. KG- RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven ...
2025 doi
-
[76]
Xiangwu Guo, Difei Gao, and Mike Zheng Shou. 2025. AUTO-Explorer: Automated Data Collection for GUI Agent. arXiv:2511.06417 [cs.AI] doi:10.48550/arXiv.2511.06417
2025 doi
- [77]
-
[78]
Fung, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu
Zhitao He, Zijun Liu, Peng Li, Yi R. Fung, Ming Yan, Ji Zhang, Fei Huang, and Yang Liu. 2025. Advanc- ing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization. arXiv:2502.14496 [cs.CL] doi:10.48550/arXiv.2502.14496
-
[79]
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, Yuxiao Dong, Ming Ding, and Jie Tang. 2024. CogAgent: A Visual Language Model for GUI Agents. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 14281–142...
2024
-
[80]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vol...
2024 doi
-
[81]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, and Dong Yu. 2025. OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization. InProceedings of the 63rd Annual Meeting of the Asso...
2025 doi
- [82]
-
[83]
Zhiyuan Hu, Shiyun Xiong, Yifan Zhang, See-Kiong Ng, Anh Tuan Luu, Bo An, Shuicheng Yan, and Bryan Hooi
- [84]
- [85]
- [86]
-
[87]
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shawn Wang, Xinchen Xu, Shuofei Qiao, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wan...
2024
-
[88]
Tian Huang, Chun Yu, Weinan Shi, Zijian Peng, David Yang, Weiqi Sun, and Yuanchun Shi. 2025. Prompt2Task: Automating UI Tasks on Smartphones from Textual Prompts.ACM Transactions on Computer-Human Interaction32, 3 (2025), 1–45. doi:10.1145/3716132
2025 doi
- [90]
-
[91]
Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai, Caiming Xiong, and Chien-Sheng Wu. 2025. GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness. arXiv:2510.00536 [cs.CL] doi:10.48550/arXiv.2510.00536
2025 doi
-
[92]
Tenghao Huang, Kinjal Basu, Ibrahim Abdelaziz, Pavan Kapanipathi, Jonathan May, and Muhao Chen. 2025. R2D2: Remembering, Replaying and Dynamic Decision Making with a Reflective Agentic Memory. InProceedings of the 63rd Annual Meeting of the Association for Computational Lingui...
2025 doi
- [93]
- [94]
-
[95]
Yiqiao Jin, Stefano Petrangeli, Yu Shen, and Gang Wu. 2025. <scp>ScreenLLM:</scp> Stateful Screen Schema for Efficient Action Understanding and Prediction. InCompanion Proceedings of the ACM on Web Conference 2025. 2008–2013. doi:10.1145/3701716.3718379
2025
-
[96]
Raghav Kapoor, Yash Parag Butala, Melisa Russak, Jing Yu Koh, Kiran Kamble, Waseem AlShikh, and Ruslan Salakhutdinov. 2024. OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.Lecture Notes in Computer Science(2024), 161–17...
2024 doi
-
[97]
Iat Long Iong, Xiao Liu, Yuxuan Chen, Hanyu Lai, Shuntian Yao, Pengbo Shen, Hao Yu, Yuxiao Dong, and Jie Tang
-
[98]
InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations)
OpenWebAgent: An Open Toolkit to Enable Web Agents on Large Language Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 72–81. doi:10.18653/v1/2024.acl-demos.8
- [99]
-
[100]
Chengyou Jia, Minnan Luo, Zhuohang Dang, Qiushi Sun, Fangzhi Xu, Junlin Hu, Tianbao Xie, and Zhiyong Wu. 2025. AgentStore: Scalable Integration of Heterogeneous Agents As Specialized Generalist Computer Assistant. InFindings of the Association for Computational Linguistics: AC...
2025 doi
-
[101]
SeokJoo Kwak, Jihoon Kim, Boyoun Kim, Jung Jae Yoon, Wooseok Jang, Jeonghoon Hong, Jaeho Yang, and Yeong-Dae Kwon. 2025. MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements. arXiv:2511.13087 [cs.AI] doi:10.48550/arXiv.2511.13087
2025 doi
-
[103]
Bradley Knox, and Kimin Lee
Juyong Lee, Dongyoon Hahm, June Suk Choi, W. Bradley Knox, and Kimin Lee. 2026. MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control.Proceedings of the AAAI Conference on Artificial Intelligence 40, 44 (2026), 37565–37573. doi:10.1609/aaai.v40i44.41090
2026 doi
-
[104]
Su Kara, Fazle Faisal, and Suman Nath. 2025. WABER: Evaluating Reliability and Efficiency of Web Agents with Existing Benchmarks. InICLR Workshop on Foundation Models in the Wild. https://www.microsoft.com/en-us/ research/publication/waber-evaluating-reliability-and-efficiency...
2025
- [105]
-
[106]
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks. InProceedings of the 62nd Annual Meeting of the Asso...
2024 doi
-
[107]
Jing Yu Koh, Stephen McAleer, Daniel Fried, and Ruslan Salakhutdinov. 2024. Tree Search for Language Model Agents. arXiv:2407.01476 [cs.AI] doi:10.48550/arXiv.2407.01476
2024 doi
-
[108]
Hongxin Li, Jingfan Chen, Jingran Su, Yuntao Chen, Li Qing, and Zhaoxiang Zhang. 2025. AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2025 doi
-
[109]
Hongxin Li, Jingran Su, Jingfan Chen, Zheng Ju, Yuntao Chen, Qing Li, and Zhaoxiang Zhang. 2025. UIPro: Unleashing Superior Interaction Capability For GUI Agents. arXiv:2509.17328 [cs.CV] doi:10.48550/arXiv.2509.17328
2025 doi
- [110]
-
[111]
Jungjae Lee, Dongjae Lee, Chihun Choi, Youngmin Im, Jaeyoung Wi, Kihong Heo, Sangeun Oh, Sunjae Lee, and Insik Shin. 2025. VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification. InProceedings of the 31st Annual International Conference on Mobile Com...
2025
- [112]
-
[113]
Sunjae Lee, Junyoung Choi, Jungjae Lee, Munim Hasan Wasi, Hojun Choi, Steve Ko, Sangeun Oh, and Insik Shin. 2024. MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation. InProceedings of the 30th Annual International Conference on Mobile Computing and ...
2024
- [114]
-
[115]
Wei Li, Fu-Lin Hsu, William Bishop, Folawiyo Campbell-Ajala, Max Lin, and Oriana Riva. 2024. UINav: A Practical Approach to Train On-Device Automation Agents. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: H...
2024 doi
-
[116]
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020. Mapping Natural Language Instructions to Mobile UI Action Sequences. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 8198–8210. doi:10.18653/v1/2020.acl-main.729
2020 doi
-
[117]
Yanda Li, Chi Zhang, Wenjia Jiang, Wanqi Yang, Bin Fu, Pei Cheng, Xin Chen, Ling Chen, and Yunchao Wei. 2024. AppAgent v2: Advanced Agent for Flexible Mobile Interactions. arXiv:2408.11824 [cs.HC] doi:10.48550/arXiv.2408. 11824
2024 doi
-
[118]
Kaixin Li, Ziyang Meng, Hongzhan Lin, Ziyang Luo, Yuchen Tian, Jing Ma, Zhiyong Huang, and Tat-Seng Chua
-
[119]
InProceedings of the 33rd ACM International Conference on Multimedia
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use. InProceedings of the 33rd ACM International Conference on Multimedia. 8778–8786. doi:10.1145/3746027.3755688
-
[120]
Tao Li, Gang Li, Zhiwei Deng, Bryan Wang, and Yang Li. 2023. A Zero-Shot Language Agent for Computer Control with Structured Reflection. InFindings of the Association for Computational Linguistics: EMNLP 2023. 11261–11274. doi:10.18653/v1/2023.findings-emnlp.753 1:44 Shengchen...
2023 doi
-
[123]
Wei Li, William Bishop, Alice Li, Chris Rawles, Folawiyo Campbell-Ajala, Divya Tyamagundlu, and Oriana Riva
-
[124]
InAdvances in Neural Information Processing Systems 37
On the Effects of Data Scale on UI Control Agents. InAdvances in Neural Information Processing Systems 37. 92130–92154. doi:10.52202/079017-2925
-
[125]
Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, WenHao Wang, Tianze Wu, Zhengxi Lu, Siheng Chen, LiLinghao, Hao Wang, Guanjing Xiong, Yong Liu, and Hongsheng Li. 2025. LLM-Power...
2025 doi
- [126]
- [127]
- [128]
- [129]
-
[130]
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun
- [131]
-
[132]
Kevin Lin, Linjie Li, Difei Gao, Qinchen Wu, Mingyi Yan, Zhengyuan Yang, Lijuan Wang, and Mike Shou. 2024. VideoGUI: A Benchmark for GUI Automation from Instructional Videos. InAdvances in Neural Information Processing Systems 37. 69329–69360. doi:10.52202/079017-2214
2024 doi
- [134]
-
[135]
Guohong Liu, Jialei Ye, Jiacheng Liu, Yuanchun Li, Wei Liu, Pengzhi Gao, Jian Luan, and Yunxin Liu. 2025. Hijacking JARVIS: Benchmarking Mobile GUI Agents against Unprivileged Third Parties. InProceedings of the 2nd International Workshop on Edge and Mobile Foundation Models. ...
2025
-
[136]
Ziwei Liu, Borui Kang, Hangjie Yuan, Zixiang Zhao, Wei Li, Yifan Zhu, and Tao Feng. 2026. Continual GUI Agents. arXiv:2601.20732 [cs.LG] doi:10.48550/arXiv.2601.20732
2026 doi
- [137]
- [138]
-
[139]
Jiarun Liu, Jia Hao, Chunhong Zhang, and Zheng Hu. 2025. WEPO: Web Element Preference Optimization for LLM-based Web Navigation.Proceedings of the AAAI Conference on Artificial Intelligence39, 25 (2025), 26614–26622. doi:10.1609/aaai.v39i25.34863
2025 doi
- [140]
- [141]
- [142]
-
[143]
Xinyi Liu, Xiaoyi Zhang, Ziyun Zhang, and Yan Lu. 2025. UI-E2I-Synth: Advancing GUI Grounding with Large- Scale Instruction Synthesis. InFindings of the Association for Computational Linguistics: ACL 2025. 15668–15684. doi:10.18653/v1/2025.findings-acl.809
2025 doi
-
[144]
Yuhang Liu, Pengxiang Li, Zishu Wei, Congkai Xie, Xueyu Hu, Xinchen Xu, Shengyu Zhang, Xiaotian Han, Hongxia Yang, and Fei Wu. 2026. InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection. In Proceedings of the 19th Conference of the European Chap...
2026 doi
- [145]
-
[146]
Yuxuan Liu, Hongda Sun, Wei Liu, Jian Luan, Bo Du, and Rui Yan. 2025. MobileSteward: Integrating Multiple App- Oriented Agents with Self-Evolution to Automate Cross-App Instructions. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1. 88...
2025
-
[147]
Xing Han Lù, Zdeněk Kasner, and Siva Reddy. 2024. WebLINX: Real-World Website Navigation with Multi-Turn Dialogue. arXiv:2402.05930 [cs.CL] doi:10.48550/arXiv.2402.05930
2024 doi
-
[148]
Pal, and Siva Reddy
Xing Han Lù, Amirhossein Kazemnejad, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, and Siva Reddy. 2025. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories. arXiv:2504.08942 [cs.LG] ...
2025 doi
- [149]
- [150]
- [151]
-
[152]
Yuheng Lu, Qian Yu, Hongru Wang, Zeming Liu, Wei Su, Yanping Liu, Yuhang Guo, Maocheng Liang, Yunhong Wang, and Haifeng Wang. 2025. TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments. InFindings of the Association for...
2025 doi
-
[154]
Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Pengxiang Zhao, Guangyi Liu, Guanjing Xiong, and Hongsheng Li. 2026. UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.Proceedings of the AAAI Conference ...
2026 doi
- [155]
- [156]
- [157]
- [158]
-
[159]
Songqin Nong, Xiaoxuan Tang, Jingxuan Xu, Sheng Zhou, Jianfeng Chen, Tao Jiang, and Wenhao Xu. 2025. CRAFT- GUI: Curriculum-Reinforced Agent For GUI Tasks. arXiv:2508.11360 [cs.AI] doi:10.48550/arXiv.2508.11360
2025 doi
- [160]
- [161]
-
[162]
Longhui Ma, Di Zhao, Siwei Wang, Zhao Lv, and Miao Wang. 2026. Beyond element-level understanding: Explicit relational understanding for GUI agents.Pattern Recognition176 (2026), 113262. doi:10.1016/j.patcog.2026.113262
2026
-
[164]
Xinbei Ma, Yiting Wang, Yao Yao, Tongxin Yuan, Aston Zhang, Zhuosheng Zhang, and Hai Zhao. 2025. Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions. InProceedings of the 63rd Annual Meeting of the Association for Computational Ling...
2025 doi
-
[165]
Xinbei Ma, Zhuosheng Zhang, and Hai Zhao. 2024. CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation. InFindings of the Association for Computational Linguistics ACL 2024. 9097–9110. doi:10.18653/v1/2024.findings-acl.539
2024 doi
- [166]
-
[167]
Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapa- dos, M
Shravan Nayak, Xiangru Jian, Kevin Qinghong Lin, Juan A. Rodriguez, Montek Kalsi, Rabiul Awal, Nicolas Chapa- dos, M. Tamer Özsu, Aishwarya Agrawal, David Vazquez, Christopher Pal, Perouz Taslakian, Spandana Gella, and Sai Rajeswar. 2025. UI-Vision: A Desktop-centric GUI Bench...
-
[168]
Ahmed, Puneet Mathur, Seunghyun Yoon, Lina Yao, Branislav Kveton, Jihyung Kil, Thien Huu Nguyen, Trung Bui, Tianyi Zhou, Ryan A
Dang Nguyen, Jian Chen, Yu Wang, Gang Wu, Namyong Park, Zhengmian Hu, Hanjia Lyu, Junda Wu, Ryan Aponte, Yu Xia, Xintong Li, Jing Shi, Hongjie Chen, Viet Dac Lai, Zhouhang Xie, Sungchul Kim, Ruiyi Zhang, Tong Yu, Mehrab Tanjim, Nesreen K. Ahmed, Puneet Mathur, Seunghyun Yoon, ...
2025
- [169]
-
[170]
Yijun Qian, Yujie Lu, Alexander Hauptmann, and Oriana Riva. 2024. Visual Grounding for User Interfaces. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track)....
2024 doi
- [171]
-
[172]
Vardaan Pahuja, Yadong Lu, Corby Rosset, Boyu Gou, Arindam Mitra, Spencer Whitehead, Yu Su, and Ahmed Hassan Awadallah. 2025. Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents. In Findings of the Association for Computational Linguistics: ...
2025 doi
- [173]
- [174]
- [175]
-
[176]
Manmatha, and Shabnam Ghadar
Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju, Kunwar Yashraj Singh, R. Manmatha, and Shabnam Ghadar. 2025. R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding. InFindings of the Association for Computational Linguistics: ACL 2025. 9669–9685. doi:10....
2025 doi
-
[177]
Pawel Pawlowski, Krystian Zawistowski, Wojciech Lapacz, Adam Wiacek, Marcin Skorupa, Sebastien Postansque, and Jakub Hoscilowicz. 2025. TinyClick: Single-Turn Agent for Empowering GUI Automation. InInterspeech 2025. 3035–3039. doi:10.21437/interspeech.2025-176
2025 doi
-
[178]
Bigham, and Amy Pavel
Yi-Hao Peng, Faria Huq, Yue Jiang, Jason Wu, Xin Yue Li, Jeffrey P. Bigham, and Amy Pavel. 2024. DreamStruct: Understanding Slides and User Interfaces via Synthetic Data Generation.Lecture Notes in Computer Science(2024), 466–485. doi:10.1007/978-3-031-72691-0_26
2024 doi
- [179]
- [180]
- [181]
-
[182]
Yucheng Shi, Wenhao Yu, Jingyuan Huang, Wenlin Yao, Wenhu Chen, and Ninghao Liu. 2025. Towards Trustworthy GUI Agents: A Survey. arXiv:2503.23434 [cs.LG] doi:10.48550/arXiv.2503.23434
2025 doi
-
[183]
Haoyi Qiu, Alexander Fabbri, Divyansh Agarwal, Kung-Hsiang Huang, Sarah Tan, Nanyun Peng, and Chien-Sheng Wu
-
[184]
InFindings of the Association for Computational Linguistics: NAACL 2025
Evaluating Cultural and Social Awareness of LLM Web Agents. InFindings of the Association for Computational Linguistics: NAACL 2025. 3978–4005. doi:10.18653/v1/2025.findings-naacl.222
2025 doi
- [185]
-
[186]
Dezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu, Yuan Cao, Ying Zhang, Wei Yang, and Tao Xie. 2024. Guardian: A Runtime Framework for LLM-Based UI Exploration. InProceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis. 958–970. doi:10.1145/3650...
2024
- [187]
- [188]
- [189]
- [190]
-
[191]
Peter Shaw, Mandar Joshi, James Cohan, Jonathan Berant, Panupong Pasupat, Hexiang Hu, Urvashi Khandelwal, Kenton Lee, and Kristina N Toutanova. 2023. From Pixels to UI Actions: Learning to Follow Instructions via Graphical User Interfaces. InAdvances in Neural Information Proc...
2023 doi
-
[192]
Jiahui Sun, Zhichao Hua, and Yubin Xia. 2025. AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents. arXiv:2503.02403 [cs.AI] doi:10.48550/arXiv.2503.02403
2025 doi
-
[193]
Liangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai, Zichen Zhu, and Kai Yu. 2022. META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI. (2022), 6699–6712. doi:10.18653/v1/2022.emnlp-main.449
2022 doi
- [194]
-
[195]
Yucheng Shi, Wenhao Yu, Zaitang Li, Yong-Lin Wang, Hongming Zhang, Ninghao Liu, Haitao Mi, and Dong Yu
- [196]
-
[197]
Kunal Singh, Shreyas Singh, and Mukund Khanna. 2025. Trishul: Towards Region Identification and Screen Hierarchy Understanding for Large VLM Based GUI Agents. In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 170–179. doi:10.1109/cvprw673...
2025
- [198]
-
[199]
Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, and Zhongmin Cai. 2024. VisionTasker: Mobile Task Au- tomation Using Vision Based UI Understanding and LLM Task Planning. InProceedings of the 37th Annual ACM Symposium on User Interface Software and Technology. 1–17. doi:10.1...
2024
- [200]
- [201]
- [202]
- [203]
- [204]
-
[205]
Xingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang, and Jing Tang. 2025. Understanding GUI Agent Localization Biases through Logit Sharpness. InFindings of the Association for Computational Linguistics: EMNLP 2025. 23361–23374. doi:10.18653/v1/2025.findings-emnlp.1268
2025 doi
-
[206]
Lucas-Andrei Thil, Mirela Popa, and Gerasimos Spanakis. 2024. Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning. InProceedings of the 39th ACM/SIGAPP Symposium on Applied Computing. 866–874. doi:10.1145/3605098.3635903
2024
-
[207]
Chan, Jikun Kang, Wenqi Wu, Filippos Christianos, Fraser Greenlee, Andy Toulis, and Marvin Purtorab
George Thomas, Alex J. Chan, Jikun Kang, Wenqi Wu, Filippos Christianos, Fraser Greenlee, Andy Toulis, and Marvin Purtorab. 2025. WebGames: Challenging General-Purpose Web-Browsing AI Agents. arXiv:2502.18356 [cs.LG] doi:10.48550/arXiv.2502.18356 Software Engineering for and w...
- [208]
- [210]
-
[211]
Karlsson, Bo An, Shuicheng Yan, and Zongqing Lu
Weihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia, Ziluo Ding, Boyu Li, Bohan Zhou, Junpeng Yue, Jiechuan Jiang, Yewen Li, Ruyi An, Molei Qin, Chuqiao Zong, Longtao Zheng, Yujie Wu, Xiaoqiang Chai, Yifei Bi, Tian- bao Xie, Pengjie Gu, Xiyun Li, Ceyao Zhang, Long Tian, Chaojie ...
- [212]
- [213]
- [214]
- [215]
- [216]
-
[217]
Sizhe Tang, Rongqian Chen, and Tian Lan. 2026. Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use Agents. arXiv:2602.02995 [cs.AI] doi:10.48550/arXiv.2602.02995
2026 doi
- [218]
-
[219]
Luyuan Wang, Yongyu Deng, Yiwei Zha, Guodong Mao, Qinmin Wang, Tianchen Min, Wei Chen, and Shoufa Chen
- [220]
- [221]
- [223]
- [224]
- [225]
-
[226]
Bryan Wang, Gang Li, and Yang Li. 2023. Enabling Conversational Interaction with Mobile UI using Large Language Models. InProceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1–17. doi:10.1145/3544548. 3580895
2023 doi
-
[227]
Bowen Wang, Xinyuan Wang, Jiaqi Deng, Tianbao Xie, Ryan Li, Yanzhe Zhang, Junli Wang, Dunjie Lu, Zicheng Gong, Gavin Li, Toh Jing Hua, Wei-Lin Chiang, Ion Stoica, Diyi Yang, Yu Su, Yi Zhang, Zhiguo Wang, Victor Zhong, and Tao Yu. 2026. Computer Agent Arena: Toward Human-Centri...
2026
-
[228]
Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang
- [229]
- [230]
- [231]
- [232]
- [233]
- [234]
-
[235]
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. AutoDroid: LLM-powered Task Automation in Android. InProceedings of the 30th Annual International Conference on Mobile Computing and Networking...
2024
- [236]
- [237]
-
[238]
WenHao Wang, Zijie Yu, Rui Ye, Jianqing Zhang, Guangyi Liu, Liang Liu, Siheng Chen, and Yanfeng Wang. 2025. FedMABench: Benchmarking Mobile GUI Agents on Decentralized Heterogeneous User Data. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Proces...
2025 doi
-
[239]
WenHao Wang, Mengying Yuan, Zijie Yu, Guangyi Liu, Rui Ye, Tian Jin, Siheng Chen, and Yanfeng Wang. 2025. MobileA3gent: Training Mobile GUI Agents Using Decentralized Self-Sourced Data from Diverse Users. InProceedings of the Fourth Workshop on Bridging Human-Computer Interact...
2025 doi
- [240]
- [241]
-
[242]
Yiqin Wang, Haoji Zhang, Jingqi Tian, and Yansong Tang. 2025. Ponder & Press: Advancing Visual GUI Agent towards General Computer Control. InFindings of the Association for Computational Linguistics: ACL 2025. 1461–1473. doi:10.18653/v1/2025.findings-acl.76
2025 doi
-
[243]
Yuanlei Wang, Liuzhou Zhang, Haohao Luo, and Ying Shen. 2025. INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent. InFindings of the Association for Computational Linguistics: EMNLP 2025. 9148–9160. doi:10.18653/v1/2025.findings-emnlp.486
2025 doi
- [244]
- [246]
- [247]
-
[248]
Qinzhuo Wu, Wei Liu, Jian Luan, and Bin Wang. 2025. ReachAgent: Enhancing Mobile Agent via Page Reaching and Operation. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Vo...
2025 doi
-
[249]
Yuyang Wanyan, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Jiabo Ye, Yutong Kou, Ming Yan, Fei Huang, Xiaoshan Yang, Weiming Dong, and Changsheng Xu. 2025. Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation. arXiv:2506.04614 [cs...
2025 doi
-
[250]
Wenyi Wu, Kun Zhou, Ruoxin Yuan, Vivian Yu, Stephen Wang, Zhiting Hu, and Biwei Huang. 2025. Auto-scaling Continuous Memory for GUI Agent. arXiv:2510.09038 [cs.AI] doi:10.48550/arXiv.2510.09038
2025 doi
-
[251]
Hao Wen, Shizuo Tian, Borislav Pavlov, Wenjie Du, Yixuan Li, Ge Chang, Shanhui Zhao, Jiacheng Liu, Yunxin Liu, Ya-Qin Zhang, and Yuanchun Li. 2025. AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation. InProceedings of the 23rd Annual International Conference on Mob...
2025
- [252]
-
[253]
Michael Wornow, Avanika Narayan, Krista Opsahl-Ong, Quinn McIntyre, Nigam Shah, and Christopher Ré. 2024. Automating the Enterprise with Foundation Models.Proceedings of the VLDB Endowment17, 11 (2024), 2805–2812. doi:10.14778/3681954.3681964
2024
-
[254]
Michael Wornow, Avanika Narayan, Ben Viggiano, Ishan Khare, Tathagat Verma, Tibor Thompson, Miguel Hernandez, Sudharsan Sundar, Chloe Trujillo, Krrish Chawla, Rongfei Lu, Justin Shen, Divya Nagaraj, Joshua Martinez, Vardhan Agrawal, Althea Hudson, Nigam Shah, and Christopher R...
2024 doi
- [255]
-
[256]
Benlong Wu, Yuang Qi, Xiuwei Shang, Weiming Zhang, Nenghai Yu, and Kejiang Chen. 2025. MMPro: A Decoupled Perception-Thinking-Execution Framework for Secure GUI Agent. InProceedings of the 33rd ACM International Conference on Multimedia. 4679–4687. doi:10.1145/3746027.3755553
2025
- [257]
- [258]
- [259]
- [260]
-
[261]
Qinzhuo Wu, Pengzhi Gao, Wei Liu, and Jian Luan. 2025. BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 4250–4272. doi:10.18653/v1/2025.emnlp-main.212
2025 doi
- [262]
-
[263]
Mingzhe Xing, Rongkai Zhang, Hui Xue, Qi Chen, Fan Yang, and Zhen Xiao. 2024. Understanding the Weakness of Large Language Model Agents within a Complex Android Environment. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 6061–6072. doi:...
2024
-
[264]
Qinzhuo Wu, Weikai Xu, Wei Liu, Tao Tan, Liujian Liujianfeng, Ang Li, Jian Luan, Bin Wang, and Shuo Shang. 2024. MobileVLM: A Vision-Language Model for Better Intra- and Inter-UI Understanding. InFindings of the Association for Computational Linguistics: EMNLP 2024. 10231–1025...
2024 doi
- [265]
- [266]
- [267]
- [268]
-
[269]
Zongru Wu, Rui Mao, Zhiyuan Tian, Pengzhou Cheng, Tianjie Ju, Zheng Wu, Lingzhong Dong, Haiyue Sheng, Zhuosheng Zhang, and Gongshen Liu. 2025. See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles. arXiv:2509.13615 [cs.AI] doi:10.4...
2025 doi
- [270]
-
[272]
Zhen Xiang, Linzhi Zheng, Yanjie Li, Junyuan Hong, Qinbin Li, Han Xie, Jiawei Zhang, Zidi Xiong, Chulin Xie, Carl Yang, Dawn Song, and Bo Li. 2025. GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning. In International Conference on Machine Learning. https://openre...
2025
- [273]
- [274]
-
[275]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for O...
2024 doi
-
[276]
Liu, Yiheng Xu, Hongjin Su, Dongchan Shin, Caiming Xiong, and Tao Yu
Tianbao Xie, Fan Zhou, Zhoujun Cheng, Peng Shi, Luoxuan Weng, Yitao Liu, Toh Jing Hua, Junning Zhao, Qian Liu, Che Liu, Leo Z. Liu, Yiheng Xu, Hongjin Su, Dongchan Shin, Caiming Xiong, and Tao Yu. 2023. OpenAgents: An Open Platform for Language Agents in the Wild. arXiv:2310.1...
- [277]
- [278]
-
[279]
Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, and Shengyu Zhang. 2025. GUI-PRA: Process Reward Agent for GUI Tasks. arXiv:2509.23263 [cs.AI] doi:10.48550/arXiv.2509.23263 1:52 Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, C...
2025 doi
- [280]
-
[281]
Hai-Ming Xu, Qi Chen, Lei Wang, and Lingqiao Liu. 2025. Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning.Proceedings of the AAAI Conference on Artificial Intelligence 39, 8 (2025), 8851–8859. doi:10.1609/aaai.v39i8.32957
2025 doi
-
[282]
Kevin Xu, Yeganeh Kordi, Tanay Nayak, Adi Asija, Yizhong Wang, Kate Sanders, Adam Byerly, Jingyu Zhang, Benjamin Van Durme, and Daniel Khashabi. 2025. TurkingBench: A Challenge Benchmark for Web Agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapt...
2025 doi
- [283]
-
[284]
Ho, Carl Yang, and Dong Yu
Ran Xu, Kaixin Ma, Wenhao Yu, Hongming Zhang, Joyce C. Ho, Carl Yang, and Dong Yu. 2025. Retrieval-augmented GUI Agents with Generative Guidelines. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 17877–17886. doi:10.18653/v1/2025.emnlp...
2025 doi
-
[285]
Tianqi Xu, Linyao Chen, Dai-Jie Wu, Yanjun Chen, Zecheng Zhang, Xiang Yao, Zhiqiang Xie, Yongchao Chen, Shilong Liu, Bochen Qian, Anjie Yang, Zhaoxuan Jin, Jianbo Deng, Philip Torr, Bernard Ghanem, and Guohao Li. 2025. CRAB: Cross-environment Agent Benchmark for Multimodal Lan...
2025 doi
-
[286]
Yifan Xu, Xiao Liu, Xinghan Liu, Jiaqi Fu, Hanchen Zhang, Bohao Jing, Shudan Zhang, Yuting Wang, Wenyi Zhao, and Yuxiao Dong. 2025. MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents. arXiv:2509.18119 [cs.LG] doi:10.48550/arXiv.2509.18119
2025 doi
-
[287]
Yifan Xu, Xiao Liu, Xueqiao Sun, Siyi Cheng, Hao Yu, Hanyu Lai, Shudan Zhang, Dan Zhang, Jie Tang, and Yuxiao Dong. 2025. AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents. InProceedings of the 63rd Annual Meeting of the Association for Computationa...
2025 doi
- [288]
- [289]
- [290]
-
[291]
Tianci Xue, Weijian Qi, Tianneng Shi, Chan Hee Song, Boyu Gou, Dawn Song, Huan Sun, and Yu Su. 2025. An Illusion of Progress? Assessing the Current State of Web Agents. arXiv:2504.01382 [cs.AI] doi:10.48550/arXiv.2504.01382
2025 doi
-
[294]
Jiaxi Yang and Haowen Hou. 2025. RWKV-UI: UI Understanding with Enhanced Perception and Reasoning. In2025 IEEE International Conference on Multimedia and Expo (ICME). 1–6. doi:10.1109/icme59968.2025.11210007
2025
-
[296]
Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, and Jianfeng Gao. 2025. Magma: A Foundation Model for Multimodal AI Agents. In2025 IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2025
- [297]
-
[298]
Pei Yang, Hai Ci, and Mike Zheng Shou. 2025. macOSWorld: A Multilingual Interactive Benchmark for GUI Agents. arXiv:2506.04135 [cs.AI] doi:10.48550/arXiv.2506.04135
2025 doi
- [299]
- [300]
-
[2021]
InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence
UIBert: Learning Generic Multimodal Representations for UI Understanding. InProceedings of the Thirtieth International Joint Conference on Artificial Intelligence. 1705–1712. doi:10.24963/ijcai.2021/235
2021 doi
- [2024]
- [2025]
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.