REVIEW 4 major objections 4 minor 1 cited by
Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This paper builds Sphinx, a multi-dimensional benchmark that decomposes mobile UI navigation into five capabilities, and finds that today's foundation models solve at most 16.6% of tasks and zero of 214 industrial testing tasks, with…
desk verdict Sphinx is a genuinely useful multi-dimensional benchmark with real industrial tasks, but the zero-success testing result rests on manually crafted evaluators that are never validated, so treat the exact numbers as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the five-capability decomposition, each measured by a dedicated evaluation toolkit. Goal understanding and app knowledge are probed with multiple-choice and binary questions; planning is measured by a completion-judgment task that asks the model to decide 'continue' or 'stop' at every step of a trajectory; grounding is measured by single-step tasks that map a natural-language instruction to the correct UI element; and instruction following is measured by invariant checks that detect repeated actions, malformed action formats, and format violations in a distraction-free focused-context setting. End-to-end success is judged by manually crafted trajectory evaluators, which check assertions about the final screen, last action, and encountered elements along the generated path, allowing alternative paths and producing partial credit through average completion proportion.
What would settle it
Have independent human raters judge a sample of the 214 WeChat testing-task trajectories that Sphinx marks as failed, allowing any completion path; if even one trajectory is accepted as satisfying the task's intent, the 0% success-rate claim and the ranking it anchors would need revision.
Extended reading notes
Core claim
Sphinx's central discovery is that the failure of foundation models in goal-based mobile UI navigation is concentrated in UI-specific capabilities rather than in general language or knowledge. In the end-to-end evaluation, the best model reaches 16.6% success rate and 21.5% average completion proportion across all tasks, and all models score 0% on the 214 WeChat testing tasks, which require longer, more context-sensitive, and path-specific sequences of actions. The multi-dimensional probes then localize the deficit: models answer goal-understanding and app-knowledge questions with accuracy above 85% for most strong models, but grounding accuracy tops out at 87.5% in text and 65.0% in vision, the average accuracy of stop decisions in planning is 40.3%, and every model repeats actions about half the time or more on end-to-end traces. These gaps cascade, so even a multimodal navigation agent built on a strong model achieves only 8.0-11.0% success on a subset of Sphinx tasks, roughly matching or falling below the plain text-agent baseline. The paper reads this as evidence that UI grounding, planning, and instruction following, not goal understanding or app knowledge, are the bottlenecks that must be trained or engineered around.
Load-bearing premise
The zero-success and ranking results depend on the manually crafted trajectory evaluators correctly judging whether a generated action sequence achieves the task's intent; the paper does not validate those evaluators against an external oracle or check inter-judge agreement, so a stricter-than-intended evaluator would lower scores across the board.
Editorial extensions
If this is right
- A direct corollary is that adding a deterministic output-format validator and a repetition suppressor to a UI navigation agent should raise end-to-end success without any model improvement, because the paper's invariants identify these as frequent failure modes.
- Because text observations from the accessibility tree strongly outperform raw or annotated screenshots, current agent designs should feed text first and treat vision as auxiliary; vision-language models are not yet competitive on this task class.
- Generic high scores on language benchmarks do not imply UI navigation competence, so downstream-specific benchmarks like Sphinx are needed for model selection in this domain.
- Agent architectures cannot compensate for the underlying model's grounding and planning limits, so further gains depend on UI-specific fine-tuning or external support systems rather than prompt-level agent design.
- The multi-dimensional scorecard allows practitioners to choose models based on the capability their task needs most, e.g. grounding accuracy for element-heavy tasks or stop-decision accuracy for open-ended exploration.
Reading between the lines
- A testable extension of the paper's own diagnosis: if an environment-level repetition blocker and format validator are added to the same chain-of-thought agent pipeline, success rates should rise toward the level implied by the capability scores; if they do not, the attribution of failure to instruction following would need revisiting.
- The lack of validation for the manually crafted trajectory evaluators means the 0% testing-task result should be stress-tested by an independent human re-judgment; until then, the zero-success figure is as much a claim about evaluator strictness as about model ability.
- The benchmark's task distribution (popular public apps plus a single industrial super-app) suggests that domain transfer of the conclusions is plausible for Android but unproven for other ecosystems such as iOS or non-mobile embodied agents.
- If the multi-dimensional bottleneck pattern generalizes, similar capability decompositions for web or desktop navigation would likely reveal the same grounding-versus-knowledge asymmetry, making Sphinx's toolkit adaptable beyond mobile.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Sphinx, a multi-dimensional benchmark for evaluating foundation models on goal-based mobile UI navigation. Sphinx contains 244 user tasks collected from 100 popular industrial apps and 214 testing tasks drawn from WeChat's internal QA process, and it evaluates five capabilities (goal understanding, app knowledge, planning, grounding, and instruction following) in addition to end-to-end success rate (SR) and average completion proportion (ACP). The authors benchmark 8 foundation models across 20 configurations, including text-only, vision, and multimodal models, and also evaluate the AppAgent UI navigation agent. Their main findings are that all models perform poorly on the benchmark, that no model solves any WeChat testing task, that text-based models outperform vision-based models, and that the primary bottlenecks are UI-specific capabilities such as grounding, planning, and instruction following. The paper also reports seven lessons learned. The benchmark and scripts are publicly available.
Significance. If its claims hold, Sphinx would be a valuable community resource: it is one of the first benchmark efforts to combine industrial-scale testing tasks with fine-grained capability evaluation, and the public release of tasks and interfaces would facilitate reproducible comparisons. The paper's multi-dimensional diagnosis of failure modes (grounding, planning, instruction following) is a useful step beyond binary pass/fail metrics, and the inclusion of WeChat QA test cases gives the benchmark genuine industrial relevance. The authors also credibly demonstrate that sophisticated agent designs such as AppAgent inherit the limitations of their underlying FMs, which is a practically important observation. However, the significance of the headline results is contingent on the validity of the paper's trajectory evaluators and on the comparability of its text and vision conditions, both of which are questionable as presented.
major comments (4)
- [§3.2.1, Table 4] The central claim that no benchmarked FM solves any of the 214 WeChat testing tasks (SR 0.0% in Table 4) is entirely defined by the manually crafted trajectory-based evaluators described in Section 3.2.1. The paper reports no validation of these evaluators: no inter-evaluator agreement, no independent oracle, no audit of false negatives, and no evidence that the handling of alternative navigation paths is complete. This concern is compounded by Section 5.1, which states that testing tasks require entering the functionality via a specified path, so a legitimate completion that uses an alternative path would be judged a failure. Because SR, ACP, and the zero-success headline are all computed from these unvalidated evaluators, evaluator strictness is a load-bearing threat: even a small rate of overly strict evaluators could change the 'none solved' result and the model ranking. I request a concrete validation study: sample a subset of trajectories (including near-misses and failures), have QA engineers independently judge task success, and report agreement rates and false-negative rates, or otherwise demonstrate that the evaluators are neither too strict nor too lenient.
- [§3.4.1, §5.1, Table 5] The paper's conclusion that 'vision modality lags significantly behind text modality' is not supported by a controlled comparison. Text-based models receive the full accessibility tree with element IDs, and the action space (Table 2) allows these models to act directly on element IDs. Vision-based models, in contrast, receive raw screenshots or SoM-labeled screenshots and must output screen coordinates. This means the text and vision conditions differ in both the information content of the observation and the difficulty of the action interface. The lower performance of vision models could be due to the lack of element IDs, the need for coordinate grounding, or the absence of semantic text, rather than to inherent visual perception limitations. I request a matched comparison, for example by providing the same element IDs to vision models as visual overlays (or by removing element IDs from the text condition), or by ablating the action space so that both modalities use equivalent output formats.
- [§5.1, Table 4] All end-to-end results are reported as single point estimates with no confidence intervals and no repeated runs. Given the stochastic nature of FM decoding, the differences that support claims such as 'larger models achieve much higher SRs and ACPs' (Section 5.1) or the ordering of GPT-4-Turbo (31.1%) versus GPT-4o (28.7%) on user tasks could be within run-to-run noise. I request that the authors report either multiple runs with variance, or statistical significance tests for the main comparisons, or at least explicitly acknowledge the absence of repeated runs as a limitation and avoid strong claims about model ordering.
- [§3.3.2, §3.3.4, §5.2] Parts of the benchmark content are generated or repaired by the same families of models that are benchmarked. Specifically, grounding instructions are generated by GPT-4o and then cleaned by human annotators (Section 3.3.4), and knowledge-probing and grounding outputs are 'repaired' by DeepSeek-V2 before scoring (Section 5.2.1, Section 5.2.4). Since GPT-4o and DeepSeek-V2 are among the evaluated models, their measured performance on these tasks could reflect familiarity with their own generated text or repair conventions rather than a general UI capability. This does not invalidate the benchmark, but it does threaten the comparability of the multi-dimensional results across models. I request that the paper either report results separately for human-authored versus model-generated task subsets, or provide evidence that the generation/repair process does not differentially benefit the generating models.
minor comments (4)
- [General] Figure 2 appears garbled in the manuscript, with series of 'uni' escape sequences instead of readable axis labels and legend text; the figure should be regenerated and the underlying rendering issue fixed.
- [§2.1, References] References [62] and [63] point to the same paper ('Understanding the weakness of large language model agents within a complex android environment') with the same authors and venue; if they are indeed meant to be the same, one should be removed, and the in-text citations to 'AndroidArena' should be made consistent.
- [§3.3.2] The text says 'an BQ' where it should read 'a BQ', and the sentence 'A BQ is a question with only two possible answers' is slightly redundant; a quick copyedit would improve readability.
- [§6, Threats to Validity] The threats-to-validity section addresses model representativeness and task representativeness but omits the most immediate threat, the correctness of the trajectory evaluators; this should be discussed even if no additional validation is performed.
Circularity Check
No significant circularity: Sphinx is an externally grounded benchmark whose SR/ACP results are computed from manually defined trajectory evaluators, not from fitted parameters or self-cited predictions.
full rationale
The paper's central claims are empirical benchmark measurements: SR and ACP are computed from trajectories produced by the evaluated FMs and checked against trajectory-based evaluators that were manually crafted from task requirements and reference trajectories. There is no equation or derivation in which a predicted quantity is defined in terms of the benchmark's own outputs, and no fitted parameter is renamed as a prediction. The use of GPT-4o to generate grounding instructions and focused-context task content, and DeepSeek-V2 to repair malformed outputs, introduces potential model-generated test bias, but it does not make the headline results reduce to the inputs by construction; the evaluated models are still measured against externally defined tasks and evaluators. The paper's citations to prior work by overlapping authors (e.g., Guardian) are contextual related-work references and are not load-bearing for the benchmark's validity. The unvalidated nature of the manually crafted evaluators is a legitimate correctness threat, but it is not circularity under the criteria of this analysis, because no claim is equivalent to its own input by definition. Accordingly, the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Goal-based mobile UI navigation can be decomposed into five independent capabilities: goal understanding, app knowledge, planning, grounding, and instruction following.
- domain assumption Manually crafted trajectory-based evaluators and reference trajectories correctly determine whether a navigation task is accomplished.
- domain assumption The accessibility tree text is a faithful and sufficient representation of the UI state for navigation.
- domain assumption GPT-4o-generated instructions and DeepSeek-V2 output repair are acceptable tools for constructing and scoring benchmark tasks.
Cite this review
Pith. "Pith review of Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation." pith.science (2026). https://pith.science/paper/AMXCYM4X
@misc{pith2026250102863,
author = {Pith},
title = {Pith review of: Beyond Pass or Fail: Multi-Dimensional Benchmarking of Foundation Models for Goal-based Mobile UI Navigation},
year = {2026},
howpublished = {\url{https://pith.science/paper/AMXCYM4X}},
note = {Machine review of arXiv:2501.02863}
}
read the original abstract
Recent advances of foundation models (FMs) have made navigating mobile applications (apps) based on high-level goal instructions within reach, with significant industrial applications such as UI testing. While existing benchmarks evaluate FM-based UI navigation using the binary pass/fail metric, they have two major limitations: they cannot reflect the complex nature of mobile UI navigation where FMs may fail for various reasons (e.g., misunderstanding instructions and failed planning), and they lack industrial relevance due to oversimplified tasks that poorly represent real-world scenarios. To address the preceding limitations, we propose Sphinx, a comprehensive benchmark for multi-dimensional evaluation of FMs in industrial settings of UI navigation. Sphinx introduces a specialized toolkit that evaluates five essential FM capabilities, providing detailed insights into failure modes such as insufficient app knowledge or planning issues. Using both popular Google Play applications and WeChat's internal UI test cases, we evaluate 8 FMs with 20 different configurations. Our results show that existing FMs universally struggle with goal-based testing tasks, primarily due to insufficient UI-specific capabilities. We summarize seven lessons learned from benchmarking FMs with Sphinx, providing clear directions for improving FM-based mobile UI navigation.
Figures
Forward citations
Cited by 1 Pith paper
-
Software Engineering for and with GUI Agent
A survey of 336 GUI-agent papers finds rapid growth alongside weak engineering support for recovery, human oversight, maintainability, and privacy, and calls for lifecycle-centered testing and governance.
Reference graph
Works this paper leans on
-
[1]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al
-
[2]
Dimitrios Alivanistos, Selene Báez Santamaría, Michael Cochez, Jan-Christoph Kalo, Emile van Krieken, and Thiviyan Thanapalasingam. 2022. Prompting as probing: Using language models for knowledge base construction. arXiv preprint arXiv:2208.11057 (2022)
arXiv 2022
-
[3]
Aliyun. 2024. Website of Qwen. https://tongyi.aliyun.com/
work page 2024
-
[4]
anthropic. 2024. Introducing computer use. https://www.anthropic.com/news/3- 5-models-and-computer-use
work page 2024
-
[5]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
2015
-
[6]
Chongyang Bai, Xiaoxue Zang, Ying Xu, Srinivas Sunkara, Abhinav Rastogi, Jindong Chen, et al. 2021. Uibert: Learning generic multimodal representations for ui understanding. arXiv preprint arXiv:2107.13731 (2021)
arXiv 2021
-
[7]
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2021. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258 (2021)
arXiv 2021
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al . 2020. Language models are few-shot learners. NIPS 33 (2020), 1877–1901
work page 2020
Show all 78 references
-
[9]
Andrea Burns, Deniz Arsan, Sanjna Agrawal, Ranjitha Kumar, Kate Saenko, and Bryan A Plummer. 2022. A dataset for interactive vision-language navigation with unknown command feasibility. In European Conference on Computer Vision . Springer, 312–328
2022
-
[10]
Jieshan Chen, Chunyang Chen, Zhenchang Xing, Xiwei Xu, Liming Zhu, Guo- qiang Li, and Jinshui Wang. 2020. Unblind your apps: Predicting natural-language labels for mobile gui components by deep learning. InProceedings of the ACM/IEEE 42nd international conference on software e...
2020
-
[11]
Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Conference’17, July 2017, Washington, DC, USA Dezhi Ran et al
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scot...
2021 arXiv
-
[12]
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2024. Mind2web: Towards a generalist agent for the web. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[13]
Felix Dobslaw, Robert Feldt, David Michaëlsson, Patrik Haar, Francisco Gomes de Oliveira Neto, and Richard Torkar. 2019. Estimating return on investment for gui test automation frameworks. In 2019 IEEE 30th International Symposium on Software Reliability Engineering (ISSRE) . ...
2019
-
[14]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024)
2024 arXiv
-
[15]
Michael D Ernst, Jake Cockrell, William G Griswold, and David Notkin. 1999. Dynamically discovering likely program invariants to support program evolution. In Proceedings of the 21st international conference on Software engineering . 213– 224
1999
-
[16]
Sidong Feng and Chunyang Chen. 2024. Prompting Is All You Need: Automated Android Bug Replay with Large Language Models. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering . 1–13
2024
-
[17]
Google. 2021. Android Accessibility Service. https://developer.android.com/ reference/android/accessibilityservice/AccessibilityService
2021
-
[18]
Nora Griffin-Shirley, Devender R Banda, Paul M Ajuwon, Jongpil Cheon, Jaehoon Lee, Hye Ran Park, and Sanpalei N Lyngdoh. 2017. A survey on the use of mobile applications for people who are visually impaired. Journal of Visual Impairment & Blindness 111, 4 (2017), 307–323
2017
-
[19]
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language under- standing. arXiv preprint arXiv:2009.03300 (2020)
2020 arXiv
-
[20]
Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)
2024 arXiv
-
[21]
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. 2023. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770 (2023)
2023 arXiv
-
[22]
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Chong Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Ruslan Salakhutdinov, and Daniel Fried
-
[23]
Kenton Lee, Mandar Joshi, Iulia Raluca Turc, Hexiang Hu, Fangyu Liu, Julian Mar- tin Eisenschlos, Urvashi Khandelwal, Peter Shaw, Ming-Wei Chang, and Kristina Toutanova. 2023. Pix2struct: Screenshot parsing as pretraining for visual lan- guage understanding. In International C...
2023
-
[24]
Toby Jia-Jun Li, Lindsay Popowski, Tom Mitchell, and Brad A Myers. 2021. Screen2vec: Semantic embedding of GUI screens and GUI components. In CHI. 1–15
2021
-
[25]
Yang Li, Jiacong He, Xin Zhou, Yuan Zhang, and Jason Baldridge. 2020. Map- ping Natural Language Instructions to Mobile UI Action Sequences. In ACL. Association for Computational Linguistics, Online
2020
-
[26]
Jun-Wei Lin, Navid Salehnamadi, and Sam Malek. 2022. Route: Roads not taken in ui testing. TOSEM (2022)
2022
-
[27]
Mario Linares-Vásquez, Carlos Bernal-Cárdenas, Kevin Moran, and Denys Poshy- vanyk. 2017. How do developers test android applications?. In 2017 IEEE In- ternational Conference on Software Maintenance and Evolution (ICSME) . IEEE, 613–622
2017
-
[28]
Aixin Liu, Bei Feng, Bin Wang, Bingxuan Wang, Bo Liu, Chenggang Zhao, Chengqi Dengr, Chong Ruan, Damai Dai, Daya Guo, et al . 2024. Deepseek- v2: A strong, economical, and efficient mixture-of-experts language model. arXiv preprint arXiv:2405.04434 (2024)
2024 arXiv
-
[29]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning. Advances in neural information processing systems 36 (2024)
2024
-
[30]
Zhe Liu, Chunyang Chen, Junjie Wang, Xing Che, Yuekai Huang, Jun Hu, and Qing Wang. 2022. Fill in the Blank: Context-aware Automated Text Input Gener- ation for Mobile GUI Testing. arXiv preprint arXiv:2212.04732 (2022)
2022 arXiv
-
[31]
Zhe Liu, Chunyang Chen, Junjie Wang, Mengzhuo Chen, Boyu Wu, Xing Che, Dandan Wang, and Qing Wang. 2024. Make llm a testing expert: Bringing human- like interaction to mobile gui testing via functionality-aware decisions. In ICSE. 1–13
2024
-
[32]
Siyu Lu, Mingzhe Liu, Lirong Yin, Zhengtong Yin, Xuan Liu, and Wenfeng Zheng
-
[33]
Meta. 2024. Llama 3.2: Revolutionizing edge AI and vision with open, customiz- able models. https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge- mobile-devices/
2024
-
[34]
Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. GAIA: a benchmark for General AI Assistants. arXiv preprint arXiv:2311.12983 (2023)
2023 arXiv
-
[35]
Maia Naftali and Leah Findlater. 2014. Accessibility in context: understanding the truly mobile experience of smartphone users with motor impairments. In Proceedings of the 16th international ACM SIGACCESS conference on Computers & accessibility. 209–216
2014
-
[36]
OpenAI. 2023. GPT-4 Technical Report
2023
-
[37]
OpenAI. 2024. Website of OpenAI. https://openai.com/
2024
-
[38]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems (2022)
2022
-
[39]
D. Ran, Y. Fu, Y. He, T. Chen, X. Tang, and T. Xie. 2024. Path Toward Elderly Friendly Mobile Apps. Computer 57, 06 (jun 2024), 29–39. https://doi.org/10. 1109/MC.2023.3322855
2024
-
[40]
Dezhi Ran, Zongyang Li, Chenxu Liu, Wenyu Wang, Weizhi Meng, Xionglin Wu, Hui Jin, Jing Cui, Xing Tang, and Tao Xie. 2022. Automated Visual Testing for Mobile Apps in an Industrial Setting. In ICSE-SEIP
2022
-
[41]
Dezhi Ran, Hao Wang, Zihe Song, Mengzhou Wu, Yuan Cao, Ying Zhang, Wei Yang, and Tao Xie. 2024. Guardian: A Runtime Framework for LLM-based UI Exploration. In ISSTA
2024
-
[42]
Dezhi Ran, Hao Wang, Wenyu Wang, and Tao Xie. 2023. Badge: Prioritizing UI Events with Hierarchical Multi-Armed Bandits for Automated UI Testing. In ICSE
2023
-
[43]
Dezhi Ran, Mengzhou Wu, Hao Yu, Yuetong Li, Jun Ren, Yuan Cao, Xia Zeng, Haochuan Lu, Zexin Xu, Mengqian Xu, et al. 2025. Sphinx. https://github.com/ PKU-ASE-RISE/Sphinx
2025
-
[44]
Christopher Rawles, Alice Li, Daniel Rodriguez, Oriana Riva, and Timothy Lilli- crap. 2024. Androidinthewild: A large-scale dataset for android device control. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[45]
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. 2023. Gpqa: A graduate-level google-proof q&a benchmark. arXiv preprint arXiv:2311.12022 (2023)
2023 arXiv
-
[46]
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al . 2025. Toward expert-level medical question answering with large language models. Nature Medicine (2025), 1–8
2025
-
[47]
Sadler, Wei-Lun Chao, and Yu Su
Chan Hee Song, Jiaman Wu, Clayton Washington, Brian M. Sadler, Wei-Lun Chao, and Yu Su. 2023. LLM-Planner: Few-Shot Grounded Planning for Embodied Agents with Large Language Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2023
-
[48]
Ting Su, Guozhu Meng, Yuting Chen, Ke Wu, Weiming Yang, Yao Yao, Geguang Pu, Yang Liu, and Zhendong Su. 2017. Guided, Stochastic Model-based GUI Testing of Android Apps. In ESEC/FSE
2017
-
[49]
Jiao Sun, Yufei Tian, Wangchunshu Zhou, Nan Xu, Qian Hu, Rahul Gupta, John Frederick Wieting, Nanyun Peng, and Xuezhe Ma. 2023. Evaluating Large Language Models on Controlled Generation Tasks.arXiv preprint arXiv:2310.14542 (2023)
2023 arXiv
-
[50]
Suresh Thummalapenta, Saurabh Sinha, Nimit Singhania, and Satish Chandra
-
[51]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971 (2023)
2023 arXiv
-
[52]
Ivan Vulić, Edoardo Maria Ponti, Robert Litschko, Goran Glavaš, and Anna Korhonen. 2020. Probing pretrained language models for lexical semantics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7222–7240
2020
-
[53]
Bryan Wang, Gang Li, Xin Zhou, Zhourong Chen, Tovi Grossman, and Yang Li. 2021. Screen2words: Automatic mobile UI summarization with multimodal learning. In UIST
2021
-
[54]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903 (2022)
2022 arXiv
-
[55]
Lili Wei, Yepang Liu, and Shing-Chi Cheung. 2016. Taming android fragmentation: Characterizing and detecting compatibility issues for android apps. InProceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering . 226–237
2016
-
[56]
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2023. Empowering LLM to use Smartphone for Intelligent Task Automation. arXiv preprint arXiv:2308.15272 (2023)
2023 arXiv
-
[57]
Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby Jia-Jun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. Autodroid: Llm-powered task automation in android. In Proceedings of the 30th Annual International Con- ference on Mobile Computing and Network...
2024
-
[58]
Hao Wen, Hongming Wang, Jiaxuan Liu, and Yuanchun Li. 2023. DroidBot-GPT: GPT-powered UI Automation for Android.arXiv preprint arXiv:2304.07061 (2023). Conference’17, July 2017, Washington, DC, USA
2023 arXiv
-
[59]
Mengzhou Wu, Hao Wang, Jun Ren, Yuan Cao, Yuetong Li, Alex Jiang, Dezhi Ran, Yitao Hu, Wei Yang, and Tao Xie. 2024. Skill-Adpative Imitation Learning for UI Test Reuse. arXiv preprint arXiv:2409.13311 (2024)
2024 arXiv
-
[60]
Mulong Xie, Zhenchang Xing, Sidong Feng, Chunyang Chen, Liming Zhu, and Xiwei Xu. 2022. Psychologically-Inspired, Unsupervised Inference of Perceptual Groups of GUI Widgets from GUI Images. arXiv preprint arXiv:2206.10352 (2022)
2022 arXiv
-
[61]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu
-
[63]
Mingzhe Xing, Rongkai Zhang, Hui Xue, Qi Chen, Fan Yang, and Zhen Xiao. 2024. Understanding the weakness of large language model agents within a complex android environment. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 6061–6072
2024
-
[64]
Shunguo Yan and PG Ramachandran. 2019. The current status of accessibility in mobile apps. ACM Transactions on Accessible Computing (TACCESS) 12, 1 (2019), 1–31
2019
-
[65]
Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao
-
[66]
arXiv:2404.07972 [cs.AI]
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. arXiv:2404.07972 [cs.AI]
-
[67]
Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023. Appagent: Multimodal agents as smartphone users.arXiv preprint arXiv:2312.13771 (2023)
2023 arXiv
-
[68]
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems 35 (2022), 20744–20757
2022
-
[69]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of thoughts: Deliberate problem solving with large language models. arXiv preprint arXiv:2305.10601 (2023)
2023 arXiv
-
[70]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629 (2022)
2022 arXiv
-
[71]
arXiv:2310.11441 [cs.CV] https://arxiv.org/abs/2310.11441
Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V. arXiv:2310.11441 [cs.CV] https://arxiv.org/abs/2310.11441
-
[72]
Prasad, and Tao Xie
Wei Yang, Mukul R. Prasad, and Tao Xie. 2013. A Grey-box Approach for Auto- mated GUI-model Generation of Mobile Applications. In FASE
2013
-
[73]
Shuyan Zhou, Frank F Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Yonatan Bisk, Daniel Fried, Uri Alon, et al . 2023. Webarena: A realistic web environment for building autonomous agents. arXiv preprint arXiv:2307.13854 (2023)
2023 arXiv
-
[77]
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. 2021. Florence: A new foundation model for computer vision. arXiv preprint arXiv:2111.11432 (2021)
2021 arXiv
-
[78]
Munazza Zaib, Wei Emma Zhang, Quan Z Sheng, Adnan Mahmood, and Yang Zhang. 2022. Conversational question answering: A survey. Knowledge and Information Systems 64, 12 (2022), 3151–3195
2022
-
[2012]
Automating test automation. In ICSE. 881–891
-
[2022]
Advances in neural information processing systems 35 (2022), 23716–23736
Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems 35 (2022), 23716–23736
2022
-
[2023]
PeerJ Computer Science 9 (2023), e1400
The multi-modal fusion in visual question answering: a review of attention mechanisms. PeerJ Computer Science 9 (2023), e1400
2023
-
[2024]
arXiv preprint arXiv:2401.13649 (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks. arXiv preprint arXiv:2401.13649 (2024)
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.