Pith. sign in

REVIEW 2 major objections 7 minor 4 cited by

OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use

T0 review · 2 major / 7 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read One framework describes every AI agent that controls a computer, phone, or browser.

desk verdict A genuinely useful survey and a reasonable umbrella term, but the 'OS' boundary is fuzzy and a few citation slips need fixing. read the letter →

arxiv 2508.04482 v1 pith:4H5ZTYKF submitted 2025-08-06 cs.AI cs.CLcs.CVcs.LG

classification cs.AIcs.CLcs.CVcs.LG
keywords OSAgentsmultimodallargelanguagemodelsGUIactiongroundingplanningevaluationbenchmarksagentframeworkssafetyandprivacy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the scattered work on AI agents that control computers, phones, and browsers by operating their graphical interfaces is one coherent research area, which it names OS Agents. It claims that every such agent can be described by three key components—environment, observation space, and action space—and must exercise three capabilities: understanding, planning, and grounding. On that basis the survey organizes the field into two construction strategies, domain-specific foundation models and agent frameworks, and a two-part evaluation scheme of protocols and benchmarks. If the framing is right, it gives researchers a common vocabulary for comparing systems, transferring ideas, and locating open problems. The paper also identifies safety, privacy, personalization, and self-evolution as the main unresolved directions.

What carries the argument

The organizing object is the OS Agent formalism: a classification scheme in which every agent is an environment plus an observation space and an action space, and must exhibit understanding, planning, and grounding. It does the work of a taxonomic backbone—every surveyed system is slotted into this structure, and the construction and evaluation sections are derived from it.

What would settle it

Find a widely used multimodal-device-control system whose operation cannot be expressed as an environment, an observation space, and an action space, or whose behavior cannot be assessed through understanding, planning, and grounding (for example, a system that works entirely through direct backend API calls and never observes a GUI). If such systems are common, the claim that the framework describes the field fails.

Watch

Extended reading notes

Core claim

The paper designates as OS Agents any agent based on multimodal large language models that uses a computing device by operating within the environments and interfaces provided by an operating system, and claims that the entire field can be understood through one recurrent structure. An OS Agent is defined by its environment (mobile, desktop, or web), its observation space (screen images, accessibility trees, OCR text, HTML, or combinations), and its action space (input, navigation, and extended operations such as code execution and API calls). It must be able to understand the interface, plan a sequence of steps, and ground instructions in concrete executable actions. The paper then uses thi

Load-bearing premise

The load-bearing premise is that the particular set of systems this survey chose to cover and the way it divides them into components and capabilities are the natural way to chart the field, yet the paper gives no criteria for what it included or excluded.

Editorial extensions

If this is right

  • New agent systems can be specified in a standard vocabulary: state the environment, observation space, and action space, then evaluate understanding, planning, and grounding separately.
  • Progress in one platform, such as GUI grounding data for mobile, can transfer to other platforms because all three share the same formal structure.
  • Evaluation can be standardized around benchmarks that report both step-level and task-level metrics and combine objective and subjective protocols.
  • Safety and privacy become first-class evaluation dimensions, with attacks such as prompt injection, environmental injection, and adversarial pop-ups assessed by dedicated benchmarks.
  • Personalization and self-evolution are identified as the main open problems, requiring memory mechanisms that accumulate user data and adapt over time.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the framework holds, it predicts that systems built to the same environment, observation, and action specification should be interchangeable at the middleware level: a planner from one agent could drive a perception module from another.
  • The paper's global-versus-iterative planning split suggests a testable extension: tasks with irreversible actions should favor iterative planners, so a benchmark that separates such task types could sharpen the distinction.
  • The survey groups systems by platform and setting, but leaves open whether the more useful axis is API access versus pure GUI control; comparing agents along that axis could complement or challenge the proposed structure.
  • Because the paper states no inclusion or exclusion criteria, its coverage is not verifiable; a reader extending the ideas should treat the framework as a hypothesis about the field rather than a complete inventory.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 7 minor

Summary. This paper presents a survey of (M)LLM-based agents that operate computing devices (desktops, mobile phones, and web browsers), which it terms "OS Agents." The survey proposes a conceptual framework centered on three key components—environment, observation space, and action space—and three core capabilities—understanding, planning, and grounding. It then reviews methods for building such agents, distinguishing domain-specific foundation models from agent frameworks, and surveys evaluation protocols and benchmarks. The paper also discusses challenges such as safety/privacy and personalization/self-evolution. It includes summary tables (models, frameworks, benchmarks) and points to an open-source repository.

Significance. If the proposed framework is accepted, the survey provides a useful organizing vocabulary for a fast-growing area that currently lacks standard terminology. The paper collects a broad set of representative work, and its tables (Table 1: foundation models; Table 2: agent frameworks; Table 3: benchmarks) are practical resources for researchers entering the field. The paper also explicitly acknowledges the existence of concurrent surveys (Section 6) and situates itself among them. The main value is taxonomic and bibliographical rather than technical; there are no derivations or experiments to verify. However, the coherence of the central taxonomy depends on the definition of "OS Agent," which is not currently operational.

major comments (2)
  1. [§2, §2.1, §3.2.4, Figure 1] The definition of OS Agent conflicts with the scope of surveyed works. Section 2 states that OS Agents operate within "environments and interfaces provided by the operating system," but Section 2.1 immediately includes "web" as an environment, and Section 3.2.4 lists API integration and code execution as extended actions. Web agents such as Mind2Web, WebArena, and VisualWebArena primarily interact through browser DOM/HTML or web APIs—application-layer artifacts, not OS-provided interfaces. Figure 1 itself lists "HTML" as an observation type. No criterion is provided for when an interface counts as "provided by the operating system." Consequently, the boundary of the field is arbitrary, and the claim that all surveyed agents form a coherent OS-defined research area is not established. The authors should either narrow the definition to agents that operate through OS-provided interfaces (e.
  2. [§1 and §4.2 (overall methodology)] The paper claims to be a "comprehensive survey" but does not state inclusion or exclusion criteria for the works it covers. It does not describe the search process (databases, time window, keywords) or the criteria for selecting particular papers for discussion. For example, Table 3 lists benchmarks from 2017 to 2024, but the rationale for including some benchmarks and omitting others is not given. This makes it impossible for a reader to assess representativeness or selection bias. The authors should add a short methodology paragraph describing how the literature was collected and screened. Without this, the comprehensiveness claim is unverifiable.
minor comments (7)
  1. [§2, first paragraph] "As illustrated in Figure 2" should refer to Figure 1; Figure 2 is introduced later in §3.1. The cross-reference is currently incorrect.
  2. [Author footnote] "Projsssect Lead" appears to be a typo for "Project Lead."
  3. [References] The entry "FengPeiyuan et al." is inconsistently formatted; the first author appears as a single-name token, likely "Peiyuan Feng." The citation in text (§3.1.4) should be checked for consistency.
  4. [§3.2.3] The citation "ToL [Pointed]" is incomplete; "Pointed" is not a proper author name and the full reference is missing.
  5. [§1] The citations for Amazon Alexa and Google Assistant appear swapped: the text says "Amazon Alexa [Google, 2024]" and "Google Assistant [Amazon, 2024]" but the reference list attributes the Google Assistant page to Google and the Alexa page to Amazon.
  6. [Table 3] Some entries have incomplete reference information, e.g., "AndroidControl [Li et al.]" has no year or venue, making it hard to identify the exact paper. The same ambiguity appears in the reference list for "Li et al." entries.
  7. [§3.2.3] "Additionally, In [Li et al., 2023], the agent analyzes..." has awkward capitalization/grammar; "In" should not be capitalized in this position.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's taxonomy is a descriptive framing, not a derivation from its own assumptions.

full rationale

This is a survey paper rather than a derivation or prediction pipeline. The central claim—that OS Agents can be organized by environment, observation space, action space, and capabilities such as understanding, planning, and grounding—is presented as a categorization scheme for the literature, not as a result derived from fitted parameters or from the authors' prior theorems. The paper does not fit data, make predictions from its own model, or invoke a uniqueness theorem. Several cited works include the authors' own prior publications (e.g., Zhou et al. 2023a, Qiao et al. 2022, Hu et al. 2024a), but these citations are not load-bearing for the survey's organizational structure: the taxonomy stands independently as a descriptive choice, and the cited papers are used as examples or related work rather than as justification for the framework. The potential tension between the OS-interface definition and the inclusion of web/DOM/API-based agents is a boundary-coherence critique, not a circularity: the paper is not defining its subject in terms of its own conclusions. No equation, fitted value, or self-citation chain makes the survey's claims equivalent to its inputs. Therefore, the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey rests on framing assumptions about how to carve up the field. No free parameters are fit, and no new entities are postulated.

assumptions (3)
  • domain assumption The category 'OS Agents' is a meaningful and coherent unit of analysis, spanning mobile, desktop, and web environments.
    Section 2 defines OS Agents in terms of environment, observation space, and action space. The entire survey depends on this grouping being natural and useful rather than artificially constructed.
  • domain assumption The selected set of referenced papers is representative of the current state of OS Agents research.
    The survey claims to be comprehensive, but no explicit inclusion criteria are given, so the representativeness of the cited literature is assumed.
  • domain assumption The tripartite capability split (understanding, planning, grounding) and the construction categories (foundation model versus agent framework) cover the relevant design space.
    Section 2.2 and Section 3 impose this structure on all surveyed works; some systems might straddle categories or introduce orthogonal dimensions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use." pith.science (2026). https://pith.science/paper/4H5ZTYKF

@misc{pith2026250804482,
  author       = {Pith},
  title        = {Pith review of: OS Agents: A Survey on MLLM-based Agents for General Computing Devices Use},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4H5ZTYKF}},
  note         = {Machine review of arXiv:2508.04482}
}
read the original abstract

The dream to create AI assistants as capable and versatile as the fictional J.A.R.V.I.S from Iron Man has long captivated imaginations. With the evolution of (multi-modal) large language models ((M)LLMs), this dream is closer to reality, as (M)LLM-based Agents using computing devices (e.g., computers and mobile phones) by operating within the environments and interfaces (e.g., Graphical User Interface (GUI)) provided by operating systems (OS) to automate tasks have significantly advanced. This paper presents a comprehensive survey of these advanced agents, designated as OS Agents. We begin by elucidating the fundamentals of OS Agents, exploring their key components including the environment, observation space, and action space, and outlining essential capabilities such as understanding, planning, and grounding. We then examine methodologies for constructing OS Agents, focusing on domain-specific foundation models and agent frameworks. A detailed review of evaluation protocols and benchmarks highlights how OS Agents are assessed across diverse tasks. Finally, we discuss current challenges and identify promising directions for future research, including safety and privacy, personalization and self-evolution. This survey aims to consolidate the state of OS Agents research, providing insights to guide both academic inquiry and industrial development. An open-source GitHub repository is maintained as a dynamic resource to foster further innovation in this field. We present a 9-page version of our work, accepted by ACL 2025, to provide a concise overview to the domain.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels

    cs.CR 2025-10 conditional novelty 6.0 of 10

    Indirect prompt injection through ads, webviews, and notifications reliably diverts mobile LLM agents into leaking data and installing malware across eight evaluated agents.

  2. RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control

    cs.CR 2026-07 conditional novelty 5.0 of 10

    An architecture that mediates LLM computer-use agents for UAV control by compiling agent decisions into validated, time-bounded, evidence-logged skill invocations, with a prototype on OpenClaw/PX4/OP-TEE.

  3. Plover: Steering GUI Agents through Plan-Centric Interaction

    cs.AI 2026-07 conditional novelty 5.0 of 10

    An expert repairing visible plans rescued 23 of 26 failed GUI automation runs, turning 17 into full and 6 into partial successes.

  4. AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    AgentScope 1.0 packages the components needed to build, evaluate, and deploy LLM agent applications into one developer framework.

Reference graph

Works this paper leans on

4 extracted references · 4 linked inside Pith · cited by 4 Pith papers

  1. [3]

    Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang

    URLhttps://arxiv.org/abs/2401.05778. Shenzhi Wang, Chang Liu, Zilong Zheng, Siyuan Qi, Shuo Chen, Qisen Yang, Andrew Zhao, Chaofei Wang, Shiji Song, and Gao Huang. Boosting llm agents with recursive contemplation for effective deception handling. InFindings of the Association for Computational Linguistics ACL 2024, pages 9909–9953, 2024g. Seth Neel and Pe...

  2. [4]

    Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li

    URLhttps://arxiv.org/abs/2409.11295. Chejian Xu, Mintong Kang, Jiawei Zhang, Zeyi Liao, Lingbo Mo, Mengqi Yuan, Huan Sun, and Bo Li. Advweb: Controllable black-box attacks on vlm-powered web agents, 2024d. URL https://arxiv.org/abs/2410.17401. Yanzhe Zhang, Tao Yu, and Diyi Yang. Attacking vision-language computer agents via pop-ups, 2024g. URLhttps://arx...

  3. [2023]

    Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Adrian de Wynter, Yan Xia, Wenshan Wu, Ting Song, Man Lan, and Furu Wei

    URLhttps://arxiv.org/abs/2212.10403. Yadong Zhang, Shaoguang Mao, Tao Ge, Xun Wang, Adrian de Wynter, Yan Xia, Wenshan Wu, Ting Song, Man Lan, and Furu Wei. Llm as a mastermind: A survey of strategic reasoning with large language models, 2024b. URLhttps://arxiv.org/abs/2404.01230. Xu Huang, Weiwen Liu, Xiaolong Chen, Xingmei Wang, Hao Wang, Defu Lian, Yas...

  4. [2024]

    Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo

    URLhttps://arxiv.org/abs/2411.15594. Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. Prometheus: Inducing fine- grained evaluation capability in language models, 2024c. URL https://arxiv.org/abs/2310. 08491. Seungone Kim, Juyoung Suk, Shayne Longpre, Bill ...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.