Pith. sign in

REVIEW 13 cited by

Cradle: Empowering Foundation Agents Towards General Computer Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.03186 v3 pith:DT2MBQDJ submitted 2024-03-05 cs.AI

classification cs.AI
keywords cradleagentsfoundationsoftwarecomplexcontrolacrosscomplete
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the success in specific scenarios, existing foundation agents still struggle to generalize across various virtual scenarios, mainly due to the dramatically different encapsulations of environments with manually designed observation and action spaces. To handle this issue, we propose the General Computer Control (GCC) setting to restrict foundation agents to interact with software through the most unified and standardized interface, i.e., using screenshots as input and keyboard and mouse actions as output. We introduce Cradle, a modular and flexible LMM-powered framework, as a preliminary attempt towards GCC. Enhanced by six key modules, Cradle can understand input screenshots and output executable code for low-level keyboard and mouse control after high-level planning, so that Cradle can interact with any software and complete long-horizon complex tasks without relying on any built-in APIs. Experimental results show that Cradle exhibits remarkable generalizability and impressive performance across four previously unexplored commercial video games, five software applications, and a comprehensive benchmark, OSWorld. Cradle is the first to enable foundation agents to follow the main storyline and complete 40-minute-long real missions in the complex AAA game Red Dead Redemption 2 (RDR2). Cradle can also create a city of a thousand people in Cities: Skylines, farm and harvest parsnips in Stardew Valley, and trade and bargain with a maximal weekly total profit of 87% in Dealer's Life 2. Cradle can not only operate daily software, like Chrome, Outlook, and Feishu, but also edit images and videos using Meitu and CapCut. Cradle greatly extends the reach of foundation agents by enabling the easy conversion of any software, especially complex games, into benchmarks to evaluate agents' various abilities and facilitate further data collection, thus paving the way for generalist agents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HumanCLAW: Can Vision-Language Models Act Through a Body?

    cs.CV 2026-07 conditional novelty 7.0 of 10

    Off-the-shelf VLMs fail closed-loop whole-body find-navigate-interact tasks (best 16.8%) because they lack embodied self-awareness, not target recognition.

  2. Long-Term Memory for VLA-based Agents in Open-World Task Execution

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    ChemBot adds dual-layer memory and future-state asynchronous inference to VLA models, enabling better long-horizon success in chemical lab automation on collaborative robots.

  3. Large Language Models are Near-Optimal Decision-Makers with a Non-Human Learning Behavior

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Across uncertainty, risk, and set-shifting tasks, LLMs generally outperformed humans and neared optimality while exhibiting distinctly non-human decision-making processes.

  4. BIMgent: Towards Autonomous Building Modeling via Computer-use Agents

    cs.AI 2025-06 conditional novelty 6.0 of 10

    BIMgent, a GUI-controlling LLM agent, completes 32% of BIM building modeling tasks end-to-end, outperforming baseline computer-use agents that complete none.

  5. ScreenExplorer: Training a Vision-Language Model for Diverse Exploration in Open GUI World

    cs.AI 2025-05 reject novelty 6.0 of 10

    A VLM trained with GRPO and a world-model curiosity reward explores a real desktop GUI more diversely than larger frozen models, but the diversity metric is nearly identical to its training reward.

  6. FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow

    cs.CL 2025-05 conditional novelty 6.0 of 10

    FullFront adds a three-task benchmark for webpage design, perception, and code generation, and finds top MLLMs still fail at fine-grained layout and interaction implementation.

  7. lmgame-Bench: How Good are LLMs at Playing Games?

    cs.AI 2025-05 conditional novelty 6.0 of 10

    lmgame-Bench turns six classic games into a scaffolded LLM evaluation suite, ranks 13 models, detects contamination, and reports RL transfer from Sokoban or Tetris to unseen games and planning tasks.

  8. Mimir: A Neuro-Symbolic Memory System with Dynamic Grounding for Embodied Agents in Interactive Environments

    cs.RO 2026-08 conditional novelty 5.0 of 10

    Separating world memory from task memory and grounding each goal in recalled evidence improves embodied-agent success rates by up to 42.5 points across 13 vision-language backbones.

  9. MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A memory-augmented LLM planner that stores and retrieves page-level summaries from past trajectories improves success rates on mobile GUI task benchmarks.

  10. AgentOrchestra: Orchestrating Multi-Agent Intelligence with the Tool-Environment-Agent(TEA) Protocol

    cs.AI 2025-06 reject novelty 5.0 of 10

    AgentOrchestra, built on the TEA protocol, reports state-of-the-art GAIA and strong HLE scores by coordinating specialized sub-agents with versioned tools, environments, and self-evolution.

  11. TextAtari: 100K Frames Game Playing with Language Agents

    cs.CL 2025-06 conditional novelty 5.0 of 10

    TextAtari is a text-based Atari benchmark for language agents; 7-8B LLMs stay below 10% of human scores in over 90% of tested conditions, and knowledge injection helps more than chain-of-thought.

  12. AppAgent-Claw: CLI Is All You Need for GUI Automation

    cs.HC 2026-04 conditional novelty 4.0 of 10

    A record-once, replay-many system converts demonstrated GUI workflows into reliable OpenClaw skills via layered visual localization and post-action validation, without runtime LLM inference.

  13. Empowering Multimodal LLMs with External Tools: A Comprehensive Survey

    cs.CV 2025-08 unverdicted novelty 2.0 of 10

    A survey paper maps how external tools are used to augment multimodal large language models across data, tasks, evaluation, and future directions.

Pith tools