Pith. sign in

REVIEW 4 cited by

Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12889 v2 pith:BTXET46L submitted 2024-09-19 cs.AI

classification cs.AI
keywords actiongamesagentsvlmsgameresearchrole-playingtasks
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, large language model (LLM)-based agents have made significant advances across various fields. One of the most popular research areas involves applying these agents to video games. Traditionally, these methods have relied on game APIs to access in-game environmental and action data. However, this approach is limited by the availability of APIs and does not reflect how humans play games. With the advent of vision language models (VLMs), agents now have enhanced visual understanding capabilities, enabling them to interact with games using only visual inputs. Despite these advances, current approaches still face challenges in action-oriented tasks, particularly in action role-playing games (ARPGs), where reinforcement learning methods are prevalent but suffer from poor generalization and require extensive training. To address these limitations, we select an ARPG, ``Black Myth: Wukong'', as a research platform to explore the capability boundaries of existing VLMs in scenarios requiring visual-only input and complex action output. We define 12 tasks within the game, with 75% focusing on combat, and incorporate several state-of-the-art VLMs into this benchmark. Additionally, we will release a human operation dataset containing recorded gameplay videos and operation logs, including mouse and keyboard actions. Moreover, we propose a novel VARP (Vision Action Role-Playing) agent framework, consisting of an action planning system and a visual trajectory system. Our framework demonstrates the ability to perform basic tasks and succeed in 90% of easy and medium-level combat scenarios. This research aims to provide new insights and directions for applying multimodal agents in complex action game environments. The code and datasets will be made available at https://varp-agent.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 7B vision-language model trained with reinforcement learning in a new four-game environment, VLM-Gym, outperforms larger proprietary models and improves its perception and reasoning abilities together.

  2. LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A hierarchical window transformer that injects image details into upsampled CLIP features and compresses them with cross-scale window attention improves MLLM fine-grained perception by 3.7% on average over LLaVA-UHD.

  3. MageBench: Bridging Large Multimodal Models to Agents

    cs.CV 2024-12 conditional novelty 6.0 of 10

    MageBench introduces a 483-scenario benchmark showing current large multimodal models are far weaker than humans at agent tasks requiring continuous visual feedback and planning.

  4. Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A multimodal knowledge graph benchmark for Monster Hunter: World, plus a training-free multi-agent graph retriever, improves MLLM accuracy on rare-domain questions from about 0.31 to 0.51 for GPT-4o.

Pith tools