Pith. sign in

REVIEW 13 cited by

GameGen-X: Interactive Open-world Game Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00769 v3 pith:HOKN4SMZ submitted 2024-11-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords videogamegenerationmodelfirstinteractiveopen-worldcontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce GameGen-X, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. This model facilitates high-quality, open-domain generation by simulating an extensive array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation. To realize this vision, we first collected and built an Open-World Video Game Dataset from scratch. It is the first and largest dataset for open-world game video generation and control, which comprises over a million diverse gameplay video clips sampling from over 150 games with informative captions from GPT-4o. GameGen-X undergoes a two-stage training process, consisting of foundation model pre-training and instruction tuning. Firstly, the model was pre-trained via text-to-video generation and video continuation, endowing it with the capability for long-sequence, high-quality open-domain game video generation. Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts. This allows the model to adjust latent representations based on user inputs, unifying character interaction and scene content control for the first time in video generation. During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated video content.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Distributed video-generation agents can create a spatiotemporally consistent shared world by tiling four views, exchanging cross-agent attention, and querying a spatial memory cache.

  2. UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A video-generation world model that warps positional encodings of memory frames to target viewpoints achieves state-of-the-art long-term consistency and camera control.

  3. WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

    cs.CV 2025-12 unverdicted novelty 6.0 of 10

    WorldPlay uses dual action representation, reconstituted context memory, and context forcing distillation to produce consistent 720p streaming video at 24 FPS for interactive world modeling.

  4. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  5. Matrix-Game: Interactive World Foundation Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 17B-parameter diffusion model generates controllable, physically consistent Minecraft video from a reference image and user actions, beating Oasis and MineWorld on a new benchmark.

  6. Video World Models with Long-term Spatial Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An autoregressive video world model with a persistent static point-cloud spatial memory and sparse episodic keyframes improves revisit consistency over point-cloud-conditioned baselines.

  7. Recurrent Autoregressive Diffusion: Global Memory Meets Local Attention

    cs.CV 2025-11 conditional novelty 5.0 of 10

    Augmenting a diffusion video transformer with an RNN memory block and frame-wise overlapping attention improves long-horizon consistency, with simple LSTM matching newer Mamba2 and TTT memory blocks.

  8. EvoCurr: Self-evolving Curriculum with Behavior Code Generation for Complex Decision-making

    cs.AI 2025-08 reject novelty 5.0 of 10

    EvoCurr couples an LLM curriculum designer with an LLM code-generating solver, but its only reported success is 1 of 5 runs and no direct baseline is shown.

  9. Impact-driven Context Filtering For Cross-file Code Completion

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    The manuscript's abstract claims a new code-completion filtering method, yet the body contains an unrelated 3D animation paper, leaving the claimed work unverifiable.

  10. RoboScape: Physics-informed Embodied World Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RoboScape jointly learns RGB video, depth, and keypoint-token consistency in one autoregressive world model, improving video quality, geometry, action control, synthetic-data policy training, and policy evaluation for...

  11. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  12. VRAG: Learning World Models for Interactive Video Generation

    cs.CV 2025-05 unverdicted novelty 5.0 of 10

    VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft a...

  13. DriveX: Omni Scene Modeling for Learning Generalizable World Knowledge in Autonomous Driving

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DriveX predicts future latent BEV features from driving video and shows consistent, modest gains on occupancy, flow, and end-to-end driving, though no code is released.

Pith tools