Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

This paper argues that the hard remaining problem in interactive game world models is not pixels but explicit game state, and it supplies a 90-hour frame-aligned dataset to attack it.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 02:49 UTC pith:RO3KQM4A

load-bearing objection A genuinely useful survey framing, a defensible central thesis, and a promising but unvalidated dataset; send to review with the condition that the dataset claims get substantiated or demoted. the 2 major comments →

arxiv 2607.14076 v1 pith:RO3KQM4A submitted 2026-07-15 cs.CV

From Pixels to States: Rethinking Interactive World Models as Game Engines

classification cs.CV
keywords interactive world modelsgame enginesgame statevideo generationaction-state-observation loopdatasetBlack Myth: Wukongframe alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper organizes the fast-growing field of interactive world models — video generators that respond to player inputs — around the action-state-observation loop that conventional game engines use: input changes an explicit state, and pixels are rendered from that state. Surveying current methods along four dimensions, it concludes the easy parts (natural input control, explorable scenes, near-real-time generation) are largely solved, while the hard parts all depend on the game state, which most models keep implicit: outcomes determined by accumulated conditions, consequences that persist when the world leaves view, and effects that surface at rule-defined moments. To make state-aware modeling possible, it contributes a data engine for Black Myth: Wukong that captures over 90 hours of boss-fight gameplay with frame-aligned keyboard/mouse inputs, engine-exported states, RGB frames, and depth maps, plus slot-structured and semantic captions. A sympathetic reader should take away that the field's next bottleneck is explicit state, and that the paper's dataset is a direct resource for that bottleneck.

Core claim

The paper's central analytical claim is that the capabilities which remain difficult in interactive game world modeling — determining interaction outcomes from accumulated game conditions, preserving consequences beyond the current view, and surfacing effects at rule-defined moments — all revolve around the game state, which most current models keep implicit. Its central data claim is the construction of a scalable data engine that collects over 90 hours of Black Myth: Wukong gameplay at 1280×720 and 30 FPS, with frame-aligned player actions, engine-exported ground-truth states, RGB frames, depth maps, and structured/semantic captions. If both claims hold, progress toward genuinely interacti

What carries the argument

The organizing lens is the recurrent action-state-observation loop of a conventional game engine: player actions update an explicit game state according to rules, and observations are rendered from the resulting state. The paper maps every method family onto one of its four requirements — player action control, game state dynamics, state-observation persistence, and real-time interactive generation. The load-bearing object is the explicit game state record (health, stamina, cooldowns, positions, animation phases) exported by the engine, which the dataset pairs with each frame; that pairing is what would let a generative model learn rule-governed transitions rather than pixel correlations.

Load-bearing premise

The dataset's value depends on the claim that engine-exported JSON records and screen-captured frames are truly frame-aligned, but the paper reports no measured alignment error, clock-drift analysis, tolerance bounds, or manual verification counts.

What would settle it

Replay the recorded inputs and exported states through the game engine and compare each re-rendered frame with the recorded RGB frame; correct alignment should keep per-frame pixel error near compression noise across all sessions, while any timestamp drift would appear as a growing offset. A second check: train a state-explicit model on the dataset and test whether it preserves an off-screen health change; failure there would weaken the state-bottleneck claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the bottleneck is explicit state, then models that maintain a state (symbolic, latent, or 3D) should outperform purely pixel-conditional generators on tasks like off-screen persistence and rule-driven outcomes.
  • Datasets with engine-exported states become a primary resource; state annotations turn from scarce to available for one AAA title, enabling state-aware training and evaluation.
  • Memory mechanisms in long-horizon generation should be grounded in state transitions, not just stored observations, since a faithful copy of the past is not a faithful estimate of the present.
  • Interaction outcome timing should be aligned with game rules (attack startup/active/recovery) rather than treated as a latency to minimize; control latency and consequence latency are different quantities.
  • Because input control, explorability, and near-real-time generation are already approaching sufficiency, the remaining research effort shifts toward state conditioning and state-consistent generation.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same instrumentation approach could be extended to other Unreal Engine titles at low marginal cost, potentially turning many commercial games into state-annotated training sources.
  • Engine-exported states could serve as automated ground truth for evaluating world-model consistency, e.g., measuring whether generated health bars or boss phases match the recorded state after off-screen intervals.
  • A natural benchmark from this data: pause a rollout for N frames, re-enter the room, and measure whether the boss appears in the correct phase; that directly tests state persistence.
  • If state-explicit training succeeds on this dataset, it would suggest that latent-state models could be trained with state supervision, closing the interpretability gap without sacrificing scalability.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. This paper proposes a unified framing for interactive world models, organizing them around the action-state-observation loop of conventional game engines. The survey is structured along four dimensions (player action control, game state dynamics, state-observation persistence, and real-time interactive generation), with existing methods grouped into representative families and analyzed for strengths and trade-offs. The paper's central thesis is that current models already provide natural input control, explorable scenes, and near-real-time generation, while the remaining hard problems—outcome determination, off-screen persistence, and rule-timed effects—all require an explicit game state, which most models keep implicit. As a complementary contribution, the paper reports a data engine for Black Myth: Wukong, collecting over 90 hours of gameplay at 1280×720 and 30 FPS with claimed frame-aligned keyboard/mouse inputs, engine-exported game states, RGB frames, and depth maps, plus slot-structured and semantic captions.

Significance. If the analytical claim is accepted, the paper provides a useful organizing lens that connects previously separate research lines (camera control, latent-action models, memory mechanisms, and streaming/distillation) into a shared vocabulary based on game-engine design. The survey portion is internally defensible: family groupings are attributed to concrete systems, and the trade-off analyses (e.g., raw motor signals underdetermine intent; stored observations can restore an outdated world) are reasonable and specific. The dataset, if real and correctly aligned, would be a substantial resource for learning explicit-state world models and would directly support the paper's central argument. The paper does not ship code or machine-checked proofs, but it presents a falsifiable data claim—frame-level alignment between pixels and game state—that can be independently validated. The main risk is that the dataset's headline property is asserted rather than demonstrated, and the paper's conclusion closely aligns with the authors' own data-collection program, so the strength of the central claim currently rests on an unvalidated resource.

major comments (2)
  1. [§4.1–§4.2, Introduction/Conclusion] The dataset is introduced and concluded as 'frame-aligned' ground truth, but the manuscript provides no empirical support for this claim. The described procedure timestamps engine-tick JSON records and OBS-captured frames from the same clock and discards samples with 'cross-stream inconsistency.' This does not establish frame-level alignment: engine records are written when a tick updates the state, while OBS timestamps when a composited frame is captured, so an unknown pipeline offset separates a frame from the state that rendered it; the ReShade split-screen (RGB + depth) adds further buffering. A constant bias or slow clock drift would not be flagged by the stated filter yet would shift or smear every pairing. The paper reports no measured alignment error, no drift analysis, no tolerance bound, and no manual verification counts. This is load-bearing because the dataset's central promi
  2. [§4, Introduction] The dataset is a headline contribution but has no supporting statistics beyond the total 'over 90 hours.' I found no breakdown by boss encounter, no player count, no clip/duration distribution, no sample-retention rate after filtering, no annotation counts, and no release URL or license. Without these numbers, readers cannot assess coverage diversity, the effect of the 'discarded in their entirety' filter, or reproducibility. Please add a dataset table with per-encounter hours, number of players and clips, filtering statistics, caption volumes, and a release plan.
minor comments (4)
  1. [§4.2] The re-encoding step is described only as 'training-friendly formats'; specify codecs, bitrates, and whether depth maps are losslessly compressed, since these affect the value of the resource.
  2. [§3.2] The claim that 'explicit states have so far served mainly as data and evaluation' is accurate for the cited systems, but the nearby statement that integration into the generation loop is 'largely unexplored' should be qualified, since AnimeGamer and PERSIST are then discussed in §3.3 as partially closing this gap.
  3. [§4.1] The text says the capture stage introduces 'minor overhead' and 'requires no additional hardware' but does not quantify the overhead. Please report typical CPU/GPU cost and any FPS impact on crowdsourced capture.
  4. [Figure 2] The pipeline diagram labels several blocks (e.g., 'Camera/Character Occlusion') that are not defined in the text; please add a short description of the occlusion criterion and the abnormal-frame filter.

Circularity Check

0 steps flagged

No significant circularity; the paper is a survey/position plus dataset, with no fitted prediction, no self-citation chain that reduces the central claim to its own inputs, and no derivation-by-definition.

full rationale

This paper contains no predictive derivation chain in the sense that would support a circularity finding. The four-dimensional framework is explicitly introduced as an organizing lens ('Taking this loop as an organizing lens, this paper examines interactive game world modeling along four dimensions'), and the concluding claim that the remaining hard capabilities 'all revolve around the game state, which most models keep implicit' is an empirical assessment of surveyed external systems, not a conclusion forced by a fitted parameter or by the definition of the framework. No equations are derived, no quantity is predicted from the collected data, and no fitted input is renamed as a prediction. Self-citations appear (Sekai [42], WildWorld [43], Yume [51,52], OmniWorld [111]), but none is load-bearing: the scarcity claim is also supported by independent references such as [6,92,30], and the Section 3.2 observation that explicit states have so far served mainly as data and evaluation is independently supported by EgoCS-400K [20] and AnimeGamer [9]. Removing the self-citations would not change the survey's characterization. The reviewer's concern about frame-alignment validation in Sections 4.1–4.2 is a real empirical/data-quality risk — no measured alignment error or clock-drift analysis is reported — but it is not a circularity: the alignment procedure is described as a concrete timestamping and filtering process, not defined in terms of the dataset's own target. No uniqueness theorem, borrowed ansatz, or renamed known result is imported from the authors' prior work. Accordingly, no specific circular step can be exhibited, and the appropriate score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

No numbers are fitted anywhere in the paper: the survey contains no derivations, and the dataset is ground-truth acquisition, not model fitting. The load-bearing premises are all domain assumptions about the correctness of the capture/alignment pipeline and about the choice of the action-state-observation loop as the organizing lens. The paper introduces no new physical or conceptual entities (explicit game state is a standard game-engine concept; the 'data engine' is a capture pipeline, not an invented entity). Unreported design choices that a reader would need include the slot-caption window length, the 'abnormal frame' filtering thresholds, and the timestamp-synchronization tolerance.

axioms (4)
  • domain assumption The recurrent action-state-observation loop of conventional game engines is the correct organizing model for generative interactive world modeling (Section 1, Section 3).
    The entire taxonomy is built on this lens; if the loop is not the right abstraction for comparing generative systems, the four dimensions lose their grounding. The paper argues rather than proves this.
  • domain assumption Engine-instrumented export at every tick yields correct and complete ground-truth game state (Section 4.1).
    'We instrument the game engine to export interaction data at every tick as a structured stream' — the dataset's value depends on these exports being accurate, untruncated, and synchronous with rendering, but no verification is reported.
  • domain assumption System-clock timestamps provide a common temporal basis that makes engine-state-to-frame alignment frame-accurate (Section 4.1, 4.2).
    Frame alignment is asserted from shared-clock timestamps; no alignment error distribution, drift analysis, or manual verification is provided. This is the load-bearing premise for the 'frame-aligned ground truth' claim.
  • domain assumption The Qwen3-VL-generated semantic captions are accurate enough to serve as supervision (Section 4.2).
    Semantic captions are produced by a large VLM with no stated human validation, quality metric, or error analysis, yet the paper proposes them as 'next-state supervision for language-based state transition.'

pith-pipeline@v1.3.0-alltime-deepseek · 15456 in / 11945 out tokens · 113396 ms · 2026-08-02T02:49:02.089309+00:00 · methodology

0 comments
read the original abstract

Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly regarded as potential next-generation game engines. Realizing a genuinely interactive game world, however, requires interaction outcomes that follow rules over evolving game conditions, consequences that persist over long horizons, and a generation loop that operates in real time. Conventional game engines realize these properties through a recurrent action-state-observation loop, in which player actions update an explicit game state according to predefined rules and observations are rendered from the resulting state. Taking this loop as an organizing lens, this paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation. For each dimension, we start from the capabilities required by an interactive game world, group existing approaches into representative families, and discuss the strengths and trade-offs of each family. Complementing this analysis, we present a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay with frame-aligned player actions, ground-truth game states, and visual observations, together with structured and semantic annotations, as a resource for state-aware game world modeling. We hope this paper offers a clear picture of where the field stands and fosters progress toward interactive game worlds.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation

    cs.CV 2026-07 conditional novelty 6.0

    Coupling explicit state prediction with MoT-style video generation raises mechanics fidelity of Street Fighter 3 rollouts by about 18.6% over stateless game world models.

Reference graph

Works this paper leans on

115 extracted references · 35 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Alonso, A

    E. Alonso, A. Jelley, V . Micheli, A. Kanervisto, A. J. Storkey, T. Pearce, and F. Fleuret. Diffusion for world modeling: Visual details matter in Atari.Advances in Neural Information Processing Systems, 37:58757–58791, 2024

  2. [2]

    A. Bar, G. Zhou, D. Tran, T. Darrell, and Y. LeCun. Navigation world models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15791–15801, 2025

  3. [3]

    Brooks, B

    T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh. Video generation models as world simulators. https:// openai.com/research/video-generation-models-as-world-simulators , 2024. OpenAI technical report

  4. [4]

    Bruce, M

    J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. InInternational Conference on Machine Learning, pages 4603–4623. PMLR, 2024

  5. [5]

    S. Cai, C. Yang, L. Zhang, Y. Guo, J. Xiao, Z. Yang, Y. Xu, Z. Yang, A. Yuille, L. Guibas, M. Agrawala, L. Jiang, and G. Wetzstein. Mixture of contexts for long video generation. InInternational Conference on Learning Representations, 2026

  6. [6]

    H. Che, X. He, Q. Liu, C. Jin, and H. Chen. GameGen-X: Interactive open-world game video generation. InInternational Conference on Learning Representations, 2025

  7. [7]

    K. Chen, D. Liang, X. Zhou, Y. Ding, X. Liu, P . Wan, and X. Bai. Out of sight but not out of mind: Hybrid memory for dynamic video world models.arXiv preprint arXiv:2603.25716, 2026

  8. [8]

    T. Chen, Z. Ding, A. Li, C. Zhang, Z. Xiao, Y. Wang, and C. Jin. Recurrent autoregressive diffusion: Global memory meets local attention.arXiv preprint arXiv:2511.12940, 2025

  9. [9]

    Cheng, Y

    J. Cheng, Y. Ge, Y. Ge, J. Liao, and Y. Shan. AnimeGamer: Infinite anime life simulation with next game state prediction. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10875–10885, 2025

  10. [10]

    Chiappa, S

    S. Chiappa, S. Racanière, D. Wierstra, and S. Mohamed. Recurrent environment simulators. In International Conference on Learning Representations, 2017

  11. [11]

    Quevedo, Q

    Decart, J. Quevedo, Q. McIntyre, S. Campbell, X. Chen, and R. Wachen. Oasis: A universe in a transformer.https://oasis-model.github.io/, 2024. Decart and Etched technical blog post

  12. [12]

    C. Deng, D. Zhu, K. Li, C. Gou, F. Li, Z. Wang, S. Zhong, W. Yu, X. Nie, Z. Song, G. Shi, and H. Fan. Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025

  13. [13]

    DreamX Team, Y. Bai, R. Chen, X. Chu, R. Dang, H. Dou, B. Gao, Q. Gu, S. Hong, J. Lei, G. Li, J. Li, R. Lin, Q. Shi, B. Song, L. Sun, J. Tang, R. Tian, J. Wang, J. Wu, P . Zhang, S. Zhang, and J. Zhu. DreamX-World 1.0: A general-purpose interactive world model.arXiv preprint arXiv:2606.16993, 2026

  14. [14]

    Esser, S

    P . Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach. Scaling rectified 8 flow transformers for high-resolution image synthesis. InInternational Conference on Machine Learning, pages 12606–12633. PMLR, 2024

  15. [15]

    R. Feng, H. Zhang, Z. Yang, J. Xiao, Z. Shu, Z. Liu, A. Zheng, Y. Huang, Y. Liu, and H. Zhang. The matrix: Infinite-horizon world generation with real-time moving control.Advances in Neural Information Processing Systems, 38, 2025

  16. [16]

    Garcin, T

    S. Garcin, T. Walker, S. McDonagh, T. Pearce, H. Bilen, T. He, K. Wang, and J. Bian. Beyond pixel context windows: Neural world simulators with persistent 3D state. InInternational Conference on Machine Learning. PMLR, 2026

  17. [17]

    Garrido, T

    Q. Garrido, T. Nagarajan, B. Terver, N. Ballas, Y. LeCun, and M. Rabbat. Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026

  18. [18]

    Z. Ge, H. Huang, M. Zhou, J. Li, G. Wang, S. Tang, and Y. Zhuang. WorldGPT: Empowering LLM as multimodal world model. InProceedings of the 32nd ACM International Conference on Multimedia, pages 7346–7355, 2024

  19. [19]

    J. Guo, Y. Ye, T. He, H. Wu, Y. Jiang, T. Pearce, and J. Bian. MineWorld: a real-time and open-source interactive world model on Minecraft.arXiv preprint arXiv:2504.08388, 2025

  20. [20]

    R. Guo, D. Liang, Y. Liu, F. Liu, T. Huang, G. P . Hancke, and R. W. H. Lau. EgoCS-400K: An egocentric gameplay dataset for world models.arXiv preprint arXiv:2606.18180, 2026

  21. [21]

    Y. Guo, C. Yang, H. He, Y. Zhao, M. Wei, Z. Yang, W. Huang, and D. Lin. End-to-end training for autoregressive video diffusion via self-resampling. InEuropean Conference on Computer Vision. Springer, 2026

  22. [22]

    Ha and J

    D. Ha and J. Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018

  23. [23]

    HaCohen, B

    Y. HaCohen, B. Brazowski, N. Chiprut, Y. Bitterman, A. Kvochko, A. Berkowitz, D. Shalem, D. Lifschitz, D. Moshe, E. Porat, et al. LTX-2: Efficient joint audio-visual foundation model.arXiv preprint arXiv:2601.03233, 2026

  24. [24]

    Hafner, T

    D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. InInternational Conference on Machine Learning, pages 2555–2565. PMLR, 2019

  25. [25]

    Hafner, J

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse control tasks through world models. Nature, 640(8059):647–653, 2025

  26. [26]

    Hafner, W

    D. Hafner, W. Yan, and T. Lillicrap. Training agents inside of scalable world models.arXiv preprint arXiv:2509.24527, 2025

  27. [27]

    S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu. Reasoning with language model is planning with world model. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8154–8173, 2023

  28. [28]

    H. He, Y. Xu, Y. Guo, G. Wetzstein, B. Dai, H. Li, and C. Yang. CameraCtrl: Enabling camera control for video diffusion models. InInternational Conference on Learning Representations, 2025

  29. [29]

    X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, et al. Matrix-Game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025

  30. [30]

    Y. He, C. D. Weilbach, M. E. Wojciechowska, Y. Sun, Y. Zhang, and F. Wood. PLAICraft: Large-scale time-aligned vision-speech-action dataset for embodied AI.arXiv preprint arXiv:2505.12707, 2025

  31. [31]

    Y. Hong, Y. Mei, C. Ge, Y. Xu, Y. Zhou, S. Bi, Y. Hold-Geoffroy, M. Roberts, M. Fisher, E. Shechtman, K. Sunkavalli, F. Liu, Z. Li, and H. Tan. RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025

  32. [32]

    InSpatio-WorldFM: An open-source real-time generative frame model.arXiv preprint arXiv:2603.11911, 2026

    InSpatio Team. InSpatio-WorldFM: An open-source real-time generative frame model.arXiv preprint arXiv:2603.11911, 2026. 9

  33. [33]

    Jiang, Y

    Y. Jiang, Y. Gu, I. W. Tsang, and M. Z. Shou. Olaf-World: Orienting latent actions for video world modeling. InInternational Conference on Machine Learning. PMLR, 2026

  34. [34]

    Kanervisto, D

    A. Kanervisto, D. Bignell, L. Y. Wen, M. Grayson, R. Georgescu, S. V . Macua, S. Z. Tan, T. Rashid, T. Pearce, Y. Cao, A. Lemkhenter, C. Jiang, G. Costello, G. Gupta, M. Tot, S. Ishida, T. Gupta, U. Arora, R. W. White, S. Devlin, C. Morrison, and K. Hofmann. World and human action models towards gameplay ideation.Nature, 638(8051):656–663, 2025

  35. [35]

    B. Kang, Y. Yue, R. Lu, Z. Lin, Y. Zhao, K. Wang, G. Huang, and J. Feng. How far is video generation from world model: A physical law perspective. InInternational Conference on Machine Learning. PMLR, 2025

  36. [36]

    S. W. Kim, Y. Zhou, J. Philion, A. Torralba, and S. Fidler. Learning to simulate dynamic environments with GameGAN. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1228–1237, 2020

  37. [37]

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y. Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P . Li, S. Li, W. Wang, W. Yu, X. Deng, Y. Li, Y. Chen, Y. Cui, Y. Peng, Z. Yu, Z. He, Z. Xu, Z. Zhou, Z. Xu, Y. Tao, ...

  38. [38]

    G. Li, B. Li, J. Chen, X. Hu, L. Zhao, and P .-T. Jiang. MagicWorld: Towards long-horizon stability for interactive video world exploration.arXiv preprint arXiv:2511.18886, 2025

  39. [39]

    J. Li, J. Tang, Z. Xu, L. Wu, Y. Zhou, S. Shao, T. Yu, Z. Cao, and Q. Lu. Hunyuan-GameCraft: High-dynamic interactive game video generation with hybrid history condition.arXiv preprint arXiv:2506.17201, 2025

  40. [40]

    R. Li, P . Torr, A. Vedaldi, and T. Jakab. VMem: Consistent interactive video scene generation with surfel-indexed view memory. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2025

  41. [41]

    R. Li, B. Yi, J. Liu, H. Gao, Y. Ma, and A. Kanazawa. Cameras as relative positional encoding.Advances in Neural Information Processing Systems, 38, 2025

  42. [42]

    Z. Li, C. Li, X. Mao, S. Lin, M. Li, S. Zhao, Z. Xu, X. Li, Y. Feng, J. Sun, Z. Li, F. Zhang, J. Ai, Z. Wang, Y. Wu, T. He, J. Pang, Y. Qiao, Y. Jia, and K. Zhang. Sekai: A video dataset towards world exploration. Advances in Neural Information Processing Systems, 38, 2025

  43. [43]

    Z. Li, Z. Meng, S. Shi, W. Peng, Y. Wu, B. Zheng, C. Li, and K. Zhang. WildWorld: A large-scale dataset for dynamic world modeling with actions and explicit state toward generative ARPG.arXiv preprint arXiv:2603.23497, 2026

  44. [44]

    Liang, L

    W. Liang, L. Yu, L. Luo, S. Iyer, N. Dong, C. Zhou, G. Ghosh, M. Lewis, W.-t. Yih, L. Zettlemoyer, and X. V . Lin. Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.T ransactions on Machine Learning Research, 2025

  45. [45]

    H. J. Lillemark, B. Huang, F. Zhan, Y. Du, and T. A. Keller. Flow equivariant world models: Structured memory for dynamic environments. InInternational Conference on Machine Learning. PMLR, 2026

  46. [46]

    J. Liu, X. Liu, K. Mei, Y. Wen, M.-H. Yang, and W. Liu. Streaming autoregressive video generation via diagonal distillation. InInternational Conference on Learning Representations, 2026

  47. [47]

    J. Liu, C. Ni, M. Liu, C. Peng, F. Wang, S. Shen, M. Pollefeys, M. Tomizuka, A. Tewari, and P . O. Kristensson. Towards interactive video world modeling: Frontiers, challenges, benchmarks, and future trends.arXiv preprint arXiv:2606.01164, 2026

  48. [48]

    K. Liu, W. Hu, J. Xu, Y. Shan, and S. Lu. Rolling forcing: Autoregressive long video diffusion in real time. InInternational Conference on Learning Representations, 2026. 10

  49. [49]

    Y. Ma, X. Liu, X. Chen, W. Liu, C. Wu, Z. Wu, Z. Pan, Z. Xie, H. Zhang, X. Yu, L. Zhao, Y. Wang, J. Liu, and C. Ruan. JanusFlow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7739–7751, 2025

  50. [50]

    Z. Ma, M. Liufu, and G. Gkioxari. Out of sight, out of mind? evaluating state evolution in video world models.arXiv preprint arXiv:2603.13215, 2026

  51. [51]

    X. Mao, Z. Li, C. Li, X. Xu, K. Ying, and K. Zhang. Yume1.5: A text-controlled interactive world generation model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7752–7761, 2026

  52. [52]

    X. Mao, S. Lin, Z. Li, C. Li, W. Peng, T. He, J. Pang, M. Chi, Y. Qiao, and K. Zhang. Yume: An interactive world generation model.arXiv preprint arXiv:2507.17744, 2025

  53. [53]

    Menapace, S

    W. Menapace, S. Lathuilière, A. Siarohin, C. Theobalt, S. Tulyakov, V . Golyanik, and E. Ricci. Playable environments: Video manipulation in space and time. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3574–3583, 2022

  54. [54]

    Menapace, S

    W. Menapace, S. Lathuilière, S. Tulyakov, A. Siarohin, and E. Ricci. Playable video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10061–10070, 2021

  55. [55]

    Menapace, A

    W. Menapace, A. Siarohin, S. Lathuilière, P . Achlioptas, V . Golyanik, S. Tulyakov, and E. Ricci. Prompt- able game models: Text-guided game simulation via masked diffusion models.ACM T ransactions on Graphics, 43(2):1–16, 2024

  56. [56]

    J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh. Action-conditional video prediction using deep networks in Atari games.Advances in Neural Information Processing Systems, 28:2863–2871, 2015

  57. [57]

    Oshima, Y

    Y. Oshima, Y. Iwasawa, M. Suzuki, Y. Matsuo, and H. Furuta. WorldPack: Compressed memory improves spatial consistency in video world modeling.arXiv preprint arXiv:2512.02473, 2025

  58. [58]

    Parker-Holder and S

    J. Parker-Holder and S. Fruchter. Genie 3: A new frontier for world models. https://deepmind. google/discover/blog/genie-3-a-new-frontier-for-world-models/ , 2025. Google DeepMind blog, August 5, 2025

  59. [59]

    Peebles and S

    W. Peebles and S. Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023

  60. [60]

    R. Po, Y. Nitzan, R. Zhang, B. Chen, T. Dao, E. Shechtman, G. Wetzstein, and X. Huang. Long-context state-space video world models. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8733–8744, 2025

  61. [61]

    Y. Qiu, Z. Zhao, W. W. Li, Y. Ziser, A. Korhonen, S. B. Cohen, and E. M. Ponti. Self-improving world modelling with latent actions.arXiv preprint arXiv:2602.06130, 2026

  62. [62]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022

  63. [63]

    Savov, N

    N. Savov, N. Kazemi, D. Zhang, D. P . Paudel, X. Wang, and L. Van Gool. StateSpaceDiffuser: Bringing long context to diffusion world models.Advances in Neural Information Processing Systems, 38, 2025

  64. [64]

    W. Sun, F. Wei, J. Zhao, X. Chen, Z. Chen, H. Zhang, J. Zhang, and Y. Lu. From virtual games to real-world play.arXiv preprint arXiv:2506.18901, 2025

  65. [65]

    J. Tang, J. Liu, J. Li, L. Wu, H. Yang, P . Zhao, S. Gong, X. Yuan, S. Shao, L. Zhang, and Q. Lu. Hunyuan- GameCraft-2: Instruction-following interactive game world model.arXiv preprint arXiv:2511.23429, 2025. 11

  66. [66]

    Z. Tong, Y. Jin, H. Lai, Z. Wang, Z. Xing, K. Cheng, H. Xu, Z. Pu, S. Zhu, R. Feng, J. Zhao, Y. Zhang, H. Tang, and L. Shao. SCOPE: Simulating cross-game operations in playable environments for FPS world models.arXiv preprint arXiv:2605.23345, 2026

  67. [67]

    Valevski, Y

    D. Valevski, Y. Leviathan, M. Arar, and S. Fruchter. Diffusion models are real-time game engines. In International Conference on Learning Representations, 2025

  68. [68]

    T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

  69. [69]

    L. Wang, Z. Chen, Y. Du, D. Yan, W. Ge, G. Shen, X. Xu, L. Wu, M. Chen, T. Xu, P . Ren, X. Tao, P . Wan, and Y.-C. Chen. A mechanistic view on video generation as world models: State and dynamics.arXiv preprint arXiv:2601.17067, 2026

  70. [70]

    R. Wang, G. Todd, Z. Xiao, X. Yuan, M.-A. Côté, P . Clark, and P . Jansen. Can language models serve as text-based world simulators? InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–17, 2024

  71. [71]

    X. Wang, X. Zhang, Z. Luo, Q. Sun, Y. Cui, J. Wang, F. Zhang, Y. Wang, Z. Li, Q. Yu, Y. Zhao, Y. Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. Wang, G. Liu, Z. He, X. Yang, J. Liu, Y. Lin, T. Huang, and Z. Wang. Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024

  72. [72]

    Z. Wang, D. Chen, Z. Xing, Z. Tong, Y. Zhang, X. Yang, and Y. Jin. ReactiveGWM: Steering NPC in reactive game world models.arXiv preprint arXiv:2605.15256, 2026

  73. [73]

    Z. Wang, Z. Liu, J. Li, K. Huang, B. Xu, F. Kang, M. An, P . Wang, B. Jiang, Y. Wei, Y. Xietian, J. Pei, L. Hu, B. Jiang, H. Xue, Z. Wang, H. Sun, W. Li, W. Ouyang, X. He, Y. Liu, Y. Li, and Y. Zhou. Matrix-Game 3.0: Real-time and streaming interactive world model with long-horizon memory.arXiv preprint arXiv:2604.08995, 2026

  74. [74]

    Z. Wang, Z. Yuan, X. Wang, Y. Li, T. Chen, M. Xia, P . Luo, and Y. Shan. MotionCtrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024

  75. [75]

    Wiedemer, Y

    T. Wiedemer, Y. Li, P . Vicol, S. S. Gu, N. Matarese, K. Swersky, B. Kim, P . Jaini, and R. Geirhos. Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025

  76. [76]

    J. Wu, S. Yin, N. Feng, X. He, D. Li, J. Hao, and M. Long. iVideoGPT: Interactive VideoGPTs are scalable world models.Advances in Neural Information Processing Systems, 37:68082–68119, 2024

  77. [77]

    R. Wu, X. He, M. Cheng, T. Yang, Y. Zhang, Z. Kang, X. Cai, X. Wei, C. Guo, C. Li, et al. Infinite-World: Scaling interactive world models to 1000-frame horizons via pose-free hierarchical memory. In International Conference on Machine Learning. PMLR, 2026

  78. [78]

    Xiang, J

    C. Xiang, J. Liu, J. Zhang, X. Yang, Z. Fang, S. Wang, Z. Wang, Y. Zou, H. Su, and J. Zhu. Geometry- aware rotary position embedding for consistent video world model.arXiv preprint arXiv:2602.07854, 2026

  79. [79]

    Xiang, Y

    J. Xiang, Y. Gu, Z. Liu, Z. Feng, Q. Gao, Y. Hu, B. Huang, G. Liu, Y. Yang, K. Zhou, D. Abrahamyan, A. Ahmad, G. Bannur, J. Chen, K. Chen, M. Deng, R. Han, X. Huang, H. Kang, Z. Liu, E. Ma, H. Ren, Y. Shinde, R. Shingre, R. Tanikella, K. Tao, D. Yang, X. Yu, C. Zeng, B. Zhou, Z. Liu, Z. Hu, and E. P . Xing. PAN: A world model for general, interactable, an...

  80. [80]

    Xiang, G

    J. Xiang, G. Liu, Y. Gu, Q. Gao, Y. Ning, Y. Zha, Z. Feng, T. Tao, S. Hao, Y. Shi, Z. Liu, E. P . Xing, and Z. Hu. Pandora: Towards general world model with natural language actions and video states.arXiv preprint arXiv:2406.09455, 2024

Showing first 80 references.