REVIEW 2 major objections 4 minor 1 cited by
This paper argues that the hard remaining problem in interactive game world models is not pixels but explicit game state, and it supplies a 90-hour frame-aligned dataset to attack it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 02:49 UTC pith:RO3KQM4A
load-bearing objection A genuinely useful survey framing, a defensible central thesis, and a promising but unvalidated dataset; send to review with the condition that the dataset claims get substantiated or demoted. the 2 major comments →
From Pixels to States: Rethinking Interactive World Models as Game Engines
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central analytical claim is that the capabilities which remain difficult in interactive game world modeling — determining interaction outcomes from accumulated game conditions, preserving consequences beyond the current view, and surfacing effects at rule-defined moments — all revolve around the game state, which most current models keep implicit. Its central data claim is the construction of a scalable data engine that collects over 90 hours of Black Myth: Wukong gameplay at 1280×720 and 30 FPS, with frame-aligned player actions, engine-exported ground-truth states, RGB frames, depth maps, and structured/semantic captions. If both claims hold, progress toward genuinely interacti
What carries the argument
The organizing lens is the recurrent action-state-observation loop of a conventional game engine: player actions update an explicit game state according to rules, and observations are rendered from the resulting state. The paper maps every method family onto one of its four requirements — player action control, game state dynamics, state-observation persistence, and real-time interactive generation. The load-bearing object is the explicit game state record (health, stamina, cooldowns, positions, animation phases) exported by the engine, which the dataset pairs with each frame; that pairing is what would let a generative model learn rule-governed transitions rather than pixel correlations.
Load-bearing premise
The dataset's value depends on the claim that engine-exported JSON records and screen-captured frames are truly frame-aligned, but the paper reports no measured alignment error, clock-drift analysis, tolerance bounds, or manual verification counts.
What would settle it
Replay the recorded inputs and exported states through the game engine and compare each re-rendered frame with the recorded RGB frame; correct alignment should keep per-frame pixel error near compression noise across all sessions, while any timestamp drift would appear as a growing offset. A second check: train a state-explicit model on the dataset and test whether it preserves an off-screen health change; failure there would weaken the state-bottleneck claim.
If this is right
- If the bottleneck is explicit state, then models that maintain a state (symbolic, latent, or 3D) should outperform purely pixel-conditional generators on tasks like off-screen persistence and rule-driven outcomes.
- Datasets with engine-exported states become a primary resource; state annotations turn from scarce to available for one AAA title, enabling state-aware training and evaluation.
- Memory mechanisms in long-horizon generation should be grounded in state transitions, not just stored observations, since a faithful copy of the past is not a faithful estimate of the present.
- Interaction outcome timing should be aligned with game rules (attack startup/active/recovery) rather than treated as a latency to minimize; control latency and consequence latency are different quantities.
- Because input control, explorability, and near-real-time generation are already approaching sufficiency, the remaining research effort shifts toward state conditioning and state-consistent generation.
Where Pith is reading between the lines
- The same instrumentation approach could be extended to other Unreal Engine titles at low marginal cost, potentially turning many commercial games into state-annotated training sources.
- Engine-exported states could serve as automated ground truth for evaluating world-model consistency, e.g., measuring whether generated health bars or boss phases match the recorded state after off-screen intervals.
- A natural benchmark from this data: pause a rollout for N frames, re-enter the room, and measure whether the boss appears in the correct phase; that directly tests state persistence.
- If state-explicit training succeeds on this dataset, it would suggest that latent-state models could be trained with state supervision, closing the interpretability gap without sacrificing scalability.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a unified framing for interactive world models, organizing them around the action-state-observation loop of conventional game engines. The survey is structured along four dimensions (player action control, game state dynamics, state-observation persistence, and real-time interactive generation), with existing methods grouped into representative families and analyzed for strengths and trade-offs. The paper's central thesis is that current models already provide natural input control, explorable scenes, and near-real-time generation, while the remaining hard problems—outcome determination, off-screen persistence, and rule-timed effects—all require an explicit game state, which most models keep implicit. As a complementary contribution, the paper reports a data engine for Black Myth: Wukong, collecting over 90 hours of gameplay at 1280×720 and 30 FPS with claimed frame-aligned keyboard/mouse inputs, engine-exported game states, RGB frames, and depth maps, plus slot-structured and semantic captions.
Significance. If the analytical claim is accepted, the paper provides a useful organizing lens that connects previously separate research lines (camera control, latent-action models, memory mechanisms, and streaming/distillation) into a shared vocabulary based on game-engine design. The survey portion is internally defensible: family groupings are attributed to concrete systems, and the trade-off analyses (e.g., raw motor signals underdetermine intent; stored observations can restore an outdated world) are reasonable and specific. The dataset, if real and correctly aligned, would be a substantial resource for learning explicit-state world models and would directly support the paper's central argument. The paper does not ship code or machine-checked proofs, but it presents a falsifiable data claim—frame-level alignment between pixels and game state—that can be independently validated. The main risk is that the dataset's headline property is asserted rather than demonstrated, and the paper's conclusion closely aligns with the authors' own data-collection program, so the strength of the central claim currently rests on an unvalidated resource.
major comments (2)
- [§4.1–§4.2, Introduction/Conclusion] The dataset is introduced and concluded as 'frame-aligned' ground truth, but the manuscript provides no empirical support for this claim. The described procedure timestamps engine-tick JSON records and OBS-captured frames from the same clock and discards samples with 'cross-stream inconsistency.' This does not establish frame-level alignment: engine records are written when a tick updates the state, while OBS timestamps when a composited frame is captured, so an unknown pipeline offset separates a frame from the state that rendered it; the ReShade split-screen (RGB + depth) adds further buffering. A constant bias or slow clock drift would not be flagged by the stated filter yet would shift or smear every pairing. The paper reports no measured alignment error, no drift analysis, no tolerance bound, and no manual verification counts. This is load-bearing because the dataset's central promi
- [§4, Introduction] The dataset is a headline contribution but has no supporting statistics beyond the total 'over 90 hours.' I found no breakdown by boss encounter, no player count, no clip/duration distribution, no sample-retention rate after filtering, no annotation counts, and no release URL or license. Without these numbers, readers cannot assess coverage diversity, the effect of the 'discarded in their entirety' filter, or reproducibility. Please add a dataset table with per-encounter hours, number of players and clips, filtering statistics, caption volumes, and a release plan.
minor comments (4)
- [§4.2] The re-encoding step is described only as 'training-friendly formats'; specify codecs, bitrates, and whether depth maps are losslessly compressed, since these affect the value of the resource.
- [§3.2] The claim that 'explicit states have so far served mainly as data and evaluation' is accurate for the cited systems, but the nearby statement that integration into the generation loop is 'largely unexplored' should be qualified, since AnimeGamer and PERSIST are then discussed in §3.3 as partially closing this gap.
- [§4.1] The text says the capture stage introduces 'minor overhead' and 'requires no additional hardware' but does not quantify the overhead. Please report typical CPU/GPU cost and any FPS impact on crowdsourced capture.
- [Figure 2] The pipeline diagram labels several blocks (e.g., 'Camera/Character Occlusion') that are not defined in the text; please add a short description of the occlusion criterion and the abnormal-frame filter.
Circularity Check
No significant circularity; the paper is a survey/position plus dataset, with no fitted prediction, no self-citation chain that reduces the central claim to its own inputs, and no derivation-by-definition.
full rationale
This paper contains no predictive derivation chain in the sense that would support a circularity finding. The four-dimensional framework is explicitly introduced as an organizing lens ('Taking this loop as an organizing lens, this paper examines interactive game world modeling along four dimensions'), and the concluding claim that the remaining hard capabilities 'all revolve around the game state, which most models keep implicit' is an empirical assessment of surveyed external systems, not a conclusion forced by a fitted parameter or by the definition of the framework. No equations are derived, no quantity is predicted from the collected data, and no fitted input is renamed as a prediction. Self-citations appear (Sekai [42], WildWorld [43], Yume [51,52], OmniWorld [111]), but none is load-bearing: the scarcity claim is also supported by independent references such as [6,92,30], and the Section 3.2 observation that explicit states have so far served mainly as data and evaluation is independently supported by EgoCS-400K [20] and AnimeGamer [9]. Removing the self-citations would not change the survey's characterization. The reviewer's concern about frame-alignment validation in Sections 4.1–4.2 is a real empirical/data-quality risk — no measured alignment error or clock-drift analysis is reported — but it is not a circularity: the alignment procedure is described as a concrete timestamping and filtering process, not defined in terms of the dataset's own target. No uniqueness theorem, borrowed ansatz, or renamed known result is imported from the authors' prior work. Accordingly, no specific circular step can be exhibited, and the appropriate score is 0.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption The recurrent action-state-observation loop of conventional game engines is the correct organizing model for generative interactive world modeling (Section 1, Section 3).
- domain assumption Engine-instrumented export at every tick yields correct and complete ground-truth game state (Section 4.1).
- domain assumption System-clock timestamps provide a common temporal basis that makes engine-state-to-frame alignment frame-accurate (Section 4.1, 4.2).
- domain assumption The Qwen3-VL-generated semantic captions are accurate enough to serve as supervision (Section 4.2).
read the original abstract
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a data-driven route toward this goal by predicting future observations conditioned on user actions, and are increasingly regarded as potential next-generation game engines. Realizing a genuinely interactive game world, however, requires interaction outcomes that follow rules over evolving game conditions, consequences that persist over long horizons, and a generation loop that operates in real time. Conventional game engines realize these properties through a recurrent action-state-observation loop, in which player actions update an explicit game state according to predefined rules and observations are rendered from the resulting state. Taking this loop as an organizing lens, this paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation. For each dimension, we start from the capabilities required by an interactive game world, group existing approaches into representative families, and discuss the strengths and trade-offs of each family. Complementing this analysis, we present a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay with frame-aligned player actions, ground-truth game states, and visual observations, together with structured and semantic annotations, as a resource for state-aware game world modeling. We hope this paper offers a clear picture of where the field stands and fosters progress toward interactive game worlds.
Forward citations
Cited by 1 Pith paper
-
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation
Coupling explicit state prediction with MoT-style video generation raises mechanics fidelity of Street Fighter 3 rollouts by about 18.6% over stateless game world models.
Reference graph
Works this paper leans on
-
[1]
Alonso, A
E. Alonso, A. Jelley, V . Micheli, A. Kanervisto, A. J. Storkey, T. Pearce, and F. Fleuret. Diffusion for world modeling: Visual details matter in Atari.Advances in Neural Information Processing Systems, 37:58757–58791, 2024
2024
-
[2]
A. Bar, G. Zhou, D. Tran, T. Darrell, and Y. LeCun. Navigation world models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15791–15801, 2025
2025
-
[3]
Brooks, B
T. Brooks, B. Peebles, C. Holmes, W. DePue, Y. Guo, L. Jing, D. Schnurr, J. Taylor, T. Luhman, E. Luhman, C. Ng, R. Wang, and A. Ramesh. Video generation models as world simulators. https:// openai.com/research/video-generation-models-as-world-simulators , 2024. OpenAI technical report
2024
-
[4]
Bruce, M
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y. Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. InInternational Conference on Machine Learning, pages 4603–4623. PMLR, 2024
2024
-
[5]
S. Cai, C. Yang, L. Zhang, Y. Guo, J. Xiao, Z. Yang, Y. Xu, Z. Yang, A. Yuille, L. Guibas, M. Agrawala, L. Jiang, and G. Wetzstein. Mixture of contexts for long video generation. InInternational Conference on Learning Representations, 2026
2026
-
[6]
H. Che, X. He, Q. Liu, C. Jin, and H. Chen. GameGen-X: Interactive open-world game video generation. InInternational Conference on Learning Representations, 2025
2025
-
[7]
K. Chen, D. Liang, X. Zhou, Y. Ding, X. Liu, P . Wan, and X. Bai. Out of sight but not out of mind: Hybrid memory for dynamic video world models.arXiv preprint arXiv:2603.25716, 2026
arXiv 2026
-
[8]
T. Chen, Z. Ding, A. Li, C. Zhang, Z. Xiao, Y. Wang, and C. Jin. Recurrent autoregressive diffusion: Global memory meets local attention.arXiv preprint arXiv:2511.12940, 2025
Pith/arXiv arXiv 2025
-
[9]
Cheng, Y
J. Cheng, Y. Ge, Y. Ge, J. Liao, and Y. Shan. AnimeGamer: Infinite anime life simulation with next game state prediction. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10875–10885, 2025
2025
-
[10]
Chiappa, S
S. Chiappa, S. Racanière, D. Wierstra, and S. Mohamed. Recurrent environment simulators. In International Conference on Learning Representations, 2017
2017
-
[11]
Quevedo, Q
Decart, J. Quevedo, Q. McIntyre, S. Campbell, X. Chen, and R. Wachen. Oasis: A universe in a transformer.https://oasis-model.github.io/, 2024. Decart and Etched technical blog post
2024
-
[12]
C. Deng, D. Zhu, K. Li, C. Gou, F. Li, Z. Wang, S. Zhong, W. Yu, X. Nie, Z. Song, G. Shi, and H. Fan. Emerging properties in unified multimodal pretraining.arXiv preprint arXiv:2505.14683, 2025
Pith/arXiv arXiv 2025
-
[13]
DreamX Team, Y. Bai, R. Chen, X. Chu, R. Dang, H. Dou, B. Gao, Q. Gu, S. Hong, J. Lei, G. Li, J. Li, R. Lin, Q. Shi, B. Song, L. Sun, J. Tang, R. Tian, J. Wang, J. Wu, P . Zhang, S. Zhang, and J. Zhu. DreamX-World 1.0: A general-purpose interactive world model.arXiv preprint arXiv:2606.16993, 2026
arXiv 2026
-
[14]
Esser, S
P . Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, D. Podell, T. Dockhorn, Z. English, K. Lacey, A. Goodwin, Y. Marek, and R. Rombach. Scaling rectified 8 flow transformers for high-resolution image synthesis. InInternational Conference on Machine Learning, pages 12606–12633. PMLR, 2024
2024
-
[15]
R. Feng, H. Zhang, Z. Yang, J. Xiao, Z. Shu, Z. Liu, A. Zheng, Y. Huang, Y. Liu, and H. Zhang. The matrix: Infinite-horizon world generation with real-time moving control.Advances in Neural Information Processing Systems, 38, 2025
2025
-
[16]
Garcin, T
S. Garcin, T. Walker, S. McDonagh, T. Pearce, H. Bilen, T. He, K. Wang, and J. Bian. Beyond pixel context windows: Neural world simulators with persistent 3D state. InInternational Conference on Machine Learning. PMLR, 2026
2026
-
[17]
Q. Garrido, T. Nagarajan, B. Terver, N. Ballas, Y. LeCun, and M. Rabbat. Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026
arXiv 2026
-
[18]
Z. Ge, H. Huang, M. Zhou, J. Li, G. Wang, S. Tang, and Y. Zhuang. WorldGPT: Empowering LLM as multimodal world model. InProceedings of the 32nd ACM International Conference on Multimedia, pages 7346–7355, 2024
2024
-
[19]
J. Guo, Y. Ye, T. He, H. Wu, Y. Jiang, T. Pearce, and J. Bian. MineWorld: a real-time and open-source interactive world model on Minecraft.arXiv preprint arXiv:2504.08388, 2025
Pith/arXiv arXiv 2025
-
[20]
R. Guo, D. Liang, Y. Liu, F. Liu, T. Huang, G. P . Hancke, and R. W. H. Lau. EgoCS-400K: An egocentric gameplay dataset for world models.arXiv preprint arXiv:2606.18180, 2026
Pith/arXiv arXiv 2026
-
[21]
Y. Guo, C. Yang, H. He, Y. Zhao, M. Wei, Z. Yang, W. Huang, and D. Lin. End-to-end training for autoregressive video diffusion via self-resampling. InEuropean Conference on Computer Vision. Springer, 2026
2026
-
[22]
D. Ha and J. Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018
Pith/arXiv arXiv 2018
-
[23]
Y. HaCohen, B. Brazowski, N. Chiprut, Y. Bitterman, A. Kvochko, A. Berkowitz, D. Shalem, D. Lifschitz, D. Moshe, E. Porat, et al. LTX-2: Efficient joint audio-visual foundation model.arXiv preprint arXiv:2601.03233, 2026
Pith/arXiv arXiv 2026
-
[24]
Hafner, T
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. InInternational Conference on Machine Learning, pages 2555–2565. PMLR, 2019
2019
-
[25]
Hafner, J
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse control tasks through world models. Nature, 640(8059):647–653, 2025
2025
-
[26]
D. Hafner, W. Yan, and T. Lillicrap. Training agents inside of scalable world models.arXiv preprint arXiv:2509.24527, 2025
Pith/arXiv arXiv 2025
-
[27]
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu. Reasoning with language model is planning with world model. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 8154–8173, 2023
2023
-
[28]
H. He, Y. Xu, Y. Guo, G. Wetzstein, B. Dai, H. Li, and C. Yang. CameraCtrl: Enabling camera control for video diffusion models. InInternational Conference on Learning Representations, 2025
2025
-
[29]
X. He, C. Peng, Z. Liu, B. Wang, Y. Zhang, Q. Cui, F. Kang, B. Jiang, M. An, Y. Ren, et al. Matrix-Game 2.0: An open-source real-time and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025
Pith/arXiv arXiv 2025
-
[30]
Y. He, C. D. Weilbach, M. E. Wojciechowska, Y. Sun, Y. Zhang, and F. Wood. PLAICraft: Large-scale time-aligned vision-speech-action dataset for embodied AI.arXiv preprint arXiv:2505.12707, 2025
arXiv 2025
-
[31]
Y. Hong, Y. Mei, C. Ge, Y. Xu, Y. Zhou, S. Bi, Y. Hold-Geoffroy, M. Roberts, M. Fisher, E. Shechtman, K. Sunkavalli, F. Liu, Z. Li, and H. Tan. RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025
arXiv 2025
-
[32]
InSpatio Team. InSpatio-WorldFM: An open-source real-time generative frame model.arXiv preprint arXiv:2603.11911, 2026. 9
Pith/arXiv arXiv 2026
-
[33]
Jiang, Y
Y. Jiang, Y. Gu, I. W. Tsang, and M. Z. Shou. Olaf-World: Orienting latent actions for video world modeling. InInternational Conference on Machine Learning. PMLR, 2026
2026
-
[34]
Kanervisto, D
A. Kanervisto, D. Bignell, L. Y. Wen, M. Grayson, R. Georgescu, S. V . Macua, S. Z. Tan, T. Rashid, T. Pearce, Y. Cao, A. Lemkhenter, C. Jiang, G. Costello, G. Gupta, M. Tot, S. Ishida, T. Gupta, U. Arora, R. W. White, S. Devlin, C. Morrison, and K. Hofmann. World and human action models towards gameplay ideation.Nature, 638(8051):656–663, 2025
2025
-
[35]
B. Kang, Y. Yue, R. Lu, Z. Lin, Y. Zhao, K. Wang, G. Huang, and J. Feng. How far is video generation from world model: A physical law perspective. InInternational Conference on Machine Learning. PMLR, 2025
2025
-
[36]
S. W. Kim, Y. Zhou, J. Philion, A. Torralba, and S. Fidler. Learning to simulate dynamic environments with GameGAN. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1228–1237, 2020
2020
-
[37]
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, K. Wu, Q. Lin, J. Yuan, Y. Long, A. Wang, A. Wang, C. Li, D. Huang, F. Yang, H. Tan, H. Wang, J. Song, J. Bai, J. Wu, J. Xue, J. Wang, K. Wang, M. Liu, P . Li, S. Li, W. Wang, W. Yu, X. Deng, Y. Li, Y. Chen, Y. Cui, Y. Peng, Z. Yu, Z. He, Z. Xu, Z. Zhou, Z. Xu, Y. Tao, ...
Pith/arXiv arXiv 2024
-
[38]
G. Li, B. Li, J. Chen, X. Hu, L. Zhao, and P .-T. Jiang. MagicWorld: Towards long-horizon stability for interactive video world exploration.arXiv preprint arXiv:2511.18886, 2025
arXiv 2025
-
[39]
J. Li, J. Tang, Z. Xu, L. Wu, Y. Zhou, S. Shao, T. Yu, Z. Cao, and Q. Lu. Hunyuan-GameCraft: High-dynamic interactive game video generation with hybrid history condition.arXiv preprint arXiv:2506.17201, 2025
Pith/arXiv arXiv 2025
-
[40]
R. Li, P . Torr, A. Vedaldi, and T. Jakab. VMem: Consistent interactive video scene generation with surfel-indexed view memory. InProceedings of the IEEE/CVF International Conference on Computer Vision, 2025
2025
-
[41]
R. Li, B. Yi, J. Liu, H. Gao, Y. Ma, and A. Kanazawa. Cameras as relative positional encoding.Advances in Neural Information Processing Systems, 38, 2025
2025
-
[42]
Z. Li, C. Li, X. Mao, S. Lin, M. Li, S. Zhao, Z. Xu, X. Li, Y. Feng, J. Sun, Z. Li, F. Zhang, J. Ai, Z. Wang, Y. Wu, T. He, J. Pang, Y. Qiao, Y. Jia, and K. Zhang. Sekai: A video dataset towards world exploration. Advances in Neural Information Processing Systems, 38, 2025
2025
-
[43]
Z. Li, Z. Meng, S. Shi, W. Peng, Y. Wu, B. Zheng, C. Li, and K. Zhang. WildWorld: A large-scale dataset for dynamic world modeling with actions and explicit state toward generative ARPG.arXiv preprint arXiv:2603.23497, 2026
arXiv 2026
-
[44]
Liang, L
W. Liang, L. Yu, L. Luo, S. Iyer, N. Dong, C. Zhou, G. Ghosh, M. Lewis, W.-t. Yih, L. Zettlemoyer, and X. V . Lin. Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.T ransactions on Machine Learning Research, 2025
2025
-
[45]
H. J. Lillemark, B. Huang, F. Zhan, Y. Du, and T. A. Keller. Flow equivariant world models: Structured memory for dynamic environments. InInternational Conference on Machine Learning. PMLR, 2026
2026
-
[46]
J. Liu, X. Liu, K. Mei, Y. Wen, M.-H. Yang, and W. Liu. Streaming autoregressive video generation via diagonal distillation. InInternational Conference on Learning Representations, 2026
2026
-
[47]
J. Liu, C. Ni, M. Liu, C. Peng, F. Wang, S. Shen, M. Pollefeys, M. Tomizuka, A. Tewari, and P . O. Kristensson. Towards interactive video world modeling: Frontiers, challenges, benchmarks, and future trends.arXiv preprint arXiv:2606.01164, 2026
Pith/arXiv arXiv 2026
-
[48]
K. Liu, W. Hu, J. Xu, Y. Shan, and S. Lu. Rolling forcing: Autoregressive long video diffusion in real time. InInternational Conference on Learning Representations, 2026. 10
2026
-
[49]
Y. Ma, X. Liu, X. Chen, W. Liu, C. Wu, Z. Wu, Z. Pan, Z. Xie, H. Zhang, X. Yu, L. Zhao, Y. Wang, J. Liu, and C. Ruan. JanusFlow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7739–7751, 2025
2025
-
[50]
Z. Ma, M. Liufu, and G. Gkioxari. Out of sight, out of mind? evaluating state evolution in video world models.arXiv preprint arXiv:2603.13215, 2026
arXiv 2026
-
[51]
X. Mao, Z. Li, C. Li, X. Xu, K. Ying, and K. Zhang. Yume1.5: A text-controlled interactive world generation model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7752–7761, 2026
2026
-
[52]
X. Mao, S. Lin, Z. Li, C. Li, W. Peng, T. He, J. Pang, M. Chi, Y. Qiao, and K. Zhang. Yume: An interactive world generation model.arXiv preprint arXiv:2507.17744, 2025
Pith/arXiv arXiv 2025
-
[53]
Menapace, S
W. Menapace, S. Lathuilière, A. Siarohin, C. Theobalt, S. Tulyakov, V . Golyanik, and E. Ricci. Playable environments: Video manipulation in space and time. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3574–3583, 2022
2022
-
[54]
Menapace, S
W. Menapace, S. Lathuilière, S. Tulyakov, A. Siarohin, and E. Ricci. Playable video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10061–10070, 2021
2021
-
[55]
Menapace, A
W. Menapace, A. Siarohin, S. Lathuilière, P . Achlioptas, V . Golyanik, S. Tulyakov, and E. Ricci. Prompt- able game models: Text-guided game simulation via masked diffusion models.ACM T ransactions on Graphics, 43(2):1–16, 2024
2024
-
[56]
J. Oh, X. Guo, H. Lee, R. L. Lewis, and S. Singh. Action-conditional video prediction using deep networks in Atari games.Advances in Neural Information Processing Systems, 28:2863–2871, 2015
2015
-
[57]
Y. Oshima, Y. Iwasawa, M. Suzuki, Y. Matsuo, and H. Furuta. WorldPack: Compressed memory improves spatial consistency in video world modeling.arXiv preprint arXiv:2512.02473, 2025
Pith/arXiv arXiv 2025
-
[58]
Parker-Holder and S
J. Parker-Holder and S. Fruchter. Genie 3: A new frontier for world models. https://deepmind. google/discover/blog/genie-3-a-new-frontier-for-world-models/ , 2025. Google DeepMind blog, August 5, 2025
2025
-
[59]
Peebles and S
W. Peebles and S. Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 4195–4205, 2023
2023
-
[60]
R. Po, Y. Nitzan, R. Zhang, B. Chen, T. Dao, E. Shechtman, G. Wetzstein, and X. Huang. Long-context state-space video world models. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 8733–8744, 2025
2025
-
[61]
Y. Qiu, Z. Zhao, W. W. Li, Y. Ziser, A. Korhonen, S. B. Cohen, and E. M. Ponti. Self-improving world modelling with latent actions.arXiv preprint arXiv:2602.06130, 2026
arXiv 2026
-
[62]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P . Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022
2022
-
[63]
Savov, N
N. Savov, N. Kazemi, D. Zhang, D. P . Paudel, X. Wang, and L. Van Gool. StateSpaceDiffuser: Bringing long context to diffusion world models.Advances in Neural Information Processing Systems, 38, 2025
2025
-
[64]
W. Sun, F. Wei, J. Zhao, X. Chen, Z. Chen, H. Zhang, J. Zhang, and Y. Lu. From virtual games to real-world play.arXiv preprint arXiv:2506.18901, 2025
Pith/arXiv arXiv 2025
-
[65]
J. Tang, J. Liu, J. Li, L. Wu, H. Yang, P . Zhao, S. Gong, X. Yuan, S. Shao, L. Zhang, and Q. Lu. Hunyuan- GameCraft-2: Instruction-following interactive game world model.arXiv preprint arXiv:2511.23429, 2025. 11
arXiv 2025
-
[66]
Z. Tong, Y. Jin, H. Lai, Z. Wang, Z. Xing, K. Cheng, H. Xu, Z. Pu, S. Zhu, R. Feng, J. Zhao, Y. Zhang, H. Tang, and L. Shao. SCOPE: Simulating cross-game operations in playable environments for FPS world models.arXiv preprint arXiv:2605.23345, 2026
Pith/arXiv arXiv 2026
-
[67]
Valevski, Y
D. Valevski, Y. Leviathan, M. Arar, and S. Fruchter. Diffusion models are real-time game engines. In International Conference on Learning Representations, 2025
2025
-
[68]
T. Wan, A. Wang, B. Ai, B. Wen, C. Mao, C.-W. Xie, D. Chen, F. Yu, H. Zhao, J. Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[69]
L. Wang, Z. Chen, Y. Du, D. Yan, W. Ge, G. Shen, X. Xu, L. Wu, M. Chen, T. Xu, P . Ren, X. Tao, P . Wan, and Y.-C. Chen. A mechanistic view on video generation as world models: State and dynamics.arXiv preprint arXiv:2601.17067, 2026
arXiv 2026
-
[70]
R. Wang, G. Todd, Z. Xiao, X. Yuan, M.-A. Côté, P . Clark, and P . Jansen. Can language models serve as text-based world simulators? InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1–17, 2024
2024
-
[71]
X. Wang, X. Zhang, Z. Luo, Q. Sun, Y. Cui, J. Wang, F. Zhang, Y. Wang, Z. Li, Q. Yu, Y. Zhao, Y. Ao, X. Min, T. Li, B. Wu, B. Zhao, B. Zhang, L. Wang, G. Liu, Z. He, X. Yang, J. Liu, Y. Lin, T. Huang, and Z. Wang. Emu3: Next-token prediction is all you need.arXiv preprint arXiv:2409.18869, 2024
Pith/arXiv arXiv 2024
-
[72]
Z. Wang, D. Chen, Z. Xing, Z. Tong, Y. Zhang, X. Yang, and Y. Jin. ReactiveGWM: Steering NPC in reactive game world models.arXiv preprint arXiv:2605.15256, 2026
Pith/arXiv arXiv 2026
-
[73]
Z. Wang, Z. Liu, J. Li, K. Huang, B. Xu, F. Kang, M. An, P . Wang, B. Jiang, Y. Wei, Y. Xietian, J. Pei, L. Hu, B. Jiang, H. Xue, Z. Wang, H. Sun, W. Li, W. Ouyang, X. He, Y. Liu, Y. Li, and Y. Zhou. Matrix-Game 3.0: Real-time and streaming interactive world model with long-horizon memory.arXiv preprint arXiv:2604.08995, 2026
Pith/arXiv arXiv 2026
-
[74]
Z. Wang, Z. Yuan, X. Wang, Y. Li, T. Chen, M. Xia, P . Luo, and Y. Shan. MotionCtrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH 2024 Conference Papers, pages 1–11, 2024
2024
-
[75]
T. Wiedemer, Y. Li, P . Vicol, S. S. Gu, N. Matarese, K. Swersky, B. Kim, P . Jaini, and R. Geirhos. Video models are zero-shot learners and reasoners.arXiv preprint arXiv:2509.20328, 2025
Pith/arXiv arXiv 2025
-
[76]
J. Wu, S. Yin, N. Feng, X. He, D. Li, J. Hao, and M. Long. iVideoGPT: Interactive VideoGPTs are scalable world models.Advances in Neural Information Processing Systems, 37:68082–68119, 2024
2024
-
[77]
R. Wu, X. He, M. Cheng, T. Yang, Y. Zhang, Z. Kang, X. Cai, X. Wei, C. Guo, C. Li, et al. Infinite-World: Scaling interactive world models to 1000-frame horizons via pose-free hierarchical memory. In International Conference on Machine Learning. PMLR, 2026
2026
- [78]
-
[79]
J. Xiang, Y. Gu, Z. Liu, Z. Feng, Q. Gao, Y. Hu, B. Huang, G. Liu, Y. Yang, K. Zhou, D. Abrahamyan, A. Ahmad, G. Bannur, J. Chen, K. Chen, M. Deng, R. Han, X. Huang, H. Kang, Z. Liu, E. Ma, H. Ren, Y. Shinde, R. Shingre, R. Tanikella, K. Tao, D. Yang, X. Yu, C. Zeng, B. Zhou, Z. Liu, Z. Hu, and E. P . Xing. PAN: A world model for general, interactable, an...
arXiv 2025
-
[80]
J. Xiang, G. Liu, Y. Gu, Q. Gao, Y. Ning, Y. Zha, Z. Feng, T. Tao, S. Hao, Y. Shi, Z. Liu, E. P . Xing, and Z. Hu. Pandora: Towards general world model with natural language actions and video states.arXiv preprint arXiv:2406.09455, 2024
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.