Pith. sign in

REVIEW 3 major objections 5 minor 65 references

Any demonstrated clip can become a reusable, frame-level action that transfers to new scenes once the same motion is seen under two different appearances.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 09:39 UTC pith:5F2JJ4C5

load-bearing objection Solid systems+representation paper: shadow pairs make dynamics identifiable by construction, with real transfer gains—but Table 1 bundles pairing with a source-asset pathway the baseline never gets. the 3 major comments →

arxiv 2607.28362 v1 pith:5F2JJ4C5 submitted 2026-07-30 cs.CV cs.AIcs.LG

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow

classification cs.CV cs.AIcs.LG
keywords interactive video world modelslatent actionsshadow pairscross-shadow predictiondynamics representationaction transferblock-causal video generationdemonstration-driven control
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Interactive video world models can explore virtual worlds but still lack a single way to specify how any kind of action should unfold frame by frame. Text and key presses are too loose; 3D tracks and pose streams are precise but family-specific and hard to get. Demonstration videos look like the answer, yet a video always shows motion through one particular look, so latents trained by reconstructing that clip entangle action with appearance and fail to transfer. ShadowDancer’s claim is that the fix is in the data, not another penalty: build shadow pairs that replay identical dynamics under independently resampled appearance, then train by predicting one video from the other. Whatever the pair resamples is discarded by construction; whatever it preserves becomes a unified dynamics latent. That latent turns any clip into a variable-length action asset that can be stored, composed, and replayed in new environments without labels, estimators, or fine-tuning, and the paper reports clear gains on transfer and long rollouts across human motion, games, robots, and camera control.

Core claim

Cross-shadow prediction on shadow pairs yields a unified dynamics representation: by construction the latent keeps only what two independently appeared renders of the same motion share, so any demonstrated clip becomes a reusable, variable-length action asset that drives frame-level control in new environments without action labels, motion estimators, or fine-tuning.

What carries the argument

Shadow pairs plus cross-shadow prediction: two frame-synchronized videos of the same dynamics under independently resampled appearance, with an encoder reading one and a decoder predicting the other, so invariance is a property of the supervision rather than a regularizer; the resulting latent conditions a block-causal video world model as reusable action assets.

Load-bearing premise

For each action family you care about, you must be able to build (or closely approximate) pairs that truly share the same frame-by-frame dynamics while fully resampling how they look—something easy in games and simulators but hard for real-world video.

What would settle it

Hold out true shadow pairs for a family, extract the latent from one video, generate from the first frame of the other, and check whether transfer metrics (reconstruction or trajectory error) and blinded rollout preference collapse toward ordinary self-reconstruction latent-action baselines; if they do, the identifying claim of the pairing fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A dynamics family becomes controllable through one interface exactly when shadow pairs can be constructed for it—no new vocabulary, labels, or per-family estimators.
  • Any post-training demonstration can join the action library at the cost of one frozen encoder pass and be replayed or chained in new scenes.
  • The same reference clip can be read as pure camera, pure scene/body/arm motion, or both, by how pairs were built and which readout head is used.
  • Interactive world models can be driven by showing rather than telling, across entertainment and simulation-based training of embodied agents.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If video editors or generators can manufacture faithful appearance-changed shadows of real footage, the same identifying signal could extend beyond engines into ordinary video.
  • Asset libraries may scale more with demonstration and generative coverage than with new end-to-end training runs, changing how interactive worlds are extended after deployment.
  • The same pairing protocol is a general recipe for making any chosen factor the controllable one, not only motion—suggesting analogous ‘shadow’ constructions for other entangled generative factors.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. ShadowDancer targets any-action, frame-level control of interactive video world models by treating demonstration videos as dense action specifications. The core claim is representational: standard latent-action training on self-reconstruction entangles dynamics with appearance, so the paper introduces (i) shadow pairs—frame-synchronized renders of identical dynamics under independently resampled appearance, built via a multi-source Shadow Library—and (ii) cross-shadow prediction, in which a latent-action model encodes one shadow and predicts the other, so that only shared dynamics survive in z. This z, together with frozen 3D-VAE source assets s, forms reusable action assets a=(z,s) that condition a block-causal video diffusion world model. A formal identification result (Theorem C.2) states conditions under which the minimal cross-shadow-sufficient statistic is the shared dynamics. Experiments report large transfer gains vs a matched Olaf-World baseline across five dynamics families (Table 1), component ablations (Table 3), latent probes (Supp. D), and ~86% average blinded 2AFC win rates on long rollouts against Olaf-World, Yume-1.5, and LingBot-World 2.0.

Significance. If the mechanism holds, the work offers a practical and conceptually clean interface for interactive world models: control by demonstration without per-family estimators, action labels, or fine-tuning, with an operational definition of controllability (a family is controllable when shadow pairs can be built). Strengths that should be credited include an explicit identification theorem with stated observability/separation assumptions (Supp. C), a source-agnostic pairing protocol spanning human motion, games, robots, and camera trajectories, factor-selective cam/dyn/full readout, and ablations plus representation probes that go beyond pure generation metrics. The limitation that real video enters only as self-pairs is stated honestly. The contribution is significant for video world models and latent-action learning if the reported transfer margins can be cleanly attributed to cross-shadow supervision rather than bundled conditioning pathways.

major comments (3)
  1. [Sec. 4.2, Table 1, Table 3, Sec. 3.3] Sec. 4.2 and Table 1 attribute large multi-family transfer gains primarily to shadow-trained latents, but the comparison confounds cross-shadow pairing with a source-asset stream the baseline lacks. The text states Olaf-World retains its original z-only design with no source-asset pathway, while ShadowDancer conditions on a=(z,s) via channel concatenation and cross-attention (Sec. 3.3, Supp. B.2). Table 3 shows assets alone already reach 15.07 PSNR vs Olaf-recipe z-only at 12.44; the contrast that isolates pairing—assets+unpaired z (a′) vs full (d)—is only +1.43 PSNR on a 12-pair subset of two families, not the five-family Table 1 protocol. For the central claim that cross-shadow prediction yields the transferable representation, Table 1 should include a matched baseline that receives the same s pathway (and, ideally, multi-head readout), and the pairing-isolation ablation should be repo
  2. [Abstract, Sec. 1, Sec. 3.4, Supp. A, Supp. C.4] The abstract and introduction market “any-action” control, while the operational scope (a family is controllable exactly when constructible shadow pairs exist) and Supp. A/C.4 make clear that identifying supervision is synthetic: real video enters only as degenerate self-pairs that do not supply the Theorem C.2 guarantee. This scoping is scientifically honest in the body but under-signaled in the title/abstract claims and in the deployment narrative (“any demonstrated clip”). Please state the synthetic-pair precondition up front when claiming any-action generality, and either quantify degradation when assets come from unpaired real/modded footage without true shadows (beyond the qualitative Fig. 5) or temper the claim to families with re-renderable dynamics.
  3. [Sec. 4.3, Table 2, Supp. D.3] Long-rollout evaluation (Sec. 4.3, Table 2) is system-level 2AFC judged by a VLM (Fable 5) over three action-centric axes, with only a 20% human audit mentioned in Supp. D.3. Because interfaces differ by design (latent assets vs text/keystrokes vs camera-pose+text), the comparison measures end-to-end command survival rather than matched information. That is a valid systems question, but the reported ~86% average win rate is load-bearing for the interactive-control claim: please report inter-annotator agreement (VLM vs human audit) per axis, raw win counts, and a sensitivity check with human-only judgments on the full set or a larger audited subset. Otherwise it is hard to know how much of Table 2 is judge noise or interface mismatch versus genuine control gains.
minor comments (5)
  1. [Fig. 2, Sec. 3.3] Fig. 2 and Sec. 3.3: the dual role of s (high-frequency motion detail vs appearance) is easy to misread as appearance leakage. A short diagram or paragraph clarifying that shadow-pair training removes the reward for copying source appearance would help.
  2. [Sec. 3] Notation: d vs D, and z vs a=(z,s), shift between “unified dynamics representation” and “action asset.” Keep one term for z alone throughout Sec. 3 and the experiments.
  3. [Sec. 3.4, Supp. D.1–D.2] Table 4 mixture weights and self-pair probability (0.5 for human body, ~1/3 overall) are important free choices; a one-sentence sensitivity pointer in the main text (even if details stay in Supp. D.7) would aid reproducibility.
  4. [Abstract, Sec. 4] Typos/formatting: “Shadow Dancer-1.github.io” spacing in the abstract; “3D-V AE” broken across lines; Fréchet rendered as “Fr´echet” in places; “Fable 5” should be identified more clearly as the judge model.
  5. [Sec. 2] Related work could briefly contrast with multi-view/invariance SSL citations already in Supp. C ([22],[54]) in the main Sec. 2, since the pairing-as-task-definition idea is central.

Circularity Check

0 steps flagged

No significant circularity: operational pairing defines the control factor by design, and the identification theorem plus empirical claims are not self-forced reductions.

full rationale

ShadowDancer’s core move is an explicit operational definition, not a disguised derivation: dynamics d is whatever a shadow-pair protocol preserves and appearance c is whatever it resamples (Eqs. 1–2, Sec. 3.1). Cross-shadow prediction then trains z to predict one render from the other (Eq. 3), so invariance is a property of the supervision rather than a penalty. Theorem C.2 / Corollary C.3 state a standard minimal-sufficiency identification result under stated assumptions A1–A3 (independent resampling, source observability, effect separation); they conclude that the coarsest cross-shadow-sufficient statistic is shared dynamics up to invertible reparameterization—i.e., they characterize what the pairing already makes identifiable, with explicit failure modes when separation or observability fails (Supp. C.4). That is not circular in the sense of claiming an independent prediction that reduces to a fitted input or to an unverified self-citation chain. Empirical transfer and rollout claims are tested against external baselines (Olaf-World, Yume-1.5, LingBot-World 2.0) on held-out pairs and blinded 2AFC, not by refitting the evaluation target. Related multi-view/SSL pairing ideas are cited as prior art (von Kügelgen et al., Gresele et al.), not as author-owned uniqueness theorems that forbid alternatives. Confounding between pairing and the source-asset stream (Table 1 vs Table 3) is a causal-isolation / correctness concern, not circularity of the derivation chain. Score 0; no circular steps.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 4 invented entities

The central empirical claim rests on a generative factorization x=R(d,c), the constructibility of independently appearance-resampled synchronized pairs, standard VAE/flow-matching training choices, and a pretrained video DiT backbone. Free parameters are ordinary ML hyperparameters rather than physics-style fitted constants. Invented entities are methodological constructs (shadow pairs, action assets, Shadow Library), not new physical objects; their 'evidence' is internal experimental performance plus a conditional identification theorem.

free parameters (4)
  • LAM latent dimension dz = 32
    Set to 32 following prior LAMs; bottleneck size directly affects what can be smuggled vs discarded and is validated only by a small budget-matched sweep.
  • β-VAE weight = 0.01
    KL pressure that favors minimal codes; chosen as 0.01, not derived.
  • self-pair probability / mixture weights = ~0.5 human self-pair; ~1/3 overall self-pairs
    Controls how much unpaired real video dilutes identifying shadow supervision (e.g., 0.5 for human body source; ~1/3 self-pairs overall); mixture percentages in Table 4 are design choices.
  • world-model conditioning gates and architecture knobs = block=3; CFG=5.0; 30 steps
    αz init, zero-init cross-attention, block size (3 latent frames), CFG 5.0, 30 denoising steps, resolutions—standard trained/chosen knobs the rollout quality depends on.
axioms (5)
  • domain assumption A video factors as x=R(d,c) with a pairing-defined split between dynamics d and appearance c.
    Sec. 3.1 Eq. (1)–(2); the entire invariance claim is relative to this operational split.
  • domain assumption Shadow pairs satisfy independent context resampling X⊥⊥(B,Y)|D, source observability D=h(X), and target kernel separation (Assumptions A1–A3).
    Supp. C; required for Theorem C.2 that the minimal cross-shadow sufficient statistic is shared dynamics.
  • standard math Standard β-VAE / flow-matching / DiT universal-approximation and training practice suffice to approach the ideal minimal representation in practice.
    Remark C.4 explicitly notes the variational channel only approximates the deterministic identification result; neural realizability uses classical UAT citations.
  • ad hoc to paper A dynamics family is controllable exactly when constructible shadow pairs exist for it ('any-action' scoped claim).
    Abstract and Sec. 3.1; converts an engineering precondition into the definition of the method's scope.
  • domain assumption VLM blinded 2AFC (Fable 5) is an adequate proxy for action control/fidelity/long-horizon quality when no ground-truth rollout exists.
    Sec. 4.3; load-bearing for the 86% win-rate headline on long rollouts.
invented entities (4)
  • Shadow pair / Shadow of a video no independent evidence
    purpose: Provide two synchronized renders of identical dynamics under independently resampled appearance so invariance is in the data.
    Core supervision construct; not a physical entity. Independent evidence is only via improved transfer probes and generations inside this paper's stack.
  • Shadow Library no independent evidence
    purpose: Source-agnostic pairing protocol and dataset spanning animation, games, robots, cameras, plus real self-pairs.
    Engineering corpus enabling the method; composition in Supp. Table 4. Not independently released as a benchmark in the manuscript.
  • Unified dynamics representation z and reusable action asset a=(z,s) no independent evidence
    purpose: Single latent interface for all dynamics families; stored variable-length control units for deployment.
    Methodological interface objects distilled by the LAM and 3D-VAE; validated by transfer/rollout experiments, not external measurement.
  • Factor-selective cam/dyn/full readout heads no independent evidence
    purpose: Separate ego-camera vs scene/body/arm dynamics within one latent space via pairing masks.
    Supp. B.1 architectural device tied to pairing protocol; evidence is internal controllability demos.

pith-pipeline@v1.2.0-daily-grok45 · 27194 in / 4255 out tokens · 78383 ms · 2026-07-31T09:39:42.344298+00:00 · methodology

0 comments
read the original abstract

We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io

Figures

Figures reproduced from arXiv: 2607.28362 by Jin Cao, Kaipeng Zhang, Zian Meng.

Figure 1
Figure 1. Figure 1: ShadowDancer learns any action from a video and its shadow. Top: shadow pairs (x, x˜) replay one dynamics under independently resampled appearance, across heterogeneous sources such as human motion, robot manipulation, and open-world gameplay, and distill it into a unified dynamics representation z1:T −1. Bottom: the same latent interface drives diverse commanded actions, spanning first- and third-person c… view at source ↗
Figure 2
Figure 2. Figure 2: Overview of ShadowDancer, taking human motion as the running example; the pipeline is identical for camera, object, and every other dynamics family. (1) A shadow pair renders one dynamics d twice, x=R(d, c) and x˜=R(d, c˜), with the remaining factors independently resampled. (2) Cross-shadow prediction trains the LAM: the encoder reads each zt from the source transition, and the decoder predicts the next s… view at source ↗
Figure 3
Figure 3. Figure 3: Qualitative action transfer against Olaf-World. Top: reference videos providing the action across four families; below: frame-synchronized generations in the new environment. Olaf-World warps subjects and leaks spurious motion; ShadowDancer re-enacts the reference dynamics faithfully. Dancer leads on every family by a wide margin. The gap is visible in [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Qualitative long action rollouts. Each comparison starts from a shared first frame and follows the same command stream; rows are methods, columns are time steps along the rollout. Where the baselines drift or degrade, ShadowDancer stays aligned. The comparison runs along three action-centric axes: ac￾tion control (is the commanded action performed at all), ac￾tion fidelity (is the specific weapon/object/mo… view at source ↗
Figure 5
Figure 5. Figure 5: Transferring unseen actions: assets recorded from a modded, unseen character (A) drive generation in a new environment (B). 4.5. Transferring Unseen Actions A latent interface is only as general as its behavior on dynam￾ics it has never seen, so we test transfer on actions introduced through game mods after training: an unseen player charac￾ter and a two-handed sword whose attack is a swing rather than a s… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

65 extracted references · 23 linked inside Pith

  1. [1]

    Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575,

    Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575,

  2. [2]

    Alemi, Ian Fischer, Joshua V

    Alexander A. Alemi, Ian Fischer, Joshua V . Dillon, and Kevin Murphy. Deep variational information bottleneck. InICLR,

  3. [3]

    Diffusion for world modeling: Visual details matter in atari.NeurIPS,

    Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos J Storkey, Tim Pearce, and Franc ¸ois Fleuret. Diffusion for world modeling: Visual details matter in atari.NeurIPS,

  4. [4]

    V-JEPA 2: Self-supervised video models enable understanding, predic- tion and planning.arXiv preprint arXiv:2506.09985, 2025

    Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-JEPA 2: Self-supervised video models enable understanding, predic- tion and planning.arXiv preprint arXiv:2506.09985, 2025. 4

  5. [5]

    Video generation models as world simulators

    Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. 2

  6. [6]

    Genie: Generative interactive environments

    Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker- Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. InICML, 2024. 2, 3

  7. [7]

    UniVLA: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025. 3

  8. [8]

    Unifying 9 precisely 3D-enhanced camera and human motion controls for video generation

    Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chao- hui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu. Unifying 9 precisely 3D-enhanced camera and human motion controls for video generation. InSIGGRAPH Asia, 2025. 2, 3, 6, 5

  9. [9]

    SkyReels-V2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074,

    Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengcheng Ma, et al. SkyReels-V2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074,

  10. [10]

    A simple framework for contrastive learning of visual representations

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geof- frey Hinton. A simple framework for contrastive learning of visual representations. InICML, 2020. 3

  11. [11]

    IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.arXiv preprint arXiv:2411.00785, 2024

    Xiaoyu Chen, Junliang Guo, Tianyu He, Chuheng Zhang, Pushi Zhang, Derek Cathera Yang, Li Zhao, and Jiang Bian. IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.arXiv preprint arXiv:2411.00785, 2024. 2, 3

  12. [12]

    villa-X: enhancing latent action modeling in vision-language-action models.arXiv preprint arXiv:2507.23682, 2025

    Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, et al. villa-X: enhancing latent action modeling in vision-language-action models.arXiv preprint arXiv:2507.23682, 2025

  13. [13]

    Moto: Latent motion token as the bridging language for learning robot manipulation from videos

    Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li, Yixiao Ge, Mingyu Ding, Ying Shan, and Xihui Liu. Moto: Latent motion token as the bridging language for learning robot manipulation from videos. InICCV, 2025. 3

  14. [14]

    Oasis: A universe in a transformer

    Decart, Julian Quevedo, Quinn McIntyre, Spruce Campbell, Xinlei Chen, and Robert Wachen. Oasis: A universe in a transformer. 2024. 2, 3

  15. [15]

    Imitating latent policies from observation

    Ashley Edwards, Himanshu Sahni, Yannick Schroecker, and Charles Isbell. Imitating latent policies from observation. In ICML, 2019. 2, 3

  16. [16]

    3D-aware implicit motion control for view-adaptive human video gener- ation.arXiv preprint arXiv:2602.03796, 2026

    Zhixue Fang, Xu He, Songlin Tang, Haoxian Zhang, Qingfeng Li, Xiaoqiang Liu, Pengfei Wan, and Kun Gai. 3D-aware implicit motion control for view-adaptive human video gener- ation.arXiv preprint arXiv:2602.03796, 2026. 2, 3

  17. [17]

    Vista: A generalizable driving world model with high fidelity and versatile controllability.NeurIPS, 2024

    Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, and Hongyang Li. Vista: A generalizable driving world model with high fidelity and versatile controllability.NeurIPS, 2024. 3

  18. [18]

    AdaWorld: Learning adaptable world mod- els with latent actions

    Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, and Chuang Gan. AdaWorld: Learning adaptable world mod- els with latent actions. InICML, 2025. 2, 3, 5, 6, 1

  19. [19]

    Infinite worlds with versatile interactions.arXiv preprint arXiv:2607.07534, 2026

    Zelin Gao, Qiuyu Wang, Jiapeng Zhu, Jingye Chen, Zichen Liu, Qingyan Bai, Jiahao Wang, Yufeng Yuan, Hanlin Wang, Yichong Lu, Ka Leong Cheng, Haojie Zhang, Jian Gao, Tian- rui Feng, Yuzheng Liu, Yao Yao, Yinghao Xu, Xing Zhu, Yujun Shen, and Hao Ouyang. Infinite worlds with versatile interactions.arXiv preprint arXiv:2607.07534, 2026. 2, 3, 7, 8

  20. [20]

    Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026

    Quentin Garrido, Tushar Nagarajan, Basile Terver, Nico- las Ballas, Yann LeCun, and Michael Rabbat. Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026. 2, 3

  21. [21]

    Google DeepMind. Veo. Model page. 1

  22. [22]

    Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Sch¨olkopf

    Luigi Gresele, Paul K. Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Sch¨olkopf. The incomplete rosetta stone problem: Identifiability results for multi-view nonlinear ICA. InProceedings of the 35th Conference on Uncertainty in Artificial Intelligence, pages 217–227, 2020. 2

  23. [23]

    Long-context autoregressive video modeling with next-frame prediction

    Yuchao Gu, Weijia Mao, and Mike Zheng Shou. Long-context autoregressive video modeling with next-frame prediction. arXiv preprint arXiv:2503.19325, 2025. 3, 6

  24. [24]

    World models.arXiv preprint arXiv:1803.10122, 2018

    David Ha and J ¨urgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018. 3

  25. [25]

    Mastering diverse domains through world models

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023. 3

  26. [26]

    Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025

    Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Cyrus Wu, Wei Li, Xuchen Song, Yang Liu, Eric Li, and Yahui Zhou. Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025. 2, 3

  27. [27]

    beta-V AE: Learning basic visual con- cepts with a constrained variational framework

    Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-V AE: Learning basic visual con- cepts with a constrained variational framework. InICLR,

  28. [28]

    RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025

    Yicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu, Yang Zhou, Sai Bi, Yannick Hold-Geoffroy, Mike Roberts, Matthew Fisher, Eli Shechtman, et al. RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025. 3

  29. [29]

    Approximation capabilities of multilayer feed- forward networks.Neural Networks, 4(2):251–257, 1991

    Kurt Hornik. Approximation capabilities of multilayer feed- forward networks.Neural Networks, 4(2):251–257, 1991. 4

  30. [30]

    Self forcing: Bridging the train-test gap in autore- gressive video diffusion.arXiv preprint arXiv:2506.08009,

    Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autore- gressive video diffusion.arXiv preprint arXiv:2506.08009,

  31. [31]

    DreamGen: Unlocking gener- alization in robot learning through video world models.arXiv preprint arXiv:2505.12705, 2025

    Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang, Jo- han Bjorck, Yu Fang, Fengyuan Hu, Spencer Huang, Kaushil Kundalia, Yen-Chen Lin, et al. DreamGen: Unlocking gener- alization in robot learning through video world models.arXiv preprint arXiv:2505.12705, 2025. 3

  32. [32]

    Tsang, and Mike Zheng Shou

    Yuxin Jiang, Yuchao Gu, Ivor W. Tsang, and Mike Zheng Shou. Olaf-world: Orienting latent actions for video world modeling.arXiv preprint arXiv:2602.10104, 2026. 2, 3, 4, 5, 6, 7, 8, 1

  33. [33]

    Miradata: A large-scale video dataset with long durations and structured captions.NeurIPS, 2024

    Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions.NeurIPS, 2024. 6, 5

  34. [34]

    Variational autoencoders and nonlinear ica: A unifying framework

    Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational autoencoders and nonlinear ica: A unifying framework. InAISTATS, 2020. 2, 3

  35. [35]

    UniSkill: Imitating human videos via cross-embodiment skill representations

    Hanjung Kim, Jaehyun Kang, Hyolim Kang, Meedeum Cho, Seon Joo Kim, and Youngwoon Lee. UniSkill: Imitating human videos via cross-embodiment skill representations. arXiv preprint arXiv:2505.08787, 2025. 3

  36. [36]

    HunyuanVideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

    Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. HunyuanVideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024. 2

  37. [37]

    Generative video motion editing with 3D point tracks

    Yao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang, Joon-Young Lee, Jia-Bin Huang, Eli Shechtman, and 10 Zhengqi Li. Generative video motion editing with 3D point tracks. InCVPR, 2026. 2, 3

  38. [38]

    Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision.arXiv preprint arXiv:2312.16256, 2023

    Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision.arXiv preprint arXiv:2312.16256, 2023. 6, 5

  39. [39]

    Flow straight and fast: Learning to generate and transfer data with rectified flow

    Xingchao Liu, Chengyue Gong, and qiang liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. InICLR, 2023. 5

  40. [40]

    Challenging common assumptions in the unsu- pervised learning of disentangled representations

    Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Sch ¨olkopf, and Olivier Bachem. Challenging common assumptions in the unsu- pervised learning of disentangled representations. InICML,

  41. [41]

    Yume-1.5: A text-controlled interactive world genera- tion model.arXiv preprint arXiv:2512.22096, 2025

    Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, Kain- ing Ying, Tong He, Jiangmiao Pang, Yu Qiao, and Kaipeng Zhang. Yume-1.5: A text-controlled interactive world genera- tion model.arXiv preprint arXiv:2512.22096, 2025. 2, 3, 7, 8

  42. [42]

    Open X-Embodiment: Robotic learning datasets and RT-X models

    Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Ab- hishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Poo- ley, Agrim Gupta, et al. Open X-Embodiment: Robotic learning datasets and RT-X models. InIEEE International Conference on Robotics and Automation (ICRA), pages 6892– 6903, 2024. 6, 5

  43. [43]

    Genie 3: A new frontier for world models

    Jack Parker-Holder, Shlomi Fruchter, et al. Genie 3: A new frontier for world models. https://deepmind.googl e/models/genie/. Blog post. 2, 3

  44. [44]

    Scalable diffusion models with transformers

    William Peebles and Saining Xie. Scalable diffusion models with transformers. InICCV, 2023. 5

  45. [45]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021. 3

  46. [46]

    Derpanis, and Kostas Daniilidis

    Oleh Rybkin, Karl Pertsch, Andrew Jaegle, Konstantinos G. Derpanis, and Kostas Daniilidis. Learning what you can do before doing anything. InICLR, 2019. 2, 3

  47. [47]

    MotionStream: Real-time video generation with interactive motion controls

    Joonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu, Jaesik Park, Eli Shechtman, and Xun Huang. MotionStream: Real-time video generation with interactive motion controls. InICLR, 2026. 2, 3

  48. [48]

    A benchmark for the evalua- tion of RGB-D SLAM systems

    J¨urgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. A benchmark for the evalua- tion of RGB-D SLAM systems. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 573–580, 2012. 6

  49. [49]

    WorldPlay: towards long- term geometric consistency for real-time interactive world modeling.arXiv preprint arXiv:2512.14614, 2025

    Wenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu, Zehan Wang, Zhenwei Wang, Yunhong Wang, Jun Zhang, Tengfei Wang, and Chunchao Guo. WorldPlay: towards long- term geometric consistency for real-time interactive world modeling.arXiv preprint arXiv:2512.14614, 2025. 2, 3

  50. [50]

    Hunyuan-GameCraft-2: Instruction- following interactive game world model.arXiv preprint arXiv:2511.23429, 2025

    Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, and Qinglin Lu. Hunyuan-GameCraft-2: Instruction- following interactive game world model.arXiv preprint arXiv:2511.23429, 2025. 3

  51. [51]

    Maniskill3: Gpu parallelized robotics simulation and rendering for generalizable embodied ai.arXiv preprint arXiv:2410.00425, 2024

    Stone Tao, Fanbo Xiang, Arth Shukla, Yuzhe Qin, Xander Hinrichsen, Xiaodi Yuan, Chen Bao, Xinsong Lin, Yulin Liu, Tse-kai Chan, Yuan Gao, Xuanlin Li, Tongzhou Mu, Nan Xiao, Arnav Gurha, Viswesh Nagaswamy Rajesh, Yong Woo Choi, Yen-Ru Chen, Zhiao Huang, Roberto Calandra, Rui Chen, Shan Luo, and Hao Su. Maniskill3: Gpu parallelized robotics simulation and r...

  52. [52]

    Advancing open-source world models.arXiv preprint arXiv:2601.20540, 2026

    Robbyant Team, Zelin Gao, Qiuyu Wang, Yanhong Zeng, Jiapeng Zhu, Ka Leong Cheng, Yixuan Li, Hanlin Wang, Yinghao Xu, Shuailei Ma, et al. Advancing open-source world models.arXiv preprint arXiv:2601.20540, 2026. 2, 3

  53. [53]

    Video- MAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.NeurIPS, 2022

    Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Video- MAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.NeurIPS, 2022. 4

  54. [54]

    Self-supervised learning with data aug- mentations provably isolates content from style

    Julius von K¨ugelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Sch ¨olkopf, Michel Besserve, and Francesco Locatello. Self-supervised learning with data aug- mentations provably isolates content from style. InAdvances in Neural Information Processing Systems, pages 16451– 16467, 2021. 2

  55. [55]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025. 2, 3, 5, 1

  56. [56]

    VGGT: Visual geometry grounded transformer

    Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual geometry grounded transformer. InCVPR, 2025. 6

  57. [57]

    Co-Evolving latent action world models.arXiv preprint arXiv:2510.26433, 2025

    Yucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao, Kaixin Wang, and Jiang Bian. Co-Evolving latent action world models.arXiv preprint arXiv:2510.26433, 2025. 3

  58. [58]

    Connectionist nonparametric regression: Mul- tilayer feedforward networks can learn arbitrary mappings

    Halbert White. Connectionist nonparametric regression: Mul- tilayer feedforward networks can learn arbitrary mappings. Neural Networks, 3(5):535–549, 1990. 4

  59. [59]

    WorldMem: Long- term consistent world simulation with memory.arXiv preprint arXiv:2504.12369, 2025

    Zeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang, Shuai Yang, Yanhong Zeng, and Xingang Pan. WorldMem: Long- term consistent world simulation with memory.arXiv preprint arXiv:2504.12369, 2025. 3

  60. [60]

    CoMo: Learning continuous latent motion from in- ternet videos for scalable robot learning.arXiv preprint arXiv:2505.17006, 2025

    Jiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu, Kai- jing Ma, Yating Wang, Gangshan Wu, Tong He, and Limin Wang. CoMo: Learning continuous latent motion from in- ternet videos for scalable robot learning.arXiv preprint arXiv:2505.17006, 2025. 2, 3, 4

  61. [61]

    Latent Action Pretraining from Videos

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu- Wei Chao, Bill Yuchen Lin, et al. Latent Action Pretraining from Videos. InICLR, 2025. 2, 3

  62. [62]

    MIND: Benchmarking memory con- sistency and action control in world models.arXiv preprint arXiv:2602.08025, 2026

    Yixuan Ye, Xuanyu Lu, Yuxin Jiang, Yuchao Gu, Rui Zhao, Qiwei Liang, Jiachun Pan, Fengda Zhang, Weijia Wu, and Alex Jinpeng Wang. MIND: Benchmarking memory con- sistency and action control in world models.arXiv preprint arXiv:2602.08025, 2026. 3

  63. [63]

    GameFactory: Creating new games with gen- erative interactive videos.arXiv preprint arXiv:2501.08325,

    Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, and Xihui Liu. GameFactory: Creating new games with gen- erative interactive videos.arXiv preprint arXiv:2501.08325,

  64. [64]

    The unreasonable effectiveness of deep features as a perceptual metric

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InCVPR, 2018. 6

  65. [65]

    replay the dynamics, resample everything else

    Shangwen Zhu, Qianyu Peng, Zhao Pu, Zhilei Shu, Xiangrui Ke, Zhaohu Xing, Zizhao Tong, Zeqing Wang, Xinyu Cui, Zian Zheng, Huangji Wang, Jian Zhao, Yeying Jin, Fan Cheng, and Ruili Feng. Incantation: Natural language as the action interface for multi-entity video world models.arXiv preprint arXiv:2605.18601, 2026. 2, 3, 7 12 ShadowDancer: Teaching Video W...