REVIEW 3 major objections 5 minor 65 references
Any demonstrated clip can become a reusable, frame-level action that transfers to new scenes once the same motion is seen under two different appearances.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-31 09:39 UTC pith:5F2JJ4C5
load-bearing objection Solid systems+representation paper: shadow pairs make dynamics identifiable by construction, with real transfer gains—but Table 1 bundles pairing with a source-asset pathway the baseline never gets. the 3 major comments →
ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Cross-shadow prediction on shadow pairs yields a unified dynamics representation: by construction the latent keeps only what two independently appeared renders of the same motion share, so any demonstrated clip becomes a reusable, variable-length action asset that drives frame-level control in new environments without action labels, motion estimators, or fine-tuning.
What carries the argument
Shadow pairs plus cross-shadow prediction: two frame-synchronized videos of the same dynamics under independently resampled appearance, with an encoder reading one and a decoder predicting the other, so invariance is a property of the supervision rather than a regularizer; the resulting latent conditions a block-causal video world model as reusable action assets.
Load-bearing premise
For each action family you care about, you must be able to build (or closely approximate) pairs that truly share the same frame-by-frame dynamics while fully resampling how they look—something easy in games and simulators but hard for real-world video.
What would settle it
Hold out true shadow pairs for a family, extract the latent from one video, generate from the first frame of the other, and check whether transfer metrics (reconstruction or trajectory error) and blinded rollout preference collapse toward ordinary self-reconstruction latent-action baselines; if they do, the identifying claim of the pairing fails.
If this is right
- A dynamics family becomes controllable through one interface exactly when shadow pairs can be constructed for it—no new vocabulary, labels, or per-family estimators.
- Any post-training demonstration can join the action library at the cost of one frozen encoder pass and be replayed or chained in new scenes.
- The same reference clip can be read as pure camera, pure scene/body/arm motion, or both, by how pairs were built and which readout head is used.
- Interactive world models can be driven by showing rather than telling, across entertainment and simulation-based training of embodied agents.
Where Pith is reading between the lines
- If video editors or generators can manufacture faithful appearance-changed shadows of real footage, the same identifying signal could extend beyond engines into ordinary video.
- Asset libraries may scale more with demonstration and generative coverage than with new end-to-end training runs, changing how interactive worlds are extended after deployment.
- The same pairing protocol is a general recipe for making any chosen factor the controllable one, not only motion—suggesting analogous ‘shadow’ constructions for other entangled generative factors.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ShadowDancer targets any-action, frame-level control of interactive video world models by treating demonstration videos as dense action specifications. The core claim is representational: standard latent-action training on self-reconstruction entangles dynamics with appearance, so the paper introduces (i) shadow pairs—frame-synchronized renders of identical dynamics under independently resampled appearance, built via a multi-source Shadow Library—and (ii) cross-shadow prediction, in which a latent-action model encodes one shadow and predicts the other, so that only shared dynamics survive in z. This z, together with frozen 3D-VAE source assets s, forms reusable action assets a=(z,s) that condition a block-causal video diffusion world model. A formal identification result (Theorem C.2) states conditions under which the minimal cross-shadow-sufficient statistic is the shared dynamics. Experiments report large transfer gains vs a matched Olaf-World baseline across five dynamics families (Table 1), component ablations (Table 3), latent probes (Supp. D), and ~86% average blinded 2AFC win rates on long rollouts against Olaf-World, Yume-1.5, and LingBot-World 2.0.
Significance. If the mechanism holds, the work offers a practical and conceptually clean interface for interactive world models: control by demonstration without per-family estimators, action labels, or fine-tuning, with an operational definition of controllability (a family is controllable when shadow pairs can be built). Strengths that should be credited include an explicit identification theorem with stated observability/separation assumptions (Supp. C), a source-agnostic pairing protocol spanning human motion, games, robots, and camera trajectories, factor-selective cam/dyn/full readout, and ablations plus representation probes that go beyond pure generation metrics. The limitation that real video enters only as self-pairs is stated honestly. The contribution is significant for video world models and latent-action learning if the reported transfer margins can be cleanly attributed to cross-shadow supervision rather than bundled conditioning pathways.
major comments (3)
- [Sec. 4.2, Table 1, Table 3, Sec. 3.3] Sec. 4.2 and Table 1 attribute large multi-family transfer gains primarily to shadow-trained latents, but the comparison confounds cross-shadow pairing with a source-asset stream the baseline lacks. The text states Olaf-World retains its original z-only design with no source-asset pathway, while ShadowDancer conditions on a=(z,s) via channel concatenation and cross-attention (Sec. 3.3, Supp. B.2). Table 3 shows assets alone already reach 15.07 PSNR vs Olaf-recipe z-only at 12.44; the contrast that isolates pairing—assets+unpaired z (a′) vs full (d)—is only +1.43 PSNR on a 12-pair subset of two families, not the five-family Table 1 protocol. For the central claim that cross-shadow prediction yields the transferable representation, Table 1 should include a matched baseline that receives the same s pathway (and, ideally, multi-head readout), and the pairing-isolation ablation should be repo
- [Abstract, Sec. 1, Sec. 3.4, Supp. A, Supp. C.4] The abstract and introduction market “any-action” control, while the operational scope (a family is controllable exactly when constructible shadow pairs exist) and Supp. A/C.4 make clear that identifying supervision is synthetic: real video enters only as degenerate self-pairs that do not supply the Theorem C.2 guarantee. This scoping is scientifically honest in the body but under-signaled in the title/abstract claims and in the deployment narrative (“any demonstrated clip”). Please state the synthetic-pair precondition up front when claiming any-action generality, and either quantify degradation when assets come from unpaired real/modded footage without true shadows (beyond the qualitative Fig. 5) or temper the claim to families with re-renderable dynamics.
- [Sec. 4.3, Table 2, Supp. D.3] Long-rollout evaluation (Sec. 4.3, Table 2) is system-level 2AFC judged by a VLM (Fable 5) over three action-centric axes, with only a 20% human audit mentioned in Supp. D.3. Because interfaces differ by design (latent assets vs text/keystrokes vs camera-pose+text), the comparison measures end-to-end command survival rather than matched information. That is a valid systems question, but the reported ~86% average win rate is load-bearing for the interactive-control claim: please report inter-annotator agreement (VLM vs human audit) per axis, raw win counts, and a sensitivity check with human-only judgments on the full set or a larger audited subset. Otherwise it is hard to know how much of Table 2 is judge noise or interface mismatch versus genuine control gains.
minor comments (5)
- [Fig. 2, Sec. 3.3] Fig. 2 and Sec. 3.3: the dual role of s (high-frequency motion detail vs appearance) is easy to misread as appearance leakage. A short diagram or paragraph clarifying that shadow-pair training removes the reward for copying source appearance would help.
- [Sec. 3] Notation: d vs D, and z vs a=(z,s), shift between “unified dynamics representation” and “action asset.” Keep one term for z alone throughout Sec. 3 and the experiments.
- [Sec. 3.4, Supp. D.1–D.2] Table 4 mixture weights and self-pair probability (0.5 for human body, ~1/3 overall) are important free choices; a one-sentence sensitivity pointer in the main text (even if details stay in Supp. D.7) would aid reproducibility.
- [Abstract, Sec. 4] Typos/formatting: “Shadow Dancer-1.github.io” spacing in the abstract; “3D-V AE” broken across lines; Fréchet rendered as “Fr´echet” in places; “Fable 5” should be identified more clearly as the judge model.
- [Sec. 2] Related work could briefly contrast with multi-view/invariance SSL citations already in Supp. C ([22],[54]) in the main Sec. 2, since the pairing-as-task-definition idea is central.
Circularity Check
No significant circularity: operational pairing defines the control factor by design, and the identification theorem plus empirical claims are not self-forced reductions.
full rationale
ShadowDancer’s core move is an explicit operational definition, not a disguised derivation: dynamics d is whatever a shadow-pair protocol preserves and appearance c is whatever it resamples (Eqs. 1–2, Sec. 3.1). Cross-shadow prediction then trains z to predict one render from the other (Eq. 3), so invariance is a property of the supervision rather than a penalty. Theorem C.2 / Corollary C.3 state a standard minimal-sufficiency identification result under stated assumptions A1–A3 (independent resampling, source observability, effect separation); they conclude that the coarsest cross-shadow-sufficient statistic is shared dynamics up to invertible reparameterization—i.e., they characterize what the pairing already makes identifiable, with explicit failure modes when separation or observability fails (Supp. C.4). That is not circular in the sense of claiming an independent prediction that reduces to a fitted input or to an unverified self-citation chain. Empirical transfer and rollout claims are tested against external baselines (Olaf-World, Yume-1.5, LingBot-World 2.0) on held-out pairs and blinded 2AFC, not by refitting the evaluation target. Related multi-view/SSL pairing ideas are cited as prior art (von Kügelgen et al., Gresele et al.), not as author-owned uniqueness theorems that forbid alternatives. Confounding between pairing and the source-asset stream (Table 1 vs Table 3) is a causal-isolation / correctness concern, not circularity of the derivation chain. Score 0; no circular steps.
Axiom & Free-Parameter Ledger
free parameters (4)
- LAM latent dimension dz =
32
- β-VAE weight =
0.01
- self-pair probability / mixture weights =
~0.5 human self-pair; ~1/3 overall self-pairs
- world-model conditioning gates and architecture knobs =
block=3; CFG=5.0; 30 steps
axioms (5)
- domain assumption A video factors as x=R(d,c) with a pairing-defined split between dynamics d and appearance c.
- domain assumption Shadow pairs satisfy independent context resampling X⊥⊥(B,Y)|D, source observability D=h(X), and target kernel separation (Assumptions A1–A3).
- standard math Standard β-VAE / flow-matching / DiT universal-approximation and training practice suffice to approach the ideal minimal representation in practice.
- ad hoc to paper A dynamics family is controllable exactly when constructible shadow pairs exist for it ('any-action' scoped claim).
- domain assumption VLM blinded 2AFC (Fable 5) is an adequate proxy for action control/fidelity/long-horizon quality when no ground-truth rollout exists.
invented entities (4)
-
Shadow pair / Shadow of a video
no independent evidence
-
Shadow Library
no independent evidence
-
Unified dynamics representation z and reusable action asset a=(z,s)
no independent evidence
-
Factor-selective cam/dyn/full readout heads
no independent evidence
read the original abstract
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals that serve one family and are hard to acquire, so precise control across diverse dynamics remains impractical. Demonstration videos are the natural remedy, specifying any dynamics frame by frame; yet a video shows its dynamics only through one particular appearance, a single shadow of the underlying dynamics, so actions learned from demonstrations transfer poorly to new scenes. ShadowDancer addresses this with two key innovations: (1) shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by our Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it; and (2) cross-shadow prediction, which learns actions by predicting one shadow from the other, so that whatever the pairing resamples is discarded by construction and whatever it preserves becomes the action, yielding a unified dynamics representation that drives a block-causal world model. Any demonstrated clip thus becomes a reusable action asset, replayed in new environments without action labels, motion estimators, or fine-tuning. Experiments demonstrate improved action transfer and long action rollout over strong latent-action and interactive world model baselines across diverse dynamics families, with an average blinded win rate of 86% in rollout comparisons. We show video results at https://ShadowDancer-1.github.io
Figures
Reference graph
Works this paper leans on
-
[1]
Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575,
Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, et al. Cosmos world foundation model platform for physical ai.arXiv preprint arXiv:2501.03575,
-
[2]
Alemi, Ian Fischer, Joshua V
Alexander A. Alemi, Ian Fischer, Joshua V . Dillon, and Kevin Murphy. Deep variational information bottleneck. InICLR,
-
[3]
Diffusion for world modeling: Visual details matter in atari.NeurIPS,
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kanervisto, Amos J Storkey, Tim Pearce, and Franc ¸ois Fleuret. Diffusion for world modeling: Visual details matter in atari.NeurIPS,
-
[4]
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, et al. V-JEPA 2: Self-supervised video models enable understanding, predic- tion and planning.arXiv preprint arXiv:2506.09985, 2025. 4
Pith/arXiv arXiv 2025
-
[5]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. 2
2024
-
[6]
Genie: Generative interactive environments
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker- Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et al. Genie: Generative interactive environments. InICML, 2024. 2, 3
2024
-
[7]
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. UniVLA: Learning to act anywhere with task-centric latent actions.arXiv preprint arXiv:2505.06111, 2025. 3
Pith/arXiv arXiv 2025
-
[8]
Unifying 9 precisely 3D-enhanced camera and human motion controls for video generation
Chenjie Cao, Jingkai Zhou, Shikai Li, Jingyun Liang, Chao- hui Yu, Fan Wang, Xiangyang Xue, and Yanwei Fu. Unifying 9 precisely 3D-enhanced camera and human motion controls for video generation. InSIGGRAPH Asia, 2025. 2, 3, 6, 5
2025
-
[9]
SkyReels-V2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074,
Guibin Chen, Dixuan Lin, Jiangping Yang, Chunze Lin, Junchen Zhu, Mingyuan Fan, Hao Zhang, Sheng Chen, Zheng Chen, Chengcheng Ma, et al. SkyReels-V2: Infinite-length film generative model.arXiv preprint arXiv:2504.13074,
-
[10]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geof- frey Hinton. A simple framework for contrastive learning of visual representations. InICML, 2020. 3
2020
-
[11]
Xiaoyu Chen, Junliang Guo, Tianyu He, Chuheng Zhang, Pushi Zhang, Derek Cathera Yang, Li Zhao, and Jiang Bian. IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI.arXiv preprint arXiv:2411.00785, 2024. 2, 3
Pith/arXiv arXiv 2024
-
[12]
Xiaoyu Chen, Hangxing Wei, Pushi Zhang, Chuheng Zhang, Kaixin Wang, Yanjiang Guo, Rushuai Yang, Yucen Wang, Xinquan Xiao, Li Zhao, et al. villa-X: enhancing latent action modeling in vision-language-action models.arXiv preprint arXiv:2507.23682, 2025
Pith/arXiv arXiv 2025
-
[13]
Moto: Latent motion token as the bridging language for learning robot manipulation from videos
Yi Chen, Yuying Ge, Weiliang Tang, Yizhuo Li, Yixiao Ge, Mingyu Ding, Ying Shan, and Xihui Liu. Moto: Latent motion token as the bridging language for learning robot manipulation from videos. InICCV, 2025. 3
2025
-
[14]
Oasis: A universe in a transformer
Decart, Julian Quevedo, Quinn McIntyre, Spruce Campbell, Xinlei Chen, and Robert Wachen. Oasis: A universe in a transformer. 2024. 2, 3
2024
-
[15]
Imitating latent policies from observation
Ashley Edwards, Himanshu Sahni, Yannick Schroecker, and Charles Isbell. Imitating latent policies from observation. In ICML, 2019. 2, 3
2019
-
[16]
Zhixue Fang, Xu He, Songlin Tang, Haoxian Zhang, Qingfeng Li, Xiaoqiang Liu, Pengfei Wan, and Kun Gai. 3D-aware implicit motion control for view-adaptive human video gener- ation.arXiv preprint arXiv:2602.03796, 2026. 2, 3
arXiv 2026
-
[17]
Vista: A generalizable driving world model with high fidelity and versatile controllability.NeurIPS, 2024
Shenyuan Gao, Jiazhi Yang, Li Chen, Kashyap Chitta, Yihang Qiu, Andreas Geiger, Jun Zhang, and Hongyang Li. Vista: A generalizable driving world model with high fidelity and versatile controllability.NeurIPS, 2024. 3
2024
-
[18]
AdaWorld: Learning adaptable world mod- els with latent actions
Shenyuan Gao, Siyuan Zhou, Yilun Du, Jun Zhang, and Chuang Gan. AdaWorld: Learning adaptable world mod- els with latent actions. InICML, 2025. 2, 3, 5, 6, 1
2025
-
[19]
Infinite worlds with versatile interactions.arXiv preprint arXiv:2607.07534, 2026
Zelin Gao, Qiuyu Wang, Jiapeng Zhu, Jingye Chen, Zichen Liu, Qingyan Bai, Jiahao Wang, Yufeng Yuan, Hanlin Wang, Yichong Lu, Ka Leong Cheng, Haojie Zhang, Jian Gao, Tian- rui Feng, Yuzheng Liu, Yao Yao, Yinghao Xu, Xing Zhu, Yujun Shen, and Hao Ouyang. Infinite worlds with versatile interactions.arXiv preprint arXiv:2607.07534, 2026. 2, 3, 7, 8
Pith/arXiv arXiv 2026
-
[20]
Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026
Quentin Garrido, Tushar Nagarajan, Basile Terver, Nico- las Ballas, Yann LeCun, and Michael Rabbat. Learning latent action world models in the wild.arXiv preprint arXiv:2601.05230, 2026. 2, 3
arXiv 2026
-
[21]
Google DeepMind. Veo. Model page. 1
-
[22]
Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Sch¨olkopf
Luigi Gresele, Paul K. Rubenstein, Arash Mehrjou, Francesco Locatello, and Bernhard Sch¨olkopf. The incomplete rosetta stone problem: Identifiability results for multi-view nonlinear ICA. InProceedings of the 35th Conference on Uncertainty in Artificial Intelligence, pages 217–227, 2020. 2
2020
-
[23]
Long-context autoregressive video modeling with next-frame prediction
Yuchao Gu, Weijia Mao, and Mike Zheng Shou. Long-context autoregressive video modeling with next-frame prediction. arXiv preprint arXiv:2503.19325, 2025. 3, 6
Pith/arXiv arXiv 2025
-
[24]
World models.arXiv preprint arXiv:1803.10122, 2018
David Ha and J ¨urgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018. 3
Pith/arXiv arXiv 2018
-
[25]
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2023. 3
Pith/arXiv arXiv 2023
-
[26]
Xianglong He, Chunli Peng, Zexiang Liu, Boyang Wang, Yifan Zhang, Qi Cui, Fei Kang, Biao Jiang, Mengyin An, Yangyang Ren, Baixin Xu, Hao-Xiang Guo, Kaixiong Gong, Cyrus Wu, Wei Li, Xuchen Song, Yang Liu, Eric Li, and Yahui Zhou. Matrix-Game 2.0: An open-source, real-time, and streaming interactive world model.arXiv preprint arXiv:2508.13009, 2025. 2, 3
Pith/arXiv arXiv 2025
-
[27]
beta-V AE: Learning basic visual con- cepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner. beta-V AE: Learning basic visual con- cepts with a constrained variational framework. InICLR,
-
[28]
RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025
Yicong Hong, Yiqun Mei, Chongjian Ge, Yiran Xu, Yang Zhou, Sai Bi, Yannick Hold-Geoffroy, Mike Roberts, Matthew Fisher, Eli Shechtman, et al. RELIC: Interactive video world model with long-horizon memory.arXiv preprint arXiv:2512.04040, 2025. 3
arXiv 2025
-
[29]
Approximation capabilities of multilayer feed- forward networks.Neural Networks, 4(2):251–257, 1991
Kurt Hornik. Approximation capabilities of multilayer feed- forward networks.Neural Networks, 4(2):251–257, 1991. 4
1991
-
[30]
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autore- gressive video diffusion.arXiv preprint arXiv:2506.08009,
-
[31]
Joel Jang, Seonghyeon Ye, Zongyu Lin, Jiannan Xiang, Jo- han Bjorck, Yu Fang, Fengyuan Hu, Spencer Huang, Kaushil Kundalia, Yen-Chen Lin, et al. DreamGen: Unlocking gener- alization in robot learning through video world models.arXiv preprint arXiv:2505.12705, 2025. 3
Pith/arXiv arXiv 2025
-
[32]
Yuxin Jiang, Yuchao Gu, Ivor W. Tsang, and Mike Zheng Shou. Olaf-world: Orienting latent actions for video world modeling.arXiv preprint arXiv:2602.10104, 2026. 2, 3, 4, 5, 6, 7, 8, 1
Pith/arXiv arXiv 2026
-
[33]
Miradata: A large-scale video dataset with long durations and structured captions.NeurIPS, 2024
Xuan Ju, Yiming Gao, Zhaoyang Zhang, Ziyang Yuan, Xintao Wang, Ailing Zeng, Yu Xiong, Qiang Xu, and Ying Shan. Miradata: A large-scale video dataset with long durations and structured captions.NeurIPS, 2024. 6, 5
2024
-
[34]
Variational autoencoders and nonlinear ica: A unifying framework
Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational autoencoders and nonlinear ica: A unifying framework. InAISTATS, 2020. 2, 3
2020
-
[35]
UniSkill: Imitating human videos via cross-embodiment skill representations
Hanjung Kim, Jaehyun Kang, Hyolim Kang, Meedeum Cho, Seon Joo Kim, and Youngwoon Lee. UniSkill: Imitating human videos via cross-embodiment skill representations. arXiv preprint arXiv:2505.08787, 2025. 3
arXiv 2025
-
[36]
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, et al. HunyuanVideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024. 2
Pith/arXiv arXiv 2024
-
[37]
Generative video motion editing with 3D point tracks
Yao-Chih Lee, Zhoutong Zhang, Jiahui Huang, Jui-Hsien Wang, Joon-Young Lee, Jia-Bin Huang, Eli Shechtman, and 10 Zhengqi Li. Generative video motion editing with 3D point tracks. InCVPR, 2026. 2, 3
2026
-
[38]
Lu Ling, Yichen Sheng, Zhi Tu, Wentian Zhao, Cheng Xin, Kun Wan, Lantao Yu, Qianyu Guo, Zixun Yu, Yawen Lu, et al. Dl3dv-10k: A large-scale scene dataset for deep learning- based 3d vision.arXiv preprint arXiv:2312.16256, 2023. 6, 5
Pith/arXiv arXiv 2023
-
[39]
Flow straight and fast: Learning to generate and transfer data with rectified flow
Xingchao Liu, Chengyue Gong, and qiang liu. Flow straight and fast: Learning to generate and transfer data with rectified flow. InICLR, 2023. 5
2023
-
[40]
Challenging common assumptions in the unsu- pervised learning of disentangled representations
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Raetsch, Sylvain Gelly, Bernhard Sch ¨olkopf, and Olivier Bachem. Challenging common assumptions in the unsu- pervised learning of disentangled representations. InICML,
-
[41]
Xiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu, Kain- ing Ying, Tong He, Jiangmiao Pang, Yu Qiao, and Kaipeng Zhang. Yume-1.5: A text-controlled interactive world genera- tion model.arXiv preprint arXiv:2512.22096, 2025. 2, 3, 7, 8
arXiv 2025
-
[42]
Open X-Embodiment: Robotic learning datasets and RT-X models
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Ab- hishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Poo- ley, Agrim Gupta, et al. Open X-Embodiment: Robotic learning datasets and RT-X models. InIEEE International Conference on Robotics and Automation (ICRA), pages 6892– 6903, 2024. 6, 5
2024
-
[43]
Genie 3: A new frontier for world models
Jack Parker-Holder, Shlomi Fruchter, et al. Genie 3: A new frontier for world models. https://deepmind.googl e/models/genie/. Blog post. 2, 3
-
[44]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InICCV, 2023. 5
2023
-
[45]
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. InICML, 2021. 3
2021
-
[46]
Derpanis, and Kostas Daniilidis
Oleh Rybkin, Karl Pertsch, Andrew Jaegle, Konstantinos G. Derpanis, and Kostas Daniilidis. Learning what you can do before doing anything. InICLR, 2019. 2, 3
2019
-
[47]
MotionStream: Real-time video generation with interactive motion controls
Joonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu, Jaesik Park, Eli Shechtman, and Xun Huang. MotionStream: Real-time video generation with interactive motion controls. InICLR, 2026. 2, 3
2026
-
[48]
A benchmark for the evalua- tion of RGB-D SLAM systems
J¨urgen Sturm, Nikolas Engelhard, Felix Endres, Wolfram Burgard, and Daniel Cremers. A benchmark for the evalua- tion of RGB-D SLAM systems. InIEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 573–580, 2012. 6
2012
-
[49]
Wenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu, Zehan Wang, Zhenwei Wang, Yunhong Wang, Jun Zhang, Tengfei Wang, and Chunchao Guo. WorldPlay: towards long- term geometric consistency for real-time interactive world modeling.arXiv preprint arXiv:2512.14614, 2025. 2, 3
Pith/arXiv arXiv 2025
-
[50]
Junshu Tang, Jiacheng Liu, Jiaqi Li, Longhuang Wu, Haoyu Yang, Penghao Zhao, Siruis Gong, Xiang Yuan, Shuai Shao, and Qinglin Lu. Hunyuan-GameCraft-2: Instruction- following interactive game world model.arXiv preprint arXiv:2511.23429, 2025. 3
arXiv 2025
-
[51]
Stone Tao, Fanbo Xiang, Arth Shukla, Yuzhe Qin, Xander Hinrichsen, Xiaodi Yuan, Chen Bao, Xinsong Lin, Yulin Liu, Tse-kai Chan, Yuan Gao, Xuanlin Li, Tongzhou Mu, Nan Xiao, Arnav Gurha, Viswesh Nagaswamy Rajesh, Yong Woo Choi, Yen-Ru Chen, Zhiao Huang, Roberto Calandra, Rui Chen, Shan Luo, and Hao Su. Maniskill3: Gpu parallelized robotics simulation and r...
Pith/arXiv arXiv 2024
-
[52]
Advancing open-source world models.arXiv preprint arXiv:2601.20540, 2026
Robbyant Team, Zelin Gao, Qiuyu Wang, Yanhong Zeng, Jiapeng Zhu, Ka Leong Cheng, Yixuan Li, Hanlin Wang, Yinghao Xu, Shuailei Ma, et al. Advancing open-source world models.arXiv preprint arXiv:2601.20540, 2026. 2, 3
Pith/arXiv arXiv 2026
-
[53]
Video- MAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.NeurIPS, 2022
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Video- MAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training.NeurIPS, 2022. 4
2022
-
[54]
Self-supervised learning with data aug- mentations provably isolates content from style
Julius von K¨ugelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Sch ¨olkopf, Michel Besserve, and Francesco Locatello. Self-supervised learning with data aug- mentations provably isolates content from style. InAdvances in Neural Information Processing Systems, pages 16451– 16467, 2021. 2
2021
-
[55]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianxiao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025. 2, 3, 5, 1
Pith/arXiv arXiv 2025
-
[56]
VGGT: Visual geometry grounded transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. VGGT: Visual geometry grounded transformer. InCVPR, 2025. 6
2025
-
[57]
Co-Evolving latent action world models.arXiv preprint arXiv:2510.26433, 2025
Yucen Wang, Fengming Zhang, De-Chuan Zhan, Li Zhao, Kaixin Wang, and Jiang Bian. Co-Evolving latent action world models.arXiv preprint arXiv:2510.26433, 2025. 3
Pith/arXiv arXiv 2025
-
[58]
Connectionist nonparametric regression: Mul- tilayer feedforward networks can learn arbitrary mappings
Halbert White. Connectionist nonparametric regression: Mul- tilayer feedforward networks can learn arbitrary mappings. Neural Networks, 3(5):535–549, 1990. 4
1990
-
[59]
WorldMem: Long- term consistent world simulation with memory.arXiv preprint arXiv:2504.12369, 2025
Zeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang, Shuai Yang, Yanhong Zeng, and Xingang Pan. WorldMem: Long- term consistent world simulation with memory.arXiv preprint arXiv:2504.12369, 2025. 3
arXiv 2025
-
[60]
Jiange Yang, Yansong Shi, Haoyi Zhu, Mingyu Liu, Kai- jing Ma, Yating Wang, Gangshan Wu, Tong He, and Limin Wang. CoMo: Learning continuous latent motion from in- ternet videos for scalable robot learning.arXiv preprint arXiv:2505.17006, 2025. 2, 3, 4
Pith/arXiv arXiv 2025
-
[61]
Latent Action Pretraining from Videos
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu- Wei Chao, Bill Yuchen Lin, et al. Latent Action Pretraining from Videos. InICLR, 2025. 2, 3
2025
-
[62]
Yixuan Ye, Xuanyu Lu, Yuxin Jiang, Yuchao Gu, Rui Zhao, Qiwei Liang, Jiachun Pan, Fengda Zhang, Weijia Wu, and Alex Jinpeng Wang. MIND: Benchmarking memory con- sistency and action control in world models.arXiv preprint arXiv:2602.08025, 2026. 3
arXiv 2026
-
[63]
Jiwen Yu, Yiran Qin, Xintao Wang, Pengfei Wan, Di Zhang, and Xihui Liu. GameFactory: Creating new games with gen- erative interactive videos.arXiv preprint arXiv:2501.08325,
-
[64]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. InCVPR, 2018. 6
2018
-
[65]
replay the dynamics, resample everything else
Shangwen Zhu, Qianyu Peng, Zhao Pu, Zhilei Shu, Xiangrui Ke, Zhaohu Xing, Zizhao Tong, Zeqing Wang, Xinyu Cui, Zian Zheng, Huangji Wang, Jian Zhao, Yeying Jin, Fan Cheng, and Ruili Feng. Incantation: Natural language as the action interface for multi-entity video world models.arXiv preprint arXiv:2605.18601, 2026. 2, 3, 7 12 ShadowDancer: Teaching Video W...
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.