Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

World action models can learn action-relevant 4D geometry in training and still deploy as the same lightweight video-action policies.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 14:22 UTC pith:Q47DYNFZ

load-bearing objection Solid training-only 4D recipe for WAMs: decayed current-frame reads plus action-aware temporal distillation give real, modest gains at flat inference cost; missing permanent-zero-g0 control is a fair gap, not a collapse. the 2 major comments →

arxiv 2607.05468 v1 pith:Q47DYNFZ submitted 2026-07-06 cs.RO cs.AI

Learning 4D Geometric Priors for Inference-Efficient World Action Models

classification cs.RO cs.AI
keywords World Action Models4D geometric priorsrobotic manipulationmulti-expert co-traininginference efficiencyvideo-action modelsgeometric distillationdecayed attention
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

World action models jointly learn how scenes evolve and how robots should act, but their video latents often track appearance more than the changing spatial relations that decide whether a grasp reaches, aligns, or holds. This paper claims those action-relevant 4D geometric priors can be injected only during training. MECo-WAM adds a lightweight 4D expert supervised by relational targets from a frozen geometry encoder, restricts temporary current-frame geometry reads that decay to zero, and distills within-frame and temporal relations with weights that favor regions tied to robot motion. At deployment every 4D piece is removed, so the observation-to-action graph and latency stay unchanged. On LIBERO, RoboTwin 2.0, and real tabletop tasks, success and spatial grounding improve without extra inference cost, which matters because precise manipulation needs geometry that pure video prediction may not keep.

Core claim

MECo-WAM establishes that action-relevant 4D geometric priors can be transferred into shared video-action representations purely at training time—via multi-expert co-training, decayed 4D read-mask attention, and action-aware temporal geometric distillation—so the deployed policy retains the original lightweight observation-to-action interface while raising manipulation success.

What carries the argument

Decayed 4D read-mask attention plus action-aware temporal geometric distillation: temporary current-frame-only geometry reads that linearly decay to zero before deployment, and relational matching of within-frame and temporal geometry weighted toward action-coupled visual tokens, supervised by frozen encoder targets.

Load-bearing premise

Temporary current-frame geometry cues from a frozen encoder, faded out before deployment, leave enough action-relevant structure inside ordinary video-action tokens to improve real manipulation success.

What would settle it

Train matched video-action models with and without the 4D path on the same data; if held-out success, action error, and depth/pose probes show no gain once the decay schedule reaches zero on LIBERO, RoboTwin, and the real cube tasks, the transfer claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Training-only 4D supervision can raise task success without any geometry module, sensor, or decoder at inference.
  • Appearance-oriented video co-training alone under-serves precise manipulation; action-conditioned geometric relations close that gap.
  • Non-pretrained world action models can match or exceed strong pretrained baselines when geometry is transferred this way.
  • Real-robot spatial grounding improves (fewer corrections, shorter completion) when geometry lives inside the deployed video-action features.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same decayed-auxiliary pattern could carry other costly signals—depth, contact, force—into policies without changing the deployed stack.
  • A frozen geometry teacher may be enough whenever the student only needs relative layout and its change, not absolute feature coordinates.
  • Existing video-action co-training pipelines could bolt on a temporary 4D branch and drop it at export without redesigning serving systems.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. MECo-WAM is a multi-expert world action model that injects action-relevant 4D geometric priors into video-action representations only during training, then removes all auxiliary geometry at deployment so the inference graph and latency match the base Fast-WAM observation-to-action path. Training adds a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder; asymmetric expert masks and decayed 4D read-mask attention (temporary current-frame g0 reads that linearly decay to zero) are intended to transfer geometry without non-causal future-geometry shortcuts; action-aware temporal geometric distillation aligns within-frame relations and their temporal changes with action-conditioned token weights. Reported results are LIBERO 98.2% average success (without embodied pretraining), RoboTwin 2.0 92.62% average, and modest real-world gains on two ARX-R5 tabletop tasks, with ablations and depth/grasp probes supporting transfer into the deployed representation.

Significance. If the transfer claim holds, the paper offers a practical and timely recipe for geometry-aware WAMs: use a frozen 4D teacher and temporary read edges as a training-time representation constraint rather than an inference-time 4D decoder or sensor stack. That addresses a real deployment tension between geometric grounding and latency. Strengths include a clear training-only design, matched ablations isolating the 4D expert, decayed read, spatial/temporal losses and action-aware weights (Table 4), multi-benchmark evaluation (LIBERO suites, RoboTwin clean/random, real robot), and representation probes that freeze the video-action backbone. The absolute gains are modest but directionally consistent and obtained without increasing inference cost, which is a useful contribution for efficient embodied policies.

major comments (2)
  1. Methodology (Decayed 4D Read-Mask Attention, Eqs. 9–10, Fig. 3) and Table 4: the central claim is that temporary current-frame g0 reads transfer action-relevant geometry into a permanently 4D-free video-action pathway. Ablations show that an isolated 4D expert barely moves RoboTwin SR (91.83→91.87) while action MSE worsens, and that decayed read access is what lifts SR and lowers MSE. That pattern is consistent with useful transfer but also with early privileged g0 side-channel shaping of shared tokens before p4d reaches 0. A load-bearing control is missing: keep the full 4D expert and L4d while permanently zeroing g0→video/action edges (or freeze the video-action path after the decay window and continue 4D-only training). Without that isolation, the modest gains cannot be cleanly attributed to pure representation transfer versus temporary privileged geometry.
  2. Real-world evaluation (Table 3, Fig. 5): only two tabletop tasks on one ARX-R5 arm, with small trial counts and high run-to-run variance (e.g., Stack Cubes SR tied with Fast-WAM at 60%; Sort Cubes gains rest on limited trials). The paper’s claim of improved real-robot spatial grounding is directionally supported by lower corrections and shorter completion time, but the evidence is too thin to carry the deployment-efficiency narrative alongside LIBERO/RoboTwin. Either expand the real-world suite (more tasks, more trials, clearer success criteria) or temper the real-world claim to a preliminary study.
minor comments (5)
  1. Figure 1 and abstract/intro: RoboTwin is reported as 92.6% while Table 2 gives 92.62%; keep one consistent rounding convention.
  2. Implementation details: free parameters (p4d schedule, λvideo/λaction/λ4d, αgeo/αtem, η, keyframe set K, 4× action-to-video ratio) are listed but sensitivity is not reported; a short appendix on schedule/weight robustness would strengthen reproducibility.
  3. Notation: expert slots use f/a/g with subscripts 0,1,2,h and predictions fp/ap/gp; a compact token glossary near Eqs. (2)–(5) would help readers track the mixed-attention sequence.
  4. Related Work: contrast with X-WAM and WAM4D is clear on deployment cost; a one-sentence statement of what is *not* claimed (no new 4D reconstruction quality, no multi-view RGB-D futures) would further bound the contribution.
  5. Figure 4/6 probes are useful; state explicitly that the DPT head and grasp analysis use only deployed video-action tokens with no 4D expert at probe time (already implied, but worth one sentence).

Circularity Check

0 steps flagged

No significant circularity: empirical training method evaluated on external benchmarks; losses and success metrics are not equivalent by construction.

full rationale

MECo-WAM is a standard empirical multi-expert training recipe. Video/action objectives are conditional flow-matching (Eqs. 20–21); the 4D branch is supervised by relational and temporal-relation losses against a frozen third-party VGGT teacher (Eqs. 11–19, 22–23). Decayed read-mask attention (Eqs. 9–10, Fig. 3) is a training schedule that is set to zero at deployment; it does not redefine the reported success rates. Claimed gains are measured on external suites (LIBERO, RoboTwin 2.0) and real-robot rollouts, with ablations (Table 4) and representation probes (Figs. 4, 6) that exclude auxiliary 4D tokens. Building on the Fast-WAM observation-to-action interface is ordinary base-model reuse, not a load-bearing self-citation that forces the result. Skeptical concerns about residual leakage from temporary g0 reads are experimental-isolation questions, not reductions of predictions to inputs by definition. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz-smuggling chain is present.

Axiom & Free-Parameter Ledger

4 free parameters · 3 axioms · 2 invented entities

The central claim rests on standard flow-matching training, a frozen external geometry encoder, and several hand-chosen schedules and loss weights that control how geometry is transferred. No new physical entities are postulated; the free parameters are ordinary training hyper-parameters whose values are stated but not derived.

free parameters (4)
  • p4d decay schedule (pstart=1.0, pend=0, Sdecay=half training)
    Linear decay of current-frame geometry read probability is chosen by hand; the claim that geometry transfers without permanent dependence rests on this schedule.
  • loss weights λvideo, λaction, λ4d, αgeo, αtem, η
    Relative weighting of video, action, spatial-relation and temporal-relation terms, plus the uniform-mixture coefficient η, are free hyper-parameters that shape the reported gains.
  • keyframe set K and action-to-video temporal ratio 4×
    Which future slots receive 4D supervision and the 33-step / 9-frame chunk design are design choices that affect the strength of the geometric signal.
  • 4D expert width d4d=512 (≈0.45 B params)
    Capacity of the auxiliary expert is selected rather than derived; ablation shows an isolated expert alone is insufficient.
axioms (3)
  • domain assumption Frozen VGGT relational features are a sufficiently accurate and action-relevant teacher of 4D geometry for manipulation.
    All geometric supervision is taken from VGGT; if those features are poorly aligned with grasp/contact geometry the transfer claim fails.
  • ad hoc to paper Asymmetric expert visibility plus current-frame-only temporary reads prevent non-causal future-geometry shortcuts into action generation.
    The mask design is asserted to block shortcuts; no formal information-flow proof is given.
  • domain assumption Conditional flow-matching on video latents and action chunks is a valid joint training objective for WAMs.
    Inherited from the Fast-WAM / video-diffusion literature and used without re-derivation.
invented entities (2)
  • decayed 4D read-mask attention no independent evidence
    purpose: Provide early geometric guidance to video/action experts then remove the dependency before deployment.
    New attention-mask schedule introduced by the paper; independent evidence is only the ablation table.
  • action-aware temporal geometric distillation no independent evidence
    purpose: Align pairwise geometric relations and their temporal changes while weighting tokens by action relevance.
    New loss construction; evidence is internal ablations and probes, not external validation of the weighting scheme.

pith-pipeline@v1.1.0-grok45 · 18473 in / 2852 out tokens · 25963 ms · 2026-07-11T14:22:52.079595+00:00 · methodology

0 comments
read the original abstract

World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.

Figures

Figures reproduced from arXiv: 2607.05468 by Chong Ma, Hanli Wang, Jianjun Zhang, Jian Zhu, Taiyi Su, Yi Xu, Zitai Huang.

Figure 1
Figure 1. Figure 1: Comparison of MECo-WAM with Baselines in [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Overview of MECo-WAM. The left side illustrates the training process, while the right side shows the inference [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Decayed read-mask attention. Rows and columns [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Depth probing of shared video-action representa [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Representative real-robot tabletop experiments on cube stacking and size-based cube sorting. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: 3D position and pose sensitivity during real-robot [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

    cs.RO 2026-07 conditional novelty 6.0

    ST-WAM adds DINOv3 semantic future prediction and history retrieval to a video-based world-action model, improving zero-shot robustness under visual distribution shifts.

Reference graph

Works this paper leans on

79 extracted references · 22 linked inside Pith · cited by 1 Pith paper

  1. [2]

    Pai, Jonas and Achenbach, Liam and Montesinos, Victoriano and Forrai, Benedek and Mees, Oier and Nava, Elvis , journal=

  2. [3]

    Ye, Angen and Wang, Boyuan and Ni, Chaojun and Huang, Guan and Zhao, Guosheng and Li, Hao and Li, Hengtao and Li, Jie and Lv, Jindi and Liu, Jingyu and others , journal=

  3. [4]

    Yuan, Tianyuan and Dong, Zibin and Liu, Yicheng and Zhao, Hang , journal=

  4. [5]

    Kim, Moo Jin and Gao, Yihuai and Lin, Tsung-Yi and Lin, Yen-Chen and Ge, Yunhao and Lam, Grace and Liang, Percy and Song, Shuran and Liu, Ming-Yu and Finn, Chelsea and others , journal=

  5. [7]

    Li, Ying and Wei, Xiaobao and Cao, Jiajun and Wang, Hao and Chi, Xiaowei and Bai, Chengyu and Sun, Qianpu and Li, Jiajun and Zhang, Xiaojie and Tang, Jian and others , journal=

  6. [8]

    Team, MotuBrain and Xiang, Chendong and Bao, Fan and Liu, Haitian and Tan, Hengkai and Bi, Hongzhe and Li, James and Liu, Jiabao and Pang, Jingrui and Jing, Kiro and others , journal=

  7. [9]

    Bi, Hongzhe and Tan, Hengkai and Xie, Shenghao and Wang, Zeyuan and Huang, Shuhe and Liu, Haitian and Zhao, Ruowen and Feng, Yao and Xiang, Chendong and Rong, Yinze and others , booktitle=

  8. [11]

    Xu, Gangwei and Zhang, Qihang and Zhou, Jiaming and Zhu, Xing and Shen, Yujun and Yang, Xin and Xu, Yinghao , journal=

  9. [12]

    Wan, Team and Wang, Ang and Ai, Baole and Wen, Bin and Mao, Chaojie and Xie, Chen-Wei and Chen, Di and Yu, Feiwu and Zhao, Haiming and Yang, Jianxiao and others , journal=

  10. [13]

    Wang, Jianyuan and Chen, Minghao and Karaev, Nikita and Vedaldi, Andrea and Rupprecht, Christian and Novotny, David , booktitle=

  11. [14]

    Mu, Yao and Chen, Tianxing and Chen, Zanxin and Peng, Shijia and Lan, Zhiqian and Gao, Zeyu and Liang, Zhixuan and Yu, Qiaojun and Zou, Yude and Xu, Mingkun and others , booktitle=

  12. [15]

    Liu, Bo and Zhu, Yifeng and Gao, Chongkai and Feng, Yihao and Liu, Qiang and Zhu, Yuke and Stone, Peter , journal=

  13. [16]

    Black, Kevin and Brown, Noah and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and Groom, Lachy and Hausman, Karol and Ichter, Brian and others , journal=

  14. [17]

    Intelligence, Physical and Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and others , journal=

  15. [18]

    Qu, Delin and Song, Haoming and Chen, Qizhi and Yao, Yuanqi and Ye, Xinyi and Ding, Yan and Wang, Zhigang and Gu, JiaYuan and Zhao, Bin and Wang, Dong and others , journal=

  16. [19]

    Pertsch, Karl and Stachowicz, Kyle and Ichter, Brian and Driess, Danny and Nair, Suraj and Vuong, Quan and Mees, Oier and Finn, Chelsea and Levine, Sergey , journal=

  17. [20]

    2025 , volume =

    Kim, Moo Jin and Pertsch, Karl and Karamcheti, Siddharth and Xiao, Ted and Balakrishna, Ashwin and Nair, Suraj and Rafailov, Rafael and Foster, Ethan P and Sanketi, Pannag R and Vuong, Quan and Kollar, Thomas and Burchfiel, Benjamin and Tedrake, Russ and Sadigh, Dorsa and Levine, Sergey and Liang, Percy and Finn, Chelsea , booktitle =. 2025 , volume =

  18. [21]

    Kim, Moo Jin and Finn, Chelsea and Liang, Percy , journal=

  19. [22]

    Liang, Zhixuan and Li, Yizhuo and Yang, Tianshuo and Wu, Chengyue and Mao, Sitong and Pei, Liuao and Nian, Tian and Zhou, Shunbo and Yang, Xiaokang and Pang, Jiangmiao and others , journal=

  20. [24]

    Zheng, Jinliang and Li, Jianxiong and Wang, Zhihao and Liu, Dongxiu and Kang, Xirui and Feng, Yuchun and Zheng, Yinan and Zou, Jiayin and Chen, Yilun and Zeng, Jia and others , journal=

  21. [25]

    2025 , url=

    Li, Peiyan and Chen, Yixiang and Wu, Hongtao and Ma, Xiao and Wu, Xiangnan and Huang, Yan and Wang, Liang and Kong, Tao and Tan, Tieniu , booktitle=. 2025 , url=

  22. [26]

    2025 , editor=

    Li, Xiaoqi and Heng, Liang and Liu, Jiaming and Shen, Yan and Gu, Chenyang and Liu, Zhuoyang and Chen, Hao and Han, Nuowei and Zhang, Renrui and Tang, Hao and Zhang, Shanghang and Dong, Hao , booktitle=. 2025 , editor=

  23. [27]

    Fan, Xianzhe and Deng, Shengliang and Wu, Xiaoyang and Lu, Yuxiang and Li, Zhuoling and Yan, Mi and Zhang, Yujia and Zhang, Zhizheng and Wang, He and Zhao, Hengshuang , journal=

  24. [28]

    Ni, Chaojun and Chen, Cheng and Wang, Xiaofeng and Zhu, Zheng and Zheng, Wenzhao and Wang, Boyuan and Chen, Tianrun and Zhao, Guosheng and Li, Haoyun and Dong, Zhehao and Zhang, Qiang and Ye, Yun and Wang, Yang and Huang, Guan and Mei, Wenjun , booktitle=

  25. [29]

    2026 , url=

    Li, Fuhao and Song, Wenxuan and Zhao, Han and Wang, Jingbo and Ding, Pengxiang and Wang, Donglin and Zeng, Long and Li, Haoang , booktitle=. 2026 , url=

  26. [30]

    Conference on Robot Learning , year=

    Generalist Robot Manipulation beyond Action Labeled Data , author=. Conference on Robot Learning , year=

  27. [31]

    Sun, Lin and Xie, Bin and Liu, Yingfei and Shi, Hao and Wang, Tiancai and Cao, Jiale , journal=

  28. [32]

    Zhang, Zhengshen and Li, Hao and Dai, Yalun and Zhu, Zhengbang and Zhou, Lei and Liu, Chenchen and Wang, Dong and Tay, Francis E. H. and Chen, Sijin and Liu, Ziwei and Liu, Yuxiao and Li, Xinghang and Zhou, Pan , booktitle=

  29. [33]

    2024 , volume=

    Zhen, Haoyu and Qiu, Xiaowen and Chen, Peihao and Yang, Jincheng and Yan, Xin and Du, Yilun and Hong, Yining and Gan, Chuang , booktitle=. 2024 , volume=

  30. [34]

    Qian, Jingjing and Han, Boyao and Shi, Chen and Xiao, Lei and Yang, Long and Shi, Shaoshuai and Jiang, Li , booktitle=

  31. [35]

    International Conference on Learning Representations , year=

    Spatially Guided Training for Vision-Language-Action Model , author=. International Conference on Learning Representations , year=

  32. [36]

    Su, Taiyi and Zhu, Jian and Wang, Tianjian and He, Youzhang and Huang, Zitai and Zhang, Jianjun and Ma, Chong and Wang, Hanyang and Zhang, Tianjiao and Yin, Munan and others , journal=

  33. [37]

    Wang, Siyin and Shi, Junhao and Fu, Zhaoyang and He, Xinzhe and Liu, Feihong and Yang, Chenchen and Zhou, Yikang and Fei, Zhaoye and Gong, Jingjing and Fu, Jinlan and others , journal=

  34. [38]

    IEEE/CVF International Conference on Computer Vision , pages=

    Vision Transformers for Dense Prediction , author=. IEEE/CVF International Conference on Computer Vision , pages=

  35. [40]

    2025 , volume =

    Hu, Yucheng and Guo, Yanjiang and Wang, Pengchao and Chen, Xiaoyu and Wang, Yen-Jen and Zhang, Jianke and Sreenath, Koushil and Lu, Chaochao and Chen, Jianyu , booktitle =. 2025 , volume =

  36. [41]

    Ma, Teli and Zheng, Jia and Wang, Zifan and Jiang, Chunli and Cui, Andy and Liang, Junwei and Yang, Shuo , journal=

  37. [42]

    2025 , volume=

    Chen, Sili and Guo, Hengkai and Zhu, Shengnan and Zhang, Feihu and Huang, Zilong and Feng, Jiashi and Kang, Bingyi , booktitle=. 2025 , volume=

  38. [43]

    Bi, H.; Tan, H.; Xie, S.; Wang, Z.; Huang, S.; Liu, H.; Zhao, R.; Feng, Y.; Xiang, C.; Rong, Y.; et al. 2026. Motus : A unified latent action world model. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 35101--35113

  39. [44]

    Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024. _0 : A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164

  40. [45]

    Chen, S.; Guo, H.; Zhu, S.; Zhang, F.; Huang, Z.; Feng, J.; and Kang, B. 2025. Video Depth Anything : Consistent Depth Estimation for Super-Long Videos. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 22831--22840

  41. [46]

    Fan, X.; Deng, S.; Wu, X.; Lu, Y.; Li, Z.; Yan, M.; Zhang, Y.; Zhang, Z.; Wang, H.; and Zhao, H. 2026. Any3D-VLA : Enhancing VLA Robustness via Diverse Point Clouds. arXiv preprint arXiv:2602.00807

  42. [47]

    Guo, J.; Li, Q.; Li, P.; Chen, Z.; Sun, N.; Su, Y.; Wang, H.; Zhang, Y.; Li, X.; and Liu, H. 2026. Unified 4D world action modeling from video priors with asynchronous denoising. arXiv preprint arXiv:2604.26694

  43. [48]

    Hu, Y.; Guo, Y.; Wang, P.; Chen, X.; Wang, Y.-J.; Zhang, J.; Sreenath, K.; Lu, C.; and Chen, J. 2025. Video Prediction Policy : A Generalist Robot Policy with Predictive Visual Representations. In International Conference on Machine Learning, volume 267, 24328--24346

  44. [49]

    Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. 2025. _ 0.5 : A Vision-Language-Action Model with Open-World Generalization. arXiv preprint arXiv:2504.16054

  45. [50]

    J.; Finn, C.; and Liang, P

    Kim, M. J.; Finn, C.; and Liang, P. 2025. Fine-Tuning Vision-Language-Action Models : Optimizing Speed and Success. arXiv preprint arXiv:2502.19645

  46. [51]

    J.; Gao, Y.; Lin, T.-Y.; Lin, Y.-C.; Ge, Y.; Lam, G.; Liang, P.; Song, S.; Liu, M.-Y.; Finn, C.; et al

    Kim, M. J.; Gao, Y.; Lin, T.-Y.; Lin, Y.-C.; Ge, Y.; Lam, G.; Liang, P.; Song, S.; Liu, M.-Y.; Finn, C.; et al. 2026. Cosmos Policy : Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163

  47. [52]

    J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E

    Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E. P.; Sanketi, P. R.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2025. OpenVLA : An Open-Source Vision-Language-Action Model. In Conference on Robot Learning, volume 270, 2679--2713

  48. [53]

    Li, F.; Song, W.; Zhao, H.; Wang, J.; Ding, P.; Wang, D.; Zeng, L.; and Li, H. 2026 a . Spatial Forcing : Implicit Spatial Representation Alignment for Vision-Language-Action Model. In International Conference on Learning Representations

  49. [54]

    Li, L.; Zhang, Q.; Luo, Y.; Yang, S.; Wang, R.; Han, F.; Yu, M.; Gao, Z.; Xue, N.; Zhu, X.; et al. 2026 b . Causal World Modeling for Robot Control. arXiv preprint arXiv:2601.21998

  50. [55]

    Li, P.; Chen, Y.; Wu, H.; Ma, X.; Wu, X.; Huang, Y.; Wang, L.; Kong, T.; and Tan, T. 2025 a . BridgeVLA : Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models. In Conference on Neural Information Processing Systems

  51. [56]

    Li, X.; Heng, L.; Liu, J.; Shen, Y.; Gu, C.; Liu, Z.; Chen, H.; Han, N.; Zhang, R.; Tang, H.; Zhang, S.; and Dong, H. 2025 b . 3DS-VLA : A 3D Spatial-Aware Vision Language Action Model for Robust Multi-Task Manipulation. In Lim, J.; Song, S.; and Park, H.-W., eds., Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, 2344-...

  52. [57]

    Li, Y.; Wei, X.; Cao, J.; Wang, H.; Chi, X.; Bai, C.; Sun, Q.; Li, J.; Zhang, X.; Tang, J.; et al. 2026 c . WAM4D : Fast 4D World Action Model via Spatial Register Tokens. arXiv preprint arXiv:2606.14048

  53. [58]

    Liang, Z.; Li, Y.; Yang, T.; Wu, C.; Mao, S.; Pei, L.; Nian, T.; Zhou, S.; Yang, X.; Pang, J.; et al. 2025. Discrete Diffusion VLA : Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies. arXiv preprint arXiv:2508.20072

  54. [59]

    Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2023. LIBERO : Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36: 44776--44791

  55. [60]

    Ma, T.; Zheng, J.; Wang, Z.; Jiang, C.; Cui, A.; Liang, J.; and Yang, S. 2026. DiT4DiT : Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control. arXiv preprint arXiv:2603.10448

  56. [61]

    Mu, Y.; Chen, T.; Chen, Z.; Peng, S.; Lan, Z.; Gao, Z.; Liang, Z.; Yu, Q.; Zou, Y.; Xu, M.; et al. 2025. RoboTwin : Dual-arm robot benchmark with generative digital twins. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27649--27660

  57. [62]

    Ni, C.; Chen, C.; Wang, X.; Zhu, Z.; Zheng, W.; Wang, B.; Chen, T.; Zhao, G.; Li, H.; Dong, Z.; Zhang, Q.; Ye, Y.; Wang, Y.; Huang, G.; and Mei, W. 2026. SwiftVLA : Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13474--13485

  58. [63]

    Pai, J.; Achenbach, L.; Montesinos, V.; Forrai, B.; Mees, O.; and Nava, E. 2025. mimic-video : Video-action models for generalizable robot control beyond vlas. arXiv preprint arXiv:2512.15692

  59. [64]

    Pertsch, K.; Stachowicz, K.; Ichter, B.; Driess, D.; Nair, S.; Vuong, Q.; Mees, O.; Finn, C.; and Levine, S. 2025. FAST : Efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747

  60. [65]

    Qian, J.; Han, B.; Shi, C.; Xiao, L.; Yang, L.; Shi, S.; and Jiang, L. 2026. GeoPredict : Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13529--13539

  61. [66]

    Qu, D.; Song, H.; Chen, Q.; Yao, Y.; Ye, X.; Ding, Y.; Wang, Z.; Gu, J.; Zhao, B.; Wang, D.; et al. 2025. SpatialVLA : Exploring spatial representations for visual-language-action model. arXiv preprint arXiv:2501.15830

  62. [67]

    Ranftl, R.; Bochkovskiy, A.; and Koltun, V. 2021. Vision Transformers for Dense Prediction. In IEEE/CVF International Conference on Computer Vision, 12179--12188

  63. [68]

    Spiridonov, A.; Zaech, J.-N.; Nikolov, N.; Van Gool, L.; and Paudel, D. P. 2025. Generalist Robot Manipulation beyond Action Labeled Data. In Conference on Robot Learning

  64. [69]

    Su, T.; Zhu, J.; Li, Y.; Ma, C.; Zhang, J.; Huang, Z.; Wang, H.; and Xu, Y. 2025. Towards high-consistency embodied world model with multi-view trajectory videos. arXiv preprint arXiv:2511.12882

  65. [70]

    Su, T.; Zhu, J.; Wang, T.; He, Y.; Huang, Z.; Zhang, J.; Ma, C.; Wang, H.; Zhang, T.; Yin, M.; et al. 2026. DeMaVLA : A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation. arXiv preprint arXiv:2605.31286

  66. [71]

    Sun, L.; Xie, B.; Liu, Y.; Shi, H.; Wang, T.; and Cao, J. 2025. GeoVLA : Empowering 3D Representations in Vision-Language-Action Models. arXiv preprint arXiv:2508.09071

  67. [72]

    Team, M.; Xiang, C.; Bao, F.; Liu, H.; Tan, H.; Bi, H.; Li, J.; Liu, J.; Pang, J.; Jing, K.; et al. 2026. MotuBrain : An advanced world action model for robot control. arXiv preprint arXiv:2604.27792

  68. [73]

    Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.-W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. 2025. Wan : Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314

  69. [74]

    Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; and Novotny, D. 2025 a . VGGT : Visual Geometry Grounded Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5294--5306

  70. [75]

    Wang, S.; Shi, J.; Fu, Z.; He, X.; Liu, F.; Yang, C.; Zhou, Y.; Fei, Z.; Gong, J.; Fu, J.; et al. 2026. World Action Models : The Next Frontier in Embodied AI. arXiv preprint arXiv:2605.12090

  71. [76]

    Wang, Y.; Li, X.; Wang, W.; Zhang, J.; Li, Y.; Chen, Y.; Wang, X.; and Zhang, Z. 2025 b . Unified Vision-Language-Action Model. arXiv preprint arXiv:2506.19850

  72. [77]

    Xu, G.; Zhang, Q.; Zhou, J.; Zhu, X.; Shen, Y.; Yang, X.; and Xu, Y. 2026. Next Forcing : Causal World Modeling with Multi-Chunk Prediction. arXiv preprint arXiv:2606.11187

  73. [78]

    Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, H.; Li, J.; Lv, J.; Liu, J.; et al. 2026 a . GigaWorld-Policy : An Efficient Action-Centered World--Action Model. arXiv preprint arXiv:2603.17240

  74. [79]

    Ye, J.; Wang, F.; Gao, N.; Yu, J.; Zhu Yangkun ; Wang, B.; Zhang, J.; Jin, W.; Fu, Y.; Zheng, F.; Chen, Y.; and Pang, J. 2026 b . Spatially Guided Training for Vision-Language-Action Model. In International Conference on Learning Representations

  75. [80]

    L.; Zhu, C.; Xiang, J.; et al

    Ye, S.; Ge, Y.; Zheng, K.; Gao, S.; Yu, S.; Kurian, G.; Indupuru, S.; Tan, Y. L.; Zhu, C.; Xiang, J.; et al. 2026 c . World action models are zero-shot policies. arXiv preprint arXiv:2602.15922

  76. [81]

    Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026. Fast-WAM : Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666

  77. [82]

    Zhang, Z.; Li, H.; Dai, Y.; Zhu, Z.; Zhou, L.; Liu, C.; Wang, D.; Tay, F. E. H.; Chen, S.; Liu, Z.; Liu, Y.; Li, X.; and Zhou, P. 2026. From Spatial to Actions : Grounding Vision-Language-Action Model in Spatial Foundation Priors. In International Conference on Learning Representations

  78. [83]

    Zhen, H.; Qiu, X.; Chen, P.; Yang, J.; Yan, X.; Du, Y.; Hong, Y.; and Gan, C. 2024. 3D-VLA : A 3D Vision-Language-Action Generative World Model. In International Conference on Machine Learning, volume 235, 61229--61245

  79. [84]

    Zheng, J.; Li, J.; Wang, Z.; Liu, D.; Kang, X.; Feng, Y.; Zheng, Y.; Zou, J.; Chen, Y.; Zeng, J.; et al. 2025. X-VLA : Soft-prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model. arXiv preprint arXiv:2510.10274