REVIEW 2 major objections 5 minor 1 cited by
World action models can learn action-relevant 4D geometry in training and still deploy as the same lightweight video-action policies.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 14:22 UTC pith:Q47DYNFZ
load-bearing objection Solid training-only 4D recipe for WAMs: decayed current-frame reads plus action-aware temporal distillation give real, modest gains at flat inference cost; missing permanent-zero-g0 control is a fair gap, not a collapse. the 2 major comments →
Learning 4D Geometric Priors for Inference-Efficient World Action Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
MECo-WAM establishes that action-relevant 4D geometric priors can be transferred into shared video-action representations purely at training time—via multi-expert co-training, decayed 4D read-mask attention, and action-aware temporal geometric distillation—so the deployed policy retains the original lightweight observation-to-action interface while raising manipulation success.
What carries the argument
Decayed 4D read-mask attention plus action-aware temporal geometric distillation: temporary current-frame-only geometry reads that linearly decay to zero before deployment, and relational matching of within-frame and temporal geometry weighted toward action-coupled visual tokens, supervised by frozen encoder targets.
Load-bearing premise
Temporary current-frame geometry cues from a frozen encoder, faded out before deployment, leave enough action-relevant structure inside ordinary video-action tokens to improve real manipulation success.
What would settle it
Train matched video-action models with and without the 4D path on the same data; if held-out success, action error, and depth/pose probes show no gain once the decay schedule reaches zero on LIBERO, RoboTwin, and the real cube tasks, the transfer claim fails.
If this is right
- Training-only 4D supervision can raise task success without any geometry module, sensor, or decoder at inference.
- Appearance-oriented video co-training alone under-serves precise manipulation; action-conditioned geometric relations close that gap.
- Non-pretrained world action models can match or exceed strong pretrained baselines when geometry is transferred this way.
- Real-robot spatial grounding improves (fewer corrections, shorter completion) when geometry lives inside the deployed video-action features.
Where Pith is reading between the lines
- The same decayed-auxiliary pattern could carry other costly signals—depth, contact, force—into policies without changing the deployed stack.
- A frozen geometry teacher may be enough whenever the student only needs relative layout and its change, not absolute feature coordinates.
- Existing video-action co-training pipelines could bolt on a temporary 4D branch and drop it at export without redesigning serving systems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MECo-WAM is a multi-expert world action model that injects action-relevant 4D geometric priors into video-action representations only during training, then removes all auxiliary geometry at deployment so the inference graph and latency match the base Fast-WAM observation-to-action path. Training adds a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder; asymmetric expert masks and decayed 4D read-mask attention (temporary current-frame g0 reads that linearly decay to zero) are intended to transfer geometry without non-causal future-geometry shortcuts; action-aware temporal geometric distillation aligns within-frame relations and their temporal changes with action-conditioned token weights. Reported results are LIBERO 98.2% average success (without embodied pretraining), RoboTwin 2.0 92.62% average, and modest real-world gains on two ARX-R5 tabletop tasks, with ablations and depth/grasp probes supporting transfer into the deployed representation.
Significance. If the transfer claim holds, the paper offers a practical and timely recipe for geometry-aware WAMs: use a frozen 4D teacher and temporary read edges as a training-time representation constraint rather than an inference-time 4D decoder or sensor stack. That addresses a real deployment tension between geometric grounding and latency. Strengths include a clear training-only design, matched ablations isolating the 4D expert, decayed read, spatial/temporal losses and action-aware weights (Table 4), multi-benchmark evaluation (LIBERO suites, RoboTwin clean/random, real robot), and representation probes that freeze the video-action backbone. The absolute gains are modest but directionally consistent and obtained without increasing inference cost, which is a useful contribution for efficient embodied policies.
major comments (2)
- Methodology (Decayed 4D Read-Mask Attention, Eqs. 9–10, Fig. 3) and Table 4: the central claim is that temporary current-frame g0 reads transfer action-relevant geometry into a permanently 4D-free video-action pathway. Ablations show that an isolated 4D expert barely moves RoboTwin SR (91.83→91.87) while action MSE worsens, and that decayed read access is what lifts SR and lowers MSE. That pattern is consistent with useful transfer but also with early privileged g0 side-channel shaping of shared tokens before p4d reaches 0. A load-bearing control is missing: keep the full 4D expert and L4d while permanently zeroing g0→video/action edges (or freeze the video-action path after the decay window and continue 4D-only training). Without that isolation, the modest gains cannot be cleanly attributed to pure representation transfer versus temporary privileged geometry.
- Real-world evaluation (Table 3, Fig. 5): only two tabletop tasks on one ARX-R5 arm, with small trial counts and high run-to-run variance (e.g., Stack Cubes SR tied with Fast-WAM at 60%; Sort Cubes gains rest on limited trials). The paper’s claim of improved real-robot spatial grounding is directionally supported by lower corrections and shorter completion time, but the evidence is too thin to carry the deployment-efficiency narrative alongside LIBERO/RoboTwin. Either expand the real-world suite (more tasks, more trials, clearer success criteria) or temper the real-world claim to a preliminary study.
minor comments (5)
- Figure 1 and abstract/intro: RoboTwin is reported as 92.6% while Table 2 gives 92.62%; keep one consistent rounding convention.
- Implementation details: free parameters (p4d schedule, λvideo/λaction/λ4d, αgeo/αtem, η, keyframe set K, 4× action-to-video ratio) are listed but sensitivity is not reported; a short appendix on schedule/weight robustness would strengthen reproducibility.
- Notation: expert slots use f/a/g with subscripts 0,1,2,h and predictions fp/ap/gp; a compact token glossary near Eqs. (2)–(5) would help readers track the mixed-attention sequence.
- Related Work: contrast with X-WAM and WAM4D is clear on deployment cost; a one-sentence statement of what is *not* claimed (no new 4D reconstruction quality, no multi-view RGB-D futures) would further bound the contribution.
- Figure 4/6 probes are useful; state explicitly that the DPT head and grasp analysis use only deployed video-action tokens with no 4D expert at probe time (already implied, but worth one sentence).
Circularity Check
No significant circularity: empirical training method evaluated on external benchmarks; losses and success metrics are not equivalent by construction.
full rationale
MECo-WAM is a standard empirical multi-expert training recipe. Video/action objectives are conditional flow-matching (Eqs. 20–21); the 4D branch is supervised by relational and temporal-relation losses against a frozen third-party VGGT teacher (Eqs. 11–19, 22–23). Decayed read-mask attention (Eqs. 9–10, Fig. 3) is a training schedule that is set to zero at deployment; it does not redefine the reported success rates. Claimed gains are measured on external suites (LIBERO, RoboTwin 2.0) and real-robot rollouts, with ablations (Table 4) and representation probes (Figs. 4, 6) that exclude auxiliary 4D tokens. Building on the Fast-WAM observation-to-action interface is ordinary base-model reuse, not a load-bearing self-citation that forces the result. Skeptical concerns about residual leakage from temporary g0 reads are experimental-isolation questions, not reductions of predictions to inputs by definition. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or ansatz-smuggling chain is present.
Axiom & Free-Parameter Ledger
free parameters (4)
- p4d decay schedule (pstart=1.0, pend=0, Sdecay=half training)
- loss weights λvideo, λaction, λ4d, αgeo, αtem, η
- keyframe set K and action-to-video temporal ratio 4×
- 4D expert width d4d=512 (≈0.45 B params)
axioms (3)
- domain assumption Frozen VGGT relational features are a sufficiently accurate and action-relevant teacher of 4D geometry for manipulation.
- ad hoc to paper Asymmetric expert visibility plus current-frame-only temporary reads prevent non-causal future-geometry shortcuts into action generation.
- domain assumption Conditional flow-matching on video latents and action chunks is a valid joint training objective for WAMs.
invented entities (2)
-
decayed 4D read-mask attention
no independent evidence
-
action-aware temporal geometric distillation
no independent evidence
read the original abstract
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable action sequences. However, existing video-action co-training methods primarily optimize appearance-oriented video latents, which may insufficiently capture the temporally evolving geometry required for precise manipulation. We propose MECo-WAM, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph. During training, MECo-WAM combines video and action experts with a lightweight 4D expert supervised by relational targets from a frozen VGGT encoder. Asymmetric expert visibility prevents non-causal shortcuts from auxiliary geometry to action generation. To transfer geometric knowledge into the deployed video-action pathway, we introduce decayed 4D read-mask attention, which provides restricted current-frame geometric guidance early in training and progressively removes this dependency. We further propose action-aware temporal geometric distillation, which aligns within-frame geometric relations and their temporal evolution while emphasizing visual regions most relevant to robot actions. At deployment, all auxiliary 4D components are removed. Experiments on LIBERO (98.2%), RoboTwin 2.0 (92.6%), and challenging real-world manipulation tasks show that MECo-WAM improves manipulation performance without increasing inference cost.
Figures
Forward citations
Cited by 1 Pith paper
-
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts
ST-WAM adds DINOv3 semantic future prediction and history retrieval to a video-based world-action model, improving zero-shot robustness under visual distribution shifts.
Reference graph
Works this paper leans on
-
[2]
Pai, Jonas and Achenbach, Liam and Montesinos, Victoriano and Forrai, Benedek and Mees, Oier and Nava, Elvis , journal=
-
[3]
Ye, Angen and Wang, Boyuan and Ni, Chaojun and Huang, Guan and Zhao, Guosheng and Li, Hao and Li, Hengtao and Li, Jie and Lv, Jindi and Liu, Jingyu and others , journal=
-
[4]
Yuan, Tianyuan and Dong, Zibin and Liu, Yicheng and Zhao, Hang , journal=
-
[5]
Kim, Moo Jin and Gao, Yihuai and Lin, Tsung-Yi and Lin, Yen-Chen and Ge, Yunhao and Lam, Grace and Liang, Percy and Song, Shuran and Liu, Ming-Yu and Finn, Chelsea and others , journal=
-
[7]
Li, Ying and Wei, Xiaobao and Cao, Jiajun and Wang, Hao and Chi, Xiaowei and Bai, Chengyu and Sun, Qianpu and Li, Jiajun and Zhang, Xiaojie and Tang, Jian and others , journal=
-
[8]
Team, MotuBrain and Xiang, Chendong and Bao, Fan and Liu, Haitian and Tan, Hengkai and Bi, Hongzhe and Li, James and Liu, Jiabao and Pang, Jingrui and Jing, Kiro and others , journal=
-
[9]
Bi, Hongzhe and Tan, Hengkai and Xie, Shenghao and Wang, Zeyuan and Huang, Shuhe and Liu, Haitian and Zhao, Ruowen and Feng, Yao and Xiang, Chendong and Rong, Yinze and others , booktitle=
-
[11]
Xu, Gangwei and Zhang, Qihang and Zhou, Jiaming and Zhu, Xing and Shen, Yujun and Yang, Xin and Xu, Yinghao , journal=
-
[12]
Wan, Team and Wang, Ang and Ai, Baole and Wen, Bin and Mao, Chaojie and Xie, Chen-Wei and Chen, Di and Yu, Feiwu and Zhao, Haiming and Yang, Jianxiao and others , journal=
-
[13]
Wang, Jianyuan and Chen, Minghao and Karaev, Nikita and Vedaldi, Andrea and Rupprecht, Christian and Novotny, David , booktitle=
-
[14]
Mu, Yao and Chen, Tianxing and Chen, Zanxin and Peng, Shijia and Lan, Zhiqian and Gao, Zeyu and Liang, Zhixuan and Yu, Qiaojun and Zou, Yude and Xu, Mingkun and others , booktitle=
-
[15]
Liu, Bo and Zhu, Yifeng and Gao, Chongkai and Feng, Yihao and Liu, Qiang and Zhu, Yuke and Stone, Peter , journal=
-
[16]
Black, Kevin and Brown, Noah and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and Groom, Lachy and Hausman, Karol and Ichter, Brian and others , journal=
-
[17]
Intelligence, Physical and Black, Kevin and Brown, Noah and Darpinian, James and Dhabalia, Karan and Driess, Danny and Esmail, Adnan and Equi, Michael and Finn, Chelsea and Fusai, Niccolo and others , journal=
-
[18]
Qu, Delin and Song, Haoming and Chen, Qizhi and Yao, Yuanqi and Ye, Xinyi and Ding, Yan and Wang, Zhigang and Gu, JiaYuan and Zhao, Bin and Wang, Dong and others , journal=
-
[19]
Pertsch, Karl and Stachowicz, Kyle and Ichter, Brian and Driess, Danny and Nair, Suraj and Vuong, Quan and Mees, Oier and Finn, Chelsea and Levine, Sergey , journal=
-
[20]
2025 , volume =
Kim, Moo Jin and Pertsch, Karl and Karamcheti, Siddharth and Xiao, Ted and Balakrishna, Ashwin and Nair, Suraj and Rafailov, Rafael and Foster, Ethan P and Sanketi, Pannag R and Vuong, Quan and Kollar, Thomas and Burchfiel, Benjamin and Tedrake, Russ and Sadigh, Dorsa and Levine, Sergey and Liang, Percy and Finn, Chelsea , booktitle =. 2025 , volume =
2025
-
[21]
Kim, Moo Jin and Finn, Chelsea and Liang, Percy , journal=
-
[22]
Liang, Zhixuan and Li, Yizhuo and Yang, Tianshuo and Wu, Chengyue and Mao, Sitong and Pei, Liuao and Nian, Tian and Zhou, Shunbo and Yang, Xiaokang and Pang, Jiangmiao and others , journal=
-
[24]
Zheng, Jinliang and Li, Jianxiong and Wang, Zhihao and Liu, Dongxiu and Kang, Xirui and Feng, Yuchun and Zheng, Yinan and Zou, Jiayin and Chen, Yilun and Zeng, Jia and others , journal=
-
[25]
2025 , url=
Li, Peiyan and Chen, Yixiang and Wu, Hongtao and Ma, Xiao and Wu, Xiangnan and Huang, Yan and Wang, Liang and Kong, Tao and Tan, Tieniu , booktitle=. 2025 , url=
2025
-
[26]
2025 , editor=
Li, Xiaoqi and Heng, Liang and Liu, Jiaming and Shen, Yan and Gu, Chenyang and Liu, Zhuoyang and Chen, Hao and Han, Nuowei and Zhang, Renrui and Tang, Hao and Zhang, Shanghang and Dong, Hao , booktitle=. 2025 , editor=
2025
-
[27]
Fan, Xianzhe and Deng, Shengliang and Wu, Xiaoyang and Lu, Yuxiang and Li, Zhuoling and Yan, Mi and Zhang, Yujia and Zhang, Zhizheng and Wang, He and Zhao, Hengshuang , journal=
-
[28]
Ni, Chaojun and Chen, Cheng and Wang, Xiaofeng and Zhu, Zheng and Zheng, Wenzhao and Wang, Boyuan and Chen, Tianrun and Zhao, Guosheng and Li, Haoyun and Dong, Zhehao and Zhang, Qiang and Ye, Yun and Wang, Yang and Huang, Guan and Mei, Wenjun , booktitle=
-
[29]
2026 , url=
Li, Fuhao and Song, Wenxuan and Zhao, Han and Wang, Jingbo and Ding, Pengxiang and Wang, Donglin and Zeng, Long and Li, Haoang , booktitle=. 2026 , url=
2026
-
[30]
Conference on Robot Learning , year=
Generalist Robot Manipulation beyond Action Labeled Data , author=. Conference on Robot Learning , year=
-
[31]
Sun, Lin and Xie, Bin and Liu, Yingfei and Shi, Hao and Wang, Tiancai and Cao, Jiale , journal=
-
[32]
Zhang, Zhengshen and Li, Hao and Dai, Yalun and Zhu, Zhengbang and Zhou, Lei and Liu, Chenchen and Wang, Dong and Tay, Francis E. H. and Chen, Sijin and Liu, Ziwei and Liu, Yuxiao and Li, Xinghang and Zhou, Pan , booktitle=
-
[33]
2024 , volume=
Zhen, Haoyu and Qiu, Xiaowen and Chen, Peihao and Yang, Jincheng and Yan, Xin and Du, Yilun and Hong, Yining and Gan, Chuang , booktitle=. 2024 , volume=
2024
-
[34]
Qian, Jingjing and Han, Boyao and Shi, Chen and Xiao, Lei and Yang, Long and Shi, Shaoshuai and Jiang, Li , booktitle=
-
[35]
International Conference on Learning Representations , year=
Spatially Guided Training for Vision-Language-Action Model , author=. International Conference on Learning Representations , year=
-
[36]
Su, Taiyi and Zhu, Jian and Wang, Tianjian and He, Youzhang and Huang, Zitai and Zhang, Jianjun and Ma, Chong and Wang, Hanyang and Zhang, Tianjiao and Yin, Munan and others , journal=
-
[37]
Wang, Siyin and Shi, Junhao and Fu, Zhaoyang and He, Xinzhe and Liu, Feihong and Yang, Chenchen and Zhou, Yikang and Fei, Zhaoye and Gong, Jingjing and Fu, Jinlan and others , journal=
-
[38]
IEEE/CVF International Conference on Computer Vision , pages=
Vision Transformers for Dense Prediction , author=. IEEE/CVF International Conference on Computer Vision , pages=
-
[40]
2025 , volume =
Hu, Yucheng and Guo, Yanjiang and Wang, Pengchao and Chen, Xiaoyu and Wang, Yen-Jen and Zhang, Jianke and Sreenath, Koushil and Lu, Chaochao and Chen, Jianyu , booktitle =. 2025 , volume =
2025
-
[41]
Ma, Teli and Zheng, Jia and Wang, Zifan and Jiang, Chunli and Cui, Andy and Liang, Junwei and Yang, Shuo , journal=
-
[42]
2025 , volume=
Chen, Sili and Guo, Hengkai and Zhu, Shengnan and Zhang, Feihu and Huang, Zilong and Feng, Jiashi and Kang, Bingyi , booktitle=. 2025 , volume=
2025
-
[43]
Bi, H.; Tan, H.; Xie, S.; Wang, Z.; Huang, S.; Liu, H.; Zhao, R.; Feng, Y.; Xiang, C.; Rong, Y.; et al. 2026. Motus : A unified latent action world model. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 35101--35113
2026
-
[44]
Black, K.; Brown, N.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; Groom, L.; Hausman, K.; Ichter, B.; et al. 2024. _0 : A Vision-Language-Action Flow Model for General Robot Control. arXiv preprint arXiv:2410.24164
Pith/arXiv arXiv 2024
-
[45]
Chen, S.; Guo, H.; Zhu, S.; Zhang, F.; Huang, Z.; Feng, J.; and Kang, B. 2025. Video Depth Anything : Consistent Depth Estimation for Super-Long Videos. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 22831--22840
2025
-
[46]
Fan, X.; Deng, S.; Wu, X.; Lu, Y.; Li, Z.; Yan, M.; Zhang, Y.; Zhang, Z.; Wang, H.; and Zhao, H. 2026. Any3D-VLA : Enhancing VLA Robustness via Diverse Point Clouds. arXiv preprint arXiv:2602.00807
Pith/arXiv arXiv 2026
-
[47]
Guo, J.; Li, Q.; Li, P.; Chen, Z.; Sun, N.; Su, Y.; Wang, H.; Zhang, Y.; Li, X.; and Liu, H. 2026. Unified 4D world action modeling from video priors with asynchronous denoising. arXiv preprint arXiv:2604.26694
Pith/arXiv arXiv 2026
-
[48]
Hu, Y.; Guo, Y.; Wang, P.; Chen, X.; Wang, Y.-J.; Zhang, J.; Sreenath, K.; Lu, C.; and Chen, J. 2025. Video Prediction Policy : A Generalist Robot Policy with Predictive Visual Representations. In International Conference on Machine Learning, volume 267, 24328--24346
2025
-
[49]
Intelligence, P.; Black, K.; Brown, N.; Darpinian, J.; Dhabalia, K.; Driess, D.; Esmail, A.; Equi, M.; Finn, C.; Fusai, N.; et al. 2025. _ 0.5 : A Vision-Language-Action Model with Open-World Generalization. arXiv preprint arXiv:2504.16054
Pith/arXiv arXiv 2025
-
[50]
Kim, M. J.; Finn, C.; and Liang, P. 2025. Fine-Tuning Vision-Language-Action Models : Optimizing Speed and Success. arXiv preprint arXiv:2502.19645
Pith/arXiv arXiv 2025
-
[51]
Kim, M. J.; Gao, Y.; Lin, T.-Y.; Lin, Y.-C.; Ge, Y.; Lam, G.; Liang, P.; Song, S.; Liu, M.-Y.; Finn, C.; et al. 2026. Cosmos Policy : Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163
Pith/arXiv arXiv 2026
-
[52]
J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E
Kim, M. J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E. P.; Sanketi, P. R.; Vuong, Q.; Kollar, T.; Burchfiel, B.; Tedrake, R.; Sadigh, D.; Levine, S.; Liang, P.; and Finn, C. 2025. OpenVLA : An Open-Source Vision-Language-Action Model. In Conference on Robot Learning, volume 270, 2679--2713
2025
-
[53]
Li, F.; Song, W.; Zhao, H.; Wang, J.; Ding, P.; Wang, D.; Zeng, L.; and Li, H. 2026 a . Spatial Forcing : Implicit Spatial Representation Alignment for Vision-Language-Action Model. In International Conference on Learning Representations
2026
-
[54]
Li, L.; Zhang, Q.; Luo, Y.; Yang, S.; Wang, R.; Han, F.; Yu, M.; Gao, Z.; Xue, N.; Zhu, X.; et al. 2026 b . Causal World Modeling for Robot Control. arXiv preprint arXiv:2601.21998
Pith/arXiv arXiv 2026
-
[55]
Li, P.; Chen, Y.; Wu, H.; Ma, X.; Wu, X.; Huang, Y.; Wang, L.; Kong, T.; and Tan, T. 2025 a . BridgeVLA : Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models. In Conference on Neural Information Processing Systems
2025
-
[56]
Li, X.; Heng, L.; Liu, J.; Shen, Y.; Gu, C.; Liu, Z.; Chen, H.; Han, N.; Zhang, R.; Tang, H.; Zhang, S.; and Dong, H. 2025 b . 3DS-VLA : A 3D Spatial-Aware Vision Language Action Model for Robust Multi-Task Manipulation. In Lim, J.; Song, S.; and Park, H.-W., eds., Conference on Robot Learning, volume 305 of Proceedings of Machine Learning Research, 2344-...
2025
-
[57]
Li, Y.; Wei, X.; Cao, J.; Wang, H.; Chi, X.; Bai, C.; Sun, Q.; Li, J.; Zhang, X.; Tang, J.; et al. 2026 c . WAM4D : Fast 4D World Action Model via Spatial Register Tokens. arXiv preprint arXiv:2606.14048
Pith/arXiv arXiv 2026
-
[58]
Liang, Z.; Li, Y.; Yang, T.; Wu, C.; Mao, S.; Pei, L.; Nian, T.; Zhou, S.; Yang, X.; Pang, J.; et al. 2025. Discrete Diffusion VLA : Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies. arXiv preprint arXiv:2508.20072
Pith/arXiv arXiv 2025
-
[59]
Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; and Stone, P. 2023. LIBERO : Benchmarking knowledge transfer for lifelong robot learning. Advances in Neural Information Processing Systems, 36: 44776--44791
2023
-
[60]
Ma, T.; Zheng, J.; Wang, Z.; Jiang, C.; Cui, A.; Liang, J.; and Yang, S. 2026. DiT4DiT : Jointly Modeling Video Dynamics and Actions for Generalizable Robot Control. arXiv preprint arXiv:2603.10448
arXiv 2026
-
[61]
Mu, Y.; Chen, T.; Chen, Z.; Peng, S.; Lan, Z.; Gao, Z.; Liang, Z.; Yu, Q.; Zou, Y.; Xu, M.; et al. 2025. RoboTwin : Dual-arm robot benchmark with generative digital twins. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27649--27660
2025
-
[62]
Ni, C.; Chen, C.; Wang, X.; Zhu, Z.; Zheng, W.; Wang, B.; Chen, T.; Zhao, G.; Li, H.; Dong, Z.; Zhang, Q.; Ye, Y.; Wang, Y.; Huang, G.; and Mei, W. 2026. SwiftVLA : Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13474--13485
2026
-
[63]
Pai, J.; Achenbach, L.; Montesinos, V.; Forrai, B.; Mees, O.; and Nava, E. 2025. mimic-video : Video-action models for generalizable robot control beyond vlas. arXiv preprint arXiv:2512.15692
Pith/arXiv arXiv 2025
-
[64]
Pertsch, K.; Stachowicz, K.; Ichter, B.; Driess, D.; Nair, S.; Vuong, Q.; Mees, O.; Finn, C.; and Levine, S. 2025. FAST : Efficient action tokenization for vision-language-action models. arXiv preprint arXiv:2501.09747
Pith/arXiv arXiv 2025
-
[65]
Qian, J.; Han, B.; Shi, C.; Xiao, L.; Yang, L.; Shi, S.; and Jiang, L. 2026. GeoPredict : Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 13529--13539
2026
-
[66]
Qu, D.; Song, H.; Chen, Q.; Yao, Y.; Ye, X.; Ding, Y.; Wang, Z.; Gu, J.; Zhao, B.; Wang, D.; et al. 2025. SpatialVLA : Exploring spatial representations for visual-language-action model. arXiv preprint arXiv:2501.15830
Pith/arXiv arXiv 2025
-
[67]
Ranftl, R.; Bochkovskiy, A.; and Koltun, V. 2021. Vision Transformers for Dense Prediction. In IEEE/CVF International Conference on Computer Vision, 12179--12188
2021
-
[68]
Spiridonov, A.; Zaech, J.-N.; Nikolov, N.; Van Gool, L.; and Paudel, D. P. 2025. Generalist Robot Manipulation beyond Action Labeled Data. In Conference on Robot Learning
2025
-
[69]
Su, T.; Zhu, J.; Li, Y.; Ma, C.; Zhang, J.; Huang, Z.; Wang, H.; and Xu, Y. 2025. Towards high-consistency embodied world model with multi-view trajectory videos. arXiv preprint arXiv:2511.12882
arXiv 2025
-
[70]
Su, T.; Zhu, J.; Wang, T.; He, Y.; Huang, Z.; Zhang, J.; Ma, C.; Wang, H.; Zhang, T.; Yin, M.; et al. 2026. DeMaVLA : A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation. arXiv preprint arXiv:2605.31286
Pith/arXiv arXiv 2026
-
[71]
Sun, L.; Xie, B.; Liu, Y.; Shi, H.; Wang, T.; and Cao, J. 2025. GeoVLA : Empowering 3D Representations in Vision-Language-Action Models. arXiv preprint arXiv:2508.09071
Pith/arXiv arXiv 2025
-
[72]
Team, M.; Xiang, C.; Bao, F.; Liu, H.; Tan, H.; Bi, H.; Li, J.; Liu, J.; Pang, J.; Jing, K.; et al. 2026. MotuBrain : An advanced world action model for robot control. arXiv preprint arXiv:2604.27792
Pith/arXiv arXiv 2026
-
[73]
Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.-W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. 2025. Wan : Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314
Pith/arXiv arXiv 2025
-
[74]
Wang, J.; Chen, M.; Karaev, N.; Vedaldi, A.; Rupprecht, C.; and Novotny, D. 2025 a . VGGT : Visual Geometry Grounded Transformer. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, 5294--5306
2025
-
[75]
Wang, S.; Shi, J.; Fu, Z.; He, X.; Liu, F.; Yang, C.; Zhou, Y.; Fei, Z.; Gong, J.; Fu, J.; et al. 2026. World Action Models : The Next Frontier in Embodied AI. arXiv preprint arXiv:2605.12090
Pith/arXiv arXiv 2026
-
[76]
Wang, Y.; Li, X.; Wang, W.; Zhang, J.; Li, Y.; Chen, Y.; Wang, X.; and Zhang, Z. 2025 b . Unified Vision-Language-Action Model. arXiv preprint arXiv:2506.19850
Pith/arXiv arXiv 2025
-
[77]
Xu, G.; Zhang, Q.; Zhou, J.; Zhu, X.; Shen, Y.; Yang, X.; and Xu, Y. 2026. Next Forcing : Causal World Modeling with Multi-Chunk Prediction. arXiv preprint arXiv:2606.11187
Pith/arXiv arXiv 2026
-
[78]
Ye, A.; Wang, B.; Ni, C.; Huang, G.; Zhao, G.; Li, H.; Li, H.; Li, J.; Lv, J.; Liu, J.; et al. 2026 a . GigaWorld-Policy : An Efficient Action-Centered World--Action Model. arXiv preprint arXiv:2603.17240
arXiv 2026
-
[79]
Ye, J.; Wang, F.; Gao, N.; Yu, J.; Zhu Yangkun ; Wang, B.; Zhang, J.; Jin, W.; Fu, Y.; Zheng, F.; Chen, Y.; and Pang, J. 2026 b . Spatially Guided Training for Vision-Language-Action Model. In International Conference on Learning Representations
2026
-
[80]
Ye, S.; Ge, Y.; Zheng, K.; Gao, S.; Yu, S.; Kurian, G.; Indupuru, S.; Tan, Y. L.; Zhu, C.; Xiang, J.; et al. 2026 c . World action models are zero-shot policies. arXiv preprint arXiv:2602.15922
Pith/arXiv arXiv 2026
-
[81]
Yuan, T.; Dong, Z.; Liu, Y.; and Zhao, H. 2026. Fast-WAM : Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666
Pith/arXiv arXiv 2026
-
[82]
Zhang, Z.; Li, H.; Dai, Y.; Zhu, Z.; Zhou, L.; Liu, C.; Wang, D.; Tay, F. E. H.; Chen, S.; Liu, Z.; Liu, Y.; Li, X.; and Zhou, P. 2026. From Spatial to Actions : Grounding Vision-Language-Action Model in Spatial Foundation Priors. In International Conference on Learning Representations
2026
-
[83]
Zhen, H.; Qiu, X.; Chen, P.; Yang, J.; Yan, X.; Du, Y.; Hong, Y.; and Gan, C. 2024. 3D-VLA : A 3D Vision-Language-Action Generative World Model. In International Conference on Machine Learning, volume 235, 61229--61245
2024
-
[84]
Zheng, J.; Li, J.; Wang, Z.; Liu, D.; Kang, X.; Feng, Y.; Zheng, Y.; Zou, J.; Chen, Y.; Zeng, J.; et al. 2025. X-VLA : Soft-prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model. arXiv preprint arXiv:2510.10274
Pith/arXiv arXiv 2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.