REVIEW 5 major objections 7 minor 4 cited by
Mobile manipulation succeeds when world-action models align temporal grain, action subspaces, and train-test conditions via latent actions, dual-level transformers, and dream forcing.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 09:20 UTC pith:R526GPKW
load-bearing objection Solid mobile-manipulation WAM systems paper: the three-alignment package is real engineering, results are multi-benchmark and ablated, but the latent-action bridge is still an unprobed assumption. the 5 major comments →
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Mobile manipulation is an alignment problem at three levels—temporal granularity, action space, and train-test consistency. ABot-M0.5 solves it by factorizing generation into video latents, then frame-level latent actions, then executable controls; by disentangling modality streams and heterogeneous action subspaces (base vs arm) inside a dual-level Mixture-of-Transformers; and by progressively training inverse dynamics on self-dreamed futures so the conditioning at training matches autoregressive inference. The result is state-of-the-art long-horizon task success and fine-grained control accuracy across simulation and real robots.
What carries the argument
The three-stage cascade zt+1 → mt → at (video latent to frame-level latent action to robot action), realized by a dual-level Mixture-of-Transformers that separates modalities and action subspaces, plus Dream Forcing that conditions action learning on model-predicted rather than ground-truth futures.
Load-bearing premise
Frame-level latent actions taken from visual frame pairs form a control-relevant, embodiment-agnostic bridge that recovers the fine contact dynamics coarse video chunks lose; if they mostly track camera or background change, the cascade fails.
What would settle it
Replace the latent-action stream with pure noise or with features from an untrained encoder, retrain under the same dual-level and dream-forcing schedule, and re-evaluate RoboCasa365 composite tasks plus real peg-insertion success; if rates stay near the full model, the bridging claim is false.
If this is right
- Long-horizon mobile success rises when inverse dynamics is trained on self-predicted futures rather than ground-truth futures.
- Separating mobility and manipulation heads reduces gradient conflict and lifts composite-task rates.
- Visual-only latent actions let motion priors transfer across robot embodiments without kinematic labels.
- The same aligned cascade improves pure tabletop and bimanual benchmarks, not only mobile ones.
- With aligned pretraining, real-robot policies can reach high success from only tens of demonstrations.
Where Pith is reading between the lines
- Dream forcing is a general recipe for closing teacher-forcing gaps in any world-model controller, not only robots.
- Dual-level action disentanglement should help any multi-frequency control stack (e.g., whole-body locomotion versus fingers).
- If latent actions truly encode contact onset, similar bridges could stabilize other autoregressive policies that currently suffer exposure bias.
- Future WAM scaling may be limited more by these three alignments than by raw parameter count.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. ABot-M0.5 is a World Action Model for language-conditioned mobile manipulation that jointly predicts future video latents, frame-level latent actions, and executable controls. The paper argues that existing WAMs are misaligned with mobile manipulation at three levels—temporal granularity (coarse video chunks vs. fine control), action structure (entangled mobility vs. manipulation), and train–test consistency (GT-conditioned inverse dynamics vs. autoregressive self-predicted futures)—and addresses them with (i) intermediate latent actions m_t extracted by a frozen ALAM-style encoder (Eq. 8; §3.2, §4.3), (ii) a dual-level Mixture-of-Transformers that separates modality streams and mobility/manipulation subspaces (§3.3), and (iii) Dream Forcing, a two-phase training scheme that conditions action prediction on self-dreamed latents (§3.4, §4.4). Training proceeds via world-model pretraining on multi-embodiment corpora, latent-action pretraining with algebraic consistency losses, then progressive SFT (teacher-forced then dream-forced). Empirically, the model reports strong results on RoboCasa365 (pretrain and target), RoboTwin 2.0, LIBERO/LIBERO-Plus, and real-robot arm tasks, with component ablations for latent-action structure, action-decoupled MoT, and Dream Forcing.
Significance. If the three-alignment thesis holds, the paper offers a concrete architectural and training recipe for extending world-action models from stationary to mobile manipulation—an important and under-served setting. The intermediate latent-action cascade, dual-level MoT, and Dream Forcing are clearly specified and, in principle, transferable. Strengths include multi-benchmark evaluation (RoboCasa365, RoboTwin clean/hard, LIBERO, LIBERO-Plus zero-shot, real robot), structured ablations (Table 7 latent-action variants; Table 8 Dream Forcing vs. continued teacher forcing; action-decoupled MoT on a composite subset), progressive train–test alignment rather than pure capacity scaling, and promised code. The contribution is primarily empirical/systems rather than theoretical; its lasting value depends on whether the reported gains are causally tied to the claimed alignments (especially what m_t encodes) and whether the mobile claims transfer beyond simulation.
major comments (5)
- §3.2, Eq. (8) and §4.3: The central temporal-granularity claim is that frozen ALAM latents m_t recover fine contact/grasp dynamics lost in coarse z_{t+1}. The paper never validates this. Algebraic losses L_add and L_rev enforce only visual-transition consistency; there is no correlation of m_t with contact onset, grasp closure, force proxies, or end-effector–object relative motion, nor a control experiment that replaces m_t with a capacity-matched intermediate stream (e.g., random or optical-flow tokens). Table 7 shows that a 3-stage cascade helps (94.0% vs. 87.6% baseline on RoboTwin Clean), but that is consistent with capacity/conditioning structure rather than contact-relevant content. Either provide a direct semantic analysis of m_t or substantially tone down the claim that the cascade specifically restores contact dynamics.
- Table 2 (§5.2): The paper reports ABot-M0.5(+Condensed Memory) as setting a new record (46.6% average) while stating that condensed memory “will be presented in our future work.” Publishing SOTA numbers for a mechanism that is neither defined nor ablated in this manuscript is not acceptable for a primary result table. Either fully specify and ablate condensed memory here, or remove those rows and base all SOTA claims on the fully described ABot-M0.5 (40.4%).
- Table 2 Composite-Unseen and Abstract/§5.2 SOTA framing: On Composite-Unseen, ABot-M0.5 scores 2.7% (7.9% with condensed memory), below Qwen-RobotManip (14.9%) and Qwen-RobotManip-Context (11.2%). The abstract and §5.2 claim state-of-the-art long-horizon mobile manipulation success without qualifying this gap. Long-horizon composite generalization is precisely where the three-alignment thesis should matter most. Please reframe SOTA claims by category (Atomic-Seen / Composite-Seen / Composite-Unseen) and discuss why Composite-Unseen remains weak.
- §5.5 Real-world experiments: All real-robot results use a fixed-base Agilex Piper 6-DoF arm; there is no mobile base, navigation phase, or viewpoint change induced by locomotion. The paper’s central problem is mobile manipulation (Abstract; §1–2; RoboCasa365). Real-world transfer therefore supports fine-grained and multi-stage arm control, not the mobility–manipulation alignment thesis. Either add real mobile-base experiments or clearly restrict the real-world claim to stationary manipulation transfer and avoid implying physical validation of mobile alignment.
- §5.4 Action-decoupled MoT: The only quantitative support is a “selected subset” of RoboCasa365 Composite-Seen (0.48 vs. 0.34) plus a training-curve figure. No full-benchmark numbers, no ablation of joint attention vs. fully separate towers, and no analysis of whether mobility/manipulation specialization actually occurs (e.g., subspace-wise error or attention patterns). Given that action-space alignment is one of the three load-bearing claims, please report full Composite-Seen/Unseen numbers with and without action-level decoupling, and ideally a diagnostic that the two heads specialize.
minor comments (7)
- Table 1 / §3.1: Notation switches between m_t, m_{0,t}, and M for latent actions; standardize early and keep one symbol.
- Eqs. (11)–(13) and (25)–(26): λ_move / λ_manip and λ_z / λ_m / λ_a are free parameters but never listed with values or sensitivity. A short hyperparameter table would help reproducibility.
- Figure 1 and Abstract: “finegrained” should be “fine-grained”; several figure captions mix “Dreamed” / “dreamed” inconsistently.
- §4.2 semantic slot allocation (2 third-person + 2 wrist): state how often slots are empty/padded on each pretraining corpus and whether padding rate correlates with downstream multi-view performance.
- Table 8: Dream Forcing gains are modest (+3.01% absolute over the 50k warm-start; +1.66% over continued teacher forcing at 60k). The text should avoid language that implies a large qualitative jump without reporting variance over seeds.
- §5.1 Implementation details are thin (optimizer, learning rates, chunk size H, number of dream denoising steps). Add a concise appendix table.
- Related work: several concurrent WAMs (Motus, Fast-WAM, Lingbot-VA, ImageWAM) are compared in tables but discussed only lightly in the introduction; a short related-work paragraph contrasting teacher forcing vs. diffusion forcing vs. Dream Forcing would help placement.
Circularity Check
Empirical systems paper: SOTA claims are measured on external benchmarks, not recovered by construction from fitted constants or self-definitional identities.
full rationale
ABot-M0.5 is an architecture-and-training paper whose central claims are empirical success rates on RoboCasa365, RoboTwin 2.0, LIBERO/LIBERO-Plus, and real-robot tasks against external VLA/WAM baselines. Intermediate latent actions are extracted by a frozen ALAM-style encoder pretrained with algebraic visual-transition losses (Ladd, Lrev, reconstruction/VQ); those losses do not include the downstream task metrics, so mt is not defined in terms of the reported success rates. Dual-level MoT and Dream Forcing are design choices whose benefit is measured by ablations (Tables 7–8, Fig. 10) and held-out task success, not forced by fitting a parameter that is then renamed a prediction. Self-citations to ALAM and prior ABot work supply reusable components (encoder, data pipeline) but do not import a uniqueness theorem that forbids alternatives or make the SOTA claim tautological. No equation equates a claimed prediction to its own fitted input. The skeptic concern that mt may encode generic visual change rather than contact dynamics is an unvalidated-assumption / correctness issue, not circularity. Score 0 is therefore the honest finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- λ_z, λ_m, λ_a and λ_move, λ_manip action subspace weights
- latent-action conditioning dropout p_drop
- shared denoising timestep τ and few-step dream sampling schedule
- four canonical multi-view semantic slots (2 third-person + 2 wrist)
- ALAM pretraining loss weights λ_vq, λ_rec, λ_perc, λ_add, λ_rev
axioms (5)
- domain assumption Conditional flow matching on video, latent-action, and action latents is a valid generative objective for robotic control.
- domain assumption Visual frame-pair transitions encode embodiment-agnostic motion intents that transfer across robot morphologies.
- domain assumption At deployment, historical chunks can be grounded in real GT observations so only the latest future chunk needs dreaming.
- ad hoc to paper Base mobility and arm manipulation are sufficiently heterogeneous that separate FFN/heads reduce gradient interference without destroying coordination via joint attention.
- standard math Standard math of CFM velocity regression and causal attention masks.
invented entities (3)
-
Frame-level intermediate latent action m_t as bridging space
no independent evidence
-
Dual-level Mixture-of-Transformers (modality + mobility/manipulation)
no independent evidence
-
Dream Forcing training paradigm (two-phase forward on self-dreamed latents)
no independent evidence
read the original abstract
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.
Forward citations
Cited by 4 Pith papers
-
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
A frozen world-action model can be steered to new tasks by adapting a lightweight memory from unlabeled human video via test-time training.
-
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
A meta-trained test-time memory lets frozen world-action models absorb unlabeled human videos and outperform in-context video conditioning on real multi-embodiment manipulation.
-
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI
A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.
-
From Foundation to Application: Improving VLA Models in Practice
LingBot-VLA 2.0 combines 60k hours of multi-embodiment pretraining data, an expanded whole-body action space, and dual-query distillation from depth and video teachers to improve VLA performance on GM-100 and long-hor...
Reference graph
Works this paper leans on
-
[1]
AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025
Pith/arXiv arXiv 2025
-
[2]
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks.arXiv preprint arXiv:1506.03099, 2015
Pith/arXiv arXiv 2015
-
[3]
Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025
Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025
Pith/arXiv arXiv 2025
-
[4]
arXiv preprint arXiv:2410.24164, 2024
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.π0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024
Pith/arXiv arXiv 2024
-
[5]
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023
Pith/arXiv arXiv 2023
-
[6]
Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2023
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2023
Pith/arXiv arXiv 2023
-
[7]
Univla: Learning to act anywhere with task-centric latent actions.RSS, 2025
Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. Univla: Learning to act anywhere with task-centric latent actions.RSS, 2025
2025
-
[8]
Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025
Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, et al. Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025
Pith/arXiv arXiv 2025
-
[9]
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025
Pith/arXiv arXiv 2025
-
[10]
Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
2025
-
[11]
Open x-embodiment: Robotic learning datasets and rt-x models.arXiv preprint arXiv:2310.08864, 2023
Open X-Embodiment Collaboration. Open x-embodiment: Robotic learning datasets and rt-x models.arXiv preprint arXiv:2310.08864, 2023
Pith/arXiv arXiv 2023
-
[12]
InternVLA-M1 Contributors. Internvla-m1: A spatially guided vision-language-action framework for generalist robot policy.arXiv preprint arXiv:2510.13778, 2025
Pith/arXiv arXiv 2025
-
[13]
Justin Cui, Jie Wu, Ming Li, Tao Yang, Xiaojie Li, Rui Wang, Andrew Bai, Yuanhao Ban, and Cho-Jui Hsieh. Self-forcing++: Towards minute-scale high-quality video generation.arXiv preprint arXiv:2510.02283, 2025
Pith/arXiv arXiv 2025
-
[14]
Fu, Stefano Ermon, Atri Rudra, and Christopher Re
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: Fast and memory-efficient exact attention with io-awareness.arXiv preprint arXiv:2205.14135, 2022
Pith/arXiv arXiv 2022
-
[15]
Robonet: Large-scale multi-robot learning.arXiv preprint arXiv:1910.11215, 2020
Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. Robonet: Large-scale multi-robot learning.arXiv preprint arXiv:1910.11215, 2020
Pith/arXiv arXiv 1910
-
[16]
Libero-plus: In-depth robustness analysis of vision-language-action models
Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. Libero-plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626, 2025
Pith/arXiv arXiv 2025
-
[17]
Yao Feng, Hengkai Tan, Xinyi Mao, Chendong Xiang, Guodong Liu, Shuhe Huang, Hang Su, and Jun Zhu. Vidar: Embodied video diffusion model for generalist manipulation.arXiv preprint arXiv:2507.12898, 2025
Pith/arXiv arXiv 2025
-
[18]
Galaxea g0.5 technical report
Galaxea Team. Galaxea g0.5 technical report. 2026. URLhttps://opengalaxea.github.io/G05/
2026
-
[19]
Xinyu Guo, Bin Xie, Wei Chai, Xianchi Deng, Tiancai Wang, Zhengxing Wu, and Xingyu Chen. Priorvla: Prior-preserving adaptation for vision-language-action models.arXiv preprint arXiv:2605.10925, 2026. 29
Pith/arXiv arXiv 2026
-
[20]
World models.arXiv preprint arXiv:1803.10122, 2018
David Ha and Jürgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018
Pith/arXiv arXiv 2018
-
[21]
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2024
Pith/arXiv arXiv 2024
-
[22]
Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026
Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026
2026
-
[23]
Chia-Yu Hung, Qi Sun, Pengfei Hong, Amir Zadeh, Chuan Li, U Tan, Navonil Majumder, Soujanya Poria, et al. Nora: A small open-sourced generalist vision language action model for embodied tasks.arXiv preprint arXiv:2504.19854, 2025
Pith/arXiv arXiv 2025
-
[24]
ABot-Claw: A foundation for persistent, cooperative, and self-evolving robotic agents
Dongjie Huo, Haoyun Liu, Guoqing Liu, Dekang Qi, Zhiming Sun, Maoguo Gao, Jianxin He, Yandan Yang, Xinyuan Chang, Feng Xiong, et al. ABot-Claw: A foundation for persistent, cooperative, and self-evolving robotic agents. arXiv preprint arXiv:2604.10096, 2026
Pith/arXiv arXiv 2026
-
[25]
Oxe-auge: A large-scale robot augmentation of oxe for scaling cross-embodiment policy learning
Guanhua Ji, Harsha Polavaram, Lawrence Yunliang Chen, Sandeep Bajamahal, Zehan Ma, Simeon Adebola, Chenfeng Xu, and Ken Goldberg. Oxe-auge: A large-scale robot augmentation of oxe for scaling cross-embodiment policy learning. arXiv preprint arXiv:2512.13100, 2025
arXiv 2025
-
[26]
Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025
Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025
Pith/arXiv arXiv 2025
-
[27]
Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2025
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2025
Pith/arXiv arXiv 2025
-
[28]
Rldx-1 technical report.arXiv preprint arXiv:2605.03269, 2026
Dongyoung Kim, Huiwon Jang, Myungkyu Koo, Suhyeok Jang, Taeyoung Kim, Beomjun Kim, Byungjun Yoon, Changsung Jang, Daewon Choi, Dongsu Han, et al. Rldx-1 technical report.arXiv preprint arXiv:2605.03269, 2026
Pith/arXiv arXiv 2026
-
[29]
Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Pith/arXiv arXiv 2024
-
[30]
Fine-tuning vision-language-action models: Optimizing speed and success
Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success. RSS, 2025
2025
-
[31]
Cosmos policy: Fine-tuning video models for visuomotor control and planning
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163, 2026
Pith/arXiv arXiv 2026
-
[32]
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martin-Martin, Chen Wang, Gabrael Levine, Wensi Ai, Benjamin Martinez, et al. Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation.arXiv preprint arXiv:2403.09227, 2024
Pith/arXiv arXiv 2024
-
[33]
Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026
Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, et al. Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026
Pith/arXiv arXiv 2026
-
[34]
Anthony Liang, Pavel Czempin, Matthew Hong, Yutai Zhou, Erdem Biyik, and Stephen Tu. Clam: Continuous latent action models for robot learning from unlabeled demonstrations.arXiv preprint arXiv:2505.04999, 2025
Pith/arXiv arXiv 2025
-
[35]
Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu, Sitong Mao, Tian Nian, Liuao Pei, Shunbo Zhou, Xiaokang Yang, Jiangmiao Pang, et al. Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies. arXiv preprint arXiv:2508.20072, 2025
Pith/arXiv arXiv 2025
-
[36]
Holobrain-0 technical report.arXiv preprint arXiv:2602.12062, 2026
Xuewu Lin, Tianwei Lin, Yun Du, Hongyu Xie, Yiwei Jin, Jiawei Li, Shijie Wu, Qingze Wang, Mengdi Li, Mengao Zhao, et al. Holobrain-0 technical report.arXiv preprint arXiv:2602.12062, 2026
arXiv 2026
-
[37]
Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2023
Pith/arXiv arXiv 2023
-
[38]
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023. 30
Pith/arXiv arXiv 2023
-
[39]
Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, and Shijian Lu. Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025
Pith/arXiv arXiv 2025
- [40]
-
[41]
Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Haiweng Xu, Chaoyi Xu, Ziheng Xi, Yuhui Fu, and Zongqing Lu. Being-h0. 7: A latent world-action model from egocentric videos.arXiv preprint arXiv:2605.00078, 2026
Pith/arXiv arXiv 2026
-
[42]
Coral: Scalable multi-task robot learning via lora experts
Yuankai Luo, Woping Chen, Tong Liang, and Zhenguo Li. Coral: Scalable multi-task robot learning via lora experts. arXiv preprint arXiv:2603.09298, 2026
arXiv 2026
-
[43]
Qi Lv, Weijie Kong, Hao Li, Jia Zeng, Zherui Qiu, Delin Qu, Haoming Song, Qizhi Chen, Xiang Deng, and Jiangmiao Pang. F1: A vision-language-action model bridging understanding and generation to actions.arXiv preprint arXiv:2509.06951, 2025
Pith/arXiv arXiv 2025
-
[44]
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied ai.IEEE Transactionson Neural Networksand Learning Systems, 2026. doi: 10.1109/TNNLS.2025. 3650584
-
[45]
Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, and Yuke Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026
arXiv 2026
-
[46]
Elucidating the exposure bias in diffusion models
Mang Ning, Mingxiao Li, Jianlin Su, Albert Ali Salah, and Itir Onal Ertugrul. Elucidating the exposure bias in diffusion models. InInternational Conference on Learning Representations, volume 2024, pages 15167–15189, 2024
2024
-
[47]
Gr00t n1.5: An improved open foundation model for generalist humanoid robots.https://research
NVIDIA. Gr00t n1.5: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_5/, 2026
2026
-
[48]
Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research
NVIDIA. Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_6/, 2026
2026
-
[49]
Gr00t n1: An open foundation model for generalist humanoid robots, 2025
NVIDIA, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You ...
Pith/arXiv arXiv 2025
-
[50]
Jonas Pai, Liam Achenbach, Victoriano Montesinos, Benedek Forrai, Oier Mees, and Elvis Nava. mimic-video: Video-action models for generalizable robot control beyond vlas.arXiv preprint arXiv:2512.15692, 2025
Pith/arXiv arXiv 2025
-
[51]
Daojie Peng, Fulong Ma, Jiahang Cao, Qiang Zhang, Xupeng Xie, Jian Guo, Ping Luo, Andrew F. Luo, Boyu Zhou, and Jun Ma. Attena+: Rectifying action inequality in robotic foundation models.arXiv preprintarXiv:2605.13548, 2026
Pith/arXiv arXiv 2026
-
[52]
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025
Pith/arXiv arXiv 2025
-
[53]
arXiv preprint arXiv:2504.16054, 2025
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.π0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025
Pith/arXiv arXiv 2025
-
[54]
Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, et al.π0.7: a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026
Pith/arXiv arXiv 2026
-
[55]
Spatialvla: Exploring spatial representations for visual-language-action model.RSS, 2025
Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, et al. Spatialvla: Exploring spatial representations for visual-language-action model.RSS, 2025. 31
2025
-
[56]
Stephane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning.arXiv preprint arXiv:1011.0686, 2011
Pith/arXiv arXiv 2011
-
[57]
Generalization in generation: A closer look at exposure bias
Florian Schmidt. Generalization in generation: A closer look at exposure bias. InProceedings of the 3rd Workshop on Neural Generation and Translation, pages 157–167, 2019
2019
-
[58]
Xiang Shi, Wenlong Huang, Menglin Zou, and Xinhai Sun. Saivla-0: Cerebrum–pons–cerebellum tripartite architecture for compute-aware vision-language-action.arXiv preprint arXiv:2603.08124, 2026
arXiv 2026
-
[59]
Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. Vla-jepa: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098, 2026
arXiv 2026
-
[60]
Habitat 2.0: Training home assistants to rearrange their habitat
Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, et al. Habitat 2.0: Training home assistants to rearrange their habitat. arXiv preprint arXiv:2106.14405, 2022
Pith/arXiv arXiv 2022
-
[61]
Interactive post-training for vision-language-action models
Shuhan Tan, Kairan Dou, Yue Zhao, and Philipp Krähenbühl. Interactive post-training for vision-language-action models. arXiv preprint arXiv:2505.17016, 2025
Pith/arXiv arXiv 2025
-
[62]
Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu, Dongjie Huo, Yandan Yang, Bin Liu, Zhejia Cai, Feng Xiong, Mu Xu, et al. Alam: Algebraically consistent latent action model for vision-language-action models.arXiv preprint arXiv:2605.10819, 2026
Pith/arXiv arXiv 2026
-
[63]
Zuojin Tang, Shengchao Yuan, Xiaoxin Bai, Zhiyuan Jing, De Ma, Gang Pan, and Bin Liu. One token per frame: Reconsidering visual bandwidth in world models for vla policy.arXiv preprint arXiv:2605.07931, 2026
Pith/arXiv arXiv 2026
-
[64]
Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025
arXiv 2025
-
[65]
Bridgedata v2: A dataset for robot learning at scale.arXiv preprint arXiv:2308.12952, 2024
Homer Walke, Kevin Black, Abraham Lee, Moo Jin Kim, Max Du, Chongyi Zheng, Tony Zhao, Philippe Hansen- Estruch, Quan Vuong, Andre He, et al. Bridgedata v2: A dataset for robot learning at scale.arXiv preprint arXiv:2308.12952, 2024
Pith/arXiv arXiv 2024
-
[66]
Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianx- iao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025
Pith/arXiv arXiv 2025
-
[67]
Qwen-vla: Unifying vision-language-action modeling across tasks, environments, and robot embodiments
Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, et al. Qwen-vla: Unifying vision-language-action modeling across tasks, environments, and robot embodiments. arXiv preprint arXiv:2605.30280, 2026
Pith/arXiv arXiv 2026
-
[68]
Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation
Kun Wu, Chengkai Hou, Jiaming Liu, Zhengping Che, Xiaozhu Ju, Zhuqin Yang, Meng Li, Yinuo Zhao, Zhiyuan Xu, Guang Yang, et al. Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation. arXiv preprint arXiv:2412.13877, 2024
Pith/arXiv arXiv 2024
-
[69]
Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation
Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, Bowen Yang, Zhe Li, Kai Zhu, Hongyu Wu, Yiheng Liu, et al. Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation. arXiv preprint arXiv:2511.17441, 2025
Pith/arXiv arXiv 2025
-
[70]
Abot-m0: Vla foundation model for robotic manipulation with action manifold learning
Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning. arXiv preprint arXiv:2602.11236, 2026
Pith/arXiv arXiv 2026
-
[71]
Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jie Li, Jindi Lv, Jingyu Liu, et al. Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026
arXiv 2026
-
[72]
Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2025
Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2025
Pith/arXiv arXiv 2025
-
[73]
World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026. 32
Pith/arXiv arXiv 2026
-
[74]
Homerobot: Open-vocabulary mobile manipulation
Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander William Clegg, John Turner, et al. Homerobot: Open-vocabulary mobile manipulation. arXiv preprint arXiv:2306.11565, 2024
Pith/arXiv arXiv 2024
-
[75]
Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models
Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, et al. Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models. arXiv preprint arXiv:2606.17846, 2026
Pith/arXiv arXiv 2026
-
[76]
Tianyuan Yuan, Zibin Dong, Yicheng Liu, and Hang Zhao. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026
Pith/arXiv arXiv 2026
-
[77]
Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang, Haitao Lin, Jingbo Zhang, Yao Mu, Xiaokang Yang, Wenjun Zeng, and Xin Jin. Imagewam: Do world action models really need video generation, or just image editing?arXiv preprint arXiv:2606.19531, 2026
Pith/arXiv arXiv 2026
-
[78]
Cot-vla: Visual chain-of-thought reasoning for vision-language-action models
Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models. InCVPR, 2025
2025
-
[79]
X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model
Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, et al. X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model. ICLR, 2025
2025
-
[80]
Acot-vla: Action chain-of-thought for vision-language-action models
Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Si Liu, and Guanghui Ren. Acot-vla: Action chain-of-thought for vision-language-action models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8152–8162, 2026
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.