Pith. sign in

REVIEW 5 major objections 7 minor 4 cited by

Mobile manipulation succeeds when world-action models align temporal grain, action subspaces, and train-test conditions via latent actions, dual-level transformers, and dream forcing.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 09:20 UTC pith:R526GPKW

load-bearing objection Solid mobile-manipulation WAM systems paper: the three-alignment package is real engineering, results are multi-benchmark and ablated, but the latent-action bridge is still an unprobed assumption. the 5 major comments →

arxiv 2607.00678 v2 pith:R526GPKW submitted 2026-07-01 cs.CV cs.RO

ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

classification cs.CV cs.RO
keywords mobile manipulationworld action modellatent actionsmixture-of-transformersdream forcinginverse dynamicstrain-test alignmentvision-language-action
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Current robot policies either react without modeling the future or build world models that are mis-matched to mobile work: they predict coarse video chunks, entangle base motion with arm control, and train action heads on ground-truth futures they never see at deployment. The paper claims these three misalignments—not mere lack of scale—are why long-horizon household tasks fail. ABot-M0.5 repairs them by inserting frame-level latent actions as a bridge between video and control, separating mobility and manipulation inside a dual-level mixture-of-transformers, and training inverse dynamics on the model’s own predicted videos (dream forcing). On mobile, bimanual, tabletop, and real-robot benchmarks the resulting cascade reports higher long-horizon success and finer contact accuracy than prior vision-language-action and world-action baselines. The practical claim is that once world modeling, action abstraction, and rollout conditioning are aligned, the same architecture transfers from stationary manipulation to true mobile manipulation.

Core claim

Mobile manipulation is an alignment problem at three levels—temporal granularity, action space, and train-test consistency. ABot-M0.5 solves it by factorizing generation into video latents, then frame-level latent actions, then executable controls; by disentangling modality streams and heterogeneous action subspaces (base vs arm) inside a dual-level Mixture-of-Transformers; and by progressively training inverse dynamics on self-dreamed futures so the conditioning at training matches autoregressive inference. The result is state-of-the-art long-horizon task success and fine-grained control accuracy across simulation and real robots.

What carries the argument

The three-stage cascade zt+1 → mt → at (video latent to frame-level latent action to robot action), realized by a dual-level Mixture-of-Transformers that separates modalities and action subspaces, plus Dream Forcing that conditions action learning on model-predicted rather than ground-truth futures.

Load-bearing premise

Frame-level latent actions taken from visual frame pairs form a control-relevant, embodiment-agnostic bridge that recovers the fine contact dynamics coarse video chunks lose; if they mostly track camera or background change, the cascade fails.

What would settle it

Replace the latent-action stream with pure noise or with features from an untrained encoder, retrain under the same dual-level and dream-forcing schedule, and re-evaluate RoboCasa365 composite tasks plus real peg-insertion success; if rates stay near the full model, the bridging claim is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Long-horizon mobile success rises when inverse dynamics is trained on self-predicted futures rather than ground-truth futures.
  • Separating mobility and manipulation heads reduces gradient conflict and lifts composite-task rates.
  • Visual-only latent actions let motion priors transfer across robot embodiments without kinematic labels.
  • The same aligned cascade improves pure tabletop and bimanual benchmarks, not only mobile ones.
  • With aligned pretraining, real-robot policies can reach high success from only tens of demonstrations.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Dream forcing is a general recipe for closing teacher-forcing gaps in any world-model controller, not only robots.
  • Dual-level action disentanglement should help any multi-frequency control stack (e.g., whole-body locomotion versus fingers).
  • If latent actions truly encode contact onset, similar bridges could stabilize other autoregressive policies that currently suffer exposure bias.
  • Future WAM scaling may be limited more by these three alignments than by raw parameter count.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 7 minor

Summary. ABot-M0.5 is a World Action Model for language-conditioned mobile manipulation that jointly predicts future video latents, frame-level latent actions, and executable controls. The paper argues that existing WAMs are misaligned with mobile manipulation at three levels—temporal granularity (coarse video chunks vs. fine control), action structure (entangled mobility vs. manipulation), and train–test consistency (GT-conditioned inverse dynamics vs. autoregressive self-predicted futures)—and addresses them with (i) intermediate latent actions m_t extracted by a frozen ALAM-style encoder (Eq. 8; §3.2, §4.3), (ii) a dual-level Mixture-of-Transformers that separates modality streams and mobility/manipulation subspaces (§3.3), and (iii) Dream Forcing, a two-phase training scheme that conditions action prediction on self-dreamed latents (§3.4, §4.4). Training proceeds via world-model pretraining on multi-embodiment corpora, latent-action pretraining with algebraic consistency losses, then progressive SFT (teacher-forced then dream-forced). Empirically, the model reports strong results on RoboCasa365 (pretrain and target), RoboTwin 2.0, LIBERO/LIBERO-Plus, and real-robot arm tasks, with component ablations for latent-action structure, action-decoupled MoT, and Dream Forcing.

Significance. If the three-alignment thesis holds, the paper offers a concrete architectural and training recipe for extending world-action models from stationary to mobile manipulation—an important and under-served setting. The intermediate latent-action cascade, dual-level MoT, and Dream Forcing are clearly specified and, in principle, transferable. Strengths include multi-benchmark evaluation (RoboCasa365, RoboTwin clean/hard, LIBERO, LIBERO-Plus zero-shot, real robot), structured ablations (Table 7 latent-action variants; Table 8 Dream Forcing vs. continued teacher forcing; action-decoupled MoT on a composite subset), progressive train–test alignment rather than pure capacity scaling, and promised code. The contribution is primarily empirical/systems rather than theoretical; its lasting value depends on whether the reported gains are causally tied to the claimed alignments (especially what m_t encodes) and whether the mobile claims transfer beyond simulation.

major comments (5)
  1. §3.2, Eq. (8) and §4.3: The central temporal-granularity claim is that frozen ALAM latents m_t recover fine contact/grasp dynamics lost in coarse z_{t+1}. The paper never validates this. Algebraic losses L_add and L_rev enforce only visual-transition consistency; there is no correlation of m_t with contact onset, grasp closure, force proxies, or end-effector–object relative motion, nor a control experiment that replaces m_t with a capacity-matched intermediate stream (e.g., random or optical-flow tokens). Table 7 shows that a 3-stage cascade helps (94.0% vs. 87.6% baseline on RoboTwin Clean), but that is consistent with capacity/conditioning structure rather than contact-relevant content. Either provide a direct semantic analysis of m_t or substantially tone down the claim that the cascade specifically restores contact dynamics.
  2. Table 2 (§5.2): The paper reports ABot-M0.5(+Condensed Memory) as setting a new record (46.6% average) while stating that condensed memory “will be presented in our future work.” Publishing SOTA numbers for a mechanism that is neither defined nor ablated in this manuscript is not acceptable for a primary result table. Either fully specify and ablate condensed memory here, or remove those rows and base all SOTA claims on the fully described ABot-M0.5 (40.4%).
  3. Table 2 Composite-Unseen and Abstract/§5.2 SOTA framing: On Composite-Unseen, ABot-M0.5 scores 2.7% (7.9% with condensed memory), below Qwen-RobotManip (14.9%) and Qwen-RobotManip-Context (11.2%). The abstract and §5.2 claim state-of-the-art long-horizon mobile manipulation success without qualifying this gap. Long-horizon composite generalization is precisely where the three-alignment thesis should matter most. Please reframe SOTA claims by category (Atomic-Seen / Composite-Seen / Composite-Unseen) and discuss why Composite-Unseen remains weak.
  4. §5.5 Real-world experiments: All real-robot results use a fixed-base Agilex Piper 6-DoF arm; there is no mobile base, navigation phase, or viewpoint change induced by locomotion. The paper’s central problem is mobile manipulation (Abstract; §1–2; RoboCasa365). Real-world transfer therefore supports fine-grained and multi-stage arm control, not the mobility–manipulation alignment thesis. Either add real mobile-base experiments or clearly restrict the real-world claim to stationary manipulation transfer and avoid implying physical validation of mobile alignment.
  5. §5.4 Action-decoupled MoT: The only quantitative support is a “selected subset” of RoboCasa365 Composite-Seen (0.48 vs. 0.34) plus a training-curve figure. No full-benchmark numbers, no ablation of joint attention vs. fully separate towers, and no analysis of whether mobility/manipulation specialization actually occurs (e.g., subspace-wise error or attention patterns). Given that action-space alignment is one of the three load-bearing claims, please report full Composite-Seen/Unseen numbers with and without action-level decoupling, and ideally a diagnostic that the two heads specialize.
minor comments (7)
  1. Table 1 / §3.1: Notation switches between m_t, m_{0,t}, and M for latent actions; standardize early and keep one symbol.
  2. Eqs. (11)–(13) and (25)–(26): λ_move / λ_manip and λ_z / λ_m / λ_a are free parameters but never listed with values or sensitivity. A short hyperparameter table would help reproducibility.
  3. Figure 1 and Abstract: “finegrained” should be “fine-grained”; several figure captions mix “Dreamed” / “dreamed” inconsistently.
  4. §4.2 semantic slot allocation (2 third-person + 2 wrist): state how often slots are empty/padded on each pretraining corpus and whether padding rate correlates with downstream multi-view performance.
  5. Table 8: Dream Forcing gains are modest (+3.01% absolute over the 50k warm-start; +1.66% over continued teacher forcing at 60k). The text should avoid language that implies a large qualitative jump without reporting variance over seeds.
  6. §5.1 Implementation details are thin (optimizer, learning rates, chunk size H, number of dream denoising steps). Add a concise appendix table.
  7. Related work: several concurrent WAMs (Motus, Fast-WAM, Lingbot-VA, ImageWAM) are compared in tables but discussed only lightly in the introduction; a short related-work paragraph contrasting teacher forcing vs. diffusion forcing vs. Dream Forcing would help placement.

Circularity Check

0 steps flagged

Empirical systems paper: SOTA claims are measured on external benchmarks, not recovered by construction from fitted constants or self-definitional identities.

full rationale

ABot-M0.5 is an architecture-and-training paper whose central claims are empirical success rates on RoboCasa365, RoboTwin 2.0, LIBERO/LIBERO-Plus, and real-robot tasks against external VLA/WAM baselines. Intermediate latent actions are extracted by a frozen ALAM-style encoder pretrained with algebraic visual-transition losses (Ladd, Lrev, reconstruction/VQ); those losses do not include the downstream task metrics, so mt is not defined in terms of the reported success rates. Dual-level MoT and Dream Forcing are design choices whose benefit is measured by ablations (Tables 7–8, Fig. 10) and held-out task success, not forced by fitting a parameter that is then renamed a prediction. Self-citations to ALAM and prior ABot work supply reusable components (encoder, data pipeline) but do not import a uniqueness theorem that forbids alternatives or make the SOTA claim tautological. No equation equates a claimed prediction to its own fitted input. The skeptic concern that mt may encode generic visual change rather than contact dynamics is an unvalidated-assumption / correctness issue, not circularity. Score 0 is therefore the honest finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 3 invented entities

Load-bearing content is architectural and empirical, not axiomatic physics. Free parameters are the usual ML knobs (loss weights, drop rates, denoising schedules, slot layout). Domain assumptions are standard for visuomotor learning (CFM, frozen latent-action encoder, multi-view slot semantics). Invented entities are the intermediate latent-action bridge and the dream-forcing conditioning regime—engineering constructs with ablation evidence inside the paper but no independent physical existence claim.

free parameters (5)
  • λ_z, λ_m, λ_a and λ_move, λ_manip action subspace weights
    Weighted sum of CFM losses across cascade stages and mobility/manipulation branches; values chosen for training stability, not derived.
  • latent-action conditioning dropout p_drop
    Ablated at 0 vs 0.2; final model uses 0 for train-inference alignment. Hand-chosen hyperparameter affecting reported success.
  • shared denoising timestep τ and few-step dream sampling schedule
    Controls noise path for CFM and Phase-A dreamed latents; efficiency/quality trade-off set by authors.
  • four canonical multi-view semantic slots (2 third-person + 2 wrist)
    Fixed camera-role allocation for heterogeneous datasets; design choice that shapes the world-model representation.
  • ALAM pretraining loss weights λ_vq, λ_rec, λ_perc, λ_add, λ_rev
    Balance reconstruction, VQ, and algebraic consistency for the frozen latent-action encoder.
axioms (5)
  • domain assumption Conditional flow matching on video, latent-action, and action latents is a valid generative objective for robotic control.
    Adopted from recent continuous-action VLA/WAM literature (§3.1–3.3); not re-proved.
  • domain assumption Visual frame-pair transitions encode embodiment-agnostic motion intents that transfer across robot morphologies.
    Core premise of intermediate latent actions and ALAM pretraining (§3.2, §4.3).
  • domain assumption At deployment, historical chunks can be grounded in real GT observations so only the latest future chunk needs dreaming.
    Justifies parallel few-step Phase-A rollout instead of full sequential self-forcing (§3.4).
  • ad hoc to paper Base mobility and arm manipulation are sufficiently heterogeneous that separate FFN/heads reduce gradient interference without destroying coordination via joint attention.
    Motivates action-level MoT; supported by ablation on RoboCasa365 composite subset but not a universal theorem (§3.3, §5.4).
  • standard math Standard math of CFM velocity regression and causal attention masks.
    Used throughout loss definitions (Eqs. 5, 9, 11–13, 25–26).
invented entities (3)
  • Frame-level intermediate latent action m_t as bridging space no independent evidence
    purpose: Capture local visual state transitions between coarse video latents and embodiment-specific controls.
    Defined via frozen encoder E_m(I_t, I_{t+1}); purpose is architectural factorization, not a new physical quantity. Independent evidence is only internal ablations.
  • Dual-level Mixture-of-Transformers (modality + mobility/manipulation) no independent evidence
    purpose: Disentangle modality streams and heterogeneous action subspaces while sharing attention.
    Architectural construct; gains shown on composite mobile tasks. No claim of independent existence outside the model.
  • Dream Forcing training paradigm (two-phase forward on self-dreamed latents) no independent evidence
    purpose: Align inverse-dynamics conditioning with autoregressive inference-time predicted videos.
    Training procedure related to self-forcing literature but specialized to closed-loop robot chunks; evidence is ablation vs teacher forcing.

pith-pipeline@v1.1.0-grok45 · 33865 in / 3873 out tokens · 41959 ms · 2026-07-12T09:20:29.372869+00:00 · methodology

0 comments
read the original abstract

Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

    cs.RO 2026-07 conditional novelty 6.0

    A frozen world-action model can be steered to new tasks by adapting a lightweight memory from unlabeled human video via test-time training.

  2. WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

    cs.RO 2026-07 conditional novelty 6.0

    A meta-trained test-time memory lets frozen world-action models absorb unlabeled human videos and outperform in-context video conditioning on real multi-embodiment manipulation.

  3. ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI

    cs.RO 2026-07 conditional novelty 6.0

    A single 8B backbone unifies spatial perception, decision making, navigation/manipulation, and progress estimation with SSR+ merging, reporting gains on most spatial benchmarks and competitive action/progress results.

  4. From Foundation to Application: Improving VLA Models in Practice

    cs.RO 2026-07 conditional novelty 4.0

    LingBot-VLA 2.0 combines 60k hours of multi-embodiment pretraining data, an expanded whole-body action space, and dual-query distillation from depth and video teachers to improve VLA performance on GM-100 and long-hor...

Reference graph

Works this paper leans on

82 extracted references · 59 linked inside Pith · cited by 3 Pith papers

  1. [1]

    Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025

    AgiBot-World-Contributors, Qingwen Bu, Jisong Cai, Li Chen, Xiuqi Cui, Yan Ding, Siyuan Feng, Shenyuan Gao, Xindong He, Xuan Hu, et al. Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems.arXiv preprint arXiv:2503.06669, 2025

  2. [2]

    Scheduled sampling for sequence prediction with recurrent neural networks.arXiv preprint arXiv:1506.03099, 2015

    Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. Scheduled sampling for sequence prediction with recurrent neural networks.arXiv preprint arXiv:1506.03099, 2015

  3. [3]

    Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025

    Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, Shuhe Huang, Haitian Liu, Ruowen Zhao, Yao Feng, Chendong Xiang, Yinze Rong, et al. Motus: A unified latent action world model.arXiv preprint arXiv:2512.13030, 2025

  4. [4]

    arXiv preprint arXiv:2410.24164, 2024

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.π0: A vision-language-action flow model for general robot control. arXiv preprint arXiv:2410.24164, 2024

  5. [5]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023

  6. [6]

    Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale.arXiv preprint arXiv:2212.06817, 2023

  7. [7]

    Univla: Learning to act anywhere with task-centric latent actions.RSS, 2025

    Qingwen Bu, Yanting Yang, Jisong Cai, Shenyuan Gao, Guanghui Ren, Maoqing Yao, Ping Luo, and Hongyang Li. Univla: Learning to act anywhere with task-centric latent actions.RSS, 2025

  8. [8]

    Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025

    Jun Cen, Chaohui Yu, Hangjie Yuan, Yuming Jiang, Siteng Huang, Jiayan Guo, Xin Li, Yibing Song, Hao Luo, Fan Wang, et al. Worldvla: Towards autoregressive action world model.arXiv preprint arXiv:2506.21539, 2025

  9. [9]

    Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025

    Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025

  10. [10]

    Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

    Cheng Chi, Zhenjia Xu, Siyuan Feng, Eric Cousineau, Yilun Du, Benjamin Burchfiel, Russ Tedrake, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

  11. [11]

    Open x-embodiment: Robotic learning datasets and rt-x models.arXiv preprint arXiv:2310.08864, 2023

    Open X-Embodiment Collaboration. Open x-embodiment: Robotic learning datasets and rt-x models.arXiv preprint arXiv:2310.08864, 2023

  12. [12]

    Internvla-m1: A spatially guided vision-language-action framework for generalist robot policy.arXiv preprint arXiv:2510.13778, 2025

    InternVLA-M1 Contributors. Internvla-m1: A spatially guided vision-language-action framework for generalist robot policy.arXiv preprint arXiv:2510.13778, 2025

  13. [13]

    Self-forcing++: Towards minute-scale high-quality video generation.arXiv preprint arXiv:2510.02283, 2025

    Justin Cui, Jie Wu, Ming Li, Tao Yang, Xiaojie Li, Rui Wang, Andrew Bai, Yuanhao Ban, and Cho-Jui Hsieh. Self-forcing++: Towards minute-scale high-quality video generation.arXiv preprint arXiv:2510.02283, 2025

  14. [14]

    Fu, Stefano Ermon, Atri Rudra, and Christopher Re

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Re. Flashattention: Fast and memory-efficient exact attention with io-awareness.arXiv preprint arXiv:2205.14135, 2022

  15. [15]

    Robonet: Large-scale multi-robot learning.arXiv preprint arXiv:1910.11215, 2020

    Sudeep Dasari, Frederik Ebert, Stephen Tian, Suraj Nair, Bernadette Bucher, Karl Schmeckpeper, Siddharth Singh, Sergey Levine, and Chelsea Finn. Robonet: Large-scale multi-robot learning.arXiv preprint arXiv:1910.11215, 2020

  16. [16]

    Libero-plus: In-depth robustness analysis of vision-language-action models

    Senyu Fei, Siyin Wang, Junhao Shi, Zihao Dai, Jikun Cai, Pengfang Qian, Li Ji, Xinzhe He, Shiduo Zhang, Zhaoye Fei, Jinlan Fu, Jingjing Gong, and Xipeng Qiu. Libero-plus: In-depth robustness analysis of vision-language-action models. arXiv preprint arXiv:2510.13626, 2025

  17. [17]

    Vidar: Embodied video diffusion model for generalist manipulation.arXiv preprint arXiv:2507.12898, 2025

    Yao Feng, Hengkai Tan, Xinyi Mao, Chendong Xiang, Guodong Liu, Shuhe Huang, Hang Su, and Jun Zhu. Vidar: Embodied video diffusion model for generalist manipulation.arXiv preprint arXiv:2507.12898, 2025

  18. [18]

    Galaxea g0.5 technical report

    Galaxea Team. Galaxea g0.5 technical report. 2026. URLhttps://opengalaxea.github.io/G05/

  19. [19]

    Priorvla: Prior-preserving adaptation for vision-language-action models.arXiv preprint arXiv:2605.10925, 2026

    Xinyu Guo, Bin Xie, Wei Chai, Xianchi Deng, Tiancai Wang, Zhengxing Wu, and Xingyu Chen. Priorvla: Prior-preserving adaptation for vision-language-action models.arXiv preprint arXiv:2605.10925, 2026. 29

  20. [20]

    World models.arXiv preprint arXiv:1803.10122, 2018

    David Ha and Jürgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018

  21. [21]

    Mastering diverse domains through world models

    Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv preprint arXiv:2301.04104, 2024

  22. [22]

    Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026

    Xun Huang, Zhengqi Li, Guande He, Mingyuan Zhou, and Eli Shechtman. Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026

  23. [23]

    Nora: A small open-sourced generalist vision language action model for embodied tasks.arXiv preprint arXiv:2504.19854, 2025

    Chia-Yu Hung, Qi Sun, Pengfei Hong, Amir Zadeh, Chuan Li, U Tan, Navonil Majumder, Soujanya Poria, et al. Nora: A small open-sourced generalist vision language action model for embodied tasks.arXiv preprint arXiv:2504.19854, 2025

  24. [24]

    ABot-Claw: A foundation for persistent, cooperative, and self-evolving robotic agents

    Dongjie Huo, Haoyun Liu, Guoqing Liu, Dekang Qi, Zhiming Sun, Maoguo Gao, Jianxin He, Yandan Yang, Xinyuan Chang, Feng Xiong, et al. ABot-Claw: A foundation for persistent, cooperative, and self-evolving robotic agents. arXiv preprint arXiv:2604.10096, 2026

  25. [25]

    Oxe-auge: A large-scale robot augmentation of oxe for scaling cross-embodiment policy learning

    Guanhua Ji, Harsha Polavaram, Lawrence Yunliang Chen, Sandeep Bajamahal, Zehan Ma, Simeon Adebola, Chenfeng Xu, and Ken Goldberg. Oxe-auge: A large-scale robot augmentation of oxe for scaling cross-embodiment policy learning. arXiv preprint arXiv:2512.13100, 2025

  26. [26]

    Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025

    Tao Jiang, Tianyuan Yuan, Yicheng Liu, Chenhao Lu, Jianning Cui, Xiao Liu, Shuiqi Cheng, Jiyang Gao, Huazhe Xu, and Hang Zhao. Galaxea open-world dataset and g0 dual-system vla model.arXiv preprint arXiv:2509.00576, 2025

  27. [27]

    Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2025

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset.arXiv preprint arXiv:2403.12945, 2025

  28. [28]

    Rldx-1 technical report.arXiv preprint arXiv:2605.03269, 2026

    Dongyoung Kim, Huiwon Jang, Myungkyu Koo, Suhyeok Jang, Taeyoung Kim, Beomjun Kim, Byungjun Yoon, Changsung Jang, Daewon Choi, Dongsu Han, et al. Rldx-1 technical report.arXiv preprint arXiv:2605.03269, 2026

  29. [29]

    Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024

  30. [30]

    Fine-tuning vision-language-action models: Optimizing speed and success

    Moo Jin Kim, Chelsea Finn, and Percy Liang. Fine-tuning vision-language-action models: Optimizing speed and success. RSS, 2025

  31. [31]

    Cosmos policy: Fine-tuning video models for visuomotor control and planning

    Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, et al. Cosmos policy: Fine-tuning video models for visuomotor control and planning. arXiv preprint arXiv:2601.16163, 2026

  32. [32]

    Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation.arXiv preprint arXiv:2403.09227, 2024

    Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martin-Martin, Chen Wang, Gabrael Levine, Wensi Ai, Benjamin Martinez, et al. Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation.arXiv preprint arXiv:2403.09227, 2024

  33. [33]

    Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026

    Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, et al. Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026

  34. [34]

    Clam: Continuous latent action models for robot learning from unlabeled demonstrations.arXiv preprint arXiv:2505.04999, 2025

    Anthony Liang, Pavel Czempin, Matthew Hong, Yutai Zhou, Erdem Biyik, and Stephen Tu. Clam: Continuous latent action models for robot learning from unlabeled demonstrations.arXiv preprint arXiv:2505.04999, 2025

  35. [35]

    Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies

    Zhixuan Liang, Yizhuo Li, Tianshuo Yang, Chengyue Wu, Sitong Mao, Tian Nian, Liuao Pei, Shunbo Zhou, Xiaokang Yang, Jiangmiao Pang, et al. Discrete diffusion vla: Bringing discrete diffusion to action decoding in vision-language-action policies. arXiv preprint arXiv:2508.20072, 2025

  36. [36]

    Holobrain-0 technical report.arXiv preprint arXiv:2602.12062, 2026

    Xuewu Lin, Tianwei Lin, Yun Du, Hongyu Xie, Yiwei Jin, Jiawei Li, Shijie Wu, Qingze Wang, Mengdi Li, Mengao Zhao, et al. Holobrain-0 technical report.arXiv preprint arXiv:2602.12062, 2026

  37. [37]

    Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le. Flow matching for generative modeling. arXiv preprint arXiv:2210.02747, 2023

  38. [38]

    Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023. 30

  39. [39]

    Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025

    Kunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan, and Shijian Lu. Rolling forcing: Autoregressive long video diffusion in real time.arXiv preprint arXiv:2509.25161, 2025

  40. [40]

    Being-h0

    Hao Luo, Ye Wang, Wanpeng Zhang, Sipeng Zheng, Ziheng Xi, Chaoyi Xu, Haiweng Xu, Haoqi Yuan, Chi Zhang, Yiqing Wang, et al. Being-h0. 5: Scaling human-centric robot learning for cross-embodiment generalization.arXiv preprint arXiv:2601.12993, 2026

  41. [41]

    Being-h0

    Hao Luo, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Haiweng Xu, Chaoyi Xu, Ziheng Xi, Yuhui Fu, and Zongqing Lu. Being-h0. 7: A latent world-action model from egocentric videos.arXiv preprint arXiv:2605.00078, 2026

  42. [42]

    Coral: Scalable multi-task robot learning via lora experts

    Yuankai Luo, Woping Chen, Tong Liang, and Zhenguo Li. Coral: Scalable multi-task robot learning via lora experts. arXiv preprint arXiv:2603.09298, 2026

  43. [43]

    F1: A vision-language-action model bridging understanding and generation to actions.arXiv preprint arXiv:2509.06951, 2025

    Qi Lv, Weijie Kong, Hao Li, Jia Zeng, Zherui Qiu, Delin Qu, Haoming Song, Qizhi Chen, Xiang Deng, and Jiangmiao Pang. F1: A vision-language-action model bridging understanding and generation to actions.arXiv preprint arXiv:2509.06951, 2025

  44. [44]

    A survey on vision-language-action models for embodied ai.IEEE Transactionson Neural Networksand Learning Systems, 2026

    Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied ai.IEEE Transactionson Neural Networksand Learning Systems, 2026. doi: 10.1109/TNNLS.2025. 3650584

  45. [45]

    Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

    Soroush Nasiriany, Sepehr Nasiriany, Abhiram Maddukuri, and Yuke Zhu. Robocasa365: A large-scale simulation framework for training and benchmarking generalist robots.arXiv preprint arXiv:2603.04356, 2026

  46. [46]

    Elucidating the exposure bias in diffusion models

    Mang Ning, Mingxiao Li, Jianlin Su, Albert Ali Salah, and Itir Onal Ertugrul. Elucidating the exposure bias in diffusion models. InInternational Conference on Learning Representations, volume 2024, pages 15167–15189, 2024

  47. [47]

    Gr00t n1.5: An improved open foundation model for generalist humanoid robots.https://research

    NVIDIA. Gr00t n1.5: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_5/, 2026

  48. [48]

    Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research

    NVIDIA. Gr00t n1.6: An improved open foundation model for generalist humanoid robots.https://research. nvidia.com/labs/gear/gr00t-n1_6/, 2026

  49. [49]

    Gr00t n1: An open foundation model for generalist humanoid robots, 2025

    NVIDIA, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You ...

  50. [50]

    mimic-video: Video-action models for generalizable robot control beyond vlas.arXiv preprint arXiv:2512.15692, 2025

    Jonas Pai, Liam Achenbach, Victoriano Montesinos, Benedek Forrai, Oier Mees, and Elvis Nava. mimic-video: Video-action models for generalizable robot control beyond vlas.arXiv preprint arXiv:2512.15692, 2025

  51. [51]

    Luo, Boyu Zhou, and Jun Ma

    Daojie Peng, Fulong Ma, Jiahang Cao, Qiang Zhang, Xupeng Xie, Jian Guo, Ping Luo, Andrew F. Luo, Boyu Zhou, and Jun Ma. Attena+: Rectifying action inequality in robotic foundation models.arXiv preprintarXiv:2605.13548, 2026

  52. [52]

    Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025

    Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025

  53. [53]

    arXiv preprint arXiv:2504.16054, 2025

    Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, et al.π0.5: a vision-language-action model with open-world generalization. arXiv preprint arXiv:2504.16054, 2025

  54. [54]

    Physical Intelligence, Bo Ai, Ali Amin, Raichelle Aniceto, Ashwin Balakrishna, Greg Balke, Kevin Black, George Bokinsky, Shihao Cao, Thomas Charbonnier, et al.π0.7: a steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026

  55. [55]

    Spatialvla: Exploring spatial representations for visual-language-action model.RSS, 2025

    Delin Qu, Haoming Song, Qizhi Chen, Yuanqi Yao, Xinyi Ye, Yan Ding, Zhigang Wang, JiaYuan Gu, Bin Zhao, Dong Wang, et al. Spatialvla: Exploring spatial representations for visual-language-action model.RSS, 2025. 31

  56. [56]

    Gordon, and J

    Stephane Ross, Geoffrey J. Gordon, and J. Andrew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning.arXiv preprint arXiv:1011.0686, 2011

  57. [57]

    Generalization in generation: A closer look at exposure bias

    Florian Schmidt. Generalization in generation: A closer look at exposure bias. InProceedings of the 3rd Workshop on Neural Generation and Translation, pages 157–167, 2019

  58. [58]

    Saivla-0: Cerebrum–pons–cerebellum tripartite architecture for compute-aware vision-language-action.arXiv preprint arXiv:2603.08124, 2026

    Xiang Shi, Wenlong Huang, Menglin Zou, and Xinhai Sun. Saivla-0: Cerebrum–pons–cerebellum tripartite architecture for compute-aware vision-language-action.arXiv preprint arXiv:2603.08124, 2026

  59. [59]

    Vla-jepa: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098, 2026

    Jingwen Sun, Wenyao Zhang, Zekun Qi, Shaojie Ren, Zezhi Liu, Hanxin Zhu, Guangzhong Sun, Xin Jin, and Zhibo Chen. Vla-jepa: Enhancing vision-language-action model with latent world model.arXiv preprint arXiv:2602.10098, 2026

  60. [60]

    Habitat 2.0: Training home assistants to rearrange their habitat

    Andrew Szot, Alex Clegg, Eric Undersander, Erik Wijmans, Yili Zhao, John Turner, Noah Maestre, Mustafa Mukadam, Devendra Chaplot, Oleksandr Maksymets, et al. Habitat 2.0: Training home assistants to rearrange their habitat. arXiv preprint arXiv:2106.14405, 2022

  61. [61]

    Interactive post-training for vision-language-action models

    Shuhan Tan, Kairan Dou, Yue Zhao, and Philipp Krähenbühl. Interactive post-training for vision-language-action models. arXiv preprint arXiv:2505.17016, 2025

  62. [62]

    Alam: Algebraically consistent latent action model for vision-language-action models.arXiv preprint arXiv:2605.10819, 2026

    Zuojin Tang, Haoyun Liu, Xinyuan Chang, Changjie Wu, Dongjie Huo, Yandan Yang, Bin Liu, Zhejia Cai, Feng Xiong, Mu Xu, et al. Alam: Algebraically consistent latent action model for vision-language-action models.arXiv preprint arXiv:2605.10819, 2026

  63. [63]

    One token per frame: Reconsidering visual bandwidth in world models for vla policy.arXiv preprint arXiv:2605.07931, 2026

    Zuojin Tang, Shengchao Yuan, Xiaoxin Bai, Zhiyuan Jing, De Ma, Gang Pan, and Bin Liu. One token per frame: Reconsidering visual bandwidth in world models for vla policy.arXiv preprint arXiv:2605.07931, 2026

  64. [64]

    Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025

    Yang Tian, Yuyin Yang, Yiman Xie, Zetao Cai, Xu Shi, Ning Gao, Hangxu Liu, Xuekun Jiang, Zherui Qiu, Feng Yuan, et al. Interndata-a1: Pioneering high-fidelity synthetic data for pre-training generalist policy.arXiv preprint arXiv:2511.16651, 2025

  65. [65]

    Bridgedata v2: A dataset for robot learning at scale.arXiv preprint arXiv:2308.12952, 2024

    Homer Walke, Kevin Black, Abraham Lee, Moo Jin Kim, Max Du, Chongyi Zheng, Tony Zhao, Philippe Hansen- Estruch, Quan Vuong, Andre He, et al. Bridgedata v2: A dataset for robot learning at scale.arXiv preprint arXiv:2308.12952, 2024

  66. [66]

    Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

    Team Wan, Ang Wang, Baole Ai, Bin Wen, Chaojie Mao, Chen-Wei Xie, Di Chen, Feiwu Yu, Haiming Zhao, Jianx- iao Yang, et al. Wan: Open and advanced large-scale video generative models.arXiv preprint arXiv:2503.20314, 2025

  67. [67]

    Qwen-vla: Unifying vision-language-action modeling across tasks, environments, and robot embodiments

    Qiuyue Wang, Mingsheng Li, Jian Guan, Jinhui Ye, Sicheng Xie, Yitao Liu, Junhao Chen, Zhixuan Liang, Jie Zhang, Xintong Hu, et al. Qwen-vla: Unifying vision-language-action modeling across tasks, environments, and robot embodiments. arXiv preprint arXiv:2605.30280, 2026

  68. [68]

    Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation

    Kun Wu, Chengkai Hou, Jiaming Liu, Zhengping Che, Xiaozhu Ju, Zhuqin Yang, Meng Li, Yinuo Zhao, Zhiyuan Xu, Guang Yang, et al. Robomind: Benchmark on multi-embodiment intelligence normative data for robot manipulation. arXiv preprint arXiv:2412.13877, 2024

  69. [69]

    Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation

    Shihan Wu, Xuecheng Liu, Shaoxuan Xie, Pengwei Wang, Xinghang Li, Bowen Yang, Zhe Li, Kai Zhu, Hongyu Wu, Yiheng Liu, et al. Robocoin: An open-sourced bimanual robotic data collection for integrated manipulation. arXiv preprint arXiv:2511.17441, 2025

  70. [70]

    Abot-m0: Vla foundation model for robotic manipulation with action manifold learning

    Yandan Yang, Shuang Zeng, Tong Lin, Xinyuan Chang, Dekang Qi, Junjin Xiao, Haoyun Liu, Ronghan Chen, Yuzhi Chen, Dongjie Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning. arXiv preprint arXiv:2602.11236, 2026

  71. [71]

    Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026

    Angen Ye, Boyuan Wang, Chaojun Ni, Guan Huang, Guosheng Zhao, Hao Li, Hengtao Li, Jie Li, Jindi Lv, Jingyu Liu, et al. Gigaworld-policy: An efficient action-centered world–action model.arXiv preprint arXiv:2603.17240, 2026

  72. [72]

    Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2025

    Seonghyeon Ye, Joel Jang, Byeongguk Jeon, Sejune Joo, Jianwei Yang, Baolin Peng, Ajay Mandlekar, Reuben Tan, Yu-Wei Chao, Bill Yuchen Lin, et al. Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2025

  73. [73]

    World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026

    Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026. 32

  74. [74]

    Homerobot: Open-vocabulary mobile manipulation

    Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav, Austin Wang, Mukul Khanna, Theophile Gervet, Tsung-Yen Yang, Vidhi Jain, Alexander William Clegg, John Turner, et al. Homerobot: Open-vocabulary mobile manipulation. arXiv preprint arXiv:2306.11565, 2024

  75. [75]

    Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models

    Haoqi Yuan, Zhixuan Liang, Anzhe Chen, Ye Wang, Haoyang Li, Pei Lin, Yiyang Huang, Zixing Lei, Tong Zhang, Jiazhao Zhang, et al. Qwen-robotmanip technical report: Alignment unlocks scale for robotic manipulation foundation models. arXiv preprint arXiv:2606.17846, 2026

  76. [76]

    Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026

    Tianyuan Yuan, Zibin Dong, Yicheng Liu, and Hang Zhao. Fast-wam: Do world action models need test-time future imagination? arXiv preprint arXiv:2603.16666, 2026

  77. [77]

    Imagewam: Do world action models really need video generation, or just image editing?arXiv preprint arXiv:2606.19531, 2026

    Yuyang Zhang, Wenyao Zhang, Zekun Qi, He Zhang, Haitao Lin, Jingbo Zhang, Yao Mu, Xiaokang Yang, Wenjun Zeng, and Xin Jin. Imagewam: Do world action models really need video generation, or just image editing?arXiv preprint arXiv:2606.19531, 2026

  78. [78]

    Cot-vla: Visual chain-of-thought reasoning for vision-language-action models

    Qingqing Zhao, Yao Lu, Moo Jin Kim, Zipeng Fu, Zhuoyang Zhang, Yecheng Wu, Zhaoshuo Li, Qianli Ma, Song Han, Chelsea Finn, et al. Cot-vla: Visual chain-of-thought reasoning for vision-language-action models. InCVPR, 2025

  79. [79]

    X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model

    Jinliang Zheng, Jianxiong Li, Zhihao Wang, Dongxiu Liu, Xirui Kang, Yuchun Feng, Yinan Zheng, Jiayin Zou, Yilun Chen, Jia Zeng, et al. X-vla: Soft-prompted transformer as scalable cross-embodiment vision-language-action model. ICLR, 2025

  80. [80]

    Acot-vla: Action chain-of-thought for vision-language-action models

    Linqing Zhong, Yi Liu, Yifei Wei, Ziyu Xiong, Si Liu, and Guanghui Ren. Acot-vla: Action chain-of-thought for vision-language-action models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8152–8162, 2026

Showing first 80 references.