AIM predicts aligned spatial value maps inside a shared video-generation transformer to produce reliable robot actions, reaching 94% success on RoboTwin 2.0 with larger gains on long-horizon and contact-rich tasks.
Rethinking the practicality of vision-language-action model: A comprehensive benchmark and an improved baseline
3 Pith papers cite this work. Polarity classification is still indexing.
3
Pith papers citing it
fields
cs.RO 3years
2026 3representative citing papers
HiMem-WAM integrates hierarchical latent actions and boundary-aware memory gates into world action models to enhance robustness and performance on memory-dependent long-horizon robotic tasks.
citing papers explorer
-
AIM: Intent-Aware Unified world action Modeling with Spatial Value Maps
AIM predicts aligned spatial value maps inside a shared video-generation transformer to produce reliable robot actions, reaching 94% success on RoboTwin 2.0 with larger gains on long-horizon and contact-rich tasks.
-
HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation
HiMem-WAM integrates hierarchical latent actions and boundary-aware memory gates into world action models to enhance robustness and performance on memory-dependent long-horizon robotic tasks.
- Can Vision-Language-Action Models Learn from Real-World Data Continually without Forgetting?