REVIEW 10 cited by
Multimodal Diffusion Transformer: Learning Versatile Behavior from Multimodal Goals
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
This work introduces the Multimodal Diffusion Transformer (MDT), a novel diffusion policy framework, that excels at learning versatile behavior from multimodal goal specifications with few language annotations. MDT leverages a diffusion-based multimodal transformer backbone and two self-supervised auxiliary objectives to master long-horizon manipulation tasks based on multimodal goals. The vast majority of imitation learning methods only learn from individual goal modalities, e.g. either language or goal images. However, existing large-scale imitation learning datasets are only partially labeled with language annotations, which prohibits current methods from learning language conditioned behavior from these datasets. MDT addresses this challenge by introducing a latent goal-conditioned state representation that is simultaneously trained on multimodal goal instructions. This state representation aligns image and language based goal embeddings and encodes sufficient information to predict future states. The representation is trained via two self-supervised auxiliary objectives, enhancing the performance of the presented transformer backbone. MDT shows exceptional performance on 164 tasks provided by the challenging CALVIN and LIBERO benchmarks, including a LIBERO version that contains less than $2\%$ language annotations. Furthermore, MDT establishes a new record on the CALVIN manipulation challenge, demonstrating an absolute performance improvement of $15\%$ over prior state-of-the-art methods that require large-scale pretraining and contain $10\times$ more learnable parameters. MDT shows its ability to solve long-horizon manipulation from sparsely annotated data in both simulated and real-world environments. Demonstrations and Code are available at https://intuitive-robots.github.io/mdt_policy/.
Forward citations
Cited by 10 Pith papers
-
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
A taxonomy of robot learning on a weights-versus-skills axis, with a five-rung self-improvement ladder whose top cell (feedback plus memory plus search) holds only a few recent systems.
-
VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning
An RGB-only diffusion policy that builds an explicit volumetric representation, distills it into spatial tokens, and conditions a multi-token decoder reaches 88.8% average success on LIBERO.
-
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
A transformer-based compressor that turns old observations into fixed-size memory tokens improves non-Markovian imitation-learning robot policies, with large gains on memory-intensive simulated tasks.
-
Robot-DIFT: Correspondence-Sensitive Diffusion Features for Contact-Rich Robot Manipulation
Distilling Stable Diffusion's decoder features into a deterministic student backbone with a multi-scale fusion network improves contact-rich manipulation success in simulation and on a real robot.
-
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies
Using tokenized proprioception plus instruction to select ~15% of visual patches matches or beats full-token VLA baselines and cuts latency by ~58%.
-
Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt
A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...
-
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
A 0.5B VLA with bridge-conditioned discrete diffusion and 2D temporal–spatial action masking reaches 95.7% LIBERO success and 4.19 CALVIN average length.
-
RynnVLA-002: A Unified Vision-Language-Action and World Model
A single model that jointly predicts robot actions and future images outperforms separate action-only and video-only models on LIBERO and real SO100 manipulation tasks.
-
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
ARFM adaptively adjusts a scaling factor in the flow-matching loss so that offline RL advantage signals are preserved while gradient variance is controlled, improving VLA robot policy fine-tuning.
-
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.
Discussion (0). Continue with ORCID to comment.