REVIEW 7 cited by
MetaMorph: Learning Universal Controllers with Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multiple domains like vision, natural language, and audio are witnessing tremendous progress by leveraging Transformers for large scale pre-training followed by task specific fine tuning. In contrast, in robotics we primarily train a single robot for a single task. However, modular robot systems now allow for the flexible combination of general-purpose building blocks into task optimized morphologies. However, given the exponentially large number of possible robot morphologies, training a controller for each new design is impractical. In this work, we propose MetaMorph, a Transformer based approach to learn a universal controller over a modular robot design space. MetaMorph is based on the insight that robot morphology is just another modality on which we can condition the output of a Transformer. Through extensive experiments we demonstrate that large scale pre-training on a variety of robot morphologies results in policies with combinatorial generalization capabilities, including zero shot generalization to unseen robot morphologies. We further demonstrate that our pre-trained policy can be used for sample-efficient transfer to completely new robot morphologies and tasks.
Forward citations
Cited by 7 Pith papers
-
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design
A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.
-
UniVideo: Unified Understanding, Generation, and Editing for Videos
UniVideo combines a frozen MLLM and a video DiT to unify video understanding, generation, in-context editing, visual prompting, and zero-shot free-form video edits under one instruction interface.
-
UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
Embodiment-Aware Diffusion Policy steers a UMI-trained diffusion policy with controller tracking-cost gradients at inference time, improving aerial manipulation success in simulation and real flights.
-
Generalized Locomotion in Out-of-distribution Conditions with Robust Transformer
A transformer with body tokenization and consistent dropout generalizes to unseen leg damages and sensor noise while trained on limited dynamics and clean observations.
-
AnyBody: A Benchmark Suite for Cross-Embodiment Manipulation
AnyBody is a benchmark suite that tests cross-embodiment manipulation generalization along interpolation, extrapolation, and composition axes, and finds zero-shot generalization to unseen robot bodies remains difficult.
-
McARL:Morphology-Control-Aware Reinforcement Learning for Generalizable Quadrupedal Locomotion
A morphology-conditioned PPO policy trained only on the Unitree Go1 transfers zero-shot in simulation to Go2, A1 and Mini Cheetah, with the best variant reaching 3.5 m/s on the Go2.
-
UniLegs: Universal Multi-Legged Robot Control through Morphology-Agnostic Policy Distillation
A two-stage teacher-student distillation produces a single Transformer policy that reaches 94.47% of specialist teacher reward on five training morphologies and 72.64% on an unseen quadruped.
Discussion (0). Continue with ORCID to comment.