REVIEW 2 cited by
Compositional Foundation Models for Hierarchical Planning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks.
Forward citations
Cited by 2 Pith papers
-
Efficient Diffusion Transformer Policies with Mixture of Expert Denoisers for Multitask Learning
MoDE, a mixture-of-experts diffusion transformer with noise-conditioned routing, reports state-of-the-art results on CALVIN and LIBERO with lower inference FLOPs than dense baselines.
-
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
Task Decodability, a k-NN measure of how separable a task is in a model's middle-layer representations, tracks and predicts in-context learning accuracy, and early-layer finetuning improves it more than late-layer finetuning.
Discussion (0). Continue with ORCID to comment.