REVIEW 11 cited by
Learning to Act from Actionless Videos through Dense Correspondences
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this work, we present an approach to construct a video-based robot policy capable of reliably executing diverse tasks across different robots and environments from few video demonstrations without using any action annotations. Our method leverages images as a task-agnostic representation, encoding both the state and action information, and text as a general representation for specifying robot goals. By synthesizing videos that ``hallucinate'' robot executing actions and in combination with dense correspondences between frames, our approach can infer the closed-formed action to execute to an environment without the need of any explicit action labels. This unique capability allows us to train the policy solely based on RGB videos and deploy learned policies to various robotic tasks. We demonstrate the efficacy of our approach in learning policies on table-top manipulation and navigation tasks. Additionally, we contribute an open-source framework for efficient video modeling, enabling the training of high-fidelity policy models with four GPUs within a single day.
Forward citations
Cited by 11 Pith papers
-
Masked Visual Actions for Unified World Modeling
A single video model finetuned on masked pixel trajectories acts as both forward and inverse robot world model, enabling policy evaluation, planning, and action extraction.
-
On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning
VM-IDM and IDM labeling coincide at infinite unlabeled data, and IDM learning is more label-efficient than BC because the ground-truth IDM is typically less complex and less stochastic than the expert policy.
-
MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment
MIND-V generates long-horizon robot manipulation videos by decomposing instructions into sub-tasks with a VLM, encoding plans into a structured bridge, and fine-tuning a video diffusion model with a V-JEPA2-based phys...
-
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
GVF-TAPE predicts future RGB-D frames from an image and text, then extracts end-effector poses to control a robot, achieving strong success rates without action-labeled data.
-
Precise Action-to-Video Generation Through Visual Action Prompts
Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.
-
RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation
A hierarchical pipeline decomposes long-horizon robot instructions into keyframes, interpolates between them, and regresses joint states from the generated video, reaching 67.4% success on simulated long-horizon tasks.
-
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.
-
Self-Consistent Model-based Adaptation for Visual Reinforcement Learning
SCMA trains a policy-agnostic observation denoiser, using a pre-trained world model as a clean-distribution reference, and shows improved visual RL performance under distractions.
-
3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model
A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.
-
GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
GEVRM combines text-guided video generation with contrastive state alignment and a diffusion policy, reporting state-of-the-art success rates on perturbed CALVIN manipulation tasks.
-
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are action-conditioned predictors of task-relevant futures; world action models couple those futures to robot actions via four paradigms: imagine-then-execute, feature-conditioned, joint, and auxiliary pr...
Discussion (0). Sign in to comment.