Pith. sign in

REVIEW 11 cited by

Learning to Act from Actionless Videos through Dense Correspondences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08576 v1 pith:H4WCIS6E submitted 2023-10-12 cs.RO cs.CVcs.LGstat.ML

classification cs.ROcs.CVcs.LGstat.ML
keywords actionapproachpolicyrobottasksvideoscorrespondencesdense
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we present an approach to construct a video-based robot policy capable of reliably executing diverse tasks across different robots and environments from few video demonstrations without using any action annotations. Our method leverages images as a task-agnostic representation, encoding both the state and action information, and text as a general representation for specifying robot goals. By synthesizing videos that ``hallucinate'' robot executing actions and in combination with dense correspondences between frames, our approach can infer the closed-formed action to execute to an environment without the need of any explicit action labels. This unique capability allows us to train the policy solely based on RGB videos and deploy learned policies to various robotic tasks. We demonstrate the efficacy of our approach in learning policies on table-top manipulation and navigation tasks. Additionally, we contribute an open-source framework for efficient video modeling, enabling the training of high-fidelity policy models with four GPUs within a single day.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Masked Visual Actions for Unified World Modeling

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single video model finetuned on masked pixel trajectories acts as both forward and inverse robot world model, enabling policy evaluation, planning, and action extraction.

  2. On the Sample Efficiency of Inverse Dynamics Models for Semi-Supervised Imitation Learning

    cs.LG 2026-02 conditional novelty 6.0 of 10

    VM-IDM and IDM labeling coincide at infinite unlabeled data, and IDM learning is more label-efficient than BC because the ground-truth IDM is typically less complex and less stochastic than the expert policy.

  3. MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment

    cs.RO 2025-12 conditional novelty 6.0 of 10

    MIND-V generates long-horizon robot manipulation videos by decomposing instructions into sub-tasks with a VLM, encoding plans into a structured bridge, and fine-tuning a video diffusion model with a V-JEPA2-based phys...

  4. Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation

    cs.RO 2025-08 conditional novelty 6.0 of 10

    GVF-TAPE predicts future RGB-D frames from an image and text, then extracts end-effector poses to control a robot, achieving strong success rates without action-labeled data.

  5. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  6. RoboEnvision: A Long-Horizon Video Generation Model for Multi-Task Robot Manipulation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A hierarchical pipeline decomposes long-horizon robot instructions into keyframes, interpolates between them, and regresses joint states from the generated video, reaching 67.4% success on simulated long-horizon tasks.

  7. AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A three-stage pipeline that turns keypoint tracks into discrete motion tokens, predicts them from action-free video, and decodes them into actions yields large few-shot and zero-shot policy improvements in robot manipulation.

  8. Self-Consistent Model-based Adaptation for Visual Reinforcement Learning

    cs.CV 2025-02 conditional novelty 6.0 of 10

    SCMA trains a policy-agnostic observation denoiser, using a pre-trained world model as a clean-distribution reference, and shows improved visual RL performance under distractions.

  9. 3DFlowAction: Learning Cross-Embodiment Manipulation from 3D Flow World Model

    cs.RO 2025-06 conditional novelty 5.0 of 10

    A diffusion world model predicts 3D optical flow as an embodiment-agnostic action plan, and constrained optimization converts the flow into robot arm actions.

  10. GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation

    cs.RO 2025-02 conditional novelty 5.0 of 10

    GEVRM combines text-guided video generation with contrastive state alignment and a diffusion policy, reporting state-of-the-art success rates on perturbed CALVIN manipulation tasks.

  11. From World Models to World Action Models: A Concise Tutorial for Robotics

    cs.RO 2026-07 unverdicted novelty 4.0 of 10

    World models are action-conditioned predictors of task-relevant futures; world action models couple those futures to robot actions via four paradigms: imagine-then-execute, feature-conditioned, joint, and auxiliary pr...

Pith tools