REVIEW 2 major objections 5 minor 1 cited by
AMPLIFY: Actionless Motion Priors for Robot Learning from Videos
T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Conditioning a policy on latent motion tokens predicted from video-only training yields 1.2-2.2x low-data gains and the first zero-action-data LIBERO generalization, this paper argues.
desk verdict A well-built latent motion token framework with a genuinely interesting zero-shot LIBERO result, but the headline human-video transfer claim needs a missing control before it can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the latent motion token: single-step velocities extracted from CoTracker tracks of a re-initialized 20 by 20 grid of points, encoded by a causally-masked transformer, quantized with Finite Scalar Quantization (FSQ) into a fixed 2048-code space, and decoded through a local-window classification head that scores each point's next motion inside a 15 by 15 pixel window around its previous location. An autoregressive transformer predicts these tokens from the current image and a language task description, and a separate cross-attention transformer decoder, the inverse dynamics model, maps image, proprioception, and predicted tokens into a distribution over a 16-step action chunk, with temporal ensembling at inference. The discrete codebook and the local-window classification turn motion prediction into a tractable classification problem; the frozen token interface is what allows the forward and inverse stages to train on entirely different datasets.
What would settle it
Construct a pair of tasks whose keypoint motion is nearly identical over the 16-frame prediction horizon but whose correct actions diverge (for example, two LIBERO-Goal tasks that differ only in which of two objects is the target, filmed so the first 0.8 seconds of motion look alike). If AMPLIFY still solves both at a rate near its reported 60% average, the motion-token interface is more informative than the 2D-ambiguity caveat suggests; if success collapses toward the near-zero inverse-only baseline, the information bottleneck sits in the 2D track representation rather than in the action head.
Extended reading notes
Core claim
The central claim is that latent motion tokens are a sufficient interface between visual dynamics and control: single-step velocities of a re-initialized 20 by 20 keypoint grid are compressed into a discrete 2048-code space, a forward model trained only on videos and task descriptions predicts the next token sequence, and an inverse model that never sees the goal decodes those tokens into an action chunk. The paper argues this decomposition cleanly separates what motion defines a task from how a robot can perform it, letting video data and interaction data scale independently. Empirically the learned dynamics are more accurate than prior keypoint and full-video-prediction baselines, and the gains concentrate exactly where action labels are scarce: few-shot learning, cross-embodiment transfer from human video, and zero-shot generalization on LIBERO suites for which no action data was ever provided.
Load-bearing premise
The pipeline stands or falls on the assumption that a re-initialized 20 by 20 grid of 2D keypoint tracks, compressed into 2048 discrete codes, preserves the information needed to choose the right action, and that the point tracker's tracks can be treated as ground truth — even though the inverse model never sees the goal and the paper's Limitations section concedes that 2D tracks leave ambiguity between actions.
Editorial extensions
If this is right
- Few-shot learning scales with video rather than action count: with only 2 demonstrations per task, AMPLIFY reaches roughly 1.94x the success rate of the keypoint baseline ATM on LIBERO.
- Action-free human video transfers to robot control: adding human demonstrations to the forward model improves real-world success by 1.32x to 1.5x over the same policy fed only robot data.
- Zero-action-data generalization is possible: trained on LIBERO-90 actions only, AMPLIFY averages 60.5% success on four unseen LIBERO suites while behavior-cloning baselines score near zero.
- Latent motion tokens function as a world-model artifact: conditioning the AVDC video predictor on them improves PSNR, LPIPS, and SSIM on BridgeData v2.
- More video, same actions: with action data held at 2 trajectories, task success rises from 0.12 with 2 training videos to 0.55 with 50, evidence that the prior keeps improving as video scales.
Reading between the lines
- A consequence the paper motivates but does not test: if the motion token is the whole task channel, the inverse model should train on undirected play or exploration data as well as on expert demos, since it never sees goals.
- The 2D ambiguity the paper concedes predicts where the method should fail — tasks whose distinguishing information sits in the goal rather than in pixel motion; the LIBERO-Goal set being the hardest zero-shot target (41%) is consistent with that, and swapping in 3D point tracking would be the direct stress test.
- The interface framing suggests motion tokens could serve as a reference channel inside other policy families, since the ablations show Gaussian, diffusion, and flow-matching action heads perform nearly identically.
- The non-monotonic video-scaling curve (0.34 at 5 videos vs 0.23 at 10) hints that the forward model's benefit depends on track diversity, not raw clip count; a scaling study over internet-scale, scene-diverse video would show whether the cross-embodiment gain saturates.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces AMPLIFY, a three-stage framework for robot learning from action-free videos. It first tokenizes CoTracker keypoint trajectories into discrete FSQ codes via an autoencoder, then trains an autoregressive forward dynamics model to predict future motion tokens from an image and task description, and finally trains an inverse dynamics model to map predicted tokens to actions. Experiments evaluate track prediction accuracy on LIBERO, BridgeData v2, and Something-Something v2; downstream policy learning in in-distribution, few-shot, cross-embodiment, and zero-shot action-label settings; and conditional video prediction. The reported results include a 3.7x reduction in track prediction MSE over ATM, 1.2-2.2x improvements in few-shot policy learning, a 1.4x average real-world improvement when human videos are added, and 60.5% average success on LIBERO target tasks with zero target-task action data.
Significance. If the results hold, AMPLIFY makes a strong contribution by demonstrating that a compact latent motion representation learned from arbitrary videos can provide a useful prior for downstream control, enabling substantial gains in low-data regimes and a novel zero-shot action-label generalization result. The paper is well-structured, includes extensive ablations, and compares against strong baselines. The few-shot and zero-shot experiments include control variants (inverse-only and w/o-tracks) that isolate the motion-token contribution, and the video-scaling experiment in Appendix F.1 is a valuable diagnostic. However, the headline cross-embodiment claim is not supported by the presented comparison, and the track-prediction evaluation split is missing.
major comments (2)
- [§3.2, Table 4; Appendix F.1] The cross-embodiment transfer claim is not isolated: the 1.4x average improvement compares an AMPLIFY policy whose forward dynamics is trained on human and robot videos against a Diffusion Policy trained only on robot demonstrations. This differs in two ways at once (motion-token conditioning and additional video data), so the gain cannot be attributed to human videos. The paper's own video-scaling experiment (Table 14) shows that adding robot videos alone raises success from 0.12 to 0.55 in a low-data LIBERO setting, demonstrating that extra video data alone can produce large gains. To support the claim that action-free human videos provide a cross-embodiment benefit, the authors should add an AMPLIFY variant whose forward dynamics is trained only on robot videos and show that the addition of human videos yields a further improvement over that control.
- [§3.1, Table 2; Appendix D] The track-prediction evaluation split is not stated anywhere in the paper. It is unclear whether the forward dynamics model is evaluated on videos that were also used for training; if it is, the reported 3.7x MSE improvement and 2.5x pixel-accuracy improvement over ATM are not evidence of generalization. The paper must specify how the videos in each dataset (LIBERO, BridgeData v2, Something-Something v2) were partitioned into training and evaluation sets for the motion tokenizer and forward dynamics model, and confirm that the evaluation rollouts are disjoint from the training rollouts.
minor comments (5)
- [Section 1, Contribution 1] The claim of 'the first latent keypoint dynamics model' should be qualified with respect to prior work such as Moto [34] and Latent Action Pretraining [33], which also learn latent motion or action representations from videos; the novelty should be positioned more precisely.
- [Section 3.2, Table 4] The sentence 'The average improvements of 1.32×, 1.4×, and 1.5×' is confusing because the table reports a single average of 0.42 vs 0.58; the per-task or per-demonstration-count ratios should be reported explicitly.
- [Section 3.2, Cross-Embodiment Transfer] The text says 'we evaluate AMPLIFY in both the few-shot setting and the full demonstration setting,' but the main text does not define the exact demonstration counts for 'All' for each task; these are only listed in Appendix Table 9, so a cross-reference would help.
- [Section 3.2, real-world results] Success rates are reported without confidence intervals despite being averaged over 10 rollouts; standard errors or per-seed results would make the comparison more rigorous.
- [Appendix D.4] The preprocessing paragraph introduces the window subscript without defining it; e.g., 'length-τ video' and 'τ length-T windows' should be defined precisely.
Circularity Check
No significant circularity: AMPLIFY's modular forward/inverse dynamics chain is independently benchmarked; the Table 4 human-video confound is an attribution gap, not a constructed equivalence.
full rationale
AMPLIFY's derivation chain is modular and not circular by construction. Motion tokens are produced by an FSQ autoencoder from CoTracker keypoint velocities (Eq. 1); the forward dynamics model predicts those tokens from image and language (Eq. 2); the inverse dynamics model maps tokens to action chunks using an external NLL action loss (Eq. 3). Each stage is trained on its own target, and none of the paper's equations define a predicted quantity as the fitted input. Downstream policy claims are checked against external baselines (Diffusion Policy, ATM, Track2Act, UniPi, BAKU, QueST) and include control variants (inverse-only, w/o tracks) that isolate the motion-token contribution. The only self-citation is QueST, which appears solely as a baseline in Table 3 and is not load-bearing for any method choice or uniqueness argument. The Limitations section openly acknowledges the 2D-track ambiguity, which is a scope caveat rather than a circular step. The cross-embodiment comparison in Table 4 lacks an AMPLIFY variant trained without human videos, so the 1.4x gain over Diffusion Policy is confounded by added robot-video data; Table 14 shows robot-video scaling alone can improve success. This is an experimental attribution gap and a correctness risk, not a construction-level circularity: the claim is not equivalent to its inputs by definition or equation. The paper is therefore self-contained against external benchmarks, with no significant circularity.
Assumptions & free parameters
free parameters (8)
- Prediction horizon T =
16
- Local window size W =
15
- FSQ codebook size =
2048
- Latent code sequence length =
16
- Hidden dimension =
768
- Transformer depths =
2 tokenizer, 8 forward, 4 inverse
- Keypoint grid size N =
400 (20x20)
- Action loss temporal discount gamma =
0.99
assumptions (6)
- domain assumption CoTracker tracks are treated as ground truth for all downstream training and evaluation.
- domain assumption A reinitialized uniform 20x20 grid captures task-relevant motion across diverse tasks and embodiments.
- domain assumption 2D pixel-space tracks suffice to disambiguate robot actions.
- domain assumption Environment dynamics are deterministic given observation and goal.
- domain assumption Human video motion priors transfer to robot action inference.
- standard math FSQ and transformer training assumptions from prior work hold as described.
invented entities (1)
-
Latent motion tokens (FSQ-quantized keypoint velocity codes)
Cite this review
Pith. "Pith review of AMPLIFY: Actionless Motion Priors for Robot Learning from Videos." pith.science (2026). https://pith.science/paper/IBGA2462
@misc{pith2026250614198,
author = {Pith},
title = {Pith review of: AMPLIFY: Actionless Motion Priors for Robot Learning from Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBGA2462}},
note = {Machine review of arXiv:2506.14198}
}
read the original abstract
Action-labeled data for robotics is scarce and expensive, limiting the generalization of learned policies. In contrast, vast amounts of action-free video data are readily available, but translating these observations into effective policies remains a challenge. We introduce AMPLIFY, a novel framework that leverages large-scale video data by encoding visual dynamics into compact, discrete motion tokens derived from keypoint trajectories. Our modular approach separates visual motion prediction from action inference, decoupling the challenges of learning what motion defines a task from how robots can perform it. We train a forward dynamics model on abundant action-free videos and an inverse dynamics model on a limited set of action-labeled examples, allowing for independent scaling. Extensive evaluations demonstrate that the learned dynamics are both accurate, achieving up to 3.7x better MSE and over 2.5x better pixel prediction accuracy compared to prior approaches, and broadly useful. In downstream policy learning, our dynamics predictions enable a 1.2-2.2x improvement in low-data regimes, a 1.4x average improvement by learning from action-free human videos, and the first generalization to LIBERO tasks from zero in-distribution action data. Beyond robotic control, we find the dynamics learned by AMPLIFY to be a versatile latent world model, enhancing video prediction quality. Our results present a novel paradigm leveraging heterogeneous data sources to build efficient, generalizable world models. More information can be found at https://amplify-robotics.github.io/.
Figures
Figures from the paper (8 more)
Forward citations
Cited by 1 Pith paper
-
3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos
Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.
Reference graph
Works this paper leans on
-
[1]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. 2019
2019
-
[2]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Nee- lakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Rad- ford, I. Sutskever, and D. Am...
1901
- [3]
-
[4]
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision, 2021. URLhttps://arxiv.org/abs/2103.00020
arXiv 2021
- [5]
-
[6]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models, 2022. URL https://arxiv.org/abs/2112. 10752
2022
-
[7]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manju- nath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...
- [8]
Show all 107 references
-
[9]
Padalkar, A
A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023
2023 arXiv
-
[11]
Pertsch, K
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025
2025 arXiv
-
[12]
L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tanner, A. Walling, H. Wang, N. Fusai, et al. Hi robot: Open-ended instruction following with hierarchical vision-language- action models.arXiv preprint arXiv:2502.19417, 2025
2025 arXiv
-
[13]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...
2024
-
[14]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An Open-Source Vision-Language-Action Model, June
-
[15]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky. $pi_0$: A Visi...
2024 arXiv
-
[16]
Y . Du, M. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel. Learning Universal Policies via Text-Guided Video Generation, Nov. 2023. URL http: //arxiv.org/abs/2302.00111. arXiv:2302.00111 [cs]
2023 arXiv
-
[17]
McCarthy, D
R. McCarthy, D. C. H. Tan, D. Schmidt, F. Acero, N. Herr, Y . Du, T. G. Thuruthel, and Z. Li. Towards Generalist Robot Learning from Internet Video: A Survey, June 2024. URL http://arxiv.org/abs/2404.19664. arXiv:2404.19664 [cs]
2024 arXiv
-
[18]
Rybkin, K
O. Rybkin, K. Pertsch, K. G. Derpanis, K. Daniilidis, and A. Jaegle. Learning what you can do before doing anything, Feb. 2019. URL http://arxiv.org/abs/1806.09655. arXiv:1806.09655 [cs, stat]. 10
2019 arXiv
-
[19]
Agarwal, A
NVIDIA, N. Agarwal, A. Ali, M. Bala, Y . Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y . Chen, Y . Cui, Y . Ding, D. Dworakowski, J. Fan, M. Fenzi, F. Ferroni, S. Fidler, D. Fox, S. Ge, Y . Ge, J. Gu, S. Gururani, E. He, J. Huang, J. Huffman, P. Jannaty, J. Jin, S. W. Kim, G....
2025 arXiv
-
[20]
J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans. Imagen video: High definition video generation with diffusion models, 2022. URLhttps://arxiv.org/abs/2210.02303
2022 arXiv
-
[21]
W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024
2024 arXiv
-
[22]
Gupta, A
Veo-Team, :, A. Gupta, A. Razavi, A. Toor, A. Gupta, D. Erhan, E. Shaw, E. Lau, F. Belletti, G. Barth-Maron, G. Shaw, H. Erdogan, H. Sidahmed, H. Nandwani, H. Moraldo, H. Kim, I. Blok, J. Donahue, J. Lezama, K. Mathewson, K. David, M. K. Lorrain, M. van Zee, M. Narasimhan, M. ...
2024
-
[23]
Bar-Tal, H
O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, G. Liu, A. Raj, et al. Lumiere: A space-time diffusion model for video generation. InSIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024
2024
-
[24]
Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training.arXiv preprint arXiv:2210.00030, 2022
2022 arXiv
-
[25]
Ghosh, C
D. Ghosh, C. Bhateja, and S. Levine. Reinforcement Learning from Passive Data via Latent Intentions, Apr. 2023. URLhttp://arxiv.org/abs/2304.04782. arXiv:2304.04782 [cs, stat]
2023 arXiv
-
[26]
Bhateja, D
C. Bhateja, D. Guo, D. Ghosh, A. Singh, M. Tomar, Q. Vuong, Y . Chebotar, S. Levine, and A. Kumar. Robotic Offline RL from Internet Videos via Value-Function Pre-Training, Sept
-
[27]
Dashora, D
N. Dashora, D. Ghosh, and S. Levine. Viva: Video-trained value functions for guiding online rl from diverse data.arXiv preprint arXiv:2503.18210, 2025
2025 arXiv
-
[28]
Sermanet, C
P. Sermanet, C. Lynch, Y . Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain. Time-contrastive networks: Self-supervised learning from video. In2018 IEEE international conference on robotics and automation (ICRA), pages 1134–1141. IEEE, 2018
2018
-
[29]
S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3m: A universal visual representa- tion for robot manipulation.arXiv preprint arXiv:2203.12601, 2022
2022 arXiv
-
[30]
T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, et al. Model-based reinforcement learning: A survey.Foundations and Trends® in Machine Learning, 16(1):1–118, 2023. 11
2023
-
[31]
Y . Hu, Y . Guo, P. Wang, X. Chen, Y .-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint arXiv:2412.14803, 2024
2024 arXiv
-
[32]
Bruce, M
J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning, 2024
2024
-
[33]
S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, et al. Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2024
2024 arXiv
-
[34]
Y . Chen, Y . Ge, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation, Dec. 2024. URL http://arxiv.org/abs/ 2412.04445. arXiv:2412.04445 [cs]
2024
-
[35]
Z. Cui, H. Pan, A. Iyer, S. Haldar, and L. Pinto. Dynamo: In-domain dynamics pretraining for visuo-motor control.Advances in Neural Information Processing Systems, 37:33933–33961, 2024
2024
-
[36]
Karaev, I
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. CoTracker: It is better to track together. 2023
2023
-
[37]
Wang, Y .-Y
Q. Wang, Y .-Y . Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely. Tracking everything everywhere all at once. InInternational Conference on Computer Vision, 2023
2023
-
[38]
Doersch, Y
C. Doersch, Y . Yang, M. Vecerik, D. Gokay, A. Gupta, Y . Aytar, J. Carreira, and A. Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10061–10072, 2023
2023
-
[39]
Y . Xiao, Q. Wang, S. Zhang, N. Xue, S. Peng, Y . Shen, and X. Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[40]
Vecerik, C
M. Vecerik, C. Doersch, Y . Yang, T. Davchev, Y . Aytar, G. Zhou, R. Hadsell, L. Agapito, and J. Scholz. RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation, Aug. 2023. URLhttp://arxiv.org/abs/2308.15975. arXiv:2308.15975 [cs]
2023 arXiv
-
[41]
Z. Qin, K. Fang, Y . Zhu, L. Fei-Fei, and S. Savarese. Keto: Learning keypoint representations for tool manipulation. In2020 IEEE International Conference on Robotics and Automation (ICRA), pages 7278–7285. IEEE, 2020
2020
-
[42]
C. Yuan, C. Wen, T. Zhang, and Y . Gao. General flow as foundation affordance for scalable robot learning.arXiv preprint arXiv:2401.11439, 2024
2024 arXiv
-
[43]
L.-H. Lin, Y . Cui, A. Xie, T. Hua, and D. Sadigh. FlowRetrieval: Flow-Guided Data Retrieval for Few-Shot Imitation Learning, Oct. 2024. URL http://arxiv.org/abs/2408. 16944. arXiv:2408.16944
2024 arXiv
-
[44]
P.-C. Ko, J. Mao, Y . Du, S.-H. Sun, and J. B. Tenenbaum. Learning to Act from Actionless Videos through Dense Correspondences.arXiv:2310.08576, 2023
2023 arXiv
-
[45]
Kareer, D
S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu. Egomimic: Scaling imitation learning via egocentric video, 2024. URL https://arxiv. org/abs/2410.24221
2024 arXiv
-
[46]
J. Gao, Z. Tao, N. Jaquier, and T. Asfour. K-VIL: Keypoints-Based Visual Imitation Learn- ing.IEEE Transactions on Robotics, 39(5):3888–3908, Oct. 2023. ISSN 1941-0468. doi: 10.1109/TRO.2023.3286074. URL https://ieeexplore.ieee.org/abstract/ document/10189175. Conference Name:...
2023
-
[47]
Fang, B.-R
X. Fang, B.-R. Huang, J. Mao, J. Shone, J. B. Tenenbaum, T. Lozano-Pérez, and L. P. Kaelbling. Keypoint Abstraction using Large Models for Object-Relative Imitation Learning, Oct. 2024. URLhttp://arxiv.org/abs/2410.23254. arXiv:2410.23254
2024 arXiv
-
[48]
C. Gao, H. Zhang, Z. Xu, C. Zhehao, and L. Shao. Flip: Flow-centric generative planning as general-purpose manipulation world model. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[49]
Manuelli, W
L. Manuelli, W. Gao, P. Florence, and R. Tedrake. kpam: Keypoint affordances for category- level robotic manipulation. InThe International Symposium of Robotics Research, pages 132–157. Springer, 2019
2019
-
[50]
Guzey, Y
I. Guzey, Y . Dai, G. Savva, R. Bhirangi, and L. Pinto. Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards, Oct. 2024. URL http://arxiv.org/abs/ 2410.23289. arXiv:2410.23289
2024 arXiv
-
[51]
M. Xu, Z. Xu, Y . Xu, C. Chi, G. Wetzstein, M. Veloso, and S. Song. Flow as the Cross-Domain Manipulation Interface, July 2024. URL http://arxiv.org/abs/2407.15208. arXiv:2407.15208 [cs]
2024 arXiv
-
[52]
Bharadhwaj, D
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation, Sept. 2024. URL https://arxiv.org/abs/2409. 16283v1
2024
-
[53]
Bharadhwaj, R
H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani. Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation.arXiv preprint arXiv:2405.01527, 2024
2024 arXiv
-
[54]
C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y . Gao, and P. Abbeel. Any-point Trajectory Mod- eling for Policy Learning, Feb. 2024. URL http://arxiv.org/abs/2401.00025. arXiv:2401.00025 [cs]
2024 arXiv
-
[55]
Hansen, X
N. Hansen, X. Wang, and H. Su. Temporal difference learning for model predictive control. In ICML, 2022
2022
-
[56]
Hansen, H
N. Hansen, H. Su, and X. Wang. TD-MPC2: Scalable, Robust World Models for Continuous Control, Mar. 2024. URL http://arxiv.org/abs/2310.16828. arXiv:2310.16828 [cs]
2024 arXiv
-
[57]
Scannell, M
A. Scannell, M. Nakhaei, K. Kujanpää, Y . Zhao, K. S. Luck, A. Solin, and J. Pajarinen. Discrete codebook world models for continuous control.arXiv preprint arXiv:2503.00653, 2025
2025 arXiv
-
[58]
Mentzer, D
F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen. Finite Scalar Quantization: VQ-V AE Made Simple, Oct. 2023. URL http://arxiv.org/abs/2309.15505. arXiv:2309.15505 [cs]
2023 arXiv
-
[59]
A. v. d. Oord, O. Vinyals, and K. Kavukcuoglu. Neural Discrete Representation Learning, May 2018. URLhttp://arxiv.org/abs/1711.00937. arXiv:1711.00937 [cs]
2018 arXiv
-
[60]
K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, June
-
[61]
Raffel, N
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020. 13
2020
-
[62]
Haldar, Z
S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learning. arXiv preprint arXiv:2406.07539, 2024
2024 arXiv
-
[63]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
-
[64]
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V . Myers, M. J. Kim, M. Du, et al. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning, pages 1723–1736. PMLR, 2023
2023
-
[65]
something something
R. Goyal, S. E. Kahou, V . Michalski, J. Materzy ´nska, S. Westphal, H. Kim, V . Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic. The "something something" video database for learning and evaluating visual common sense,
-
[66]
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023
2023 arXiv
-
[67]
X. Gu, C. Wen, W. Ye, J. Song, and Y . Gao. Seer: Language instructed video prediction with latent diffusion models.arXiv preprint arXiv:2303.14897, 2023
2023 arXiv
-
[69]
A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg. QueST: Self-Supervised Skill Abstractions for Learning Continuous Control, Sept. 2024. URL http://arxiv.org/abs/2407. 15840. arXiv:2407.15840 [cs]
2024 arXiv
-
[70]
T. D. Ngo, P. Zhuang, C. Gan, E. Kalogerakis, S. Tulyakov, H.-Y . Lee, and C. Wang. Delta: Dense efficient long-range 3d tracking for any video.arXiv preprint arXiv:2410.24211, 2024
2024 arXiv
-
[71]
Misra, A
D. Misra, A. Saran, T. Xie, A. Lamb, and J. Langford. Towards principled representation learning from videos for reinforcement learning.arXiv preprint arXiv:2403.13765, 2024
2024 arXiv
-
[72]
M. Yang, D. Schuurmans, P. Abbeel, and O. Nachum. Dichotomy of control: Separating what you can control from what you cannot.arXiv preprint arXiv:2210.13435, 2022
2022 arXiv
-
[73]
S. Park, D. Ghosh, B. Eysenbach, and S. Levine. Hiql: Offline goal-conditioned rl with latent states as actions.Advances in Neural Information Processing Systems, 36:34866–34891, 2023
2023
-
[74]
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play.arXiv preprint arXiv:2302.12422, 2023
2023 arXiv
-
[75]
L. Chen, S. Bahl, and D. Pathak. PlayFusion: Skill Acquisition via Diffusion from Language- Annotated Play. InProceedings of The 7th Conference on Robot Learning, pages 2012–2029. PMLR, Dec. 2023. URL https://proceedings.mlr.press/v229/chen23c. html. ISSN: 2640-3498
2012
-
[76]
Lynch, M
C. Lynch, M. Khansari, T. Xiao, V . Kumar, J. Tompson, S. Levine, and P. Sermanet. Learning Latent Plans from Play. InProceedings of the Conference on Robot Learning, pages 1113–1132. PMLR, May 2020. URL https://proceedings.mlr.press/v100/lynch20a. html. ISSN: 2640-3498
2020
-
[77]
Bharadhwaj, A
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V . Kumar. Zero-shot robot manipulation from passive human videos.arXiv preprint arXiv:2302.02011, 2023
2023 arXiv
-
[78]
Xiong, Q
H. Xiong, Q. Li, Y .-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg. Learning by watching: Physical imitation of manipulation skills from human videos. In2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7827–7834. IEEE, 2021. 14
2021
-
[79]
K. Shaw, S. Bahl, and D. Pathak. Videodex: Learning dexterity from internet videos. In Conference on Robot Learning, pages 654–665. PMLR, 2023
2023
-
[80]
Mandikal and K
P. Mandikal and K. Grauman. Dexvip: Learning dexterous grasping with human hand pose priors from video. InConference on Robot Learning, pages 651–661. PMLR, 2022
2022
-
[81]
S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild.arXiv preprint arXiv:2207.09450, 2022
2022 arXiv
-
[82]
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak. Affordances from human videos as a versatile representation for robotics. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023
2023
-
[83]
Mendonca, S
R. Mendonca, S. Bahl, and D. Pathak. Structured world models from human videos.arXiv preprint arXiv:2308.10901, 2023
2023 arXiv
-
[84]
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al. Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. InConference on Robot Learning, pages 1303–1315. PMLR, 2022
2022
-
[85]
Schmeckpeper, O
K. Schmeckpeper, O. Rybkin, K. Daniilidis, S. Levine, and C. Finn. Reinforcement learning with videos: Combining offline observations with interaction.arXiv preprint arXiv:2011.06507, 2020
2011 arXiv
-
[86]
Bhateja, D
C. Bhateja, D. Guo, D. Ghosh, A. Singh, M. Tomar, Q. Vuong, Y . Chebotar, S. Levine, and A. Kumar. Robotic offline rl from internet videos via value-function pre-training.arXiv preprint arXiv:2309.13041, 2023
2023 arXiv
-
[87]
Ghosh, C
D. Ghosh, C. A. Bhateja, and S. Levine. Reinforcement learning from passive data via latent intentions. InInternational Conference on Machine Learning, pages 11321–11339. PMLR, 2023
2023
-
[88]
Bruce, M
J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, an...
2024 arXiv
-
[89]
Escontrela, A
A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y . Lee, D. Hafner, and P. Abbeel. Video prediction models as rewards for reinforcement learning.arXiv preprint arXiv:2305.14343, 2023
2023 arXiv
-
[90]
Huang, G
T. Huang, G. Jiang, Y . Ze, and H. Xu. Diffusion reward: Learning rewards via conditional video diffusion.arXiv preprint arXiv:2312.14134, 2023
2023 arXiv
-
[91]
Schmidt and M
D. Schmidt and M. Jiang. Learning to act without actions. InThe Twelfth International Conference on Learning Representations (ICLR), 2024
2024
-
[92]
Y . Wen, J. Lin, Y . Zhu, J. Han, H. Xu, S. Zhao, and X. Liang. Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation, 2024. URL https: //arxiv.org/abs/2411.09153
2024 arXiv
-
[93]
Liang, R
J. Liang, R. Liu, E. Ozguroglu, S. Sudhakar, A. Dave, P. Tokmakov, S. Song, and C. V ondrick. Dreamitate: Real-world visuomotor policy learning via video generation, 2024. URL https: //arxiv.org/abs/2406.16862
2024 arXiv
-
[94]
Jaegle, F
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira. Perceiver: General Perception with Iterative Attention, June 2021. URL http://arxiv.org/abs/2103. 03206. arXiv:2103.03206 [cs, eess]. 15
2021 arXiv
-
[95]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[96]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[97]
A. Veit, M. Wilber, and S. Belongie. Residual Networks Behave Like Ensembles of Rel- atively Shallow Networks, Oct. 2016. URL http://arxiv.org/abs/1605.06431. arXiv:1605.06431 [cs]
2016 arXiv
-
[98]
Lipman, R
Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling, 2023. URLhttps://arxiv.org/abs/2210.02747
2023 arXiv
-
[99]
Bjorck, F
NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, ...
2025 arXiv
-
[100]
C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, Mar. 2023. URL http://arxiv.org/ abs/2303.04137. arXiv:2303.04137 [cs]. 16 Appendix In this document, we provide detailed supplementary m...
2023 arXiv
-
[106]
Require:DatasetsV,R 1:Preprocess keypoint tracksκ t inV 2:Learn latent motion encoding to compressκ t into discrete tokensz t using Eq
Use Action-Free Videos (and Expert Demonstrations) to learn how observations evolve with respect to a goal 17 Algorithm 1AMPLIFYTraining. Require:DatasetsV,R 1:Preprocess keypoint tracksκ t inV 2:Learn latent motion encoding to compressκ t into discrete tokensz t using Eq. 1 3...
-
[107]
Put the Rubik’s Cube on the Box
Use Interaction Data (Undirected and Demonstrations) to learn how a sequence of observations maps to a sequence of actions This decomposition effectively decouplestask understanding(the sequence of observations that correspond to a goal) andtask execution(translating a referen...
-
[108]
MSE: We take normalized track predictions in the range [−1,1] and compute MSE between predicted(x, y)values for each point and the corresponding ground-truth point from CoTracker
-
[109]
3.∆ AUC: This metric was originally introduced by works presenting point tracking [ 38, 36] and later used for track prediction [53]
Pixel Accuracy: We measure the (normalized) percentage of predictions that are pixel-perfect compared to ground-truth. 3.∆ AUC: This metric was originally introduced by works presenting point tracking [ 38, 36] and later used for track prediction [53]. The metric is computed a...
-
[110]
ground truth
video generation model by conditioning it on the motion tokens produced by the forward dynamics model. These motion tokens serve as a conditioning signal that guides video prediction based on the expected dynamics. To condition the video generation model, motion tokens generat...
- [2016]
-
[2017]
URLhttps://arxiv.org/abs/1706.04261
-
[2020]
URL https://proceedings.neurips.cc/paper_files/paper/2020/ file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
- [2024]
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.