REVIEW 4 major objections 5 minor 3 cited by
ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Adding a self-supervised objective that predicts future observations in latent space to an action-chunking policy improves the policy's world model and keeps task success at or above supervised baselines.
desk verdict A genuinely new combination of action chunking and JEPA-style latent observation prediction, cleanly described and honestly limited, but the world-model claim rests on a partly circular probe and the abstract oversells the numbers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is the joint-embedding predictive head: a predictor that turns the context encoder's output plus learnable mask tokens into a sequence of abstract future-observation embeddings, matched against targets from an exponentially moving average target encoder. This is JEPA, a Joint-Embedding Predictive Architecture, which predicts in latent space to discard irrelevant detail. The second load-bearing idea is chunking: predicting a whole sequence of n future states and n future actions rather than one step, which is what lets the same representation support both the action policy and the world model. The cross-attention conditioning keeps the predictor's cost linear in context length, which the paper cites as an efficiency advantage over self-attention over all inputs.
What would settle it
On a task suite where a held-out sensor stream, such as depth images, is recorded alongside the trained proprioceptive states, train ACT-JEPA and an action-only baseline with identical data, then probe both frozen encoders to reconstruct the held-out stream. If the JEPA-trained encoder does not reconstruct the held-out stream substantially better, the claim that latent observation prediction improves world-model understanding fails.
Extended reading notes
Core claim
The central result is that a policy trained to predict both actions and future abstract observations develops a better internal representation of environment dynamics than a policy trained on actions alone. In ACT-JEPA, the context encoder takes the current image, proprioceptive state, and a task label and outputs a single context token; the target encoder embeds a sequence of future proprioceptive states; the predictor, conditioned on the context through cross-attention, predicts those future embeddings from learnable mask tokens; and the action decoder maps the same context plus mask tokens to a chunk of future actions. The model is trained end-to-end by summing an L1 action loss and an L1 latent-observation loss, with the target encoder updated by exponential moving average of the context encoder. The paper's headline measurements are that the frozen context encoder supports lower reconstruction error than the baseline's encoder, with RMSE 3.696 versus 4.219 and ATE 6.521 versus 7.393, and that the full policy reaches 91.6% success, slightly above the strongest baseline at 91.1%. This is presented as evidence that jointly predicting actions and abstract observation sequences improves policy representation and world-model understanding while remaining competitive on the task.
Load-bearing premise
The paper's case for an improved world model rests on a probe that reconstructs the same proprioceptive states the model was trained to predict, so the probe may reward memorizing the training target rather than genuine understanding of environment dynamics.
Editorial extensions
If this is right
- If the central claim is right, policies trained with a joint action-plus-latent-observation objective should transfer to downstream tasks better than action-only behavior cloning, as the probing results indicate.
- The pretraining experiment suggests that large unlabeled observation collections can build the policy's context encoder, with expert action labels needed only in a short fine-tuning phase.
- The parity in success rate implies that adding a world-model objective does not force a trade-off against task performance, at least in the low-data regime tested.
- Because the model predicts whole observation sequences rather than a single future frame, its world model should be richer than single-step dynamics models, which should matter more in complex, long-horizon tasks.
Reading between the lines
- The probe's target modality is the same proprioceptive states used in the training objective, so the reported representation gain may partly reflect memorization of the prediction target; a held-out modality would test whether the world model generalizes.
- The architecture's cross-attention predictor scales linearly in context length, so a natural extension is to test whether longer prediction horizons and longer observation histories continue to improve the representation.
- The paper's closing speculation suggests that abstract state representations could transfer across robot embodiments with similar limb configurations; that is testable by pretraining on one embodiment and fine-tuning on another.
- In the current benchmark most tasks saturate near 100% success, so the practical advantage of the world-model objective is more likely to appear on harder or more diverse tasks than the ones used here.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes ACT-JEPA, a policy-learning architecture that combines imitation learning (action chunking, as in ACT) with a Joint-Embedding Predictive Architecture (JEPA). The model is trained end-to-end with two losses: an action-reconstruction loss and an observation-prediction loss in latent space, where the latter is intended to build a world model. The authors formulate three hypotheses: (H1) predicting abstract observation sequences improves representation quality and environment-dynamics understanding, (H2) such representations transfer to action-sequence prediction, and (H3) ACT-JEPA performs comparably to established supervised baselines. Experiments on Meta-World (15 tasks) compare ACT-JEPA against a regression-based transformer (RBC) and ACT, reporting a 12.4% RMSE improvement and 11.8% ATE improvement on a probing task, a 91.6% vs. 91.1% success rate difference, and a decreasing action-reconstruction loss in a pretraining/fine-tuning protocol. The authors conclude that all three hypotheses are supported and that ACT-JEPA offers an efficient policy-representation learning method.
Significance. If the central claims were fully supported, ACT-JEPA would be a valuable contribution: it cleanly integrates a self-supervised JEPA-style objective into an imitation-learning policy, proposes a concrete pretraining/fine-tuning protocol for representation transfer, and is evaluated against a strong baseline (ACT) in a low-data regime. The paper is clearly structured, the architecture is described in enough detail to be reproducible, and the authors state that code and data will be released. However, the significance is substantially undercut by the gap between the abstract's headline claims and the reported numbers, and by the circularity of the H1 probing evaluation. The H2 experiment is a promising idea but is currently reported only through a proxy training loss. The core architectural idea is worth pursuing, but the present evidence does not yet support the strong conclusions the paper draws.
major comments (4)
- [Abstract, Section 5.1, Section 5.3] The abstract claims 'up to 40% improvement in world model understanding and up to 10% higher task success rate,' but the full text reports a 12.4% RMSE improvement and an 11.8% ATE improvement in Section 5.1 (Table 1) and a 0.5 percentage-point success-rate difference (91.6% vs. 91.1% in Table 2), which is within the reported error bars (91.1 ± 1.4 vs. 91.6 ± 1.7). These are numerically inconsistent, and the phrase 'up to 40%' does not correspond to any number reported in the experimental results. The abstract should be revised to match the actual reported magnitudes, and the success-rate claim should be phrased as 'comparable' rather than 'up to 10% higher.' This is load-bearing because the paper's framing and the scientific record of its claims depend on the accuracy of these numbers.
- [Section 4.3.1, Section 3.3, Section 5.1] The probing experiment used to support H1 is circular for the claimed conclusion. The probe trains a decoder to reconstruct proprioceptive state sequences from the frozen context encoder, but the observation loss in Section 3.3 (L_obs) explicitly trains the same context encoder (via the predictor and target encoder) to predict abstract proprioceptive sequences. Thus the probe measures how well the representation retains information about its own training target, not whether the model has learned general environment dynamics or can generalize beyond the training distribution. The improvement over ACT is unsurprising, since ACT is not trained to predict proprioceptive states at all. To support H1, the authors should probe a modality or a prediction task that is not used in training (e.g., image-frame reconstruction or action-conditioned future observation prediction), or demonstrate generalization to a held-out task with different dynamics. As written, the claim that 'ACT-JEPA significantly outperforms the baselines' on 'understanding environment dynamics' is not supported by this experiment.
- [Section 4.3.2, Section 5.2, Figure 2] The evidence for H2 (that predicting abstract observation sequences generalizes to action prediction) is currently only a decreasing action-reconstruction loss during fine-tuning. This is a weak proxy for actual policy quality or generalization, and the protocol alternates pretraining with one-epoch fine-tuning using a newly initialized decoder, so the trend could partly reflect optimization dynamics rather than representation quality. The paper does not report task success rates or any held-out action-prediction metric after fine-tuning. The authors should either report downstream task performance after a full fine-tuning stage or present other evidence that the pretrained representations benefit action prediction, such as a control experiment that trains the same decoder from a randomly initialized or ACT-pretrained encoder.
- [Abstract, Section 4.1, Table 2] The abstract states that ACT-JEPA is evaluated 'in different environments and across multiple tasks,' but the experimental section only describes and reports results for Meta-World (Section 4.1, Table 2). No results are shown for any other environment. This is a mismatch between the claimed scope and the presented evidence. If other environments were evaluated, their results should be reported; otherwise, the abstract and introduction should be revised to state that the evaluation is limited to a single environment suite.
minor comments (5)
- [Title page] The authors' affiliations contain typos: 'Faculy of Techical Sciences' should be 'Faculty of Technical Sciences.'
- [Section 1] In the contributions list, 'different decision-masking tasks' should be 'different decision-making tasks.'
- [Section 3.2.1] The sentence 'Then, each modality is encoded with a different modality-specific function' is followed by 'It receives an image and the proprioceptive state, along with a task a task label,' which contains a duplicated article and is awkwardly phrased.
- [Section 3.3] The notation for the losses is inconsistent: the observation loss is written as L_obs (possibly with subscripts) and the action loss as L_act, but the symbols are not defined precisely in one place, and the subscript formatting is inconsistent (e.g., 'L_obs' vs. 'L_obs'). A single, clearly defined notation would improve readability.
- [Section 5.1] The caption of Table 1 states 'We observe that ACT-JEPA significantly outperforms the baselines,' but no statistical significance test is reported. Given the small number of seeds (3 training seeds), the word 'significantly' should either be supported by an appropriate test or replaced with 'consistently outperforms.'
Circularity Check
The H1 world-model claim is circular: the probing task reconstructs the exact modality (proprioceptive states) that the observation loss trains the encoder to predict, so the reported RMSE/ATE improvements are by construction.
-
self definitional
[Section 3.3 (observation loss) and Section 4.3.1/5.1 (H1 probing experiment)]
"Therefore, the loss penalizes the distance between the predicted abstract observations and their corresponding targets. ... The decoder head is trained to reconstruct proprioceptive state sequences directly from the frozen representations."
The observation loss optimizes the context encoder to predict abstract proprioceptive-state sequences produced by the target encoder from those same proprioceptive states. The H1 probe then freezes the context encoder and trains a decoder to reconstruct proprioceptive-state sequences from it. The probing target is therefore the same modality as the training target, so a model trained with the observation objective is expected to beat a model without it on this probe. The reported 12.4% RMSE and 11.8% ATE improvements reflect the fact that the encoder was explicitly optimized to retain information about proprioceptive states, not independent evidence of understanding environment dynamics or generalization beyond the training distribution.
full rationale
We identify one substantive circular step. The H1 evaluation in Section 4.3.1 probes the frozen context encoder by decoding proprioceptive-state sequences, while Section 3.3's observation loss is trained to predict abstract representations of those same proprioceptive-state sequences. The probe improvement is therefore a consequence of the training objective rather than independent evidence of world-model understanding; the paper's 'up to 40% improvement in world model understanding' and Section 5.1's internal-world-model conclusion rest on this circular measure. This is partial circularity, not total: H3 compares ACT-JEPA to ACT/RBC on task success in Meta-World, an external benchmark with no dependence on the observation-loss target, and H2's transfer-to-action experiment, while lacking a control, is not itself circular. No load-bearing self-citation or imported-uniqueness issues were found. The score reflects that one of the three headline claims (improved world model) reduces by construction to the training target, while the policy-performance comparison remains externally grounded.
Assumptions & free parameters
free parameters (3)
- chunk size n (prediction horizon)
- loss weighting between action and observation objectives =
1:1 (equal sum)
- architecture hyperparameters (transformer dimensions, depth, learning rate, batch size, epochs)
assumptions (4)
- standard math Transformer attention and L1 losses are appropriate for policy learning and representation learning.
- domain assumption Proprioceptive states are a sufficient target modality for learning a useful world model.
- domain assumption The probing task (reconstructing proprioceptive sequences from frozen representations) measures world model understanding.
- domain assumption A decreasing action-loss during fine-tuning after pretraining indicates that pretrained representations transfer to action prediction.
Cite this review
Pith. "Pith review of ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning." pith.science (2026). https://pith.science/paper/VEFGXT4G
@misc{pith2026250114622,
author = {Pith},
title = {Pith review of: ACT-JEPA: Novel Joint-Embedding Predictive Architecture for Efficient Policy Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/VEFGXT4G}},
note = {Machine review of arXiv:2501.14622}
}
read the original abstract
Learning efficient representations for decision-making policies is a challenge in imitation learning (IL). Current IL methods require expert demonstrations, which are expensive to collect. Additionally, they are not explicitly trained to understand the environment. Consequently, they have underdeveloped world models. Self-supervised learning (SSL) offers an alternative, as it can learn a world model from diverse, unlabeled data. However, most SSL methods are inefficient because they operate in raw input space. In this work, we propose ACT-JEPA, a novel architecture that unifies IL and SSL to enhance policy representations. It is trained end-to-end to jointly predict 1) action sequences and 2) latent observation sequences. To learn in latent space, we utilize Joint-Embedding Predictive Architecture, which allows the model to filter out irrelevant details and learn a robust world model. We evaluate ACT-JEPA in different environments and across multiple tasks. Our results show that it outperforms the strongest baseline in all environments. ACT-JEPA achieves up to 40% improvement in world model understanding and up to 10% higher task success rate. Finally, we show that predicting latent observation sequences effectively generalizes to predicting action sequences. This work demonstrates how integrating IL and SSL leads to efficient policy representation learning, an improved world model, and a higher task success rate.
Forward citations
Cited by 3 Pith papers
-
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
SJEPA learns latent predictive states whose transitions are compact symbolic laws with a regularized neural residual, and demonstrates simpler, less divergent pendulum dynamics than post-hoc symbolic fitting.
-
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
FactorJEPA splits a video prediction model into layout, agent, and interaction channels with a visibility gate, and a new DENSEWORLD dataset tests it on crowded Indian city scenes.
-
EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation
Action latents supervised by both control reconstruction and environment-conditioned visual consequences outperform trajectory-centric tokenizers for robotic manipulation, especially long-horizon tasks.
Reference graph
Works this paper leans on
-
[1]
and Kaiser, Lukasz and Polosukhin, Illia , booktitle =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , booktitle =. Attention Is All You Need , year =
-
[2]
A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27 , year =
LeCun, Yann , month =. A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27 , year =
work page 2022
-
[3]
Masked Autoencoders Are Scalable Vision Learners , url =
Kaiming He and Xinlei Chen and Saining Xie and Yanghao Li and Piotr Dollar and Ross Girshick , booktitle =. Masked Autoencoders Are Scalable Vision Learners , url =. doi:10.1109/cvpr52688.2022.01553 , pages =
arXiv 2022
-
[4]
Philipp Wu and Arjun Majumdar and Kevin Stone and Yixin Lin and Igor Mordatch and P. Abbeel and A. Rajeswaran , doi =. Masked Trajectory Models for Prediction, Representation, and Control , url =. Proceedings of the 40th International Conference on Machine Learning , month =
-
[5]
Decision transformer: Reinforcement learning via sequence modeling , volume =
Chen, Lili and Lu, Kevin and Rajeswaran, Aravind and Lee, Kimin and Grover, Aditya and Laskin, Misha and Abbeel, Pieter and Srinivas, Aravind and Mordatch, Igor , journal =. Decision transformer: Reinforcement learning via sequence modeling , volume =
-
[6]
Tong, Zhan and Song, Yibing and Wang, Jue and Wang, Limin , journal =. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training , volume =
-
[7]
A-JEPA: Joint-Embedding Predictive Architecture Can Listen , url =
Fei, Zhengcong and Fan, Mingyuan and Huang, Junshi , doi =. A-JEPA: Joint-Embedding Predictive Architecture Can Listen , url =
-
[8]
Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture , url =
Mahmoud Assran and Quentin Duval and Ishan Misra and Piotr Bojanowski and Pascal Vincent and Michael Rabbat and Yann LeCun and Nicolas Ballas , booktitle =. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture , url =. doi:10.1109/cvpr52729.2023.01499 , month =
Show all 50 references
-
[9]
Navigation World Models , url =
Amir Bar and Gaoyue Zhou and Danny Tran and Trevor Darrell and Yann LeCun , booktitle =. Navigation World Models , url =. doi:10.1109/cvpr52734.2025.01472 , month =
2025
-
[10]
Sparsh: Self-supervised touch representations for vision-based tactile sensing , url =
Carolina Higuera and Akash Sharma and Chaithanya Krishna Bodduluri and Taosha Fan and Patrick Lancaster and Mrinal Kalakrishnan and Michael Kaess and Byron Boots and Mike Lambeta and Tingfan Wu and Mustafa Mukadam , doi =. Sparsh: Self-supervised touch representations for visi...
-
[11]
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning , url =
Gaoyue Zhou and Hengkai Pan and Yann LeCun and Lerrel Pinto , doi =. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning , url =. Proceedings of the 42nd International Conference on Machine Learning , month =
-
[12]
_0 : A Vision-Language-Action Flow Model for General Robot Control , url =
Kevin Black and Noah Brown and Danny Driess and Adnan Esmail and Michael Equi and Chelsea Finn and Niccolo Fusai and Lachy Groom and Karol Hausman and Brian Ichter and Szymon Jakubczak and Tim Jones and Liyiming Ke and Sergey Levine and Adrian Li-Bell and Mohith Mothukuri and ...
-
[13]
DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control , url =
Zichen Jeff Cui and Hengkai Pan and Aadhithya Iyer and Siddhant Haldar and Lerrel Pinto , doi =. DynaMo: In-Domain Dynamics Pretraining for Visuo-Motor Control , url =. Neural Information Processing Systems , month =
-
[14]
BAKU: An Efficient Transformer for Multi-Task Policy Learning , url =
Siddhant Haldar and Zhuoran Peng and Lerrel Pinto , doi =. BAKU: An Efficient Transformer for Multi-Task Policy Learning , url =. Neural Information Processing Systems , month =
-
[15]
Gopalakrishnan and Karol Hausman and Alexander Herzog and Jasmine Hsu and Julian Ibarz and Brian Ichter and A
Anthony Brohan and Noah Brown and Justice Carbajal and Yevgen Chebotar and Joseph Dabis and Chelsea Finn and K. Gopalakrishnan and Karol Hausman and Alexander Herzog and Jasmine Hsu and Julian Ibarz and Brian Ichter and A. Irpan and Tomas Jackson and Sally Jesmonth and Nikhil ...
-
[16]
Choromanski and Tianli Ding and Danny Driess and Kumar Avinava Dubey and Chelsea Finn and Peter R
Anthony Brohan and Noah Brown and Justice Carbajal and Yevgen Chebotar and K. Choromanski and Tianli Ding and Danny Driess and Kumar Avinava Dubey and Chelsea Finn and Peter R. Florence and Chuyuan Fu and Montse Gonzalez Arenas and K. Gopalakrishnan and Kehang Han and Karol Ha...
-
[17]
Levine and Chelsea Finn , booktitle =
Tony Zhao and Vikash Kumar and S. Levine and Chelsea Finn , booktitle =. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware , url =. doi:10.15607/RSS.2023.XIX.016 , month =
2023 doi
-
[18]
Diffusion policy: Visuomotor policy learning via action diffusion , url =
Cheng Chi and Zhenjia Xu and Siyuan Feng and Eric Cousineau and Yilun Du and Benjamin Burchfiel and Russ Tedrake and Shuran Song , doi =. Diffusion policy: Visuomotor policy learning via action diffusion , url =. The International Journal of Robotics Research , month =
-
[19]
World Models , url =
Ha, David and Schmidhuber, Jürgen , doi =. World Models , url =
-
[20]
Learning and Leveraging World Models in Visual Representation Learning , url =
Garrido, Quentin and Assran, Mahmoud and Ballas, Nicolas and Bardes, Adrien and Najman, Laurent and LeCun, Yann , doi =. Learning and Leveraging World Models in Visual Representation Learning , url =
-
[21]
Using ChatGPT to annotate a dataset: A case study in intelligent tutoring systems , url =
Vujinović, Aleksandar and Luburić, Nikola and Slivka, Jelena and Kovačević, Aleksandar , doi =. Using ChatGPT to annotate a dataset: A case study in intelligent tutoring systems , url =. Machine Learning with Applications , month =
-
[22]
Imitation Learning: A Survey of Learning Methods , url =
Hussein, Ahmed and Gaber, Mohamed Medhat and Elyan, Eyad and Jayne, Chrisina , doi =. Imitation Learning: A Survey of Learning Methods , url =. ACM Computing Surveys , month =
-
[23]
Revisiting Feature Prediction for Learning Visual Representations from Video , url =
Adrien Bardes and Quentin Garrido and Jean Ponce and Xinlei Chen and Michael Rabbat and Yann LeCun and Mido Assran and Nicolas Ballas , doi =. Revisiting Feature Prediction for Learning Visual Representations from Video , url =. Trans. Mach. Learn. Res. , month =
-
[24]
AnySkin: Plug-and-Play Skin Sensing for Robotic Touch , url =
Raunaq Bhirangi and Venkatesh Pattabiraman and Enes Erciyes and Yifeng Cao and Tess Hellebrekers and Lerrel Pinto , booktitle =. AnySkin: Plug-and-Play Skin Sensing for Robotic Touch , url =. doi:10.1109/icra55743.2025.11128638 , month =
2025
-
[25]
doi:10.48550/arxiv.2411.02479 , publisher =
Lambeta, Mike and Wu, Tingfan and Sengül, Ali and Most, Victoria Rose and Black, Nolan and Sawyer, Kevin and Qi, Haozhi and Sohn, Alexander and Taylor, Byron and Tydingco, Norb and Kammerer, Gregg and Khatha, Jake and Jenkins, Kurt and Most, Kyle and Stein, Neal and Chavira, R...
-
[26]
Jin and Shafiullah, Nur Muhammad Mahi and Pinto, Lerrel , booktitle =
Lee, Seungjae and Wang, Yibin and Etukuru, Haritheja and Kim, H. Jin and Shafiullah, Nur Muhammad Mahi and Pinto, Lerrel , booktitle =. Behavior Generation with Latent Actions , url =. doi:10.48550/arxiv.2403.03181 , month =
-
[27]
FiLM: Visual Reasoning with a General Conditioning Layer , url =
Ethan Perez and Florian Strub and Harm De Vries and Vincent Dumoulin and Aaron Courville , doi =. FiLM: Visual Reasoning with a General Conditioning Layer , url =. Proceedings of the AAAI Conference on Artificial Intelligence , month =
-
[28]
Maxime Oquab and Timothée Darcet and Théo Moutakanni and Huy V. Vo and Marc Szafraniec and Vasil Khalidov and Pierre Fernandez and Daniel Haziza and Francisco Massa and Alaaeldin El-Nouby and Mido Assran and Nicolas Ballas and Wojciech Galuba and Russell Howes and Po-Yao Huang...
-
[29]
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Tianhe Yu and Deirdre Quillen and Zhanpeng He and Ryan Julian and Karol Hausman and Chelsea Finn and Sergey Levine , booktitle =. Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning. , url =. doi:10.48550/arXiv.1910.10897 , month =
-
[30]
Behavior Transformers: Cloning k modes with one stone , url =
Nur Muhammad (Mahi) Shafiullah and Zichen Jeff Cui and Ariuntuya Altanzaya and Lerrel Pinto , doi =. Behavior Transformers: Cloning k modes with one stone , url =. Neural Information Processing Systems , month =
-
[31]
Evaluating egomotion and structure-from-motion approaches using the TUM RGB-D benchmark , author=. Proc. of the Workshop on Color-Depth Camera Fusion in Robotics at the IEEE/RJS International Conference on Intelligent Robot Systems (IROS) , volume=
-
[33]
Homer Rich Walke and Kevin Black and Tony Z. Zhao and Quan Vuong and Chongyi Zheng and Philippe Hansen-Estruch and Andre Wang He and Vivek Myers and Moo Jin Kim and Max Du and Abraham Lee and Kuan Fang and Chelsea Finn and Sergey Levine , doi =. BridgeData V2: A Dataset for Ro...
-
[34]
VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning , url =
Han Lin and Tushar Nagarajan and Nicolas Ballas and Mido Assran and Mojtaba Komeili and Mohit Bansal and Koustuv Sinha , doi =. VEDIT: Latent Prediction Architecture For Procedural Video Representation Learning , url =. International Conference on Learning Representations , month =
-
[35]
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think , url =
Sihyun Yu and Sangkyung Kwak and Huiwon Jang and Jongheon Jeong and Jonathan Huang and Jinwoo Shin and Saining Xie , doi =. Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think , url =. International Conference on Learning Represent...
-
[36]
Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud , url =
Ayumu Saito and Prachi Kudeshia and Jiju Poovvancheri , booktitle =. Point-JEPA: A Joint Embedding Predictive Architecture for Self-Supervised Learning on Point Cloud , url =. doi:10.1109/wacv61041.2025.00714 , month =
2025
-
[37]
A Survey on Deep Generative Models for Robot Learning From Multimodal Demonstrations , url =
Julen Urain and Ajay Mandlekar and Yilun Du and Nur Muhammad "Mahi" Shafiullah and Danfei Xu and Katerina Fragkiadaki and Georgia Chalvatzaki and Jan Peters , doi =. A Survey on Deep Generative Models for Robot Learning From Multimodal Demonstrations , url =. IEEE Transactions...
-
[38]
Proceedings of the 41st International Conference on Machine Learning , month=
Genie: Generative interactive environments , author=. Proceedings of the 41st International Conference on Machine Learning , month=
-
[39]
Whole-body Humanoid Robot Locomotion with Human Reference , url =
Qiang Zhang and Peter Cui and David Yan and Jingkai Sun and Yiqun Duan and Gang Han and Wen Zhao and Weining Zhang and Yijie Guo and Arthur Zhang and Renjing Xu , booktitle =. Whole-body Humanoid Robot Locomotion with Human Reference , url =. doi:10.1109/iros58592.2024.1080145...
-
[40]
Balakrishna and Suraj Nair and Rafael Rafailov and E
Moo Jin Kim and Karl Pertsch and Siddharth Karamcheti and Ted Xiao and A. Balakrishna and Suraj Nair and Rafael Rafailov and E. Foster and Grace Lam and Pannag R. Sanketi and Quan Vuong and Thomas Kollar and Benjamin Burchfiel and Russ Tedrake and Dorsa Sadigh and Sergey Levin...
-
[41]
NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration , url =
Ajay Sridhar and Dhruv Shah and Catherine Glossop and Sergey Levine , booktitle =. NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration , url =. doi:10.1109/icra57147.2024.10610665 , month =
2024
- [42]
-
[43]
Octo: An Open-Source Generalist Robot Policy , url =
Ghosh, Dibya and Walke, Homer and Pertsch, Karl and Black, Kevin and Mees, Oier and Dasari, Sudeep and Hejna, Joey and Kreiman, Tobias and Xu, Charles and Luo, Jianlan and Tan, You and Chen, Lawrence and Vuong, Quan and Xiao, Ted and Sanketi, Pannag and Sadigh, Dorsa and Finn,...
-
[44]
International conference on machine learning , pages=
Generative pretraining from pixels , author=. International conference on machine learning , pages=. 2020 , organization=
2020
-
[45]
OpenAI blog , volume=
Language Models are Unsupervised Multitask Learners , author=. OpenAI blog , volume=
-
[46]
Advances in neural information processing systems , volume=
Denoising diffusion probabilistic models , author=. Advances in neural information processing systems , volume=
-
[47]
Scalable Diffusion Models with Transformers , url =
Peebles, William and Xie, Saining , booktitle =. Scalable Diffusion Models with Transformers , url =. doi:10.1109/ICCV51070.2023.00387 , month =
2023
-
[48]
Learning Interactive Real-World Simulators , url =
Sherry Yang and Yilun Du and Seyed Kamyar Seyed Ghasemipour and Jonathan Tompson and Leslie Pack Kaelbling and Dale Schuurmans and Pieter Abbeel , doi =. Learning Interactive Real-World Simulators , url =. International Conference on Learning Representations , month =
-
[49]
GAIA-1: A Generative World Model for Autonomous Driving , url =
Hu, Anthony and Russell, Lloyd and Yeo, Hudson and Murez, Zak and Fedoseev, George and Kendall, Alex and Shotton, Jamie and Corrado, Gianluca , doi =. GAIA-1: A Generative World Model for Autonomous Driving , url =
-
[50]
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning , url =
Assran, Mido and Bardes, Adrien and Fan, David and Garrido, Quentin and Howes, Russell and Mojtaba and Komeili and Muckley, Matthew and Rizvi, Ammar and Roberts, Claire and Sinha, Koustuv and Zholus, Artem and Arnaud, Sergio and Gejji, Abha and Martin, Ada and Hogan, Francois ...
-
[51]
Demonstrating GPU Parallelized Robot Simulation and Rendering for Generalizable Embodied AI with ManiSkill3 , url =
Stone Tao and Fanbo Xiang and Arth Shukla and Yuzhe Qin and Xander Hinrichsen and Xiaodi Yuan and Chen Bao and Xinsong Lin and Yulin Liu and Tse-kai Chan and Yuan Gao and Xuanlin Li and Tongzhou Mu and Nan Xiao and Arnav Gurha and Viswesh Nagaswamy Rajesh and Yong Woo Choi and...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.