REVIEW 3 major objections 6 minor 3 cited by
Zero-Shot Visual Generalization in Robot Manipulation
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that a manipulation policy trained on fixed images can generalize zero-shot to new lighting, colors, and backgrounds by routing every observation through a discrete, disentangled latent space, and that the same mechanism…
desk verdict Useful empirical scaling of ALDA to manipulation and diffusion policies, but the central causal claim about associative latents is under-ablated and the evidence lacks error bars; still deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the associative latent dynamics model, whose association step is $z^d_j = \mathrm{Softmax}(\beta\, \mathrm{Sim}(z_j, V_j)) \odot V_j$: each dimension of the continuous encoder output is compared with a fixed set of scalar code values and replaced by a weighted mixture (effectively the nearest code) to form the discrete latent $z^d$. A reconstruction loss and a commitment loss train the encoder and codebooks so that this discrete code is both informative and factorized; at test time the same snapping operation is what maps novel visuals back to familiar latents. The second mechanism, learned canonicalization, uses a lightweight equivariant network $C(o)$ to rotate an input image into a canonical pose before the frozen pretrained policy consumes it, and the policy is finetuned to match the actions it would have taken on the original, unrotated image.
What would settle it
Run a trained policy on a sequence of synthetic out-of-distribution images that change only a single, task-irrelevant visual factor (for example, the background image), and record both the success rate and the per-dimension distance between $z^d$ and the nearest codebook value. If success collapses while a task-relevant latent dimension moves, or if the latent drifts continuously for perturbations the policy is claimed to survive, the association step is not forcing the representation in-distribution and the paper's mechanism is falsified.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that the association step of its associative latent disentanglement method—mapping each continuous latent coordinate to a discrete scalar codebook by attention-weighted similarity—acts as a test-time filter that forces any out-of-distribution observation back into the support of the training distribution before the policy reads it. Because the latent codes are trained to be factorized, irrelevant variations such as background content or lighting occupy separate dimensions from task-relevant factors like the cube's position, so snapping a perturbed image to its nearest in-distribution code preserves the information needed to act. The paper demonstrates this with a reinforcement-learning agent and with a diffusion-policy imitation-learning agent: both outperform their respective baselines on a suite of visual perturbations in simulation, and the imitation variant succeeds on a physical robot under changed lighting, a gray cube, and distractor objects. A separate finetuning step, built on learned canonicalization, keeps success high when images are rotated in discrete steps of 45, 30, or 15 degrees. The paper also records boundary conditions: changing table color at test time collapses performance unless table and object colors were independently randomized during training.
Load-bearing premise
The method assumes that the discrete lookup tables of visual features learned during training already cover every factor of variation that matters, so any new image can be mapped back to a familiar entry without losing task-relevant information.
Editorial extensions
If this is right
- A policy trained on one fixed camera scene can be deployed under changed lighting, backgrounds, and object colors without domain randomization or data augmentation.
- The gains transfer from reinforcement learning to imitation learning: conditioning a diffusion-based action generator on the disentangled latents preserves high success where standard diffusion behavior cloning and transformer-based action chunking fail, especially on precise pick tasks.
- Any pretrained vision-based policy can be made invariant to discrete camera rotations by a short finetuning step with a lightweight canonicalizer, without changing the policy architecture.
- Data diversity still matters: the paper's table-color collapse shows that if two visual factors are correlated in the training set, changing one at test time can break the mechanism; randomizing those factors independently during training restores generalization.
- The structured representation is compatible with stronger downstream actors: the paper's long-horizon pushing results suggest that a future, stronger base policy would inherit the visual generalization gains.
Reading between the lines
- A direct test of the mechanism is to feed out-of-distribution frames through the encoder and measure whether the snapped latent $z^d$ coincides exactly with a training-time code; the paper's account predicts zero drift on perturbations the policy survives, and visible drift exactly where it fails.
- The table-color failure points to a general diagnostic: whenever two factors of variation are correlated in the training set, the method should fail when either factor changes alone, and a latent-traversal analysis should show both factors moving along a single codebook dimension.
- Because the association step sits in the observation encoder, the same recipe should transfer to non-diffusion action heads such as transformers or MLP policies, so an inexpensive ablation is to keep the encoder and codebooks fixed and swap only the action generator.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends Associative Latent DisentAnglement (ALDA) from reinforcement learning to imitation learning, proposing ALDA-DP (ALDA combined with Diffusion Policy) for vision-based robot manipulation. It introduces a ManiSkill3-based visual generalization benchmark (MVGB) with distracting backgrounds, random colors, and random lighting, and reports simulation results for ALDA-SAC and ALDA-DP against SAC, SAC-AE, TD-MPC2, Diffusion Policy, and ACT. The authors also propose a learned-canonicalization finetuning procedure intended to make pretrained policies invariant to discrete planar image rotations, and they evaluate ALDA-DP on a real Franka arm under lighting, color, and distractor perturbations. The central claim is that disentangled representations paired with associative latent dynamics provide strong zero-shot visual generalization without domain randomization or augmentation.
Significance. If the central claim is substantiated, the paper would offer a practical alternative to domain randomization and augmentation for visual generalization in manipulation, and its extension of ALDA to diffusion-based behavior cloning would be a useful bridge between representation learning and modern imitation learning. The strengths are the breadth of the evaluation, the large numbers of rollouts used in simulation, the large margins on PickCube, the real-robot validation, and the low-cost learned-canonicalization finetuning (at most 7 minutes for ALDA-DP and 15 minutes for ALDA-SAC on C24). However, the current evidence does not isolate the associative-latent-dynamics mechanism from the rest of the representation-learning objective, the statistical reporting lacks error bars and seed counts, and the equivariant-adaptation objective in Eq. (4) is inconsistent with Algorithm 1. These issues are load-bearing for the paper's claims, so the significance is high conditional on their resolution.
major comments (3)
- [Section 3.1, Eq. (2); Section 5] The causal claim that 'associative latent dynamics' are responsible for the reported generalization is not isolated by any experiment in the paper. ALDA-DP differs from Diffusion Policy by adding the entire ALDA objective (reconstruction loss, commitment loss, activation penalties, and the codebook projection), and ALDA-SAC differs from SAC-AE in the same composite way; the RL-block comparison does not transfer to the BC setting. Section 6's table-color collapse is the failure mode predicted by the Eq. (2) association mechanism when codebook coverage is incomplete, so it does not resolve the attribution question. Please add a controlled ablation that keeps the ALDA objective but replaces the discrete codebook association in Eq. (2) with a continuous latent bottleneck, or otherwise removes only the association step, and report it on the same MVGB variations.
- [Section 4, Figure 4, Table 1, Table 2] The empirical claims are reported without error bars, confidence intervals, or the number of independent training seeds; aggregating 1000 or 500 rollouts from a single policy run does not quantify seed-to-seed variability. Table 2's real-world results are over 20 trials, and the 'Basic' condition reports exactly 80.0 for all three methods, which is hard to interpret without a description of how trials were randomized and whether the identical number is a coincidence or an artifact of reporting. Please report means with standard deviations or 95% confidence intervals, state the number of seeds for each simulation method, and clarify the real-world trial protocol.
- [Section 3.3, Eq. (4), Algorithm 1] The displayed objective in Eq. (4) writes π(a | l(f(o))), while Algorithm 1 computes the policy on o_canon = C_φ(o); if C is the canonicalizer, the canonicalized observation should appear in the policy and latent arguments. As written, Eq. (4) does not match Algorithm 1, and the pseudocode does not show the inverse group action ρ'(C(o)) that Eq. (1) requires. Please reconcile the equations with the algorithm and define explicitly how a discrete rotation of the input is undone before the policy's action is produced.
minor comments (6)
- [Table 1] Please clarify whether the C_n rows average over rotations including 0 degrees; the 'None' row is listed separately, so the current wording leaves it ambiguous which rotations enter the reported cyclic-group averages.
- [Section 4.2, Table 2] The Directed Light column notation is ambiguous; also, middle and right lighting entries are 0.0 for every method, so the text should state explicitly that these conditions are at floor for all methods and that ALDA-DP's left-light 70.0 exceeds ACT's 55.0.
- [Figure 4] Consider including a numeric table or value labels in the figure, since the bar heights alone cannot be read precisely; adding error bars would also help.
- [General] Please add a code and data availability statement; the current manuscript provides videos but no code, and the reproducibility of the simulation benchmark would be greatly improved by releasing the MVGB configuration and the ALDA-DP implementation.
- [Appendix B, Eq. (2)] The appendix says 'We use the negative L1 distance as our similarity function', but Eq. (2)'s Sim(·,·) is not written out; please expand the similarity function explicitly.
- [Section 3.3] The text says the goal is to make π equivariant to group actions on z, while the stated aim is invariance of the policy under camera rotations; please clarify whether the action representation is transformed by ρ' during training and evaluation.
Circularity Check
No circularity found: the paper's claims are empirical evaluations against external baselines, and the cited ALDA prior work is a component with independent prior validation rather than a fitted target renamed as prediction.
full rationale
The paper's central claims are empirical. ALDA-DP and ALDA-SAC are trained on fixed demonstration or replay-buffer data and evaluated on held-out visual perturbations; no test-time success value is used as a training target, and no parameter is fitted to the benchmark metrics and then relabeled as a prediction. The codebook association in Eq. (2) is a fixed architectural mechanism, not a quantity fitted to the generalization results, so the claim that it 'forces' out-of-distribution representations in-distribution is a mechanism description rather than a derivation that reduces to its own output. The reliance on the authors' prior ALDA paper [23] is a normal component citation: that prior work is a separate, already-validated method, and this paper contributes new manipulation and real-robot results that do not depend solely on the citation for their evidential force. The equivariant adaptation is explicitly attributed to external works [30, 31] and is evaluated on held-out rotations rather than being defined into existence. The absence of an ablation isolating the codebook mechanism is an experimental design limitation, not a circular reduction, and the reported table-color failure is an honestly disclosed limitation that further confirms the results are not manufactured by construction. Overall, the derivation chain is self-contained as an empirical study, with no load-bearing self-citation or fitted-input-called-prediction pattern.
Assumptions & free parameters
free parameters (4)
- beta (codebook sharpness)
- w1, w2 (reconstruction and commitment weights) =
1.0, 0.1
- lambda_theta, lambda_phi (encoder/decoder activation penalties) =
0.1
- Number of latents |zd|, values per latent |V| =
20 and 20 (sim), 10 and 12 (SAC real)
assumptions (4)
- domain assumption Disentangled latent representations are sufficient for visual generalization in manipulation
- domain assumption ALDA's associative latent dynamics (from the authors' own prior work) function as claimed, projecting OOD observations to in-distribution latents
- domain assumption ManiSkill3 simulator visual perturbations are representative of real-world distribution shifts
- domain assumption Learned canonicalization with a surrogate equivariant network can be effectively finetuned on robot policies
Cite this review
Pith. "Pith review of Zero-Shot Visual Generalization in Robot Manipulation." pith.science (2026). https://pith.science/paper/OD7XECQ3
@misc{pith2026250511719,
author = {Pith},
title = {Pith review of: Zero-Shot Visual Generalization in Robot Manipulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/OD7XECQ3}},
note = {Machine review of arXiv:2505.11719}
}
read the original abstract
Training vision-based manipulation policies that are robust across diverse visual environments remains an important and unresolved challenge in robot learning. Current approaches often sidestep the problem by relying on invariant representations such as point clouds and depth, or by brute-forcing generalization through visual domain randomization and/or large, visually diverse datasets. Disentangled representation learning - especially when combined with principles of associative memory - has recently shown promise in enabling vision-based reinforcement learning policies to be robust to visual distribution shifts. However, these techniques have largely been constrained to simpler benchmarks and toy environments. In this work, we scale disentangled representation learning and associative memory to more visually and dynamically complex manipulation tasks and demonstrate zero-shot adaptability to visual perturbations in both simulation and on real hardware. We further extend this approach to imitation learning, specifically Diffusion Policy, and empirically show significant gains in visual generalization compared to state-of-the-art imitation learning methods. Finally, we introduce a novel technique adapted from the model equivariance literature that transforms any trained neural network policy into one invariant to 2D planar rotations, making our policy not only visually robust but also resilient to certain camera perturbations. We believe that this work marks a significant step towards manipulation policies that are not only adaptable out of the box, but also robust to the complexities and dynamical nature of real-world deployment. Supplementary videos are available at https://sites.google.com/view/vis-gen-robotics/home.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 3 Pith papers
-
A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces
DPA-FTG couples a 5 Hz diffusion-based task selector with a 60 Hz force-reactive GRU decoder, improving safe task success over diffusion baselines on bimanual compliant sheet separation.
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.
-
OpenTie: Open-vocabulary Sequential Rebar Tying System
A claimed training-free rebar tying pipeline based on point clouds and open-vocabulary detection, but the reported evaluation is too vague to verify the claimed 90% success.
Reference graph
Works this paper leans on
-
[1]
T. Mu, Z. Ling, F. Xiang, D. Yang, X. Li, S. Tao, Z. Huang, Z. Jia, and H. Su. Maniskill: Gen- eralizable manipulation skill benchmark with large-scale demonstrations. In J. Vanschoren and S. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, vir-...
work page 2021
-
[2]
V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac gym: High performance GPU based physics simulation for robot learning. In J. Vanschoren and S. Yeung, edi- tors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets...
work page 2021
-
[3]
P. Katara, Z. Xian, and K. Fragkiadaki. Gen2sim: Scaling up robot learning in simu- lation with generative models. InIEEE International Conference on Robotics and Au- tomation, ICRA 2024, Yokohama, Japan, May 13-17, 2024, pages 6672–6679. IEEE, 2024. doi:10.1109/ICRA57147.2024.10610566. URLhttps://doi.org/10.1109/ICRA57147. 2024.10610566
arXiv 2024
-
[4]
M. Andrychowicz, B. Baker, M. Chociej, R. J ´ozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba. Learning dexterous in-hand manipulation.Int. J. Robotics Res., 39(1), 2020. doi:10.1177/0278364919887447. URLhttps://doi.org/10.1177/0278364919887447
-
[5]
A. Almuzairee, N. Hansen, and H. I. Christensen. A recipe for unbounded data augmentation in visual reinforcement learning.RLJ, 1:130–157, 2024. URLhttps://rlj.cs.umass. edu/2024/papers/Paper26.html
work page 2024
-
[6]
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang. Solving rubik’s cube with a robot hand.CoRR, abs/1910.07113, 2019. URLhttp://arxiv.org/abs/1910.07113
arXiv 1910
-
[7]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model. In P. Agrawal, O. Kroemer, and W. Burgard, editors,Conference on Robot Learning, 6-9 November 202...
2024
-
[8]
Brohan, N
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K. Lee, S. Levine, Y . Lu, U. Malla, D. Man- junath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, ...
2023
Show all 73 references
-
[9]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V . Vanhoucke, H. T. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. San- keti, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, ...
2023
-
[10]
Pertsch, K
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. FAST: efficient action tokenization for vision-language-action models.CoRR, abs/2501.09747, 2025. doi:10.48550/ARXIV .2501.09747. URLhttps://doi.org/10. 48550/arXiv.2501.09747
-
[11]
Nayebi, R
A. Nayebi, R. Rajalingham, M. Jazayeri, and G. R. Yang. Neural foundations of men- tal simulation: Future prediction of latent representations on dynamic scenes. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Proc...
2023
-
[12]
Abbas and S
A. Abbas and S. Deny. Progress and limitations of deep networks to recognize objects in un- usual poses. In B. Williams, Y . Chen, and J. Neville, editors,Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications o...
2023 doi
- [13]
-
[14]
J. J. DiCarlo, D. Zoccolan, and N. C. Rust. How does the brain solve visual object recognition? Neuron, 73(3):415–434, 2012
2012
-
[15]
T. E. Behrens, T. H. Muller, J. C. Whittington, S. Mark, A. B. Baram, K. L. Stachenfeld, and Z. Kurth-Nelson. What is a cognitive map? organizing knowledge for flexible behavior. Neuron, 100(2):490–509, 2018
2018
-
[16]
J. J. Bakermans, J. Warren, J. C. Whittington, and T. E. Behrens. Constructing future behavior in the hippocampal formation through composition and replay.Nature Neuroscience, pages 1–12, 2025
2025
-
[17]
W. Tang, H. Chang, C. Liu, S. Perez-Hernandez, W. Y . Zheng, J. Park, A. Oliva, and A. Fernandez-Ruiz. A hippocampal population code for rapid generalization.bioRxiv, pages 2025–03, 2025
2025
-
[18]
Higgins, L
I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner. beta-vae: Learning basic visual concepts with a constrained variational frame- work. In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, A...
2017
-
[19]
K. Hsu, W. Dorrell, J. C. R. Whittington, J. Wu, and C. Finn. Disentanglement via latent quantization. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, 11 editors,Advances in Neural Information Processing Systems 36: Annual Conference on Neu- ral Informa...
2023
-
[20]
Locatello, D
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszko- reit, A. Dosovitskiy, and T. Kipf. Object-centric learning with slot attention. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Ad- vances in Neural Information Processin...
2020
-
[21]
Yarats, A
D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus. Improving sample ef- ficiency in model-free reinforcement learning from images. InThirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Art...
2021 doi
-
[22]
Higgins, A
I. Higgins, A. Pal, A. A. Rusu, L. Matthey, C. P. Burgess, A. Pritzel, M. M. Botvinick, C. Blundell, and A. Lerchner. DARLA: improving zero-shot transfer in reinforcement learn- ing. In D. Precup and Y . W. Teh, editors,Proceedings of the 34th International Confer- ence on Mac...
2017
- [23]
-
[24]
Hansen, H
N. Hansen, H. Su, and X. Wang. Stabilizing deep q-learning with convnets and vision trans- formers under data augmentation. 2021
2021
-
[25]
C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 20...
2023 doi
-
[26]
Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu. 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors,Robotics: Science and Systems XX, Delft, The Netherlands, July 15...
2024 doi
-
[27]
T. Ke, N. Gkanatsios, and K. Fragkiadaki. 3d diffuser actor: Policy diffusion with 3d scene representations. In P. Agrawal, O. Kroemer, and W. Burgard, editors,Conference on Robot Learning, 6-9 November 2024, Munich, Germany, volume 270 ofProceedings of Machine Learning Resear...
2024
-
[28]
D. Wang, R. Walters, and R. Platt. $\mathrm{SO}(2)$-equivariant reinforcement learning. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URLhttps://openreview.net/forum?id= 7F9cOhdvfk_. 12
2022
-
[29]
K. Chen, X. Chen, Z. Yu, M. Zhu, and H. Yang. Equidiff: A conditional equivariant diffu- sion model for trajectory prediction. In25th IEEE International Conference on Intelligent Transportation Systems, ITSC 2022, Macau, China, October 8-12, 2022, pages 746–751. IEEE, 2023. do...
2022
-
[30]
A. K. Mondal, S. S. Panigrahi, O. Kaba, S. Mudumba, and S. Ravanbakhsh. Equiv- ariant adaptation of large pretrained models. In A. Oh, T. Naumann, A. Glober- son, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Infor- mation Processing Systems 36: Annual Confere...
2023
-
[31]
S. Kaba, A. K. Mondal, Y . Zhang, Y . Bengio, and S. Ravanbakhsh. Equivariance with learned canonicalization functions. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors,International Conference on Machine Learning, ICML 2023, 23- 29 July 2...
2023
-
[32]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms.CoRR, abs/1707.06347, 2017. URLhttp://arxiv.org/abs/1707.06347
2017 arXiv
-
[33]
Schulman, S
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz. Trust region policy opti- mization. In F. R. Bach and D. M. Blei, editors,Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 ofJMLR Workshop a...
2015
-
[34]
V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In M. Balcan and K. Q. Weinberger, editors,Proceedings of the 33nd International Conference on Ma- chine Learning, ICML ...
2016
-
[35]
Haarnoja, A
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum en- tropy deep reinforcement learning with a stochastic actor. In J. G. Dy and A. Krause, ed- itors,Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholms...
2018
-
[36]
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra. Continuous control with deep reinforcement learning. In Y . Bengio and Y . LeCun, editors,4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Ri...
2016
-
[37]
Fujimoto, H
S. Fujimoto, H. van Hoof, and D. Meger. Addressing function approximation error in actor- critic methods. In J. G. Dy and A. Krause, editors,Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm ¨assan, Stockholm, Sweden, July 10-15, 2018...
2018
-
[38]
V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Ried- miller. Playing atari with deep reinforcement learning.CoRR, abs/1312.5602, 2013. URL http://arxiv.org/abs/1312.5602
2013 arXiv
-
[39]
Hafner, T
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. In8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URLhttps:// openreview.net/foru...
2020
-
[40]
Hafner, T
D. Hafner, T. P. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 201...
2019
-
[41]
Hansen, H
N. Hansen, H. Su, and X. Wang. TD-MPC2: scalable, robust world models for continuous control. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps://openreview.net/ forum?id=Oxh5CstDJU
2024
- [42]
-
[43]
B. Tang, M. A. Lin, I. Akinola, A. Handa, G. S. Sukhatme, F. Ramos, D. Fox, and Y . S. Narang. Industreal: Transferring contact-rich assembly tasks from simulation to reality. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daeg...
2023 doi
-
[44]
B. Tang, I. Akinola, J. Xu, B. Wen, A. Handa, K. V . Wyk, D. Fox, G. S. Sukhatme, F. Ramos, and Y . S. Narang. Automate: Specialist and generalist assembly policies over diverse geome- tries. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors,Robotics: Science and...
2024 doi
-
[45]
Handa, A
A. Handa, A. Allshire, V . Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. V . Wyk, A. Zhurkevich, B. Sundaralingam, and Y . S. Narang. Dextreme: Transfer of agile in-hand manipulation from simulation to reality. InIEEE International Conference on Robotics and A...
2023
-
[46]
Tassa, Y
Y . Tassa, Y . Doron, A. Muldal, T. Erez, Y . Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. P. Lillicrap, and M. A. Riedmiller. Deepmind control suite.CoRR, abs/1801.00690, 2018. URLhttp://arxiv.org/abs/1801.00690
2018 arXiv
-
[47]
Mandlekar, D
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. In A. Faust, D. Hsu, and G. Neumann, editors,Conference on Robot Learn...
2021
-
[48]
Haldar, J
S. Haldar, J. Pari, A. Rai, and L. Pinto. Teach a robot to FISH: versatile imitation from one minute of demonstrations. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors, Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi: 10.1...
2023 doi
-
[49]
J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Ad- vances in Neural Information Processing Systems 33: Annual Conference on Neu- ral Information Processing Systems 2020, NeurIPS ...
2020
-
[50]
Y . Zeng, M. Suganuma, and T. Okatani. Inverting the generation process of denoising diffusion implicit models: Empirical evaluation and a novel method. InIEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025, Tucson, AZ, USA, February 26 - March 6, 2025, pa...
2025
-
[51]
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenRe- view.ne...
2021
-
[52]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674– 10685. IEEE, 2022. doi:1...
2022
-
[53]
Zhang, M
X. Zhang, M. Chang, P. Kumar, and S. Gupta. Diffusion meets dagger: Supercharging eye- in-hand imitation learning. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors, Robotics: Science and Systems XX, Delft, The Netherlands, July 15-19, 2024, 2024. doi: 10.15607/R...
2024 doi
-
[54]
Pearce, T
T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V . Macua, S. Z. Tan, I. Momennejad, K. Hofmann, and S. Devlin. Imitating human behaviour with diffusion models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rw...
2023
-
[55]
A. Vaswani. Attention is all you need.Advances in Neural Information Processing Systems, 2017
2017
-
[56]
Shridhar, L
M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. In K. Liu, D. Kulic, and J. Ichnowski, editors,Conference on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, volume 205 ofProceedings of Machine Lea...
2022
-
[57]
Goyal, J
A. Goyal, J. Xu, Y . Guo, V . Blukis, Y . Chao, and D. Fox. RVT: robotic view transformer for 3d object manipulation. In J. Tan, M. Toussaint, and K. Darvish, editors,Conference on Robot Learning, CoRL 2023, 6-9 November 2023, Atlanta, GA, USA, volume 229 ofProceedings of Mach...
2023
-
[58]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi:10.15607/ R...
2023 doi
-
[59]
Garrido, N
Q. Garrido, N. Ballas, M. Assran, A. Bardes, L. Najman, M. Rabbat, E. Dupoux, and Y . LeCun. Intuitive physics understanding emerges from self-supervised pretraining on natural videos. CoRR, abs/2502.11831, 2025. doi:10.48550/ARXIV .2502.11831. URLhttps://doi.org/ 10.48550/arX...
-
[60]
J. Pari, N. M. M. Shafiullah, S. P. Arunachalam, and L. Pinto. The surprising effectiveness of representation learning for visual imitation. In K. Hauser, D. A. Shell, and S. Huang, ed- itors,Robotics: Science and Systems XVIII, New York City, NY, USA, June 27 - July 1, 2022,
2022
-
[61]
Samborska, J
V . Samborska, J. L. Butler, M. E. Walton, T. E. Behrens, and T. Akam. Complementary task representations in hippocampus and prefrontal cortex for generalizing the structure of problems.Nature Neuroscience, 25(10):1314–1326, 2022
2022
-
[62]
W. Sun, M. Advani, N. Spruston, A. Saxe, and J. E. Fitzgerald. Organizing memories for generalization in complementary learning systems.Nature neuroscience, 26(8):1438–1448, 2023
2023
-
[63]
J. C. R. Whittington, W. Dorrell, S. Ganguli, and T. Behrens. Disentanglement with biolog- ical constraints: A theory of functional cell types. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net,
2023
-
[64]
K. Hsu, J. I. Hamid, K. Burns, C. Finn, and J. Wu. Tripod: Three complementary inductive biases for disentangled representation learning. InForty-first International Conference on Ma- chine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https...
2024
-
[65]
Higgins, N
I. Higgins, N. Sonnerat, L. Matthey, A. Pal, C. P. Burgess, M. Bosnjak, M. Shanahan, M. M. Botvinick, D. Hassabis, and A. Lerchner. SCAN: learning hierarchical compositional visual concepts. In6th International Conference on Learning Representations, ICLR 2018, Vancou- ver, BC...
2018
-
[66]
J. J. Hopfield. Neural networks and physical systems with emergent collective computational abilities.Proceedings of the national academy of sciences, 79(8):2554–2558, 1982
1982
-
[67]
Ramsauer, B
H. Ramsauer, B. Sch ¨afl, J. Lehner, P. Seidl, M. Widrich, L. Gruber, M. Holzleitner, T. Adler, D. P. Kreil, M. K. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter. Hopfield networks is all you need. In9th International Conference on Learning Representations, ICLR 2021, ...
2021
-
[68]
Hoover, Y
B. Hoover, Y . Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. Zaki, and D. Krotov. Energy transformer.Advances in Neural Information Processing Systems, 36, 2024
2024
-
[69]
B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba. Places: A 10 million image database for scene recognition.IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 2017. 16 A Latent Traversals Figure 5: Traversing the disentangled latent space of ALDA-DP t...
2017
-
[2018]
URLhttps://openreview.net/forum?id=rkN2Il-RZ
-
[2022]
URLhttps://doi.org/10.15607/RSS.2022
doi:10.15607/RSS.2022.XVIII.010. URLhttps://doi.org/10.15607/RSS.2022. XVIII.010
2022 doi
-
[2023]
URLhttps://openreview.net/forum?id=9Z_GfhZnGH
-
[2713]
URLhttps://proceedings.mlr.press/v270/kim25c.html
PMLR, 2024. URLhttps://proceedings.mlr.press/v270/kim25c.html
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.