Pith. sign in

REVIEW 3 major objections 6 minor 3 cited by

Zero-Shot Visual Generalization in Robot Manipulation

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper claims that a manipulation policy trained on fixed images can generalize zero-shot to new lighting, colors, and backgrounds by routing every observation through a discrete, disentangled latent space, and that the same mechanism…

desk verdict Useful empirical scaling of ALDA to manipulation and diffusion policies, but the central causal claim about associative latents is under-ablated and the evidence lacks error bars; still deserves peer review. read the letter →

arxiv 2505.11719 v1 pith:OD7XECQ3 submitted 2025-05-16 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords zero-shotvisualgeneralizationrobotmanipulationdisentangledrepresentationlearningassociativelatentdynamicsdiffusionpolicyimitationlearnedcanonicalizationequivariance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to show that a robot policy can generalize to new visual conditions—changed lighting, background clutter, or object colors—without domain randomization, data augmentation, or a larger dataset, provided observations are first compressed into a discrete, disentangled latent space. The central claim is that when a trained encoder sees an out-of-distribution image, an associative step snaps each latent dimension back to the closest codebook value, so the policy effectively acts on a familiar, in-distribution representation. The authors extend this idea from reinforcement learning to imitation learning by conditioning a diffusion-based action generator on the disentangled latents, and they report large gains over standard diffusion behavior cloning and transformer-based action chunking on precise pick tasks in simulation, plus success on a real robot under several perturbations. They also introduce a finetuning procedure, adapted from learned-canonicalization methods, that makes a pretrained policy invariant to planar image rotations. If the claim holds, it offers a route to visual robustness that does not depend on enumerating every possible scene at training time.

What carries the argument

The load-bearing object is the associative latent dynamics model, whose association step is $z^d_j = \mathrm{Softmax}(\beta\, \mathrm{Sim}(z_j, V_j)) \odot V_j$: each dimension of the continuous encoder output is compared with a fixed set of scalar code values and replaced by a weighted mixture (effectively the nearest code) to form the discrete latent $z^d$. A reconstruction loss and a commitment loss train the encoder and codebooks so that this discrete code is both informative and factorized; at test time the same snapping operation is what maps novel visuals back to familiar latents. The second mechanism, learned canonicalization, uses a lightweight equivariant network $C(o)$ to rotate an input image into a canonical pose before the frozen pretrained policy consumes it, and the policy is finetuned to match the actions it would have taken on the original, unrotated image.

What would settle it

Run a trained policy on a sequence of synthetic out-of-distribution images that change only a single, task-irrelevant visual factor (for example, the background image), and record both the success rate and the per-dimension distance between $z^d$ and the nearest codebook value. If success collapses while a task-relevant latent dimension moves, or if the latent drifts continuously for perturbations the policy is claimed to survive, the association step is not forcing the representation in-distribution and the paper's mechanism is falsified.

Watch

Extended reading notes

Core claim

On its own terms, the paper's discovery is that the association step of its associative latent disentanglement method—mapping each continuous latent coordinate to a discrete scalar codebook by attention-weighted similarity—acts as a test-time filter that forces any out-of-distribution observation back into the support of the training distribution before the policy reads it. Because the latent codes are trained to be factorized, irrelevant variations such as background content or lighting occupy separate dimensions from task-relevant factors like the cube's position, so snapping a perturbed image to its nearest in-distribution code preserves the information needed to act. The paper demonstrates this with a reinforcement-learning agent and with a diffusion-policy imitation-learning agent: both outperform their respective baselines on a suite of visual perturbations in simulation, and the imitation variant succeeds on a physical robot under changed lighting, a gray cube, and distractor objects. A separate finetuning step, built on learned canonicalization, keeps success high when images are rotated in discrete steps of 45, 30, or 15 degrees. The paper also records boundary conditions: changing table color at test time collapses performance unless table and object colors were independently randomized during training.

Load-bearing premise

The method assumes that the discrete lookup tables of visual features learned during training already cover every factor of variation that matters, so any new image can be mapped back to a familiar entry without losing task-relevant information.

Editorial extensions

If this is right

  • A policy trained on one fixed camera scene can be deployed under changed lighting, backgrounds, and object colors without domain randomization or data augmentation.
  • The gains transfer from reinforcement learning to imitation learning: conditioning a diffusion-based action generator on the disentangled latents preserves high success where standard diffusion behavior cloning and transformer-based action chunking fail, especially on precise pick tasks.
  • Any pretrained vision-based policy can be made invariant to discrete camera rotations by a short finetuning step with a lightweight canonicalizer, without changing the policy architecture.
  • Data diversity still matters: the paper's table-color collapse shows that if two visual factors are correlated in the training set, changing one at test time can break the mechanism; randomizing those factors independently during training restores generalization.
  • The structured representation is compatible with stronger downstream actors: the paper's long-horizon pushing results suggest that a future, stronger base policy would inherit the visual generalization gains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the mechanism is to feed out-of-distribution frames through the encoder and measure whether the snapped latent $z^d$ coincides exactly with a training-time code; the paper's account predicts zero drift on perturbations the policy survives, and visible drift exactly where it fails.
  • The table-color failure points to a general diagnostic: whenever two factors of variation are correlated in the training set, the method should fail when either factor changes alone, and a latent-traversal analysis should show both factors moving along a single codebook dimension.
  • Because the association step sits in the observation encoder, the same recipe should transfer to non-diffusion action heads such as transformers or MLP policies, so an inexpensive ablation is to keep the encoder and codebooks fixed and swap only the action generator.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper extends Associative Latent DisentAnglement (ALDA) from reinforcement learning to imitation learning, proposing ALDA-DP (ALDA combined with Diffusion Policy) for vision-based robot manipulation. It introduces a ManiSkill3-based visual generalization benchmark (MVGB) with distracting backgrounds, random colors, and random lighting, and reports simulation results for ALDA-SAC and ALDA-DP against SAC, SAC-AE, TD-MPC2, Diffusion Policy, and ACT. The authors also propose a learned-canonicalization finetuning procedure intended to make pretrained policies invariant to discrete planar image rotations, and they evaluate ALDA-DP on a real Franka arm under lighting, color, and distractor perturbations. The central claim is that disentangled representations paired with associative latent dynamics provide strong zero-shot visual generalization without domain randomization or augmentation.

Significance. If the central claim is substantiated, the paper would offer a practical alternative to domain randomization and augmentation for visual generalization in manipulation, and its extension of ALDA to diffusion-based behavior cloning would be a useful bridge between representation learning and modern imitation learning. The strengths are the breadth of the evaluation, the large numbers of rollouts used in simulation, the large margins on PickCube, the real-robot validation, and the low-cost learned-canonicalization finetuning (at most 7 minutes for ALDA-DP and 15 minutes for ALDA-SAC on C24). However, the current evidence does not isolate the associative-latent-dynamics mechanism from the rest of the representation-learning objective, the statistical reporting lacks error bars and seed counts, and the equivariant-adaptation objective in Eq. (4) is inconsistent with Algorithm 1. These issues are load-bearing for the paper's claims, so the significance is high conditional on their resolution.

major comments (3)
  1. [Section 3.1, Eq. (2); Section 5] The causal claim that 'associative latent dynamics' are responsible for the reported generalization is not isolated by any experiment in the paper. ALDA-DP differs from Diffusion Policy by adding the entire ALDA objective (reconstruction loss, commitment loss, activation penalties, and the codebook projection), and ALDA-SAC differs from SAC-AE in the same composite way; the RL-block comparison does not transfer to the BC setting. Section 6's table-color collapse is the failure mode predicted by the Eq. (2) association mechanism when codebook coverage is incomplete, so it does not resolve the attribution question. Please add a controlled ablation that keeps the ALDA objective but replaces the discrete codebook association in Eq. (2) with a continuous latent bottleneck, or otherwise removes only the association step, and report it on the same MVGB variations.
  2. [Section 4, Figure 4, Table 1, Table 2] The empirical claims are reported without error bars, confidence intervals, or the number of independent training seeds; aggregating 1000 or 500 rollouts from a single policy run does not quantify seed-to-seed variability. Table 2's real-world results are over 20 trials, and the 'Basic' condition reports exactly 80.0 for all three methods, which is hard to interpret without a description of how trials were randomized and whether the identical number is a coincidence or an artifact of reporting. Please report means with standard deviations or 95% confidence intervals, state the number of seeds for each simulation method, and clarify the real-world trial protocol.
  3. [Section 3.3, Eq. (4), Algorithm 1] The displayed objective in Eq. (4) writes π(a | l(f(o))), while Algorithm 1 computes the policy on o_canon = C_φ(o); if C is the canonicalizer, the canonicalized observation should appear in the policy and latent arguments. As written, Eq. (4) does not match Algorithm 1, and the pseudocode does not show the inverse group action ρ'(C(o)) that Eq. (1) requires. Please reconcile the equations with the algorithm and define explicitly how a discrete rotation of the input is undone before the policy's action is produced.
minor comments (6)
  1. [Table 1] Please clarify whether the C_n rows average over rotations including 0 degrees; the 'None' row is listed separately, so the current wording leaves it ambiguous which rotations enter the reported cyclic-group averages.
  2. [Section 4.2, Table 2] The Directed Light column notation is ambiguous; also, middle and right lighting entries are 0.0 for every method, so the text should state explicitly that these conditions are at floor for all methods and that ALDA-DP's left-light 70.0 exceeds ACT's 55.0.
  3. [Figure 4] Consider including a numeric table or value labels in the figure, since the bar heights alone cannot be read precisely; adding error bars would also help.
  4. [General] Please add a code and data availability statement; the current manuscript provides videos but no code, and the reproducibility of the simulation benchmark would be greatly improved by releasing the MVGB configuration and the ALDA-DP implementation.
  5. [Appendix B, Eq. (2)] The appendix says 'We use the negative L1 distance as our similarity function', but Eq. (2)'s Sim(·,·) is not written out; please expand the similarity function explicitly.
  6. [Section 3.3] The text says the goal is to make π equivariant to group actions on z, while the stated aim is invariance of the policy under camera rotations; please clarify whether the action representation is transformed by ρ' during training and evaluation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper's claims are empirical evaluations against external baselines, and the cited ALDA prior work is a component with independent prior validation rather than a fitted target renamed as prediction.

full rationale

The paper's central claims are empirical. ALDA-DP and ALDA-SAC are trained on fixed demonstration or replay-buffer data and evaluated on held-out visual perturbations; no test-time success value is used as a training target, and no parameter is fitted to the benchmark metrics and then relabeled as a prediction. The codebook association in Eq. (2) is a fixed architectural mechanism, not a quantity fitted to the generalization results, so the claim that it 'forces' out-of-distribution representations in-distribution is a mechanism description rather than a derivation that reduces to its own output. The reliance on the authors' prior ALDA paper [23] is a normal component citation: that prior work is a separate, already-validated method, and this paper contributes new manipulation and real-robot results that do not depend solely on the citation for their evidential force. The equivariant adaptation is explicitly attributed to external works [30, 31] and is evaluated on held-out rotations rather than being defined into existence. The absence of an ablation isolating the codebook mechanism is an experimental design limitation, not a circular reduction, and the reported table-color failure is an honestly disclosed limitation that further confirms the results are not manufactured by construction. Overall, the derivation chain is self-contained as an empirical study, with no load-bearing self-citation or fitted-input-called-prediction pattern.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The paper contributes an empirical combination of existing components. It introduces no new physical entities or mathematical axioms. The key assumptions are that the prior ALDA mechanism works as described, and that the canonicalization finetuning transfers to policies.

free parameters (4)
  • beta (codebook sharpness)
    Controls separation in the attention-based association, set without a stated value in the paper (Appendix B).
  • w1, w2 (reconstruction and commitment weights) = 1.0, 0.1
    Chosen manually for all experiments (Appendix B).
  • lambda_theta, lambda_phi (encoder/decoder activation penalties) = 0.1
    Chosen manually (Appendix B).
  • Number of latents |zd|, values per latent |V| = 20 and 20 (sim), 10 and 12 (SAC real)
    Hand-set hyperparameters (Table 3, Appendix E).
assumptions (4)
  • domain assumption Disentangled latent representations are sufficient for visual generalization in manipulation
    The paper relies on this from prior work to justify the ALDA objective (Sections 1 and 3.1).
  • domain assumption ALDA's associative latent dynamics (from the authors' own prior work) function as claimed, projecting OOD observations to in-distribution latents
    The paper uses this mechanism without re-deriving it (Section 3.1, Eq. 2).
  • domain assumption ManiSkill3 simulator visual perturbations are representative of real-world distribution shifts
    The simulation benchmark is used as a proxy for real-world generalization (Section 4).
  • domain assumption Learned canonicalization with a surrogate equivariant network can be effectively finetuned on robot policies
    This is adapted from image classification work and assumed to transfer to control (Section 3.3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Zero-Shot Visual Generalization in Robot Manipulation." pith.science (2026). https://pith.science/paper/OD7XECQ3

@misc{pith2026250511719,
  author       = {Pith},
  title        = {Pith review of: Zero-Shot Visual Generalization in Robot Manipulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OD7XECQ3}},
  note         = {Machine review of arXiv:2505.11719}
}
read the original abstract

Training vision-based manipulation policies that are robust across diverse visual environments remains an important and unresolved challenge in robot learning. Current approaches often sidestep the problem by relying on invariant representations such as point clouds and depth, or by brute-forcing generalization through visual domain randomization and/or large, visually diverse datasets. Disentangled representation learning - especially when combined with principles of associative memory - has recently shown promise in enabling vision-based reinforcement learning policies to be robust to visual distribution shifts. However, these techniques have largely been constrained to simpler benchmarks and toy environments. In this work, we scale disentangled representation learning and associative memory to more visually and dynamically complex manipulation tasks and demonstrate zero-shot adaptability to visual perturbations in both simulation and on real hardware. We further extend this approach to imitation learning, specifically Diffusion Policy, and empirically show significant gains in visual generalization compared to state-of-the-art imitation learning methods. Finally, we introduce a novel technique adapted from the model equivariance literature that transforms any trained neural network policy into one invariant to 2D planar rotations, making our policy not only visually robust but also resilient to certain camera perturbations. We believe that this work marks a significant step towards manipulation policies that are not only adaptable out of the box, but also robust to the complexities and dynamical nature of real-world deployment. Supplementary videos are available at https://sites.google.com/view/vis-gen-robotics/home.

Figures

Figures reproduced from arXiv: 2505.11719 by the authors.

Figure 1
Figure 1. Behavior cloning with disentangled representations and associative latent dynamics [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of ALDA + Diffusion Policy (ALDA-DP). ALDA-DP jointly learns a factorized representation of the image observation while training the policy. The diffusion model denoises actions conditioned on this representation. work in neuroscience finds evidence that the hippocampus in mice achieves rapid generalization through disentangled memory representations [17], providing a biologically plausible motivation for t… view at source ↗
Figure 3
Figure 3. ManiSkill3 visual generalization tasks. Left to right: random lighting, random cube color, distracting background (DBG), DBG + random cube color, DBG + random lighting + random cube color. To investigate how our method scales to manipulation tasks, we choose ManiSkill3’s [1] tabletop manipulation suite. ManiSkill3 is a high-throughput simulator with support for GPU parallelization. It comes with demonstrations for i… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Average success rate of ALDA-SAC and ALDA-DP compared to various RL and BC [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Traversing the disentangled latent space of ALDA-DP trained on real demonstrations and [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: ALDA-SAC method diagram. The original ALDA-SAC implementation in [23] assumes as input a stack of observations ∈ R k×C×H×W . The ALDA model is trained to independently encode and reconstruct each individual frame and disentangle the latent representation according to t…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces

    cs.RO 2026-08 conditional novelty 6.0 of 10

    DPA-FTG couples a 5 Hz diffusion-based task selector with a 60 Hz force-reactive GRU decoder, improving safe task success over diffusion baselines on bimanual compliant sheet separation.

  2. DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization

    cs.RO 2025-11 conditional novelty 5.0 of 10

    Fusing RGB and point-cloud inputs with training-time modality dropout plus cross-attention makes a diffusion visuomotor policy markedly more robust to visual and spatial shifts than unimodal or naively fused baselines.

  3. OpenTie: Open-vocabulary Sequential Rebar Tying System

    cs.RO 2025-08 reject novelty 4.0 of 10

    A claimed training-free rebar tying pipeline based on point clouds and open-vocabulary detection, but the reported evaluation is too vague to verify the claimed 90% success.

Reference graph

Works this paper leans on

73 extracted references · 48 canonical work pages · cited by 3 Pith papers

  1. [1]

    T. Mu, Z. Ling, F. Xiang, D. Yang, X. Li, S. Tao, Z. Huang, Z. Jia, and H. Su. Maniskill: Gen- eralizable manipulation skill benchmark with large-scale demonstrations. In J. Vanschoren and S. Yeung, editors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, vir-...

  2. [2]

    Makoviychuk, L

    V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac gym: High performance GPU based physics simulation for robot learning. In J. Vanschoren and S. Yeung, edi- tors,Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets...

  3. [3]

    Katara, Z

    P. Katara, Z. Xian, and K. Fragkiadaki. Gen2sim: Scaling up robot learning in simu- lation with generative models. InIEEE International Conference on Robotics and Au- tomation, ICRA 2024, Yokohama, Japan, May 13-17, 2024, pages 6672–6679. IEEE, 2024. doi:10.1109/ICRA57147.2024.10610566. URLhttps://doi.org/10.1109/ICRA57147. 2024.10610566

  4. [4]

    Andrychowicz, B

    M. Andrychowicz, B. Baker, M. Chociej, R. J ´ozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba. Learning dexterous in-hand manipulation.Int. J. Robotics Res., 39(1), 2020. doi:10.1177/0278364919887447. URLhttps://doi.org/10.1177/0278364919887447

  5. [5]

    Almuzairee, N

    A. Almuzairee, N. Hansen, and H. I. Christensen. A recipe for unbounded data augmentation in visual reinforcement learning.RLJ, 1:130–157, 2024. URLhttps://rlj.cs.umass. edu/2024/papers/Paper26.html

  6. [6]

    Akkaya, M

    OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang. Solving rubik’s cube with a robot hand.CoRR, abs/1910.07113, 2019. URLhttp://arxiv.org/abs/1910.07113

  7. [7]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model. In P. Agrawal, O. Kroemer, and W. Burgard, editors,Conference on Robot Learning, 6-9 November 202...

  8. [8]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K. Lee, S. Levine, Y . Lu, U. Malla, D. Man- junath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, ...

Show all 73 references
  1. [9]

    Zitkovich, T

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V . Vanhoucke, H. T. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. San- keti, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, ...

  2. [10]

    Pertsch, K

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. FAST: efficient action tokenization for vision-language-action models.CoRR, abs/2501.09747, 2025. doi:10.48550/ARXIV .2501.09747. URLhttps://doi.org/10. 48550/arXiv.2501.09747

  3. [11]

    Nayebi, R

    A. Nayebi, R. Rajalingham, M. Jazayeri, and G. R. Yang. Neural foundations of men- tal simulation: Future prediction of latent representations on dynamic scenes. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Information Proc...

  4. [12]

    Abbas and S

    A. Abbas and S. Deny. Progress and limitations of deep networks to recognize objects in un- usual poses. In B. Williams, Y . Chen, and J. Neville, editors,Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications o...

  5. [13]

    Perin and S

    A. Perin and S. Deny. On the ability of deep networks to learn symmetries from data: A neural kernel theory.CoRR, abs/2412.11521, 2024. doi:10.48550/ARXIV .2412.11521. URL https://doi.org/10.48550/arXiv.2412.11521

  6. [14]

    J. J. DiCarlo, D. Zoccolan, and N. C. Rust. How does the brain solve visual object recognition? Neuron, 73(3):415–434, 2012

  7. [15]

    T. E. Behrens, T. H. Muller, J. C. Whittington, S. Mark, A. B. Baram, K. L. Stachenfeld, and Z. Kurth-Nelson. What is a cognitive map? organizing knowledge for flexible behavior. Neuron, 100(2):490–509, 2018

  8. [16]

    J. J. Bakermans, J. Warren, J. C. Whittington, and T. E. Behrens. Constructing future behavior in the hippocampal formation through composition and replay.Nature Neuroscience, pages 1–12, 2025

  9. [17]

    W. Tang, H. Chang, C. Liu, S. Perez-Hernandez, W. Y . Zheng, J. Park, A. Oliva, and A. Fernandez-Ruiz. A hippocampal population code for rapid generalization.bioRxiv, pages 2025–03, 2025

  10. [18]

    Higgins, L

    I. Higgins, L. Matthey, A. Pal, C. P. Burgess, X. Glorot, M. M. Botvinick, S. Mohamed, and A. Lerchner. beta-vae: Learning basic visual concepts with a constrained variational frame- work. In5th International Conference on Learning Representations, ICLR 2017, Toulon, France, A...

  11. [19]

    K. Hsu, W. Dorrell, J. C. R. Whittington, J. Wu, and C. Finn. Disentanglement via latent quantization. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, 11 editors,Advances in Neural Information Processing Systems 36: Annual Conference on Neu- ral Informa...

  12. [20]

    Locatello, D

    F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszko- reit, A. Dosovitskiy, and T. Kipf. Object-centric learning with slot attention. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Ad- vances in Neural Information Processin...

  13. [21]

    Yarats, A

    D. Yarats, A. Zhang, I. Kostrikov, B. Amos, J. Pineau, and R. Fergus. Improving sample ef- ficiency in model-free reinforcement learning from images. InThirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Thirty-Third Conference on Innovative Applications of Art...

  14. [22]

    Higgins, A

    I. Higgins, A. Pal, A. A. Rusu, L. Matthey, C. P. Burgess, A. Pritzel, M. M. Botvinick, C. Blundell, and A. Lerchner. DARLA: improving zero-shot transfer in reinforcement learn- ing. In D. Precup and Y . W. Teh, editors,Proceedings of the 34th International Confer- ence on Mac...

  15. [23]

    Batra and G

    S. Batra and G. S. Sukhatme. Zero-shot generalization of vision-based RL without data aug- mentation.CoRR, abs/2410.07441, 2024. doi:10.48550/ARXIV .2410.07441. URLhttps: //doi.org/10.48550/arXiv.2410.07441

  16. [24]

    Hansen, H

    N. Hansen, H. Su, and X. Wang. Stabilizing deep q-learning with convnets and vision trans- formers under data augmentation. 2021

  17. [25]

    C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 20...

  18. [26]

    Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu. 3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors,Robotics: Science and Systems XX, Delft, The Netherlands, July 15...

  19. [27]

    T. Ke, N. Gkanatsios, and K. Fragkiadaki. 3d diffuser actor: Policy diffusion with 3d scene representations. In P. Agrawal, O. Kroemer, and W. Burgard, editors,Conference on Robot Learning, 6-9 November 2024, Munich, Germany, volume 270 ofProceedings of Machine Learning Resear...

  20. [28]

    D. Wang, R. Walters, and R. Platt. $\mathrm{SO}(2)$-equivariant reinforcement learning. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net, 2022. URLhttps://openreview.net/forum?id= 7F9cOhdvfk_. 12

  21. [29]

    K. Chen, X. Chen, Z. Yu, M. Zhu, and H. Yang. Equidiff: A conditional equivariant diffu- sion model for trajectory prediction. In25th IEEE International Conference on Intelligent Transportation Systems, ITSC 2022, Macau, China, October 8-12, 2022, pages 746–751. IEEE, 2023. do...

  22. [30]

    A. K. Mondal, S. S. Panigrahi, O. Kaba, S. Mudumba, and S. Ravanbakhsh. Equiv- ariant adaptation of large pretrained models. In A. Oh, T. Naumann, A. Glober- son, K. Saenko, M. Hardt, and S. Levine, editors,Advances in Neural Infor- mation Processing Systems 36: Annual Confere...

  23. [31]

    S. Kaba, A. K. Mondal, Y . Zhang, Y . Bengio, and S. Ravanbakhsh. Equivariance with learned canonicalization functions. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors,International Conference on Machine Learning, ICML 2023, 23- 29 July 2...

  24. [32]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms.CoRR, abs/1707.06347, 2017. URLhttp://arxiv.org/abs/1707.06347

  25. [33]

    Schulman, S

    J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz. Trust region policy opti- mization. In F. R. Bach and D. M. Blei, editors,Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 ofJMLR Workshop a...

  26. [34]

    V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In M. Balcan and K. Q. Weinberger, editors,Proceedings of the 33nd International Conference on Ma- chine Learning, ICML ...

  27. [35]

    Haarnoja, A

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum en- tropy deep reinforcement learning with a stochastic actor. In J. G. Dy and A. Krause, ed- itors,Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholms...

  28. [36]

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra. Continuous control with deep reinforcement learning. In Y . Bengio and Y . LeCun, editors,4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Ri...

  29. [37]

    Fujimoto, H

    S. Fujimoto, H. van Hoof, and D. Meger. Addressing function approximation error in actor- critic methods. In J. G. Dy and A. Krause, editors,Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsm ¨assan, Stockholm, Sweden, July 10-15, 2018...

  30. [38]

    V . Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Ried- miller. Playing atari with deep reinforcement learning.CoRR, abs/1312.5602, 2013. URL http://arxiv.org/abs/1312.5602

  31. [39]

    Hafner, T

    D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. In8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReview.net, 2020. URLhttps:// openreview.net/foru...

  32. [40]

    Hafner, T

    D. Hafner, T. P. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson. Learning latent dynamics for planning from pixels. In K. Chaudhuri and R. Salakhutdinov, editors, Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 201...

  33. [41]

    Hansen, H

    N. Hansen, H. Su, and X. Wang. TD-MPC2: scalable, robust world models for continuous control. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps://openreview.net/ forum?id=Oxh5CstDJU

  34. [42]

    Ha and J

    D. Ha and J. Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2018

  35. [43]

    B. Tang, M. A. Lin, I. Akinola, A. Handa, G. S. Sukhatme, F. Ramos, D. Fox, and Y . S. Narang. Industreal: Transferring contact-rich assembly tasks from simulation to reality. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daeg...

  36. [44]

    B. Tang, I. Akinola, J. Xu, B. Wen, A. Handa, K. V . Wyk, D. Fox, G. S. Sukhatme, F. Ramos, and Y . S. Narang. Automate: Specialist and generalist assembly policies over diverse geome- tries. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors,Robotics: Science and...

  37. [45]

    Handa, A

    A. Handa, A. Allshire, V . Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. V . Wyk, A. Zhurkevich, B. Sundaralingam, and Y . S. Narang. Dextreme: Transfer of agile in-hand manipulation from simulation to reality. InIEEE International Conference on Robotics and A...

  38. [46]

    Tassa, Y

    Y . Tassa, Y . Doron, A. Muldal, T. Erez, Y . Li, D. de Las Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, T. P. Lillicrap, and M. A. Riedmiller. Deepmind control suite.CoRR, abs/1801.00690, 2018. URLhttp://arxiv.org/abs/1801.00690

  39. [47]

    Mandlekar, D

    A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. In A. Faust, D. Hsu, and G. Neumann, editors,Conference on Robot Learn...

  40. [48]

    Haldar, J

    S. Haldar, J. Pari, A. Rai, and L. Pinto. Teach a robot to FISH: versatile imitation from one minute of demonstrations. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors, Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi: 10.1...

  41. [49]

    J. Ho, A. Jain, and P. Abbeel. Denoising diffusion probabilistic models. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors,Ad- vances in Neural Information Processing Systems 33: Annual Conference on Neu- ral Information Processing Systems 2020, NeurIPS ...

  42. [50]

    Y . Zeng, M. Suganuma, and T. Okatani. Inverting the generation process of denoising diffusion implicit models: Empirical evaluation and a novel method. InIEEE/CVF Winter Conference on Applications of Computer Vision, WACV 2025, Tucson, AZ, USA, February 26 - March 6, 2025, pa...

  43. [51]

    Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenRe- view.ne...

  44. [52]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674– 10685. IEEE, 2022. doi:1...

  45. [53]

    Zhang, M

    X. Zhang, M. Chang, P. Kumar, and S. Gupta. Diffusion meets dagger: Supercharging eye- in-hand imitation learning. In D. Kulic, G. Venture, K. E. Bekris, and E. Coronado, editors, Robotics: Science and Systems XX, Delft, The Netherlands, July 15-19, 2024, 2024. doi: 10.15607/R...

  46. [54]

    Pearce, T

    T. Pearce, T. Rashid, A. Kanervisto, D. Bignell, M. Sun, R. Georgescu, S. V . Macua, S. Z. Tan, I. Momennejad, K. Hofmann, and S. Devlin. Imitating human behaviour with diffusion models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rw...

  47. [55]

    A. Vaswani. Attention is all you need.Advances in Neural Information Processing Systems, 2017

  48. [56]

    Shridhar, L

    M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. In K. Liu, D. Kulic, and J. Ichnowski, editors,Conference on Robot Learning, CoRL 2022, 14-18 December 2022, Auckland, New Zealand, volume 205 ofProceedings of Machine Lea...

  49. [57]

    Goyal, J

    A. Goyal, J. Xu, Y . Guo, V . Blukis, Y . Chao, and D. Fox. RVT: robotic view transformer for 3d object manipulation. In J. Tan, M. Toussaint, and K. Darvish, editors,Conference on Robot Learning, CoRL 2023, 6-9 November 2023, Atlanta, GA, USA, volume 229 ofProceedings of Mach...

  50. [58]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. In K. E. Bekris, K. Hauser, S. L. Herbert, and J. Yu, editors,Robotics: Science and Systems XIX, Daegu, Republic of Korea, July 10-14, 2023, 2023. doi:10.15607/ R...

  51. [59]

    Garrido, N

    Q. Garrido, N. Ballas, M. Assran, A. Bardes, L. Najman, M. Rabbat, E. Dupoux, and Y . LeCun. Intuitive physics understanding emerges from self-supervised pretraining on natural videos. CoRR, abs/2502.11831, 2025. doi:10.48550/ARXIV .2502.11831. URLhttps://doi.org/ 10.48550/arX...

  52. [60]

    J. Pari, N. M. M. Shafiullah, S. P. Arunachalam, and L. Pinto. The surprising effectiveness of representation learning for visual imitation. In K. Hauser, D. A. Shell, and S. Huang, ed- itors,Robotics: Science and Systems XVIII, New York City, NY, USA, June 27 - July 1, 2022,

  53. [61]

    Samborska, J

    V . Samborska, J. L. Butler, M. E. Walton, T. E. Behrens, and T. Akam. Complementary task representations in hippocampus and prefrontal cortex for generalizing the structure of problems.Nature Neuroscience, 25(10):1314–1326, 2022

  54. [62]

    W. Sun, M. Advani, N. Spruston, A. Saxe, and J. E. Fitzgerald. Organizing memories for generalization in complementary learning systems.Nature neuroscience, 26(8):1438–1448, 2023

  55. [63]

    J. C. R. Whittington, W. Dorrell, S. Ganguli, and T. Behrens. Disentanglement with biolog- ical constraints: A theory of functional cell types. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net,

  56. [64]

    K. Hsu, J. I. Hamid, K. Burns, C. Finn, and J. Wu. Tripod: Three complementary inductive biases for disentangled representation learning. InForty-first International Conference on Ma- chine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024. URL https...

  57. [65]

    Higgins, N

    I. Higgins, N. Sonnerat, L. Matthey, A. Pal, C. P. Burgess, M. Bosnjak, M. Shanahan, M. M. Botvinick, D. Hassabis, and A. Lerchner. SCAN: learning hierarchical compositional visual concepts. In6th International Conference on Learning Representations, ICLR 2018, Vancou- ver, BC...

  58. [66]

    J. J. Hopfield. Neural networks and physical systems with emergent collective computational abilities.Proceedings of the national academy of sciences, 79(8):2554–2558, 1982

  59. [67]

    Ramsauer, B

    H. Ramsauer, B. Sch ¨afl, J. Lehner, P. Seidl, M. Widrich, L. Gruber, M. Holzleitner, T. Adler, D. P. Kreil, M. K. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter. Hopfield networks is all you need. In9th International Conference on Learning Representations, ICLR 2021, ...

  60. [68]

    Hoover, Y

    B. Hoover, Y . Liang, B. Pham, R. Panda, H. Strobelt, D. H. Chau, M. Zaki, and D. Krotov. Energy transformer.Advances in Neural Information Processing Systems, 36, 2024

  61. [69]

    B. Zhou, A. Lapedriza, A. Khosla, A. Oliva, and A. Torralba. Places: A 10 million image database for scene recognition.IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 2017. 16 A Latent Traversals Figure 5: Traversing the disentangled latent space of ALDA-DP t...

  62. [2018]

    URLhttps://openreview.net/forum?id=rkN2Il-RZ

  63. [2022]

    URLhttps://doi.org/10.15607/RSS.2022

    doi:10.15607/RSS.2022.XVIII.010. URLhttps://doi.org/10.15607/RSS.2022. XVIII.010

  64. [2023]

    URLhttps://openreview.net/forum?id=9Z_GfhZnGH

  65. [2713]

    URLhttps://proceedings.mlr.press/v270/kim25c.html

    PMLR, 2024. URLhttps://proceedings.mlr.press/v270/kim25c.html

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.