Pith. sign in

REVIEW 4 major objections 5 minor 3 cited by

Object-Centric Representations Improve Policy Generalization in Robot Manipulation

T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Object-centric vision lifts robot manipulation generalization

desk verdict Useful benchmark, but the headline causal claim about object-centric structure is confounded with robot-data pretraining and an unreported DINOv2 dense baseline. read the letter →

arxiv 2505.11563 v1 pith:JOBZIBO3 submitted 2025-05-16 cs.RO cs.AIeess.IV

classification cs.ROcs.AIeess.IV
keywords object-centricrepresentationsslotattentionrobotmanipulationpolicygeneralizationimitationlearningvisualrepresentationdistributionshiftVIDEOSAUR
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to test whether object-centric representations—encodings that split an image into a small set of entity slots rather than a global vector or dense patch map—give robot manipulation policies better generalization under visual shifts such as new lighting, textures, and distractors. Across two simulation benchmarks and a five-task real-world setup, the authors find that slot-based encoders, especially their robot-pretrained VIDEOSAUR variant, outperform dense and global baselines in out-of-distribution conditions and match or beat them in-domain. The headline result is a 70% real-world success rate for VIDEOSAUR* versus 50% for the best dense baseline. The authors argue this makes object-centric encoders a promising direction for designing robust visuomotor policies in real-world environments.

What carries the argument

The load-bearing mechanism is Slot Attention with a frozen DINOv2 vision backbone, as realized in DINOSAUR and extended in VIDEOSAUR. Slot Attention is an iterative cross-attention module that compresses N dense patch features into K slot vectors through a softmax renormalization over slots, so each slot specializes on one entity. VIDEOSAUR adds a transformer predictor that initializes slots at time t from slots at t−1 plus a temporal consistency loss, and the paper's VIDEOSAUR* retrains the slot-attention module on a mixture of robot manipulation videos to align the slots with manipulation dynamics. These slot vectors are fed, frozen, into transformer-based policy heads (BAKU in simulation, ACT in the real world), replacing the usual global or dense feature input.

What would settle it

Run the VIDEOSAUR* pipeline with the slot-attention module replaced by a mean-pooling or learned pooling of the same frozen DINOv2 features, keeping the temporal transformer and robot-mixture pretraining; if that pooled variant still reaches roughly 70% real-world success and similar LIBERO numbers, the paper's attribution of the gains to object-centric structure is falsified.

Watch

Extended reading notes

Core claim

The central claim is that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining. In the paper's comparison, VIDEOSAUR*—a slot-attention video model with a DINOv2 backbone, a temporal transformer predictor, and slot attention retrained on a mixture of 188k robot trajectories—achieves the highest average success in LIBERO-90 and in the real-world suite, reaching 70% success versus 50% for the best dense baseline, while remaining competitive in MetaWorld. The paper attributes this to the inductive bias of slot attention: decomposing the scene into discrete entities lets the policy ignore task-irrelevant background and stay robust to appearance changes. The authors also report that robot-data pretraining and temporal dynamics modeling each contribute large gains, with VIDEOSAUR* beating DINOSAUR* by 9 and 26 points in LIBERO and the real-world suite, respectively.

Load-bearing premise

The load-bearing premise is that the performance gap comes from the object-centric inductive bias itself, but the best model differs from the dense baselines in several simultaneous ways—a DINOv2 backbone, a temporal transformer, and slot-attention pretraining on 188k robot trajectories—so no single factor is isolated, and calling those robot datasets 'not task-specific' is itself an assumption.

Editorial extensions

If this is right

  • If the central claim holds, robot policy designers can expect slot-based encoders to provide more robust performance under lighting, texture, and distractor shifts than dense or global encoders.
  • Pretraining the slot-attention module on large robot video collections is a key lever: VIDEOSAUR* adds 10–13 points in mean success over VIDEOSAUR across the three environments.
  • Temporal dynamics in the object-centric encoder matter: VIDEOSAUR* beats DINOSAUR* by 9 points in LIBERO and 26 points in the real-world suite when both use the same robot-mixture pretraining.
  • Object-centric models remain competitive in-domain while winning out-of-distribution, so switching to them does not sacrifice standard performance in these benchmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A matched ablation that replaces the slot-attention module with pooling of the same frozen DINOv2 features, keeping the temporal transformer and robot-mixture pretraining, would settle whether the object-centric structure itself or the extra components cause the gains.
  • The 'without task-specific pretraining' phrasing is definitionally fragile: all three pretraining sources are manipulation datasets, so a stricter reading is that the gains survive when the objective is reconstruction rather than action prediction.
  • The paper's own slot visualizations suggest that grounding slots semantically, for example with affordances, could reduce distractor capture and is a natural next test for the approach.
  • If object-centric encoders are adopted more widely, the practical benchmark to watch is whether they continue to dominate when dense baselines are given the same backbone, data, and temporal modeling.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper investigates whether object-centric representations (OCRs), specifically slot-based encoders DINOSAUR and VIDEOSAUR, improve the learning and generalization of visuomotor policies compared with global and dense visual representations. The authors introduce robot-pretrained variants DINOSAUR* and VIDEOSAUR* obtained by training the slot-attention module on a mixture of BridgeData V2, Fractal, and DROID, and evaluate all models on MetaWorld, LIBERO-90, and five real-world LeRobot tasks under in-domain and shifted conditions (distractors, textures, lighting). The central claim is that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining, with VIDEOSAUR* reported as the strongest method.

Significance. If the central claim is established, the paper would provide a useful benchmark and a concrete argument for object-centric inductive biases in robotic manipulation, supported by a unified framework, frozen-encoder comparisons across simulation and real hardware, and open-source release of the evaluation code. The paper also contributes a practical pretraining recipe for slot-attention models on large robot datasets and reports task-level performance. However, the evidence as presented does not yet isolate the OCR inductive bias from several confounds, and some headline numbers rest on very small real-world evaluation budgets.

major comments (4)
  1. [Section 4, 'Robotic pre-training'; Table 6] The headline comparison between VIDEOSAUR* and the dense baselines is confounded by the robot-mixture pretraining. Table 6 shows that pretraining on the robot mixture improves VIDEOSAUR from 0.77 to 0.86 in LIBERO and from 0.58 to 0.70 in the real-world setup, gains of the same magnitude as the reported OCR advantage over dense baselines. Without a dense or global encoder trained on the same robot-mixture data, the observed gains cannot be attributed specifically to object-centric structure rather than to pretraining-data alignment. Please add an ablation that trains a non-OCR baseline on the same robot mixture, or otherwise remove the attribution of these gains to the OCR inductive bias.
  2. [Appendix E, 'Baselines details'] The DINOv2 dense representation is mentioned but never reported: the appendix states that 'the Global representation was always outperforming the other alternative,' yet no dense DINOv2 numbers are shown anywhere. Since VIDEOSAUR* uses a DINOv2 ViT-B14 backbone with slot attention on top, the missing DINOv2 dense baseline is exactly the control needed to determine whether slot attention adds anything over the frozen DINOv2 patch features. Please report the DINOv2 dense result in all tables, or justify its omission with explicit numbers.
  3. [Section 5.1; Table 3; Figure 3] The real-world results are based on only 10 rollouts per task and are reported without error bars or per-seed variation, despite the stated protocol of three random seeds. Table 3 shows differences such as VIDEOSAUR* 0.44 overall versus VIDEOSAUR 0.40 which are likely within the noise of 10 rollouts per condition, and Figure 3 has no error bars at all. The claim that OCRs 'consistently' generalize better in the real world needs confidence intervals, more rollouts, or per-seed results; otherwise the 70% versus 50% headline comparison is not robustly supported.
  4. [Abstract; Section 5.1; Table 2] The claim that OCRs outperform dense and global representations 'even without task-specific pretraining' is not supported by the presented comparisons. The robot-mixture pretraining is manipulation-specific pretraining, so the starred models do not satisfy the 'without task-specific pretraining' condition. Moreover, in Table 2 the non-robot-pretrained OCR models, DINOSAUR at 0.46 and VIDEOSAUR at 0.41, do not beat Theia at 0.47 on MetaWorld overall. Please either qualify the claim to distinguish robot-pretrained and non-robot-pretrained OCR variants, or provide evidence that non-robot-pretrained OCRs consistently surpass strong dense baselines.
minor comments (5)
  1. [Section 3, Eq. (1)] The dimensions in Eq. (1) are inconsistent: the text defines Q in R^{NxD} but writes K in R^{KxD}, while Slot Attention normally projects queries from K slots and keys from N features; please clarify the notation so that the softmax dimensions match the description.
  2. [Appendix G and Appendix H] The captions of Figure 7 and Figure 8 say '12 slots' and '8 slots', respectively, while Table 4 specifies 10 slots for both DINOSAUR and VIDEOSAUR; please reconcile these numbers.
  3. [Section 3, 'Object-centric representation for videos'] There are several typos in this section, including 'Convolutionnal' and 'alse' for 'also'; a proofreading pass would improve readability.
  4. [Table 8] In Table 8, the VIDEOSAUR row is cited as '[18]' but should reference [41]; please correct the citation.
  5. [Section 5.2, paragraph 2] The sentence 'In real-world evaluations as can be seen in Table 3' refers to the table that follows, but the preceding sentence also references Table 2; please make the table references unambiguous.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the paper's claims are supported by external benchmark evaluations rather than by fitting or self-referential derivation.

full rationale

This is an empirical benchmarking paper, not a derivation. Its central claim—that OCR-based policies outperform dense and global representations in generalization settings—is supported by measured success rates on MetaWorld, LIBERO, and a real-world LeRobot suite, with frozen encoders and held-out rollouts. No fitted parameter is renamed as a prediction: the robot-mixture pretraining of DINOSAUR* and VIDEOSAUR* is a method variant, and Table 6 exposes rather than conceals its effect. There is no load-bearing self-citation: DINOSAUR and VIDEOSAUR are prior external works; no uniqueness theorem is invoked; no target result appears as an assumption. The 'even without task-specific pretraining' claim is tested by the unstarred OCR variants (e.g., VIDEOSAUR reaches 0.58 real-world vs Theia 0.32), so it is not definitional. The main concerns are empirical confounds and reporting gaps: VIDEOSAUR* differs from dense baselines in backbone, temporal transformer, and robot-data pretraining, and Appendix E reports that the DINOv2 global representation was shown because it outperformed its dense variant, so the dense DINOv2 comparison is omitted. These are experimental-control issues, not circular reasoning. The Limitations section candidly lists failures (slots capturing background and distractors, no dynamics alignment, limited scale) and contains no assertion of circularity. Therefore no circular step can be exhibited.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The benchmark rests on several hand-set architecture choices and domain assumptions that are not independently justified, most importantly the use of frozen encoders, DINOv2 as the OCR backbone, the 10-slot configuration, and the robot-mixture pretraining selected after single-source ablations. No invented entities are introduced.

free parameters (4)
  • Number of slots K = 10
    Hand-set for all OCR models; the number of objects in a scene varies across tasks, so a fixed K=10 may limit or bias the slot decomposition.
  • Slot size = 128
    Dimension of each slot vector; chosen from the original DINOSAUR/VIDEOSAUR papers and used without task-specific tuning.
  • Slot Attention iterations = 3
    Number of iterative refinement steps in Slot Attention; hand-set and identical across OCR models, affecting how cleanly slots separate.
  • Robot-mixture pretraining composition = Balanced mixture of BridgeData V2, Fractal, DROID (188k trajectories)
    The mixture was selected after comparing single-source ablations in Table 6, which introduces a selection bias for the downstream results.
assumptions (4)
  • domain assumption Frozen pretrained visual encoders are a sufficient basis for comparing representations
    All encoders are kept frozen during policy learning; if fine-tuning were allowed, the ranking of representations could change. Entered in Section 3 'Policy training'.
  • domain assumption DINOv2 ViT-B14 is an appropriate backbone for OCR models and provides features comparable to the DINOv2 baseline
    DINOSAUR and VIDEOSAUR use DINOv2 features, while the DINOv2 baseline is reported only as a global pooled vector; the comparison assumes backbone quality is not the main confound. Entered in Appendix A and Table 1.
  • domain assumption The robot-mixture pretraining is not task-specific tuning
    The paper claims gains 'even without task-specific pretraining', but BridgeData, Fractal, and DROID are robot manipulation datasets, so the assumption that this is not task-specific is questionable. Entered in Section 4 'Robotic pre-training'.
  • domain assumption Real-world tasks with 10 rollouts per task provide stable success-rate estimates
    Success rates in Table 3 are reported without error bars; the assumption that 10 rollouts is sufficient underlies the real-world claims. Entered in Section 5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Object-Centric Representations Improve Policy Generalization in Robot Manipulation." pith.science (2026). https://pith.science/paper/JOBZIBO3

@misc{pith2026250511563,
  author       = {Pith},
  title        = {Pith review of: Object-Centric Representations Improve Policy Generalization in Robot Manipulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JOBZIBO3}},
  note         = {Machine review of arXiv:2505.11563}
}
read the original abstract

Visual representations are central to the learning and generalization capabilities of robotic manipulation policies. While existing methods rely on global or dense features, such representations often entangle task-relevant and irrelevant scene information, limiting robustness under distribution shifts. In this work, we investigate object-centric representations (OCR) as a structured alternative that segments visual input into a finished set of entities, introducing inductive biases that align more naturally with manipulation tasks. We benchmark a range of visual encoders-object-centric, global and dense methods-across a suite of simulated and real-world manipulation tasks ranging from simple to complex, and evaluate their generalization under diverse visual conditions including changes in lighting, texture, and the presence of distractors. Our findings reveal that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining. These insights suggest that OCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.

Figures

Figures reproduced from arXiv: 2505.11563 by the authors.

Figure 1
Figure 1. Overall architecture. We use a set of pre-trained visual models with different structures of latent space - global, dense and object-centric - (a) as input to a policy model for robotic manip￾ulation learning (b). We showcase the benefits of Object-Centric Representations (VIDEOSAUR, DINOSAUR) over different visual models on the final performance of policies and the generaliza￾tion capabilities in simulation and rea… view at source ↗
Figure 2
Figure 2. In simulation, we use MetaWorld [46], a well-established benchmark comprising tabletop manip￾ulation tasks performed with a Sawyer robotic arm. MetaWorld offers a controlled setting with standardized tasks and supports structured generalization testing, making it a suitable baseline to evaluate representation performance in simple, single-object scenarios. To challenge the scalability of object-centric representatio… view at source ↗
Figure 3
Figure 3. Overall success rate. Mean success rate over all tasks for each visual model on the three environments, e.g., MetaWorld (left), LIBERO (middle) and Real using LeRobot (right). DI￾NOSAUR* and VIDEOSAUR* have been pretrained over robot data mixture. 5.1 Q1: Do OCRs improve manipulation policy learning? [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Real world setup. Our setup is based on the LeRobot library. We use a SO-100 arm on a tabletop environment with two realsense cameras, one with a top overview of the scene and a second with a side view [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Overview tasks real-world. From left to right: Banana bowl, Open drawer, Pick coffee, Close drawer, Fold bag. E Baselines details This section details the 7 state-of-the-art visual models we compare in our experiments. It covers different architecture types (Convolutio…
Figure 6
Figure 6. Figure 6: Overview of different generalization levels. From left to right: new distractors, new textures of the table, new lighting conditions. Top row: Metaworld. Bottom row: Real-World and time-contrastive learning, making it particularly effective for robotic tasks that requi…
Figure 7
Figure 7. Figure 7: Slots visualization. Visualization of a set of slots (12 slots) extracted from VIDEOSAUR* model on the easy distractor setup in Metaworld 19 [PITH_FULL_IMAGE:figures/full_fig_p019_7.png]
Figure 8
Figure 8. Figure 8: Failure slots visualization. Visualization of a set of slots (8 slots) extracted from VIDEOSAUR* model on the hard distractor setup in Metaworld. The slot that should handle the hammer to perform the task also capture the brown box, leading to noise in the subsequent s…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning

    cs.RO 2026-07 conditional novelty 6.0 of 10

    FORGE decouples robotic tool-use into keypoint trajectory prediction from action-free data and action grounding from limited demonstrations, achieving over 2X improvement in functional generalization to unseen tools.

  2. STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation

    cs.RO 2026-01 conditional novelty 5.0 of 10

    STORM uses a two-stage, text-guided slot attention module on frozen DINOv2 features to improve robot manipulation success and generalization to visual distractors in simulated benchmarks.

  3. Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

    cs.RO 2025-08 conditional novelty 4.0 of 10

    A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.

Reference graph

Works this paper leans on

63 extracted references · 9 canonical work pages · cited by 3 Pith papers

  1. [1]

    Haldar, Z

    S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learning,

  2. [2]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Man- junath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...

  3. [3]

    O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy, 2024. URL https: //arxiv.org/abs/2405.12213

  4. [4]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model, 2024. URL https://arxiv.org/abs/2406.09246

  5. [5]

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3m: A universal visual represen- tation for robot manipulation, 2022. URL https://arxiv.org/abs/2203.12601

  6. [6]

    Majumdar, K

    A. Majumdar, K. Yadav, S. Arnaud, Y . J. Ma, C. Chen, S. Silwal, A. Jain, V .-P. Berges, P. Abbeel, J. Malik, D. Batra, Y . Lin, O. Maksymets, A. Rajeswaran, and F. Meier. Where are we in the search for an artificial visual cortex for embodied intelligence?, 2024. URL https://arxiv.org/abs/2303.18240. 10

  7. [7]

    Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training, 2023. URL https: //arxiv.org/abs/2210.00030

  8. [8]

    Shang, K

    J. Shang, K. Schmeckpeper, B. B. May, M. V . Minniti, T. Kelestemur, D. Watkins, and L. Her- lant. Theia: Distilling diverse vision foundation models for robot learning, 2024. URL https://arxiv.org/abs/2407.20179

Show all 63 references
  1. [9]

    Radosavovic, T

    I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell. Real-world robot learn- ing with masked visual pre-training, 2022. URL https://arxiv.org/abs/2210.03109

  2. [10]

    Jiang, Y

    G. Jiang, Y . Sun, T. Huang, H. Li, Y . Liang, and H. Xu. Robots pre-train robots: Manipulation- centric robotic representation from large-scale robot dataset.arXiv preprint arXiv:2410.22325, 2024

  3. [11]

    Parisi, A

    S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta. The unsurprising effectiveness of pre-trained vision models for control, 2022. URL https://arxiv.org/abs/2203.03580

  4. [12]

    Burns, Z

    K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman. What makes pre-trained visual representations successful for robust manipulation?, 2023. URLhttps://arxiv.org/ abs/2312.12444

  5. [13]

    B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. Building machines that learn and think like people, 2016. URL https://arxiv.org/abs/1604.00289

  6. [14]

    Greff, S

    K. Greff, S. van Steenkiste, and J. Schmidhuber. On the binding problem in artificial neural networks, 2020. URL https://arxiv.org/abs/2012.05208

  7. [15]

    Kroemer, S

    O. Kroemer, S. Niekum, and G. Konidaris. A review of robot learning for manipulation: Challenges, representations, and algorithms, 2020. URL https://arxiv.org/abs/1907. 03146

  8. [16]

    Bengio, A

    Y . Bengio, A. Courville, and P. Vincent. Representation learning: A review and new perspec- tives, 2014. URL https://arxiv.org/abs/1206.5538

  9. [17]

    Locatello, D

    F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf. Object-centric learning with slot attention, 2020. URL https: //arxiv.org/abs/2006.15055

  10. [18]

    Seitzer, M

    M. Seitzer, M. Horn, A. Zadaianchuk, D. Zietlow, T. Xiao, C.-J. Simon-Gabriel, T. He, Z. Zhang, B. Sch ¨olkopf, T. Brox, and F. Locatello. Bridging the gap to real-world object- centric learning, 2023. URL https://arxiv.org/abs/2209.14860

  11. [19]

    Yoon, Y .-F

    J. Yoon, Y .-F. Wu, H. Bae, and S. Ahn. An investigation into pre-training object-centric repre- sentations for reinforcement learning, 2023. URL https://arxiv.org/abs/2302.04419

  12. [20]

    Heravi, A

    N. Heravi, A. Wahid, C. Lynch, P. Florence, T. Armstrong, J. Tompson, P. Sermanet, J. Bohg, and D. Dwibedi. Visuomotor control in multi-object scenes using object-aware representations,

  13. [21]

    Haramati, T

    D. Haramati, T. Daniel, and A. Tamar. Entity-centric reinforcement learning for object manip- ulation from pixels, 2024. URL https://arxiv.org/abs/2404.01220

  14. [22]

    Watters, L

    N. Watters, L. Matthey, M. Bosnjak, C. P. Burgess, and A. Lerchner. Cobra: Data-efficient model-based rl through unsupervised object discovery and curiosity-driven exploration, 2019. URL https://arxiv.org/abs/1905.09275

  15. [23]

    T. Kipf, G. F. Elsayed, A. Mahendran, A. Stone, S. Sabour, G. Heigold, R. Jonschkowski, A. Dosovitskiy, and K. Greff. Conditional object-centric learning from video, 2022. URL https://arxiv.org/abs/2111.12594. 11

  16. [24]

    Zhang, A

    C. Zhang, A. Gupta, and A. Zisserman. Is an object-centric video representation beneficial for transfer?, 2022. URL https://arxiv.org/abs/2207.10075

  17. [25]

    X. Chen, H. Fan, R. Girshick, and K. He. Improved baselines with momentum contrastive learning, 2020. URL https://arxiv.org/abs/2003.04297

  18. [26]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers, 2021. URL https://arxiv.org/abs/ 2104.14294

  19. [27]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A...

  20. [28]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020

  21. [29]

    Grauman, A

    K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Radosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, D. Batra, V . Cartillier, S. Crane, T. D...

  22. [30]

    Dasari, M

    S. Dasari, M. K. Srirama, U. Jain, and A. Gupta. An unbiased look at datasets for visuo-motor pre-training, 2023. URL https://arxiv.org/abs/2310.09289

  23. [31]

    Hamdan and F

    S. Hamdan and F. G ¨uney. Carformer: Self-driving with learned object-centric representations,

  24. [32]

    Mosbach, J

    M. Mosbach, J. N. Ewertz, A. Villar-Corrales, and S. Behnke. Sold: Slot object-centric latent dynamics models for relational manipulation learning from pixels, 2025. URL https:// arxiv.org/abs/2410.08822

  25. [33]

    B. Wang, L. Li, J. Zhang, Y . Nakashima, and H. Nagahara. Explainable image recognition via enhanced slot-attention based classifier, 2024. URL https://arxiv.org/abs/2407. 05616

  26. [34]

    URL https://arxiv.org/abs/2407.15843

  27. [35]

    C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner. Monet: Unsupervised scene decomposition and representation, 2019. URL https://arxiv. org/abs/1901.11390

  28. [36]

    Jiang, F

    J. Jiang, F. Deng, G. Singh, and S. Ahn. Object-centric slot diffusion, 2023. URL https: //arxiv.org/abs/2303.10834. 12

  29. [37]

    Kabra, D

    R. Kabra, D. Zoran, G. Erdogan, L. Matthey, A. Creswell, M. Botvinick, A. Lerchner, and C. P. Burgess. Simone: View-invariant, temporally-abstracted object representations via unsu- pervised video decomposition, 2021. URL https://arxiv.org/abs/2106.03849

  30. [38]

    Singh, F

    G. Singh, F. Deng, and S. Ahn. Illiterate dall-e learns to compose, 2022. URL https:// arxiv.org/abs/2110.11405

  31. [39]

    G. F. Elsayed, A. Mahendran, S. van Steenkiste, K. Greff, M. C. Mozer, and T. Kipf. Savi++: Towards end-to-end object-centric learning from real-world videos, 2022. URL https:// arxiv.org/abs/2206.07764

  32. [40]

    Z. Wu, J. Hu, W. Lu, I. Gilitschenski, and A. Garg. Slotdiffusion: Object-centric generative modeling with diffusion models, 2023. URL https://arxiv.org/abs/2305.11281

  33. [41]

    Zadaianchuk, M

    A. Zadaianchuk, M. Seitzer, and G. Martius. Object-centric learning for real-world videos by predicting temporal feature similarities, 2023. URLhttps://arxiv.org/abs/2306.04829

  34. [42]

    Didolkar, A

    A. Didolkar, A. Zadaianchuk, A. Goyal, M. Mozer, Y . Bengio, G. Martius, and M. Seitzer. Zero-shot object-centric representation learning, 2024. URL https://arxiv.org/abs/ 2408.09162

  35. [43]

    Singh, Y .-F

    G. Singh, Y .-F. Wu, and S. Ahn. Simple unsupervised object-centric learning for complex and naturalistic videos, 2022. URL https://arxiv.org/abs/2205.14065

  36. [44]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023. URL https://arxiv.org/abs/2304.13705

  37. [45]

    Warner, A

    B. Warner, A. Chaffin, B. Clavi ´e, O. Weller, O. Hallstr ¨om, S. Taghadouini, A. Gallagher, R. Biswas, F. Ladhak, T. Aarsen, N. Cooper, G. Adams, J. Howard, and I. Poli. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long conte...

  38. [46]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin. Attention is all you need, 2023. URL https://arxiv.org/abs/1706.03762

  39. [47]

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310, 2023

  40. [48]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition, 2015. URL https://arxiv.org/abs/1512.03385

  41. [49]

    T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforce- ment learning, 2021. URL https://arxiv.org/abs/1910.10897

  42. [50]

    Walke, K

    H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V . Myers, K. Fang, C. Finn, and S. Levine. Bridgedata v2: A dataset for robot learning at scale, 2024. URL https://arxiv.org/abs/2308.12952

  43. [51]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...

  44. [52]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. URL https: //arxiv.org/abs/2010.11929

  45. [53]

    T.-Y . Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Doll´ar. Microsoft coco: Common objects in context, 2015. URL https: //arxiv.org/abs/1405.0312

  46. [54]

    N. Xu, L. Yang, Y . Fan, D. Yue, Y . Liang, J. Yang, and T. Huang. Youtube-vos: A large-scale video object segmentation benchmark, 2018. URL https://arxiv.org/abs/1809.03327

  47. [55]

    Russakovsky, J

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. Imagenet large scale visual recogni- tion challenge, 2015. URL https://arxiv.org/abs/1409.0575

  48. [56]

    Todorov, T

    E. Todorov, T. Erez, and Y . Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages 5026–

  49. [57]

    A. Xie, L. Lee, T. Xiao, and C. Finn. Decomposing the generalization gap in imitation learning for visual robotic manipulation, 2023. URL https://arxiv.org/abs/2307.03659

  50. [58]

    Cadene, S

    R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf. Lerobot: State- of-the-art machine learning for real-world robotics in pytorch. https://github.com/ huggingface/lerobot, 2024

  51. [62]

    Mandlekar, D

    A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. In Conference on Robot Learning (CoRL) , 2021. A Implementation detail...

  52. [63]

    Figure 7: Slots visualization

    Note that the model has never seen the provided data before as it has been pre-trained on frozen to learn subsequent policy. Figure 7: Slots visualization. Visualization of a set of slots (12 slots) extracted from VIDEOSAUR* model on the easy distractor setup in Metaworld 19 H...

  53. [2023]

    URL https://arxiv.org/abs/2205.06333

  54. [2024]

    URL https://arxiv.org/abs/2406.07539

  55. [5033]

    doi:10.1109/IROS.2012.6386109

    IEEE, 2012. doi:10.1109/IROS.2012.6386109

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.