REVIEW 4 major objections 5 minor 3 cited by
Object-Centric Representations Improve Policy Generalization in Robot Manipulation
T0 review · 4 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Object-centric vision lifts robot manipulation generalization
desk verdict Useful benchmark, but the headline causal claim about object-centric structure is confounded with robot-data pretraining and an unreported DINOv2 dense baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is Slot Attention with a frozen DINOv2 vision backbone, as realized in DINOSAUR and extended in VIDEOSAUR. Slot Attention is an iterative cross-attention module that compresses N dense patch features into K slot vectors through a softmax renormalization over slots, so each slot specializes on one entity. VIDEOSAUR adds a transformer predictor that initializes slots at time t from slots at t−1 plus a temporal consistency loss, and the paper's VIDEOSAUR* retrains the slot-attention module on a mixture of robot manipulation videos to align the slots with manipulation dynamics. These slot vectors are fed, frozen, into transformer-based policy heads (BAKU in simulation, ACT in the real world), replacing the usual global or dense feature input.
What would settle it
Run the VIDEOSAUR* pipeline with the slot-attention module replaced by a mean-pooling or learned pooling of the same frozen DINOv2 features, keeping the temporal transformer and robot-mixture pretraining; if that pooled variant still reaches roughly 70% real-world success and similar LIBERO numbers, the paper's attribution of the gains to object-centric structure is falsified.
Extended reading notes
Core claim
The central claim is that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining. In the paper's comparison, VIDEOSAUR*—a slot-attention video model with a DINOv2 backbone, a temporal transformer predictor, and slot attention retrained on a mixture of 188k robot trajectories—achieves the highest average success in LIBERO-90 and in the real-world suite, reaching 70% success versus 50% for the best dense baseline, while remaining competitive in MetaWorld. The paper attributes this to the inductive bias of slot attention: decomposing the scene into discrete entities lets the policy ignore task-irrelevant background and stay robust to appearance changes. The authors also report that robot-data pretraining and temporal dynamics modeling each contribute large gains, with VIDEOSAUR* beating DINOSAUR* by 9 and 26 points in LIBERO and the real-world suite, respectively.
Load-bearing premise
The load-bearing premise is that the performance gap comes from the object-centric inductive bias itself, but the best model differs from the dense baselines in several simultaneous ways—a DINOv2 backbone, a temporal transformer, and slot-attention pretraining on 188k robot trajectories—so no single factor is isolated, and calling those robot datasets 'not task-specific' is itself an assumption.
Editorial extensions
If this is right
- If the central claim holds, robot policy designers can expect slot-based encoders to provide more robust performance under lighting, texture, and distractor shifts than dense or global encoders.
- Pretraining the slot-attention module on large robot video collections is a key lever: VIDEOSAUR* adds 10–13 points in mean success over VIDEOSAUR across the three environments.
- Temporal dynamics in the object-centric encoder matter: VIDEOSAUR* beats DINOSAUR* by 9 points in LIBERO and 26 points in the real-world suite when both use the same robot-mixture pretraining.
- Object-centric models remain competitive in-domain while winning out-of-distribution, so switching to them does not sacrifice standard performance in these benchmarks.
Reading between the lines
- A matched ablation that replaces the slot-attention module with pooling of the same frozen DINOv2 features, keeping the temporal transformer and robot-mixture pretraining, would settle whether the object-centric structure itself or the extra components cause the gains.
- The 'without task-specific pretraining' phrasing is definitionally fragile: all three pretraining sources are manipulation datasets, so a stricter reading is that the gains survive when the objective is reconstruction rather than action prediction.
- The paper's own slot visualizations suggest that grounding slots semantically, for example with affordances, could reduce distractor capture and is a natural next test for the approach.
- If object-centric encoders are adopted more widely, the practical benchmark to watch is whether they continue to dominate when dense baselines are given the same backbone, data, and temporal modeling.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether object-centric representations (OCRs), specifically slot-based encoders DINOSAUR and VIDEOSAUR, improve the learning and generalization of visuomotor policies compared with global and dense visual representations. The authors introduce robot-pretrained variants DINOSAUR* and VIDEOSAUR* obtained by training the slot-attention module on a mixture of BridgeData V2, Fractal, and DROID, and evaluate all models on MetaWorld, LIBERO-90, and five real-world LeRobot tasks under in-domain and shifted conditions (distractors, textures, lighting). The central claim is that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining, with VIDEOSAUR* reported as the strongest method.
Significance. If the central claim is established, the paper would provide a useful benchmark and a concrete argument for object-centric inductive biases in robotic manipulation, supported by a unified framework, frozen-encoder comparisons across simulation and real hardware, and open-source release of the evaluation code. The paper also contributes a practical pretraining recipe for slot-attention models on large robot datasets and reports task-level performance. However, the evidence as presented does not yet isolate the OCR inductive bias from several confounds, and some headline numbers rest on very small real-world evaluation budgets.
major comments (4)
- [Section 4, 'Robotic pre-training'; Table 6] The headline comparison between VIDEOSAUR* and the dense baselines is confounded by the robot-mixture pretraining. Table 6 shows that pretraining on the robot mixture improves VIDEOSAUR from 0.77 to 0.86 in LIBERO and from 0.58 to 0.70 in the real-world setup, gains of the same magnitude as the reported OCR advantage over dense baselines. Without a dense or global encoder trained on the same robot-mixture data, the observed gains cannot be attributed specifically to object-centric structure rather than to pretraining-data alignment. Please add an ablation that trains a non-OCR baseline on the same robot mixture, or otherwise remove the attribution of these gains to the OCR inductive bias.
- [Appendix E, 'Baselines details'] The DINOv2 dense representation is mentioned but never reported: the appendix states that 'the Global representation was always outperforming the other alternative,' yet no dense DINOv2 numbers are shown anywhere. Since VIDEOSAUR* uses a DINOv2 ViT-B14 backbone with slot attention on top, the missing DINOv2 dense baseline is exactly the control needed to determine whether slot attention adds anything over the frozen DINOv2 patch features. Please report the DINOv2 dense result in all tables, or justify its omission with explicit numbers.
- [Section 5.1; Table 3; Figure 3] The real-world results are based on only 10 rollouts per task and are reported without error bars or per-seed variation, despite the stated protocol of three random seeds. Table 3 shows differences such as VIDEOSAUR* 0.44 overall versus VIDEOSAUR 0.40 which are likely within the noise of 10 rollouts per condition, and Figure 3 has no error bars at all. The claim that OCRs 'consistently' generalize better in the real world needs confidence intervals, more rollouts, or per-seed results; otherwise the 70% versus 50% headline comparison is not robustly supported.
- [Abstract; Section 5.1; Table 2] The claim that OCRs outperform dense and global representations 'even without task-specific pretraining' is not supported by the presented comparisons. The robot-mixture pretraining is manipulation-specific pretraining, so the starred models do not satisfy the 'without task-specific pretraining' condition. Moreover, in Table 2 the non-robot-pretrained OCR models, DINOSAUR at 0.46 and VIDEOSAUR at 0.41, do not beat Theia at 0.47 on MetaWorld overall. Please either qualify the claim to distinguish robot-pretrained and non-robot-pretrained OCR variants, or provide evidence that non-robot-pretrained OCRs consistently surpass strong dense baselines.
minor comments (5)
- [Section 3, Eq. (1)] The dimensions in Eq. (1) are inconsistent: the text defines Q in R^{NxD} but writes K in R^{KxD}, while Slot Attention normally projects queries from K slots and keys from N features; please clarify the notation so that the softmax dimensions match the description.
- [Appendix G and Appendix H] The captions of Figure 7 and Figure 8 say '12 slots' and '8 slots', respectively, while Table 4 specifies 10 slots for both DINOSAUR and VIDEOSAUR; please reconcile these numbers.
- [Section 3, 'Object-centric representation for videos'] There are several typos in this section, including 'Convolutionnal' and 'alse' for 'also'; a proofreading pass would improve readability.
- [Table 8] In Table 8, the VIDEOSAUR row is cited as '[18]' but should reference [41]; please correct the citation.
- [Section 5.2, paragraph 2] The sentence 'In real-world evaluations as can be seen in Table 3' refers to the table that follows, but the preceding sentence also references Table 2; please make the table references unambiguous.
Circularity Check
No circularity found: the paper's claims are supported by external benchmark evaluations rather than by fitting or self-referential derivation.
full rationale
This is an empirical benchmarking paper, not a derivation. Its central claim—that OCR-based policies outperform dense and global representations in generalization settings—is supported by measured success rates on MetaWorld, LIBERO, and a real-world LeRobot suite, with frozen encoders and held-out rollouts. No fitted parameter is renamed as a prediction: the robot-mixture pretraining of DINOSAUR* and VIDEOSAUR* is a method variant, and Table 6 exposes rather than conceals its effect. There is no load-bearing self-citation: DINOSAUR and VIDEOSAUR are prior external works; no uniqueness theorem is invoked; no target result appears as an assumption. The 'even without task-specific pretraining' claim is tested by the unstarred OCR variants (e.g., VIDEOSAUR reaches 0.58 real-world vs Theia 0.32), so it is not definitional. The main concerns are empirical confounds and reporting gaps: VIDEOSAUR* differs from dense baselines in backbone, temporal transformer, and robot-data pretraining, and Appendix E reports that the DINOv2 global representation was shown because it outperformed its dense variant, so the dense DINOv2 comparison is omitted. These are experimental-control issues, not circular reasoning. The Limitations section candidly lists failures (slots capturing background and distractors, no dynamics alignment, limited scale) and contains no assertion of circularity. Therefore no circular step can be exhibited.
Assumptions & free parameters
free parameters (4)
- Number of slots K =
10
- Slot size =
128
- Slot Attention iterations =
3
- Robot-mixture pretraining composition =
Balanced mixture of BridgeData V2, Fractal, DROID (188k trajectories)
assumptions (4)
- domain assumption Frozen pretrained visual encoders are a sufficient basis for comparing representations
- domain assumption DINOv2 ViT-B14 is an appropriate backbone for OCR models and provides features comparable to the DINOv2 baseline
- domain assumption The robot-mixture pretraining is not task-specific tuning
- domain assumption Real-world tasks with 10 rollouts per task provide stable success-rate estimates
Cite this review
Pith. "Pith review of Object-Centric Representations Improve Policy Generalization in Robot Manipulation." pith.science (2026). https://pith.science/paper/JOBZIBO3
@misc{pith2026250511563,
author = {Pith},
title = {Pith review of: Object-Centric Representations Improve Policy Generalization in Robot Manipulation},
year = {2026},
howpublished = {\url{https://pith.science/paper/JOBZIBO3}},
note = {Machine review of arXiv:2505.11563}
}
read the original abstract
Visual representations are central to the learning and generalization capabilities of robotic manipulation policies. While existing methods rely on global or dense features, such representations often entangle task-relevant and irrelevant scene information, limiting robustness under distribution shifts. In this work, we investigate object-centric representations (OCR) as a structured alternative that segments visual input into a finished set of entities, introducing inductive biases that align more naturally with manipulation tasks. We benchmark a range of visual encoders-object-centric, global and dense methods-across a suite of simulated and real-world manipulation tasks ranging from simple to complex, and evaluate their generalization under diverse visual conditions including changes in lighting, texture, and the presence of distractors. Our findings reveal that OCR-based policies outperform dense and global representations in generalization settings, even without task-specific pretraining. These insights suggest that OCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 3 Pith papers
-
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning
FORGE decouples robotic tool-use into keypoint trajectory prediction from action-free data and action grounding from limited demonstrations, achieving over 2X improvement in functional generalization to unseen tools.
-
STORM: Slot-based Task-aware Object-centric Representation for robotic Manipulation
STORM uses a two-stage, text-guided slot attention module on frozen DINOv2 features to improve robot manipulation success and generalization to visual distractors in simulated benchmarks.
-
Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges
A survey that taxonomizes robotic manipulation policies trained by imitation learning, traces their evolution, and compiles benchmark comparisons.
Reference graph
Works this paper leans on
-
[1]
Haldar, Z
S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learning,
-
[2]
A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Man- junath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...
arXiv 2023
-
[3]
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy, 2024. URL https: //arxiv.org/abs/2405.12213
arXiv 2024
-
[4]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model, 2024. URL https://arxiv.org/abs/2406.09246
arXiv 2024
-
[5]
S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3m: A universal visual represen- tation for robot manipulation, 2022. URL https://arxiv.org/abs/2203.12601
arXiv 2022
-
[6]
A. Majumdar, K. Yadav, S. Arnaud, Y . J. Ma, C. Chen, S. Silwal, A. Jain, V .-P. Berges, P. Abbeel, J. Malik, D. Batra, Y . Lin, O. Maksymets, A. Rajeswaran, and F. Meier. Where are we in the search for an artificial visual cortex for embodied intelligence?, 2024. URL https://arxiv.org/abs/2303.18240. 10
arXiv 2024
-
[7]
Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training, 2023. URL https: //arxiv.org/abs/2210.00030
arXiv 2023
- [8]
Show all 63 references
-
[9]
Radosavovic, T
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell. Real-world robot learn- ing with masked visual pre-training, 2022. URL https://arxiv.org/abs/2210.03109
2022 arXiv
-
[10]
Jiang, Y
G. Jiang, Y . Sun, T. Huang, H. Li, Y . Liang, and H. Xu. Robots pre-train robots: Manipulation- centric robotic representation from large-scale robot dataset.arXiv preprint arXiv:2410.22325, 2024
2024 arXiv
-
[11]
Parisi, A
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta. The unsurprising effectiveness of pre-trained vision models for control, 2022. URL https://arxiv.org/abs/2203.03580
2022 arXiv
-
[12]
Burns, Z
K. Burns, Z. Witzel, J. I. Hamid, T. Yu, C. Finn, and K. Hausman. What makes pre-trained visual representations successful for robust manipulation?, 2023. URLhttps://arxiv.org/ abs/2312.12444
2023 arXiv
-
[13]
B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman. Building machines that learn and think like people, 2016. URL https://arxiv.org/abs/1604.00289
2016 arXiv
-
[14]
Greff, S
K. Greff, S. van Steenkiste, and J. Schmidhuber. On the binding problem in artificial neural networks, 2020. URL https://arxiv.org/abs/2012.05208
2020 arXiv
-
[15]
Kroemer, S
O. Kroemer, S. Niekum, and G. Konidaris. A review of robot learning for manipulation: Challenges, representations, and algorithms, 2020. URL https://arxiv.org/abs/1907. 03146
2020
-
[16]
Bengio, A
Y . Bengio, A. Courville, and P. Vincent. Representation learning: A review and new perspec- tives, 2014. URL https://arxiv.org/abs/1206.5538
2014 arXiv
-
[17]
Locatello, D
F. Locatello, D. Weissenborn, T. Unterthiner, A. Mahendran, G. Heigold, J. Uszkoreit, A. Dosovitskiy, and T. Kipf. Object-centric learning with slot attention, 2020. URL https: //arxiv.org/abs/2006.15055
2020 arXiv
-
[18]
Seitzer, M
M. Seitzer, M. Horn, A. Zadaianchuk, D. Zietlow, T. Xiao, C.-J. Simon-Gabriel, T. He, Z. Zhang, B. Sch ¨olkopf, T. Brox, and F. Locatello. Bridging the gap to real-world object- centric learning, 2023. URL https://arxiv.org/abs/2209.14860
2023 arXiv
-
[19]
Yoon, Y .-F
J. Yoon, Y .-F. Wu, H. Bae, and S. Ahn. An investigation into pre-training object-centric repre- sentations for reinforcement learning, 2023. URL https://arxiv.org/abs/2302.04419
2023 arXiv
-
[20]
Heravi, A
N. Heravi, A. Wahid, C. Lynch, P. Florence, T. Armstrong, J. Tompson, P. Sermanet, J. Bohg, and D. Dwibedi. Visuomotor control in multi-object scenes using object-aware representations,
-
[21]
Haramati, T
D. Haramati, T. Daniel, and A. Tamar. Entity-centric reinforcement learning for object manip- ulation from pixels, 2024. URL https://arxiv.org/abs/2404.01220
2024 arXiv
-
[22]
Watters, L
N. Watters, L. Matthey, M. Bosnjak, C. P. Burgess, and A. Lerchner. Cobra: Data-efficient model-based rl through unsupervised object discovery and curiosity-driven exploration, 2019. URL https://arxiv.org/abs/1905.09275
2019 arXiv
-
[23]
T. Kipf, G. F. Elsayed, A. Mahendran, A. Stone, S. Sabour, G. Heigold, R. Jonschkowski, A. Dosovitskiy, and K. Greff. Conditional object-centric learning from video, 2022. URL https://arxiv.org/abs/2111.12594. 11
2022 arXiv
-
[24]
Zhang, A
C. Zhang, A. Gupta, and A. Zisserman. Is an object-centric video representation beneficial for transfer?, 2022. URL https://arxiv.org/abs/2207.10075
2022 arXiv
-
[25]
X. Chen, H. Fan, R. Girshick, and K. He. Improved baselines with momentum contrastive learning, 2020. URL https://arxiv.org/abs/2003.04297
2020 arXiv
-
[26]
Caron, H
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers, 2021. URL https://arxiv.org/abs/ 2104.14294
2021 arXiv
-
[27]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A...
2024 arXiv
-
[28]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020
2021 arXiv
-
[29]
Grauman, A
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Radosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, D. Batra, V . Cartillier, S. Crane, T. D...
2022 arXiv
-
[30]
Dasari, M
S. Dasari, M. K. Srirama, U. Jain, and A. Gupta. An unbiased look at datasets for visuo-motor pre-training, 2023. URL https://arxiv.org/abs/2310.09289
2023 arXiv
-
[31]
Hamdan and F
S. Hamdan and F. G ¨uney. Carformer: Self-driving with learned object-centric representations,
-
[32]
Mosbach, J
M. Mosbach, J. N. Ewertz, A. Villar-Corrales, and S. Behnke. Sold: Slot object-centric latent dynamics models for relational manipulation learning from pixels, 2025. URL https:// arxiv.org/abs/2410.08822
2025 arXiv
-
[33]
B. Wang, L. Li, J. Zhang, Y . Nakashima, and H. Nagahara. Explainable image recognition via enhanced slot-attention based classifier, 2024. URL https://arxiv.org/abs/2407. 05616
2024
-
[34]
URL https://arxiv.org/abs/2407.15843
-
[35]
C. P. Burgess, L. Matthey, N. Watters, R. Kabra, I. Higgins, M. Botvinick, and A. Lerchner. Monet: Unsupervised scene decomposition and representation, 2019. URL https://arxiv. org/abs/1901.11390
2019 arXiv
-
[36]
Jiang, F
J. Jiang, F. Deng, G. Singh, and S. Ahn. Object-centric slot diffusion, 2023. URL https: //arxiv.org/abs/2303.10834. 12
2023 arXiv
-
[37]
Kabra, D
R. Kabra, D. Zoran, G. Erdogan, L. Matthey, A. Creswell, M. Botvinick, A. Lerchner, and C. P. Burgess. Simone: View-invariant, temporally-abstracted object representations via unsu- pervised video decomposition, 2021. URL https://arxiv.org/abs/2106.03849
2021 arXiv
-
[38]
Singh, F
G. Singh, F. Deng, and S. Ahn. Illiterate dall-e learns to compose, 2022. URL https:// arxiv.org/abs/2110.11405
2022 arXiv
-
[39]
G. F. Elsayed, A. Mahendran, S. van Steenkiste, K. Greff, M. C. Mozer, and T. Kipf. Savi++: Towards end-to-end object-centric learning from real-world videos, 2022. URL https:// arxiv.org/abs/2206.07764
2022 arXiv
-
[40]
Z. Wu, J. Hu, W. Lu, I. Gilitschenski, and A. Garg. Slotdiffusion: Object-centric generative modeling with diffusion models, 2023. URL https://arxiv.org/abs/2305.11281
2023 arXiv
-
[41]
Zadaianchuk, M
A. Zadaianchuk, M. Seitzer, and G. Martius. Object-centric learning for real-world videos by predicting temporal feature similarities, 2023. URLhttps://arxiv.org/abs/2306.04829
2023 arXiv
-
[42]
Didolkar, A
A. Didolkar, A. Zadaianchuk, A. Goyal, M. Mozer, Y . Bengio, G. Martius, and M. Seitzer. Zero-shot object-centric representation learning, 2024. URL https://arxiv.org/abs/ 2408.09162
2024 arXiv
-
[43]
Singh, Y .-F
G. Singh, Y .-F. Wu, and S. Ahn. Simple unsupervised object-centric learning for complex and naturalistic videos, 2022. URL https://arxiv.org/abs/2205.14065
2022 arXiv
-
[44]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023. URL https://arxiv.org/abs/2304.13705
2023 arXiv
-
[45]
Warner, A
B. Warner, A. Chaffin, B. Clavi ´e, O. Weller, O. Hallstr ¨om, S. Taghadouini, A. Gallagher, R. Biswas, F. Ladhak, T. Aarsen, N. Cooper, G. Adams, J. Howard, and I. Poli. Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long conte...
2024 arXiv
-
[46]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polo- sukhin. Attention is all you need, 2023. URL https://arxiv.org/abs/1706.03762
2023 arXiv
-
[47]
B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowl- edge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310, 2023
2023 arXiv
-
[48]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition, 2015. URL https://arxiv.org/abs/1512.03385
2015 arXiv
-
[49]
T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine. Meta-world: A benchmark and evaluation for multi-task and meta reinforce- ment learning, 2021. URL https://arxiv.org/abs/1910.10897
2021 arXiv
-
[50]
Walke, K
H. Walke, K. Black, A. Lee, M. J. Kim, M. Du, C. Zheng, T. Zhao, P. Hansen-Estruch, Q. Vuong, A. He, V . Myers, K. Fang, C. Finn, and S. Levine. Bridgedata v2: A dataset for robot learning at scale, 2024. URL https://arxiv.org/abs/2308.12952
2024 arXiv
-
[51]
Khazatsky, K
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...
2024 arXiv
-
[52]
Dosovitskiy, L
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. URL https: //arxiv.org/abs/2010.11929
2021 arXiv
-
[53]
T.-Y . Lin, M. Maire, S. Belongie, L. Bourdev, R. Girshick, J. Hays, P. Perona, D. Ramanan, C. L. Zitnick, and P. Doll´ar. Microsoft coco: Common objects in context, 2015. URL https: //arxiv.org/abs/1405.0312
2015 arXiv
-
[54]
N. Xu, L. Yang, Y . Fan, D. Yue, Y . Liang, J. Yang, and T. Huang. Youtube-vos: A large-scale video object segmentation benchmark, 2018. URL https://arxiv.org/abs/1809.03327
2018 arXiv
-
[55]
Russakovsky, J
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei. Imagenet large scale visual recogni- tion challenge, 2015. URL https://arxiv.org/abs/1409.0575
2015 arXiv
-
[56]
Todorov, T
E. Todorov, T. Erez, and Y . Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages 5026–
2012
-
[57]
A. Xie, L. Lee, T. Xiao, and C. Finn. Decomposing the generalization gap in imitation learning for visual robotic manipulation, 2023. URL https://arxiv.org/abs/2307.03659
2023 arXiv
-
[58]
Cadene, S
R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf. Lerobot: State- of-the-art machine learning for real-world robotics in pytorch. https://github.com/ huggingface/lerobot, 2024
2024
-
[62]
Mandlekar, D
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. In Conference on Robot Learning (CoRL) , 2021. A Implementation detail...
2021
-
[63]
Figure 7: Slots visualization
Note that the model has never seen the provided data before as it has been pre-trained on frozen to learn subsequent policy. Figure 7: Slots visualization. Visualization of a set of slots (12 slots) extracted from VIDEOSAUR* model on the easy distractor setup in Metaworld 19 H...
-
[2023]
URL https://arxiv.org/abs/2205.06333
-
[2024]
URL https://arxiv.org/abs/2406.07539
-
[5033]
doi:10.1109/IROS.2012.6386109
IEEE, 2012. doi:10.1109/IROS.2012.6386109
2012
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.