REVIEW 5 major objections 6 minor 173 references
Advances in Transformers for Robotic Applications: A Review
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This review argues that Transformers have become a mainstream architecture across robotic perception, planning, and control, advancing human-robot interaction, long-horizon planning, and system performance.
desk verdict Broad, current-through-2024 survey of transformers in robotics; the organizing claim is fine, but citation errors and broken references make it unsafe as a reference and not yet worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the self-attention mechanism introduced by the original Transformer, which lets every element of an input sequence attend to every other element in parallel, independent of sequence length. In this review it is the shared engine behind all three adoption routes: attention over language tokens gives robots instruction-following, attention over visual tokens gives perception and grounding, and attention over sequences of states, actions, and rewards gives reinforcement learning and planning. The review's organizing distinction among foundation models, deep-reinforcement-learning hybrids, and perception-planning-control systems is the map that shows this one mechanism spreading across robotics.
What would settle it
Run a systematic search of robotics papers from 2022 through 2024 using the review's topical keywords and check whether transformer-based methods appear in a substantial share of perception, planning, and control systems; then attempt to reproduce the headline numbers on the original benchmarks, including TransformerMPC's 6.8x to 34.9x speedups, MarineFormer's 20% success-rate gain, and the 92.73% fault-instruction detection rate. If transformer use is concentrated in a narrow slice of robotics or those numbers do not reproduce, the review's characterization of adoption is unsubstantiated.
Extended reading notes
Core claim
The paper's central claim is that in robotics, Transformers are being adopted in three major ways: as pretrained foundation models facilitating human-robot interaction and generalization, as transformer variants integrated with deep reinforcement learning for long-horizon planning, and as components that enhance perception, planning, and control systems. It presents evidence from generalist robot policies such as RT-1, RT-2, OpenVLA, Octo, and π0, from sequence-modeling RL methods such as the Decision Transformer and Trajectory Transformer, and from perception and control systems such as CLIP, SAM, ViNT, and TransformerMPC. The paper reports representative results, including a 92.73% detection rate for faulty instructions in construction human-robot interaction, a 20% success-rate improvement for MarineFormer, and 6.8x to 34.9x speedups from TransformerMPC, as evidence of this trend.
Load-bearing premise
The review runs no experiments and uses no systematic search criteria, so its central picture rests on the assumption that the quantitative results reported in the cited primary papers are accurate and that the informally chosen set of papers fairly represents the field.
Editorial extensions
If this is right
- If the central claim is correct, transformer-based components will continue displacing CNNs, RNNs, and classical planners as the default building blocks for robotic systems.
- Shared multi-embodiment datasets such as Open X-Embodiment become a critical resource, so data collection, standardization, and embodiment balance are where the field's bottlenecks now sit.
- Sequence-modeling formulations of reinforcement learning make long-horizon tasks more tractable, allowing robot learning systems to be pretrained offline and then fine-tuned online.
- Transformer-based control accelerators such as TransformerMPC can cut runtime by an order of magnitude while preserving constraint satisfaction, making real-time deployment on robots more practical.
- Zero-shot and few-shot generalization is treated as achievable for manipulation in controlled settings, while the review also warns that it is not yet guaranteed in unstructured, outdoor, or safety-critical environments.
Reading between the lines
- Going beyond the paper: if this trend continues, robotics may follow the same trajectory as natural language processing, converging on a few large generalist policies pretrained on pooled data and then fine-tuned per task.
- The headline numbers in the review come from heterogeneous tasks, metrics, and baselines, so direct comparison across the three routes is not possible without a standardized evaluation protocol.
- Because the review selected papers informally, the most direct test of its broad-adoption thesis is a systematic literature search; a search that found transformer use concentrated in manipulation and largely absent from field robotics would require qualifying that thesis.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a narrative review of transformer architectures for robotic applications. After a background section on the original transformer and its efficient, multimodal, and sparse variants, it organizes robotic applications into three threads: pre-trained foundation models (including vision-language-action models and human-robot interaction), transformer-based deep reinforcement learning variants, and transformer use in perception, planning, and control. It closes with challenges and future directions. The paper presents no new experiments or derivations; its central descriptive claim is that transformers are now broadly adopted across robotic perception, planning, and control, which is consistent with the current literature.
Significance. If the survey's attributions were reliable, the paper would serve as a useful broad overview of a fast-moving area, with a sensible tripartite organization and up-to-date coverage (through late 2024) of models such as RT-2, OpenVLA, Octo, π0, and TransformerMPC. The authors explicitly acknowledge limitations in Section 4 (dataset imbalance, sim-to-real gap, safety, memory footprint) and propose concrete future directions such as quantization for resource-constrained deployment and cross-embodiment training. The main weakness is that several specific quantitative and architectural claims are mis-cited or unverifiable from the reference list; because a review's value depends on readers being able to trace claims to sources, these issues are load-bearing and must be corrected before the paper can be relied upon as a survey.
major comments (5)
- [§3.2, references [101]/[102]] The same work (Parisotto and Salakhutdinov, 'Efficient transformers in reinforcement learning using actor-learner distillation') appears as both [101] and [102]. Reference [101] is then used for two incompatible claims: that transformers are effective as world models, and that 'MarineFormer [101] ... demonstrated 20% success rate improvement' for marine environments. The cited paper is neither primarily a world-model paper nor MarineFormer; MarineFormer appears to be a different work that is missing from the bibliography. These mis-citations make specific quantitative claims untraceable and must be fixed, either by correcting the citations or by inserting the correct references.
- [§2.3.1, Table 1 (Efficient Transformers)] The table row 'Synthesizer [138]' cites reference [138], which is Tay et al., 'Sparse Sinkhorn Attention.' The Synthesizer architecture is a separate paper (Tay et al., 'Synthesizer: Rethinking Self-Attention for Transformer Models') and is not in the bibliography, while the Sinkhorn Transformer is correctly associated with [138] in Table 2 of Section 2.3.3. The two attention mechanisms are conflated; please replace the citation and/or add the Synthesizer reference.
- [§2.3.3, Table 2 (Sparse and Adaptive Transformers)] The table lists 'Routing Transformer [121]' twice with different asymptotic complexities: O(n log n) and O(n^{1.5} d^*). This is internally inconsistent, and neither entry is adequately justified. Please consolidate the duplicate entries into a single row and state the complexity reported in the original Routing Transformer paper (Roy et al., 2020), or explain why two different forms are listed.
- [§3.3.2, Planning] The sentence 'Gated Transformer-XL [23] for Long-Term Memory' cites reference [23], which is Transformer-XL by Dai et al.; the Gated Transformer-XL model is reference [103] (Parisotto et al., 'Stabilizing Transformers for Reinforcement Learning'), already correctly cited in Section 3.2. In the same subsection, the ViNT model is discussed ('ViNT, a Transformer-based model for visual navigation, demonstrates significant potential...') without any citation at all. Please correct the attribution and add a citation for ViNT.
- [§3.1, Human-Robot Collaboration] The paragraph on human-robot collaboration reports specific success rates: '92.73% detection of fault instructions by humans' [104] and '99.95% success rate in simulation' [105]. From the reference list alone, I cannot confirm that these exact numbers appear in the cited papers, since [104] and [105] are construction-HRI papers but no page numbers, sections, tables, or experiment descriptions are given. Because these figures are the only quantitative evidence in that paragraph, please verify the attributions or qualify the claims as reported in the cited sources.
minor comments (6)
- [§2.3.2 and §3.1] Both sections contain unresolved 'see Figure ??' cross-references (after the ViLBERT description and after the OpenVLA description, respectively); these must be fixed before publication.
- [§2.3.2, Table 3] The table entry 'VisualGPT [74]' is mis-cited: reference [74] is the VisualBERT paper, which is correctly named in the bullet list immediately below the table. Please correct the table entry or add the actual VisualGPT reference.
- [§4] There is a typo in the last paragraph: 'zero-short generalization' should read 'zero-shot generalization.'
- [§3.2] The phrase 'having low inference times while maintaining low sample effeciency of transformers' appears to state the opposite of the intended meaning; the actor-learner distillation approach is designed to retain the sample efficiency of transformers, so please rephrase to 'high sample efficiency' or similar.
- [Figure 2 caption] The caption is grammatically tangled and the subfigure references are unclear: 'while Franka has the most number of scenes see 2a x-Arm and Google Robot have the biggest contribution to trajectory data see 2b and 2c' does not map cleanly to panels (a)–(e). Please rewrite the caption to describe each panel explicitly.
- [References] Reference [93] is an incomplete URL-like entry containing a sentence fragment rather than a full citation, and reference [59] has a garbled author list and ordering. Both need to be reformatted according to the journal style.
Circularity Check
No circularity found: the review makes no predictions, fits no parameters, derives no equations from its own claims, and its load-bearing evidence is external cited work.
full rationale
This paper is a narrative literature review. It contains no derivation chain, no fitted parameters, no empirical predictions, and no mathematical claims that reduce to its own inputs. The central claim — that Transformers are being adopted in robotics through foundation models, DRL variants, and perception/planning/control — is supported by numerous citations to external primary papers and to other surveys. None of those citations are self-citations by the present authors, and none are invoked as a uniqueness theorem or as a forced choice that would make the review's taxonomy self-justifying. The survey's limitations (dataset imbalance, sim-to-real gaps, safety) are stated as open challenges, not as derived results. Weaknesses such as possible misattribution of specific numbers concern fidelity to external sources, which is a correctness/accuracy issue, not circularity. A result that depends on the accuracy of other people's experiments is still an external, non-circular dependence. Therefore, no circular step can be exhibited, and the appropriate score is 0.
Assumptions & free parameters
assumptions (1)
- domain assumption Quantitative results reported in the cited primary papers are accurate and reproducible as stated.
Cite this review
Pith. "Pith review of Advances in Transformers for Robotic Applications: A Review." pith.science (2026). https://pith.science/paper/TAL5TDBK
@misc{pith2026241210599,
author = {Pith},
title = {Pith review of: Advances in Transformers for Robotic Applications: A Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/TAL5TDBK}},
note = {Machine review of arXiv:2412.10599}
}
read the original abstract
The introduction of Transformers architecture has brought about significant breakthroughs in Deep Learning (DL), particularly within Natural Language Processing (NLP). Since their inception, Transformers have outperformed many traditional neural network architectures due to their "self-attention" mechanism and their scalability across various applications. In this paper, we cover the use of Transformers in Robotics. We go through recent advances and trends in Transformer architectures and examine their integration into robotic perception, planning, and control for autonomous systems. Furthermore, we review past work and recent research on use of Transformers in Robotics as pre-trained foundation models and integration of Transformers with Deep Reinforcement Learning (DRL) for autonomous systems. We discuss how different Transformer variants are being adapted in robotics for reliable planning and perception, increasing human-robot interaction, long-horizon decision-making, and generalization. Finally, we address limitations and challenges, offering insight and suggestions for future research directions.
Figures
Reference graph
Works this paper leans on
-
[101]
Efficient transformers in reinforcement learn- ing using actor-learner distillation, 2021
Emilio Parisotto and Ruslan Salakhutdinov. Efficient transformers in reinforcement learn- ing using actor-learner distillation, 2021
work page 2021
-
[138]
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Sparse sinkhorn attention. arXiv preprint arXiv:2002.11296 , 2020
arXiv 2002
-
[102]
Efficient transformers in reinforcement learn- ing using actor-learner distillation
Emilio Parisotto and Ruslan Salakhutdinov. Efficient transformers in reinforcement learn- ing using actor-learner distillation. ArXiv, abs/2104.01655, 2021
arXiv 2021
-
[121]
Efficient content- based sparse attention with routing transformers, 2020
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. Efficient content- based sparse attention with routing transformers, 2020
work page 2020
-
[23]
Le, and Ruslan 17 Salakhutdinov
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan 17 Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length con- text, 2019
2019
-
[103]
Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Caglar Gulcehre, Sid- dhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matthew M. Botvinick, Nicolas Heess, and Raia Hadsell. Stabilizing transformers for reinforcement learning, 2019
work page 2019
-
[104]
Somin Park, Carol C. Menassa, and Vineet R. Kamat. Integrating large language models with multimodal virtual reality interfaces to support collaborative human-robot construc- tion work, 2024
work page 2024
-
[105]
Somin Park, Xi Wang, Carol C. Menassa, Vineet R. Kamat, and Joyce Y. Chai. Natu- ral language instructions for intuitive human interaction with robotic assistants in field construction work. Automation in Construction , 161:105345, May 2024
work page 2024
Show all 173 references
-
[1]
Generating out-of-distribution scenarios using language models, 2024
Erfan Aasi, Phat Nguyen, Shiva Sreeram, Guy Rosman, Sertac Karaman, and Daniela Rus. Generating out-of-distribution scenarios using language models, 2024
2024
-
[2]
Constrained policy opti- mization, 2017
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. Constrained policy opti- mization, 2017
2017
-
[3]
Pranav Agarwal, Aamer Abdul Rahman, Pierre-Luc St-Charles, Simon J. D. Prince, and Samira Ebrahimi Kahou. Transformers in reinforcement learning: A survey, 2023
2023
-
[4]
Lstm inefficiency in long-term dependencies regression problems
Safwan Mahmood Al-Selwi, Mohd Fadzil, Said Jadid Abdulkadir, and Amgad Muneer. Lstm inefficiency in long-term dependencies regression problems. Journal of Advanced Research in Applied Sciences and Engineering Technology , 2023
2023
-
[5]
Advances in medical im- age analysis with vision transformers: A comprehensive review
Reza Azad, Amirhossein Kazerouni, Moein Heidari, Ehsan Khodapanah Aghdam, Amir Molaei, Yiwei Jia, Abin Jose, Rijo Roy, and Dorit Merhof. Advances in medical im- age analysis with vision transformers: A comprehensive review. Medical image analysis , 91:103000, 2023
2023
-
[6]
The option-critic architecture, 2016
Pierre-Luc Bacon, Jean Harb, and Doina Precup. The option-critic architecture, 2016
2016
-
[7]
data2vec: A general framework for self-supervised learning in speech, vision and language, 2022
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A general framework for self-supervised learning in speech, vision and language, 2022
2022
-
[8]
Peters, and Arman Cohan
Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer, 2020
2020
-
[9]
π0: A vision-language-action flow model for general robot control
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, Jam...
-
[10]
Josh C. Bongard. Probabilistic robotics. sebastian thrun, wolfram burgard, and dieter fox. (2005, mit press.) 647 pages. Artificial Life, 14:227–229, 2008
2005
-
[11]
Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasm...
2023
-
[12]
Rt-1: Robotics transformer for real-world control at scale, 2023
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov,...
2023
-
[13]
Lewis, and Satinder Singh
Ethan Brooks, Logan Walls, Richard L. Lewis, and Satinder Singh. Large language models can implement policy iteration, 2023
2023
-
[14]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[15]
Plancritic: Formal planning with human feedback, 2024
Owen Burns, Dana Hughes, and Katia Sycara. Plancritic: Formal planning with human feedback, 2024
2024
-
[16]
Radfar, Athanasios Mouchtaris, Brian King, and Siegfried Kun- zmann
Feng-Ju Chang, Martin H. Radfar, Athanasios Mouchtaris, Brian King, and Siegfried Kun- zmann. End-to-end multi-channel transformer for speech recognition.ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 5884–5888, 2021
2021
-
[17]
Decision transformer: Reinforcement learning via sequence modeling, 2021
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling, 2021
2021
-
[18]
Clip2scene: Towards label-efficient 3d scene understanding by clip, 2023
Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li, Yue- nan Hou, Yu Qiao, and Wenping Wang. Clip2scene: Towards label-efficient 3d scene understanding by clip, 2023
2023
-
[19]
Generating long sequences with sparse transformers, 2019
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers, 2019
2019
-
[20]
Rethinking attention with performers, 2022
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking attention with performers, 2022
2022
-
[21]
Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Ir- pan, Alexander Khazatsky, Anant Rai,...
2024
-
[22]
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
2023
-
[24]
Fu, Stefano Ermon, Atri Rudra, and Christopher R´ e
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R´ e. Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
2022
-
[25]
Universal transformers, 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers, 2019
2019
-
[26]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics , 2019
2019
-
[27]
DiPietro and Gregory Hager
Robert S. DiPietro and Gregory Hager. Deep learning: Rnns and lstm. 2020
2020
-
[28]
Scaling cross- embodied learning: One policy for manipulation, navigation, locomotion and aviation, 2024
Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. Scaling cross- embodied learning: One policy for manipulation, navigation, locomotion and aviation, 2024
2024
-
[29]
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...
2021
-
[30]
Addressing some limitations of transformers with feedback memory, 2021
Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. Addressing some limitations of transformers with feedback memory, 2021
2021
-
[31]
Model-agnostic meta-learning for fast adaptation of deep networks, 2017
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks, 2017
2017
-
[32]
Foundation models in robotics: Ap- plications, challenges, and the future
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, Brian Ichter, Danny Driess, Jiajun Wu, Cewu Lu, and Mac Schwager. Foundation models in robotics: Ap- plications, challenges, and the fu...
-
[33]
D. Fox, W. Burgard, and S. Thrun. Markov localization for mobile robots in dynamic environments. Journal of Artificial Intelligence Research , 11:391–427, November 1999
1999
-
[34]
Generalized decision transformer for offline hindsight information matching, 2022
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu. Generalized decision transformer for offline hindsight information matching, 2022
2022
-
[35]
Physically grounded vision-language models for robotic ma- nipulation, 2024
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Ma- jumdar, and Dorsa Sadigh. Physically grounded vision-language models for robotic ma- nipulation, 2024
2024
-
[36]
Learning in reality: A case study of stanley, the robot that won the darpa challenge
Christian Glaser and Philipp Hennig. Learning in reality: A case study of stanley, the robot that won the darpa challenge. 2012
2012
-
[37]
Rvt: Robotic view transformer for 3d object manipulation
Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox. Rvt: Robotic view transformer for 3d object manipulation. ArXiv, abs/2306.14896, 2023
2023 arXiv
-
[38]
Metamorph: Learning universal controllers with transformers, 2022
Agrim Gupta, Linxi Fan, Surya Ganguli, and Li Fei-Fei. Metamorph: Learning universal controllers with transformers, 2022
2022
-
[39]
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018. 18
2018
-
[40]
Transformers in medical image analysis: A review
Kelei He, Chen Gan, Zhuoyuan Li, Islem Rekik, Zihao Yin, Wen Ji, Yang Gao, Qian Wang, Junfeng Zhang, and Dinggang Shen. Transformers in medical image analysis: A review. ArXiv, abs/2202.12165, 2022
2022 arXiv
-
[41]
Henry, Onyeka Emebob, and Conrad Asotie Omonhinmin
Emerald U. Henry, Onyeka Emebob, and Conrad Asotie Omonhinmin. Vision transform- ers in medical imaging: A review. ArXiv, abs/2211.10043, 2022
2022 arXiv
-
[42]
Convolutional vision transformer as a path following controller for omnidirectional robots
Sandesh Hiremath, Cheng-Yi Huang, Argtim Tika, and Naim Bajcinca. Convolutional vision transformer as a path following controller for omnidirectional robots. In 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages 16633–16639, 2024
2024
-
[43]
Long short-term memory.Neural Computation, 9:1735–1780, 1997
Sepp Hochreiter and J¨ urgen Schmidhuber. Long short-term memory.Neural Computation, 9:1735–1780, 1997
1997
-
[44]
On transforming re- inforcement learning with transformers: The development trajectory
Shengchao Hu, Li Shen, Ya Zhang, Yixin Chen, and Dacheng Tao. On transforming re- inforcement learning with transformers: The development trajectory. IEEE Transactions on Pattern Analysis and Machine Intelligence , 46(12):8580–8599, 2024
2024
-
[45]
Toward general-purpose robots via foundation models: A survey and meta-analysis, 2024
Yafei Hu, Quanting Xie, Vidhi Jain, Jonathan Francis, Jay Patrikar, Nikhil Keetha, Se- ungchan Kim, Yaqi Xie, Tianyi Zhang, Hao-Shu Fang, Shibo Zhao, Shayegan Omidshafiei, Dong-Ki Kim, Ali akbar Agha-mohammadi, Katia Sycara, Matthew Johnson-Roberson, Dhruv Batra, Xiaolong Wang...
2024
-
[46]
Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022
2022
-
[47]
Hpcneu- ronet: Advancing neuromorphic audio signal processing with transformer-enhanced spik- ing neural networks
Murat Isik, Hiruna Vishwamith, Kayode Inadagbo, and Ismail Can Dikmen. Hpcneu- ronet: Advancing neuromorphic audio signal processing with transformer-enhanced spik- ing neural networks. 2024 4th Interdisciplinary Conference on Electrics and Computer (INTCEC), pages 1–7, 2023
2024
-
[48]
A comprehensive survey on applications of transformers for deep learning tasks
Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Najat Drawel, Gaith Rjoub, and Witold Pedrycz. A comprehensive survey on applications of transformers for deep learning tasks. ArXiv, abs/2306.07303, 2023
2023 arXiv
-
[49]
Perceiver: General perception with iterative attention, 2021
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention, 2021
2021
-
[50]
Bc-z: Zero-shot task generalization with robotic imita- tion learning, 2022
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. Bc-z: Zero-shot task generalization with robotic imita- tion learning, 2022
2022
-
[51]
When to trust your model: Model-based policy optimization, 2021
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: Model-based policy optimization, 2021
2021
-
[52]
Offline reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. In Advances in Neural Information Processing Systems, 2021
2021
-
[53]
Flair: Feeding via long-horizon acquisition of realistic dishes, 2024
Rajat Kumar Jenamani, Priya Sundaresan, Maram Sakr, Tapomayukh Bhattacharjee, and Dorsa Sadigh. Flair: Feeding via long-horizon acquisition of realistic dishes, 2024
2024
-
[54]
A survey of robot intelligence with large language models
Hyeongyo Jeong, Haechan Lee, Changwon Kim, and Sungtae Shin. A survey of robot intelligence with large language models. Applied Sciences, 14(19), 2024. 19
2024
-
[55]
Le, Yunhsuan Sung, Zhen Li, and Tom Duerig
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision, 2021
2021
-
[56]
Highly accurate protein structure prediction with alphafold
John Jumper, Richard Evans, Alexander Pritzel, and et al. Highly accurate protein structure prediction with alphafold. Nature, 596:583–589, 2021
2021
-
[57]
Trans- formers are rnns: Fast autoregressive transformers with linear attention, 2020
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran¸ cois Fleuret. Trans- formers are rnns: Fast autoregressive transformers with linear attention, 2020
2020
-
[58]
Segment anything in high quality, 2023
Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu. Segment anything in high quality, 2023
2023
-
[59]
Real-world robot applications of foundation models: a review
Andrew Gambardella Jiaxian Guo Chris Paxton Kento Kawaharazuka, Tatsuya Mat- sushima and Andy Zeng. Real-world robot applications of foundation models: a review. Advanced Robotics, 38(18):1232–1254, 2024
2024
-
[60]
Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Sid- dharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Ja- son Ma, Patrick Tree ...
2024
-
[61]
What and when to explain? on-road evaluation of explanations in highly automated vehicles
Gwangbin Kim, Dohyeon Yeo, Taewoo Jo, Daniela Rus, and SeungJun Kim. What and when to explain? on-road evaluation of explanations in highly automated vehicles. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. , 7(3), September 2023
2023
-
[62]
Openvla: An open-source vision-language-action model, 2024
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla...
2024
-
[63]
A survey on integration of large language models with intelligent robots
Yeseung Kim, Dohyun Kim, Jieun Choi, Jisang Park, Nayoung Oh, and Daehyung Park. A survey on integration of large language models with intelligent robots. Intelligent Service Robotics, 17(5):1091–1107, August 2024
2024
-
[64]
Berg, Wan-Yen Lo, Piotr Doll´ ar, and Ross Girshick
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll´ ar, and Ross Girshick. Segment anything, 2023
2023
-
[65]
Reformer: The efficient transformer, 2020
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer, 2020. 20
2020
-
[66]
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research , 32(11):1238–1274, 2013
2013
-
[67]
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021
Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021
2021
-
[68]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM , 60:84 – 90, 2012
2012
-
[69]
Conservative q-learning for offline reinforcement learning, 2020
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning, 2020
2020
-
[70]
Phong Le and Willem H. Zuidema. Quantifying the vanishing gradient and long distance dependency problem in recursive neural networks and recursive lstms. InRep4NLP@ACL, 2016
2016
-
[71]
Fnet: Mixing tokens with fourier transforms, 2022
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. Fnet: Mixing tokens with fourier transforms, 2022
2022
-
[72]
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1334–1342, 2016
2016
-
[73]
Landman, and S
Jun Li, Junyu Chen, Yucheng Tang, Ce Wang, Bennett A. Landman, and S. Kevin Zhou. Transforming medical imaging with transformers? a comparative review of key properties, current progresses, and future perspectives. Medical image analysis , 85:102762, 2022
2022
-
[74]
Visualbert: A simple and performant baseline for vision and language, 2019
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language, 2019
2019
-
[75]
Grounded language-image pre-training, 2022
Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jian- feng Gao. Grounded language-image pre-training, 2022
2022
-
[76]
Shapegrasp: Zero-shot task-oriented grasping with large language models through geometric decomposition, 2024
Samuel Li, Sarthak Bhagat, Joseph Campbell, Yaqi Xie, Woojun Kim, Katia Sycara, and Simon Stepputtis. Shapegrasp: Zero-shot task-oriented grasping with large language models through geometric decomposition, 2024
2024
-
[77]
A survey on transformers in reinforcement learning, 2023
Wenzhe Li, Hao Luo, Zichuan Lin, Chongjie Zhang, Zongqing Lu, and Deheng Ye. A survey on transformers in reinforcement learning, 2023
2023
-
[78]
Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020
2020
-
[79]
Lillicrap, Jonathan J
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yu- val Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning, 2019
2019
-
[80]
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. A survey of transformers. AI Open, 3:111–132, 2022
2022
-
[81]
Instruction-following agents with multimodal transformer, 2023
Hao Liu, Lisa Lee, Kimin Lee, and Pieter Abbeel. Instruction-following agents with multimodal transformer, 2023. 21
2023
-
[82]
Constrained decision transformer for offline safe reinforcement learning
Zuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen, Wenhao Yu, Tingnan Zhang, and Ding Zhao. Constrained decision transformer for offline safe reinforcement learning. In An- dreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, edito...
2023
-
[83]
Few-shot subgoal plan- ning with language models, 2022
Lajanugen Logeswaran, Yao Fu, Moontae Lee, and Honglak Lee. Few-shot subgoal plan- ning with language models, 2022
2022
-
[84]
Uncertainty-aware hybrid paradigm of nonlinear mpc and model- based rl for offroad navigation: Exploration of transformers in the predictive model
Faraz Lotfi, Khalil Virji, Farnoosh Faraji, Lucas Berry, Andrew Holliday, David Meger, and Gregory Dudek. Uncertainty-aware hybrid paradigm of nonlinear mpc and model- based rl for offroad navigation: Exploration of transformers in the predictive model. In 2024 IEEE Internatio...
2024
-
[85]
Multi- agent actor-critic for mixed cooperative-competitive environments, 2020
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi- agent actor-critic for mixed cooperative-competitive environments, 2020
2020
-
[86]
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019
2019
-
[87]
Mega: Moving average equipped gated attention, 2023
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. Mega: Moving average equipped gated attention, 2023
2023
-
[88]
Transformers are sample efficient world models
Vincent Micheli, Eloi Alonso, and Franccois Fleuret. Transformers are sample efficient world models. ArXiv, abs/2209.00588, 2022
2022 arXiv
-
[89]
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015
2015
-
[90]
Integrating reinforcement learning with foundation models for autonomous robotics: Methods and perspectives, 2024
Angelo Moroncelli, Vishal Soni, Asad Ali Shahid, Marco Maccarini, Marco Forgione, Dario Piga, Blerina Spahiu, and Loris Roveda. Integrating reinforcement learning with foundation models for autonomous robotics: Methods and perspectives, 2024
2024
-
[91]
Jane Mulligan and Gregory Z. Grudic. Editorial for journal of field robotics—special issue on machine learning based robotics in unstructured environments. Journal of Field Robotics, 23, 2006
2006
-
[92]
Data-efficient hierarchical reinforcement learning, 2018
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine. Data-efficient hierarchical reinforcement learning, 2018
2018
-
[93]
B. A. Newman et al. Many household tasks can be posed as multi-object rearrangement tasks, but solutions to these problems often target a single, hand defined solution or are ... Carnegie Mellon University , page 4, 2024. Retrieved from https://www.andrew.cmu. edu/~bnewman1/data
2024
-
[94]
Newman, Pranay Gupta, Kris Kitani, Yonatan Bisk, Henny Admoni, and Chris Paxton
Benjamin A. Newman, Pranay Gupta, Kris Kitani, Yonatan Bisk, Henny Admoni, and Chris Paxton. Degustabot: Zero-shot visual preference estimation for personalized multi- object rearrangement, 2024
2024
-
[95]
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart J Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the 16th International Conference on Machine Learning (ICML) , pages 278–287, 1999. 22
1999
-
[96]
Nils J. Nilsson. Shakey the robot. 1984
1984
-
[97]
Deep explo- ration via bootstrapped dqn, 2016
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep explo- ration via bootstrapped dqn, 2016
2016
-
[98]
Voicepilot: Harnessing llms as speech interfaces for physically assistive robots
Akhil Padmanabha, Jessie Yuan, Janavi Gupta, Zulekha Karachiwalla, Carmel Majidi, Henny Admoni, and Zackory Erickson. Voicepilot: Harnessing llms as speech interfaces for physically assistive robots. In Proceedings of the 37th Annual ACM Symposium on User Interface Software an...
2024
-
[99]
Lifelong robot learning with human assisted language planners, 2023
Meenal Parakh, Alisha Fong, Anthony Simeonov, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Lifelong robot learning with human assisted language planners, 2023
2023
-
[100]
Lifelong robot learning with human assisted language planners
Meenal Parakh, Alisha Fong, Anthony Simeonov, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Lifelong robot learning with human assisted language planners. In 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages 523–529, 2024
2024
-
[106]
Dexhub and dart: Towards internet scale robot data collection, 2024
Younghyo Park, Jagdeep Singh Bhatia, Lars Ankile, and Pulkit Agrawal. Dexhub and dart: Towards internet scale robot data collection, 2024
2024
-
[107]
Efros, and Trevor Darrell
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction, 2017
2017
-
[108]
Transformers in the real world: A survey on nlp applications
Narendra Patwardhan, Stefano Marrone, and Carlo Sansone. Transformers in the real world: A survey on nlp applications. Information, 14(4), 2023
2023
-
[109]
Chun-Cheng Peng and G. D. Magoulas. Sequence processing with recurrent neural net- works. In Encyclopedia of Artificial Intelligence , 2009
2009
-
[110]
Miller, and Sebastian Riedel
Fabio Petroni, Tim Rockt¨ aschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexan- der H. Miller, and Sebastian Riedel. Language models as knowledge bases?, 2019
2019
-
[111]
Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher R´ e
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher R´ e. Hyena hierarchy: Towards larger convolutional language models, 2023. 23
2023
-
[112]
Learning transferable visual models from natural language supervision, 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021
2021
-
[113]
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan. Improving language understanding by generative pre-training. 2018
2018
-
[114]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1– 67, 2020
2020
-
[115]
Zero-shot text-to-image generation, 2021
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation, 2021
2021
-
[116]
Rlds: an ecosystem to generate, share and use datasets in reinforcement learning, 2021
Sabela Ramos, Sertan Girgin, L´ eonard Hussenot, Damien Vincent, Hanna Yakubovich, Daniel Toyama, Anita Gergely, Piotr Stanczyk, Raphael Marinier, Jeremiah Harmsen, Olivier Pietquin, and Nikola Momchev. Rlds: an ecosystem to generate, share and use datasets in reinforcement le...
2021
-
[117]
Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S
Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S. Khan. Glamm: Pixel grounding large multimodal model, 2024
2024
-
[118]
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018
2018
-
[119]
Sam 2: Segment anything in images and videos, 2024
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨ adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll´ ar, and Christoph Fei...
2024
-
[120]
Generalist agents
Scott Reed, Suraj Nair, Felix Hill, Oriol Vinyals, Nando de Freitas, et al. Generalist agents. arXiv preprint arXiv:2205.06175 , 2022
2022 arXiv
-
[122]
Rumelhart, Geoffrey E
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representa- tions by back-propagating errors. Nature, 323:533–536, 1986
1986
-
[123]
Jordan, and Pieter Abbeel
John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. Trust region policy optimization, 2017
2017
-
[124]
Proxi- mal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proxi- mal policy optimization algorithms, 2017
2017
-
[125]
Behavior transformers: Cloning k modes with one stone, 2022
Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, and Lerrel Pinto. Behavior transformers: Cloning k modes with one stone, 2022
2022
-
[126]
Fahad Shamshad, Salman Hameed Khan, Syed Waqas Zamir, Muhammad Haris Khan, Munawar Hayat, Fahad Shahbaz Khan, and H. Fu. Transformers in medical imaging: A survey. Medical image analysis , 88:102802, 2022. 24
2022
-
[127]
Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network
Alex Sherstinsky. Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network. ArXiv, abs/1808.03314, 2018
2018 arXiv
-
[128]
Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn
Lucy Xiaoyang Shi, Zheyuan Hu, Tony Z. Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn. Yell at your robot: Improving on-the-fly from language corrections, 2024
2024
-
[129]
Perspectives and prospects on trans- former architecture for cross-modal tasks with language and vision
Andrew Shin, Masato Ishii, and Takuya Narihira. Perspectives and prospects on trans- former architecture for cross-modal tasks with language and vision. International Journal of Computer Vision , 130:435 – 454, 2021
2021
-
[130]
Flava: A foundational language and vi- sion alignment model, 2022
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vi- sion alignment model, 2022
2022
-
[131]
Adaptive attention span in transformers, 2019
Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. Adaptive attention span in transformers, 2019
2019
-
[132]
Videobert: A joint model for video and language representation learning, 2019
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning, 2019
2019
-
[133]
Plate: Visually-grounded planning with transformers in procedural tasks
Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg. Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automation Letters, 7(2):4924–4930, April 2022
2022
-
[134]
Eva-clip: Improved training techniques for clip at scale, 2023
Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale, 2023
2023
-
[135]
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. Sequence to sequence learning with neural networks. ArXiv, abs/1409.3215, 2014
2014 arXiv
-
[136]
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018
2018
-
[137]
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. ACM Computing Surveys , 55:1 – 28, 2020
2020
-
[139]
Octo: An open-source generalist robot policy, 2024
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo...
2024
-
[140]
Robots that use language
Stefanie Tellex, Nakul Gopalan, Hadas Kress-Gazit, and Cynthia Matuszek. Robots that use language. Annu. Rev. Control. Robotics Auton. Syst. , 3:25–55, 2020
2020
-
[141]
Vision-and- dialog navigation, 2019
Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. Vision-and- dialog navigation, 2019
2019
-
[142]
Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia M
Sebastian Thrun, Michael Montemerlo, Hendrik Dahlkamp, David Stavens, Andrei Aron, James Diebel, Philip W. Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia M. Oakley, Mark Palatucci, Vaughan R. Pratt, Pascal Stang, Sven Strohband, Cedric Dupont, Lars-Erik J...
2006
-
[143]
Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation, 2024
Marcel Torne, Anthony Simeonov, Zechu Li, April Chan, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation, 2024
2024
-
[144]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023
2023
-
[145]
Audio transformers: Transformer architectures for large scale audio understanding
Prateek Verma and Jonathan Berger. Audio transformers: Transformer architectures for large scale audio understanding. adieu convolutions. ArXiv, abs/2105.00335, 2021
2021 arXiv
-
[146]
Apricot: Active preference learning and constraint-aware task planning with llms, 2024
Huaxiaoyue Wang, Nathaniel Chin, Gonzalo Gonzalez-Pumariega, Xiangwan Sun, Neha Sunkara, Maximus Adrian Pace, Jeannette Bohg, and Sanjiban Choudhury. Apricot: Active preference learning and constraint-aware task planning with llms, 2024
2024
-
[147]
Large language models for robotics: Opportunities, challenges, and perspectives, 2024
Jiaqi Wang, Zihao Wu, Yiwei Li, Hanqi Jiang, Peng Shu, Enze Shi, Huawen Hu, Chong Ma, Yiheng Liu, Xuhui Wang, Yincheng Yao, Xuan Liu, Huaqin Zhao, Zhengliang Liu, Haixing Dai, Lin Zhao, Bao Ge, Xiang Li, Tianming Liu, and Shu Zhang. Large language models for robotics: Opportun...
2024
-
[148]
Li, Madian Khabsa, Han Fang, and Hao Ma
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self- attention with linear complexity, 2020
2020
-
[149]
Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models, 2023
Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, and Daniela Rus. Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models, 2023
2023
-
[150]
A trajectory is worth three sentences: multimodal transformer for offline reinforcement learning
Yiqi Wang, Mengdi Xu, Laixi Shi, and Yuejie Chi. A trajectory is worth three sentences: multimodal transformer for offline reinforcement learning. In Robin J. Evans and Ilya Shpitser, editors, Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence...
2023
-
[151]
Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition, 2021
Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang. Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition, 2021
2021
-
[152]
Challengesand solutions for autonomous ground robot scene understanding and navigation in unstructured outdoor environments: A review
Liyana Wijayathunga, Alexander Rassau, and Douglas Chai. Challengesand solutions for autonomous ground robot scene understanding and navigation in unstructured outdoor environments: A review. Applied Sciences, 2023
2023
-
[153]
Greedy hierarchical variational autoencoders for large-scale video prediction, 2021
Bohan Wu, Suraj Nair, Roberto Martin-Martin, Li Fei-Fei, and Chelsea Finn. Greedy hierarchical variational autoencoders for large-scale video prediction, 2021
2021
-
[154]
Tidybot: personalized robot assistance with large language models
Jimmy Wu, Rika Antonova, Adam Kan, Marion Lepert, Andy Zeng, Shuran Song, Jean- nette Bohg, Szymon Rusinkiewicz, and Thomas Funkhouser. Tidybot: personalized robot assistance with large language models. Autonomous Robots, 47(8):1087–1102, November 2023
2023
-
[155]
Transformers in medical image segmentation: A review
Han Xiao, Li Li, Qi yu Liu, Xiuhong Zhu, and Qihang Zhang. Transformers in medical image segmentation: A review. Biomed. Signal Process. Control., 84:104791, 2023
2023
-
[156]
Nystr¨ omformer: A nystr¨ om-based algorithm for approximating self- attention, 2021
Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. Nystr¨ omformer: A nystr¨ om-based algorithm for approximating self- attention, 2021
2021
-
[157]
A joint modeling of vision-language-action for target-oriented grasping in clutter, 2024
Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yue Wang, and Rong Xiong. A joint modeling of vision-language-action for target-oriented grasping in clutter, 2024. 26
2024
-
[158]
Peng Xu, Xiatian Zhu, and David A. Clifton. Multimodal learning with transformers: A survey, 2023
2023
-
[159]
A review of recurrent neural networks: Lstm cells and network architectures
Yong Yu, Xiaosheng Si, Changhua Hu, and Jian xun Zhang. A review of recurrent neural networks: Lstm cells and network architectures. Neural Computation, 31:1235–1270, 2019
2019
-
[160]
Trans- former in reinforcement learning for decision-making: A survey
Weilin Yuan, Jiaxing Chen, Shaofei Chen, Dawei Feng, Zhenzhen Hu, and Peng Li. Trans- former in reinforcement learning for decision-making: A survey. October 2023
2023
-
[161]
Big bird: Transformers for longer sequences, 2021
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences, 2021
2021
-
[162]
Merlot: Multimodal neural script knowledge models, 2021
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models, 2021
2021
-
[163]
Fanlong Zeng, Wensheng Gan, Yongheng Wang, Ning Liu, and Philip S. Yu. Large language models for robotics: A survey, 2023
2023
-
[164]
Distilling and retrieving generalizable knowledge for robot manipulation via language corrections, 2024
Lihan Zha, Yuchen Cui, Li-Heng Lin, Minae Kwon, Montserrat Gonzalez Arenas, Andy Zeng, Fei Xia, and Dorsa Sadigh. Distilling and retrieving generalizable knowledge for robot manipulation via language corrections, 2024
2024
-
[165]
Hirt: Enhancing robotic control with hierarchical robot transformers, 2024
Jianke Zhang, Yanjiang Guo, Xiaoyu Chen, Yen-Jen Wang, Yucheng Hu, Chengming Shi, and Jianyu Chen. Hirt: Enhancing robotic control with hierarchical robot transformers, 2024
2024
-
[166]
Pointclip: Point cloud understanding by clip, 2021
Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip, 2021
2021
-
[167]
Decision transformer as a foundation model for partially observable continuous control, 2024
Xiangyuan Zhang, Weichao Mao, Haoran Qiu, and Tamer Ba¸ sar. Decision transformer as a foundation model for partially observable continuous control, 2024
2024
-
[168]
Encoder-decoder models in sequence-to-sequence learning: A survey of rnn and lstm approaches
Yunong Zhang. Encoder-decoder models in sequence-to-sequence learning: A survey of rnn and lstm approaches. Applied and Computational Engineering , 2023
2023
-
[169]
Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid
Tony Z. Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid. Aloha unleashed: A simple recipe for robot dexterity, 2024
2024
-
[170]
Online decision transformer
Qinqing Zheng, Amy Zhang, and Aditya Grover. Online decision transformer. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning , volume 162 of Proceedings o...
2022
-
[171]
Comparative study of sequence-to-sequence models: From rnns to trans- formers
Jiancong Zhu. Comparative study of sequence-to-sequence models: From rnns to trans- formers. Applied and Computational Engineering , 2024
2024
-
[172]
Transformermpc: Accelerating model predictive control via transformers, 2024
Vrushabh Zinage, Ahmed Khalil, and Efstathios Bakolas. Transformermpc: Accelerating model predictive control via transformers, 2024
2024
-
[173]
Atabay A. A. Ziyaden, Amir Yelenov, and Alexander Pak. Long-context transformers: A survey. 2021 5th Scientific School Dynamics of Complex Networks and their Applications (DCNA), pages 215–218, 2021. 27
2021
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.