Pith. sign in

REVIEW 5 major objections 6 minor 173 references

Advances in Transformers for Robotic Applications: A Review

T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This review argues that Transformers have become a mainstream architecture across robotic perception, planning, and control, advancing human-robot interaction, long-horizon planning, and system performance.

desk verdict Broad, current-through-2024 survey of transformers in robotics; the organizing claim is fine, but citation errors and broken references make it unsafe as a reference and not yet worth refereeing. read the letter →

arxiv 2412.10599 v1 pith:TAL5TDBK submitted 2024-12-13 cs.RO cs.AI

classification cs.ROcs.AI
keywords Transformersroboticsfoundationmodelsvision-language-actiondeepreinforcementlearningperceptionplanningcontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review aims to establish that Transformers are now a broadly adopted component in robotics, spanning perception, planning, and control rather than remaining confined to one subdomain. It organizes the field into three adoption routes: pretrained foundation models that give robots language understanding and zero-shot generalization; transformer variants integrated with deep reinforcement learning that treat decision-making as sequence modeling and enable long-horizon planning; and transformer-based systems that improve perception, planning, and control performance. The paper argues that these routes share a common engine, the self-attention mechanism, which lets a model relate every part of an input sequence to every other part regardless of distance. If the picture is correct, roboticists should expect transformer-based architectures to become the default scaffold for new autonomous systems.

What carries the argument

The load-bearing object is the self-attention mechanism introduced by the original Transformer, which lets every element of an input sequence attend to every other element in parallel, independent of sequence length. In this review it is the shared engine behind all three adoption routes: attention over language tokens gives robots instruction-following, attention over visual tokens gives perception and grounding, and attention over sequences of states, actions, and rewards gives reinforcement learning and planning. The review's organizing distinction among foundation models, deep-reinforcement-learning hybrids, and perception-planning-control systems is the map that shows this one mechanism spreading across robotics.

What would settle it

Run a systematic search of robotics papers from 2022 through 2024 using the review's topical keywords and check whether transformer-based methods appear in a substantial share of perception, planning, and control systems; then attempt to reproduce the headline numbers on the original benchmarks, including TransformerMPC's 6.8x to 34.9x speedups, MarineFormer's 20% success-rate gain, and the 92.73% fault-instruction detection rate. If transformer use is concentrated in a narrow slice of robotics or those numbers do not reproduce, the review's characterization of adoption is unsubstantiated.

Watch

Extended reading notes

Core claim

The paper's central claim is that in robotics, Transformers are being adopted in three major ways: as pretrained foundation models facilitating human-robot interaction and generalization, as transformer variants integrated with deep reinforcement learning for long-horizon planning, and as components that enhance perception, planning, and control systems. It presents evidence from generalist robot policies such as RT-1, RT-2, OpenVLA, Octo, and π0, from sequence-modeling RL methods such as the Decision Transformer and Trajectory Transformer, and from perception and control systems such as CLIP, SAM, ViNT, and TransformerMPC. The paper reports representative results, including a 92.73% detection rate for faulty instructions in construction human-robot interaction, a 20% success-rate improvement for MarineFormer, and 6.8x to 34.9x speedups from TransformerMPC, as evidence of this trend.

Load-bearing premise

The review runs no experiments and uses no systematic search criteria, so its central picture rests on the assumption that the quantitative results reported in the cited primary papers are accurate and that the informally chosen set of papers fairly represents the field.

Editorial extensions

If this is right

  • If the central claim is correct, transformer-based components will continue displacing CNNs, RNNs, and classical planners as the default building blocks for robotic systems.
  • Shared multi-embodiment datasets such as Open X-Embodiment become a critical resource, so data collection, standardization, and embodiment balance are where the field's bottlenecks now sit.
  • Sequence-modeling formulations of reinforcement learning make long-horizon tasks more tractable, allowing robot learning systems to be pretrained offline and then fine-tuned online.
  • Transformer-based control accelerators such as TransformerMPC can cut runtime by an order of magnitude while preserving constraint satisfaction, making real-time deployment on robots more practical.
  • Zero-shot and few-shot generalization is treated as achievable for manipulation in controlled settings, while the review also warns that it is not yet guaranteed in unstructured, outdoor, or safety-critical environments.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Going beyond the paper: if this trend continues, robotics may follow the same trajectory as natural language processing, converging on a few large generalist policies pretrained on pooled data and then fine-tuned per task.
  • The headline numbers in the review come from heterogeneous tasks, metrics, and baselines, so direct comparison across the three routes is not possible without a standardized evaluation protocol.
  • Because the review selected papers informally, the most direct test of its broad-adoption thesis is a systematic literature search; a search that found transformer use concentrated in manipulation and largely absent from field robotics would require qualifying that thesis.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. This manuscript is a narrative review of transformer architectures for robotic applications. After a background section on the original transformer and its efficient, multimodal, and sparse variants, it organizes robotic applications into three threads: pre-trained foundation models (including vision-language-action models and human-robot interaction), transformer-based deep reinforcement learning variants, and transformer use in perception, planning, and control. It closes with challenges and future directions. The paper presents no new experiments or derivations; its central descriptive claim is that transformers are now broadly adopted across robotic perception, planning, and control, which is consistent with the current literature.

Significance. If the survey's attributions were reliable, the paper would serve as a useful broad overview of a fast-moving area, with a sensible tripartite organization and up-to-date coverage (through late 2024) of models such as RT-2, OpenVLA, Octo, π0, and TransformerMPC. The authors explicitly acknowledge limitations in Section 4 (dataset imbalance, sim-to-real gap, safety, memory footprint) and propose concrete future directions such as quantization for resource-constrained deployment and cross-embodiment training. The main weakness is that several specific quantitative and architectural claims are mis-cited or unverifiable from the reference list; because a review's value depends on readers being able to trace claims to sources, these issues are load-bearing and must be corrected before the paper can be relied upon as a survey.

major comments (5)
  1. [§3.2, references [101]/[102]] The same work (Parisotto and Salakhutdinov, 'Efficient transformers in reinforcement learning using actor-learner distillation') appears as both [101] and [102]. Reference [101] is then used for two incompatible claims: that transformers are effective as world models, and that 'MarineFormer [101] ... demonstrated 20% success rate improvement' for marine environments. The cited paper is neither primarily a world-model paper nor MarineFormer; MarineFormer appears to be a different work that is missing from the bibliography. These mis-citations make specific quantitative claims untraceable and must be fixed, either by correcting the citations or by inserting the correct references.
  2. [§2.3.1, Table 1 (Efficient Transformers)] The table row 'Synthesizer [138]' cites reference [138], which is Tay et al., 'Sparse Sinkhorn Attention.' The Synthesizer architecture is a separate paper (Tay et al., 'Synthesizer: Rethinking Self-Attention for Transformer Models') and is not in the bibliography, while the Sinkhorn Transformer is correctly associated with [138] in Table 2 of Section 2.3.3. The two attention mechanisms are conflated; please replace the citation and/or add the Synthesizer reference.
  3. [§2.3.3, Table 2 (Sparse and Adaptive Transformers)] The table lists 'Routing Transformer [121]' twice with different asymptotic complexities: O(n log n) and O(n^{1.5} d^*). This is internally inconsistent, and neither entry is adequately justified. Please consolidate the duplicate entries into a single row and state the complexity reported in the original Routing Transformer paper (Roy et al., 2020), or explain why two different forms are listed.
  4. [§3.3.2, Planning] The sentence 'Gated Transformer-XL [23] for Long-Term Memory' cites reference [23], which is Transformer-XL by Dai et al.; the Gated Transformer-XL model is reference [103] (Parisotto et al., 'Stabilizing Transformers for Reinforcement Learning'), already correctly cited in Section 3.2. In the same subsection, the ViNT model is discussed ('ViNT, a Transformer-based model for visual navigation, demonstrates significant potential...') without any citation at all. Please correct the attribution and add a citation for ViNT.
  5. [§3.1, Human-Robot Collaboration] The paragraph on human-robot collaboration reports specific success rates: '92.73% detection of fault instructions by humans' [104] and '99.95% success rate in simulation' [105]. From the reference list alone, I cannot confirm that these exact numbers appear in the cited papers, since [104] and [105] are construction-HRI papers but no page numbers, sections, tables, or experiment descriptions are given. Because these figures are the only quantitative evidence in that paragraph, please verify the attributions or qualify the claims as reported in the cited sources.
minor comments (6)
  1. [§2.3.2 and §3.1] Both sections contain unresolved 'see Figure ??' cross-references (after the ViLBERT description and after the OpenVLA description, respectively); these must be fixed before publication.
  2. [§2.3.2, Table 3] The table entry 'VisualGPT [74]' is mis-cited: reference [74] is the VisualBERT paper, which is correctly named in the bullet list immediately below the table. Please correct the table entry or add the actual VisualGPT reference.
  3. [§4] There is a typo in the last paragraph: 'zero-short generalization' should read 'zero-shot generalization.'
  4. [§3.2] The phrase 'having low inference times while maintaining low sample effeciency of transformers' appears to state the opposite of the intended meaning; the actor-learner distillation approach is designed to retain the sample efficiency of transformers, so please rephrase to 'high sample efficiency' or similar.
  5. [Figure 2 caption] The caption is grammatically tangled and the subfigure references are unclear: 'while Franka has the most number of scenes see 2a x-Arm and Google Robot have the biggest contribution to trajectory data see 2b and 2c' does not map cleanly to panels (a)–(e). Please rewrite the caption to describe each panel explicitly.
  6. [References] Reference [93] is an incomplete URL-like entry containing a sentence fragment rather than a full citation, and reference [59] has a garbled author list and ordering. Both need to be reformatted according to the journal style.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the review makes no predictions, fits no parameters, derives no equations from its own claims, and its load-bearing evidence is external cited work.

full rationale

This paper is a narrative literature review. It contains no derivation chain, no fitted parameters, no empirical predictions, and no mathematical claims that reduce to its own inputs. The central claim — that Transformers are being adopted in robotics through foundation models, DRL variants, and perception/planning/control — is supported by numerous citations to external primary papers and to other surveys. None of those citations are self-citations by the present authors, and none are invoked as a uniqueness theorem or as a forced choice that would make the review's taxonomy self-justifying. The survey's limitations (dataset imbalance, sim-to-real gaps, safety) are stated as open challenges, not as derived results. Weaknesses such as possible misattribution of specific numbers concern fidelity to external sources, which is a correctness/accuracy issue, not circularity. A result that depends on the accuracy of other people's experiments is still an external, non-circular dependence. Therefore, no circular step can be exhibited, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Review paper; no numbers are fitted to data and no new entities are introduced. The only substantive assumption is that the surveyed primary literature is accurately reported.

assumptions (1)
  • domain assumption Quantitative results reported in the cited primary papers are accurate and reproducible as stated.
    The review repeats performance numbers such as 92.73% fault instruction detection and 99.95% simulation success (Section 3.1), a 20% success rate improvement for MarineFormer (Section 3.2), and TransformerMPC speedups of 6.8x to 34.9x (Section 3.3) without independent verification.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Advances in Transformers for Robotic Applications: A Review." pith.science (2026). https://pith.science/paper/TAL5TDBK

@misc{pith2026241210599,
  author       = {Pith},
  title        = {Pith review of: Advances in Transformers for Robotic Applications: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TAL5TDBK}},
  note         = {Machine review of arXiv:2412.10599}
}
read the original abstract

The introduction of Transformers architecture has brought about significant breakthroughs in Deep Learning (DL), particularly within Natural Language Processing (NLP). Since their inception, Transformers have outperformed many traditional neural network architectures due to their "self-attention" mechanism and their scalability across various applications. In this paper, we cover the use of Transformers in Robotics. We go through recent advances and trends in Transformer architectures and examine their integration into robotic perception, planning, and control for autonomous systems. Furthermore, we review past work and recent research on use of Transformers in Robotics as pre-trained foundation models and integration of Transformers with Deep Reinforcement Learning (DRL) for autonomous systems. We discuss how different Transformer variants are being adapted in robotics for reliable planning and perception, increasing human-robot interaction, long-horizon decision-making, and generalization. Finally, we address limitations and challenges, offering insight and suggestions for future research directions.

Figures

Figures reproduced from arXiv: 2412.10599 by the authors.

Figure 1
Figure 1. The Original ”Vanilla” Transformer Architecture [144]. Visual representation taken [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. The Open X-Embodiment(OXE) Dataset, graphs are from the original paper. a) The [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Visual representation of OpenVLA architecture, generating low-level robot action [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

173 extracted references · 56 canonical work pages

  1. [101]

    Efficient transformers in reinforcement learn- ing using actor-learner distillation, 2021

    Emilio Parisotto and Ruslan Salakhutdinov. Efficient transformers in reinforcement learn- ing using actor-learner distillation, 2021

  2. [138]

    Sparse sinkhorn attention

    Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Sparse sinkhorn attention. arXiv preprint arXiv:2002.11296 , 2020

  3. [102]

    Efficient transformers in reinforcement learn- ing using actor-learner distillation

    Emilio Parisotto and Ruslan Salakhutdinov. Efficient transformers in reinforcement learn- ing using actor-learner distillation. ArXiv, abs/2104.01655, 2021

  4. [121]

    Efficient content- based sparse attention with routing transformers, 2020

    Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier. Efficient content- based sparse attention with routing transformers, 2020

  5. [23]

    Le, and Ruslan 17 Salakhutdinov

    Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, and Ruslan 17 Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length con- text, 2019

  6. [103]

    Francis Song, Jack W

    Emilio Parisotto, H. Francis Song, Jack W. Rae, Razvan Pascanu, Caglar Gulcehre, Sid- dhant M. Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, Matthew M. Botvinick, Nicolas Heess, and Raia Hadsell. Stabilizing transformers for reinforcement learning, 2019

  7. [104]

    Menassa, and Vineet R

    Somin Park, Carol C. Menassa, and Vineet R. Kamat. Integrating large language models with multimodal virtual reality interfaces to support collaborative human-robot construc- tion work, 2024

  8. [105]

    Menassa, Vineet R

    Somin Park, Xi Wang, Carol C. Menassa, Vineet R. Kamat, and Joyce Y. Chai. Natu- ral language instructions for intuitive human interaction with robotic assistants in field construction work. Automation in Construction , 161:105345, May 2024

Show all 173 references
  1. [1]

    Generating out-of-distribution scenarios using language models, 2024

    Erfan Aasi, Phat Nguyen, Shiva Sreeram, Guy Rosman, Sertac Karaman, and Daniela Rus. Generating out-of-distribution scenarios using language models, 2024

  2. [2]

    Constrained policy opti- mization, 2017

    Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. Constrained policy opti- mization, 2017

  3. [3]

    Pranav Agarwal, Aamer Abdul Rahman, Pierre-Luc St-Charles, Simon J. D. Prince, and Samira Ebrahimi Kahou. Transformers in reinforcement learning: A survey, 2023

  4. [4]

    Lstm inefficiency in long-term dependencies regression problems

    Safwan Mahmood Al-Selwi, Mohd Fadzil, Said Jadid Abdulkadir, and Amgad Muneer. Lstm inefficiency in long-term dependencies regression problems. Journal of Advanced Research in Applied Sciences and Engineering Technology , 2023

  5. [5]

    Advances in medical im- age analysis with vision transformers: A comprehensive review

    Reza Azad, Amirhossein Kazerouni, Moein Heidari, Ehsan Khodapanah Aghdam, Amir Molaei, Yiwei Jia, Abin Jose, Rijo Roy, and Dorit Merhof. Advances in medical im- age analysis with vision transformers: A comprehensive review. Medical image analysis , 91:103000, 2023

  6. [6]

    The option-critic architecture, 2016

    Pierre-Luc Bacon, Jean Harb, and Doina Precup. The option-critic architecture, 2016

  7. [7]

    data2vec: A general framework for self-supervised learning in speech, vision and language, 2022

    Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli. data2vec: A general framework for self-supervised learning in speech, vision and language, 2022

  8. [8]

    Peters, and Arman Cohan

    Iz Beltagy, Matthew E. Peters, and Arman Cohan. Longformer: The long-document transformer, 2020

  9. [9]

    π0: A vision-language-action flow model for general robot control

    Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch, Lucy Xiaoyang Shi, Jam...

  10. [10]

    Josh C. Bongard. Probabilistic robotics. sebastian thrun, wolfram burgard, and dieter fox. (2005, mit press.) 647 pages. Artificial Life, 14:227–229, 2008

  11. [11]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control, 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasm...

  12. [12]

    Rt-1: Robotics transformer for real-world control at scale, 2023

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov,...

  13. [13]

    Lewis, and Satinder Singh

    Ethan Brooks, Logan Walls, Richard L. Lewis, and Satinder Singh. Large language models can implement policy iteration, 2023

  14. [14]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  15. [15]

    Plancritic: Formal planning with human feedback, 2024

    Owen Burns, Dana Hughes, and Katia Sycara. Plancritic: Formal planning with human feedback, 2024

  16. [16]

    Radfar, Athanasios Mouchtaris, Brian King, and Siegfried Kun- zmann

    Feng-Ju Chang, Martin H. Radfar, Athanasios Mouchtaris, Brian King, and Siegfried Kun- zmann. End-to-end multi-channel transformer for speech recognition.ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 5884–5888, 2021

  17. [17]

    Decision transformer: Reinforcement learning via sequence modeling, 2021

    Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. Decision transformer: Reinforcement learning via sequence modeling, 2021

  18. [18]

    Clip2scene: Towards label-efficient 3d scene understanding by clip, 2023

    Runnan Chen, Youquan Liu, Lingdong Kong, Xinge Zhu, Yuexin Ma, Yikang Li, Yue- nan Hou, Yu Qiao, and Wenping Wang. Clip2scene: Towards label-efficient 3d scene understanding by clip, 2023

  19. [19]

    Generating long sequences with sparse transformers, 2019

    Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. Generating long sequences with sparse transformers, 2019

  20. [20]

    Rethinking attention with performers, 2022

    Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. Rethinking attention with performers, 2022

  21. [21]

    Embodiment Collaboration, Abby O’Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Herzog, Alex Ir- pan, Alexander Khazatsky, Anant Rai,...

  22. [22]

    Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

  23. [24]

    Fu, Stefano Ermon, Atri Rudra, and Christopher R´ e

    Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher R´ e. Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022

  24. [25]

    Universal transformers, 2019

    Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. Universal transformers, 2019

  25. [26]

    Bert: Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics , 2019

  26. [27]

    DiPietro and Gregory Hager

    Robert S. DiPietro and Gregory Hager. Deep learning: Rnns and lstm. 2020

  27. [28]

    Scaling cross- embodied learning: One policy for manipulation, navigation, locomotion and aviation, 2024

    Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. Scaling cross- embodied learning: One policy for manipulation, navigation, locomotion and aviation, 2024

  28. [29]

    An image is worth 16x16 words: Transformers for image recognition at scale, 2021

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at...

  29. [30]

    Addressing some limitations of transformers with feedback memory, 2021

    Angela Fan, Thibaut Lavril, Edouard Grave, Armand Joulin, and Sainbayar Sukhbaatar. Addressing some limitations of transformers with feedback memory, 2021

  30. [31]

    Model-agnostic meta-learning for fast adaptation of deep networks, 2017

    Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks, 2017

  31. [32]

    Foundation models in robotics: Ap- plications, challenges, and the future

    Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, Brian Ichter, Danny Driess, Jiajun Wu, Cewu Lu, and Mac Schwager. Foundation models in robotics: Ap- plications, challenges, and the fu...

  32. [33]

    D. Fox, W. Burgard, and S. Thrun. Markov localization for mobile robots in dynamic environments. Journal of Artificial Intelligence Research , 11:391–427, November 1999

  33. [34]

    Generalized decision transformer for offline hindsight information matching, 2022

    Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu. Generalized decision transformer for offline hindsight information matching, 2022

  34. [35]

    Physically grounded vision-language models for robotic ma- nipulation, 2024

    Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Ma- jumdar, and Dorsa Sadigh. Physically grounded vision-language models for robotic ma- nipulation, 2024

  35. [36]

    Learning in reality: A case study of stanley, the robot that won the darpa challenge

    Christian Glaser and Philipp Hennig. Learning in reality: A case study of stanley, the robot that won the darpa challenge. 2012

  36. [37]

    Rvt: Robotic view transformer for 3d object manipulation

    Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox. Rvt: Robotic view transformer for 3d object manipulation. ArXiv, abs/2306.14896, 2023

  37. [38]

    Metamorph: Learning universal controllers with transformers, 2022

    Agrim Gupta, Linxi Fan, Surya Ganguli, and Li Fei-Fei. Metamorph: Learning universal controllers with transformers, 2022

  38. [39]

    Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018

    Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018. 18

  39. [40]

    Transformers in medical image analysis: A review

    Kelei He, Chen Gan, Zhuoyuan Li, Islem Rekik, Zihao Yin, Wen Ji, Yang Gao, Qian Wang, Junfeng Zhang, and Dinggang Shen. Transformers in medical image analysis: A review. ArXiv, abs/2202.12165, 2022

  40. [41]

    Henry, Onyeka Emebob, and Conrad Asotie Omonhinmin

    Emerald U. Henry, Onyeka Emebob, and Conrad Asotie Omonhinmin. Vision transform- ers in medical imaging: A review. ArXiv, abs/2211.10043, 2022

  41. [42]

    Convolutional vision transformer as a path following controller for omnidirectional robots

    Sandesh Hiremath, Cheng-Yi Huang, Argtim Tika, and Naim Bajcinca. Convolutional vision transformer as a path following controller for omnidirectional robots. In 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages 16633–16639, 2024

  42. [43]

    Long short-term memory.Neural Computation, 9:1735–1780, 1997

    Sepp Hochreiter and J¨ urgen Schmidhuber. Long short-term memory.Neural Computation, 9:1735–1780, 1997

  43. [44]

    On transforming re- inforcement learning with transformers: The development trajectory

    Shengchao Hu, Li Shen, Ya Zhang, Yixin Chen, and Dacheng Tao. On transforming re- inforcement learning with transformers: The development trajectory. IEEE Transactions on Pattern Analysis and Machine Intelligence , 46(12):8580–8599, 2024

  44. [45]

    Toward general-purpose robots via foundation models: A survey and meta-analysis, 2024

    Yafei Hu, Quanting Xie, Vidhi Jain, Jonathan Francis, Jay Patrikar, Nikhil Keetha, Se- ungchan Kim, Yaqi Xie, Tianyi Zhang, Hao-Shu Fang, Shibo Zhao, Shayegan Omidshafiei, Dong-Ki Kim, Ali akbar Agha-mohammadi, Katia Sycara, Matthew Johnson-Roberson, Dhruv Batra, Xiaolong Wang...

  45. [46]

    Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022

    Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. Language models as zero-shot planners: Extracting actionable knowledge for embodied agents, 2022

  46. [47]

    Hpcneu- ronet: Advancing neuromorphic audio signal processing with transformer-enhanced spik- ing neural networks

    Murat Isik, Hiruna Vishwamith, Kayode Inadagbo, and Ismail Can Dikmen. Hpcneu- ronet: Advancing neuromorphic audio signal processing with transformer-enhanced spik- ing neural networks. 2024 4th Interdisciplinary Conference on Electrics and Computer (INTCEC), pages 1–7, 2023

  47. [48]

    A comprehensive survey on applications of transformers for deep learning tasks

    Saidul Islam, Hanae Elmekki, Ahmed Elsebai, Jamal Bentahar, Najat Drawel, Gaith Rjoub, and Witold Pedrycz. A comprehensive survey on applications of transformers for deep learning tasks. ArXiv, abs/2306.07303, 2023

  48. [49]

    Perceiver: General perception with iterative attention, 2021

    Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention, 2021

  49. [50]

    Bc-z: Zero-shot task generalization with robotic imita- tion learning, 2022

    Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn. Bc-z: Zero-shot task generalization with robotic imita- tion learning, 2022

  50. [51]

    When to trust your model: Model-based policy optimization, 2021

    Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine. When to trust your model: Model-based policy optimization, 2021

  51. [52]

    Offline reinforcement learning as one big sequence modeling problem

    Michael Janner, Qiyang Li, and Sergey Levine. Offline reinforcement learning as one big sequence modeling problem. In Advances in Neural Information Processing Systems, 2021

  52. [53]

    Flair: Feeding via long-horizon acquisition of realistic dishes, 2024

    Rajat Kumar Jenamani, Priya Sundaresan, Maram Sakr, Tapomayukh Bhattacharjee, and Dorsa Sadigh. Flair: Feeding via long-horizon acquisition of realistic dishes, 2024

  53. [54]

    A survey of robot intelligence with large language models

    Hyeongyo Jeong, Haechan Lee, Changwon Kim, and Sungtae Shin. A survey of robot intelligence with large language models. Applied Sciences, 14(19), 2024. 19

  54. [55]

    Le, Yunhsuan Sung, Zhen Li, and Tom Duerig

    Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V. Le, Yunhsuan Sung, Zhen Li, and Tom Duerig. Scaling up visual and vision-language representation learning with noisy text supervision, 2021

  55. [56]

    Highly accurate protein structure prediction with alphafold

    John Jumper, Richard Evans, Alexander Pritzel, and et al. Highly accurate protein structure prediction with alphafold. Nature, 596:583–589, 2021

  56. [57]

    Trans- formers are rnns: Fast autoregressive transformers with linear attention, 2020

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran¸ cois Fleuret. Trans- formers are rnns: Fast autoregressive transformers with linear attention, 2020

  57. [58]

    Segment anything in high quality, 2023

    Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu, Yu-Wing Tai, Chi-Keung Tang, and Fisher Yu. Segment anything in high quality, 2023

  58. [59]

    Real-world robot applications of foundation models: a review

    Andrew Gambardella Jiaxian Guo Chris Paxton Kento Kawaharazuka, Tatsuya Mat- sushima and Andy Zeng. Real-world robot applications of foundation models: a review. Advanced Robotics, 38(18):1232–1254, 2024

  59. [60]

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Sid- dharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, Peter David Fagan, Joey Hejna, Masha Itkina, Marion Lepert, Yecheng Ja- son Ma, Patrick Tree ...

  60. [61]

    What and when to explain? on-road evaluation of explanations in highly automated vehicles

    Gwangbin Kim, Dohyeon Yeo, Taewoo Jo, Daniela Rus, and SeungJun Kim. What and when to explain? on-road evaluation of explanations in highly automated vehicles. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. , 7(3), September 2023

  61. [62]

    Openvla: An open-source vision-language-action model, 2024

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. Openvla...

  62. [63]

    A survey on integration of large language models with intelligent robots

    Yeseung Kim, Dohyun Kim, Jieun Choi, Jisang Park, Nayoung Oh, and Daehyung Park. A survey on integration of large language models with intelligent robots. Intelligent Service Robotics, 17(5):1091–1107, August 2024

  63. [64]

    Berg, Wan-Yen Lo, Piotr Doll´ ar, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Doll´ ar, and Ross Girshick. Segment anything, 2023

  64. [65]

    Reformer: The efficient transformer, 2020

    Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. Reformer: The efficient transformer, 2020. 20

  65. [66]

    Reinforcement learning in robotics: A survey

    Jens Kober, J Andrew Bagnell, and Jan Peters. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research , 32(11):1238–1274, 2013

  66. [67]

    Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021

    Ilya Kostrikov, Denis Yarats, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels, 2021

  67. [68]

    Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. Communications of the ACM , 60:84 – 90, 2012

  68. [69]

    Conservative q-learning for offline reinforcement learning, 2020

    Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. Conservative q-learning for offline reinforcement learning, 2020

  69. [70]

    Phong Le and Willem H. Zuidema. Quantifying the vanishing gradient and long distance dependency problem in recursive neural networks and recursive lstms. InRep4NLP@ACL, 2016

  70. [71]

    Fnet: Mixing tokens with fourier transforms, 2022

    James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon. Fnet: Mixing tokens with fourier transforms, 2022

  71. [72]

    End-to-end training of deep visuomotor policies

    Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages 1334–1342, 2016

  72. [73]

    Landman, and S

    Jun Li, Junyu Chen, Yucheng Tang, Ce Wang, Bennett A. Landman, and S. Kevin Zhou. Transforming medical imaging with transformers? a comparative review of key properties, current progresses, and future perspectives. Medical image analysis , 85:102762, 2022

  73. [74]

    Visualbert: A simple and performant baseline for vision and language, 2019

    Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. Visualbert: A simple and performant baseline for vision and language, 2019

  74. [75]

    Grounded language-image pre-training, 2022

    Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jian- feng Gao. Grounded language-image pre-training, 2022

  75. [76]

    Shapegrasp: Zero-shot task-oriented grasping with large language models through geometric decomposition, 2024

    Samuel Li, Sarthak Bhagat, Joseph Campbell, Yaqi Xie, Woojun Kim, Katia Sycara, and Simon Stepputtis. Shapegrasp: Zero-shot task-oriented grasping with large language models through geometric decomposition, 2024

  76. [77]

    A survey on transformers in reinforcement learning, 2023

    Wenzhe Li, Hao Luo, Zichuan Lin, Chongjie Zhang, Zongqing Lu, and Deheng Ye. A survey on transformers in reinforcement learning, 2023

  77. [78]

    Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020

    Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. Oscar: Object-semantics aligned pre-training for vision-language tasks, 2020

  78. [79]

    Lillicrap, Jonathan J

    Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yu- val Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning, 2019

  79. [80]

    A survey of transformers

    Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. A survey of transformers. AI Open, 3:111–132, 2022

  80. [81]

    Instruction-following agents with multimodal transformer, 2023

    Hao Liu, Lisa Lee, Kimin Lee, and Pieter Abbeel. Instruction-following agents with multimodal transformer, 2023. 21

  81. [82]

    Constrained decision transformer for offline safe reinforcement learning

    Zuxin Liu, Zijian Guo, Yihang Yao, Zhepeng Cen, Wenhao Yu, Tingnan Zhang, and Ding Zhao. Constrained decision transformer for offline safe reinforcement learning. In An- dreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett, edito...

  82. [83]

    Few-shot subgoal plan- ning with language models, 2022

    Lajanugen Logeswaran, Yao Fu, Moontae Lee, and Honglak Lee. Few-shot subgoal plan- ning with language models, 2022

  83. [84]

    Uncertainty-aware hybrid paradigm of nonlinear mpc and model- based rl for offroad navigation: Exploration of transformers in the predictive model

    Faraz Lotfi, Khalil Virji, Farnoosh Faraji, Lucas Berry, Andrew Holliday, David Meger, and Gregory Dudek. Uncertainty-aware hybrid paradigm of nonlinear mpc and model- based rl for offroad navigation: Exploration of transformers in the predictive model. In 2024 IEEE Internatio...

  84. [85]

    Multi- agent actor-critic for mixed cooperative-competitive environments, 2020

    Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, Pieter Abbeel, and Igor Mordatch. Multi- agent actor-critic for mixed cooperative-competitive environments, 2020

  85. [86]

    Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019

    Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks, 2019

  86. [87]

    Mega: Moving average equipped gated attention, 2023

    Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. Mega: Moving average equipped gated attention, 2023

  87. [88]

    Transformers are sample efficient world models

    Vincent Micheli, Eloi Alonso, and Franccois Fleuret. Transformers are sample efficient world models. ArXiv, abs/2209.00588, 2022

  88. [89]

    Human-level control through deep reinforcement learning

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, et al. Human-level control through deep reinforcement learning. Nature, 518(7540):529–533, 2015

  89. [90]

    Integrating reinforcement learning with foundation models for autonomous robotics: Methods and perspectives, 2024

    Angelo Moroncelli, Vishal Soni, Asad Ali Shahid, Marco Maccarini, Marco Forgione, Dario Piga, Blerina Spahiu, and Loris Roveda. Integrating reinforcement learning with foundation models for autonomous robotics: Methods and perspectives, 2024

  90. [91]

    Jane Mulligan and Gregory Z. Grudic. Editorial for journal of field robotics—special issue on machine learning based robotics in unstructured environments. Journal of Field Robotics, 23, 2006

  91. [92]

    Data-efficient hierarchical reinforcement learning, 2018

    Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine. Data-efficient hierarchical reinforcement learning, 2018

  92. [93]

    B. A. Newman et al. Many household tasks can be posed as multi-object rearrangement tasks, but solutions to these problems often target a single, hand defined solution or are ... Carnegie Mellon University , page 4, 2024. Retrieved from https://www.andrew.cmu. edu/~bnewman1/data

  93. [94]

    Newman, Pranay Gupta, Kris Kitani, Yonatan Bisk, Henny Admoni, and Chris Paxton

    Benjamin A. Newman, Pranay Gupta, Kris Kitani, Yonatan Bisk, Henny Admoni, and Chris Paxton. Degustabot: Zero-shot visual preference estimation for personalized multi- object rearrangement, 2024

  94. [95]

    Policy invariance under reward transformations: Theory and application to reward shaping

    Andrew Y Ng, Daishi Harada, and Stuart J Russell. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the 16th International Conference on Machine Learning (ICML) , pages 278–287, 1999. 22

  95. [96]

    Nils J. Nilsson. Shakey the robot. 1984

  96. [97]

    Deep explo- ration via bootstrapped dqn, 2016

    Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy. Deep explo- ration via bootstrapped dqn, 2016

  97. [98]

    Voicepilot: Harnessing llms as speech interfaces for physically assistive robots

    Akhil Padmanabha, Jessie Yuan, Janavi Gupta, Zulekha Karachiwalla, Carmel Majidi, Henny Admoni, and Zackory Erickson. Voicepilot: Harnessing llms as speech interfaces for physically assistive robots. In Proceedings of the 37th Annual ACM Symposium on User Interface Software an...

  98. [99]

    Lifelong robot learning with human assisted language planners, 2023

    Meenal Parakh, Alisha Fong, Anthony Simeonov, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Lifelong robot learning with human assisted language planners, 2023

  99. [100]

    Lifelong robot learning with human assisted language planners

    Meenal Parakh, Alisha Fong, Anthony Simeonov, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Lifelong robot learning with human assisted language planners. In 2024 IEEE International Conference on Robotics and Automation (ICRA) , pages 523–529, 2024

  100. [106]

    Dexhub and dart: Towards internet scale robot data collection, 2024

    Younghyo Park, Jagdeep Singh Bhatia, Lars Ankile, and Pulkit Agrawal. Dexhub and dart: Towards internet scale robot data collection, 2024

  101. [107]

    Efros, and Trevor Darrell

    Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell. Curiosity-driven exploration by self-supervised prediction, 2017

  102. [108]

    Transformers in the real world: A survey on nlp applications

    Narendra Patwardhan, Stefano Marrone, and Carlo Sansone. Transformers in the real world: A survey on nlp applications. Information, 14(4), 2023

  103. [109]

    Chun-Cheng Peng and G. D. Magoulas. Sequence processing with recurrent neural net- works. In Encyclopedia of Artificial Intelligence , 2009

  104. [110]

    Miller, and Sebastian Riedel

    Fabio Petroni, Tim Rockt¨ aschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexan- der H. Miller, and Sebastian Riedel. Language models as knowledge bases?, 2019

  105. [111]

    Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher R´ e

    Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher R´ e. Hyena hierarchy: Towards larger convolutional language models, 2023. 23

  106. [112]

    Learning transferable visual models from natural language supervision, 2021

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021

  107. [113]

    Improving language understanding by generative pre-training

    Alec Radford and Karthik Narasimhan. Improving language understanding by generative pre-training. 2018

  108. [114]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1– 67, 2020

  109. [115]

    Zero-shot text-to-image generation, 2021

    Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation, 2021

  110. [116]

    Rlds: an ecosystem to generate, share and use datasets in reinforcement learning, 2021

    Sabela Ramos, Sertan Girgin, L´ eonard Hussenot, Damien Vincent, Hanna Yakubovich, Daniel Toyama, Anita Gergely, Piotr Stanczyk, Raphael Marinier, Jeremiah Harmsen, Olivier Pietquin, and Nikola Momchev. Rlds: an ecosystem to generate, share and use datasets in reinforcement le...

  111. [117]

    Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S

    Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker, Salman Khan, Hisham Cholakkal, Rao M. Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S. Khan. Glamm: Pixel grounding large multimodal model, 2024

  112. [118]

    Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018

    Tabish Rashid, Mikayel Samvelyan, Christian Schroeder de Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson. Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning, 2018

  113. [119]

    Sam 2: Segment anything in images and videos, 2024

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨ adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll´ ar, and Christoph Fei...

  114. [120]

    Generalist agents

    Scott Reed, Suraj Nair, Felix Hill, Oriol Vinyals, Nando de Freitas, et al. Generalist agents. arXiv preprint arXiv:2205.06175 , 2022

  115. [122]

    Rumelhart, Geoffrey E

    David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. Learning representa- tions by back-propagating errors. Nature, 323:533–536, 1986

  116. [123]

    Jordan, and Pieter Abbeel

    John Schulman, Sergey Levine, Philipp Moritz, Michael I. Jordan, and Pieter Abbeel. Trust region policy optimization, 2017

  117. [124]

    Proxi- mal policy optimization algorithms, 2017

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proxi- mal policy optimization algorithms, 2017

  118. [125]

    Behavior transformers: Cloning k modes with one stone, 2022

    Nur Muhammad Mahi Shafiullah, Zichen Jeff Cui, Ariuntuya Altanzaya, and Lerrel Pinto. Behavior transformers: Cloning k modes with one stone, 2022

  119. [126]

    Fahad Shamshad, Salman Hameed Khan, Syed Waqas Zamir, Muhammad Haris Khan, Munawar Hayat, Fahad Shahbaz Khan, and H. Fu. Transformers in medical imaging: A survey. Medical image analysis , 88:102802, 2022. 24

  120. [127]

    Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network

    Alex Sherstinsky. Fundamentals of recurrent neural network (rnn) and long short-term memory (lstm) network. ArXiv, abs/1808.03314, 2018

  121. [128]

    Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn

    Lucy Xiaoyang Shi, Zheyuan Hu, Tony Z. Zhao, Archit Sharma, Karl Pertsch, Jianlan Luo, Sergey Levine, and Chelsea Finn. Yell at your robot: Improving on-the-fly from language corrections, 2024

  122. [129]

    Perspectives and prospects on trans- former architecture for cross-modal tasks with language and vision

    Andrew Shin, Masato Ishii, and Takuya Narihira. Perspectives and prospects on trans- former architecture for cross-modal tasks with language and vision. International Journal of Computer Vision , 130:435 – 454, 2021

  123. [130]

    Flava: A foundational language and vi- sion alignment model, 2022

    Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela. Flava: A foundational language and vi- sion alignment model, 2022

  124. [131]

    Adaptive attention span in transformers, 2019

    Sainbayar Sukhbaatar, Edouard Grave, Piotr Bojanowski, and Armand Joulin. Adaptive attention span in transformers, 2019

  125. [132]

    Videobert: A joint model for video and language representation learning, 2019

    Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. Videobert: A joint model for video and language representation learning, 2019

  126. [133]

    Plate: Visually-grounded planning with transformers in procedural tasks

    Jiankai Sun, De-An Huang, Bo Lu, Yun-Hui Liu, Bolei Zhou, and Animesh Garg. Plate: Visually-grounded planning with transformers in procedural tasks. IEEE Robotics and Automation Letters, 7(2):4924–4930, April 2022

  127. [134]

    Eva-clip: Improved training techniques for clip at scale, 2023

    Quan Sun, Yuxin Fang, Ledell Wu, Xinlong Wang, and Yue Cao. Eva-clip: Improved training techniques for clip at scale, 2023

  128. [135]

    Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. Sequence to sequence learning with neural networks. ArXiv, abs/1409.3215, 2014

  129. [136]

    Reinforcement Learning: An Introduction

    Richard S Sutton and Andrew G Barto. Reinforcement Learning: An Introduction. MIT Press, 2nd edition, 2018

  130. [137]

    Efficient transformers: A survey

    Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. Efficient transformers: A survey. ACM Computing Surveys , 55:1 – 28, 2020

  131. [139]

    Octo: An open-source generalist robot policy, 2024

    Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Lawrence Yunliang Chen, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo...

  132. [140]

    Robots that use language

    Stefanie Tellex, Nakul Gopalan, Hadas Kress-Gazit, and Cynthia Matuszek. Robots that use language. Annu. Rev. Control. Robotics Auton. Syst. , 3:25–55, 2020

  133. [141]

    Vision-and- dialog navigation, 2019

    Jesse Thomason, Michael Murray, Maya Cakmak, and Luke Zettlemoyer. Vision-and- dialog navigation, 2019

  134. [142]

    Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia M

    Sebastian Thrun, Michael Montemerlo, Hendrik Dahlkamp, David Stavens, Andrei Aron, James Diebel, Philip W. Fong, John Gale, Morgan Halpenny, Gabriel Hoffmann, Kenny Lau, Celia M. Oakley, Mark Palatucci, Vaughan R. Pratt, Pascal Stang, Sven Strohband, Cedric Dupont, Lars-Erik J...

  135. [143]

    Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation, 2024

    Marcel Torne, Anthony Simeonov, Zechu Li, April Chan, Tao Chen, Abhishek Gupta, and Pulkit Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation, 2024

  136. [144]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need, 2023

  137. [145]

    Audio transformers: Transformer architectures for large scale audio understanding

    Prateek Verma and Jonathan Berger. Audio transformers: Transformer architectures for large scale audio understanding. adieu convolutions. ArXiv, abs/2105.00335, 2021

  138. [146]

    Apricot: Active preference learning and constraint-aware task planning with llms, 2024

    Huaxiaoyue Wang, Nathaniel Chin, Gonzalo Gonzalez-Pumariega, Xiangwan Sun, Neha Sunkara, Maximus Adrian Pace, Jeannette Bohg, and Sanjiban Choudhury. Apricot: Active preference learning and constraint-aware task planning with llms, 2024

  139. [147]

    Large language models for robotics: Opportunities, challenges, and perspectives, 2024

    Jiaqi Wang, Zihao Wu, Yiwei Li, Hanqi Jiang, Peng Shu, Enze Shi, Huawen Hu, Chong Ma, Yiheng Liu, Xuhui Wang, Yincheng Yao, Xuan Liu, Huaqin Zhao, Zhengliang Liu, Haixing Dai, Lin Zhao, Bao Ge, Xiang Li, Tianming Liu, and Shu Zhang. Large language models for robotics: Opportun...

  140. [148]

    Li, Madian Khabsa, Han Fang, and Hao Ma

    Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma. Linformer: Self- attention with linear complexity, 2020

  141. [149]

    Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models, 2023

    Tsun-Hsuan Wang, Alaa Maalouf, Wei Xiao, Yutong Ban, Alexander Amini, Guy Rosman, Sertac Karaman, and Daniela Rus. Drive anywhere: Generalizable end-to-end autonomous driving with multi-modal foundation models, 2023

  142. [150]

    A trajectory is worth three sentences: multimodal transformer for offline reinforcement learning

    Yiqi Wang, Mengdi Xu, Laixi Shi, and Yuejie Chi. A trajectory is worth three sentences: multimodal transformer for offline reinforcement learning. In Robin J. Evans and Ilya Shpitser, editors, Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence...

  143. [151]

    Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition, 2021

    Yulin Wang, Rui Huang, Shiji Song, Zeyi Huang, and Gao Huang. Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition, 2021

  144. [152]

    Challengesand solutions for autonomous ground robot scene understanding and navigation in unstructured outdoor environments: A review

    Liyana Wijayathunga, Alexander Rassau, and Douglas Chai. Challengesand solutions for autonomous ground robot scene understanding and navigation in unstructured outdoor environments: A review. Applied Sciences, 2023

  145. [153]

    Greedy hierarchical variational autoencoders for large-scale video prediction, 2021

    Bohan Wu, Suraj Nair, Roberto Martin-Martin, Li Fei-Fei, and Chelsea Finn. Greedy hierarchical variational autoencoders for large-scale video prediction, 2021

  146. [154]

    Tidybot: personalized robot assistance with large language models

    Jimmy Wu, Rika Antonova, Adam Kan, Marion Lepert, Andy Zeng, Shuran Song, Jean- nette Bohg, Szymon Rusinkiewicz, and Thomas Funkhouser. Tidybot: personalized robot assistance with large language models. Autonomous Robots, 47(8):1087–1102, November 2023

  147. [155]

    Transformers in medical image segmentation: A review

    Han Xiao, Li Li, Qi yu Liu, Xiuhong Zhu, and Qihang Zhang. Transformers in medical image segmentation: A review. Biomed. Signal Process. Control., 84:104791, 2023

  148. [156]

    Nystr¨ omformer: A nystr¨ om-based algorithm for approximating self- attention, 2021

    Yunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan, Glenn Fung, Yin Li, and Vikas Singh. Nystr¨ omformer: A nystr¨ om-based algorithm for approximating self- attention, 2021

  149. [157]

    A joint modeling of vision-language-action for target-oriented grasping in clutter, 2024

    Kechun Xu, Shuqi Zhao, Zhongxiang Zhou, Zizhang Li, Huaijin Pi, Yue Wang, and Rong Xiong. A joint modeling of vision-language-action for target-oriented grasping in clutter, 2024. 26

  150. [158]

    Peng Xu, Xiatian Zhu, and David A. Clifton. Multimodal learning with transformers: A survey, 2023

  151. [159]

    A review of recurrent neural networks: Lstm cells and network architectures

    Yong Yu, Xiaosheng Si, Changhua Hu, and Jian xun Zhang. A review of recurrent neural networks: Lstm cells and network architectures. Neural Computation, 31:1235–1270, 2019

  152. [160]

    Trans- former in reinforcement learning for decision-making: A survey

    Weilin Yuan, Jiaxing Chen, Shaofei Chen, Dawei Feng, Zhenzhen Hu, and Peng Li. Trans- former in reinforcement learning for decision-making: A survey. October 2023

  153. [161]

    Big bird: Transformers for longer sequences, 2021

    Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. Big bird: Transformers for longer sequences, 2021

  154. [162]

    Merlot: Multimodal neural script knowledge models, 2021

    Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. Merlot: Multimodal neural script knowledge models, 2021

  155. [163]

    Fanlong Zeng, Wensheng Gan, Yongheng Wang, Ning Liu, and Philip S. Yu. Large language models for robotics: A survey, 2023

  156. [164]

    Distilling and retrieving generalizable knowledge for robot manipulation via language corrections, 2024

    Lihan Zha, Yuchen Cui, Li-Heng Lin, Minae Kwon, Montserrat Gonzalez Arenas, Andy Zeng, Fei Xia, and Dorsa Sadigh. Distilling and retrieving generalizable knowledge for robot manipulation via language corrections, 2024

  157. [165]

    Hirt: Enhancing robotic control with hierarchical robot transformers, 2024

    Jianke Zhang, Yanjiang Guo, Xiaoyu Chen, Yen-Jen Wang, Yucheng Hu, Chengming Shi, and Jianyu Chen. Hirt: Enhancing robotic control with hierarchical robot transformers, 2024

  158. [166]

    Pointclip: Point cloud understanding by clip, 2021

    Renrui Zhang, Ziyu Guo, Wei Zhang, Kunchang Li, Xupeng Miao, Bin Cui, Yu Qiao, Peng Gao, and Hongsheng Li. Pointclip: Point cloud understanding by clip, 2021

  159. [167]

    Decision transformer as a foundation model for partially observable continuous control, 2024

    Xiangyuan Zhang, Weichao Mao, Haoran Qiu, and Tamer Ba¸ sar. Decision transformer as a foundation model for partially observable continuous control, 2024

  160. [168]

    Encoder-decoder models in sequence-to-sequence learning: A survey of rnn and lstm approaches

    Yunong Zhang. Encoder-decoder models in sequence-to-sequence learning: A survey of rnn and lstm approaches. Applied and Computational Engineering , 2023

  161. [169]

    Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid

    Tony Z. Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid. Aloha unleashed: A simple recipe for robot dexterity, 2024

  162. [170]

    Online decision transformer

    Qinqing Zheng, Amy Zhang, and Aditya Grover. Online decision transformer. In Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvari, Gang Niu, and Sivan Sabato, editors, Proceedings of the 39th International Conference on Machine Learning , volume 162 of Proceedings o...

  163. [171]

    Comparative study of sequence-to-sequence models: From rnns to trans- formers

    Jiancong Zhu. Comparative study of sequence-to-sequence models: From rnns to trans- formers. Applied and Computational Engineering , 2024

  164. [172]

    Transformermpc: Accelerating model predictive control via transformers, 2024

    Vrushabh Zinage, Ahmed Khalil, and Efstathios Bakolas. Transformermpc: Accelerating model predictive control via transformers, 2024

  165. [173]

    Atabay A. A. Ziyaden, Amir Yelenov, and Alexander Pak. Long-context transformers: A survey. 2021 5th Scientific School Dynamics of Complex Networks and their Applications (DCNA), pages 215–218, 2021. 27

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.