Pith. sign in

REVIEW 3 major objections 6 minor 300 references

Efficient Sensorimotor Learning for Open-world Robot Manipulation

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The dissertation argues that regularity in demonstrations is what makes data-efficient, generalizable robot manipulation possible.

desk verdict A well-written dissertation whose published systems are worth taking seriously, but whose central 'regularity causes efficiency' claim is a framing, not an experimentally isolated result. read the letter →

arxiv 2505.06136 v1 pith:MFE2KRC2 submitted 2025-05-07 cs.RO cs.AI

classification cs.ROcs.AI
keywords robotmanipulationopen-worldimitationlearningdataefficiencyregularityobject-centricrepresentationvideolifelong
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This dissertation tries to establish that data-efficient, generalizable robot manipulation is achievable when learning systems exploit the regular patterns already present in demonstrations. It names three such regularities—object, spatial, and behavioral—and argues that each enables a different form of learning: object-centric policies from a few teleoperation demos, video imitation from a single human demonstration, and continual learning by reusing discovered skills. A sympathetic reader would take the central claim to be that these regularities, not the particular network architectures, are the cause of the observed data efficiency. If that claim holds, it points toward personal robots that ordinary users can teach with a few demos or a video rather than large curated datasets.

What carries the argument

The machinery that carries the argument is the notion of regularity itself, divided into three named kinds: object regularity (semantics and function of objects persist across appearance and viewpoint), spatial regularity (task success is determined by invariant spatial relations between objects and manipulator), and behavioral regularity (manipulation decomposes into recurring primitive behaviors). Operationally, the machinery includes object-centric representations (region proposals, segmented point clouds), the Open-world Object Graph (a keyframe graph whose nodes are object point clouds plus a hand node and whose edges mark contact relations) used for video imitation, and hierarchical behavioral cloning with a skill library, with continual skill discovery so that past skills can be reused on new tasks. These are the mechanism by which the paper converts small demonstration sets into generalizable policies.

What would settle it

Train the same transformer policy on the same teleoperation demonstrations twice—once with object-centered point-cloud tokens and once with equal-sized patch tokens over the raw image—and evaluate generalization to new backgrounds, cameras, and object variants in simulation. If the gap between the two is not systematically in favor of the object-centered version across tasks, the causal role of object regularity is not supported.

Watch

Extended reading notes

Core claim

The dissertation's central claim is that a robot can learn generalizable, closed-loop manipulation policies from data quantities that would normally be considered far too small—tens of teleoperation demonstrations or a single human video—because physical demonstrations are dense with three reusable regularities: object regularity, spatial regularity, and behavioral regularity. It treats these not as properties of any algorithm but as inherent properties of the physical world, and it argues that the correct design move is to build neural policies that let these regularities do the work, e.g., object proposals and segmented point clouds for object regularity, keyframe-based object graphs for spatial regularity, and discovered skill libraries for behavioral regularity. If read sympathetically, the dissertation's seven method chapters are one extended demonstration of this principle.

Load-bearing premise

The load-bearing premise is that the measured data efficiency comes from the three regularities the dissertation names, not from the particular network architectures or foundation models the systems happen to use; no experiment holds architecture and data fixed while varying the regularity prior alone.

Editorial extensions

If this is right

  • From tens of space-mouse demonstrations, closed-loop visuomotor policies trained with object-centric priors can generalize to new object placements, backgrounds, camera angles, and unseen instances of familiar categories.
  • A robot with no task-specific action labels can imitate a manipulation skill from a single human video, because the task is represented as object-centric keyframe plans that capture invariant spatial relations.
  • The same spatial-regularity machinery transfers to humanoid robots with bimanual dexterous hands, substantially outperforming object-location-only retargeting.
  • By discovering reusable skills from past demonstrations, a robot can be trained on a sequence of tasks without catastrophic forgetting, improving average success over continuous learning.
  • A benchmark generated by procedural task generation allows these lifelong-learning claims to be evaluated quantitatively across many tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If regularity is the causal factor, then a task-independent measure of regularity—for instance, how consistently segmentation or keypoint tracking persists across demonstrations—could predict in advance how many demonstrations a new task needs; the dissertation does not construct such a measure.
  • The spatial-regularity framing suggests cross-embodiment transfer should succeed without teleoperation data on the target robot; one testable extension is training on human video and deploying on a mobile manipulator with different kinematics.
  • Associating discovered skills with language labels would let a user command a personal robot by naming a skill, turning the skill library into a spoken interface; this extends the behavioral-regularity idea to human-robot interaction.
  • The framework predicts that any method that injects the same priors into a larger or smaller backbone will retain its data efficiency, so the regularity framing could transfer to future foundation-model policies; this is an inference, not a claim in the dissertation.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This dissertation proposes a methodology for open-world robot manipulation organized around the notion of "regularity," defined as statistical regularities in demonstration data: object regularity, spatial regularity, and behavioral regularity. It presents seven systems: VIOLA and GROOT for object-centric imitation learning; ORION and OKAMI for imitation from a single human video; BUDS and LOTUS for continual skill discovery; and the LIBERO benchmark. The abstract claims that leveraging regularity is the key to data-efficient learning and generalization. Most technical chapters are drawn from peer-reviewed publications, with simulation and real-robot evaluations, ablation studies, and detailed appendices.

Significance. If the central claim were established, the dissertation would provide a unifying conceptual framework for data-efficient manipulation and a strong set of reusable systems. The individual systems are nontrivial, validated on real hardware, and supported by baseline comparisons and ablations. The appendices are unusually detailed, and the authors explicitly state limitations of the video-imitation setting. The weakness is that the advertised scientific principle, that regularity is the cause of efficiency, is never isolated experimentally; the chapters demonstrate that each system works, but not that the defined regularities are the operative variable. This gap matters because the abstract's causal claim is the dissertation's distinctive contribution beyond the individual published papers.

major comments (3)
  1. [Abstract; Section 2.5; Table 3.1] The load-bearing claim that "the key" to efficient sensorimotor learning lies in regularity is not supported by the experiments. No study holds architecture, data, and compute fixed while varying only whether the regularity prior is present. The closest evidence, VIOLA-Patch in Table 3.1, shows that removing the object-proposal prior does not consistently harm performance: on Stacking, VIOLA-Patch matches VIOLA in Canonical (71.2 vs 71.3) and exceeds it in Background-Change (41.4 vs 38.6), while only underperforming clearly on the long-horizon Kitchen task. Figure 4.5 likewise varies several design choices at once, so the gains cannot be attributed specifically to object regularity as defined in Section 2.5. Either an experiment that isolates a regularity prior, or a revision that explicitly reframes the abstract and Section 1 as proposing a perspective rather than a demonstrated causal mechanism, is needed.
  2. [Section 2.5.1–2.5.3] The definitions of the three regularities are co-extensive with the design choices of the methods that are supposed to exploit them: object regularity is instantiated by object proposals and segmentation, spatial regularity by keyframe plans and object graphs, and behavioral regularity by skill clustering. Because there is no independent measure of "regularity" and no condition in which the same architecture operates without the regularity prior, the attribution is not falsifiable. Section 2.5 itself acknowledges that the contribution is "a holistic perspective," not the proposition of the regularities. This is internally consistent, but it conflicts with the stronger causal language in the abstract. The manuscript should either add a controlled manipulation or consistently soften the causal claims.
  3. [Sections 5.1, 5.2.1, 6.1.1] The "open-world" claim for video imitation is substantially narrower than the term suggests. ORION requires an RGB-D video, a single human hand, tabletop scenes, and a user-provided list of English object descriptions (Section 5.2.1); OKAMI requires the upper body and both hands to be visible and a static camera (Section 6.1.1). Section 5.1 concedes that a solution to the full problem "is beyond the scope of our work or any existing work." The evaluations cover seven and six tasks, respectively. These are useful contributions, but the results should be presented as evidence for a restricted version of open-world imitation, not for the general problem stated in the introduction.
minor comments (6)
  1. [Section 2.2] In the sentence "we refer to the policies as sensorimotor policies or visuomotor policies interchangeability," the word should be "interchangeably."
  2. [Equation (2.3)] The indicator notation 1(i = k) is used before k is introduced and is never defined; please add a definition or rewrite the equation so that the role of k is clear.
  3. [Section 2.3.1] The definition st ≡ o≤t is followed by a stray "s" at the end of the displayed formula; please clean up the typesetting.
  4. [Figure 2.2] The formulas for FWT_m, NBT_m, and AUC_m are hard to read because overlines and subscripts are easily confused; please reformat and define r_m,m and r̄_i explicitly in the caption.
  5. [Section 1.2] The contribution list claims "the first end-to-end closed-loop neural network policy that can make coffee autonomously," but no evidence is provided for the "first" claim and the related work does not discuss coffee-making systems; please substantiate or soften this claim.
  6. [Chapters 5 and 6] Real-robot evaluations use 15 and 12 trials per task, respectively, without confidence intervals or statistical tests; given the small sample sizes, some reported differences may not be significant, so the presentation would be stronger with uncertainty quantification.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the systems are evaluated against external benchmarks and real-robot success criteria, and the regularity framing is explicitly an organizing perspective rather than an input from which results are derived.

full rationale

The dissertation does not contain a derivation in which a fitted quantity is renamed as a prediction, nor does it invoke a self-citation to force its central choice. Each methods chapter (VIOLA, GROOT, ORION, OKAMI, BUDS/LOTUS) reports task success rates against external baselines in simulation and on physical robots, with success determined by object/goal configurations independent of the model's own objective. The 'regularity' concept in Section 2.5 is explicitly disclaimed as a novel proposition: 'our contribution does not lie in the proposition of these regularities—they are inherent properties of the physical world, which the field has tapped into. Instead, we provide a holistic perspective.' The abstract's causal claim about regularity is not experimentally isolated (no condition varies only the regularity prior while holding architecture and data fixed), but that is a support weakness, not circularity. The self-citations to LIBERO and to the FWT/NBT/AUC metrics are evaluation infrastructure: the benchmark is a published community resource and the metrics are standard lifelong-learning measures, so they do not make the empirical claims equivalent to their own inputs. No specific equation or fitted parameter reduces to the paper's definitions, so no circular step can be quoted.

Assumptions & free parameters 4 free parameters · 6 assumptions · 2 invented entities

The ledger lists hyperparameters that affect the pipelines (Q, K, keyframe sensitivity, masking ratio), domain assumptions about the three regularities and about the reliability of foundation models, and the OOG and reference-plan representations introduced inside the methods. No free constant is derived in the paper; these are all empirical inputs.

free parameters (4)
  • Q: number of top object proposals in VIOLA = 20 (simulation; marginal recall increase 1% per 5 beyond 20)
    Selected in Section 3.2.2 based on recall saturation in Table 3.3; it determines the object-centric representation input to the transformer policy.
  • K: number of discovered skills in BUDS/LOTUS = Varies per task; evaluated with different K in Table 7.2
    Skill library size is set by clustering thresholds and must be tuned per dataset; performance depends on K (Table 7.2), and LOTUS relies on cross-task skill merging.
  • Keyframe detection sensitivity in ORION/OKAMI = Not reported numerically
    Standard changepoint detection (Section 5.2.3) has a sensitivity parameter that determines the number of keyframes in the manipulation plan; plan quality directly affects action synthesis (Equation 5.1).
  • Random masking ratio in GROOT = Not reported numerically
    Ablation in Figure 4.5 shows random masking is needed for camera-shift robustness; the ratio is a hyperparameter tuned per task.
assumptions (6)
  • standard math Contextual MDP formulation with sparse reward, universal transition dynamics, and surjective context mapping
    Section 2.1 models open-world manipulation as a CMDP where transition probabilities are independent of task context; this is a standard modeling assumption, not derived.
  • domain assumption Object regularity: object semantics and within-category functionality persist despite changes in appearance, lighting, background, and camera viewpoint
    Stated in Section 2.5.1 as the basis for VIOLA and GROOT; supported by examples, not by measurement.
  • domain assumption Spatial regularity: successful task completion is determined by spatial relations between task-relevant objects, and these relations are invariant across different manipulator embodiments
    Stated in Section 2.5.2 and used by ORION/OKAMI to transfer from human videos to robots; the invariance across embodiments is assumed.
  • domain assumption Behavioral regularity: long-horizon manipulation tasks decompose into recurring primitive skills that are shared across tasks
    Stated in Section 2.5.3; BUDS and LOTUS assume such recurring intervals exist and can be discovered from unsegmented demonstrations.
  • domain assumption Foundation models (Detic RPN, SAM, DINOv2, XMem, Cutie, CoTracker, HaMeR, GPT-4V) provide sufficiently reliable object localization, tracking, and semantics out of the box
    Every method in Parts I and II depends on these pretrained models; failure rates such as 'missed tracking' in ORION (Figure 5.4) show the assumption is only partially true.
  • domain assumption Task-relevant object descriptions (ORION) or VLM-generated object lists (OKAMI) are complete and correct
    ORION requires a human-provided comma-separated list of object descriptions (Section 5.2.1); OKAMI trusts GPT-4V output (Section 6.1.2). Errors in these lists would break plan generation.
invented entities (2)
  • Open-world Object Graph (OOG)
    purpose: Graph-based representation of a keyframe state in video imitation, containing object point clouds, a hand node, keypoint trajectories, and contact edges; used by ORION to generate a manipulation plan from a human video.
    Introduced in Section 5.2.2; evaluated only through task success inside the ORION pipeline, with no external falsifiable prediction attached to the representation itself.
  • Reference plan in OKAMI
    purpose: Sequence of steps, each containing target and reference object point clouds and an SMPL-H pose trajectory segment, used to guide object-aware retargeting onto a humanoid robot.
    Introduced in Section 6.1.2; its validity is internal to the OKAMI pipeline and no independent measurement is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficient Sensorimotor Learning for Open-world Robot Manipulation." pith.science (2026). https://pith.science/paper/MFE2KRC2

@misc{pith2026250506136,
  author       = {Pith},
  title        = {Pith review of: Efficient Sensorimotor Learning for Open-world Robot Manipulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MFE2KRC2}},
  note         = {Machine review of arXiv:2505.06136}
}
read the original abstract

This dissertation considers Open-world Robot Manipulation, a manipulation problem where a robot must generalize or quickly adapt to new objects, scenes, or tasks for which it has not been pre-programmed or pre-trained. This dissertation tackles the problem using a methodology of efficient sensorimotor learning. The key to enabling efficient sensorimotor learning lies in leveraging regular patterns that exist in limited amounts of demonstration data. These patterns, referred to as ``regularity,'' enable the data-efficient learning of generalizable manipulation skills. This dissertation offers a new perspective on formulating manipulation problems through the lens of regularity. Building upon this notion, we introduce three major contributions. First, we introduce methods that endow robots with object-centric priors, allowing them to learn generalizable, closed-loop sensorimotor policies from a small number of teleoperation demonstrations. Second, we introduce methods that constitute robots' spatial understanding, unlocking their ability to imitate manipulation skills from in-the-wild video observations. Last but not least, we introduce methods that enable robots to identify reusable skills from their past experiences, resulting in systems that can continually imitate multiple tasks in a sequential manner. Altogether, the contributions of this dissertation help lay the groundwork for building general-purpose personal robots that can quickly adapt to new situations or tasks with low-cost data collection and interact easily with humans. By enabling robots to learn and generalize from limited data, this dissertation takes a step toward realizing the vision of intelligent robotic assistants that can be seamlessly integrated into everyday scenarios.

Figures

Figures reproduced from arXiv: 2505.06136 by the authors.

Figure 1.1
Figure 1.1. Overview of the chapter dependencies. An arrow connection means that [PITH_FULL_IMAGE:figures/full_fig_p023_1_1.png] view at source ↗
Figure 2.1
Figure 2.1. Intra-task generalization involves four dimensions of variations: background [PITH_FULL_IMAGE:figures/full_fig_p028_2_1.png] view at source ↗
Figure 2.2
Figure 2.2. This figure visualizes the three lifelong long metrics evaluated over a task [PITH_FULL_IMAGE:figures/full_fig_p031_2_2.png] view at source ↗
Figures from the paper (43 more)
Figure 2.3
Figure 2.3. Figure 2.3: We show the devices used for collecting demonstrations through either [PITH_FULL_IMAGE:figures/full_fig_p036_2_3.png]
Figure 2.4
Figure 2.4. Figure 2.4: We use an example to illustrate object regularity. Objects of interest from the left scene appear and behave regularly despite visual variations, no matter how the background, lighting, and camera angle change. 2.5.2 Spatial Regularity Spatial understanding of the wo…
Figure 2.5
Figure 2.5. Figure 2.5: We use an example to illustrate spatial regularity. In this example, we consider a task goal of having a coffee mug on top of a table mat. Despite variations in object locations or different robot embodiments, the task goal is achieved as long as the spatial relation…
Figure 2.6
Figure 2.6. Figure 2.6: We use an example to illustrate behavioral regularity. The figure shows three different tasks. Behaviors in each task can be decomposed into primitives, abstracted in the central diagram. Squares of the same color refer to a recurring primitive. ( corresponds to a di…
Figure 2.7
Figure 2.7. Figure 2.7: This figure shows joint configurations, task space commands, and the base [PITH_FULL_IMAGE:figures/full_fig_p044_2_7.png]
Figure 2.8
Figure 2.8. Figure 2.8: This figure shows joint configurations, task space commands, and the base [PITH_FULL_IMAGE:figures/full_fig_p044_2_8.png]
Figure 3.1
Figure 3.1. Figure 3.1: VIOLA Overview. VIOLA first obtains a set of general object proposals from raw visual observations. It extracts object features from the proposals to build the object-centric representation. The transformer-based policy uses multi-head self￾attention to reason over t…
Figure 3.2
Figure 3.2. Figure 3.2: VIOLA Model Architecture. At time t, VIOLA computes the per-step features ht using the top Q object proposals. Then, it constructs the object-centric representation zt by composing per-step features from the last ∆tH + 1 time-step observations along with their tempor…
Figure 3.3
Figure 3.3. Figure 3.3: Visualization of the initial and goal configurations for real-world tasks. [PITH_FULL_IMAGE:figures/full_fig_p054_3_3.png]
Figure 3.4
Figure 3.4. Figure 3.4: Success rates (%) in real robot tasks. The quantitative evaluation in Fig￾ure 3.4 shows that VIOLA outperforms BC￾RNN by 46.7% success rate on average. Qual￾itatively, we observe that the VIOLA policy can robustly grasp K-cups or open the coffee machine in the Make-C…
Figure 3.5
Figure 3.5. Figure 3.5: Visualization of top-3 regions weighted most by transformer attention. [PITH_FULL_IMAGE:figures/full_fig_p060_3_5.png]
Figure 4.1
Figure 4.1. Figure 4.1: GROOT overview. GROOT learns closed-loop visuomotor policies from demonstrations under a single setup, and generalizes to new setups with unseen condi￾tions, namely different visual distractions, changed camera angles, and new objects. leveraging the 3D-aware propert…
Figure 4.2
Figure 4.2. Figure 4.2: GROOT Model Architecture. GROOT leverages an interactive seg￾mentation model, S2M, to obtain a single-frame annotation from demonstrators. Then a Video Object Segmentation model, XMem, propagates segmentation masks across time frames. The object masks are then back-p…
Figure 4.3
Figure 4.3. Figure 4.3: Visualization of simulation tasks, for Canonical, Background(Easy), Background(Hard), Camera(Easy), and Camera(Hard). 4.2.1 Experimental Setup We use both simulation and real-robot experiments to evaluate GROOT poli￾cies. Task Designs. We use three simulation tasks b…
Figure 4.4
Figure 4.4. Figure 4.4: Visualization of objects used in real-robot experiments. In each image, the [PITH_FULL_IMAGE:figures/full_fig_p071_4_4.png]
Figure 4.5
Figure 4.5. Figure 4.5: Changes in success rates (%) with design choices for GROOT. Pick Place Cup Stamp The Paper Take the Mug Put The Mug On The Coaster Roll The Stamp 0.0 0.2 0.4 0.6 0.8 1.0 Success Rates Canonical Camera-Shift Background-Change New-Object [PITH_FULL_IMAGE:figures/full_…
Figure 5.1
Figure 5.1. Figure 5.1: Overview of ORION. ORION tackles the problem of imitating manip￾ulation from single human video demonstrations. ORION first extracts a sequence of Open-World Object Graphs (OOGs), where each OOG models a keyframe state with task-relevant objects and hand information.…
Figure 5.2
Figure 5.2. Figure 5.2: Plan Generation. ORION generates a manipulation plan from a given video V for subsequent action synthesis. ORION first tracks the objects and keypoints across the video frames. Then, keyframes are identified based on the velocity statistics of the keypoint trajectori…
Figure 5.3
Figure 5.3. Figure 5.3: Action Synthesis. ORION first localizes task-relevant objects at test time and retrieves the matched OOG from the generated manipulation plan. Then ORION uses the retrieved OOGs to predict the object motions by first computing global registration of object point clou…
Figure 5.4
Figure 5.4. Figure 5.4: Visualization of evaluation tasks. Each block corresponds to one task. The [PITH_FULL_IMAGE:figures/full_fig_p089_5_4.png]
Figure 5.5
Figure 5.5. Figure 5.5: Experimental Evaluation of ORION Policies. (a) Experimental com￾parison between ORION and the two baselines, namely Hand-Motion-Imitation and Dense-Correspondence. (b) Ablation study on using videos of the same task recorded in three different settings. We select the…
Figure 5.6
Figure 5.6. Figure 5.6: (a) Visualization of initial and final frames of the three videos of the [PITH_FULL_IMAGE:figures/full_fig_p093_5_6.png]
Figure 6.1
Figure 6.1. Figure 6.1: This chapter focuses on enabling a human user to teach the humanoid robot [PITH_FULL_IMAGE:figures/full_fig_p096_6_1.png]
Figure 6.2
Figure 6.2. Figure 6.2: OKAMI Model Overview. OKAMI is a two-staged method that enables a humanoid robot to imitate a manipulation task from a single human video. In the first stage, OKAMI generates a reference plan using GPT-4V and large vision models for subsequent manipulation. In the se…
Figure 6.3
Figure 6.3. Figure 6.3: Visualization of initial and final frames of both human demonstrations and [PITH_FULL_IMAGE:figures/full_fig_p102_6_3.png]
Figure 6.4
Figure 6.4. Figure 6.4: Experimental Evaluation of OKAMI Policies. (a) Evaluation of OKAMI over all six tasks, including the success rates and the quantification of failed trials, separated by failure mode. (b) Evaluation of OKAMI using videos from different demonstrations. Demonstrator 1 i…
Figure 6.5
Figure 6.5. Figure 6.5: The initial and end frames of videos performed by different human demon [PITH_FULL_IMAGE:figures/full_fig_p106_6_5.png]
Figure 6.6
Figure 6.6. Figure 6.6: Success rates (%) of learned visuomotor policies on Sprinkle-salt and Bagging using 50 and 100 trajec￾tories, respectively. KL weight 10 chunk size 60 hidden dimension 512 batch size 45 feedforward dimension 3200 epochs 25000 learning rate 5e-5 temporal weighting 0.0…
Figure 7.1
Figure 7.1. Figure 7.1: BUDS Overview. BUDS constructs hierarchical task structures of demonstration sequences in a bottom-up manner, from which mid-level temporal seg￾ments are discovered for discovering and learning sensorimotor skills. tion of demonstration segments [127], our approach d…
Figure 7.2
Figure 7.2. Figure 7.2: Hierarchical Visuomotor Policy in BUDS. Given an observation image of the workspace, the meta-controller selects the skill index and generates the latent subgoal vector ωt . Then, the selected sensorimotor skill generates action at (end-effector displacements and gri…
Figure 7.3
Figure 7.3. Figure 7.3: Visualization of the four simulation tasks and one real robot task used in our [PITH_FULL_IMAGE:figures/full_fig_p119_7_3.png]
Figure 7.4
Figure 7.4. Figure 7.4: BUDS skill segmentation visualization. achieving a 56% success rate. The performance is on par with the performance of our simulation evaluations, showing that BUDS generalizes well to real-world data and physical hardware. We also evaluate the most competitive basel…
Figure 7.5
Figure 7.5. Figure 7.5: We visualize the percentage of skills in each task. We also show three [PITH_FULL_IMAGE:figures/full_fig_p125_7_5.png]
Figure 8.1
Figure 8.1. Figure 8.1: LOTUS Overview. LOTUS is a lifelong imitation learning algorithm through continual skill discovery. LOTUS starts from the base task stage, where it builds an initial library of sensorimotor skills. In the subsequent lifelong task stage, LOTUS continuously discovers n…
Figure 8.2
Figure 8.2. Figure 8.2: Method Overview. LOTUS consists of two processes: continual skill discovery with open-world perception and hierarchical policy learning with the skill library. For continual skill discovery, we obtain temporal segments from demonstrations using hierarchical clusterin…
Figure 8.3
Figure 8.3. Figure 8.3: Visualization of skill discovery results in [PITH_FULL_IMAGE:figures/full_fig_p138_8_3.png]
Figure 9.1
Figure 9.1. Figure 9.1: Top: Libero has four procedurally-generated task suites: LIBERO￾Spatial, LIBERO-Object, and LIBERO-Goal have 10 tasks each and require transferring knowledge about spatial relationships, objects, and task goals; LIBERO￾100 has 100 tasks and requires the transfer of e…
Figure 9.2
Figure 9.2. Figure 9.2: Libero’s procedural generation pipeline: Extracting behavioral templates from a large-scale human activity dataset (1), Ego4D, for generating task instruc￾tions (2); Based on the task description, selecting the scene and generating the PDDL description file (3) that …
Figure 9.3
Figure 9.3. Figure 9.3: LIBERO-Spatial [PITH_FULL_IMAGE:figures/full_fig_p148_9_3.png]
Figure 9.4
Figure 9.4. Figure 9.4: LIBERO-Object [PITH_FULL_IMAGE:figures/full_fig_p148_9_4.png]
Figure 9.5
Figure 9.5. Figure 9.5: LIBERO-Goal. tasks in simulation. We use this pipeline to create 130 standardized tasks and conduct a comprehensive set of experiments on policy and algorithm designs. This benchmark 134 [PITH_FULL_IMAGE:figures/full_fig_p148_9_5.png]
Figure 9.6
Figure 9.6. Figure 9.6: LIBERO-100. serves as a testbed for designing algorithms that exploit behavioral regularity, especially in the context of lifelong robot learning. Beyond its contribution to this dissertation, our Libero benchmark is designed to support multiple general research dire…
Figure 11.1
Figure 11.1. Figure 11.1: Human-Robot Coevolution presents a long-term research theme based on [PITH_FULL_IMAGE:figures/full_fig_p170_11_1.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 8 canonical work pages

  1. [1]

    https://en.wikipedia.org/wiki/Unimate, 2024

    Unimate. https://en.wikipedia.org/wiki/Unimate, 2024. Accessed: 2024-01-29

  2. [2]

    https://en.wikipedia.org/wiki/ENIAC, 2024

    Eniac. https://en.wikipedia.org/wiki/ENIAC, 2024. Accessed: 2024-04-02

  3. [3]

    Do as i can, not as i say: Grounding language in robotic af- fordances

    Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic af- fordances. arXiv preprint arXiv:2204.01691 , 2022

  4. [4]

    Code as policies: Language model programs for embodied control

    Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 9493–9500. IEEE, 2023

  5. [5]

    Moka: Open- vocabulary robotic manipulation through mark-based visual prompting

    Fangchen Liu, Kuan Fang, Pieter Abbeel, and Sergey Levine. Moka: Open- vocabulary robotic manipulation through mark-based visual prompting. arXiv preprint arXiv:2403.03174, 2024

  6. [6]

    Pivot: It- erative visual prompting elicits actionable knowledge for vlms

    Soroush Nasiriany, Fei Xia, Wenhao Yu, Ted Xiao, Jacky Liang, Ishita Das- gupta, Annie Xie, Danny Driess, Ayzaan Wahid, Zhuo Xu, et al. Pivot: It- erative visual prompting elicits actionable knowledge for vlms. arXiv preprint arXiv:2402.07872, 2024

  7. [7]

    204 Open x-embodiment: Robotic learning datasets and rt-x models

    Abhishek Padalkar, Acorn Pooley, Ajinkya Jain, Alex Bewley, Alex Herzog, Alex Irpan, Alexander Khazatsky, Anant Rai, Anikait Singh, Anthony Brohan, et al. 204 Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023

  8. [8]

    Droid: A large-scale in-the-wild robot manipulation dataset

    Alexander Khazatsky, Karl Pertsch, Suraj Nair, Ashwin Balakrishna, Sudeep Dasari, Siddharth Karamcheti, Soroush Nasiriany, Mohan Kumar Srirama, Lawrence Yunliang Chen, Kirsty Ellis, et al. Droid: A large-scale in-the-wild robot manipulation dataset. arXiv preprint arXiv:2403.12945 , 2024

Show all 300 references
  1. [9]

    Event structure in perception and concep- tion

    Jeffrey M Zacks and Barbara Tversky. Event structure in perception and concep- tion. Psychological bulletin, 127(1):3, 2001

  2. [10]

    On the binding problem in artificial neural networks

    Klaus Greff, Sjoerd Van Steenkiste, and J¨ urgen Schmidhuber. On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208 , 2020

  3. [11]

    Building machines that learn and think like people

    Brenden M Lake, Tomer D Ullman, Joshua B Tenenbaum, and Samuel J Gersh- man. Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017

  4. [12]

    Core knowledge

    Elizabeth S Spelke and Katherine D Kinzler. Core knowledge. Developmental science, 10(1):89–96, 2007

  5. [13]

    Motion perception and the scene statistics of motion

    Tal Tversky. Motion perception and the scene statistics of motion. The University of Texas at Austin, 2008

  6. [14]

    On the opportunities and risks of foundation models

    Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258, 2021

  7. [15]

    Training language models to follow instructions with human feedback

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems , 35:277...

  8. [16]

    Segment anything

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. arXiv preprint arXiv:2304.02643 , 2023

  9. [17]

    Dinov2: Learning robust visual features without supervision

    Maxime Oquab, Timoth´ ee Darcet, Th´ eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023

  10. [18]

    Markov decision processes

    Martin L Puterman. Markov decision processes. Handbooks in operations research and management science , 2:331–434, 1990

  11. [19]

    Gti: Learning to generalize across long-horizon tasks from human demon- strations

    Ajay Mandlekar, Danfei Xu, Roberto Martın-Martın, Silvio Savarese, and Li Fei- Fei. Gti: Learning to generalize across long-horizon tasks from human demon- strations. In RSS, 2020

  12. [20]

    Generalization through hand-eye coordination: An action space for learning spatially-invariant visuomotor control

    Chen Wang, Rui Wang, Ajay Mandlekar, Li Fei-Fei, Silvio Savarese, and Danfei Xu. Generalization through hand-eye coordination: An action space for learning spatially-invariant visuomotor control. In 2021 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems (IR...

  13. [21]

    Contextual markov decision processes

    Assaf Hallak, Dotan Di Castro, and Shie Mannor. Contextual markov decision processes. arXiv preprint arXiv:1502.02259 , 2015

  14. [22]

    Mutex: Learning unified policies from multimodal task specifications

    Rutav Shah, Roberto Mart´ ın-Mart´ ın, and Yuke Zhu. Mutex: Learning unified policies from multimodal task specifications. In 7th Annual Conference on Robot Learning, 2023

  15. [23]

    Gradient episodic memory for con- tinual learning

    David Lopez-Paz and Marc’Aurelio Ranzato. Gradient episodic memory for con- tinual learning. Advances in neural information processing systems , 30, 2017

  16. [24]

    Don’t forget, there is more than forgetting: new metrics for continual learning

    Natalia D´ ıaz-Rodr´ ıguez, Vincenzo Lomonaco, David Filliat, and Davide Maltoni. Don’t forget, there is more than forgetting: new metrics for continual learning. arXiv preprint arXiv:1810.13166 , 2018. 206

  17. [25]

    Libero: Benchmarking knowledge transfer for lifelong robot learning

    Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning. arXiv preprint arXiv:2306.03310, 2023

  18. [26]

    Overcoming catastrophic forgetting in neural networks

    James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al. Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of scien...

  19. [27]

    Continual learning and catastrophic forgetting

    Gido M van de Ven, Nicholas Soures, and Dhireesha Kudithipudi. Continual learning and catastrophic forgetting. arXiv preprint arXiv:2403.05175 , 2024

  20. [28]

    Deep recurrent q-learning for partially observable mdps

    Matthew Hausknecht and Peter Stone. Deep recurrent q-learning for partially observable mdps. In 2015 aaai fall symposium series , 2015

  21. [29]

    Gaussian mixture models

    Douglas A Reynolds et al. Gaussian mixture models. Encyclopedia of biometrics, 741(659-663), 2009

  22. [30]

    What matters in learning from offline human demonstrations for robot manipu- lation

    Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Ro- hun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Mart´ ın-Mart´ ın. What matters in learning from offline human demonstrations for robot manipu- lation. arXiv preprint arXiv:2108.03298 , 2021

  23. [31]

    Hierarchical imitation and reinforcement learning

    Hoang Le, Nan Jiang, Alekh Agarwal, Miroslav Dud´ ık, Yisong Yue, and Hal Daum´ e. Hierarchical imitation and reinforcement learning. InICML, pages 2917– 2926, 2018

  24. [32]

    Factor graphs for robot perception

    Frank Dellaert, Michael Kaess, et al. Factor graphs for robot perception. Foun- dations and Trends® in Robotics, 6(1-2):1–139, 2017

  25. [33]

    Imagenet: A large-scale hierarchical image database

    Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 207

  26. [34]

    Processing of representations in declarative and procedural working memory

    Alessandrada Silva Souza, Klaus Oberauer, Miriam Gade, and Michel D Druey. Processing of representations in declarative and procedural working memory. Quarterly Journal of Experimental Psychology , 65(5):1006–1033, 2012

  27. [35]

    A unified approach for motion and force control of robot ma- nipulators: The operational space formulation

    Oussama Khatib. A unified approach for motion and force control of robot ma- nipulators: The operational space formulation. IEEE Journal on Robotics and Automation, 3(1):43–53, 1987

  28. [36]

    Pink: Python inverse kinematics based on Pinocchio, 2024

    St´ ephane Caron, Yann De Mont-Marin, Rohan Budhiraja, Seung Hyeon Bang, Ivan Domrachev, and Simeon Nedelchev. Pink: Python inverse kinematics based on Pinocchio, 2024. URL https://github.com/stephane-caron/pink

  29. [37]

    Mink: Python inverse kinematics based on MuJoCo, July 2024

    Kevin Zakka. Mink: Python inverse kinematics based on MuJoCo, July 2024. URL https://github.com/kevinzakka/mink

  30. [38]

    Deoxys: A modular, real-time controller library for robot learn- ing

    Yifeng Zhu. Deoxys: A modular, real-time controller library for robot learn- ing. https://github.com/UT-Austin-RPL/deoxys_control, 12 2022. Accessed: 2022-12-31

  31. [39]

    Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data

    Ajay Mandlekar, Fabio Ramos, Byron Boots, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Dieter Fox. Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data. In2020 IEEE International Conference on Robotics and Automa...

  32. [40]

    Deep imitation learning for complex manipulation tasks from virtual reality teleoperation

    Tianhao Zhang, Zoe McCarthy, Owen Jow, Dennis Lee, Xi Chen, Ken Goldberg, and Pieter Abbeel. Deep imitation learning for complex manipulation tasks from virtual reality teleoperation. In 2018 IEEE International Conference on Robotics and Automation (ICRA) , pages 5628–5635. IEEE, 2018

  33. [41]

    A reduction of imitation learning and structured prediction to no-regret online learning

    St´ ephane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 627–635. JMLR Workshop and ...

  34. [42]

    Object-aware regularization for addressing causal confusion in imitation learning

    Jongjin Park, Younggyo Seo, Chang Liu, Li Zhao, Tao Qin, Jinwoo Shin, and Tie- Yan Liu. Object-aware regularization for addressing causal confusion in imitation learning. Advances in Neural Information Processing Systems , 34, 2021

  35. [43]

    Causal confusion in imita- tion learning

    Pim De Haan, Dinesh Jayaraman, and Sergey Levine. Causal confusion in imita- tion learning. Advances in Neural Information Processing Systems , 32, 2019

  36. [44]

    Fight- ing copycat agents in behavioral cloning from observation histories

    Chuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman, and Yang Gao. Fight- ing copycat agents in behavioral cloning from observation histories. Advances in Neural Information Processing Systems, 33:2564–2575, 2020

  37. [45]

    Exploring the limitations of behavior cloning for autonomous driving

    Felipe Codevilla, Eder Santana, Antonio M L´ opez, and Adrien Gaidon. Exploring the limitations of behavior cloning for autonomous driving. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 9329–9338, 2019

  38. [46]

    Query-efficient imitation learning for end-to- end autonomous driving

    Jiakai Zhang and Kyunghyun Cho. Query-efficient imitation learning for end-to- end autonomous driving. arXiv preprint arXiv:1605.06450 , 2016

  39. [47]

    Imitation learning via off-policy distribution matching

    Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson. Imitation learning via off-policy distribution matching. arXiv preprint arXiv:1912.05032 , 2019

  40. [48]

    Self-supervised correspon- dence in visuomotor policy learning

    Peter Florence, Lucas Manuelli, and Russ Tedrake. Self-supervised correspon- dence in visuomotor policy learning. IEEE Robotics and Automation Letters , 5 (2):492–499, 2019

  41. [49]

    Graph-structured visual imitation

    Maximilian Sieb, Zhou Xian, Audrey Huang, Oliver Kroemer, and Katerina Fragkiadaki. Graph-structured visual imitation. In Conference on Robot Learn- ing, pages 979–989, 2020

  42. [50]

    Mask r-cnn

    Kaiming He, Georgia Gkioxari, Piotr Doll´ ar, and Ross Girshick. Mask r-cnn. In Proceedings of the IEEE international conference on computer vision , pages 2961–2969, 2017. 209

  43. [51]

    Faster r-cnn: Towards real-time object detection with region proposal networks

    Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in neural information processing systems, 28, 2015

  44. [52]

    Cascade r-cnn: Delving into high quality object detection

    Zhaowei Cai and Nuno Vasconcelos. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6154–6162, 2018

  45. [53]

    Objects as points.arXiv preprint arXiv:1904.07850, 2019

    Xingyi Zhou, Dequan Wang, and Philipp Kr¨ ahenb¨ uhl. Objects as points.arXiv preprint arXiv:1904.07850, 2019

  46. [54]

    Detecting twenty-thousand classes using image-level supervision

    Xingyi Zhou, Rohit Girdhar, Armand Joulin, Phillip Kr¨ ahenb¨ uhl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. arXiv preprint arXiv:2201.02605, 2022

  47. [55]

    Viola: Imitation learn- ing for vision-based manipulation with object proposal priors

    Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu. Viola: Imitation learn- ing for vision-based manipulation with object proposal priors. arXiv preprint arXiv:2210.11339, 2022

  48. [56]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems , 30, 2017

  49. [57]

    Deep residual learning for image recognition

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016

  50. [58]

    End-to-end training of deep visuomotor policies

    Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. End-to-end training of deep visuomotor policies. Journal of Machine Learning Research , 17 (39):1–40, 2016

  51. [59]

    Layer normalization

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. Layer normalization. arXiv preprint arXiv:1607.06450 , 2016. 210

  52. [60]

    Bert: Pre- training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre- training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018

  53. [61]

    Image augmentation is all you need: Regularizing deep reinforcement learning from pixels

    Denis Yarats, Ilya Kostrikov, and Rob Fergus. Image augmentation is all you need: Regularizing deep reinforcement learning from pixels. In International conference on learning representations, 2021

  54. [62]

    Critic regularized regression

    Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Sprin- genberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, et al. Critic regularized regression. Advances in Neural Information Pro- cessing Systems, 33:7768–7778, 2020

  55. [63]

    robosuite: A modu- lar simulation framework and benchmark for robot learning

    Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Mart´ ın-Mart´ ın, Abhishek Joshi, Kevin Lin, Soroush Nasiriany, and Yifeng Zhu. robosuite: A modu- lar simulation framework and benchmark for robot learning. In arXiv preprint arXiv:2009.12293, 2020

  56. [64]

    Bag of tricks for image classification with convolutional neural networks

    Tong He, Zhi Zhang, Hang Zhang, Zhongyue Zhang, Junyuan Xie, and Mu Li. Bag of tricks for image classification with convolutional neural networks. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 558–567, 2019

  57. [65]

    Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation

    Rishabh Jangir, Nicklas Hansen, Sambaran Ghosal, Mohit Jain, and Xiaolong Wang. Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation. IEEE Robotics and Automation Letters , 2022

  58. [66]

    Neural discrete representation learning

    Aaron Van Den Oord, Oriol Vinyals, et al. Neural discrete representation learning. Advances in neural information processing systems , 30, 2017

  59. [67]

    Learning to play by imitating humans

    Rostam Dinyari, Pierre Sermanet, and Corey Lynch. Learning to play by imitating humans. arXiv preprint arXiv:2006.06874 , 2020. 211

  60. [68]

    Vision-based multi-task manipulation for inexpensive robots using end- to-end learning from demonstration

    Rouhollah Rahmatizadeh, Pooya Abolghasemi, Ladislau B¨ ol¨ oni, and Sergey Levine. Vision-based multi-task manipulation for inexpensive robots using end- to-end learning from demonstration. In 2018 IEEE international conference on robotics and automation (ICRA) , pages 3758–37...

  61. [69]

    An image is worth 16x16 words: Transformers for image recognition at scale

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arX...

  62. [70]

    Deep spatial autoencoders for visuomotor learning

    Chelsea Finn, Xin Yu Tan, Yan Duan, Trevor Darrell, Sergey Levine, and Pieter Abbeel. Deep spatial autoencoders for visuomotor learning. In 2016 IEEE Inter- national Conference on Robotics and Automation (ICRA) , pages 512–519. IEEE, 2016

  63. [71]

    Random erasing data augmentation

    Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. Random erasing data augmentation. In Proceedings of the AAAI conference on artificial intelligence, volume 34, pages 13001–13008, 2020

  64. [72]

    Learning generalizable manipulation policies with object-centric 3d representations

    Yifeng Zhu, Zhenyu Jiang, Peter Stone, and Yuke Zhu. Learning generalizable manipulation policies with object-centric 3d representations. In 7th Annual Con- ference on Robot Learning, 2023

  65. [73]

    Object-centric learning with slot attention

    Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahen- dran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, and Thomas Kipf. Object-centric learning with slot attention. Advances in Neural Information Pro- cessing Systems, 33:11525–11538, 2020

  66. [74]

    Modular interactive video ob- ject segmentation: Interaction-to-mask, propagation and difference-aware fusion

    Ho Kei Cheng, Yu-Wing Tai, and Chi-Keung Tang. Modular interactive video ob- ject segmentation: Interaction-to-mask, propagation and difference-aware fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5559–5568, 2021. 212

  67. [75]

    Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model

    Ho Kei Cheng and Alexander G Schwing. Xmem: Long-term video object seg- mentation with an atkinson-shiffrin memory model. In European Conference on Computer Vision, pages 640–658. Springer, 2022

  68. [76]

    Open3d: A modern library for 3d data processing

    Qian-Yi Zhou, Jaesik Park, and Vladlen Koltun. Open3d: A modern library for 3d data processing. arXiv preprint arXiv:1801.09847 , 2018

  69. [77]

    Modern robotics

    Kevin M Lynch and Frank C Park. Modern robotics. Cambridge University Press, 2017

  70. [78]

    Masked autoencoders for point cloud self-supervised learning

    Yatian Pang, Wenxiao Wang, Francis EH Tay, Wei Liu, Yonghong Tian, and Li Yuan. Masked autoencoders for point cloud self-supervised learning. In Com- puter Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part II , pages 604–621. ...

  71. [79]

    Pointnet++: Deep hierarchical feature learning on point sets in a metric space

    Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J Guibas. Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Advances in neural information processing systems, 30, 2017

  72. [80]

    Multi-view masked world models for visual robotic manipulation

    Younggyo Seo, Junsu Kim, Stephen James, Kimin Lee, Jinwoo Shin, and Pieter Abbeel. Multi-view masked world models for visual robotic manipulation. arXiv preprint arXiv:2302.02408, 2023

  73. [81]

    Vision- based manipulators need to also see from their hands

    Kyle Hsu, Moo Jin Kim, Rafael Rafailov, Jiajun Wu, and Chelsea Finn. Vision- based manipulators need to also see from their hands. InInternational Conference on Learning Representations, 2022

  74. [82]

    Learning generalizable robotic reward functions from” in-the-wild” human videos

    Annie S Chen, Suraj Nair, and Chelsea Finn. Learning generalizable robotic reward functions from” in-the-wild” human videos. arXiv preprint arXiv:2103.16817, 2021

  75. [83]

    R3m: A universal visual representation for robot manipulation

    Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:2203.12601, 2022. 213

  76. [84]

    Vip: Towards universal visual reward and representation via value-implicit pre-training

    Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training. arXiv preprint arXiv:2210.00030 , 2022

  77. [85]

    Mimicplay: Long-horizon imitation learning by watching human play

    Chen Wang, Linxi Fan, Jiankai Sun, Ruohan Zhang, Li Fei-Fei, Danfei Xu, Yuke Zhu, and Anima Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play. In Conference on Robot Learning, pages 201–221. PMLR, 2023

  78. [86]

    Learning by watching: Physical imitation of manipu- lation skills from human videos

    Haoyu Xiong, Quanzhou Li, Yun-Chun Chen, Homanga Bharadhwaj, Samarth Sinha, and Animesh Garg. Learning by watching: Physical imitation of manipu- lation skills from human videos. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 7827–78...

  79. [87]

    Imitation Learning from Observation

    Faraz Torabi. Imitation Learning from Observation . PhD thesis, University of Texas at Austin, 2021. PhD Thesis

  80. [88]

    Vision-based manipula- tion from single human video with open-world object graphs

    Yifeng Zhu, Arisrei Lim, Peter Stone, and Yuke Zhu. Vision-based manipula- tion from single human video with open-world object graphs. arXiv preprint arXiv:2405.20321, 2024

  81. [89]

    Grounding dino: Marry- ing dino with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, et al. Grounding dino: Marry- ing dino with grounded pre-training for open-set object detection. In European Conference on Computer Vision , pages 38–55. Sprin...

  82. [90]

    Putting the object back into video object segmentation

    Ho Kei Cheng, Seoung Wug Oh, Brian Price, Joon-Young Lee, and Alexander Schwing. Putting the object back into video object segmentation. arXiv preprint arXiv:2310.12982, 2023

  83. [91]

    Cotracker: It is better to track together

    Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Cotracker: It is better to track together. In European Conference on Computer Vision , pages 18–35. Springer, 2024. 214

  84. [92]

    Robotap: Tracking ar- bitrary points for few-shot visual imitation

    Mel Vecerik, Carl Doersch, Yi Yang, Todor Davchev, Yusuf Aytar, Guangyao Zhou, Raia Hadsell, Lourdes Agapito, and Jon Scholz. Robotap: Tracking ar- bitrary points for few-shot visual imitation. arXiv preprint arXiv:2308.15975 , 2023

  85. [93]

    You only demon- strate once: Category-level manipulation from single visual demonstration

    Bowen Wen, Wenzhao Lian, Kostas Bekris, and Stefan Schaal. You only demon- strate once: Category-level manipulation from single visual demonstration. arXiv preprint arXiv:2201.12716, 2022

  86. [94]

    Optimal detection of change- points with a linear computational cost

    Rebecca Killick, Paul Fearnhead, and Idris A Eckley. Optimal detection of change- points with a linear computational cost. Journal of the American Statistical As- sociation, 107(500):1590–1598, 2012

  87. [95]

    Reconstructing hands in 3d with transformers.arXiv preprint arXiv:2312.05251, 2023

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3d with transformers.arXiv preprint arXiv:2312.05251, 2023

  88. [96]

    Robust reconstruction of indoor scenes

    Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun. Robust reconstruction of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5556–5565, 2015

  89. [97]

    Zero-shot robot manipulation from passive human videos

    Homanga Bharadhwaj, Abhinav Gupta, Shubham Tulsiani, and Vikash Ku- mar. Zero-shot robot manipulation from passive human videos. arXiv preprint arXiv:2302.02011, 2023

  90. [98]

    Learning to act from actionless videos through dense correspondences

    Po-Chen Ko, Jiayuan Mao, Yilun Du, Shao-Hua Sun, and Joshua B Tenenbaum. Learning to act from actionless videos through dense correspondences. arXiv preprint arXiv:2310.08576, 2023

  91. [99]

    Ditto: Demonstration imitation by trajectory transformation

    Nick Heppert, Max Argus, Tim Welschehold, Thomas Brox, and Abhinav Valada. Ditto: Demonstration imitation by trajectory transformation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 7565–

  92. [100]

    Robust plan- ning for multi-stage forceful manipulation

    Rachel Holladay, Tom´ as Lozano-P´ erez, and Alberto Rodriguez. Robust plan- ning for multi-stage forceful manipulation. The International Journal of Robotics Research, 43(3):330–353, 2024

  93. [101]

    Human-like motion of a humanoid robot arm based on a closed-form solution of the inverse kinematics problem

    Tamim Asfour and R¨ udiger Dillmann. Human-like motion of a humanoid robot arm based on a closed-form solution of the inverse kinematics problem. InProceed- ings 2003 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2003)(Cat. No. 03CH37453) , volume 2...

  94. [102]

    Retargetting motion to new characters

    Michael Gleicher. Retargetting motion to new characters. Proceedings of the 25th annual conference on Computer graphics and interactive techniques , 1998

  95. [103]

    Whole-body geo- metric retargeting for humanoid robots

    Kourosh Darvish, Yeshasvi Tirupachuri, Giulio Romualdi, Lorenzo Rapetti, Diego Ferigo, Francisco Javier Andrade Chavez, and Daniele Pucci. Whole-body geo- metric retargeting for humanoid robots. In 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids) , pa...

  96. [104]

    Expressive whole-body control for humanoid robots

    Xuxin Cheng, Yandong Ji, Junming Chen, Ruihan Yang, Ge Yang, and Xiao- long Wang. Expressive whole-body control for humanoid robots. arXiv preprint arXiv:2402.16796, 2024

  97. [105]

    Task model of lower body motion for a biped humanoid robot to imitate human dances

    Shinichiro Nakaoka, Atsushi Nakazawa, Fumio Kanehiro, Kenji Kaneko, Mit- suharu Morisawa, and Katsushi Ikeuchi. Task model of lower body motion for a biped humanoid robot to imitate human dances. In2005 IEEE/RSJ International Conference on Intelligent Robots and Systems , page...

  98. [106]

    Online human walking imitation in task and joint space based on quadratic programming

    Kai Hu, Christian Ott, and Dongheui Lee. Online human walking imitation in task and joint space based on quadratic programming. In2014 IEEE International Conference on Robotics and Automation (ICRA) , pages 3458–3464. IEEE, 2014

  99. [107]

    Nonparametric motion retargeting for humanoid robots on shared latent space

    Sungjoon Choi, Matthew KXJ Pan, and Joohyung Kim. Nonparametric motion retargeting for humanoid robots on shared latent space. In Robotics: Science and Systems, 2020. 216

  100. [108]

    Human mo- tion reconstruction and synthesis of human skills

    Emel Demircan, Thor Besier, Samir Menon, and Oussama Khatib. Human mo- tion reconstruction and synthesis of human skills. In Advances in Robot Kinemat- ics: Motion in Man and Machine: Motion in Man and Machine , pages 283–292. Springer, 2010

  101. [109]

    Okami: Teaching humanoid robots manipulation skills through single video imitation

    Jinhan Li, Yifeng Zhu, Yuqi Xie, Zhenyu Jiang, Mingyo Seo, Georgios Pavlakos, and Yuke Zhu. Okami: Teaching humanoid robots manipulation skills through single video imitation. In 8th Annual Conference on Robot Learning , 2024

  102. [110]

    The 2019 davis challenge on vos: Unsupervised multi-object segmentation

    Sergi Caelles, Jordi Pont-Tuset, Federico Perazzi, Alberto Montes, Kevis-Kokitsi Maninis, and Luc Van Gool. The 2019 davis challenge on vos: Unsupervised multi-object segmentation. arXiv preprint arXiv:1905.00737 , 2019

  103. [111]

    Out of sight, still in mind: Reasoning and planning about unobserved objects with video tracking enabled memory models

    Yixuan Huang, Jialin Yuan, Chanho Kim, Pupul Pradhan, Bryan Chen, Li Fuxin, and Tucker Hermans. Out of sight, still in mind: Reasoning and planning about unobserved objects with video tracking enabled memory models. arXiv preprint arXiv:2309.15278, 2023

  104. [112]

    Open-world object manipulation using pre-trained vision-language models

    Austin Stone, Ted Xiao, Yao Lu, Keerthana Gopalakrishnan, Kuang-Huei Lee, Quan Vuong, Paul Wohlhart, Sean Kirmani, Brianna Zitkovich, Fei Xia, et al. Open-world object manipulation using pre-trained vision-language models. In Conference on Robot Learning, pages 3397–3417. PMLR, 2023

  105. [113]

    Expressive body capture: 3d hands, face, and body from a single image

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognitio...

  106. [114]

    Learning fine-grained bimanual manipulation with low-cost hardware

    Tony Z Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023. 217

  107. [115]

    Open- television: Teleoperation with immersive active visual feedback

    Xuxin Cheng, Jialong Li, Shiqi Yang, Ge Yang, and Xiaolong Wang. Open- television: Teleoperation with immersive active visual feedback. In 8th Annual Conference on Robot Learning, 2024

  108. [116]

    Between mdps and semi- mdps: A framework for temporal abstraction in reinforcement learning

    Richard S Sutton, Doina Precup, and Satinder Singh. Between mdps and semi- mdps: A framework for temporal abstraction in reinforcement learning. Artificial intelligence, 112(1-2):181–211, 1999

  109. [117]

    Temporal abstraction in reinforcement learning, 2000

    Doina Precup. Temporal abstraction in reinforcement learning, 2000

  110. [118]

    Feudal networks for hierarchi- cal reinforcement learning

    Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu. Feudal networks for hierarchi- cal reinforcement learning. In ICML, 2017

  111. [119]

    Multi-level discovery of deep options

    Roy Fox, Sanjay Krishnan, Ion Stoica, and Ken Goldberg. Multi-level discovery of deep options. arXiv:1703.08294, 2017

  112. [120]

    Variational intrinsic control

    Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra. Variational intrinsic control. arXiv:1611.07507, 2016

  113. [121]

    Skill discovery in continuous reinforcement learning domains using skill chaining

    George Konidaris and Andrew Barto. Skill discovery in continuous reinforcement learning domains using skill chaining. NIPS, 22:1015–1023, 2009

  114. [122]

    Option discovery using deep skill chaining

    Akhil Bagaria and George Konidaris. Option discovery using deep skill chaining. In ICLR, 2019

  115. [123]

    Expanding motor skills using relay networks

    Visak CV Kumar, Sehoon Ha, and C Karen Liu. Expanding motor skills using relay networks. In CoRL, 2018

  116. [124]

    Diversity is all you need: Learning skills without a reward function.arXiv:1802.06070, 2018

    Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine. Diversity is all you need: Learning skills without a reward function.arXiv:1802.06070, 2018

  117. [125]

    Learning an embedding space for transferable robot skills

    Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller. Learning an embedding space for transferable robot skills. In ICLR, 2018. 218

  118. [126]

    Dynamics-aware unsupervised discovery of skills

    Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar, and Karol Hausman. Dynamics-aware unsupervised discovery of skills. In ICLR, 2019

  119. [127]

    Skill acquisition from human demonstration using a hidden markov model

    Geir E Hovland, Pavan Sikka, and Brenan J McCarragher. Skill acquisition from human demonstration using a hidden markov model. In ICRA, volume 3, 1996

  120. [128]

    Learn- ing and generalization of complex tasks from unstructured demonstrations

    Scott Niekum, Sarah Osentoski, George Konidaris, and Andrew G Barto. Learn- ing and generalization of complex tasks from unstructured demonstrations. In IROS. IEEE, 2012

  121. [129]

    Learning robot skills with temporal vari- ational inference

    Tanmay Shankar and Abhinav Gupta. Learning robot skills with temporal vari- ational inference. In ICML, pages 8624–8633, 2020

  122. [130]

    Skid raw: Skill discovery from raw trajectories

    Daniel Tanneberg, Kai Ploeger, Elmar Rueckert, and Jan Peters. Skid raw: Skill discovery from raw trajectories. RA-L, 6(3):4696–4703, 2021

  123. [131]

    Compositional imitation learning: Ex- plaining and executing one task at a time

    Thomas Kipf, Yujia Li, Hanjun Dai, Vinicius Zambaldi, Edward Grefenstette, Pushmeet Kohli, and Peter Battaglia. Compositional imitation learning: Ex- plaining and executing one task at a time. arXiv:1812.01483, 2018

  124. [132]

    Taco: Learning task decomposition via temporal alignment for control

    Kyriacos Shiarlis, Markus Wulfmeier, Sasha Salter, Shimon Whiteson, and In- gmar Posner. Taco: Learning task decomposition via temporal alignment for control. In International Conference on Machine Learning , pages 4654–4663. PMLR, 2018

  125. [133]

    Bottom-up skill discovery from unseg- mented demonstrations for long-horizon robot manipulation

    Yifeng Zhu, Peter Stone, and Yuke Zhu. Bottom-up skill discovery from unseg- mented demonstrations for long-horizon robot manipulation. IEEE Robotics and Automation Letters, 7(2):4126–4133, 2022

  126. [134]

    Making sense of vision and touch: Learning multimodal representations for contact-rich tasks

    Michelle A Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Krishnan Srinivasan, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg. Making sense of vision and touch: Learning multimodal representations for contact-rich tasks. IEEE Transactions on Robotics, 2020. 219

  127. [135]

    Training products of experts by minimizing contrastive di- vergence

    Geoffrey E Hinton. Training products of experts by minimizing contrastive di- vergence. Neural computation, 14(8):1771–1800, 2002

  128. [136]

    Auto-encoding variational bayes

    Diederik P Kingma and Max Welling. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013

  129. [137]

    Action recognition by hierarchical mid-level action elements

    Tian Lan, Yuke Zhu, Amir Roshan Zamir, and Silvio Savarese. Action recognition by hierarchical mid-level action elements. In ICCV, pages 4552–4560, 2015

  130. [138]

    A tutorial on spectral clustering

    Ulrike Von Luxburg. A tutorial on spectral clustering. Statistics and computing , 17(4):395–416, 2007

  131. [139]

    Online bayesian changepoint detection for articulated motion models

    Scott Niekum, Sarah Osentoski, Christopher G Atkeson, and Andrew G Barto. Online bayesian changepoint detection for articulated motion models. In ICRA. IEEE, 2015

  132. [140]

    Lifelong robot learning

    Sebastian Thrun and Tom M Mitchell. Lifelong robot learning. Robotics and autonomous systems, 15(1-2):25–46, 1995

  133. [141]

    Lotus: Continual imita- tion learning for robot manipulation through unsupervised skill discovery

    Weikang Wan, Yifeng Zhu, Rutav Shah, and Yuke Zhu. Lotus: Continual imita- tion learning for robot manipulation through unsupervised skill discovery. arXiv preprint arXiv:2311.02058, 2023

  134. [142]

    Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning

    Sanjay Krishnan, Animesh Garg, Sachin Patil, Colin Lea, Gregory Hager, Pieter Abbeel, and Ken Goldberg. Transition state clustering: Unsupervised surgical trajectory segmentation for robot learning. IJRR, 2017

  135. [143]

    Silhouettes: a graphical aid to the interpretation and valida- tion of cluster analysis

    Peter J Rousseeuw. Silhouettes: a graphical aid to the interpretation and valida- tion of cluster analysis. Journal of computational and applied mathematics , 20: 53–65, 1987

  136. [144]

    On tiny episodic memories in continual learning

    Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajan- than, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. On tiny episodic memories in continual learning. arXiv preprint arXiv:1902.10486 , 2019. 220

  137. [145]

    Mixture density networks

    Christopher M Bishop. Mixture density networks. 1994

  138. [146]

    Learning transferable visual models from natural language supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning , p...

  139. [147]

    Contin- ual lifelong learning in natural language processing: A survey

    Magdalena Biesialska, Katarzyna Biesialska, and Marta R Costa-Jussa. Contin- ual lifelong learning in natural language processing: A survey. arXiv preprint arXiv:2012.09823, 2020

  140. [148]

    Online continual learning in image classification: An empirical survey

    Zheda Mai, Ruiwen Li, Jihwan Jeong, David Quispe, Hyunwoo Kim, and Scott Sanner. Online continual learning in image classification: An empirical survey. Neurocomputing, 469:28–51, 2022

  141. [149]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceed- ings of the IEEE/CVF Conference on Computer Visio...

  142. [150]

    Pddl-the planning domain definition language

    Drew McDermott, Malik Ghallab, Adele Howe, Craig Knoblock, Ashwin Ram, Manuela Veloso, Daniel Weld, and David Wilkins. Pddl-the planning domain definition language. 1998

  143. [151]

    Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments

    Sanjana Srivastava, Chengshu Li, Michael Lingelbach, Roberto Mart´ ın-Mart´ ın, Fei Xia, Kent Elliott Vainio, Zheng Lian, Cem Gokmen, Shyamal Buch, Karen Liu, et al. Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments. In ...

  144. [152]

    Dynamic movement primitives-a framework for motor control in humans and humanoid robotics

    Stefan Schaal. Dynamic movement primitives-a framework for motor control in humans and humanoid robotics. In Adaptive motion of animals and machines , pages 261–280. Springer, 2006

  145. [153]

    Learning motor primitives for robotics

    Jens Kober and Jan Peters. Learning motor primitives for robotics. In 2009 IEEE International Conference on Robotics and Automation , pages 2112–2118. IEEE, 2009

  146. [154]

    Probabilistic movement primitives

    Alexandros Paraschos, Christian Daniel, Jan R Peters, and Gerhard Neumann. Probabilistic movement primitives. Advances in neural information processing systems, 26, 2013

  147. [155]

    Us- ing probabilistic movement primitives in robotics

    Alexandros Paraschos, Christian Daniel, Jan Peters, and Gerhard Neumann. Us- ing probabilistic movement primitives in robotics. Autonomous Robots, 42(3): 529–551, 2018

  148. [156]

    From play to policy: Conditional behavior generation from uncurated robot data

    Zichen Jeff Cui, Yibin Wang, Nur Muhammad, Lerrel Pinto, et al. From play to policy: Conditional behavior generation from uncurated robot data. arXiv preprint arXiv:2210.10047, 2022

  149. [157]

    Learning and re- trieval from prior data for skill-based imitation learning

    Soroush Nasiriany, Tian Gao, Ajay Mandlekar, and Yuke Zhu. Learning and re- trieval from prior data for skill-based imitation learning. In6th Annual Conference on Robot Learning, 2022

  150. [158]

    Cliport: What and where pathways for robotic manipulation

    Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Cliport: What and where pathways for robotic manipulation. In Conference on Robot Learning, pages 894–

  151. [159]

    Perceiver-actor: A multi-task transformer for robotic manipulation

    Mohit Shridhar, Lucas Manuelli, and Dieter Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. In Conference on Robot Learning , pages 785–799. PMLR, 2023. 222

  152. [160]

    Dexcap: Scalable and portable mocap data collection system for dexterous manipulation

    Chen Wang, Haochen Shi, Weizhuo Wang, Ruohan Zhang, Li Fei-Fei, and C Karen Liu. Dexcap: Scalable and portable mocap data collection system for dexterous manipulation. arXiv preprint arXiv:2403.07788 , 2024

  153. [161]

    Learning visuotactile skills with two multifingered hands

    Toru Lin, Yu Zhang, Qiyang Li, Haozhi Qi, Brent Yi, Sergey Levine, and Jitendra Malik. Learning visuotactile skills with two multifingered hands. arXiv preprint arXiv:2404.16823, 2024

  154. [162]

    3d diffusion policy

    Yanjie Ze, Gu Zhang, Kangning Zhang, Chenyuan Hu, Muhan Wang, and Huazhe Xu. 3d diffusion policy. arXiv preprint arXiv:2403.03954 , 2024

  155. [163]

    One-shot visual imitation via attributed waypoints and demonstration augmentation

    Matthew Chang and Saurabh Gupta. One-shot visual imitation via attributed waypoints and demonstration augmentation. In 2023 IEEE International Con- ference on Robotics and Automation (ICRA) , pages 5055–5062. IEEE, 2023

  156. [164]

    One-shot composi- tion of vision-based skills from demonstration

    Tianhe Yu, Pieter Abbeel, Sergey Levine, and Chelsea Finn. One-shot composi- tion of vision-based skills from demonstration. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 2643–2650. IEEE, 2019

  157. [165]

    Demonstrate once, imitate immediately (dome): Learning visual servoing for one- shot imitation learning

    Eugene Valassakis, Georgios Papagiannis, Norman Di Palo, and Edward Johns. Demonstrate once, imitate immediately (dome): Learning visual servoing for one- shot imitation learning. In 2022 IEEE/RSJ International Conference on Intelli- gent Robots and Systems (IROS) , pages 8614...

  158. [166]

    Coarse-to-fine imitation learning: Robot manipulation from a single demonstration

    Edward Johns. Coarse-to-fine imitation learning: Robot manipulation from a single demonstration. In 2021 IEEE international conference on robotics and automation (ICRA), pages 4613–4619. IEEE, 2021

  159. [167]

    Learning multi-stage tasks with one demon- stration via self-replay

    Norman Di Palo and Edward Johns. Learning multi-stage tasks with one demon- stration via self-replay. In Conference on Robot Learning , pages 1180–1189. PMLR, 2022. 223

  160. [168]

    Learning latent plans from play

    Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet. Learning latent plans from play. In Confer- ence on robot learning, pages 1113–1132. PMLR, 2020

  161. [169]

    Relay policy learning: Solving long-horizon tasks via imitation and rein- forcement learning

    Abhishek Gupta, Vikash Kumar, Corey Lynch, Sergey Levine, and Karol Haus- man. Relay policy learning: Solving long-horizon tasks via imitation and rein- forcement learning. In CoRL, pages 1025–1037, 2020

  162. [170]

    Learning multi-arm manipulation through collab- orative teleoperation

    Albert Tung, Josiah Wong, Ajay Mandlekar, Roberto Mart´ ın-Mart´ ın, Yuke Zhu, Li Fei-Fei, and Silvio Savarese. Learning multi-arm manipulation through collab- orative teleoperation. In ICRA, 2021

  163. [171]

    Hindsight experience replay

    Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba. Hindsight experience replay. In NIPS, 2017

  164. [172]

    One-shot imi- tation learning

    Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba. One-shot imi- tation learning. Advances in neural information processing systems , 30, 2017

  165. [173]

    Watch and match: Supercharging imitation with regularized optimal transport

    Siddhant Haldar, Vaibhav Mathur, Denis Yarats, and Lerrel Pinto. Watch and match: Supercharging imitation with regularized optimal transport. In Confer- ence on Robot Learning, pages 32–43. PMLR, 2023

  166. [174]

    Teach a robot to fish: Versatile imitation from one minute of demonstrations

    Siddhant Haldar, Jyothish Pari, Anant Rai, and Lerrel Pinto. Teach a robot to fish: Versatile imitation from one minute of demonstrations. arXiv preprint arXiv:2303.01497, 2023

  167. [175]

    View: Visual imitation learning with waypoints

    Ananth Jonnavittula, Sagar Parekh, and Dylan P Losey. View: Visual imitation learning with waypoints. arXiv preprint arXiv:2404.17906 , 2024

  168. [176]

    Learning multi-step manipulation tasks from a single human demonstration

    Dingkun Guo. Learning multi-step manipulation tasks from a single human demonstration. arXiv preprint arXiv:2312.15346 , 2023. 224

  169. [177]

    Learning to select and generalize striking movements in robot table tennis

    Katharina M¨ ulling, Jens Kober, Oliver Kroemer, and Jan Peters. Learning to select and generalize striking movements in robot table tennis. The International Journal of Robotics Research, 32(3):263–279, 2013

  170. [178]

    Masked visual pre-training for motor control

    Tete Xiao, Ilija Radosavovic, Trevor Darrell, and Jitendra Malik. Masked visual pre-training for motor control. arXiv preprint arXiv:2203.06173 , 2022

  171. [179]

    Cacti: A framework for scalable multi-task multi- scene visual imitation learning

    Zhao Mandi, Homanga Bharadhwaj, Vincent Moens, Shuran Song, Aravind Ra- jeswaran, and Vikash Kumar. Cacti: A framework for scalable multi-task multi- scene visual imitation learning. arXiv preprint arXiv:2212.05711 , 2022

  172. [180]

    Scaling robot learning with semantically imagined experience

    Tianhe Yu, Ted Xiao, Austin Stone, Jonathan Tompson, Anthony Brohan, Su Wang, Jaspiar Singh, Clayton Tan, Jodilyn Peralta, Brian Ichter, et al. Scaling robot learning with semantically imagined experience. arXiv preprint arXiv:2302.11550, 2023

  173. [181]

    Genaug: Retarget- ing behaviors to unseen situations via generative augmentation

    Zoey Chen, Sho Kiami, Abhishek Gupta, and Vikash Kumar. Genaug: Retarget- ing behaviors to unseen situations via generative augmentation. arXiv preprint arXiv:2302.06671, 2023

  174. [182]

    Transporter networks: Rearranging the visual world for robotic manipula- tion

    Andy Zeng, Pete Florence, Jonathan Tompson, Stefan Welker, Jonathan Chien, Maria Attarian, Travis Armstrong, Ivan Krasin, Dan Duong, Vikas Sindhwani, et al. Transporter networks: Rearranging the visual world for robotic manipula- tion. arXiv preprint arXiv:2010.14406 , 2020

  175. [183]

    Synergies between affordance and geometry: 6-dof grasp detection via implicit representa- tions

    Zhenyu Jiang, Yifeng Zhu, Maxwell Svetlik, Kuan Fang, and Yuke Zhu. Synergies between affordance and geometry: 6-dof grasp detection via implicit representa- tions. arXiv preprint arXiv:2104.01542 , 2021

  176. [184]

    On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline

    Nicklas Hansen, Zhecheng Yuan, Yanjie Ze, Tongzhou Mu, Aravind Rajeswaran, Hao Su, Huazhe Xu, and Xiaolong Wang. On pre-training for visuo-motor control: Revisiting a learning-from-scratch baseline. arXiv preprint arXiv:2212.05749 , 2022. 225

  177. [185]

    Deep object pose estimation for semantic robotic grasping of household objects

    Jonathan Tremblay, Thang To, Balakumar Sundaralingam, Yu Xiang, Dieter Fox, and Stan Birchfield. Deep object pose estimation for semantic robotic grasping of household objects. arXiv preprint arXiv:1809.10790 , 2018

  178. [186]

    6-dof pose estimation of household objects for robotic manipulation: An accessible dataset and benchmark

    Stephen Tyree, Jonathan Tremblay, Thang To, Jia Cheng, Terry Mosier, Jef- frey Smith, and Stan Birchfield. 6-dof pose estimation of household objects for robotic manipulation: An accessible dataset and benchmark. arXiv preprint arXiv:2203.05701, 2022

  179. [187]

    Object-centric task and motion planning in dynamic environments

    Toki Migimatsu and Jeannette Bohg. Object-centric task and motion planning in dynamic environments. IEEE Robotics and Automation Letters , 5(2):844–851, 2020

  180. [188]

    Deep object-centric policies for autonomous driving

    Dequan Wang, Coline Devin, Qi-Zhi Cai, Fisher Yu, and Trevor Darrell. Deep object-centric policies for autonomous driving. In 2019 International Conference on Robotics and Automation (ICRA) , pages 8853–8859. IEEE, 2019

  181. [189]

    Deep object- centric representations for generalizable robot learning

    Coline Devin, Pieter Abbeel, Trevor Darrell, and Sergey Levine. Deep object- centric representations for generalizable robot learning. In 2018 IEEE Interna- tional Conference on Robotics and Automation (ICRA) , pages 7111–7118. IEEE, 2018

  182. [190]

    Monet: Unsupervised scene decomposition and representation

    Christopher P Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, and Alexander Lerchner. Monet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390 , 2019

  183. [191]

    Visuomotor control in multi-object scenes using object-aware representations

    Negin Heravi, Ayzaan Wahid, Corey Lynch, Pete Florence, Travis Armstrong, Jonathan Tompson, Pierre Sermanet, Jeannette Bohg, and Debidatta Dwibedi. Visuomotor control in multi-object scenes using object-aware representations. arXiv preprint arXiv:2205.06333 , 2022

  184. [192]

    Emerging properties in self-supervised vision 226 transformers

    Mathilde Caron, Hugo Touvron, Ishan Misra, Herv´ e J´ egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision 226 transformers. In Proceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021

  185. [193]

    what” and “where

    Junyao Shi, Jianing Qian, Yecheng Jason Ma, and Dinesh Jayaraman. Plug-and- play object-centric representations from “what” and “where” foundation models. In ICRA, 2024

  186. [194]

    Keypoint action tokens enable in-context imitation learning in robotics

    Norman Di Palo and Edward Johns. Keypoint action tokens enable in-context imitation learning in robotics. arXiv preprint arXiv:2403.19578 , 2024

  187. [195]

    Vl-bert: Pre-training of generic visual-linguistic representations

    Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. Vl-bert: Pre-training of generic visual-linguistic representations. arXiv preprint arXiv:1908.08530, 2019

  188. [196]

    Uniter: Universal image-text representation learning

    Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. Uniter: Universal image-text representation learning. In European conference on computer vision , pages 104–120. Springer, 2020

  189. [197]

    Graph inverse reinforcement learning from diverse videos

    Sateesh Kumar, Jonathan Zamora, Nicklas Hansen, Rishabh Jangir, and Xiaolong Wang. Graph inverse reinforcement learning from diverse videos. In Conference on Robot Learning, pages 55–66. PMLR, 2023

  190. [198]

    Structurenet: Hierarchical graph networks for 3d shape gen- eration

    Kaichun Mo, Paul Guerrero, Li Yi, Hao Su, Peter Wonka, Niloy Mitra, and Leonidas J Guibas. Structurenet: Hierarchical graph networks for 3d shape gen- eration. arXiv preprint arXiv:1908.00575 , 2019

  191. [199]

    Planning for multi-object manipulation with graph neural network relational classifiers

    Yixuan Huang, Adam Conkey, and Tucker Hermans. Planning for multi-object manipulation with graph neural network relational classifiers. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 1822–1829. IEEE, 2023. 227

  192. [200]

    Nerp: Neural rearrangement planning for unknown objects

    Ahmed H Qureshi, Arsalan Mousavian, Chris Paxton, Michael C Yip, and Dieter Fox. Nerp: Neural rearrangement planning for unknown objects. arXiv preprint arXiv:2106.01352, 2021

  193. [201]

    Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs

    Yifeng Zhu, Jonathan Tremblay, Stan Birchfield, and Yuke Zhu. Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs. In 2021 IEEE International Conference on Robotics and Automation (ICRA) , pages 6541–6548. IEEE, 2021

  194. [202]

    Human-to-robot imitation in the wild

    Shikhar Bahl, Abhinav Gupta, and Deepak Pathak. Human-to-robot imitation in the wild. arXiv preprint arXiv:2207.09450 , 2022

  195. [203]

    Imitation from observation: Learning to imitate behaviors from raw video via context translation

    YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Imitation from observation: Learning to imitate behaviors from raw video via context translation. In 2018 IEEE International Conference on Robotics and Automation (ICRA) , pages 1118–1125. IEEE, 2018

  196. [204]

    Third-person visual imitation learning via decoupled hierarchical controller

    Pratyusha Sharma, Deepak Pathak, and Abhinav Gupta. Third-person visual imitation learning via decoupled hierarchical controller. Advances in Neural In- formation Processing Systems, 32, 2019

  197. [205]

    Avid: Learning multi-stage tasks via pixel-level translation of human videos

    Laura Smith, Nikita Dhawan, Marvin Zhang, Pieter Abbeel, and Sergey Levine. Avid: Learning multi-stage tasks via pixel-level translation of human videos. arXiv:1912.04443, 2019

  198. [206]

    Xskill: Cross embodiment skill discovery

    Mengda Xu, Zhenjia Xu, Cheng Chi, Manuela Veloso, and Shuran Song. Xskill: Cross embodiment skill discovery. In Conference on Robot Learning, pages 3536–

  199. [207]

    Learning fabric manipulation in the real world with human videos

    Robert Lee, Jad Abou-Chakra, Fangyi Zhang, and Peter Corke. Learning fabric manipulation in the real world with human videos. arXiv preprint arXiv:2211.02832, 2022. 228

  200. [208]

    Learning dexterity from human hand motion in internet videos

    Kenneth Shaw, Shikhar Bahl, Aravind Sivakumar, Aditya Kannan, and Deepak Pathak. Learning dexterity from human hand motion in internet videos. The International Journal of Robotics Research , 43(4):513–532, 2024

  201. [209]

    Screwmimic: Bimanual imitation from human videos with screw space projection

    Arpit Bahety, Priyanka Mandikal, Ben Abbatematteo, and Roberto Mart´ ın- Mart´ ın. Screwmimic: Bimanual imitation from human videos with screw space projection. arXiv preprint arXiv:2405.03666 , 2024

  202. [210]

    Towards generalizable zero-shot manipulation via translating human interaction plans

    Homanga Bharadhwaj, Abhinav Gupta, Vikash Kumar, and Shubham Tulsiani. Towards generalizable zero-shot manipulation via translating human interaction plans. arXiv preprint arXiv:2312.00775 , 2023

  203. [211]

    Ridm: Reinforced inverse dynamics modeling for learning from a single observed demonstration

    Brahma S Pavse, Faraz Torabi, Josiah Hanna, Garrett Warnell, and Peter Stone. Ridm: Reinforced inverse dynamics modeling for learning from a single observed demonstration. IEEE Robotics and Automation Letters , 5(4):6262–6269, 2020

  204. [212]

    Adversarial imitation learning from video using a state observer

    Haresh Karnan, Faraz Torabi, Garrett Warnell, and Peter Stone. Adversarial imitation learning from video using a state observer. In 2022 International Con- ference on Robotics and Automation (ICRA) , pages 2452–2458. IEEE, 2022

  205. [213]

    Imitation learning from video by leveraging proprioception

    Faraz Torabi, Garrett Warnell, and Peter Stone. Imitation learning from video by leveraging proprioception. In Proceedings of the 28th International Joint Con- ference on Artificial Intelligence , pages 3585–3591, 2019

  206. [214]

    Generative adversarial imitation from observation

    Faraz Torabi, Garrett Warnell, and Peter Stone. Generative adversarial imitation from observation. In Imitation, Intent, and Interaction (I3) Workshop at ICML 2019, June 2019

  207. [215]

    Behavioral cloning from obser- vation

    Faraz Torabi, Garrett Warnell, and Peter Stone. Behavioral cloning from obser- vation. In Proceedings of the 27th International Joint Conference on Artificial Intelligence, pages 4950–4957, 2018. 229

  208. [216]

    Learning human-to-humanoid real-time whole-body teleoper- ation

    Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleoper- ation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8944–8951. IEEE, 2024

  209. [217]

    Hierarchical quadratic programming: Fast online humanoid-robot motion generation

    Adrien Escande, Nicolas Mansard, and Pierre-Brice Wieber. Hierarchical quadratic programming: Fast online humanoid-robot motion generation. The International Journal of Robotics Research , 33(7):1006–1028, 2014

  210. [218]

    Berkeley humanoid: A research platform for learning-based control

    Qiayuan Liao, Bike Zhang, Xuanyu Huang, Xiaoyu Huang, Zhongyu Li, and Koushil Sreenath. Berkeley humanoid: A research platform for learning-based control. arXiv preprint arXiv:2407.21781 , 2024

  211. [219]

    A multimode teleoperation framework for humanoid loco- manipulation: An application for the icub robot

    Luigi Penco, Nicola Scianca, Valerio Modugno, Leonardo Lanari, Giuseppe Oriolo, and Serena Ivaldi. A multimode teleoperation framework for humanoid loco- manipulation: An application for the icub robot. IEEE Robotics & Automation Magazine, 26(4):73–82, 2019

  212. [220]

    Whole body motion control frame- work for arbitrarily and simultaneously assigned upper-body tasks and walking motion

    Doik Kim, Bum-Jae You, and Sang-Rok Oh. Whole body motion control frame- work for arbitrarily and simultaneously assigned upper-body tasks and walking motion. Modeling, Simulation and Optimization of Bipedal Walking , pages 87–98, 2013

  213. [221]

    Multi-contact motion retargeting from human to humanoid robot

    Alessandro Di Fava, Karim Bouyarmane, Kevin Chappellet, Emanuele Ruffaldi, and Abderrahmane Kheddar. Multi-contact motion retargeting from human to humanoid robot. In 2016 IEEE-RAS 16th international conference on humanoid robots (humanoids), pages 1081–1086. IEEE, 2016

  214. [222]

    Human to robot whole-body motion transfer

    Miguel Arduengo, Ana Arduengo, Adri` a Colom´ e, Joan Lobo-Prat, and Carme Torras. Human to robot whole-body motion transfer. In 2020 IEEE-RAS 20th In- ternational Conference on Humanoid Robots (Humanoids), pages 299–305. IEEE, 2021. 230

  215. [223]

    Team janus humanoid avatar: A cybernetic avatar to embody human telepresence

    R Cisneros, M Benallegue, K Kaneko, H Kaminaga, G Caron, A Tanguy, R Singh, L Sun, A Dallard, C Fournier, et al. Team janus humanoid avatar: A cybernetic avatar to embody human telepresence. In Toward Robot Avatars: Perspectives on the ANA Avatar XPRIZE Competition, RSS Worksh...

  216. [224]

    Telexistence cockpit for humanoid robot control

    Susumu Tachi, Kiyoshi Komoriya, Kazuya Sawada, Takashi Nishiyama, Toshiyuki Itoko, Masami Kobayashi, and Kozo Inoue. Telexistence cockpit for humanoid robot control. Advanced Robotics, 17(3):199–217, 2003

  217. [225]

    Humanoid dynamic synchronization through whole-body bilateral feedback teleoperation

    Joao Ramos and Sangbae Kim. Humanoid dynamic synchronization through whole-body bilateral feedback teleoperation. IEEE Transactions on Robotics, 34 (4):953–965, 2018

  218. [226]

    Bilateral hu- manoid teleoperation system using whole-body exoskeleton cockpit tablis

    Yasuhiro Ishiguro, Tasuku Makabe, Yuya Nagamatsu, Yuta Kojio, Kunio Kojima, Fumihito Sugai, Yohei Kakiuchi, Kei Okada, and Masayuki Inaba. Bilateral hu- manoid teleoperation system using whole-body exoskeleton cockpit tablis. IEEE Robotics and Automation Letters, 5(4):6419–6426, 2020

  219. [227]

    Humanoid teleoperation using task-relevant haptic feedback

    Firas Abi-Farrajl, Bernd Henze, Alexander Werner, Michael Panzirsch, Christian Ott, and M´ aximo A Roa. Humanoid teleoperation using task-relevant haptic feedback. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 5010–5017, 2018. doi: 1...

  220. [228]

    Nimbro avatar: Interactive immersive telepresence with force-feedback telemanipulation

    Max Schwarz, Christian Lenz, Andre Rochow, Michael Schreiber, and Sven Behnke. Nimbro avatar: Interactive immersive telepresence with force-feedback telemanipulation. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 5312–5319. IEEE, 2021

  221. [229]

    Deep imitation learning for humanoid loco-manipulation through human teleoperation

    Mingyo Seo, Steve Han, Kyutae Sim, Seung Hyeon Bang, Carlos Gonzalez, Luis Sentis, and Yuke Zhu. Deep imitation learning for humanoid loco-manipulation through human teleoperation. In IEEE-RAS International Conference on Hu- manoid Robots (Humanoids), 2023. 231

  222. [230]

    Virtual reality teleoperation of a humanoid robot using markerless hu- man upper body pose imitation

    Matthias Hirschmanner, Christiana Tsiourti, Timothy Patten, and Markus Vincze. Virtual reality teleoperation of a humanoid robot using markerless hu- man upper body pose imitation. in 2019 ieee-ras 19th international conference on humanoid robots (humanoids), 2019

  223. [231]

    Online telemanipulation framework on humanoid for both manipulation and imitation

    Daegyu Lim, Donghyeon Kim, and Jaeheung Park. Online telemanipulation framework on humanoid for both manipulation and imitation. 2022 19th In- ternational Conference on Ubiquitous Robots (UR) , pages 8–15, 2022. URL https://api.semanticscholar.org/CorpusID:250577582

  224. [232]

    Hu- manplus: Humanoid shadowing and imitation from humans

    Zipeng Fu, Qingqing Zhao, Qi Wu, Gordon Wetzstein, and Chelsea Finn. Hu- manplus: Humanoid shadowing and imitation from humans. arXiv preprint arXiv:2406.10454, 2024

  225. [233]

    Perpetual humanoid control for real-time simulated avatars

    Zhengyi Luo, Jinkun Cao, Kris Kitani, Weipeng Xu, et al. Perpetual humanoid control for real-time simulated avatars. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision , pages 10895–10904, 2023

  226. [234]

    Amp: Adversarial motion priors for stylized physics-based character control

    Xue Bin Peng, Ze Ma, Pieter Abbeel, Sergey Levine, and Angjoo Kanazawa. Amp: Adversarial motion priors for stylized physics-based character control. ACM Transactions on Graphics (ToG), 40(4):1–20, 2021

  227. [235]

    Motiongpt: Human motion as a foreign language

    Biao Jiang, Xin Chen, Wen Liu, Jingyi Yu, Gang Yu, and Tao Chen. Motiongpt: Human motion as a foreign language. Advances in Neural Information Processing Systems, 36, 2024

  228. [236]

    Optimization- based locomotion planning, estimation, and control design for the atlas humanoid robot

    Scott Kuindersma, Robin Deits, Maurice Fallon, Andr´ es Valenzuela, Hongkai Dai, Frank Permenter, Twan Koolen, Pat Marion, and Russ Tedrake. Optimization- based locomotion planning, estimation, and control design for the atlas humanoid robot. Autonomous robots, 40:429–455, 2016. 232

  229. [237]

    Dynamic movement primitive based motion retargeting for dual-arm sign lan- guage motions

    Yuwei Liang, Weijie Li, Yue Wang, Rong Xiong, Yichao Mao, and Jiafan Zhang. Dynamic movement primitive based motion retargeting for dual-arm sign lan- guage motions. In 2021 IEEE International Conference on Robotics and Au- tomation (ICRA), pages 8195–8201. IEEE, 2021

  230. [238]

    Constructing skill trees for reinforcement learning agents from demon- stration trajectories

    George Dimitri Konidaris, Scott Kuindersma, Andrew G Barto, and Roderic A Grupen. Constructing skill trees for reinforcement learning agents from demon- stration trajectories. In NIPS, 2010

  231. [239]

    Modeling long-horizon tasks as sequential interaction landscapes

    S¨ oren Pirk, Karol Hausman, Alexander Toshev, and Mohi Khansari. Modeling long-horizon tasks as sequential interaction landscapes. arXiv:2006.04843, 2020

  232. [240]

    Learning manipulation graphs from demonstrations using multimodal sensory signals

    Zhe Su, Oliver Kroemer, Gerald E Loeb, Gaurav S Sukhatme, and Stefan Schaal. Learning manipulation graphs from demonstrations using multimodal sensory signals. In ICRA. IEEE, 2018

  233. [241]

    Real- time multisensory affordance-based control for adaptive object manipulation

    Vivian Chu, Reymundo A Gutierrez, Sonia Chernova, Andrea L Thomaz, Vivian Chu, Reymundo A Gutierrez, Sonia Chernova, and Andrea L Thomaz. Real- time multisensory affordance-based control for adaptive object manipulation. In ICRA. IEEE, 2019

  234. [242]

    The senses considered as percep- tual systems, volume 2

    James Jerome Gibson and Leonard Carmichael. The senses considered as percep- tual systems, volume 2. Houghton Mifflin Boston, 1966

  235. [243]

    A bottom-up approach for pancreas segmentation using cascaded su- perpixels and (deep) image patch labeling

    Amal Farag, Le Lu, Holger R Roth, Jiamin Liu, Evrim Turkbey, and Ronald M Summers. A bottom-up approach for pancreas segmentation using cascaded su- perpixels and (deep) image patch labeling. IEEE Transactions on Image Process- ing, 26(1):386–399, 2016

  236. [244]

    Efficient parameter-free clustering using first neighbor relations

    Saquib Sarfraz, Vivek Sharma, and Rainer Stiefelhagen. Efficient parameter-free clustering using first neighbor relations. In CVPR, 2019. 233

  237. [245]

    Temporally-weighted hierarchical clustering for unsupervised action segmentation

    M Saquib Sarfraz, Naila Murray, Vivek Sharma, Ali Diba, Luc Van Gool, and Rainer Stiefelhagen. Temporally-weighted hierarchical clustering for unsupervised action segmentation. arXiv:2103.11264, 2021

  238. [246]

    Stacked hourglass networks for human pose estimation

    Alejandro Newell, Kaiyu Yang, and Jia Deng. Stacked hourglass networks for human pose estimation. In ECCV, 2016

  239. [247]

    Cornernet: Detecting objects as paired keypoints

    Hei Law and Jia Deng. Cornernet: Detecting objects as paired keypoints. In ECCV, 2018

  240. [248]

    Bottom-up object detec- tion by grouping extreme and center points

    Xingyi Zhou, Jiacheng Zhuo, and Philipp Krahenbuhl. Bottom-up object detec- tion by grouping extreme and center points. In CVPR, pages 850–859, 2019

  241. [249]

    A robust layered control system for a mobile robot

    Rodney Brooks. A robust layered control system for a mobile robot. IEEE Journal on Robotics and Automation , 2(1):14–23, 1986

  242. [250]

    Intelligence without representation

    Rodney A Brooks. Intelligence without representation. Artificial Intelligence, 47 (1-3):139–159, 1991

  243. [251]

    A hierarchical architecture for behavior- based robots

    Monica N Nicolescu and Maja J Matari´ c. A hierarchical architecture for behavior- based robots. In AAMAS, 2002

  244. [252]

    Voyager: An open-ended embodied agent with large language models

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar. Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291 , 2023

  245. [253]

    A deep hierarchical approach to lifelong learning in minecraft

    Chen Tessler, Shahar Givony, Tom Zahavy, Daniel Mankowitz, and Shie Mannor. A deep hierarchical approach to lifelong learning in minecraft. In Proceedings of the AAAI conference on artificial intelligence , volume 31, 2017

  246. [254]

    Embodied lifelong learning for task and motion planning

    Jorge A Mendez, Leslie Pack Kaelbling, and Tom´ as Lozano-P´ erez. Embodied lifelong learning for task and motion planning. arXiv preprint arXiv:2307.06870, 2023. 234

  247. [255]

    Evaluating continual learning on a home robot

    Sam Powers, Abhinav Gupta, and Chris Paxton. Evaluating continual learning on a home robot. arXiv preprint arXiv:2306.02413 , 2023

  248. [256]

    Progressive neural networks

    Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell. Progressive neural networks. arXiv preprint arXiv:1606.04671 , 2016

  249. [257]

    Lifelong learning with dynamically expandable networks

    Jaehong Yoon, Eunho Yang, Jeongtae Lee, and Sung Ju Hwang. Lifelong learning with dynamically expandable networks. arXiv preprint arXiv:1708.01547 , 2017

  250. [258]

    Packnet: Adding multiple tasks to a single network by iterative pruning

    Arun Mallya and Svetlana Lazebnik. Packnet: Adding multiple tasks to a single network by iterative pruning. In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 7765–7773, 2018

  251. [259]

    Compacting, picking and growing for unforgetting continual learning

    Ching-Yi Hung, Cheng-Hao Tu, Cheng-En Wu, Chien-Hung Chen, Yi-Ming Chan, and Chu-Song Chen. Compacting, picking and growing for unforgetting continual learning. Advances in Neural Information Processing Systems , 32, 2019

  252. [260]

    Firefly neural architecture descent: a general approach for growing neural networks

    Lemeng Wu, Bo Liu, Peter Stone, and Qiang Liu. Firefly neural architecture descent: a general approach for growing neural networks. Advances in Neural Information Processing Systems, 33:22373–22383, 2020

  253. [261]

    Lifelong reinforcement learning with modulating masks

    Eseoghene Ben-Iwhiwhu, Saptarshi Nath, Praveen K Pilly, Soheil Kolouri, and Andrea Soltoggio. Lifelong reinforcement learning with modulating masks. arXiv preprint arXiv:2212.11110, 2022

  254. [262]

    Riemannian walk for incremental learning: Understanding forgetting and intransigence

    Arslan Chaudhry, Puneet K Dokania, Thalaiyasingam Ajanthan, and Philip HS Torr. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Proceedings of the European Conference on Computer Vision (ECCV), pages 532–547, 2018

  255. [263]

    Progress & com- 235 press: A scalable framework for continual learning

    Jonathan Schwarz, Wojciech Czarnecki, Jelena Luketina, Agnieszka Grabska- Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell. Progress & com- 235 press: A scalable framework for continual learning. In International Conference on Machine Learning, pages 4528–4537. PMLR, 2018

  256. [264]

    Continual learning with recursive gradient optimiza- tion

    Hao Liu and Huaping Liu. Continual learning with recursive gradient optimiza- tion. arXiv preprint arXiv:2201.12522 , 2022

  257. [265]

    Efficient lifelong learning with a-gem

    Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed El- hoseiny. Efficient lifelong learning with a-gem. arXiv preprint arXiv:1812.00420 , 2018

  258. [266]

    Dark experience for general continual learning: a strong, simple base- line

    Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. Dark experience for general continual learning: a strong, simple base- line. Advances in neural information processing systems , 33:15920–15930, 2020

  259. [267]

    Deep reinforcement learning amidst lifelong non-stationarity

    Annie Xie, James Harrison, and Chelsea Finn. Deep reinforcement learning amidst lifelong non-stationarity. arXiv preprint arXiv:2006.10701 , 2020

  260. [268]

    Lifelong robotic reinforcement learning by retaining experiences

    Annie Xie and Chelsea Finn. Lifelong robotic reinforcement learning by retaining experiences. In Conference on Lifelong Learning Agents , pages 838–855. PMLR, 2022

  261. [269]

    Polytask: Learning unified policies through behavior distillation

    Siddhant Haldar and Lerrel Pinto. Polytask: Learning unified policies through behavior distillation. arXiv preprint arXiv:2310.08573 , 2023

  262. [270]

    Continual vision-based reinforcement learning with group sym- metries

    Shiqi Liu, Mengdi Xu, Peide Huang, Xilun Zhang, Yongkang Liu, Kentaro Oguchi, and Ding Zhao. Continual vision-based reinforcement learning with group sym- metries. In 7th Annual Conference on Robot Learning , 2023

  263. [271]

    Modular lifelong reinforce- ment learning via neural composition

    Jorge A Mendez, Harm van Seijen, and Eric Eaton. Modular lifelong reinforce- ment learning via neural composition. arXiv preprint arXiv:2207.00429 , 2022

  264. [272]

    Fast lifelong adaptive inverse reinforcement 236 learning from demonstrations

    Letian Chen, Sravan Jayanthi, Rohan R Paleja, Daniel Martin, Viacheslav Za- kharov, and Matthew Gombolay. Fast lifelong adaptive inverse reinforcement 236 learning from demonstrations. In Conference on Robot Learning , pages 2083–

  265. [273]

    Experience replay for continual learning

    David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne. Experience replay for continual learning. Advances in Neural Information Processing Systems, 32, 2019

  266. [274]

    The mnist database of handwritten digit images for machine learning research

    Li Deng. The mnist database of handwritten digit images for machine learning research. IEEE Signal Processing Magazine , 29(6):141–142, 2012

  267. [275]

    Learning multiple layers of features from tiny images

    Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009

  268. [276]

    Core50: a new dataset and benchmark for continuous object recognition

    Vincenzo Lomonaco and Davide Maltoni. Core50: a new dataset and benchmark for continuous object recognition. In Conference on Robot Learning, pages 17–26. PMLR, 2017

  269. [277]

    Glue: A multi-task benchmark and analysis platform for natural language understanding

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. Glue: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv:1804.07461 , 2018

  270. [278]

    Superglue: Learning feature matching with graph neural networks

    Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabi- novich. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 4938–4947, 2020

  271. [279]

    Playing atari with deep reinforcement learning

    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602 , 2013

  272. [280]

    Open-ended learning leads to generally capable agents

    Open Ended Learning Team, Adam Stooke, Anuj Mahajan, Catarina Barros, Charlie Deck, Jakob Bauer, Jakub Sygnowski, Maja Trebacz, Max Jaderberg, Michael Mathieu, et al. Open-ended learning leads to generally capable agents. arXiv preprint arXiv:2107.12808 , 2021. 237

  273. [281]

    Vizdoom: A doom-based ai research platform for visual reinforce- ment learning

    Micha l Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Ja´ skowski. Vizdoom: A doom-based ai research platform for visual reinforce- ment learning. In 2016 IEEE conference on computational intelligence and games (CIG), pages 1–8. IEEE, 2016

  274. [282]

    Continual world: A robotic benchmark for continual reinforcement learn- ing

    Maciej Wo lczyk, Michal Zajkac, Razvan Pascanu, Lukasz Kuci’nski, and Piotr Milo’s. Continual world: A robotic benchmark for continual reinforcement learn- ing. In Neural Information Processing Systems, 2021

  275. [283]

    Cora: Benchmarks, baselines, and metrics as a platform for continual reinforce- ment learning agents

    Sam Powers, Eliot Xing, Eric Kolve, Roozbeh Mottaghi, and Abhinav Gupta. Cora: Benchmarks, baselines, and metrics as a platform for continual reinforce- ment learning agents. arXiv preprint arXiv:2110.10067 , 2021

  276. [284]

    Leveraging procedu- ral generation to benchmark reinforcement learning

    Karl Cobbe, Chris Hesse, Jacob Hilton, and John Schulman. Leveraging procedu- ral generation to benchmark reinforcement learning. In International conference on machine learning , pages 2048–2056. PMLR, 2020

  277. [285]

    Minihack the planet: A sandbox for open-ended reinforcement learn- ing research

    Mikayel Samvelyan, Robert Kirk, Vitaly Kurin, Jack Parker-Holder, Minqi Jiang, Eric Hambro, Fabio Petroni, Heinrich K¨ uttler, Edward Grefenstette, and Tim Rockt¨ aschel. Minihack the planet: A sandbox for open-ended reinforcement learn- ing research. arXiv preprint arXiv:2109...

  278. [286]

    Alfred: A benchmark for interpreting grounded instructions for everyday tasks

    Mohit Shridhar, Jesse Thomason, Daniel Gordon, Yonatan Bisk, Winson Han, Roozbeh Mottaghi, Luke Zettlemoyer, and Dieter Fox. Alfred: A benchmark for interpreting grounded instructions for everyday tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern ...

  279. [287]

    F-siol-310: A robotic dataset and benchmark for few-shot incremental object learning

    Ali Ayub and Alan R Wagner. F-siol-310: A robotic dataset and benchmark for few-shot incremental object learning. In 2021 IEEE International Conference on Robotics and Automation (ICRA) , pages 13496–13502. IEEE, 2021. 238

  280. [288]

    Openloris-object: A robotic vision dataset and benchmark for lifelong deep learning

    Qi She, Fan Feng, Xinyue Hao, Qihan Yang, Chuanlin Lan, Vincenzo Lomonaco, Xuesong Shi, Zhengwei Wang, Yao Guo, Yimin Zhang, et al. Openloris-object: A robotic vision dataset and benchmark for lifelong deep learning. In 2020 IEEE international conference on robotics and automa...

  281. [289]

    Architecture matters in contin- ual learning

    Seyed Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Timothy Nguyen, Razvan Pascanu, Dilan Gorur, and Mehrdad Farajtabar. Architecture matters in contin- ual learning. arXiv preprint arXiv:2202.00275 , 2022

  282. [290]

    Disentangling transfer in continual reinforcement learning

    Maciej Wo lczyk, Michal Zajkac, Razvan Pascanu, Lukasz Kuci’nski, and Pi- otr Milo’s. Disentangling transfer in continual reinforcement learning. ArXiv, abs/2209.13900, 2022

  283. [291]

    Memory efficient continual learning with transformers

    Beyza Ermis, Giovanni Zappella, Martin Wistuba, Aditya Rawal, and Cedric Archambeau. Memory efficient continual learning with transformers. Advances in Neural Information Processing Systems , 35:10629–10642, 2022

  284. [292]

    Meta-world: A benchmark and evaluation for multi- task and meta reinforcement learning

    Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta-world: A benchmark and evaluation for multi- task and meta reinforcement learning. In Conference on robot learning , pages 1094–1100. PMLR, 2020

  285. [293]

    Causalworld: A robotic manipulation benchmark for causal structure and transfer learning

    Ossama Ahmed, Frederik Tr¨ auble, Anirudh Goyal, Alexander Neitz, Yoshua Ben- gio, Bernhard Sch¨ olkopf, Manuel W¨ uthrich, and Stefan Bauer. Causalworld: A robotic manipulation benchmark for causal structure and transfer learning. arXiv preprint arXiv:2010.04296, 2020

  286. [294]

    Rl- bench: The robot learning benchmark & learning environment

    Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. Rl- bench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020. 239

  287. [295]

    Behavior-1k: A benchmark for embodied ai with 1,000 every- day activities and realistic simulation

    Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivas- tava, Roberto Mart´ ın-Mart´ ın, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al. Behavior-1k: A benchmark for embodied ai with 1,000 every- day activities and realistic simulation. In Confer...

  288. [296]

    Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations

    Tongzhou Mu, Zhan Ling, Fanbo Xiang, Derek Yang, Xuanlin Li, Stone Tao, Zhiao Huang, Zhiwei Jia, and Hao Su. Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations. arXiv preprint arXiv:2107.14483 , 2021

  289. [297]

    Maniskill2: A unified benchmark for generalizable manipulation skills

    Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al. Maniskill2: A unified benchmark for generalizable manipulation skills. arXiv preprint arXiv:2302.04659, 2023

  290. [298]

    Com- posuite: A compositional reinforcement learning benchmark

    Jorge A Mendez, Marcel Hussing, Meghna Gummadi, and Eric Eaton. Com- posuite: A compositional reinforcement learning benchmark. arXiv preprint arXiv:2207.04136, 2022

  291. [299]

    Research on human-robot coevolution for a richer soci- ety

    The University of Tokyo. Research on human-robot coevolution for a richer soci- ety. https://www.u-tokyo.ac.jp/adm/uci/en/projects/ai/project_00005. html, 2020. Accessed: 2025-03-30

  292. [300]

    Depth anything: Unleashing the power of large-scale unlabeled data

    Lihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu, Jiashi Feng, and Heng- shuang Zhao. Depth anything: Unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10371–10381, 2024

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.