REVIEW 3 cited by
Imitation Learning: Progress, Taxonomies and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Imitation learning aims to extract knowledge from human experts' demonstrations or artificially created agents in order to replicate their behaviors. Its success has been demonstrated in areas such as video games, autonomous driving, robotic simulations and object manipulation. However, this replicating process could be problematic, such as the performance is highly dependent on the demonstration quality, and most trained agents are limited to perform well in task-specific environments. In this survey, we provide a systematic review on imitation learning. We first introduce the background knowledge from development history and preliminaries, followed by presenting different taxonomies within Imitation Learning and key milestones of the field. We then detail challenges in learning strategies and present research opportunities with learning policy from suboptimal demonstration, voice instructions and other associated optimization schemes.
Forward citations
Cited by 3 Pith papers
-
Few-Shot Neuro-Symbolic Imitation Learning for Long-Horizon Planning and Acting
A neuro-symbolic system learns symbolic task rules and neural control policies from as few as five demonstrations and generalizes to larger unseen task instances.
-
Force-Based Robotic Imitation Learning: A Two-Phase Approach for Construction Assembly Tasks
Force-based human demonstrations, collected through a haptic VR-robot setup, improved simulated pipe-insertion policy learning over visual demonstrations.
-
Swarm Behavior Cloning
An ensemble of behavior-cloned policies trained with a hidden-activation alignment penalty produces more consistent actions and higher mean episode returns than standard ensemble behavior cloning.
Discussion (0). Continue with ORCID to comment.