{"id":"9a11a7f2-54c7-4ea3-a6d0-0536a88e046a","arxiv_id":"2506.13498","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that organizes imitation learning research for contact-rich robot tasks into teaching, learning, sensing, and application categories, and maps current challenges and future directions.","lead":"This paper reviews recent research on teaching robots contact-rich skills, such as assembly, wiping, and surgery, by having them imitate human demonstrations. It organizes the field into teaching methods, learning algorithms, sensing modalities, and application domains, and identifies data scarcity and simulation-to-reality transfer as key open challenges.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'first survey' claim is internally contradicted: Sec. 1 asserts no IL survey exists for contact-rich tasks, yet Sec. 5.2 cites An et al. (2025), a survey of imitation learning for dexterous manipulation, a canonical contact-rich domain.","rationale":"The reader's weakest assumption was that the selected papers are representative enough to support the survey's trend claims and first-survey status, given no search strategy or inclusion criteria. My stress-test identifies a more specific and more damaging version of the same concern: the negative novelty claim is not just unverified but internally contradicted by a prior survey cited in the paper itself. This strengthens the case for revision and re-scoping, but does not change the overall verdict: the survey remains a useful, readable organization of IL methods for contact-rich manipulation, and the appropriate outcome is CONDITIONAL acceptance after the first-survey claim is either substantiated with a reproducible search protocol and explicit comparison to prior surveys, or softened. I set agreement_with_reader to 'partial' because the reader's weakest assumption focused on representativeness and reproducibility, whereas my concern targets the factual correctness of the 'no survey exists' claim through an internal inconsistency; the two overlap but are not identical. The concrete test is designed to settle the concern directly: an independent check of An et al. (2025) and a bibliographic search would establish whether the paper is truly the first survey in this space or merely one of several overlapping surveys.","tokens_in":45296,"tokens_out":2912,"duration_ms":29506,"concrete_test":"Retrieve the full text of arXiv:2504.03515 (An et al., 2025) and code its scope: does it survey imitation learning algorithms applied to dexterous manipulation tasks with sustained physical contact (e.g., in-hand manipulation, grasping, insertion)? Then run a bibliographic search restricted to publications before 2025-06-16 with queries such as ('imitation learning' AND 'contact-rich' AND 'survey') across Scopus, Web of Science, and Semantic Scholar. If even one prior survey covers IL for contact-rich tasks, the Sec. 1 negative claim is false and must be revised and the contribution re-scoped. If An et al. (2025) excludes force/tactile contact-rich aspects and no other prior survey is found, the claim can stand with an explicit scope caveat and a description of the search protocol.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's novelty rests on the statement in Sec. 1 that 'No survey exists that investigates imitation learning research in contact-rich tasks.' This is a falsifiable negative claim, and the manuscript itself provides evidence against it. In Sec. 5.2, the authors cite An et al. (2025), 'Dexterous manipulation through imitation learning: A survey' (arXiv:2504.03515). Dexterous manipulation -- in-hand reorientation, grasping, insertion, and other multi-contact tasks -- is inherently contact-rich by the paper's own definition in Sec. 2.1 ('continuous and complex interactions between the robot and its environment, often requiring sophisticated control of forces'). If An et al. (2025) surveys imitation learning methods for such tasks, the 'no survey exists' claim is false as stated, and the central contribution becomes an incremental taxonomy rather than the first systematic organization of the field. The absence of any search strategy or inclusion criteria (noted by the reader) makes the negative claim even harder to defend: it is not merely unsupported, it is contradicted by the authors' own reference list. At minimum, the paper must delimit 'contact-rich tasks' so as to exclude dexterous manipulation, or explicitly distinguish its scope from An et al. (2025) and any similar prior survey.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper surveys imitation learning (IL) for contact-rich robotic tasks. It organizes the area into data collection (teaching methods and sensory modalities), learning algorithms (behavior cloning, dynamic movement primitives, generative methods, inverse RL, offline RL, and other approaches), available datasets, and application domains (industrial, household/service, and healthcare robots). The paper proposes a 2x2 taxonomy that crosses online/offline teaching with online/offline learning, and it claims to be the first survey specifically devoted to IL for contact-rich tasks.","tokens_in":45696,"tokens_out":3283,"duration_ms":33555,"significance":"If its scope and organization are made precise, the survey would be a useful entry point for researchers entering this subfield. Its strengths include a broad collection of recent work, a sensible separation of teaching and learning concepts, a useful review of DMP variants for contact-rich manipulation, a concrete list of public datasets, and attention to force and tactile modalities that are often underrepresented in general IL surveys. Several sections, particularly the DMP and multi-modal IL discussions, are informative and well grounded in the literature. However, the central novelty claim of being the first IL survey for contact-rich tasks is not adequately delimited or supported, and the absence of a reported search strategy weakens the reproducibility of the survey's coverage.","major_comments":[{"comment":"The claim that 'No survey exists that investigates imitation learning research in contact-rich tasks' is contradicted by the authors' own reference list: Sec. 5.2 cites An et al. (2025), a survey of imitation learning for dexterous manipulation (arXiv:2504.03515). Under the paper's own Sec. 2.1 definition of contact-rich manipulation as involving continuous and complex physical interactions and sophisticated force control, dexterous manipulation is a canonical contact-rich domain. The authors must either delimit 'contact-rich tasks' explicitly to exclude dexterous manipulation or explain how the present survey differs from An et al. (2025); otherwise the paper's central contribution is not the first survey but an overlapping taxonomy.","section":"Sec. 1"},{"comment":"No search strategy, database list, inclusion or exclusion criteria, or quality filter is described anywhere in the paper. The contributions list in Sec. 1 promises a 'systematic organization' of existing research, but the reader cannot verify that the selection of papers in Secs. 3-5 is representative or comprehensive. The claim to be the first survey is a negative claim that is especially sensitive to this omission: without a defined search window and inclusion criteria, the paper cannot support either its first-survey status or its trend statements. I recommend adding a methodology subsection or explicitly characterizing the coverage as a selective overview rather than a systematic survey.","section":"Secs. 3-5"},{"comment":"Sections 4.4.1 (Adversarial Imitation Learning) and 4.4.2 (Generative Adversarial Imitation Learning) are near-verbatim duplicates. Both describe the same GAN-based discriminator/generator framework, both cite Ho and Ermon (2016), and both claim effectiveness on pick-and-place and assembly tasks with the same citation (Li and Zou, 2023). Additionally, Sec. 3.2 classifies GAIL as an example of 'demo-augmented reinforcement learning' with online learning, while Sec. 4.4.2 treats GAIL as an IRL variant. This duplication and inconsistent classification directly weaken the survey's organizational contribution and must be corrected.","section":"Secs. 4.4.1 and 4.4.2"}],"minor_comments":[{"comment":"The citation 'Neville, 1985' appears to be Neville Hogan's impedance control paper, but the author name is given as 'Neville H' rather than the standard 'Hogan, N.'; this should be corrected in both the text and the reference list.","section":"Sec. 2.1 and reference list"},{"comment":"The reference 'PJ and BDO, 1971' is not a usable bibliographic entry; it should be replaced with the actual authors and title of the inverse optimal control paper being cited.","section":"Sec. 4.4"},{"comment":"The headings contain stray spacing: 'V ariational AutoEncoder' and 'F oundation models' should be 'Variational AutoEncoder' and 'Foundation models'.","section":"Secs. 4.3.1-4.3.2"},{"comment":"The offline reinforcement learning section is not, by itself, imitation learning; a few sentences explaining why demo-initialized or demo-augmented offline RL methods are included in an IL survey would clarify the scope, especially since Sec. 3.2 already distinguishes demo-augmented RL from direct imitation.","section":"Sec. 4.6"},{"comment":"Several in-text citations lack page numbers or venue details (e.g., 'Englert and Toussaint, 2017' has a DOI but no venue; 'PJ and BDO, 1971' is incomplete). A final pass to standardize the reference format would improve usability.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The main risk for the editor is the paper's framing as the first survey of imitation learning for contact-rich tasks. This negative claim is load-bearing and is contradicted by a survey already cited in the authors' own bibliography. The paper has genuine value as a selective overview, especially in the DMP and multimodal sections, but the authors need to either narrow the claimed scope or provide a meaningful comparison with prior IL surveys of dexterous/contact-rich manipulation. The self-citation density is not disqualifying, but in the absence of an inclusion methodology it makes the selection appear less neutral than a systematic survey would require."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time if you need a map of imitation learning for contact-rich manipulation, but the paper overstates its novelty. The central claim in Sec. 1 that no survey exists on IL for contact-rich tasks is false as stated. Sec. 5.2 cites An et al. (2025), a survey of imitation learning for dexterous manipulation — a domain that is contact-rich by the paper's own definition in Sec. 2.1. The authors need to either scope the claim or acknowledge An et al. and explain the difference.\n\nWhat works: the 2x2 taxonomy (online/offline teaching crossed with online/offline learning) is a decent organizing frame and would help a newcomer. The DMP section is thorough and correctly distinguishes reformulation vs. parallelization. The multi-modal IL section covers force, tactile, vision, and language with reasonable depth. Most summaries are accurate.\n\nSoft spots: the duplicated AIL/GAIL subsections (4.4.1 and 4.4.2) are nearly verbatim — one of them should go. The malformed citations are embarrassing: 'Neville 1985' should be Hogan's impedance control paper, and 'PJ and BDOA 1971' is probably Kalman's inverse optimal control work; both need fixing. No search strategy or inclusion criteria is given, so 'systematic organization' overstates what is an editorial selection. This also makes the negative claim about prior surveys harder to defend.\n\nNone of this is fatal. The survey covers the right ground and would be a solid reference if the authors fix the first-survey claim, delete the duplicate, and clean the bibliography. I would not desk-reject it; a serious referee can deal with these issues in a revision. But I would not cite it in its current form.","headline":"A useful, readable survey of imitation learning for contact-rich tasks, but its central 'first survey' claim is contradicted by its own reference list.","tokens_in":46079,"tokens_out":1973,"would_cite":false,"duration_ms":17836,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey maps imitation learning for contact-rich robot tasks, arguing that force and touch, not just vision, carry the skill and that no prior survey has covered this intersection.","keywords":["imitation learning","contact-rich manipulation","learning from demonstration","tactile sensing","dynamic movement primitives","behavior cloning","foundation models","robot manipulation survey"],"falsifier":"A reproducible literature search with explicit inclusion and exclusion criteria for surveys covering both imitation learning and contact-rich manipulation published before June 2025 would settle the first-survey claim: if it returns a prior survey with the same scope, the central claim fails, and if it returns none, the claim stands. A second check would take the survey's own corpus and test whether its trend statements—for example, that behavior cloning and foundation models dominate recent contact-rich work—survive a quantitative tally over the full set of cited papers.","tokens_in":45079,"feed_emoji":"🤖","tokens_out":5496,"duration_ms":54661,"temperature":0.7,"pith_summary":"The paper claims to be the first survey devoted specifically to imitation learning for contact-rich robotic tasks, and it organizes the field into a map of demonstration collection, learning algorithms, datasets, and applications. A sympathetic reader would care because contact-rich skills—assembly, polishing, surgery, household chores—are where robots still fail, and human demonstrations encode tacit force and compliance knowledge that is hard to specify by hand. The survey's organizing insight is that these tasks are nonlinear, sensitive to tiny positional deviations, and only partially visible through cameras, so force and tactile feedback are not optional extras but necessary complements to vision. If the map is accurate, it gives researchers and practitioners a common vocabulary for choosing teaching methods and learning algorithms, and it clarifies where the field's real bottlenecks lie.","feed_headline":"First survey maps imitation learning for contact-rich robot tasks","feed_subtitle":"Why it matters: force and touch, not just vision, carry the skill—and the map shows where methods still fall short.","key_machinery":"The load-bearing organizing object is the taxonomy of teaching and learning. Online versus offline teaching describes where demonstration trajectories come from—directly operating the robot, remote control, virtual reality, or sensor observation of human movement—while online versus offline learning describes when the policy updates, either during execution with feedback or from a stored dataset. Crossing these two dichotomies produces the four method families that structure the entire survey, and this frame carries the argument that imitation learning for contact-rich tasks is not a single technique but a design space. The companion mechanism is the modality argument: because contact-rich tasks are partially observable under vision and the contact point is often occluded, force and tactile sensing must be added to position and vision, and the dataset and application sections show how those modalities enter different learning algorithms.","core_discovery":"The survey's central claim is that no prior survey has investigated imitation learning for contact-rich tasks, and that this paper fills that gap by systematically organizing current research along two axes: how demonstrations are collected (online teaching, such as kinesthetic teaching, teleoperation, and VR-based teaching, versus offline observation of human movement) and when policies are learned (online during execution or offline from stored data). Crossing these axes yields four named categories—interactive imitation, demo-augmented reinforcement learning, direct imitation, and observational learning—which the paper uses to classify methods throughout. It further argues that contact-rich imitation learning is fundamentally multimodal: position, force, vision, and tactile signals each contribute different information, and fusion of several modalities is needed because interaction forces cannot be inferred from images alone. The survey also reviews datasets and benchmarks, showing that large general manipulation datasets exist but that touch-specific data remain smaller and mostly research-grade, and it identifies three future directions: dual-process hierarchical architectures, multimodal sensing, and improved simulation-to-reality transfer.","pith_inferences":["My inference: if the first-survey claim is correct, then the field's most pressing missing artifact is a shared evaluation benchmark for contact-rich imitation learning, and the datasets and applications collected here could serve as a seed for building one.","My inference: the taxonomy suggests a testable scaling relationship—methods that combine force observation during teaching with online correction during learning should outperform vision-only behavior cloning on precision tasks such as peg-in-hole with micrometer clearance.","My inference: continued progress in foundation models is likely to shift the bottleneck from policy learning to demonstration capture, making cheap force and tactile teleoperation hardware at least as important as the learning algorithms themselves."],"forward_implications":["The four-way teaching/learning taxonomy gives a practitioner a direct way to select a method: use interactive imitation when real-time corrections are available, direct or observational learning when working from stored data, and adversarial or demo-augmented reinforcement learning when reward design is the obstacle.","The modality analysis implies that adding force or tactile sensing to a vision-only demonstration pipeline is a necessity rather than an enhancement for tasks where contact forces determine success, such as assembly or surgery, because occlusion hides the contact interface.","The dataset review anchors expectations about scale: large trajectories exist for general manipulation, but contact-specific tactile datasets are smaller and mostly research-based, meaning data availability, not algorithm choice, may be the practical ceiling for contact-rich imitation learning.","The survey's stated future directions—dual-process hierarchical architectures, multimodal sensing, and sim-to-real transfer—identify where the next advances are likely to come from if the current trend claims are correct."],"supporting_citations":[{"why":"The closest prior survey; it reviews reinforcement learning rather than imitation learning for contact-rich tasks, establishing the gap this survey claims to fill.","marker":"Elguea-Aguinaco et al. (2023)"},{"why":"Foundational survey of robot learning from demonstration that established key paradigms and methodologies for the field.","marker":"Argall et al. (2009)"},{"why":"Survey specifically addressing imitation learning in robotic manipulation, synthesizing approaches and challenges in transferring human skills to robots.","marker":"Fang et al. (2019)"},{"why":"Explores the interconnections between reinforcement learning, imitation learning, and transfer learning, framing the relative roles of each.","marker":"Hua et al. (2021)"},{"why":"Comprehensive survey of interactive imitation learning, emphasizing the role of human-robot interaction in skill acquisition.","marker":"Celemin et al. (2022)"},{"why":"Survey on deep generative models for learning from multimodal demonstrations, supplying the multimodal and generative backdrop.","marker":"Urain et al. (2024)"},{"why":"Systematic survey of robot manipulation in contact, showing that contact-rich manipulation surveys exist but do not focus on imitation learning.","marker":"Suomalainen et al. (2022)"},{"why":"Extensive review of robot learning for manipulation, providing the structured analysis of challenges and representations that this survey builds upon.","marker":"Kroemer et al. (2021)"}],"fun_headline_variants":["First map of imitation learning for contact-rich tasks","Imitation learning for contact-rich tasks: the missing survey","Survey: contact-rich imitation learning needs force and touch","Contact-rich imitation learning: first systematic review","Imitation learning meets tactile robotics: a survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey assumes that the papers it selected in Sections 3 through 5 are representative enough of the broader field to support its trend claims and its status as the first survey, yet it does not describe any reproducible search strategy, inclusion criteria, or quality filter.","fun_headline_variants_meta":{"raw":{"variants":["First map of imitation learning for contact-rich tasks","Imitation learning for contact-rich tasks: the missing survey","Survey: contact-rich imitation learning needs force and touch","Contact-rich imitation learning: first systematic review","Imitation learning meets tactile robotics: a survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000187,"raw_usage":{"total_tokens":1280,"prompt_tokens":851,"completion_tokens":429,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":467,"completion_tokens_details":{"reasoning_tokens":356}},"tokens_in":467,"tokens_out":429,"duration_ms":4151,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:59:41.146990+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reproducible literature search with explicit inclusion and exclusion criteria for surveys covering both imitation learning and contact-rich manipulation published before June 2025 would settle the first-survey claim: if it returns a prior survey with the same scope, the central claim fails, and if it returns none, the claim stands. A second check would take the survey's own corpus and test whether its trend statements—for example, that behavior cloning and foundation models dominate recent contact-rich work—survive a quantitative tally over the full set of cited papers.","supporting_citations":[],"review_version":2}