{"id":"7f055c07-d540-4754-bba1-6ad8533fc7d0","arxiv_id":"2412.18516","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A systematic review of 22 papers maps explanation-generation methods for autonomous robots and finds no standardized way to evaluate them.","lead":"This paper reviews 22 studies on how autonomous robots explain their behavior, mapping the types of explanations, the skills explained, and how explanations are tested. It finds that text-based explanations for navigation dominate but that standardized evaluation methods are missing, a gap that matters for anyone building robots people can trust.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Seed-set inclusion biases the quantitative trends: removing the six auto-included semi-reference papers raises the evaluation rate from 50% to 62.5%, so the 'only half evaluate' claim is not robust.","rationale":"The reader's weakest assumption is that the search and selection process yields a representative sample; my analysis locates a concrete, testable failure mode inside that assumption: the automatic inclusion of six non-random semi-reference papers. The manuscript's own methodology section makes this visible: Section II-A describes the semi-reference construction, Section II-C3 uses it for validation, and Section II-E1 (IC1) forces it into the final set. The quantitative claims in Section IV are then computed on a sample where 27% of the items were preselected rather than discovered. I verified from Tables 8-9 that the evaluation-rate finding is the most affected: full set 11/22 (50%), non-seed subset 10/16 (62.5%), and the self-authored seed exclusion alone gives 11/20 (55%). The other main trends (textual dominance, navigation as most explained) persist after removing the seeds, so the central map of the field is probably qualitatively correct. For this reason I do not ask for a harsher verdict than CONDITIONAL; the paper should add a sensitivity analysis or qualify the quantitative statements. I find no basis for rejection or for claiming the authors were methodologically deceptive: the procedure is transparent and reproducible, but the inference from this sample to the field is weaker than the text implies. A single re-computation of the tables with and without the seed papers would settle whether the concern is material.","tokens_in":16582,"tokens_out":8111,"duration_ms":69471,"concrete_test":"Recompute the Section IV frequency tables after deleting the six semi-reference final papers (AD4, AD5, AD12, AD14, AD21, AD22), and separately after deleting only the two self-authored seeds (AD12, AD22). Compare the evaluation-method proportion (reported as 11/22 = 50%) and the skill/format shares with the full-set values. If the evaluation proportion rises above 55% in either variant, the conclusion that only about half of the literature evaluates explanations should be reworded or qualified in Section V; if it stays near 50%, the seed-set concern is not load-bearing for that claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's central frequency claims rest on a 22-paper final set in which 6 papers (AD4, AD5, AD12, AD14, AD21, AD22) are semi-reference items automatically accepted through inclusion criterion IC1 (Section II-E1). The semi-reference set was not randomly sampled: it was assembled from a preliminary search, authors' prior work, and colleague recommendations (Section II-A), and two of its items are co-authored by the reviewers. Because the same set was used to derive keywords and to validate the search (Section II-C3, 70% recall threshold), the validation confirms only that the queries recover papers that helped build the queries; it does not establish coverage of the wider XAR literature. This makes the Section IV frequencies and Section V conclusions sensitive to seed-set composition. A concrete symptom: Table 9 shows 11/22 papers (50%) proposing an evaluation method. Removing the six seed papers leaves 16 papers of which 10 (62.5%) propose evaluation; removing only the two self-authored seeds (AD12, AD22, both coded False) gives 11/20 (55%). Navigation and textual-format dominance survive this removal, but the 'only about half evaluate' claim and the related claim of a missing shared evaluation methodology are materially weaker once the preselected seeds are excluded. The paper does not report an excluded-papers list or a sensitivity analysis, so readers cannot tell how much of the landscape description is an artifact of the selection procedure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a systematic literature mapping of methods for generating explanations in autonomous robots, following PRISMA and the recommendations of Kitchenham et al. From 803 initial records retrieved in five databases plus a semi-reference set, 22 papers survive duplicate removal, title/abstract and full-text screening, and a quality threshold of 7.0/10.0. The authors extract information about application domain, experimental environment, robot skills, explanation type and format, whether explanations are question-driven or real-time, and whether an evaluation method is proposed. The main descriptive findings are that social robotics is the most common domain, navigation is the most frequently explained skill, text-based descriptive and causal explanations dominate, and only about half of the included studies propose an evaluation method, leading to conclusions about the lack of shared definitions and robust evaluation methodology.","tokens_in":16881,"tokens_out":6025,"duration_ms":51420,"significance":"If the sample is representative, the review would be a useful and focused complement to broader XAI surveys, with a transparent PRISMA flow, documented search strings in Appendix A, a quality checklist, and detailed extraction tables. The distinction between question-driven explanation, real-time explanation, and having an explicit evaluation method is a useful analytic grid for organizing the XAR literature. However, the value of the review hinges entirely on selection validity, and the current methodology does not establish that the 22 papers are representative of the wider literature. The quantitative conclusions, especially the evaluation-gap claim, are sensitive to the inclusion of semi-reference seed papers, two of which are co-authored by the authors. These concerns are addressable, and the provided tables would support the needed sensitivity analysis.","major_comments":[{"comment":"The search validation step is circular in a way that is load-bearing for the review's quantitative claims. Keywords and search strings are derived from the eight semi-reference papers (Section II-A), and the search is declared valid if at least 70% of that same set is retrieved. Because six of the eight semi-reference papers are automatically admitted through inclusion criterion IC1, the validation only shows that the queries recover papers that helped build them; it does not demonstrate coverage of the broader XAR literature. This matters because the frequency claims about domains, skills, explanation formats, and evaluation rates in Sections IV and V are all computed on the final set whose composition is partly fixed by this circular step.","section":"Section II-C3 (search validation) and Section II-E1 (IC1)"},{"comment":"The statement that 'only half of the articles propose a method for evaluating the explanations' is not robust to the seed-set composition. Removing the six semi-reference papers that entered through IC1 (AD4, AD5, AD12, AD14, AD21, AD22) changes the evaluation rate from 11/22 (50%) to 10/16 (62.5%); removing only the two self-authored papers (AD12 and AD22), both coded False for evaluation, yields 11/20 (55%). Navigation dominance and textual-format dominance survive these exclusions, but the conclusion that the field particularly lacks evaluation methodology is materially weakened. The manuscript should report the PRISMA-mandated list of excluded full-text papers with reasons and provide a sensitivity analysis, since readers currently cannot assess how much of the landscape description is an artifact of the selection procedure.","section":"Table 9 and Section IV (evaluation rate)"},{"comment":"The quality cutoff of 7.0 out of 10.0 is presented as balancing 'quality and relevance' without a justification, and the threshold description is internally inconsistent: Section II-F says papers scoring 'equal to or below the threshold value' are rejected, while Section III-C says articles with a score 'lower than 7.0' were discarded. Seven papers are discarded at this stage, so the final set and every subsequent trend depend on this arbitrary parameter. The authors should justify the chosen threshold and test whether the main conclusions are stable under alternative thresholds, for example 6.0 or 8.0.","section":"Section II-F (quality threshold) and Section III-C"},{"comment":"The search strings are not uniform across databases. Search String 1 (WoS) has no mandatory explanation-generation terms, while the IEEE, Scopus, and Springer strings include terms such as 'providing explanation', 'explanation generation', 'generate explanations', and 'answering questions'. Consequently, the reported semi-reference recall is not a test of a single search protocol, and the uneven constraints may bias which papers are retrieved from each database. The authors should explain the rationale for these differences and discuss their possible effect on the sample and on the recall-based validation.","section":"Section II-D and Appendix A (search strings)"}],"minor_comments":[{"comment":"The heading 'Skills del robot' is in Spanish and should be translated to English.","section":"Section II-G2"},{"comment":"The table contains typographical errors: 'Y es/No' should be 'Yes/No' and 'Real Word' should be 'Real World'.","section":"Table 6"},{"comment":"The text states that 6 of 8 semi-reference papers correspond to 'approximately 70%', but the exact value is 75%; the wording should be corrected.","section":"Section III-A"},{"comment":"There are typos in the full text: 'V arious' should be 'Various' and 'Spacial Robotics' should be 'Spatial Robotics'.","section":"Section I and Section II-G1"},{"comment":"The phrase 'This criteria is defined' should be 'These criteria are defined'; additionally, 'criteria' is used as a singular noun in one place in Section II-C2.","section":"Section II-E"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of a robotics venue and provides a useful set of extraction tables. The seed-set inclusion and circular validation are the decisive methodological concerns; they appear to be design flaws rather than deliberate bias, but the 'only half evaluate' conclusion is likely to be cited and should not be published without a robustness check. If the authors provide an excluded-papers list and a sensitivity analysis, or alternatively soften the evaluation-gap claim to reflect the sensitivity, the paper could become acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent, transparent systematic review of methods for generating explanations in autonomous robots. The new contribution is the 22-paper map, the extraction tables, and the PRISMA-based protocol. The qualitative conclusions—text explanations dominate, navigation is the most explained skill, and the field lacks a shared evaluation methodology—are plausible and align with earlier surveys. The paper does several things well: search strings are given in full, the PRISMA diagram is standard, and the authors disclose their semi-reference set instead of hiding it.\n\nThe soft spots are the ones flagged in the stress test, and they are real. The semi-reference set was assembled from a preliminary search, the authors' prior work, and colleague tips, then used both to derive keywords and to validate the search. That is circular. Two of the eight seeds are co-authored by the reviewers, and those two end up in the final set. Because IC1 auto-accepts all seeds, the quantitative frequencies in Section IV are sensitive to seed-set composition. The evaluation-rate claim is the clearest symptom: 11/22 papers (50%) propose an evaluation method in Table 9, but removing the six seed papers leaves 16 papers with 10 evaluations (62.5%); removing only the two self-authored seeds gives 11/20 (55%). Navigation and textual-format dominance survive the removal, but the 'only half evaluate' claim—which is load-bearing for the 'lack of methodology' conclusion—is materially weaker. The absent excluded-papers list and sensitivity analysis make it hard for a reader to judge how much of the landscape description is selection artifact.\n\nThat said, I don't think the central qualitative conclusion collapses. Even at 62.5%, the evaluation methods are heterogeneous and mostly ad hoc; the lack of a standardized methodology remains a fair reading of the evidence. The 7.0 quality threshold is arbitrary but not obviously harmful, and the divergent search strings across databases are a coverage concern, not a fatal one.\n\nWho this is for: someone entering XAR research or looking for a starting map of explanation-generation methods. It's not a practice-changer, but it's a serviceable survey. I'd send it to peer review with a request for major revision: report a sensitivity analysis, list excluded papers, justify the threshold, and soften the 'half' wording. The paper's honest about its process, and the core findings are likely correct, so it deserves referee time.","headline":"A transparent systematic review with a real but correctable seed-set bias; the central qualitative claims likely hold, but the 'only half evaluate' number is an artifact of the selection procedure.","tokens_in":17380,"tokens_out":4323,"would_cite":true,"duration_ms":37885,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review of 803 candidate papers, narrowed to 22, claims that robot-explainability research is dominated by text-based descriptive and causal explanations, centers on navigation in social robotics, and that only about half of the…","keywords":["explainability","eXplainable Autonomous Robot (XAR)","human-robot interaction","explanation generation","systematic literature review","robotics","trustworthy AI","survey"],"falsifier":"Re-running the same research questions with an independently built keyword set, identical search strings across all five databases, and without the reviewers' own papers in the validation set would either reproduce the reported distributions—text-dominant, navigation-heavy, only half evaluating—or not; any large deviation would settle whether the current map is an artifact of the sample.","tokens_in":16416,"feed_emoji":"🤖","tokens_out":6401,"duration_ms":56105,"temperature":0.7,"pith_summary":"The paper tries to establish a map of the current research on eXplainable Autonomous Robots (XAR): how researchers generate explanations for robot behavior, what those explanations look like, and where the open gaps are. By running a structured search across five literature databases, screening 803 candidate papers down to 22, and coding each final paper with a fixed extraction form, the authors claim that text-based descriptive and causal explanations dominate, that social robotics is the most common application domain, and that navigation is the most frequently explained robot skill. They also claim that only about half of the studies evaluate the explanations they produce, and that the field lacks both a shared definition of explainability and a standard methodology for assessing explanation quality. A sympathetic reader would take this as evidence that explanation generation for robots is a real but immature area: many techniques exist, yet the field has not agreed on what counts as a good explanation or how to measure it.","feed_headline":"Half of robot-explainability studies skip evaluation","feed_subtitle":"A 22-paper systematic review maps the field: text explanations lead, navigation dominates, and quality checks are missing.","key_machinery":"The load-bearing machinery is a systematic literature-mapping pipeline rather than a single mathematical object. An eight-paper semi-reference set supplies the keyword vocabulary and a validation check requiring that at least 70 percent of those seed papers appear in the search results; database-specific search strings, inclusion and exclusion criteria, and a quality checklist with a 7.0-out-of-10.0 threshold shrink 803 raw hits to a 22-paper corpus. A fixed extraction form then codes each paper for application domain, robot skill, environment, explanation type, explanation format, real-time generation, question-driven generation, and whether the explanations are evaluated.","core_discovery":"The review's central claim is that the literature on generating explanations for autonomous robots is diverse but not yet consolidated. Among the 22 papers that survived the quality filter, the dominant pattern is a dialog-based system in which the robot answers a \"why\"-type question with a textual, descriptive or causal explanation; contrastive and multimodal explanations are rarer. Social robotics is the most studied application domain, navigation the most studied skill, and simulation the most common experimental environment. Most importantly, only 11 of the 22 papers propose any method for evaluating the explanations they generate, and the metrics that do appear are ad hoc, such as user satisfaction, response time, accuracy, and similarity. From this the paper concludes that the concept of explanation itself lacks standardization and that a common evaluation methodology is needed before the field can determine how adequate a robot's explanation is.","pith_inferences":["The dominance of simulation experiments and the 2022 publication dip the authors attribute to pandemic restrictions suggest that some of the map's shape reflects practical constraints on testing with physical robots, not just scientific interest; a future review could separate these effects.","If evaluation continues to be absent or ad hoc, the fastest lever for field progress is probably not a new explanation algorithm but a shared evaluation benchmark, since the review indicates the bottleneck is assessment rather than generation.","Because the keyword vocabulary was derived from a small seed set that includes two papers by the review team, the map might shift if the same research questions were run from a different seed set; running two parallel reviews would quantify that sensitivity.","The same coding scheme could be applied to adjacent domains such as autonomous vehicles or medical robots to test whether text-based, navigation-centered explanations are a general property of explainable autonomy or an artifact of the human-robot interaction and social-robotics literature."],"forward_implications":["New work on robot explanations can expect the default audience to be a user asking \"why\" in a dialogue, so designing for question-driven interaction will be easier to compare with the existing literature.","Researchers wanting to make a distinctive contribution have room in multimodal and contrastive explanations, which appear far less often than plain text.","Any claim that a robot explanation is good is currently contestable, because about half of the surveyed papers provide no evaluation and the metrics that do appear vary across studies.","Because navigation and social robotics dominate, experiments that explain manipulation, perception, or decision-making test less charted ground.","The absence of a shared taxonomy of explanation types means that future work should first classify its explanation before claiming novelty."],"supporting_citations":[{"why":"Defines the eXplainable Autonomous Robot (XAR) concept that the review takes as its object of study.","marker":"[1]"},{"why":"Prior systematic review of explainable agents and robots that this review positions itself against, establishing the gap of a robotics-focused analysis.","marker":"[5]"},{"why":"Supplies the standard reporting guideline whose stages structure the identification, screening, and eligibility phases.","marker":"[6]"},{"why":"Supplies the systematic-review and literature-mapping methodology the paper follows.","marker":"[7]"},{"why":"Supplies the 70-percent rule used to judge whether the search results recover enough of the semi-reference set.","marker":"[17]"},{"why":"Supplies the original quality checklist that the paper adapts into its 7.0-out-of-10.0 inclusion threshold.","marker":"[18]"}],"fun_headline_variants":["Most robot-explainability papers skip evaluation","Robot explanation research: evaluation rare","Half of robot explainability studies lack testing","Systematic review: robot explanations poorly evaluated"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review's frequency claims rest on the assumption that its search-and-screening pipeline returned a representative sample of the eXplainable Autonomous Robot literature, an assumption anchored in a hand-picked eight-paper seed set, two of which were written by the reviewers themselves, and in search strings that differed from one database to the next.","fun_headline_variants_meta":{"raw":{"variants":["Most robot-explainability papers skip evaluation","Robot explanation research: evaluation rare","Half of robot explainability studies lack testing","Systematic review: robot explanations poorly evaluated"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1413,"prompt_tokens":881,"completion_tokens":532,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":497,"completion_tokens_details":{"reasoning_tokens":479}},"tokens_in":497,"tokens_out":532,"duration_ms":5098,"temperature":1.0,"reasoning_tokens":479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:43:01.205566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-running the same research questions with an independently built keyword set, identical search strings across all five databases, and without the reviewers' own papers in the validation set would either reproduce the reported distributions—text-dominant, navigation-heavy, only half evaluating—or not; any large deviation would settle whether the current map is an artifact of the sample.","supporting_citations":[{"cited_title":"Anjomshoae, A","cited_arxiv_id":null,"evidence_quote":"Prior systematic review of explainable agents and robots that this review positions itself against, establishing the gap of a robotics-focused analysis."},{"cited_title":"Moher, A","cited_arxiv_id":null,"evidence_quote":"Supplies the standard reporting guideline whose stages structure the identification, screening, and eligibility phases."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the systematic-review and literature-mapping methodology the paper follows."},{"cited_title":"Zhang, M","cited_arxiv_id":null,"evidence_quote":"Supplies the 70-percent rule used to judge whether the search results recover enough of the semi-reference set."},{"cited_title":"Dybå and T","cited_arxiv_id":null,"evidence_quote":"Supplies the original quality checklist that the paper adapts into its 7.0-out-of-10.0 inclusion threshold."}],"review_version":1}