REVIEW 4 major objections 5 minor 36 references
Generating Explanations for Autonomous Robots: a Systematic Review
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This review of 803 candidate papers, narrowed to 22, claims that robot-explainability research is dominated by text-based descriptive and causal explanations, centers on navigation in social robotics, and that only about half of the…
desk verdict A transparent systematic review with a real but correctable seed-set bias; the central qualitative claims likely hold, but the 'only half evaluate' number is an artifact of the selection procedure. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a systematic literature-mapping pipeline rather than a single mathematical object. An eight-paper semi-reference set supplies the keyword vocabulary and a validation check requiring that at least 70 percent of those seed papers appear in the search results; database-specific search strings, inclusion and exclusion criteria, and a quality checklist with a 7.0-out-of-10.0 threshold shrink 803 raw hits to a 22-paper corpus. A fixed extraction form then codes each paper for application domain, robot skill, environment, explanation type, explanation format, real-time generation, question-driven generation, and whether the explanations are evaluated.
What would settle it
Re-running the same research questions with an independently built keyword set, identical search strings across all five databases, and without the reviewers' own papers in the validation set would either reproduce the reported distributions—text-dominant, navigation-heavy, only half evaluating—or not; any large deviation would settle whether the current map is an artifact of the sample.
Extended reading notes
Core claim
The review's central claim is that the literature on generating explanations for autonomous robots is diverse but not yet consolidated. Among the 22 papers that survived the quality filter, the dominant pattern is a dialog-based system in which the robot answers a "why"-type question with a textual, descriptive or causal explanation; contrastive and multimodal explanations are rarer. Social robotics is the most studied application domain, navigation the most studied skill, and simulation the most common experimental environment. Most importantly, only 11 of the 22 papers propose any method for evaluating the explanations they generate, and the metrics that do appear are ad hoc, such as user satisfaction, response time, accuracy, and similarity. From this the paper concludes that the concept of explanation itself lacks standardization and that a common evaluation methodology is needed before the field can determine how adequate a robot's explanation is.
Load-bearing premise
The review's frequency claims rest on the assumption that its search-and-screening pipeline returned a representative sample of the eXplainable Autonomous Robot literature, an assumption anchored in a hand-picked eight-paper seed set, two of which were written by the reviewers themselves, and in search strings that differed from one database to the next.
Editorial extensions
If this is right
- New work on robot explanations can expect the default audience to be a user asking "why" in a dialogue, so designing for question-driven interaction will be easier to compare with the existing literature.
- Researchers wanting to make a distinctive contribution have room in multimodal and contrastive explanations, which appear far less often than plain text.
- Any claim that a robot explanation is good is currently contestable, because about half of the surveyed papers provide no evaluation and the metrics that do appear vary across studies.
- Because navigation and social robotics dominate, experiments that explain manipulation, perception, or decision-making test less charted ground.
- The absence of a shared taxonomy of explanation types means that future work should first classify its explanation before claiming novelty.
Reading between the lines
- The dominance of simulation experiments and the 2022 publication dip the authors attribute to pandemic restrictions suggest that some of the map's shape reflects practical constraints on testing with physical robots, not just scientific interest; a future review could separate these effects.
- If evaluation continues to be absent or ad hoc, the fastest lever for field progress is probably not a new explanation algorithm but a shared evaluation benchmark, since the review indicates the bottleneck is assessment rather than generation.
- Because the keyword vocabulary was derived from a small seed set that includes two papers by the review team, the map might shift if the same research questions were run from a different seed set; running two parallel reviews would quantify that sensitivity.
- The same coding scheme could be applied to adjacent domains such as autonomous vehicles or medical robots to test whether text-based, navigation-centered explanations are a general property of explainable autonomy or an artifact of the human-robot interaction and social-robotics literature.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a systematic literature mapping of methods for generating explanations in autonomous robots, following PRISMA and the recommendations of Kitchenham et al. From 803 initial records retrieved in five databases plus a semi-reference set, 22 papers survive duplicate removal, title/abstract and full-text screening, and a quality threshold of 7.0/10.0. The authors extract information about application domain, experimental environment, robot skills, explanation type and format, whether explanations are question-driven or real-time, and whether an evaluation method is proposed. The main descriptive findings are that social robotics is the most common domain, navigation is the most frequently explained skill, text-based descriptive and causal explanations dominate, and only about half of the included studies propose an evaluation method, leading to conclusions about the lack of shared definitions and robust evaluation methodology.
Significance. If the sample is representative, the review would be a useful and focused complement to broader XAI surveys, with a transparent PRISMA flow, documented search strings in Appendix A, a quality checklist, and detailed extraction tables. The distinction between question-driven explanation, real-time explanation, and having an explicit evaluation method is a useful analytic grid for organizing the XAR literature. However, the value of the review hinges entirely on selection validity, and the current methodology does not establish that the 22 papers are representative of the wider literature. The quantitative conclusions, especially the evaluation-gap claim, are sensitive to the inclusion of semi-reference seed papers, two of which are co-authored by the authors. These concerns are addressable, and the provided tables would support the needed sensitivity analysis.
major comments (4)
- [Section II-C3 (search validation) and Section II-E1 (IC1)] The search validation step is circular in a way that is load-bearing for the review's quantitative claims. Keywords and search strings are derived from the eight semi-reference papers (Section II-A), and the search is declared valid if at least 70% of that same set is retrieved. Because six of the eight semi-reference papers are automatically admitted through inclusion criterion IC1, the validation only shows that the queries recover papers that helped build them; it does not demonstrate coverage of the broader XAR literature. This matters because the frequency claims about domains, skills, explanation formats, and evaluation rates in Sections IV and V are all computed on the final set whose composition is partly fixed by this circular step.
- [Table 9 and Section IV (evaluation rate)] The statement that 'only half of the articles propose a method for evaluating the explanations' is not robust to the seed-set composition. Removing the six semi-reference papers that entered through IC1 (AD4, AD5, AD12, AD14, AD21, AD22) changes the evaluation rate from 11/22 (50%) to 10/16 (62.5%); removing only the two self-authored papers (AD12 and AD22), both coded False for evaluation, yields 11/20 (55%). Navigation dominance and textual-format dominance survive these exclusions, but the conclusion that the field particularly lacks evaluation methodology is materially weakened. The manuscript should report the PRISMA-mandated list of excluded full-text papers with reasons and provide a sensitivity analysis, since readers currently cannot assess how much of the landscape description is an artifact of the selection procedure.
- [Section II-F (quality threshold) and Section III-C] The quality cutoff of 7.0 out of 10.0 is presented as balancing 'quality and relevance' without a justification, and the threshold description is internally inconsistent: Section II-F says papers scoring 'equal to or below the threshold value' are rejected, while Section III-C says articles with a score 'lower than 7.0' were discarded. Seven papers are discarded at this stage, so the final set and every subsequent trend depend on this arbitrary parameter. The authors should justify the chosen threshold and test whether the main conclusions are stable under alternative thresholds, for example 6.0 or 8.0.
- [Section II-D and Appendix A (search strings)] The search strings are not uniform across databases. Search String 1 (WoS) has no mandatory explanation-generation terms, while the IEEE, Scopus, and Springer strings include terms such as 'providing explanation', 'explanation generation', 'generate explanations', and 'answering questions'. Consequently, the reported semi-reference recall is not a test of a single search protocol, and the uneven constraints may bias which papers are retrieved from each database. The authors should explain the rationale for these differences and discuss their possible effect on the sample and on the recall-based validation.
minor comments (5)
- [Section II-G2] The heading 'Skills del robot' is in Spanish and should be translated to English.
- [Table 6] The table contains typographical errors: 'Y es/No' should be 'Yes/No' and 'Real Word' should be 'Real World'.
- [Section III-A] The text states that 6 of 8 semi-reference papers correspond to 'approximately 70%', but the exact value is 75%; the wording should be corrected.
- [Section I and Section II-G1] There are typos in the full text: 'V arious' should be 'Various' and 'Spacial Robotics' should be 'Spatial Robotics'.
- [Section II-E] The phrase 'This criteria is defined' should be 'These criteria are defined'; additionally, 'criteria' is used as a singular noun in one place in Section II-C2.
Circularity Check
The semi-reference set builds, validates, and seeds the final sample, making the search validation and the 'only half evaluate' conclusion partly self-referential.
-
self definitional
[Section II-A and II-C3; applied in Section III-A]
"The semi-reference set serves several purposes. On one hand, it allows for the extraction of keywords and helps establish the direction of the research. On the other hand, it will serve as a validation method for the search process. ... This criterion requires that at least 70% of the semi-reference papers appear in the search results."
The same eight papers are both the source of the keywords and search strings (Sections II-B and II-D) and the target of the 70% recall check (Section II-C3). A search built from the seed set will naturally retrieve much of that same set, so the reported 'valid' result (6 of 8 recalled, Section III-A) is a self-consistency check, not an independent coverage test. The validation therefore confirms only that the queries reproduce the papers used to construct them; it does not establish coverage of the wider XAR literature.
-
other
[Section II-E1 (Table 3, IC1), Section III-B2, Section IV, Table 9]
"IC1 The paper is part of the semi-reference set ... Once again, all articles from the semi-reference set were accepted according to the inclusion criterion IC1. ... The data analysis indicates that only half of the articles propose a method for evaluating the explanations."
Six of the 22 final papers are auto-admitted because they are the very semi-reference papers used to derive keywords and validate the search (AD4, AD5, AD12, AD14, AD21, AD22); two of these (AD12, AD22) are co-authored by the reviewers. Since AD12 and AD22 are both coded as proposing no evaluation method (Table 9), the conclusion that 'only half of the articles' evaluate explanations is partly produced by the inclusion rule itself. Recomputing without the six auto-included seeds changes the evaluation rate from 11/22 (50%) to 10/16 (62.5%); removing only the two self-authored seeds gives 11/20 (55%). The frequency landscape in Section IV and the SQ3 answer in Section V are therefore not robust to the seed-set composition, even though the navigation and textual-format trends survive.
full rationale
The review's core descriptive results—text-based explanations dominate, navigation is the most frequently explained skill, and social robotics is the most common domain—are empirical codings of the 22 included papers and do not reduce to the search expressions; they are also broadly consistent with earlier XAI/XAR surveys. The circularity is localized to the review's self-referential selection machinery. The same eight-paper semi-reference set is used to extract keywords (Section II-B), to build the search strings (Section II-D), to validate the search via a 70% recall threshold (Section II-C3), and then to populate the final sample through inclusion criterion IC1 (Section II-E1). Because the search is validated against the very papers that supplied its terms, the 6/8 recall reported in Section III-A is a self-consistency check rather than independent evidence of coverage. The auto-inclusion of seed papers also matters numerically: six of the 22 final items are seeds, two of them co-authored by the reviewers, and both are coded as not proposing an evaluation method. The claim that only half of the analyzed articles evaluate explanations (Section IV, SQ3 in Section V) moves from 50% to 62.5% when the six auto-included seeds are removed, so a headline gap-claim is partially an artifact of the seed-set inclusion rule. This is genuine but partial circularity: the navigation and text-format trends survive seed removal, so the central landscape description still has independent content.
Assumptions & free parameters
free parameters (2)
- Quality threshold =
7.0 out of 10
- Publication year cutoff =
2015
assumptions (4)
- domain assumption Searching title, abstract, and keywords captures the relevant XAR literature.
- domain assumption The semi-reference set of 8 papers is representative of the XAR field.
- domain assumption The adapted quality checklist from Dybå and Dingsøyr is a valid measure of study quality and relevance.
- domain assumption The 70% semi-reference retrieval criterion is a sufficient validation of the search.
Cite this review
Pith. "Pith review of Generating Explanations for Autonomous Robots: a Systematic Review." pith.science (2026). https://pith.science/paper/FJY6TYOQ
@misc{pith2026241218516,
author = {Pith},
title = {Pith review of: Generating Explanations for Autonomous Robots: a Systematic Review},
year = {2026},
howpublished = {\url{https://pith.science/paper/FJY6TYOQ}},
note = {Machine review of arXiv:2412.18516}
}
read the original abstract
Building trust between humans and robots has long interested the robotics community. Various studies have aimed to clarify the factors that influence the development of user trust. In Human-Robot Interaction (HRI) environments, a critical aspect of trust development is the robot's ability to make its behavior understandable. The concept of an eXplainable Autonomous Robot (XAR) addresses this requirement. However, giving a robot self-explanatory abilities is a complex task. Robot behavior includes multiple skills and diverse subsystems. This complexity led to research into a wide range of methods for generating explanations about robot behavior. This paper presents a systematic literature review that analyzes existing strategies for generating explanations in robots and studies the current XAR trends. Results indicate promising advancements in explainability systems. However, these systems are still unable to fully cover the complex behavior of autonomous robots. Furthermore, we also identify a lack of consensus on the theoretical concept of explainability, and the need for a robust methodology to assess explainability methods and tools has been identified.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[2]
S. Kumar, S. Sarraf, A. K. Kar, and P . V . Ilavarasan, ‘‘A study of explainable artificial intelligence: A systematic literature review of the applications,’’ IoT, Big Data and AI for Improving Quality of Everyday Life: Present and Future Challenges: IOT, Data Science and Artificial Intelligence Technolo- gies, pp. 243–259, 2023
work page 2023
-
[3]
G. Vilone and L. Longo, ‘‘Explainable artificial intelligence: a systematic review,’’arXiv preprint arXiv:2006.00093, 2020
arXiv 2006
-
[4]
S. Wallkötter, S. Tulli, G. Castellano, A. Paiva, and M. Chetouani, ‘‘Ex- plainable embodied agents through social cues: a review,’’ ACM Transac- tions on Human-Robot Interaction (THRI) , vol. 10, no. 3, pp. 1–24, 2021
work page 2021
-
[5]
S. Anjomshoae, A. Najjar, D. Calvaresi, and K. Främling, ‘‘Explainable agents and robots: Results from a systematic literature review,’’ in 18th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2019), Montreal, Canada, May 13–17, 2019 . International Foundation for Autonomous Agents and Multiagent Systems, 2019, pp. 1078–1088
work page 2019
-
[6]
D. Moher, A. Liberati, J. Tetzlaff, and D. G. Altman, ‘‘Preferred reporting items for systematic reviews and meta-analyses: The PRISMA statement,’’ International Journal of Surgery , vol. 8, no. 5, pp. 336– 341, 2010. [Online]. Available: https://linkinghub.elsevier.com/retrieve/ pii/S1743919110000403
work page 2010
-
[7]
B. A. Kitchenham, D. Budgen, and P . Brereton, Evidence-Based Software Engineering and Systematic Reviews , 0th ed. Chapman and Hall/CRC, Nov. 2015. [Online]. Available: https://www.taylorfrancis.com/books/ 9781482228663
work page 2015
-
[8]
M. A. González-Santamarta, L. Fernández-Becerra, D. Sobrín-Hidalgo, A. M. Guerrero-Higueras, I. González, and F. J. R. Lera, Using Large Language Models for Interpreting Autonomous Robots Behaviors . Springer Nature Switzerland, 2023, p. 533–544. [Online]. Available: http://dx.doi.org/10.1007/978-3-031-40725-3_45
-
[9]
Z. Han, D. Giger, J. Allspaw, M. S. Lee, H. Admoni, and H. A. Y anco, ‘‘Building the Foundation of Robot Explanation Generation Using Behavior Trees,’’ ACM Transactions on Human-Robot Interaction , vol. 10, no. 3, pp. 26:1–26:31, Jul. 2021. [Online]. Available: https: //dl.acm.org/doi/10.1145/3457185
doi:10.1145/3457185 2021
Show all 36 references
-
[10]
Sakai and T
T. Sakai and T. Nagai, ‘‘Explainable autonomous robots: a survey and perspective,’’ Advanced Robotics , vol. 36, no. 5-6, pp. 219–238, Mar
-
[11]
Halilovic and F
A. Halilovic and F. Lindner, ‘‘Visuo-textual explanations of a robot’s navigational choices,’’ inCompanion of the 2023 ACM/IEEE International Conference on Human-Robot Interaction, 2023, pp. 531–535
2023
-
[12]
Arnold, D
T. Arnold, D. Kasenberg, and M. Scheutz, ‘‘Explaining in time: Meeting interactive standards of explanation for robotic systems,’’ J. Hum.-Robot Interact. , vol. 10, no. 3, jul 2021. [Online]. Available: https://doi.org/10.1145/3457183
2021 doi
-
[13]
Fernández-Becerra, M
L. Fernández-Becerra, M. A. González-Santamarta, D. Sobrín-Hidalgo, Á. M. Guerrero-Higueras, F. J. R. Lera, and V . M. Olivera, ‘‘Accountability and explainability in robotics: A proof of concept for ros 2-and nav2-based mobile robots,’’ in Computational Intelligence in Securi...
2023
-
[14]
Sakai, T
T. Sakai, T. Nagai, and K. Abe, ‘‘Implementation and Evaluation of Algorithms for Realizing Explainable Autonomous Robots,’’ IEEE Access, pp. 1–1, 2023. [Online]. Available: https://ieeexplore.ieee.org/ document/10210548/
2023
-
[15]
Borgo, M
R. Borgo, M. Cashmore, and D. Magazzeni, ‘‘Towards Providing Explanations for AI Planner Decisions,’’ 2018. [Online]. Available: https://arxiv.org/abs/1810.06338
2018 arXiv
-
[16]
Schardt, M
C. Schardt, M. B. Adams, T. Owens, S. Keitz, and P . Fontelo, ‘‘Utilization of the PICO framework to improve searching PubMed for clinical questions,’’BMC Medical Informatics and Decision Making , vol. 7, no. 1, p. 16, Dec. 2007. [Online]. Available: https://bmcmedinformdecism...
2007 doi
-
[17]
Zhang, M
H. Zhang, M. A. Babar, and P . Tell, ‘‘Identifying relevant studies in software engineering,’’ Information and Software Technology , vol. 53, no. 6, pp. 625–637, Jun. 2011. [Online]. Available: https://linkinghub. elsevier.com/retrieve/pii/S0950584910002260
2011
-
[18]
Dybå and T
T. Dybå and T. Dingsøyr, ‘‘Strength of evidence in systematic reviews in software engineering,’’ in Proceedings of the Second ACM-IEEE International Symposium on Empirical Software Engineering and Measurement, ser. ESEM ’08. New Y ork, NY , USA: Association for Computing Machi...
2008
-
[19]
Halilovic and S
A. Halilovic and S. Krivic, ‘‘The influence of a robot’s personality on real- time explanations of its navigation,’’ inInternational Conference on Social Robotics. Springer, 2023, pp. 133–147
2023
-
[20]
Han and H
Z. Han and H. Y anco, ‘‘Communicating missing causal information to explain a robot’s past behavior,’’ ACM Transactions on Human-Robot Interaction, vol. 12, no. 1, pp. 1–45, 2023
2023
-
[21]
Rosenthal, P
S. Rosenthal, P . Vichivanives, and E. Carter, ‘‘The impact of route descrip- tions on human expectations for robot navigation,’’ ACM Transactions on Human-Robot Interaction (THRI), vol. 11, no. 4, pp. 1–19, 2022
2022
-
[22]
F. Cruz, R. Dazeley, P . V amplew, and I. Moreira, ‘‘Explainable robotic systems: Understanding goal-driven actions in a reinforcement learning scenario,’’Neural Computing and Applications, vol. 35, no. 25, pp. 18 113– 18 130, 2023
2023
-
[23]
Hu and T
S. Hu and T. Nagai, ‘‘Explainable autonomous robots in continuous state space based on graph-structured world model,’’ Advanced Robotics, vol. 37, no. 16, pp. 1025–1041, 2023
2023
-
[24]
BOGA TARKAN and E
A. BOGA TARKAN and E. ERDEM, ‘‘Explanation generation for multi- modal multi-agent path finding with optimal resource utilization using answer set programming,’’ Theory and Practice of Logic Programming , vol. 20, no. 6, p. 974–989, 2020
2020
-
[25]
Mualla, I
Y . Mualla, I. Tchappi, T. Kampik, A. Najjar, D. Calvaresi, A. Abbas-Turki, S. Galland, and C. Nicolle, ‘‘The quest of parsimonious xai: A human-agent architecture for explanation formulation,’’ Artificial intelligence, vol. 302, p. 103573, 2022
2022
-
[26]
Zakershahrak, Z
M. Zakershahrak, Z. Gong, N. Sadassivam, and Y . Zhang, ‘‘Online expla- nation generation for planning tasks in human-robot teaming,’’ in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 6304–6310
2020
-
[27]
Mota and M
T. Mota and M. Sridharan, ‘‘Answer me this: constructing disambiguation queries for explanation generation in robotics,’’ in2021 IEEE International Conference on Development and Learning (ICDL). IEEE, 2021, pp. 1–8
2021
-
[28]
Gong and Y
Z. Gong and Y . Zhang, ‘‘Behavior explanation as intention signaling in human-robot teaming,’’ in 2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN) . IEEE, 2018, pp. 1005–1011
2018
-
[29]
Schroeter, F
N. Schroeter, F. Cruz, and S. Wermter, ‘‘Introspection-based explainable reinforcement learning in episodic and non-episodic scenarios,’’ 2022. [Online]. Available: https://arxiv.org/abs/2211.12930
2022 arXiv
-
[30]
X. Gao, R. Gong, Y . Zhao, S. Wang, T. Shu, and S.-C. Zhu, ‘‘Joint mind modeling for explanation generation in complex human-robot collaborative tasks,’’ in 2020 29th IEEE international conference on robot and human interactive communication (RO-MAN). IEEE, 2020, pp. 1119–1126
2020
-
[31]
Brandao, G
M. Brandao, G. Canal, S. Krivić, and D. Magazzeni, ‘‘Towards providing explanations for robot motion planning,’’ in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2021, pp. 3927– 3933
2021
-
[32]
S. Chen, K. Boggess, and L. Feng, ‘‘Towards transparent robotic planning via contrastive explanations,’’ in2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 6593–6598
2020
-
[33]
Olivares-Alarcos, S
A. Olivares-Alarcos, S. Foix, and G. Alenyà, ‘‘Knowledge representation for explainability in collaborative robotics and adaptation,’’ 2021
2021
-
[34]
Brandao, A
M. Brandao, A. Coles, and D. Magazzeni, ‘‘Explaining path plan opti- mality: Fast explanation methods for navigation meshes using full and incremental inverse optimization,’’ in Proceedings of the International Conference on Automated Planning and Scheduling, vol. 31, 2021, pp. 56– 64
2021
-
[35]
Rosenthal, S
S. Rosenthal, S. P . Selvaraj, and M. M. V eloso, ‘‘V erbalization: Narration of autonomous robot experience.’’ in IJCAI, vol. 16, 2016, pp. 862–868. VOLUME 11, 2023 13 Author et al.: Preparation of Papers for IEEE TRANSACTIONS and JOURNALS DAVID SOBRÍN-HIDALGO received the B....
2016
-
[2020]
His research interests include working with robotic systems and Large Language Models (LLMs)
During his master’s studies, he completed an external internship in 2020 with the Robotics Group of the University of Leon. His research interests include working with robotic systems and Large Language Models (LLMs). In 2022, he began his Ph.D. studies in the Pro- duction and...
2020
-
[2022]
Available: https://www.tandfonline.com/doi/full/10.1080/ 01691864.2022.2029720
[Online]. Available: https://www.tandfonline.com/doi/full/10.1080/ 01691864.2022.2029720
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.