Pith. sign in

REVIEW 4 major objections 5 minor 27 references

"What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read This paper claims that a multimodal interpreter which narrates and visualizes an autonomous vehicle's actions raises passenger trust, with average trust gains above 8% and ordinary-driving trust rising by up to 30%.

desk verdict The system is sensible but the evaluation does not support the headline: the before and after groups seem to have different sizes (19 vs 14) and no inferential statistics are reported. read the letter →

arxiv 2501.05322 v2 pith:H6KQ2KMU submitted 2025-01-09 cs.HC

classification cs.HC
keywords autonomousvehiclespassengertrustmultimodalinterfaceexplainableartificialintelligencelargelanguagemodelbird'seyeviewhuman-centereddesignuserstudy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that riders trust autonomous vehicles more when the car explains itself through a multimodal system built around human preferences. The proposed Human-centered Multimodal Interpreter (HMI) combines bird's-eye view, map, and text displays with voice answers generated by a prompt-tuned large language model, and it lets passengers choose their preferred feedback channel. In a simulator study with 19 valid respondents, the paper reports that self-reported distrust fell from 47.4% to 14.2% after using the system, that average trust rose by over 8%, and that the share of trust or complete trust in ordinary driving rose by over 30% to 85.7%. If true, this matters because public distrust is a central barrier to autonomous vehicle adoption, and the result suggests that inexpensive explanation layers could meaningfully reduce that barrier.

What carries the argument

The load-bearing mechanism is the Human-centered Multimodal Interpreter (HMI) itself: a prompt-engineered large language model voice channel plus a visual panel combining bird's-eye view (BEV), map, and text. The BEV provides a top-down view of surrounding objects, the map provides route context, the text display narrates current actions, and the LLM answers passenger questions in concise speech. The prompt design carries the argument because it breaks explanations into three scenario-triggered levels ('No Explain', 'Explain How', 'Explain How and Why'), keeping the audio short and context-appropriate. The mechanism is what changes between the two sessions in the user study, so the entire trust comparison rides on it.

What would settle it

Run the same simulator study with two counterbalanced groups (HMI first versus no-HMI first) and identical participant sets; if the no-HMI-after group shows a comparable trust rise, or if the HMI-first group gains no more than the no-HMI-first group, the reported effect is an order or demand artifact rather than the system itself.

Watch

Extended reading notes

Core claim

The central discovery is that a context-sensitive, multimodal explanation system measurably increases passengers' self-reported trust in autonomous driving. On the paper's own terms, the HMI system significantly boosts passenger trust in AVs, increasing average trust levels by over 8%, with trust in ordinary environments rising by up to 30%. The effect is strongest in ordinary environments, where the proportion of participants reporting trust or complete trust reaches 85.7%, while overall distrust drops from 47.4% to 14.2%. The authors also find that trust in the large language model's voice explanations increases after actual use, that trust in the bird's-eye view display declines by over 20% in low-visibility conditions after HMI introduction, and that participants increasingly prefer combined BEV-plus-LLM explanations as scenarios become more complex (64.3%, 71.4%, and 78.6% across ordinary, low visibility, and emergency braking).

Load-bearing premise

The load-bearing premise is that the trust rise between the first session (no HMI) and the later session (with HMI) is caused by the HMI system itself, not by practice, fatigue, or a desire to please the experimenter, since the HMI condition always came second and the reported percentages may not come from identical respondent sets.

Editorial extensions

If this is right

  • An autonomous vehicle that can explain its actions in real time through a multimodal interface should raise passenger trust, with average gains above 8% and ordinary-driving trust gains of roughly 30%.
  • The share of passengers reporting trust or complete trust in ordinary environments rises to 85.7%, making routine rides the clearest target for explanation systems.
  • LLM voice explanations earn more trust after actual use than before, while BEV-only trust can drop in low-visibility conditions, so the system is not uniformly beneficial across all modalities and scenarios.
  • As driving complexity increases, passengers increasingly prefer combined BEV-plus-LLM explanations over either alone, supporting a scenario-adaptive multimodal design.
  • If the effect is real, explanation systems could be prioritized in ordinary driving conditions, where the largest trust gains appear with the least dramatic scenario.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the effect transfers to production vehicles, adding an explanation layer of natural-language voice plus a top-down display could improve rider trust without changing the vehicle's sensing or driving stack, making it a relatively cheap retrofit.
  • A natural next experiment is to replace questionnaire scores with behavioral measures such as ride-acceptance rates, takeover frequency, or physiological arousal, because self-reported trust does not always match what riders actually do.
  • The finding that passengers increasingly prefer combined BEV-plus-LLM explanations as scenarios grow complex suggests an adaptive system that switches modalities by detected scenario complexity would outperform a fixed interface; this is directly testable in the same simulator setup.
  • Because all participants were young to middle-aged university affiliates, the trust gains may not transfer to elderly riders or those with less formal education, a limitation the paper itself acknowledges.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper presents a Human-centered Multimodal Interpreter (HMI) for autonomous vehicles, combining a Bird's Eye View (BEV) display, a map, a text display, and a fine-tuned LLM-based voice interaction. A preliminary survey of 30 participants informed the design. The central empirical claim is that a user study with 20 volunteers (19 valid data points) shows the HMI 'significantly boosts' passenger trust in autonomous vehicles, with average trust rising by over 8% and trust in ordinary environments rising by up to 30%. The study compares self-reported trust questionnaires before and after experiencing three driving scenarios, with the HMI present only in the second session.

Significance. If the central claim were supported, the paper would make a useful contribution to human-centered explainability for autonomous driving, particularly in combining visual, textual, and auditory explanations and in adapting explanations to driving scenarios. The system design is thoughtful and the authors report a diverse participant sample across seven countries. However, the empirical evidence for the headline claim is not currently convincing: the reported data contain an internal inconsistency in sample sizes, no inferential statistics are provided, and the experimental design confounds the HMI condition with session order. Because the paper's value rests on this user study, the lack of a sound statistical and experimental basis is a load-bearing issue.

major comments (4)
  1. [§5.4, Table 2] The before and after percentages in Table 2 imply different numbers of participants. The 'before' column sums to 100% with 19 respondents (10.5% = 2, 21.1% = 4, 15.8% = 3, 47.4% = 9, 5.3% = 1), while the 'after' column sums to 100% with 14 respondents (7.1% = 1, 64.3% = 9, 21.4% = 3). This contradicts the statement in §5.4 that 19 valid data points were obtained with one dropout. The before and after groups therefore have different sizes, so the reported change in trust is not a paired within-subject comparison and may reflect attrition rather than the HMI effect.
  2. [§5.4, Abstract] No inferential statistics are reported anywhere in the results. The word 'significantly' in the abstract is unsupported: there is no paired test, no confidence interval, no effect size, and no significance level. Please provide an appropriate statistical analysis (e.g., Wilcoxon signed-rank test for ordinal trust ratings with paired data, or a mixed model) or remove all causal and inferential language from the claims.
  3. [§5.3, Procedure] The experimental design always presents the HMI condition in a second session after a 24-hour interval, with no counterbalancing and no control condition. This fully confounds order, learning, familiarity, and demand effects with the presence of the HMI. The claim that the HMI system 'boosts passenger trust' cannot be separated from these confounds in the current design. The Limitations section (§5.5) acknowledges simulator realism and demographic skew but does not mention these design and statistical threats.
  4. [Abstract, §5.4, Table 2] The claimed 'average trust levels rising by over 8%' is not derivable from the reported data. From Table 2, the mean before score is (1×2 + 2×4 + 3×3 + 4×9 + 5×1)/19 = 60/19 ≈ 3.16, and the mean after score is (2×1 + 3×1 + 4×9 + 5×3)/14 = 56/14 = 4.00, an increase of 0.84 points, or about 26.6% relative to the before mean. If the 8% refers to a different calculation (e.g., percentage-point change on the 1–5 scale), this should be stated explicitly; as written, the headline number is not reproducible from the tables.
minor comments (5)
  1. [Title] The title contains a typo: 'V ehicles' should be 'Vehicles'.
  2. [§2.1] There is an unresolved citation marker '[?]' at the end of the first paragraph; the reference list should be completed.
  3. [§4.1, §5.2] The simulation platform is inconsistently named: §4.1 says AirSim, §5.2 first says 'Calar software' and then says 'CARLA-based simulations.' Please clarify which simulator was used.
  4. [§5.4, Table 4] In the Low Visibility column, the percentages 21.4% + 7.1% + 71.4% sum to 99.9%; either adjust the rounding or add a note that values are rounded to one decimal place.
  5. [§5.4] The text describes scores of 'three or less' as 'did not trust autonomous driving,' but a score of 3 on a 1–5 Likert scale is usually neutral; please clarify the scale labels and the interpretation of each score.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the trust outcome is an external self-report, and the preliminary survey informs design without being reused as the outcome measure.

full rationale

The paper contains no derivation chain in which a predicted quantity reduces by construction to an input. The HMI system is built from a preliminary survey (§3, e.g., "90% of testers indicating that their trust in autonomous driving(AD) would increase if there were a system to explain driving behaviors"), but that survey is not the outcome measure used to validate the system. The central empirical claim is a before/after comparison of self-reported trust questionnaire scores (§5.3–§5.4), which is an external measurement collected after system use, not an algebraic consequence of the design choices. The only in-paper citations to the authors' own prior work (refs [7] and [12]) are not invoked to justify the trust conclusion, and no uniqueness theorem or imported ansatz is used. Statistical weaknesses—non-counterbalanced order, the apparent n discrepancy between Tables 2 and 3, and the absence of inferential tests—are concerns about causal validity and reporting, not circularity, because the outcome variable is not defined in terms of the system's parameters. No step was found in which an equation, parameter, or stated assumption is equivalent to the target result by definition. The honest finding is therefore no significant circularity, with a score of 1 reflecting only minor self-citation in the reference list that is not load-bearing.

Assumptions & free parameters 1 free parameters · 2 assumptions · 0 invented entities

No mathematical model or fitted constants drive the results; the ledger reflects the measurement assumptions of the user study rather than free parameters in a derivation.

free parameters (1)
  • Distrust cutoff = score ≤ 3 on 5-point Likert scale
    The paper defines distrust as a trust score of three or less (§5.4) to compute the 47.4% to 14.2% reduction; this threshold is chosen by the authors without justification.
assumptions (2)
  • domain assumption Self-reported Likert-scale trust ratings measure passenger trust in AVs
    The study's outcome variable is a single self-report item; no validated trust scale or behavioral measures are used.
  • domain assumption Simulator scenarios evoke trust responses comparable to real road conditions
    Acknowledged in §5.5; simulator fidelity limits generalization.

how reviews work

0 comments
Cite this review

Pith. "Pith review of "What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles." pith.science (2026). https://pith.science/paper/H6KQ2KMU

@misc{pith2026250105322,
  author       = {Pith},
  title        = {Pith review of: "What's Happening"- A Human-centered Multimodal Interpreter Explaining the Actions of Autonomous Vehicles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/H6KQ2KMU}},
  note         = {Machine review of arXiv:2501.05322}
}
read the original abstract

Public distrust of self-driving cars is growing. Studies emphasize the need for interpreting the behavior of these vehicles to passengers to promote trust in autonomous systems. Interpreters can enhance trust by improving transparency and reducing perceived risk. However, current solutions often lack a human-centric approach to integrating multimodal interpretations. This paper introduces a novel Human-centered Multimodal Interpreter (HMI) system that leverages human preferences to provide visual, textual, and auditory feedback. The system combines a visual interface with Bird's Eye View (BEV), map, and text display, along with voice interaction using a fine-tuned large language model (LLM). Our user study, involving diverse participants, demonstrated that the HMI system significantly boosts passenger trust in AVs, increasing average trust levels by over 8%, with trust in ordinary environments rising by up to 30%. These results underscore the potential of the HMI system to improve the acceptance and reliability of autonomous vehicles by providing clear, real-time, and context-sensitive explanations of vehicle actions.

Figures

Figures reproduced from arXiv: 2501.05322 by the authors.

Figure 2
Figure 2. Example of HMI Visual Interface BEV is a technology that provides an overhead perspec￾tive of the vehicle and its surroundings. It displays various objects, such as vehicles and pedestrians, from a compre￾hensive, top-down view. The primary purpose of BEV is to offer an integrated global view, enabling users to gain a thorough understanding of their environment and the dy￾namics of the situation. Additionally, BEV e… view at source ↗
Figure 3
Figure 3. Examples of HMI Voice Interaction We fine-tuned the speech output of a LLM through prompt engineering to prevent excessive noise during car travel. Research [17] shows that an overload of speech con￾tent can lead to user irritation and anxiety, negatively im￾pacting the overall experience. Consequently, we adjusted the LLM to ensure that the speech output remains concise and clear. Our prompt engineering design is f… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references · 26 canonical work pages

  1. [1]

    Fear of self-driving cars on the rise

    AAA. Fear of self-driving cars on the rise. AAA Newsroom,

  2. [2]

    Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions

    Shahin Atakishiyev, Mohammad Salameh, Hengshuai Yao, and Randy Goebel. Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions. arXiv preprint arXiv:2112.11561, 2021. 2

  3. [3]

    Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions

    Shahin Atakishiyev, Mohammad Salameh, Hengshuai Yao, and Randy Goebel. Explainable artificial intelligence for autonomous driving: A comprehensive overview and field guide for future research directions. IEEE Access, 2024. 2

  4. [4]

    Research on the influence and mechanism of human–vehicle moral matching on trust in autonomous vehicles

    Na Chen, Yao Zu, and Jing Song. Research on the influence and mechanism of human–vehicle moral matching on trust in autonomous vehicles. Frontiers in Psychology, 14:1071872,

  5. [5]

    J. K. Choi and Y . G. Ji. Investigating the importance of trust on adopting an autonomous vehicle. International Journal of Human–Computer Interaction, 31(10):692–702, 2015. 1, 2

  6. [6]

    The role of human-automation consensus in multiple unmanned vehicle scheduling

    Mary L Cummings, Andrew Clare, and Christin Hart. The role of human-automation consensus in multiple unmanned vehicle scheduling. Human Factors, 52(1):17–27, 2010. 4

  7. [7]

    Energy-efficient Hybrid Model Predictive Trajectory Planning for Autonomous Electric Vehicles

    Fan Ding, Xuewen Luo, Gaoxuan Li, Hwa Hui Tew, Junn Yong Loo, Chor Wai Tong, AS Bakibillah, Ziyuan Zhao, and Zhiyu Tao. Energy-efficient hybrid model pre- dictive trajectory planning for autonomous electric vehicles. arXiv preprint arXiv:2411.06111, 2024. 2

  8. [8]

    Why did the ai make that decision? towards an explainable artificial intelligence (xai) for autonomous driving systems

    Jiqian Dong, Sikai Chen, Mohammad Miralinaghi, Tiantian Chen, Pei Li, and Samuel Labi. Why did the ai make that decision? towards an explainable artificial intelligence (xai) for autonomous driving systems. Transportation Research Part C: Emerging Technologies, 156:104358, 2023. 1, 2, 5

Show all 27 references
  1. [9]

    V oice user interface interaction design re- search based on user mental model in autonomous vehicle

    Yuemeng Du, Jingyan Qin, Shujing Zhang, Sha Cao, and Jinhua Dou. V oice user interface interaction design re- search based on user mental model in autonomous vehicle. In Human-Computer Interaction. Interaction Technologies: 20th International Conference, HCI International 2018...

  2. [10]

    Adaptive user-centered multimodal interac- tion towards reliable and trusted automotive interfaces

    Amr Gomaa. Adaptive user-centered multimodal interac- tion towards reliable and trusted automotive interfaces. In Proceedings of the 2022 International Conference on Multi- modal Interaction, pages 690–695, 2022. 2

  3. [11]

    Ai in autonomous vehicles: Opportunities, challenges, and regula- tory implications

    Nirvikar Katiyar, Abhay Shukla, Namita Chawla, Mr Raju Singh, Sudhir Kumar Singh, and Mohd Faraz Husain. Ai in autonomous vehicles: Opportunities, challenges, and regula- tory implications. Educational Administration: Theory and Practice, 30(4):6255–6264, 2024. 1, 2

  4. [12]

    Pkrd-cot: A unified chain-of-thought prompting for multi-modal large language models in au- tonomous driving, 2024

    Xuewen Luo, Fan Ding, Yinsheng Song, Xiaofeng Zhang, and Junnyong Loo. Pkrd-cot: A unified chain-of-thought prompting for multi-modal large language models in au- tonomous driving, 2024

  5. [13]

    Multimodal con- trol system for autonomous vehicles using speech and ges- ture recognition

    Takuma Nakagawa and Norihide Kitaoka. Multimodal con- trol system for autonomous vehicles using speech and ges- ture recognition. Journal of the Acoustical Society of Amer- ica, 140(4 Supplement):2963–2964, 2016. 2

  6. [14]

    Omeiza, K

    D. Omeiza, K. Kollnig, H. Web, M. Jirotka, and L. Kunze. Why not explain? effects of explanations on human percep- tions of autonomous driving. IEEE International Conference on Advanced Robotics and Its Social Impacts (ARSO), pages 194–199, 2021. 1, 2

  7. [15]

    Learning to map vehicles into bird’s eye view

    Andrea Palazzi, Guido Borghi, Davide Abati, Simone Calderara, and Rita Cucchiara. Learning to map vehicles into bird’s eye view. In Image Analysis and Processing- ICIAP 2017: 19th International Conference, Catania, Italy, September 11-15, 2017, Proceedings, Part I 19, pages 233–

  8. [16]

    Text-based information design for in-vehicle displays: a sys- tematic review

    Nattaporn Phongphaew and Arisara Jiamsanguanwong. Text-based information design for in-vehicle displays: a sys- tematic review. Transportation research part F: traffic psy- chology and behaviour, 103:442–459, 2024. 1, 2

  9. [17]

    Mobile interac- tion: automatically adapting audio output to users and con- texts on communication and media control scenarios

    Tiago Reis, Lu ´ıs Carric ¸o, and Carlos Duarte. Mobile interac- tion: automatically adapting audio output to users and con- texts on communication and media control scenarios. InUni- versal Access in Human-Computer Interaction. Intelligent and Ubiquitous Interaction Environme...

  10. [18]

    why should i trust you?

    Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. “why should i trust you?” explaining the predictions of any classifier. Proceedings of the 22nd ACM SIGKDD Interna- tional Conference on Knowledge Discovery and Data Min- ing (KDD ’16), pages 1135–1144, 2016. 1, 2

  11. [19]

    Enhancing trust in autonomous vehicles through intelligent user interfaces that mimic human behavior

    Peter AM Ruijten, Jacques MB Terken, and Sanjeev N Chan- dramouli. Enhancing trust in autonomous vehicles through intelligent user interfaces that mimic human behavior. Mul- timodal Technologies and Interaction, 2(4):62, 2018. 1, 2

  12. [20]

    Airsim: High-fidelity visual and physical simulation for autonomous vehicles

    Shital Shah, Debadeepta Dey, Chris Lovett, and Ashish Kapoor. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Re- sults of the 11th International Conference , pages 621–635. Springer, 2018. 4

  13. [21]

    Improvement of au- tonomous vehicles trust through synesthetic-based multi- modal interaction

    Xiaofeng Sun and Yimin Zhang. Improvement of au- tonomous vehicles trust through synesthetic-based multi- modal interaction. IEEE Access, 9:28213–28223, 2021. 2

  14. [22]

    Reli- able and transparent in-vehicle agents lead to higher behav- ioral trust in conditionally automated driving systems

    Skye Taylor, Manhua Wang, and Myounghoon Jeon. Reli- able and transparent in-vehicle agents lead to higher behav- ioral trust in conditionally automated driving systems. Fron- tiers in Psychology, 14:1121622, 2023. 2

  15. [23]

    Ad- vancing explainable autonomous vehicle systems: A com- prehensive review and research roadmap

    Sule Tekkesinoglu, Azra Habibovic, and Lars Kunze. Ad- vancing explainable autonomous vehicle systems: A com- prehensive review and research roadmap. arXiv preprint arXiv:2404.00019, 2024. 5

  16. [24]

    Trust in automated vehicles: constructs, psychological processes, and assessment

    Francesco Walker, Yannick Forster, Sebastian Hergeth, Jo- hannes Kraus, William Payre, Philipp Wintersberger, and Marieke Martens. Trust in automated vehicles: constructs, psychological processes, and assessment. Frontiers in Psy- chology, 14:1279271, 2023. 1, 2

  17. [25]

    i’d like an explanation for that!

    G. Wiegand, M. Eiband, M. Haubelt, and H. Hussmann. “i’d like an explanation for that!” exploring reactions to unex- pected autonomous driving. 22nd International Conference on Human-Computer Interaction with Mobile Devices and Services (MobileHCI ’20), pages Article 36, 1–11, 2020. 1

  18. [26]

    Making pre-trained language models end-to-end few-shot learners with con- trastive prompt tuning

    Ziyun Xu, Chengyu Wang, Minghui Qiu, Fuli Luo, Runxin Xu, Songfang Huang, and Jun Huang. Making pre-trained language models end-to-end few-shot learners with con- trastive prompt tuning. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining...

  19. [27]

    Review and challenge: High defini- tion map technology for intelligent connected vehicle

    Mengmeng Yang, Kun Jiang, Benny Wijaya, Tuopu Wen, Jinyu Miao, Jin Huang, Cao Zhong, Wei Zhang, Huixian Chen, and Diange Yang. Review and challenge: High defini- tion map technology for intelligent connected vehicle. Fun- damental Research, 2024. 5 8

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.