Pith. sign in

REVIEW 1 cited by

Explaining Agent Behavior with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10346 v1 pith:EZA4U3HN submitted 2023-09-19 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords behavioragentexplanationslanguageagentsapproachhumanlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Intelligent agents such as robots are increasingly deployed in real-world, safety-critical settings. It is vital that these agents are able to explain the reasoning behind their decisions to human counterparts, however, their behavior is often produced by uninterpretable models such as deep neural networks. We propose an approach to generate natural language explanations for an agent's behavior based only on observations of states and actions, agnostic to the underlying model representation. We show how a compact representation of the agent's behavior can be learned and used to produce plausible explanations with minimal hallucination while affording user interaction with a pre-trained large language model. Through user studies and empirical experiments, we show that our approach generates explanations as helpful as those generated by a human domain expert while enabling beneficial interactions such as clarification and counterfactual queries.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI

    cs.LG 2025-06 conditional novelty 5.0 of 10

    An in-context LLM framework that produces dual expert and non-expert explanations, evaluated on well-being clustering with a user study and LIME-alignment metrics.

Pith tools