Pith. sign in

REVIEW 1 cited by

Explainable Multi-Agent Reinforcement Learning for Temporal Queries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.10378 v1 pith:ODTY2TJM submitted 2023-05-17 cs.AI

classification cs.AI
keywords marlapproachexplanationsqueryuseragentsmulti-agentproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As multi-agent reinforcement learning (MARL) systems are increasingly deployed throughout society, it is imperative yet challenging for users to understand the emergent behaviors of MARL agents in complex environments. This work presents an approach for generating policy-level contrastive explanations for MARL to answer a temporal user query, which specifies a sequence of tasks completed by agents with possible cooperation. The proposed approach encodes the temporal query as a PCTL logic formula and checks if the query is feasible under a given MARL policy via probabilistic model checking. Such explanations can help reconcile discrepancies between the actual and anticipated multi-agent behaviors. The proposed approach also generates correct and complete explanations to pinpoint reasons that make a user query infeasible. We have successfully applied the proposed approach to four benchmark MARL domains (up to 9 agents in one domain). Moreover, the results of a user study show that the generated explanations significantly improve user performance and satisfaction.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning

    cs.LG 2025-02 conditional novelty 4.0 of 10

    Neurons in DRL agents are matched to short Boolean formulas over hand-defined state predicates, with anecdotal perturbation evidence that these matches reflect real behavior.

Pith tools