Pith. sign in

REVIEW 1 cited by

Understanding Reward Ambiguity Through Optimal Transport Theory in Inverse Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12055 v1 pith:7XTURZHJ submitted 2023-10-18 cs.LG cs.SYeess.SYmath.OC

classification cs.LGcs.SYeess.SYmath.OC
keywords rewardambiguityfunctionsgeometricbehaviorscentralchallengesexpert
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In inverse reinforcement learning (IRL), the central objective is to infer underlying reward functions from observed expert behaviors in a way that not only explains the given data but also generalizes to unseen scenarios. This ensures robustness against reward ambiguity where multiple reward functions can equally explain the same expert behaviors. While significant efforts have been made in addressing this issue, current methods often face challenges with high-dimensional problems and lack a geometric foundation. This paper harnesses the optimal transport (OT) theory to provide a fresh perspective on these challenges. By utilizing the Wasserstein distance from OT, we establish a geometric framework that allows for quantifying reward ambiguity and identifying a central representation or centroid of reward functions. These insights pave the way for robust IRL methodologies anchored in geometric interpretations, offering a structured approach to tackle reward ambiguity in high-dimensional settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning

    eess.SY 2025-06 reject novelty 5.0 of 10

    A cooperative MARL algorithm that regularizes each agent's policy toward the Sinkhorn barycenter of the team's visitation distributions, with a claimed but insufficiently proven geometric convergence guarantee.

Pith tools