Pith. sign in

REVIEW 1 cited by

Align Your Intents: Offline Imitation Learning via Optimal Transport

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13037 v2 pith:V4GDBS37 submitted 2024-02-20 cs.LG cs.AI

classification cs.LGcs.AI
keywords learningofflineoptimalimitationrewardtransportagentailot
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Offline Reinforcement Learning (RL) addresses the problem of sequential decision-making by learning optimal policy through pre-collected data, without interacting with the environment. As yet, it has remained somewhat impractical, because one rarely knows the reward explicitly and it is hard to distill it retrospectively. Here, we show that an imitating agent can still learn the desired behavior merely from observing the expert, despite the absence of explicit rewards or action labels. In our method, AILOT (Aligned Imitation Learning via Optimal Transport), we involve special representation of states in a form of intents that incorporate pairwise spatial distances within the data. Given such representations, we define intrinsic reward function via optimal transport distance between the expert's and the agent's trajectories. We report that AILOT outperforms state-of-the art offline imitation learning algorithms on D4RL benchmarks and improves the performance of other offline RL algorithms by dense reward relabelling in the sparse-reward tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Is Optimal Transport Necessary for Inverse Reinforcement Learning?

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Simple nearest-neighbor and segment-matching reward functions match or beat Optimal Transport based Inverse RL across 32 benchmarks.

Pith tools