Pith. sign in

REVIEW 3 cited by

Active Preference-Based Gaussian Process Regression for Reward Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.02575 v2 pith:PJXOUXDE submitted 2020-05-06 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords rewardfunctionsapproachdemonstrationslearningpreference-basedchallengesdifficult
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Designing reward functions is a challenging problem in AI and robotics. Humans usually have a difficult time directly specifying all the desirable behaviors that a robot needs to optimize. One common approach is to learn reward functions from collected expert demonstrations. However, learning reward functions from demonstrations introduces many challenges: some methods require highly structured models, e.g. reward functions that are linear in some predefined set of features, while others adopt less structured reward functions that on the other hand require tremendous amount of data. In addition, humans tend to have a difficult time providing demonstrations on robots with high degrees of freedom, or even quantifying reward values for given demonstrations. To address these challenges, we present a preference-based learning approach, where as an alternative, the human feedback is only in the form of comparisons between trajectories. Furthermore, we do not assume highly constrained structures on the reward function. Instead, we model the reward function using a Gaussian Process (GP) and propose a mathematical formulation to actively find a GP using only human preferences. Our approach enables us to tackle both inflexibility and data-inefficiency problems within a preference-based learning framework. Our results in simulations and a user study suggest that our approach can efficiently learn expressive reward functions for robotics tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Generalizing Preference-based Reinforcement Learning: a Rationality Model for Incomparability

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A Bradley-Terry-style rationality model with an incomparability score based on utility-difference standard deviation recovers multi-dimensional rewards and Pareto frontiers from trajectory comparisons that include inc...

  2. Residual Reward Models for Preference-based Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Combining a hand-designed or learned prior reward with a preference-trained residual improves sample efficiency and final performance in preference-based reinforcement learning.

  3. Learning Implicit Social Navigation Behavior using Deep Inverse Reinforcement Learning

    cs.RO 2025-01 conditional novelty 4.0 of 10

    S-MEDIRL, a deep inverse RL method with a bilateral filtering smoothing loss and demonstration extrapolation, learns to yield and avoid deadlock in a narrow crossing, reaching about 92% success.

Pith tools