Pith. sign in

REVIEW 1 cited by

FuRL: Visual-Language Models as Fuzzy Rewards for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.00645 v2 pith:JZFBZH74 submitted 2024-06-02 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords rewardtasksfurlfine-tuningfuzzylearningmethodmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work, we investigate how to leverage pre-trained visual-language models (VLM) for online Reinforcement Learning (RL). In particular, we focus on sparse reward tasks with pre-defined textual task descriptions. We first identify the problem of reward misalignment when applying VLM as a reward in RL tasks. To address this issue, we introduce a lightweight fine-tuning method, named Fuzzy VLM reward-aided RL (FuRL), based on reward alignment and relay RL. Specifically, we enhance the performance of SAC/DrQ baseline agents on sparse reward tasks by fine-tuning VLM representations and using relay RL to avoid local minima. Extensive experiments on the Meta-world benchmark tasks demonstrate the efficacy of the proposed method. Code is available at: https://github.com/fuyw/FuRL.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preference VLM: Leveraging VLMs for Scalable Preference-Based Reinforcement Learning

    cs.LG 2025-02 conditional novelty 5.0 of 10

    PrefVLM combines VLM-generated trajectory preferences with selective human feedback and inverse-dynamics VLM adaptation, matching PEBBLE on five Meta-World tasks with up to 2x fewer human labels.

Pith tools