Pith. sign in

REVIEW 2 cited by

LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.08958 v1 pith:2VRRBXUX submitted 2023-12-14 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords modelsagentfoundationlearningteachersenvironmentfeedbackframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a framework that leverages foundation models as teachers, guiding a reinforcement learning agent to acquire semantically meaningful behavior without human feedback. In our framework, the agent receives task instructions grounded in a training environment from large language models. Then, a vision-language model guides the agent in learning the multi-task language-conditioned policy by providing reward feedback. We demonstrate that our method can learn semantically meaningful skills in a challenging open-ended MineDojo environment while prior unsupervised skill discovery methods struggle. Additionally, we discuss observed challenges of using off-the-shelf foundation models as teachers and our efforts to address them.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Think Twice, Act Once: A Co-Evolution Framework of LLM and RL for Large-Scale Decision Making

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A training-time LLM that refines actions and reshapes rewards, co-trained with an RL controller through a shared buffer, reports large gains over RL and LLM baselines on three L2RPN power grid challenges.

  2. Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback

    cs.AI 2025-06 conditional novelty 5.0 of 10

    EXIF repeatedly has a teacher agent explore an environment, relabel the exploration as tasks, train a student agent on it, and use the student's failures to guide the next round, improving 7B-8B agents in Webshop and Crafter.

Pith tools