Pith. sign in

REVIEW 5 cited by

A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14660 v1 pith:CS5ZHYPL submitted 2024-10-18 cs.LG

classification cs.LG
keywords rewardfeedbackcodetaskscardfunctiontrajectorybetter
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large Language Models (LLMs) have shown significant potential in designing reward functions for Reinforcement Learning (RL) tasks. However, obtaining high-quality reward code often involves human intervention, numerous LLM queries, or repetitive RL training. To address these issues, we propose CARD, a LLM-driven Reward Design framework that iteratively generates and improves reward function code. Specifically, CARD includes a Coder that generates and verifies the code, while a Evaluator provides dynamic feedback to guide the Coder in improving the code, eliminating the need for human feedback. In addition to process feedback and trajectory feedback, we introduce Trajectory Preference Evaluation (TPE), which evaluates the current reward function based on trajectory preferences. If the code fails the TPE, the Evaluator provides preference feedback, avoiding RL training at every iteration and making the reward function better aligned with the task objective. Empirical results on Meta-World and ManiSkill2 demonstrate that our method achieves an effective balance between task performance and token efficiency, outperforming or matching the baselines across all tasks. On 10 out of 12 tasks, CARD shows better or comparable performance to policies trained with expert-designed rewards, and our method even surpasses the oracle on 3 tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilinguality in LLM-Designed Reward Functions for Restless Bandits: Effects on Task Performance and Fairness

    cs.CL 2025-01 conditional novelty 6.0 of 10

    English prompts produce better LLM-designed reward functions for RMABs than Hindi, Tamil, or Tulu prompts, and complex prompts worsen performance and fairness across languages.

  2. Reward Evolution with Graph-of-Thoughts: A Bi-Level Language Model Framework for Reinforcement Learning

    cs.RO 2025-09 conditional novelty 5.0 of 10

    RE-GoT combines graph-of-thoughts planning in LLMs with VLM feedback from rollout videos to automatically write and refine RL reward functions, beating prior LLM-based reward design on RoboGen and ManiSkill2.

  3. An Automated Reinforcement Learning Reward Design Framework with Large Language Model for Cooperative Platoon Coordination

    cs.LG 2025-04 conditional novelty 5.0 of 10

    PCRD, an LLM-based framework, automatically writes and evolves reward functions for multi-agent platoon coordination, averaging about 10% higher objective values than human-designed rewards in the tested simulator.

  4. A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

    cs.CR 2025-04 conditional novelty 5.0 of 10

    A large collaborative survey organizes LLM and LLM-agent safety issues into a full-stack lifecycle framework from data preparation to deployment.

  5. Multi-agent Embodied AI: Advances and Future Directions

    cs.AI 2025-05 conditional novelty 3.0 of 10

    A survey that maps multi-agent embodied AI methods and benchmarks across control, learning, and generative-model categories, and lists open challenges.

Pith tools