REVIEW 4 cited by
Beyond Imitation: Leveraging Fine-grained Quality Signals for Alignment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Alignment with human preference is a desired property of large language models (LLMs). Currently, the main alignment approach is based on reinforcement learning from human feedback (RLHF). Despite the effectiveness of RLHF, it is intricate to implement and train, thus recent studies explore how to develop alternative alignment approaches based on supervised fine-tuning (SFT). A major limitation of SFT is that it essentially does imitation learning, which cannot fully understand what are the expected behaviors. To address this issue, we propose an improved alignment approach named FIGA. Different from prior methods, we incorporate fine-grained (i.e., token or phrase level) quality signals that are derived by contrasting good and bad responses. Our approach has made two major contributions. Firstly, we curate a refined alignment dataset that pairs initial responses and the corresponding revised ones. Secondly, we devise a new loss function can leverage fine-grained quality signals to instruct the learning of LLMs for alignment. Extensive experiments have demonstrated the effectiveness of our approaches by comparing a number of competitive baselines.
Forward citations
Cited by 4 Pith papers
-
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Jointly training a Bradley-Terry preference head and a multi-attribute regression head on a shared embedding improves reward-model robustness to reward hacking and boosts multi-objective scoring performance.
-
T-REG: Preference Optimization with Token-Level Reward Regularization
T-REG adds contrastive-prompt self-generated token-level rewards as a weighted regularizer to DPO and SimPO, improving Alpaca Eval 2 length-controlled win rate by up to 3.8% and Arena-Hard win rate by up to 4.4% over ...
-
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
A contrastive token-scoring method identifies 'critical tokens' in incorrect reasoning traces and penalizes them in DPO, yielding small but consistent accuracy gains on math benchmarks.
-
Learning Explainable Dense Reward Shapes via Bayesian Optimization
Reward shaping based on SHAP/LIME token attributions, with weights optimized by Bayesian optimization, improves RLHF training speed and downstream win rates while preserving the optimal policy.
Discussion (0). Continue with ORCID to comment.