Pith. sign in

REVIEW 2 cited by

Vanishing Gradients in Reinforcement Finetuning of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.20703 v3 pith:G6A2ILVN submitted 2023-10-31 cs.LG cs.AIcs.CLstat.ML

classification cs.LGcs.AIcs.CLstat.ML
keywords rewarddeviationexpectedfinetuninggradientgradientssmallstandard
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Pretrained language models are commonly aligned with human preferences and downstream tasks via reinforcement finetuning (RFT), which refers to maximizing a (possibly learned) reward function using policy gradient algorithms. This work identifies a fundamental optimization obstacle in RFT: we prove that the expected gradient for an input vanishes when its reward standard deviation under the model is small, even if the expected reward is far from optimal. Through experiments on an RFT benchmark and controlled environments, as well as a theoretical analysis, we then demonstrate that vanishing gradients due to small reward standard deviation are prevalent and detrimental, leading to extremely slow reward maximization. Lastly, we explore ways to overcome vanishing gradients in RFT. We find the common practice of an initial supervised finetuning (SFT) phase to be the most promising candidate, which sheds light on its importance in an RFT pipeline. Moreover, we show that a relatively small number of SFT optimization steps on as few as 1% of the input samples can suffice, indicating that the initial SFT phase need not be expensive in terms of compute and data labeling efforts. Overall, our results emphasize that being mindful for inputs whose expected gradient vanishes, as measured by the reward standard deviation, is crucial for successful execution of RFT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

    cs.LG 2025-05 conditional novelty 5.0 of 10

    DGRO decouples the KL regularization coefficient in reward optimization into two hyperparameters and shows strong reasoning results, though its reward-variance ablation is confounded.

  2. Learning Explainable Dense Reward Shapes via Bayesian Optimization

    cs.LG 2025-04 conditional novelty 5.0 of 10

    Reward shaping based on SHAP/LIME token attributions, with weights optimized by Bayesian optimization, improves RLHF training speed and downstream win rates while preserving the optimal policy.

Pith tools