TGDPO modifies DPO by weighting each token's log-ratio with an external token-level reward, and reports win-rate gains over DPO and SimPO on three instruction-following benchmarks.
Ultrafeedback: Boosting language models with scaled AI feedback
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
TGDPO modifies DPO by weighting each token's log-ratio with an external token-level reward, and reports win-rate gains over DPO and SimPO on three instruction-following benchmarks.