Pith. sign in

{severe level} harm question:

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.CL 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Data-adaptive Safety Rules for Training Reward Models

cs.CL · 2025-01-26 · conditional · novelty 6.0

Selecting the five safety rules with largest response discrepancy maximizes mutual information under stated assumptions, and a reward model trained with this adaptive labeling achieves top RewardBench safety.

citing papers explorer

Showing 1 of 1 citing paper.

  • Data-adaptive Safety Rules for Training Reward Models cs.CL · 2025-01-26 · conditional · none · ref 5

    Selecting the five safety rules with largest response discrepancy maximizes mutual information under stated assumptions, and a reward model trained with this adaptive labeling achieves top RewardBench safety.