R1-RE uses GRPO reinforcement learning with format and accuracy rewards to make a 7B model reason through annotation guidelines, improving cross-domain relation classification.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
R1-RE: Cross-Domain Relation Extraction with RLVR
R1-RE uses GRPO reinforcement learning with format and accuracy rewards to make a 7B model reason through annotation guidelines, improving cross-domain relation classification.