Pith. sign in

Evalchemy: Automatic evals for llms, 2024

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CL 1

years

2025 1

verdicts

REJECT 1

roles

background 1

polarities

background 1

representative citing papers

citing papers explorer

Showing 1 of 1 citing paper.

  • Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models cs.CL · 2025-06-05 · reject · none · ref 6

    RLSC, a method that uses a language model's self-confidence as reward, is shown to improve math benchmark accuracy, but the results are undermined by training on the AIME test set and the method reduces to known self-distillation.